Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
cpnwaugha
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
cpnwaugha
27d ago
Formatting the command block for easy copy-paste. Via the CLI: pip install vlmrun vlmrun gw models vlmrun config set --api-key 'vlmrun' # anon-user, rate-limited vlmrun gw chat <doc>.pdf -m zai-org/glm-ocr vlmrun gw cha
2.
▲
by
cpnwaugha
28d ago
Yeah, GPT 5.6 Sol is very good. Generally, Gemini models remain SOTA for VLM tasks with 3.7-flash at the top. --- That said, considering variables like cost (say, over 100k PDF pages) and accuracy requirements (e.g., construction documents
3.
▲
by
cpnwaugha
28d ago
With a one-line change, you can switch between open-weight OCR VLMs (DeepSeek OCR 2, GLM-OCR, dots.mocr, Paddle OCR VL, PP-OCRv6, etc.) and process 100K+ pages for under $60 on VLM Run Gateway > one OpenAI-compatible endpoint for open-we
4.
▲
by
cpnwaugha
28d ago
Here's a Colab link to explore the gateway: https://colab.research.google.com/drive/1RkuVIyuc5Po-UlcSlFy...
5.
▲
by
cpnwaugha
2mo ago
mm can spot charts on a screen. did you try it?
6.
▲
by
cpnwaugha
2mo ago
Awesome!
7.
▲
by
cpnwaugha
2mo ago
Sirus (Qax) is 90% built with LLM coding agents (mix). I should say that I'm deeply involved in the process - architecting and reviewing (and sometimes revising) to ensure the code is clean and aligned with my preferences. one approach
8.
▲
by
cpnwaugha
2mo ago
The comments in this post strongly validate the need for reliable video processing and understanding with VLMs. While you can use Gemini or other local VLMs, the real challenge is token efficiency, accuracy, and coverage. For example, how d
9.
▲
Show HN: Wikicube: instant wiki for any GitHub repo
(wikicube.vercel.app)
1 points
by
cpnwaugha
3mo ago
|
0 comments
10.
▲
by
cpnwaugha
3mo ago
Two approaches for passing binary files, say a 2-hour lecture video, to a VLM: 1. Send it whole: accurate, but slow to encode and process. 2. Keyframe extraction: fast, but could miss what matters. There's no right or wrong approach he
11.
▲
Mm – Unix tools (find/cat/grep) rebuilt for the multimodal era
(vlm.run)
1 points
by
cpnwaugha
5mo ago
|
0 comments
12.
▲
by
cpnwaugha
1y ago
We're helping litigation attorneys and insurance claims adjusters turn video into insights in seconds
13.
▲
Atlas Video Chat
(veedo.ai)
3 points
by
cpnwaugha
1y ago
|
1 comments
14.
▲
Transform your videos into shareable slides
(veedo.ai)
2 points
by
cpnwaugha
1y ago
|
1 comments
15.
▲
by
cpnwaugha
1y ago
AI-generated, with summaries and key moments Tired of sending 40-minute videos that no one watches to the end? We found that converting long videos into short-form slide (and PDF) content dramatically improves engagement and retention. With