Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
gillesjacobs
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
gillesjacobs
6d ago
Main takeaway: Average cost per attempt, without → with RTK: Claude/Fable: $1.72 → $1.64 (~5% cheaper) DeepSeek: $0.115 → $0.121 (~5% more expensive) Almost all Claude savings came from a single task. Excluding it, savi
2.
▲
by
gillesjacobs
1mo ago
He was an immoral man taking full advantage of an immoral system.
3.
▲
by
gillesjacobs
1mo ago
The custom complications on WearOs watch faces do it for me. I can see my investment portfolio value change right from my wrist.
4.
▲
by
gillesjacobs
1mo ago
Cool, there is no check on the direction of the camera. Various camera's I checked do not seem to point in the direction of the sun during totality. It does filter candidate cameras down though.
5.
▲
by
gillesjacobs
1mo ago
And I'd rather see the Mona Lisa.
6.
▲
by
gillesjacobs
2mo ago
https://youtu.be/uJblcC4lKYw This video benchmarks slop-style indicators with different skill/prompt solutions including the STE skill vs. George Orwell's six rules of writing prompt: Orwell came out on top overal
7.
▲
by
gillesjacobs
2mo ago
They asked to put EU buildings on there to grab power and cultural capital. The EU is centralizing and consolidating by any means necessary.
8.
▲
by
gillesjacobs
2mo ago
Because their fear-based marketing gave the US gov justification in blocking them for a while. They wisely didn't do that for Opus 5.
9.
▲
by
gillesjacobs
2mo ago
Same person that was mocking the hands in image generation in 2023, is the same person that was saying 'hands are fixed but it can't generate "the red dog jumps over the jump rope held by the blue pelican while juggling 5 bal
10.
▲
by
gillesjacobs
2mo ago
In ML, you want to test general capability of a model (generalizability), because you want it to perform well on unseen tasks. In that benchmark, the literal reference is leaking through web search, the agent can see the matching real codeb
11.
▲
by
gillesjacobs
2mo ago
Except the only need Chat control is fulfilling is the governments and bureaucrats need to control and check all civilian communications.
12.
▲
by
gillesjacobs
3mo ago
> this is wildly incorrect in an academic context. Worked for me in my academic career ¯\_(ツ)_/¯ Of course if you suspect a work is poignant, you should read it, then cite it.
13.
▲
by
gillesjacobs
3mo ago
Which work has more value: the abstract description of a catalogue of potential model architectures or their validated application trained on real data? In the Schmidhuber case their is 20 years and a chain of countless other works in betwe
14.
▲
by
gillesjacobs
3mo ago
Of course, but if you haven't read them you also shouldn't cite them. And that's where Schmidhuber goes off the rails: publicly shaming published papers into citing you isn't good academic practice. It's bullying.
15.
▲
by
gillesjacobs
4mo ago
With some caveats, you wouldn't be able to connect two 4k monitors to a dock without TB5.
16.
▲
by
gillesjacobs
5mo ago
https://web.archive.org/web/20260402155236/https://www.redha... Archive URL to original paper
17.
▲
by
gillesjacobs
6mo ago
I stand corrected, that is pretty scummy. I bet Moonshot is going to make them open their wallets to avoid legal trouble.
18.
▲
by
gillesjacobs
6mo ago
They probably licensed it. Still a bit deceptive not to mention it on the model card/blog post, but companies whitelabel all the time without mentioning. It goes against the ML community ethos to obscure it, but is common branding prac
19.
▲
by
gillesjacobs
6mo ago
Cursor is mostly an IDE / coding-agent harness company. So it probably makes sense for them not to train their own base model, but instead license something like Kimi and fine-tune it for their own harness and workflows. Their moat loo
20.
▲
by
gillesjacobs
8mo ago
I liked his cartoons and he did no wrong.
21.
▲
by
gillesjacobs
9mo ago
You're underselling this as a process manager, it could also be a productivity tool with some prompt changes; Determine procrastination apps: games, non-professional chat, video streaming and kill it.
22.
▲
by
gillesjacobs
9mo ago
https://archive.ph/awvmJ
23.
▲
by
gillesjacobs
11mo ago
They save money by cheap labour and batching large quantities for analysis. For the consumer this means long wait times and potentially expired DNA samples. I tried two samples with Nebula, waited 11 months total. Both samples failed. Got a
24.
▲
by
gillesjacobs
1y ago
Nice framing for PMs, but technically it is way too rosy. MCP is real but still full of low utility services and security issues, so “skills as plug-ins” is not production ready. A2A protocols were only just announced this year (Google, etc
25.
▲
by
gillesjacobs
1y ago
I am always on the lookout for new document extraction tools, but can't seem to find any benchmarks for PageIndex-OCR. There are several like OmniDocBench and readoc. So... Got benchmark?
26.
▲
by
gillesjacobs
1y ago
Extracting structure and elements from HTML should be trivial and probably has multiple libraries in your programming language of choice. Be happy you have machine-readable semantic documents, that's best-case scenario in NLP. I used t
27.
▲
by
gillesjacobs
1y ago
A suspicious lack of any performance metrics on the many standard RAG/QA benchmarks out there, except for their highly fine-tuned and dataset-specific MAFIN2.5 system. I would love the see this approach vs. a similarly well-tuned struc
28.
▲
by
gillesjacobs
1y ago
Had many a friend in the Belgian hacker scene who were threatened with legal action after responsible disclosure. To my knowledge, these threats always remained empty: if there is one thing more expensive than engineering a fix, it is start
29.
▲
by
gillesjacobs
1y ago
It doesn't do it magically. The "tools" an LLM agent calls to create responses are typically REST APIs for these services. Previously, many companies gated these APIs but with the MCP AI hype they are incentivized to expose w
30.
▲
by
gillesjacobs
1y ago
Looks very cool! I prototyped something like this with build123d for Python and Cursor + OCP VSCode plugin. Build123d is too new with too little examples out there, unlike OpenSCAD. I can only get it to generate good code with largr reasoni
More ›