Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
couAUIA
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
couAUIA
25d ago
haiku 4.5 top 7, that benchmark is absolute crap
2.
▲
Benchmark Driven Development
(charlesazam.com)
2 points
by
couAUIA
1mo ago
|
0 comments
3.
▲
Stop Reasoning Blindly. Benchmark driven development
(charlesazam.com)
2 points
by
couAUIA
1mo ago
|
1 comments
4.
▲
by
couAUIA
1mo ago
How I tackle complex problems with benchmark driven development.
5.
▲
by
couAUIA
2mo ago
Yes I agree, but I actually did a lot more runs, with different prompts, different times ect... And each time /goal had a small or insignificant impact
6.
▲
by
couAUIA
2mo ago
well thank you so much for this
7.
▲
by
couAUIA
2mo ago
A deepdive on the /goal effect on a problem literally made for this.
8.
▲
Fable 5 vs. GPT-5.6 Sol on an NP-Hard Problem: Does /goal help?
(charlesazam.com)
257 points
by
couAUIA
2mo ago
|
125 comments
9.
▲
Gemini 3.5 Pro delays due to coding performance, upgraded Flash model in testing
(9to5google.com)
2 points
by
couAUIA
2mo ago
|
0 comments
10.
▲
by
couAUIA
2mo ago
It seems that it is very strong at Frontend, I wonder how good are its multimodal capabilities. This is very important for me and not well represented in benchmarks in my opinion.
11.
▲
Kimi K3 is ranked 3rd on artificial analysis, only 2 points behind Sol
(artificialanalysis.ai)
8 points
by
couAUIA
2mo ago
|
2 comments
12.
▲
I ran an AI nuclear engineering department for a week
(charlesazam.com)
2 points
by
couAUIA
3mo ago
|
0 comments
13.
▲
White House lifts ban on Anthropic models
(ft.com)
2 points
by
couAUIA
3mo ago
|
1 comments
14.
▲
by
couAUIA
3mo ago
Please use the sharing tools found via the share button at the top or side of articles. Copying articles to share with others is a breach of FT.com T&Cs and Copyright Policy. Email licensing@ft.com to buy additional rights. Subscribers
15.
▲
Hardware Engineering as Code
(github.com)
1 points
by
couAUIA
3mo ago
|
0 comments
16.
▲
Scientific documents should be written in Python (2022)
(github.com)
1 points
by
couAUIA
3mo ago
|
0 comments
17.
▲
Mistral Compute? I hear Mistral Cloud
(mistral.ai)
13 points
by
couAUIA
4mo ago
|
0 comments
18.
▲
Proton Pass for AI Agents
(proton.me)
3 points
by
couAUIA
4mo ago
|
0 comments
19.
▲
by
couAUIA
4mo ago
LoL they added "Copilot AI Model Providers" in githubstatus and it has 100% up time. Thanks for pointing out that nobody is using that thing
20.
▲
Ask HN: Should I continue this project ? (Being able to change AI harness)
(github.com)
1 points
by
couAUIA
5mo ago
|
1 comments
21.
▲
Evaluate Your Own RAG: Why Best Practices Failed Us
(charlesazam.com)
2 points
by
couAUIA
7mo ago
|
0 comments
22.
▲
GLM-5 topped the coding benchmarks. Then I used it
(charlesazam.com)
5 points
by
couAUIA
7mo ago
|
1 comments
23.
▲
by
couAUIA
7mo ago
TL;DR: GLM-5 tops coding benchmarks. I tested it on an unpublished NP-hard optimization problem (KIRO) and 89-task Terminal-Bench. Best case: competitive. Typical case: 30% invalid output, every trial timed out, and two identical runs could
24.
▲
I benchmarked 4 coding agents on an NP-hard problem I solved 8 years ago
(charlesazam.com)
3 points
by
couAUIA
7mo ago
|
2 comments
25.
▲
by
couAUIA
7mo ago
I gave an unpublished fiber network optimization problem to Claude Code, Codex, Gemini CLI, and Mistral. The score is total fiber length (lower is better). A good human solution in 30 minutes: ~40,000. My best after days of C++: 34,123. Giv
26.
▲
I Accidentally Rebuilt OpenHands from Scratch – Here's What I Learned
(huggingface.co)
2 points
by
couAUIA
9mo ago
|
0 comments
27.
▲
Evaluate Your Own RAG, Why Best Practices Failed Us
(huggingface.co)
2 points
by
couAUIA
11mo ago
|
1 comments
28.
▲
by
couAUIA
11mo ago
At Jimmy, we're developing France's first Small Modular Reactor. Our engineers needed to quickly search and extract insights from thousands of complex scientific PDFs—nuclear research papers, regulatory documents, multilingual con