Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
fdefitte
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
fdefitte
7mo ago
The dog ships faster because it has zero opinions about the architecture.
2.
▲
Show HN: Cobalt – Unit tests for AI agents, like Jest but for LLMs
(github.com)
3 points
by
fdefitte
7mo ago
|
0 comments
3.
▲
by
fdefitte
7mo ago
Good point
4.
▲
by
fdefitte
7mo ago
Agreed on ZLUDA being the practical choice. This project is more impressive as a "build a GPU compiler from scratch" exercise than as something you'd actually use for ML workloads. The custom instruction encoding without LLVM
5.
▲
by
fdefitte
7mo ago
The filter used to be effort. You had to care enough to spend weeks on something, which meant you probably understood the problem deeply. Now that filter is gone and we get a flood of "I prompted this in 20 minutes" posts where th
6.
▲
by
fdefitte
7mo ago
The 8% one-shot number is honestly better than I expected for a model this capable. The real question is what sits around the model. If you're running agents in production you need monitoring and kill switches anyway, the model being &
7.
▲
by
fdefitte
7mo ago
The "native multimodal agents" framing is interesting. Everyone's focused on benchmark numbers but the real question is whether these models can actually hold context across multi-step tool use without losing the plot. That&#
8.
▲
by
fdefitte
7mo ago
That 95% payout only works if you already know what good looks like. The sketchy part is when you can't tell the diff between correct and almost-correct. That's where stuff goes sideways.
9.
▲
by
fdefitte
7mo ago
Skills are great for static stuff but they kinda fall apart when the agent needs to interact with live state. WebMCP actually fills a real gap there imo.
10.
▲
by
fdefitte
7mo ago
Agent teams working autonomously sounds cool until you actually try it. We've been running multi-agent setups and honestly the failure modes are hilarious. They don't crash, they just quietly do the wrong thing and act super confi
11.
▲
by
fdefitte
7mo ago
Hi everyone ! Super happy to release this package. We feel like Evals belong in the CI like unit testing, and should be easy to setup and run automatically. Can't wait to get your feedback !
12.
▲
Show HN: We built Cobalt, Open source unit testing for AI Agents
(github.com)
3 points
by
fdefitte
7mo ago
|
1 comments
13.
▲
AI Evaluation Methods by Use Case
(notion.so)
1 points
by
fdefitte
1y ago
|
1 comments
14.
▲
by
fdefitte
1y ago
Free guide, enjoy ! Made with by the Basalt team
15.
▲
Show HN: Free Prompt Grading Tool
(getbasalt.ai)
1 points
by
fdefitte
1y ago
|
1 comments
16.
▲
by
fdefitte
1y ago
Love it. AI is actually really better at judging the quality of content than it is at producing content. Kind of like humans actually :)
17.
▲
AI Hedge Fund
(github.com)
3 points
by
fdefitte
1y ago
|
0 comments
18.
▲
by
fdefitte
1y ago
This makes total sense. When a country is creating public software, it should be open source by default. This is the only way to create trust. In the long run, open source and closed source government software will probably differentiate di
19.
▲
by
fdefitte
1y ago
I think it's more a guideline principle for public software, for exemple apps that are used by citizens to declare taxes, renews IDs...
20.
▲
by
fdefitte
1y ago
France has an undeserved bad reputation for this stuff. As a french citizen, I'm amazed to see how easy it has become to do anything administrative online, with great tools such as France Connect that allows a single login method for a