Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
afro88
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
afro88
16d ago
It's not the techs fault. This person chose to use it this way. They could have also carved out 20 mins at the start of their workday to do the same thing
2.
▲
by
afro88
22d ago
> something that can beat a turing test w/o sweating Ok I'll be that guy. It's pretty easy to figure out if you're talking to an LLM now we know it's tics, failure modes, jailbreak techniques etc
3.
▲
by
afro88
23d ago
> Give me the cold hard truth. Does this actually work to get better answers, or does it now just say "the cold hard truth" instead of "the honest take"
4.
▲
by
afro88
24d ago
I mean... that's actually amazing advice. Not because they would grow up to create browser startups. But because they would grow up to create web startups that succeed because of very fundamental of how the web is rendered. Which is pa
5.
▲
by
afro88
26d ago
Where?
6.
▲
by
afro88
1mo ago
> I would expect an ethics team to build frameworks that can help train/eval the model that the company spend millions of dollars and months training is going to be aligned to the ethical stances the company chooses. Which like any
7.
▲
by
afro88
1mo ago
I'm not a fan of Zuckerberg in the least, but one area a super intelligent lawyer would be fine at is being drowned in court filings and paperwork The bigger problem with his argument IMO is that even with open models, it's still
8.
▲
by
afro88
1mo ago
I'm all for this and I'm keen to try it. I certainly don't want to take away from sharing another neat use case for learning. But > What you get is a beautiful animation that is 100% accurate and free of hallucinations 100
9.
▲
by
afro88
1mo ago
Well go on then...
10.
▲
by
afro88
1mo ago
> The beauty of intelligence at this cost (even if it's not SOTA) is that it opens a whole bunch of new use cases. Test failure in CI? Have the bot automatically propose a fix, its cheap enough that you can discard it w/h issue
11.
▲
by
afro88
1mo ago
You can't do that with tools either. Skills are basically prompts - they're not analogous to tools or MCPs. I'm not sure what your point is
12.
▲
by
afro88
1mo ago
Gary doesn't argue it's hype though. He argues 2 things: other people are getting carried away with the result, and we don't know enough about how it was reached to know where it falls on the impressive scale. He literally sa
13.
▲
by
afro88
2mo ago
I've got a few of these. There's one in particular that I use quite often and have for about a year, vibed for myself: it's a chat interface that walks you through processing an emotional or difficult moment, following a proc
14.
▲
by
afro88
2mo ago
Are they in 2026? I haven't had an issue with json and LLMs in a long while
15.
▲
by
afro88
2mo ago
IMO a much better test would be designs that aren't AI to begin with. Much more useful to see how well a model can html an image design without slopping it up
16.
▲
by
afro88
2mo ago
> Only an encrypted blind relay to allow for shared editing. The relay doesn't see any of the data. Would love to know more about how this works then? Is it more or less encrypted P2P?
17.
▲
by
afro88
2mo ago
Is this a quote from a book? Beautifully written
18.
▲
by
afro88
2mo ago
When did that happen with Codex? I thought that was a Claude Code thing
19.
▲
by
afro88
2mo ago
It's not about figuring out if it's LLM written though. The style is hard to read and annoying. With the kind of sentences GP was talking about it's actually harder to get the substance.
20.
▲
by
afro88
2mo ago
Curious whether you were just bare asking it questions, or whether you provided it with lessons one by one with instruction that the lesson is the baseline truth etc
21.
▲
by
afro88
2mo ago
This has been the case since the early days. Aider had a bunch of code to be very forgiving with formatting of tool calls (file editing in particular at first). It's just the nature of the beast. It surprises me that Pi doesn't ha
22.
▲
by
afro88
3mo ago
Maybe I'm too optimistic, but given appropriate skills and references (not just for writing but also reviewing) and intelligent use of subagents for isolated reviews and checks, you can lengthen the leash a bit. But you still need to p
23.
▲
by
afro88
3mo ago
We selected PRs (real ones we merged over the 6 months prior) and have an "LLM as judge" score how close the AI generated code is to the PR. Same as how other benchmarks do it, but it's with tasks we actually do and code we h
24.
▲
by
afro88
3mo ago
Similar result on our kotlin coding benchmark at work. It measures how close agents can get to a small mergable PR (according to my team). 20 tasks of varying difficulty, with 5 attempts each, LLM as judge to evaluate accuracy (same outcome
25.
▲
by
afro88
3mo ago
I'd love to read about the predictions that have been wrong (genuinely)
26.
▲
by
afro88
3mo ago
I wonder if there's a way to include data that's so unique you can prove it was trained on and sue later
27.
▲
by
afro88
3mo ago
> The dynamic of agent codes human reviews does seem like the only sane one for the foreseeable future. Even Anthropic themselves still fall back to this. Do they? I saw some crazy stat from the guy who built claude code that he was push
28.
▲
by
afro88
3mo ago
This is a branching point. One dev would find someone else and convince them to approve it. Another would redo the task (code is cheap now, right?) in a PR stack that can actually be reviewed, cleaned up etc. I hope they were the latter.
29.
▲
by
afro88
4mo ago
That's an example of why it would be useful for someone to actually do it. A random commenter on HN is one thing. A direct comparison on a brand new app that isn't part of any training is another
30.
▲
by
afro88
4mo ago
It's very addictive when you're working on something cool and the agents are iterating nicely. Instead of browsing reddit / HN / instagram etc during downtime, I find it much more fun to build something.
More ›