Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
wolttam
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
wolttam
7d ago
If you loop an entire transformer model on itself, that seems like by-definition hidden reasoning. If the output of the model is its reasoning trace, and you simply feed that back into the model again at inference time instead of outputting
2.
▲
by
wolttam
7d ago
Just my intuition about it but it does seem like a data issue. V4 flash and V4 pro feel very similar, which would make sense if they were pre-trained on largely the same corpus. All that would suggest to me is that V4 Flash is capable of ab
3.
▲
by
wolttam
8d ago
Brought to you by the company that gave an LLM access to directly reset user passwords, all you had to do was manipulate the LLM.
4.
▲
by
wolttam
8d ago
That IS very cool. I can’t help but appreciate the insight and forethought that goes into the decision to construct such massive waterworks projects.
5.
▲
by
wolttam
8d ago
I don’t get the push on infra. Focus on the models
6.
▲
by
wolttam
9d ago
But it will be automated long before then
7.
▲
by
wolttam
12d ago
~10B tokens a month is pretty typical overall input/output usage from my own experience and other developer accounts I've seen
8.
▲
by
wolttam
15d ago
Unfortunately it demonstrates effectively zero reason to use this model over, say, GLM 5.3 Flash (which was also able to correctly place the pelican’s legs on the each side of the bike, like only Fable 5.1 xhigh was able to do here) I sti
9.
▲
by
wolttam
15d ago
Don’t worry. Folks are entirely too concerned about the distinction between human made and LLM made. In a short handful of years, it will be incredibly rare to find something an LLM hasn’t been involved with
10.
▲
by
wolttam
17d ago
I think the solution is for the POW being done by the clients to *actually benefit the site owner*. Users remain just as mildly annoyed as with Anubis, but maybe a bit less knowing that the work they’re doing benefits the site owner/au
11.
▲
by
wolttam
21d ago
That will hurt latency and latency is very important for good tensor-parallelism performance
12.
▲
by
wolttam
21d ago
It depends how you use the models. These small models work great for developers who prefer to stay more in the loop, and only task the model with things that can really only be interpreted in one way. Not to mention, they’re great for self-
13.
▲
by
wolttam
21d ago
Hopefully you also bought a switch
14.
▲
by
wolttam
21d ago
The recent and slightly smaller DSv4 Flash is also GLM 5.2 equivalent (or close enough)
15.
▲
by
wolttam
24d ago
HN will automatically do that to your title
16.
▲
by
wolttam
25d ago
The model has zero awareness of MCP, it’s the harness’ job to talk to the MCP server and simply present the model with the tools just like any other tool. The only giveaway to the model about where the tools come from is the ‘mcp__’ prefix
17.
▲
by
wolttam
26d ago
I think I might have been a bit heavy handed in my comment as I was rebutting the idea that only tech people can tell. I suspect it helps to have been exposed to a lot of earlier model writing, which was even more sloppy and had more of the
18.
▲
by
wolttam
26d ago
Neither. Performance of all models is incredibly spikey.
19.
▲
by
wolttam
26d ago
Certainly not. Anybody who cares about language to any reasonable degree surely notices and is repulsed by heavily AI generated content.
20.
▲
by
wolttam
26d ago
No one’s mentioned robots, so… robots. VLA models, etc.
21.
▲
by
wolttam
27d ago
Go ask a (capable, preferably open) LLM. If it doesn’t give you a direct answer it will give you enough of a starting point to start asking more questions. (This is not financial advice)
22.
▲
by
wolttam
27d ago
They are different categories. Brain rot isn’t an imagined thing and we aren’t confused about where it comes from
23.
▲
by
wolttam
27d ago
This is probably getting close to the right interface for these things, and that's coming from someone who is religious about using terminals, neovim, etc.
24.
▲
by
wolttam
28d ago
It’s also exactly what the tech enables. It’s not hard to imagine there being hundreds of thousands of different harness projects, if not more.
25.
▲
by
wolttam
28d ago
Pretty big self-own
26.
▲
by
wolttam
28d ago
That's just it - most use-cases are pretty basic, and if you don’t know what you need then Postgres is probably a great place to start. If you’re just starting out, keep things simple. Otherwise, you probably already know exactly why y
27.
▲
by
wolttam
29d ago
I would have thought the Amazon tax is “free shipping” that has altered the pricing of every item, and now B&M retail is picking up the same pricing despite a much lower shipping overhead Ads too, I guess
28.
▲
by
wolttam
1mo ago
+ issues + pull requests. There’s some nice to haves in there, that you’d need some other service or convention to manage
29.
▲
by
wolttam
1mo ago
I’m getting good at reading LLM-speak (which is probably sad) It means two “layers” of information access/validation. Sounds like it first made a little map of model <-> benchmark, then went and filled in the score boxes. Definit
30.
▲
by
wolttam
1mo ago
I think an argument can be made that some software has gotten more secure. I don’t think it can be denied that the models do show an ability to find security vulnerabilities that may have otherwise been missed
More ›