Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
anon373839
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
anon373839
3d ago
I will say that LLMs are somewhat unlike guns in that they emit text.
2.
▲
by
anon373839
3d ago
Privacy is a great reason, but independence is another. It’s very nice knowing that you’re going to get the same reliable product every time you call the model. Nothing is going to change unless you decide to change it.
3.
▲
by
anon373839
3d ago
It’s absolutely laughable that he refers to METR as if they were neutral observers. They are ex-Anthropic employees and others with direct financial interests in Anthropic.
4.
▲
by
anon373839
4d ago
It is costly, especially right now. I don’t think you can make a case for it on cost savings! The throughput in a single stream is about 50 tokens/sec (a bit less for prose, a bit more for code due to speculative draft acceptance rate
5.
▲
by
anon373839
4d ago
> If someone put in frontier AI models from like .... last june I guess? in a box and let me run it with "decent" token throughput I would be happy. You can have that! Qwen 3.8 Flash-Next is ~Opus 4.6 and runs nicely on a DGX S
6.
▲
by
anon373839
4d ago
No, it’s not the active parameters. Qwen 3.8 Flash has 6B active and it smokes both models.
7.
▲
ChatGPT: "Super shady" data controls settings
(twitter.com)
2 points
by
anon373839
7d ago
|
1 comments
8.
▲
by
anon373839
7d ago
https://xcancel.com/edoardocontente/status/20976073405059281...
9.
▲
by
anon373839
8d ago
I don’t use ChatGPT anymore, but when I did, the training opt-out setting was frequently silently resetting itself to off.
10.
▲
by
anon373839
8d ago
> GLM 5.3-flash released last week, and that means Project Glasswing and Daybreak are running out of time. Cheap models capable of dangerous hacking are now available to anyone, without the normal safeguards for refusing malicious action
11.
▲
by
anon373839
11d ago
Claude is different from S3. AWS doesn’t need to rifle through your files to stay ahead of the competition or to mine them for business ideas because the core business is overvalued and rapidly commoditizing. AI labs, on the other hand, hav
12.
▲
by
anon373839
12d ago
Hallucinations are very damaging to a model’s utility. But doesn’t the Omniscience Index focus on knowledge-based queries? To me, using LLMs for their memorized knowledge is very 2023 and suboptimal. IMO, what really makes a model useful is
13.
▲
by
anon373839
12d ago
Jensen Huang?
14.
▲
by
anon373839
13d ago
Ah, I see Anthropic is back at that “saving humanity” game…
15.
▲
by
anon373839
13d ago
Sebastian Raschka posted about this architecture: > A lot of hype around OpenAI's Astra model here on my timeline today. Apparently, this goes back to a new article from The Information, which said Astra is a "recurrent depth o
16.
▲
by
anon373839
15d ago
Apt.
17.
▲
by
anon373839
15d ago
Macs have excellent generation speed, and the new Ultra will positively smash that at 1.2TB/sec of bandwidth. For example, that new 176B parameter Qwen model would generate tokens at ~200 tokens/sec. Macs don’t have very good pref
18.
▲
by
anon373839
17d ago
> The idea of releasing a frontier model without RL is frightening In case you were not aware, strong base models (no post training at all) have been available for quite some time now. Including ones that eclipse “scary” frontier models
19.
▲
by
anon373839
19d ago
You and I have the same machine. Do you mind sharing the model ID you're using? I'm on oMLX also but I haven't seen anything above ~20 tok/s out of 3.8 27B, even with MTP and generating code.
20.
▲
by
anon373839
20d ago
Yes; my point was not related to any of that.
21.
▲
by
anon373839
20d ago
Route A is good for rapid, disposable prototyping. If you ever have a stray thought, “I wonder how this would work if the whole paradigm were turned sideways”, you now have a chance to preview a “working” version of your idea. If you like i
22.
▲
by
anon373839
20d ago
> The bitter lesson is about hand tuned AI vs computational general methods. However in truth today’s AI uses both. We have general compute heavy models which require narrow expert instructions The models are not even really trained bitt
23.
▲
by
anon373839
21d ago
Does anyone have an idea how this might perform on a DGX Spark at longer contexts? I've been trying to investigate their performance with these medium-sized MoE models, but I'm seeing a lot of incomplete and conflicting informatio
24.
▲
by
anon373839
21d ago
That's an interesting paper, but there is virtually no discussion of reasoning behaviors or optimization for long-horizon tasks (i.e., all of the recent advances in LLMs that people care about). The evaluation methodology also is prett
25.
▲
by
anon373839
21d ago
How about proof that black-box distillation can deliver these results without a very sophisticated RL pipeline doing the heavy lifting?
26.
▲
by
anon373839
23d ago
It's fun to imagine that it could be GLM 5.3-Flash. Between GLM 4 and 5, the flagship's total parameters doubled and the active parameters went up 25%. GLM 4.7-Flash was 30B / 3B active. If this model were 60B / 4B activ
27.
▲
by
anon373839
23d ago
Agreed. I think LLMs are best used as pair programmers or typists for users who already know what they’re doing. Or as tutors for users who want to learn. Vibe coding is mostly garbage. But it can be useful for creating instant, disposable
28.
▲
by
anon373839
23d ago
This is such a tough problem. Anthropic would need access to some kind of technology that could, like, intelligently handle unforeseen circumstances and nuances. Yeah, that’s definitely not something we should expect of them.
29.
▲
by
anon373839
24d ago
The issue is that, if there is a “good enough” point approximately here, it is only a matter of time before models become small and efficient enough not to need all those data centers. Though, it should be good for companies that sell compu
30.
▲
by
anon373839
25d ago
GPT-OSS 20B didn’t really merit the fanfare even when it was released; it’s definitely not competitive now. Even the 120B version has been well eclipsed by smaller LLMs at this point. The last version of Qwen 27B/35B was better, and no
More ›