Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Bolwin
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
1.
▲
by
Bolwin
6d ago
Distillation requires you to have the actual logits of each token from the teacher model, which in practice means having the model itself. What you're describing is just synthetic data. Note Anthropic misused the term in their post abo
2.
▲
by
Bolwin
7d ago
I don't think you know what distill means
3.
▲
by
Bolwin
7d ago
This and their newer Aura A1 I've wanted to get cause of the compact size, but they seem to lack support for US 5g bands
4.
▲
by
Bolwin
10d ago
As someone who's done a lot of llm fiction, that reads as pretty typical slop, and very human centered, nothing like a museum
5.
▲
by
Bolwin
14d ago
I highly doubt it's behind in practice, except for Anthropic
6.
▲
by
Bolwin
18d ago
Cache hit % on openrouter is not a good metric, it's mainly driven by openrouter's own provider juggling than the providers themselves
7.
▲
by
Bolwin
23d ago
LLMs overuse em-dashes and use them in specific ways that are annoying to read. Stringing together clauses unnecessarily, spacing words out instead of using a comma etc. I don't kind em-dashes in decent human writing, they're plea
8.
▲
by
Bolwin
24d ago
Glm had made vision models in the past. Look up GLM 5v. The only question now is if it's 5.3v, 5.4/5.5 or a dedicated flash/vision model
9.
▲
by
Bolwin
24d ago
They already do, extensively. Its not Shakespeare, in fact, it sucks at prose and creativity, like most newer llms. But people are not as alike as you think. I doubt I share your unique preferences. That said, I don't spend much time t
10.
▲
by
Bolwin
1mo ago
It's an opt in feature in openrouter to get a 1% discount.
11.
▲
by
Bolwin
1mo ago
Love the X logo that goes to bluesky
12.
▲
by
Bolwin
1mo ago
Might be in part because multiple times I've seen signs/poster etc only good the agent to contradict it once you get there, adding or removing requirements. So you stop trusting it. A whiteboard I might trust because it was likely
13.
▲
by
Bolwin
1mo ago
It doesn't. It's called preserved reasoning and every recent reasoning model does it
14.
▲
by
Bolwin
1mo ago
More than 8 primary schools in a small town seems a lot no?
15.
▲
by
Bolwin
1mo ago
Yeah and it's degraded significantly since then. Older llms were still mostly language focused and had a lot of latent knowledge about things like writing styles. Now it's crowded out in favor of agenetic work, programming etc.
16.
▲
by
Bolwin
1mo ago
> Muse Spark 1.2 is available today in Muse Code and in Meta Model API with expanded global access Wasn't the previous one us only? This is probably the biggest part of the post Anyone know if muse code is open source?
17.
▲
by
Bolwin
1mo ago
Two things 1. If we're using native harnesses, I'd have preferred you use kimi code, not opencode 2. The variation in the two kimi providers just shows how you can't trust n = 1 trials
18.
▲
by
Bolwin
2mo ago
What about your examples has an llm tell? I don't trust pangram 4 much. There was a post here earlier confirming it fails for many others.
19.
▲
by
Bolwin
2mo ago
It feels like that should be a golden opportunity for competitors, but every time a competitor makes a decent replacement, the big tech company either buys it or briefly invests in their product again to make it good enough and since they h
20.
▲
by
Bolwin
2mo ago
Why is the title "Africa" and not Morocco
21.
▲
by
Bolwin
2mo ago
Screams it in fact
22.
▲
by
Bolwin
2mo ago
Jeez way to ruin of the few remaining joys of flying. Lock everyone in a tin can. You will experience reality through screens only and you will enjoy it
23.
▲
by
Bolwin
2mo ago
Llms use them a lot more than humans, including this blog post. Like all slop. There's a reason it's called slop and it's not because of restraint
24.
▲
by
Bolwin
2mo ago
I like this. It feels like more a person talking and less like a prepared speech.
25.
▲
by
Bolwin
2mo ago
For a fair comparison, you should compare to K3 (which AA has not tested yet unfortunately) and GPT 5.6 Sol also on medium or the closest equivalent
26.
▲
by
Bolwin
2mo ago
This is a lot of words to say "switch models instead of writing a plan file and starting a new session" which I do anyway. That said, it takes me a while to reads plans and often I'll take a break, by when the cache has expir
27.
▲
by
Bolwin
2mo ago
X.com is not publicly accessible. I wish people would stop using it as a source
28.
▲
by
Bolwin
2mo ago
Every time a competitor for acquired, sold out, or changed business models, Kilo made a snarky post. All companies are the same
29.
▲
by
Bolwin
2mo ago
Where are you getting supposedly. It does worse in most benchmarks
30.
▲
by
Bolwin
2mo ago
Claude is very cache friendly, however there have been some inconsistencies with non anthropic endpoints that led to cache breakages
More ›