Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jug
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
1.
▲
by
jug
13d ago
And without this harness it scores about 62%+, a dramatic improvement over even Fable 5.1 at 30%. I thought it just bears saying for context.
2.
▲
by
jug
27d ago
Political? It's using a Chinese LLM trait. If it only refused to talk about a particular kind of soccer, we'd probe it with that instead. The goal is not to discuss politics, the goal is to find out which model it is. What's
3.
▲
by
jug
28d ago
From what I'm seeing I do believe TikTok is worse.
4.
▲
by
jug
1mo ago
I really like the combo 5.6 Luna & Sol for price and performance and would be perfectly happy if they stayed here for a moment without mucking about with sidegrades that I think AI evolution has often felt like lately.
5.
▲
by
jug
1mo ago
This is true and is only becoming more important the more they improve. I am already moving to checking so they're at least somewhat following the status quo and otherwise prioritizing price and platform. I think this will be an emergi
6.
▲
by
jug
2mo ago
If that's marketed well it feels like it should cause a system shock like R1 did. It would also be interesting to see the reaction with code models becoming so good already i.e. cost efficient models aren't necessarily invalidated
7.
▲
by
jug
2mo ago
Yes, I've seen this too and how Luna xhigh is so good that Terra doesn't really serve a purpose because beyond that you can continue at Sol medium. This can be the most cost efficient way, and especially now!
8.
▲
by
jug
2mo ago
Kimi K3 is fairly cheap per token but thinks like a madman with poor self esteem.
9.
▲
by
jug
2mo ago
So is this how Opus 5 ran it? https://arcprize.org/results/anthropic-claude-opus-5
10.
▲
by
jug
2mo ago
In my opinion, next step is to cut down on reasoning tokens while maintaining intelligence. The Chain of Thought and looping can still be an issue with these Chinese models. They in fact said K3 would improve in the area but it's still
11.
▲
by
jug
2mo ago
I agree and this is why I think open models will win in the end. There is just so much to gain on being 10% behind the curve. Especially when the curve is far beyond your needs.
12.
▲
by
jug
2mo ago
This will also make it harder to compete because it holds true for everyone, not just you. I am already, this year, seeing vibe coders put out some decent stuff on App Stores but the problem is marketing it. You no longer automatically stan
13.
▲
by
jug
2mo ago
Yeah I've noted this behavior with best in class open weight models. They said K3 would have token efficiency improvements and I was hoping especially solving the thinking loop issue that plagued K2.x but even if this release helped so
14.
▲
by
jug
2mo ago
I think they're less and less advertised as true generalists these days, as they pivot to profits that obviously lie (for the time being) first and foremost in agentic coding. It's no longer unusual to see regressions in terms of
15.
▲
by
jug
2mo ago
Addiction due to the dopamine hits of occasional struggle and then churning out apps that work: https://leaddev.com/ai/ai-coding-is-addictive-engineers-are-...
16.
▲
by
jug
2mo ago
This article is also related to exhausting AI through generating pressure and posted here recently: AI coding is addictive. Engineers are paying the price https://leaddev.com/ai/ai-coding-is-addictive-engineers-are-...
17.
▲
by
jug
2mo ago
You probably have it backwards. It's Grok that is shoving right wing ideology down your throat. Research has shown that without specific guidance to otherwise, LLM's tend to be slightly left leaning by default. There are some theo
18.
▲
by
jug
2mo ago
Yet, in a month we'll be fine. We were fine with Anthropic naming models by music. I'm sure celestial bodies will be OK too. Larger = better. It's simple. As for the why? Marketing, making products feel "fresh", exc
19.
▲
by
jug
2mo ago
Free tier of Google Gemini can summarize and let you ask questions about pasted YT links.
20.
▲
by
jug
3mo ago
Alternative 1 isn’t all that unlikely given Opus 4.8 couldn’t do this. So it’s a recently possible hack. Not something LLM corps have been blindsided by for years. I also strongly recommend RTFA in this case, namely ”The honest part, read b
21.
▲
by
jug
3mo ago
I often feel like we're nowadays mostly pushing AI developments in the ways of finetuning differences. Like how new editions of Claude are tuned for agentic coding which might even be detrimental if you're using it for non-agentic
22.
▲
by
jug
4mo ago
”If you don't cannibalize yourself, someone else will." — Steve Jobs
23.
▲
by
jug
4mo ago
Looks like an ongoing theme and a very poor benchmark. Not at all the claims I expected.
24.
▲
by
jug
4mo ago
It's also very surprising to me. This whole deal where humans instantly started taking AI answers at face value, as sources standing on their own legs, or delegating their own mind to a third party, not even a human, but an algorithm.
25.
▲
by
jug
4mo ago
While oldest source of it, note that the 86-DOS v0.1-C binaries are even earlier (and v0.34 has also been found) than this v1.00 source and can be downloaded and used in an emulator. :-) https://arstechnica.com/gadgets/
26.
▲
by
jug
4mo ago
This is a risk although then this is fortunately a model that isn't tied to Chinese hosting. But indeed something to consider if using straight DeepSeek.com.
27.
▲
by
jug
4mo ago
I found this thought provoking and just had to see how the new Gemini 3.5 Flash reasoned about this (I find it fun to go meta on modern AI like this), and I'm happy that I did! Also as an opportunity to trial this recent model. https:
28.
▲
by
jug
4mo ago
I think that's what the Omniscience Index is for: https://artificialanalysis.ai/evaluations/omniscience#aa-omn... It rewards correct answers and penalizes hallucinations, and finally no reward for refusing to answ
29.
▲
by
jug
4mo ago
They have now been released on e.g Hugging Face with model suffixes "-assistant".
30.
▲
by
jug
5mo ago
Shouldn't one use e.g a Wolfram Alpha MCP endpoint for math in AI? From what I've seen on even premium non-quantized models, I would never ever trust the innate ability of a LLM to calculate.
More ›