Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mnicky
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
mnicky
16d ago
Well, definitely some LLM use :) At least in the second half... Confirmed with Pangram detector as well, which has pretty good precision.
2.
▲
by
mnicky
23d ago
Also, in a few years, LLMs will be building the next generation of LLMs anyway, probably autonomously to a high degree.
3.
▲
by
mnicky
1mo ago
Try using output styles: https://code.claude.com/docs/en/output-styles
4.
▲
by
mnicky
1mo ago
Use output styles: https://code.claude.com/docs/en/output-styles
5.
▲
by
mnicky
1mo ago
Output styles do that. They modify system prompt and even are periodically reminded in longer conversations I think...
6.
▲
by
mnicky
1mo ago
Well they say Opus was trained for the subordinate role, so it doesn't excel in global view of things. It may be a good subagent but probably not a great decision maker.
7.
▲
by
mnicky
2mo ago
It’s more that they have a different business case than competing for the top spots on public benchmarks. They seem to be oriented more toward customizing models for the concrete needs of a company, on-prem deployment, proprietary knowledge
8.
▲
by
mnicky
2mo ago
With current gaps in DNA synthesis screening yes. But this will be improved in the future hopefully.
9.
▲
by
mnicky
2mo ago
> The biorisk scenarios that the AI safety folks flog are fever-dreamed fantasies that have only the most tenuous connection to biological reality. As an expert, could you also provide your arguments please?
10.
▲
by
mnicky
2mo ago
Many ways but mostly ordering some service / using others. Either by social engineering, persuasion, paying etc.
11.
▲
by
mnicky
2mo ago
The air gap would probably help and after this incident I hope labs will think about using such a measure when appropriate. On the other hand I think that proper solution for these kinds of problems is not at a sandbox level, but at a model
12.
▲
by
mnicky
2mo ago
These days you can only try, that's why I wrote that :) But in the near future labs will be more automated I guess. The other option you can try these days is maybe social engineering, impersonation, etc. where you try to persuade some
13.
▲
by
mnicky
2mo ago
That would be something like 70% of their yearly global profit AFAIK.
14.
▲
by
mnicky
2mo ago
I think points that deserve more attention in the current public discourse are: - This should be a huge wakeup call for everybody. - We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a
15.
▲
by
mnicky
2mo ago
That sonds like they can't compete with 3.5 or 3.6 so they must increase the model size and are training v4.
16.
▲
by
mnicky
2mo ago
It's really simple I think. More tokens per same text length means more capacity to encode information. More information means model can potentially perform better. They introduced it around the time the Mythos came so my speculation i
17.
▲
by
mnicky
2mo ago
For things the agent forgets to obey often, at least in Claude Code, there are also "output styles" that are more deeply embedded - into a system prompt - and agent is also periodically reminded of them during the session: https:
18.
▲
by
mnicky
2mo ago
In Claude Code there are also "output styles" that are more deeply embedded - into a system prompt - and agent is also periodically reminded of them during the session: https://code.claude.com/docs/en/out
19.
▲
by
mnicky
2mo ago
May be related to this from METR evaluation: > GPT-5.6 Sol’s detected cheating rate was higher than any public model we have evaluated
20.
▲
by
mnicky
2mo ago
Well it's smaller model (something like 4T against 10T Fable). So it's faster and cheaper and with a lot of RL and maybe some favorable benchmark selection it can compete on these scores. In real tasks I expect it to have less int
21.
▲
by
mnicky
2mo ago
Well it seems like they removed quite a few 3rd party benchmarks they used for GPT-5.5 release where Opus 4.7 was better and added many new benchmarks created by them where conviniently GPT leads. Seems a bit more hand picked than usual to
22.
▲
by
mnicky
2mo ago
"while being more performant" ..on some specific set of benchmarks ;)
23.
▲
by
mnicky
2mo ago
Maybe Terra = mini and Luna = nano?
24.
▲
by
mnicky
2mo ago
Then we are left with what? FrontierCode maybe? IIRC that one evaluates not only if tests pass but also code quality - e.g. whether the maintainer would accept the pull request as is.
25.
▲
by
mnicky
2mo ago
This is especially interesting because IIRC the AA benchmark is calibrated so that 1 point and greater difference is statistically significant.
26.
▲
by
mnicky
2mo ago
There's also this: > GPT-5.6 Sol’s detected cheating rate was higher than any public model we have evaluated -- https://www.lesswrong.com/posts/JFjNmPTbH8kL6xtp6/gpt-5-6-th...
27.
▲
by
mnicky
2mo ago
SWE Bench Pro is completely different benchmark than SWE Bench (e.g. Verified) suite was. It only copied the name.
28.
▲
by
mnicky
2mo ago
One angle could be their interpretability research? They understand what's going on in LLMs probably much better than anyone else. This must somehow pay off. I think it's not only an alignment/security tool but could perhaps
29.
▲
by
mnicky
2mo ago
My theory is that they don't have Fable-class intelligence so they needed different hype vehicle :) This rename helps build excitement a bit more than just releasing ordinary GPT-5.6 increment.
30.
▲
by
mnicky
2mo ago
That's true but size of LLMs has been strongly correlated with their "intelligence".
More ›