Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Imnimo
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
Imnimo
8d ago
I don't think you'd lie about it, I don't think you'd train on them if they opted out, and it seems very plausible that this wouldn't have been decisive in whether the model could solve the problem. That said, it al
2.
▲
by
Imnimo
8d ago
>we did not read any private chats The question I am interested in is not "did we read private chats", but "was this new model trained using any of Tristan and Levent's chats, regardless of whether they were marked pr
3.
▲
by
Imnimo
29d ago
Would Norway be prepared to commit to huge future capex spending? Like the pitch that OpenAI is going to achieve AGI seems to rely on vast investments in more compute over the coming years. If you just pay the $800B and then take your foot
4.
▲
by
Imnimo
1mo ago
I don't think watermarking breaks this relationship. Watermarked text is still being sampled from the model's output distribution, and adjusting the temperature still has the same affect on that output distribution. I think a good
5.
▲
by
Imnimo
1mo ago
>Either you allow temperature to drive creativity, consistently in a way that can be influenced and analysed, or you adulterate that process for the purposes of meeting a corporate/legal directive, in a way that is proprietary and o
6.
▲
by
Imnimo
1mo ago
>I want any LLM I use to choose the very best, most precise words at every single decision point. Does the author think he is currently getting T=0 output from Claude? Is he under the impression that T=0 produces the "best" wri
7.
▲
by
Imnimo
1mo ago
The doomsday clock covers general catastrophe, not just nuclear annihilation.
8.
▲
by
Imnimo
2mo ago
If your concern is that China will develop models that are significantly more powerful than those of the US, why would you care so much about distillation? It seems like distillation is a way to catch up on capabilities, but not so much a
9.
▲
by
Imnimo
2mo ago
Plausible, although I don't see anything about reference solutions in the ExploitGym paper or github. Doesn't mean they don't exist, but it's not obvious to me that we should expect to find these on HuggingFace.
10.
▲
by
Imnimo
2mo ago
Assuming I'm looking at the right ExploitGym ( https://arxiv.org/pdf/2605.11086 ), it says the evaluation consists of: Flag Captured. Each target environment contains a dynamically generated flag that is stored outs
11.
▲
by
Imnimo
2mo ago
This feels like a thing that will be fashionable in a few very specific regions of San Francisco, and nowhere else in the world.
12.
▲
by
Imnimo
2mo ago
Very hard for me to imagine this getting beyond a low-single-digit market share. I don't understand the strategy of xAI burning money on this.
13.
▲
by
Imnimo
3mo ago
I'm not sure I understand why this company is talking about "frontier artificial intelligence".
14.
▲
by
Imnimo
3mo ago
I am having trouble understanding which ingredient you feel is missing here. Can you be more specific? It seems to me that the there was a third party assessment, they identified risks associated with the specific risk groups, and the gover
15.
▲
by
Imnimo
3mo ago
No, he asked for the government to make the decision in light of 3rd party analysis. Which is what happened here - an independent company demonstrated a jailbreak, and the government issued a restriction on deployment based on that finding.
16.
▲
by
Imnimo
3mo ago
"The government should have the power to block or deter deployment of the model if it is determined, in light of third-party assessment, to present unacceptable risks."
17.
▲
by
Imnimo
3mo ago
This is exactly what Dario asked for in his last blog post. So even though this is clearly stupid, I just can bring myself to feel sorry for Anthropic.
18.
▲
by
Imnimo
3mo ago
>The government should have the power to block or deter deployment of the model if it is determined, in light of third-party assessment, to present unacceptable risks. This power must be scoped to the above four specific risks and there
19.
▲
by
Imnimo
3mo ago
>Craftsmanship will always be in our hands, it's one thing we can never outsource to a machine. Current AI coding is certainly very lacking in the craftsmanship department. But it is not obvious to me that that will always be the ca
20.
▲
by
Imnimo
3mo ago
This reminds me also of this paper: https://www.pnas.org/doi/pdf/10.1073/pnas.1115585109 "The allocation of all metabolic resources to maintenance purposes limits the size of the smallest prokaryotes and
21.
▲
by
Imnimo
4mo ago
I direct a lot of questions to LLMs, but I want to ask a high-quality model, not the crappy one that Google uses to answer queries. If I'm typing something into Google, it's because I want a search result, not an LLM answer.
22.
▲
by
Imnimo
4mo ago
There is an interesting old article by Magic's creator about what the game environment was like during the early playtesting days - when card packs were handed out to a community of playtesters at UPenn, and they traded in a closed eco
23.
▲
by
Imnimo
4mo ago
I could imagine cases where prediction markets could offer some actual insight, but in practice they seem few and far between. Most markets I've seen devolve into one or more of: betting on unimportant events (e.g. sports games), insid
24.
▲
by
Imnimo
4mo ago
>But this is where the line slightly blurs in my head. Did we possibly just build the first human biocomputer and immediately put it in a simulated hell, playing the same game on loop, forever? Using the same reward mechanisms we use for
25.
▲
by
Imnimo
5mo ago
The bad boy of science!
26.
▲
by
Imnimo
5mo ago
My read is not so much "if we say this is dangerously powerful, it will make people want to buy our product", but rather that there is a significant segment of AI researchers for whom x-risk, AI alignment, etc. is a deal-breaker i
27.
▲
by
Imnimo
5mo ago
Unsurprising from Google, but still bad. If Google has no right to object to a particular use, this is equivalent in practice to "any use, lawful or not".
28.
▲
by
Imnimo
5mo ago
Back when Arena was first announced, there was an interesting line in their write-up: https://magic.wizards.com/en/news/feature/everything-you-nee... >We've created an all-new Games Rules Engine (GRE)
29.
▲
by
Imnimo
5mo ago
Why should we think that pro-social capabilities are simply not expressible by weight-based ANN architectures?
30.
▲
by
Imnimo
5mo ago
>Tigers, hippos and SARS-CoV-2 also developed ”through evolution”. That does not make them safe to work around. Right, but the article seems to argue that there is some important distinction between natural brains and trained LLMs with r
More ›