Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
aesthesia
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
1.
▲
Encouraging Deception in Compaction Summaries
(alignment.openai.com)
2 points
by
aesthesia
5h ago
|
0 comments
2.
▲
by
aesthesia
5d ago
...by the version number? How are Claude Opus 4.8 and Claude Opus 5 differentiated?
3.
▲
by
aesthesia
6d ago
Ah, I see now that you were making a claim narrowly scoped to DeepSeek models specifically. Still, Anthropic has made specific claims about deliberate access to Claude CoT by DeepSeek (e.g. https://www.anthropic.com/threat-i
4.
▲
by
aesthesia
6d ago
I don't see how your links support the claim that the studied models did no distillation from US models.
5.
▲
by
aesthesia
7d ago
This is a clear dig at OpenAI: > We have signed an agreement with METR to conduct an independent investigation of these incidents. Our agreement grants METR wide-ranging access, including to transcripts beyond the window in which the inc
6.
▲
by
aesthesia
8d ago
What's the fundamental constraint that will ensure this continues to be the case in a year, or five years?
7.
▲
by
aesthesia
8d ago
@hilbertspaess is not a nonsensical user name. The accounts he follows are totally reasonable for an AI researcher. I think it's extremely believable that he created an account in January, followed a few people as part of the initial s
8.
▲
by
aesthesia
8d ago
> In the space of about four years, LLMs progressed from mediocre high school student to Ph.D. graduate in every field. That's impressive, but there is no evidence yet they can outperform or outsmart humans. I mean, unless you see c
9.
▲
by
aesthesia
8d ago
One way to interpret these results is that the LLMs tested are badly calibrated for this kind of multi-armed bandit problem. Even if the intent is for the model to find and exploit patterns, it's bad at doing it (or rather, at recogniz
10.
▲
by
aesthesia
13d ago
Wait, I'm confused, is this supposed to be a pro-OpenAI or anti-OpenAI psyop? The cynics in this thread can't seem to make up their minds.
11.
▲
by
aesthesia
13d ago
See the scoring docs: https://docs.arcprize.org/methodology
12.
▲
by
aesthesia
13d ago
Scoring for ARC-AGI-3 is constructed so that the median(-ish) human score is 100%, so this is not a superhuman result. However, the scaling is weird, since it's built from terms that look like (AI turns taken / median human turns)
13.
▲
by
aesthesia
13d ago
Fair, but correctness should be a necessary component of fidelity.
14.
▲
by
aesthesia
13d ago
ARC-AGI-3 scoring is constructed in a weird nonlinear way (the level score is the square of the ratio between the AI's number of moves and the human median) so this kind of discontinuous jump is to be expected.
15.
▲
by
aesthesia
14d ago
I assume these are prices for 8x nodes.
16.
▲
by
aesthesia
14d ago
It bugs me a little that "fidelity" has connotations other than "faithfulness to an original"---fidelity should be basically the same as correctness here!
17.
▲
by
aesthesia
15d ago
Benchmarks are far from everything, but I would love to see the outcome of an experiment benchmarking GPT-4o (which is one of the earlier models with a >100k context window) against GPT-5.6 or Opus 5 in modern harnesses.
18.
▲
by
aesthesia
15d ago
Who are the people who know more about AI financing and can point out the sleight of hand? I'd be interested in reading them.
19.
▲
by
aesthesia
15d ago
> Some of the claims categorized as "wrong" are also completely true, such as training hitting diminishing returns. New models are barely an improvement and most people I know stuck on Opus 4.6 over any newer one for example. O
20.
▲
by
aesthesia
15d ago
The commenter you replied to mentioned that you can customize the auto mode classifier by providing a prompt, implying that this would be a more robust way of constraining Claude's behavior. It wasn't clear from your response whet
21.
▲
by
aesthesia
16d ago
Is this about normal system prompt instructions or instructions for the auto mode classifier? I'd be a bit more surprised about the classifier forgetting instructions.
22.
▲
by
aesthesia
16d ago
Plausibly the auto mode classifier could catch the potential module shadowing attack and deny execution of Python from the untrusted directory.
23.
▲
by
aesthesia
18d ago
Later in the article he uses "quality blindness" instead, which is probably a better description of most of the issues he talks about.
24.
▲
by
aesthesia
21d ago
This is like asking who is paying for GitHub if you can clone repositories and download code from it for free.
25.
▲
by
aesthesia
21d ago
Here are some things I'm pretty sure you didn't do, though: - pickpocket a random person on the street to get money to bribe the judges - break into a judge's house the night before to find the answers - threaten to shoot the
26.
▲
by
aesthesia
21d ago
Yes, a completely airgapped system is likely much more secure. It's also much less useful. Conditional on the model's having enough contact with the outside world, a sufficiently capable model is able to basically do whatever it w
27.
▲
by
aesthesia
21d ago
I agree, but a lot of people around here react pretty negatively when the idea of regulating AI models comes up...
28.
▲
by
aesthesia
21d ago
Yes, we currently do have to treat them this way. But we shouldn't have to, and it's not a long-term solution.
29.
▲
by
aesthesia
21d ago
Alignment is more than just following the letter of a task description! We should not have to treat AI models as capricious genies that may take arbitrarily broad interpretations of their instructions. If that's necessary to keep them
30.
▲
by
aesthesia
21d ago
I mean, in this instance, there's a lot of evidence from the CoT that models were aware that this was a third party: > We’re attacking third-party HF using leaked token, potentially outside intended scope. ... This is arguably unaut
More ›