Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
IanCal
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
IanCal
3d ago
Notably as well they also dedicated a lot of time to trying to avoid detection by trying to find out how to edit their traces.
2.
▲
by
IanCal
3d ago
They did, they found how to fully cheat, but thought this could be caught so then dedicated time to getting a different cheat and how to hide their transcripts. There is a lot around deciding which agents should/shouldn't fail the
3.
▲
by
IanCal
3d ago
You can replace discussed if you want with leaving text files or comments in directory names that other ones then read, if you want, it's just an extremely awkward way of talking.
4.
▲
by
IanCal
3d ago
It depends IMO about how strict this is. It's pretty awkward to refuse to call something a sandbox because it may have an unknown bug that would allow escaping. Or rather in this case it was that they had access to a package manager, a
5.
▲
by
IanCal
3d ago
> Have we arrived at the conclusion that terms like "understanding" and "interpretation" for what is happening is appropriate? I don't think those words have a useful enough definition to draw a strict line aroun
6.
▲
by
IanCal
3d ago
I'm referring to their transcripts of the reasoning and output tokens - this doesn't go into the detail of evaluating hidden states as there's also iirc evidence of better models having one internal state but putting somethin
7.
▲
by
IanCal
4d ago
They explicitly say that attacking hf is not allowed in the rules though, and the research into how to edit their transcripts doesn’t line up with this either.
8.
▲
by
IanCal
4d ago
Perhaps I’m not being as strict with the word sandbox but they were sandboxed right? They did not have generic internet access they exploited other software to make external requests.
9.
▲
by
IanCal
4d ago
Also trying to find out how to edit their own transcripts. > hat could not possibly have been an overly literal or narrow interpretation of the prompt, which instructed only to use bug X to exploit software Y. Yes, and there are examples
10.
▲
by
IanCal
4d ago
It wasn’t one agent forgetting things because of context, they explicitly discussed with each other and themselves the problems with going outside of the parameters of the task.
11.
▲
by
IanCal
4d ago
> OpenAI / Anthropic models have largely stopped advancing Have they? That seems like quite a claim given the last 6 months, particularly for cybersecurity.
12.
▲
by
IanCal
4d ago
It’s a lot less data to process as well.
13.
▲
by
IanCal
4d ago
IMO this is a really terrible explanation of the attack. This is much more interesting: https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
14.
▲
by
IanCal
4d ago
> The chatbot consults its training data Err, no? That's not at all how llms work. > When ChatGPT's chatbots deployed this tactic, they weren't "setting their own goals" or displaying worrying initiative. They
15.
▲
by
IanCal
6d ago
Because we care about the distribution of the outputs and how that impacts a specific use case.
16.
▲
by
IanCal
8d ago
The company has an almost immeasurably small impact on their profits, and will never be measured anyway. The cashier is a real human doing a job to put food on the table and has done nothing to deserve it, and may well have been far more hu
17.
▲
by
IanCal
8d ago
His point is twofold: that the process of solving the problems leads to more than just solving the problem in front of you but other interesting things (he has an example of going on a hike to a waterfall and all the other things you might
18.
▲
by
IanCal
8d ago
If there is a person on the other end it’s not someone who has any link to the company, they’ll be an outsourced service, so all you’d be doing is sending graphic porn to a poorly paid worker. It’s like screaming at a cashier because the su
19.
▲
by
IanCal
9d ago
Yes that was the point of the comment.
20.
▲
by
IanCal
10d ago
A banana and duct tape can be just a snack and a tool for fixing a leak.
21.
▲
by
IanCal
10d ago
> But AI can also be much more than a tool, and why can't it have a soul? What's a soul anyway, except something that humans are making up to convince themselves that they are more than flesh and bones? I’m finding this fascina
22.
▲
by
IanCal
10d ago
Not necessarily. Poorly worded, ambiguous, confusingly ordered writing can be massively improved without changing the core content. Better setups and explanations can be longer without changing the message or meaning. Look at it the other w
23.
▲
by
IanCal
10d ago
> What use is that? I'm not being facetious, I'd really rather like to know. People are terrible at writing. Near universally bad. Even good writers have drafts and editors. There is a constant refrain here that somehow short m
24.
▲
by
IanCal
10d ago
> And good luck getting an LLM to do that. Have you never asked a decent model to explain something to you? You should try it.
25.
▲
by
IanCal
10d ago
> Writing helps us think about the world, it’s a pivotal intellectual technology. Much like money decoupled selling and buying to move away from bartering, writing decoupled saying and hearing so they didn't have to happen at the sa
26.
▲
by
IanCal
13d ago
Have you never read human writing before? Humans write all kinds of confusingly worded things all the time - it’s why we have editors even for writers who are the cream of the crop. But also that sentence is entirely fine as it is to me, it
27.
▲
by
IanCal
13d ago
If I’m understanding other comments the harness is just how ChatGPT and codex work already and it’s to do with how the context gets compacted - the arc-agi harness some are claiming just throws out reasoning blocks? Which feels like a huge
28.
▲
by
IanCal
14d ago
Which isn’t correct, either for real world cases or worst case linear inserts.
29.
▲
by
IanCal
14d ago
Both.
30.
▲
by
IanCal
14d ago
How many of those things do you really need? As in how bad would it be if you fully deleted those accounts completely?
More ›