Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
puppystench
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
puppystench
5mo ago
They ran a bounty on Kaggle last year but with $500k in payouts and with all results open and publishable. https://www.kaggle.com/competitions/openai-gpt-oss-20b-red-t... With only $25k in payouts and everything locked
2.
▲
by
puppystench
5mo ago
The Claude UI still only has "adaptive" reasoning for Opus 4.7, making it functionally useless for scientific/coding work compared to older models (as Opus 4.7 will randomly stop reasoning after a few turns, even when prompte
3.
▲
by
puppystench
5mo ago
In the announcement webpage: >For API developers, gpt-5.5 will soon be available in the Responses and Chat Completions APIs at $5 per 1M input tokens and $30 per 1M output tokens, with a 1M context window.
4.
▲
by
puppystench
5mo ago
For API usage, GPT-5.5 is 2x the price of GPT-5.4, ~4x the price of GPT-5.1, and ~10x the price of Kimi-2.6. Unfortunately I think the lesson they took from Anthropic is that devs get really reliant and even addicted on coding agents, and t
5.
▲
by
puppystench
5mo ago
Does this mean Claude no longer outputs the full raw reasoning, only summaries? At one point, exposing the LLM's full CoT was considered a core safety tenet.
6.
▲
by
puppystench
5mo ago
I believe you're right, it's an issue of the model misinterpreting things that sound like user message as actual user messages. It's a known phenomenon: https://arxiv.org/abs/2603.12277
7.
▲
by
puppystench
5mo ago
>Several people questioned whether this is actually a harness bug like I assumed, as people have reported similar issues using other interfaces and models, including chatgpt.com. One pattern does seem to be that it happens in the so-call