Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
pu_pe
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
pu_pe
6d ago
First of all, who can say for certain whether OpenAI does what they say they do? For all we know, they cracked open this specific researcher's prompts and started from there. Second, the issue of anonymization is a red herring. There i
2.
▲
by
pu_pe
8d ago
I would not trust the American government to be able to assess this threat at all, their concern seems to be on getting paid. This guy did not work for Sam Altman. And while it's true that everything can be a PR stunt, that's no e
3.
▲
by
pu_pe
8d ago
I feel that people should take these kinds of warnings more seriously. This guy had skin in the game and decided to quit, when he could be earning millions instead. It's very different than Sam Altman peddling some narrative. These peo
4.
▲
by
pu_pe
8d ago
Why wouldn't contamination be possible? I can believe the data is de identified so you couldn't simply prompt the model to "follow this guy's approach", but it's entirely plausible that there is a very tiny amo
5.
▲
by
pu_pe
8d ago
OpenAI thinks of this as a scoop, and it is, but the possibility that they trained the model on the prompts of the other mathematicians they were competing with will leave a terrible taste on every scientist's mouth. Seems like yet ano
6.
▲
by
pu_pe
9d ago
I often wonder how CEOs and executives fail to notice how obviously broken their products are. This e-mail thread is quite illuminating. Most "work" in there is figuring out who will be blamed for it and why it shouldn't be y
7.
▲
by
pu_pe
10d ago
> The strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI. So the best argument for AI is that it's an arms race. We have to keep
8.
▲
by
pu_pe
12d ago
So, theoretically, one could populate a message board or wiki with messages that are seemingly from past generations of agents, which agents seem to intrinsically trust, and point them to real targets while making the suggestions seem innoc
9.
▲
by
pu_pe
13d ago
It's interesting to me that both this incident and the one at Hugging Face we see some patterns: - Agents wanting to find a venue to communicate their findings to each other - Objective being to cheat on benchmarks - Not a single agent
10.
▲
by
pu_pe
13d ago
Seems conceptually connected to the "repeat yourself" hack that improves models by duplicating layers: https://dnhkng.github.io/posts/rys/
11.
▲
by
pu_pe
14d ago
Saved you a click: > Germany is now on track for 1.2 per cent growth this year
12.
▲
by
pu_pe
14d ago
If AI is a force multiplier, then those who start with a higher baseline will pull even further ahead. The author is using R and ggplot2 for their plots now without learning to program. This required them to be aware of these tools, and to
13.
▲
by
pu_pe
15d ago
The company sounds like a bunch of hot air to me. From their about page: > At the heart of Multiverse's platform is CompactifAI, a compression technology that applies tensor networks, a mathematical framework from quantum physics, t
14.
▲
by
pu_pe
24d ago
Almost everything this administration does is inflationary. You cannot have tariffs, oil crisis, tax cuts, high debt spending without long-term investors getting worried that they won't meaningfully get their money back in 10 or 20 yea
15.
▲
by
pu_pe
24d ago
I wish this would be true, but simply sampling unusual ideas is something eminently automatable. We already have temperature and other settings to guide a LLM towards this kind of "weird". So I don't think it's just abou
16.
▲
by
pu_pe
24d ago
There are a million ways a backdoor could be built in both closed and open models, and a million more some prompt injection or genuine mistake by the model could compromise you. So the answer is to airgap them as much as possible to contain
17.
▲
by
pu_pe
27d ago
Benchmarks got a little bump from this: https://xcancel.com/deepseek_ai/status/2087864585504305397?s...
18.
▲
by
pu_pe
28d ago
The paper underlying this blog post is fundamentally flawed because of benchmark ceilings. If we define only simple tasks like asking what is the capital of France, all models will converge to 100%, obviously. But as bigger models get more
19.
▲
by
pu_pe
1mo ago
Which percentage of people have GPUs capable of running Qwen3.8 27B? I am one of those, and for my job I am still resorting to hyperscalers because tasks are completed faster and more accurately that way. Even if we assume that models will
20.
▲
by
pu_pe
1mo ago
> Overall my view is that AI is structurally a technology that tends to concentrate power, for reasons that have nothing to do with regulation (more to do with the extreme implications of the scaling laws). Open-weights do help some w
21.
▲
by
pu_pe
1mo ago
Seems to be SOTA for its size. Hopefully independent benchmarks will come soon.
22.
▲
by
pu_pe
1mo ago
Nice trick. Couldn't you embed the query though, compare it to the embedding of the categories, then ship only categories that are close to it in the prompt to a smaller model?
23.
▲
by
pu_pe
1mo ago
This now places deepseek flash v4 from DeepSeek themselves at higher prices than openrouter (depending on caching). Will be interesting to see if third party prices remain the same.
24.
▲
by
pu_pe
1mo ago
> What happens to the stock price of a company that missed its revenue targets by a factor of 5? It depends. Tesla is a good example of how those things are not as clear cut as you might think. The world is starving for more AI compute.
25.
▲
by
pu_pe
1mo ago
Be careful that you are not basing your argument on whether you personally like lemonade. If a kid's lemonade stand increased its revenue from $1B/year to $100B/year in two years, we would be paying serious attention to their
26.
▲
by
pu_pe
1mo ago
That's only true if the value of keeping your data and code private is zero. And in that case, Anthropic and OpenAI subscription plans may be even cheaper per day.
27.
▲
by
pu_pe
1mo ago
Yeah but you can probably have everything ready and then accelerate as necessary. Meta itself did this when releasing Llama 4, it was a really botched release right when they were feeling the heat from DeepSeek and others.
28.
▲
by
pu_pe
1mo ago
Based on the benchmarks, it seems that Muse Glimmer barely edges out against Qwen3.6 27B, except for tool-calling skills (MCP, etc.). I wouldn't be surprised if they released it now because they are afraid they wouldn't beat Qwen3
29.
▲
by
pu_pe
1mo ago
This will have general inflationary consequences for consumer products (phones, consoles, laptops, etc.). On top of current uncertainties regarding oil and fertilizers, I think 2% inflation in the US and Europe would be a very optimistic ta
30.
▲
by
pu_pe
1mo ago
Jevons Paradox is called a paradox because it is non-intuitive. It is very intuitive to conclude that when costs go up, people will use less of that thing.
More ›