Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ertgbnm
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
1.
▲
by
ertgbnm
13d ago
This is a very mundane release compared to GPT-4 and GPT-5. I think they probably scaled back a bit after the lukewarm response to the GPT-5 announcement. But it still very weird that there wasn't even a livestream,
2.
▲
by
ertgbnm
13d ago
The boring question is whether there is a claudeflare issue impacting them all. The fun question is whether this a rogue AI taking control of the compute of all of the organizations.
3.
▲
by
ertgbnm
21d ago
Didn't AISI literally report exactly that regarding Claude last month?
4.
▲
by
ertgbnm
1mo ago
It's a failure on all fronts. I think it's an indictment of Flock in the same way that full body scans are a scam. If you are always scanning for illnesses/crimes you will inevitably find tons and tons of false positives with
5.
▲
by
ertgbnm
1mo ago
Yeah if they were so committed to doing good, you'd think the billions of dollars they have spent would bear some evidence of that by now.
6.
▲
by
ertgbnm
1mo ago
I think you nailed it. If the author was a billionaire in the Netherlands he might not notice the differences as they do in the article.
7.
▲
by
ertgbnm
1mo ago
This has been a frequent problem on fiction writing on royal road. Author promises that they don't use AI to write and then proceed to use AI to generate the cover for the series. They might be telling the truth, but I am not going to
8.
▲
by
ertgbnm
2mo ago
I've had the feeling that labs aren't pelicanmaxxing specifically but that they do have some sort of RL environment for SVGs that they are letting the AIs overcook in. Specifically I'm thinking of the gemini 3.1 pro annoucnem
9.
▲
by
ertgbnm
3mo ago
Once you lose, you have lost. Ok, but how does that help us predict when something will lose?
10.
▲
by
ertgbnm
4mo ago
> with AI-generated content excluded from pre-training. > without distillation from third-party models sounds like zero unless they are lying.
11.
▲
by
ertgbnm
4mo ago
My point is that if I made someone "smarter" they wouldn't suddenly know "What day, month, and year was Carrie Underwood’s album “CryPretty” certified Gold by the RIAA?" which is an example of a question in the Simp
12.
▲
by
ertgbnm
4mo ago
Knowledge benchmarks can't really be improved upon via distillation or RL. It requires those facts be added to the training corpus and for the model to memorize them better. Neither distillation or RL really do that and thus we shouldn
13.
▲
by
ertgbnm
4mo ago
"Shark attacks correlate strongly with ice cream sales" is an entirely true statement that some would argue is also misleading. Misleading should be removed as a category and replaced with a better hedge like "not sure"
14.
▲
by
ertgbnm
4mo ago
My threshold for asking for help might be a little higher than the median, but, I like this operating style personally. Maybe it's just how I was raised, but the thought of not trying to figure something out for myself first is unthink
15.
▲
by
ertgbnm
4mo ago
The null hypothesis isn't just the opposite of whatever your opposition believes. For LLMs the null hypothesis would be that there is no relationship between the input and output tokens. Something that is so obviously not true that it&
16.
▲
by
ertgbnm
4mo ago
Sending an AI response to a question that someone asks you is insulting because it's a bit like sending them a link to letmegooglethat where it just animates typing the question you have into google. I think it's only appropriate
17.
▲
by
ertgbnm
4mo ago
If your AI alignment strategy is so fickle that it breaks if people simply discuss potential problems with the strategy then you didn't really have an alignment strategy to begin with.
18.
▲
by
ertgbnm
4mo ago
It makes scams like that scalable. Once you discover one vector of scamming an AI bookkeeper, you can scam all of the users of that AI, using your own AI to scale it for you.
19.
▲
by
ertgbnm
5mo ago
Well did they sell the website too? Or just the domain? Because the domain doesn't generate ad revenue, the original website did. Like just because I sell the domain name for my blog doesn't mean you also get the content of my blo
20.
▲
by
ertgbnm
5mo ago
I think most of the people who pick blue would be empathic, loving people that are just kind of bad at game theory. I don't think I want to live in a world in which they all died out.
21.
▲
by
ertgbnm
5mo ago
The downside of redding is that some portion of the world probably dies and you now have to live in that worse world that if you and 50% of the rest of the world has just blued, would not have happened.
22.
▲
by
ertgbnm
5mo ago
can't wait for "our worst and dumbest model yet"
23.
▲
by
ertgbnm
5mo ago
Instagram follows is not a good way to hire football players but it's probably a good way to hire instagram influencers. The football analogy is a little unfair because VCs are investing in more than just a company's ability to &q
24.
▲
by
ertgbnm
5mo ago
Agreed. I think the starting comparison actually works here. It's a bit like the automobile. The advice of "just don't" doesn't work for cars. It takes a deliberate effort on every scale of society to accomplish, it
25.
▲
by
ertgbnm
5mo ago
It's going for a rendition of the leaning tower of Lire.
26.
▲
by
ertgbnm
6mo ago
Most breakthroughs that are published are for efficiency because most breakthroughs that are published are for open source.' All the foundation model breakthroughs are hoarded by the labs doing the pretraining. That being said, RL reas
27.
▲
by
ertgbnm
6mo ago
Does the data not support a 2X increase in packages? Pre-ChatGPT, in ~2020, there were about 5,000 new packages per month. Starting in 2025 (the actual year agents took off), there is a clear uptick in packages that is consistently about 10
28.
▲
by
ertgbnm
6mo ago
It's less about performance and more about ecosystem lockin. It's a bit like imperial vs metric units. Why would you ever chose to learn imperial if you had the option to only ever use metric to begin with?
29.
▲
by
ertgbnm
6mo ago
shhh just surf that deadly tsunami bro
30.
▲
by
ertgbnm
6mo ago
That would be great if journals bothered publishing replication studies. But since they don't, researchers can't get adequate funding to perform them, and since they can't perform them, they don't exist. We can't lo
More ›