Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
beering
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
beering
7d ago
Literally every famous open math problem has had >1 mathematicians ask ChatGPT to solve it. Probably greater than >1000 if you count randos. There is no math problem that OpenAI/Anthropic can solve that didn’t have users already
2.
▲
by
beering
9d ago
That is addressed in the article.
3.
▲
by
beering
9d ago
Is it surprising that different groups are working on the same problems? With each new model generation, the LLMs get good enough to solve a new small fraction of open problems. Of course the problems that get solved are going to be the sam
4.
▲
by
beering
12d ago
Depending on what you’re doing, there are circuit simulators that can verify your work. You can have the LLM drive the simulation but of course then you have to trust that it’s setting up and interpreting the simulation correctly.
5.
▲
by
beering
13d ago
5.6 Sol can already do this with two caveats: 1. It’s too slow for real-time games. To play mario, you’d need to step frame by frame like a TAS. I don’t know if Gruntz has real-time elements or not. 2. It will be expensive. You won’t get ve
6.
▲
by
beering
14d ago
There is simply no level of announcement that won’t have people complaining. What is so important of having a livestream?
7.
▲
by
beering
18d ago
+1 I’ve had this conversation with so many non-techies. They assume Codex can only do coding.
8.
▲
by
beering
18d ago
It can debug, yes, but capability ranges from godlike for algorithmic issues to mediocre for subtle UI things. We are pretty close to your described app singularity if the app is within GPT’s wheelhouse.
9.
▲
by
beering
21d ago
ItMs because Claude sprinkles these words as flavoring without aiding understanding. It feels like Claude thinks of metaphors that don’t actually mean anything (or maybe only makes sense to itself).
10.
▲
by
beering
26d ago
My local grocery chain does the same HHHH strategy: 1. Get new customers 2. Retain them as customers 3. Track what the customers do 4. Keep internal operations confidential Honestly I dislike them but they are the closest store to me.
11.
▲
by
beering
27d ago
> Changes created by Codex had fewer comments in Ruby/Ruby on Rails code. I liked that a lot, and I will soon share some experiments I ran on this. Why is fewer comments a good thing?
12.
▲
by
beering
27d ago
Not really. Selling the book onward does seem legally dubious but legally nothing (yet) prevents you from storing the book in a warehouse. obviously it’s cheaper to dispose of them.
13.
▲
by
beering
27d ago
This is not true. Google Books does not destroy books. No court has ruled that you must destroy books to legally keep a digital copy.
14.
▲
by
beering
1mo ago
Yeah. And the detector can’t even tell you “X% chance this is watermarked” because it doesn’t know the input distribution. It can only tell you “Y% chance that an unwatermakred text would score this high” and how many history professors und
15.
▲
by
beering
1mo ago
On average the distribution of selected watermarked tokens is the same as the original distribution. You can test this experimentally. I think there is a possible weakness in the context of the watermarker but that is not your claim iiuc.
16.
▲
by
beering
1mo ago
Well, it cant be that he is super worried on behalf of people who publish AI slop. That’s not a credible motivation. In fact, he complained a lot about the new ChatGPT app so I can’t believe your claim that he is not using AI. Seems like he
17.
▲
by
beering
1mo ago
It’s a strawman argument because if the LLM is really just “proofreading” for you, there will be little or no watermarked text in your writing. Not enough to trip the watermark detector.
18.
▲
by
beering
1mo ago
No, you are jumping to conclusions about how watermarking works. This is some audiophile thinking that because your RNG is “pure”, you get text with an expansive soundstage or whatever. Intuitively this may be true or false depending on you
19.
▲
by
beering
1mo ago
No, watermark detection is not binary, you get a real number. You decide on a threshold when looking for the watermark. This is the problem - by random chance, some human text will be detected as watermarked. You can turn the detection thre
20.
▲
by
beering
1mo ago
That comment merely says quality must be compromised. It doesn’t make it clear why that must be true. Empirical study seems to say that quality is not compromised, and looking at various proposed schemes, it seems intuitively true.
21.
▲
by
beering
1mo ago
> which inherently compromises quality. I don’t see how this follows? Tokens are chosen randomly. If you choose tokens with a different RNG in the same distribution, you’re still getting equally good or bad tokens.
22.
▲
by
beering
1mo ago
Google has A/B tested watermarking on millions of responses. They say they observed no difference in user behavior.
23.
▲
by
beering
1mo ago
Exactly. The watermark is proportional to how much text is AI generated. Either the AI really just “fixed some typos” (not enough AI content to hide a watermark) or the AI did most of the writing (enough AI content to hide a watermark).
24.
▲
by
beering
1mo ago
The watermark doesn’t change the distribution, only per-token selection. I think not understanding that is the source of most people’s FUD.
25.
▲
by
beering
1mo ago
I commonly see people online say something along the lines of, “instead of curing cancer we got a stupid chatbot.” Dario is simply responding to what people say they want.
26.
▲
by
beering
1mo ago
It’s an econ paper. What do you want? A page of Claude slop with the small caps subtitles?
27.
▲
by
beering
1mo ago
Gotta build some personal benchmarks if you don’t trust the public ones. But the well-known public ones, despite their flaws, are generally high signal on model intelligence.
28.
▲
by
beering
1mo ago
You (or anyone else) can just benchmark and compare. If they were serving a dumber model it would be trivially detectable.
29.
▲
by
beering
1mo ago
You are right, reasoning is unrelated to tokens per second.
30.
▲
by
beering
1mo ago
I’m so excited for driverless cars, and I wonder how this will affect the shift towards services like Waymo. If the choice of whether to take Waymo vs Uber comes down to cost, unions seem like they would shift the usage towards Waymo. (or W
More ›