6 ms·
Based on Tao’s description of how the proof came about - a human is taking results backwards and forwards between two separate AI tools and using an AI tool to
by MyFirstSass 8mo ago
Based on Tao’s description of how the proof came about - a human is taking results backwards and forwards between two separate AI tools and using an AI tool to fill in gaps the human found?
I don’t think it can really be said to have occurred autonomously then?
Looks more like a 50/50 partnership with a super expert human one the one side which makes this way more vague in my opinion - and in line with my own AI tests, ie. they are pretty stupid even OPUS 4.5 or whatever unless you're already an expert and is doing boilerplate.
EDIT: I can see the title has been fixed now from solved to "more or less solved" which is still think is a big stretch.
- D-Machine 8mo agoYou're understanding correctly, this is back and forth between Aristotle and ChatGPT and a (very smart) user.
- MyFirstSass 8mo agoI'm not sure i understand the wild hype here in this thread then. Seems exactly like the tests at my company where even frontier models are revealed to be very expensive rubber ducks, but completely fails with non experts or anything novel or math heavy. Ie. they mirror the intellect of the user but give you big dopamine hits that'll lead you astray.
- Davidzheng 8mo agoThe proof is ai generated?
- MyFirstSass 8mo agoEh? The text reads: "Aristotle integrates three main components: a Lean proof search system, an informal reasoning system that generates and formalizes lemmas, and a dedicated geometry solver" Not saying it's not an amazing setup, i just don't understand the word "AI" being used like this when it's the setup / system that's brilliant in conjunction with absolute experts.
- deleted 8mo ago[deleted]
- kortex 8mo agoThat's literally AI though. AI has been around formally since 1956. https://en.wikipedia.org/wiki/Dartmouth_workshop https://en.wikipedia.org/wiki/Dartmouth_workshop AI != AGI != neural networks != LLMs But Tao did mention ChatGPT so i believe LLMs were involved at least partially.
- markusde 8mo agoYes, the contributions of the people promoting the AI should be considered, as well as the people who designed the Lean libraries used in-the-loop while the AI was writing the solution. Any talk of "AGI" is, as always, ridiculous. But speaking as a specialist in theorem proving, this result is pretty impressive! It would have likely taken me a lot longer to formalize this result even if it was in my area of specialty.
- falcor84 8mo ago> Any talk of "AGI" is, as always, ridiculous. How did you arrive at "ridiculous"? What we're seeing here is incredible progress over what we had a year ago. Even ARC-AGI-2 is now at over 50%. Given that this sort of process is also being applied to AI development itself, it's really not clear to me that humans would be a valuable component in knowledge work for much longer.
- feastingonslop 8mo agoExcellent! Humans can then spend their time on other activities, rather than get bogged down in the mundane.
- navels 8mo agoOther activites such as the sublime pursuit of truth and beauty . . . aka mathematics ;-)
- latexr 8mo agoNot going to happen as long as the society we live in has this big of a hard on for capitalism and working yourself to the bone is seen as a virtue. Every time there’s a productivity boost, the newly gained free time is immediately consumed by more work. It’s a sick version of Parkinson’s law where work is infinite. https://en.wikipedia.org/wiki/Parkinson%27s_law https://en.wikipedia.org/wiki/Parkinson%27s_law
- deleted 8mo ago[deleted]
- HDThoreaun 8mo ago"the more interesting capability revealed by these events is the ability to rapidly write and rewrite new versions of a text as needed, even if one was not the original author of the argument." From the Tao thread. The ability to quickly iterate on research is a big change because "This is sharp contrast to existing practice where....large-scale reworking of the paper often avoided due both to the work required and the large possibility of introducing new errors."
- SecretDreams 8mo ago> Ie. they mirror the intellect of the user but give you big dopamine hits that'll lead you astray. This hits so true to home. Just today in my field a manager without expertise in a topic gave me an AI solution to something I am an expertise in. The AI was very plainly and painfully wrong, but it comes down to the user prompting really poorly. When I gave a el formulated prompt to the same topic, I got the correct answer on the first go.
- anthem2025 8mo ago[dead]
- jacquesm 8mo agoThis accurately mirrors my experience. It never - so far - has happened that the AI brought any novel insight at the level that I would see as an original idea. Presumably the case of TFA is different but the normal interaction is that that the solution to whatever you are trying to solve is a millimeter away from your understanding and the AI won't bridge that gap until you do it yourself and then it will usually prove to you that was obvious. If it was so obvious then it probably should have made the suggestion... Recent case: I have a bar with a number of weights supported on either end: |---+-+-//-+-+---| What order and/or arrangement or of removing the weights would cause the least shift in center-of-mass? There is a non-obvious trick that you can pull here to reduce the shift considerably and I was curious if the AI would spot it or not but even after lots of prompting it just circled around the obvious solutions rather than to make a leap outside of that box and come up with a solution that is better in every case. I wonder what the cause of that kind of blindness is.
- jiggawatts 8mo agoThat problem is not clearly stated, so if you’re pasting that into an AI verbatim you won’t get the answer you’re looking for. My guess is: first move the weights to the middle, and only then remove them. However “weights” and “bar” might confuse both machines and people into thinking that this is related to weight lifting, where there’s two stops on the bar preventing the weights from being moved to the middle.
- jacquesm 8mo agoThe problem is stated clearly enough that humans that we ask the question of will sooner or later see that there is an optimum and that that optimum relies on understanding. And no, the problem is not 'not clearly stated'. It is complete as it is and you are wrong about your guess. And if machines and people think this is related to weight lifting then they're free to ask follow up questions. But even in the weight lifting case the answer is the same.
- red75prime 8mo agoIllusion of transparency. You are imagining yourself asking this question, while standing in the gym and looking at the bar (or something like this). I, for example, have no idea how the weights are attached and which removal actions are allowed. Yeah, LLMs have a tendency to run with some interpretation of a question without asking follow-up questions. Probably, it's a consequence of RLHFing them in that way.
- krzat 8mo agoIn other words, LLMs work best when *you are absolutely right" and "this is a very insightful question" are actually true.
- encyclopedism 8mo agoLots of users seem to think LLM's think and reason so this sounds wonderful. A mechanical process isn't thinking, certainly it does NOT mirror human thinking. The processes being altogether different.
- EA-3167 8mo agoDo you have any idea how many people here have paychecks that depend on the hype, or hope to be in that position? They were the same way for Crypto until it stopped being part of the get-rich-quick dream.
- NooneAtAll3 8mo agohttps://www.erdosproblems.com/forum/thread/728#post-2808 https://www.erdosproblems.com/forum/thread/728#post-2808 > There seems to be some confusion on this so let me clear this up. No, after the model gave its original response, I then proceeded to ask it if it could solve the problem with C=k/logN arbitrarily large. It then identified for itself what both I and Tao noticed about it throwing away k!, and subsequently repaired its proof. I did not need to provide that observation. so it was literally "yo, your proof is weak!" - "naah, watch this! [proceeds to give full proof all on its own]" I'd say that counts
- adityaathalye 8mo agoExactly "The Geordi LaForge Paradox" of "AI" systems. The most sophisticated work requires the most sophisticated user, who can only become sophisticated the usual way --- long hard work, trial and error, full-contact kumite with reality, and a degree of devotion to the field.
- Yeask 8mo agoIs a good economic decision to hype a bit the importance of the LLM$.
- deleted 8mo ago[deleted]
- Davidzheng 8mo agoDo you need to be a super expert to find gaps in proofs? Debatable
- jasonfarnon 8mo agoI had the impression Tao/community weren't even finding the gaps, since they mentioned using an automatic proof verifier. And that the main back and forth involved re-reading Erdos' paper to find out the right problem Erdos intended. So more like 90/10 LLM/human. Maybe I misread it.
- NewsaHackO 8mo agoThis is what I got from Tao's post as well.
- mmphosis 8mo agoThis website was made by Thomas Bloom, a mathematician who likes to think about the problems Erdős posed. Technical assistance with setting up the code for the website was provided by ChatGPT -from the FAQ
- Tenobrus 8mo agostrongly think you should go read the thread to get a sense of the level of expertise and amount of effort put in by the humans involved: https://www.erdosproblems.com/forum/thread/728#post-2852 https://www.erdosproblems.com/forum/thread/728#post-2852
- dpacmittal 8mo agoThere's a lot more detail in this reddit post from the author - https://www.reddit.com/r/OpenAI/comments/1q6yw5g/how_we_used_gpt52_to_solve_an_erdos_problem/ https://www.reddit.com/r/OpenAI/comments/1q6yw5g/how_we_used...
- naasking 8mo ago> EDIT: I can see the title has been fixed now from solved to "more or less solved" which is still think is a big stretch. "solved more or less autonomously by AI" were Tao's exact words, so I think we can trust his judgment about how much work he or the AI did, and how this indicates a meaningful increase in capabilities.