15 ms·
“Erdos problem #728 was solved more or less autonomously by AI”
- observationist 8mo agoReconfiguring existing proofs in ways that have been tedious or obscured from humans, or using well framed methods in novel ways, will be done at superhuman speeds, and it'll unlock all sorts of capabilities well before we have to be concerned about AGI. It's going to be awesome to see what mathematicians start to do with AI tools as the tools become capable of truly keeping up with what the mathematicians want from the tools. It won't necessarily be a huge direct benefit for non-mathematicians at first, because the abstract and complex results won't have direct applications, but we might start to see millenium problems get taken down as legitimate frontier model benchmarks. Or someone like Terence Tao might figure out how to wield AI better than anyone else, even the labs, and use the tools to take a bunch down at once. I'm excited to see what's coming this year.
- malux85 8mo agoThis is what has excited me for many years - the idea I call "scientific refactoring" What happens if we reason upwards but change some universal constants? What happens if we use Tao instead of Pi everywhere, these kind of fun questions would otherwise require an enormous intellectual effort whereas with the mechanisation and automation of thought, we might be able to run them and see!
- stouset 8mo ago> What happens if we use Tao instead of Pi everywhere Literally nothing other than mild convenience. It’s just 2pi.
- lapetitejort 8mo agoCall me a mathematical extremist but I think pi should equal 6.28... and tau, which looks like half of pi, should equal 3.14...
- measurablefunc 8mo agoIn 1897, the Indiana General Assembly attempted to legislate a new value for pi, proposing it be defined as 3.2, which was based on a flawed mathematical proof. This bill, known as the Indiana pi bill, never became law due to its incorrect assertions and the prior proof that squaring the circle is impossible: https://en.wikipedia.org/wiki/Indiana_pi_bill https://en.wikipedia.org/wiki/Indiana_pi_bill
- measurablefunc 8mo agoYou're forgetting that some equations have π/2 so on balance nothing will change. It will be the same number of symbols.
- ogogmad 8mo agoI don't think it's just the sheer number of symbols. It's also the fact that the symbol τ means "turn". So you can say "quarter-turn" instead of π/2. I'm not sure why that point gets lost in these discussions. And personally, I think of the set of fundamental mathematical objects as having a unique and objective definition. So, I get weirdly bothered by the offset in the Gamma function.
- deleted 8mo ago[deleted]
- chmod775 8mo agoI can write a sed command/program that replaces every occurence of PI with TAU/2 in LaTeX formulas and it'll take me about 30 minutes. The "intellectual effort" this requires is about 0. Maybe you meant Euler's number? Since it also relates to PI, it can be used and might actually change the framework in an "interesting way" (making it more awkward in most cases - people picked PI for a reason).
- saulpw 8mo agoYeah but you also have to replace all (2*tau/2) with tau, and 4*(tau/2)^2 with tau^2, etc etc...
- observationist 8mo agoI think they mean in a more general way - thinking with tau instead of pi might shift the context in terms of another method or problem solving algorithm, or there might be obscure or complex uses of tau or pi that haven't cross-fertilized in the literature - where it might be natural to think of clever extensions or use cases in one context but not the other, and those extensions and extrapolations will be apparent to AI, within reach of a tedious and exhaustive review of existing literature. I think what they were getting at is something like this: The application of existing ideas that simply haven't been applied in certain ways because it's too boring or obvious or abstract for humans to have bothered with, but AI can plow through a year's worth of human drudgery in a day or a month or so, and that sort of "brute force" won't require any amazing new technical capabilities from AI.
- sublinear 8mo ago* Tau
- HardCodedBias 8mo agoThink of how this opened up EM: https://ddcolrs.wordpress.com/2018/01/17/maxwells-equations-from-20-to-4/ https://ddcolrs.wordpress.com/2018/01/17/maxwells-equations-...
- kridsdale3 8mo agoNot just for math, but ALL of Science suffers heavily from a problem of less than 1% of the published works being capable of being read by leading researchers. Google Scholar was a huge step forward for doing meta-analysis vs a physical library. But agents scanning the vastness of PDFs to find correlations and insights that are far beyond human context-capacity will I hope find a lot of knowledge that we have technically already collected, but remain ignorant of.
- newyankee 8mo agoExactly, and I think not every instance can be claimed to be a hallucination, there will be so much latent knowledge they might have explored. It is likely we might see some AlphaGo type new styles in existing research workflows that AI might work out if there is some verification logic. Humans could probably never go into that space, or may be none of the researchers ever ventured there due to different reasons as progress in general is mostly always incremental.
- zozbot234 8mo agoGoogle Scholar is still ignoring a huge amount of scholarship that is decades old (pre-digital) or even centuries old (and written in now-unused languages that ChatGPT could easily make sense of).
- semi-extrinsic 8mo agoThis idea is just ridiculous to anyone who's worked in academia. The theory is nice, but academic publishing is currently in the late stages of a huge death spiral. In any given scientific niche, there is a huge amount of tribal knowledge that never gets written down anywhere, just passed on from one grad student to the rest of the group, and from there spreads by percolation in the tiny niche. And papers are never honest about the performance of the results and what does not work, there is always cherry picking of benchmarks/comparisons etc. There is absolutely no way you can get these kinds of insights beyond human context capacity that you speak of. The information necessary does not exist in any dataset available to the LLM.
- 8mo ago
- ogogmad 8mo agoI'm using LLMs to rewrite every formula featuring the Gamma function to instead use the factorial. Just let "z!" mean "Gamma(z+1)", substitute everywhere, and simplify. Then have the AI rewrite any prose.
- sublinear 8mo agoI agree only with the part about reconfiguring existing proofs. That's the value here. It is still likely very tedious to confirm what the LLMs say, but at least it's better than waiting for humans to do this half of the work. For all topics that can be expressed with language, the value of LLMs is shuffling things around to tease out a different perspective from the humans reading the output. This is the only realistic way to understand AI enough to make it practical and see it gain traction. As much as I respect Tao, I feel like his comments about AI usage can be misleading without carefully reading what he is saying in the linked posts.
- deleted 8mo ago[deleted]
- roadside_picnic 8mo ago> It is still likely very tedious to confirm what the LLMs say, A large amount of Tao's work is around using AI to assist in creating Lean proofs. I'm generally on the more skeptical side of things regarding LLMs and grand visions, but assisting in the creation of Lean proofs is a huge area of opportunity for LLMs and really could change mathematics in fundamental ways. One naive belief many people have is that proofs should be "intelligible" but it's increasingly clear this is not the case. We have proofs that are gigabytes (I believe even terabytes in some cases) in size, but we know they are correct because they check in Lean. This particular pattern of using state of the art work in two different areas (LLMs and theorem proving) absolutely has the potentially to fundamentally change how mathematics is done. There's a great picture on pp 381 of Type Theory and Formal Proof where you can easily see how LLMs can be placed in two of the most tricky parts of that diagram to solve. Because the work is formally verified we can throw out entire classes of LLM problems (like hallucinations). Personally I think strongly typed language, with powerful type systems are also the long term ideal coding with LLMs (but I'm less optimistic about devs following this path).
- zozbot234 8mo ago> I don't believe that's what's happening in this specific example (and am probably wrong), but this is where a lot of Tao's enthusiasm lies. It absolutely is. With the twist that ChatGPT 5.2 can now also "explain" an AI-generated Lean proof in human-readable terms. This is a game changer, because "refactoring" can now become end-to-end: if the human explanation of a Lean proof is hard to grok and could be improved, you can test changes directly on the formal text and check that the proof still goes through for the original statement.
- ComplexSystems 8mo agoIf this isn't AGI, what is? It seems unavoidable that an AI which can prove complex mathematical theorems would lead to something like AGI very quickly.
- mkl 8mo agoThis is very narrow AI, in a subdomain where results can be automatically verified (even within mathematics that isn't currently the case for most areas).
- threethirtytwo 8mo agoNarrow AI? I’m not saying it’s AGI but this is not a narrow AI it’s a general AI given a narrow problem. ChatGPT.
- gf000 8mo agoIn a very specialized setup, in tandem with a verifier. Just because a specialized human placed in an F-16 can fly at Mach 2.0, doesn't mean humans in general can fly.
- threethirtytwo 8mo agoAn apt analogy. A human is a general intelligence that can fly with an F-16. What happens when we put an artificial general intelligence in an F-16? That's what happened here with this proof.
- mkl 8mo agoNot really. A completely unintelligent autopilot can fly an F-16. You cannot assume general intelligence from scaffolded tool-using success in a single narrow area.
- threethirtytwo 8mo ago
- xorcist 8mo ago> Reconfiguring existing proofs in ways that have been tedious or obscured from humans, To a layman, that doesn't sound like very AI-like? Surely there must be a dozen algorithms to effectively search this space already, given that mathematics is pretty logical?
- tombert 8mo agoI actually know about this a bit since it was part of what I was studying with my incomplete PhD. Isabelle has had the "Sledgehammer" tool for quite awhile [1]. It uses solvers like z3 to search and apply a catalog of proof strategies and then try and construct a proof for your main proof or any remaining subtasks that you have to complete. It's not perfect but it's remarkably useful (even if it does sometimes give you proofs that import like ten different libraries and are hard to read). I think Coq has Coqhammer but I haven't played with that one yet. [1] https://isabelle.in.tum.de/dist/doc/sledgehammer.pdf https://isabelle.in.tum.de/dist/doc/sledgehammer.pdf
- matu3ba 8mo ago1 Does this mean that Sledgehammer and Coqhammer offer concolic testing based on an input framework (say some computing/math system formalization) for some sort of system execution/evaluation or does this only work for hand-rolled systems/mathematical expressions? Sorry for my probably senseless questions, as I'm trying to map the computing model of math solvers to common PL semantics. Probably there is better overview literature. I'd like to get an overview of proof system runtime semantics for later usage. 2 Is there an equivalent of fuzz testing (of computing systems) in math, say to construct the general proof framework? 3 Or how are proof frameworks (based on ideas how the proof could work) constructed? 4 Do I understand it correct, that math in proof systems works with term rewrite systems + used theory/logic as computing model of valid representation and operations? How is then the step semantic formally defined?
- wizzwizz4 8mo agoThese questions are hard to understand. 1. Yes, the Sledgehammer suite contains three testing systems that I believe are concolic (quickcheck, nitpick, nunchaku), but they only work due to hand-coded support for the mathematical constructs in question. They'd be really inefficient for a formalised software environment, because they'd be operating at a much lower-level than the abstractions of the software environment, unless dedicated support for that software environment were provided. 2. Quickcheck is a fuzz-testing framework, but it doesn't help to construct anything (except as far as it constructs examples / counterexamples). Are you thinking of something that automatically finds and defines intermediary lemmas for arbitrary areas of mathematics? Because I'm not aware of any particular work in that direction: if computers could do that, there'd be little need for mathematicians. 3. By thinking really hard, then writing down the mathematics. Same way you write computer programs, really, except there's a lot more architectural work. (Most of the time, the computer can brute-force a proof for you, so you need only choose appropriate intermediary lemmas.) 4. I can't parse the question, but I suspect you're thinking of the meta-logic / object-logic distinction. The actual steps in the term rewriting are not represented in the object logic: the meta-logic simply asserts that it's valid to perform these steps. (Not even that: it just does them, accountable to none other than itself.) Isabelle's meta-logic is software, written in a programming language called ML.
- Davidzheng 8mo agoI don't think there's a real boundary between reconfiguring existing proofs and combining existing methods and "truly novel" math
- D-Machine 8mo agoThis is great, there is still so much potential in AI once we move beyond LLMs to specialized approaches like this. EDIT: Look at all the people below just reacting to the headline and clearly not reading the posts. Aristotle (https://arxiv.org/abs/2510.01346 https://arxiv.org/abs/2510.01346) is key here folks. EDIT2: It is clear much of the people below don't even understand basic terminology. Something being a transformer doesn't make it an LLM (vision transformers, anyone) and if you aren't training on language (e.g. AlphaFold, or Aristotle on LEAN stuff), it isn't a "language" model.
- XCSme 8mo ago> beyond LLMs to specialized approached Do you mean that in this case, it was not a LLM?
- D-Machine 8mo agoIt could not be done without Aristotle (https://arxiv.org/pdf/2510.01346 https://arxiv.org/pdf/2510.01346), as clearly described in Tao's posts.
- TeMPOraL 8mo agoAristotle is an LLM system.
- D-Machine 8mo ago"Aristotle integrates three main components: a Lean proof search system, an informal reasoning system that generates and formalizes lemmas, and a dedicated geometry solver" It is far more than an LLM, and math != "language".
- TeMPOraL 8mo ago> Aristotle integrates three main components (...) The second one being backed by a model. > It is far more than an LLM It's an LLM with a bunch of tools around it, and a slightly different runtime that ChatGPT. It's "only" that, but people - even here, of all places - keep underestimating just how much power there is in that. > math != "language". How so?
- tachim 8mo agoYou can try out Aristotle yourself today https://aristotle.harmonic.fun/ https://aristotle.harmonic.fun/. No more waitlist!
- dang 8mo agoThis deserves a HN thread in its own right! Do you want to submit it and email hn@ycombinator.com so we can put it in the SCP (https://news.ycombinator.com/item?id=26998308 https://news.ycombinator.com/item?id=26998308)? Edit: I just realized from https://news.ycombinator.com/item?id=46296801 https://news.ycombinator.com/item?id=46296801 that you're the CEO! - in that case maybe you, or whoever you think most appropriate from your organization, could submit it along with a text description of what it is, and what is the easiest and/or most fun way to try it out?
- tachim 8mo agoSure! Should this be a "Show HN" or some other type of post?
- zamadatix 8mo agoAbsolutely, you've made something new you want to show us that we can try out. dang once posted some tips about making these types of submissions https://news.ycombinator.com/item?id=22336638 https://news.ycombinator.com/item?id=22336638 I'd recommend reading first though. Edit: Also https://hn.algolia.com/?dateRange=all&page=2&prefix=true&query=Show%20HN&sort=byPopularity&type=story https://hn.algolia.com/?dateRange=all&page=2&prefix=true&que... for the most popular Show HNs. Don't be discouraged we like personal/open source projects most often, Obsidian made #1
- svat 8mo ago- Minor nit: The documentation mentions "uvx aristotlelib@latest aristotle" but that doesn't work; it should be "uvx --from aristotlelib@latest aristotle" - It took me a minute or two of clicking around to figure out that the (only?) way to use it is to create an API key, then start aristotle in the terminal and interact with it there. It could be more obvious I think. - Your profile links to http://www.cs.stanford.edu/~tachim/ http://www.cs.stanford.edu/~tachim/ which doesn't work; should be http://cs.stanford.edu/~tachim/ http://cs.stanford.edu/~tachim/ (without the www) (I think Stanford broke something recently for the former not to work.)
- lwansbrough 8mo agoCan anyone with specific knowledge in a sophisticated/complex field such as physics or math tell me: do you regularly talk to AI models? Do feel like there's anything to learn? As a programmer, I can come to the AI with a problem and it can come up with a few different solutions, some I may have thought about, some not. Are you getting the same value in your work, in your field?
- ceh123 8mo agoContext: I finished a PhD in pure math in 2025 and have transitioned to being a data scientist and I do ML/stats research on the side now. For me, deep research tools have been essential for getting caught up with a quick lit review about research ideas I have now that I'm transitioning fields. They have also been quite helpful with some routine math that I'm not as familiar with but is relatively established (like standard random matrix theory results from ~5 years ago). It does feel like the spectrum of utility is pretty aligned with what you might expect: routine programming > applied ML research > stats/applied math research > pure math research. I will say ~1 year ago they were still useless for my math research area, but things have been changing quickly.
- posed 8mo agoDo you use LLM models? Or something else?
- jacquesm 8mo agoI don't have a degree in either physics or math, but what AI helps me to do is to stay focused on the job before me rather than to have to dig through a mountain of textbooks or many wikipedia pages or scientific papers trying to find an equation that I know I've seen somewhere but did not register the location of and did not copy down. This saves many days, every day. Even then I still check the references once I've found it because errors can and do slip into anything these pieces of software produce, and sometimes quite large ones (those are easy to spot though). So yes, there is value here, and quite a bit but it requires a lot of forethought in how you structure your prompts and you need to be super skeptical about the output as well as able to check that output minutely. If you would just plug in a bunch of data and formulate a query and would then use the answer in an uncritical way you're setting yourself up for a world of hurt and lost time by the time you realize you've been building your castle on quicksand.
- bgwalter 8mo ago[flagged]
- Arainach 8mo agoWhether powered by human or computer, it is usually easier (and requires far fewer resources) to verify a specific proof than to search for a proof to a problem.
- bgwalter 8mo agoProfessors elsewhere can verify the proof, but not how it was obtained. My assumption was that the focus here is on how "AI" obtains the proof and not on whether it is correct. There is no way to reproduce this experiment in an unbiased, non-corporate, academic setting.
- deleted 8mo ago[deleted]
- perching_aix 8mo agoWhat bias? It seems to me that in your view the sheer openness to evaluate LLM use, anecdotally or otherwise, is already a bias. I don't see how that's sensible, given that to evaluate the utility of something, it's necessary to accept the possibility of that utility existing in the first place. On the other hand, if this is not just me strawmanning you, your rejection of such a possibility is absolutely a bias, and it inhibits exploration. To willfully conflate finding such an exploration illegitimate with the findings of someone who thinks otherwise as illegitimate, strikes me as extremely deceptive. I don't appreciate being forced to think with someone else's opinion covertly laundered in very much. And no, Tao's comments do not meet this same criteria, as his position is not covert, but explicit.
- stevenhuang 8mo ago[flagged]
- 8mo ago
- svat 8mo agoFor context, Terence Tao started a wiki page titled “AI contributions to Erdős problems”: https://github.com/teorth/erdosproblems/wiki/AI-contributions-to-Erd%C5%91s-problems https://github.com/teorth/erdosproblems/wiki/AI-contribution... (as mentioned in an earlier post https://mathstodon.xyz/@tao/115818402639190439 https://mathstodon.xyz/@tao/115818402639190439) — even relative to when he started this page less than two weeks ago (Dec 31), the current result (for problem [728]) represents a milestone: it is the first green in Section 1 of that wiki page.
- pama 8mo agoVery interesting that the vast majority of proofs formalized by AI (section 6) were only completed in the last few months. Exciting times ahead!
- somecontext 8mo agoSee https://xenaproject.wordpress.com/2025/12/05/formalization-of-erdos-problems/ https://xenaproject.wordpress.com/2025/12/05/formalization-o... for a blog post about that.
- leggothrow 8mo agoThis almost implies mathematicians aren’t some ungodly geniuses if something as absolutely dumb as an LLM can solve these problems via blind pattern matching. Meanwhile I can’t get Claude code to fix its own shit to save my life.
- embedding-shape 8mo ago> Meanwhile I can’t get Claude code to fix its own shit to save my life. Maybe this should give you some hint to that you're trying to use it in a different way than others?
- Davidzheng 8mo agoYou're right we're not
- sponnath 8mo agoThere are "ungodly geniuses" within mathematics but no one is saying every mathematician is an "ungodly genius". The quality of results you get from an LLM can vary greatly depending on the environment you place it in and the context you provide it. This isn't to say it's your fault Claude Code can't fix whatever issue you're having.
- oytis 8mo agoAs I understand, a lot of mathematics, at least the part about solving problems, is basically back and forth between exploration (which involves pattern matching) and formalising. We've basically solved formalising a while ago, and now LLMs are getting better and better at exploration. If you think about it, it's also what a lot of other intellectual activity looks like, at least in STEM.
- libraryofbabel 8mo ago2026 should be interesting. This stuff is not magic, and progress is always going to be gradual with solutions to less interesting or "easier" problems first, but I think we're going to see more milestones like this with AI able to chip away around the edges of unsolved mathematics. Of course, that will require a lot of human expertise too: even this one was only "solved more or less autonomously by AI (after some feedback from an initial attempt)". People are still going to be moving the goalposts on this and claiming it's not all that impressive or that the solution must have been in the training data or something, but at this point that's kind of dubiously close to arguing that Terence Tao doesn't know what he's talking about, which to say the least is a rather perilous position. At this point, I think I'm making a belated New Years resolution to stop arguing with people who are still staying that LLMs are stochastic parrots that just remix their training data and can never come up with anything novel. I think that discussion is now dead. There are lots of fascinating issues to work out with how we can best apply LLMs to interesting problems (or get them to write good code), but to even start solving those issues you have to at least accept that they are at least somewhat capable of doing novel things. In 2023 I would have bet hard against us getting to this point ("there's no way chatbots can actually reason their way through novel math!"), but here we are are three years later. I wonder what comes next?
- zozbot234 8mo agoUh, this was exactly a "remix" of similar proofs that most likely were in the training data. It's just that some people misunderestimate how compelling that "remix" ability can be, especially when paired with a direct awareness of formal logical errors in one's attempted proof and how they might be addressed in the typical case.
- libraryofbabel 8mo agoThen what sort of math problem would be a milestone for you where an AI was doing something novel? Or are you just saying that solving novel problems involves remixing ideas? Well, that's true for human problem solving too.
- 8mo ago
- thomasahle 8mo agoIt took Andrew Wiles 7 years of intense work to solve Fermat's Last Theorem. The METR institute predicts that the length of tasks AI agents can complete doubles every 7 months. We should expect it to take until 2033 before AI solves Clay Institute-level problems with 50% reliability.
- kelseyfrog 8mo agoThat's exactly why the Millennium Prize Problem Bench[1] was created. 1. https://mppbench.com/ https://mppbench.com/
- thomasahle 8mo agoThat's amazing :D
- zozbot234 8mo agoThere is an ongoing effort to formalize a modern, streamlined proof of FLT in Lean, with all the needed prereqs. It's estimated that it will take approx. 5 years, but perhaps AI will lead to some meaningful speedup.
- pfdietz 8mo agoWhat I'm hoping to see is high volume automated formalization of the math literature, with the goal of formalizing (or finding flaws in) the entire thing. And once we have that formalized corpus, it's all set up as training data for moving forward.
- zozbot234 8mo agoWe can't really have across-the-board formalization of the math literature without getting the basics done first (including the whole undergrad curriculum) which is what the mathlib folks are working on. It will in fact be interesting to see if AI can meaningfully speed up that work (although they seem to be bottlenecked on review and merging at the moment, not new contribs per se. So a "coding" AI workflow may be a bit of a closer fit.)
- cultofmetatron 8mo agoI remember seeing a documentary where there was a bit about some guy who' life's work was computing pi to 30 digits. Imagine all that time to do what my computer can do in less than a second + a day or two to write the code using the algorithm he used. 10 min if you use newton's
- tonygrue 8mo agoYou’re likely thinking of the Veritasium episode https://youtu.be/gMlf1ELvRzc?si=Qwevl2GwHCzSFcsQ https://youtu.be/gMlf1ELvRzc?si=Qwevl2GwHCzSFcsQ
- MyFirstSass 8mo agoBased on Tao’s description of how the proof came about - a human is taking results backwards and forwards between two separate AI tools and using an AI tool to fill in gaps the human found? I don’t think it can really be said to have occurred autonomously then? Looks more like a 50/50 partnership with a super expert human one the one side which makes this way more vague in my opinion - and in line with my own AI tests, ie. they are pretty stupid even OPUS 4.5 or whatever unless you're already an expert and is doing boilerplate. EDIT: I can see the title has been fixed now from solved to "more or less solved" which is still think is a big stretch.
- D-Machine 8mo agoYou're understanding correctly, this is back and forth between Aristotle and ChatGPT and a (very smart) user.
- MyFirstSass 8mo agoI'm not sure i understand the wild hype here in this thread then. Seems exactly like the tests at my company where even frontier models are revealed to be very expensive rubber ducks, but completely fails with non experts or anything novel or math heavy. Ie. they mirror the intellect of the user but give you big dopamine hits that'll lead you astray.
- Davidzheng 8mo agoThe proof is ai generated?
- MyFirstSass 8mo agoEh? The text reads: "Aristotle integrates three main components: a Lean proof search system, an informal reasoning system that generates and formalizes lemmas, and a dedicated geometry solver" Not saying it's not an amazing setup, i just don't understand the word "AI" being used like this when it's the setup / system that's brilliant in conjunction with absolute experts.
- esafak 8mo agoHow are academics going to assess AI-coauthored research for appointment and promotion?
- Davidzheng 8mo agoDw, by next 3 year AI itself will be better than as coauthor
- markusde 8mo agoVery cool to see how far things have come with this technology! Please remember that this is a theorem about integers that is subject to a fairly elementary proof that is well-supported by the existing Mathlib infrastructure. It seems that the AI relies on the symbolic proof checker, and the proofs that it is checking don't use very complex definitions in this result. In my experience, proofs like this which are one step removed from existing infra are much much more likely to work. Again though, this is really insanely cool!!
- dnw 8mo agoI really want to see if someone can prompt out a more elegant proof of Fermat's Last Theoremthan, compared to that of Wiles's proof.
- maxwells-daemon 8mo agoI work at Harmonic, the company behind Aristotle. To clear up a few misconceptions: - Aristotle uses modern AI techniques heavily, including language modeling. - Aristotle can be guided by an informal (English) proof. If the proof is correct, Aristotle has a good chance at translating it into Lean (which is a strong vote of confidence that your English proof is solid). I believe that's what happened here. - Once a proof is formalized into Lean (assuming you have formalized the statement correctly), there is no doubt that the proof is correct. This is the core of our approach: you can do a lot of (AI-driven) search, and once you find the answer you are certain it's correct no matter how complex the solution is. Happy to answer any questions!
- xiphias2 8mo agoFirst congrats! Sometimes when I'm using new LLMs I'm not sure if it’s a step forward or just benchmark hacking, but formalized math results always show that the progress is real and huge. When do you think Harmonic will reach formalizing most (even hard) human written math? I saw an interview with Christian Szegedy (your competitor I guess) that he believes it will be this year.
- maxwells-daemon 8mo agoThank you! It depends on the topic. Some fields (algebra, number theory) are covered well by Lean's math library, and so I think we are already there; I recommend trying Aristotle for yourself to see how reliably it can formalize these theorems! In other fields (topology, probability, linear algebra), many key definitions are not in Mathlib yet, so you will struggle to write down the theorem itself. (But in some cases, Aristotle can define the structure you are talking about on the fly!) This is not an intrinsic limitations of Lean, it's just that nobody has taken the time to formalize much of those fields yet. We hope to dramatically accelerate this process by making it trivial to prove lemmas, which make up much of the work. For now, I still think humans should write the key definitions and statements of "central theorems" in a field, to ensure they are compatible with the rest of the library.
- 8mo ago
- deleted 8mo ago[deleted]
- emsign 8mo agoSounds to me the actual work was done in the discussions with ChatGPT by the researchers.
- maximgeorge 8mo ago[dead]
- shevy-java 8mo agoSkynet 3.0 is annoying.
- remix2000 8mo agoSo far it's more like Slopnet for the most part
- demirbey05 8mo agohttps://news.ycombinator.com/item?id=46550836 https://news.ycombinator.com/item?id=46550836 Another view on that.
- zkmon 8mo agoWhen Deep Blue beat Kaspaorov, it was not the end of career for human players. But since mathematics is not a sport with human players, what are the career prospects for mathematicians or mathematics-like fields?
- benrutter 8mo agoI think its worth saying two things: 1. This result is very far from showing something like "human mathematicians are no longer needed to advance mathematics". 2. Even if it did show that, as long as we need humans trained in understanding maths, since "professional mathematicians" are mostly educators, they probably aren't going anywhere.
- zkmon 8mo ago> ... are mostly educators, they probably aren't going anywhere Educator business survived so far, only because they provided in-person interactive knowledge transfer and credentials - both were not possible by static sources of knowledge such as libraries and internet. But now all that is possible without involvement of human teachers.
- Davidzheng 8mo agoI wouldn't say professional mathematicians are mostly educators. The educating that mathematicians do even at graduate level to non-future-mathematicians can mostly be done (not fully at parity due to depth of understanding that we accumulate but close) by non professional mathematicians. Most of the education is to other current/future mathematicians in my limited opinion.
- becquerel 8mo agoTao's broad project, which he has spoken about a few times, is for mathematics to move beyond the current game of solving individual theorems to being able to make statements about broad categories of problems. So not 'X property is true for this specific magma' but 'X property is true for all possible magmas', as an example I just came up with. He has experimented with this via crowdsourcing problems in a given domain on GitHub before, and I think the implications of how to use AI here are obvious.
- deleted 8mo ago[deleted]
- MORPHOICES 8mo ago[dead]
- aziis98 8mo agoThe erdos problem website tells the theorem is formalized in Lean but on the mathlib project there is just the theorem statement with a sorry. Does someone know where I can find the lean proof? I don't know maybe it's in some random pull request I didn't find. Edit: Found it here https://github.com/plby/lean-proofs/blob/main/src/v4.24.0/ErdosProblems/Erdos728b.lean https://github.com/plby/lean-proofs/blob/main/src/v4.24.0/Er...
- snowmobile 8mo agoDigging through the PDFs on Google Drive, this seems to be (one of) the generated proofs. I may be misunderstanding something, but 1400 lines of AI-generated code seems a very good place for some mistake in the translation to sneak in https://github.com/plby/lean-proofs/blob/main/src/v4.24.0/ErdosProblems/Erdos728b.lean https://github.com/plby/lean-proofs/blob/main/src/v4.24.0/Er... Though I suppose if the problem statement in Lean is human-generated and there are no ways to "cheat" in a Lean proof, the proof could be trusted without understanding it
- dylanz 8mo agoDoes it work on cryptography? Can it find out the methods behind the fourth Kryptos problem?
- muldvarp 8mo agoEveryone who works for a living is about to have a really bad time.
- kittikitti 8mo agoThis is a great achievement for AI! I quickly read through the thread but found that Tao's page on Github to be easier to comprehend, https://github.com/teorth/erdosproblems/wiki/AI-contributions-to-Erd%C5%91s-problems https://github.com/teorth/erdosproblems/wiki/AI-contribution... It classifies the advancements based on the level of AI input. In particular, the entry in Table 1 related to the original post has both a green and yellow light, reflecting the skepticism from others.
- mehdi1964 8mo agoIf AI can rewrite and formalize proofs this way, do we risk losing the human intuition behind the arguments? Or is it just a tool to explore math faster?
- spopejoy 8mo agoMuch of the discussion here seems focused on the Lean part/correctness, but it sure looks like for Tao its the rapid iteration on the _paper_ that's the important part: > ... to me, the more interesting capability revealed by these events is the ability to rapidly write and rewrite new versions of a text as needed, even if one was not the original author of the argument. > This is sharp contrast to existing practice where the effort required to produce even one readable manuscript is quite time-consuming, and subsequent revisions (in response to referee reports, for instance) are largely confined to local changes (e.g., modifying the proof of a single lemma), with large-scale reworking of the paper often avoided due both to the work required and the large possibility of introducing new errors. However, the combination of reasonably competent AI text generation and modification capabilities, paired with the ability of formal proof assistants to verify the informal arguments thus generated, allows for a much more dynamic and high-multiplicity conception of what a writeup of an argument is, with the ability for individual participants to rapidly create tailored expositions of the argument at whatever level of rigor and precision is desired. Of course this implies that the math works which is the Aristotle part, and that's great ... but this rebuts the "but this isn't AI by itself, this is AI and a bunch of experts working hard, nothing to see here": right, well even "experts working hard" fail to iterate on the paper which significantly hinders research progress.