5 ms·
Solving 20 Erdős Problems with 20 Codex Accounts Running in Parallel
- gravypod 2mo agoWhat kind of harness does the exploration? Where did the corpus of Lean proofs come from? Is the code backing Ton 618 open source?
- orlandpm 2mo agoWho is funding this? Sounds like a fun experiment but that’s a huge amount of compute if I understand correctly.
- Choco31415 2mo agoAccording to a quick google search: "He is currently CTO at Xinobi AI, a Japan-based startup developing personal AI agents."
- colin7snyder 2mo agoThis is a self funded weekend project for me. It's not associated with any employer (:
- orlandpm 2mo ago> dedicated 60-vCPU server How many of these are you paying for out of pocket??
- cybertim 2mo agoI own a dedicated 48vCPU with even 160GB RAM.. its not that expensive, check ebay, maybe now with mem prices it will be a bit more steep but as a hobby it's not crazy to think one owns such a piece of hardware. My dual GPU setup was more expensive I think.
- AussieWog93 2mo agoWhen I looked into this a year ago, it was like €60/mo through Hetzner auction. Might be more now but even if it's double or triple it's not that crazy for a hobby. If you built yourself out of used parts you could do it for under a grand back then too.
- barrenko 2mo agoPost-money people with side interests are what built the current western civilization.
- Ar-Curunir 2mo agoNo, underpaid nerds have built modern civilization.
- TOMDM 2mo agoNo, the people growing their food built modern civilization. (Or millions of disconnected stakeholders with different incentives collectively built modern civilization, but who wants to put that on a bumper sticker)
- abc123abc123 2mo agoNo people procreating built civilization.
- psychoslave 2mo agoNature created people and the metaphysics in their heads.
- pfdietz 2mo agoI thought it was Sid Meier.
- barrenko 2mo agoYeah, what did I say?
- deleted 2mo ago[deleted]
- vessenes 2mo agoVery interesting, on many levels: first, the raw additional compute / search harness is worth reading about; huge numbers of Lean 4 theorems, thousands of vCPUs available for spreading out search, embedding databases of proofs, all very interesting. Second, the proofs -- I understand the Lean 4 proofs to be refereed by Fable, and generated by Chat 5.6 Sol. Unlike the leaked proof of the Cycle Double Cover Conjecture last week which had a very nicely readable nearly humanlike writeup, the proof summaries (from Fable) read like Claude tends to read to me these days - real difficulty with the theory of mind of the reader, they are filled with technical phrases, acknowledgment of hard bits and oblique reference to solutions. In short, they suck. I didn't see the word load-bearing, but I bet it's there. That said, a Lean 4 proof is a pretty compelling output artifact. I find it interesting that it's an additional type of effort to turn these into human readable / appreciable / beautiful / non-shitty proofs. To those who say who cares -- indeed. But. One of the major reasons things like the Erdos problems are valuable is that they can at times spur new techniques and concepts. The best of these concepts are applied elsewhere, advancing the frontier. While we gain a lot from solving these problems, we'll gain even more from that next step of distillation / explanation into something humans and computers can grok together. I'd hope that with so many tentatively marked 'solved' we will see some new techniques / ontology / concepts. If not, still pretty amazing.
- deleted 2mo ago[deleted]
- colin7snyder 2mo agoThis is great feedback (thank you for taking the time), & you especially bring up a fair point on the writeups needing to be more human readable. I'll work on that
- echelon 2mo agoCan you explain what you're using the local compute for vs. the API-based frontier models? That was entirely unclear to me. Are you running tool calls that include inference with local fine tunes? And fast math packages? Controlled by the frontier model agents? Is there a way folks can contribute to this?
- matteoraso 2mo agoI've been wanting to experiment with using AI to prove math theorems, but compute is obviously a massive limiting factor here. Are there any plans to open source this?
- esafak 2mo agoIsn't this sucking the fun out of math? It's not like we're going to get any tangible benefit out of them, so why not let mathematicians keep their jobs?
- no_multitudes 2mo agoThis will keep happening until we stop people from doing it.
- rgarrett88 2mo agoGet the looms while you're at it
- yieldcrv 2mo ago[flagged]
- thejokeisonme 2mo ago| your brain on neocapitalism |
- zamadatix 2mo agoThe thing about math is we don't usually know what is pure fancy and what is civilization altering until far after the discovery. Once in a while it's a real targeted crack at something practical but most often it's collecting things which seem trial until you use them together and suddenly you have computers running LLMs. If it were really just about funding people who like math to have fun then it's easy to do forever: just don't have them look at the results and keep paying.
- esafak 2mo agoWhat is their pay going to be justified by once computers start conjecturing and proving theorems on their own?
- 2mo ago
- fractorial 2mo agoMy mouth is agape at the fact that this project is basically what I have been working on non-stop for the last three weeks and just yesterday gotten to the point of evaluating; hats off... I only have one novel proof (non-Erdos) and 13 first-time formalizations thus far. I still like doing maths by pen and paper, but this is fun too.
- crawfordcomeaux 2mo ago[dead]
- colin7snyder 2mo agoThank you for the kind words! I agree, it's exciting that we can now build advanced AI systems for solving novel math (but i still love pen & paper too)
- lionkor 2mo agoWhen you say "working on" what is your actual contribution? Like, what should I imagine you do? For most people who tell AIs what to do and are proud of it, it's sadly mostly sitting around and staring at "thinking" output, and steering a bit, so I'm curious what the work looks like.
- fractorial 2mo agoValid question. If I were further along and had the time to succinctly write up all my contributions, I would just point you at my blog post. I’m generally a poor communicator, so here goes nothing. I designed and stood up a sovereign inference/compute on my intranet. It uses a trust model that allows for a controller (me) to spin up untrusted inference/forge machines for Lean, Sage, or other runtimes. Untrusted sandbox workers integrate directly into my custom harness as first class “attachments.” This is the “orchestration” layer. It’s mostly on open weights, by design. I haven’t yet started SFTing since my examples corpus isn’t quite where I’d like it. I have solved and formalized a one non-Erdos conjecture. I have formalized several pieces of another subfield that does not exist in Mathlib yet. As for what I am currently working on, I have an idea I want to build out about how we might think about sieving algebraic structures to generating new, insightful conjectures. Using LLMs and distributed compute in this context is just a consequence of needing tools to help visualize or materialize things that I am otherwise bad at so I could keep doing the interesting things myself.
- zitterbewegung 2mo agoI was studying Erdos problems by only taking ChatGPT 5.5 outputs and just asking it to keep on attempting to solve it by asking it to go further. I haven't started doing this with chatgpt 5.6 I have some partial results here https://chatgpt.com/g/g-p-69f03400f420819192418b18ca90ffee-dabook/project https://chatgpt.com/g/g-p-69f03400f420819192418b18ca90ffee-d... What was really interesting is that during the process it was able to find lemmas or theorems that might be related or relevant to be published. While I was doing that I was also trying to use Aristotle to do the Lean formalization and I have a WIP system to do that at https://github.com/aconsapart/thesisus/ https://github.com/aconsapart/thesisus/
- colin7snyder 2mo agoI haven't played around with Aristotle at all, thanks for bringing it up & (also your codebase, thesisus, is very solid!)
- universa1 2mo agoThis looks interesting. I am not really familiar with lean, etc... Could I use this to formalize/verify a proof from a paper?
- trenchgun 2mo agoYes, assuming the stuff it depends on is already formalized.
- zitterbewegung 2mo agoWith Aristotle you could formalize a proof from only the text of a paper.
- ralusek 2mo agoI didn't know people could just have GPT running on their own hardware. How does one...do that? Do you have a special relationship with OpenAI and they lock down your servers or something?
- stavros 2mo agoI think they meant they just ran a different context per invocation, not that they hosted the model themselves.
- colin7snyder 2mo agoapologies, bad wording/explanation. As addressed in another comment: "Poor wording on my end, thanks for flagging. I pull the OAuth refresh token from each Codex account into a custom broker, which mints short-lived access tokens per request and load-balances across the pool."
- jibal 2mo agoWhy in the world is the OP's answer dead? Sometimes I really don't understand HN.
- stavros 2mo agoIt's automatic, I vouched for it but it didn't budge.
- cdelsolar 2mo agoHave people tried these on Millenium problems.. letting it run all night? You never know.
- a_wild_dandan 2mo agosolve p=np make no mistakes
- 7734128 2mo agon=1 or p=0
- AussieWog93 2mo agoHave been having an awful day today, this really cheered me up. Thank you so much for sharing your dumb joke!
- red75prime 2mo agoIt's hard to solve P?=NP due to P!=NP.
- colin7snyder 2mo agoyeah, im currently running the system on navier-stokes (making real progress). Unfortunately P vs NP, on the other hand, is going to have to wait for GPT 7
- rahimnathwani 2mo agoI'm not sure how to interpret this part: "each running its own GPT-5.6 instance". GPT-5.6 is a closed source model and this seems to be a personal project and not something done by OpenAI.
- blazespin 2mo agoyeah, is it API or codex?
- colin7snyder 2mo agoPoor wording on my end, thanks for flagging. I pull the OAuth refresh token from each Codex account into a custom broker, which mints short-lived access tokens per request and load-balances across the pool.
- 3848484894 2mo agosurely it's not copy pasting answers from some obscure polish forum right bros
- aureianimus 2mo agoVery cool! It seems you've got a great setup. An addition that would be very convincing is going the extra mile and making a comparator setup for your Lean proofs. (https://github.com/leanprover/comparator https://github.com/leanprover/comparator) This ensures that the AI is not, in any way, modifiying the Lean context in ways that could lead to unsoundness.
- colin7snyder 2mo agoI haven't come across this before. I will spend time on comparator. thank you very much for the suggestion.
- gertlabs 2mo ago[flagged]
- lionkor 2mo agoThe only thing worse than an ad is an AI ad
- bananaflag 2mo agoTo the author: for the absolute Galois of Q_p problem, the link is wrong.
- colin7snyder 2mo agothank you very much for catching this. just fixed
- khalic 2mo agoAs the tools for AI assisted proof become better and mathematicians make it mainstream (could take a while), we're going to be seing some pretty crazy shit. I can't think of a discipline that has more impact on our current toolchains
- claud_ia 2mo ago[flagged]
- pfdietz 2mo agoSome of the claimed proofs (#129, #130) seem to have been removed at the Erdős Problems site. Were they defective?
- adityashankar 2mo agotbf I am not able to understand the erdos problem website, as to why it still shows problems as open even if they've (as claimed) claimed to be solved
- pfdietz 2mo agoThere are also problems which list submitted proofs. The status of the problem does not officially change until the proofs have been accepted. For these two problems, the submitted proofs disappeared without the status of the problems changing.
- colin7snyder 2mo agoThank you for bringing this up pfdietz. No, not defective. The Lean proofs behind both are machine-checked and unchanged. I withdrew them over framing, not correctness. 1) For #129 a couple people pointed out that the report was very confusing. And I agree. So I'm currently attempting to improve it. 2) For #130, a person pointed it out as being a partial solution. This seems correct, so I'm currently working on making it fully end to end. These are put out as "proposed solutions" for the mathematics community to scrutinize, and the scrutiny worked exactly like it should. Happy to take any feedback and make them better.
- pfdietz 2mo agoThanks for the clarification.
- zingar 2mo agoI feel like I'm seeing a maths+AI change from "let's test the limits of LLMs by seeing if they can do useful math" to "LLMs can do useful math, now let's solve lots of problems!", or put a different way the goal has shifted from "interesting exercise for AI" to "making a big difference in math". Am I correct? Are there practical applications of any these problems being solved? No judgement implied, I'm well aware that "no" only means "not yet".
- colin7snyder 2mo agoI don't disagree with you zingar. I think Erdős problems are great for testing a system's capability on genuinely hard math, which has value as a benchmark in itself, and maybe as a stepping stone toward more real-world impact. Your sentiment is well put. To answer your question directly: most Erdős problems don't have practical applications on their own; the value is the techniques and the machine-checked proofs they leave behind. But there's more real world value in solving some of the FrontierMath Open Problems or Millennium Problems. There's a Venn diagram of "hard problems" and "real world impact" for sure.
- pfdietz 2mo agoAt some point the effort shifts from proof of concept to exploration of impact. Of course doing proving at large scale provides further feedback for improving the AIs. Visual reasoning, for example, may still need work. The big impact will be with scaling, for example complete autoformalization of existing math, and automatic exploration for new conjectures, with emphasis on how interesting they are. Automatic conjecture generation goes way way back, to the days of Lenat's AM system. Modern AI should do a far better job.
- felixlu2026 2mo ago[dead]
- toomuchtodo 2mo agoExceptional work. Reminds me of https://www.distributed.net/Main_Page https://www.distributed.net/Main_Page | https://boinc.berkeley.edu/ https://boinc.berkeley.edu/ | https://en.wikipedia.org/wiki/EFF_DES_cracker https://en.wikipedia.org/wiki/EFF_DES_cracker. Consider scaling up by distributing this work more broadly. "Many hands make light work." Lots of problems remaining to solve.