9 ms·
Speaking as a postdoc in math, I must say that this is rather exciting. This is outside of my field, but the companion remarks document is quite digestible. It
by trostaft 4mo ago
Speaking as a postdoc in math, I must say that this is rather exciting. This is outside of my field, but the companion remarks document is quite digestible. It appears as though the proof here fairly inspired by results in literature, but the tweaks are non-trivial. Or, at least to me, they appear to be substantial to where I would consider the entire publication novel and exciting.
Many of my colleagues and I have been experimenting with LLMs in our research process. I've had pretty great success, though fairly rarely do they solve my entire research question outright like this. Usually, I end up with a back and forth process of refinements and questions on my end until eventually the idea comes apparent. Not unlike my traditional research refinement process, just better. Of course, I don't have access to the model they're using =) .
Nevertheless, one thing that struck me in this writeup, was the lack of attribution in the quoted final response from the model. In a field like math, where most research is posted publicly and is available, attribution of prior results is both social credit and how we find/build abstractions and concentrate attention. The human-edited paper naturally contains this. I dug through the chain-of-thought publication and did actually find (a few of) them. If people working on these LLMs are reading, it's very important to me that these are contained in the actual model output.
One more note: the comments on articles like these on HN and otherwise are usually pretty negative / downcast. There's great reason for that, what with how these companies market themselves and how proponents of the technology conduct themselves on social media. Moreover, I personally cannot feel anything other than disgust seeing these models displace talented creatives whose work they're trained on (often to the detriment of quality). But, for scientists, I find that these tools address the problem of the exploding complexity barrier in the frontier. Every day, it grows harder and harder to contain a mental map of recent relevant progress by simple virtue of the amount being produced. I cannot help but be very optimistic about the ambition mathematicians of this era will be able to scale to. There still remain lots of problems in current era tools and their usage though.
- School-Cotton 4mo agoWhy would it excite you, rather than terrifying you? The better LLMs get at math, the closer the expertise you spent your whole life building is to being worthless. Along with all the rest of what humans find meaningful and fulfilling.
- CamperBob2 4mo agoWhat's happening is the verbal/linguistic equivalent of the invention of calculus. No intellectual field will ever be the same again. Who wouldn't find that exciting, and want to experience it?
- rogerrogerr 4mo agoPeople who enjoy thinking. Ya know, the "intellectual" part.
- aroman 4mo agoThis is the beginning of thinking, not the end...
- windexh8er 4mo agoIt depends. If you are in a disadvantaged class it is very likely going to err towards a dismal result long term. However if you are a privileged intellectual these models can accelerate and expand your horizon. It isn't the end, surely. It is, however, both impressive and depressive simultaneously and that perspective only depends on your point of view.
- doctorwho42 4mo agoBut when the bar to entry is beyond expertise in a field or subfield, how does an individual ever hope to attain an unexplored space to explore? It may be the beginning of thinking, but to many who view things on a longer timeline. It starts to look like it will breakdown the frameworks of which are required to get to that position. Otherwise, you just end up retreading explored ground. This removing the joy of discovery from any humans hand/mind.
- 8n4vidtmkvmk 4mo agoRecently I've found my mind reawaken. It's about asking good questions now. The models can find the answers, but you have to know what to ask. Sometimes the model is wrong and you have to challenge it to find an alternative. Being able to explore problem spaces quickly is interesting.
- isotypic 4mo agoI cannot quite share your enthusiasm. The clearest analogy that I can think of to try to explain why I feel this way is that it seems there will eventually be a phantom textbook of all of mathematics contained in the weights of an LLM; every definition, every proof, etc; and the role of a mathematician is going to be reduced towards reading certain parts of this phantom textbook (read: prompting an LLM to generate a proof or explore some problem) and sharing the resulting text with others, which of course anybody else could have found if they simply also knew the right point of the textbook. To be blunt, this seems incredibly uninteresting to me. I enjoy learning mathematics, sure, but I just don't find much inherent meaning in reading a textbook or a paper. The meaning comes from the taking those ideas and applying them to my own problems, be it a direct proof of a conjecture or coming up with the right framework or tools for those conjectures. But, of course, in this future, those proofs and frameworks are already in the textbook. So what's the point? If someone cared about these answers in the first place, they probably could have found the right prompt to extract it from this phantom textbook anyways. You could argue for there being work still like marginal improvements and applying the returned proof to other scenarios as happened in this case, but as above, what is really there to do if this is already in the phantom textbook somewhere and you just need to prompt better? The mathematicians in this case added to the exposition of the proof, but why wouldn't the phantom textbook already have good enough exposition in the first place? I think my complete dismissal of the value of things like extending the proofs from an LLM or improving exposition is too strong -- there is value in both of them, and likely will always be -- but it would still represent a sharp change in what a mathematician does that I don't think I am excited for. I also don't think this phantom textbook is contained even in the weights of whatever internal model was used here just yet (especially since as some of the mathematicians in the article pointed out, a disproof here did not need to build any new grand theories), but it really does seem to me it eventually will be, and I can't help but find the crawl towards that point somewhat discouraging.
- ted_dunning 4mo agoIn Erdös idiosyncratic nomenclature, all the best proofs are "in the book" and it was always a joyful thing to not only find a proof, but to find the proof that is in the book. Who cares if it is God's book or the machine's Xeroxed copy?
- inciampati 4mo agoI am also using these models to accelerate scientific discovery. Yes, they are making all the difference at the frontier. At least, they feel they are. The messy thing is that we still need to communicate with each other and that's not getting dramatically faster or better. As you note the models need to be built so they do more work to participate in our communication economy. Or we will do so much, alone, to get nowhere fast because so much of our behavior is still bound up in old (good, tested, but clunky) ways of building shared knowledge.
- utopiah 4mo ago> I am also using these models to accelerate scientific discovery. Yes, they are making all the difference at the frontier. Can you please expand on how you do so?
- xbmcuser 4mo agoThis is the main thing that I keep harping about that human knowledge is too vast today for a person or even a group of people and llm will change that many discoveries that require serendipity in the past will be more likely than ever
- pojzon 4mo agoThere is one issue with this. When noone can prove or disprove what AI came up with. Currently we can live with it because someone can review that work. Soon we wont be able.
- aspenmartin 4mo agoWe can use things like Lean for math proofs, verification is tractable.
- yourapostasy 4mo agoI want to hear James Burke of Connections [1] and his writer team to have a free wheeling discussion for a few hours on what they see will happen with LLM’s making these connections with more conscious intent a lot easier. The awesome compression of knowledge aspect of LLM’s is a far undersold aspect of the technology. [1] https://en.wikipedia.org/wiki/Connections_(British_TV_series) https://en.wikipedia.org/wiki/Connections_(British_TV_series...
- Nashooo 4mo agoI don't think you can look at the current AI landscape and say any aspect of it is undersold.
- serf 4mo agoI totally disagree. AI is overhyped, but the demos that have impressed me most are things/elements that very few people are touching and that no one seems to be talking about. the hype is where the money is, as is always -- marketing & porn. Both touched heavily by AI already.
- doctorpangloss 4mo ago> Every day, it grows harder and harder to contain a mental map of recent relevant progress by simple virtue of the amount being produced. I cannot help but be very optimistic about the ambition mathematicians of this era will be able to scale to. There still remain lots of problems in current era tools and their usage though. Always, always always, the problem with research and development is leadership, not insufficient supportive technology. It is a political problem, there is absolutely, positively no shortage of technologies to support research. Your optimism is totally misplaced. The NSF funding cuts have negatively impacted math more than AI has benefitted it. And guess who supports the administration that cut NSF funding? The people who ousted the PhDs from OpenAI.
- virgildotcodes 4mo agoI think we’re looking at a new class of wonderful machines that can potentially make meaningful contributions to the sciences and maybe even humanity as a whole, in addition to far more insidious and destructive capabilities. You are right to point out that the ones who fully own and pilot the machines all belong to the “fuck science and humanity as a whole” group. So the likely outcomes don’t look good. Echoes the early promise of the internet vs the eventual state and consequences of it, although seemingly primed for far more dire and deeply penetrating consequences.
- doctorpangloss 4mo ago> I think we’re looking at a wonderful machine that can potentially make meaningful contributions to the sciences and maybe even humanity as a whole. That's true. But. Maybe you've seen the Oppenheimer movie, there is a moment where Oppenheimer shakes Teller's hand, basically after the guy ruins Oppenheimer's life in a completely immature betrayal. That's what people are angry about, the academy community is Oppenheimer's wife asking, why the fuck did you shake his hand? At least regarding leadership and funding, I don't know if it's a matter of likely or unlikely outcomes. It's just facts: these guys are collaborators. The commenter might very well have zero graduate students starting next year. What pisses me off is the utter obliviousness that STEM people have about how deeply political their work is. And perhaps this is the real reckoning for the mathematics community. Not the possibility that AI is going to replace their jobs, it's not going to do that. But that having these intensely myopic and disagreeable personalities mean that basically zero leadership skills have been nurtured in the mathematics community. You cannot name a single politician who is a mathematician. You have to be elected to have power in this country, it's that simple, there are way more billionaires than there are presidents! Leadership is far more scarce. So that's why these disputes matter, and while it's great that people engage on Hacker News about it, it's intensely disappointing that "reduced science funding is really bad" gets downvoted. That is a result of Hacker News's emphasis on this very 2010s view that it wants to be a place where the math nerds gather (in @dang's words) - he doesn't get that the quality of the discourse was caused by great leadership at many political and academic levels. Nobody credits how much better leaders were during Y Combinator's biggest success stories, or how much we overvalue the intellectual powers of math because it makes money as opposed to enlightening our view of the world.
- energy123 4mo agoTerence Tao gave a recent talk about this issue (lack of attribution). He called it the decoupling of implicit and explicit goals. AI is only good at solving the explicit goals for now, and humans don't have the bandwidth or the institutions to know how to integrate AI into the field. https://youtu.be/Uc2zt198U_U?si=OkwO3xT8-zhSABwh https://youtu.be/Uc2zt198U_U?si=OkwO3xT8-zhSABwh
- sigmoid10 4mo agoThat is an odd summary of the talk. He was talking about how the explicit goal of solving a problem is kind of becoming trivialized, but the abundance of 100-page AI generated proofs will not help the implicit goal of furthering human understanding, because we lack the bandwidth to really digest them. Adhering to things like (human-focused) academic etiquette is a different problem and can probably easily be solved by just giving the model the right context. But having humanity keep up with AI insights into math and science is something we might have to give up eventually. Or at least whoever does will be far ahead of us as a society, because most people's lives will only be affected by the explicit results.
- energy123 4mo agoAn odd summary of a talk you didn't even listen to? He explicitly mentioned references and attribution as a special case of implicit goals.
- sigmoid10 4mo agoDo you really think these models lack the intelligence or language capabilities to handle human etiquette? They can't "read the room" yet because they lack modalities and people don't give them the right context. That's the issue. But I have no doubt that what you two describe here will be solved very soon. And yet the actual implicit goal of all this will need humanity to rethink its priorities.
- computerex 4mo agoI feel like that’s already becoming true. I sometimes work on problems/projects where the AI agent is definitely more qualified than me to call the shots. For example, this library here for deep learning is 100% ai generated and far beyond my technical capabilities. https://github.com/computerex/dlgo https://github.com/computerex/dlgo
- shalmanese 4mo ago> But, for scientists, I find that these tools address the problem of the exploding complexity barrier in the frontier. Every day, it grows harder and harder to contain a mental map of recent relevant progress by simple virtue of the amount being produced. AI is going to both help and hinder this process though. At the end of the day, mathematics is mostly a social process at this point. The goal is not raw number of theorems proven, it’s how proving theorems affects the working operational models of mathematicians. Only a rare few new theorems in mathematics nowadays have direct real world applicability. If AI produced legitimate theoretical breakthroughs at a pace mathematicians are unable to absorb, then the impact will be neutral to negative.
- dyauspitr 4mo ago> Only a rare few new theorems in mathematics nowadays have direct real world applicability. I am no mathematician and very naïve about this, but in a world that is rapidly becoming extremely calculation and network dependent that sounds hard to believe. > If AI produced legitimate theoretical breakthroughs at a pace mathematicians are unable to absorb, then the impact will be neutral to negative. I think the idea here is that all mathematicians will just be using AI for their future work so they don’t really have to absorb it as long as it’s in the training data.
- mkl 4mo ago> > Only a rare few new theorems in mathematics nowadays have direct real world applicability. > I am no mathematician and very naïve about this, but in a world that is rapidly becoming extremely calculation and network dependent that sounds hard to believe. I am a mathematician. It is true. The key is we're talking about new theorems, and direct, current real world applicability. Some theorems that have no applicability now may in the future, as theory often precedes applications by a long way and the usefulness is likely to come from other things built on top of the new maths, and a lot of pure maths will never have direct real world applications but contributes to our overall understanding.
- deleted 4mo ago
- bandrami 4mo agoI am curious if LLMs are better at some kinds of problems than others. IIRC this and another big recent one were cases of the LLM producing a counterexample to a conjecture.
- ricardobayes 4mo agoIMO, it's due to some problems being better documented, with more well-documented, previous research available. LLMs don't really create novel mathematics, they mostly "connect the dots". LLMs by design are not coming up with anything new, unless by statistical probability, aka "brute forcing". I don't want to minimize LLMs capabilities, it's pretty cool they are doing this, and it's useful from a research point of view. But it's important to set expectations.
- jeremyjh 4mo ago> LLMs don't really create novel mathematics, they mostly "connect the dots". That is not what the mathematicians are saying. I don't have the knowledge to evaluate this myself, but a number of mathematicians - for example, in the SP - are saying it goes further than that - they really do introduce novel ideas. Of course everything is based on and inspired by some previous work, but that is true of all human mathematics as well. LLMs that have been trained through reinforcement learning on mathematics are NOT simply token predictors. Only base models can be accurately described that way. They have learned how to do mathematics. They have learned to do coding. Its really amazing we're three years into instruct models and such a large part of Hacker News still does not understand the most basic facts about this field.
- jdub 4mo agoReinforcement learning perturbs the model such that the token prediction process (inference) tends towards the desired result.
- colordrops 4mo agoMaybe I'm misunderstanding how these models work, but isn't it more the responsibility of the harness and its prompts rather than the model itself to make sure that a result is generated with explicit sources?
- PaulRobinson 4mo agoProbably. "All" a model is doing is predicting the next words, based on the statistical distribution of words it has seen similar to the ones read/produced so far. We push a model towards a particular set of distributions through context. If I ask a model "What is the capital of France?", there is a non-zero chance it goes down the dad joke answer of "The letter F". The far more likely option is "Paris", because the joke appears much less often in training material, but if I wanted to be absolutely sure of getting a consistent geography answer I'd address that with additional context. We can add context via prompts, RAG, agents, skills and so on. However, when training a model, we select the material. We could show it a lot more geography information (or dad jokes!), and skew the statistical distribution in the direction we wanted. We could also decide to design the system prompt towards the direction we prefer - which the user would interpret as "the model" - and so nudge the context model-wide. We can also construct the interaction to iterate on context with a specific framing and call it "reasoning". In this specific example, you could therefore solve the problem by a) training skewed towards mathematical papers, which likely degrades performance in general and likely for the specific case too, b) train the user to provide better context/prompts for mathematical work, shifting the workload to them which feels very "a la 2024", c) publish agents and skills that are tailored to mathematics work (very "a la 2026"), d) tweak the system prompt for when the model is doing mathematics work, which the user would see as "the model" doing the change, but you and I might look under the hood and say that is in the harness or a specific type of prompt, or e) add "reasoning" execution that is set to focus on mathematical formatting, or f) a mixture of the above. Right now we're probably looking at agents and skills. I think over time we're going to see smaller models targets towards domains with a mixture of all of it, where some of this sits at user configurable levels, and some is "baked in" via training, system prompts and execution modes, but from a user perspective it's all just "the model".
- peepee1982 4mo ago
- teiferer 4mo ago> Every day, it grows harder and harder to contain a mental map of recent relevant progress by simple virtue of the amount being produced. And by opening the door to LLM-generated results, you'll see greater and greater amounts without any hope of ever navigating this field again without machine help. It's a little like a software project which more and more gets extended by a AI agents with less and less review by human software engineers and in the end the complexity and spaghetti design are so incomprehensible by humans that the maintenance requires an AI agent. The risk is that math as a whole (the field itself) will experience that effect.
- biztos 4mo agoI'm no mathematician but it seems like if this happens, we get to a quite intriguing place as a species. Say we achieve interstellar travel, but nobody actually knows how it works. Or we cure cancer, but the "cure" requires a microrobotic implant, and it runs as a blackbox AI, and only the other AIs can make one, and there's no guarantee they will know how to make one tomorrow. Or we solve global warming but it requires giant cooling machines running 24/7 and again, nobody knows how it works, but with the added bonus that the planet is cooked if they ever stop working.
- f055 4mo agoSo I guess sci-fi movies were right all along. Nobody in Star Wars knows how hyperspace travel works, it just works. The little robots know everything but almost no human bothers to care. People just carry on with their bickering lives while the bots whiz in the background, and these robots are astonished at human inefficiency every single time, but rarely do anything about it. And people are still people.
- cyclopeanutopia 4mo agoThat's only because movies like Star Wars are not sci-fi movies, but more like westerns in space.
- 4mo ago
- qnleigh 4mo agoCan you describe what the reaction to these results has been like in your department? Obviously many people are excited, but what else? How do grad students feel about this? Are any professors getting worried about becoming obsolete?
- csheehan10 4mo agoI am a PhD student in mathematical statistics. The people I have spoken too think this is very exciting and cool. There is also a sense of unease about what this will mean in terms of being a mathematician, and what effects it will have on our future employment. I am also a little worried about what it means for your training as a junior PhD. Often you would try and solve a problem your advisor thinks is doable that they assign to you as a learning exercise. It may be more and more difficult to find problems that a junior PhD can solve but that AI can not. Tim Gowers has written about that here: https://gowers.wordpress.com/2026/05/08/a-recent-experience-with-chatgpt-5-5-pro/ https://gowers.wordpress.com/2026/05/08/a-recent-experience-...
- fragmede 4mo agoHN comments for Gowers post. https://news.ycombinator.com/item?id=48071262 https://news.ycombinator.com/item?id=48071262
- anonymousDan 4mo agoOne concern is that it will become more and more challenging to conduct cutting edge maths research without substantial resources only available at very rich institutions (to pay for state of the art AI assistants).
- ontouchstart 4mo ago> I dug through the chain-of-thought publication and did actually find (a few of) them. If people working on these LLMs are reading, it's very important to me that these are contained in the actual model output. This is a very important point, especially when the output is from a non-deterministic random walk with some unknown probability distribution.
- julianozen 4mo agoNice response to read
- Mikhail_K 4mo ago> But, for scientists, I find that these tools address the problem of the > exploding complexity barrier in the frontier. They do the opposite by locking the results the produce within the slop presentation that needs more AI to comprehend.
- JohnHammersley 4mo agoYes, I share your optimism overall, although I think it is raising a question of what the future role of the researcher is (much like the current debate on developer roles). I attended a conference on AI for maths and open science a few weeks ago, and was struck by just how many examples of AI-supported solutions there already are. Virtually every speaker had an example of either their own use of (often the frontier) AI models in solving a problem that was previously too hard (for various definitions of hard). I wrote up a few notes [1], and most of the speaker videos are available via the conference website [2]. [1] https://scholarlyfutures.substack.com/p/ai-and-the-practical-scientist https://scholarlyfutures.substack.com/p/ai-and-the-practical... [2] https://www.newton.ac.uk/event/ooew11/ https://www.newton.ac.uk/event/ooew11/