10 ms·
GPT-5.6 used a prompt to close a 30-year gap in convex optimization
- baal80spam 2mo agoWaiting for comments saying that LLMs can't produce anything new and general goalpost moving.
- qsera 2mo agoFrom the post lol >So I wouldn't really say that this result is using or creating some fundamentally new techniques in convex geometry or optimization theory. What this means from my perspective is that if a result is attainable with existing techniques, modern AI methods will be able to solve those problems. I don't think researchers in math/TCS will be made obsolete, but I think it will instead no longer make sense to work on any low-hanging, or even medium-hanging (you know what I mean) fruit. We'll be needed for problems where actual novel approaches are needed.
- monster_truck 2mo agoso it seems like The New Big Question In Math is How's It Hanging, Brother?
- WA 2mo agoIf knowledge is a Swiss cheese, LLMs can help fill the holes, but not make the cheese bigger.
- peddling-brink 2mo agoToday maybe. I disagree in the long term. While they’ll never have the same subjective experience as humans, what stops an LLM from applying similar lines of thought* in a manner that results in a novel conjecture? They are prediction machines, and so are we in a way. We can give them nearly limitless resources to scale their predictive capabilities. We have billions of years of training baked in. They distill directly from our knowledge and can walk down paths that no human has before. It’s silly to say they’ll never do anything novel. At their current capabilities, it sounds like they are already capable of being a specific type is research assistant. What will that look like in 10-20 years?
- seiferteric 2mo agoThey also have ability to go deep and wide in a way that humans just can't. We have limits, get tired, distracted and biased where AI does not. I think there a lot of problem where all the information needed to solve them is there, but we just can't put the pieces together. Like no matter how many people you throw at some problems, you hit human limits and more people won't help, but AI will because it is just relentless.
- qsera 2mo ago>biased where AI does not. AI can be totally biased... The fact that it can spout bullshit all day long to a human who can be tired and would actually act on the said bullshit, is not very comforting... For example, an LLM could confidently declare something a tired human would take as a fact, but would backfire in a real world.
- seiferteric 2mo agoNot really the kind of biased I meant though. There was a recent article about a AI disproving I think an Erdos conjecture by doing similar things humans have tried, but it was much messier and less "beautiful". I think it is a common bias in science and math that things should be "beautiful" but there is no real reason to think that.
- qarl2 2mo ago> While they’ll never have the same subjective experience as humans You state this as a fact - are you aware the question is unresolved? EDIT: I'd love to know why you're downvoting me for stating a known fact.
- skeledrew 2mo agoFear spreads.
- peddling-brink 2mo agoOk, I guess never is a real long time and eventually when we merge our consciousness with the machines we may have the same subjective experience. I can confidently state that GPT-5.6 Sol is not experiencing the same reality as me. They _might_ be "experiencing" and I personally think they are, but their reality and experience is not the same as ours.
- ben_w 2mo agoFamously, all of maths is axioms and tautologies, so I'm not sure this will assuage any professional mathematicians currently having an existential crisis. Maths was already infinite, it's still infinite, but who wants to spend all their lives changing rooms inside Hilbert's Hotel?
- throw310822 2mo agoThe author explains he's an expert in the domain and that he had worked sporadically on the problem for about a year, also with the help of previous LLMs. So whatever he means by "I wouldn't really say that this result is using or creating some fundamentally new techniques" it doesn't mean that the result was trivial. Also, says it might not make sense to work on low or even medium hanging fruits in the future- and I bet that's by far the largest share of work for most mathematicians. Sure, it's not a breakthrough that opens new roads in mathematics- is this where the goalpost has moved now?
- tripleee 2mo agothis is a fairly bleak outlook even when you're trying to make it sound the opposite. Only the cream of the crop talent will have value going on? Most of us aren't Terence Tao
- greenhat76 2mo agoOh brother I can tell you didn't read the entire article.
- qarl2 2mo agoHEH. Don't know why you're getting downvoted. It's painfully obvious that there is a vicious AI backlash now, where every amazing advancement is met with denial and loathing. Oh wait, sorry, I do know why you're getting downvoted. Fear.
- piloto_ciego 2mo agoA lot of people who thought they were special and “better” than mere blue-collar workers are realizing that in fact;l they are the same just working with a different medium.
- qarl2 2mo agoI see a trend line from anti-AI back to anti-evolution to vitalism all the way to Galileo. Humans have a deep need to be special magic flowers - and they can't stand it when science eventually shows them they're not.
- emp17344 2mo ago[flagged]
- qarl2 2mo agoI hope you understand you're proving my point. Ad hominem doesn't fly well around here.
- dudul 2mo agoHumans just want to be able to work and feed their families. Notice how this is something that is never addressed for example in Asimov's writings where robots do everything. How do people do anything to be able to afford these robots? How do Solarians make money to maintain their massively gigantic estates? It is never explained, but we always jump straight to the "robots do everything, everything is free, nobody works". How about the in between?
- ewe42 2mo ago[flagged]
- smokel 2mo agoLean is the Mizar here. For those who have no clue what this is about, Mizar [1] was an early automated theorem prover. Can't wait for HN to add AI features to explain concepts in the sideline, and autovoting. [1] https://en.wikipedia.org/wiki/Mizar_system https://en.wikipedia.org/wiki/Mizar_system
- robinzfc 2mo agoMizar is an early theorem prover. It still exists, see the 2025 issue of Formalized Mathematics journal [1] that publishes math articles formally verified by Mizar (since 1990). [1] https://reference-global.com/issue/FORMA/33/1 https://reference-global.com/issue/FORMA/33/1
- jdw64 2mo agoWhat I'm feeling is that there's a need to study how to use AI well. I've seen professors using AI, and it was amazing. In that sense, I think AI prompt input will become stratified. In the past, implementation skills were very important, but these days, concepts feel more important this is one of those things. It's not that AI brings equality, but rather that the output varies depending on how much background knowledge you have. You could call it a stratification of input I'm starting to feel like there's no place left for programmers like me who focus on quickly churning out MVPs.
- hilariously 2mo ago[dead]
- slifin 2mo agoI think there's a lot of interesting things to the side of development that don't get the resources they deserve Debuggers, testing techniques, testing layers Essentially things that could be used to ground your ai back to reality and work good for humans too
- redsocksfan45 2mo ago[dead]
- semiquaver 2mo agoYou’re at least 18 months out of date claiming that prompting will be the new hot skill. Turns out LLMs are also good at prompting other LLMs.
- aprilthird2021 2mo agoAnd yet in this case a human prompted the LLM for this result, not another LLM
- brookst 2mo agoAh, but who prompts the prompters?
- 2mo ago
- applfanboysbgon 2mo agoTwo points: - Hasn't been peer reviewed yet, so take with a grain of salt. This applies to all claimed proofs, not just AI-generated ones. Even humans hallucinate proofs too! - The prompt is on page 27 here[1]. It is ten pages of advanced mathematics priming the model in the right direction, apparently informed by a year of prior research. That doesn't invalidate the result if it is genuine, but it is worth noting that this wasn't a matter of "ChatGPT, solve this unsolved problem. Make no mistakes." and required substantial domain expertise and human research beforehand. [1]https://arxiv.org/pdf/2607.13335 https://arxiv.org/pdf/2607.13335
- throwthrowuknow 2mo agoSaying “solve this problem” doesn’t get good results most of the time with humans either, it’s entirely underspecified so the person assigned that problem may solve it in a variety of unacceptable ways or not at all or perhaps worse solve the wrong problem because you weren’t clear about its definition. This actually happens all the time. What matters is the ability to communicate clearly and with precision as well as the “harness” which for humans is procedure, training, planning and management.
- applfanboysbgon 2mo ago> Saying “solve this problem” doesn’t get good results most of the time with humans either Sure. That is not even remotely the point I was getting at. Already we see the thread filling up with comments about how human skills are irrelevant, using a mathematics PhD applying his expert skills in a way that the people who are saying that could never have done to justify their inane conclusion.
- camdenreslink 2mo agoThe subtext of this whole post (or at least a subtext that some might read), is "we don't need mathematicians/programmers anymore" or "we will need much fewer mathematicians/programmers". So the fact that this result required a year of prior research and a 10 page prompt of specialized knowledge goes against that subtext. You still needed the human just as much to get to the result, and the LLM ended up being a tool to find the last bit.
- mw67 2mo agoCrazy how intelligence is cheap, efficient and commonplace now. We humans better refocusing our energy on our core values/principles, given most of our skills are becoming irrelevant
- esafak 2mo agoOnce we figure out the pesky problem of how we're going to pay for housing, food, and healthcare.
- timcobb 2mo agoI can't stop wondering myself.... I'm writing some software with AI and wondering, why am I doing this? Will anyone need this? Will anyone have money to buy this? Best I've come up with is we'll need to be adopted by technofeudlaist overlords to be our patrons like in the roman days
- skeke 2mo ago[flagged]
- deleted 2mo ago[deleted]
- georgemcbay 2mo ago> Best I've come up with is we'll need to be adopted by technofeudlaist overlords to be our patrons like in the roman days Continually progressing AI (combined with our current socioeconomic systems) throws a lot of uncertainty into our mid to long term future, but I don't think this is going to be what happens. There are billions more of "us" than of "them", people don't respond well en masse to a drastic worsening of their societal status and "they" are lagging very far behind on building their robot armies. If we poorly navigate this transition the outcome should be worrying them more than it worries us.
- 2mo ago
- rakel_rakel 2mo ago> I don't think researchers in math/TCS will be made obsolete, but I think it will instead no longer make sense to work on any low-hanging, or even medium-hanging (you know what I mean) fruit. We'll be needed for problems where actual novel approaches are needed. I wonder how this compares to what we see happening with "juniors" in software development? In math research, do you also get the training for the profession from working on the low hanging fruits for a while, to then move to the medium-hanging, and later go on to work on previously unsolved stuff?
- skybrian 2mo agoThis apparently required a 10-page prompt. It seems like someone needs to know enough to write it?
- jvanderbot 2mo agoYeah, back to the gold-in-gold out use of LLMs.
- bredren 2mo agoI was thinking this past week I have gotten so lazy w my prompting via CLIs. Back in the before I had put such discipline into my prompting and supporting context. Now I’m like, “look here and here and here are some tools, and /skill /skill okay go.” Or “restate this request in your own words and enrich it as appropriate handling any gaps. Okay go”
- Quothling 2mo agoWe're also at the point where you can roll out context to your entire organisation. I created an app for our m365 Cowork and deployed it to everyone who develops software. It does a couple of things, but it main knows our compliance policies and can guide developers through writing the documentation needed for NIS2 compliance. It also guardrails against non-approved packages, and helps developers find alternatives, or if none can be reasonably found, how to get a new package/dependency approved (or rejected). A few months back this would be something every developer kind of did on their own. Maybe they shared skills, we certainly encouraged it and tried to do all the change management things, but nobody really had the same versions of the skills. Which was horrible in the deployment pipelines, something like the compliance documentation often had to go back and forth several times before it could be approved. Now it's just there, for everyone. In a year or two, I expect a lot of these things to have become even more standardized. So that we don't even really have to build our own apps, but can simply use the ones in the catalog with minimal configuration (and that config will likely only be necessary because I'm from a tiny country that nobody will maintain standards for).
- throwatdem12311 2mo ago[flagged]
- karahime 2mo agoIt's interesting to see the old "Why would we go to space when there are still uncured diseases" show up in a place like this. Science and discovery are singular, all discovery aids all discovery.
- awaythrow9191 2mo agoThe demographics of HN have changed drastically over the past 10-15 years. I don't want to be the "back in my day", but back in my day, there were a lot more technical people here, and politics was much smaller. Now there's a ton more people, ton more politics, and a lot less "hey here's something really cool I built, it's like rsync but nice"
- deleted 2mo ago[deleted]
- ianm218 2mo agoCancer is also bottleknecked by a lot more than just intelligence. If you have 100 of the smartest PHd students working on a cancer problem you have to wait for funding, lab experiments, and clinical trials etc. Math is deterministic and requires nothing like that.
- esafak 2mo agoHave you not heard of things like AlphaFold?
- slashdave 2mo agoLLMs work within the world of what has been written. That is, what is known. And cancer is not a single disease that can be cured with one therapy.
- elhart05 2mo ago[dead]
- oulipo 2mo agoExcept solving problem is probably the least (even though it's important) interesting thing in research... The most interesting thing in research is finding new questions, that we understand and that we know why they are important. And that's something that humans need to do (by definition)
- dash2 2mo agoI keep hearing this but lots of maths problems are practically important! We want to know the answer because it will be useful for applied science, or statistics, or engineering. It’s not all just about knowledge for its own sake.
- a_imho 2mo agoIf I recall correctly there was a proposed proof to the abc conjecture by Mochizuki https://en.wikipedia.org/wiki/Abc_conjecture#Claimed_proofs https://en.wikipedia.org/wiki/Abc_conjecture#Claimed_proofs which was rejected due to being rather inpenetrable to humans. Shouldn't this be an ideal target for LLMs?
- anorwell 2mo agoIt was rejected for being wrong (or most charitably, incomplete).
- lg5689 2mo agoThere was recently an announcement that a group trying to formalize it found a gap exactly where other mathematicians were pointing. So to the extent there was any doubt, it should be gone now--the proof was incorrect. But I agree LLMs have a lot of potential for checking proofs--both informally (they can read quickly and find gaps) and formally (by attempting to formalize).
- 7373737373 2mo agoSimilarly, I'd love to see LLMs create a formal proof of the https://en.wikipedia.org/wiki/Classification_of_finite_simple_groups https://en.wikipedia.org/wiki/Classification_of_finite_simpl...
- charlieyu1 2mo agoI’d like to see four color conjecture and an elementary proof of FLT.
- _alternator_ 2mo agoI know a bit about this field. This conjecture reads as somewhat more niche than the cyclic double cover conjecture recently proved by OpenAI, but nevertheless represents a real contribution. You want to know how long it takes to solve an optimization problem, in this case over convex, lipschitz functions. (The restriction to a spherical domain is not really a restriction, you can just change variables for any bounded domain.) Anyway, showing upper bounds on time complexity is "easy" because it's just the runtime of your algorithm. Showing (nontrivial) lower bounds is usually much harder because it requires constraining all algorithms. This proof apparently shows that the lower bound time complexity is equal to the time complexity of an existing 30-year old algorithm: it requires Omega(d^2) function evaluations to solve over this class of functions. My gut says likely implies that d is the minimal number of evaluations if you have a gradient oracle because you can approximate a gradient with d function evaluations, but I'm not sure how hard it is to make that rigorous.
- LPisGood 2mo agoIt should be noted that optimization of a convex bounded lipschitz function is exactly what most modern statistical learning (AI) models are based on.
- hodgehog11 2mo agoVery confused by this comment. The older (poorer) parts of the ML literature focus on models with convex and (gradient-)Lipschitz objectives, but that's not representative of reality, not even close. Modern objectives for AI models are famously nonconvex (catastrophically, from the point of view of classical optimisation theory), and that's where the interesting research is.
- _alternator_ 2mo agoI'd push back on this. Most of the core optimization techniques (eg, ADAM, stochastic gradient descent) are straight out of the convex optimization literature. Generally you need to use optimizers that work well on convex objectives because near minimizers, functions tend to be convex. (Proof by contradiction: a non-convex point has a strict descent direction.) The fact that neural networks are highly nonconvex has encouraged a lot of research, but it's more of the kind aimed at resolving tension: these methods are probably good for convex functions, why do they continue to work for nonconvex problems, and are there tweaks we can make to improve them in that setting? It's not a lot of de novo theory; more standing on the shoulders of giants, etc etc.
- luciana1u 2mo ago[flagged]
- spwa4 2mo agoThe problem is that we're going to have another deepseek moment when someone uses GLM or Kimi K3 to do this.
- paytonjjones 2mo agoWhat was the first DeepSeek moment? (genuine question, I'm out of the loop on what you mean)
- theragra 2mo agoI guess when cheap Chinese model was very close to SoTA models
- spwa4 2mo ago... and now another one when a (not so cheap tbh) Chinese model matched OpenAI GPT Sol and Anthropic Fable/GPT 4.8 (some scores below, some scores above) ... Oh and exceeded Google's best internal model capability enough to have them cancel the next Gemini release. Probably. But I'm 95% certain that's exactly what happened. This is the first time I really fear for the future of Google, because OpenAI matched or exceeded the experience of Google search ... and now a Chinese company matched OpenAI ... followed by Google failing to catch up. That really made me take pause. It indicates the the "secret sauce" of OpenAI and Anthropic is ... nothing, really. Or none of it matters, except access to the hardware needed for 1.5 trillion parameter models, which Chinese firms now have as well. It means there is no "AI takeoff" where other firms can't match OpenAI/Anthropic models. And Google, inventors of transformers just missed the takeoff. There's more secret sauce, but clearly a great engineering firm can find it in less than 3 months. The fact that this Chinese model has similar cost to OpenAI and ballpark cost compared to Anthropic while we can be quite sure they're minimizing the cost (that's what China is doing everywhere else) also raises questions about the true cost of serving comparable models and with that about the future profitability of OpenAI and Anthropic.
- bananaflag 2mo agoJanuary 2025, release of DeepSeek R1, the first open reasoning model. There was a lot of panic then that it was done with very few resources.
- threethirtytwo 2mo agoGenuine question: If you still or did think LLMs are just stochastic parrots that just summarize everything and have no form of creativity, what do you think after seeing results like this? I'm very curious how people reconcile their fear/hatred of AI with actual objective reality. This is actually what interests me most about the whole AI thing. How we tell ourselves what we tell ourselves.
- barnacs 2mo agoI hold my stance that LLMs are stochastic parrots. Making the parrots ever more complex and training on ever more data produced by intelligent, creative beings may make them more useful or convincing but does at no point give rise to intelligence or creativity.
- beering 2mo agoWith such high standards, most HN commenters also do not have intelligence nor creativity. I don’t think we can set the bar that high.
- tctcd6 2mo agoComical human arrogance...
- qnleigh 2mo agoI won't touch creativity, but if this and other results like it do not demonstrate intelligence, what does? How was it able to solve problems that specialist mathematicians have tried and failed to solve for years?
- barnacs 2mo agoMathematics is a language. If anything, it's much more well defined and formal than most others. Train on enough examples and statistical autocomplete gets you places. I'm surprised how anyone would even consider this intelligence?
- in-silico 2mo ago
- d4rkp4ttern 2mo agoIn the Reddit post there was clarification that this was done with Sol Pro not Ultra - curious what is everyone’s mental model of the difference. My understanding is that ChatGPT Pro is effectively a multi agent system, or somehow uses multiple LLMs in parallel and selects a best answer. And Ultra is more similar to Claude-Code UltraCode where the main agent can choose to create a dynamic JS workflow that deterministically orchestrates multiple agents to handle different parts of a task and have adversarial checkers etc. Is that more or less the difference? Any substantiating sources would be great to see.
- gizmodo59 2mo agoultra in codex is just a way to run multi agent system, pro is similar to other pro models like 5.5
- sashank_1509 2mo agoThis is all a depressing and bleak future that I don’t look forward to. One solution is to ban LLM’s, to artificially create a demand for human thought, that just feels like living in an artificially constructed zoo. Another solution is humans don’t do anything that AI can do better , / doesn’t need the human touch. So I suppose we will all become artists, sportsmen or politicians, the only jobs that will remain except for select few. Maybe this is ok, I don’t know. Another solution is we find a way to mind-meld with AI so that human + Ai >> AI alone. This is dystopian, who gets to decide who mind melds with AI, how much will it cost etc etc. For the stupid copes that the prompt required human ingenuity, let me first add that the author used GPT5.6 to write most of the prompt. He just gave some mild direction. That amount of direction does not require deep expertise and the expertise required will keep falling with time, eventually an undergrad can create this loop and then maybe a high school student. And prompt engineering / loop engineering nonsense is not real. Calling it engineering is a psy-op because it is something simple, imprecise and future models will be much better at it than you. In fact, in the future the most likely outcome is you tell the agent what you want (I want this app, or I want this theorem solved) and it will set up the loop, or loop of loops and use all its computing effort to come up with a result. This is completely dystopian to a human life.
- slashdave 2mo agoThere is more to human life than programming and math proofs.
- tripleee 2mo agounfortunately many won't get to experience it much because we'll be stressed and struggling to make a living
- kevincox 2mo agoThe problem isn't that we are gaining more knowledge and better technology. The problem is that we are allowing the rich to use the technology to subjugate the majority. If AI improves human productivity so much that millions of people no longer need to work that should be an incredible thing. But the flawed structure of our society punishes those people rather than freeing them to persue endeavors that interest then.
- sdwvit 2mo agoNot yet peer reviewed
- ck2 2mo agocould machine-learning even handle a TEN PAGE PROMPT just a year ago? this is changing my mind, at least about experts using advanced tools like any profession where it's like the magic of watching a lifetime of hard-earned skill at work > After seeing OpenAI’s CDC result, I wrote a much more elaborate prompt following the same general methodology. My prompt is about ten pages long and attached at the end of the preprint (see collection of links below). There is a lot baked into this prompt, on approaches to try and also on how exactly the model should proceed, but it's built exactly in the style of OpenAI's CDC prompt. One note is that I gave it a relatively small error requirement, to prove the quadratic lower bound under order d⁻⁴ accuracy. > After 148 minutes, GPT-5.6 Sol Pro returned a proposed proof resolving the quadratic dimension dependence at accuracy of order d⁻³. After checking things myself, I formally verified the proof in Lean, and it passed the formal verification check.
- nilamo 2mo agoIs this interesting? AI does what we made it to do, news at 8?
- ChrisArchitect 2mo agoNon-reddit: https://medium.com/@kerger.p/an-ai-assisted-breakthrough-in-convex-optimization-an-optimization-problem-dating-back-30-years-a-db5c631119de https://medium.com/@kerger.p/an-ai-assisted-breakthrough-in-... (https://news.ycombinator.com/item?id=48939768 https://news.ycombinator.com/item?id=48939768)
- Kirillekko 2mo agoThe amount of people several months ago stating that no one cares about the "unsolved" mathematical problems that AI is able to solve is funny.
- Noe2097 2mo agoCan't wait for GPT to prove that P=NP (or not)!
- stfnon 2mo agonah, scientist with the name Shakey Onail found this and all creds are given to LLMs is crazy
- VonTum 2mo agoDo you have a link to Shakey's publication?
- charlieyu1 2mo agoI tried using AI to solve some advanced math problems. One thing I see is that they can throw an enormous amount of brute force into a problem. When mathematical logic can be brute forced we will see some interesting advances.
- macwhisperer 2mo agoBasically, he proved that *information is power.* If you don't know which way to go (the subgradient), you're gonna be calculating forever!
- YeGoblynQueenne 2mo agoSo if you dig down a bit it turns out the author had been trying to solve that problem for a year with GPT 5.4 and 5.5 and he fed all that information to the prompt he gave to Sol Pro which may or may not had direct access to the author's chat history. So the claimed "148 minutes" was really "a year plus 148 minutes". Moreover, it seems the prompt included the technique used to solve the problem: https://old.reddit.com/r/math/comments/1uxj3cy/after_openais_cdc_proof_announcement_gpt56_used_a/oxyn68j/ https://old.reddit.com/r/math/comments/1uxj3cy/after_openais... In the prompt I basically just throw all reasonable approaches at it, without making a big distinction for what to explore most, and these approaches would all be reasonable for someone who knows the area. Sol helped me with the prompt as well, for which I gave it the CDC prompt, some ideas and specifications, a crystal clear problem description, and then modified things slightly myself after. One thing I do wonder is how much it accessed memory of previous chats, since as mentioned I had worked with 5.5 and 5.4 on this previously, and the main construction is not so different from something I discussed there. But, the function class max of affine functions that worked in the end was also in my prompt, so I'm not totally sure. So it's not clear to me the degree to which "GPT-5.6 used a prompt" to close the gap etc, or the author basically did all the work himself and assigned it to GPT-5.6 out of enthusiasm.
- dwohnitmok 2mo ago> Moreover, it seems the prompt included the technique used to solve the problem: I don't believe this is true. The author sent techniques he used, but I don't believe any of those were ultimately what GPT-5.6 used. GPT-5.6 also provided the Lean formalization, which was not provided at all by the author.
- goldylochness 2mo agosounds like a feature, not a bug it might imply using weaker models to attempt the problem first is a good supplementary prompt to a more advanced model trying to do the same
- phillip_kerger 2mo agoI (author of the original post and paper) can add a few things here: 1. My previous approaches with GPT 5.5 were really not very sophisticated in terms of my input. I threw the problem at it, and just kept encouraging it to go iterate through ideas without any success. 2. The approaches that are in the prompt, though they will seem cryptic to someone not in the field, are relatively natural ideas. In fact, the construction that worked was something that even 5.5 initially looked at but was just too weak to see how to make it work. From my view, I would have never gotten this result myself. Imagine you are telling a contractor to build the empire state building, and you say: "You should explore approaches that can include building materials like steel, wood, concrete, or clay, and any combinations of those. You can use arcs, columns, supportive beams, and anything else you can think of to solve load-bearing issues. Do not stop until you've completed a viable plan to construct the empire state building." And then the contractor shows you the finished empire state building using reinforced concrete and steel beams with all kinds of crazy ways of making everything stable; that's kinda how I feel.
- Eridanus2 2mo agoGenuinely asking... How do you get chatgpt to work for 148 mins when I can't get gemini to think for even a minute at a time?
- seizethecheese 2mo agoChatGPT pro will work 20-30 min in my experience pretty regularly in my experience. Never had 148min but it seems plausibly in the very right tail.
- phillip_kerger 2mo agoclear specifications for what counts as completing the tasks, and an explicit list of what does not count as completing the task, and clearly stating to not return until the task has been completed. I've had agents run for almost 12h!
- marshray 2mo agoIt says I've been "blocked by network security". MSIE on Windows. Reddit must really be circling the drain.
- pastakatsu 2mo agoit blocks vpns and tor
- marshray 2mo agoNot using either. Just my US-based home ISP that works for everything else.
- pona-a 2mo agoWe have a non peer reviewed proof in a niche area of mathematics claiming to have been "co-written" by an LLM. What are the comments about? Lamenting or celebrating humanity's intellectual death... Very insightful.
- bjt12345 2mo agoI can't see this article as Reddit doesn't allow me to view it. Do we have to use Reddit though? It's a horrible website.
- rurban 2mo agoYou have to replace www. with old. Then it works. When I submitted the very same link, HN replaced the good old prefix with the bad prefix www.
- feiz45607 2mo ago[flagged]
- shivam27cool 2mo ago[dead]