11 ms·
ChatGPT 5.6 Sol Pro believes that the proof is sound. Usually it’s very good at determining if proofs are correct and their mistakes (a friend of mine is a top
by noname120 2mo ago
ChatGPT 5.6 Sol Pro believes that the proof is sound. Usually it’s very good at determining if proofs are correct and their mistakes (a friend of mine is a top mathematician researcher and confirmed): https://chatgpt.com/share/6a515ead-b464-83ed-b85c-c8674f56ead3 https://chatgpt.com/share/6a515ead-b464-83ed-b85c-c8674f56ea...
Personally this gives me additional confidence that this is the real deal.
- stavros 2mo agoOf course it believes the proof is sound, it wrote it. If you want to check an LLM's output, you should use a different LLM.
- noname120 2mo agoYour comment is not substantiated at all.
- stavros 2mo agoIf you'd ever tried to get an LLM to review its own code, you'd know.
- teravor 2mo agoif you get the same session that wrote the code to review it the poor results are entirely deserved. and if you get a different instance to review the code then you would know that it works rather well.
- amluto 2mo agoNo, the comment is right. The prompt had GPT-5.6 reviewing the proof, and the result, unsurprisingly, survives review by GPT-5.6.
- gf000 2mo agoGiven a new context, why couldn't the same model have a decent shot at reviewing some results? It's not like they identify whether this output is from them and then go "yeah correct", that's not how they work.
- amluto 2mo agoIt’s the other way around. The prompt instructed a GPT-5.6 agent to try to make a proof that would survive review by a GPT-5.6 subagent. If there were some defect that would cause the reviewer subagent to accept an incorrect proof, then one might imagine that someone else asking the same model to review the same proof would give the same result. And the proof generation process might even be biased to find such an incorrect proof.
- hoppp 2mo agoUse a human maybe. Only people can really verify clankers. Can't trust anything LLM since it will confidently lie too. It can't take responsibility for verification so it can't verify.