5 ms·
Did you try it with high reasoning effort?
by hypoxia 1y ago
Did you try it with high reasoning effort?
- ares623 1y agoSorry, not directed at you specifically. But every time I see questions like this I can’t help but rephrase in my head: “Did you try running it over and over until you got the results you wanted?”
- chairmansteve 1y agoOr... "Did you try a room full of chimpanzees with typewriters?"
- SequoiaHope 1y agoWhat you describe is a person selecting the best results, but if you can get better results one shot with that option enabled, it’s worth testing and reporting results.
- ares623 1y agoI get that. But then if that option doesn't help, what I've seen is that the next followup is inevitably "have you tried doing/prompting x instead of y"
- theshackleford 1y ago> I get that. But then if that option doesn't help, what I've seen is that the next followup is inevitably "have you tried doing/prompting x instead of y" Maybe I’m misunderstanding, but it sounds like you’re framing a completely normal proces (try, fail, adjust) as if it’s unreasonable? In reality, when something doesn’t work, it would seem to me that the obvious next step is to adapt and try again. This does not seem like a radical approach but instead seems to largely be how problem solving sort of works? For example, when I was a kid trying to push start my motorcycle, it wouldn’t fire no matter what I did. Someone suggested a simple tweak, try a different gear. I did, and instantly the bike roared to life. What I was doing wasn’t wrong, it just needed a slight adjustment to get the result I was after.
- ares623 1y agoI get trying and improving until you get it right. But I just can't make the bridge in my head around 1. this is magic and will one-shot your questions 2. but if it goes wrong, keep trying until it works Plus, knowing it's all probabilistic, how do you know, without knowing ahead of time already, that the result is correct? Is that not the classic halting problem?
- theshackleford 1y ago> I get trying and improving until you get it right. But I just can't make the bridge in my head around > 1. this is magic and will one-shot your questions 2. but if it goes wrong, keep trying until it works Ah that makes sense. I forgot the "magic" part, and was looking at it more practically.
- ares623 1y agoTo clarify on the “learn and improve” part, I mean I get it in the context of a human doing it. When a person learns, that lesson sticks so errors and retries are valuable. For LLMs none of it sticks. You keep “teaching” it and the next time it forgets everything. So again you keep trying until you get the results you want, which you need to know ahead of time.
- furyofantares 1y agoSomething I've experienced with multiple new model releases is plugging them into my app makes my app worse. Then I do a bunch of work on prompts and now my app is better than ever. And it's not like the prompts are just better and make the old model work better too - usually the new prompts make the old model worse or there isn't any change. So it makes sense to me that you should try until you get the results you want (or fail to do so). And it makes sense to ask people what they've tried. I haven't done the work yet to try this for gpt5 and am not that optimistic, but it is possible it will turn out this way again.
- Art9681 1y agoIt can be summarized as "Did you RTFM?". One shouldn't expect optimal results if the time and effort wasn't invested in learning the tool, any tool. LLMs are no different. GPT-5 isn't one model, it's 6: gpt-5, gpt-5 mini, gpt-nano. Each takes high|medium|low configurations. Anyone who is serious about measuring model capability would go for the best configuration, especially in medicine. I skimmed through the paper and I didnt see any mention of what parameters they used other than they use gpt-5 via the API. What was the reasoning_effort? verbosity? temperature? These things matter.
- dcre 1y agoThis is not a good analogy because reasoning models are not choosing the best from a set of attempts based on knowledge of the correct answer. It really is more like what it sounds like: “did you think about it longer until you ruled out various doubts and became more confident?” Of course nobody knows quite why directing more computation in this way makes them better, and nobody seems to take the reasoning trace too seriously as a record of what is happening. But it is clear that it works!
- aprilthird2021 1y ago> Of course nobody knows quite why directing more computation in this way makes them better, and nobody seems to take the reasoning trace too seriously as a record of what is happening. But it is clear that it works! One thing it's hard to wrap my head around is that we are giving more and more trust to something we don't understand with the assumption (often unchecked) that it just works. Basically your refrain is used to justify all sorts of odd setup of AIs, agents, etc.
- dcre 1y agoTrusting things to work based on practical experience and without formal verification is the norm rather than the exception. In formal contexts like software development people have the means to evaluate and use good judgment. I am much more worried about the problem where LLMs are actively misleading low-info users into thinking they’re people, especially children and old people.
- brendoelfrendo 1y agoBad news: it doesn't seem to work as well as you might think: https://arxiv.org/pdf/2508.01191 https://arxiv.org/pdf/2508.01191 As one might expect, because the AI isn't actually thinking, it's just spending more tokens on the problem. This sometimes leads to the desired outcome but the phenomenon is very brittle and disappears when the AI is pushed outside the bounds of its training. To quote their discussion, "CoT is not a mechanism for genuine logical inference but rather a sophisticated form of structured pattern matching, fundamentally bounded by the data distribution seen during training. When pushed even slightly beyond this distribution, its performance degrades significantly, exposing the superficial nature of the “reasoning” it produces."
- deleted 1y ago[deleted]