10 ms·
You say you can have increasingly nuanced discussions with stronger models. What I say is, when I asked Claude why he applied a certain change I didn't underst
by sixtram 3mo ago
You say you can have increasingly nuanced discussions with stronger models.
What I say is, when I asked Claude why he applied a certain change I didn't understand, and boy, it was a small change, he said he "reasoned from first principles" based on the code paths. But it didn't work, and when I asked, "Okay, describe the steps of your reasoning from first principles," it literally answered that it had just made it up.
So, nuanced discussions with models, I don't buy it.
- solenoid0937 3mo agoPosts like this are meaningless without more context - the model you're using, the harness, the initial prompt and context. Fable is better than most staff engineers at my FAANG.
- hrmon 3mo agoBut staff engineers take "responsibility"
- sn0n 3mo agoMaybe I’m missing something, but he talks about charm and tasks (repos on his GitHub). Charm being his harness, and tasks being one of his skills. Idk, maybe I’m mistaken from reading the article… https://github.com/taoeffect https://github.com/taoeffect
- maccard 3mo ago> Fable is better than most staff engineers at my FAANG. While this wouldn’t entirely surprise me, my experience is just not that. Using Claude and fable, it regularly (poorly) recreates features that exist inside our codebase. Sure, I could give way more initial context but at a certain point I’ve given so much context that I would have been faster writing the code myself, or I could have literally handed it to even a fresh graduate to write.
- well_ackshually 3mo agoFable will definitely be the one on call when it inevitably breaks down from the pile of shit slop it wrote at 5AM, don't worry <3
- solenoid0937 3mo agoWe already use AI for oncall and it works better than our humans most of the time.
- OtomotO 3mo agoIncluding you?
- Toutouxc 3mo ago> Fable is better than most staff engineers at my FAANG. That’s genuinely disturbing.
- semilin 3mo ago"Nuanced discussion" doesn't necessarily mean the sort one would have with a human. Statistical apologies are never going to be meaningful. One could edit nonsense into the context window and the model would attempt to rationalize it. The models are smart but you need to use them in a way that makes sense for what they are.
- doctoboggan 3mo agoYou can never ask why a model did a certain thing, or what it was "thinking" when it said something - just like you can't ask a human which neurons were firing when they had a certain thought. The information just isn't available at that level. You absolutely can have deep nuanced discussions with LLMs however, you just need to better understand their strengths and weaknesses.
- loose-cannon 3mo agoDude, these two things are not at all analogous: 1. Asking a model why it did a certain thing, and 2. Expecting a human to say which neuron fired in their response.
- lambdaone 3mo agoEven asking a human being why they did a certain thing is questionable. The research on choice blindness seems like a pretty definitive debunking of post-hoc rationalization: https://en.wikipedia.org/wiki/Introspection_illusion#Choice_blindness https://en.wikipedia.org/wiki/Introspection_illusion#Choice_...
- loose-cannon 3mo agoI'm not sure what point you're trying to make. In science and engineering, being able to provide justification is a core skill. The comparison we should be making is against the human practitioners who are trained in their fields. There will always be a distribution of ability. Saying that there's evidence that people are capable of providing post-hoc rationalization doesn't say anything about the ability of experts to produce well thought out responses (in their respective fields) that don't immediately fall apart under scrutiny.
- lambdaone 3mo agoStructured thinking and deliberation are indeed important, but you can also make LLMs do structured "thinking" if you work hard enough, and generate quite plausible reasoned arguments with valid real-world results, and you can get them to write down their working as they go. But as research has shown, it's not "true" thinking, just pattern matching at a higher level, and eventually runs out of steam.[0] But you only have to drill down a couple more layers and you are back in the void again; do you have any proof that your own thinking, no matter how structured and accurate, is anything other than pattern-matching at a sufficiently much higher level at which you are incapable of seeing it as such? I think we will be finding some very interesting things out soon using the combination of LLMs and theorem provers, as demonstrated by Terence Tao's recent work.[1] A cheetah is not a motorbike is not an aircraft is not a rocket. [0] https://arxiv.org/abs/2506.06941 https://arxiv.org/abs/2506.06941 [1] https://arxiv.org/abs/2603.12744 https://arxiv.org/abs/2603.12744
- sothatsit 3mo ago"Nuanced discussions" is more about describing a design to a model, asking the model to critique your design and ask you for clarifications, and then you providing those clarifications and the model "getting it" and proceeding to additional levels of detail before implementation. In particular the models being able to highlight concerns you have not yet thought about is a pretty good sign of this. Fable is noticeably better at this compared to Opus. I was not talking about models making mistakes. Mistakes, and then models making up justifications for those mistakes, is a failure mode of any LLM, and Fable is no different in that regard. Newer models might make less mistakes, or at least make less egregious mistakes, but they still make mistakes.
- dolebirchwood 3mo ago> he :/
- cpursley 3mo ago[flagged]
- atq2119 3mo agoWe can point out mistakes that feel rather grating without assuming intent behind them. I agree that their use of "he" is likely because they're not a native speaker, especially because they're arguing against the capabilities of LLMs. That doesn't make it inherently wrong to point out the mistake when it's so intertwined with the deeper discussion here, especially given the fact that some (hopefully few) people do build relationships with LLMs.
- weakfish 3mo ago> turd bucket autist I’d be more willing to engage with your argument in good faith without inflammatory language like this. Try and meet people where they are and these conversations become easier.
- recroad 3mo agoThat may be true but it’s still capable of nuanced discussions.