8 ms·
As I posted in another comment, I found Fable to be substantially more powerful than any previous model. However, this isn't just an ungrounded opinion - I uplo
by Tossrock 3mo ago
As I posted in another comment, I found Fable to be substantially more powerful than any previous model. However, this isn't just an ungrounded opinion - I uploaded my full session transcript and code created working on a very complex implementation, so people can judge for themselves, if they're interested: https://tossrock.substack.com/p/36-hours-with-fable https://tossrock.substack.com/p/36-hours-with-fable
- koobyverse 3mo agoOh wow this is quite interesting, thanks for sharing.
- shshnsnnsma 3mo agoThis is very cool, thank you for the write-up. What caught my eye is the complexity you assign to a project like this. It’s hairy but I wouldn’t call it super complicated. I find that super interesting to be honest because it probably means that it is really hard and I am just used to this shit now and it all looks doable to me now. I never think of anything as “complex”, certainly not my own work and I always think what other people do is so much more impressive but I’m starting to realize it might be a me-issue. I worked on some pretty hairy nonsense like say a DB replication solution but I still think it was just tangly, not complex like say a particle collider. Maybe I also need to call my work super complex and highly abstract. Now that I think of it I have a history of not being taken seriously while others with easy shit get credits.
- Lutger 3mo agoImposter syndrome maybe? In a way, nothing is complex at the point where you have untangled it, by definition. Software development is, after all, the art of untangling complexity. The real challenge is (re-)imagining something in the simplest way that fits the goal you are given. When you have arrived there, everything seems obvious and simple. But not everybody could have done it.
- Tossrock 3mo agoThanks, and I can definitely relate to not wanting to assign complexity to one's own work. I think the trick there is that, once you know how to do something, it doesn't seem hard, even if acquiring the knowledge and skills to do it is itself quite a challenge. And I agree that, in some senses, it's not /that/ hard - I mean I'm not proving P=NP, here. It's a software engineering problem, with existing solutions. That said, there is a spectrum of difficulty, even within software engineering problems with existing solutions. Fizzbuzz is less complex than distributed systems. This particular problem strikes me as rather difficult, and one way you can tell (beyond the stuff I mention in the post around serialization, UI paradigms, meta applications, etc) is that earlier models /couldn't/ do it. Which is why Fable being able to, when they could not, was so exciting to me.
- varjag 3mo agoInteresting. I tried Fable vs Codex 5.5 xhigh on three different cases. 1. A resource leak with unknown cause. Both of them zoomed onto the same potential issue and proposed almost identical patches. Fable missed an edge case that Codex handled correctly. 2. Review of a SPICE model. Models had different comments, none substantial. Both missed important issues that were simulated inadequately. Clearly a valley where they are undertrained. 3. An open research problem in CS, presented as a codebase with documentation and performance metrics over datasets. Both were spinning wheels. Which can certainly mean the whole approach had run its course but older models were not able to identify the previous round of improvement either. I liked the prose coming out of Fable more: it was almost like if Obama was giving tech speeches. By actual solution metrics however they both appear in the same place, naturally with the caveat that we didn't really have more time with Fable to compare further.
- mirsadm 3mo agoTo me it feels like they're basically tweaking these things around the edges. I'm not seeing any difference in capability just preference. This has been the case for a while.
- kingkongjaffa 3mo agoMost people thought Fable had more 'taste' than Opus, there was certainly a better quality of writing that felt more 'smart human' and not 'stochastic parrot stringing sentences together'.
- _heimdall 3mo agoThat makes sense, its seemed to me for a while now the competing product is the harness not the model itself.
- Lerc 3mo ago>2. Review of a SPICE model. Models had different comments, none substantial. Both missed important issues that were simulated inadequately. Clearly a valley where they are undertrained. When models miss things, there is always the possibility that it has the capability to identify the issues but it is misevaluating the level of analysis that you want it to do. The fine tuning will have them targeting a balance of subjective opinions of what is appropriate. To go beyond broad demographic guessing the model really needs to 'get to know you' to know what it means when you specifically request an action. Without that information about you it has to weigh your words against the level of sophistication it expects a standard user is able to express.
- NetOpWibby 3mo agoGreat post. I miss Fable.
- tasuki 3mo ago> code created working on a very complex implementation I always find it amusing when people claim "a very complex implementation". Sometimes it's a hard problem, other times an easy one. Either way that's not for you to judge. And the implementation being complex... is that a good thing? Wouldn't a simple implementation be better? It reminded me of the parable of two programmers.
- enraged_camel 3mo ago>> Either way that's not for you to judge. Says who? If you find something complex, you can just say that it's complex. I don't get what the objection is.
- tasuki 3mo agoPeople have so vastly different opinions of what constitutes a complex problem that it carries no meaning.
- cognitiveinline 3mo agowhy is it not for the author to judge, you can disagree with their judgement, but they have brought the receipts to back the claim
- Tossrock 3mo agoI go a lot more into why this was a complex problem in the post, but the short version is, I had it finish the implementation of a meta-application (an application that creates other applications), which has substantial irreducible complexity.
- tasuki 3mo agoFair. To be honest I didn't read your (probably very good, judging by the comments here) post.
- teekert 3mo agoYou guys are getting Fable?
- KronisLV 3mo agoAt least someone is bringing receipts! I think LLM discussions could use a lot of this, both ways - to see what works and also what doesn't work. Still wouldn't help with circumstances where models might be secretly getting dumbed down during peak load, but at least it's something!
- l1ng0 3mo agoYou write to the AI as if it were a person. From my point of view it looks like a fair bit of extra typing and extra tokens. Is there a reason you include things like your emotional response and use a very chatty tone? Do you find this seems to alter responses?
- icholy 3mo agoI don't want LLM usage to inadvertently change the way I communicate with people.
- l1ng0 3mo agoLots of interesting answers here, but I'm particularly intrigued people feel this would affect how they talk to humans. I guess I see AI the same way I see a compiler. I've never worried that I'll end up using code syntax in human conversations. I tend to write to the AI using language I'd use in technical documentation: concise, detailed, unambiguous as possible.
- Yiin 3mo agoI do the same, and it's mostly because I use one type of human communication to both communicate with people and to provide inputs to llms - and I'd rather not have to "mode-switch" between the two, so keeping same style of mannerism is easier to manage as it lets me focus on my requests instead of thinking how to sound more robotic to save tokens.
- blanched 3mo agoI had a coworker who occasionally clearly wouldn't mode-switch from LLM to person mode when asking me questions over slack, which was very jarring. They were normally were personable and friendly, so it was obvious when it happened. Grammar and niceties went out the window. I briefly felt like I was roleplaying an LLM!
- jcims 3mo agoSame. I still say please and thank you as well. It's not for the LLM, it's for me.
- varispeed 3mo agoI would maybe be impressed if it created the code from scratch. It is using the ready made framework, probably it has also learned the code that is using it. What is so impressive about it? You could have done something like this easily with older models. I personally found Mythos to be mediocre. Way worse performance than I remember when using Opus 4.6 before it was nerfed.
- flatline 3mo agoA nit: did you go from Opus 4.5 to Fable? One of the big questions in my mind is how much of a real change Fable is over the existing models. Opus 4.5 -> 4.8 was also a major capability increase.
- Tossrock 3mo agoI've been using 4.6, 4.7 and 4.8 since each was released. I agree 4.5 => 4.8 is a jump in capability, but from my perspective was nothing like the jump from Opus to Fable. I encourage you to read the transcripts and form your own opinions, though!
- chatmasta 3mo agoWhat tool did you use to export the transcript as HTML?
- Tossrock 3mo agoI had claude create one, it's in the same repo as the transcript: https://github.com/Tossrock/claude_transcripts/ https://github.com/Tossrock/claude_transcripts/
- epolanski 3mo agoYes it was great, but it also was stubbornly overtrained to go from prompt to solution on its own. Since Opus 4.6 each following model has been increasingly worse at assisting me, and turned me into the assistant. Maybe I'm struggling to cope with the vibe coding thing but it was so frustrating to ask it to investigate X (where X was easy to find by connecting dots in code) and see it working 10 minutes writing endless stuff in /tmp. More than once I asked it similar investigation tasks and it proceeded to fix stuff (while not understanding properly the context of the business). Was it brilliant? Yes. But it truly felt a major paradigm shift in human-llm interaction which I struggled with. I'm increasingly certain they are too RLed to go from one prompt to solution, and that there are no meaningful tasks aimed at multi turn dialogue and user assistance. It really felt like in a league of its own when it came to vibe coding, but light years away the usefulness of GPT 5.5 pro.