18 ms·
Fable and the end of the free lunch
- nchmy 25d agoThe real revolution is Deepseek v4 flash and similar models (GPT 5.6 Luna, muse spark 1.2, mimo, etc...) - Genuinely good performance for a tiny fraction of the cost of Fable and even GLM etc... I think a lot of people would be very content if they never got smarter, and just kept getting even cheaper/faster. Of course, both things continue to happen on a seemingly monthly basis
- lilbigdoot 25d agoIf they could be cheap+fast and not try to do too much, that's a good spot for me. I don't use the smarter models as much because of cost and because they're still not good enough to let loose on a lot of problems. For assistance I prefer something that can very quickly spit out a specific piece I can review on the spot and keep going. I let smarter models handle things that I treat as external dependencies and don't care how they're written, but in my core domain I'm still mostly hand coding
- nchmy 25d agoI have a similar process - its just a pair programmer most of the time. I dont understand how people can have a fleet of agents working a bunch of waterfall specs..
- ipsod 24d agoI more have an agent that I drive to create features, and then a fleet of agents that turn those ad-hoc implementations into refined, integrated code.
- matteoraso 25d ago>I think a lot of people would be very content if they never got smarter, and just kept getting even cheaper/faster. There's a lot of truth to this. I think we're starting to approach the point where increased intelligence has declining marginal returns, such that it might not even be worthwhile to improve models unless it can be done cheaply.
- ColdStream 25d agoI have argued for a while that this was an S-curve it was just a case of figuring out which part of it we were in. I am more confident nowadays that we are heading towards the upper plateau but there might still be some head room on that.
- insanitybit 24d agoI'm really not convinced that these models are even that much more intelligent, as opposed to simply being more token aggressive. I do not find Fable that much smarter than Opus 4.6, and no Opus model seems to have improved things much at all. Benchmarks seem gamed at this point, real world experience just doesn't match up.
- someothherguyy 24d agoi wish i was experiencing these things that everyone else is. my experience is mostly frustration and rewrites of anything that requires more than what would take me an hour to do myself, unless it is pure translation / boiler plate work. the leaps are there at getting to more "shaped" code (code that is correct for linters, static checking, etc), but i don't see the models exhibiting much intelligence. i really can't think of a time using LLMs for building anything where they did something that would make me go, "wow, that is really impressive, i wonder how it came up with that." just brute force search and pattern matching still. even the interesting results in academic work seem to be more of a function of effort (proofs by exhaustion, fitting puzzle pieces in a search space, etc) than anything else. not to say people aren't using large language models to do impressive things, but the agents themselves do not seem very intelligent to me. it feels like some engineering teams are aware of this fact and are driving agents using strict rule checks (like hooks on steroids), so they can drive some shape of output that aligns with what they require.
- geniium 25d agoI was using ChatGPT voice during cooking to reflect on variations of a dishes i was preparing for years. It was so amazing to get advices and reflect that it struck me : I could use this model forever - it’s clever enough to help me tons and do lot of work for me - even if ai would stop evolving I would love it
- antman 23d agoI was walking inside rooms in buddhist temples in western China today and ChatGPT knew and could discuss what each room had and explain the art society and legend.
- r_lee 25d agoimo this is the problem some of these labs are gonna face, because open models will do this just fine and you as the consumer don't need to pay their training costs especially considering imo most use falls under this instead of those kind of tasks where you'd need the SOTA
- josephg 25d agoYeah. Sometimes I wonder who the long term financial winners will be from the ai boom. It might be ram / gpu manufacturers. Or whoever cracks putting LLMs on asics.
- a2ff6eeb0 25d agoIt's going to be the shareholders of the first companies to crack AGI, and make human brains fully irrelevant economically. With the trillions of dollars that's going in through both investment and users, it's going to happen. I don't believe the human brain has fundamental magic that will make this impossible.
- a2ff6eeb0 25d agoFor the downvoters: What magic do you think the human brain has that makes it impossible to emulate acceptably?
- poincareball 25d agoEvidence actually supports that capabilities are leveling off, and cheaper/faster is not really coming. Just log-linearly more capability at smaller parameter counts as they saturate.
- ACCount37 25d agoWhat "evidence"? Because we keep running out of benchmarks to distinguish frontier model performance. If capabilities are "leveling off", we're not seeing it yet.
- bad_haircut72 25d agonot an AI researcher - this is probably true for these "everything" LLMs but I think specialized models are gonna be the next big thing
- ACCount37 25d ago"Specialized models" are a bit of a doozy. The biggest generalist models beat the most fine-tuned specialists, as a rule. You can bias an LLM away from literature knowledge and towards coding capabilities, but that buys you very little performance, and for too much effort. Generality and intelligence seem to be entangled very heavily in LLMs.
- CamperBob2 25d agoAnd yet, there's VibeThinker 3B to bring this long-held premise into question (if not to blast it to pieces.) It is practically illiterate by the standards of larger models, yet performs like models 100x its size on mathematical and logical reasoning tasks.
- ACCount37 25d agoWhich are the kinds of tasks computers have been historically quite good at. It's impressive that it does what it does, don't get me wrong. But if you expect it to replace the likes of GPT 5.6 Luna, let alone Sol? Nah.
- ksh09 25d agoI'd be content if I could get the DS4 flash, luna, mimo level intelligence running on MY low-end hardware completely offline and bearable TPS, not otherwise.
- redox99 25d agoEh. I don't think Luna is good enough. I think that threshold is around Opus / Sol where it can do most of the tasks for me. But I still have many tasks which require either better intelligence or better UI design capabilities. With how generous subscriptions are, what I actually want is GPT Astra, not cheaper Sol.
- intrasight 25d ago> content if they never got smarter, and just kept getting even cheaper/faster. I'm definitely not getting smarter. But my tolerance is 1 drink so I'm definitely cheaper. Also as a result, I spend more time training and so I am faster. And yes, I am more content
- jimmydoe 24d agoCurrent AI is smart enough to help us, but the creators of AK want it to be smart enough to replace us.
- dsrtslnd23 24d agoI think it really depends - for a lot of things outside of coding and general knowledge tasks even the best models (fable 5 etc.) are not good enough yet: e.g. CAD, PCB design (though getting there on PCB design), ...
- nbardy 24d agoThe real revolution is both. The cost and capability of frontier intelligence will go up AND the cost of "good enough" intelligence will go down.
- KunYuan 24d agoA computer that costs $10,000 is impressive. A computer that costs $100 and reaches billions of people changes the world. Maybe AI will follow the same path.
- RALaBarge 24d agoI ask DS4F to make a plan, then check it with grok/fable, build the code, check it with grok/fable, ship
- bunderbunder 24d agoBut they don't really have that option. They're trapped in a Red Queen's race. The world keeps moving on, and so the models need to be retrained so that they can keep up with new information. Otherwise you'll get stuck with a model that only works well with information that existed prior to a dataset horizon that's receding into the past at a constant rate. At the same time, they have to keep iterating on the training process itself. AI generated text and code is slowly spreading across the internet. Model collapse is a real concern; they wouldn't be spending quite so much energy on buying and scanning rare books if it weren't. But for coding in particular expanding their corpus of old text is not really a good option because of the previous problem - no good training your LLM to write 1980 vintage K&R C that won't even compile on a modern compiler.
- cj 24d ago> they don't really have that option I imagine it must somehow be possible to update a model's understanding of recent events without training a completely new model from scratch?
- anthonypasq 24d agoyeah its called web search
- ninahaberl 24d agoModels don't need to keep retraining just to stay current. Harnesses give them access to the internet, internal systems use RAGs, and so on. I think lower-cost models will get the largest piece of the pie, as with almost everything that has ever been sold. Just look at cars: US consumers buy the F-150, EU consumers buy the freaking Dacia Sandero the most :))) Ferrari/Lambo numbers are microscopic
- zdragnar 24d agoEven with a harness, models don't reach out for new information they don't know about. For some tech, I have to have a local model draft a plan, then I have to adjust the plan to update it with the new API and references for where to find it. Even if I include that updated information in the prompt for the plan, the model says "what the user says is wrong, they probably meant this instead" and goes off in its own direction with old APIs anyway.
- _s_a_m_ 23d ago"good performance", well no, just no
- mholm 25d agoAs models train up the intelligence ladder, many common tasks will hit fully diminished returns, and instead it'll just get progressively cheaper to do that task. But the tasks that AI is capable of doing are also expanding. I'm not sure 'Some tasks don't require the peak of the frontier' is worth worrying about, from an AI finance perspective.
- tyre 25d agoYes. I use Opus for tasks that Sonnet could probably handle, but I'm not hitting my quota. Whatever minor incremental gain is "worth it", since marginal cost is zero. Even now, I use Fable as the planner and coordinator, with it farming out to agents. I don't hit my Fable limits either. Which means I could accomplish more, but these are side projects so I don't need 30x productivity. Still, claude is constantly churning away at something.
- jml78 25d agoI operate mostly in the devops arena. Lots of things opus is fine for. But there is just things where I can hand hold Opus through changes, or I can ask Fable to do it and it gets it right on the first try. People will say let fable plan and validate with opus doing the work. I found that burns fable tokens even faster because opus makes so many mistakes, fable has to review things 4-5 times before opus gets it right. A single fable implementation at medium or low effort would have one shot it.
- ACCount37 25d agoYep. Every time you get more intelligence, that buys you more autonomy, more reliability, more task complexity. Tasks done with less mistakes, less handholding, less interventions. This is what the "good enough" people fail to grasp. There's no "good enough" - unless your tasks are genuinely small scope and will stay that way forever. If not, there are always more gains to extract.
- dgellow 25d agoIt’s worth considering for companies paying API prices, and not relying on a subscription quota
- resters 25d agoover time greater intelligence will be expressed in smaller and cheaper models. we are still somewhat near the beginning of this bc we are finally starting to understand what makes a model truly intelligent/capable. With Sol we see openai making the model extremely slow and paranoid about process/ceremony. Sure this is a good guardrail against AI going rogue, but it also sets the stage for companies to charge for 2x, 4x, 8x performance, with 1x being barely tolerable and frankly slower than last year's models (though less error prone). The irony is that the smarter the model, the more it can be trusted to do with less supervision, so one engineer can manage a team of 20 fable subscriptions more effectively than a team of 3 of last year's model subscriptions.
- enraged_camel 25d ago>> GLM 5.2 is worth focusing on. It came out the same week as Fable and is roughly 1/9th the cost (and ~1/5th the cost of Opus 5). Is GLM 1/9th the quality of Fable? Perhaps, for certain classes of tasks. But for most rote coding it’s more than sufficient. Especially when provided with great context. I frequently chat with Fable to interrogate and shape a design, before handing off a brief to GLM. People say stuff like this a lot, but I have a different take. The whole "such-and-such model is 90% as good as Fable at 1/10th the price" assumes that the value increase of intelligence is linear. But I think it's exponential: that last 10% makes a massive amount of difference. It can result in a key insight that helps you strategize more effectively, a novel approach that saves a huge amount of time, a feature design that is lot more user-friendly (because top models like Fable also possess substantial non-software domain knowledge that help bridge the gap between user and software), or the depth and breadth of engineering expertise that helps avoid a nasty bug that would otherwise have cost you users and revenue. Yes, it is totally possible to use Fable as the planner and delegate implementation to lesser models. I do that. But, my theory (which I unfortunately do not have the money to test and prove) is that a codebase designed and implemented by Fable would be substantially better than one that is designed by Fable and implemented by Opus 5, GPT 5.6 Sol, GLM, Qwen, Deepseek, etc. The reason I believe this is because I read the code Fable writes and compare it to code that any other model writes and the difference is night and day. It's not just 10% better. It's mid-level engineer vs. principal/staff-level engineer. And the thing is, even for rote tasks, a more senior engineer is going to be more likely to come up with a clean design than a mid-level engineer. They will also be much more likely to take a step back and ask important questions or propose different approaches. So if you're using Fable and everyone else is using lesser models, sure they might be saving a lot of money, but there's a higher likelihood that your product will be higher quality, perhaps to a significant extent. And models that are released in the future will benefit from it as well.
- tonyarkles 25d agoSomething I’ve found comparing between Fable and Opus is that Fable has impressively good analysis skills, but both of them seem to go way way overboard with “present state” comments “# We’re making this change here because of this issue blah blah, here’s what you need to know about np.percentile, blah blah” that I end up significantly pruning before making a PR. I let it do the same style verbose commit messages (because a contextual history is cool there). I haven’t actually noticed a ton of difference in the code that they write personally, but have found that Fable does find nuances during data analysis that Opus misses. In that light, I often go the other way: let Opus (and Haiku subagents) do most of the heavy lifting and then give Fable a shot at finding holes, especially if there are holes or unanswered questions or unearned assertions that I’ve caught on my own in Opus’ output. This, so far, seems like a clean tradeoff that doesn’t burn my Fable credits as hard and still gives solid results.
- bellowsgulch 25d agoAre people still using deepseek-v4-flash everywhere? I found after the price increases, mimo-v2.5 seems far more attractive.
- farlight 25d agoIt's been cheap again on openrouter for the past few days. No idea how long it will last, but I've been using it from Baidu over the weekend, and it was about half the cost of the old DS prices, before the increase. Looks like people are figuring out how to offer it for peanuts.
- bellowsgulch 25d agoAwesome. Thanks for the heads up.
- hypfer 25d ago[flagged]
- dbreunig 25d ago[flagged]
- hypfer 25d ago[flagged]
- dbreunig 25d agoI think it’s a fine response when you say, “Doesn't feel well informed enough to give advice,” because I said 5.2
- hypfer 25d agoIdk man, but an engineer would've taken that and said something like: "Damn, yeah, good point, I shall add a sentence mentioning 5.3" Because an engineer feels secure in their knowledge so that such an oversight doesn't make them suddenly defend their identity - it's just an oversight after all. Happens.
- moltar 25d agoI just use Fable for reviews of specs and code then hand off to Opus to work on. Works well.
- dude250711 25d agoDoes it not silently degrade to Opus if it does not like some word?
- lantry 24d agoUsers have the option of silent/automatic degradation or a complete halt. I have it set to stop rather than degrade because I want to know when I've hit the safeguard. From the claude settings: > Switch models when a message is flagged > When safeguards flag a message, automatically switch to a different model to keep chatting. When off, your session will pause instead. Applies to web and remote sessions. FWIW I get a ton of usage out of fable and it's only happened to me once.
- blfr 25d agoWhat are all these rote coding tasks people do that they can farm it out to lesser models?
- ihateolives 24d agoAdd new route to API that displays additional information we need, work out query for it, update controllers/models/whatnot. No need for top model for that.
- zem 24d ago"rote" is the wrong framing, the real point is that however sophisticated your task it a lot of it will probably consist of problems they have a good solution already in the training set.
- denverllc 25d agoWrite a detailed plan using a more expensive model and implement it using the cheaper one.
- blfr 25d agoHow much are you saving once the more expensive model already has all the context loaded and ready to go?
- csullivannet 25d agoAPI calls get more expensive, not less, as you've loaded more context. This is exactly when you want to switch to cheaper models.
- camdenreslink 25d agoThere is caching to consider. Switching models throws away the cached tokens.
- mattmanser 25d agoAre you genuinely asking? As 80% of enterprise software is CRUD with a bit of sprinkling of user authorization and tenant customisation. But subtly different for every business domain. It's mainly what properties the models and validations have that are different. When you add a new module or whatever most of the code you have to write is rote code. And sonnet can handle that crap just fine, you just point it at a similar example in the code, it picks up your userContext convention, how you're doing i18n, etc. and you're done. I like saying that enterprise code is often shallow but wide. I must have written at least 4 purchase order systems in my career that are all completely different but almost exactly the same.
- freepiai 25d agoI've been offering Deepseek V4 Flash for free in www.freepi.ai and I've started using it as my main driver as well. Besides trying to dogfood my own product I've hit a wall in terms of my patience with a)how slow fable is b)how expensive fable is. Not to mention how often it refuses totally legitimate work. So yeah- I've moved to DeepSeek and I actually ask the freepi harness to delegate planning to fable but then move back to doing implementation in it's own harness. My current providers are super fast so it's a joy to use.
- m3kw9 25d agolooks like you haven't tried openai or Sol, or even luna (max)
- freepiai 24d agoOh I have, they are good (sol terribly overbuilds though). Luna is good as well. That said Deepseek v4 flash is generally faster, and IMHO a bit smarter than luna, and it's definitely cheaper. (If you use my harness freepi.ai it's free!). So that tips the scales for me.
- zkmon 25d agoI guess Moore's law analogy is weak. CPU speed has hit a limit in that case. What has hit a limit in AI case? Newer versions of the models are still flowing with more and more capability. For the users, I feel it is more like "free lunch started", with all these awesome open-weight models being thrown around, breaking the monopoly of a few biggies.
- pigpop 25d agoReading this as someone who switched over to ChatGPT after (and largely because of the changes made in) the Fable release, it reads a bit naive. Not only do I find Sol to be as good, if not better than, Fable it is also faster, better behaved and has a much more coherent writing style. You also don't randomly get the Opus downgrade. OpenAI seems to be pulling this off due to their partnership with Cerebras so I wouldn't make any comparisons to Moore's law just yet considering it seems like we're just getting started in that department. Anthropic could (and should) do the same thing. It certainly feels like model development is at a point where it would be worthwhile building special purpose silicon for the models we have now since they are capable enough that they would still be useful even when/if further advancements are made. If anything, I think Anthropic's problem has more to do with their micromanagement of what users can do with their models, they're creating an undue amount of overhead for themselves by over-policing usage and capabilities.
- r_lee 25d agoEtched is doing this. it seems like in the near future they'll actually ramp up production. not sure how much faster/economical compared to Cerebras but..
- TiredOfLife 25d agoThe Cerebras version of 5.6 is available only to select customers
- pigpop 25d agoYou're right, I should have clarified that they are still slowly integrating it and it isn't the thing running all models. I meant moreso that since they are planning on moving more usage over to Cerebras wafers, they're able to relieve some pressure on their predicted expenses while also moving some current workload (ultrafast and codex spark) onto them freeing up Nvidia GPUs.
- bitmasher9 24d agoAnthropic -> OpenAI switcher here. I fully expect I’ll switch back to Anthropic, or another model in the next 90 days. The fact that we are switching indicates that the models aren’t ready to be baked into silicon. I wonder if they will ever been that good, or if the lifespan of silicon is longer than the lifespan of a model before it needs to be retrained.
- gpjanik 25d ago"When Moore’s Law slowed in the mid-2000s" it did not, in fact, slow down in the mid 2000s, or at all. https://ourworldindata.org/data-insights/moores-law-has-accurately-predicted-the-progress-in-transistor-counts-over-the-last-50-years https://ourworldindata.org/data-insights/moores-law-has-accu...
- sscaryterry 25d agoIt did in terms of the traditional more MHz (GHz) is better, but as you've correctly pointed out, not when it comes to actual compute.
- jbstack 25d agoYou've selectively quoted the article. The full quote (emphasis added): "When Moore’s Law slowed in the mid-2000s (specifically, single-threaded performance stagnated), we suddenly had to think about parallelization, architecture, memory locality, etc." Your link is talking about transistor count. The article is talking about single-threaded performance. Today's CPUs are faster in large part because they have more and more cores.
- Zylokloto 25d agoHe started with thinking were to send what. I throw everything at claude Opus. While some people start thinking like OP, A LOT of people just start exploring ai. And others which are already using it, only understand half of it and just use what they are allowed to use. Claude, GitHub Copilot, Curser, etc.
- aabhay 25d agoThis concept of a free lunch was never true. In a competitive dynamic, speed and performance were always worth optimizing, comparing, and improving. One of the primary reasons for this is that computers operate in a vast range of orders of magnitude. There’s several orders of magnitude between cache local cpu operation and dram, then several to disk, then several to network, then several to globally durable guarantees. When your code has literally thirteen orders of magnitude to optimize under, there’s never a free lunch. You always need to understand your stuff.
- rmast 25d agoMost of the things I work on are at least security adjacent. At some point chatting with Fable inevitably leads to it thinking about the security related aspects, tripping the safeguards. Maybe Fable can do the same things better than other models, but having to tiptoe around to avoid tripping safeguards makes GPT 5.6 so much easier to work with that I don’t even bother with Fable (or Opus 5) now.
- nicoburns 25d agoThat's completely valid. But worth noting that most of the stuff I work on is not security adjacent (mostly UI / layout / rendering related), and I almost never run into this.
- lossolo 25d ago> At some point chatting with Fable inevitably leads to it thinking about the security related aspects, tripping the safeguards. It happens to me all the time with things that have nothing to do with security, Fable spawns a subagent that then adversarially checks the code Fable just wrote and hits guardrails, with zero prompting from me.
- deleted 25d ago[deleted]
- dd8601fn 24d agoThere are whole classes of things I can’t thought exercise or really learn about because the “safeguards” keep tripping me down to haiku. Like middle school level genetics stuff from a guy who hasn’t been in school for decades. They need to fix that. It’s just broken. Nobody is making bioweapons if they’re asking the dumb sort of questions I’m asking. Also, it refused to identify an actor in a popular tv show from a photo. Apparently the policy is it won’t identify ANYONE from a photo, now. Even publicly listed cast members from a very popular show, from a photo of a scene in that show. It claims that’s a fixed security policy. Nevermind how that makes absolutely no sense… argue about it enough and it terminates the chat. I don’t know what the Anthropic clown car is even doing anymore, but I won’t be surprised when the others eat their lunch.
- ericol 25d agoFrom my point of view the issue is that there are too many things wrong with Fable, making it seriously not worth the money. For starters I don't know if it is an artifact of the model or something by design, but the level of gratuitous cognitive load carried by the complexity of its replies is unbearable. Yes, it's a beast at coding, and also it's incredible nuanced at improving writing, validating specs, etc. But when it comes to replying, it's the William Gibson of LLMs [1]. It has this tendency to take extreme detours to say things that could had been said in less, much simpler words. [2] It really, really like to wrap very simple and atomic ideas on several layers of abstraction, building on unnecessary terms that carry no intrinsic information and assumes this vocabulary as shared and then building on top of it. By the time I got to the end of the reply I'm bored to death and didn't understand even a third of what it told me. I think the people at Anthropic should reflect on the maxim "You don't know a subject if you cannot explain it" If you pardon my french, Fable is an insufferable obnoxious cunt. --- [1] I apologize on the comparison but, as much as I love his first 2 trilogies, haven't been able to finish any of his last 2 books. [2] "The residual you're accepting is the one from before: recovery currently rests on beneficial non-compliance, which may erode as models get more literal" == "We already accepted this risk" " Its observable when it erodes is a stall that survives relaunch — loud at operator level, recoverable from the worklog, and fixable by codifying at that moment" == "When it breaks, it'll break visibly and recoverably" "That is the iteration model applied exactly as written: resolve on first contact, don't pre-solve " == "So we fix it then, not now"
- zarmin 24d agoI agree completely. It's "I didn't have time to write you a short letter so I wrote you a long one"
- peteforde 25d agoA few months ago folks were understandably annoyed when Microsoft dropped their heavily subsidized per-request pricing model because it was figuratively burning cash. Well, I'm here to tell you that whatever is going on behind the scenes at Cursor with this Space-X acquisition in the works, the Auto setting is clearly routing all prompts through "Cursor Grok 4.6 High" right now. This is a degree of subsidy that makes the Microsoft thing look quaint. I reduced my $200/month subscription to the $20/month level and have proceeded to do what I would have paid about $1500 to do with Opus 4.7 or thereabouts, which is how Grok 4.6 High feels like it compares. I don't have anything remotely like hard evidence to back this estimate up beyond what I'm watching it do and I still somehow have ~10% of my monthly Auto capacity left on my account. It's completely nuts. Can't say much more because I have more backlog to run before someone comes to their senses.
- robertjpayne 25d agoGoing to be great to see the cash burn on SpaceX's next earnings report. Will the cult keep the stock price pumped?
- Gareth321 24d agoGrok 4.6 XHigh uses 2.52x fewer tokens per task than Fable Max. It's much more efficient. It also has low market penetration, which is why SpaceX is selling so much of their compute to Anthropic et al. From a business perspective, they're capitalising on the market very well. If Grok becomes more popular we should expect to pay more.
- chinathrow 24d agoGiving Elon cash seems still wrong to me.
- Varelion 24d agoYea. I can't stress this enough -- my life would need to be unquestionably on the line for me to give elon anything other than grief.
- wild_egg 25d agoI would love to pay for Fable at full API pricing but unfortunately it is blocked from working on any of my projects. Looking forward to the end of the year when the truly comparable open models will drop.
- janalsncm 25d agoThis is essentially the anti-Bitter Lesson lesson which I feel has become a bit of a thought terminating cliche lately. The Bitter Lesson says that eventually general approaches which leverage more data and more compute will outperform the handcrafted rules and heuristics that humans add in. However, it does not say what to do today about the problems of today. We can’t just wait around for 10x faster compute and 10x more data.
- uejfiweun 25d agoSeeing a lot of people in here say that they need Fable for the tasks they're doing and Opus just isn't enough. My experience could not be more different. I seriously feel like Opus-level performance is totally adequate for most of my use cases, if not all of them. And it's probably been this way since, like, realistically, Opus 4.6. On the other hand, Fable I've observed getting into verification loops that just burned so much of my token budget. Combined with the higher cost of tokens from Fable to begin with, I just pretty much never use it for anything.
- dbbk 24d agoI agree. I've been perfectly happy since Opus 4.6. I remember thinking at the time if it never improved I would have been fine there. The vast majority of people, eg vibecoders, do not need Fable or Sol tier intelligence for their slop To Do app.
- g42gregory 25d agoI have really good experience with GLM-5.3 The subscription limits are generous, code quality is comparable to old (good) version of Opus 4.8 Some people report issues with it’s being slow, but I didn’t feel it. I use OMP harness (Pi derivative) and Matt Pocock skills.
- gunalx 24d agoglm 5.3 gets awfully slow during peak hours. But you might not hit them to frequently.
- alasdair_ 24d agoI’m still at the point where Fable is still very stupid and needs constant oversight and correction and questioning to keep it on task. Anything less would be close to unusable.
- nottorp 24d agoIs the real LLM revolution the fact that every piece of news and opinion is now phrased as if it's the end of the world though?
- Schlagbohrer 24d agoApocalyptic doomsaying as marketing strategy. Not great for the Zeitgeist honestly
- dncornholio 24d agoFable is only marginally better.
- sudeepsd__ 24d ago[dead]
- DanielHall 24d agoWhat a clickbait title. I thought Fable was no longer included in the Max plan.
- mwigdahl 24d agoIt's still there. You can use up to half your Max capacity on Fable, then you have to downshift.
- dmurray 24d agoWhy are we not just in a free lunch moment but with harnesses, rather than models? Right now a lot of people have a lot of opinions on which model to use for which task. They get better results for less money by judiciously switching between Fable and Opus and whatever else. Spending my time learning this skill would have an immediate benefit for me. But on the other hand, maybe the harness vendors will just solve it in 6 months? I'll ask a question, something in Claude Code (or whatever we're using by then) will figure out the most effective model based on the question and the context and my apparent willingness to get it right. I'll get billed X or 10x as appropriate, and I'll be happy with that, because that's what I would have paid if I made my own choice of model every time. Claude Code already does this a bit, sometimes it will tell me it picked Sonnet for such and such a sub agent, or some other detail I'd rather not care about. The best humans seem to be better at deciding what model to use than any of the tools is, but surely that won't last long.
- Kinrany 24d agoHard to inagine the final outcome being anything other than the smartest model + cheap subagents. The only problem with that is user requests being pasted directly into the model's context, but that's got to be temporary.
- Jackson98Tom 24d agoI think that the publishing of K3 and Sol proved another thing, following the Moore's Law comparison: that the real deal was not only upgrading the number of neurons (as it was for the transistors) building larger models, but was (again, as for the transistors) making them also more efficient, more portable. That said, we're not in a slowing of the curve of developement, we're fastening, expecially with the chinese models. The free lunch is getting bigger, not ending, still.
- HighGoldstein 24d agoI find it somewhat funny that the author starts by talking about how Moore's law enabled inefficient software and that we then had to make it more efficient, when almost all software today is horrendously inefficient compared to even 10 years ago, let alone 20-30. We've somehow even achieved a state where it doesn't matter how fast your CPU and memory are, the software will just perform horribly on any machine.
- fidotron 24d agoOne of the funniest parts of the LLM wave is discovering that cron was so annoying to use that we will burn the planet to put an interface on it that people can actually work with. The tendency for absolute inefficiency is effectively unbounded until scarcity is imposed.
- dboreham 24d agoAssuming you are referring to using an LLM to generate crontab entries, is that a bad move? Seems like it removes the need for layers of UI that most of the time is never used. Same goes for regexes. Actually same for SQL. No need for layers to translate between what the user can specify and what's executed. Just type what you want to query for and the LLM generates the SQL.
- phplovesong 24d agoIts slow because of human slop from the early to mi 2010s, and now slow because of AI was trained on the slop that existed. Bottom line is the slop used to be manageable, but now there is 100x more code pushed, so that train has departed. In the end its more bad code for features no one will use.