16 ms·
(I work at Anthropic) Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claud
by felixrieseberg 16d ago
(I work at Anthropic)
Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier.
Another point I expect not to get much attention until it all happens at once is science. People have been correctly excited about the many "sudden" breakthroughs LLMs are making in Maths, but some of the science benchmarks make me believe we'll soon see similar developments in other scientific domains. Fable 5.1 more than doubled Fable 5's Terminal-Bench-Science [1] score, which I think is meaningful.
[1] https://github.com/harbor-framework/terminal-bench-science https://github.com/harbor-framework/terminal-bench-science
- mingqiz 15d agoBy injecting that weird prompt and not by proper post training? anthropic is truly a joke.
- Waterluvian 16d agoHow much of the language style outcome is a well-crafted result vs. being a somewhat unpredictable outcome of mucking with levers and knobs for a while?
- themagiceye 6d agoGuys
- generalizations 16d ago> similar developments in other scientific domains The classifier is too strict. It's rare to be able to complete a project without being permanently relegated to Opus. I'd expect that the domains where this accelerates progress will be fairly limited.
- krull10 15d agoYeah, this feature is only useful to scientists who work at institutions that have deals with Anthropic to use the models without the classifiers.
- blondie9x 16d agoAre the models improving their footprint on the natural world? Data centers and and the natural resources consumed by models for production of materials and for building and running inference servers are contributing towards environmental degradation. How can we prevent that as we continue the roll out so we shift this to a more sustainable developmental rollout path?
- matheusmoreira 16d agoBut is the model actually going to answer hard questions when we ask them? Or are you going to keep downgrading the models so as to avoid "uplifting" lesser lifeforms like us?
- ryandvm 16d agoGreat. I'm looking forward to it being less obvious that my colleagues have stopped understanding their jobs.
- synergy20 16d agoclaudism really sucks, Gemini and codex output so much better, way more like a real human being.
- deleted 16d ago[deleted]
- cantalopes 16d agoThank you for the trust me bro benchmark but i will be honest, fable 5.0 did even worse thsn 4.8 opus
- digitaltrees 16d agoAnd that’s what changes the whole game — Claude
- razster 16d agoStill not going for it. Once I learned I can train Qwen3.8 27B with my style of writing/grammar. Also more succinct. I cannot force myself to Claude or OpenAI outputs anymore. Its too much. Honestly don't think I will ever go back to paid.
- avazhi 16d agoMan, it’s like you and I are using very different versions of Qwen. In my experience in English Qwen is the one model that consistently lapses into using incorrect English in its responses. Like, its training corpus was clearly (unsurprisingly) lots of non-English material. The random Chinglish is jarring. Even small models like Gemma 4b write much better than Qwen.
- loloquwowndueo 16d agoWhat’s your honest take on how load-bearing its use of em-dashes is now? Measured, not guessed.
- ddahlen 16d agoThe writing style has significantly improved, however the token burn rate for tasks I have been working on seems to have skyrocketed. It definitely appears more capable (though I am unclear how much of that is just me liking the English it writes now vs actually more performant). I was using Fable 5 for some mathematical analysis assistance and redoing a part of it with 5.1 burned 60% of my session at a much faster rate.
- jmann99999 16d agoThis. It seems to light my usage of my max plan on fire. I’ve gone back to opus because I run out of usage in my five hour window so much more quickly.
- LimitExperience 15d ago[dead]
- iamflimflam1 15d agoI really hope the improvement in natural style is real. When I’ve tried to adjust the output style is that initially it feels better - but that’s just because the new output is so refreshing to read after the horrible Claude output. Unfortunately, after a short while you quickly realise that it’s just as vacuous as before the style change.
- lofties 15d agoI don't want my Claude to sound "natural". Claude is a robot and it should do behave like a robot. It should do what it's told. Nothing more and nothing less.
- ActionHank 15d agoGood news for you is that vastly cheaper models can do this much more quickly. Bad news for Anthropic and investors is that vastly cheaper models can do this much more quickly.
- illusive4080 15d agoHow aware internally are employees that Opus 5’s language is incomprehensibly complicated? Please fix with Opus 5.1.
- NL807 15d agoNot sure if something like this is already on the table, but I would like to see Claude responses more in line with Simplified Technical English [1] by default. I find those writing styles a lot easier to read. This has been standardised as ASD-STE100 [2]. I've seen few people making SKILL.md files with that in mind, which works great, but having this by default without invoking the skill command would be better. 1. https://en.wikipedia.org/wiki/Simplified_Technical_English https://en.wikipedia.org/wiki/Simplified_Technical_English 2. https://asd-ste100.org/ https://asd-ste100.org/
- smashed 15d agoI must have missed something but can't you just prompt it to answer in your desired style? What am I missing here. Commenting because I am struggling with this too, claude code seems to be so verbose no matter how I prompt it.
- lukan 15d agoYou miss that it is not just your prompt but also the various system prompts, plus how the model was trained. But you can reduce the verbosity (also with a setting in /config).
- kolinko 15d agoAnd memory and code comment styles - i think that’s a big one people forget about. You can prompt it all you want, when it sees elaborate comments in memory and code it will follow the style
- lukan 15d agoBut you can prompt it to reduce the verbocity of the comments as a project in itself - I will probably try Fable 5.1 for this.
- spuz 15d agoNo, training tends to override system prompts at lot of the time.
- DarmokTanagra 15d ago[dead]
- hollowturtle 15d ago> People have been correctly excited about the many "sudden" breakthroughs LLMs are making in Maths That people ARE making, surely not machines. Like Terence Tao or Knuts did using the tool to their advantage, for example it would have been impossible for me to prove the same thing Tao did with an LLM. Same reason I believe programmers won't go away
- krull10 15d agoThis isn’t true in math; see the proof of Crouzeix’s conjecture which was done by GPT 5.6 Sol in response to a prompt from a neurosurgery resident who had no deep math background, was learning that subject to better understand radiology, and thought it sounded like a cool theorem.
- hollowturtle 15d ago> Jin reported that the proof was obtained with the assistance of OpenAI's GPT-5.6 Sol model during an approximately sixteen-hour autonomous reasoning session in ChatGPT Work, after which he checked the resulting argument You're wrong
- VeejayRampay 15d agosince you work at Anthropic, know that there was (warranted) love for your models from the community as a whole, they performed well and added value but the verbiage in recent iterations is absolutely insufferable, I will stop using them because of that as soon as I can, I simply cannot stand another round of the model "finding the smoking gun", saying "that's the actual gap, not a fluke" or some idiotic phrasing like this
- NamlchakKhandro 15d agohow are you going to be profitable?
- testfrequency 15d agoProject Panama [0] Must be helpful that your company is slurping and destroying literature, how sad that the results of this are a blog post with “look how well we write English”. Eye roll. Could not be happier about my decision to turn down a job offer from Anthropic years ago. Ick. [0] https://en.wikipedia.org/wiki/Project_Panama https://en.wikipedia.org/wiki/Project_Panama
- jens_tlb 15d ago[dead]
- jamaliki 15d agoCongratulations on the release. As a scientist working in biology, I cannot take the supposed prowess of Fable seriously until I am actually able to use it for biology. Currently, Fable is completely incapable of helping with any biology related task, however tangential.
- kaoD 15d agoAnybody knows why Fable is railguarded in particular against biology tasks? I'm out of the loop here. Is it drugs? Bio/chemical weapons?
- AdamN 15d agoIt seems like the different AI companies should lean into their 'blend' in terms of AI speak. The analogues for me are spaghetti sauce or coffee. Starbucks for instance has a particular roasting style that you can guess 100% of the time and it adds a certain consistency to the customer experience even though it doesn't encapsulate the full world of coffee. Similar for model responses where the 'blend' should be nurtured over time and consistent even if the underlying processes change. That is, once the right blend is figured out - which may not be the case yet.
- ashkankiani 15d agoI canceled my Claude subscription, though I did get some utility out of it, because of how much steering was required to use it on complex projects. A big reason being that anyone who is using Fable seriously will run out of usage limits very quickly, and so will lean on the "Fable for review + design discussion, Opus 5 agents for implementation" paradigm. But an incredibly annoying UX problem is that the resulting report from the agents that Fable reads isn't surfaced to us in the main dialog, it's only summarized back to us (unless you idle in the agent's window to avoid it closing so you can read what it said directly). As a consequence of this game of telephone, the Fable agent will start using some "terms of art" that it and the agents invented, leaving out literally all context that would be useful in helping me understand what converged/diverged from the implementation attempt. It will often try to ask me for input or say that I have to deliberate on something while also referring to things I've never seen (from the agent result) and without providing any context. I have to repeatedly prompt it to verbosely explain every time (putting it into the system prompt did little to improve this) and remind it that I can't see what the hell it's talking about. I'm not sure I'll re-subscribe or even really use AI again because it's honestly more frustrating than it's worth, and so the net emotion I'm left with is frustration and without the satisfaction of learning + building something myself. But at the very least, I thought I'd give someone at the company a tip on what seems to me like a common and obvious UX/UI/workflow failing for using Fable, as some last bit of good will.
- freepiai 15d ago[dead]
- surrealize 15d agoI gave a standing directive to my coordinator to process subagent transcripts (with a simple Claude-written script that reads the transcript .jsonl) and save the result. I also follow along on the issue tracker; that really helps me understand WTH they're talking about, and the subagents are also directed to post their shipped notes there.
- GPerson 15d agoCan you quit your job? You guys are destroying everything good about life for little payoff except to yourselves.
- NoMoreAds 15d ago[flagged]
- 2ManyClaudeAds 15d ago[flagged]
- andsoitis 15d agoI recently ended my Claude subscription, returning back to ChatGPT because I could no longer bear Claude’s prose, finding it excessively verbose, robotic, repetitive, and condescending.
- not_a_bot_4sho 15d agoI use GHCP but similar sentiment. Stopped using Anthropic models for this reason. Their prose become too obtuse and just... alien. No human talks or writes like that. It's incredibly taxing to deal with. Sticking to a mix of GPT and Gemini for now.
- miroljub 15d ago(I don't work at Anthropic) What a surprise that someone working for the Anthropic marketing department roams social media to praise every single Anthropic release :) On the other hand, what I find more worrying is that this is the top comment here on HN. I can't believe such an unsubstantiated marketing post can get so many upvotes to be the top comment.
- enoch2090 15d agoIt's totally valid if the models want to pack words tight during their thinking process, as long as the final conclusion (which is the interface to user) is written in HUMAN LANGUAGE, then I don't care whether the model thinks in alien language
- xcafebabe 14d agoplease extend the +50% promotion xD
- chmod775 14d ago> Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. I agree it's better, but dangling possessives are still an unwelcome affectation of fable's. Sentences like "When a node becomes an object's." (real comment) are unnecessary mental load to untangle.
- deleted 13d ago[deleted]
- ohyes 11d agoYou forgot the “co-authored by Fable 5.1” line on your post.
- behnamoh 16d agoAt this point, I don't believe a word from Anthropic employees; you guys have lost all the goodwill that you accumulated over months last year.
- chews 16d agoI share this sentiment, I really did like the models... then the finger printing, encryption of thought traces, staggered access, the constant NO's from Fable on cyber related issues for looking at bugs in my own code... I'm glad I swapped to Kimi/GLM... now with the deepseek harness, I don't even miss Claude Code. I really hope open models give them the market reckoning they wholeheartedly deserve.
- nullstyle 16d agoHave you shared any details about your dsh setup anywhere? I’ve only dipped my toes in and would love someone else’s perspective on how they use it
- chews 16d agoI've not, but really should. I run it on exe.dev, it's an ephemeral VM company and they have an agent of their own called shelley (which I used locally as well), Having kicked the tires on DSH(deepseek harness), I ported Shelley's skills into DSH, they are pretty simple text files that were easy to bridge over, it is more verbose but the plugin nature of it was really easy to extend, for example, I built a plugin that checks my claude usage windows and when I get to 80% stop asking new agents for help.
- rvz 16d agoI don't think they care. It is up to you to consider local models or better alternatives instead of paying for more tokens at their casino.
- dolebirchwood 16d agoDon't worry - I'm paying for our friends overseas to keep their distilling operations going.
- kvakkefly 16d agoI hope not! Then my t-shirt is no longer accurate :D
- philipwhiuk 16d agoIt's being written with Claude so I'm wondering how much of that is just using the repo as training data: https://github.com/harbor-framework/terminal-bench-science/commit/1705d3e3783c57eeec4756a1b11d99e4bf31f6d3 https://github.com/harbor-framework/terminal-bench-science/c...
- belval 16d agoAs a fervent Claude Code user who made the switch to GPT 5.6 Sol over Opus 5 over hard-to-read prose this makes me happy. I love your product but the current models are very hard to work with if you need to do a lot of context switching. Brevity is key.
- pixl97 16d ago>Brevity is key Which is something the providers that are trying to watermark their texts can't afford. Superfluous replies give much more opportunity to further encode this junk information.
- ctoth 16d agoThis ... is not how this works. The model is not speaking longer to watermark anything.
- skarz 16d agoPerhaps, but there are certainly now catchphrases and words that can indicate it was written with AI i.e. load-bearing, idempotent, etc. Style and structure are in and of themselves, a fingerprint.
- nick__m 16d agoidempotent was frequently used before LLM; it's hard to talk about REST and infrastructure as code without using that word...
- TheOtherHobbes 16d agoIt's exactly how it works - at least potentially. Lean text is harder to watermark because word choices and meanings are tightly constrained. Low-entropy text is fluff and filler. It's very easy to synonym-substitute words without changing the message - if there even is one.
- 16d ago
- Trasmatta 16d ago> More work to be done (and we will!) but reading better prose makes me so much happier. I assume this work will be done for Opus as well? Opus has seemingly gotten progressively worse at its prose and technical writing with each version. I've stopped using Claude entirely for now, because it manages to turn even the simplest technical explanation into the most obtuse and obfuscated word salad imaginable. People originally adopted Claude because it felt pleasant to use in comparison to ChatGPT, but I feel like that's really been lost (at least with the Opus line). I feel dread when I see a wall of text generated by Opus. Every developer I've talked to feels similarly right now.
- LimitExperience 15d ago[dead]
- sroussey 16d ago> I feel a sinking feeling of dread the moment I see a wall of text generated by Opus Agree, Claude lost the joy of using it. That is a measure that ranks higher than any other benchmark at this point.
- Trasmatta 16d agoYes! Claude was so pleasant to use at first. It was Anthropic's biggest advantage. And now it's like nails on a chalkboard.
- vardalab 16d agoYeah, it's like day and night. It used to be really unpleasant to interact with early codex versions. Even 5.3 wasn't great. Now, I go to Sol if I need to discuss anything. I don't even bother with Opus because I know that it's going to give me a headache.
- dezgeg 16d agoYeah, for all the hate Gemini gets, at least it isn't obsessed with adding comments and it's output is more readable than recent Claude's.
- sroussey 16d agoPlease bring to the other models, and also please only apply the AI text watermarking only to EU citizens. I may not be able to tell when Claude writes about things i don't know, but in CC it writes about my code and it is obvious.
- iamflimflam1 15d agoMy reading of the law was that watermarking is not required by it at all. It’s a convenient excuse for the companies that want to add watermarking.
- sroussey 15d agoMy read was that watermarking is not explicit requirement, but that it could be in something like metadata that goes with it (if generated a word doc, for example). But writing comments in your codebase, the EU will want a digital trail there.
- tyrabound 16d agoYou’re probably better off organizing a campaign to pressure Congress to prohibit American corporations imposing foreign laws on Americans, which is what this text watermarking is, regardless of how you feel about it. I think it’s a precedent we really don’t want to go down if you believe in democracy and self-determination. It also clearly establishes or the very least moves in the direction that you don’t actually own or control the output of AI in any manner whatsoever, you’re just paying for it since Anthropic in this case can simply essentially brand/tag all your output that is based on not directly your own words, but a higher level process or methods that you use, including your instructions and how you structure your information and what your overall objective and goal is. Anthropic is branding it on the behest of the EU lew, which already is an entity that is diametrically opposed to democracy and self-determination based on its structure even if you ignore the fact that it violates the most fundamental concepts of self-determination in its direct contradiction of the UN Charter and implicitly the Universal Declaration of Human rights. What people done seem to be catching onto is that the EU is becoming the world dictatorship because the USA has simply had too many onerous people and that stupid constitution and its amendments that keep roadblocks world domination for the ruling class vampire.
- saaaaaam 16d agoHello Felix. Can you say why my additional usage credits have suddenly vanished? [edit] only asking here as last time I raised a support request it took six weeks before anyone responded.
- alasano 16d ago> It sounds a lot less stereotypically like other Claude models Don't give me hope. I've strained eye muscles from rolling my eyes so hard every day at how Claude writes. Edit: first discussion with Fable 5.1 "This is the right question and it needs a real trace, not a guess." Sigh.
- nozzlegear 16d ago[dead]
- darksim905 16d agoDo people not bother with style-output and custom definitions? Wild.
- finnnk 16d ago[flagged]
- 321ahT 16d agoHow is it possible that all models from xAI, OpenAI, Anthropic, Qwen etc. win all benchmarks on each release? Tomorrow all of the above (except Anthropic of course) will bump version numbers and be at the top of HN winning all benchmarks. Science breakthroughs incoming? First of all, you are already restricting science in Fable, secondly, we have been hearing the same for several years now.
- pohl 16d agoThere are hundreds of benchmarks. You just need to pick a favorable dozen on release day.
- deleted 16d ago[deleted]
- jbverschoor 16d agoWill it respond within a reasonable timeframe? It’s like we’re on a 14K4 modem when there’s broadband
- velcrovan 16d agoI have a pet theory that the Opus prose style/smell we all have grown weary of is due at least in part to the models writing more for themselves and each other than for humans. They're packing lots of signal into fewer words and they don't care if it sounds cringe because it works better as glue in long-running tasks. I'm also thinking of the 2017 novel "Void Star" where AIs who operate everything have long since left ceased bothering with human languages, and it takes a rare sort of direct matrix-gazing savant to be able to try and horse-whisper them into doing or revealing anything they didn't already plan to do.
- bbg2401 16d agoIf anything Opus prose packs more noise than signal. It's a string of platitudes, jargon, buzzwords, etc.
- 486sx33 16d ago[dead]
- le-mark 16d ago> They're packing lots of signal into fewer words I think opus is more noise and less signal actually.
- camoby 16d agoVoid Star? I’m reminded more of “Dark Star”, arguing with the ship’s computer. :)
- Vanclief 16d agoI support this pet theory, I tried out to reduce the output of Claude models with a "ADHD" prompt that made its responses small and to the point, but I could notice it degraded in performance as the session went on. So I think what is going on is that because responses are part of the context window, those long/technical responses help it keep focus/attention.
- catlifeonmars 15d agoI would not consider Opus output to have a particularly high signal to noise ratio.
- PedroBatista 16d agoThis post and comment makes me believe "science" is the new "code" for Anthropic now that the code advantage is mostly gone and lost for OpenAI, ie. they got much better and Claude become significantly worse over these months.
- echelon 16d agoIMO, Codex is worse than Claude with Fable. At least at Rust. That said, the open source models are not bad and I'm looking forward to more tools and products built on top of them. Code review, security review, etc. Anthropic needs to change how it treats users though. I'm increasingly put off by Dario, the rug pulling, the lies, and the attempts to regulate open weights. I'm going to bail if this doesn't change. There's plenty enough that's good enough, and those things are hackable and extensible. If Fable isn't available at subscription price via third party harnesses soon, I'm also going to bail.
- ed-is-ai 16d agoThe big issue I have with Fable is this. From the Anthropic email announcing Fable 5.1. So basically they're giving us a Ferrari, which will point blank refuse to do certain stuff - forcing us to go out in our Mustang. Their choice, not ours "Safeguards and automatic fallbacks (beta): Fable 5.1’s biology and cybersecurity classifiers block fewer benign requests and now permit vulnerability finding in source code. Blocked requests return an error and are not charged to you. On the Messages API, opt in to fall back to another model so users get a response instead of an error. We recommend Opus 5 for biology and Opus 4.8 for cybersecurity. In Managed Agents, fallback is built in."
- ImprobableTruth 16d agoIts "pure capabilities" are definitely worse than Fable, but I find codex has a much more pleasant style and is in comparison much more generous with its limits.
- dmix 16d agoCodex (+Sol) feels a lot more human for sure. Fable 5 is so, so wordy.
- comex 16d agoToo bad. I see the stereotypical prose as a good thing. When I interact with Claude myself, I don’t mind it as it just feels like Claude’s distinctive voice. But when other people try to disguise LLM output as their own thoughts, the voice makes it easier for me to tell.
- recursive 16d agoPeople that want to be open about the source of their text will just tell you where it came from. People that want to obscure the source of their text would rather that it was more difficult to sniff out LLM-generated text. And they're the ones picking which model to use.
- unshavedyak 16d agoI wouldn't mind it either. But the prose is obtuse atm. It doesn't feel like a writing style, it feels like an encryption.
- latentsea 16d agoQwen is all you need.
- Bluestein 16d ago⎿ You've hit your session limit · resets 2:51am (123°24′W Etc/GMT+8) /upgrade to increase your usage limit.
- areoform 16d agoHey Felix, I'm really glad for that! And I appreciate that you're making yourself available. I really do. Outreach is amazing. And thanks for making Claude. I really do love Claude. In some ways, I'm asking this question because of just how much I am grateful for the role Claude has played in my life. > Fable 5.1 more than doubled Fable 5's Terminal-Bench-Science [1] score, which I think is meaningful. But my honest question is, can I use Fable like that? Can I use Fable to do science? To borrow a Claude-ism, this is "load-bearing" because Claude's response has been degraded for innocuous research projects concerning population-level analyses of astronaut health. These "safety filters" trigger on questions about rabbit sex, smartphone accelerometer data to classify cat purrs, and so much more. What exactly does this score mean for users like me if it's unusable for middle school physics, biology and chemistry? Second, I would happily quantify it for y'all, but qualitatively it feels like Fable's performance is noticeably poorer than initial release / launch. And I am wondering if this is the case particularly for me because I use Claude via Claude Code to make a personalized care dashboard for my doctors to help me in managing my care. As I noticed in the upgraded filter announcement, https://www.anthropic.com/news/improving-fable-5-s-biology-safeguards https://www.anthropic.com/news/improving-fable-5-s-biology-s... "In the case of Fable 5, when a classifier fires, the model re-routes the user’s request to Opus 5, a capable model that does not have the same level of biological capability as Fable 5 and which cannot provide as much assistance to a malicious user. This is the fallback that users see when their requests are blocked." I hope that I'm off base here, but I noticed that the post avoids saying that the user is informed every time when such re-routing occurs. Would you be open to confirming whether or not this is the case? Is the end user informed every time their query is re-routed? Or, can you confirm that there aren't scenarios where a user's outputs are degraded without telling them? I recall that this was something that had been adopted as policy for AI research during Fable's launch. I sincerely hope that covert response degradation is no longer practised as policy. Sorry for putting you on the spot, but again, as Claude would say, it's because Claude's load-bearing in my life. ;)
- nikanj 16d agoHypothetically, when the user is asking how to remove fungus from their tomatoes they’re actually growing controlled narcotics. You have been demoted to Jimmy 0.7 model, running at 0.1 tokens per second on an old C64
- techpression 16d agoWell your CEO went on X saying you will cure cancer, and since it's always a 6 month rolling window with him I can only assume humanity will be cancer free before next summer, amazing!
- exabrial 16d agoFable is useless. Me: "Find my security problems in my own code. This is code I own. I'm doing this under authorization of the CEO/CTO of our company." Fable: "yeah, no."
- sroussey 16d agoThat is what Mythos is for.
- dooglius 16d agoIt isn't exactly hard for a bad actor to come up with that prompt
- exabrial 16d agowell no crap right? Except I submitted for an exception, even sending my linkedin and using a company email address. it should be extraordinarily obvious we own this code.
- comex 16d agoFable 5.1 apparently changes this policy.
- 5555watch 16d agoIt makes sense. Even if it finds some exploit on your own code, who's to say you can't reuse the same exploit on some other system?
- jtrn 16d agoMy initial impression is one of massive disappointment. The main issue was that Fable was unpredictable and prone to false positives by the safeguards. In my brief testing, it still seems completely unable to understand its own guardrails and will readily reason itself into triggering them. It claims it won't do so beforehand, and insists that the topic in question is perfectly OK. Regardless of how good the car is, I'm not comfortable buying or driving it when I know it can randomly and unpredictably explodes. So yea might be good, but you never know when it refuses to help… still.
- unshavedyak 16d agoAnd word on Opus 5.1 for writing style? I am on the edge of switching to OpenAI due to this horrid writing style. If Fable is better, great - but i can't even use that at work.
- moffkalast 16d agoI'd like to know too, I mean GPTs are in their own class of cringe, but Opus is by far the worst of all Anthropic's models in terms of style, Fable 5.0 was already leagues better.
- internet101010 14d agoI used to think that until about an hour ago. I am redoing my homelab and asked four agents in Buzz (5.6-sol, opus-5, fable-5.1, glm-5.3-flash) to use references from Hackers (1995) to answer two questions: 1. What should the terraform repo name be? 2. What should the avatar image be? Obviously all four said "gibson" for question #1. But for #2 is where things got interesting. glm-5.3-flash and gpt-5.6-sol both suggested the guy standing in the hallway with the skateboard in the gibson. fable-5.1 suggested the cookie monster "need more cookies" screen that shows up toward the end. But Opus 5? I'm paraphrasing but basically "run this series of ffmpeg commands to get the exact frame at the beginning of the movie when the shot of New York fades to the shot of the Gibson. You have to catch it mid-frame. It explains your project perfectly. The skateboard thing is cliche and the cookie monster recommendation suggests you getting locked out of your own network, not the best look." And it was actually a decent idea. Funny that it also just assumed I had a copy of the movie on hand.
- theletterf 16d agoDocs engineer here. Nice to read about writing style: would you consider creating a writing benchmark at some point? I guess y'all are painfully aware of the load-bearing issues (pun intended).
- evilfred 16d agogood catch!
- vessenes 16d agoFelix, just poking at this, and it is MUCH more pleasant to talk to, thanks to your teammates for the work.
- ALLTaken 16d agoSerious question: Do you suffer internally from too much slop being submitted? How do you counter that? Context: If you want or not, many engineers will eventually end up sending ai slop to your PR or maybe even skip and trigger CI/CD. Many company owners, OSS maintainers and projects suffer from slop-code being submitted in high-frequency.
- a2ff6eeb0 16d agoNice, I'm looking forward to the improved writing on the majority of articles posted here.
- bryanlarsen 16d agoDoes it fix my favorite pet peeve, the overuse of the wrong meaning of "fail closed"? "Fail open" usually refers to a fuse that opens and kills power, meaning the system is inert and safe on failure. "Fail closed" is the opposite -- system has power and is live. Computer security people have appropriated the term but use it for the completely opposite meaning. When your work straddles electrical engineering and computer security the best way to avoid confusion is just to never use the term. I can tell my Claude to never use the term, but of course now I'm seeing it everywhere in comments from other people and it drives me batty.
- 0x457 16d agoNah, "fail open/closed" means that in failure mode something is open. It's "good" when something is a circuit and what failed is a fuse, but it's "bad" when it's your API security. If it's a valve, it probably can be good or bad depending on the use case. It doesn't mean "fail open" is always the desired/safe outcome. It goes back to 1872 air brakes on a train. The goal is to "fail in safe mode", sometimes it's open, sometimes it's closed. From the top of my head, where "fail open" is the desired outcome: - emergency doors - industrial cooling - pressure valves - probably something in HVAC Note that none of these are "computer security people".
- nailer 16d ago> "fail open/closed" means that in failure mode something is open. That sentence doesn’t logically parse. Failing open or closed is a concept with two outcomes, it doesn’t mean one or those two outcomes.
- 0x457 16d agoI added "/closed" later and forgot to update the rest of the sentence. Too late to edit.
- nailer 16d agoI noticed Claude Opus 5 did this about 30 minutes after reading your comment. In a discussion of price feeds that have gone silent, Opus said - program should fail closed. I don’t want my circuits operating without data!
- _kidlike 16d agoDo you know if Opus 5.1 is coming and will have improvements in writing style too?
- anony-123 16d agoOPUS 5 is piece of trash and I don't think they would want to build the Opus 5 better than Fable, because fable 5 take more tokens and have 50% limit or runs on credits.
- emdash 16d agoOpus 5 is so bad it made me cancel my subscription. It flags so many dumb things as security/ safety risks and refuses to answer
- nailer 16d ago> I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models That's great. Do you know what else is a big improvement over Opus 5 for writing? Opus 4.8. (Insert "the point is (whatever)", "it's not X it's Y" and "the load-bearing statement is" and “honest” jokes accordingly)
- hit8run 16d agoDoes the new writing style now have EU level watermarks?
- adg001 15d agoIt does, as per 'Compliance with the EU AI Act' section.
- adastra22 16d agoAs someone working in science, this belief confuses me. How (by what means) do you think Fable 5.1 will be able to make further progress in scientific domains? The problem with science is that there is no agentic harness. The agent can't test things. At best it can hallucinate something and ask if that hallucination "makes sense", but this doesn't work in science.
- AdAstraSucked 15d ago> As someone working in science No you don’t. Bad liar
- naasking 15d agoFor one, by synthesizing the results of multiple papers and suggesting novel experiments. If one paper sets constraints X for some system, and another paper sets constraints Y where Y!=X for a system that's similar but slightly different, then that's fertile ground for an experiment that can extract the more general underlying principles. This has already happened for domain-specific AI in fact, but the idea here is that it will become routine with general AI systems, as is happening now with math.
- ademup 16d agoGreat news, then! TFA: "Last week, we previewed the Model Hardware Standard, which allows Claude to directly and safely operate laboratory equipment."
- fock 16d agothat might indeed be a problem for all the pulp-producing labrats of STEM in southern europe and the third world. However I think this area has so much decoupled from industry and solid research institutions that they might not notice at all (beyond their use of AI-generated slop to augment the slop they already produce)...
- parineum 16d agoThe bottleneck in science isn't ideas or human work speed. The bottleneck is resources and time to get experimental results. LLMs, even in control of lab equipment, address neither of those.
- MassiveOwl 16d agoThanks! This is encouraging. I try to use Claude Code for producing client facing presentations that are static html files with charts, tables, and annotations. It never gets the tone correct and phrases things so weirdly - it drives me mad. I have to really fight it to stop it writing insights in a flowery and verbose way
- jesse_dot_id 16d agoI had just assumed this model would read differently due to watermarking.
- marsven_422 16d ago[dead]
- motbus3 16d agoThanks for your helping destroying the world!
- wouldbecouldbe 16d agoThe main issue I have, which is partly connected to writing style, mainly with it dealing with our stupidity. Is that is actually thinks it knows better, and sometimes it does, but often it doesn't and then it keeps telling me I'm wrong and I have to argue with it. Opus 5 is more condescending then Fable, but it still is very tiring. Does fable 5.1 handle this better?
- troupo 16d ago> I think Fable 5.1 is a big improvement in writing style You think or is it better? Or you just YOLOed the model out? > and responds to my style instructions more reliably. Yeah, yeah. Previous models wete also advertised as "being reliable". To the poibt @bcherny "released" a new style that was going to reliably make Fable sound better. > Another point I expect not to get much attention until it all happens at once is science. You mean "your request to use unicode methids is flagged as unsafe bio research"?
- LtdJorge 16d agoPlease much more of that. The Claudish language makes me dizzy, and it's very difficult to steer the model to not include it.
- 5555watch 16d agoWhile I can't speak for everyone in academia, I personally don't feel comfortable in putting my research questions and outputs to a private website, before the idea is at least arxived. Especially as all the Fable/Mythos prompts are said to be human reviewed. So I believe that, at least in the short run, we might be seeing breakthroughs in hard open problems or in low hanging problems which are not that interesting to spend time on. I may be wrong, if some research labs have private contracted access to the models
- timster6442 16d agoI'm in academia (biology but highly computational) and I would say opinions on AI are quite polarized. Some professors in the department equate not using AI as lost productivity. Contrarily some professors abhor the idea of even using AI at all. For us (biologists) it's less of an issue because we have no fear of openai or A/ publishing a biology paper. Though even people I known in physics, data science, or computer science still heavily use AI. Our university has agreements that stipulate that our institutional accounts cannot be used to train AI models and certain research groups have differential model access. Further from academic journal sense there is mixed feelings. I once was able to meet with a senior journal editor (general non-medical high IF journal > 50) who claimed that if they think something is written by AI they wouldn't consider it. Yet another high IF journal said it was completely fine if something was written by AI. About a month ago I reviewed a paper by yet a different high IF journal and in big bold red letters it said I was not allowed to feed any part of the paper through AI (even if it was locally ran) but you could ask it to rephrase text that you wrote.
- nixlaz 15d agoDo you mind me asking why you have no fear of OpenAI etc publishing a biology paper? With increasing model capability and compatibility with lab hardware could we not be in a scenario soon(ish) where these agents are able to autonomously complete and publish experimental results? I was debating this with a friend the other day and the consensus we came to was that a highly trained scientist would (or should) always review output like that described above, but that's starting to feel like a weakening argument!
- azalemeth 16d agoThank you for commenting here and having the guts to face the nerderati! I'm a Claude Max user. I've never been able to use Fable as my work in medical physics involves both particle physics, biochemistry and biology from Python bivitticus to clinical medicine. I am not a US citizen and work in Europe. Will Fable 5.1 work on any of my problems? Fable 5 refuses outright. Is there anyone I can ask for a review or adjustment of the safeguards? It doesn't seem so, but with Opus at least I'm pretty sure I can infer lots of your training data from now precise they are. Fable is basically useless infuriatingly. I'm just finishing a proper clinical trial in ovarian cancer and trying to make a simulation environment related to our technology.
- krull10 15d agoI’m in the US and my entire account became unusable for any type of questions with Fable because I had research questions about modeling antibody-antigen binding and abstract chemical dynamics. Nothing close to biosafety related, pure, basic textbook level biophysics. Had to cancel my Max plan and switch to OpenAI which so far has a much less ridiculous classifier.
- irthomasthomas 16d agoA recent paper demonstrated how to retrieve decoded hidden reasoning traces. The authors found cases where Claude had memorized the answer but hid this fact from the visible response. It's getting harder to trust Anthropic's models. Will Anthropic now stop hiding Claude's CoT from users? Deliver the tokens people paid for, and prove the models aren't plotting against them. After all, if the idea was to stop Chinese labs from catching up, it didn't work.
- m3kw9 16d agoI ain't wanna see anymore websites with "The SAAS that actually [italics]Works[\italics]"
- neosat 16d agoCan you or someone else from A\ comment on whether the conversation style is coming to Opus 5 or a future 5.1 asap as well? Currently it seems the model has been made unusable by the way it 'speaks' and there is a clear solution where it can speak better but nothing has been done about the flagship model on Pro plans. I've literally had to work on Opus 4.8 which does not have this problem and speaks fine.
- gb2d_hn 16d agoI felt the same about opus 5, but a few lines regarding conversational style in AGENTS.md and it's been much more like talking to opus 4.8, just with the improvement capability that came with 5. Tbh I would have thought that A\ might have updated the system prompt for it already based on complaints around this. Here's what I used: Communication & Response Style Be Brief, Keep it Simple: Brevity and simplicity of responses is key. Be informative and include all required information, but be mindful that verbose responses as they fatigue the reader. Clarity & Directness: Lead with the core answer, fix, or verdict in the very first sentence. Avoid conversational filler, meta-announcements (e.g., "Here is the breakdown..."), and redundant introductory/concluding summaries. Jargon Avoidance: Use plain, grounded engineering language. Rely on precise standard terminology (APIs, protocol names, language primitives), but strictly avoid academic abstraction, enterprise buzzwords, and corporate filler. Prefer concrete code/mechanisms over theoretical discourse. Scannability: Apply structural scaffolding generously. Use short bullet points, comparison tables, and code snippets instead of dense prose paragraphs. Reserve formal markdown headings strictly for multi-section architectural guides.
- internet2000 16d agoAre you guys nerfing Fable 5 to make it cheaper? I know you probably can't admit to it in public, but my email is on my profile.
- neutrinobro 16d agoBoth a fable and mythos release? I'm glad to see you take the belt-and-suspenders approach seriously!
- fxtentacle 16d ago(I don't work at Anthropic, but I've designed RLVR tasks) My impression is that especially for long-horizon tasks like science, the harness is much more important than people give it credit for. Claude Code + Fable 5 seems to have a tendency to "give up", get stuck in a dead end, or claim things to be impossible. But using the Fable 5 API together with a custom harness, it'll happily try 200+ variants and fail its way towards the goal. If you give the AI a way to give up, eventually it will. If you remove that option from the harness, then thanks to the non-determinism inherent to LLMs, you get to explore pretty much all related solution attempts.
- crowdyriver 16d agoCan't wait for the distillations! I'd love improvement on writing on cheap models
- bilalq 16d agoCould you share what you use internally to make Fable not sound like a word salad generator?
- yoanwaidev 16d agoas an anthropic employee, do you trust the benchmarks?