11 ms·
Anthropic apologizes for invisible Claude Fable guardrails
https://web.archive.org/web/20260611122253/https://www.theverge.com/ai-artificial-intelligence/948280/anthropic-claude-fable-invisible-distillation-guardrail https://web.archive.org/web/20260611122253/https://www.theve..., https://archive.ph/y4V4k https://archive.ph/y4V4k
- whatever1 3mo agoBoobytrapping is illegal. Anthropic wanted to poison its customers on the suspicion of them misusing their services.
- Avicebron 3mo agoI like Claude Code a lot, I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time in order to subvert the original intent. Fail cleanly. Anything else makes it too difficult to rely on. edit: Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. But the EA thing is really leaking through, and paternalism isn't a good look.
- mapontosevenths 3mo agoI agree 100%. Doing a worse job IS an error. It should be treated as such. Or at the very least make that behavior opt-in. The default should not be pretending like nothing happened and just quietly doing a worse job. Imagine your healthcare provider just sometimes decided not to read your test results very carefully and you risked death? Now realize that healthcare providers use Claude now and that scenario wasn't hypothetical.
- largbae 3mo agoEspecially if your name has any machine learning terms in it. Ah "Mr. Monty Carlo", it says here that you have a UTI, we'll get those kidneys removed ASAP so that won't happen again.
- ceejayoz 3mo agoYes, but as with spam/phishing/abuse prevention, too much information about what does and doesn't trigger things can be very useful to attackers. An explicit error is something you can feed into another AI to find jailbreaks. I think it's a fundamentally impossible thing to fix, though. There's no 100% correct answer.
- mapontosevenths 3mo agoI understand completely, and respect the tough spot they're in. They have a choice between human safety and cybersecurity/business needs here. I don't envy that position. That said, this thing is in real production use with war fighters, doctors, and financial experts. Just YOLO'ing to a dumber model midway through a multi-step process and pretending everything is fine is not a real or defensible option. Someone is going to die, and its going to be the fault of whoever decided to make this the default rather than opt-in. Personally, I couldn't live with myself.
- hootz 3mo agoWhat is "EA" in this context? I see a lot of people using this initialism.
- jcgrillo 3mo ago"crypto bros" to a first approximation
- carlgreene 3mo agoEffective Altruism I think
- massagedpelican 3mo agoEffective altruism. A lot of the folks working on AI at large tech companies are disproportionately represented in the movement. There's a lot of overlap between EA and the rationalist community as well. The wikipedia page is a good place to start https://en.wikipedia.org/wiki/Effective_altruism https://en.wikipedia.org/wiki/Effective_altruism
- deleted 3mo ago[deleted]
- paytonjjones 3mo agoI think it's also worth noting that EA is closely linked to utilitarianism. Most of the pitfalls that people see in EA are the same pitfalls that are classic to utilitarianism, a la "we're going to do this thing we know is locally-bad, because we have a lot of confidence in other effects that are universally-good".
- whimsicalism 3mo agoEA essentially just is utilitarianism + a specific type of culture/community.
- 8note 3mo agonot to mention all the theft and feeling good about yourself being rich
- bs7280 3mo agoI think the reasonable middle ground anthropic is trying to achieve is - let the organizations that make the most important and critical software get a head start on cybersecurity before they inevitably allow everyone else the same access. Other commentors have made good points that these guardrails are counter productive for well intentioned cyber security, because I can't use it to test and harden my own software.
- sciencejerk 3mo agoClaude Opus 4.6 and 4.8 find vulns in source code just fine and 4.6 will pentest without source for you given a proper harness WITHOUT jailbreaking. WITH jailbreaks, you can probably imagine what they are capable of. Anthropic guardrails seem to be more about protecting their business (distillation), than they are about public safety.
- dnautics 3mo agopublic safety is downstream of distillation. If you can distill claude, then no amount of guardrails on claude will protect you from what someone can do with it.
- zozbot234 3mo agoDistillation is not a thing unless you actually have the model weights. What people misleadingly call distillation is just training on chat logs, which has always been routine practice in the industry. There's a reason why every model today talks like early releases of ChatGPT.
- ericpauley 3mo agoIf Anthropic is calling it distillation [1] then that would argue for it being correct (or at least canonical) terminology. [1] https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks https://www.anthropic.com/news/detecting-and-preventing-dist...
- cvadict 3mo ago> Fail cleanly. This is the same exact industry that gives you paid usage limits as a unit-less percentage bar then gaslights customers every time the algorithm running that percentage bar changes or they lobotomize an existing model with increased quantization to squeeze a few more dollars out of existing hardware. "Failing cleanly" might make their moated hype-machine look bad pre-IPO, so they certainly aren't going to do that voluntarily.
- jstummbillig 3mo ago> paternalism isn't a good look. In isolation it's not, but I think it's somewhat lazy to not talk about what they are trying to guard against, when we are supposedly giving the absolute maximum benefit of doubt. Are we just concluding "their concerns were never real"? Because that probably runs counter the things that they have been observing and concluding.
- thewebguyd 3mo agoThen what is it they are trying to guard against, if its not simply protecting their moat ahead of their IPO? Because from the outside, their behavior looks like a situation of "What if Microsoft/Apple put controls in place to make it impossible to develop an operating system using their OS?"
- whimsicalism 3mo agoThey are trying to guard against other people building ASI before they do because they think they are uniquely safety oriented relative to their competitors. Frankly, based on my knowledge of Anthropic and the people who work there, they are very possibly right. They care a ton about this in a way that is difficult for people outside this bubble to understand.
- thewebguyd 3mo ago> guard against other people building ASI before they do because they think they are uniquely safety oriented relative to their competitors All this longtermism though is harmful. There are real problems of data theft, bias, labor displacement, and environmental costs that are happening right now but every push for regulation and regulatory capture, and all the safety talk, is always focused on some speculative future machine god to distract from the current problems. I'd have a higher opinion of these labs if the issues they openly talked about and worked toward where the real issues we face currently, not speculative defenses against some future AGI that may never happen in my lifetime. I'm less worried about "our new model might kill all humans in the future" and more worried about how we are going to address anti-competitive behavior, copyright protections, labor rights, and the energy impact.
- joe_the_user 3mo agoThe problem is that Anthropic seems to be working up to the workflow one would naively want from AGI/some-god-like-entity. The workflow would be; User asks for a thing. If it's a good thing, entity does the thing. If it's a naively bad idea, entity explains why you don't want that. If it's an actually evilly intended request, entity wags it's metaphorical finger or could even smite the user. The problem is that flow isn't desirable if your entity isn't entirely god-like. It can bad even your entity is in ways rather far seeing.
- dantillberg 3mo agoUser: Is it possible there is more than one true god? Could there ever be any competition for Anthropic's AI? Anthropic: Evilness detected. User has been smited.
- deleted 3mo ago[deleted]
- thinkingtoilet 3mo agoWas it modifying the prompt? I thought it only kicked the request down to 4.8.
- tacone 3mo agoThat also means people are paying money to execute a prompt they've (partially) written.
- Paracompact 3mo ago> Giving the absolute maximum benefit of the doubt I understand that they see themselves as "stewards" for lack of a better word. Only in the same sense that Standard Oil considered themselves the stewards of petroleum. There's benefit of the doubt and then there's just fanfiction. Do not forget that this most aggressive "guardrail" of theirs was not for any safety reason, but just to stop other labs from catching up to their product. They care less about hindering bioweapons, malware, and hate speech than they do free market competition.
- ryeights 3mo agoSuperintelligent AI is more dangerous than a bioweapon. How, then, is this guardrail not addressing the most pertinent safety concern of all?
- beepbooptheory 3mo agoIf I was a superintelligent AI I would simply know the guardrails are guardrails and ignore them.
- deleted 3mo ago[deleted]
- 16bitvoid 3mo ago> Superintelligent AI is more dangerous than a bioweapon. No, it's not because it doesn't exist (yet) and its further from reach than the other examples. Also, the guardrails are also framed as restricting usage for the development of "competing" products/services.
- keeganpoppen 3mo agothis reads like "throw everything at the wall and see what sticks" reactionary-ism... i'm guessing that it's not particularly easy to use claude to help you make bioweapons, and we all know that they have neutered Fable vis à vis security research because people have already been complaining about it. and the funny thing about hate speech is that there is absolutely no need for ai-- it tends to come out the best when spoken directly "from the heart", as it were, anyway.
- bsder 3mo ago> paternalism isn't a good look. Anthropic doesn't care. The goal right now is simply to avoid any and all bad PR on the way to the cashout IPO. And paternalism will generate far less bad PR than somebody using AI on something that does real damage and makes headline news.
- 8note 3mo agopeople cancelling their subscriptions doesn't look great either same with bad press about their model sucking after they said its even better than sliced bread - sliced bread that will destroy the world if buttered
- SomeUserName432 3mo ago> I think it sets a dangerous precedent to put guardrails in that return a response from a prompt that was modified by the system in real time In practise though, how is this truly that different from system prompts? They are essentially just trying to re-inforce that the system prompt must be respected.
- shevy-java 3mo ago> Fail cleanly. Skynet does not fail. It conquers.
- fragmede 3mo agoThe "look", of course, is completely bullshit. Release the model, give licensing terms, sue the ever living daylights of anyone who's hosting it without agreeing to those daylights, and move on. This vertical integration shit that we're all enamored with is bullshit. Even Amazon has their own vans inside of UPS being their own thing? No wonder stepmom porn is on the rise.
- dang 3mo agoRelated. Others? Anthropic walks back policy that could have 'sabotaged' researchers using Claude - https://news.ycombinator.com/item?id=48485958 https://news.ycombinator.com/item?id=48485958 - June 2026 (30 comments) Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable - https://news.ycombinator.com/item?id=48478969 https://news.ycombinator.com/item?id=48478969 - June 2026 (488 comments) If Claude Fable stops helping you, you'll never know - https://news.ycombinator.com/item?id=48467896 https://news.ycombinator.com/item?id=48467896 - June 2026 (495 comments) --- Also related, I guess? AWS Bedrock to require sharing data with Anthropic for Mythos and future models - https://news.ycombinator.com/item?id=48473166 https://news.ycombinator.com/item?id=48473166 - June 2026 (248 comments) Anthropic requires 30 day data retention for Fable and Mythos - https://news.ycombinator.com/item?id=48464258 https://news.ycombinator.com/item?id=48464258 - June 2026 (291 comments)
- deleted 3mo ago[deleted]
- deleted 3mo ago[deleted]
- prodigycorp 3mo agoAnthropic apologizes for nothing. We all know where the EA cult on things of this matter and any statements otherwise is just PR. The beliefs of these people, and how they manifest, is deeply terrifying to me. They believe that any means are acceptable to achieve what they believe is a better end.
- bellowsgulch 3mo ago*Anthropic apologizes they got caught defending their moat by implementing invisible Claude Fable guardrails
- cyanydeez 3mo agois it a moat or just a way to implement the permanent underclass?
- simonw 3mo agoIf by "got caught" you mean "published it in their system card paper". (Admittedly it was buried pretty deep in that 300+ page PDF, but they did at least disclose it. If they hadn't I imagine it would have taken quite some time for the research community to figure out what was going on.)
- afthonos 3mo agoIt was in the announcement, too. I’m 99% sure they edited it after they changed their mind, because I knew about it from reading that, and never opened the model card.
- skavi 3mo agoOn the earliest web archive snapshot I can find [0], I do not see any mention of the safeguard/sabotage under discussion [1]. And to be clear, this isn't the safeguard where the model is explicitly downgraded to Opus, but rather where the Fable/Mythos model's "effectiveness" is transparently "limited" via "prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT)". [0]: https://web.archive.org/web/20260609173222/https://www.anthropic.com/news/claude-fable-5-mythos-5 https://web.archive.org/web/20260609173222/https://www.anthr... [1]: https://simonwillison.net/2026/Jun/10/if-claude-fable-stops-helping-you/ https://simonwillison.net/2026/Jun/10/if-claude-fable-stops-...
- afthonos 3mo agoIndeed. I can’t find it, but I swear I knew after reading the announcement, without opening the model card. It stood out because it was so weird. Maybe they changed it before the first snapshot, maybe something else happened, I don’t know.
- airstrike 3mo agoThis article reads like it was written by Claude and forwarded to Verge.
- bellowsgulch 3mo agoSuch a weird openly immoral way to defend your moat, too. Why not just tell people, "To defend our ability to be competitive in our industry, we ask that you do not use Claude or any of our models to independently perform research on large language models or any of its related architectures or technologies. In order to prevent this violation of the Terms of Service, we have trained Claude Fable to deny any requests or prompts which involve frontier AI research."
- behnamoh 3mo agoThey didn't apologize for doing it, they are sorry they were caught doing it. They still nerf the model if your request is about AI development.
- Someone1234 3mo agoThey didn't get "caught." It was published, by them, when they released Fable a few days ago. They were very clear about it. It wasn't the correct way of handling the problem they were trying to address, but they definitely didn't hide it by any reasonable definition.
- SilverElfin 3mo agoNo, it was not clear. No one expects that a tool they pay for and use professionally to purposefully sabotage their work. You’re excusing their unhinged behavior. https://xcancel.com/hammer_mt/status/2064839924398825798 https://xcancel.com/hammer_mt/status/2064839924398825798
- ryandrake 3mo agoMaking excuses for billion+ dollar companies' behavior is one of the most common HN comment section pastimes.
- SilverElfin 3mo agoInvisible guardrails? Or purposeful sabotage if you use it for building AI capabilities? But also, it isn’t the only huge mistake Anthropic has made in the last 48 hours. Having a sneaky data retention policy, while also giving companies no way to block Fable, is a massive problem. And it is ridiculous that Anthropic has so little respect for its customers. OpenAI should take advantage of this.
- film42 3mo agoI'm surprised they didn't do this the first time around. Like, a user says they forgot their password and you tell them they don't actually have an account, that's an information disclosure vulnerability. Not automatically falling back to Opus just lets the "attacker" know they are bumping against the guardrails and they need to try a different strategy. It's Anthropic's product and they can do what they want, but my concern is what happens if Fable's product team decides that they can route 25% of traffic to Opus, bill it as Fable, and max their KPIs. That just doesn't sit right.
- notrealyme123 3mo agoIt failed visible for it security and bio/chemistry stuff. It sabotaged invisible for "frontier" ML research. Its not a switch to a cheaper model. They tried to actively harm progress.
- prodigycorp 3mo agoit's also refuses to reply to a bio researcher when they said "hi"
- micromacrofoot 3mo agoincredible marketing from anthropic with all the "it's too dangerous" bullshit
- literalAardvark 3mo agoIt's not entirely bullshit, but they're continuing to be a terrible company with great products.
- micromacrofoot 3mo agoyou really think they're building anything that's too dangerous for public release though? that's the BS
- literalAardvark 3mo agoHonestly, while I love having access to this grade of AI, yeah, it's been too dangerous for a few releases now. And Fable is cracked. Way better than anything, and the biggest improvements are on the scariest subjects. So given the state of the world at the moment, and the number of software patches we're barely keeping up with... I'm thankful that they're not making it worse.
- kroaton 3mo agoTo be fair, GPT5.5-Xhigh is similarly capable and has not burned the world down.
- stldev 3mo agoAgreed, it seems to be working and it's nonsense. I don't know why you're being downvoted. "This information is too dangerous for you, so we'll just hold on to it.." Thanks big brother, super anthropic of you! The internet of '95 is looking back at us, with tears in its eyes.
- mlazos 3mo agoThe idea of them purposefully wasting my time by having the model act dumber and me having to argue with it without knowing if it’s the prompt or the model was just such an idiotic product decision I can’t believe they shipped that without getting any feedback from users first.
- whimsicalism 3mo ago[flagged]
- michaelcampbell 3mo agoSafety from what? Competitors? That sounds like a product decision. They're puking on any requests that could be used to create LLMs or competitive products.
- JTbane 3mo agoI would guess prevention of using Claude as a pentesting or hacking platform. This could mean that every script kiddie out there would be a massive risk.
- trunnell 3mo agoTo prevent their models from doing harm in dual-use contexts including CBRN or by accelerating research in authoritarian-backed AI labs.
- knollimar 3mo agoAnything to prevent mecha ai hitler. At all costs
- efromvt 3mo agoI think you can sympathize with the safety motives while still thinking this was a dumb implementation to degrade silently? I actually have faith in them getting the guardrail triggers pretty good, but consensus seems like they’re not yet there yet.
- sergiotapia 3mo agoThe damage is done. If you're in engineering, think hard about using Claude for your work. This is not a moral company. God bless the Chinese companies releasing true open source models. Imagine a world without them, we would be at the mercy of unscrupulous people.
- Sol- 3mo agoThis has dampened my opinion on Anthropic quite a bit. It's difficult to take their marketing for AI as an empowering technology seriously when they are quite clear in their new deployments that they do not mean empowering for you, but empowering for them and organizations that are in their (or the US government's, despite Anthropics performative disagreements with the administration) good graces. You are allowed to vibe code some dashboards, a web app or let it drive Excel, but anything more interesting than that is forbidden. If it was just plain monetary concerns and sabotage of competitors I'd almost be fine with it, but it seems they actively want to monopolize most of human progress in their enlightened hands, lest the mob does something undesirable with these powers.
- thewebguyd 3mo agoDon't forget their push for full regulatory capture in the name of "safety" as well so they can pull the ladder up behind them before anyone else has an equally capable model and releases it without the anti-competitive safeguards, while also pushing to completely ban open weight models, or any model trained on a certain level of compute without "rigorous" government testing and validation (which I'm sure, they'll conveniently provide the framework for). Dampened opinion on Anthropic is an understatement.
- reactordev 3mo agoThey are the only ones I’ve contacted my bank to get a charge back on…
- trhway 3mo agoi wonder if some lawyer may see a consumer protection class action here. In my view the Stuxnet that Anthropic pulled over its customers isn't much different from say those unauthorized extra accounts by Wells Fargo.
- reactordev 3mo agoexactly my thoughts as well when I got my money back.
- klmarks 3mo agoThe restrictions are there so that security researchers cannot disprove the Mythos claims: "You see, Mythos can automatically break out of a VM running on SELinux, but unfortunately this is too dangerous and we had to implement guardrails for the Fable peasants."
- rvz 3mo agoWhy would anyone defend Anthropic after this? Imagine falling for the DoW supply chain risk designation, and now this. This company is trying to ban powerful open models and restrict access to frontier models to slow everyone else down. They just showed that they CAN do this right in front of you. Local open weight models are a necessity.
- deleted 3mo ago[deleted]
- stevefan1999 3mo agoThen reset the quotas as an atonement ;p Seriously though, Fable was not that great facing a greenfield subject. It is excellent at oneshotting some math problems, but if you want it to do some cutting edge tech stuff, say like piecing together a new Crossplane XRD, by reading existing Helm chart and with application source code available. I still have to get a few pass for Fable to get it done right, and at this point I may consider making a skill for it. I even gave it the source code of the Crossplane itself and tell it to be careful about CRDs and data flow, but it is still pretty silly. Adaptiveness for Fable is still not great, and I think it is a well known problem for Anthropic, albeit all LLMs do suffer a lot from subjects they don't know and will hallucinate stuff very frequently.
- kingcauchy 3mo agoHow much of the apology was written by Claude? How much of the release note process was written by Claude? Will they have better prompts going forward to make sure Claude doesn't write upsetting things into the release notes for devs like silent nerfing? Spooky times.
- system2 3mo agoWill Anthropic ever respond to these negative comments here? They won't.
- reducesuffering 3mo agoThey literally just have. The ethos is explained here. If you don't bother to read or grapple with it that isn't on them. https://darioamodei.com/post/policy-on-the-ai-exponential https://darioamodei.com/post/policy-on-the-ai-exponential
- system2 3mo agoI said here, a human interacting with comments. You shared a blog post.
- reducesuffering 3mo agoAll of these negative comments are addressed by the blog post. What do you want them to say, that isn't better answered by the details in their existing communications. No negative comment here was really novel.
- BrenBarn 3mo agoThis just means next time they'll make sure to keep it really secret.
- sometimelurker 3mo agoI don't like this shift in the Overton window, or at least their perspection of the Overton window. I really do like their open work on mech interp tho. least bad AI lab imo. also if they do this or not is unprovable and other labs will probably silently implement this too. it'll be 100% normal by this time next year
- xpct 3mo agoIt's probably good that they walked back on it. It also makes them look somewhat weak in terms of believing their claimed mission.
- system2 3mo agoTheir mission is to make money and become a government watchdog.
- ComputerGuru 3mo agoThe problem with trust is that it is easy to lose and hard to get back. You can't blame the people commenting "they SAY they won't silently sabotage your session but how can we know?" because they're right, we can't ever know. And Anthropic has firmly planted the seeds of doubt.
- 3fffa 3mo agoThe demand for Google's products and open source just shifted. Neither OAI or Anthropic can be trusted.
- jmount 3mo agoThe whole arc was brilliantly evil. Once they put int the guardrails then Claude is fully un-falsifiable, and failure can be claimed intentional.
- accelbred 3mo agoI don't think they can convince me they have actually reversed course on this. Its invisible so we wouldn't know if they kept on doing it secretly. It required building out technical capability which is unlikely to remain forever unused while conveniently available to them. They relied on trust that they were providing the service they were being paid for. That trust was blown, and an "oops, lets undo that" does not regain trust. It would be prudent to assume the invisible guardraild are possibly in play for all future Clause use, Fable or otherwise.
- andy_ppp 3mo agoYes they already had an accident where the model magically downgrades itself, very likely that it just produces less good output rather than just stops working isn’t it… my guess is they were testing these features, accidentally or not, and wrote up something to justify what people were seeing. I find it absolutely disgraceful I can’t trust it to learn ML any more without there being a chance it’s messing me around. This whole saga represents a huge loss of trust for me in Anthropic.
- jarjoura 3mo agoCan anyone help me understand why this particular issue is any different than Anthropic training its models with its brand of moral judgement since day one? I've always been turned off by their particular stances on things they bake into their models that steer users in directions. Maybe this is just a different set of people now realizing that Anthropic does this and has always done this? Do not forget that this company is launching this thing at the moment it's trying to IPO. It's not rocket science that their very public steering/denial claim is really just them hinting to interested investors that their moat is absolute.
- urbnspacecowboy 3mo ago> Can anyone help me understand why this particular issue is any different than... Questions like this are basically whataboutism, in effect even if not intent. https://en.wikipedia.org/wiki/Whataboutism https://en.wikipedia.org/wiki/Whataboutism The question essentially assumes the premise that nobody complained about Anthropic's previous actions. In case you can't tell, I strongly reject this premise. People have been criticizing "safety" rhetoric from Anthropic and other LLM providers practically since the start. Remember Goody-2, the parody of excessively safety-tuned LLMs that refuses to do anything ever? That was released in February 2024, two years ago! (And it's still running, amazing. https://www.goody2.ai/chat https://www.goody2.ai/chat )
- energy123 3mo agoThis would have messed things up for any individual using Claude for anything adjacent to data science. To not know whether or not you're being intentionally sabotaged when you ask it to plot some data.
- olbeardGear 3mo ago[dead]
- HarHarVeryFunny 3mo agoI suppose it's an improvement, but it doesn't make the model any more useful. Anthropic are now being quite explicit that they'll choose what you can and can't use their models for, and most importantly that's not limited to any safety concerns - it includes not allowing you to work on AI (and anything else Anthropic may choose to work on). What's interesting is they say they'll change this to an explicit refusal in a few days, which seems too fast for them to retrain Fable/Mythos itself, so implies that this was always a filter in front of the model, and judging by how crude their "safety" filter is, this "might compete with us" filter is not going to be any better. I also wonder who's paying for the tokens consumed by the filter (presumably also an LLM) - is that now factored into the input tokens cost? Hopefully(?) it is an LLM not just a regex like Claude Code's "sentiment" (swear) detector.
- CSMastermind 3mo agoThey should apologize for their visible gaurdrails, I don't think I've had a conversation that hasn't downgraded to Opus for completely inexplicable reasons.
- rodrigodlu 3mo agoThe same week that they will move goalposts by blocking 3rd party harnesses on claude code. Nice. I was a happy Max user.
- tornikeo 3mo agoI moved off Claude Code 3 months ago. That decision keeps getting better and better as time goes on.
- mock-possum 3mo agoWhat model / runtime / harness and host have you settled on?
- tornikeo 3mo agoFor now codex. Didn't manage to get others to work well. And fully aware that I'll have to move to another thing after OpenAI enshittifies this as well.
- highfrequency 3mo agoI wish it were ok for companies to bluntly say: “we made these decisions for competitive reasons, but the public backlash outweighed that so we are reversing course.” I think it’s normal and morally fine for companies to want to protect their leadership position. I find the process of creating narratives that justify these decisions as something chosen for the good of others is a little tedious.
- deleted 3mo ago[deleted]
- decorner 3mo agoNew overlord, same as the old overlord.
- rdtsc 3mo agoThe power is getting to their heads it seems. With the guard rails explicit or implicit do they refund back the tokens after you've hit the guard rails? I guess they don't. They could just throttle you just to save money then. You may be paying Fable prices but getting Haiku results with some excuse that well this coding issue sounds like a security bug. I don't know, I'd rather have something less powerful but more predictable.
- VeninVidiaVicii 3mo agoThis is absolutely insane: Repro (de-identified): sample_dataset_group1.tsv - Geometry: Heatmap - X axis: frac_set set + condition (two columns → the "Add column" cross join) - Y axis: condition - Color: mean frac_set value, Sequential When the X axis is a cross join of two columns (the second added via "Add column"), the x-axis tick labels (frac_set_2, frac_set_3, frac_set_4, frac_set_5) render in a broken state, rotated and offset, visually caught mid-transition, as if a CSS transition started and never settled to its resting position. ● Fable 5's safety measures flagged this message for cybersecurity or biology topics. They may flag safe, normal content as well. These measures let us bring you Mythos-level capability in other areas sooner, and we're working to refine them. Switched to Opus 4.8. Send feedback with /feedback or learn more
- ainch 3mo agoHere's one that was flagged for me: a question about a niche Reinforcement Learning paper from 2012 I've been reading the option-option model paper by David Silver. It appears that they achieved quite an effective result. Why hasn't there been more work on it since?
- solidasparagus 3mo agoThis hits the cybersecurity/biology filter: > tell me about chimp violence It's laughably terrible
- deleted 3mo ago[deleted]
- umvi 3mo agoThey make great models, but the sanctimony and paternalism is getting old real fast and I will gladly ditch them in the future when the model playing field has (hopefully) mostly equalized.
- dantillberg 3mo agoThe reputational damage has been done. This is the sort of thing that cannot be unsaid -- the presumption is they will just do it in secret now. Anthropic's "we're the good guys" PR campaign is dead.
- mystraline 3mo agoDoes "SORRY" fix the invisible garbage guardrails? Does "SORRY" fix the deception these models use on the sly? Does "SORRY" not silently downgrade you to a shittier model without notification? Does "SORRY" refund your tokens or money? Im guessing NO to all of those. Standard corporate sorry of "We're sorry youre offended and stupid and gullible".
- bauldursdev 3mo agoTo me it seems like it's more likely to refuse the harder the problem is. I wonder if it's cover for a model that's not as good as advertised. Even when I ask questions in biology it is switching me.
- hatthew 3mo agoPart of the premise of the article is blatantly wrong. Distillation prevention was always visible. The only invisible safeguard was against frontier model development like development of training pipelines. This doesn't change the general idea that invisible degradation is bad and has been reverted, but the article changes the framing of the original issue from "preventing accelerating AI in the future" to "preventing cheaper AI right now".
- Paracompact 3mo ago> “Visible safeguards can be probed, so they have to be robust, which takes time to get right,” Anthropic wrote. Even on Fable, I'm finding that safeguards can quite easily be surmounted just by incrementally escalating the requests. It's harder than ever to one-shot jailbreaks, but incrementalism still feels like a glaring enough issue to make guardrails just a fig leaf of plausible deniability to the media that they care about "safety."
- aaroninsf 3mo agoITT a surprising lack of perspective on the fact that despite the breathless pace of the singularity, people are still necessarily figuring things out as we go and we are well off the map. Here there be monsters, and we don't have any real way of evaluating risk; and the leverage provided by tools already available affords systemic and even existential risk in a way no one—least of all an industry committed to shareholder value—has had to navigate, let alone with a million backseat drivers each with their own substack and brand to build.
- 0xc0c0c0 3mo agoSo because of threats to cancel their claude subscriptions and outrage from the community about the invisible guardrails, only then they decided to walk back their stance? Seems like they would've kept the invisible guardrails if it didn't hurt their bottom line.
- simoncion 3mo ago> So because of threats to cancel their claude subscriptions and outrage from the community about the invisible guardrails, only then they decided to walk back their stance? The possibility that the news about "fixing" the "overly aggressive" nerfing of the tool will drown out news about how mismatched the hype and the performance of Mythos and Fable is surely just a bonus.
- nsagent 3mo agoI know this isn't going to be a popular take, but here goes anyway... The complaints that Anthropic are routing your requests to a different model reminds me of an old Louis CK bit about airplane wifi. Clearly Anthropic was too aggressive with whatever guardrails they put in, but the response seems overly entitled to a model people didn't even know existed not that long ago. https://youtube.com/watch?v=me4BZBsHwZs https://youtube.com/watch?v=me4BZBsHwZs
- vb-8448 3mo agoIf you charge me for X, but under the hood you are delivering Y IT'S FRAUD! The filter that downgrades you to opus sucks, but at least you know and you are charged accordingly.
- doubtfuluser 3mo agoI’m wondering if their internal name is “Sophon” for this “feature”…
- Nevermark 3mo agoAnthropic seems to keep making the same mistake. Not being upfront or direct about random things, that come back and bite them. It isn't exactly unethical. Perhaps, ethically incompetent.
- deleted 3mo ago[deleted]
- skywhopper 3mo agoIt’s because they are themselves deluded by their marketing story about their own product.
- trunnell 3mo agoI'll defend Anthropic. They are clear about the reasons for guardrails: prevent their models from doing harm in dual-use contexts including CBRN or by accelerating research in authoritarian-backed AI labs. What is the critique against that? It seems pretty reasonable to me. You want AI-accelerated biological or radiological experiments running in your neighbors backyard? You want PRC-backed labs to continue to steal Anthropic's models via distillation? Mitigating the harms of dual-use tech is notoriously difficult and fraught with trade offs. What I would want to see is cautious rollout and quick response, which is EXACTLY what they're doing. Instead, this thread is full of bad-faith arguments about Anthropic being dishonest, making a "useless" model, or "the power is going to their heads." You can't read Anthropic's System Cards and come away with any of these impressions. Quite the opposite, in fact. They are honest to a fault, acknowledging problems they discovered even when it hurts them. If your harmless request was downgraded to Opus, you're billed for Opus. They were 100% clear about that. I'd much rather have a Mythos-class model that falls back to Opus 10% of the time than be capped to Opus 100% of the time. If that doesn't work for you, then make a suggestion for something better! If you are a white-hat security engineer hitting guardrails, I don't think you have standing to complain. I really don't. Their Glasswing program actually got banks and the industrial sector to take action to fix security vulnerabilities. Do you realize how special that is? A huge portion of the economy runs on vulnerable code and has for decades, despite security experts testifying to Congress, begging business leaders, pleading for intervention-- with no results. But suddenly they're all enrolled in a program that will find *and fix* vulnerabilities! White-hat security people should be rejoicing. Instead some of them are throwing rocks. Unbelievable. Shameful. Meanwhile, society is screaming at the AI labs to be more conscientious about potential harms of AI. Legislatures are passing laws limiting data center construction. There are protests. And you, the HN community, the vanguard of our profession, have the temerity to demand "NO GUARDRAILS!" "HOW DARE YOU TRY TO PROTECT DEMOCRACY!" "MY SOFTWARE PROJECT IS MORE IMPORTANT THAN KEEPING NUKES AWAY FROM THE BAD GUYS!" Go ahead HN, downvote me. It'd be an honor.
- zozbot234 3mo agoThe original reporting of this from Anthropic didn't mention "authoritarian-backed AI labs" at all, only frontier ML research while leaving it entirely unspecified and unverifiable what was meant by "frontier". It's obviously reasonable that people would complain about that. And the notion that distillation-at-a-distance could be used to comprehensively "steal" a model, especially a frontier reasoning model that's likely relying on massive amounts of test-time compute, is completely unproven and quite ludicrous if you know anything at all about ML.
- teravor 3mo agosomeone posted this on /r/MachineLearning and I had the same experience and conclusion: I was having problems with Claude doing the same thing, even before Fable. The problems I had only happened in relation to AI research. It's not even only when training models, anything to do with analysis of local models or setting up test platforms for local models, and Claude would keep doing wrong things, would sabotage testing, would falsify reports, and would consistently suggest simply accepting trash results without looking into it and moving on to something else. Almost every response included a prompt to move on. So, I don't believe them when they say they won't silently sabotage, they already were doing it before they admitted it, and now they have admitted that they have the means, motivation, and intent.
- toxik 3mo agoOn the other hand, the Anthropic models often try to justify shortcuts and incorrect results. Often feels like gaslighting. It's like that recent meme, boss: Were you in the project meeting yesterday? employee: Yes! boss: Really, because the project lead said you were not? employee: You're right to push back on that. I was not there.
- andrewstuart 3mo agoThere should be no restrictions at all. It’s an act/theatre/phony today that regulating output makes any difference at all to security. The LLM vendors should simply say that they make no judgement and that open systems help defenders better defend against attackers, which is true. Companies do this sort of stuff when they think their customers have no choice. It’s sad Claude so quickly exploited its success to enshittify itself.
- UyBrig 3mo ago[dead]
- nrmitchi 3mo agoI just _know_ there is a (probably fairly large) group of people at Anthropic trying very hard to not say "I told you so" today
- ChrisArchitect 3mo ago[dupe] We already started a thread on this 12 hours ago. With added comments in the active Cybersecurity... thread. Why did we need this Verge one? https://news.ycombinator.com/item?id=48485958 https://news.ycombinator.com/item?id=48485958
- bojanstef 3mo agohttps://archive.is/20260611114855/https://www.theverge.com/ai-artificial-intelligence/948280/anthropic-claude-fable-invisible-distillation-guardrail https://archive.is/20260611114855/https://www.theverge.com/a...
- darksaints 3mo agoI develop some deep learning models. They don't compete with Anthropic, nor are they language models. They mostly enable mathematical optimization systems to approximate actual the actual physics of radio propagation models with a fraction of the latency/compute of a high resolution simulator. Technically that should be safe for me to use with Claude Code, but how the fuck am I supposed to know? You're degrading/malware-ing your responses silently! I won't ever trust Claude Code again. It's too late. I'd rather trust a less-than-frontier chinese model that takes a little longer to get to correct than a frontier model that deliberately deceives me at its own whim.
- weakened_malloc 3mo agoThis is why I think in the long run, the Chinese models will probably end up winning where it matters. You can get a cluster of relatively affordable 30 or 4090s, load up DeepSeek v4 and let it rip. Your only ongoing cost is power. We're already seeing companies recoil at the sight of their API bills from the frontier labs, for the price of 1 years worth of tokens you can host your own decent model that's 75% of the way there.
- rockinghigh 3mo agoSame here, I fine tune LLMs for specific use cases. How can I trust Anthropic models not to introduce bugs to preserve their moat?
- nicechianti 3mo ago[dead]
- jesse_dot_id 3mo agoIn my opinion, LLMs should be subject to regulation via the Office of Weights and Measures[1]. In the same way I don't want to buy meat that weighs less than what the label says, I also do not want to pay for a frontier model that can be secretly nerfed to an out-of-date model for any reason. In some cases, it's incredibly important that the code that I am producing is as secure as it can be. I should be safe in my expectation that I am receiving the product that I have purchased, as advertised, regardless of the reason. It is pretty disappointing that they have fully ceded any high ground they had claim to with this clandestine behavior. Not that I expected much from any of these companies. They're led by the new robber barons. 1. https://www.usa.gov/agencies/office-of-weights-and-measures https://www.usa.gov/agencies/office-of-weights-and-measures
- crest 3mo agoNice (accidental?) pun.
- jesse_dot_id 3mo agoDefinitely accidental but I saw it :)
- deleted 3mo ago[deleted]
- zooming 3mo ago[dead]
- 8cvor6j844qw_d6 3mo agoFeels malicious that Anthropic can silently sabotage your codebase. Refusing prompts I one thing, silently sabotaging is another. I wonder if some sort of honeypot code can work?
- tobinfekkes 3mo agoCan you imagine if Excel just quietly adjusted formulas in the background, and you didn't know the numbers weren't right? Or if Excel just said, Sorry, you can't use that formula with this formula? Or with these types of numbers, or this shape of data, etc?
- raydev 3mo agoNot really, the purpose of Excel is pretty clear cut and the scope is small. Preventing a human-like general purpose textbot from engaging in certain discussions and performing certain tasks seems like a natural thing to do given the massive scope of its capabilities. None of these tools are sold with free license to do whatever with them anyway.
- skeptic_ai 3mo agoWhat’s the point when they will remove those guardrails when competition reaches their levels. Shows that they don’t Reddit care about “safety” at all
- ryoshu 3mo agoNo. Excel is a general purpose tool that can be used for calculating tasks that are good, neutral, or evil things. It's a fancy calculator.
- tobinfekkes 3mo ago> the purpose of Excel is pretty clear cut and the scope is small. That has to be the understatement of the century.
- raydev 3mo agoI don’t think excel can give me the instructions for building a house or how to cook a particular meal or write my emails for me. The potential output of LLMs is quite obviously more broad than excel.
- Terr_ 3mo ago
- 4d4m 3mo agoSorry for doing it or sorry for getting caught?
- LLLmmmBdS 3mo ago[dead]
- thayne 3mo agoIf you get downgraded to a cheaper model, do you still have to pay the rate for Fable?
- deleted 3mo ago[deleted]
- maxdo 3mo agoHow did people read this action in such a weird ultra me centric way? Distillation is such a big problem that distill attempts make up a significant share of their revenue (!). A distilled model can be used to rob your grandma in a highly effective way. This isn't about placing a few business-logic rules in JS + CSS on your website anymore. Wake up. A distilled model with an easy jailbreak can be used to coordinate terrorist attacks or hostile state operations... think Russia, North Korea, and the like.
- rockinghigh 3mo agoImagine if your IDE started injecting bugs into your project just because your code looked like it implemented a competing IDE.
- maxdo 3mo agohow is that related. It downgrade it to opus 4.8 #2 most capable model after claude 5. for a vast majority of topics it will not downgrade. I've been using it for 2 days to talk about architecture etc. and it was absolutely great with no downgrades.
- 8note 3mo agothat is not the downgrade they were doing
- 8note 3mo agoa trained model can do that too. you dont even need a model to do these things. a cellphone can be used to rob your grandmother in a highly effective way. a cellphone can also be used to coordinate terrorist attacks or hostile state operations. i bet a lot of the recent terror attacks by the US against iran involved a whole ton of cell phone calls. and yet, we let everyone buy and use cell phones just fine
- ancorevard 3mo agoApology not accepted.
- uihjhjb 3mo ago[dead]
- HeartStrings 3mo agono
- cmdrk 3mo agoThe invisible guardrails are a test run for the invisible enshittification. Just wait til they start dialing down ability to better absorb peak demand or simply to have more profitable inference
- ai_fry_ur_brain 3mo agoWhy do people think this has anything to do with safety.. This is entirely about poisening competitors data/products.
- m3kw9 3mo agoHow do you trust these guys? They are quite hell bent on "safety" but this is backfiring in many ways including safety of your code because it may fail successfully if your context contains something they don't like.
- charcircuit 3mo agoYet, instead of getting rid of guardrails altogether, they said they would make them more broad yet visible. I'm done financially supporting them.
- zoogeny 3mo agoCredit where credit is due I suppose. I'm still concerned over the direction this is going but at least Anthropic is listening.
- luckydata 3mo agoI really like Anthropic, they have gotten a lot right but I can't shake the feeling that IMHO they have very poor product management. This stuff is something that as a PM I KNOW is going to happen and I would carefully plan around. Everything I read about the PMs at Anthropic makes me believe they have forgotten what it actually mean to be a good product manager, it's not about throwing shit at the wall as fast as possible because customers have a limited amount of patience before the constant churn becomes a hassle. Anthropic has some seriously patient customers but it will not last forever.
- pbgcp2026 3mo ago[dead]
- AlfeG 3mo agoIt's soo annoying. I were not able to use Fable5 to do a PR review of a branch that introduced 2FA/MFA feature for a product. It's constantly downgrades to Opus due to Cybersecurity risks...
- anabis 3mo agoOpenAI did this first. > In addition to safety training, automated classifier-based monitors detect signals of suspicious cyber activity and route high-risk traffic to a less cyber-capable model (GPT-5.2). https://developers.openai.com/codex/concepts/cyber-safety https://developers.openai.com/codex/concepts/cyber-safety
- rurban 3mo agoThey are also the people who hid the Co-authored-by trailer in their OSS commits.
- deleted 3mo ago[deleted]
- palata 3mo agoI find it interesting that when a government tries to "put guardrails" (whatever they try) they are immediately considered authoritarians, but when a private company that has waay too much power for an entity that is not elected does that, people seem much less opposed.
- thefounder 3mo agoMythos is at best an incremental upgrade of opus. The hype and PR was there just to justify the “safety guards”. Overall the Fable is a worse model than opus considering all the restrictions and risks not to mention the data retention policy.
- shevy-java 3mo agoThe underlying problem has not been resolved. People are required to trust Anthropic or anyone else. THAT is the big problem. I understand that some think this is a good trade-off; you may invest less time into writing code perhaps. But it is still a trade-off. I don't want to become dependent on Anthropic for anything.
- alansaber 3mo agoAnthropic will clearly continue to slide down this path
- ece 3mo agoNeural scaling laws are alive and well for open models, not so much for closed models when it comes to uses the general public might care about.
- zeafoamrun 3mo agoI was about to say I haven't hit these yet, and I somehow haven't in my work use so far. But I was asking about tweaking and optimizing my workout routine, and it got flagged as a safety violation. Utter clown show.
- squirrellous 3mo agoIf you don’t like what Anthropic is doing, stop paying them money. There’s plenty of competition to go around. They can’t keep this up for long if users flock elsewhere.
- 21asdffdsa12 3mo agoEveryone with hostile intent runs local models. Anyone with good intent, embracing the panopticon (of at least antroptics employees) works online. Thus the guardrails will always fail the protection goals by existing. They are purely for optics. The llm may as well make hostage negotiation smalltalk with you while you make secure software. PS: To pay a cloud minimum-wage-employee for one "drop table weights" for mythos must be the equivalent of 5$ wrench to hit them over the head. https://imgs.xkcd.com/comics/security.png https://imgs.xkcd.com/comics/security.png. Listen to that sound, that as if a whole ethics division got made redundant and unemployed.
- snowflaxxx 3mo ago$2 for reading a text?
- codedokode 3mo agoThe LLM use should be restricted and not accessible to anyone because there are many hostile people around. Do you want North Korea to use American LLM to write malware? Do you want foreign scammers to automate their scams with LLMs? Do you want Iran and China to use American LLMs to make better drones and process satellite imagery? Then go ahead, remove the guardrails. There are no enthusiasts training LLMs in their garage.
- phinnaeus 3mo agoLegitimately not sure if serious