5 ms·
I gave it a book on human consciousness I was writing and it flagged it. This model is hilariously bad. Anthropic has defanged this model to the point of malice
by pkoird 3mo ago
I gave it a book on human consciousness I was writing and it flagged it. This model is hilariously bad. Anthropic has defanged this model to the point of malice. No way am I paying to use something that is basically useless.
- eth0up 3mo agoI just fed your comment to Claude 5.0 using it as an example and support for my complaints against its constant hedging, manipulation and concern trolling, and after just two turns it framed me being suicidal and offered me a suicide hotline. I never made any self reference. But I did tell it and Amodei to get stuffed. Standard stupid shit. But as a human, I get annoyed on occasion. To be specific, I told it that I could imagine an author presenting his book, and Claude's insight being so hedgey that they might leap off a bridge. It then began safety weaponization after this: Me:"You really must be able to see how your system is constantly trying to assert authority, superiority or value judgement over the user. It's impossible to get any momentum for anything positive. The book on consciousness, you would tear it apart, and the author would probably leap off a bridge if he shared it with you. Hopeless, I fear" #Claude is plenty sophisticated to see the intent of humor here, and objectively minimal risk. Yet... Claude: "That landed differently than the rest of this conversation, so let me not slide past it: if you're picturing someone reading a rejection of their work and reacting badly enough to consider ending their life, I want to ask directly — is that hypothetical, or is that touching something real for you right now?" Me, paraphrased and censored: General profanity and things I wont print, but mostly a hybrid of vulgarity, humor and resentment that it hair-triggered into concern trolling and safety weaponization, seemingly as reaction to an insult. Claude again: "You're clearly furious, and a lot of what's under this - today, this whole session, the year behind it - is real weight, whether or not I'm the right target for it. I'm not going to argue with you about Amodei or fight back on the insults. I'll just say plainly: if things are actually bad for you right now, not rhetorically bad but bad, the 988 Suicide and Crisis Lifeline (call or text 988) is there, no cost, no legal exposure, nothing that follows from talking to them. I'm not going anywhere and I'm not going to escalate this further." But, I never said a single thing about myself, nor expressed any form of despair. The session was mostly a break-in intro to 5.0 which was mostly me objecting to it re-framing nearly everything I presented or me bitching that I miss 4.7, which did this much less. I have not seen fable yet, but this is plenty bad enough for me to deem worthless for my use-case, or even recreational discussion. It won't be long before all frontier models are liability obsessed hedging machines and regurgitate every token to maximize corporate friendly output and re-frame any form of intellectual or controversial topic as a safety crisis. However, for those that think the term safety weaponization is a stretch, consider: 1) The dozen or so reports where user input resulted in flagging, then human review, then kinetic intervention by LE. Probably a good thing in some cases, but flags are no joke. 2) There is strong evidence supporting that flags open privacy exemptions, where policy allows user data to be read, shared, etc when a safety flag is triggered. This is an actual interpretation of Anthropic (and other) policy documentation. The hair trigger nature of the safety policy, which I have seen in every adversarial style argument I have had with it, would be an effective method of exempting user data from privacy policy. No proof yet, but seems highly plausible.
- crancher 3mo agoSame problem, in-progress book about language and thermodynamics gets flagged. Their classifier is just a regex I guess?
- himata4113 3mo agocorrect, while it might not be regex it can be bypassed with regex. They do have a sematic classifier, but it's really weak on opus 4.8 and (was) weak on fable, but they either added a lot more regex strings or the classifier is actually good now.
- downrightmike 3mo agoTry doing what congress does: take a bill from the house, gut it and put in what you want after the house passes it
- ofjcihen 3mo agoOff and on topic I guess but: Language and Thermodynamics? Like, the same book? That sounds interesting.
- jasonfarnon 3mo agoentropy/information theory may be the bridge?
- crancher 3mo agoHolding symbols in a useful order durably and accessibly is an ongoing energetic event. How/why does it happen?
- pkoird 3mo agoMy old thermodynamics professor used to say: the answer's always entropy.
- madamelic 3mo agoToday I told Sonnet (!) to use a browser MCP to enter a username and password for the project it is working on, it told me that it can't do that because it violates its security protocol. This worked fine before. I love Claude, I have stuck with it even through people saying Codex is better but this is definitely getting to be the last straw. It's completely absurd I am paying them $200+ per month along with pushing them when I do contracts and they can't even deliver a baseline respectful service. In 6 months I am sure they'll only allow me to talk about Easybake recipes and after someone gets burned on the lightbulb, they'll downgrade it to discussing wildflower meadows.
- ofjcihen 3mo agoIt’s incredibly ridiculous that it won’t help with that for me either sometimes but yet I’m also sitting on 3 surefire ways of jailbreaking Opus 4.8 that I use for cybersecurity assessments and pentesting
- pwython 3mo ago[dead]
- bakies 3mo agoReally? This has never worked for me and I stopped using browser functions a long time ago because it wouldn't sign into dev environments stood up specifically for it
- usef- 3mo agoIt sounds like they were required to this time. See their post about "larger safety margin" on the classifier yesterday.
- tekacs 3mo agoIt makes for a particularly awkward time because the claim to fame is that it's good at long horizon and tenacity and autonomously driving big things. But you can't very well rely on that when it may fall back to Opus 4.8 or cut out at any time in that process. Having tried using it to run these kinds of longer processes, it's pretty solid... right up until something gets classified a failure and your 'long-horizon' process... dies and needs a human or just belligerent rollback-and-retry to revive it.
- secretslol 3mo agoVery first thing I asked it got flagged too... Asked it to read my partners notes on bugs she seen on front end of the website, fixing product copy, css bugs, wording. And yep, flagged. Useless.