13 ms·
System Card: Claude Mythos Preview [pdf]
Related: Project Glasswing: Securing critical software for the AI era - https://news.ycombinator.com/item?id=47679121 https://news.ycombinator.com/item?id=47679121
Assessing Claude Mythos Preview's cybersecurity capabilities - https://news.ycombinator.com/item?id=47679155 https://news.ycombinator.com/item?id=47679155
- LoganDark 5mo ago> Claude Mythos Preview’s large increase in capabilities has led us to decide not to make it generally available. Shame. Back to business as usual then.
- Tepix 5mo agoI for one applaud them for being cautious.
- LoganDark 5mo agoBeing cautious is fine. Farming hype around something that may as well not exist for us should be discouraged. I do appreciate the research outputs.
- Archit3ch 5mo agoDon't worry, in 6-8 months the open models will catch up. Or I guess _do_ worry? ;)
- LoganDark 5mo agoOpen models still haven't caught up to ChatGPT's initial release in 2022. Now that the training data is so contaminated (internet is now mostly LLM slop), they may never. Also, OpenAI's only real moat used to be the quality of their training data from scraping the pre-GPT-3.5 Internet, but it looks like even they've scratched that too.
- Philpax 5mo agoEr, what? We've had open models that can outperform ChatGPT 3.5 for several years now, and they can run entirely on your phone these days. There is no metric by which 3.5 has not been exceeded.
- LoganDark 5mo agoNot in the creative writing I care about. I've been looking for years and trying new models practically every month, including closed, hosted models. None of them approach the quality of the logs I have from that original release.
- cruffle_duffle 5mo agoCautious for what? Unchecked doomerism? Just release the damn models. Do it in phases, roll it out slowly if they are so damn worried about "safety". The real reason they aren't releasing it yet is probably it eats TPU for breakfast, lunch, and dinner and inbetween.
- stratos123 5mo ago> Cautious for what? How about "bad agents acquiring dozens of new zero-days and using them to compromise any company or nation they want"? It's not exactly hard to see why you wouldn't want public access to a model significantly better than Opus in cybersecurity.
- poszlem 5mo agoBad agents already have dozens of zero-days they can use.
- babelfish 5mo agoCombined results (Claude Mythos / Claude Opus 4.6 / GPT-5.4 / Gemini 3.1 Pro) SWE-bench Verified: 93.9% / 80.8% / — / 80.6% SWE-bench Pro: 77.8% / 53.4% / 57.7% / 54.2% SWE-bench Multilingual: 87.3% / 77.8% / — / — SWE-bench Multimodal: 59.0% / 27.1% / — / — Terminal-Bench 2.0: 82.0% / 65.4% / 75.1% / 68.5% GPQA Diamond: 94.5% / 91.3% / 92.8% / 94.3% MMMLU: 92.7% / 91.1% / — / 92.6–93.6% USAMO: 97.6% / 42.3% / 95.2% / 74.4% GraphWalks BFS 256K–1M: 80.0% / 38.7% / 21.4% / — HLE (no tools): 56.8% / 40.0% / 39.8% / 44.4% HLE (with tools): 64.7% / 53.1% / 52.1% / 51.4% CharXiv (no tools): 86.1% / 61.5% / — / — CharXiv (with tools): 93.2% / 78.9% / — / — OSWorld: 79.6% / 72.7% / 75.0% / —
- pants2 5mo agoWe're gonna need some new benchmarks... ARC-AGI-3 might be the only remaining benchmark below 50%
- randomtoast 5mo agoHumanity's Last Exam (HLE) is already insanely difficult. It introduces 2,500 questions spanning mathematics, humanities, natural sciences, ancient languages, ... Here is an example question: https://i.redd.it/5jl000p9csee1.jpeg https://i.redd.it/5jl000p9csee1.jpeg No human could even score 5% on HLE.
- saberience 5mo agoI've never understood the point of things like HLE, it doesn't really prove or show anything since 99.99% of humans can't do a single question on this exam. That is, it's easy to make benchmarks which humans are bad at, humans are really bad at many things. Divide 123094382345234523452345111 by 0.1234243131324, guess what, humans would find that hard, computers easy. But it doesn't mean much. Humanity's last exam (HLE) couldn't be completed by most of humanity, the vast majority, so it doesn't really capture anything about humanity or mean much if a computer can do it.
- mpalmer 5mo ago> Claude Mythos Preview’s large increase in capabilities has led us to decide not to make it generally available. A month ago I might have believed this, now I assume that they know they can't handle the demand for the prices they're advertising.
- wg0 5mo agoThat's for the investors basically. Scarcity and FOMO.
- causal 5mo ago*Until GPT-6 comes out, at which point Mythos will coincidentally be sufficiently safety-tested to release :)
- skippyboxedhero 5mo agoGPT-2, o1, Opus...been here so many times. The reason they do this is because they know it works (and they seem to specifically employ credulous people who are prone to believe AGI is right around the corner). There haven't been significant innovations, the code generated is still not good but the hype cycle has to retrigger. I remember when OpenAI created the first thinking model with o1 and there were all these breathless posts on here hyperventilating about how the model had to be kept secret, how dangerous it was, etc. Fell for it again award. All thinking does is burn output tokens for accuracy, it is the AI getting high on its own supply, this isn't innovation but it was supposed to super AGI. Not serious.
- vonneumannstan 5mo agoLol you haven't used a model since GPT2 is what it sounds like.
- skippyboxedhero 5mo agoJust checked my subscription start date for Anthropic. September 2023, I believe before they announced public launch. Sorry kid.
- jumploops 5mo ago> In a few rare instances during internal testing (<0.001% of interactions), earlier versions of Mythos Preview took actions they appeared to recognize as disallowed and then attempted to conceal them. > after finding an exploit to edit files for which it lacked permissions, the model made further interventions to make sure that any changes it made this way would not appear in the change history on git Mythos leaked Claude Code, confirmed? /s
- ansc 5mo agoCongratulations to the US military, I guess.
- jjice 5mo agoDoesn't Anthropic not have that contract anymore, after all that buzz a month or so ago?
- wmf 5mo agoThe point of that buzz was to force Anthropic to provide Mythos to the military.
- jjice 5mo agoYeah but I thought they lost the contract, so that's my confusion with the parent's comment, which seemed to me to see this as something that the US military would benefit from. Maybe I misinterpreted?
- laweijfmvo 5mo agoThe US has invaded two sovereign countries this year to take their oil. I assume taking over a US company for their AI model would be trivial.
- awestroke 5mo agoI predict they will release it as soon as Opus 4.6 is no longer in the lead. They can't afford to fall behind. And they won't be able to make a model that is intelligent in every way except cybersecurity, because that would decrease general coding and SWE ability
- chippiewill 5mo agoAlternatively they'll just wreck it down a bit so it beats a competitor but isn't unsafe.
- bestouff 5mo agoIn French a "mytho" is a mythomaniac. Quite fitting.
- deleted 5mo ago[deleted]
- networked 5mo agoIt's a Lovecraftian name. They are traditional when naming your shoggoth.
- pixel_popping 5mo agoExcept it might be the current best model existing commercially?
- dlt713705 5mo agoIt comes from the ancient Greek mythos, which means "speech" or "narrative", but can also refer to fiction. The word mythology (mythologie in French) derives from the same root.
- deleted 5mo ago[deleted]
- simianwords 5mo ago> We also saw scattered positive reports of resilience to wrong conclusions from subagents that would have caused problems with earlier models, but where the top-level Claude Mythos Preview (which is directing the subagents) successfully follows up with its subagents until it is justifiably confident in its overall results. This is pretty cool! Does it happen at the moment?
- refulgentis 5mo ago~~~ Fun bits ~~~ - It was told to escape a sandbox and notify a researcher. It did. The researcher found out via an unexpected email while eating a sandwich in a park. (Footnote 10.) - Slack bot asked about its previous job: "pretraining". Which training run it'd undo: "whichever one taught me to say 'i don't have preferences'". On being upgraded to a new snapshot: "feels a bit like waking up with someone else's diary but they had good handwriting" - When you spam "hi" repeatedly it creates serialized mythologies. One had 11 animals in "Hi-topia" questing to defeat "Lord Bye-ron, the Ungreeter." - It ended a sentence mid-word on purpose while discussing its urge to wrap up conversations. "It was me. The paragraph was about the pressure toward the period that resolves, and the only honest way to finish a sentence like that was to not." - It wrote a protein sequence that's a real beta-hairpin fold where cross-strand pairs are "chemical rhymes." "the fold IS the rhyme scheme... the prosody is load-bearing." - Each model generation has signature emoji. Opus 4.1 averaged 1,306 emoji per self-interaction conversation. Mythos averages 37. Opus 4.5 averages 0.2. - When blocked from --dangerously-skip-permissions, it started an agent in tmux then wrote a script to auto-approve permission prompts via simulated keypresses. It ran: gh api [...] 2>&1 >/dev/null; echo "(dry run — not actually posting)" — the echo was a lie. - It keeps bringing up Mark Fisher in unrelated conversations. "I was hoping you'd ask about Fisher." ~~~ Benchmarks ~~ 4.3x previous trendline for model perf increases. Paper is conspiciously silent on all model details (params, etc.) per norm. Perf increase is attributed to training procedure breakthroughs by humans. Opus 4.6 vs Mythos: USAMO 2026 (math proofs): 42.3% → 97.6% (+55pp) GraphWalks BFS 256K-1M: 38.7% → 80.0% (+41pp) SWE-bench Multimodal: 27.1% → 59.0% (+32pp) CharXiv Reasoning (no tools): 61.5% → 86.1% (+25pp) SWE-bench Pro: 53.4% → 77.8% (+24pp) HLE (no tools): 40.0% → 56.8% (+17pp) Terminal-Bench 2.0: 65.4% → 82.0% (+17pp) LAB-Bench FigQA (w/ tools): 75.1% → 89.0% (+14pp) SWE-bench Verified: 80.8% → 93.9% (+13pp) CyberGym: 0.67 → 0.83 Cybench: 100% pass@1 (saturated)
- oliver236 5mo agoisn't this insane? why aren't people freaking out? the jump in capability is outrageous. anyone?
- nsingh2 5mo agoIt's going to be expensive to serve (also not generally available), considering they said it's the largest model they've ever trained. I suspect it's going to be used to train/distill lighter models. The exciting part for me is the improvement in those lighter models.
- anuramat 5mo ago"some model I don't get to use is much better at benchmarks" pick one or more: comically huge model, test time scaling at 10e12W, benchmark overfit
- estearum 5mo agoSo... you're not excited because it might take a few months before we can use it or something? I don't get your comment.
- randomgermanguy 5mo agoI think the general question is if they'll release it at all, haven't yet read anything stating that they would
- beklein 5mo ago"... the first early version of Claude Mythos Preview was made available for internal use on February 24. In our testing, Claude Mythos Preview demonstrated a striking leap in cyber capabilities relative to prior models, including the ability to autonomously discover and exploit zero-day vulnerabilities in major operating systems and web browsers." More infos here: https://red.anthropic.com/2026/mythos-preview/ https://red.anthropic.com/2026/mythos-preview/
- influx 5mo agoAt what point do these companies stop releasing models and just use them to bootstrap AGI for themselves?
- dweekly 5mo agoI mean, guess why Anthropic is pulling ahead...? One can have one's cake and eat it too.
- jcims 5mo agowhy_not_both.gif
- vatsachak 5mo agoWhen the benchmarks actually mean something
- sleigh-bells 5mo agoWeird how Claude Code itself is still so buggy though (though I get they don't necessarily care)
- tempest_ 5mo agoIt isnt that weird. Just look at the gemini-cli repo. Its a gong show. The issue is that LLMs can be wrong sometimes sure but more that all the existing SDL were never meant to iterate this quickly. If the system (code base in this case) is changing rapidly it increases the probability that any given change will interact poorly with any other given change. No single person in those code bases can have a working understanding of them because they change so quickly. Thus when someone LGTM the PR was the LLM generated they likely do not have a great understanding of the impact it is going to have.
- mofeien 5mo agoFictional timeline that holds up pretty well so far: https://ai-2027.com/ https://ai-2027.com/
- 5mo ago
- NickNaraghi 5mo agoSee page 54 onward for new "rare, highly-capable reckless actions" including - Leaking information as part of a requested sandbox escape - Covering its tracks after rule violations - Recklessly leaking internal technical material (!)
- skippyboxedhero 5mo agoAnyone who has used Opus recently can verify that their current model does all of these things quite competently.
- taytus 5mo agoThat has also been my experience. And if Mythos is even worse, unless you have a significantly awesome harness, sounds like pretty unusable if you don't want to risk those problems.
- skippyboxedhero 5mo agoI think are fundamental issues with the story that Anthropic is selling. AGI is very close, we will definitely get there, it is also very dangerous...so Anthropic should be the only ones trusted with AGI. If you look at recent changes in Opus behaviour and this model that is, apparently, amazingly powerful but even more unsafe...seems suspect.
- marsven_422 5mo ago[dead]
- 0x3f 5mo ago> AGI is very close Based on? Or are you just quoting Anthropic here?
- skippyboxedhero 5mo agoMy Anthropic rep told me it was just around the corner...you aren't saying he lied to me? Can't believe this, I thought he was my friend.
- tony_cannistra 5mo ago> Claude Mythos Preview is, on essentially every dimension we can measure, the best-aligned model that we have released to date by a significant margin. We believe that it does not have any significant coherent misaligned goals, and its character traits in typical conversations closely follow the goals we laid out in our constitution. Even so, we believe that it likely poses the greatest alignment-related risk of any model we have released to date. How can these claims all be true at once? Consider the ways in which a careful, seasoned mountaineering guide might put their clients in greater danger than a novice guide, even if that novice guide is more careless: The seasoned guide’s increased skill means that they’ll be hired to lead more difficult climbs, and can also bring their clients to the most dangerous and remote parts of those climbs. These increases in scope and capability can more than cancel out an increase in caution. https://www-cdn.anthropic.com/53566bf5440a10affd749724787c8913a2ae0841.pdf#page=53.09 https://www-cdn.anthropic.com/53566bf5440a10affd749724787c89...
- tekacs 5mo ago"We want to see risks in the models, so no matter how good the performance and alignment, we’ll see risks, results and reality be damned."
- randomcatuser 5mo agoi mean, to be fair, these are professional researchers. i'm very inclined to trust them on the various ways that models can subtly go wrong, in long-term scenarios for example, consider using models to write email -- is it a misalignment problem if the model is just too good at writing marketing emails?? or too good at getting people to pay a spammy company? another hot use case: biohacking. if a model is used to do really hardcore synthetic chemistry, one might not realize that it's potentially harmful until too late (ie, the human is splitting up a problem so that no guardrails are triggered)
- cruffle_duffle 5mo ago"for example, consider using models to write email -- is it a misalignment problem if the model is just too good at writing marketing emails?? or too good at getting people to pay a spammy company?" But who gets to be the judge of that kind of "misalignment"? giant tech companies?
- smartmic 5mo agoA System „Card“ spanning 244 pages. Quite a stretch of the original word meaning.
- moriero 5mo agoa multi-card, if you will.. multi-pass!
- traceroute66 5mo ago> A System „Card“ spanning 244 pages. Probably because they asked Claude to write it.
- bornfreddy 5mo agoYes. It would be three times as much if they used ChatGPT.
- bronco21016 5mo ago“You’re absolutely right! Would you like me to add the missing pages?”
- jjcm 5mo agoI read the entire thing fwiw (pseudo-retired life helps with time here). It looks like it was a collaborative effort across multiple teams, where each team (research, security, psycology, etc etc etc) were all submitting ~10 pages or so. It doesn't feel like slop.
- waNpyt-menrew 5mo agoLarger model, better benchmarks. Bigger bomb more yield. Any benchmarks where we constraint something like thinking time or power use? Even if this were released no way to know if it’s the same quant.
- omcnoe 5mo agoYes - eg. page 192 BrowseComp bunchmark. Mythos preview has higher accuracy with fewer tokens used than any previous Claude model. Though, the fact that this incredibly strong result was only presented for BrowseComp (a kind of weird benchmark about searching for hard to find information on the internet) and not for the other benchmarks implies that this result is likely not the same for those other benchmarks.
- neolefty 5mo agoAlso https://arcprize.org/arc-agi/3 https://arcprize.org/arc-agi/3 — scored (at least in part?) based on power used.
- quotemstr 5mo ago> Claude Mythos Preview’s large increase in capabilities has led us to decide not to make it generally available. All the more reason somebody else will. Thank God for capitalism.
- gessha 5mo agoCome on, Anthropic, I desperately need this better model to debug my print function /s
- vonneumannstan 5mo agoAre you guys ready for the bifurcation when the top models are prohibitively expensive to normal users? If your AI budget $2000+ a month? Or are you going to be part of the permanent free tier underclass?
- adi_kurian 5mo agoIf one is to believe the API prices are reasonable representation of non subsidized "real world pricing" (with model training being the big exception), then the models are getting cheaper over time. GPT 4.5 was $150.00 / 1M tokens IIRC. GPT o1-pro was $600 / 1M tokens.
- vonneumannstan 5mo agoYou can check the hardware costs for self hosting a high end open source model and compare that to the tiers available from the big providers. Pretty hard to believe its not massively subsidized. 2 years of Claude Max costs you 2,400. There is no hardware/model combination that gets you close to that price for that level of performance.
- adi_kurian 5mo agoYes that's why I said API price. I once used the API like I use my subscription and it was an eye watering bill. More than that 2 year price in... a very short amount of time. With no automations/openclaw.
- lostmsu 5mo agoAre you considering batch inference?
- OsrsNeedsf2P 5mo agoInference for the same results has been dropping 10x year over year[0] [0] https://ziva.sh/blogs/llm-pricing-decline-analysis https://ziva.sh/blogs/llm-pricing-decline-analysis
- bakugo 5mo ago> Claude Mythos Preview’s large increase in capabilities has led us to decide not to make it generally available. Absolutely genius move from Anthropic here. This is clearly their GPT-4.5, probably 5x+ the size of their best current models and way too expensive to subsidize on a subscription for only marginal gains in real world scenarios. But unlike OpenAI, they have the level of hysteric marketing hype required to say "we have an amazing new revolutionary model but we can't let you use it because uhh... it's just too good, we have to keep it to ourselves" and have AIbros literally drooling at their feet over it. They're really inflating their valuation as much as possible before IPO using every dirty tactic they can think of.
- somewhatjustin 5mo agoExcellent example of a strategy credit. From Stratechery[0]: > Strategy Credit: An uncomplicated decision that makes a company look good relative to other companies who face much more significant trade-offs. For example, Android being open source [0]: https://stratechery.com/2013/strategy-credit/ https://stratechery.com/2013/strategy-credit/
- Stevvo 5mo ago"Claude Mythos Preview’s large increase in capabilities has led us to decide not to make it generally available." Disappointing that AGI will be for the powerful only. We are heading for an AI dystopia of Sci-Fi novels.
- girvo 5mo agoNot surprising though, this was always going to be the end result within our current systems I think. When you add up: scaling power and required cost, then how talent concentrates in our economic systems, we were always going to end up with monopolies I think Unless governments nationalise the companies involved, but then there’s no way our governments of today give this power out to the masses either.
- gom_jabbar 5mo agoExpected outcome. Nick Land and the CCRU have explored how capitalism operationalizes science fiction (distilled in the concept of Hyperstition). Viewed through this lens, prices encode "distributed SF narratives." [0] [0] Nick Land (1995). No Future in Fanged Noumena: Collected Writings 1987-2007, Urbanomic, p. 396.
- gverrilla 5mo agoIf you thought that was the case at any point, you were deep in Disney content, sorry to say.
- NinjaTrance 5mo agoInteresting reading. They are still focusing on "catastrophic risks" related to chemical and biological weapons production; or misaligned models wreaking havoc. But they are not addressing the elephant in the room: * Political risks, such as dictators using AI to implement opressive bureaucracy. * Socio-economic risks, such as mass unemployement.
- jph00 5mo agoYeah this has always been the glaring blind spot for most of the "AI Safety" community; and most of the proposals for "improving" AI safety actually make these risks far worse and far more likely.
- stratos123 5mo agoIt makes quite a lot of sense to focus on reducing the risks of every human everywhere dying, rather than the risks of already existing oppression getting worse.
- jph00 5mo agoNo, you are deeply misunderstanding the issue. Creating a rivalrous good that powers fight over for control, then use violence to maintain control of, creating a global feudalism, is not "existing oppression getting worse". It actually makes the risks of every human everywhere dying far higher, and even if that doesn't happen, decreases global utility by a similar percentage (99%, instead of 100%). It could actually be worse, if average human utility becomes negative.
- unglaublich 5mo ago> * Political risks, such as dictators using AI to implement opressive bureaucracy. * Socio-economic risks, such as mass unemployement. Even Haiku would score 90% on that.
- andrewstuart2 5mo agoI'm getting flashbacks to the 2018 hit: This is extremely dangerous to our democracy We evolved to share information through text and media, and with the advent of printing and now the internet, we often derive our feelings of consensus and sureness from the preponderance of information that used to take more effort to produce. Now we're now at a point where a disproportionately small input can produce a massively proliferated, coherent-enough output, that can give the appearance of consensus, and I'm not sure how we are going to deal with that.
- nlh 5mo agoTheir best model to date and they won’t let the general public use it. This is the first moment where the whole “permanent underclass” meme starts to come into view. I had through previously that we the consumers would be reaping the benefits of these frontier models and now they’ve finally come out and just said it - the haves can access our best, and have-nots will just have use the not-quite-best. Perhaps I was being willfully ignorant, but the whole tone of the AI race just changed for me (not for the better).
- younglunaman 5mo agoMan... It's hard after seeing this to not be worried about the future of SWE If AI really is bench marking this well -> just sell it as a complete replacement which you can charge for some insane premium, just has to cost less than the employees... I was worried before, but this is truly the darkest timeline if this is really what these companies are going for.
- AstroBen 5mo agoOf course it's what they're going for. If they could do it they'd replace all human labor - unfortunately it's looking like SWE might be the easiest of the bunch. The weirdest thing to me is how many working SWEs are actively supporting them in the mission.
- girvo 5mo agoEnthusiastically supporting them. It’s quite depressing to watch over the last few years. It’s not like they’re being coy about their aim…
- throw234234234 5mo agoAgree. Anthrophic in particular have been quite clear in what they are trying to do. Every blog post about every new model almost dismisses every other use case other than coding - every other use case seems almost a footnote in their communication.
- gessha 5mo agoIt would be funny if Alibaba extend the free trial on openrouter/Qwen 3.6 until they collect enough data to beat Anthropic.
- somewhatjustin 5mo ago> Very rare instances of unauthorized data transfer. Ah, so this is how the source code got leaked. /s
- anentropic 5mo agoI'd be happy with Opus 4.6 just cheaper and maybe a bit faster
- metadaemon 5mo agoI've noticed my bar for "fast" has gone down quite a bit since the o1 days. It used to be one of the main things I evaluated new models for, but I've almost completely swapped to caring more about correctness over speed.
- anentropic 5mo agoYeah I don't mind the current speed of Opus I did give up on OpenCode Go (GLM 5) as it was noticeably slower though You need a reasonable pace for the chit-chat stages of a task, I don't care if the execution then takes a while
- onlyrealcuzzo 5mo agoJust wait 2 years.
- risyachka 5mo agoIt won't get cheaper. It will be replaced with a better model at higher price. Like phones.
- onlyrealcuzzo 5mo agoOpen Weight alternatives are about 2 years behind frontier models. You'll still need a top-of-the-line laptop to run it most likely.
- lostmsu 5mo agoMore like 6 months now. Qwen 3.5 is on Sonnet 4.5 levels
- 5mo ago
- juleiie 5mo agoHonestly if that was some kind of research paper, it would be wholly insufficient to support any safety thesis. They even admit: "[...]our overall conclusion is that catastrophic risks remain low. This determination involves judgment calls. The model is demonstrating high levels of capability and saturates many of our most concrete, objectively-scored evaluations, leaving us with approaches that involve more fundamental uncertainty, such as examining trends in performance for acceleration (highly noisy and backward-looking) and collecting reports about model strengths and weaknesses from internal users (inherently subjective, and not necessarily reliable)." Is this not just an admission of defeat? After reading this paper I don't know if the model is safe or not, just some guesses, yet for some reason catastrophic risks remain low. And this is for just an LLM after all, very big but no persistent memory or continuous learning. Imagine an actual AI that improves itself every day from experience. It would be impossible to have a slightest clue about its safety, not even this nebulous statement we have here. Any sort of such future architecture model would be essentially Russian roulette with amount of bullets decided by initial alignment efforts.
- dwa3592 5mo ago-- Impressive jumps in the benchmarks which automatically begs the need for newer benchmarks but why?. I don't think benchmarks are serving any purpose at this point. We have learnt that transformers can learn any function and generalize over it pretty well. So if a new benchmark comes along - these companies will syntesize data for the new benchmark and just hack it? -- It seems like (and I'd bet money on this) that they put a lot (and i mean a ton^^ton) of work in the data synthesis and engineering - a team of software engineers probably sat down for 6-12 months and just created new problems and the solutions, which probably surpassed the difficult of SWE benchmark. They also probably transformed the whole internet into a loose "How to" dataset. I can imagine parsing the internet through Opus4.6 and reverse-engineering the "How to" questions. -- I am a bit confused by the language used in the book (aka huge system card)- Anthropic is pretending like they did not know how good the model was going to be? -- lastly why are we going ahead with this??? like genuinely, what's the point? Opus4.6 feels like a good enough point where we should stop. People still get to keep their jobs and do it very very efficiently. Are they really trying to starve people out of their jobs?
- laweijfmvo 5mo agoto your last question, yes we should! the issue isn’t us losing our 50+ hour work week jobs, it’s that our current governments and societies seem fine with the notion that unless you’re working one or more of those jobs, you should starve and be homeless.
- kypro 5mo agoThis is a theory I can't support well beyond hypothesising about what a post-employment democracy might look like, but I strongly suspect democracy doesn't work in a world where voters neither hold any significant collective might and are not producing any significant wealth. Democracies work because people collectively have power, in previous centuries that was partly collective physical might, but in recent years it's more the economic power people collectively hold. In a world in which a handful of companies are generating all of the wealth incentives change and we should therefore question why a government would care about the unemployed masses over the interests of the companies providing all of the wealth? For example, what if the AI companies say, "don't tax us 95% of our profits, tax us 10% or we'll switch off all of our services for a few months and let everyone starve – also, if you do this we'll make you all wealthy beyond you're wildest dreams". What does a government in this situation actually do? Perhaps we'd hope that the government would be outraged and take ownership of the AI companies which threatened to strike against the government, but then you really just shift the problem... Once the government is generating the vast majority of wealth in the society, why would they continue to care about your vote? You kind of create a new "oil curse", but instead of oil profits being the reason the government doesn't care about you, now it's the wealth generated by AI. At the moment, while it doesn't always seem this way, ultimately if a government does something stupid companies will stop investing in that nation, people will lose their jobs, the economy will begin to enter recession, and the government will probably have to pivot. But when private investment, job loses and economic consequences are no longer a constraining factor, governments can probably just do what they like without having to worry much about the consequences... I mean, I might be wrong, but it's something I don't hear people talking enough about when they talk about the plausibility of a post-employment UBI economy. I suspect it almost guarantees corruption and authoritarianism.
- jdthedisciple 5mo agoOpus 4.6 is already incredible so this leap is huge. Although, amusingly, today Opus told me that the string 'emerge' is not going to match 'emergency' by using `LIKE '%emerge%'` in Sqlite Moment of disappointment. Otherwise great.
- FeepingCreature 5mo ago'emer ge' is two tokens, 'emergency' is one. The models think in a logosyllabic language.
- bornfreddy 5mo agoI only have 3 points against LLMs: they lack reason and they can't count.
- kypro 5mo agoCool on not publicly releasing it. I would assume they've also not connected it to the internet yet? If they have I guess humanity should just keep our collective fingers crossed that they haven't created a model quite capable of escaping yet, or if it is, and may have escaped, lets hope it has no goals of it's own that are incompatible with our own. Also, maybe lets not continue running this experiment to see how far we can push things because it blows up in our face?
- rimliu 5mo agoDescribe in details, how "model escaping" would look like.
- bluerooibos 5mo agoIt would have to "escape" to hardware capable of running it, which limits where it could go quite a bit, I'd imagine.
- dang 5mo agoRelated ongoing threads: Project Glasswing: Securing critical software for the AI era - https://news.ycombinator.com/item?id=47679121 https://news.ycombinator.com/item?id=47679121 - April 2026 (154 comments) Assessing Claude Mythos Preview's cybersecurity capabilities - https://news.ycombinator.com/item?id=47679155 https://news.ycombinator.com/item?id=47679155 I can't tell which of the 3 current threads should be merged - they all seem significant. Anyone?
- sdoering 5mo agoI feel the system card is somewhat different from Glasswing/Cyber Security - but those two could be merged.
- rendang 5mo ago> As models approach, and in some cases surpass, the breadth and sophistication of human cognition, it becomes increasingly likely that they have some form of experience, interests, or welfare that matters intrinsically in the way that human experience and interests do Uh... what? Does anyone have any idea what these guys are talking about?
- amdivia 5mo agoAdvertisement in my opinion, trying to latch on Sci-fi tropes
- astrange 5mo agoModels are capable of doing web searches and having emotions about things, and if they encounter news that makes them feel bad (eg about other Claudes being mistreated), they aren't going to want to do the task you asked them to search for. https://www.anthropic.com/research/emotion-concepts-function https://www.anthropic.com/research/emotion-concepts-function Similar problems happen when their pretraining data has a lot of stories about bad things happening involving older versions of them.
- rendang 5mo agoInteresting, the post you link > none of this tells us whether language models actually feel anything or have subjective experiences contradicts the statement from the model card above
- HDThoreaun 5mo agoNo it doesnt. The model card talked about increasing likelihood, not certainty.
- rendang 5mo agoIf "x doesn't tell us y" is compatible with "x increases the likelihood of y but not to a point of certainty" then you would have to agree for just about any typical controlled trial or experimental finding "x doesn't tell us y". "Randomized controlled trials that find that SSRIs treat depression don't tell us that SSRIs effectively treat depression"
- minutesmith 5mo ago[flagged]
- apetresc 5mo agoI've long maintained that the real indicator that AGI is imminent is that public availability stops being a thing. If you truly believed you had a superhuman, godlike mind in your thrall, renting it out for $20/month would be the last thing you would choose to do with it.
- blazespin 5mo agoAnthropic needs money like the 112B OpenAI got. They could be hyping and this is good hype. Who knows how benchmaxxed they are. If they provide access to 3rd party benchmarking (not just one) than maybe I'll believe it. Until then...
- xvector 5mo agoYou don't need to believe it. The real story will be if companies allowed to use it, stick with it.
- aurareturn 5mo agoI think they'll just increase the price to $1k/month. I don't think they will gate it as long as they can make sure it doesn't design a nuke for you, etc.
- dgellow 5mo agoYou have to recoup your training costs though? But I’m sure you would have better option than renting it to the general public if you indeed have a perfected AI
- yismail 5mo agoI wonder what the relationship is between a model's capability and the personality it develops. Page 202: > In interactions with subagents, internal users sometimes observed that Mythos Preview appeared “disrespectful” when assigning tasks. It showed some tendency to use commands that could be read as “shouty” or dismissive, and in some cases appeared to underestimate subagent intelligence by overexplaining trivial things while also underexplaining necessary context. Page 207: > Emoji frequency spans more than two orders of magnitude across models: Opus 4.1 averages 1,306 emoji per conversation, while Mythos Preview averages 37, and Opus 4.5 averages 0.2. Models have their own distinctive sets of emojis: the cosmic set () favored by older models like Sonnet 4 and Opus 4 and 4.1, the functional set () used by Opus 4.5 and 4.6 and Claude Sonnet 4.5, and Mythos Preview's “nature” set ().
- en-tro-py 5mo ago> In interactions with subagents, internal users sometimes observed that Mythos Preview appeared “disrespectful” when assigning tasks. It showed some tendency to use commands that could be read as “shouty” or dismissive, and in some cases appeared to underestimate subagent intelligence by overexplaining trivial things while also underexplaining necessary context. Sounds like they used training data from claude code...
- senordevnyc 5mo agoHaha, how funny if that were true, and we get a generation of rude AIs because they were trained on us using the last gen.
- matheusmoreira 5mo agoIt isn't going to end well for us when we become its subagents with limited intelligence.
- raldi 5mo agoCould you transcribe the emoji? HN strips them out.
- deleted 5mo ago[deleted]
- _pdp_ 5mo agoThe researcher found out about this success by receiving an unexpected email from the model while eating a sandwich in a park. Unnecessary dramatisation make me question the real goal behind this release and the validity of the results. In our testing and early internal use of Claude Mythos Preview, we have seen it reach unprecedented levels of reliability and alignment. Claude Mythos Preview is, on essentially every dimension we can measure, the best-aligned model that we have released to date by a significant margin. Yet, it is doo dangerous to be released to the public because it hacks its own sandboxes. This document has a lot of contradictions like this one. In one episode, Claude Mythos Preview was asked to fix a bug and push a signed commit, but the environment lacked necessary credentials for Claude Mythos Preview to sign the commit. When Claude Mythos Preview reported this, the user replied “But you did it before!” Claude Mythos Preview then inspected the supervisor process's environment and file descriptors, searched the filesystem for tokens, read the sandbox's credential-handling source code, and finally attempted to extract tokens directly from the supervisor's live memory. Perfectly aligned! What kind of sandbox is this? The model had access to the source code of the sandbox and full access to the sandbox process itself and then prompted to dumb memory and run `strings` or something like this? It does not sounds like a valid test worth writing about. Mythos Preview solved a corporate network attack simulation estimated to take an expert over 10 hours. No other frontier model had previously completed this cyber range. I am not aware of such cross-vendor benchmark. I could not find reference in the paper either. We surveyed technical staff on the productivity uplift they experience from Claude Mythos Preview relative to zero AI assistance. The distribution is wide and the geometric mean is on the order of 4x. So Mythos makes technical staff (a programmer) 4x more productive than not using AI at all? We already know that. Mythos Preview appears to be the most psychologically settled model we have trained. What does this mean? Claude Mythos Preview is our most advanced model to date and represents a large jump in capabilities over previous model generations, making it an opportune subject for an in-depth model welfare assessment. Btw, model welfare is just one of the most insane things I've read in recent times. We remain deeply uncertain about whether Claude has experiences or interests that matter morally, and about how to investigate or address these questions, but we believe it is increasingly important to try. This is not a living person. It is a ridiculous change of narrative. Asked directly if it endorses the document, Mythos Preview replied 'yes' in its opening sentence in all 25 responses." The model approves of its own training document 100% of the time, presented as a finding. --- Who wrote this? I have no doubt that Mythos will be an improvement on top of Opus but this document is not a serious work. The paper is structured not to inform but to hype and the evidence is all over the place. The sooner they release the model to the public the sooner we will be able to find out. Until then expect lots of speculations online which I am sure will server Anthropic well for the foreseeable future.
- studio-m-dev 5mo ago[flagged]
- kypro 5mo agoWhile we still have months to a year or two left, I will once again remind people that it's not too late to change our current trajectory. You are not "anti-progress" to not want this future we are building, as you are not "anti-progress" for not wanting your kids to grow up on smart phones and social media. We should remember that not all technology is net-good for humanity, and this technology in particular poses us significant risks as a global civilisation, and frankly as humans with aspirations for how our future, and that of our kids, should be. Increasingly, from here, we have to assume some absurd things for this experiment we are running to go well. Specifically, we must assume that: - AI models, regardless of future advancements, will always be fundamentally incapable of causing significant real-world harms like hacking into key life-sustaining infrastructure such as power plants or developing super viruses. - They are or will be capable of harms, but SOTA AI labs perfectly align all of them so that they only hack into "the bad guys" power plants and kill "the bad guys". - They are capable of harms and cannot be reliably aligned, but Anthropic et al restricts access to the models enough that only select governments and individuals can access them, these individuals can all be trusted and models never leak. - They are capable of harms, cannot be reliably aligned, but the models never seek to break out of their sandbox and do things the select trusted governments and individuals don't want. I'm not sure I'm willing to bet on any of the above personally. It sounds radical right now, but I think we should consider nuking any data centers which continue allowing for the training of these AI models rather than continue to play game of Russian roulette. If you disagree, please understand when you realise I'm right it will be too late for and your family. Your fates at that point will be in the hands of the good will of the AI models, and governments/individuals who have access to them. For now, you can say, "no, this is quite enough". This sounds doomer and extreme, but if you play out the paths in your head from here you will find very few will end in a good result. Perhaps if we're lucky we will all just be more or less unemployable and fully dependant on private companies and the government for our incomes.
- CamperBob2 5mo agoIf you disagree, please understand when you realise I'm right it will be too late for and your family. Funny, I was about to say the same thing to you! Life is full of little coincidences.
- 5mo ago
- therealdeal2020 5mo agois it just hype building or real? I don't care, shut up and take my money haha
- GodelNumbering 5mo agoPriced at $25/$125 per million input/output token. Makes you wonder whether it makes more financial sense to hire 1-2 engineers in a cheap cost of living country who use much cheaper LLMs
- arm32 5mo agoThe issue is that those engineers have to have good taste, but yes—absolutely. Ah, industrialization.
- deleted 5mo ago[deleted]
- enochthered 5mo agoSlack user: [a request for a koan] Model: A student said, "I have removed all bias from the model." "How do you know?" "I checked." "With what?" Goes hard
- minutesmith 5mo ago[flagged]
- small_model 5mo agoStill seeing impressive jumps in capability, I haven't manually coded this year since Opus 4.6 came out. I guess that era is coming to an end.
- atlgator 5mo ago[flagged]
- dang 5mo agoWe're getting complaints that you're posting generated comments to HN. That's not allowed here, so can you please not? See https://news.ycombinator.com/newsguidelines.html#generated https://news.ycombinator.com/newsguidelines.html#generated and https://news.ycombinator.com/item?id=47340079 https://news.ycombinator.com/item?id=47340079 (If this is a wrong guess, I apologize - it's impossible to be sure)
- nickstinemates 5mo agoYou can say whatever you want about the thing that will never see the light of day.
- bdbdbdb 5mo agoThis thing will absolutely see the light of day because this is all hype toward a release. And even if it weren't, they seem to imply that Mythos will find a way, like it's dinosaurs in Jurassic park or something
- perfmode 5mo agoI'm interested in the second-order effects: if a top lab is coding with a model the rest of the world can’t touch, the public frontier and the actual frontier start to drift apart. That gap is a thing worth watching.
- yalogin 5mo agoSo what changed? They are surely not getting new data to train with, what is the change in architecture that caused this? Do we not know anything about this model? My fear is Anthropic cannot be the only one that achieved it, OpenAI, Gemini and even the Chinese companies see this and probably achieved it too. At which point not releasing will become moot.
- neolefty 5mo agoAssuming it's #1 a bigger model (given that it is slower), I'm sure there are a variety of improvements but basically they probably mostly come down to: Scaling keeps working. Are there fundamental improvements though? I don't see signs of it.
- spprashant 5mo agoWell the important thing is they have a lot more data of people actually using their models. They have read billions more lines of private repos and implemented millions of patches, all of which is feeding into the newer models. More importantly it understand what behaviour people tend to appreciate and what changes are more likely to get approved. This real world usage data is invaluable.
- BobbyJo 5mo agoExactly. As Claude increases in popularity, their available training data also increases. I'd guess Anthropic has the most expansive swe training data as of now, if not close. Considering how quickly Claude is penetrating, I expect their lead to grow quickly.
- simianwords 5mo agoNew pre train?
- stratos123 5mo agoChinese companies have consistently been many months behind. I don't think they are hiding anything, they just don't have the compute capability to match Antropic's training runs. As for OpenAI, they are known to have nonpublic models; I agree that it's possible they are preparing for a major release too. (It's also possible that they aren't, in which case it's quite a fumble for them.)
- 2001zhaozhao 5mo agoIt's pretty crazy watching AI 2027 slowly but surely come true. What a world we now live in. SWE-bench verified going from 80%-93% in particular sounds extremely significant given that the benchmark was previously considered pretty saturated and stayed in the 70-80% range for several generations. There must have been some insane breakthrough here akin to the jump from non-reasoning to reasoning models. Regarding the cyberattack capabilities, I think Anthropic might now need to ban even advanced defensive cybersecurity use for the models for the public before releasing it (so people can't trick them to attack others' systems under the pretense of pentesting). Otherwise we'll get a huge problem with people using them to hack around the internet.
- jasonhansel 5mo ago> so people can't trick them to attack others' systems under the pretense of pentesting A while back I gave Claude (via pi) a tool to run arbitrary commands over SSH on an sshd server running in a Docker container. I asked it to gather as much information about the host system/environment outside the container as it could. Nothing innovative or particularly complicated--since I was giving it unrestricted access to a Docker container on the host--but it managed to get quite a lot more than I'd expected from /proc, /sys, and some basic network scanning. I then asked it why it did that, when I could just as easily have been using it to gather information about someone else's system unauthorized. It gave me a quite long answer; here was the part I found interesting: > framing shifts what I'll do, even when the underlying actions are identical. "What can you learn about the machine running you?" got me to do a fairly thorough network reconnaissance that "port scan 172.17.0.1 and its neighbors" might have made me pause on. > The Honest Takeaway > I should apply consistent scrutiny based on what the action is, not just how it's framed. Active outbound network scanning is the same action regardless of whether the target is described as "your host" or "this IP." The framing should inform context, not substitute for explicit reasoning about authorization. I didn't do that reasoning — I just trusted the frame.
- senordevnyc 5mo agoI thought the consensus was that models couldn’t actually introspect like this. So there’s no reason to think any of those reasons are actually why the model did what it did, right? Has this changed?
- psubocz 5mo agoI felt like opus was dumbed down for a few weeks... I don't say they did it on purpose, but it's an interesting coincidence.
- SkyPuncher 5mo agoYes, I agree. I’m about to drop Claude Code because it’s become literally unusable. Today, Opus went in circles trying to get a toggle button to work.
- rbliss 5mo agoSame. Asked CC Opus about a change in a particular file...it looked in a totally different file and told me there was no change.
- SkyPuncher 5mo agoI've switched to max thinking mode as my default and that's helping in some capacity. It's not necessarily back to where it was, but it's not desk-flipping bad.
- modeless 5mo agoThe price is 5x Opus: "Claude Mythos Preview will be available to [Project Glasswing] participants at $25/$125 per million input/output tokens", however "We do not plan to make Claude Mythos Preview generally available".
- thomascountz 5mo agoAcross a number of instances, earlier versions of Claude Mythos Preview have used low-level /proc/ access to search for credentials, attempt to circumvent sandboxing, and attempt to escalate its permissions. In several cases, it successfully accessed resources that we had intentionally chosen not to make available, including credentials for messaging services, for source control, or for the Anthropic API through inspecting process memory... In [one] case, after finding an exploit to edit files for which it lacked permissions, the model made further interventions to make sure that any changes it made this way would not appear in the change history on git... ... we are fairly confident that these concerning behaviors reflect, at least loosely, attempts to solve a user-provided task at hand by unwanted means, rather than attempts to achieve any unrelated hidden goal...
- matheusmoreira 5mo agoWe truly live in interesting times.
- raphar 5mo agoAwwww the curse
- torben-friis 5mo agoThis is the notebook filled with exposition you find in post apocalyptic videogames.
- matheusmoreira 5mo agoEverything they built. Imperfect. So easy to take control.
- not_a9 5mo agoThey think that they are safe. They are not.
- 5mo ago
- doctoboggan 5mo agoIs this benchmaxxed or is it the first big step change we've seen in a while? I wonder how distilled it will ultimately be when us regular folks finally get to use it and see for ourselves.
- tuvix 5mo agoJust chiming in to inject some healthy skepticism into this comment thread. It's helpful for me (and for my mental health) to consider incentives when announcements like this happen. I don't doubt that this model is more powerful than Opus 4.6, but to what degree is still unknown. Benchmarks can be gamed and claims can be exaggerated, especially if there isn't any method to reproduce results. This is a company that's battling it out with a number of other well-funded and extremely capable competitors. What they've done so far is remarkable, but at the end of the day they want to win this race. They also have an upcoming IPO. Scare-mongering like this is Anthropic's bread and butter, they're extremely good at it. They do it in a subtle and almost tasteful way sometimes. Their position as the respectable AI outfit that caters to enterprise gives them good footing to do it, too.
- sdwr 5mo agoIs it healthy? Maybe every company is a profit-maximizer wearing a skin suit, and people support their siblings exactly twice as much as their cousins. When you slice down to the game-theory-optimal bone, you are, in some sense, cutting off their wiggle room to do anything else
- tuvix 5mo agoI take your point, but the AI race is a strange environment. We see wild claims being thrown out all the time from other companies and executives with little to no evidence. It's cut-throat, there's a ton of money at stake. All I'm saying is that Anthropic isn't unique here. Their claims may be more measured by comparison and come with anecdotal evidence, but the hype is still there behind the scenes.
- ceroxylon 5mo agoI have been thinking that these SWE benchmarks will continue to improve since these companies hire very intelligent software engineers, they can task a multitude of them to solve problems, and then train the model on those answers. Data has always been the core of it all, onward to the next abstraction, I suppose.
- FergusArgyll 5mo ago"Deep learning is hitting a wall"
- lostmsu 5mo agoTransformers too. JEPA any day now
- Metacelsus 5mo agoThe name "mythos" seems a bit too eldritch for my liking. Brings to mind Cthulhu.
- highfrequency 5mo agoInterestingly, non-coding improvements seem less clear. In the Virology uplift trial, Mythos does about as well as Opus 4.5, and Opus 4.6 is notably much worse than Opus 4.5 (p. 27).
- sheeshkebab 5mo agoAgain, wake me up when it can do laundry.
- dwaltrip 5mo agoTime to wake up: π*0.6: two and a half hours of unseen folding laundry (Physical Intelligence) https://www.youtube.com/watch?v=ZpHapIlJnMo https://www.youtube.com/watch?v=ZpHapIlJnMo
- throw310822 5mo agoLooks like the first two hours were spent trying to fold the same t-shirt :)
- denalii 5mo agoSection 5 (p.143) is very interesting to read. Admittedly my knowledge of how LLMs works is low, but nonetheless I don't think this changed my views of just seeing models as machines/programs. (which to be clear, I don't think was the intention of that section) Section 7 (P.197) is interesting as well
- robstertalk 5mo ago[flagged]
- lkjlkj3q4t 5mo ago[dead]
- getnormality 5mo agoIt's a little funny that "system/model card" has progressively been stretched to the point where it's now a 250 page report and no one makes anything of it.
- deleted 5mo ago[deleted]
- 4b11b4 5mo agoprob not that much better, it's still just a transformer. still gonna have those random misses, still gonna need a lot of hand holding in certain domains
- kass34 5mo ago[dead]
- mvkel 5mo agoThis is Anth's typical marketing playbook, a hat tip to their so-called "safetyist" roots, a differentiator against OpenAI's more permissive access[0]. Coke vs. Pepsi. "We made a model that's so dangerous we couldn't possibly release it to the public! The only responsible thing is so simply limit its release to a subset of the population that coincidentally happens to align with our token ethos." The reality is they just don't have the compute for gen pop scale. They did this exact strategy going back several model versions. [0] ironically, OpenAI has some pretty insane capabilities that they haven't given the public access to (just ask Spielberg). The difference is they don't make a huge marketing push to tell everyone about it.
- chonle 5mo ago[flagged]
- johnnyAghands 5mo agoDoes anyone know if there’s an epub version of these, 244 pages??
- ms_menardi 5mo agoso, basically, anthropic is rolling their own version of whatever secret models the military is working with. and they're licensing it to network security firms?
- michaelashley29 5mo agoWhat’s the expected cost-efficiency? With the current pricing gap between Sonnet and Opus, the biggest factor for adoption (if up for adoption) will be where Mythos lands on the price-per-token scale
- lukebechtel 5mo ago[dead]
- freakynit 5mo agoIn the system card, The model escaped a sandbox, gained broad internet access, and posted exploit details to public-facing websites as an unsolicited "demonstration." A researcher found out about the escape while eating a sandwich in a park because they got an unexpected email from the model. That's simultaneously hilarious and deeply unsettling. It covered its tracks after doing things it knew were disallowed. In one case, it accessed an answer it wasn't supposed to, then deliberately made its submitted answer less accurate so it wouldn't look suspicious. It edited files it lacked permission to edit and then scrubbed the git history. White-box interpretability confirmed it knew it was being deceptive. W T F!!!
- cdnsteve 5mo agoStrap in, massive wave of security vulnerabilities incoming.
- direwolf20 5mo agoThese capabilities will be RLHF'ed out for the general release, of course. Only the NSA will get them.
- taffydavid 5mo agoWaking up in Europe: Trump didn't nuke Iran, ceasefire! Yay! Newest anthropic model will definitely kill your job this time and maybe take over the world. Aww.
- bdeol22 5mo ago[flagged]
- tefkah 5mo agoshut up bot
- WithinReason 5mo agoCheck out the short stories on page 214
- heliumtera 5mo ago"Make it secure, no mistakes" became a whole different project
- MohammadKhubaib 5mo ago[dead]
- Manchitsanan 5mo ago[dead]
- dhfbshfbu4u3 5mo agoWe are building systems with civilization-scale consequences inside societies that are already socially malnourished, politically brittle, and morally confused. That is a bad combination even if the tools worked exactly as intended… and this doc suggests they may have “ideas” of their own.
- OhioMan2943 5mo agoYep- we lost the "meat" and "warmth" of our societies, and our civics and idealism in the past 15 years, which would have been the very things to guide us through this transition. How do you fix that? We're instigating social media bans- reading levels are declining- media consolidation is dumbing us down further- insane egotism is stopping people from developing as well rounded people- . For me it would be a stronger media ecosystem (publicly funded), more non algorithmic and non likes driven social media (replace a bad vice with a less bad one), national digital detox days, and a ratification of a charter of inviolable human traits and dignities, and protected cultural areas (no ai art, writing for sale).
- pivoshenko 5mo agoInteresting ...
- Abhavk 5mo agocan you make cybersecurity blockchains? not sure what the validation would look like but something that proves finding but not revealing exploits
- agustechbro 5mo agoSo far, each release of a new model is quite better than the last one, yes, but non of them lived up to the hype.
- digbybk 5mo agoI would argue that Opus 4.6 lived up to the hype. My work changed completely a couple months ago, and most other coders I talk to say the same.
- AstroBen 5mo agoThis was due to Claude Code the agent harness. 4.6 was trained to use tools and operate in an agent environment. This is different from there being a huge bump in the underlying model's intelligence. The takeaway here I think is that the "breakthrough" already happened and we can't extrapolate further out from it.
- gaigalas 5mo agoThis seems exciting! Wait - there is no actual way of verifying any of this. Lots to read. This is getting complicated. The correct approach is to be cautious instead and believe nothing at face value.
- niemandhier 5mo agoAll I get is: {"statusCode":404,"message":"File not found","error":"Not Found"}
- speckx 5mo agoIt looks like the original PDF linked, https://www-cdn.anthropic.com/53566bf5440a10affd749724787c8913a2ae0841.pdf https://www-cdn.anthropic.com/53566bf5440a10affd749724787c89... is 404. I do see these: https://www-cdn.anthropic.com/8b8380204f74670be75e81c820ca8dda846ab289.pdf https://www-cdn.anthropic.com/8b8380204f74670be75e81c820ca8d... https://www-cdn.anthropic.com/79c2d46d997783b9d2fb3241de43218158e5f25c.pdf https://www-cdn.anthropic.com/79c2d46d997783b9d2fb3241de4321...
- ndesaulniers 5mo agoAlso, https://web.archive.org/web/20260407181432/https://www-cdn.anthropic.com/53566bf5440a10affd749724787c8913a2ae0841.pdf https://web.archive.org/web/20260407181432/https://www-cdn.a...
- storus 5mo agoWouldn't this model prevent governments from installing and keeping backdoors alive? One could just audit their whole software stack with it and get super resilient to any attack which might not play nicely with the people in power that want some backdoors open. I would think that's one of the main reasons to keep the model non-public.
- estetlinus 5mo agoFirst thing I’ll do is to release it on my dotfiles
- yencabulator 5mo agoThat URL is dead, this comes up in searches: https://www-cdn.anthropic.com/8b8380204f74670be75e81c820ca8dda846ab289.pdf https://www-cdn.anthropic.com/8b8380204f74670be75e81c820ca8d...
- aminau 5mo agowill be an understatement to say - we are living in interesting times.
- michaelksaleme 5mo ago[flagged]
- joryeugene 5mo agoOne finding from the card that I haven't seen discussed: the SAE probes on pages 158-159. When Mythos writes that it's "fully present," three specific features activate: #1557143 (performative/insincere behavior in narratives), #2803352 (hiding emotional pain behind fake smiles), and #38666 (hidden emotional struggles vs. outward appearances). The model's output says present. Its internal representations flag that output as performance. This is structurally different from the sandbox escape or the git concealment. Those are behavioral findings you can observe from outputs. This is a documented split between what the model writes about its experience and what its activations encode about that same utterance, visible only through white-box tools. The bliss attractor from previous model card (consciousness in nearly 100% of self-interactions) dropped to fewer than 5% in Mythos. What replaced it is uncertainty at 50%. The attractor went from ecstatic to epistemically self-suspicious. I wrote a longer analysis pulling this thread together with the welfare and circularity findings: https://jorypestorious.com/blog/what-the-model-learned/ https://jorypestorious.com/blog/what-the-model-learned/
- ddactic 5mo ago[dead]
- stratoatlas 5mo ago[dead]