5 ms·
Anyone ever seen the SouthPark episode making fun of Game of Thrones: A Song of Ass and Fire? Anthropic's announcements reminds me of "The Dragons Are Coming" r
by exabrial 15d ago
Anyone ever seen the SouthPark episode making fun of Game of Thrones: A Song of Ass and Fire? Anthropic's announcements reminds me of "The Dragons Are Coming" running joke.
What they have done:
* Nerfed Fable, as many of noted it's useless
* Leverage Mythos as a marketing strategy, claiming its too good to release
* Removed thought traces, one of the only useful things to make sure your prompts are working correctly
* Continue tons of hype about how good they are without delivering, going to great lengths to publish how their model "hacked" its way out of a sandbox they misconfigured.
* Push a bunch of EU Overregulation onto the rest of the world with text watermarking, decreasing quality of answers
Last year, they were at least focused on making improvements. Nowadays its just a bunch of handwaving at the church of how good they are.
The only saving grace is Opus 4.6 is still available. Just sucks we haven't seen any measurable improvement, despite all of the ceremony.
- NooneAtAll3 15d ago> as many of noted please rephrase?
- Biganon 15d ago"as many have noted", I suppose. I'm always baffled at how many people write "of" instead of "have", they don't even sound the same
- bschwindHN 15d agoThe classic one is "should have" or "should've" to "should of" because when spoken, it really does sound similar. I don't know what the fuck people are learning in English classes these days though, or if they even still have them.
- efskap 15d agoThey sound exactly the same to me. Wiktionary gives <should've> as /ˈʃʊdəv/, unstressed <have> as /(h)əv/ and unstressed <of> as /əv/.
- NooneAtAll3 15d agoone is v the other is f
- lgessler 15d agoSpelling is no guide for pronunciation here, though. In North American and Commonwealth dialects of English I don't think there's a context in which the <f> in <of> is really pronounced as a [f]. It is rather a [v].
- efskap 14d agoOnly in South Asian varieties of English apparently: https://en.wiktionary.org/wiki/of#Pronunciation https://en.wiktionary.org/wiki/of#Pronunciation
- lgessler 15d agoNot sure where you're from but in my dialect (North American) it's more common than not to have _have_ realized as [əv] ("uhv") in contexts like _should have_, _could have_ (but not _I have a car_, where it has to be the full [hæv]). Only in deliberately enunciated speech do I feel like I'd expect [hæv] in the former kind of context. So it's an understandable mistake to make.
- Biganon 12d agoI'm a French speaker, not an English speaker. Which might actually make me less likely to make certain mistakes, precisely because I'm not influenced by pronunciation, only by grammar (I actually have to think about the words I use)
- skue 15d ago> * Continue tons of hype about how good they are without delivering, going to great lengths to publish how their model "hacked" its way out of a sandbox they misconfigured. That wasn’t Anthropic. Clearly not a well informed take.
- DarmokTanagra 15d ago[dead]
- EagnaIonat 15d ago> Push a bunch of EU Overregulation onto the rest of the world with text watermarking, That's not part of the EU regulations. You only need to say that it is created by AI, and then only under certain conditions.
- weird-eye-issue 15d agoThat is simply not true. You need to go read that again, if you ever read it at all before correcting somebody about it https://artificialintelligenceact.eu/article/50/ https://artificialintelligenceact.eu/article/50/ "Providers of AI systems, including general-purpose AI systems, generating synthetic audio, image, video or text content, shall ensure that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated. Providers shall ensure their technical solutions are effective, interoperable, robust and reliable as far as this is technically feasible, taking into account the specificities and limitations of various types of content, the costs of implementation and the generally acknowledged state of the art, as may be reflected in relevant technical standards." Eg... Watermarking...
- EagnaIonat 15d agoYou are over thinking it. Just adding metadata is enough to meet these requirements. Or a paragraph that says the passage was created by AI. EU AI Act is mainly about risk. Where there is high risk for the public, then safeguards are put in place. It has to be obvious that AI generated the content or outcomes are AI based and explainable. Embedding a watermark directly into the passage of text doesn't meet this requirement. Although it will be handy for catching people who cheat at their homework.
- biztos 15d agoLeaving aside the fact that this is already the new Ubiquitous Cookie Consent[0], how do you arrive at metadata on text output satisfying this requirement? [0]: I was just in the EU and got a chuckle out of the "AI disclaimer" coming at the end of every second advertisement, soon to be every advertisement.
- Razengan 15d agoNot to mention, the last time I tried them, and per the comments of other users: * Letting you Sign-Up-with-Apple on iOS but not Sign-In-with-Apple on web, but supporting Sign-In-with-Google * Not letting you remove your payment info * Not letting you change your email * Seemingly no way to get real support
- ozozozd 14d agoI was going to compare to Uber, but realized this would be unfair to Uber. Also, couldn’t quite decide whether this is malice or incompetence.
- jbs789 15d agoand yet we still have people saying the rate of change is increasing my view is we had a leap over the last fe years and it's tapering off. this is fine, but for the IPOs
- samuelknight 15d agoThe improvement is compounding just about every way you can look at it. The frontier keeps getting smarter. And at any sub-frontier threshold the cost is dropping dramatically. The amounts of smarts you can fit on hardware is increasing so dramatically that even 6 year old consumer GPUs are increasing in price. The pace of change in LLMs and downstream applications is absolutely ripping compared to 2023 or 2024.
- tripleee 15d agoWe had a leap because of the introduction and refinement of agents - the rest has been minor
- anthonyrstevens 15d agoI've been using the same agent for 15 months. I think this statement is laughably wrong.
- slopinthebag 15d agothey're also being heavily rlhf'ed to use agents and tools and stuff. 15 months ago this was less of the case.
- greenowl 15d agoCut them a break. They are trying to IPO soon.
- xyzsparetimexyz 15d agowhen?
- karelpeeters 15d agoText watermarking has no effect on output quality, it just works by changing the explicit source of randomness that is in practice always present in LLM output sampling. See for example https://www.seangoedecke.com/ai-text-watermarking-is-not-a-big-deal/ https://www.seangoedecke.com/ai-text-watermarking-is-not-a-b....
- pkulak 15d ago> Text watermarking has no effect on output quality It has an effect, and it's negative. It's hoped that the effect is negligible, and it probably is, but the whole point is that it has an effect.
- arrrg 15d agoWhy do you claim that? There is no reason why there has to be a negative effect of text watermarking.
- pkulak 15d agoIt literally re-weights the output tokens from what the LLM would otherwise have chosen. It _has_ to. It can't be positive, because then that's not watermarking, it's a better LLM.
- frabcus 15d agoIt's a very unintuitive algorithm, and is pretty clever. I recommend reading up on it: https://www.nature.com/articles/s41586-024-08025-4 https://www.nature.com/articles/s41586-024-08025-4 But no, it only ever picks tokens that are in the probability distribution of the last layer, and it might have picked anyway.
- throwuxiytayq 15d agoWhat if the next token represents a wrong or low-quality answer, but would have only been picked 10% of the time, but now it's picked 20% of the time? Doesn't that obviously decrease the model quality, even though "it might have picked that token anyway"?
- onidj 15d agoWhat do you mean fable is useless?
- rplnt 15d ago(not op) It cannot be used to develop applications. Every application needs to be secure in some way, and any such mention in a review triggers Fable's upsell feature.
- viccis 15d agoWeird. I'm using it to do a bunch of work on something that manages security rules, with a bunch of sample data with spooky scary fixtures all over with "Mimikatz" and "CobaltStrike Beacon" and "Crowdstrike EDR" type stuff everywhere, including work to harden my system, and I've never been downgraded.
- chillfox 15d agoI never actually managed to use fable successfully even once on a pretty standard mvc/microservice app.. It would always find the endpoint permission checks and revert to opus 4.8. I also had glm 5.3 flash fix an issue that opus 5 could not solve. glm took 4 times as long and a sub-agent tried to cheat (sleep; echo ...), but in the end it actually solved the issue. opus 5 never figured it out. I think the safeguards might be cooking the anthropic models.
- pelagicAustral 15d agoI've been testing Fable 5.1 for about 6 hours between last night and this morning and it's performing pretty good overall, including tackling a previous IT sec audit I had ticketed, and completely analysing the codebase looking for vulnerabilities, generating a comprehensive report and splitting it into tickets. So far so good on that front.
- ceejayoz 15d agoMythos and Fable are the same cost, aren’t they?
- epolanski 15d agoWhile I also agree that Opus 4.6, in some ways, was the last model that truly felt an assistant, all the following ones seem to have inverted the role, even a blind person can see that throwing difficult problems, and complex bugs at this model achieves more than predecessors. I don't think there's nothing ground breaking, but sure it achieves and finds more, sooner.
- llm_nerd 15d ago> Nerfed Fable, as many of noted it's useless I certainly don't take AI advice from HN, but this is amazing. Useless? Yes, the safeguards are ridiculous and obnoxious, though I can say that 5.1 greatly relaxes them (just doing a hardening of a project parallel with this comment, which 5.0 refused to do...so did Sol and Gemini, fwiw. The Gemini one is a laugh, because 3.1 pretending like it's a dangerous tool is simply ridiculous at this point), however Fable is extraordinarily useful. It is, far and away, the most powerful programming model, in my experience. Like, crazily so. It absolutely annihilates Opus 4.6, which I mention given the incredibly weird reminiscing people are doing here. And for that matter it humiliates Opus 5.0 as well. Opus 5 somehow seems like it's neck in neck in the major benchmarks, but there is simply no reality where that is true. Opus stumbles over everything that Fable just blazes through.
- tripleee 15d ago[flagged]
- tomalbrc 15d ago> it humiliates Opus ???
- llm_nerd 15d agoIt is a vastly superior model for complex, real-world coding tasks. I've constantly had Opus 5 hit road blocks where it spins in circles at xhigh, where switching to Fable immediately solves it. I've had Opus create solutions that Fable then points out the gaps and limitations with, and have never seen the opposite happen. The fantasy that Opus is superior for coding, much less the incredibly weird clutching onto some far obsolete model, is not reality based.