27 ms·
Ars Technica being caught using LLMs that hallucinated quotes by the author and then publishing them in their coverage about this is quite ironic here. Even on
by Springtime 7mo ago
Ars Technica being caught using LLMs that hallucinated quotes by the author and then publishing them in their coverage about this is quite ironic here.
Even on a forum where I saw the original article by this author posted someone used an LLM to summarize the piece without having read it fully themselves.
How many levels of outsourcing thinking is occurring to where it becomes a game of telephone.
- trollbridge 7mo agoThe amount of effort to click an LLM’s sources is, what, 20 seconds? Was a human in the loop for sourcing that article at all?
- phire 7mo agoHumans aren't very diligent in the long term. If an LLM does something correctly enough times in a row (or close enough), humans are likely to stop checking its work throughly enough. This isn't exactly a new problem we do it with any bit of new software/hardware, not just LLMs. We check its work when it's new, and then tend to trust it over time as it proves itself. But it seems to be hitting us worse with LLMs, as they are less consistent than previous software. And LLM hallucinations are partially dangerous, because they are often plausible enough to pass the sniff test. We just aren't used to handling something this unpredictable.
- potatoman22 7mo agohttps://en.wikipedia.org/wiki/Automation_bias https://en.wikipedia.org/wiki/Automation_bias
- Waterluvian 7mo agoIt’s a core part of the job and there’s simply no excuse for complacency.
- jatora 7mo agoThere's not a human alive that isnt complacent in many ways.
- pixl97 7mo agoThe words on the page are just a medium to sell ads. If shit gets ad views then producing shit is part of the job... unless you're the one stepping up to cut the checks.
- Marsymars 7mo agoArs also sells ad-free subscriptions.
- intended 7mo agoThis is a first degree expectation of most businesses. What the OP pointed out is a fact of life. We do many things to ensure that humans don’t get “routine fatigue”- like pointing at each item before a train leaves the station to ensure you don’t eyes glaze over during your safety check list. This isn’t an excuse for the behavior. Its more about what the problem is and what a corresponding fix should address.
- zahlman 7mo agoThere's a weird inconsistency among the more pro-AI people that they expect this output to pass as human, but then don't give it the review that an outsourced human would get.
- vidarh 7mo agoThe irony is that while from perfect, an LLM-based fact-checking agent is likely to be far more dilligent (but still needs human review as well) by nature of being trivial to ensure it has no memory of having done a long list of them (if you pass e.g. Claude a long list directly in the same context, it is prone to deciding the task is "tedious" and starting to take shortcuts). But at the same time, doing that makes it even more likely the human in the loop will get sloppy, because there'll be even fewer cases where their input is actually needed. I'm wondering if you need to start inserting intentional canaries to validate if humans are actually doing sufficiently torough reviews.
- kortilla 7mo agoThe source would just be the article, which the Ars author used an LLM to avoid reading in the first place.
- prussia 7mo agoThe kind of people to use LLM to write news article for them tend not to be the people who care about mundane things like reading sources or ensuring what they write has any resemblance to the truth.
- adamddev1 7mo agoThe problem is that the LLM's sources can be LLM generated. I was looking up some health question and tried clicking to see the source for one of the LLMs claim. The source was a blog post that contained an obvious hallucination or false elaboration.
- kmeisthax 7mo agoIf a human had enough time to check all the sources they wouldn't have been using an LLM to write for them.
- giobox 7mo agoMore than ironic, it's truly outrageous, especially given the site's recent propensity for negativity towards AI. They've been caught red-handed here doing the very things they routinely criticize others for. The right thing to do would be a mea-culpa style post and explain what went wrong, but I suspect the article will simply remain taken down and Ars will pretend this never happened. I loved Ars in the early years, but I'd argue since the Conde Nast acquisition in 2008 the site has been a shadow of its former self for a long time, trading on a formerly trusted brand name that recent iterations simply don't live up to anymore.
- netsharc 7mo agoProbably "one bad apple", soon to be fired, tarred and feathered...
- zahlman 7mo agoIf Kyle Orland is about to be fingered as "one bad apple" that is pretty bad news for Ars.
- JumpCrisscross 7mo ago“Kyle Orland has been the Senior Gaming Editor at Ars Technica since 2012” [1]. [1] https://arstechnica.com/author/kyle-orland/ https://arstechnica.com/author/kyle-orland/
- rectang 7mo agoThere are apparently two authors on the byline and it’s not hard to imagine that one may be more culpable than the other. You may be fine with damning one or the other before all the facts are known, zahlman, but not all of us are.
- sho_hn 7mo agoI don't read their comment as implying this. It might in fact hint at the opposite; it's far more likely for the less senior author to get thrown under the bus, regardless of who was lazy.
- llbbdd 7mo agoHonestly frustrating that Scott chose not to name and shame the authors. Liability is the only thing that's going to stop this kind of ugly shit.
- rectang 7mo agoThere is no need to rush to judgment on the internet instant-gratification timescale. If consequences are coming for journalist or publication, they are inevitable. We’ll know more in only a couple days — how about we wait that long before administering punishment?
- llbbdd 7mo agoIt's not rushing to judgement, the judgement has been made. They published fraudulent quotes. Bubbling that liability up to Arse Technica is valuable for punishing them too but the journalist is ultimately responsible for what they publish too. There's no reason for any publication to ever hire them again when you can hire ChatGPT to lie for you. EDIT: And there's no plausible deniability for this like there is for typos, or maligned sources. Nobody typed these quotes out and went "oops, that's not what Scott said". Benj Edwards or Kyle Orland pulled the lever on the bullshit slot machine and attacked someone's integrity with the result. "In the past, though, the threat of anonymous drive-by character assassination at least required a human to be behind the attack. Now, the potential exists for AI-generated invective to infect your online footprint."
- rectang 7mo agoWe do not yet know just how the story unfolded between the two people listed on the byline. Consider the possibility that one author fabricated the quotes without the knowledge of the other. The sin of inadequate paranoia about a deceptive colleague is not the same weight as the sin of deception. Now to be clear, that’s a hypothetical and who knows what the actual story is — but whatever it is, it will emerge in mere days. I can wait that long before throwing away two lives, even if you can’t. > Bubbling that liability up to Arse Technica is valuable for punishing them Evaluating whether Ars Technica establishes credible accountability mechanisms, such as hiring an Ombud, is at least as important as punishing individuals.
- epistasis 7mo agoYikes I subscribed to them last year on the strength of their reporting in a time where it's hard to find good information. Printing hallucinated quotes is a huge shock to their credibility, AI or not. Their credibility was already building up after one of their long time contributors, a complete troll of a person that was a poison on their forums, went to prison for either pedophilia or soliciting sex from a minor. Some serious poor character judgement is going on over there. With all their fantastic reporters I hope the editors explain this carefully.
- singpolyma3 7mo agoTBF even journalists who interview people for real and take notes routinely quite them saying things they didn't say. The LLMs make it worse, but it's hardly surprising behaviour from them
- epistasis 7mo agoIt's surprising behavior to come from Ars Technica. But also when journalists misquote it's through a different phrasing of something that Pepe have actually said, sometimes with different emphasis or eve meaning. But of the people I've known who have been misquoted it's always traceable to something they actually did say.
- pmontra 7mo agoI knew first hand about a couple of news in my life. Both were reported quite incorrectly. That was well before LLMs. I assume that every news is quite inaccurate, so I read/hear them to get the general gist of what happened, then I research the details if I care about them.
- justinclift 7mo ago> Their credibility was already building up ... Don't you mean diminishing or disappearing instead of building up? Building up sounds like the exact opposite of what I think you're meaning. ;)
- 7mo ago
- sho_hn 7mo agoAlso ironic: When the same professionals advocating "don't look at the code anymore" and "it's just the next level of abstraction" respond with outrage to a journalist giving them an unchecked article. Read through the comments here and mentally replace "journalist" with "developer" and wonder about the standards and expectations in play. Food for thought on whether the users who rely on our software might feel similarly. There's many places to take this line of thinking to, e.g. one argument would be "well, we pay journalists precisely because we expect them to check" or "in engineering we have test-suites and can test deterministically", but I'm not sure if any of them hold up. The "the market pays for the checking" might also be true for developers reviewing AI code at some point, and those test-suites increasingly get vibed and only checked empirically, too. Super interesting to compare.
- rsynnott 7mo ago> When the same professionals advocating "don't look at the code anymore" and "it's just the next level of abstraction" respond with outrage to a journalist giving them an unchecked article. I doubt, by and large, that it's the same people. Just as this LLM misquoting is journalistic malpractice, "don't look at the code anymore" is engineering malpractice.
- ffsm8 7mo agoWhile I don't subscribe to the idea that you shouldn't look at the code - it's a lot more plausible for devs because you do actually have ways to validate the code without looking at it. E.g you technically don't need to look at the code if it's frontend code and part of the product is a e2e test which produces a video of the correct/full behavior via playwright or similar. Same with backend implementations which have instrumentation which expose enough tracing information to determine if the expected modules were encountered etc I wouldn't want to work with coworkers which actually think that's a good idea though
- Pay08 7mo agoIf you tried this shit in a real engineering principle, you'd end up either homeless or in prison in very short order.
- sphars 7mo agoAurich Lawson (creative director at Ars) posted a comment[0] in response to a thread about what happened, the article has been pulled and they'll follow-up next week. [0]: https://arstechnica.com/civis/threads/journalistic-standards.1511650/post-44249741 https://arstechnica.com/civis/threads/journalistic-standards...
- _HMCB_ 7mo agoIt’s funny they say the article “may have” run afoul of their journalistic standards. May have is carrying a lot of weight there.
- llbbdd 7mo agoThe article "may have" drawn too much attention to how little they care.
- arduanika 7mo agoEquivalently: Our standards "may have" been low enough that this was just fine, actually.
- pseudalopex 7mo agoSaying may have during an investigation was unremarkable.
- usefulposter 7mo agoJust like in the original thread that was wiped (https://news.ycombinator.com/item?id=47012384 https://news.ycombinator.com/item?id=47012384), Ars Subscriptors continue to display lack of reading comprehension and jump to defending Condé Nast. All threads have since been locked: https://arstechnica.com/civis/threads/journalistic-standards.1511650/ https://arstechnica.com/civis/threads/journalistic-standards... https://arstechnica.com/civis/threads/is-there-going-to-be-a-mea-culpa-article-about-the-after-a-routine-code-rejection-an-ai-agent-published-a-hit-piece-on-someone-by-name-article.1511656/ https://arstechnica.com/civis/threads/is-there-going-to-be-a... https://arstechnica.com/civis/threads/um-what-happened-to-the-article-after-a-routine-code-rejection-an-ai-agent-published-a-hit-piece-on-someone-by-name.1511658/ https://arstechnica.com/civis/threads/um-what-happened-to-th...
- deleted 7mo ago[deleted]
- neya 7mo agoArs Technica has always trash even before LLMs and is mostly an advertisement hub for the highest bidder
- Lerc 7mo agoHas it been shown or admitted that the quotes were hallucinations, or is it the presumption that all made up content is a hallucination now?
- Pay08 7mo agoYou could read the original blog post...
- Lerc 7mo agoHow could that prove hallucinations? It could only possibly prove that they are not. If the quotes are in the original post then they are not hallucinations. If they are not in the post they could be caused by something is not a LLM. Misquotes and fabricated quotes have existed long before AI, And indeed, long before computers.
- DonHopkins 7mo ago[dead]
- Lerc 7mo agoThere is no goalpost moving here. I read the article. My claim is as it has always been. If we accept that the misquotes exist it does not follow that they were caused by hallucinations? To tell that we would still need additional evidence. The logical thing to ask would be; Has it been shown or admitted that the quotes were hallucinations?
- Dylan16807 7mo agoYou've deeply misunderstood their argument in some way I can't quite figure out. It's simple. We know the quotes are fake, but we don't know for sure if they're hallucinations. The blog post does not resolve this uncertainty. And yes other answers are reasonably plausible. You said in another comment that they're "retreating" and "refusing to read" and... no. Your insults are not justified at all.
- 7mo ago
- usefulposter 7mo agoIncredible. When Ars pull an article and its comments, they wipe the public XenForo forum thread too, but Scott's post there was archived. Username scottshambaugh: https://web.archive.org/web/20260213211721/https://arstechnica.com/civis/threads/after-a-routine-code-rejection-an-ai-agent-published-a-hit-piece-on-someone-by-name.1511649/ https://web.archive.org/web/20260213211721/https://arstechni... >Scott Shambaugh here. None of the quotes you attribute to me in the second half of the article are accurate, and do not exist at the source you link. It appears that they themselves are AI hallucinations. The irony here is fantastic. Instead of cross-checking the fake quotes against the source material, some proud Ars Subscriptors proceed to defend Condé Nast by accusing Scott of being a bot and/or fake account. EDIT: Page 2 of the forum thread is archived too. This poster spoke too soon: >Obviously this is massive breach of trust if true and I will likely end my pro sub if this isnt handled well but to the credit of ARS, having this comment section at all is what allows something like this to surface. So kudos on keeping this chat around.
- bombcar 7mo agoThis is just one of the reasons archiving is so important in the digital era; it's key to keeping people honest.
- Imustaskforhelp 7mo agoYes, Wayback machine/archive.org is one of the best websites on the whole world wide web.
- webXL 7mo agoAgreed and that's why there's an incentive to DDoS it and degrade the quality. Are there any p2p backup solutions?
- bombcar 7mo agoThere are some various attempts, the problem is reliability - not that they're always up, but how do you trust them? If archive.org shows a page at a date, you presume it is true and correct. If I provide a PDF of a site at a date, you have no reason to believe I didn't modify the content before PDFing it.
- 0xbadcafebee 7mo ago> How many levels of outsourcing thinking is occurring to where it becomes a game of telephone How do you know quantum physics is real? Or radio waves? Or just health advice? We don't. We outsource our thinking around it to someone we trust, because thinking about everything to its root source would leave us paralyzed. Most people seem to have never thought about the nature of truth and reality, and AI is giving them a wake-up call. Not to worry though. In 10 years everyone will take all this for granted, the way they take all the rest of the insanity of reality for granted.
- DonHopkins 7mo ago[dead]
- JPKab 7mo agoI just wish people would remember how awful and unprofessional and lazy most "journalists" are in 2026. It's a slop job now. Ars Technica, a supposedly reputable institution, has no editorial review. No checks. Just a lazy slop cannon journalist prompting an LLM to research and write articles for her. Ask yourself if you think it's much different at other publications.
- joquarky 7mo agoI would assume that most who were journalists 10 years ago have now either gone independent or changed careers The ones that remain are probably at some extreme on one or more attributes (e.g. overworked, underpaid) and are leaning on genAI out of desperation.
- troyvit 7mo agoI work with the journalists at a local (state-wide) public media organization. It's night and day different from what is described at ars. These are people who are paid a third (or less) of what a sales engineer at meta makes. We have editorial review and ban LLMs for any editorial work except maybe alt-text if I can convince them to use it. They're over-worked, underpaid, and doing what very few people here (including me) have the dedication to do. But hey, if people didn't hate journalists they wouldn't be doing their job.
- moomin 7mo agoIronically, if you actually know what you’re doing with an LLM, getting a separate process to check the quotations are accurate isn’t even that hard. Not 100% foolproof, because LLM, but way better than the current process of asking ChatGPT to write something for you and then never reading it before publication.
- Springtime 7mo agoThe wrinkle in this case is the author blocked AI bots from their site (doesn't seem to be a mere robots.txt exclusion from what I can tell), so if any such bot were trying to do this it may have not been able to read the page to verify, so instead made up the quotes. This is what the author actually speculated may have occurred with Ars. Clearly something was lacking in the editorial process though that such things weren't human verified either way.
- seanhunter 7mo agoIt’s fascinating that on the one hand Ars Technica didn’t think the article was worth writing (so got an LLM to do it) but expect us to think it’s worth reading. Then some people don’t think it’s worth reading (so get an LLM to do it) but think somehow we will think it’s not worth reading the article but is worth reading the llm summary. Feel like you can carry on that process ad infinitum always going for a smaller and smaller audience who are somehow willing to spend less and less effort (but not zero).