13 ms·
New arXiv policy: 1-year ban for hallucinated references
- bigfishrunning 4mo agoGood; academic literature is in crisis because of all of the slop. Forcing some consequences on easily-detectable hallucinations can only be a good thing
- tengwar2 4mo agoIt's not just AI, though. I did a doctorate in physics about 40 years back, and bad references were a problem back then.
- lucb1e 4mo agoIn what way? Surely something like the source not quite saying what was cited, or mixing up citations, rather than inventing them outright?
- miki123211 4mo agoThat, and mixing reference details from multiple sources and messing it up. Let's say you read a paper on Arxiv but cite the version that was submitted to a journal or conference, without realizing that the authors made changes to the version they submitted and forgot to upload them to Arxiv.
- ls612 4mo agoThis deserves a lifetime ban for being a bad researcher obviously. (/s)
- tengwar2 4mo agoIn physics, references which just didn't exist. That could be that the author made it up, but often it's because they transcribed the reference from another paper without reading it - we know because a few people have deliberately introduced fake references to trace how far they would go. The reasons are not the same as for AI, but the problem they produce is the same. References which don't accurately reflect the quoted material seem more common in other subjects.
- add-sub-mul-div 4mo agoYes and ffs arrows kill people too but we don't bring that up every time we talk about what to do with guns.
- dualvariable 4mo agoDoesn't matter if it is AI hallucinations or entirely human scientific fraud, the problem is the same, and the solution works fine for both cases. If you can't validate that your bibliography is full of real articles, you shouldn't get published. LLMs have just poured gasoline on the fire.
- wrs 4mo ago"Bad", like, you literally just made them up? I hope that would have been a problem.
- tengwar2 4mo agoWell, I certainly didn't make them up. But it was common to follow a reference and find that there was no paper on the other end.
- asdff 4mo agoImagine how bad they are now then.
- noobermin 4mo agoWhich is why the angry replies on Twitter from AI hype accounts is so funny. You should get penalised for fake references and profanity in your submissions, even if you wrote your slop longhand. I don't know why anyone would have an issue with this policy.
- Ozzie-D 4mo ago[flagged]
- imenani 4mo agohttps://xcancel.com/tdietterich/status/2055000956144935055 https://xcancel.com/tdietterich/status/2055000956144935055
- JumpCrisscross 4mo ago> Our Code of Conduct states that by signing your name as an author of a paper, each author takes full responsibility for all its contents, irrespective of how the contents were generated (Dieterrich, T. G.)
- incognition 4mo agocoauthors about to get roasted
- lou1306 4mo agoTo be a coauthor on a preprint that you have not submitted, you have to actively "claim" it (using a password given to the author who submitted). It's on you to double-check before claiming.
- emil-lp 4mo agoIs that your definition or theirs? I can't see that in the code of conduct.
- lou1306 4mo agoI don't think they mention it in the CoC, but anyone can upload an ArXiv preprint listing literally anybody else as a coauthor, so it is only logical that only "confirmed" coauthors should be affected.
- btown 4mo ago> The penalty is a 1-year ban from arXiv followed by the requirement that subsequent arXiv submissions must first be accepted at a reputable peer-reviewed venue. This is incredibly good for science. arXiv is free, but it's a privilege not a right! I'm not seeing this clearly listed on https://info.arxiv.org/help/policies/index.html https://info.arxiv.org/help/policies/index.html so it's possible this is planned but not live yet - or perhaps I'm not digging deeply enough? As a certain doctor once said: the whole point of the doomsday machine is lost if you keep it a secret!
- dataflow 4mo ago> This is incredibly good for science. I disagree. It's just one darn hallucinated citation for heaven's sake, not fraud or something. It doesn't account for the substance or quality of their work at all. A one-year ban seems plenty sufficient for a minor first time mistake like this. People make mistakes and a good fraction of them can learn from those mistakes. There's no need to permanently cripple someone's ability to progress their life or contribute to humanity just because an AI hallucinated a reference one time in their life. That's punitive instead of rehabilitative.
- ajkjk 4mo agoIt's not the kind of mistake that is possible unless you're engaging in fraud anyway.
- dataflow 4mo ago> It's not the kind of mistake that is possible unless you're engaging in fraud anyway. Seriously? You can't fathom an honest researcher asking for AI to find a citation they know exists, and the AI inserting or modifying a citation incorrectly without them realizing? If you find evidence of fraud by all means lay down the hammer. Using a single hallucinated citation like it's some kind of ironclad proxy just because you think they must be committing fraud is insane.
- 4mo ago
- random3 4mo agoIt seems a good idea to ban cheating, but how hard is it, especially in new reasoning/agents contexts to validate references? The deeper question is whether legitimate AI generated results are allowed or not? Test - In the extreme - think proof of Riemann Hypothesis autonomously generated (end to end) formally proven - is it allowed or not?
- pointlessone 4mo agoIt is allowed as long as it’s verified. The thread specifically points out that if authors can’t be arsed to simply proofread their text the rest can not be trusted either. It’s a simple heuristic against low quality submissions, not an anti-ai measure.
- Ifkaluva 4mo agoThis is not about banning cheating, it’s about banning inaccurate information.
- Retric 4mo agoYou don’t need to solve everything, catching a few thousand non existent citations with such a policy is on its own a net benefit.
- pinkmuffinere 4mo ago> think proof of Riemann Hypothesis autonomously generated (end to end) formally proven - is it allowed or not? Sorry to be rude, but this seems like a dumb question. I want science to progress. A primary purpose of these journals is to progress science. A full proof of the Riemann Hypothesis progresses science. I don't care how it was produced, if Hitler is coauthor, etc, I just care that it is correct. Whether the authors should be rewarded for whatever methods they used can be a separate question.
- kingstnap 4mo agoTerence Tao had a nice talk from the Future of Mathematics conference posted yesterday [0] that shapes a lot of my own feelings on this matter. The short of it is he argues how first to correctness shouldn't be the only goal / isn't a great optimisation incentive. Presentation and digestibility of correct results is a missing 1/3 when you've finished generation and verification. I completely agree with him. You don't just need an AI generated proof of the Reimann Hypothesis. You would really like it to be intentional and structured for others to understand. A really beautiful quote I learned of in the talk is this: > "We are not trying to meet some abstract production quota of definitions, theorems, and proofs. The measure of our success is whether what we do enables people to understand and think more clearly and effectively about math." - William Thurston [0] https://www.youtube.com/watch?v=Uc2zt198U_U https://www.youtube.com/watch?v=Uc2zt198U_U
- squirrelon 4mo agoHad a colleague submit a paper with literal AI slop left in the text, got hit with a nasty revision request. Check your drafts before you submit, people. The reviewers will find it.
- miki123211 4mo agoAlso check your LaTeX comments, Arxiv makes those publicly visible!!! I'm a screen reader user and usually read papers as raw TeX. I've seen everything: slurs, demeaning comments towards reviewers and professors, admissions of fraud, instructions to coauthors to commit further fraud before paper submission to mask the earlier fraud... it's all there. There's far less of it than I would think, definitely <1% of papers, but it's there. I think it would be useful to run an LLM anti-fraud pass on the TeX source of all new arxiv papers. It wouldn't catch everything, but it would catch some of the dumbest fraudsters. On the positive side, you can also find stronger claims that didn't survive review, additional explanations that didn't make the cut due to the conference's page limit, as well as experimental results that the authors felt weren't really worth including. Those need to be approached with an abundance of caution, but are genuinely useful sometimes.
- SchemaLoad 4mo agoSad the suggestion here is to just disguise the slop to make it harder for reviewers to spot rather than not submitting slop to begin with.
- MinimalAction 4mo agoThere needs be to a careful vetting before such adverse actions. If somebody includes a name and pushed it without express permission, does everyone get the ban? I agree that implemented the right way, this is good.
- vasco 4mo agoPlus afaik you can add any co-author you want without validation. So you can ban everyone on arxiv with one paper with one sentence.
- lou1306 4mo agoAs I mentioned in another thread: To be a coauthor on a preprint that you have not submitted, you have to actively "claim" it (using a password given to the author who submitted). It's on you to double-check before claiming. I surely hope that only "confirmed" coauthors will get the ban, it's only logical.
- deleted 4mo ago[deleted]
- noobermin 4mo agoSeeing the usual LLM hypers angry replying to this on twitter is such a tell. Just like the comments on the LLM poisoning articles, some people just can't accept that some people don't like LLMs and get upset when you put any amount of hindrance to their rapid acceptance.
- lbrito 4mo agoCrazy that this is graytexted. So basically HN consensus is that we need to be hyper and accelerate llm adoption everywhere. Bonkers. At the same time peak hn
- ethin 4mo agoIt's also hilarious that they complain about this because, from what I've seen, most LLM hypers will talk about something being irrelevant or taken over by AI with no understanding of what that something really is or involves.
- lou1306 4mo ago> some people don't like LLMs It's not even that they "don't like LLMs". They just don't like academic fraud! If references were fabricated with a Markov chain it would be just as bad!
- tdeck 4mo agoIt's hard for me to even understand their perspective. Researching references for a published academic paper isn't some incidental busywork task, it's supposed to be a core part of doing research which is the core of the job. If you don't have sympathy for someone who, say, paid a person on Fiverr to cook up a paper rather than writing it themselves and then didn't even bother to check the references, why is using an LLM and not checking any better?
- boccaff 4mo agoThere is a lot of "throw it against the wall, and if it sticks, write it up" empirical work against benchmarks. It leads to post-hoc rationalization of the work and browser plugins using LLMs to find references for work that is already written. It is a bureaucratic view about "you need a citation for this", where people misunderstand the citation as a checkbox, instead of "you need to substantiate this claim, as I, the reviewer, do not accept this as a fact".
- jszymborski 4mo agoShould be more harsh in my opinion.
- rgmerk 4mo agoGood. If it’s not worth your time to check the output of your LLM carefully, it’s not worth my time to read it.
- djoldman 4mo agoUnfortunately, it's probably not worth your time to read 99% of arxiv papers, LLM generated or otherwise. Ever pick a random one and really dive in?
- rgmerk 4mo agoAgreed. There was already too much human generated slop in academia. And I’m not talking about good faith research that didn’t pan out, I mean research that is completely useless for any other purpose other than convincing a casual observer that the authors are doing research.
- gammalost 4mo agoWell, yeah, 99% of arXiv papers were not written for me or you. They were written for someone who works in a niche within a niche. That's (in my view) the beauty of research.
- nullc 4mo agoIt's been pretty eye opening watching Craig Wright (of bitcoin fakery fame) flooding out LLM generated 'academic' papers and even having some of them accepted. He's toast if SSRN were to adopt a similar policy.
- hyunwoo222 4mo ago[flagged]
- mks_shuffle 4mo agoWhile this is certainly a welcome step, I hope there is more work done to fix the underlying problem of easily creating correct BibTeX entries for the cited papers. Citations for any given paper can come from a wide range of journals with various publishers, conferences, and preprints. The same paper can be available from multiple sources with varying details, e.g. arXiv and the conference website. Tools like Zotero have certainly made it significantly easier to extract citations from webpages of publication, but I still find issues with the extracted BibTeX details. While author names and titles are often extracted correctly, I still have to manually ensure that details like publication venue, year, volume number, page number, URL, etc. are extracted correctly and also shown correctly in LaTeX format. Different publications can use different citation styles. This can unfortunately lead to taking shortcuts with AI-generated citation data due to the lack of an easy and unified approach to extract consistent citation data. I am not sure whether hallucinated citations are being generated in the main manuscript or in a separate BibTeX file, so I may be a bit off in my understanding.
- flexagoon 4mo agoNote that Zotero also has a free online tool to generate citations in any format or BibTeX files from a URL/DOI/ISBN/... https://zbib.org/ https://zbib.org/
- gucci-on-fleek 4mo agoFun fact: if an article has a DOI, you can just use curl to get a BibTeX entry. An example using one of my articles: $ curl -L "https://doi.org/10.47397/tb/43-1/tb133chernoff-widows" -H 'Accept: application/x-bibtex' @article{Chernoff_2022, title={Automatically removing widows and orphans with <tt>lua-widow-control</tt>}, volume={43}, ISSN={0896-3207}, url={http://dx.doi.org/10.47397/tb/43-1/tb133chernoff-widows}, DOI={10.47397/tb/43-1/tb133chernoff-widows}, number={1}, journal={TUGboat}, publisher={TeX Users Group}, author={Chernoff, Max}, year={2022}, pages={28–39} } This is the exact same method that Zotero uses internally, so this won't ever give you better results, but I still find it kinda neat.
- jeremie_strand 4mo ago[flagged]
- jimmygrapes 4mo agoAs of yet no comments here seem to address the "reputable" condition. Reputable review is based on what criteria?
- ElenaDaibunny 4mo agohow will they detect hallucinated refs at scale? Manual spot checks? Automated DOI verification? The policy seems right but enforcement is the hard part.
- fc417fc802 4mo agoHowever difficult it might be right now it's only going to get easier. Anyway I don't think proactive enforcement is the point. Rather now they have an official method by which to address incidents that are brought to their attention.
- ddosmax556 4mo agoEnforcement is secondary and is allowed to take weeks / months / never at all if nobody reads the paper. It's about being able to ban if an issue arrises; not about keeping the database strictly clean.
- druub 4mo agoWhat are reasonable alternatives to arXiv? It has become increasinbgly slow. Techrxiv?
- az226 4mo agoNext, for AI papers, a reproducibility requirement. So much code and details are fudged and paper's cannot be reproduced. Ran the training with some other config, or other data, etc. to make their mechanism or intervention seem better.
- az226 4mo agoHurray!
- cyclecycle 4mo agoThis has become such a problem in scholarly publishing that we have a business that provides citation checking https://groundedai.company/ https://groundedai.company/ that we've been buidling for a couple of years now
- gizajob 4mo agoWhat’s the hallucination rate of your AI?
- cyclecycle 4mo agoSo far we basically just provide a very rule-based approach and try not use LLMs as much as possible. So we extract and parse the citations using various ML and rule-based approaches, and carry out a bunch of predetermined queries and do various fuzzy matching approaches on the metadata components, and have a bunch of rules around risk levels of things we should have found/matched based on what type of source it is, which venue we should have found it in, etc. So there are absolutely a bunch of tasks that could be evaled/benchmarked, but "hallucination rate" isn't particularly applicable/interesting as a metric of how good the tool is that said, we do use various LLMs (mostly local, fine-tuned, small, for things like NER/parsing/metadata comparison, etc.). and they can and do hallucinate, but we have very hard constraints on the validation, so any extraction results that don't match 1:1 back to the input text are discarded for example. so again, rather than hallucination risk we prefer hard constraints
- thatjoeoverthr 4mo agoNo mercy to brain slugs.
- Foivos 4mo agoI just wish to anyone who is against this policy to be forced to review a paper that turns out to be unedited AI slop. Reviewers are experts volunteers who do it for free. It is incredibly frustrating to have spent 4 hours reading a paper where you try your best to make sense of what the authors are trying to prove just to realize that it is hallucinations. The authors should value the time of the reviewers higher than their own time. So, if you include AI nonsense in your paper, it is insulting.
- scirob 4mo agoGreat, it's so easy to automate checking ref super bad to not check
- soraminazuki 4mo agoIt's not unexpected, but still sad to see so many comments opposing even the smallest step against low-effort fraud in academic publications. Is this what hacker culture has been reduced to in the age of the slop era? Open hostility against science and engineering?
- skydhash 4mo agoSomething something about not understanding the problem when your salary depends on it…
- clearstack 4mo agothe hallucination problem is especially bad in financial AI. one made-up revenue number and someone makes a bad investment. the actual 10-K is the only source of truth for company data
- kevincox 4mo agoThat sounds like way less of a problem than science. If you take bad sources you hurt yourself financially. For science (especially anything medicine related) you can hurt many others physically.
- clearstack 4mo agofair at the individual level. but now these tools are used in quant models and at retail scale — one systematically hallucinated metric can affect a lot of real portfolios
- qarl 4mo agoI don't understand. I have standing instructions with my agents: do not commit anything without review by at least two sub agents. The chance of two agents hallucinating at the same time in the same way is almost nil. Why do people still do it?
- vwkd 4mo agoPreviously: „Generate a bibliography.“ Now: „Generate a bibliography. Make no mistakes.“
- bigfishrunning 4mo ago"You are an expert research assistant that never makes mistakes. Please write a..."
- lbrito 4mo agoAs another comment said (not sure if sarcastically or not - impossible to tell): "just spin up another agent(s) to check the work of the first!"
- VerifiedReports 4mo agoFabricated, not "hallucinated."
- theuniverseson 4mo ago[flagged]
- kramit1288 4mo agothats good but difficult to identify if its pure AI generated or not. identifying itself can be hallicunation.