8 ms·
AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights
- makeitrain 5mo agoVibe resume?
- johndhi 5mo agoAnother way to phrase this might be that LLMs make better resumes no?
- jezzamon 5mo agoIn text generation, LLM language is full of very emphatic phrases. At a surface level it might sound stronger. But as a human reader, it's not necessarily better
- budoso 5mo agoIf that were the case they would select the ones generated by other models at a similar rate to the ones they generated themselves.
- delecti 5mo agoYou'd have to define "better". All this shows is that LLMs generate resumes that fit the heuristics LLMs use to judge resumes. And that makes sense, but isn't necessarily a given.
- mrktf 5mo agoOr in other words: LLM it is optimizing function which is generated by same LLM, think you have random variable y, where generator sin(x+r) and your optimizer trying to fit function sin(x+unkown1) + unknown2 ("unknown" function) - it is obvious that will find best fit.
- rectang 5mo agoBy one metric, yes! If you are a candidate who wants to be hired, and your target employers use LLMs to filter resumes, then an LLM-generated resume that the employer LLM-powered resume filters favor is "better" — as in "more likely to get you the job".
- mathgeek 5mo ago*for getting past ATS reviews.
- Emanation 5mo agoWhere I work, my boss decided to make an application that uses AI to score long text field entries to ensure required information is present. The AI lacks the ability to extract nuance and implicit information, which means entires end up being long winded and repeatitive. For each requirement its looking for, it must be explicity expressed-- it's quite unnatural, and almost feels like solving a puzzle, to which the obvious solution is to write a comment, then give it and the AI feedback to a failing comment to AI, so it can generate the proper structure the rubric-AI is looking for. LLMs are statistically driven, and I can only imagine having the AI rewrite the comment produces a result that's more statistically fitting to the model than if any given human were to write it. So, it might mean, yeah, LLMs are better at writing resumes that the LLM can successfully classify-- are they better for a human to consume? Who knows.
- nottorp 5mo agoEasy then. Apply N times, each time with a resume generated by a different LLM. No human is going to notice anyway. Or add a N+1 resume written by yourself in which you describe your strategy, just in case.
- zipy124 5mo agoDo you really believe no human is going to read your resume at some point in the process and notice the classic AI tells? Further de-duplication is rather easy, and will likely see you black-listed by competant organisations.
- cl0ckt0wer 5mo agoThe only resumes that make it past the ai to a human are ai generated
- zipy124 5mo agoRather unlikely to be the case, supported by the original article itself here, since if your statement was to be the case they would find that the human generated resume is 100% less likely to be shortlisted.
- stingraycharles 5mo agoObviously it’s not 100% of all human resumes are going to be filtered out, but it’s quite damning that human resumes are more likely to be filtered out just because they didn’t LLM-ify it.
- Esophagus4 5mo agoWhen I’m hiring, a human recruiter (or the hiring manager) reads most resumes. For us, there is some sorting by basic keyword analysis and we start near the top, but there is no proverbial black box that rejects candidates outright. If candidates are ignored by humans, it’s not because AI rejected them, it’s because we are starting with candidates earlier in the list and might not make it to applicant 537.
- randomdrake 5mo agoI wonder if this extends to training models on new content as well. Are we creating a cyclical information-consumption and training situation in which models being trained are more likely to pick up on and reference content created by themselves or by other LLMs than by other humans?
- logicalfails 5mo agoI suspect this is more a function of the corporate sanitization of language within the models. When I have passed my resume through the models for refinement, it often sanitizes some of the more easy going or simpler wording. It expands the vocabulary, makes it more dense, and uses more corpo speak in the bullets and formatting. Each model likely has its own biases in terms of what constitutes correct corporate speak, and it chooses the resumes that best fit this. Ultimately, I suspect it's more a function of model saying "this grammer, syntax structure, and formatting is most aligned with what is correct corporate language, so flag as high quality".
- charliebwrites 5mo agoAnecdata, sample size of one: When I was looking for my next role after being laid off, I didn’t get much of a response with my human handmade resume despite my experience Just for kicks, I asked ChatGPT to “Analyze my resume and give it a score for what percentage it was in” then I asked it to revise it to make it score as high as possible I still tweaked and fact checked it but after I started sending that out, I got a much higher hit rate than before But who knows, maybe the market changed, was a better time of year, etc I still had to pass interviews and prove my worth. But it probably helped me get my foot in the door
- fuzzy_biscuit 5mo agoI've done as you described and then edited it down to sound human again.
- amelius 5mo agoI suppose the HR folks gave you a "+1 knows how to use AI".
- ben_w 5mo agoSome will, others openly say on the job ad they will fail you for using AI.
- dawnerd 5mo agoI know if I got a resume from someone that had obviously used AI to generate it, it would be a pass.
- bell-cot 5mo agoWhat if your own HR's LLM didn't send you any other kind?
- drillsteps5 5mo agoBefore the resume ends up in the hiring manager's inbox it needs to be picked by the recruiter from literally hundreds of others. The recruiter uses HR software to determine the match (usually the percentage), and then picks top 5% or top 20 or whatever highest ranked resumes. Guess what's doing the ranking.
- benashford 5mo agoIntuitively this feels obvious. Content generated by the model will be shaped by its training, therefore when reading it back it will resonate with that same training and have a positive view as a result. Human when preparing a CV: "Make my CV more professional" LLM many days later presenting a report to HR: "This CV is really professional" There's probably more to it than that of course. But it justifies my personal policy of using a different LLM family for code review tasks than for code generation tasks. To avoid the "marking your own homework" problem.
- gzread 5mo agoAnd not in human-interpretable ways. An LLM was told to behave in a certain way and then output random numbers. When the numbers were pasted to another LLM instance, it also behaved that way. I wish I remembered more about that study or had a link to it - it was fascinating.
- mnicky 5mo agoWasn't it this one? Article: https://alignment.anthropic.com/2025/subliminal-learning/ https://alignment.anthropic.com/2025/subliminal-learning/ Paper: https://arxiv.org/abs/2507.14805 https://arxiv.org/abs/2507.14805
- sb057 5mo agoWell yeah, LLMs generate resumes (and other text) that they judge as superior to alternative plausible texts. Why would that judgement change just because a different instance hasn't seen it before? To anthropomorphize it, it's like having a hiring manager write a resume, get amnesia, and then have to judge it among other resumes.
- bendergarcia 5mo agoI wouldn’t put it past these tech companies to prefer ai outputs to encourage ai inputs
- Ekaros 5mo agoSeems like obvious thing. If LLM have some weights involved on what is good resume to write there is very likely correlation to what would be good resume to rate. And this is probably a even good thing, at least from model quality perspective. Model itself should rate highly whatever it produces. There should be correlation between output and review of same output.
- hyperpape 5mo agoI'll copy what I wrote on LinkedIn (note: I read roughly 25 pages, which is half the paper, and read it quickly)[0]: "If I read the paper correctly, they don’t actually show that LLMs prefer resumes they generate. Their actual method seems to be taking a human written resume, deleting the executive summary, having an LLM rewrite the executive summary based on the rest of the resume and then having another LLM rate the executive summary without the rest of the resume. That’s likely to massively overstate any real impact, if you can even rely on it capturing a real effect. I really wonder if I read that correctly, because I can’t come up with a justification for that study design." [0] I couldn't help but mildly copy-edit before pasting here. Edit: yes, the authors present a reason for their design, and an ideal version of my comment would've said that. I do not consider it much of a justification. See below: https://news.ycombinator.com/item?id=47987256#47987727 https://news.ycombinator.com/item?id=47987256#47987727.
- delusional 5mo ago[flagged]
- nearbuy 5mo agoI assume they meant they can't come up with a reasonable justification.
- delusional 5mo agoI doubt it since they, admittedly, didn't read it. The question he posed, about the paper, is answered in that very same paper. He has structured his whole reply to have the tone of uncovering the hidden caveat in the small print that invalidates the paper, when it's actually a straightforwardly stated assumption in their methodology section.
- lunchbucket 5mo agoNow that they've confirmed that was in fact what they meant, how have your views on this exchange changed?
- AlexB138 5mo agoThis may lead to some interesting gamesmanship. For instance, if I am applying to a company, and I know they use a certain applicant tracking system, and I know that ATS uses a certain model provider for its filter, I should then use that model to write the version of my resume I send to the company.
- mft_ 5mo agoGood observation. There are so many versions of the future that just become an LLM arms race.
- einpoklum 5mo ago> As artificial intelligence (AI) tools become widely adopted, large language models (LLMs) are increasingly involved ... [in] ... decision-making processes That's the problem right there.
- bendergarcia 5mo agoAbsolutely! I don’t think people are really considering the full effects of just letting ai be the middle man. I mean Sam Altman basically said this is what he wants Gwen he said intelligence is a commodity no?
- mpurbo 5mo agoAt this point, all these are becoming almost like comedy.
- bendergarcia 5mo agoWe are without our consent introducing a party in between people. The models become the arbiters of who does and does not get a job. It feels problematic.
- bendergarcia 5mo agoAnd I feel the common response of: well just use the model that’s available. Ai is and will probably always be resource constrained and profit driven, that means we will eventually see a world where poor people have worse resumes than rich people and there really won’t be any way around it because the man in the middle has the final say
- adrianN 5mo agoNot too long ago I bet resumes that were printed from a computer were preferred to resumes typed on a typewriter. What happened was that computers became commodities. It is reasonable to assume that LLMs will become commodified too.
- YurgenJurgensen 5mo agoThat would hardly be surprising. Monospaced fonts make natural language a pain to read, so what that would prove is that well-presented resumes are preferred to poorly-presented ones. This case is different, as the LLM output isn’t measurably better than the human output (unless you have a particular love of bland corpo-speak).
- celdon25 5mo agoThis is a terrible way to soften an obvious alignment failure with AI rollout.
- sneak 5mo agoWe already did that when we all created LinkedIn accounts.
- ekianjo 5mo ago
- jimnotgym 5mo agoI just guessed that and got Copilot to rewrite my profile on the internal HR system. I also got a job spec benchmarked higher by getting Copilot to write it with that exact aim given in the prompt
- fecalmatter 5mo agoi straight up lied about my work experience we are exactly the same
- rogermarley 5mo agoI think resumes will eventually (or have already) become obsolete in tech. The SNR is so low, they offer very thin filtering value. Even taking the tiny bits of the resume that are "hard signal", like GPA, certifications, prior roles, etc, it doesn't translate into their performance in the initial screening interview. This is why what I think the industry sorely needs is examination consortia. Rather than trying to guess capability from the name of the university they went to, leading tech companies creating standardized tests in various fields, and your test scores form your "resume", so that developers can just focus on improving their scores rather than wasting time on resume/application/repetitive-screening toil.
- indiv0 5mo agoEventually even a system like that can be gamed, similarly to how Leetcode-maxxing and the like sprung up in response to typical SV interview questions. Studying for the job becomes studying for the test becomes studying for the pre-test test.
- aDyslecticCrow 5mo ago> standardized tests in various fields This is itself a massively difficult problem. Standardised tests are bad indicator of topic understanding. (setting aside the massive incentive for blatant cheating) You're effectively advocating for leetcode being effective hiring tool, which many would highly criticize.
- rogermarley 5mo agoBut I think even if it were purely leetcode-like, devs would actually be quite happy with this, since at least you'd only have to do it once and then it's re-usable for every application. At the end of the day it doesn't really matter what our opinions of good screening are, but what the salary-payers are. Personally I just rely on live (& conversational) task-based coding tests.
- aDyslecticCrow 5mo ago> our opinions of good screening are I want competent and skilled coworkers. I care about our hiring process, and the hiring process of where I apply. Many modern screening processes are abysmal, and a abysmal screening process is reflected in the company and culture over time. My experience of university exams makes it very clear that studying for test and studying to understand a topic are two different goals that collide or even contradict. I dont want to hire anyone that studdied for the test instead of the topic. Placing any higher stakes on the test result encurrage the wrong behaviour and filters the wrong people. I have friend who failed physics because they spent all their time writing their own kernel for mips assembly. And plenty of classmates who aced the exam by memorising prior year question examples. who would you hire?
- jamiecurle 5mo agodisclaimer: Not a lawyer, but studying towards CIPP/E. You'd make no friends doing it, but as I understand it, for those that have GDPR as a statutory right then under "[Article 22 - Automated individual decision-making, including profiling][0]" you can request to know if your CV was screened by AI and what (and this is key) "meaningful human interaction" led to that decision. Technically this falls under a data subject access request and so a response is mandatory (but who really is going to enforce that - ICO / <insert your data protection agency here> probably isn't). Companies can't just smash a button and claim meaningful interaction, it has to be, well, meaningful and smashing a "nope" button obviously isn't meaninful. If it turns out that it was only AI that screened it you can request a human review. Do not hold your breath. Again, you'd make no friends doing it, but sooner or later a test case will emerge to generate some case law around "AI said no" because employment, or lack of because AI says no, does have significant impact on a human. [0]: https://gdpr.algolia.com/gdpr-article-22 https://gdpr.algolia.com/gdpr-article-22
- noprocrasted 5mo agoThe issue is that indeed, nobody is going to enforce that.
- jonahs197 5mo agoWill people snap over this?
- ilia-a 5mo agoSeems kinda obvious, given that most large recruiting firms/hr use algos to analyze resumes and AI written version likely do a better job at hitting keywords/structure algos/llms pick up on...
- embedding-shape 5mo agoYou'll find the same is true if you have two different LLMs first independently come up with a plan for an implementation, then ask each one of them to say which one of the two designs/plans are the best. They're much more likely to favor the plans generated from the same model, rather than from other models. I'm sure, internally, this somehow makes sense, but it's worth thinking about if you're doing the whole "ask N models for voting/rating N plans to find the best" charade.
- SeriousM 5mo agoThat's why I let the LM write it's own AGENT.md or SAFESPOT.md because it "knows" best how to write it so it can resume next time without issues. Is hits the same spot as that I would take other notes than anyone else and no one could follow them as easily than I do. Everyone leaves the "of course" parts out of the notes if it's for the own use.
- jqpabc123 5mo agoRepeat after me --- it makes no sense to try and prompt a language prediction engine to display good judgment.
- Der_Einzige 5mo agoThis is extremely obvious to anyone whose read other papers. There's tons of papers showing LLMs prefer their own outputs. It's a big enough problem that LLM-as-judge has to be a different LLM from the LLM you are testing in papers.
- ryeguy_24 5mo agoDoes anyone know of any HR departments actually using LLMs for scoring, selection, extraction, classification or any real use cases? I'm curious to hear about it and how they are using it.
- redbonsai 5mo agoThere's an AI layer built into most ATS systems as well as LinkedIn and Indeed
- oogetyboogety 5mo agoWe were told by hr NY has strict state laws against this
- bjourne 5mo agoThe only test that has worked 100% of the time for me is to read the candidate's code. Two hours is enough to precisely estimate the candidate's qualities as a software developer. I never understood why companies waste time with tests and quizzes because since it is so easy for me it should be just as easy for other software developers too. Of course, a candidate may be a jerk or unfit for other reasons, but ranking them on a software developer hot-or-not scale is not very difficult.
- noprocrasted 5mo agoJust like they'll send you an LLM'd resume, they will send you LLM'd code.
- bjourne 5mo agoConceptually no different from copy-pasting someone else's code.
- parentheses 5mo agoReading only the abstract: LLMs prefer output of their own generation over humans or even other models. This is a very good reason to avoid using model-generated data to train future models. We'd be deepening this bias by continuing to do that, essentially forcing society to reshape their output using LLMs to increase engagement. This feels like a form of enshittification that doesn't just touch one product but all of society.
- visarga 5mo agoWhen classifying resumes it is better to use the LLM as a feature extractor, think of 10-20 features you base your decision on, and extract them by LLM. The LLM only needs to do lower level task of question answering. Then you fit a classical ML model (xgboost for example) on the extracted features, based on company triage data points. This way you don't rely on the biases in the model, you can decide what criteria to use and how to judge cases without retraining the LLM. The feature extractor is generic, and the actual triage model is a toy you can retrain in seconds on new data points. It is also much more explainable, you can see how features influence decisions.
- aDyslecticCrow 5mo agoI'd rather my employers just does the classic of shredding random 80% and looking at the remainder properly.
- cyberax 5mo agoAh, the good old "we don't need unlucky losers here" strategem.
- drillsteps5 5mo agoThat's what people on both side have been doing for at least couple years already. Recruiters scan resumes for the best match with LLMs, candidates use the same LLMs (there's only like 3 of them) to tweak their resume for better match. I don't know what research you need to see why that makes sense.
- yagi0x00 5mo agoThis indicates that resumes created by the same model may have an advantage over those created by other model, so I suppose technically you may have a small advantage if an insider tells you the resume parsing tool is powered by Gemini as opposed to the other models. My broader discomfort is that we are still learning about model biases while human biases are arguably better understood, and I don't like the ethics of rejecting a person based on criteria I don't fully understand.
- drillsteps5 5mo agoI wasn't saying that this is the optimal solution (it clearly is not). I was saying that it makes perfect sense for both sides - HR has their work automated and candidates have better chance to be noticed - and therefore became a common practice in many places. The well has been already poisoned, to survive you have to get in on the action. Don't want to play this game? Make connections, set up the network, and use it to get/stay employed.
- aDyslecticCrow 5mo agoIt further makes expecting or spending the effort hand writing a proper introduction useless. Which then undermine the entire purpose of it.
- idopmstuff 5mo agoEven if we take this to be true, I'm not sure that it really matters? It's comparing two resumes with the same information and picking one of the two. That's obviously a situation that would never occur in actual hiring. This doesn't demonstrate anything at all that indicates that LLMs would incorrectly preference LLM-written resumes in the real world. It'd be interesting to do the same thing but with two resumes that are almost identical. One is slightly better (an extra year of experience or a specific note of some skill that is relevant to the role), and the other slightly worse one is written by an LLM. If the reviewing LLM picks the worse one in that case, you're potentially establishing a bias that would matter. As it stands this experiment just seems contrived and pointless.
- neom 5mo agoI did a very hack job version of this as your question got me curious, but again, not at all a rigorous test. I took my resume and had an LLM re-write the exec summary, then changed the names of the business to comparable ones and gave one a couple more years experience, then prompted: "2 Candidates that are very similar on paper, If you had to pick just based on these 2 resumes for a GM & Marketing of a series A robotics startup, who would you pick? Just reply with the name." It picked the one with the LLM exec summary and less experience, then I said: "Interesting, less experience, why did you pick him?" It said: "Honestly? Looking back at the two, I shouldn't have answered that confidently. They're nearly identical on paper — same career arc, same era, same kind of advisory portfolio, both went from community roles → hosting/cloud operator → founder → VC partner. The scale claims are even comparable (Vultr "hundreds of millions ARR" vs. DigitalOcean "$500M+ ARR," both 10 → hundreds of employees). If I'm being rigorous about why John edged it for me: The DigitalOcean tenure is more legible and verifiable — "2nd largest cloud provider," reporting to the CEO, specific scope (devrel, people ops, product ops). Joshua's Vultr description is a bit vaguer on role and outcome. For a GM hire at a Series A robotics company, where you're betting on someone to actually run the operating function, I lean toward the resume that shows clearer operational ownership rather than a marketing/CMO title. John also has Tiny Mile (sidewalk delivery robots) and Starship-adjacent exposure — wait, actually Joshua has Starship Technologies as an advisor, which is the more relevant robotics signal. That cuts the other way. Truthfully, this is a coin flip and I called it too quickly. You'd want to dig into references, what they actually did day-to-day at Vultr/DO, and how they think about hardware/capex businesses before picking." Opus 4.7 and again, pretty hack job test but I was curious.
- bdangubic 5mo agoMy new CV contains 37 emdashes
- booleandilemma 5mo agoHR departments aren't using LLMs to select candidates for jobs are they?
- interstice 5mo ago"I'm not just good, I'm amazing"
- abubakir1997 5mo agoVery interesting.
- ivansmf 5mo agoI suspect the entire industry uses "auto-raters", where an agent instance is used to scores the agent's output. The idea is similar in intent as using adversarial networks to train image generation, minus the human labelers. Raising the scores of the auto-rater then becomes the metric teams optimize, and it is no wonder the end result is that the agent scores its own generated content the highest.
- analog8374 5mo agoThis means that LLM human resource departments will only hire LLMs. Which is kind of beautiful.
- aykutseker 5mo agoThe uncomfortable part is that this is probably rational behavior for both sides. Employers use models to filter resumes, candidates optimize resumes for those models, and suddenly the resume is no longer written for a human at all.
- deleted 5mo ago[deleted]
- skeledrew 5mo agoPretty straight forward IMO. The model is looking for particular qualities in a given resume, and strives to ensure the qualities it looks for is present in resumes it creates. Humans do the exact same thing (unless forced by something like DEI, etc to do otherwise), so I see nothing noteworthy here.
- deleted 5mo ago[deleted]
- onlyrealcuzzo 5mo agoFurther, LLMs consistently think LLM written content is "good". Ask an LLM to write some design doc for you, wait until you get one that's very bad, send it to other LLMs and get their feedback, they will typically have good things to say. Compare that to a very well written document you have. They will typically have a lot more bad things to say, even if the premise is solid. Someone should study this. LLMs clearly have a lot of value. But IMO this is very interesting and points out a weakness that's not entirely clear what the full ramifications of it are. I suspect LLMs also have a major bias to code they write. Take something universally considered to be well written like Redis, feed it to an LLM for feedback. They'll probably find much to pick apart (and a lot of it may be flat out wrong). Feed the same LLM some clearly garbage LLM repository. Do they have a similar response as they do with design? Do they treat language different than code, and they're just susceptible to the way they write regular language that's different from logical code? Or do they have the same problem? Has anyone done this?
- mcv 5mo agoTimely topic for me. My CV had grown to 7 pages, and I kept reading everywhere that it should be no more than 2, so I asked Gemini to rewrite it. Took a lot of time, because Gemini loves to exaggerate everything, but I'm quite happy with the result. The first couple of recruiters I sent it to preferred my old 7 page CV. I guess they're not using enough AI yet.
- samagragune 5mo ago[dead]
- cyberax 5mo agoAs always, XKCD is prescient here: https://xkcd.com/2237/ https://xkcd.com/2237/
- danielodievich 5mo agoSo just to test, loaded qwen/qwen3-v1-30b locally, and fed my 100% human-written resume and asked it "Make this resume more professional". Mucho bullets came out. My sentence "I specialized in enterprise data modeling and worked on Cost of Goods Sold optimizations across entire customer base." became a bullet sentence "Specialized in enterprise data modeling and performance optimization, driving $5M+ in recurring cost savings across the customer base.". The $5M+ sure sounds awesome, and clearly the corpus of resumes lean towards metrics, but its not true and I didn't ask the model to make up numbers. Oh and it awarded me a "Bachelor of Science in Computer Science from University of California, Berkeley | 1996 – 1998" out of thin air. My resume has a SDE job between 1996 -1998. Oh man.
- voncheese 5mo agoOh man is right! The making stuff up is going to make this problem even bigger. There will be people that correct those hallucinations, in that scenario it’s “only” the applicants time that is wasted. There will be other people that don’t correct those hallucinations, in that scenario the best case outcome is wasted time for the applicants and interviewers (who find the mistake later). The worst case scenario is people are hired who aren’t capable of doing the job and that’s all kinds of messy and inefficient for all.
- oytis 5mo agoThat makes sense to me. "Write me a good CV" and "find me CVs in this pile" are kind of two queries for the same underlying data.
- appz3 5mo ago[flagged]