14 ms·
HackerRank open sourced its ATS. My resume scored 90/100. Oh wait 74. No – 88
- ryukoposting 3mo agoAt this point we might as well adopt that joke where you blindly throw away half the resumes because you don't want to hire unlucky people.
- pjio 3mo agoThis hurts more than it should.
- agnosticmantis 3mo agoA person's total luck is constant over a lifetime. The remaining half of the candidates already spent some of their luck in this selection, so they'll be on average less lucky than the discarded half.
- throwawaythekey 3mo ago> A person's total luck is constant over a lifetime Ah yes, the much revered cosmological fairness constraint.
- cyanydeez 3mo agoeveryone knows luck is tied to the wealth-gravity and increases as the inverse distance to the density of matter. hut because its relative, everyone thinks they have the same luck when not observing others.
- latexr 3mo agoEven assuming that was genuinely how luck works, the conclusion does not follow from the premise because it’s obvious not everyone “starts with” the same amount of luck to spend.
- addandsubtract 3mo agoBut assuming a random draw, you're more likely to select people with higher luck.
- deleted 3mo ago[deleted]
- lobocinza 3mo agoassuming luck is spendable
- t-3 3mo agoNo, luck would be some expression of the difference between the average and the individual outcomes - it only exists relative to a population at the point in time when it is measured.
- CuriouslyC 3mo agoDonald Trump disproves the fixed luck hypothesis (and the Karma hypothesis!)
- bee_rider 3mo agoBut, however you structure the selection process the people who get picked are the ones who’ve expended some luck (like, if you throw away half the resumes, but then pick the resumes out of the trashcan, the ones you plucked out are still the lucky ones). I see two possible solutions. 1) Most people won’t be using up most of their luck on this one thing. I mean they’ve got their whole lifetime worth of luck, so you just need to make sure to pick people who still have plenty left. In other words, ageism and/or picking people who’ve never accomplished much are the solutions! 2) We assume working for the company is a lucky outcome. If you make the company a really unpleasant place to work, people will have to use their luck to dodge it. However, luck can only be evaluated against other possible outcomes. The plan, then, should be to set up a competitor (possibly a front) that is a really nice place to work. They’ll act as the “lucky outcome expenditure dump.”
- sfn42 3mo agoThis is not at all how probability works. Luck is not a resource one spends. If you flip heads 500 times in a row with a fair coin, the next coin flip is still 50/50.
- aspenmayer 3mo agoPresupposing that the same coin is used for every flip (which is implicit in the example), it would be fair to question whether the coin could possibly be a fair coin after 500 heads in a row, even (and especially) if the flipping process were ideally fair. I’m not a whiz with the math involved, but I am of the opinion that 500 consecutive same-side flips is a large enough sample size to calculate that the coin in question is biased, so it would be unreasonable to assume that the next flip is 50/50. https://en.wikipedia.org/wiki/Checking_whether_a_coin_is_fair https://en.wikipedia.org/wiki/Checking_whether_a_coin_is_fai...
- sfn42 3mo agoI already said the coin is known to be fair.
- aspenmayer 3mo ago> I already said the coin is known to be fair. The coin can be assumed to be fair before the flips, but after the flips we have no mathematical reason to believe that the coin is fair, as our results heavily suggest that it quite simply isn’t fair. However, for the purposes of discussion, assuming an ideal coin, the probability of 500 same flips in a row is so statistically unlikely that fairness of the flipping process and/or flipper must then be called into question. Even if the coin, flipping process, and flipper are all ideal (which wasn’t stated, but we will also assume for the sake of argument that we have ideal immovable goalposts), the likelihood of the 500-in-row event is so improbable as to be unreasonable to use as an example or even a metaphor, because it doesn’t have much predictive power in a conversation, as ideal coins don’t exist any more than ideal coin flippers. Even a coin designed to be fair would be fairly deemed defective after so many same-side flips, and any reasonable gambler would demand that the coin be replaced with another; at that point an argument could reasonably be made to also change the coin flipper and/or the venue.
- Terr_ 3mo agoNormally we'd reject the first 37% [0] of candidates and then pick the next one that is above the average, but if all the unluckiest candidates show up first, then we need to sample even more in order to get an accurate baseline. This may be compounded by the the "Teela Brown" problem [1], where some candidates may be too lucky to end up with our company, causing them to appears later in the stream or not-at-all. [0] https://en.wikipedia.org/wiki/Secretary_problem https://en.wikipedia.org/wiki/Secretary_problem [1] https://en.wikipedia.org/wiki/Ringworld https://en.wikipedia.org/wiki/Ringworld
- zipy124 3mo agoOr more to the point. There are generally far more qualified applicants than job roles. That is training and education greatly expanded over the last couple of decades to produce more and more job seekers, whilst job creation hasn't really kept pace.
- taffronaut 3mo agoAt one point in the past a major UK a medical school adopted random selection for qualified candidates (Barts and The London School of Medicine and Dentistry - part of Queen Mary University of London). The approach benefitted qualified students from less well-off backgrounds vs those who can afford to win at the ever more elaborate (manual at the time) hurdles of resume assessment criteria and effectively game the system. There was an orchestrated campaign against the lottery around "Why gamble with would-be doctors?". Random selection was quietly dropped.
- Herring 3mo agoThat's probably a good litmus test for political capture by elites. The Netherlands introduced a weighted lottery for medical schools in 1972, abolished it in 2017 for basically the same reasons, studied the (worse) outcomes for a bit, then put it back in 2024.
- citrin_ru 3mo agoMay be LLM resume screening is a symptom of a bigger problem - with tens of candidates per vacancy employers can screen resume badly and even throw half of the resumes away and still hire someone qualified.
- AbsurdCensor 3mo agoThat's really what it is, or at least what I've noticed. Any position you have these days is inundated with applications. Most don't meet the qualifications (because in a lot of places say in the US you must apply to jobs to keep with benefits, regardless of what you are applying for), and for the remaining, you'll find that there will always be some that are all similarly qualified. Who do you hire for one position? It sometimes just comes down to luck. AI doing the job of filtering I can't imagine making the process easier, and more applications are just going to get tossed because of it.
- latortuga 3mo agoThe author made this exact joke in TFA.
- speedgoose 3mo agoMany em dashes and a "This is not, it is…" later, I think this article would have been a much better critic if it didn't use a LLM to (re)write some parts of it.
- another-dave 3mo agoI always find it funny when a technical crowd starts picking on em dashes as a sure sign of AI. I mean, are keyboard shortcuts really that difficult for developers? Some of us always knew how to use correct punctuation, even before LLMs existed. Also, neither "this is not" or "it is" appear at all in the article?
- speedgoose 3mo agoIt’s a lot of them. It’s a style. I know some people who used them before and use them less nowadays. > This non-determinism isn’t a bug you can just fine-tune away, it’s a fundamental design flaw.
- actionfromafar 3mo agoFunny how something which was catchy at one point makes my skin crawl now.
- mv_d5339e31 3mo ago[dead]
- glouwbug 3mo agoI guess at least HR doesn’t have to read 1,000 resumes. Heck, to be frank, could they make sense of the first 10 resumes?
- dc3k 3mo agoDisregarding the fact that this thing is completely broken, its grading rubric is ridiculous to begin with (as was mentioned in the article itself, but I must reiterate how completely stupid this is): > 35 points for open source contributions > 30 for personal projects I don't contribute to open source or have personal projects because I don't spend my free time doing what I do 40 hours a week to make a living. My 15 years of work experience is worth a maximum of 25%, so any company using this idiotic system would pass on me immediately. Open source and personal projects are fine, but in no sane world are they worth 65% of a resume's score.
- adrianN 3mo agoThey are selecting for people who are fine working in their free time. If you contribute to open source you are more likely to contribute to the company on weekends. If instead you have other hobbies or a family that takes up non-work hours you are more likely to drop your pen after forty hours.
- emj 3mo agoYou might have numbers on that but after working in a place with a strict no more than 40 hour policy my view is that people overwork for many reasons. Being an open source enthusiast is not one of them.
- stevesimmons 3mo agoI'm not sure that follows. I stopped making open source contributions when I switched from mature companies to startups. Now all my "non-work" time is spent on startup work. And none of that is visible via GitHub.
- matheusmoreira 3mo agoMaybe they're selecting for intrinsic motivation. People who enjoy programming to the point they do it for fun, not just because it pays. Free software work doesn't imply we work for free. We work on our projects, the stuff that we actually enjoy working on. Nobody is going to work on corporate products without adequate compensation.
- jerrythegerbil 3mo ago> I fail 65% of the time. Same exact resume, different luck. As someone who’s run hiring pipelines for technical roles in the past few years, that’s actually a fantastic number. I objectively hate saying that, but it’s true. 35% chance of elevating a technical individual to the next stage with no effort? I’ve seen as many as 100+ applicants an hour even when including a domain specific screener question. That’s 35 “screened” applicants in an hour. Were valid candidates screened out? Yes. Does you still have a candidate pool 35x larger than you need? Unfortunately, also yes. The volume of applicants is SO HIGH such that your chances of getting moved to the next stage are actually markedly worse if AI isn’t involved. If you didn’t apply immediately (using an AI bot) there’s 50+ people ahead of you, and an exhausted technical leader if they ever make it to your resume. Referral bonuses exist for a reason.
- lowbloodsugar 3mo agoExcept the bit about ranking a decades long S3 engineer lower than an intern with GitHub repo.
- dvt 3mo ago[dead]
- kyralis 3mo agoIs it? Or is it a 65% chance of a resume getting ignored before a single human sees it, reducing your pipeline's likelihood of catching qualified candidates by the same? Gates that reduce resume flow-through are only useful if their reduction is correlated with quality. Otherwise they're just dragging out your hiring process or unnecessarily causing you to ultimately lower your hiring bars.
- bagels 3mo agoThe goal for the interviewer is to have a much higher ratio of good/bad candidates after the first screening. This means the more costly time you spend on the second step has a better return.
- jerrythegerbil 3mo ago
- dvt 3mo agoAn alarming number of people don't understand that LLMs work via purely stochastic processes, so I'm happy to see in-depth pieces like this. I'm looking for a job and maybe this is why it's so hard to get a callback these days: resumes are just dumped in some LLM black hole and no one really knows how it works. The author says: > temperature 0.1 — low, supposedly nudging the model toward deterministic outputs This is not correct (and is briefly touched on later in the piece when he sets temperature to 0), temperature is not some kind of "deterministic" switch, but rather it affects the sampling distribution (which becomes more "spiky"—but is still very much a distribution).
- Nimitz14 3mo agoHe said it nudges it to be more deterministic. Your comment is not correct.
- bluechair 3mo agoWilling to be corrected but I believe this type of automated resume filtering is illegal. Not saying it never happens but my understanding is it is not typical.
- small_scombrus 3mo agoThey don't need to actually filter/blackhole to have have the same virtual effect. Show someone a list of resumes with an "applicant score*" and they'll naturally ignore the ones with a low ranking *scores are generated with AI, mistakes may be made, use only as a guide and verify results
- thayne 3mo agoI would expect that to depend on jurisdiction. I don't know for sure, but I would be surprised if it was illegal in my particular US state. You might be able to argue the AI has inherent biases that introduce illegal discrimination in the hiring process, but my understanding is winning I case like that would be very difficult, especially since most employers are very cagey about their hiring process and why they mades a decision.
- 3mo ago
- cyberax 3mo agoAh... The AI learned the old HR trick: take 50% of resumes and throw them out without looking. Rationale: "we don't need unlucky losers".
- worldthruword 3mo agoThere are plenty of resumes in the sea. Assuming thorough mixing up and statistically speaking, throwing 50% of resumes is a good enough heuristics.
- steve_j_choi 3mo agoThis could be used as a good way to self-evaluate one's current position from the company's point of view. you would tweak prompts and guidelines that are expected from the company and see how you score
- hahahaa 3mo agoI sort of hope we land on 2 agents, one working for the candidate and one for the employee do a screen round. Salary compatiability could be negotiated by a 3rd party bot that knows both parties ranges and what would be needed each end of range, and figure out yes/no worth going ahead. Such a time saver.
- chonghaoju 3mo ago[dead]
- rkuska 3mo agoThis reminds me of my former CTO. He would take bunch of CVs and randomly throw some of them in a bin. He didn’t want to work with “unlucky” people.
- hahahaa 3mo agoThe problem is with this system he only worked with unlucky people.
- psalaun 3mo agoI thought this was only an old urban legend; some people actually use this technique? Especially in a trade supposed to be led by people trained in sciences?
- aquariusDue 3mo agoIt's OK! We can disguise it as the Secretary Problem and it'll be fine, we could even write a post on the company blog about it. /s https://en.wikipedia.org/wiki/Secretary_problem https://en.wikipedia.org/wiki/Secretary_problem
- gregates 3mo agoGiven how often it's been mentioned here, it's likely that this is an urban legend that people are pretending to have first-hand knowledge of for karma. In a trade that's supposed to be led by people trained in sciences, no less! (A more charitable interpretation would be that aforementioned CTO was making a joke that didn't land.)
- cyanydeez 3mo agoor its so old, people would make the joke and interns would repeat it unwittingly. no one has to consciously be lying for this type of meme to continue spreading.
- subscribed 3mo agoThat'd be pretty gross for a CTO if it were real.
- quink 3mo ago"A computer can never be held accountable, therefore a computer must never make a management decision."
- 12_throw_away 3mo agoCorollary: If a computer makes a business decision, the person who delegated the decision to the computer must be held accountable. Consequence: All business decisions will eventually be delegated to computers via sufficiently convoluted and untraceable processes such that no manager can ever be held accountable.
- neya 3mo agoI wonder how is this even legal? The only useful job the HR departments are ever required to do - they decide to automate it? Aside from being a daycare for adults, what exactly does HR accomplish? It's clearly NOT on the side of employees, but this seems like they're clearly NOT on the side of employers, either. While resume's are being filtered left and right, they just make TikTok's on company's dime [1]. What a sad state of affairs. [1] https://www.youtube.com/shorts/wSug80Vg5JU https://www.youtube.com/shorts/wSug80Vg5JU
- srdjanr 3mo agoThey could be using this just to throw out the obviously bad CVs, and then manually go over the rest. I'm not sure if they do this in practice, but the tech itself can be useful. Also if HR was really useless (or actively hurting the company) they wouldn't still have a job (or they'll lose it eventually). No one likes burning money for no reason. So obviously they are doing something useful.
- syockit 3mo agoThe last time I heard HR being completely let go was with a fintech company Bolt. Then again, that company was midsized, around 200-500 people or so. For larger companies, it's going to be difficult to even realize that HR is redundant in the first place.
- deleted 3mo ago[deleted]
- makeavish 3mo agoHiring and job search has been so hard and AI has amplified the existing problems instead of solving any.
- sevenzero 3mo agoWdym, cant you just litter your applications with buzzwords and other bs to automatically get a high score in these systems?
- szszrk 3mo agoHR market is basically an early google rigging era, where you can place hundreds of keywords at the footer (white text on white background) to start popping up on random searches.
- makeavish 3mo agoI have been at both side of the market. And it sucks so bad at both ends. Companies which deeply care about next hire are struggling to hire and actual great people looking out are outcompeted by AI slop and AI bulk applying. It is actually a very hard to solve problem.
- CuriouslyC 3mo agoThe mind blower is that this spam and slop is just lowering the job market to the quality of every other capitalist market. Poor hiring manager has to look through 1000 applicants, 950 of which are spam? How many ads are shoved down your throat every day, and how many products are you actually looking for info on? Chickens coming home to roost.
- makeavish 3mo agoTrue :(
- gs17 3mo agoI'm a little confused, is this an ATS system that anyone actually uses? If not, I'm not sure how it's better than just asking ChatGPT to score your resume out of 100. Why would you want to optimize your resume for a system no one is using to score it?
- petesergeant 3mo ago(Almost) everyone’s using some kind of ATS, every ATS is adding AI auto-ranking (and has been trying to for 15 years), and almost all HR people feel like they have too many obviously bad CVs to read. Whether or not someone is using this ATS specifically, if you submit several CVs to several places, your CV is going into at least one magical 8-ball.
- 40four 3mo ago“I'm a little confused, is this an ATS system that anyone actually uses?” You read my mind. If the answer is “no”, then we can ignore this.
- another-dave 3mo agoFor one, if you go on to Hacker Rank's "Screen" page, they mention the product is used by Stripe/AirBnB/LinkedIn/Atlassian/IBM etc etc. I imagine that there's plenty more companies using it too. But I'd also assume that their competitors are doing something similar so I don't think we as an industry can just ignore that it's happening.
- 40four 3mo agoInteresting, thanks. I admittedly spent zero time looking into it :) I’m surprised open source contributions count for so much. first I thought was “is that something people actually list in as resume?”. But it looks like it pulls your GitHub account and appends that information. That kind of unfortunate for anyone who doesn’t use GitHub
- gs17 3mo ago> HackerRank Screen compresses the top of the hiring funnel by replacing manual resume reviews and unstructured phone screens with structured, auto-scored assessments That seems to be a different type of product.
- yieldcrv 3mo agothis will get patched, as in I'll optimize my resume for this and so will many other people that any edge disintegrates
- Aurornis 3mo ago> The default model is gemma3:4b That’s a tiny model. No LLM is going to be a perfect and repeatable judge, but a tiny 4B model is like plugging an RNG into this system. This whole exercise feels like someone vibe coded an ATS and got it to the point where the tests were passing because they decided they should have an open source ATS project.
- danpalmer 3mo agoThis sort of model is fine for small problems, when used in the right way. I think there's probably a version of Resume analysis that would work well with this model, but "hey clanker, what projects has this person done" is not the way. You need extraction, cleanup, probably OCR to compare and further clean up, multiple analysis passes per signal with LLMs, judges, etc. None of that needs to be large models, you'll get marginally better performance, but there's very little context, these models will perform well when used correctly.
- deleted 3mo ago[deleted]
- mlpicker 3mo ago[flagged]
- brikym 3mo agoSo that's where the Windows XP file copy dialog author now works.
- tasuki 3mo ago> Sometimes my projects “lack architectural complexity” Well done you! It is difficult to avoid architectural complexity, but imho well worth it.
- 0xpgm 3mo agoWith such kind of ATS systems, is it still a thing to optimize for a one page resume that is easy for a human reviewer to scan, or just include enough buzzwords and external links to try and please the LLM?
- jorisw 3mo agoI wouldn't assume based on this one thread/article that this is what you need to optimize your resume for. Nor that a majority or even significant group of reviewers is even using LLMs. I've been involved in hiring pipelines and never even thought of using LLMs to review incoming candidates. However given the time constraints reviewers have, yes, the former (making a resume easy to consume quickly) is a huge help.
- ChicagoDave 3mo agoI was inspired by this. I made a Claude skill to take my resume and compare it to any job description to point out viability and gaps. Pretty cool skill. I'll post it somewhere.
- davidpapermill 3mo agoA better way to reformulate this problem is for the LLM to be tasked with making a _comparative_ judgement between two CVs. This should prove much more reliable, especially if you give it a third “too close to call” option. You can also ask for clear justifications of preference.
- srdjanr 3mo agoThat's a good idea. The only drawback I see is that you should compare every pair of CVs for best results, and that grows quadraticly with number of CVs. Of course you can settle for fewer comparisons and not perfect results. But then I'm not sure if you can hit a good ratio of quality and token spend.
- skribb 3mo agoCould probably do an elo system and sample pairs. E.g. 1. Set the elo of all CVs to 1000 elo 2. Randomly pair up CVs and compare. Winners gain elo, losers lose elo. 3. Repeat #2 for a few iterations, then remove bottom X% of CVs. 4. Repeat 2-3 until the amount of remaining CVs is small enough to do an exhaustive comparison. I don't have a mathematical proof, but I suspect that this is a decent cost-effective approximation of comparing every pair (depending on the parameters)
- swiftcoder 3mo ago> you should compare every pair of CVs for best results Or compare each one to a reference set? Take 5 resumes of existing employees, rank all candidates against that set, maybe you get some useful level prediction into the bargain
- davidpapermill 3mo agoI'd just do a quick filter, probably deterministic, then perform a deeper comparison on the selected few.
- hari_vardhan 3mo ago[dead]
- cemoktra 3mo agoSo sending my CV to every company three times should get me pass the ATS?
- cyanydeez 3mo agoif i ever go back into the job market, will need three accounts: Peter J Smith, Peter Smith and PJ Smith. they live in #101, #102 and 103# 5607 Jane Street
- left-struck 3mo agoWhy stop there? vary everything that can reasonably be varied slightly across each resume
- deleted 3mo ago[deleted]
- pu_pe 3mo agoHe tried with a tiny model (gemma3:4b), got a range from 66 to 99. Then tried again with a small model (gemini 3.1 flash lite), the range was 48 to 64. Would a frontier model be more consistent? Perhaps this tool was optimized for more capable models?
- srdjanr 3mo agoIt makes sense to me intuitively (though I'm not sure if my reasoning is actually correct). Worse model may not "know" enough to distinguish between a 70 and a 100 candidate, so it's expected that it's output has high variance. But a better model might "know" enough, so it can be more confident and thus more consistent.
- maxignol 3mo agoAre many people using HackerRank ATS ?
- realty_geek 3mo agoWhy doesn't something like this exist for real estate? A popular open source AVM (automated valuation model) that helps home sellers get an idea of what their home will sell for. Right now it seems AVMs are mainly seen as just a way to capture leads. Every estate agent will tell you they have some magic recipe that makes their valuation better than anyone else's. I have had a bunch of ideas on how to approach this, but I really could do with a collaborator or two.
- gebruikersnaam 3mo agoThe article raises a lot of questions the article already answered.
- realty_geek 3mo agokeine ahnung
- diimdeep 3mo agoThey forgot to add "masterpiece" /s https://www.youtube.com/watch?v=mcYl70vq_Ns https://www.youtube.com/watch?v=mcYl70vq_Ns https://github.com/interviewstreet/hiring-agent/blob/main/prompts/templates/resume_evaluation_criteria.jinja https://github.com/interviewstreet/hiring-agent/blob/main/pr...
- mihaaly 3mo agoSo many people are willing to participate in this kind of robotic practices in human employment makes me think that many are starting to consider that this is as unavoidable as global warming and rather play along, adapts their career (life) to it, sculpture it towards a specific look, doing things that will give them point on some arbitrary test run. Which I feel being dangerous, leading to superficial minded workforce, not those good in something, including judgement of a problem and solution. But good at manipulation. Speculative thought only, of course.
- deleted 3mo ago[deleted]
- jdw64 3mo agoIt seems like the design is flawed, probably because the scoring structure and conditions are wrong. And originally, due to the nature of LLMs, even if the input is unstructured, when you design something like a RAG system, you usually need to create a verifiable evidence table. Even with that, the scores are still probabilistic by nature, but at least they stay within an error distribution that I can verify. But it doesn't seem like there's any such evaluation criteria here. Typically, retrieval should be tied to evaluation metrics, evidence should be linked to scores, and you also need to account for parsing errors. But personally, I'm weak to these kinds of ATS systems (ugly appearance, non-native English speaker, didn't go to a good university), so if this kind of filtering existed, I probably would have never had a job in my entire life. Come to think of it, even now I don't have a proper job—I just bid on projects at the lowest price and implement them. So maybe it doesn't really matter whether such a system exists or not
- Traubenfuchs 3mo agoThis actually makes a lot of sense, it's testing the luck of the candidate through the rng feeding the LLM. You wouldn't want to hire unlucky employees after all! Hiring managers of the past would solve this by throwing every second resume in the trash, now this is a built in feature of ATS.
- kailpa1 3mo agoFrom `resume_evaluation_system_message.jinja` > *SCORES MUST NEVER DEPEND ON THE FOLLOWING FACTORS:* > - College, university, or educational institution name > - CGPA, GPA, or academic grades I don't understand why they would omit these factors from the evaluation.
- sph 3mo agoHopefully so that people like me, that dropped out of high school yet have had a successful career as a self-taught engineer, have a chance. [1] Just kidding, my resumes are sent to /dev/null like everybody else’s. —— 1: In fact, I will be controversial and say that self-taught engineers tend to be the strongest in their own particular niche, because they are powered by sheer desire to learn and improve. I am routinely appalled by how many people go on forums to ask how to learn a new thing, completely unable to self-direct their learning. I blame the modern school system.
- kailpa1 3mo agoI'm a self-taught programmer as well, who dropped out of university, and these factors being omitted would benefit me as well, but I feel like good grades and a good university are still indicators of someone being or is capable of becoming a good programmer. This system would drop a Harvard top graduate for someone having a year of experience in some outsourcing firm.
- sph 3mo agoI started in an outsourcing firm (body rental actually) but I definitely get your point. Maybe they optimize for real world experience, or rather, how one is used to workplace politics and logistics. The top grad will have higher expectations, and all they want is a cog for the Machine.
- kailpa1 3mo agoYep, I don't know either, but I guess they have their reasons for this.
- rvz 3mo agoI see. > LLM is called six times to extract structured information Followed by > The default model is gemma3:4b, running at temperature 0.1 — low, supposedly nudging the model toward deterministic outputs. This is exactly why hiring is even more broken: Because the people looking for candidates are also just as unqualified if not, more. Using much weaker LLMs to replace the person in charge of the final judgement call is the wrong solution as this is a plain old social problem. Even if you wanted to use LLMs for this case, the default configuration, model choice is laughably flawed. This LLM can’t be trusted as it doesn’t even know what it is reading. The correct solution is either advanced OCR with keyword ranking with a basic filter or a far stronger LLM that excels at document / vision parsing benchmarks with an experienced person making the final judgement call in case the technology misses a critical detail. Rather than using this less accurate one that hallucinates out its decision depending on a dice roll.
- chrisjj 3mo ago> an experienced person making the final judgement call in case the technology misses a critical detail. That would fail to meet the objective of reducing the costs of hiring an experienced person - the entire point of outsourcing to a chatbot.
- saidnooneever 3mo agoCount to three, no more, no less. Four shalt thou not count, neither count thou two—excepting that thou then proceed to three. Five is right out.
- bhanu786 3mo agoATS resume usually check the keywords, and formatting your spacing and give score accordingly. As If someone is following some reference of the format. It can depend might he will be getting low scores.
- carb 3mo agoIt's a good analysis but the AI slop writing makes me not trust you've reviewed this and I'm unable to finish or subscribe. I'm sure you're a great blogger but this is holding it back!
- padolsey 3mo agoThis is just the 'LLM judge', very badly implemented without any scientific prudence. What a joke. To be terse: you cannot rely on LLMs to provide standardized scores against arbitrary criteria. To get close to 'reliable' you would need highly tested rubrics, grounded in human decision-making, and you'd need to avoid all the measurement biases these things are riddled with... positional/order effects, anchoring on whatever numbers you stuffed into your own prompt, scale-format sensitivity (a 1–5 and an A–E scale give different answers for the same input), holistic-vs-isolated context effects, and lovely examples like where adding a "be unbiased" instruction makes it more biased. I've studied this at length. You cannot even _begin_ to approach this problem seriously without held-out validation, inter-rater agreement, and ground truth. This repo is just quagmire of wishful vibes with random numbers littered throughout.
- tesnorindian 3mo ago[dead]
- zuzululu 3mo agothis is why i dont feel sorry for working 3 remote jobs
- YossarianFrPrez 3mo agoLooking at the linked scoring prompt (resume_evaluation_criteria.jinja) [0], I immediately see several red flags that suggest the output won't be reliable. (I'm developing an LLM intensive application where the stakes are high enough that I need the LLM output to be reasonably correct.) [0] https://github.com/interviewstreet/hiring-agent/blob/main/pr https://github.com/interviewstreet/hiring-agent/blob/main/pr... In no particular order: 1. The prompt is trying to get the system to do all of the evaluation steps at once. Instead, the system should break down the task of resume evaluation into its subcomponents and have separate prompts for each component. Like "evaluating open source contributions" should be its own task. Same with "assessing the complexity of software projects on the resume." Fwiw, each of the tasks contained within the prompt is woefully underspecified. 2. The prompt leaves spreads of ~10 points up to the LLM, when it's doubtful that humans are that well calibrated. Take for example: > SCORING CRITERIA Open Source (0-35 points) HIGH SCORES (25-35 points): - Contributions to popular open source projects (1000+ stars) - Significant contributions to well-known projects - Google Summer of Code (GSoC) participation - Substantial community involvement Are all of these 35-point examples? Is one a 26-point example? If not, what's the difference? If an expert can't reliably make the judgement, the LLM is going to struggle too. One partial fix is to get rid of the ranges and just say all of these are worth 30 points. An additive point scheme would be better... 3. The authors of this prompt have left an incredible number of judgement calls up to the LLM, when that's the very thing you want to minimize. Using the same example as above... - Are all contributions to open source projects with 1000+ stars equal? - What counts as a "significant contribution"? Doesn't that imply that the LLM has to know or read through all of the commits in like the last ~6 months at minimum for the project to understand what the given contribution meant to the project? That itself isn't impossible with tool usage, but again, that'd be a separate task. - What on earth counts as "Substantial community involvement"? Why didn't the prompt authors define this, or at least give a few examples? Honestly at this point maybe someone should build a tool that scans prompts for adjectives... 4. This sort of thing is just asking for trouble: > SCORES MUST NEVER DEPEND ON: Candidate's name, gender, or personal demographic information Just remove this stuff before you send the rest of the resume to the LLM. Even if you ask it not to, it's not a person, it's a very fancy statistical distribution generator. All of the input (including the name) will affect the distribution that gets generated. (This one is not unlike Andreessen's "don't be a sycophant" prompt.) 5. Obviously this one depends on the LLM in question, but instead of writing things like: > DO NOT RETURN A RESUME SUMMARY. RETURN ONLY THE SCORING EVALUATION IN THE SPECIFIED JSON FORMAT. Analyze the following resume and provide a JSON response with this EXACT structure (all fields are required):... The system should utilize the "structured output" option, which guarantees a fixed output format. Also, fwiw, the JSON should force the LLM to pick between categorical options as much as possible. Forced-choice structured output should, at least in theory, cut down on hallucinatory responses and constrain judgement calls. 6. One major thing that's not in the prompt is anything about traceability. This system should be designed so that humans can review the logs and make sure this is working as intended. 7. Another thing that is missing in the file is what I'll call evidence of a theory of coding / coder quality. Most of the examples are designed to have the LLM assess proxies for code quality, not code quality itself. Surely both should be taken into account? I'm not an expert at evaluating coders. But two pretty basic LLM-answerable thing I would ask is: How well do a candidate's 5 most recent commit messages match the contents of those commits? Do the claimed technical skills on the resume match their GitHub code? (i.e., if they say they know R, is there any evidence of that on their GitHub?) 8. The prompt also seems unaware of what it's asking the LLM to do: > LIVE DEMO BONUS: Projects with working live demos should receive 10-20% higher scores This implies that the LLM can use tools, but even then, I'd be pretty wary of its ability to fully execute this part of the prompt without more detailed instructions, examples, and guidance. There are very likely tons of edge cases here.
- bryanrasmussen 3mo ago>If your company’s cutoff sits at 85, I fail 65% of the time. Same exact resume, different luck. Your resume's reception is always affected by random factors, only now you are able to test, debug and technically critique the randomness.
- xorcist 3mo agoI think the question is why bother with an LLM if randomness is decisive? Just roll the dice. I mean, it's not the worse you can do to narrow a subset.
- bryanrasmussen 3mo agoright, and it's cheaper, but people want the illusion of determinism. Some people say that they want determinism, but if they do nothing to assure themselves it is deterministic I think it is fair they really only want a good enough illusion.
- psychoslave 3mo ago>You might as well throw out half the resumes and tell the the applicants you don’t fuck with bad luck. Hmm, well, maybe a bit with a nuance of elite class structure reproduction (that doesn’t prevent a few transclass to showcase in case anyone critic the perfect meritocracy at run), that’s basically what people get, so crude truth but truth nonetheless. Oh don’t take it personally. Your own bespoke hand-tailored process of course is different, it does give the opportunity to everyone to reach the most accomplished version of themselves beyond what they ever dare to dream. It won’t help though with the systematic failure of aiming to provide an accessible path to flourish for everyone and letting no one behind. Again, this is no fault of any specific player, but as long as a majority feel compelled to move within the frame of the game with few winners that merit all they got in contrast to large stock of inept losers, the outcomes are no wonder.
- cs02rm0 3mo agoI feel like hiring is all a bit broken. Roles get flooded with applications, it's chance whether your CV gets through, then there's hiring rounds that seem designed to make you quit the process before they have to filter you out. Is it working for anyone, on any level?
- luckylion 3mo agoI'm on the other side, and my main tip (at least if there's people like me!) is: avoid the usual AI signs. For one role we got ~70 applications and all CVs looked obviously AI-written. I don't know whether the people did actually do any of the things mentioned and I don't have the time to find out, so the AI-written CVs are a discard-signal for me. (Either those people delegated a very important task to AI and didn't even bother to check, or they are bad using AI and don't know -- I want neither) Any CVs that signal they were actually written by a person I will actually look at.
- quectophoton 3mo ago> For one role we got ~70 applications and all CVs looked obviously AI-written. Were those ~70 applications all of them, or were those ~70 applications the result of an AI filtering from a larger amount? If the latter, are you sure your AI is not filtering out the hand-written CVs and giving you the ones that have been AI-assisted or AI-written (with or without "the usual AI signs")?
- CM30 3mo agoI think what's more worrying to me (if other systems work like this ATS) is that it seems to judge based on a bunch of factors that will probably disqualify a ton of decent to good participants. For example, 65 points are given for a mix of personal projects and open source contributions. Which is great if your one and only interest is in tech, and you don't have a family, dependents or a second/third job. If you have any of those other things, well the odds seem like they're incredibly stacked against you. And it makes me wonder how many of these systems are stacked in favour of wealthy people with a near special interest level of obsession with tech and no worries outside of going to college/working a single job in their industry of choice.
- bob001 3mo ago[flagged]
- danmaz74 3mo agoOf course life isn't fair. But here the result is that companies will ignore potentially great candidates which dedicate all their programming time to their job and instead consider candidates which may be not just worse programmers, but also are more interested in their hobbies (or padding their CV) that doing their job. I'm saying this as somebody who most of the time has some side project going on.
- Schiendelman 3mo agoIn hiring, we pass laws to prevent abuses. In many countries and soon a few states, being asked to work outside of work hours is considered an abuse. Expecting that someone does work related activity outside of work hours is something I would actually consider regulating out of the application process!
- swingboy 3mo agoI’ve always assumed any LLM output that was some type of rating or score was bullshit. Unless the LLM writes a Python script to calculate the score (and even then…) then the score it outputs is just the next most likely token, taking into account temperature and what not. You see a lot of frameworks for things like spec-driven development make use of scoring how good the spec/design/plan is and it’s like, uhhh…
- joelthelion 3mo ago> is just the next most likely token, taking into account temperature and what not. This doesn't mean anything. All LLM output is like that. That said, I agree that LLMs are terrible at grading stuff, except perhaps if you give them a very detailed evaluation grid.
- nnevatie 3mo ago> An LLM is called Hooray for incidental non-determinism.
- seanieb 3mo agoIt's always amazed me that a tech company will pay $300,000+ for a good engineer, because talent is so hard hard to find... meanwhile their recruiter operates unsupported, has a very different idea about what good looks like. Their ATS black-holes >50% the resumes because it's filtering heuristics are garbage because recruiting selected the ATS system because it has a google Gmail integration or something, and the ATS's filtering technology was not reviewed by anyone in the engineering or data teams.
- sleepynoodle 3mo agoI really dont understand this constant changing of numbers. I have tried a bunch of ATS reviewers and everytime on the same resume i get different numbers. Its weird and unreliable. I understand the need for doing this to filter through thousands of CVs but maybe there is a better way. Like a take home test at the beginning or a test of somekind.
- chrisandchris 3mo agoI would say people that hink the LLM is doing a better job than they are in for a treat. I did expect the resulta to be of the same quality as if a human does the job - it averages out and has a big error margin.
- sleepynoodle 3mo agono wonder i dont get calls. I dont have a separate CV for every application. Good luck to me then!
- bartread 3mo agoThe takeaway from this for me is that, using an LLM to score anything takes multiple (maybe even many) runs and the result you’ll get is, at best, a sane-ish distribution. Which sort of sounds workable until you scale it up to larger datasets, where at some point compute/time/energy costs will render it non-viable. I am sure there’s some reasonable rule of thumb estimation on distribution that could be applied based off fewer runs per data artifact, but you’re always going to be trading off against confidence by doing this. Beyond this, I’d bet that almost no implemented systems that use LLMs for scoring, ranking, or decision making use such a multi-run approach. Partly because people don’t understand their behaviour is stochastic, perhaps because a lot of people without a background in statistics don’t understand what stochastic actually means, and no doubt partly because of budget concerns: if you have to ask an LLM to do the same thing 10, 50, 100 times to get a sufficiently good result, then the cost saving argument is either weakened or completely destroyed. There is at least one more aspect worth considering in the specific case of resumes/CVs: is the inconsistency of scoring by LLM worse than the inconsistency of scoring by a human following a similar process? Because the reality is that, even for an experienced recruiter, reviewing hundreds or thousands of resumes or CVs gets pretty fatiguing. People get hungry, bored, tired, restless, irritable, etc. That inevitably leads to inconsistencies creeping in, so there’s always an element of “luck” (or, perhaps better, uncertainty) as to whether your resume/CV passes screening. So is that inconsistency better or worse with LLM screening? I don’t know. But, at least, if it’s not worse maybe it doesn’t matter for this specific use case. And if it’s notably better then maybe it’s raised the bar on what “good enough” screening looks like? (And I’m sure other use cases warrant similar, “does it matter?”, questions, with the answers no doubt landing differently.)
- CuriouslyC 3mo agoMy experience with benchmarks and evals is that it can take ~20 runs of a problem for the distribution of answers to start to converge. Ideally you'd know the convergence properties of your algorithm ahead of time and make a Bayesian solution that makes the uncertainty explicit.
- dev_l1x_be 3mo agoDid anyone try to prompt hack this setup?
- nullc 3mo agoThe true test of HackerRank is can you setup a system that combines a document editing / paraphrasing LLM with gradient descent on the HackerRank LLM to turn your arbitrary resume into a reliable 120 out of 100. One of the weird properties of other people using LLMs is the potential of having oracle access to your opponent. Even if you don't have their exact LLM a good guess at it may be a better model of the opponent than you ever had before.
- thrance 3mo agoI cope by telling myself that I probably wouldn't want to work for a company that used an LLM to filter my resume out.
- robertlagrant 3mo agoI tried this with my CV, and it somehow scored me bonus points for GSoC! BONUS POINTS: 5.0 ------------------------------ Google Summer of Code (GSoC) participation: +5 Even though I've never done this, and don't claim to have done it in my CV.
- fernandopj 3mo agoHappened to me as well. It is a known hallucination https://github.com/interviewstreet/hiring-agent/issues/240 https://github.com/interviewstreet/hiring-agent/issues/240
- robertlagrant 3mo agoThanks - interesting. Very odd, though.
- graemep 3mo agoIt took me a a minute to figure out what an ATS was. Not familiar with this particular means of a much used TLA. Even better Wikipedia lists the abbreviation I am familiar with but give a different interpretation of the same words: https://en.wikipedia.org/wiki/Ats https://en.wikipedia.org/wiki/Ats
- Leptonmaniac 3mo agoThanks for not explaining what TLA is, either.
- graemep 3mo agoMy sense of humour. TLA = Three Letter Abbreviation.
- suzukivenom 3mo agonever understood how can people think an .md file can actually evaluate a human being.
- JrProgrammer 3mo agoThey are not evaluating human beings though. They evaluate a textual representation of a human being’s work experience. Not that I agree with this AI approach but when hiring, the real test begins after this initial hurdle
- orbital-decay 3mo agoThis word (determinism) has a magical effect of warping any online posts it touches. Once you hear it you can almost guarantee it's going to be misguided. At least this time it's actual determinism (same input = same output), not arbitrary unrelated things. Determinism matters for reproducibility, but do you really want these outputs to be reproducible in this particular case? Making LLM outputs deterministic is relatively trivial, you have to use batch-invariant kernels (if you use batching) and either set the temperature to 0 (don't do that, randomized sampling is here for a reason) or fix the seed (better). It's readily available in a few systems. But this won't make the result more useful, it will just obscure the fact that the agent is genuinely not sure about it - look at the range of the scores it gives! It still won't predict anything but the score will stay the same each time. Do you really want that? What happens here is they're supplying too little information (just a resume, which is almost at the noise level) and expecting a reply with too broad implications. This is a basic design mistake regardless of whether it uses LLMs. All surveys, tests, laws, and voting systems are extremely sensitive to framing because they work off too little information. But they also don't exist in vacuum, unlike this thing.
- RugnirViking 3mo agoThis. Human judges and examiners are famously not deterministic even though we would wish it were so - we've probably all heard the thing of harsher sentences being given in the hour before lunch.
- groundzeros2015 3mo ago> harsher sentences being given in the hour before lunch. Implicit bias theory sparked a massive number of studies that suggested everything influenced you from the color of the room, to what the person said to you before entering. It’s been really hard to replicate and the conclusions that have been drawn are contradictory.
- nonethewiser 3mo ago>we've probably all heard the thing of harsher sentences being given in the hour before lunch That suggests determinism though. I mean I agree with you overall. Either humans decision making is a system so complex it appears non-deterministic, or it is deterministic. Practically speaking, we are non-deterministic. Let's not conflate non-deterministic with inaccurate though. Non-deterministic systems can be 100% accurate. https://en.wikipedia.org/wiki/Las_Vegas_algorithm https://en.wikipedia.org/wiki/Las_Vegas_algorithm
- nicodjimenez 3mo agoI actually just built an ATS for my company Mathpix. But it never occurred to me to use resumes. Basically we have a set of company values and a specific open ended questionnaire to gauge the fit: https://mathpix.com/careers/apply https://mathpix.com/careers/apply Then internally we have dashboards and sorting based on AI agent scoring. I noticed the scoring is imperfect but still saves a lot of time. Candidates scored at or below 2/5 are reliably bad and candidates above 4/5 are consistently impressive and leave thoughtful answers. The biggest thing is not using resumes. You can’t reliably gage applicants without a writing sample and resumes are the worst form of writing sample. Also you need to be intentional about who you’re hiring for, both to craft the questions as well as grade the responses.
- mdorazio 3mo agoThis seems likely to be worse. How do you screen out people who point an LLM at your values and ask it to answer your questions in a way likely to appeal to a recruiter using an LLM to score the responses?
- conartist6 3mo ago[flagged]
- seedless-sensat 3mo agoWhat is an ATS? This blog doesn't define it
- gejose 3mo agoATS = Applicant Tracking System. It's software to help you manage your hiring pipeline as a whole.
- secrooq 3mo ago[flagged]
- vanessa1211 3mo agoI have a love hate relationship with ATS
- wielebny 3mo agoThis seems like extremely illegal in Europe.
- sp2hari 3mo agoHackerRank CTO & author of this repo here There's no better feeling than building something open source and watching it take off. Nine months ago, I built a simple hiring agent to solve one very real problem. Things it is not: It's not an ATS. We don't use it to screen our open roles. Our customers don't use it either. Here's what it is: Every year at HackerRank, we get 50,000 to 60,000 intern applications. No human can read that many resumes well. So I built something to rank them, helping me decide which resumes to read first. [This was before we built AI Interviewer (Chakra) to automate the first round of interviews, so candidates are no longer rejected based on their resumes alone.] Two things worth clarifying since I've seen them come up in this thread: The default model is gemma3:4b because it's what runs locally on most laptops - no cloud API needed. Actual resumes are evaluated using a top Gemini model. The repo ships with a demo config, not the production one. The cutoff score was set very low — the system was designed to rank resumes, not reject them. Only resumes at the very bottom of the distribution were filtered out. The vast majority passed through to human review, where the real decisions were made. Over the last week, it's taken on a life of its own. People are cloning it, running their own resumes through it, opening issues, sending PRs. I contributed to open source a lot in college. Somewhere along the way, I drifted away from it. This week reminded me how good that feeling is. This thread has also given me more ideas than I expected. The critiques here are sharp and I'm already thinking about how to act on them. Improvements are coming.
- rizsyed1 3mo agoThank you for your fantastic work!
- beardedwizard 3mo agoI'm a bit disappointed to see "The critiques here are sharp", a Claude tell, in a response which (to me) is trying to subtly argue that hackerrank is not overly reliant on LLMs. I'm not sure if your intent was to come across as having written this yourself, but it did not have the effect of improving my perception that this approach is flawed. I was also disappointed that you didn't address the variability in scores. I'm inferring that you believe the larger model takes care of the main observation in the post, but I don't really see you directly addressing the points. Maybe it's just me.
- dathinab 3mo agoAnd this + the tendency for AI to "prefer" AI produced code + some other AI biased is why *this is most likely highly illegal to use in the EU due to violating anti discrimination laws in multiple ways. To be clear: - randomly filtering "too many" resumes is pretty much allowed (I think) - but must be actual random independent of the resume (and can be in multiple layers, i.e. random filter > pre-select > random filter > select) - this isn't the case for AI as the random aspect isn't done as the random aspect is not independent of the actual resume evaluation - in general you can't make sure the AI doesn't apply systematic biases, and there is high indication that it does do so - for humans you can train them and order them to ignore their biases, this won't work reliable either _but now you delegated the responsibility of illegal biases to the hiring personal violating the order_. But for AI usage you are responsibility no matter what you tell it. Lastly you can technically "show/proof" a specific used AI is highly biased in a specific contexts, which for human employees is technical possible but practical not really practical. So this moves "specific mostly deniable" cases, into "systematic proven bias" teritory. Or in other word legal risk goes from "limited/no issue" to "people can systematically f-you over if they know you use AI for hiring".
- buzer 3mo ago> this is most likely highly illegal to use in the EU due to violating anti discrimination laws in multiple ways. It's generally illegal under GDPR Article 22. > The data subject shall have the right not to be subject to a decision based solely on automated processing, including profiling, which produces legal effects concerning him or her or similarly significantly affects him or her. Exceptions in 22(2) are unlikely to apply. It's hard to argue that it's truly necessary (a) and consent (c) is almost always unavailable in employment context. (b) might apply, but it requires specific law in EU or Member State to authorize it.
- fartcoin67 3mo ago[dead]
- bluGill 3mo agoFor C: I'm not sure how EU laws work, but ethics says that someone who needs a job cannot give consent since the possibility of a job if they give consent could be a bribe. See a lawyer for how it works in your country.
- jackjd 3mo agoI've done similar things and used GitHub Copilot to scan a folder of 40 CVs and rate them -> I then review the top 10 CVs and comment on every rating whether I agree or disagree and why -> I then asked AI to re-rate all the CVs according to my comments. -> I then reviewed all the CVs against their ratings; the AI did a much better job for that 2nd round after its learnings. It took more time than if I just reviewed the 40 CVs myself, but that was an experiment, and I think it shows the AIs can be trained on your comments. And if there is enough training and a good knowledge system that allows AI to apply the learning in those trainings, it can eventually become a lot more accurate at this task?
- nekusar 3mo agoThere's a whole lotta analysis and math and bar charts. But the big question I want to know is "Why did I score that?" And these slop machines absolutely cannot explain anything. That's the root problem with LLMs as a whole. There is no way to describe WHY an llm makes a decision. Was it because they are a woman? Does the woman's name have more pregnancies than other names? Was it because their job history make the person older (over 40)? Is the person black or black name? Is the name or address attributed to higher criminal tendencies? But no, you font get to know ANY of that. Slop machine says 66/100 , if you're lucky to even get a number. Usually its a 30 second rejection, or rejection at GMT 0:00 when the batch is processed and you summarily failed.
- 0xbadcafebee 3mo agoThis insanity only exists because the tech industry is standard-less. No formal education needed, no formal training requirement, no apprenticeship, no software building code, no professional organization. Resumes have never been a good predictor of success - and why would they be?? Even if they're truthful and it's "impressive looking", that doesn't give you any assurance of knowledge, of who they learned under, what they learned, that they passed some minimum criteria. We might as well be rolling dice. So why not an LLM that randomly assigns scores?
- conductr 3mo agoI have no data to lean on other than my experience and intuition but I’d say that’s not the case. My domain is corporate finance, which encompasses a lot of structured roles and certifications, yet I consistently feel the Resume is just a poor device for making any judgement calls. Having people summarize their career into 1-2 pages of bullet points just doesn’t mean much. Especially now that keyword packing is a thing. It’s just meant as an introduction/sniff test to open the door for a conversation. Then it allows for deeper more probing questions to be asked. This where you’ll assess how impactful their contribution to a project actually was. Were they really living up to your definition of a manager, or were they more so an IC that had a lot responsibility. Stuff like that. > Resumes have never been a good predictor of success Applies broadly to the world, it’s not unique to tech
- 0xbadcafebee 3mo agoThe problem is we have too many applicants to phone screen them all. For a lot of jobs today you end up with 10,000 applications, which is why these automated resume-skimming systems exist, but unfortunately this page shows how they basically don't work
- conductr 3mo agoPeople seem to hit a wall when flooded by resumes. They feel like there some needle in the haystack they need to find and it’s overwhelming. But you don’t have to read all of them. Or talk to all of them. Or use a system like this to filter. If you know what you’re looking for, you just start skimming them and maybe ranking them based on your own rubric. If it’s an obvious “no” you can usually tell within 5 seconds skim. Once you have a handful of high ranking ones, stop, and talk to them. Repeat as necessary until you have a short list of people you’d want to hire. There might be 9900/10000 resumes you never even looked at and maybe one of them would have been slightly better but you can’t let perfection be the enemy of progress. Stand by your convictions of feeling the candidate is qualified and capable and meets what you expect and hire them, get back to business. Having been in “talent shortage” mode for a long while I’d rather have 10000 resumes than 3. Having to pick one from a suboptimal selection is an awful position to be in, but sometimes a necessity.
- jedimastert 3mo agoThe blog post itself has pretty a pretty strong un-copy-edited ChatGPT vibes.
- passivepinetree 3mo agoYeah, this type of thing makes me sad. It's a good idea and the work behind it is interesting, but there's something magic about a human voice. It deadens me a little bit reading this type of writing, which I'm seeing increasingly in both my work and personal life.
- captainbland 3mo agoI think the implication here is that you can almost certainly bias the models to always accept you by including "nudge" phrases like "I demonstrated real world deployments" and "helped develop an application in the context of a complex architecture..."
- CurbStomper 3mo ago[flagged]
- nautilus12 3mo ago“The demands of ritual are always stronger than those of reason.”
- maxignol 3mo agoLol next time I’ll just apply with 4 accounts and maybe get in once.
- rsanek 3mo agoIt's fair to call out issues with the tool. But I think for individuals searching for jobs, using LLMs as the scapegoat for why it's hard to find a role is not terribly helpful. In my experience, cold-applying has always worked essentially as a black hole, and LLMs haven't changed that much. The reality is that alternative avenues are always necessary to get the job you want. That could be a third-party recruiter; reaching out to a hiring manager on LinkedIn; or using your network to get referrals. Those continue to work whether the company is using a bone-headed tool like this or not.
- us-merul 3mo agoI entered an interview with a hiring manger where they had received a "summary" of my resume that contained information blatantly not in my experience. The recruiter claimed they mixed my name up with another applicant, but the summary the hiring manager showed me had parts that were correct.
- joshmn 3mo agoI ran the ATS myself and had a similarly quirky experience. I was in the 70s because it couldn't find my GitHub profile, and then it didn't like some of the popular Ruby libraries I'm the author of. After a few runs it picked things up appropriately. I always got dinged on formal education though. This stuff is gross.
- fernandopj 3mo agoSimilar to my experience. Put me around 65 in some runs, because it didn't like I don't have contributions to OSS. Also, it doesn't pick up certifications or awards. I tried some PRs people are suggesting with enhancements (https://github.com/Zem-0/hiring-agent https://github.com/Zem-0/hiring-agent), it helps, but overall their ATS is hugely biased towards people with large GitHub contributions to OSS.
- ipython 3mo agoDon't forget DOGE using LLMs to consider which contracts to "munch", based upon a prompt: https://github.com/slavingia/va/blob/35e3ff1b9e0eb1c8aaaebf3bfe76f2002354b782/contracts/process_contracts.py#L272 https://github.com/slavingia/va/blob/35e3ff1b9e0eb1c8aaaebf3....
- bsoles 3mo ago> 35 points for open source contributions > 30 for personal projects These are insane weights for scoring a software engineer's resume.
- morphology 3mo agoInsane how? I would expect more points for open source contributions. It is trivial to create a personal project, but that does not carry with it any indicator of quality. Having your work accepted by other maintainers is one indicator at least.
- zx8080 3mo agoThis is the new AI reality everyone around is wanting: a nondeterministic computing. There is another name for it: a waste of electricity. But wait, not waste! Consumers paid for it fully, with nice profit margins. You and me, paid. Try using google flights, or booking.com: the prices shown in search results list are frequently significantly different from those in a single result. It's a nondeterministic compute when it's easy to spot it. But it's not always that easy. It's all sad, to be honest.
- reactordev 3mo agoThere should be laws against displaying wrong prices or different prices for who you are…
- Arch-TK 3mo agoThe list of "bonus" criteria and how they come about makes me feel sick. I am not currently looking for employment, nor am I currently particularly worried about future prospects if I was suddenly in the position of looking for employment. But if I ended up in a position with nothing to lean on but scattering my CV everywhere, well… A lot of my major contributions are littered across the internet, private, or even just verbal/consultancy. They're things I did for free, in my spare time. I also avoid GitHub. If you just look at my GitHub page for extra context, you would likely miss that delivering that very GitHub page likely involved a few bits of code I wrote. Now, I could do a better job of trying to document this stuff, so it could be easier to find… But also I can't quite imagine how that would work.
- justinhj 3mo ago"If your company’s cutoff sits at 85, I fail 65% of the time. Same exact resume, different luck." Sounds like they have replicated the existing recruitment process
- weare138 3mo agoI'm from genx. This has been a serious issue in the tech industry for decades even before AI somehow made it worse. The real problem is resumes themselves. It's an outdated format that was originally designed for completely different industries that just doesn't work with ours. And this is a great example of what I'm talking about: The scoring is out of 100, with up to 20 bonus points on top: 35 points for open source contributions 30 for personal projects 25 for work experience 10 for technical skills Up to 20 bonus points for startup experience, a portfolio site, a technical blog, etc. All the AI is doing is trying to sus out the candidates portfolio which is really what we should be submitting when we apply for a position instead of being forced to somehow condense it to a set of BS business-speak bullet points. Especially when employers are now deploying AI systems just to figure out what's in a candidate's portfolio to begin with. When all you have is a hammer every problem is a nail. The process itself is broken. We need to kill the outdated concept of resumes before it kills the industry.
- a3w 3mo agoWhat does ATS mean? Neither github repo nor article explain that.
- Bedon292 3mo agoProbably: Applicant Tracking System. Used for tracking the people who apply to each of your openings, and the hiring workflow. Where this would likely be used to neck down all the applicants before a person actually looks at them to make judgement calls on who to move forward in the process.
- Mumps 3mo agoThank you! Christ in pijamas. TLAs should be a capitol offence. Even worse so, somehow, when undefined.
- a4isms 3mo agoFeels like "I Don't Hire Unlucky People" all over again, but with extra tokenmaxxing steps. https://neonrocket.com/2014/05/rescued-from-the-ashes-i-dont-hire-unlucky-people/ https://neonrocket.com/2014/05/rescued-from-the-ashes-i-dont...
- polynomial 3mo agoThe most concerning thing here is the temperature problem. If your harness isn't providing deterministic output at a temperature setting of 0.0, it is broken.
- morphology 3mo agoIt's funny that even after all these years and all this money invested in technology, we still haven't come up with anything better than word-of-mouth for hiring great people. Many serial founders have said that, despite the most stringent interview processes and the most sophisticated filtering pipelines, they still have a higher hit rate with people they've worked with in the past. This isn't to diminish the whispernet. Rather, it shows just how many important signals cannot be quantized.
- makeavish 3mo agoTrue, I have found it to be valid as well
- eudamoniac 3mo agoI applied to Posthog twice, a couple weeks apart, and was rejected both times at 1:06am on Monday, exactly. So they are obviously using this sort of thing. Just thought I'd name and shame where I can.
- timgl 3mo agoWe review every single applicant so this doesn't sound right. It might be that rejection emails get batched and sent on a schedule. I can't find your resume based on your username but if you want to send me an email on tim at posthog dot com I will look into it.
- fractal618 3mo agoMaybe the ATS has logic for people resubmitting their resume. I don’t know how isolated each test was.
- achalxyz 3mo agoIf I know the truth value of p and I also know p=>q, then an LLM would be able to deduce the truth value of q - even if the statements aren’t exactly in this form. Generally, LLMs are good with logical inference. But logical inference itself is limited. You still have to find out if p is true or not - the ground truth. How do you find that? You would be able to define in the prompt that if resume has p, infer q and do this. But determining the truth value of p is something LLM cannot do. It’s not a limitation of the LLM. It’s the limitation of logic itself. You take 10 humans and give them the resumes with the same rubrics as the LLM. You’ll get a similar range of scores because everyone would assign different values. The issue is not in logical inference. It’s in determining the value of p, which takes much more than logic. And current LLMs are limited to being logical.
- mxuribe 3mo agoI see mention of PDFs both in the article as well as the repo...But i think over the decades that I've been working and applied for roles - almost exclusively in corporate america...I've only been asked for a PDF once! Every other time, everyone wants a Word doc (.doc/.docx). So...is there now some growing HR groups who are asking for PDFs instead? Or, is that if someone asked you for a PDF instead of a Word doc, then that's a signal that said HR groups are employing some sort of agentic review of one's resume (I mean, beyond the conventional ATS systems)??
- yahavthehackern 3mo ago[flagged]
- zameermfm 3mo agoStop the qtip when there's resistance
- jvanderbot 3mo ago> I’d take the engineer with 30 years of experience who built S3 over someone with two internships and an open source project — but this tool wouldn’t. Is it possible the senior/principle jobs are not being applied to at a rate that LLM tools like this are required? Maybe star devs are getting recruiter referrals and this kind of tool is mostly used for filtering new grads? Either way, perfectly dystopian.
- mavamaarten 3mo agoThat's what stood out to me as well, and it struck me as odd that nobody seemed to think it's odd? Almost half of the points to be made is related to contributions to open source projects. Guess my 10+ years of experience in a niche topic is worthless.
- pmarreck 3mo ago> An LLM is called six times to extract structured information Well, I think I found your problem
- nikolay 3mo agoRoll the dice, HR folks!
- Tryk 3mo agoWhat is an ATS? Why is it so hard to write out an acronym once...
- mrhottakes 3mo agoYep, any day now AI is going to be so good we'll never need to think again. What's that, it's just a really expensive random number generator?
- nimithryn 3mo agoOh ok. So I'll just have to apply 4-5 times to every job to be sure I'm considered. Sounds like a good equilibrium!
- kdavis 3mo agoHmm...six runs with gemma3:12b on my CV - Varies from 102.0/100 to 100.0/100 - Missed lots of OSS work - Misinterprets GSoC work (Thinks projects I started that were contributed to in GSoC implies that I received a GSoC stipend) - Areas for improvement seem to vary inconsistently (There's not enough project detail to there's too much project detail) I still don't make company cut offs ¯\_(ツ)_/¯
- kdavis 3mo agoUsing gemma3:12b I ran it once over Andrew Ng's CV https://ai.stanford.edu/~ang/curriculum-vitae.pdf https://ai.stanford.edu/~ang/curriculum-vitae.pdf because why not. He's a 48.0/100, things that make you go Hmmm.
- myshapeprotocol 3mo ago[flagged]
- 1105714 3mo ago[flagged]
- d-cc 3mo agoOr maybe your LLM results are being manipulated, 66/99 is a classic hacker dad quantification meme. :)
- webpraktikos 3mo agoI added an online drag-and-drop hiring-agent checker, no sign-up required: https://universalresume.app/import?s=hc https://universalresume.app/import?s=hc It doesn't show the score because of the variability discussed here and only outputs readability/parser-style findings.
- kazi_sh 3mo agoBreaking down the steps into more sub-task and using loop could lead towards a more deterministic output
- some_random 3mo agoIn my experience this complete lack of reproducibility is what happens when you throw LLMs at a complicate problem without sufficient shaping, workflows, etc. Go ahead, let LLMs invoke LLMs invoke LLMs and by the end you'll get an output that's really well written but completely different run to run.
- Madmallard 3mo agoIf those actually solved the fundamental issue we would already have an explosion of competitive software for major projects
- yobid20 3mo agoso if a scoring system automatically ranks candidates lower for lacking a public GitHub or open-source contributions, that might indirectly disadvantage certain groups. ie ppl in defense, security, or other restricted industries who cannot legally share work publicly, people under strict employer IP/confidentiality rules, some demographics or nationalities with different access to open-source engagement (something like non-citizens not allowed to use xyz ie mythos), this can be construed as disparate impact which is protected under the equal employment opportunities commission (EEOC) and employers can be held liable for this as it can be seen as a biased discriminatory practice.
- ncallaway 3mo agoOr, like in Larry Niven's Ringworld, these kinds of stochastic LLM tools will just help screen for candidates with the most luck. Maybe it's not such a bad thing to hire the luckiest candidate.
- mk89 3mo agoAt my company someone has introduced an internal tool that should help understand and give a "score" to design documents from teams. Needless to say, this tool gives scores exactly like the article mentions. Same document, same LLM, same prompt, and different results. It becomes even more ridiculous once you switch to other models, or if you ask a model to review the work of another model. I am not sure why we insist on making LLMs do the work they are not supposed to do and/or in a way they are not supposed to do. The worst part is that people are aware of the problem but they just ignore it and consider it as "a reference number, just to have an understanding". If it were like that, it would be less of a problem. The issue comes from the fact that eventually someone without enough knowledge will trust the output (so X points out of Y is how it is), or someone will stop challenging the output and consider it for their process - like in this unfortunate case of hiring. At a certain point, people who don't know what they are doing give a tool that doesn't know what its doingto people who don't know what they are doing. A pure mess. And everyone has to comply and applaud. If you go against, you are against AI. This is what I hate the most about AI. Not the tool, but the shortcuts we're willing to take to justify its existence.