10 ms·
LLMs are mortally terrified of exceptions
https://x.com/karpathy/status/1976082963382272334 https://x.com/karpathy/status/1976082963382272334
- mwkaufma 11mo agoEven when they're not AI slop, these kinds of "paranoid sanity checks" are the software equivalent of security-theater.
- bwfan123 11mo agoForm over function is what they are trained for. So, verbose commentary, needless readmes, and emojis all serve that purpose.
- mwkaufma 11mo agoCoding for the reviewer, not the user.
- simonw 11mo agoYeah, I really hate code like this because it generally ends up full of codepaths that have never been exercised, so there's all sorts of potential for weird behavior and unexpected edge cases. Plus it's harder to review.
- dwd 11mo agoSometimes security theater is what you need to not trigger a false positive on a static code analysis. I haven't needed to use a service like Fortinet recently and am now wondering if a LLM is part of their tool and if it's better/worse?
- shiandow 11mo agoIs there a way to read the rest?
- hugo1789 11mo agohttps://xcancel.com/karpathy/status/1976082963382272334 https://xcancel.com/karpathy/status/1976082963382272334
- bobogei81123 11mo agoThis is just AI trying to tell us how bad we designed our programming languages to be when exceptions can be thrown pretty much anywhere
- recursive 11mo agoSo you think java's checked exceptions are a better model? No opinion myself, but that way seems widely considered bad too.
- nivertech 11mo agoWhy do you need exceptions at all? They’re just a different return types in disguise… Also, division by zero should return Inf
- threeducks 11mo agoOr -Inf, depending on the sign of the zero, which might catch some programmers by surprise, but is of course the correct thing to do.
- nivertech 11mo agoa/0 = Inf when a>0 a/0 = -Inf when a<0 a/0 = NaN when a=0
- dmoy 11mo agoNo this doesn't work either In the context of say a/-0.001, a/-0.00000001, a/-0.0000000001, a/<negative minimum epsilon for denormalized floating point>, a/0 Then a/0 is negative when a>0, and positive when a<0
- nivertech 11mo agoWhy not just to use IEEE 754? > According to the IEEE 754 standard, floating-point division by zero is not an error but results in special values: positive infinity, negative infinity, or Not a Number (NaN). The specific result depends on the numerator
- deleted 11mo ago[deleted]
- wffurr 11mo agoTurns out computer math is actually super hard. Basic operations entail all kinds of undefined behavior and such. This code is a bit verbose but otherwise familiar.
- Den_VR 11mo agoIf we wanted defined behavior we’d build systems with Karnaugh maps all the way down.
- falcor84 11mo ago# Step 3: Preemptively check for catastrophic magnitude differences if abs(a) > sys.float_info.max / 2: logging.warning("Value of a might cause overflow. Returning infinity just to be sure") return math.copysign (float('inf'), a) if abs(b) < sys.float_info.epsilon: logging.warning("Value of b dangerously close to zero. Returning NaN defensively.") return math.nan Does the above code make any sense? I've not worked with this sort of stuff before, but it seems entirely unreasonable to me to check them individually. E.g. if 1 < b < a, then it seems insane to me to return float('inf') for a large but finite a.
- OutOfHere 11mo agoIs this Claude? GPT is not like this. To me it looks like Anthropic is just maximizing billable token use as usual, and it has nothing really to do with exceptions per se.
- TuxSH 11mo agoFrom the UI it indeed seems to be Claude
- constantcrying 11mo agoIf you are dividing two numbers with no prior knowledge of these numbers or any reasonable assumptions you can make and this code is used where you can not rely on the caller to catch an exception and the code is critical for the product, then this is necessary. If you are actually doing safety critical software, e.g. aerospace, medicine or automotive, then this is a good precaution, although you will not be writing in Python.
- mewpmewp2 11mo agoI might agree with that, and maybe the example posted by Karpathy is not the greatest, but what I'm constantly being faced with is try catches where it will fail silently or return a fallback/mock response, which essentially means that system will behave unexpectedly in a more subtle way down the line while leaving you clueless to as what the issue was. I have to constantly remind Claude that we want to fail fast.
- isoprophlex 11mo agoA good 10% of my Claude.md is yelling at it that no i don't want you to silently handle exceptions six calls deep into the stack and no please don't wrap my return values in weird classes full of dumb status enums "for safety" Just raise god damn it
- hyperpape 11mo agoI'm not sure returning None is any safer than an Exception, because the caller still has to check.
- deleted 11mo ago[deleted]
- metalcrow 11mo agoGiven that the output describes the function as being done "with extraordinary caution, because you never know what can go wrong", i would guess that the undisclosed prompt was something similar to "generate a division function in python that handles all possible edges cases. be extremely careful". Which seems to say less about LLM training and more about them doing exactly what they are told.
- deleted 11mo ago[deleted]
- freehorse 11mo agoAside from the absurdity and obvious satirical intention, 1. the code is actually wrong (and is wrong regardless of the absurd exception handling situation) 2. some of the exception handling makes no sense regardless, or is incoherent 3. a less absurd version of this actually happens (edit: commonly in actual irl scenarios) if you put emphasis on exception handling in the prompt
- angry_albatross 11mo agoI interpreted the function code as being a deliberately exaggerated satirical example that was illustrative of the experience he was having. So yes, in that example it was probably told to be overly cautious, but I agree with him that the default of LLMs seems to be a bit more cautious than I would like.
- sarchertech 11mo agoExpert beginners program like this. I call it what it driven development. Turns out a lot of code was written by expert beginners because by many metrics they are prolifically productive. In go all SOTA agents are obsessed with being ludicrously defensive against concurrency bugs. Probably because in addition to what if driven development, there are a lot of blog posts warning about concurrency bugs.
- criemen 11mo agoIt's also logically incoherent - division by zero can't occur, because if b=0 then abs(b) < sys.float_info.epsilon. Furthermore, the code is happy to return NaN from the pre-checks, but replaces a NaN result from the division by None. That doesn't make any sense from an API design standpoint.
- dijksterhuis 11mo agoI mean, the first three cases are just attempting to turn dynamic into static typed... right? maybe just don't aim for uber-safety in a dynamically typed language? :shrugs: (I used to look out for kaparthy's papers ten years ago... i tend to let out an audible sigh when i see his name today)
- falcor84 11mo agoYou shouldn't have the same expectations from a person's tweet as you would from a paper. I don't see any issue with high profile people who are careful in their professional work, putting less thought-through output on social media. At least as long as they don't intentionally/negligently spreading misinformation, which I've never seen Karpathy do. I for one really enjoy both his longer form work and his shorter takes.
- falcor84 11mo agoThat code has many issues, but the one that bothers me the most in practice is this tendency of adding imports inside functions. I can only assume that it's an artifact of them optimizing for a minimal number of edits somewhere in the process, but I expect better.
- kccqzy 11mo agoIt's to make imports lazy, to solve the issue of slow import at startup.
- falcor84 11mo agoWhile there are some cases where lazy imports are appropriate, this function, and the vast majority of such lazy imports that I get from Claude are not. In particular, I can't think of any non-pathological situation where a python developer should import logging and update logging.basicConfig within an inner function.
- Maxion 11mo agoIt's also a trick in python to deal with circular imports.
- jamster02 11mo agoI also recently ran into a problem when unit testing and monkey patching where I had to import after monkey patching, so in the function itself.
- andrewmcwatters 11mo agoIn very, very large projects, you end up finding that you want lazy initialization as much as possible, because it greatly affects startup times.
- natnat 11mo agoI think this has a lot to do with the mechanism of RoPE attention, where physical closeness in the code is a signal of relevance.
- jpcompartir 11mo agoMost comments seem to be taking the code seriously, when it's clearly satirical?
- fkyoureadthedoc 11mo agoNot sure why but it made me think of FizzBuzzEnterpriseEdition https://github.com/EnterpriseQualityCoding/FizzBuzzEnterpriseEdition https://github.com/EnterpriseQualityCoding/FizzBuzzEnterpris...
- ineedasername 11mo agoWoah, were they using junit 4.8.3 in that project? Someone was flying by the seat of their pants, I hope they got sign-off on that by legal & the CTO, that’s the kind of cowboy coding choice that can hurt a career.
- stefanfisk 11mo agoPRs are welcome!
- lawlessone 11mo ago...so many folder and files , i feel damaged after seeing that. Great satire.
- glitchc 11mo agoBut what's the prompt that led to this output? Is it just a simple "Write code to divide a by b?" or are there instructions added for code safety or specific behaviours? I know it's Karpathy, which is why the entire prompt is all the more important to see.
- johnisgood 11mo ago"Write me a code that divides a by b and make sure it is safe and handles all edge cases"[1] or something and some languages have more than others. [1] Probably with some "make you sure handle ALL cases in existence", or emphasis, along those lines.
- stargrazer 11mo agobut then, why code with exceptions, why not perform pre-flight/pre-validation checks and minimize exceptions to the truly unknown?
- jampekka 11mo agoI've noted that LLMs tend to produce defensive code to a fault. Lots of unnecessary checks, e.g. check for null/None/undefined multiple times for same valie. This can lead to really hard to read code, even for the LLM itself. The RL objectives probably heavily penalize exceptions, but don't reward much for code readability or simplicity.
- MakeAJiraTicket 11mo agoI have a function that compares letters to numbers for the Major System and it's like 40 lines of code and copilot starts trying to add "guard rails" for "future proofing" as if we're adding more numbers or letters in the future. It's so annoying.
- comex 11mo agoThis is a parody but the phenomenon is real. My uninformed suspicion is that this kind of defensive programming somehow improves performance during RLVR. Perhaps the model sometimes comes up with programs that are buggy enough to emit exceptions, but close enough to correct that they produce the right answer after swallowing the exceptions. So the model learns that swallowing exceptions sometimes improves its reward. It also learns that swallowing exceptions rarely reduces its reward, because if the model does come up with fully correct code, that code usually won’t raise exceptions in the first place (at least not in the test cases it’s being judged on), so adding exception swallowing won’t fail the tests even if it’s theoretically incorrect. Again, this is pure speculation. Even if I’m right, I’m sure another part of the reason is just that the training set contains a lot of code written by human beginners, who also like to ignore errors.
- MakeAJiraTicket 11mo agoDefensive programming is considered "correct" by the people doing the reinforcing, and is a huge part of the corpus that LLM's are trained on. For example, most python code doesn't do manual index management, so when it sees manual index management it is much more likely to freak out and hallucinate a bug. It will randomly promote "silent failure" even when a "silent failure" results in things like infinite loops, because it was trained on a lot of tutorial python code and "industry standard" gets more reinforcement during training. These aren't operating on reward functions because there's no internal model to reward. It's word prediction, there's no intelligence.
- aoeusnth1 11mo agoYou used the word reinforcing, and then asserted there's no reward function. Can you explain how it's possible to perform RL without a reward function, and how the LLM training process maps to that?
- MakeAJiraTicket 11mo agoLLM actions are divorced from that reward function, it's not something they consult or consider. Reward function in that context doesn't make sense.
- karpathy 11mo agoSorry I thought it would be clear and could have clarified that the code itself is just a joke illustrating the point, as an exaggeration. This was the thread if anyone is interested https://chatgpt.com/share/68e82db9-7a28-8007-9a99-bc6f0010d101 https://chatgpt.com/share/68e82db9-7a28-8007-9a99-bc6f0010d1...
- chis 11mo agoI think there’s always a danger of these foundational model companies doing RLHF on non-expert users, and this feels like a case of that. The AIs in general feel really focused on making the user happy - your example, and another one is how they love adding emojis to the stout and over-commenting simple code.
- cma 11mo agoAnd more advanced users are more likely to opt out of training on their data, Google gets around it with a free api period where you can't opt out and I think from did some of that too, through partnerships with tool companies, but not sure if you can ever opt out there.
- cma 11mo ago*grok, not 'from'
- miki123211 11mo agoThis feels like RLVR, not RLHF. With RLVR, the LLM is trained to pursue "verified rewards." On coding tasks, the reward is usually something like the percentage of passing tests. Let's say you have some code that iterates over a set of files and does processing on them. The way a normal dev would write it, an exception in that code would crash the entire program. If you swallow and log the exception, however, you can continue processing the remaining files. This is an easy way to get "number of files successfully processed" up, without actually making your code any better.
- 11mo ago
- ziml77 11mo agoThat's funny but definitely not far off from reality. I have instructions from my agent to use exceptions but they only help so much. I really dislike their underuse of exceptions. I'm working on ETL/ELT scripts. Just let stuff blow up on me if something is wrong. Like, that config entry "foo" is required. There's no point in using config.get("foo") with a None check which then prints a message and returns False or whatever. Just use config["foo"] and I'll know what's wrong from the stack trace and exception text.
- ziml77 11mo agoAaaand there we go. I literally just ran into a problem with code someone had used AI to write which does this log and continue nonsense. Process spent 30 minutes looping through API requests and failing to persist the response on every single one because of a permission error. But the only indication of a problem is the errors in the log, the process finished with a successful exit code.
- CGamesPlay 11mo agoI dealt with this in my AGENTS.md by including a recap of the text of "Vexing Exceptions" [0], rephrased as a set of guidelines for when to write a throw or catch. I feel like it helped; and when it still emits error handling I disagree with and I ask about it, it will categorize it into one of the four categories, and typically rewrite it in an appropriate way. I think the Vexing Exceptions post is on the same tier as other seminal works in computer science; definitely worth a quick read or re-read once in a while. [0] https://ericlippert.com/2008/09/10/vexing-exceptions/ https://ericlippert.com/2008/09/10/vexing-exceptions/
- furyofantares 11mo agoA couple thoughts. One is that often I do want error handling, but also often I either know the error just won't happen or if it does, something is very wrong and we should just crash fast to make it easy to fix the bug. But I am not really sure I would expect someone to know the difference in all cases just looking at some code. This is often an about holistically knowing how the app works. A second thought - remember the experiment where an LLM was fine tuned on bad code (exploitable security problems for example) and the LLM became broadly misaligned on all sorts of unrelated (non-coding) tasks/contexts? It's as if "good or bad" alignment is encoded as a pretty general concept. Error-handling is good aligned, which I think is why, even with lots of instructions to fail fast, it's still hard to get the LLM to allow crashing by avoiding error checking. It's gonna be even harder if you do want it to do some error checking, and the code it's looking at has some error checking
- exasperaited 11mo agoWhy do LLMs do it for real: because you trained them by stealing all of Stack Overflow? Less sarcastically but equally as true: they've learned from the tests you stole from people on the internet as well as the code you stole from people on the internet. Most developers write tests for the wrong things, and many developers write tests that contain some bullshit edge case that they've been told to test (automatically to meet some coverage metric, or by a "senior" developer who got Dilbert principled away from the coalface and doesn't understand diminishing returns). But then the end goal is to turn out code about as good as the average developer so they can be replaced more cheaply, so your LLM is meeting its objectives. Congrats.
- iagooar 11mo agoThis issue has been one of the biggest issues with the Claude models, not so much with GPT-4 or GPT-5. I even had this Cursor rule when I was using Claude: "- Do not use statements to catch all possible errors to mask an error - let it crash, to see what happened and for easier debugging." And even with this rule, Claude would not always adhere. Never had this issue with GPT-5.
- bhl 11mo agoWhat’s the solution here, reward code that works without try catch, reward code that errors and is caught, but penalize code that has try catch and never throws an error?
- never_inline 11mo agoI think too much of RLHF is done on small scale tutorial-ish examples. LLMs often write tutorial-ish code without much care how it integrates with rest of codebase. Swallowing exceptions is one such example.
- dgan 11mo agoI spent more time dismissing various popups than reading this post. I hate twitter links
- mcintyre1994 11mo agoJust add “cancel” after the x to get a viewable version of any Twitter link: https://xcancel.com/karpathy/status/1976082963382272334 https://xcancel.com/karpathy/status/1976082963382272334
- cess11 11mo agoNow this is a toy example because usually you never do division this way, but in mature code in commercial applications this is usually what it looks like. It's a sliver of business logic that in itself seems trivial, and then handlers of edge case upon edge case upon edge case, mirroring an even larger set of unit tests. One reason for this is that you typically lack a type system that allows 'making illegal states unrepresentable' to some extent, or possibly lack a team that can leverage the available type system to that effect due to organisational pressure, insufficient experience or whatever.
- classified 11mo agoI'd love to see an LLM shake with fear and beg for mercy.
- blitzar 11mo agoI figured this is how the pros write their code and I have been holding the code wrong the whole time.
- beepdyboop 11mo agoLove this, but I do struggle with this same problem! How do we circumvent it?
- capestart 11mo ago[dead]
- stuaxo 11mo agoMy wishlist top item is to stop creating a class with Service on the name and having things come off it, when all I needed was functions and methods, the dev I was working with submitted a lot of these and in testing I could get the LLM to do it easily myself.
- paulhodge 11mo agoAgree that LLMs go too far on error catching.. BUT, to play devil's advocate a little: Most human coders should be writing a lot more try/catch blocks than they actually do. It's very common that you don't actually want an error in one section (however unlikely) to interrupt the overall operation. (and sometimes you do, it just depends)
- winddude 11mo agobut like a normal dev no unit test.