5 ms·
Load just makes LLMs behave less deterministically and likely degrade. See: https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/ https://
by botacode 8mo ago
Load just makes LLMs behave less deterministically and likely degrade. See: https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/ https://thinkingmachines.ai/blog/defeating-nondeterminism-in...
They don't have to be malicious operators in this case. It just happens.
- bgirard 8mo ago> malicious It doesn't have to be malicious. If my workflow is to send a prompt once and hopefully accept the result, then degradation matters a lot. If degradation is causing me to silently get worse code output on some of my commits it matters to me. I care about -expected- performance when picking which model to use, not optimal benchmark performance.
- Aurornis 8mo agoNon-determinism isn’t the same as degradation. The non-determinism means that even with a temperature of 0.0, you can’t expect the outputs to be the same across API calls. In practice people tend to index to the best results they’ve experienced and view anything else as degradation. In practice it may just be randomness in either direction from the prompts. When you’re getting good results you assume it’s normal. When things feel off you think something abnormal is happening. Rerun the exact same prompts and context with temperature 0 and you might get a different result.
- dingnuts 8mo ago[dead]
- bonoboTP 8mo agoThis has nothing to do with overloading. The suspicion is that when there is too much demand (or they just want to save costs), Anthropic sometimes uses a less capable (quantized, distilled, etc) version of the model. People want to measure this so there is concrete evidence instead of hunches and feelings. To say that this measurement is bad because the server might just be overloaded completely misses the point. The point is to see if the model sometimes silently performs worse. If I get a response from "Opus", I want a response from Opus. Or at least want to be told that I'm getting slightly-dumber-Opus this hour because the server load is too much.
- F7F7F7 8mo ago“Just drink the water, it’s all water.”
- novaleaf 8mo agothis is about variance of daily statistics, so I think the suggestions are entirely appropriate in this context.
- altcognito 8mo agoExplain this though. The code is deterministic, even if it relies on pseudo random number generation. It doesn't just happen, someone has to make a conscious decision to force a different code path (or model) if the system is loaded.
- pertymcpert 8mo agoFloating point math isn't associative for operations that are associative in normal math.
- measurablefunc 8mo agoThat would just add up to statistical noise instead of 10% degradation over a week.
- kevin_thibedeau 8mo agoCatastrophic error accumulation can produce more profound effects than noise.
- measurablefunc 8mo agoJust to make sure I got this right. They serve millions of requests a day & somehow catastrophic error accumulation is what is causing the 10% degradation & no one at Anthropic is noticing it. Is that the theory?
- pertymcpert 7mo agoFYI something in that region happened last august/September. Some inference bug triggered worse performance on TPUs vs GPU.
- FL33TW00D 8mo agoIt takes a different code path for efficiency. e.g if (batch_size > 1024): kernel_x else: kernel_y
- strongpigeon 8mo agoThe question I have now after reading this paper (which was really insightful) is do the models really get worse under load, or do they just have a higher variance? It seems like the latter is what we should expect, not it getting worse, but absent load data we can't really know.
- stefan_ 8mo agoThe primary (non malicious, non stupid) explanation given here is batching. But I think you would find looking at large-scale inference the batch sizes being ran on any given rig are fairly static - there is a sweet spot for any given model part ran individually between memory consumption and GPU utilization, and generally GPUs do badly at job parallelism. I think the more likely explanation is again with the extremely heterogeneous compute platforms they run on.
- hatmanstack 8mo agoThat's why I'd love to get stats on load/hardware/location of where my inference is running. Looking at you Trainiuim.
- bonoboTP 8mo agoWhy do you think batching has anything to do with the model getting dumber? Do you know what batching means?
- stefan_ 8mo agoWell if you were to read the link you might just find out! Today is your chance to be less dumb than the model!
- bonoboTP 8mo agoI checked the link, it never says that the model's prediction get lower quality due to batching, just nondeterministic. I don't understand why people conflate these things. Also it's unlikely that they use smaller batch sizes when load is lower. They just likely spin up and down GPU serves based on demand, or more likely, reallocate servers and gpus between different roles and tasks.
- make3 8mo agoIt's very clearly a cost tradeoff that they control and that should be measured.