13 ms·
ECC matters
- wagslane 6y agoIt really does. I did a write-up recently on it as I was diving in and understanding the benefits: https://qvault.io/2020/09/17/very-basic-intro-to-elliptic-curve-cryptography/ https://qvault.io/2020/09/17/very-basic-intro-to-elliptic-cu...
- avianes 6y agoBe careful not to confuse ECC memory with ECC encryption. ECC memory = memory with Error-Correcting Code ECC encryption = Elliptic Curve Cryptography
- linsomniac 6y agoThis reminds me of last year we ordered a new $14K server, it arrived and we ran it through our burn-in process which included running memtest86 on it, and it would, after around 7 hours, generate errors. Support was only interested if their built-in memory tester, which even on it's most thorough, would only run for ~3 hours, would show errors, which it wouldn't. IIRC, the BMC was logging "correctable memory errors", but I may be misremembering that. "We've run this test on every server we've gotten from you, including several others that were exactly the same config as this, this is the only one that's ever thrown errors". Usually support is really great, but they really didn't care in this case. We finally contacted sales. "Uh, how long do we have to return this server for a refund?" All of a sudden support was willing to ship us out a replacement memory module (memtest86 identified which slot was having the problem), which resolved the problem. They were all too willing to have us go to production relying on ECC to handle the memory error.
- scottlamb 6y ago> They were all too willing to have us go to production relying on ECC to handle the memory error. Good call in not accepting this. Even ignoring the possibility you have a double-bit error that causes a crash, or a triple-bit error that maybe can't be detected, frequent ECC errors are problematic. I've encountered machines that consistently ran my software horribly slowly. I don't remember specifics, but let's say at least 100X latency of other machines for similar operations. When I dug in, I found these machines had a huge amount of correctable memory errors. The correction apparently degrades performance significantly. I'm not sure exactly why, but I guess there's an MCE trap to report the memory error, and perhaps that path is slow.
- indolering 6y agoMy favorite example is a bit flip altering election results: https://www.wnycstudios.org/podcasts/radiolab/articles/bit-flip https://www.wnycstudios.org/podcasts/radiolab/articles/bit-f...
- b0rsuk 6y agoI browsed some online listings for ECC memory modules, and they seem to be sold one module at a time. Standard DDR4 modules are sold in pairs, to benefit from dual channel mode. Does ECC memory support dual channel??
- tgbugs 6y agoA relevant Bryan Cantrill talk segment on this, which heightens the paranoia around this. Namely, firmware hiding correctable errors and only reporting uncorrectable errors. https://www.youtube.com/watch?t=2104&v=fE2KDzZaxvE https://www.youtube.com/watch?t=2104&v=fE2KDzZaxvE
- willis936 6y agoWhenever this topic comes up I wonder how much more resilient are CPU registers compared to DRAM.
- kozak 6y agoI'm about to write some code that will allocate a random buffer, take a checksum of it, and just sit on the buffer, periodically checksuming it again until a bit flips. Or maybe even allocate a buffer of zeros and wait until a non-zero appears in it.
- trissylegs 6y agoWhen I chose my PC parts when Ryzen first came out I tried to get ECC parts. The RAM was obtainable, the problem was that no motherboards had ECC support at the time. I hope the situation has improved by the time I get my next motherboard/cpu upgrade.
- raghavtoshniwal 6y agoOnce trained a GPT2 model to do text-gen on Linus’ emails. Boy there were some choice angry rants and non-sensical technical jargon that was generated
- sally1620 6y agoLinux is accusing Intel of killing ECC intentionally. But that is not really the case, they just wanted people to pay up. If you care about ECC, you pay for Xeon. Majority of consumers don't run critical applications on their devices, so they are happy with a cheap device that may crash once in a while. AMD is only changing the game because they are trying to undercut Intel. They have been putting pro features into all of their CPUs including over-clocking, extra PCIE lanes and ECC. Honestly, what is the point of bullet-proof hardware when the software reliability (at least on consumer devices) has gone down to two nines.
- fomine3 6y agoECC isn't enough to be bulletproof but improves reliability for well known relatively unreliable parts. Extra theoretical cost for ECC should be accepted for most of computer users. It also helps developing cheaper RAM technology (see what's happened on SSD).
- Dylan16807 6y agoIntel had to kill consumer ECC as part of making it a feature that people can "pay up" for. That's very intentional. > AMD is only changing the game because they are trying to undercut Intel. They have been putting pro features into all of their CPUs including over-clocking, extra PCIE lanes and ECC. You are correct to call them a corporation. AMD is not your friend, but they are the good actor in this fight.
- srtjstjsj 6y agoI guess Linus's recent project to communicate more respectfully didn't pan out.
- Noxmiles 6y agoI was reading it and thought: wow, this guy is absolutely right! Great things he's talking about. After reading it, i saw it was Linux Torvalds :D
- MisterTea 6y ago> ECC availability matters a lot - exactly because Intel has been instrumental in killing the whole ECC industry with it's horribly bad market segmentation. The phrase that strikes me is "horribly bad market segmentation". I agree 100%. Remember when the Pentium/pro/2/3 could operate in single and dual socket configurations with ECC? The same CPU that plugged into your low end consumer board could also plug into a high end server/workstation board. All you needed was the right motherboard.
- greyhair 6y agoECC is required on mission critical hardware. I have spent 36 years fielding embedded devices in core network (D1/E1, SONET, ROADM/MPLS, Cellular basestation) and I will tell you that large ECC covered memory arrays always show small numbers of correctable error events over the course of a year. I have seen, over the course of my career, exactly one controller card replaced early in the field, because it started throwing excessive recoverable ECC events over time, until it hit a threshold of 10x the average of a typical board. On the order of ten recoverable ECC events per month instead of one event per month. I have never observed a logged non-correctable ECC event in the field. In the lab, yes, but never in fielded equipment. If you are fine with your PC experiencing one or two bits flipped in memory every month, then you really don't need ECC. That is the question you need to answer. For mission critical systems? ECC is a requirement.
- kensai 6y ago“ECC availability matters a lot - exactly because Intel has been instrumental in killing the whole ECC industry with it's horribly bad market segmentation.” Its. There, I finally corrected Linus Torvalds in something. :))
- raverbashing 6y agoYeah I'm always annoyed with this kind of mistake. Especially as non-native speakers should know better than the native ones (which usually don't give a f.). Now the point about internally doing ECC is an interesting one, could be a way out of this mess. And apparently ECC is more available in AMD land
- touisteur 6y agoI think it's available for customer SKUs on AMD and not just for servers like in 'Xeon-land'... How I've wanted an ECC-ready NUC...
- jeffbee 6y agoThe AMD parts all have the ECC feature but the platform support outside of EPYC may as well not exist. Most motherboards for the Ryzen segment don’t do it properly or don’t do it at all, some support it but aren’t capable of reporting events to the operating system which is dumb. Ryzen laptops don’t have it either. Closest you can come to a nuc with ecc is I think a mini server equipped with one of the four-core i3 parts that have ecc.
- touisteur 6y agoAh thanks for the real-world check on AMD Ryzen and ECC. Sad state of affairs really, when so many things are integrated in the SOC, why skimp on this... As for NUCs, I thought there were some Atom chips with ECC but not in NUC form factor ? Something shiny from Logic Supply might do then...
- erkkie 6y ago
- IgorPartola 6y agoI wish this was more of a cohesive argument. He says he thinks it’s important and points to row-hammer problems but doesn’t explain why. Probably because the audience it was written for already knows the arguments of why, but this is not the best argument. If in doubt, get ECC. Do your own research on how it works and why. This post won’t explain it, just will blame Intel (probably rightfully so).
- eloy 6y agoHe does explain it: > We have decades of odd random kernel oopses that could never be explained and were likely due to bad memory. And if it causes a kernel oops, I can guarantee that there are several orders of magnitude more cases where it just caused a bit-flip that just never ended up being so critical. It might be false, but I think it's a reasonable assumption.
- IgorPartola 6y agoTo someone on HN who isn’t familiar with what ECC does that explains nothing about how ECC works and how it could have prevented these situations. Or how often they really happen.
- simias 6y agoThe problem is that, if you don't have ECC to detect the errors, it's very hard to know what exactly caused a random, non-reproducible crash. Especially in kernel mode where there's little memory protection and basically any driver could be writing anywhere at any time. I can understand Linus's frustration from that point of view: without ECC RAM when you get some super weird crash report where some pointer got corrupted for no apparent reason you can't be sure if it's was just a random bitflip or if it's actually hiding a bigger problem.
- andi999 6y agoYou could run memtest on a pc without ecc for a couple of days and to estimate the error rate, or not?
- simias 6y agoI used to be pretty skeptical of ECC for consumer-grade hardware, mainly because I felt that I'd always prefer cheaper/more RAM over ECC RAM even if it meant that I'd get a couple of crash every year due to rogue bitflips. For servers it's a different story, but for a desktop I'm fine dealing with some instability for better performance. But these days with the RAM density being so high and bitflipping attacks being more than a theoretical threat it seems like there's really no good reason not to switch to ECC everywhere.
- fctorial 6y ago> cheaper/more RAM It's faster too.
- ekianjo 6y ago> no good reason not to switch to ECC everywhere. Not all CPUs support ECC however.
- loeg 6y ago(Intel)
- josefx 6y agoJust Intel fucking over security by making ECC a non feature on consumer grade hardware - wouldn't be surprised if it was just a single bit flipped in a feature mask.
- jjeaff 6y agoWell, with as common as a bunch of people in this thread seem to think bit flips are, it should just be a matter of time until that bit gets flipped on your cpu and activates the ecc feature.
- josefx 6y agoThat bit probably is either burned in or stored with the firmware in something more permanent than RAM. Modern RAM has the issue that it is optimized for capacity and speed to a point where state changes can leak into nearby bits.
- freeqaz 6y agoI bought ECC RAM for my laptop and it definitely was about 4x the price. It's valuable to me for a few reasons -- peace of mind being a big one. Bit flips happen and are real. I really wish ECC was plentiful and not brutally expensive!
- jjeaff 6y agoYou should be able to check logs for corrected errors, right? I'm guessing you won't find any.
- temac 6y agoNote that the price is mostly due to market segmentation, in your case most of it by the laptop vendor (of course some for Intel, but not that much compared to the laptop vendor) Xeon with ECC are not that overpriced compared with similar Core without. Likewise, RAM sticks with ECC are cheap to produce (basically just one more chip to populate per side per module). Likewise soldered RAM would simply add maybe $10 or $20 of extra chips.
- bitcharmer 6y agoThis is the first time I hear about a laptop that supports ECC memory. Could you please share the make and model?
- lb1lf 6y ago-My boss has a Xeon Dell - a 7550, methinks - luggable. It is filled to the gunwales with ECC RAM. Cost him the equivalent of $7k or so. Eeek.
- dijit 6y agoI have a Dell Precision 5520 (chassis of an XPS 15) which has a Xeon and ECC memory. Finding a memory upgrade seems difficult though.
- markonen 6y agoI was looking at getting the Xeon-based NUC recently and one of the reasons I decided against it was that ECC SO-DIMMs seem to be a really marginal product. If you want ECC, something that takes full-size DIMMs seems much easier to buy memory for.
- londons_explore 6y agoI simply care that my computer executes code perfectly. Let's settle on "one instance of unintended behaviour per hundred years" for that metric. If it needs ECC memory to do that, then fit it with ECC memory. If there are other ways to achieve that (for example deeper dram cells to be more robust to cosmic rays) that's fine too. Just meet the reliability spec - I don't care how.
- simias 6y agoThen you'll have to pay a huge primer for that privilege. I can assure you that your standard computer components are not rated for century-scale use. That's why I've always been on the fence with this ECC thing. For servers it's vital because you need stability and security. For desktops I think that for a long time it was fine without ECC. If I have to chose between having, say, 30% more RAM or avoid a potential crash once a year, I'll probably take the additional RAM. The problem is that now these problem can be exploited by malicious code instead of just merely happening because of cosmic rays. That's the main argument in favour of ECC IMO, the rest is just a tradeoff to consider.
- ClumsyPilot 6y agoBut it isn't just a crash, it's also silent data corruption that will never be detected
- dev_tty01 6y agoThis. How many user documents have memory flip errors introduced that are never detected? Impossible to say, but it is not a small number given the world-wide use of DRAM. Most are in trivial and unimportant documents, but some aren't...
- simias 6y agoIt can be a concern, that's true, but personally most of the stuff I edit end up checked into a git repository or something similar. And I mean, we all spend all day editing test messages and comments and files on non-ECC hardware, yet bitflip-induced corruption is rare enough that I can't say that I've witnessed a single instance of it in my life, despite spending a good chunk of it looking at screens. It's just not a problem that occurs in practice in my experience. If you're compiling the release build of a critical piece of software, you probably want ECC. If you're building the dev version of your webapp or writing an email to your boss, you'll probably survive without it.
- sys_64738 6y agoECC memory is predominantly used in servers where failure absolutely must be identified and logged. The desktop market to a lesser extent due to lack of mission critical tasks being run from there.
- dijit 6y agoThere are situations though, where you’re working on a document and the documents “save” format is a memory dump. Corruption for things of that type (Adobe RAW for example) would remove data. It might present itself as a 1pixel colour difference, but it could be more damaging (incorrect finances, in accounting software for example). Software trusts memory; but memory can lie. That’s dangerous.
- MaxBarraclough 6y agoThat's an interesting point. In an extreme case, an order or money transfer might be placed for an incorrect quantity, or to an incorrect recipient.
- KingMachiavelli 6y agoWell maybe. Rather than having to trust memory completely, it would just be better to use a binary format where each bit is verifiable so then at least a single bit flip would be immediately obvious. For example, a bit flip in a TLS session causes the whole session to fail rather than a random page element to change.
- mark-r 6y agoThat's the principle behind Gray Code counting: https://en.wikipedia.org/wiki/Gray_code https://en.wikipedia.org/wiki/Gray_code
- Dylan16807 6y agoNot really. Gray codes are designed so that if you're counting correctly, only one bit flips at a time. But if you flip the wrong bit by accident, you'll end up with a completely different number, no way to tell a problem happened. If you want to detect a bit flip, use parity.
- amelius 6y agoDoes Apple use ECC in its M1 laptop?
- alexwillner 6y agoAt least some kernel log messages imply that the M1 might support ECC: https://eclecticlight.co/2020/12/09/what-happens-when-an-m1-mac-starts-up/ https://eclecticlight.co/2020/12/09/what-happens-when-an-m1-...
- dijit 6y agoNo. It uses a unified package of LPDDR4x SDRAM
- my123 6y agoLPDDR4X systems with ECC exist, but it indeed looks like Apple M1 systems aren't one...
- graeme 6y agoThis is my one worry. I have an imac pro and anecdotally it has been a LOT more reliable than my old macbook pro. The imac pro has ecc.
- dijit 6y agoI beg this, every time this conversation comes up it’s the same answer “I don’t see a problem”. It’s so easy to chalk these kind of errors to other issues, a little corruption here, a running program goes bezerk there- could be a buggy program or a little accidental memory overwrite. Reboot will fix it. But I ran many thousands of physical machines, petabytes of RAM, I tracked memory flip errors and they were _common_; common even in: less dense memory, in thick metal enclosures surrounded by mesh. Where density and shielding impacts bitflips a lot. My own experience tracking bitflips across my fleet led me to buy a Xeon laptop with ECC memory (precision 5520) and it has (anecdotally) been significantly more reliable than my desktop.
- lighttower 6y agoCan you get decent battery life with this ecc memory in a laptop?
- dijit 6y agoYes. ECC memory uses only marginally more power than non-ECC memory. And memory isn’t the largest consumer of battery life by a country mile. Screen, Wi-Fi, and to a much lesser extent (unless under load) the CPU are the most major culprits of low battery life.
- indolering 6y agoIt can actually reduce power consumption, because refresh rates don't need to be so high: https://media-www.micron.com/-/media/client/global/documents/products/white-paper/ecc_for_mobile_devices_white_paper.pdf?la=en-in&rev=b892d7cdf84a495d83b2560b2fe523ef https://media-www.micron.com/-/media/client/global/documents...
- loeg 6y agoYeah, it's real obnoxious of Intel to silo ECC support off into the Xeon line, isn't it? I switched to ECC memory in 2013 or 2014 with a Xeon E3 (fundamentally a Core i7 without the ECC support fused off) and of course a Xeon-supporting motherboard (with weird "server board" quirks: e.g., no on-board sound device). I love that AMD doesn't intentionally break ECC on its consumer desktop platforms and upgraded to the Threadripper in 2017.
- phh 6y agoI don't know if ECC is that important, but reliability of RAM (or any storage) feels pretty crazy to me. 128GB being refreshed every second for a month error requires that the per-bit refresh process has a reliability of 99.9999999999999999% to be flawless. Considering we are dealing with quantum effects (which are inherently probabilistic), I wouldn't trust myself to design anything like that. Now back to ECC, I'll probably be corrected, but I don't think ECC helps gain more than two order of magnitudes, so we still need incredibly reliable RAM. If we move to ECC RAM by default everywhere, aren't we simply going to get less reliable RAM at the end?
- bitcharmer 6y agoA system on Earth, at sea level, with 4 GB of RAM has a 96% percent chance of having a bit error in three days without ECC RAM. With ECC RAM, that goes down to 1.67e-10 or about one chance in six billions. So I'd say ECC is not only important but insanely impactful. There's a reason why many organizations don't even want to hear about getting rigs with non-ECC memory.
- tomxor 6y agoI like when people back up their claims with numbers, but would you mind describing roughly what that 96% probability of error is based upon? I understand altitude has some kind of proportionality to cosmic ray exposure, and number of bits will multiply the probability of an error.. I'm presuming there is also an inherent error rate to DRAM separate from environment. But what are those numbers.
- bitcharmer 6y agoApologies, you're totally right. I should have linked to the source: http://lambda-diode.com/opinion/ecc-memory#:~:text=A%20system%20on%20Earth%2C%20at,one%20chance%20in%20six%20billions http://lambda-diode.com/opinion/ecc-memory#:~:text=A%20syste....
- tomxor 6y ago
- dboreham 6y agoYou don't need to look at kernel crashes to speculate about bus and memory errors -- just check the logs on a few systems that do have ecc. Pretty soon you'll see correctable errors being reported.
- maddyboo 6y agoI don’t know much about this topic, but is it possible that ECC memory is more prone to single bit errors than non-ECC memory because there is less pressure on companies to minimize such errors? If this were the case, it would skew the data.
- justin66 6y agoThere are 12.5% more memory cells for a given module size, which equals more targets to possibly be flipped by cosmic rays. It’s not crazy to think that modules of equivalent quality (same brand, same chip part numbers) would experience a greater incidence of that kind of single bit flip (which would be corrected on the ECC modules). If a manufacturer were shipping chips prone to bit flipping because of slightly radioactive packaging, as happened at times in the past, you might see something similar. But you’ve got it backwards about the incentives. A manufacturer has less incentive to deliberately ship a defective part in the case of ECC modules. If the modules consistently log ECC errors, they can easily be identified and returned under warranty to the manufacturer. A consumer is much less likely to identify an intermittent problem with a non-ECC part.
- 1996 6y agoLinus is absolutely right. I am trying to get a laptop with dual NVMe (for ZFS) and ECC RAM. I can't get that, at all - even without the other fancy things I would like such as a 4k OLED with pen/touchscreen. In 2020, even the Dell XPS stopped shipping OLED (goodbye dear 7390!) I will gladly give my money to anyone who sells AMD laptop with ECC. Hopefully, it will show there's demand for "high end yet non bulky laptops"
- miahi 6y agoLenovo P53 has 3 NVMe slots, 4k OLED with touchscreen (and optional pen) and up to 128GB ECC RAM if you choose the Xeon processor. It's big and heavy, but it exists. I hope AMD will create a better market for the ECC laptop memory (right now it's hard to find + expensive).
- 1996 6y agoI know- I had my eye on this very model, as you can even add a mSata on the WWAN slot to get a 4th drive. Unfortunately, Lenovo is not selling the P53 anymore, which is exactly why I say I can't get that even in a "bulky" version.
- nix23 6y agoI always have that conversation when ZFS comes up. Some peoples think ZFS NEEDS ECC, but in fact ZFS needs ECC much as every single one FS in Linux. And every single reliable Machine needs ECC.
- paulie_a 6y agoThere was a great defcon talk a while back regarding using ECC. The concept was called "dns jitter" Basically you can register domains using small bit differences for domains and start getting email and such for that domain If I recall correctly the example given was a variation of microsoft.com All because so much equipment doesn't use ECC
- zx2c4 6y agoVoila http://media.blackhat.com/bh-us-11/Dinaburg/BH_US_11_Dinaburg_Bitsquatting_WP.pdf http://media.blackhat.com/bh-us-11/Dinaburg/BH_US_11_Dinabur...
- tyoma 6y agoThere were some great follow up talks as well! It turns out a viable attack vector was also MX records. And there was the guy who registered kremlin.re ( versus kremlin.ru ).
- jeffbee 6y agomiclosoft.com is only one bit away from microsoft.com. Used to see these problems all the time when I worked on gmail. At Google even with ECC everywhere there wasn't enough systematic error detection and correction to prevent the global database of monitoring metrics from filling up with garbage. /rpc/server/count was supposed to exist but also in there would be /lpc/server/count and /rpc/sdrver/count and every other thing. Reminded me daily of the terrors of flipped bits.
- thu2111 6y agoAhaha. Reminds me of when I worked there. One day a large service tanked in some datacenter because BigTable replication in that location just stopped. Digging in, it turned out the BigTable should have been replicating from YQ but had started trying to use QQ instead, which didn't exist. Q being one bit away from Y. Or it was something like that, I don't remember exactly. There'd been a bit flip in the exact part of memory that contained the name of the database cluster to replicate from!
- deleted 6y ago[deleted]
- louwrentius 6y agoECC matters, even on the desktop, it's not even a discussion, to me. If you think it doesn't matter: how do you know? If you don't run with ECC memory, you'll never know if memory was corrupted (and recovered). That blue screen, that sudden reboot, that program crashing. That corrupted picture of your kid. Who knows. I'll tell you, who knows. God damn every sysadmin (or the modern equivalent) can tell you how often they get ECC errors. And at even a small scale you'll encounter them. I have, on servers and even on an SAN Storage controller, for crying out loud. If you care about your data, use ECC memory in your computers.
- supernovae 6y agoI've got nearly 30 years of experience and not once has non ECC memory lead to corruption. Maybe a crash, maybe a panic, maybe a kernel dump... But.. in all my time operating servers over 3 decades, it's always been bad drivers, bad code and problematic hardware that's caused most of my headaches. Have i seen ECC error correction in logs? yeah.. I don't advocate against it but, i've found for most people you design around multiple failure scenarios more than you design around preventing specific ones. Take the average web app - you run it on 10 commodity systems and distribute the load.. if one crashes, so what. Chances are, a node will crash for many more reasons other than memory issues. If you have an app that requires massive amounts of ram or you do put all of your begs in one basket, then ECC makes sense... I just know i like going horizontal and I avoid vertical monoliths.
- ptx 6y ago> if one crashes, so what Crashes might not matter, but silent data corruption does. The owner/user of that data will care when they eventually discover that it at some point mysteriously got corrupted.
- louwrentius 6y agoThe problem with memory corruption is not just crashes, those are the more benign outcomes. The real killer is data corruption. Houw would you even begin to know that data is corrupted until it is too late?
- spacedcowboy 6y agoSeems likely that “bad ram” was the reason for the recent AT&T fiber issues, given that 1 bit was being flipped reliably in data packets [1] [1]: https://twitter.com/catfish_man/status/1335373029245775872?lang=en https://twitter.com/catfish_man/status/1335373029245775872?l...
- SV_BubbleTime 6y agoI think you meant seems unlikely
- p_l 6y agoI have had in the past encountered an issue where line card was stripping exactly one bit of address data. Don't know of the follow up investigation, but it probably wasn't TCAM
- type0 6y agoConsumer awareness about ECC needs to be better, with recent security implications I simply can't understand why more motherboard manufacturers don't support it on AMD. Intel of course is all to blame on the blue side, I stopped buying their overpriced Xeons because of this.
- rajesh-s 6y agoGood point on the need for awareness! The industry has convinced the average user of consumer hardware that PPA (Power,Performance,Area) is all that needs to get better with generational improvements. Hoping that the concerning aspects of security and reliability that have come to light in the recent past changes this.
- z3t4 6y agoMemory often comes with lifetime guarantees. If they had ECC it would be much easier to detect bad memory...
- jkuria 6y agoFor those, like me, wondering what ECC is, here's an explanation: https://www.tomshardware.com/reviews/ecc-memory-ram-glossary-definition,6013.html https://www.tomshardware.com/reviews/ecc-memory-ram-glossary...
- petermcneeley 6y agoI would also add that Row Hammer Attacks are much harder on ECC. When I first tried to replicate the row hammer attack I was not getting any results. Turns out I was doing this on ECC. On non ECC memory the same test easily replicated the row hammer attack. https://en.wikipedia.org/wiki/Row_hammer https://en.wikipedia.org/wiki/Row_hammer
- zdw 6y agoGood news is that for DDR5, ECC is a required part of the spec and should be a feature of every module: https://www.anandtech.com/show/15912/ddr5-specification-released-setting-the-stage-for-ddr56400-and-beyond https://www.anandtech.com/show/15912/ddr5-specification-rele...
- bradfa 6y agoI read it to say that on die ecc is recommended but that dimm-wide ecc is still optional. And now you have 8 bits of ecc per 32 data versus older DDR having 8 bits of ecc per 64 data. Hence the cost for dimm-wide ecc is going up.
- toast0 6y agoOn die ECC is great for increasing reliability, if all else is equal, but if it doesn't report to the memory controller, and if the memory controller doesn't report to the OS, I think it will be worse than status quo, because all else won't be equal. With no feedback, systems are going to continue to run on the edge, but now detectable failures will all be multi-bit; because single bit errors are hidden.
- cududa 6y agoHuh? Why would the memory controller not be updated accordingly? Also I have no idea about Linux or Mac, but Windows has had ECC support and active management for decades?
- indolering 6y agoIt's part of the firmware first trend of fixing things at the firmware level before reporting problems up the stack. This makes it a real nightmare for systems integrators to do root cause analysis.
- mlyle 6y agoNormally, ECC has meant just the DIMM stores some extra bits, and the memory controller itself implements ECC-- writing the extra parity, and recovering when errors emerge (and halting when non-recoverable errors happen). DDR5 includes on-die ECC, where the RAM fixes the errors before sending them over the memory bus. This means if the bus between the processor and ram corrupts the bits-- tough luck, they're still corrupted. And it's unclear whether we're going to get the quality of memory error reporting that we're used to or get the desired halt-on-non-recoverable error behavior (I've not been able to obtain/read the DDR5 specification as yet).
- cbanek 6y agoAs someone who has had to read thousands of random game crash reports from all over the interwebs (you know when Windows says you might want to send that crash log? like that), I totally agree. Of all the things to be worried about, like OS bugs, bad hardware configuration, etc. bad memory is one of those really troubling things. You look at the code and say "it's can't make it here, because this was set" but when you can't trust your memory you can't trust anything. And as the timeline goes to infinity, you may also get one of these reports and be asked to fix it... good luck.
- lighttower 6y agoSomeone reads those reports!?! Wow, how do I write them to ensure someone who reads them takes them seriously?
- xmodem 6y agoThe best way is to submit the same crash report from thousands of different locations, repeatedly
- apankrat 6y agoAye. I have an assert in the code that fronts a very pedantic test of the context. In all cases when this assert was tripped (and reported) an overnight memtest86 test surfaced RAM issues. - Edit - Also, bit flips in the non-ECC memory are _the_ cause of the "bitrot" phenomenon. That is when you write out X to a storage device, but you get Y when you read it back. A common explanation is that the corruption happens _at rest_. However all drives from the last 30+ years have FEC support, so in reality the only way a bit rot can happen is if the data is damaged _in transit_, while in RAM, on the way to/from the storage media. So, if you ever decide if to get an ECC RAM, get it. It's very much worth it.
- srtjstjsj 6y agoBitrot in human memory is the same. Memories change during the process of recalling them, not while they are in "storage".
- otterley 6y agoD. J. Bernstein (of qmail/daemontools fame) spoke of it over a decade ago as well. https://cr.yp.to/hardware/ecc.html https://cr.yp.to/hardware/ecc.html
- slim 6y agothese days he's more famous for the NaCl crypto library
- loup-vaillant 6y agoFor which bit flips are even more relevant: EdDSA has this nasty tendency of leaking the private key if the wrong bits are flipped (there are papers on fault injection attacks). People who sign lots of stuff all the time, say Let's Encrypt, could conceivably gain some piece of mind with ECC. (Note: EdDSA is still much much better than ECDSA, most notably because it's easier to implement correctly.)
- otterley 6y agoAbout 1/3 of Google's machines and 8% of Google's DIMMs in their fleet suffer at least one correctible memory error per year: http://www.cs.toronto.edu/~bianca/papers/sigmetrics09.pdf http://www.cs.toronto.edu/~bianca/papers/sigmetrics09.pdf
- jjeaff 6y agoWhich means, assuming google is running very large machines with lots of memory that one might expect a single correctable error once every 6-10 years on your average workstation of small server. That's generously assuming your workstation has 1/3 as much memory as the average google server.
- Nebasuke 6y agoGoogle does not use very large or even large machines for most of their fleet. You can quickly see in the paper this is for 1, 2, and 4 GB RAM machines (in 2006-2008).
- otterley 6y agoThese were state of the art 13 years ago. It’s not safe to extrapolate from this paper that they aren’t using servers having significantly more memory today. Thirteen years is more than 2 full depreciation intervals.
- tpetry 6y agoWith a single bit flip on 8% of the dimms you only need 12.5 dimms in your workstation to have one bit flip every year. Not everyone has that much dimms, but at least 4 is pretty normal. So in average every 3 years for every workstation. But i don‘t know how relevant these metrics from 2009 are. Did memory got better or worse compared to 2009 for bit flips?
- JoeAltmaier 6y agoECC works if done right. Accessing a memory location can fix bit-flips (ECC is a 'correcting' code). But systems that don't regularly visit every memory location, can accumulate risk. Those dark corners of RAM can eventually get double-bit errors and be uncorrectable. So an OS might 'wash' RAM during idle moments, reading every location in a round-robin manner to get ECC to kick in and auto-correct. Doesn't matter how fast (1M every hour or whatever) as long as somehow ECC has a chance to work.
- jacquesm 6y agoInteresting, similar to scrubbing raid arrays. How often do those double bitflips appear though? You'd have to have a pretty long running server for that to be a problem, no?
- jeffbee 6y agoAccording to Google's old paper on the subject, about 1% of their machines suffered from an uncorrectable (i.e. multi-bit) error in a year.
- musingsole 6y agoA double-bit error in many cases is fine. If the error is at least detectable at the time of a read, your protection worked. What's scary is a triple-flip event. Most of those will still look like corrupted data, but if it happens to flip into looking like a fixable, single-bit error, you're out of luck and won't even know it.
- a1369209993 6y ago> Most of those will still look like corrupted data, Not if you're using a typical 72-bit SECDED code[0]. You have two error indicators: a summary parity bit (even number of errors: 0,2,etc vs odd number of errors: 1,etc), and a error index: 0 for no errors, or the bitwise xor of the locations each bit error. For a triple error at bits a,b, and c, you'll have summary parity of 1 (odd number of errors, assumed to be 1), and a error index of a^b^c, in the range 0..127, of which 0..71[1] (56.25%, a clear albeit not overwhelming majority) will correspond to legitimate single-bit errors. 0: https://en.wikipedia.org/wiki/Hamming_code#Hamming_codes_with_additional_parity_(SECDED) https://en.wikipedia.org/wiki/Hamming_code#Hamming_codes_wit... 1: or 72 out of 128 anyway; the active bits might not all be assigned contiguous indexes starting from zero, but it doesn't change the probability and it's simpler to analyse if summary is bit 0 and index bit i is substrate bit 2^i.
- rahimiali 6y agoI have trouble parsing information from this rant. Is someone willing to translate this into an argument (a string of facts tied by logical steps)?
- mark-r 6y ago1. Linux sometimes has crashes, not due to software errors but because of memory glitches. 2. ECC would prevent memory glitches. 3. ECC is hard to find on desktop PCs because Intel uses the feature to differentiate desktop CPUs from server CPUs, so it can charge more for servers. 4. Even when someone like AMD makes the feature available, the market doesn't have ECC DRAM modules or motherboards readily available because Intel killed the demand for it.
- johnklos 6y agoFrom the fortune database: As far as we know, our computer has never had an undetected error. -- Weisert
- mauri870 6y agoIn case the page os not loading, refer to the wayback machine[1] for a copy [1] https://web.archive.org/web/*/https://www.realworldtech.com/forum/?threadid=198497&curpostid=198647 https://web.archive.org/web/*/https://www.realworldtech.com/...
- musingsole 6y agoIt's a shame we don't have ECC for individuals. How many of society's bugs come from someone wandering around with a bit flipped?
- FartyMcFarter 6y agoDoes anyone know why ECC memory requires the CPU to support it? Naively, I can understand why error reporting has dependencies on other parts of the system, but it would seem possible for error correction to work transparently.
- TomVDB 6y agoI think the memory just provides additional storage bits to detect the issue, but doesn't contain the logic. This is in line with all technical parameters of DRAM: everything must be as cheap as possible, and all the difficult parts are moved to the memory controller. Which is the right thing to do, because you can share one memory controller with multiple DRAM chips.
- wmf 6y agoHistorically the detection and correction is performed in the memory controller not the DRAM.
- toast0 6y agoAs implemented today, ECC is a feature of the memory controller. You need special ram, because instead of 8 parallel rams per bank, you need 9, and all the extra data lines to go to the controller. Modern CPUs have integrated memory controllers, so that's why the CPU needs to support it. Correction without reporting isn't great; anyway, you need a reporting mechanism for uncorrectable errors, or all you've done is ensure any memory errors you do experience are worse.
- fomine3 6y agoError correcting and reporting is better, but even only correcting is better than non-ECC. I wonder this compromise could be accepted by Intel.
- deleted 6y ago[deleted]
- vlovich123 6y agoA couple of years ago there was advancements that claimed to make Rowhammer work on ECC RAM even with DDR4 [1]. Is that no longer a concern for some reason? I would think the only guaranteed solutions to Rowhammer are actually cryptographic digests and/or guard pages. [1] https://www.zdnet.com/article/rowhammer-attacks-can-now-bypass-ecc-memory-protections/ https://www.zdnet.com/article/rowhammer-attacks-can-now-bypa...
- theevilsharpie 6y agoECC isn't a direct mitigation against Rowhammer attacks, as memory errors caused by three or more flipped bits would still go undetected (unless you're using ChipKill, but that's a rare setup). However, flipped three bits simultaneously isn't trivial, and the attempts that flip fewer bits will be detected and logged.
- rajesh-s 6y agoRight! Section 1.3 of this publication discusses possible mitigations for the row hammer problem and where ECC fits in https://users.ece.cmu.edu/~omutlu/pub/rowhammer-summary.pdf https://users.ece.cmu.edu/~omutlu/pub/rowhammer-summary.pdf
- GregarianChild 6y agoThe paper you cite is from 2014 and the mitigations discussed there have all been circumvented. [1] is from 2020 and a better read for Rowhammer mitigation. [1] J. S. Kim et al, Revisiting RowHammer: An Experimental Analysis of Modern DRAM Devices and Mitigation Techniques https://arxiv.org/abs/2005.13121 https://arxiv.org/abs/2005.13121
- rajesh-s 6y agoThanks for pointing that out!
- GregarianChild 6y ago
- KingMachiavelli 6y agoIs there such a thing as 'software' ECC where a segment in memory also has a checksum stored in memory and the CPU just verifies it when the memory segment is accessed? It would be a lot slower than real ECC but it could just be used for operations that would be especially vulnerable to bit flips. It would also not know for certain if the memory segment of data or the memory segment holding the checksum was corrupted besides their relative sizes (checksum is much smaller so more unlikely to have had a bit flip in it's memory region).
- a1369209993 6y agoActually... there is a word of memory that you already have to read every time you access a region of memory: the page table entry for that region. If you have 64-byte cache lines, that's 64 lines per (4KB) page, so you could load a second 64-bit word from the page table[0], and use that as a parity bit for each cache line, storing it back on write the same way you store active and dirty bits in the PTE proper. Actual E[correcting]C would require inflating the effective PTEs from 8(orginal)-16(parity) bytes to about 64(7 bits per line, insufficient)-128(15, excessive), which is probably untenable, but you could at least get parity checks this way. There's also the obvious tactic of just storing every logical 64-bit word as 128 bits of physical memory, which gives you room for all kinds of crap[1], at the expense of halving your effective memory and memory bandwidth. 0: This is extremely cheap since you're loading a 64- vs 128-bit value, with no extra round trip time and still fits in a cache line, so you're likely just paying extra memory use from larger page tables. 1: Offhand, I think you could fit triple or even quadruple error correction into that kind of space (there's room for eight layers of SECDED, but I don't remember how well bit-level ECC scales).
- temac 6y agoIntel has some recent patents on that.
- MarkusWandel 6y agoThis is one justified Linus rant! My personal history includes data loss twice because of defective RAM, and many more RAMs discarded after the now obligatory overnight run of MemTest86+ (these were all secondhand RAMs - I would never buy a new one without a refund guarantee). My very first "PC" still had the ECC capability and I used it. My own now very dated rant on the subject: http://wandel.ca/homepage/memory_rant.html http://wandel.ca/homepage/memory_rant.html
- mixmastamyk 6y agoA few years back memtest86 wouldn’t run on newer machines, has that been fixed?
- MarkusWandel 6y agoWouldn't know, I don't run newer machines. But since it's a boot option on Fedora disks, I imagine it would run.
- salmon 6y agoYou bought used RAM DIMMs and were surprised that they failed?
- MarkusWandel 6y agoUsed computers that have RAM in them. But as I wrote, two of those computers were brand new with new RAMs in them.
- arendtio 6y agoIt would be interesting to see how many more kernel oops appear on machines without ECC compared to those with ECC.
- aborsy 6y agoFor the average user, what’s the impact of bit flips in memory in practical terms? I am not talking about servers dealing with critical data. Suppose that I maintain a repository (documents, audio and video), one copy in a ZFS-ECC system and one in an ext4-nonECC system. Would I notice a difference between these two copies after 5-10 years? That tells us if ECC matters for most people.
- throwaway9870 6y agoThis isn't about disk storage, this is about DRAM. A bit flip in DRAM might corrupt data, but could also cause random crashes and system hangs. That generally matters to everyone.
- deleted 6y ago[deleted]
- theevilsharpie 6y ago> For the average user, what’s the impact of bit flips in memory in practical terms? The most likely impact (other than nothing, if bits are flipped in unused memory) is program crashes or system lock-ups for no apparent reason.
- JumpCrisscross 6y agoWhat is the status of ECC on Macs?
- CalChris 6y agoiMac Pro which has Xeon M. There's a good chance that will go away with the new Apple Silicon iMac Pro due out this year. MacRumors roundup article doesn't mention ECC. https://www.macrumors.com/roundup/imac/ https://www.macrumors.com/roundup/imac/
- MAXPOOL 6y agoWell shit. I run some large ML models in my home PC and I get NaN's and some out of range floats every month or so. I have spent hours debugging but doing the same computation with the same random seeds does not recreate the problem. How about GPU's and their GDDR SDRAM? Do they have parity bits?
- deleted 6y ago[deleted]
- layer8 6y agoSome pro-level Nvidia GPUs have ECC RAM, they are very expensive though. I don’t think regular gaming GPUs have parity, due to the extra cost, performance impact (probably minor but measurable) and irrelevance for gaming.
- vbezhenar 6y agoCheap pro-level GPUs don't have ECC RAM either. And it's not easy to find out, it might be buried somewhere.
- ratiolat 6y agoI have: Asus PRIME A520M-K Motherboard 2x M391A2K43DB1-CVF (Samsung 16GiB ECC Unbuffered RAM) AMD Ryzen 5 3600 I specifically was looking for bang for buck, low(er) wattage and ECC.
- IanCutress 6y agoThose AMD motherboards with consumer CPUs are a bit iffy. They run ECC memory, but it's hard to tell if it is running in ECC mode. Even some of the tools that identify ECC is running will say it is, even when it isn't, because the motherboard will report it is, even when it isn't. ECC isn't a qualified metric on the consumer boards, hence all the confusion.
- unixhero 6y agoFantastic burn by Linus Torvalds whom also had some skin in the CPU game. Offtopic, I wonder if he trawls that site regularly. And eventually I wonder, is he here also? :)
- knorker 6y agoI have multiple times postponed buying new computers for YEARS, because I'm waiting for intel to get their head out of their ass and actually let me buy something that does ECC for desktop. (incl laptops) I would have bought computers when I "wanted one". Now I buy them when I need one. Because buying a non-ECC computer just feels like buying a defective product. In the last 10 years I would have bought TWICE as many computers if they hadn't segmented their market. Fuck intel. I sense that Linus self-censored himself in this post, and like me is even angrier than the text implies.
- vbezhenar 6y agoThere are plenty of Xeons which are suitable for desktops and there are plenty of laptops with Xeons. Price is not nice though.
- skibbityboop 6y agoHave you finally stopped buying Intel? Current Ryzens are a much better CPU anyhow, just dump Intel and be happy with your ECC and everything else.
- knorker 6y agoI'm in the market for a new laptop (since a few years). Is there something like the X1 carbon but with ECC?
- nostrademons 6y agoI still remember Craig Silverstein being asked what his biggest mistake at Google was and him answering "Not pushing for ECC memory." Google's initial strategy (c. 2000) around this was to save a few bucks on hardware, get non-ECC memory, and then compensate for it in software. It turns out this is a terrible idea, because if you can't count on memory being robust against cosmic rays, you also can't count on the software being stored in that memory being robust against cosmic rays. And when you have thousands of machines with petabytes of RAM, those bitflips do happen. Google wasted many man-years tracking down corrupted GFS files and index shards before they finally bit the bullet and just paid for ECC.
- tyoma 6y agoFigure this is as good of a time as any to ask this: There are many various DRAMs in a server (say, for disk cache). Has Google or anyone who operates at a similar scale seen single bit errors in these components?
- deleted 6y ago[deleted]
- gh02t 6y agoThe supercomputing community has looked at some of the effect on different parts of the GPU. https://ieeexplore.ieee.org/abstract/document/7056044 https://ieeexplore.ieee.org/abstract/document/7056044
- bsder 6y agoThis is as old as computing and predates Google. When America Online was buying EV6 servers as fast as DEC could produce them, they used to see about about 1 double bit error per day across their server farm that would reboot the whole machine. DRAM has only gotten worse--not better.
- sitkack 6y agoYes. Bit flips (for all reasons) occur in buses, registers, caches, etc. Anything that has state can have state changed incorrectly. This is why filesystems like ZFS exist and storage formats have pervasive checksums.
- elgfare 6y agoFor those out of the loop like me, ECC does indeed stand for error correcting code. https://en.m.wikipedia.org/wiki/ECC_memory https://en.m.wikipedia.org/wiki/ECC_memory
- wicket 6y agoOver the years, I don't think I've ever been able to explain to anyone that their memory error could have been caused a cosmic ray without being laughed at.
- jhoechtl 6y agoI definitely do not want Linus Torvalds yelling at me in that tone --- but reading his utterings is certainly entertaining.
- rafaelturk 6y agoLittle bit offtopic: Again seems that Intel? what?! is the one lowering the bar.
- belzebalex 6y agoAsked myself, would it be possible to build a Geiger counter with RAM?
- _0ffh 6y agoPlease someone correct me if I'm wrong, but as far as I can remember memory with extra capacity for error detection used to be a rather common thing on early PCs. That really only changed a couple of decades in, in order to be able to offer lower prices to home users who didn't know or care about the difference. Probably about the time, or earlier, when with some hard disk manufacturers megabytes suddenly shrunk to 10^6 bytes (before kibibytes or mebibytes where a thing, btw).
- qwerty456127 6y agoECC should be everywhere. It seems outrageous to me almost no laptops have ECC.