19 ms·
Exploiting System Management Mode with a very long interrupt
- nazgulsenpai 1mo agoI'm amused at the lengths the readme goes to in order to drive home the fact that this needs to be a LOOOOOOOOOOOOOOOOOOOONG instruction, including the unnecessarily long code block illustration. The topic is interesting anyway, but that makes it way more entertaining.
- BadBadJellyBean 1mo agoDo you think a short instruction is okay or does it need to be long? The instructions were a bit unclear in that regard :D
- nazgulsenpai 1mo agoOnly if the short instruction is incredibly long.
- deleted 1mo ago[deleted]
- londons_explore 1mo agoUnclear why there is a 1 second timeout at all. Presumably the patch for that will be to make it an infinity timeout.
- xxpor 1mo agoCan this be patched? Is there a chance it's a hw watchdog that you can't fix in microcode?
- ramses0 1mo agoLooks like it's ~4 billion (2^32) crossover counter?
- toast0 1mo agosystem management mode does a lot of stuff, some of which is time critical. If your system is overheating and one of the cores is stuck off in the weeds, it's probably better to get on with the thermal response rather than waiting forever. Also, the System Management Interrupts are supposed to return to normal processing in some finite timespan; a timeout bounds the wait time.
- PunchyHamster 1mo agoIf it is critical it should not be running on same cores
- Geezus_42 1mo agoUsing the example above, if a CPU core is overheating, can you down clock that core using and instruction run on another core? I don't actually know that much about how the hardware actually works at that level, so I am genuinely asking.
- userbinator 1mo agoIIRC clock control doesn't need to be done in SMM, especially PROCHOT which is a hardwired thermal shutdown. That said, power management is one of the things SMM was originally designed for, so it may be used for some of that.
- inigyou 1mo agoThere's a hardwired emergency shutoff but you probably want the BIOS to set the fans to maximum long before the computer just shuts off.
- quotemstr 1mo agoIt could also react to hitting the timeout with a hard reset. Annoying perhaps, but at least safe. Ancient principle of system design is that when you must fail, it is better to fail safe than fail deadly even when it's annoying in the short term.
- justusthane 1mo agoThe author’s take on this in the Mitigations section makes sense to me: > Remove the timeout, and a legitimately stuck core hangs the platform on the first SMI. Increase the timeout, and you kill performance on many-core platforms that are forced to quiesce all cores every SMM entry. It's not clear what the best path forward is, or if there is even a path forward at all.
- inigyou 1mo agoReads like AI.
- justusthane 1mo agoNot to me, but who knows
- cryptonector 1mo agoI would expect a way to interrupt super-long-running instructions would be the better option, even if it was not fully backwards-compatible (say your process executing long-running instructions gets killed).
- kmeisthax 1mo ago...huh, I was wondering why serial machine code prankster xoreaxeaxeax was keeping lists of extremely long-running instructions. Hopefully this is at least only possible in kernel mode, right? Right?!
- xxpor 1mo agoMaybe with vfio/igb_uio/uio_pci_generic? Still root level access.
- tptacek 1mo agoIs it really a long running instruction? I mean, obviously yes, but what makes it slow is that it's doing an MMIO copy from a slow source. It's like a read(2) system call being "slow" because the fd is associated with a socket to the moon.
- tuetuopay 1mo agoIt's an instruction in the sense that timing boundaries are x86 instruction boundaries, which is what the security model bases itself on. So yeah, not an instruction in the strict CPU sense (microcode + micro-ops), but in the useful sense.
- kmeisthax 1mo agoA read that happens to touch a particular torment nexus fd is still a long-running syscall, even if the syscall servicing routine itself is not long-running. The underlying problem is that program code that is "in a syscall" or "in an instruction" is in a special state for which interruption might not be possible or implemented well[0]. [0] Remember ITS and the PC2 problem?
- touisteur 1mo agoI have a reproducible way to have a pwrite syscall on a specific SSD on a specific machine take 15+ seconds and completely block any syscall related to that SSD by any other thread or core during that amount of time. I tried and couldn't preempt it either (sched_fifo and preempt kernel options). I should have a look soon with Intel PT to check whether it's on the same instruction every time :)
- mike_hearn 1mo agoThe designers of the firmware anticipate this attack but punt it to the vendor, apparently: // // Platform implementor should choose a timeout value appropriately: [snip] // - The timeout value must be longer than longest possible IO operation in the system
- Liftyee 1mo agoI don't know much about the specifics of CPU architecture apart from the existence of assembly and different modes. Either way the explanation was still entertaining and interesting. smiiiiiiii
- hyperhello 1mo agoSMM calls for a timeout because it wants everything to be between instructions pro forma. So there’s a very long instruction on a core, but after it completes, the core does stop, right? It seems like to make this into an attack you’d have to a very long instruction that also somehow interacts with the thing the SMM is doing, while it’s doing it.
- mirashii 1mo agoIf I’ve understood correctly, what your missing here is that the first core in SMM tells the second to join it in SMM, times out on the wait, does its thing and exits, but then the second core joins SMM after the first has exited, so now the first core is running outside SMM, second core in SMM, so first core can attack the second.
- Hyperlisk 1mo agoRelated repo from them, mentioned in the readme as well: https://github.com/xoreaxeaxeax/asm-hall-of-shame https://github.com/xoreaxeaxeax/asm-hall-of-shame > Instruction latency analysis usually focuses on performance optimization—making code run as fast as possible. The Assembly Hall of Shame takes the opposite approach: searching for the absolute floor of single-instruction performance. Fun stuff!
- inigyou 1mo agoAnd already posted to HN a few days ago which is surely where this post originated from.
- PunchyHamster 1mo agoIt's nice to see SMM is as terrible idea now as it was at moment of conception. All coz they can't be arsed to put a tiny management core separate from the rest and save a penny
- quotemstr 1mo agoARM has EL3, which is basically the same thing. There's nothing inherently wrong with the CPU having multiple privilege levels. The problem with SMM has always been its user-hostile opaque implementation, not that the technical mechanism exists.
- inigyou 1mo agoTo be fair it originates from 386. Cores didn't get much smaller and you don't want your CPU to be 1.5 times as expensive.
- cheschire 1mo agoAlmost nothing from this GitHub profile posted until the last four days. From a meta perspective what is going on? What am I missing? Why is this GitHub profile suddenly getting massive attention and making front page so frequently?
- _def 1mo agoThis rosenbridge repos commits claim to be 8 years old https://github.com/xoreaxeaxeax/rosenbridge https://github.com/xoreaxeaxeax/rosenbridge
- cheschire 1mo agoYep there are links as far back as 11 years ago posted here. But I’m saying why suddenly in four days is this GitHub profile linked in lots of front page threads? Is it just that one thread brought attention and several people are slowly digesting the other repos on that profile? Or is there another meta reason?
- bri3d 1mo agoHe gave a talk at Defcon.
- ironhaven 1mo agoDEFCON the fun hacking conference in las Vegas that happened last weekend.
- dnautics 1mo agoxoreax has a famous video showing how to find hidden x86 secret instructions that... who knows who asked the manufacturer to put there. the story of the hardware setup (pxe booted via terminal pos iirc) to find it is epic because certain opcode faults could brick the machine under normal automation conditions so each pos had to be monitored and have its physical on/off tapped to do a "manual" hard reboot
- rft 1mo agoChris just has some fun projects. Sandsifter was pretty well known back then due to finding that VIA x86 opcode, the movfuscator is just an amazing piece of mostly useless engineering and the fun reverse engineering psychological warfare was my first contact with his projects. I guess someone stumbled on one of his projects and others clicked through to his other projects. Many of them are a perfect fit for HN, no wonder they got posted.
- quotemstr 1mo ago> The code waits for all cores to enter SMM, or for up to 1 second, whichever occurs first. See, this is why the mantra that all blocking operations should have a timeout is stupid and short-sighted no matter how many times junior devs and AIs bleat it in code review. Continuing after arbitrary timeouts usually violates invariants, and failing after arbitrary timeouts introduces hard-to-debug failures under load. Better for the system to hang so you can debug it --- and maybe reboot as a whole via a watchdog --- than for the code to say "Oh, this operation is supposed to be done after one second, but isn't. Situation normal, everything fine. We continue." No. That situation is very much not fine.
- deleted 1mo ago[deleted]
- wtallis 1mo ago> failing after arbitrary timeouts introduces hard-to-debug failures under load What needs to fail here is the instruction doing insanely slow MMIO. That's not going to be too hard to debug; none of the examples of suitably slow instructions are anywhere close to reasonable, and a fault on a vmovdqu in MMIO address space is a big red flag. And this attack requires enough ridiculous behavior from coordinating software beyond just the single super-slow instruction that it's hard to imagine any reasonable workload being affected if this case starts causing a fault.
- neerajsi 1mo agoI initially agreed with your idea, but realized the problem. At the bus/inter agent communication level, the CPU has sent a read request and is expecting a response. These protocols are usually synchronous with no clear cancellation semantics. There are probably core resources tracking then expected response and if you just freed one of those up and ended the instruction with an exception, you could later have what appears to be an unsolicited response. This dynamic probably repeats between the core and the pci root complex and then again between the root complex and the device implementing the mmio. Severing the request from the response is probably too complicated for such an unusual case.
- engzaanin 1mo agoThe timeout idea is interesting. If firmware can strictly bound SMM execution time, would that actually eliminate this class of attack, or just turn it into a crash/DoS instead?
- inigyou 1mo agoSomeone would definitely notice if SMM took a long time. It would look like the core hung. The exploit is hanging another core outside SMM.
- codedokode 1mo agoTechnically this is not a vulnerability because you need to be root. I would rather call it "taking back control of your hardware". SMM is an evil thing because the user cannot control it or look into SMM memory region. Why do CPU vendors implement a mode that cannot be controlled by the user? Obviously to use it in user-hostile purposes (software copying prevention and reporting, DRM, government access backdoors, etc.), I see no other explanation.
- xp84 1mo agoI have to agree with your overall take (only caveat being that I know too little about hardware to know if it's well-founded). The incentives are such that if it's possible to make hardware that's cryptographically locked into being aligned against the interests of its supposed owner, that's exactly what will be, which is why we now have two fully-closed systems (Google Play Services and iOS), one 99%-closed one (macOS -- Apple controlling the 'notarization' signing and showing their willingness to use it for petty reasons proves macOS is closed), and one clearly marching toward the same basic idea (Windows). What's more depressing is, even if suddenly every court agreed with me, we'd just transition overnight into a leasing paradigm, where vendors would cease to sell devices, only rent them to us. "As the device owner, should we not have the right to govern its use to only responsible purposes and protect it from 'mAlWaRe'?" And the devices would be quickly 'accepted' by the market, as "unmanaged" devices would be locked out of everything, just like you can't use banking apps, streaming apps, or even the McDonald's app, on a rooted/jailbroken phone today.
- matheusmoreira 1mo ago> Why do CPU vendors implement a mode that cannot be controlled by the user? It's even worse than that. It's gotten to the point that "operating systems" aren't actually operating the system anymore. Linux is just the "user OS", a tiny blip on the overall system schematics. Just some app to be sandboxed away from the real system. https://youtu.be/36myc8wQhLo https://youtu.be/36myc8wQhLo
- inigyou 1mo agoI'm glad someone apart from me loves this talk!
- cryptonector 1mo agoWell, shit.
- just60sec 1mo ago[dead]
- podocarp 1mo agoOne question, the firmware the repo shows seems to come from tianocore. How would you know what firmware your OEM is using? What if they have some other kind of configuration? Also I'm kind of surprised that the firmware can execute instructions like this, I was under the impression that after booting the firmware basically has finished it's job and is never touched again.