4 ms·
Great article and analysis as always, thanks! Somewhat crazy to remember that a (as you argue) minor CPU erretum made world wide headlines. So many worse ones o
by mras0 2y ago
Great article and analysis as always, thanks! Somewhat crazy to remember that a (as you argue) minor CPU erretum made world wide headlines. So many worse ones out there (like you mention from Intel) but others as well, that are completely forgotten.
For the Pentium, I'm curious about the FPU value stack (or whatever the correct term is) rework they did. It's been a long time, but didn't they do some kind of early "register renaming" thing that had you had to manually manage doing careful fxchg's?
- lallysingh 2y agoAFAIK, the FPU was a stack calculator. So you pushed things on and ran calculations on the stack. https://en.wikibooks.org/wiki/X86_Assembly/Floating_Point https://en.wikibooks.org/wiki/X86_Assembly/Floating_Point
- Sesse__ 2y agoIt's only a stack machine in front, really. Behind-the-scenes, it's probably just eight registers (the stack is a fixed size, it doesn't spill to memory or anything).
- lallysingh 2y agoDefinitely was 8 regs: https://intranetssn.github.io/www.ssn.net/twiki/pub/CseIntranet/CseAEC6504/8087.pdf https://intranetssn.github.io/www.ssn.net/twiki/pub/CseIntra... also where you'd see 'long double'
- Sesse__ 2y agoYes, internally fxch is a register rename—_and_ fxch can go in the V-pipe and takes only one cycle (Pentium has two pipes, U and V). IIRC fadd and fmul were both 3/1 (three cycles latency, one cycle throughput), so you'd start an operation, use the free fxch to get something else to the top, and then do two other operations while you were waiting for the operation to finish. That way, you could get long strings of FPU operations at effectively 1 op/cycle if you planned things well. IIRC, MSVC did a pretty good job of it, too. GCC didn't, really (and thus Pentium GCC was born).
- ack_complete 2y agoFMUL could only be issued every other cycle, which made scheduling even more annoying. Doing something like a matrix-vector multiplication was a messy game of FADD/FMUL/FXCH hot potato since for every operation one of the arguments had to be the top of the stack, so the TOS was constantly being replaced. Compilers got pretty good at optimizing straight line math but were not as good at cases where variables needed to be kept in the stack during a loop, like a running sum. You had to get the order of exchanges just right to preserve stack order across loop iterations. The compilers at the time often had to spill to memory or use multiple FXCHs at the end of the loop.
- Sesse__ 2y ago> FMUL could only be issued every other cycle, which made scheduling even more annoying. Huh, are you sure? Do you have any documentation that clarifies the rules for this? I was under the impression that something like `FMUL st, st(2) ; FXCH st(1), FMUL st, st(2)` would kick off two muls in two cycles, with no stall.
- Tuna-Fish 2y agoAgner Fog's manuals are clear on this. Only the last of FMUL's 3 cycles can overlap with another FMUL. You can immediately overlap with a FADD.