5 ms·
> Opening the binary with Binary Ninja revealed that clang had already managed to leverage the SSE registers. X86-64 uses SSE registers for all floating point
by bigcheesegs 6y ago
> Opening the binary with Binary Ninja revealed that clang had already managed to leverage the SSE registers.
X86-64 uses SSE registers for all floating point operations. I'm not sure that the author realized that they were looking at an -O0 binary. -O0 does not do vectorization (or anything else for that matter).
- slavik81 6y agoLooking at it on Godbolt, it doesn't really leverage SSE on -O3, either. You can get a reasonable grasp of whether it's using SSE effectively or not just by looking at the instruction names. mulss: multiplication of a single single-precision floating point value. mulsd: multiplication of a single double-precision floating point value. mulps: multiplication of a packed group of single-precision floating point values. mulpd: multiplication of a packed group of double-precision floating point values. If you're mostly seeing -ps suffixes only on moves and shuffles, you're looking at code that is not being vectorized. (And, actually, if you're seeing a lot of shuffles, that's also a good sign its not well-vectorized.) Incidentally, if you're seeing unexpected -sd suffixes, those are often due to unintended conversions between float and double. They can have a noticeable effect on performance, especially if you end up calling the double versions of math functions (as they often use iterative algorithms that need more iterations to achieve double-precision). I'm linking GCC output, because it's simpler to follow, but you see more or less the same struggle with Clang. https://godbolt.org/z/XtVqsU https://godbolt.org/z/XtVqsU
- gorgoiler 6y agoOff topic: I teach compilers in high school and godbolt.org looks amazing, thanks for the link!
- skavi 6y agoWow, really cool that there are high schools teaching compilers.
- pjmlp 6y agoYes, in Portuguese high schools you can do a technical education during the last three years (10 - 12), that gets you ready for the job market. I did mine with focus on informatics during in the late 80's/early 90's. Brief overview of three years subjects, besides the usual high school stuff. Graphics programming, compilers, databases, MS-DOS, UNIX (Xenix back then), Networking (Novell Netware), OS development. Languages that we got to use for different kinds of assignments during those three years, GW-Basic, Turbo Basic, Turbo Pascal 5.5, Turbo C 2.0/K&R C, Turbo C++ 1.0, Dbase III+, Clipper Summer '87 and OOP variant Clipper 5.x, 8086 and 68000 Assembly. The high school I took it on still offers this, naturally updated to more modern stacks and teaching subjects.
- slavik81 6y agoFor comparison, my Canadian high school offered a "Teach Yourself C++ in 30 Days" book that you could study for up to 10 hours in the optional Computers course. If you chose that module, by the end of the first class, you would be more knowledgeable on the subject than any teacher in the school. (It was actually an excellent school; they just did not care about computing. Nevertheless, I'm quite jealous of those kids with such an interesting option available to them.)
- pjmlp 6y agoYes, unfortunately education varies a lot across the globe. This is why we have Raspberry PIs now, even though such boards have existed for years. It all started as an effort to reboot UK high school computing education that was stuck into teaching Office.
- a-priori 6y agoAlso when I was in high school in Canada in about 2001-2003, I ended up taking over the class and teaching C++ because I’d already learned enough of it in my spare time that knew it better than the teacher (he liked Pascal better).
- rumanator 6y agoWow that's some next level stuff. How do high schoolers cope with the topic?
- gorgoiler 6y agoVery well! Its extremely simplified: we introduce the hierarchy of high level Python, low level C, assembler mnemonics and binary machine code and look at examples of each. But prior to that I have a few “virtual machines” that we use as compiler targets, where the VM is a robot finger that accepts the left right up down and press key commands, and the compilation step is to convert a string like HELLO into a series of robot commands. So no branching. No labels or repeatable units of code. The example gets them warmed up to the idea of converting ideas in high levels to simpler code at lower levels, for simple machines to execute. Towards the end we look at (but don’t dive too deep into) real world compilers. What does print(hello 2+3) look like in mach-O 64 assembler? Answer: erm quite a lot of gibberish but the ADDL is visible, and we can change it to SUBL and get “hello -1” to print :) Personally, the hardest parts of compiling for me to understand were the steps after lexing. Moving through a grammar to actually do things. Having everything in Python helps this a lot, as you can see how parsing some source code is just a way of triggering other code to execute. Apologies for the hand waving. I have a CS degree so I promise it’s not quite as vague as I make it out to be!
- K0nserv 6y agoThis post is really topical for me. I spent hours yesterday trying to write explicit SIMD code[0] for my the vector dot product in my raytracer and all I managed to do was slow the code down about 20-30%. The code generated by Rust from the naive solution uses ss instructions mostly whereas my two tries using `mm_dp_ps` and `mm_mul_ps` and `mm_hadd_ps` where both significantly slower even though it results in fewer instructions. I suspect that the issue is that for a single dot product the overhead of loading in and out of mm128 registers is more cost than it's worth. Naive Rust version output .cfi_startproc pushq %rbp .cfi_def_cfa_offset 16 .cfi_offset %rbp, -16 movq %rsp, %rbp .cfi_def_cfa_register %rbp vmovss (%rdi), %xmm0 vmulss (%rsi), %xmm0, %xmm0 vmovsd 4(%rdi), %xmm1 vmovsd 4(%rsi), %xmm2 vmulps %xmm2, %xmm1, %xmm1 vaddss %xmm1, %xmm0, %xmm0 vmovshdup %xmm1, %xmm1 vaddss %xmm1, %xmm0, %xmm0 popq %rbp retq My handwritten version with `mm_mul_ps` and `mm_hadd_ps` .cfi_startproc pushq %rbp .cfi_def_cfa_offset 16 .cfi_offset %rbp, -16 movq %rsp, %rbp .cfi_def_cfa_register %rbp vmovaps (%rdi), %xmm0 vmulps (%rsi), %xmm0, %xmm0 vhaddps %xmm0, %xmm0, %xmm0 vhaddps %xmm0, %xmm0, %xmm0 popq %rbp retq Intuatively it feels like my version should be faster but it isn't. In this code I changed the the struct from 3 f32 components to an array with 4 f32 elements to avoid having to create the array during computation itself, the code also requires specific alignment not to segfault which I guess might also affected performance. 0: https://github.com/k0nserv/rusttracer/commits/SIMD-mm256-dp-ps https://github.com/k0nserv/rusttracer/commits/SIMD-mm256-dp-...
- slavik81 6y agoI'm actually rather bad at this, but my understanding is that the horizontal operations are relatively slow. The easiest way to get throughput out of SIMD is to have a structure representing 4 points, with a vec4 of your X values, a vec4 of your Y values, and a vec4 of your Z values. Then you can do 4 dot products easily and efficiently, using only a handful of vertical packed instructions. (If figuring out how to effectively use a structure like that sounds difficult and annoying, that's because it is.)
- K0nserv 6y agoYeah that was my conclusion too. I don't think I have any cases where the need to perform the dot product between multuple paris of vectors arise however, at least not anywhere in the hot loop where it would help. Raytracers tend to use a lot of dot products followed by some checks then another dot product but there's a strictly sequential process in these algorithms.
- moosedev 6y agoYeah, use of SSE registers does not imply SIMD, since x87 is gone in x86-64, so even scalar FP has to use SSE registers. The asm snippets for v::operator*() in the "Optimization level 1" section use scalar SSE arithmetic only (mulss). (There's some use of movaps to move data around, but it's a stretch to call that SIMD.) I think the "leverage" sentence you quoted and the "with SIMD taken care of" one shortly after are maybe a bit misleading, since the asm snippets there don't really demonstrate SIMD.
- amluto 6y ago> since x87 is gone in x86-64, so even scalar FP has to use SSE registers. No, it’s still there. What’s actually going on is that all x86-64 CPUs support SSE2, so there is little reason to use x87 in 64-bit code. (You can use it for 80-bit precision. OTOH, for most purposes, 80-bit precision is actively harmful, and x87 is an incredible mess, so almost no one wants it.)
- moosedev 6y agoYou're right - thanks for the correction!
- ethelward 6y ago> 80-bit precision is actively harmful How comes? Unexpected clipping when converting back and forth to 64bits?
- amluto 6y agoExactly. The conversations happen at unexpected and unpredictable times depending on when the compiler needs to spill registers, which has surprising effects.
- fabiensanglard 6y agoYou are correct. Mārtiņš Možeiko pointed out that I had been too hasty when the article came out (https://twitter.com/mmozeiko/status/1257574246462570497 https://twitter.com/mmozeiko/status/1257574246462570497). To conclude SIMD is leveraged when XMM registers are used is wrong. What I should have looked for are packed instructions.
- mappu 6y agoIt would be interesting to see how an ispc version performs, if it is able to extract any more CPU parallelism.