6 ms·
Of course you can make compressed work. E.g. you fetch 66 bytes instead of 64. Hell, Intel/AMD manage to make x86 fairly fast. But it's definitely more awkward
by timhh 2mo ago
Of course you can make compressed work. E.g. you fetch 66 bytes instead of 64. Hell, Intel/AMD manage to make x86 fairly fast.
But it's definitely more awkward and has costs throughout the CPU.
I would be really surprised if the lower code density is worse than the improvement due to everything being nicely aligned. Especially because Qualcomm had actual data that it isn't (if you add new instructions with the extra coding space you free up).
- camel-cdr 2mo ago> E.g. you fetch 66 bytes instead of 64 Not really, you would fetch fewer bytes with RVC [2, page 9], because the code density is better. > I would be really surprised if the lower code density is worse than the improvement due to everything being nicely aligned. This is very hard to quantify. > Especially because Qualcomm had actual data that it isn't (if you add new instructions with the extra coding space you free up). I've liked the back and forth slides bellow. Though I want to bring up to things regrading the Qualcomm slides: > [RVC] Performance benefit is modest > • Best case: 2-3% speedup I recently benchmark compiling programs with a rva23 clang build and clang compiled for rva23-without-C and got a 10% performance improvement from RVC on the SpacemiT X100. I also have no idea how they got those numbers. (not that they are wildly implausible, it's just not transparent) > Improving Android Code Size In the last presentation they show how you can add a +-64M 32-bit long jump instruction to improve codesize in large binaries, like those in android. I want to point out, that the JAL opcode has enough space left (7/8th) to encode a long jump of a same range and there is a proposal for a 32-bit +32M -12M 32-bit long jump: https://github.com/riscv/riscv-isa-manual/blob/zijfal/src/unpriv/zijfal.adoc https://github.com/riscv/riscv-isa-manual/blob/zijfal/src/un... [1] https://lists.riscv.org/g/tech-profiles/attachment/321/0/A%20case%20to%20remove%20the%20C%20extension%20from%20app%20profiles,%20part%202%20-%20Profiles%20TG%2020231005.pdf https://lists.riscv.org/g/tech-profiles/attachment/321/0/A%2... [2] https://lists.riscv.org/g/tech-profiles/attachment/353/0/RISCV-20231004-C.pdf https://lists.riscv.org/g/tech-profiles/attachment/353/0/RIS... [3] https://lists.riscv.org/g/tech-profiles/attachment/378/0/Response%20to%20SiFive%20C%20Presentation.pdf https://lists.riscv.org/g/tech-profiles/attachment/378/0/Res... [4] https://lists.riscv.org/g/tech-profiles/attachment/400/0/AOSP%20Compression.pdf https://lists.riscv.org/g/tech-profiles/attachment/400/0/AOS...
- Taniwha 2mo agoYou don't really fetch 66 bytes instead of 64, what real implementations do is read cache lines (of whatever size) and hold on to 2 bytes from the previous cache line if there was 1/2 a 32-bit instruction at the end of the previous cache line (the ISA has the 16/32-bit tag in the lower byte so you know how big an instruction will be even if you've only seen half of it)
- inkyoto 2mo ago> […] and hold on to 2 bytes from the previous cache line if there was 1/2 a 32-bit instruction at the end of the previous cache line […] Well, and that is the worst case scenario from the performance standpoint since, if a 32-bit instruction is spans a page boundary, it will result in a page fault stalling the instruction decoder. It might be acceptable in implementations not sensitive to such an overhead (e.g. embedded solutions) but is wholly unacceptable in high performance scenarios.
- brucehoult 2mo agox86 seems to get by. And how is an instruction spanning a page boundary and causing a page fault any worse than an instruction NOT spanning a page boundary and the next instruction causing the page fault instead? As Paul said, if a 4 byte instruction spans a cache line/page boundary then you just hang on to the last 2 bytes of the page (first 2 bytes of that instruction) and decode them along with the instructions in that next cache line / page. The only time it could possibly make a difference is if that spanning instruction is a jump to somewhere else AND that instruction could somehow have fit entirely in the previous page. If there was no C extension then that next instruction would NOT be entirely in the previous page, it would be somewhere well into the next page, and that next page would have been required to be fetched much sooner. The C extension typically allows 30% to 50% more functionality to fit in each VM page. Also, Qualcomm's proposed new instructions did not in fact use the freed-up space from not having C. They fit into other unused parts of the ISA. I don't object to the new instructions Qualcomm suggested. I'd be perfectly happy to see them ratified and added to a future standard (even to RVA23 if they'd chosen to pursue that, but they didn't). What I and others objected to was dropping the C extension from RVA23, or any future RVA-series, overnight given that RVA20 and RVA22 already existed with the C extension. There will come a time when some RISC-V extensions will be retired and replaced, and it's entirely possible that C might be one of them, but there is currently no mechanism to do that, and when there is I'd expect that it would be done with a 10 or 12 year deprecation period, minimum. NEVER overnight between one standard and the next one. Which wouldn't have helped Qualcomm with their Nuvia core anyway. Anyway, Qualcomm had now bought Ventana, which has engineers who know how to support the C extension with high performance, and they already had high performance RISC-V cores doing so. So problem solved.