12 ms·
Increasingly the performance limit for modern CPUs is the amount of data you can feed through a single core: basically memcpy() speed. On most x86 cores the lim
by eliasdejong 8mo ago
Increasingly the performance limit for modern CPUs is the amount of data you can feed through a single core: basically memcpy() speed. On most x86 cores the limit is around 6 GB/s and about 20 GB/s for Apple M chips.
When you see advertised numbers like '200 GB/s' that is total memory bandwidth, or all cores combined. For individual cores, the limit will still be around 6 GB/s.
This means even if you write a perfect parser, you cannot go faster. This limit also applies to (de)serializing data like JSON and Protobuf, because those formats must typically be fully parsed before a single field can be read.
If however you use a zero-copy format, the CPU can skip data that it doesn't care about, so you can 'exceed' the 6 GB/s limit.
The Lite³ serialization format I am working on aims to exploit exactly this, and is able to outperform simdjson by 120x in some benchmarks as a result: https://github.com/fastserial/lite3 https://github.com/fastserial/lite3
- dehrmann 8mo ago> 6 GB/s Samsung is selling NVMe SSDs claiming 14 GB/s sequential read speed.
- johncolanduoni 8mo agoSequential read speed is attainable while still having a (small) number of independent sequential cursors. The underlying SSD translation layer will be mapping to multiple banks/erase blocks anyway, and those are tens of megabytes each at most (even assuming a 'perfect' sequential mapping, which is virtually nonexistent). So you could be reading 5 files sequentially, each only producing blocks at 3GB/s. A not totally implausible access pattern for e.g. a LSM database, or object store.
- eliasdejong 8mo ago> 14 GB/s Yes, those numbers are real but only in very short bursts of strictly sequential reads, sustained speeds will be closer to 8-10 GB/s. And real workloads will be lower than that, because they contain random access. Most NVMe drivers on Linux actually DMA the pages directly into host memory over the PCIe link, so it is not actually the CPU that is moving the data. Whenever the CPU is involved in any data movement, the 6 GB/s per core limit still applies.
- yusyusyus 8mo agoDMAing as opposed to what?
- imtringued 8mo agoLetting the CPU burn cycles by accessing memory mapped device memory instead of using DMA like in the good old days? What even is this question?
- wtallis 8mo ago> accessing memory mapped device memory instead of using DMA like in the good old days That's not actually an option for NVMe devices, is it? The actual NAND storage isn't available for direct memory mapping, and the only way to access it is to set up DMA transfers.
- loeg 8mo agoIssuing extremely slow PCIe reads? GP probably meant “all” rather than “most.”
- yusyusyus 8mo agoyep that makes more sense.
- saidnooneever 8mo agoi think he means simply the cpu speed bottleneck doesnt apply as it related to dma which doesnt 'transfer data' via the cpu directly. its phrased a bit weird.
- jauntywundrkind 8mo agoI feel like you are pulling all sorts of nonsense out of nowhere. Your numbers seem all made up. 6GB/s seems outlandishly tiny. Your justifications are not really washing. Zen4 here shows single core as, at absolute worst behavior, dropping to 57GB/s. Basically 10x what you are spinning. You are correct in that memory limits are problematic, but we also have had technology like Intel's Direct Data IO (2011) that lets the CPU talk to peripherals without having to go through main memory at all (big security disclosure on that in 2019, yikes). AMD is making what they call "Smart Data Cache Injection" which similarly makes memory speed not gating. So even if you do divide the 80GB/s memory speed across 16 chips on desktop and look at 5GB/s, that still doesn't have to tell the whole story. https://chipsandcheese.com/p/amds-zen-4-part-2-memory-subsystem-and-conclusion https://chipsandcheese.com/p/amds-zen-4-part-2-memory-subsys... https://nick-black.com/dankwiki/index.php/DDIO https://nick-black.com/dankwiki/index.php/DDIO As for SSD, for most drives, it's true true that they cannot sustain writes indefinitely. They often write in SLC mode then have to rewrite, re-pack things into denser storage configurations that takes more time to write. They'll do that in the background, given the chance, so it's often not seen. But write write write and the drive won't have the time. Thats very well known, very visible, and most review sites worth a salt test for it and show that sustained write performance. Some drives are much better than others. Even still, an Phison E28 will let you keep writing at 4GB/s until just before the drive is full full full. https://www.techpowerup.com/review/phison-e28-es/6.html https://www.techpowerup.com/review/phison-e28-es/6.html Drive reads don't have this problem. When review sites benchmark, they are not benchmarking some tiny nanosliver of data. Common benchmark utilities will test sustained performance, and it doesn't suddenly change 10 seconds in or 90 seconds in or whatever. These claims just don't feel like they're straight to me.
- wmf 8mo agoAny code that's reading/writing to SSD needs to use multiple cores. The SSD is faster than a single CPU core.
- vlovich123 8mo agoThat doesn’t sound right. A single core should more than fast enough to saturate IOPs (particularly with iouring) unless you’re doing something insane like a lot of small writes. A write of 16mib or 32mob should still be about 1 ssd iop - more CPUs shouldn’t help(and in fact should be slower if you have 2 16mib IOPs vs 1 32mib iop)
- wmf 8mo agoDo you want to process that data, or just let it hang out in memory?
- deleted 8mo ago[deleted]
- 15155 8mo agoSSDs are not faster than a DMA core.
- johncolanduoni 8mo agoWhat is the nature of the architectural limit here? The bus between an individual core and the caches and/or memory controller?
- eliasdejong 8mo agoThe limit is the number of outstanding cache line requests to the memory controller. CPUs have a fixed number of slots for this, around 10-12 usually. Intel calls them LFBs (Line Fill Buffers) and AMD MSHRs (Miss Status Holding Registers). When the slots are filled, the CPU can issue no more requests and has to wait for them to complete. Apple M chips (probably) have more slots and the memory is physically packaged together with the CPU, so they get better numbers.
- foota 8mo agoI assume these must be really expensive? Otherwise it seems like a great way to improve throughput on low concurrency tasks.
- IgorPartola 8mo agoAt least in older CPUs the caches were SRAM (static RAM). It is complicated but requires no refreshing. DRAM is basically just a capacitor per bit and capacitors leak so you constantly have to refresh the entire memory space. When the CPU sends a request to RAM, the memory controller might be too busy refreshing the soon to decay parts to actually respond right away. And if I recall correctly when you read from DRAM you destroy what was there so the process is to read it, then write it back, then send the answer to the CPU which is just a lot of steps. But the price and die size difference is huge so we use GB or TB levels of DRAM and MB levels of SRAM.
- foota 8mo agoWouldn't this bound the overall memory bandwidth, not the per core bandwidth? I've sort of assumed that just providing more line fill buffers wouldn't be sufficient, and that the number of LFB is chosen in tandem with a number of other things, but I'm not sure what the other things are (that is, just increasing the # of LFB might not be meaningful without also increasing XYZ).
- Nathanba 8mo agocool, do you think it's possible to add a schema mode to lite3 to remove the message size tradeoff? I think most people will still want to use lite3 with hard schemas during both serialization and deserialization. It's nice that it also works in a schemaless mode though.
- eliasdejong 8mo agoBeing schemaless is deliberate design decision as it eliminates the need for managing and building schema files. By not requiring schema, messages are always readable to arbitrary consumers. If you want schema, it must be done by the application through runtime type checking. All messages contain type information. Though I do see the value of adding pydantic-like schema checking in the future. EDIT: Regarding message size, Lite³ does demand a message size penalty for being schemaless and fully indexed at the same time. Though if you are using it in an RPC / streaming setting, this can be negated through brotli/zstd/dict compression.
- eru 8mo ago> By not requiring schema, messages are always readable to arbitrary consumers. That sounds a bit silly.. > All messages contain type information. That would (partially) enable what you were suggesting in the sentence I quoted first. But that's orthogonal to being schema-full or schema-less.
- digdugdirk 8mo agoPydantic was the first thing I thought of when I saw this. The possibilities are very intriguing. Do you have any thoughts/recommendations for someone if they were to try making a pydantic interface layer for lite3?
- eliasdejong 8mo agoThe library contains iterator functions that can be used to recursively traverse a message from the root and discover the entire structure. Then every value can be checked for defined constraints appropriately while at the same time validating the traversal paths. Also a schema description format is needed, perhaps something like JSONSchema can be used. Doing this by definition requires processing an entire message which can be a little slow, so it is best done in C and exposed through CPython bindings / CFFI.
- lunixbochs 8mo agoyour single core numbers seem way too low for peak throughput on one core, unless you stipulate that all cores are active and contending with each other for bandwidth e.g. dual channel zen 1 showing 25GB/s on a single core https://stackoverflow.com/a/44948720 https://stackoverflow.com/a/44948720 I wrote some microbenchmarks for single-threaded memcpy zen 2 (8-channel DDR4) naive c: 17GB/s non-temporal avx: 35GB/s Xeon-D 1541 (2-channel DDR4, my weakest system, ten years old) naive c: 9GB/s non-temporal avx: 13.5GB/s apple silicon tests (warm = generate new source buffer, memset(0) output buffer, add memory fence, then run the same copy again) m3 naive c: 17GB/s cold, 41GB/s warm non-temporal neon: 78GB/s cold+warm m3 max naive c: 25GB/s cold, 65GB/s warm non-temporal neon: 49GB/s cold, 125GB/s warm m4 pro naive c: 13.8GB/s cold, 65GB/s warm non-temporal neon: 49GB/s cold, 125GB/s warm (I'm not actually sure offhand why asi warm is so much faster than cold - the source buffer is filled with new random data each iteration, I'm using memory fences, and I still see the speedup with 16GB src/dst buffers much larger than cache. x86/linux didn't have any kind of cold/warm test difference. my guess would be that it's something about kernel page accounting and not related to the cpu) I really don't see how you can claim either a 6GB/s single core limit on x86 or a 20GB/s limit on apple silicon
- eliasdejong 8mo ago> A key feature of this code is that it skips CPU cache when copying Are those numbers also measured while skipping the CPU cache?
- lunixbochs 8mo agonaive c is just a memcpy. non-temporal uses the streaming instructions.
- nine_k 8mo agoAs much as I can understand a Zen 5 CPU core can run two AVX512 operations per clock (1024 bits) + 4 integer operations per clock (which use up FPU circuitry in the process), so additional 256 bits. At 4 GHz, this is 640 GB/s. I suppose that in real life such ideal condition do not occur, but it shows how badly the CPU is limited by its memory bandwidth for streaming tasks. Its maximum memory-read bandwidth is 768 bits per clock. only 60% of its peak bit-crunching performance. DRAM bandwidth is even more limiting. And this is a single core of at least 12 (and at most 64).
- hamandcheese 8mo agoLite claims that it can be modified in-place, but I'm curious how that works with variable-length structures like strings?
- eliasdejong 8mo agoIf the new value is equal size or smaller, it will overwrite the old value in-place. If it is larger, then the new value is appended to the buffer and the index structure is updated to point to the new location. In the case of append, the old value still lives inside the buffer but is zeroed out. This means that if you keep replacing variable-sized elements, over time the buffer will fragment. You can vacuum a message by recursively writing it from the root to a new buffer. This will clear out the unused space. This operation can be delayed for as long as you like. If you are only setting fixed-size values like integers or floats, then the buffer never grows as they are always updated in-place.
- eru 8mo agoInteresting. Sounds like you are getting copy-on-write-with-sharing for growing sizes and in-place updates when your data shrinks?
- tiffanyh 8mo ago> On most x86 cores the limit is around 6 GB/s and about 20 GB/s for Apple M chips. What makes M-series have 3x the bandwidth (per core), over x86?
- MindSpunk 8mo agoM-series have a substantially wider memory bus allowing much higher throughput. It's not really an x86/M-series thing, rather it's a packaging limitation. Apple integrates the memory into the same package as the CPU, the vast majority of x86 CPUs are socketed with socketed memory. Apple are able to push a wider bus at higher frequencies because they aren't limited by signal integrity problems you encounter trying to do the same over sockets and motherboard traces. x86 CPUs like the Ryzen AI Max 395+, when packaged without socketed memory, are able to push equally wide busses at high frequencies.
- Analemma_ 8mo agoThe sibling comment has the correct, more detailed answer, but the high-level answer is that the M-series chips are SoCs with all the RAM on-die. That lets you push way more data than you can over a bus out to socketed memory. The tradeoff is that it's non-upgradeable, but (contra some people who claim this is only a cash-grab by Apple to prevent RAM upgrades) it's worth it for the bandwidth.
- Scaevolus 8mo agoApple had soldered DDR ram for a long time that was no faster than any other laptop. It's only with the Apple Silicon M1 that it started being notably higher bandwidth.
- vlovich123 8mo agoThere were always technical benefits like lower power consumption iirc
- eru 8mo ago> The tradeoff is that it's non-upgradeable, but (contra some people who claim this is only a cash-grab by Apple to prevent RAM upgrades) it's worth it for the bandwidth. That, and if you come at it from the phone / tablet or even laptop angle: most people are quite ok just buying their computing devices pre-assembled and not worrying about upgrading them. You just buy a new one when the old one fails or you want an upgrade. Similar to how cars these days are harder to repair for the layman, but they also need much less maintenance. The guy who was always tinkering away with his car in old American sitcoms wasn't just a trope, he was truth-in-television. Approximately no one has to do that anymore with modern cars.
- rattray 8mo agoDoes capn proto have similar properties?
- eliasdejong 8mo agoYes, any zero-copy format in general will have this advantage because reading a value is essentially just a pointer dereference. Most of the message data can be completely ignored, so the CPU never needs to see it. Only the actual data accessed counts towards the limit. Btw: in my project README I have benchmarks against Cap'N Proto & Google Flatbuffers.
- eru 8mo agoHave you benchmarked against Rust's rkyv, too?
- eliasdejong 8mo agoNo, I have not benchmarked yet against other languages. Rkyv is Rust only. One primary difference is that Rkyv does not support in-place mutation. So any modification of a message requires full reserialization, unlike Lite³.
- eru 8mo agoRkyv supports a limited amount of in-place mutation. But not in general like your approach does.
- squirrellous 8mo agoWould you mind sharing what problems motivated Lite? Curious what are the typical use cases for selective reading / in place modification of serialized data. My understanding is that for cases that really want all of the fields, the zero-copy solutions aren’t much better than JSON / protobuf, so these are solutions to different problems.
- eliasdejong 8mo agoThe primary motivations were performance requirements and frustration with schema formats. The ability to mutate serialized data allows for 2 things: 1) Services exchanging messages and making small changes each time e.g. adding a timestamp without full reserialization overhead. 2) A message producer can keep a message 'template' ready and only change the necessary fields each time before sending. As a result, serializing becomes practically free from a performance perspective.
- squirrellous 8mo agoThanks. These are actually achievable with a custom protocol buffer implementation, but I agree at that point one may as well create a new format. FWIW the other use case I was expecting is something like database queries where you’d select one particular column out of a json-like blob and want to avoid deserialization.
- zozbot234 8mo agoOn quite a few recent chips (including, AIUI, Apple M series) you can only saturate memory bandwidth by resorting to the iGPU (which has access to unified memory), CPU cores on their own won't do it. It means that using the iGPU as a blitter for huge in-memory transfers and for all throughput-limited computation (including such things as parallel parsing or de/compression workloads) is now the technically advisable choice, provided that this can be arranged. > If however you use a zero-copy format, the CPU can skip data that it doesn't care about, so you can 'exceed' the 6 GB/s limit. Of course the "skipping" is by cachelines. A cacheline is effectively a self-contained block of data from a memory throughput perspective, once you've read any part of it the rest comes for free.
- woooooo 8mo ago> If however you use a zero-copy format, the CPU can skip data that it doesn't care about, so you can 'exceed' the 6 GB/s limit. You still have to load a 64-byte cache line at a time, and most CPUs do some amount of readahead, so you'll need a pretty large "blank" space to see these gains, larger than typical protobufs.
- auselen 8mo agoHow do you measure/calculate 6GB/s?
- mgaunard 8mo agoQuite easy to outperform a parsing library when you're not actually doing any parsing work and just memory-mapping pre-parsed data... That being said storing trees as serializable flat buffers is definitely useful, if only because you can release them very cheaply.
- eliasdejong 8mo agoImagine if you measured the speed of beer delivery by the rate at which beer cans can be packed/unpacked from truck pallets. But then somebody shows up with a tanker truck and starts pumping beer directly in and out. You might argue this is 'unfair' because the tanker is not doing any packing or unpacking. But then you realize it was never about packing speed in the first place. It was about delivering beer.
- StilesCrisis 8mo agoThis is actually a good analogy; the beer cans are self-contained and ready for the customer with zero extra work. The beer delivered by the tanker still needs to be poured into individual glasses by the bartender, which is slow and tedious.
- cake-rusk 8mo agoHe probably meant beer kegs. Memory map-able data is closer to sending beer cans since both are in a ready to use format.
- brunoborges 8mo ago> This limit also applies to (de)serializing data like JSON and Protobuf, because those formats must typically be fully parsed before a single field can be read. Which file formats allow partial parsing?
- perching_aix 8mo agoAnything that is both streamable and seekable?
- mort96 8mo agoYou just need to encode the size of values in bytes to make it possible to partially parse the format. Imagine the following object: { "attributes": [ .. some really really long array of whatever .. ], "name": "Bob" } In JSON, if you want to extract the "name" property, you need to parse the whole "attributes" array. However, if you encoded the size of the "attributes" array in bytes, a parser could look at the key "attributes", decide that it's not interested, and jump past it without parsing anything. You'd typically want some kind of binary format, but for illustration purposes, here's an imaginary XML representation of the JSON data model which achieves this: <object content-size-in-bytes="8388737"> <array key="attributes" content-size-in-bytes="8388608"> .. some 8 MiB large array of values .. </array> <string key="name" content-size-in-bytes="3">Bob</string> </object> If this was stored in a file, you could use fseek to seek past the "attributes" array. If it was compressed, or coming across a socket, you'd need more complicated mechanisms to seek past irrelevant parts of the object.
- brunoborges 8mo agoYeah, I was thinking of binary formats as the only solution, but your XML example is perfect. Thank you.
- quadrature 8mo agoFor what it’s worth simdjson now has an on demand api that lets you skip over keys that you don’t need.
- mort96 8mo agoWhich is great, but a JSON parser fundamentally can't avoid looking at every byte. You can't jump to the next key, you have to parse your way to the next key.
- 1vuio0pswjnm7 8mo agoPardon the ignorance, but is there a reason, or reasons, that netstrings/bencode is not included in the list of formats against which Lite^3 is tested