5 ms·
See also: https://news.ycombinator.com/item?id=32788840 https://news.ycombinator.com/item?id=32788840 TigerBeetle is designed to keep running even if all machi
by _vvhw 4y ago
See also: https://news.ycombinator.com/item?id=32788840 https://news.ycombinator.com/item?id=32788840
TigerBeetle is designed to keep running even if all machines are experiencing radioactive levels of local storage corruption, or else shutdown safely when it detects that it must. We use automated testing to test TigerBeetle with read/write storage fault injection levels as high as 20-30%. On the other hand, this invariant is not typically given by other engines, per the storage fault research that's come out of UW-Madison the past few years. For example, “Protocol-Aware Recovery for Consensus-Based Storage”.
While other engines may have incredible test suites built up over years and years, they were also designed mostly before the advent of autonomous Deterministic Simulation Testing (think Jepsen except you can speed up time and replay bugs, for example, that would otherwise take 10 years to manifest in real time), which is a showstopper.
Finally, we wanted TigerBeetle to be highly available and distributed. TigerBeetle can run across 3 availability zones with 2 replicas in each, with seamless failover. You can stay running even if you lose a whole AZ plus another replica, thanks to Heidi Howard's Flexible Quorums. Again, this problem is not as simple as simply slapping on RAFT for distribution, because “Redundancy Does Not Imply Fault-Tolerance”, and because RAFT makes concessions around storage faults and dueling leaders for the sake of the readability of the paper, that we didn't want to make for TigerBeetle's actual implementation, hence our choice of Viewstamped Replication (MIT, '88, '12).