5 ms·
How does TernFS compare to CephFS and why not CephFS, since it is also tested for the multiple Petabyte range?
by ttfvjktesd 1y ago
How does TernFS compare to CephFS and why not CephFS, since it is also tested for the multiple Petabyte range?
- rostayob 1y ago(Disclaimer: I'm one of the authors of TernFS and while we evaluated Ceph I am not intimately familiar with it) Main factors: * Ceph stores both metadata and file contents using the same object store (RADOS). TernFS uses a specialized database for metadata which takes advantage of various properties of our datasets (immutable files, few moves between directories, etc.). * While Ceph is capable of storing PBs, we currently store ~600PBs on a single TernFS deployment. Last time we checked this would be an order of magnitude more than even very large Ceph deployments. * More generally, we wanted a system that we knew we could easily adapt to our needs and more importantly quickly fix when something went wrong, and we estimated that building out something new rather than adapting Ceph (or some other open source solution) would be less costly overall.
- mgrandl 1y agoThere are definitely insanely large Ceph deployments. I have seen hundreds of PBs in production myself. Also your usecase sounds like something that should be quite manageable for Ceph to handle due to limited metadata activity, which tends to be the main painpoint with CephFS.
- kachapopopow 1y agoCeph is more of: here's a raw block of data, do whatever the hell you want with it, not really good for immutable data.
- mgrandl 1y agoWell sure you would have to enforce immutability at the client side.
- kachapopopow 1y agoIt's more that it has all the systems to allow mutability which add a lot of overhead when used as an immutable system.
- rostayob 1y agoI'm not fully up to date since we looked into this a few years ago, at the time the CERN deployments of Ceph were cited as particularly large examples and they topped out at ~30PB. Also note that when I say "single deployment" I mean that the full storage capacity is not subdivided in any way (i.e. there are no "zones" or "realms" or similar concepts). We wanted this to be the case after experiencing situations where we had significant overhead due to having to rebalance different storage buckets (albeit with a different piece of software, not Ceph). If there are EB-scale Ceph deployments I'd love to hear more about them.
- mgrandl 1y agoThere are much larger Ceph clusters, but they are enterprise owned and not really publicly talked about. Sadly I can’t share what I personally worked on.
- rostayob 1y agoThe question is whether there are single Ceph deployments are that large. I believe Hetzner uses Ceph for its cloud offering, and that's probably very large, but I'd imagine that no single tenant is storing hundreds of PBs in it. So it's very easy to shard across many Ceph instances. In our use-case we have a single tenant which stores 100s of PBs (and soon EBs).
- ttfvjktesd 1y agoDigital Ocean is also using Ceph[1]. I think these cloud providers could easily have 100s of PBs Clusters at their size, but it's not public information. Even smaller company's (< 500 employees) in today's big data collection age often have more than 1 PB of total data in their enterprise pool. Hosters like Digital Ocean hosts thousands of these companies. I do think that Ceph will hit performance issues at that size and going into the EB range will likely require code changes. My best guess would be that Hetzner, Digital Ocean and similar, maintain their own internal fork of Ceph and have customizations that tightly addresses their particular needs. [1]: https://www.digitalocean.com/blog/why-we-chose-ceph-to-build-block-storage https://www.digitalocean.com/blog/why-we-chose-ceph-to-build...
- eps 1y agoLast point is an extremely important advantage that is often overlooked and denigrated. But having a complex system that you know inside-out because you made it from scratch pays in gold in the long term.
- _jsmh 1y agoAny compression at the filesystem level?
- rostayob 1y agoNo, we have our custom compressor as well but it's outside the filesystem.
- jleahy 1y agoThe seamless realtime intercontinental replication is a key feature for us, maybe the most important single feature, and AFAIK you can’t do that with Ceph (even if Ceph could scale to our original 10 exabyte target in one instance).
- cmdrk 1y agoCephFS implements a (fully?) POSIX filesystem while it seems that TernFS makes tradeoffs by losing permissions and mutability for further scale. Their docs mention they have a custom kernel module, which I suppose is (today) shipped out of tree. Ceph is in-tree and also has a FUSE implementation. The docs mention that TernFS also has its own S3 gateway, while RADOSGW is fully separate from CephFS.
- jcul 1y agoMy (limited) understanding is that cephfs, RGW (S3), RBD (block device) are all different things using the same underlying RADOS storage. You can't mount and access RGW S3 objects as cephfs or anything, they are completely separate (not counting things like goofys, s3fs etc.), even if both are on the same rados cluster. Not sure if TernFS differs there, would be kind of nice to have the option of both kinds of access to the same data.
- KaiserPro 1y agoCeph isn't that well suited for high performance. its also young and more complex than you'd want it to be (ie you get a block storage system, which you then have to put a FS layer on after.) if you want performance, then you'll probably want lustre, or GPFS, or if you're rich a massive isilon system.