8 ms·
How we replaced Elasticsearch and MongoDB with Rust and RocksDB
- maelito 1y agoI wonder if this could help Photon, the open source ElasticSearch/OpenSearch search engine for OSM data. It's a mini-revolution in the OSM world, where most apps have a bad search experience where typos aren't handled. https://github.com/komoot/photon https://github.com/komoot/photon
- hyc_symas 1y agoA system built on LMDB will work better for this use case than RocksDB. And OSM Express already uses it. https://wiki.openstreetmap.org/wiki/OSM_Express https://wiki.openstreetmap.org/wiki/OSM_Express
- sophia01 1y agoThey're not open sourcing it though?
- pbowyer 1y agoDoesn't sound like it, but it's a nice writeup of the tools they stitched together. For someone to copy and open source... hopefully :)
- cicloid 1y agoTempted, specially for switching H3 instead of S2… I prototyped a similar solution a couple of weeks ago, so I could probably do a second pass
- ellenhp 1y agoWhat's wrong with S2? H3 is so much more complex for very little gain from what I can tell.
- ellenhp 1y agoThere are a few piece of this that rely on proprietary data, especially the FastText training step, so that's a dead-end unfortunately (would love to be proven wrong). I'd consider subbing in a small bert model with a classifier head for something FOSS without access to tons of user data, but then you lose the ability to serve high qps.
- mips_avatar 1y agoI guess not having that would only breaking forward geocoding from an address?
- ellenhp 1y agoMy guess is that they're using FastText for semantic search, so it's more likely to break queries like "coffee near me" than address search, the latter likely being handled by tantivy. For context, I've also written a geocoder [0] based on tantivy. :) [0] https://github.com/ellenhp/airmail https://github.com/ellenhp/airmail
- mips_avatar 1y agoWow Airmail looks awesome. Have you ever benchmarked it on latency? I'm working on geocoding solutions for AI agents so quick tool calls is really important.
- j_kao 1y agoIt's a bit difficult at the moment, given we have a lot of proprietary data at the moment and a lot of the logic follows it. I'm hoping we can get it to a state where it can be indexed and serving OSM data but that is going to take some time. That being said, we are currently working on getting our Google S2 Rust bindings open-sourced. This is a geo-hashing library that makes it very easy to write a reverse geocoder, even from a point-in-polygon or polygon-intersection perspective.
- mips_avatar 1y agoCould you write a photon replacement if you had that? I would love to spend less per month running photon for my project.
- deleted 1y ago[deleted]
- softwaredoug 1y agoIt’s interesting as someone in the search space how many companies are aiming to “replace Elasticsearch”
- mikeocool 1y agoIn my experience, the care and feeding that goes into an Elastic Search cluster feels like it's often substantially higher than that involved in the primary data store, which has always struck me as a little odd (particularly in cases where the primary data store is an RDBMS). I'd be very happy to use simpler more bulletproof solutions with a subset of ES's features for different use cases.
- dewey 1y agoTo add another data point: After working with ES for the past 10 years in production I have to say that ES is never giving us any headaches. We've had issues with ScyllaDB, Redis etc. but ES is just chugging along and just works. The one issue I remember is: On ES 5 we once had an issue early on where it regularly went down, turns out that some _very long_ input was being passed into the search by some scraper and killed the cluster.
- everfrustrated 1y agoHow big is the team that looks after it?
- dewey 1y agoNobody is actively looking after it. Good alerting + monitoring and if there's an alert like a node going down because of some Kubernetes node shuffling or a version upgrade that has to be performed one of our few infra people will do that. It's really not something that needs much attention in my experience.
- itpragmatik 1y agohow many clusters, how many indexes and how many documents per index? do you use self hosted es or aws managed opensearch?
- pm90 1y agoSlightly meta, but I find its a good sign that we're back to designing and blogging about in-house data storage systems/ Query engines again. There was an explosion of these in the 2010's which seemed to slow down/refocus on AI recently.
- 8n4vidtmkvmk 1y agoIs it good? What's left to innovate on in this space? I don't really want experimental data stores. Give me something rock solid.
- cfors 1y agoI don't disagree that rock solid is a good choice, but there is a ton of innovation necessary for data stores. Especially in the context of embedding search, which this article is also trying to do. We need database that can efficiently store/query high-dimensional embeddings, and handle the nuance of real-world applications as well such as filtered-ANN. There is a ton of innovation in this space and it's crucial to powering the next generation architectures of just about every company out there. At this point, data-stores are becoming a bottleneck for serving embedding search and I cannot understate that advancements in this are extremely important for enabling these solutions. This is why there is an explosion of vector-databases right now. This article is a great example of where the actual data-providers are not providing the solutions companies need right now, and there is so much room for improvement in this space.
- whakim 1y agoI do not think data stores are a bottleneck for serving embedding search. I think the raft of new-fangled vector db services (or pgvector or whatever) can be a bottleneck because they are mostly optimized around the long tail of pretty small data. Real internet-scale search systems like ES or Vespa won’t struggle with serving embedding search assuming you have the necessary scale and time/money to invest in them.
- 1y ago
- jothirams 1y agoIs horizondb publicly available for us to try as well..
- trimbo 1y agoThis article is lacking detail. For example, how is the data sharded, how much time between indexing and serving, and how does it handle node failure, and other distributed systems questions? How does the latency compare? Etc. etc.
- reactordev 1y agoI mean, anything could replace elasticsearch, but can it actually? It sounds like they had the wrong architecture to start with and they built a database to handle it. Kudos. Most would have just thrown cache at it or fine tuned a readonly postgis database for the geoip lookups. Without benchmarks it’s just bold claims we’ll have to ascertain.
- brunohaid 1y agoBit thin on details and not looking like they’ll open source it, but if someone clicked the post because they’re looking for their “replace ES” thing: Both https://typesense.org/ https://typesense.org/ and https://duckdb.org/ https://duckdb.org/ (with their spatial plugin) are excellent geo performance wise, the latter now seems really production ready, especially when the data doesn’t change that often. Both fully open source including clustered/sharded setups. No affiliation at all, just really happy camper.
- mcdonje 1y agoNot sure what they'll opensource. The rust code? They're calling it a DB, but they described an entire stack.
- jjordan 1y agoTypesense is an absolute beast, and it has a pretty great dev experience to boot.
- porridgeraisin 1y agoCan you share what makes it better than competitors? And what's great about the dev experience? Did you use their cloud offering? The marketing material looks great, but I want to hear a user's experience.
- brunohaid 1y agoFor me it's a combination 1) solid foundational choices all along, no bolted on vanity features or constant rewrites chasing the latest trend, with everything well documented and 2) incredibly responsive founding team, so you get very quick answers from the people actually building it.
- sureglymop 1y agoThese are great. I am eternally grateful that projects like this are open source, I do however find it hard to integrate them into your own projects. A while ago I tried to create something that has duckdb + its spatial and SQLite extensions statically linked and compiled in. I realized I was a bit in over my head when my build failed because both of them required SQLite symbols but from different versions.
- kosolam 1y agoSide note 1: ES can also be embedded in your app (on the JVM). Note 2: I actually used RocksDB to solve many use cases and it’s quite powerful and very performant. If anything from this post take this, it’s open source and a very solid building block. Note 3: I would like to test drive quickwit as an ES replacement. Haven’t got the time yet.
- j_kao 1y ago1 - I think if we were sticking with the JVM, I do wonder if Lucene would be the right choice in that case 2 - It's a great tool with a lot of tuneability and support! 3 - We've been using it for K8s logs and OTEL (with Jaeger). Seems good so far, though I do wonder how the future of this will play out with the $DDOG acquisition.
- vips7L 1y agoI really enjoy embedding things in the vm. I run a discord bot with a few thousand users with embedded H2. Recently I’ve been looking at trying to embed keycloak (or something similar) for some other apps.
- kosolam 1y agoI did that with ES to squeeze performance but IIRC it didn’t really produce meaningful results. Otherwise, for most use cases an integration is better imho than embedding stuff that is when you have a full software service such as keycloack or ES. Rocksdb and h2 are tailor made as embedded libraries
- mexxixan 1y agoWould love to know how they scaled it. Also, what happens when you lose the machine and the local db? I imagine there are backups but they should have mentioned it. Even with backups how do you ensure zero data loss.
- tracker1 1y agoNice... it's cool to see how different companies are putting together best fit solutions. I'm also glad that they at least started out with off the shelf apps instead of jumping to something like a bespoke solution early on. Quickwit[1] looks interesting, found via Tantivity reference. Kind of like ES w/ Lucene. 1. https://github.com/quickwit-oss/quickwit https://github.com/quickwit-oss/quickwit
- francoismassot 1y agoit's tantivy :)
- 9cb14c1ec0 1y agoClicked because of Elasticsearch, then wondered why I hadn't known of radar.com before. Just the autocomplete at a reasonable price that I need.
- darqis 1y agoSearching for HorizonDB I find a Python project on github. I'm guessing it's closed source *aas only?
- lisbbb 1y agoI can see ditching Mongo, but what's bad about ElasticSearch? Too expensive in some way? Isn't RocksDB just the db engine for Kafka?
- 0xbadcafebee 1y agoRocks is a fork of Level, and Level is well known for data corruption and other bugs. They are both "run at production scale", but at least back when I worked on stuff that used Level, nobody talked publicly about all the toil spent on cleaning up and repairing Level to keep the services based on it running. Whenever you see an advertisement like this (these posts are ads for the companies publishing them), they will not be telling you the full truth of their new stack, like the downsides or how serious they can be (if they've even discovered them yet). It's the same for tech talks by people from "big name companies". They are selling you a narrative.
- Jweb_Guru 1y agoRocksDB diverged from LevelDB a long time ago at this point and has had extensive work done on it by both industry and academia. It's not a toy database like LevelDB was. I can't speak to the problems they're supposedly hiding in their stack, but they are unlikely to come from RocksDB.
- KAdot 1y agoThis is not my experience. I've been running RocksDB for 4 years on thousands of machines, each storing terabytes of data, and I haven't seen a single correctness issue caused by RocksDB.
- dboreham 1y agoThese are not the same kinds of things.
- nekitamo 1y agoI've used RocksDB a lot in the past and am very satisfied with it. It was helpful building a large write-heavy index where most of the data had to be compressed on disk. I'm wondering if anyone here has experience with LMDB and can comment on how they compare? https://www.symas.com/mdb https://www.symas.com/mdb I'm looking at it next for a project which has to cache and serve relatively small static data, and write and look up millions of individual points per minute.
- hyc_symas 1y agoLMDB is for read-heavy workloads. The opposite of RocksDB. RocksDB can use thousands of file descriptors at once, on larger DBs. Makes it unsuitable for servers that may also need to manage thousands of client connections at once. LMDB uses 2 file descriptors at most; just 1 if you don't use its lock management, or if you're serving static data from a readonly filesystem. RocksDB requires extensive configuration to tune properly. LMDB doesn't require any tuning.
- pianoben 1y agoLol I "love" that the first benefit this company lists in their jobs page is "In-Office Culture". Do people actually believe that having to commute is a benefit?
- aflag 1y agoI rather commute than WFH. So yeah, people do. Maybe not all the people, but certainly some people.
- 01HNNWZ0MV43FF 1y agoIn-office culture would be dope if there were actual benefits to an office like maybe Learning from smart people, making friends, free food and drinks, a DDR machine My last office job had none of that. Instead it was just sort of like a depressing scaled up version of my home office
- LtWorf 1y agoMy office has some nice perks! 1. It's extremely cold and dark! I must wear extra clothes when going inside and I get depressed at wasting a day of nice weather in what looks like a WW1 bunker. 2. Terrible accessibility for disabled people! (such as myself) 3. Filthy toilets! 4. Internet is slower than at home! 5. Half the team lives somewhere else so all meetings are on teams anyway! 6. They couldn't afford a decent headset so I get pain in my head after 5 minutes, but I don't have a laptop so I can't move to a meeting room. The HR really can't understand why after all these great perks I insist on wanting to work from home. I am such an illogical person!
- throw738338 1y agoFriend works at office that allows dogs. Her workplace is one big dog toilet! She is expected to clean it (she is not toilet cleaner). She get sexually assaulted, when her boss shoved his dog into her crotch! There were some hospitalisations from work related injuries... Regular bullying, threats of violence.... Lovely office culture!
- tinyhouse 1y agofastText? last time I checked it wasn't even maintained.
- feverzsj 1y agoSounds like all they need is Postgres or just Sqlite.
- benjiro 1y agoYep, their goal was to use a monolite solution, instead of a clustered elasticsearch. Postgres + pg_search (= tantivy) will have gotten them there for 80%. Sure, postgres really needs a plugin storage engine for better SSD support (see orioledb). But creating your own database for your own company, is just silly. There is a lot of money in the database market, and everybody wants to do their own thing to tied customers down to those databases. And that is their main goal.
- tapirl 1y agoIt is weird to include "Rust" (a language) in the title. Readers might wonder what is replaced by Rust? Elasticsearch or MongoDB?