9 ms·
Elasticsearch was never a database
- unethical_ban 8mo agoI work in infosec and several popular platforms use elasticsearch for log storage and analysis. I would never. Ever. Bet my savings on ES being stable enough to always be online to take in data, or predictable in retaining the data it took in. It feels very best-effort and as a consultant, I recommend orgs use some other system for retaining their logs, even a raw filesystem with rolling zips, before relying on ES unless you have a dedicated team constantly monitoring it.
- toenail 8mo agoDunno, I've had three node clusters running very stable for years. Which issues did you have that require a full team?
- unethical_ban 8mo agoTo be fair, I think it is chronically underprovisioned clusters that get overwhelmed by log forwarding. I wasn't on the team that managed the ELK stack a decade ago, but I remember our SOC having two people whose full time job was curating the infrastructure to keep it afloat. Now I work for a company whose log storage product has ES inside, and it seems to shit the bed more often than it should - again, could be bugs, could be running "clusters" of 1 or 2 instead of 3.
- PedroBatista 8mo agoEven most toy databases "built in a weekend" can be very stable for years if: - No edge-case is thrown at them - No part of the system is stressed ( software modules, OS,firmware, hardware ) - No plug is pulled Crank the requests to 11 or import a billion rows of data with another billion relations and watch what happens. The main problem isn't the system refusing to serve a request or throwing "No soup for you!" errors, it's data corruption and/or wrong responses.
- kentm 8mo agoDo you happen to know if ES was the only storage? Its been almost 8 years, but if I was building a log storage and analysis system, then I'd push the logs to S3 or some other object store and build an ES index off of that S3 data. From the consumer's perspective, it may look like we're using ES to store the data, but we have a durable backup to regenerate ES if necessary.
- lillesvin 8mo agoSearchable snapshots in Elasticsearch can be backed by S3 and they perform very well. No need to store the data on hot nodes any longer than it takes for the index to do a rollover, and from then it's all S3.
- cyberpunk 8mo agoMeh i run hundreds of es nodes, its gotten a lot more friendly these days, but yes it can be a bit unforgiving at times. Turns out running complicated large distributed systems requires a bit more than a ./apply, who would have guessed it?
- 1_1xdev1 8mo agoYou have to slap something durable and a queue in front of it. Elastic’s own consultants will tell you this …
- yencabulator 8mo ago> I work in infosec and several popular platforms use elasticsearch for log storage and analysis. Storing logs in ElasticSearch is just stupid, as it does not preserve order: https://logstash.jira.com/browse/LOGSTASH-192 https://logstash.jira.com/browse/LOGSTASH-192
- lvspiff 8mo agoEverything is a database if you believe hard enough Feel like the christmas story kid -- >simplicity, and world-class performance, get started with XXXXXXXX. A crummy commercial?
- Quarrelsome 8mo agoram is a database, you just need bigger capacitors.
- marcosdumay 8mo agoThat's literally the original MySQL philosophy. And it was good for a lot of things.
- PedroBatista 8mo agoI really never understood how people could store very important information in ES like it was a database. Even if they don't understand what ES is and what a "normal" database is, I'm sure some of those people run into issues where their "db" got either corrupted of lost data even when testing and building their system around it. This is and was general knowledge at the time, it was no secret that from time to time things got corrupted and indexes needed to be rebuilt. Doesn't happen all the time, but way greater than zero times and it's understandable because Lucene is not a DB engine or "DB grade" storage engine, they had other more important things to solve in their domain. So when I read stories of data loss and things going South, I don't have sympathy for anyone involved other than the unsuspecting final clients. These people knew or more or less knew and choose to ignore and be lazy.
- kentm 8mo ago> I really never understood how people could store very important information in ES like it was a database. I agree. Its been a while since I touched it, but as far as I can remember ES has never pretended to be your primary store of information. It was mostly juniors that reached for it for transaction processing, and I had to disabuse them of the notion that it was fit for purpose there. ES is for building a searchable replica of your data. Every ES deployment I made or consulted sourced its data from some other durable store, and the only thing that wrote to it were replication processes or backfills.
- vjerancrnjak 8mo agoThey market it as a general purpose store. Successfully, even though hc cs wizards wouldn’t touch it ever, c suite likes it Best example is IoT marketing, as if it can handle the load without bazillion shards, and since when does a text engine want telemetry
- WASDx 8mo agoI've managed a 100+ node cluster for years without seeing any corruption. Where are you getting this from?
- 8mo ago
- alittletooraph2 8mo ago[dead]
- toenail 8mo agoI think elastic always clearly documented to expect "eventual consistency", they never claimed to be a "database" in the sense that tfa defines.
- xeraa 8mo agoFirst step of a marketing campaign: Claim something never said and then tell everyone why it's wrong ;)
- cess11 8mo agoIt's not so much that Elastic is saying it as a lot of people doing the supposed wrong the advert-article describes. I've seen some examples of people using ES as a database, which I'd advise against for pretty much the reasons TFA brings up, unless I can get by on just a YAGNI reasoning.
- xeraa 8mo agoIt will also depend a lot on the type of data: Logs are an easy yes. Something that required multi-document transactions (unless you're able to structure it differently) is a harder tradeoff. Though loss of ACKed documents shouldn't really be a thing any more.
- roywiggins 8mo ago> Elastic has been working on this gap. The more recent ES|QL introduces a similar feature called lookup joins, and Elastic SQL provides a more familiar syntax (with no joins). But these are still bound by Lucene’s underlying index model. On top of that, developers now face a confusing sprawl of overlapping query syntaxes (currently: Query DSL, ES|QL, SQL, EQL, KQL), each suited to different use cases, and with different strengths and weaknesses. I suppose we need a new rule, "Any sufficiently successful data store eventually sprouts at least one ad hoc, informally-specified, inconsistency-ridden, slow implementation of half of a relational database"
- kayo_20211030 8mo ago... and then becomes an email client (https://en.wikipedia.org/wiki/Jamie_Zawinski#Zawinski%27s_Law https://en.wikipedia.org/wiki/Jamie_Zawinski#Zawinski%27s_La...). A two-fer. lol.
- all2 8mo agoIt seems like everything converges on either LISP or emacs.
- esafak 8mo agoICYMI https://en.wikipedia.org/wiki/Greenspun's_tenth_rule https://en.wikipedia.org/wiki/Greenspun's_tenth_rule
- wasting_time 8mo agoICYMI expands to "in case you missed it", ICYMI.
- virgil_disgr4ce 8mo agoICYMI expands to ... wait, shit
- patates 8mo agoI see... why am I...
- speedgoose 8mo agoAccenture managed to build a data platform for my company with Elasticsearch as the primary database. I raised concerns early during the process but their software architect told me they never had any issues. I assume he didn’t lie. I was only an user so I didn’t fight and decided to not make my work rely on their work.
- CuriouslyC 8mo agoElastic feels about as much like a primary data store as Mongo, FWIW.
- victor106 8mo ago> Accenture They messed up a $30 million dollar project big time at a previous company. My cto swore to never recommend them
- 9rx 8mo agoI've seen some mess-ups in my life, but they started sticking out like a sore thumb long, long, long, long before anywhere close to $30 million was spent on it. What does a $30 million dollar mess-up look like?
- nwallin 8mo agoI am not OP and am not speaking for them. "A $30 million mess-up" can look like (at least) two things. It can be $30 million was spent on a project that earned $0 revenue and was ultimately canceled, or it can look like $x was spent on a project to win a $30 million contract but a competitor won the contract instead.
- rawgabbit 8mo agoTeams of consultants on site, some remote, and many offshore. Tons of documents are created and many environments and DevOps pipelines are stood up. First code release is when the people who push buttons touch the system for the first time. It is crap. Several more code releases attempt to make the system usable. Eventually another consultant or two are brought to evaluate the project and they say the project violated every best practice and common sense rule. Most egregiously the internal stakeholders who voiced serious concerns at the beginning of the project were dismissed or forced out etc.
- cluckindan 8mo ago”That means a recently acknowledged write may not show up until the next refresh.” Which is why you supply the parameter refresh: ”wait_for” in your writes. This forces a refresh and waits for it to happen before completing the request. ”schema migrations require moving the entire system of record into a new structure, under load, with no safety net” Use index aliases. Create new index using the new mapping, make a reindex request from old index to new one. When it finishes, change the alias to point to the new index. The other criticisms are more valid, but not entirely: for example, no database ”just works” without carefully tuning the memory-related configuration for your workload, schema and data.
- nkmnz 8mo agoIt took me years before I started tuning the memory-related configuration of postgres for workload, schema and data, in any way. It "just works" for the first ten thousand concurrent users.
- cluckindan 8mo agoWell, most people working on a car don’t have a car lift: it only makes sense when you need to safely work on a large volume of cars. If you only work on one or two, a jack and a pile of wood works just fine.
- nkmnz 8mo agoPlease don't move the goal post. Writing `no database ”just works” without (...)` is gatekeeping behavior, creating an image of complexity that for most use cases - especially for those starting out - just doesn't exist.
- cluckindan 8mo agoIn fairness, it doesn’t exist for Elasticsearch either.
- 8mo ago
- stefanon 8mo agoYep!
- this_user 8mo agoI mean, it is called "ElasticSEARCH", not "Elasticdatabase".
- _joel 8mo agoMySQL isn't mine either, it's Larry Ellison's.
- rpdillon 8mo agoWell, "My" is the name of the author's daughter, rather than a reference to who owns it.
- 8mo ago
- gmuslera 8mo ago... for a particular, opinionated definition of what a database should be.
- aaroninsf 8mo agoWe use ES like a DB, but, not with SQL; and most importantly, it's not the source of truth/primary store. It's operational truth and best-effort.
- jfengel 8mo agoNo, of course not. But the question is, do you need a database? A database is a big proposition: transactions, indexes, query processing, replication, distribution, etc. A fair number of use cases are just "Take this data and give it back to me when I ask for it". ES (or any other not-a-database) might not be a full-bore DBMS. But it might be what you need.
- immibis 8mo agoRule of thumb: Whenever you think you don't need relational database features, you will later discover why you do. The one thing relational databases don't have, that you might need, is scaling. Maintaining data consistency implies a certain level of non-concurrency. Conversely, maintaining perfect concurrency implies a certain level of data inconsistency.
- deleted 8mo ago[deleted]
- 9rx 8mo agoThe other thing relational databases don't have, that you are definitely going to need, is a practical implementation. You could maybe consider Rel if you have a particular type of workload, but, realistically, just use a tablational database. It will be a lot easier and is arguably better.
- wwarner 8mo agoThese drawbacks are all true, but sometimes storing directly to elastic is still the best way forward.
- throw_m239339 8mo agoIt has an index? It has data that can be queried with indexes? it is a database. PERIOD. Let's not turn the word database into a buzzword. It should obviously NOT be a "main" database but part of an ETL pipeline for search purposes for instance.
- 9rx 8mo ago> Let's not turn the word database into a buzzword. It is much too late for that, but you're right that we'd be wise to put effort into undoing that. This is exactly how you end up with people using Elasticsearch as a primary datastore. When someone hears that they need a database, a database is what you are going to see them pick. If we regularly used the proper terminology with appropriate specificity then those without the deep technical knowledge required to understand all the different kinds of databases and the tradeoffs that come them are able to narrow their search to the solutions that fit within the specification.
- CodeCompost 8mo agoYes but is it webscale? (Obviously I'm referring to a famous YouTube video on the subject)
- trgn 8mo agoit now has dedicated index types for logs and metrics with all kinds of sugar and tweaks in default behavior, they should introduce a new one called "database" that's acid.
- vedhant 8mo agoOfcourse it is not meant as a primary database. What baffles me is that people use it as log storage. As an application scales, storage and querying logs become the bottleneck if elasticsearch is used. I was dealing with a system that could afford only 1 week of log retention!
- SlightlyLeftPad 8mo agoLogs are always notoriously expensive to store and also are notorious for accidentally exposing PII, API/private/db keys, etc. They should generally only be stored for a relatively short period of time at scale. In fact, to remain compliant to CCPA, 28 days is the safe number for most things. Metrics are much more efficient and are the tool of choice for longer term storage and debugging.
- lillesvin 8mo agoWhat kind of storage do you have backing your Elasticsearch? And how have you configured sharding and phase rollover in your indices? I work with a cluster that holds 500+ TB logs (where most are stored for a year and some for 5 years because of regulations) in searchable snapshots backed by a locally hosted S3 solution. I can do filtering across most of the data in less than 10 seconds. Some especially gnarly searches may take around 60-90 seconds on the first run as the searchable snapshots are mounted and cached, but subsequent searches in the cached dataset are obviously as fast as any other search in hot data. Obviously Elasticsearch isn't without its quirks and drawbacks, but I have yet to come across anything that performs better and is more flexible for logs — especially in terms of architectural freedom and bang-for-the-buck.
- largbae 8mo agoAre the folks still using ES simply unaware of the performance advantages of ClickHouse, or is there some use case that ES covers that CH is still missing?
- wpaladin 8mo agoFull-text search. It was only added to Clickhouse relatively recently, and is still in Beta. It's a core feature of ES from the beginning. https://clickhouse.com/docs/engines/table-engines/mergetree-family/textindexes https://clickhouse.com/docs/engines/table-engines/mergetree-...
- zacksiri 8mo agoI never thought of Elasticsearch as a database and always designed systems around what elasticsearch is supposed to be an index based document store for used with search. I think their API is great and have had amazing results with it. Their recent innovations around quantization (bbq) has been amazing for my use case building an agentic movie database for discovering movies and personalized movie recommendations. There are benefits to not using your database for everything, even if it adds a bit of complexity by introducing another dependency. If the benefits out weigh the cost of complexity reaching for elastic has almost always been worth it for me.
- almosthere 8mo agoI've always generally used some other data source and ran a spark job to populate elastic. For live data, just have a database trigger through a message queue populate ES.
- almosthere 8mo agoMost people in this entire discussion don't really even understand the analytic queries are and think ES was for full text search. Picard Hand Over Face
- forinti 8mo agoES and Postgresql integrate nicely with FDW, so you can have the best of both worlds.
- jamesgresql 8mo agoI know it sounds obvious, but some people are pretty determined to us it that way!