5 ms·
Anyone have some insight into situations where running map reduce on redis makes more sense than other software like the traditional hadoop?
by grantjgordon 14y ago
Anyone have some insight into situations where running map reduce on redis makes more sense than other software like the traditional hadoop?
- seiji 14y agoHadoop is a bloated pile of elephant poo. Any and all alternatives are welcome. Disco (http://discoproject.org/ http://discoproject.org/) is popular in some parts of the mapreducesphere.
- achompas 14y agoI hear the above comment about Hadoop a lot. Can you explain why?
- seiji 14y agoI'd love to, but it would take about an hour to run through everything. Here's a short version: There's a collective ecosystem problem of fragmented applications, not-quite-right command line utilities, web interfaces that look like they were designed in 1995, noisy log files people actually have to read constantly, and cross coupling of dependencies that make keeping a cluster live for production use a full time job. There's the programming problem of nobody actually writing hadoop mapreduce code because it's impossibly complicated. Everybody uses hive and pig and half a dozen other tools to compile to pre-templated java classes (this knocks off 5% to 30% of your performance if you could do it by hand). It hasn't grown because it's so amazing, performant, and company saving. It grows because people jumped on a fad wagon then got stuck with having a few hundred TB in HDFS. The lack of a competing project with equal mindshare and battle-testedness doesn't foster any competition. It's the mysql of distributed processing systems. It works (mostly), but it breaks (in a few dozen known ways), so people keep adding features and building on top of it.
- cgh 14y agoI'd also just like to say: NameNode = single point of failure. I worked on a contract for a large, very well-known social networking company a while back who refused to consider Hadoop because of this.
- MichaelSalib 14y agoseiji pretty much nails it. Hadoop seems to have come out of a weird culture. It is a distributed system with a single point of failure (name node) because its designers insisted on avoiding Paxos (distributed systems are too hard so we'll just make a broken-by-design protocol instead). Another example is that a lot of the database code built on top of Hadoop is designed around one Java hashmap per row which really limits performance. There are all sorts of oddities and you can mostly work around them but it is...exhausting, and I spend a lot of time thinking "surely there must be a better way".
- jamii 14y ago> surely there must be a better way http://www.spark-project.org/ http://www.spark-project.org/
- sqrt17 14y agoWait, so Zookeeper (= distributed consensus thingie that I think implements the Paxos algorithm) is a Hadoop project but not actually used in Hadoop mapreduce?
- mumrah 14y agoThat's correct. I believe they are using it in some new "high availability" stuff coming down the road
- grantjgordon 14y agoThanks for your insightful comments! I appreciate that you took the time to back up your opinion by distilling your thoughts into something quickly digestible. Have you heard of any other projects outside of disco that are more performant than hadoop when used for similar applications?
- fsaintjacques 14y agoUsing disco here, very happy with it.
- grantjgordon 14y agoMind sharing how long you've been using it and how it compares to hadoop in your opinion? I'm very interesting in hearing your experience.
- achompas 14y agoSame here, I'm really interested in hearing about Disco and potential benefits/costs vs. Hadoop.
- heynemann 14y agoThe reason I wrote r³ is because I was a little overwhelmed by how complex disco is to administer and scale. r³ was designed from the ground up to adhere to HTTP. That means it's pretty easy to scale using our old and well-proven techniques: caching and load-balancing.
- sitkack 14y agoI <B disco. It is so well designed and easy to run. Truly, love it.
- hogu 14y agoIf you have alot of data, and network IO is a big issue, you'll want to use something like hadoop (or disco) becuase they come with an integrated distributed file system and they preserve data locality. If you don't have that much data, MR on redis is fine