9 ms·
A Deep Dive into Solid Queue for Ruby on Rails
- nkraft11 1y agoI briefly worked at a YC company that was a ruby shop. Their answer to every performance problem was to stick it on a queue. There were, I don’t know, dozens of them. Then they decided they needed to be multi-region, because reasons. But the queues weren’t set up to be multi-region, so they built an entirely new service that’s job was to decide which queue in which region jobs needed to go on. So now you had jobs crisscrossing datacenters and tracking any issue became literally impossible. Massively turned me off to both that company and ruby in general.
- cortesoft 1y agoDon't let one company's misuse of a language turn you off the entire thing!
- teaearlgraycold 1y agoRuby and the age of “I don’t care what type this variable is, it quacks like a duck!” is over and dead. Improvements to type systems have shown there is a better way to do software development.
- ch4s3 1y agoMaybe, the cycle has happened before and maybe come back around again. Dynamic typing is really nice when most of your data looks like bags of strings. Compilers and tools just don’t add a lot when you’re passing around glorified blobs of stringy json-like stuff. Type gymnastics can eat a lot of time where you could otherwise be shipping something useful.
- riffraff 1y agoI would argue when you're passing around stringy JSON-like thingies is when typing is most useful :) You're not going to misuse an API that takes a Person or Cart, but mixing up two hashes cause you used two different strings as keys can happen easily. (I do think dynamic typing is mostly fine, but I do wish ruby had optional static typing with some nice syntax instead of RBS)
- Lio 1y agoIt’s on the way thankfully. I’m really excited about Sorbet getting behind the new RBS-inline comment syntax and the prospect of both runtime and static analysis as optional tools when needed.
- ch4s3 1y ago> You're not going to misuse an API that takes a Person or Cart, but mixing up two hashes cause you used two different strings as keys can happen easily. This is more or less trivial to catch and fix, I'm just not sure a type system is worth it's weight for that kind of case.
- jaredsohn 1y agoI haven't really felt a need for static typing in my job's Rails app (small team, and I've been working in this codebase a really long time) but I think LLMs can be a huge help for automating the type gymnastics.
- AstroBen 1y agoStatic typing isn't free. Dynamic typing is perfectly viable to build any sized software - there's living proof, far from dead So we're only left with personal opinion
- sbarre 1y agoOr the standard "it depends" answer that everyone eventually realizes is the only correct answer. ;-)
- RangerScience 1y agoTypescript is literally a language describing what the quack “sounds like” so you can attempt to ensure any particular variable makes those kinds of quacks. Typescript doesn’t care that it also quacks like a dog. Plus, Ruby has lots of easy ways for you to check typing, if you want to.
- trevorhinesley 1y agoNeither sticking everything into a queue nor going multi-region are Ruby’s fault.
- rubyfan 1y agoYeah this is not the fault of ruby. Sounds more like bad choices that could be made with any language or framework.
- charcircuit 1y agoCulture around a language influences what choices are made.
- trevorhinesley 1y agoI’ve never felt like “throw everything into a queue” was a mindset within the Ruby community, nor have we done that at my companies. And multi-region is a business decision.
- morkalork 1y agoDoesn't Ruby, like Python, have a GIL? I always found that one is enough to encourage some "premature scalable architecture"
- trevorhinesley 1y agoIt does have a GIL. You’re not wrong, but by that same logic, there’s pitfalls when using multi-threading as well, even in languages where it’s native (e.g., Elixir). Regardless, in my experience, when you run into scenarios that need queueing, multi-threading, etc., you need to know what you’re doing.
- Lio 1y agoThat depends on the Ruby implementation. MRI (CRuby) has a GVL which is why you might use a forking web server like Puma or Pitchfork. JRuby and TruffleRuby though have true multi-threading and no GVL. I’ve used the Concurrent Ruby library with JRuby and Tomcat quite a bit and find works very well for what I need.
- aaronblohowiak 1y ago
- throwaway493943 1y agoI'm going to be brave (but still use a throwaway) and ask the dumb question - what is wrong with putting things in queues to help with performance problems? If some endpoint is too slow to return a response to the frontend within a reasonable time, enqueueing it via a worker makes sense to me. That doesn't cover all performance issues but it handles a lot of them. You should also do things like optimize SQL queries, cache in redis or the db, perhaps run multiple threads within an endpoint, etc. but I don't see anything wrong with specifically having dozens of workers/queues. We have that in my work's Rails app. Happy to hear how I can do things better if I'm missing something.
- teyc 1y agoQueues have several problems - if the caller is http it may timeout and retry, leading to more jobs being queued - the caller may no longer care because it took so long and the work is wasted - if the caller is called from a queue it can cause cascades - you can fill a disk up and crash the system
- crowcroft 1y agoTo me the question is, so what's a better alternative? At least queues can be designed to handle timeouts, errors, and flakey APIs.
- rubyfan 1y agoYou’re not missing anything and are correct in that there are plenty of reasons to use queues and defer work that can be handled asynchronously outside of a request/response. This is not specific to ruby or any language for that matter. The parent indicated the cross region dynamic required extra routing logic and introduced debugging problems.
- zdragnar 1y agoThere's two primary areas that I've seen teams get bitten by this personally: 1) Designers don't understand that things are going to happen async, and the UI ends up wanting to make assumptions that everything is happening in real time. Even if it works with the current design, it's one small change away from being impossible to implement. This is a general difficulty with working in eventually consistent systems, but if you're putting something in a queue because you're too lazy to optimize (rather than the natural complexity of the workload demanding it) you're going to be hurting yourself unnecessarily. 2) Errors get swallowed really easily. Instead of being properly reported to the team and surfaced to a user in a timely manner, the default setting of some configurations to just keep retrying the job later means if you're not monitoring closely you'll end up with tens of thousands of jobs retrying over and over at various intervals.
- pmontra 1y agoWe were offloading to jobs every long running activity in the Elixir/Phoenix project I've been working on years ago. There is no other way. The response to a web request must complete in a short time and free the server for further requests. We solved debugging by sending all log lines to a centralized server. We were running on the Google cloud. We were not multiregion though. My current Rails project uses sidekiq a lot to send mail, generate PDFs, any activity that does not have to necessarily complete before we return the response. We keep the interactive web app up to date by websockets and with callbacks for clients using our public API. I don't think we would have done it differently in any other language. By the way, we built our slimmer version of sidekiq for Elixir because the language plus the OTP libraries have a lot of functionality but we still need to persist jobs, retry them even after a complete reboot, exponential back off, etc.
- yxhuvud 1y agoBad architecture can happen in any language. I don't see how the language choice could ever protect you against the described structural problem you built. Also you will see that the answer to most actual performance problems tend to be queues even in other languages. At least in mature places - mostly because it is possible to inspect what a queue is doing. Though it will of course be a problem if it is part of a big spaghetti architecture.
- deleted 1y ago[deleted]
- maineagetter 1y agoAny performance comparisons to Oban?
- pqdbr 1y agoGreat writeup. I'd love to know more about how the Supervisor works, and how it. "fork[s] a separate process for each supervised worker/dispatcher/scheduler". In a Rails app served with Puma, I've always had a hard time understanding what would be the canonical way for having a loop doing some periodic work. I know Puma has plugin support but I don't see much documentation there. Forking a process / threads is something that we're used having Rails / Puma take care for us. Pressed for time and without having time to deep dive, we ended up settling with sidekiq-cron, and it's been serving us so nicely.
- hschne 1y agoUnder the hood, it uses good ol' fork and keeps track of the generated process IDs. It's surprisingly simple. You can check out the relevant source here: https://github.com/rails/solid_queue/blob/main/lib%2Fsolid_queue%2Fsupervisor.rb#L77-L89 https://github.com/rails/solid_queue/blob/main/lib%2Fsolid_q...
- hschne 1y agoAuthor here. What a pleasant surprise to see this on HN! Happy to answer any questions.
- dzonga 1y agothanks for this write up. some of us are happy rails users so more rails content is always welcome
- j0rd72 1y agoThanks for the write-up, I feel like I understand Solid Queue quite well, now. I suppose my primary question is: What does this do better than Sidekiq+Redis; or, why should I convert my Sidekiq jobs to use Solid Queue? I'm curious also if there are comparisons of performance anywhere. All-in-all, though, it looks technically quite promising!
- zem 1y agoman, that brings back memories - very early in my career I tried to use postgres as a task queue, thinking that with O(hundreds) of jobs it wasn't worth setting up something like rabbitmq. sadly I knew pretty much nothing about db design and the performance was horrible, ended up ripping it out and installing rabbitmq after all (and having a whole new set of headaches with random rabbit admin issues but at least when it worked it was fast)