7 ms·
A Celery-like Python Task Queue in 55 Lines of Code
- cschmidt 13y agoHaving a way to pickle code objects and their dependencies is a huge win, and I'm angry I hadn't heard of PiCloud earlier. That's a nice use of the cloud library, without using the PiCloud service. Unfortunately, the PiCloud service itself is shutting down on February 25th (or thereabouts).
- jknupp 13y agoAh, I'm sorry to hear that. Hopefully, they make their library code more readily available before closing their doors.
- caidan 13y agoLooks like the PiCloud team is joining dropbox but according to their blog post "The PiCloud Platform will continue as an open source project operated by an independent service, Multyvac (http://www.multyvac.com/ http://www.multyvac.com/)." Source: http://blog.picloud.com/2013/11/17/picloud-has-joined-dropbox/ http://blog.picloud.com/2013/11/17/picloud-has-joined-dropbo...
- cschmidt 13y agoYes, the Multyvac launch has been delayed until February 19. They are supporting much of the PiCloud functionality, but not function publishing, which I used quite a lot. (That was a way you could "publish" a function to PiCloud, and then call it from a RESTful interface. It was a nice way to decouple my computational code from my website, which has very different dependencies.) I fear it is more oriented toward the use case of doing long running scientific jobs, rather than short Celery-like jobs. I hope for the best.
- est 13y agohttp://docs.python.org/2/library/multiprocessing.html#sharing-state-between-processes http://docs.python.org/2/library/multiprocessing.html#sharin... Why don't anyone build Celery and Redis alternative using this?
- Iftheshoefits 13y agoI can't speak for Celery, as I've not used it very much. It may be straightforward to write a functional Redis alternative in pure python using this library, but I would certainly have questions/concerns about performance.
- SaberTail 13y agoCelery uses `billiard`, which is a fork of multiprocessing. That doesn't help with communicating with a distributed worker pool, though.
- mzs 13y agoThere have been a few times I did things in python that ended-up being a terrible PITA, multiprocessing was one of those. Basically after fork, you should call exec, but multiprocessing doesn't. So many things worked just fine in Linux and then completely fell apart on FreeBSD, OSX, and Windows. I think a lot of this has been fixed since then by using a Manger and a Pool.
- agentultra 13y agoI'd recommend looking at alternative serialization formats. Pickle is a security risk that programmers writing distributed systems in Python should be educated about.
- jonesetc 13y agoI understand the risk is basically because you're evaling when unpickling. What formats are safe then?
- jknupp 13y agoPickle (or any use of eval) is a security risk only if you're using it in the context of untrusted code. Basically any distributed task queue is going to have that risk if it can execute arbitrary code.
- michaelmior 13y agoPickle doesn't really use eval, but there is still the potential for users to execute arbitrary code[1]. JSON, YAML, MessagePack, etc are safe in this respect (assuming a well-implemented parsing library) because all the parser does is convert the data into simple data structures. [1] http://lincolnloop.com/blog/playing-pickle-security/ http://lincolnloop.com/blog/playing-pickle-security/
- jonesetc 13y ago
- m0th87 13y agoFor an alternative, check out RQ: http://python-rq.org/ http://python-rq.org/ We use it in production and it's been rock-solid. The documentation is sparse but the source is easy to follow.
- marban 13y agoIt's great but as far as i can remember doesn't support P3. Also, it'd be nice to use it with Mongo instead of Redis.
- jmagnusson 13y agoActually Py3-support landed 6 months ago https://github.com/nvie/rq/pull/239 https://github.com/nvie/rq/pull/239
- volker48 13y agoThis is the first thing I thought of when I saw this post. I use RQ extensively and it's great.
- tonymillion 13y agoAlthough Celery can use it, why is Amazon SQS treated as a second class citizen in python background worker systems? I've yet to find/see a background worker pool that played nicely (properly) with SQS.
- jbaiter 13y agoAre there any non-distributed task queues for Python? I need something like this for a tiny web application that just needs a queue for background tasks that is persisted, so tasks can resume in case the application crashes/restarts. Installing Redis or even ZeroMQ seems kind of excessive to me, given that the application runs on a Raspberry Pi and serves maximum 5 users at a time.
- asksol 13y agoI wouldn't really classify this example as a "task queue", it resembles more the RPC pattern really (and I would guess that this example does not persist the jobs in any way). Celery does have a very experimental filesystem based transport that could be used for your use case. I don't know if it fits on a Raspberry PI but celery does not have very high memory/space requirements (c.f. other Python libraries).
- jknupp 13y agoZeroMQ takes care of the queueing. Though I didn't delve into it in depth in this example, you can create quite sophisticated broker-less distributed systems pretty easily with ZeroMQ.
- asksol 13y agoRight, I understand that ZeroMQ is used to send and receive messages, what I mean is that there's no persistency involved so the tasks will not survive a system restart.
- rch 13y agoI like these one-off projects that Jeff is doing, but it would be particularly instructive to see one, or a combination, make it to 'real' status.
- jknupp 13y agoCheckout sandman: www.sandman.io or www.github.com/jeffknupp/sandman
- thruflo 13y agoI scratched an itch in this space to create, in Python, a web hook task queue. I wrote it up here http://ntorque.com http://ntorque.com -- would love to know if the rationale makes sense...
- dangayle 13y agoThanks Jeff. As someone else mentioned, I love these little projects that demonstrate the basics of what the big projects actually do. Makes it much easier to understand the big picture.
- dmunoz 13y agoAbsolutely. I'm always pleased when documentation includes some pseudocode for what the system generally does, without the overhead of configuration, exceptional control flow, etc. It's not always possible with large systems, but makes it a lot easier to see the forest, not the trees, in even mid-sized code bases.