Skip to content

Performance

Six Redis-backed Python queues, one machine, one sitting, three workloads. This page is the long version: what was run, what came out, and where the numbers flatter us and shouldn’t.

The suite is open and reproducible, including the row where we lose: python-task-queue-benchmarks.

  • 1 worker process, 10 concurrent tasks, 3 interleaved rounds — reported as mean ± std.
  • Rounds are interleaved and the starting library rotates, so no library sits at the same position twice and machine drift is spread across all six rather than landing on whoever runs last.
  • Three scenarios:
    • io_task — 1,000 tasks, each a 100 ms sleep (asyncio.sleep for async libs, time.sleep for sync).
    • cpu_task — 1,000 tasks, each 1,000 SHA-256 hashes over 1 KiB inputs.
    • noop_task — 20,000 tasks that do nothing, so what is measured is the queue’s own cost to move a task and nothing else.
  • Enqueuing is timed separately from processing, so a fast producer is never credited to the worker’s throughput.
  • Machine: 8-core / 16-thread x86-64, 15 GB RAM, CPython 3.13, Redis 7.4.
  • Versions: arq 0.28, Taskiq 0.12.4, Streaq 6.5.0, Celery 5.5.3, Dramatiq 2.1.0, and ArdiQ from development — the producer fixes behind these numbers ship in the next release; 0.5.0 from PyPI measures about 12% lower on CPU work.

A task that does nothing takes the work out of the measurement and leaves the queue. This is the scenario the Rust core is built for.

QueueThroughput (tasks/s)Latency p99Memory
ArdiQ 🦀2,895 ±496,820 ms ±12734 MB 🪶
Taskiq2,051 ±999,292 ms91 MB
Dramatiq1,787 ±7310,492 ms56 MB
Streaq1,179 ±5615,626 ms52 MB
Celery861 ±3921,885 ms50 MB
arq789 ±5223,500 ms30 MB

41% ahead of the next queue, and the tightest tail in the field. Read those two together: the p99 is not just lower, it is lower by 26% while ArdiQ’s round-to-round spread stays at ±127 ms.

As milliseconds of overhead per task, which is what this really measures:

QueueOverhead per task
ArdiQ0.345 ms
Taskiq0.488 ms
Dramatiq0.560 ms
Streaq0.848 ms
Celery1.161 ms
arq1.268 ms

ArdiQ crosses the Rust/Python boundary twice per task — once to start it, once to return its result — and still spends less per task than the pure-Python queues that never cross anything. That is the loop and the Redis I/O being off the GIL.

The cpu_task body hashes under the GIL, so for every single-process queue the task body is serial on one core. What varies is the framing around it.

QueueThroughput (tasks/s)Latency p99Memory
ArdiQ 🦀394.3 ±4.72,480 ms ±5133 MB 🪶
Taskiq356.4 ±18.52,185 ms ±1,03291 MB
Streaq322.3 ±3.43,032 ms ±2048 MB
arq282.9 ±6.13,466 ms ±7530 MB
Celery12.574,909 ms51 MB
Dramatiq12.575,781 ms55 MB

11% faster than Taskiq, on a third of its memory. The p99 column is a tie worth being honest about: Taskiq’s mean tail is lower here, but it moves by ±1,032 ms between rounds against ArdiQ’s ±51, and the two swap places from one run to the next. Call it even.

With 1,000 tasks at concurrency 10 and a 100 ms sleep, the arithmetic ceiling is 100 tasks/s. Everything near it is waiting on the network, not on the queue.

QueueThroughput (tasks/s)% of ceilingMemory
Taskiq96.997%91 MB
ArdiQ 🦀95.395%34 MB 🪶
Dramatiq94.194%56 MB
Streaq91.892%48 MB
arq87.688%30 MB
Celery71.271%51 MB

Second, by 1.7%, in the one scenario where there is almost nothing left to win. Both queues are up against the sleep, not against each other — the interesting column here is the third one.

How long the producer is blocked staging work — the number that matters when the thing enqueuing is a web request.

QueueTasks staged per second20,000 tasks take
ArdiQ 🦀81,6360.2 s
Streaq27,4330.7 s
Dramatiq2,7117.4 s
Taskiq2,5038.0 s
Celery1,88610.6 s
arq99920.0 s

ArdiQ and Streaq stage batches through a bulk API (enqueue_many); the other four await one round trip per task, which is their only option — except Celery, which has group(...).apply_async() that this suite does not use. The clean comparison is ArdiQ against Streaq: 3× faster, both batching. Against the rest, part of the gap is our API and part is how the suite calls theirs.

QueueDispatch (tasks/s per MB)CPU (tasks/s per MB)I/O (tasks/s per MB)
ArdiQ 🦀87.211.92.84
Dramatiq31.80.231.69
arq26.09.32.88
Streaq22.76.61.90
Taskiq22.53.91.06
Celery17.30.251.39

On dispatch, ArdiQ does 2.7× the work per megabyte of the next queue, and 3.9× Taskiq’s. arq edges it on I/O by under 2% — the two are tied there, and arq is 33% slower on CPU work and 3.7× slower on dispatch.

  • ⚡ 41% more tasks a second than the next queue on dispatch, 11% more on CPU work.
  • 🪶 33 MB per worker — a third of Taskiq’s, two thirds of Streaq’s. Only arq is lighter, by 3 MB, and it is last or second-to-last on all three.
  • 🎯 The tightest tail in the field on dispatch — p99 6,820 ms ±127 against 9,292.
  • 📈 95% of the arithmetic ceiling on I/O work, 1.7% off the lead.
  • 🚀 0.2 s to stage 20,000 tasks, against 8 s for the next queue with a comparable API.
  • The GIL caps in-process CPU work for every Python queue, ArdiQ included. Your task body is serial per worker. Scale CPU-bound work out with more worker processes.
  • Never compare across runs or machines. Only the libraries within one table are comparable. On this machine the same unchanged code has measured 5–8% apart on different evenings, which is larger than several of the gaps above.
  • The enqueue table is not apples-to-apples for four of the six. See the note under it.
  • CPU parallelism isn’t measured here. All libraries run one worker; this suite measures per-task overhead, not multi-core scaling.
  • These numbers are ArdiQ’s development build. The fixes behind them are committed and tested but not yet released; PyPI’s 0.5.0 measures about 12% lower on CPU work and roughly half on dispatch.

Raw per-iteration samples, the full methodology, and every scenario we run live in the benchmark repo.

Made bytay.dev