RabbitMQ load testing: publisher confirms, consumer throughput and queue depth
In RabbitMQ trouble usually starts quietly: consumers cannot keep up, the queue grows, the broker approaches its memory or disk limit and slows the publishers down, and the slowness then shows in the services that publish. A RabbitMQ load test pushes the message path at a realistic rate and asks: at what confirm latency does the broker accept this rate, can the consumers consume at the same rate, and does the queue depth stay flat or keep growing.
What are we measuring?
- Publish latency: with publisher confirms on, the time from sending a message to the broker's confirm. Persistent messages and quorum queues are slower because they are written to disk and to replicas before the confirm; use the production settings in the test.
- Consumer throughput: messages consumed and acknowledged per second. The prefetch value and the ack mode set this rate directly.
- Queue depth: the number of messages waiting in the queue. A small, steady depth is normal; a depth that keeps growing under load means the consumers cannot keep up, and eventually leads to the broker's memory or disk alarm.
- Flow control: when the broker struggles it throttles connections, and under an alarm it blocks publishers entirely. This shows as a sudden jump in publish latency.
- Errors: unroutable messages (that reach no queue), messages the broker rejects (nack), permission errors and timeouts.
Shaping the load
On the publisher side the target is a rate ("300 orders a second"), so use a constant arrival rate: the send rate does not drop when the broker slows down, and the latency shows as it is. In a closed model with a fixed number of VUs the publishers slow down with the broker and the problem hides as a quiet drop in throughput (load test types).
On the consumer side match the count to your production consumer instances. If the point is to see whether the queue stays in balance, try a publish rate a little below and a little above the consumers' real capacity in two runs; to find the rate at which consumers fall behind, raise the rate in steps (breakpoint testing). To see whether the depth grows, the load has to stay steady for at least 10–15 minutes.
Common mistakes
- Publishing without confirms. Without publisher confirms a publish is almost instant, but you do not know whether the broker took the message; the test looks faster than it is. If production uses confirms, use them in the test.
- Auto-ack and unlimited prefetch. These two flatter the consumer and lose messages when a consumer crashes. In the test mimic the prefetch and ack mode your service uses.
- A queue without consumers. A test that only publishes grows the queue without end and can push the broker into a memory alarm. Add a consumer scenario, or give the queue a length limit and a TTL.
- A shared queue. Do not send test messages to a queue that production consumers read. Use a separate exchange, queue or vhost for the test.
- Average latency. Disk writes and replica sync cause rare but large spikes; look at p95 and p99 (the p95 and p99 guide).
With Spitfire
First add a RabbitMQ (AMQP) connection under Connections: the address (amqp://host:5672/vhost; the vhost is the address's path), username, password and TLS if needed (turning it on makes the address amqps). The password is stored encrypted and the test names only the connection; a password written into the address is refused. By default the VUs of each runner share one AMQP connection and every VU opens its own channels, like the instances of a service; the Separate connection per VU option gives every VU a connection of its own. Connecting and opening channels are not part of the step's duration. Spitfire does not declare exchanges or queues: the exchanges, queues and bindings the test uses must already exist. The AMQP step has three actions:
- publish: sends a message to an exchange with a routing key; body, content type, headers and persistence (
persistent) can be set. The channel is in confirm mode: the step's duration is the time to the broker's confirm. The message is sentmandatory, so a message that reaches no queue counts as a failure, and so does one the broker rejects. Values can use{{$uuid}},{{$randInt 10 5000}}or variables from a CSV data file. An empty exchange is the default exchange: the routing key is the queue's name. - consume: the VU joins the queue as a consumer with prefetch 1 and takes one message per step; the message is acknowledged after the measurement. The step's duration is the wait for the next message; when the wait (10 s by default) runs out, the step fails with a timeout. The body goes to checks and extraction (with JSONPath when it is JSON);
exchange,routing-key,redelivered,content-type,correlation-id,message-idand the message's own headers can be read as headers. - rpc: for request-reply: sends the message with RabbitMQ's direct reply-to and a correlation id per request, and waits for the reply with the same correlation id. The duration is the round trip; this action does not wait for a publisher confirm.
A test that publishes 300 order events a second, persistent, to the orders exchange while 10 consumers consume the orders.loadtest queue; the p95 of the publish confirm must stay under 50 ms and the error rate under 0.1%. The exchange, the queue and the binding between them must be created first. The test was checked with spitfire validate.
{
"name": "RabbitMQ: sipariş kuyruğu",
"scenarios": [
{ "name": "yayinci",
"executor": { "type": "constant-arrival-rate", "rate": 300, "timeUnit": "1s",
"duration": "10m", "preAllocatedVUs": 20, "maxVUs": 100 },
"steps": [ { "id": "publish", "name": "Sipariş yayınla", "protocol": "amqp", "connection": "rabbit",
"amqp": { "action": "publish", "exchange": "orders", "routingKey": "order.created",
"persistent": true, "contentType": "application/json",
"body": "{\"orderId\":\"{{$uuid}}\",\"amount\":{{$randInt 10 5000}}}" } } ] },
{ "name": "tuketici",
"executor": { "type": "constant-vus", "vus": 10, "duration": "10m" },
"steps": [ { "id": "consume", "name": "Sipariş tüket", "protocol": "amqp", "connection": "rabbit",
"amqp": { "action": "consume", "queue": "orders.loadtest", "wait": "5s" },
"checks": [ { "type": "jsonPath", "path": "$.orderId", "op": "exists" } ] } ] }
],
"thresholds": [
{ "metric": "req_duration", "filter": { "step": "publish" }, "expr": "p(95)<50" },
{ "metric": "req_failed", "expr": "rate<0.001" }
]
}During the run you watch per-step publish and consume rates, p95, p99, data sent and received, and error kinds (permission, refused connection, unroutable or rejected message, timeout) live. The consume step's duration is not end-to-end latency: when messages pile up in the queue the next one is already waiting and the duration gets short. So consume durations suddenly getting shorter can be a sign that the queue has started to grow.
Spitfire does not measure queue depth itself. If you already collect queue and broker metrics with RabbitMQ's Prometheus plugin, add those queries to a Prometheus connection under Observability: the run page's Backend tab shows them aligned with the load stages (OpenTelemetry root cause). Spitfire's client speaks AMQP 0-9-1; AMQP 1.0 and the stream protocol are not supported.
Spitfire installs on Docker or Kubernetes with one command; every testing feature and protocol is open in the free edition.