MQTT load testing: thousands of devices, connections, QoS and message latency
What stresses an MQTT broker is often not the message rate but the number of connections: tens of thousands of devices, each sending a small message every few seconds, mean tens of thousands of sessions, subscriptions and keep-alives the broker holds in memory. An MQTT load test builds that picture and looks for answers to three questions: how many concurrent devices does the broker accept, how long does a message's acknowledgement take at that load, and what happens when every device reconnects at once after an outage.
What are we measuring?
- Concurrent connections: the number of clients connected to the broker at once. Each connection costs memory, a file descriptor and keep-alive traffic; the limit often comes here before CPU.
- Connect rate: how many new connections are accepted per second. The TLS handshake and authentication are expensive; thousands of devices reconnecting together after a network outage (a reconnect storm) is a bottleneck that a steady load never shows.
- Publish latency: it depends on QoS. At QoS 0 the client writes the message and moves on without an answer from the broker; at QoS 1 you measure the time to the broker's PUBACK, at QoS 2 to the PUBCOMP at the end of a four-step handshake. Use the QoS your devices really use.
- Fan-out: a message is copied to every matching subscriber. A few services subscribed to a broad filter such as
devices/+/telemetrymultiply the messages the broker sends. - Errors: refused connections, authentication failures, timeouts and sessions the broker closes.
Shaping the load
In MQTT the natural model is "one virtual user per device": each VU opens its own connection, sends a message at intervals and stays connected. The message rate follows from the device count and the interval: 500 devices sending every 5 seconds is 100 messages a second. Give the interval a little randomness (say 4–6 s) instead of a fixed value; otherwise every device sends at the same moment and you get waves that do not happen in reality.
The ramp-up sets the connect rate. Reaching 500 devices in 2 minutes is about 4 connections a second; reaching the same number in 10 seconds is a reconnect storm test. Run them as separate tests: one shows the broker under steady load, the other how it recovers after an outage (spike testing). To find the device count where the broker breaks, raise it in steps (breakpoint testing).
Common mistakes
- The same client id. In MQTT a second client connecting with the same client id drops the first one's session. Give every device its own id; if you run two tests against the same broker at once, give their ids different prefixes.
- Many messages from few connections. Sending 10,000 messages a second over 10 connections stresses the broker very differently from ten thousand devices. Keep the device count close to reality and set the rate with the interval.
- Retained messages and persistent sessions. Messages the test sends with
retainand sessions without clean session stay on the broker after the test. Use a separate topic tree for the test and plan the cleanup. - The load generator's limit. Every connection is a socket on the machine generating the load too. For thousands of devices raise the file descriptor limit and spread the load over several machines; otherwise you measure the load generator, not the broker.
- Average latency. Broker latency has a long tail; look at p95 and p99 (the p95 and p99 guide).
With Spitfire
First add an MQTT connection under Connections: broker addresses (tcp://host:1883, ssl://host:8883 or ws://host:8083/mqtt), client id, username and password, keep-alive (30 s by default), clean session (on by default) and TLS if needed (CA, client certificate). The password is stored encrypted and the test names only the connection; a password written into the address is refused. The client id is a template, spitfire-{{__VU}} by default: the VU number is unique across every scenario and runner of a run.
Every VU keeps its own MQTT client, just like a device. The client connects on the first step and stays connected until the VU stops; connecting is not part of the step's duration, but a failed connect fails the step. If the connection drops, the VU reconnects on its next step and renews its subscriptions. The MQTT step has three actions:
- publish: sends the payload to the topic at QoS 0, 1 or 2, with
retainif you like. The step's duration is writing the message at QoS 0, the time to PUBACK at QoS 1 and to PUBCOMP at QoS 2. Topic and payload can use{{__VU}},{{$randInt 18 30}},{{$timestamp}}or variables from a CSV data file. - subscribe: the VU subscribes to the topic filter once (
+and#are valid only here) and takes one incoming message per step. The step's duration is the wait for the next message, not the time from publish to receipt; when the wait (10 s by default) runs out, the step fails with a timeout. Incoming messages wait in a buffer of 256 messages per VU and filter, and the oldest is dropped when it is full. The payload goes to checks and extraction (with JSONPath when it is JSON);topic,qos,retained,duplicateandmessage-idcan be read as headers. - request: for request-reply: the VU subscribes to
responseTopic, sends the message and takes the first message on that topic as the reply. The duration is the round trip. There is no correlation id; the responding service has to know which topic to answer on, and each VU should get its own reply topic (such asreplies/{{__VU}}).
A test where 500 devices connect over 2 minutes, then for 10 minutes each sends telemetry at QoS 1 every 4–6 seconds (about 100 messages a second), while two listeners watch every device's topic. The p95 of the publish acknowledgement must stay under 100 ms and the error rate under 0.1%. The test was checked with spitfire validate.
{
"name": "MQTT: cihaz telemetrisi",
"scenarios": [
{ "name": "cihazlar",
"executor": { "type": "ramping-vus", "startVUs": 0,
"stages": [ { "duration": "2m", "target": 500 }, { "duration": "10m", "target": 500 },
{ "duration": "30s", "target": 0 } ] },
"steps": [ { "id": "telemetry", "name": "Telemetri gönder", "protocol": "mqtt", "connection": "broker",
"mqtt": { "action": "publish", "topic": "devices/{{__VU}}/telemetry", "qos": 1,
"payload": "{\"device\":{{__VU}},\"temp\":{{$randInt 18 30}},\"ts\":{{$timestamp}}}" },
"thinkTime": { "min": "4s", "max": "6s" } } ] },
{ "name": "dinleyici",
"executor": { "type": "constant-vus", "vus": 2, "duration": "12m30s" },
"steps": [ { "id": "listen", "name": "Telemetriyi dinle", "protocol": "mqtt", "connection": "broker",
"mqtt": { "action": "subscribe", "topic": "devices/+/telemetry", "qos": 1, "wait": "5s" },
"checks": [ { "type": "jsonPath", "path": "$.temp", "op": "exists" } ] } ] }
],
"thresholds": [
{ "metric": "req_duration", "filter": { "step": "telemetry" }, "expr": "p(95)<100" },
{ "metric": "req_failed", "expr": "rate<0.001" }
]
}During the run you watch per-step message rates, p95, p99, data sent and received, and error kinds (authentication, refused connection, timeout) live; the listener's checks confirm the shape of incoming messages. Instead of opening thousands of devices from one machine, spread the load over several runners: Spitfire splits the VUs across runners and the total device count stays as you defined it. If you would rather fix the message rate, you can also build the device scenario with an arrival-rate executor; then a VU is a pool of senders rather than a device.
Spitfire's MQTT client speaks MQTT 3.1.1; features specific to MQTT 5 (user properties, the response topic and correlation data properties) are not available in the test.
Spitfire installs on Docker or Kubernetes with one command; every testing feature and protocol is open in the free edition.