NATS and AWS SQS/SNS load testing: acknowledgements and end-to-end latency
In messaging systems "request time" is two different things: how long the sender waits for an acknowledgement, and how long the message takes to reach a receiver. An order service may write to JetStream in 3 ms while the handlers have fallen behind and the order is handled only 40 seconds later. The job of a NATS or SQS load test is to measure these two times separately and show which one breaks at which load.
What are we measuring?
- Publish acknowledgement: a core NATS publish waits for no acknowledgement; its time is writing to the client buffer and says little about the server. On JetStream the time runs to the stream's PubAck, on SQS to
SendMessage's answer: the moment the message is stored. - Request-reply: in NATS request-reply, the round trip, which is the responding service itself.
- End-to-end latency: the time from sending a message to receiving it. When a backlog starts to build, this is the first number to break; acknowledgement times may still look fine.
- Errors: requests nobody listens to, streams or queues that do not exist, credential and permission errors, AWS throttling and timeouts.
Shaping the load
In messaging the load is not "how many users" but "how many messages a second". Build the sending side with a fixed arrival rate, and the consuming side with as many VUs as there are real handlers. When the send rate exceeds the consuming capacity, end-to-end latency grows minute by minute; to see it, run the test for at least ten minutes rather than a few, and raise the rate in steps.
Common mistakes
- Using the real service's consumer or queue. A message acknowledged on a work-queue stream, or received and deleted on SQS, is gone: the test consumes the real service's messages. Use a consumer name and a queue of the test's own.
- Forgetting the SQS bill. SQS and SNS charge per request; a ten-minute test at a hundred messages a second is hundreds of thousands of requests. Use a separate account, or try the test first against LocalStack or ElasticMQ.
- Measuring latency across clocks. When the sender and the receiver are different machines, a few milliseconds of clock difference creep into the end-to-end latency. Measure it on messages sent and received on the same machine.
- Looking at the average. Queue latency has a long tail; look at p95 and p99 (the p95 and p99 guide).
With Spitfire: NATS
Add a NATS connection under Connections: servers, authentication (username and password, token, NKey seed or a .creds file; stored encrypted), TLS and, if needed, a JetStream domain. A runner's VUs share 4 connections by default; to test the connection limit, each VU can open its own. The NATS step has five actions:
- Publish: core publish; no acknowledgement.
- Wait for a message: the VU subscribes to the subject once (
orders.*,orders.>); with a queue group, messages are shared among the VUs. - Request/reply: the time is the round trip; with no listener the error is
no_responders. - Publish to JetStream: the time runs to the PubAck; with no stream for the subject the error is
no_stream. - Consume from JetStream: takes a message from a durable pull consumer and acknowledges it. The consumer is shared by the runner's VUs; when missing, one starting from new messages is created, and an existing one is used as it is.
With Spitfire: SQS and SNS
Add an AWS connection: a region and an access key (stored encrypted), or the machine's own credentials (environment variables, a profile, an IAM role). Only an installation admin can turn on the latter, since everyone using the connection acts with the installation's AWS identity. A custom endpoint is set for LocalStack, ElasticMQ or a VPC endpoint. The SQS step has three actions: Send (message group and deduplication ids on FIFO queues, a 0–15 minute delay, message attributes), Receive and delete with long polling (or leave the message; it returns to the queue when its visibility timeout ends) and Publish to SNS. Queues are named by name or URL.
End-to-end latency
Spitfire stamps the messages it sends: the Spitfire-Ts header on NATS, the spitfire-ts attribute on SQS. Receive steps measure nats_e2e_latency and sqs_e2e_latency for messages the same runner sent in the same run: both ends read the same clock, so no clock difference creeps in. For SNS delivering to an SQS queue, the latency is still measured when raw message delivery is on. You write thresholds on these metrics like on any other.
A 10-minute test where 200 orders a second are published to JetStream, an invoice for each is put on an SQS queue, and 20 handlers consume both. The JetStream acknowledgement's p95 must stay under 20 ms, NATS end-to-end latency's p95 under 50 ms and SQS's under 1 second. The test was checked with spitfire validate.
{
"name": "Sipariş akışı: NATS ve SQS",
"scenarios": [
{ "name": "siparisler",
"executor": { "type": "constant-arrival-rate", "rate": 200, "timeUnit": "1s",
"duration": "10m", "preAllocatedVUs": 50, "maxVUs": 200 },
"steps": [
{ "id": "place", "name": "Siparişi yayınla", "protocol": "nats", "connection": "bus",
"nats": { "action": "jsPublish", "subject": "orders.{{__VU}}", "payload": "{\"vu\": {{__VU}}}" } },
{ "id": "queue", "name": "Faturayı kuyruğa at", "protocol": "sqs", "connection": "aws-test",
"sqs": { "action": "send", "queue": "invoices-loadtest", "body": "{\"order\": {{__ITER}}}" } } ] },
{ "name": "isleyiciler",
"executor": { "type": "constant-vus", "vus": 20, "duration": "10m" },
"steps": [
{ "id": "handle", "name": "Siparişi al", "protocol": "nats", "connection": "bus",
"nats": { "action": "jsConsume", "subject": "orders.>", "stream": "ORDERS", "consumer": "spitfire", "wait": "5s" } },
{ "id": "take", "name": "Faturayı al", "protocol": "sqs", "connection": "aws-test",
"sqs": { "action": "receive", "queue": "invoices-loadtest", "wait": "10s" } } ] }
],
"thresholds": [
{ "metric": "req_duration", "filter": { "step": "place" }, "expr": "p(95)<20" },
{ "metric": "nats_e2e_latency", "expr": "p(95)<50" },
{ "metric": "sqs_e2e_latency", "expr": "p(95)<1000" },
{ "metric": "req_failed", "expr": "rate<0.001" }
]
}Spitfire installs on Docker or Kubernetes with one command; every testing feature and protocol is open in the free edition.