Distributed load testing on Kubernetes: runners, scaling and multiple locations
Past a certain load a single machine starts measuring itself instead of the system under test: the CPU fills up, sockets and ephemeral ports run out, the network card saturates and latency grows inside the load generator. Distributed load testing splits the load across several machines (runners) and merges the results into one run. Kubernetes is a natural place for this: runners are pods, their number changes with one command, and they run close to the system under test.
When do you need distributed load?
- When the load generator saturates: latencies measured while the runner's CPU is high cannot be trusted. The rule is simple: if the machine generating the load is struggling, you are seeing the machine, not the result.
- With many connections: every virtual user is at least one socket; tens of thousands of WebSocket or MQTT connections run into one machine's file descriptor and port limits.
- From several locations: if your users are in different cities or cloud regions, generating load from there brings network latency and CDN/load balancer behaviour into the result.
Common mistakes
- Averaging percentiles. The average of two runners' p95s is not the p95 of all requests. The right result is computed by merging every runner's measurements (the p95 and p99 guide).
- The same nodes as the system under test. If runner pods run on the same nodes as the application's pods, the two compete for CPU and network; the test distorts its own result. Put the runners on a separate node pool or a separate cluster.
- Throttling at the CPU limit. Kubernetes throttles a pod that hits its CPU limit, and that adds latency to requests inside the runner. Give runners enough CPU and watch their usage through the run.
- An egress bottleneck. If traffic leaving the cluster goes through a NAT gateway or a single egress IP, connection tracking tables and port counts limit the load; on the target side all the load seems to come from one IP and may hit rate limits.
With Spitfire
Spitfire's install script puts everything into the current kubectl context: PostgreSQL, the controller (web UI and API) and the runners, with network policies between them. The controller is a single pod; the runners are a Deployment whose size is set with --runners at install and changed later with kubectl scale (below: an install with 3 runners, then scaling to 8). Every component runs from one image (amd64 and arm64). All the options, and the update and backup steps, are in the installation guide.
curl -fsSL https://spitfire.tr/install.sh | bash -s -- kubernetes --namespace loadtest --runners 3
kubectl -n loadtest scale deploy/spitfire-runner --replicas=8- The load is split exactly. When you start a run you pick runners by count, by label, one by one or by location; the load is split across them, and the total VU count and arrival rate stay exactly as defined. In arrival-rate models each iteration belongs to exactly one runner, so no work is done twice.
- The result is one run. Runners send their measurements to the controller every second; the controller merges them and computes percentiles such as p95 and p99 over every measurement, never averaging the runners' values. The results page also breaks them down by step and by location.
- Pods per run. Instead of keeping runners up all the time you can start them for the run (
--runners 0 --on-demand-runners 20): when a run picked by count asks for more runners than are idle, the rest start as pods of the same image, the run waits in the queue until they connect (usually 20–60 seconds), and the pods are deleted when it ends; at most N pods at once. The Run dialog suggests a runner count from the test's peak VUs (one pod per 1,000 VUs). The controller gets permission for pods in its own namespace only. - Sizing. Runner pods are limited to 2 CPUs and 1 GiB of memory by default. A runner warns when its file descriptor limit is too low for its VU count, and the findings after a run point out a saturated load generator too. Tune VUs per pod to the protocol and the think time: HTTP users with long think times and an arrival rate that never waits use the same CPU very differently.
Multiple locations
Runners in another cluster, data centre or cloud region connect to the controller outbound, not inbound: over TLS, pinning the controller's key. No port is opened in the runner's network; it only needs a connection out to the controller's gRPC port. The Runners page gives the registration key and the install command; a key can be bound to a location.
curl -fsSL https://spitfire.tr/install.sh | bash -s -- runner \
--controller spitfire.example.com:8471 --token sfrun_… --ca-pin sha256:… --runners 2When you start a run you spread the load over locations by percentage (Istanbul 60%, Frankfurt 40%, for example). The results page shows each location's request count, rate, p95, p99 and errors separately; a threshold can target one location ("filter": { "location": "frankfurt" }). If one location's p95 is clearly higher than another's, the findings after the run say so. How many runners and locations a run can use depends on the license (pricing).
With Spitfire on Kubernetes the data stays on your infrastructure: tests, results and connection passwords are in the PostgreSQL in your cluster, and runners connect to the system under test directly (self-hosted load testing). To find the load where the system breaks, run a breakpoint test with distributed load.
Spitfire installs on Docker or Kubernetes with one command; every testing feature and protocol is open in the free edition.