Load, stress, spike and soak testing: the differences and when to use which
"Performance testing" is not one test but a family of tests that answer different questions. Do we carry the expected traffic? Where do we break? Do we recover after a sudden campaign spike? Is there a memory leak after hours of running? Each question needs a different load shape. This guide covers the four basic types, the question each answers and how to set each up.
Two ideas first: VUs and request rate
You can describe load in two ways. In the virtual user (VU) model a fixed number of users loop through the scenario; when the system slows down each user sends fewer requests, so the load drops by itself (a closed model). In the request rate model you set how many iterations start each second; when the system slows down new requests keep arriving at the same rate (an open model). Real users keep arriving without knowing the system is slow, so for capacity questions the rate model usually gives a more honest answer. When the user journey and session behaviour matter, the VU model is more natural.
In Spitfire you choose them with the executor type: constant-vus and ramping-vus are the VU model, constant-arrival-rate and ramping-arrival-rate the rate model.
Load test: does it carry the expected traffic?
A load test runs the system at the normal or peak-hour traffic you expect and checks whether the result meets your targets. The question is not "at how many users does it break" but "at the load we expect, is p95 under 500 ms and the error rate under 1%". It is the test you run most often: before a release, in CI, in scheduled nightly runs.
Typical shape: a few minutes of ramp-up, a long enough hold at the target load, and a ramp-down. The ramp-up matters: connection pools, caches and JIT compilation behave differently in the first minutes, and counting the first seconds skews the result. Staying 10–20 minutes at the target load is enough to separate short fluctuations from the result.
"executor": { "type": "ramping-vus",
"stages": [
{ "duration": "5m", "target": 300 },
{ "duration": "20m", "target": 300 },
{ "duration": "5m", "target": 0 }
] }Define success with thresholds such as p(95)<500 and rate<0.01. Without thresholds a load test produces a chart, not a decision. How to read percentiles: the p95 and p99 guide.
Stress test: what happens beyond the expected?
A stress test pushes the load above the expected level and looks at how the system behaves under pressure. You learn two things: where the limit is and how the system fails past it. A good system degrades gracefully: latency grows, some requests get 429 or 503, and it recovers when the load drops. A bad one fails in a cascade: timeouts trigger retries, queues fill up, and the service does not come back even after the load is gone.
A specific, measurable form of stress testing is the breakpoint test: the load climbs step by step, each step is judged against the criteria, and the test stops at the first step that fails. The result is one number: the highest rate at which the system met its criteria. More: the breakpoint testing guide.
Spike test: what happens under a sudden surge?
A spike test drives the load very high within seconds, holds it briefly and brings it back down. A campaign launch, traffic after a push notification, half-time in a match or a sudden wave after a news story is the question here. What to look at: does autoscaling kick in fast enough, do connection pools and queues hold, and how long does latency take to return to normal after the spike?
The rate model fits a spike better, because in the VU model the load drops by itself as the system slows down and softens the spike. Leave some time at normal load before and after it; that is the only way to see the recovery.
"executor": { "type": "ramping-arrival-rate", "startRate": 100, "timeUnit": "1s",
"preAllocatedVUs": 100, "maxVUs": 1000,
"stages": [
{ "duration": "5m", "target": 100 },
{ "duration": "30s", "target": 1000 },
{ "duration": "3m", "target": 1000 },
{ "duration": "30s", "target": 100 },
{ "duration": "10m", "target": 100 }
] }Soak test: what happens after hours?
A soak (endurance) test holds a normal or slightly higher load for a long time, for hours. It catches what short tests cannot see: memory leaks, disks or logs filling up, connections that are never closed, growing tables and indexes, cache behaviour that degrades over time, hourly or daily batch jobs colliding with live traffic. The symptom is usually a slow slope: p95 at 180 ms in the first hour and 260 ms in the fourth.
In a soak test look not only at the load side but at the system's resources: memory, connection count, disk and queue depth. What matters is less the overall p95 than its trend over time. Run length depends on the license: 15 min on the free edition, 4 h on the Growth plans, 8 h on Scale, unlimited on Enterprise.
Which one, when?
| Type | Question | Typical length |
|---|---|---|
| Load test | Does it meet targets at the expected traffic? | 10–30 min |
| Stress / breakpoint | Where is the limit, how does it fail past it? | 15–60 min |
| Spike | Does it absorb a surge and recover? | 10–20 min |
| Soak | Does it degrade after hours? | 2–8+ h |
In practice the order is usually: a load test at the expected load first (the baseline), then a breakpoint test to see where the limit is, a spike test before a campaign or launch, and a soak test before major releases. Put the load test in CI or a nightly schedule and regressions reach you before your users: performance testing in CI/CD.
Whatever the type: four practical rules
- Know the environment. A result from an environment smaller than production is not production's capacity; write the differences (pod count, database size, caches) next to the result.
- Use realistic data. A test with the same user and the same product every time is served from cache. Use different users from a CSV and random ids.
- Watch the load generator. If the runner's CPU is saturated, the latency you measure is the load generator's, not your system's.
- Change one thing at a time. Change both code and configuration between two runs and you cannot tell where the difference came from.
With Spitfire
In Spitfire all four types use the same test definition; only the load model changes. In the web editor you pick the scenario's executor (ramping-vus, ramping-arrival-rate…) and stages, while the steps and thresholds stay the same. Several scenarios can run side by side: for example a constant background load and a spike scenario that starts at minute four (startTime). The "Constant request rate and a spike" file on the example tests page is exactly that.
The Run dialog runs the same test at a different scale (50% = half the VUs and rates) or against another environment; you do not write a separate test for stress. Find the breaking point climbs one scenario's load in steps and states the capacity. A finished run's findings list the slowest step, kinds of errors and throughput that stopped following the load, with their numbers, and a regression against the baseline run if there is one.
Spitfire installs on Docker or Kubernetes with one command; every testing feature and protocol is open in the free edition.