Breakpoint testing: finding your system's capacity
The answer to "how much load can our system take?" is not a chart but a number: the highest request rate at which it still meets the criteria you set (for example p95 under 500 ms and an error rate under 1%). A breakpoint test finds that number by climbing the load in steps. This guide covers how to set the test up, how to read the result and how to take it into capacity planning.
Why stepped load?
If you raise the load continuously and linearly (a ramp), somewhere on the chart you see latency climb, but it is hard to say at which load: the system is under constantly changing load, and queues and caches never settle. With stepped load each step is a short ramp followed by a constant hold. Each step is judged only on the numbers of its own hold, so you can make an exact statement such as "at 200/s, p95 was 290 ms".
Setting it up
- Write the criteria first. Breaking does not mean "the system crashed"; it means "it missed the criteria". A number found without a p95 target and an error-rate limit is open to interpretation. On percentiles: the p95 and p99 guide.
- Climb by rate. Raising iterations per second instead of VUs keeps the load from dropping by itself as the system slows down; the result reads directly as "operations per second".
- Choose start, step and upper bound sensibly. Start from a load you know is comfortable; keep the step at the resolution you need the answer in (too small makes the test long, too large finds the capacity coarsely).
- One scenario, a realistic mix. Capacity depends on the mix: a number found with only list requests does not hold for checkout-heavy traffic. Keep the step ratios in the scenario close to real traffic.
- Do not let the load generator be the limit. Allocate enough VUs and runners, or the limit you find is the load generator's, not the system's.
Reading the result
Criteria: p95 < 500 ms and error rate < 1%. Start 140, step 30, up to 260 iterations/s.
| Step | p95 | Error rate | Result |
|---|---|---|---|
| 140/s | 180 ms | 0.1% | ✓ met |
| 170/s | 210 ms | 0.2% | ✓ met |
| 200/s | 290 ms | 0.4% | ✓ met |
| 230/s | 640 ms | 2.8% | ✗ p95 and error rate |
Result: met the criteria at 200 iterations/s, not at 230. The real capacity lies between the two.
Also look at what broke at the failing step. If only latency rose, the system is saturated and requests are queuing. If errors started, their kind matters: a 429 may point to a rate limit, a 503 to a load balancer or circuit breaker, timeouts to a pool that filled up. If throughput stopped following the load (target 230/s, achieved 205/s), the system can no longer keep up with the incoming work.
The breaking point is a number measured from outside; it does not say which resource saturated inside. For that you need the system's own metrics: CPU, memory, connection pool, queue depth. Repeating the test on different days matters too: a few percent between two runs in the same environment is normal, large differences mean the environment changed.
When and how often?
A breakpoint test is not something you run once and forget. Repeat it after anything that changes capacity: a major release, a new dependency, a change in the database schema or indexes, a change in infrastructure size, a campaign that will raise the expected traffic. Keep the results with their dates; seeing how capacity changes across releases tells far more than a single measurement. Before a campaign, climb at least a few steps beyond the target traffic, since forecasts can be wrong.
From capacity to a plan
The number you found holds for one environment and one instance (pod, server) count. Estimating how many instances a target traffic needs is possible, but it is an estimate resting on assumptions: that instances scale linearly up to the breaking point and that shared dependencies (database, cache, queues, outside services) do not saturate first. Since the real capacity lies between the step that held and the one that failed, the estimate should be a range. As long as linear scaling has not been measured, confirm the plan with a second test on more instances.
With Spitfire
In Spitfire you do not write a separate test. In the Run dialog choose Find the breaking point; one scenario's load climbs in steps from a start, by a step, up to a maximum (each step a 10 s ramp and a hold). Every step is judged on its own numbers alone, that step's p95 and error rate. The first step that fails stops the run, and the run page states the capacity with the step table and why: "met the criteria at 200 iterations/s, not at 230". When the climb stopped because the virtual users ran out while latency was fine, it says this is not the system's limit. The test itself does not change.
On a finished breakpoint run's page, the Capacity planning card takes how many instances served the system during the run and the number of concurrent users to plan for, and answers with something like "≈ 18 instances (range 15–18)". The card says it is an estimate, not a measurement, next to every assumption it rests on; its confidence is never above "medium", because linear scaling itself was not measured.
With your systems' OpenTelemetry data connected, the resources that jumped most at the failing step (CPU, memory, connection pool, queue) are ranked against the step before and written in the finding as "at the same time": OpenTelemetry: closer to the root cause.
Spitfire installs on Docker or Kubernetes with one command; every testing feature and protocol is open in the free edition.