What is p95? How to read p95, p99 and latency percentiles
The number you see most in performance test results is p95: "p95 420 ms". It means 95% of the requests were answered in 420 ms or less and 5% took longer. This guide explains what percentiles are, why the average misleads, why p99 matters and how to write thresholds.
What is a percentile?
Imagine sorting every response time of a run from smallest to largest. p50 (the median) is the value in the middle: half the requests were faster, half slower. p90 is the value at the end of the first 90%, p95 at the end of the first 95%, p99 at the end of the first 99%. So p99 = 1200 ms means one request in a hundred took longer than 1.2 seconds.
Percentiles answer "what did the typical user get" (p50) and "what did the unlucky user get" (p95, p99) separately. Latency distributions are almost never symmetric: most requests are fast and a small share are very slow. Without looking at that long tail you cannot tell how the system feels to its users.
Why the average misleads
Ten response times (ms), sorted: 80 · 85 · 90 · 92 · 95 · 98 · 100 · 105 · 900 · 2400
Average 405 ms, median (p50) 96.5 ms, slowest 2400 ms. The average describes neither the typical request nor the slow ones: eight users out of ten got about 100 ms while two waited seconds.
A few very slow requests pull the average up, yet it still does not tell how many people got a slow answer. The other way round, many fast requests (cached responses, or errors that return quickly) can pull the average down and hide a real problem. That is why performance targets are written with percentiles, not averages.
p95 or p99?
They say different things. p95 covers most users' experience and is stable; it is meaningful with fewer samples and reacts less to noise. It is a good start for release comparisons and CI thresholds. p99 shows the tail: GC pauses, lock waits, queuing for a connection pool, calls close to their timeout. Problems in the tail are usually the first symptoms to show as the load grows.
Do not wave p99 away: when one page or API call fans out to many services behind it, the chance that the user hits at least one slow call grows fast. If each of 20 independent calls has a 1% chance of being beyond p99, the chance that at least one is beyond it is 1 − 0.99²⁰ ≈ 18%. A backend p99 feels close to a p80 for the user.
A practical rule: set the target with p95 and put a looser upper bound on p99. Mind the sample count when reading p99: in 200 requests, p99 is a number set by the two slowest requests.
Four traps when reading percentiles
- Percentiles do not average. Adding two runners' or two time windows' p95 and dividing by two does not give a p95. A correct p95 is computed from all samples (or from histograms that can be merged).
- The overall p95 hides the steps. With many fast list requests and few slow checkout requests, the overall p95 looks fine. Look at critical steps separately and give them their own thresholds.
- Errors can be fast. 500s that return at once improve p95. Always read latency together with the error rate.
- A closed model can hide the tail. In the VU model users send requests less often as the system slows down; queued requests are never sent, so they are never measured. For capacity questions the arrival-rate model reduces this effect.
How to write a threshold
A threshold turns a percentile into a decision: "passed if p95 is under 500 ms and the error rate under 1%". Derive the target from the user experience and the business need, not from today's result; a threshold written from today's result only catches regressions. Then complement it with a comparison against a baseline run: p95 rising 20% is a signal even when no threshold breaks.
Giving every endpoint the same target is not right either. A search the user waits for and a report generated in the background need different limits; a critical step such as checkout usually gets the tightest target.
Product list: p95 < 300 ms · Search: p95 < 500 ms · Checkout: p95 < 800 ms, p99 < 2 s · All requests: error rate < 1%. These values are a starting example; derive yours from your own user experience targets.
With Spitfire
Spitfire shows p50, p90, p95 and p99 live, second by second, with per-step and per-location breakdowns. Thresholds live in the test definition and turn a run into a pass/fail result; they can be scoped to a step, a scenario or a location:
"thresholds": [
{ "metric": "req_duration", "expr": "p(95)<500" },
{ "metric": "req_duration", "expr": "p(99)<1500" },
{ "metric": "req_duration", "filter": { "step": "checkout" }, "expr": "p(95)<300" },
{ "metric": "req_failed", "expr": "rate<0.01" }
]A finished run's findings state the slowest step and a long latency tail with their numbers; mark a run as the baseline and later runs are compared with it on p95, p99, error rate and request rate. The same thresholds become an exit code in CI: performance testing in CI/CD. To find the load at which the system misses its p95 target: breakpoint testing.
Spitfire installs on Docker or Kubernetes with one command; every testing feature and protocol is open in the free edition.