Spitfire

Version: 0.21.0This documentation is for Spitfire 0.21.0.

Load model and thresholds

The Test editor page covers what a test does. This page covers how much load is applied and when a run counts as passed or failed: the load models (executors), stages, multiple scenarios, splitting the load across runners and locations, license limits, thresholds, and ready-made patterns for smoke, load, stress, spike, soak and breaking point (breakpoint) tests.

What it is for

  • For each scenario it sets how many VUs run, or how many iterations start per second, and how that changes over time.
  • From the values you enter in the editor it draws a plan chart: you see the run's duration, the maximum VU count and (for open models) the total number of iterations before running anything.
  • Thresholds decide the result at the end of the run (and during it, if you want): Passed or Thresholds failed. In CI, a failed threshold gives exit code 99.

When to use it

  • When setting the load of a new test for the first time.
  • When you want to run the same test for different purposes (a quick smoke test, a long soak test, a capacity measurement).
  • When results look "too slow" or "too many errors", to make sure the test's load really is the load you meant (e.g. with Dropped iterations in an open model, the load did not reach its target).

Core concepts

Concept Meaning
VU (virtual user) A user that runs the scenario's steps in order. Each VU has its own variables and (by default) its own cookies.
Iteration One run of all of a scenario's steps by a VU, from top to bottom.
Closed model (constant-vus, ramping-vus) The VU count is fixed; each VU starts a new iteration as soon as it finishes one (after think time). When the target slows down, fewer iterations start, so the load drops by itself.
Open model (constant-arrival-rate, ramping-arrival-rate) Iterations start on a timetable, however slow the target is. Without a free VU, an iteration is dropped (dropped_iterations); it is never queued, so latency isn't hidden.
Iteration-count models (shared-iterations, per-vu-iterations) The test ends when a given number of iterations is done (at Max duration at the latest).
RPS Requests per second. With N steps per iteration, RPS ≈ iterations/s × N.
Tip

A closed model answers "N users at the same time", an open model answers "what happens when N requests arrive per second". Real users keep arriving without waiting for the system to slow down, so for capacity and latency targets (SLAs) an open model gives the more faithful result.

Load models

The load model is picked per scenario in the Load model list under the Load heading of the Scenarios & steps tab. The help text under the list describes the selected model in one sentence. Duration fields take values such as 30s, 2m, 1m30s, 500ms; an invalid duration shows "Examples: 30s, 2m, 1m30s, 500ms".

Constant VUs

Constant VUs (constant-vus): a fixed number of VUs run iterations back to back for the duration.

Field JSON Description
VUs vus 1 or more.
Duration duration For example 5m.
json
"executor": { "type": "constant-vus", "vus": 50, "duration": "10m" }

Ramping VUs

Ramping VUs (ramping-vus): the VU count changes in stages. It is the default for new tests.

Field JSON Description
Start VUs startVUs The VU count at the start of the run (usually 0).
Graceful ramp-down gracefulRampDown When the VU count goes down, the time surplus VUs get to finish their running iterations. Default 30s.
Stages stages Each row: "Duration to reach Target VUs". The target must be a whole number.
json
"executor": {
  "type": "ramping-vus",
  "startVUs": 0,
  "stages": [
    { "duration": "2m", "target": 100 },
    { "duration": "10m", "target": 100 },
    { "duration": "1m", "target": 0 }
  ]
}

Constant arrival rate

Constant arrival rate (constant-arrival-rate): the given number of iterations start in every timeUnit.

Field JSON Description
Iterations rate Iterations to start per timeUnit (may be a decimal).
per (timeUnit) timeUnit Default 1s. rate: 30, timeUnit: "1m" = 30 iterations per minute.
Duration duration For example 10m.
Pre-allocated VUs preAllocatedVUs VUs prepared at the start of the run; 1 or more.
Max VUs maxVUs The most VUs that may be started if needed; not below the pre-allocated VUs.
json
"executor": {
  "type": "constant-arrival-rate",
  "rate": 200,
  "timeUnit": "1s",
  "duration": "10m",
  "preAllocatedVUs": 100,
  "maxVUs": 400
}

The hint in the form says how to size the VUs: "You need about latency × rate VUs (e.g. 200 ms × 100/s ≈ 20 VUs). With fewer, iterations are dropped." More precisely: VUs needed ≈ iterations per second × the length of one iteration in seconds (all steps and think times included). If your iteration has 3 steps and takes 1.5 seconds in total, 200 iterations per second need about 300 VUs, and more when the target slows down. Keep Max VUs at two to three times that number.

Ramping arrival rate

Ramping arrival rate (ramping-arrival-rate): the iteration rate changes in stages. Suited to capacity and stress tests.

Field JSON Description
Start rate startRate The rate at the start of the run (iterations per timeUnit).
Rate unit (timeUnit) timeUnit Default 1s.
Stages stages Each row: "Duration to reach Target iterations/unit".
Pre-allocated VUs, Max VUs preAllocatedVUs, maxVUs As for the constant arrival rate.
json
"executor": {
  "type": "ramping-arrival-rate",
  "startRate": 10,
  "timeUnit": "1s",
  "preAllocatedVUs": 50,
  "maxVUs": 500,
  "stages": [
    { "duration": "2m", "target": 100 },
    { "duration": "5m", "target": 100 },
    { "duration": "2m", "target": 300 },
    { "duration": "5m", "target": 300 }
  ]
}

Shared iterations

Shared iterations (shared-iterations): a fixed number of iterations is shared by the VUs; a faster VU does more. The test ends when the work is done or when Max duration is up. For data loading, smoke tests and one-off jobs.

Field JSON Description
VUs vus 1 or more.
Iterations in total iterations Not below the VU count.
Max duration maxDuration Default 10m. Stops at this time even if iterations remain.
json
"executor": { "type": "shared-iterations", "vus": 10, "iterations": 1000, "maxDuration": "15m" }

Per-VU iterations

Per-VU iterations (per-vu-iterations): every VU runs exactly the given number of iterations; the test ends when they are done (at Max duration at the latest). For scenarios where every user does the same flow the same number of times.

json
"executor": { "type": "per-vu-iterations", "vus": 20, "iterations": 5, "maxDuration": "10m" }

This example runs 20 × 5 = 100 iterations in total.

Load model form and plan chartLoad model form and plan chart

How stages work

In the ramping-vus and ramping-arrival-rate models the Stages list gives the shape of the load over time:

  • Each stage has a Duration and a Target. The load moves linearly from the value at the stage's start to the target over that duration.
  • The first stage starts from Start VUs (or Start rate); every later stage starts from the previous stage's target.
  • A stage with the same target as the previous one holds the load (a plateau).
  • A last stage with target 0 brings the load down (ramp-down).
  • Add stage adds a row (default duration 30s, same target as the previous row); the cross at the right of a row deletes it.

Example: startVUs: 0, stages 2m → 100, 10m → 100, 1m → 0: a 2-minute ramp-up from 0 to 100 VUs, 10 minutes at 100 VUs, a 1-minute ramp-down to 0. 13 minutes in total.

The plan chart

As soon as you enter the values, a chart titled Plan: VU (closed and iteration-count models) or Plan: iterations/s (open models) is drawn under the form. Under it is a summary line, e.g. "13m · up to 100 VUs", or for open models "… · 126,000 iterations in total". For iteration-count models the duration is a ceiling: "up to 10m · 20 VUs".

The Load profile card on the test page shows the same for the saved test: each scenario's model, maximum VUs and duration, the Total duration and the number of thresholds.

Load profile card on the test pageLoad profile card on the test page

Which load model should I pick

Your goal Model Why
A target like "200 users at the same time" Constant VUs, or Ramping VUs with a plateau stage Gives the concurrent user count directly.
An SLA/capacity target like "500 requests per second" Constant arrival rate The rate holds even when the target slows down; dropped iterations are visible.
Raising the load step by step to see where the system breaks Ramping arrival rate, or Find the breaking point in the Run dialog In an open model the load does not drop as the target slows down.
A sudden traffic spike (campaign, push notification) Ramping VUs or Ramping arrival rate with short stages A stage of a few seconds gives a sudden increase.
An endurance (soak) test lasting hours Constant VUs or Constant arrival rate, long duration Slow problems such as memory leaks or exhausted connection pools.
A quick check (smoke test) Shared iterations (e.g. 1 VU, 10 iterations) or Constant VUs with 1–2 VUs You see quickly whether the system and the test work.
Loading 10,000 records once Shared iterations Ends when the work is done. Use it with a unique data file (see Test data).
Every user doing the flow N times Per-VU iterations Every VU does the same number of iterations.

Step by step setting up the load

  1. Open the Scenarios & steps tab in the editor and select the scenario with the buttons at the top.
  2. Pick the model in the Load model list under the Load heading. The fields change with the model.
  3. Fill in the fields. Write durations as 30s, 5m.
  4. For ramping models, define the stages with Add stage.
  5. For open models, size Pre-allocated VUs and Max VUs with the calculation in Constant arrival rate.
  6. Look at the plan chart and summary line under the form: are the duration and maximum VUs what you expect? Are they within the license limits (see License limits)?
  7. If needed, set how long after the run starts the scenario begins with Start delay.
  8. Define the thresholds on the Thresholds (N) tab (below).
  9. Save. You can start the first run small by setting Load scale to e.g. 10% in the Run dialog; the test does not change (see Runs and results).

Several scenarios

All scenarios of a test run at the same time, each with its own load model. This is how you build a realistic traffic mix:

json
{
  "version": 1,
  "name": "Shop traffic",
  "variables": { "base": "https://shop.example.com" },
  "options": {},
  "scenarios": [
    {
      "name": "browse",
      "executor": { "type": "constant-arrival-rate", "rate": 150, "timeUnit": "1s", "duration": "15m", "preAllocatedVUs": 100, "maxVUs": 400 },
      "steps": [
        { "id": "home", "name": "Home page", "protocol": "http", "request": { "method": "GET", "url": "{{base}}/" },
          "checks": [ { "type": "status", "op": "eq", "value": 200 } ] }
      ]
    },
    {
      "name": "checkout",
      "startTime": "1m",
      "executor": { "type": "constant-arrival-rate", "rate": 10, "timeUnit": "1s", "duration": "14m", "preAllocatedVUs": 20, "maxVUs": 100 },
      "steps": [
        { "id": "cart", "name": "Cart", "protocol": "http", "request": { "method": "GET", "url": "{{base}}/api/cart" },
          "checks": [ { "type": "status", "op": "eq", "value": 200 } ] }
      ]
    }
  ],
  "thresholds": [
    { "metric": "req_duration", "filter": { "scenario": "checkout" }, "expr": "p(95)<800" },
    { "metric": "req_failed", "expr": "rate<0.01" }
  ]
}
  • The checkout scenario starts 1 minute after the run starts, with startTime: "1m" (Start delay in the editor).
  • The run's total duration is the end of the longest scenario (15 minutes here).
  • The plan chart is drawn per scenario; the Load profile card on the test page also shows each scenario on its own line.
  • To apply a threshold to a single scenario, use filter.scenario; this filter is not in the form and is written on the JSON tab.

Splitting across runners and locations

The VU counts and rates in a test are the run's totals. When a run is spread over several runners or locations, Spitfire splits that total between the runners: "The load is split across the selected runners; total VU counts and arrival rates stay exactly as defined." You don't change the test for the number of runners.

  • In By location mode you give each location's share as a percentage (100% in total); a location's share is split evenly across its runners.
  • In the shared-iterations model each runner's share of iterations follows its share of VUs; the total is still exactly the test's number.
  • Unique and sequential data files are split between runners too; no two runners use the same row (see Test data).
  • Options → When a runner is lost: Stop the run (default) or Continue with the rest (less load).

Runner selection and the location split are done in the Run dialog; see Runs and results.

License limits

The license limits one run's maximum VU count, requests per second, duration, and number of runners and locations. In the free edition these limits are shown at the top of the Run dialog: per run at most 100 VUs, 500 requests per second, 15 minutes, 1 runner, 1 location; one run at a time. Paid plans' limits are shown on the License page.

A run over the limits does not start, and the reason is shown, e.g.: "This run needs 150 concurrent VUs; the free edition allows up to 100." What you can do:

  1. Lower Load scale in the Run dialog (e.g. 60%); the test does not change.
  2. Or reduce the VU/rate/duration values in the editor.

Limits are computed from the plan before the run starts:

  • VUs: the sum of the maximum VUs of the scenarios running at the same time. For open models this is Max VUs (maxVUs); setting it much higher than needed can hit the VU limit.
  • Request rate: for open models, the target iteration rate × the scenario's number of steps (not counting Run once per VU steps). For closed models the rate cannot be known beforehand, so during the run the runners slow requests down to stay under the limit.
  • Duration: the duration on the plan chart (the end of the longest scenario).

Defining thresholds

A threshold is a pass/fail criterion of the run. Thresholds are defined on the editor's Thresholds (N) tab. The tab's help text: "If a threshold is crossed the run ends as "thresholds failed" (exit code 99 in CI). With "stop when crossed" the run is cut short."

Step by step:

  1. Click the Thresholds (N) tab.
  2. Click Add threshold. The new row comes with Response time (ms) and p(95)<500.
  3. Pick the metric in the Metric list on the left. The expression box is filled with an example expression for that metric.
  4. Type the criterion in the Expression box (e.g. p(95)<300).
  5. In the Scope list, pick where the criterion applies: All requests, a step (Step: scenario / step name) or a location (Location: name).
  6. If the run should stop as soon as the criterion is crossed, tick Stop when crossed and type the warm-up time in the wait (30s) box next to it.
  7. Save.

Thresholds tab: metric, expression, scope and Stop when crossedThresholds tab: metric, expression, scope and Stop when crossed

Metrics

Name in the form Metric Kind Example expression
Response time (ms) req_duration trend p(95)<500
Failed request rate req_failed rate rate<0.01
Check pass rate checks rate rate>0.99
Request count / rate reqs counter rate>100 (per second), count>10000
Iteration duration (ms) iteration_duration trend p(95)<2000
Dropped iterations dropped_iterations counter count<1
Rows returned/affected (SQL/Mongo) rows trend avg>0
Time to first byte (HTTP, ms) http_req_waiting trend p(95)<300
Connecting (HTTP, ms) http_req_connecting trend p(95)<50
Iteration count / rate iterations counter count>100
Active VUs vus gauge value<500
WebSocket connect (ms) ws_connecting trend p(95)<200
SSE time to first event (ms) sse_time_to_first_event trend p(95)<1000

The list also has the gRPC stream metrics (grpc_time_to_first_message, grpc_message_latency, grpc_message_gap, grpc_messages_received, grpc_messages_sent) and the Kafka metrics (kafka_e2e_latency, kafka_consumer_lag).

Expressions

An expression has the form <stat> <operator> <number>. Operators: <, <=, >, >=, ==, !=.

Stat Meaning For which metric kind
p(95), p(99), p(99.9) Percentile (above 0, at most 100) trend
avg, min, max, med Average, minimum, maximum, median trend (min/max also on gauges)
count Total count trend, counter
rate Ratio (0–1), or per-second count for a counter rate, counter
value Last value gauge

Durations are in milliseconds, rates between 0 and 1: a 1% error rate is rate<0.01, not rate<1.

Tip

The average (avg) hides a small share of slow requests. For the latency users experience, use p(95) or p(99).

Scope and filters

Scope JSON What it measures
All requests no filter All scenarios and steps together
Step: … "filter": { "step": "create_order" } Only that step (the step id)
Location: … "filter": { "location": "frankfurt" } Only that location's runners
(JSON only) "filter": { "scenario": "checkout" } Only that scenario
(JSON only) "filter": { "check": "status == 200" } On the checks metric, only the check with that name

Filters can be combined (scenario + step + check), but a location filter is used alone.

json
"thresholds": [
  { "metric": "req_duration", "expr": "p(95)<500" },
  { "metric": "req_failed", "expr": "rate<0.01", "abortOnFail": true, "delayAbortEval": "1m" },
  { "metric": "req_duration", "filter": { "step": "create_order" }, "expr": "p(99)<1500" },
  { "metric": "checks", "filter": { "check": "status == 200" }, "expr": "rate>0.995" },
  { "metric": "req_duration", "filter": { "location": "frankfurt" }, "expr": "p(95)<800" },
  { "metric": "dropped_iterations", "expr": "count<1" }
]

Stop when crossed and warm-up

  • A threshold with Stop when crossed (abortOnFail) stops the run the moment it is crossed during the run. On the run page the threshold gets an "aborts" badge, and a "stopped the run" badge if it did.
  • wait (30s) (delayAbortEval): "Warm-up: won't stop before this". In the first seconds of a run, latency can be high while connections and caches warm up; no stop decision is made before this time.
  • Thresholds are computed on the total from the start of the run until that moment. The warm-up time only postpones the stop decision; the final evaluation at the end of the run still includes the requests of the warm-up period. To leave the warm-up out of the result entirely, see the Warm-up pattern.

Reading threshold results

  • During the run, a threshold with no data yet counts as passing. At the end of the run a threshold with nothing to measure (its filtered step never ran) fails and shows "no data": it does not go green on something it did not measure.
  • Counters are the exception: when no iteration was dropped, count<1 on dropped_iterations passes.
  • The Thresholds card on the run page shows each threshold's measured value and result; failed thresholds are also listed at the top of the page ("N thresholds failed:").
  • With no thresholds: "No thresholds defined; the run always counts as passed."

Thresholds card on the run pageThresholds card on the run page

Test patterns

The patterns below are built on the same test by changing only the load model.

Smoke test

Goal: do the test and the system work? 1 VU, a few iterations.

json
"executor": { "type": "shared-iterations", "vus": 1, "iterations": 10, "maxDuration": "2m" }

Load test

Goal: latency and error rate at the expected normal load. Ramp-up, long plateau, ramp-down.

json
"executor": {
  "type": "ramping-vus", "startVUs": 0,
  "stages": [ { "duration": "3m", "target": 200 }, { "duration": "20m", "target": 200 }, { "duration": "2m", "target": 0 } ]
}

Stress test

Goal: what happens above the expected load? Raise the rate step by step in an open model and hold at each step.

json
"executor": {
  "type": "ramping-arrival-rate", "startRate": 50, "timeUnit": "1s",
  "preAllocatedVUs": 100, "maxVUs": 1000,
  "stages": [
    { "duration": "2m", "target": 100 }, { "duration": "5m", "target": 100 },
    { "duration": "2m", "target": 200 }, { "duration": "5m", "target": 200 },
    { "duration": "2m", "target": 400 }, { "duration": "5m", "target": 400 }
  ]
}

Spike test

Goal: resilience to a sudden jump, and recovery afterwards.

json
"executor": {
  "type": "ramping-vus", "startVUs": 10,
  "stages": [
    { "duration": "2m", "target": 10 },
    { "duration": "10s", "target": 500 },
    { "duration": "2m", "target": 500 },
    { "duration": "10s", "target": 10 },
    { "duration": "3m", "target": 10 }
  ]
}

Soak test

Goal: problems that appear over hours (memory, connection pool, disk). Normal load for a long time. Mind the license's run duration limit.

json
"executor": { "type": "constant-arrival-rate", "rate": 100, "timeUnit": "1s", "duration": "4h", "preAllocatedVUs": 100, "maxVUs": 300 }

Finding the breaking point

To find the highest rate the system can take you don't need to change the test: Find the breaking point (capacity test) in the Run dialog raises the selected scenario's load step by step and judges each step on p95 and error rate; the first step that fails stops the run.

  1. Click Run on the test page.
  2. Tick Find the breaking point (capacity test) in the dialog.
  3. Fill in the fields (defaults in brackets):
    • Scenario: the scenario whose load is raised (the others do not run in this run).
    • Start (10), Step (10), Up to (200): iterations/s.
    • Step length (60 s): the measured part of each step; at least 10 seconds. There are also 10 s transitions between steps (not measured).
    • p95 limit (500 ms) and Error rate limit (1%).
    • VU limit (auto).
  4. The line under the form states the plan: "20 steps, about 24 minutes at most (10 s between steps)." There can be at most 60 steps.
  5. Click Start.

Find the breaking point form in the Run dialogFind the breaking point form in the Run dialog

The test's thresholds do not apply in this run (otherwise they would stop it themselves at the very breaking point you are looking for), and the test definition does not change. The Capacity test card on the run page states the result, e.g.: "The system met the criteria at 200 iterations/s (200 requests/s); at the 230 iterations/s step it did not." If a step stopped while latency was fine because there were not enough VUs to start the iterations, it says so plainly: "…but this is not the system's limit: there were not enough virtual users to start the iterations. Raise the VU limit and try again."

Warning

The steps' total duration and top rate are subject to the license limits too. The default settings (20 steps, about 24 minutes) exceed the free edition's 15-minute duration limit; lower Up to or raise Step.

Warm-up

There are two ways to keep the first minutes, when caches, connection pools and the JIT warm up, out of the result:

  1. Postpone the stop decision: on thresholds with Stop when crossed, set wait (e.g. 2m) to the warm-up time. The result evaluation still covers the whole run.
  2. Make the warm-up a separate scenario: a low-load warmup scenario runs at the start of the run, the main scenario starts later with Start delay, and the thresholds are filtered to the main scenario only:
json
{
  "version": 1,
  "name": "Checkout load test (with warm-up)",
  "variables": { "base": "https://shop.example.com" },
  "options": {},
  "scenarios": [
    {
      "name": "warmup",
      "executor": { "type": "constant-vus", "vus": 5, "duration": "2m" },
      "steps": [
        { "id": "warmup_home", "name": "Warm-up", "protocol": "http", "request": { "method": "GET", "url": "{{base}}/" } }
      ]
    },
    {
      "name": "main",
      "startTime": "2m",
      "executor": { "type": "constant-arrival-rate", "rate": 100, "timeUnit": "1s", "duration": "10m", "preAllocatedVUs": 50, "maxVUs": 300 },
      "steps": [
        { "id": "main_home", "name": "Home page", "protocol": "http", "request": { "method": "GET", "url": "{{base}}/" },
          "checks": [ { "type": "status", "op": "eq", "value": 200 } ] }
      ]
    }
  ],
  "thresholds": [
    { "metric": "req_duration", "filter": { "scenario": "main" }, "expr": "p(95)<400" },
    { "metric": "req_failed", "filter": { "scenario": "main" }, "expr": "rate<0.01" }
  ]
}

Common problems

Symptom Cause Fix
...executor.maxVUs: must be at least preAllocatedVUs Max VUs < Pre-allocated VUs. Make Max VUs at least as large as the pre-allocated VUs.
...executor.preAllocatedVUs: must be positive Pre-allocated VUs is 0 in an open model. Enter at least 1; for sizing see Constant arrival rate.
...executor.stages: at least one stage is required A ramping model without stages. Add at least one stage with Add stage.
...stages[1].target: must be a whole number A decimal target (12.5) in Ramping VUs. A VU target must be a whole number. Decimal rates are valid only in arrival-rate models.
...stages[0].duration: must be positive An empty or 0s stage duration. Enter a duration such as 30s, 2m.
...executor.iterations: must be at least the number of VUs Iterations in total < VU count. Raise the iterations or lower the VUs.
"N dropped" on the run page (Dropped iterations) No free VU was left in an open model; the target slowed down or Max VUs is low. Raise Max VUs; if latency rose, the system is nearing its limit. Add a count<1 threshold on dropped_iterations.
RPS lower than expected (closed model) As the target slows, VUs start fewer iterations; think time lowers RPS too. Switch to an open model if you have a rate target.
invalid threshold (examples: p(95)<500, rate<0.01, avg<=200) Wrong expression syntax (p95<500, < 500ms). Write p(95)<500; don't write units.
req_failed is a rate metric; use one of rate A stat that does not fit the metric kind (p(95) on the error rate). Use one of the stats in the table.
unknown step id … The step id in the threshold filter is not in the test (the step was deleted, or its name was written). Pick the step again in the Scope list; in JSON, write the step's id, not its name.
no step has a check named … A check filter points at a check name that does not exist. Check the name on the run page's Checks card or in the JSON.
a location filter cannot be combined with scenario, step or check location combined with another filter. Make the location threshold a separate row.
A threshold failed with "no data" The filtered step never ran in the run, or always hit connection errors. Check the step id and the run's Steps table.
"This run needs …; the free edition allows up to …" License limit. Follow License limits.
  • Test editor: defining scenarios, steps, extractors and checks.
  • Test data: how data files are split between runners.
  • Runs and results: the Run dialog, watching live, reading results.