Performance testing in CI/CD: a load test gate in the pipeline
Performance regressions usually do not come from one big change but from many small ones adding up: a JOIN added to a query, a cache removed, a JSON response that grew. A load test run once before a release catches them late, and finding the change behind them gets harder. Bringing performance tests into the CI/CD pipeline lets you see a regression on the same day as the change.
Where in the pipeline, and how much?
Running an hour-long soak test on every commit is neither possible nor needed. The setup that works is usually three layers:
- On every pull request: a short load test of 3–10 minutes against staging or an ephemeral environment, at the expected load or a fraction of it. The goal is not to measure capacity but to catch a clear regression: per-step p95 and error rate, thresholds and a comparison with a baseline run.
- Every night: a longer, more realistic load test with the full data set, several scenarios and several locations if needed. The result is in the team's channel in the morning.
- Before a release: a breakpoint test, a spike and, if needed, a soak (test types). These usually do not block the pipeline; they feed the release decision.
Dealing with noise
The biggest enemy of performance tests in CI is noise: the same code can give results 5–10% apart in two runs. If the gate breaks every day for no reason, the team stops trusting it. To reduce noise:
- Run against a dedicated, fixed-size environment where possible, not one shared with other load.
- Start every run with the same data; leave a short ramp-up for caches to warm.
- Keep absolute thresholds realistic and somewhat loose (for example 20% above the p95 target); catch small drifts with a baseline comparison and a tolerance.
- Give critical steps their own thresholds; the overall p95 alone hides a step's regression (the p95 and p99 guide).
Security and permissions
The credential the pipeline gets should do only what it needs. Use a revocable token limited to running tests, not the load testing tool's admin password, and keep it as a pipeline secret. Do not run data-changing tests against production from a PR pipeline.
With Spitfire
You save the test once in Spitfire; the pipeline runs it on the controller by name. Create a personal token under API tokens in the profile menu and keep it as the pipeline secret SPITFIRE_TOKEN. The token acts with its user's current role and groups and may only read, validate, start and stop runs. Runs it starts are tagged ci, noted with the pipeline and audited with the token's name.
# .github/workflows/load-test.yml
on: [pull_request]
permissions:
contents: read
pull-requests: write # the PR comment
jobs:
load-test:
runs-on: ubuntu-latest
steps:
- env:
SPITFIRE_URL: https://spitfire.example.com
SPITFIRE_TOKEN: ${{ secrets.SPITFIRE_TOKEN }}
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: |
curl -fsSL "$SPITFIRE_URL/api/v1/runner-dist/spitfire-cli-linux-amd64.gz" | gunzip > spitfire && chmod +x spitfire
./spitfire cloud run "Checkout flow" --env staging --pr-comment --summary-md spitfire.mdspitfire cloud run waits for the run and reports the result as an exit code; the pipeline step passes or fails on it:
| Exit code | Meaning |
|---|---|
0 | passed |
99 | thresholds (or performance budgets) failed |
97 | a threshold aborted the run (abortOnFail) |
2 | the test writes data and --confirm-writes was not given |
1 | other errors (for example the concurrent-run limit is full) |
A comparative run (--compare A=test,B=dev) exits by its verdict: 0 better or no difference, 98 worse, 97 invalid, 99 arm B's thresholds. More: version comparison.
Common options: --env staging runs against one of the test's environments, --scale 0.5 halves the load, --var key=value changes a variable for that run, -l istanbul=60:2 -l frankfurt=40 splits the run across locations, -o summary.json keeps the summary and --report report.pdf the report as build artifacts. The CLI also ships in the image: docker run --rm -e SPITFIRE_URL -e SPITFIRE_TOKEN algebransoft/spitfire spitfire cloud run …
Nightly runs do not need a pipeline
The nightly layer does not need a pipeline: with Schedule on the test page the test runs every night or every weekday morning. When a threshold breaks, a run fails or a scheduled run does not start, Slack, Teams, SMS or e-mail hears about it. When a run is clearly worse than its baseline run (p95, p99, error rate, requests/s) an alert goes out even before any threshold breaks, so a slowly accumulating regression shows before users notice.
The performance diff on a pull request
--pr-comment writes a summary on the pull request (GitHub Actions) or merge request (GitLab CI): each step's p95, p99 and error rate against the test's baseline run, with ⚠️ where it got worse beyond the tolerance. Every push updates the same comment. --summary-md writes the same Markdown to a file, and on GitHub to the job page too. On the test page you set a performance budget per step (for example POST /orders p95 < 300 ms); a breached budget fails the run like a threshold (exit code 99). The test page charts the step p95 trend of recent runs by version tag or commit. Only the CI job talks to GitHub or GitLab; the controller never does.
Pass/fail on thresholds in CI and every protocol are open in the free edition too. PR comments and performance budgets come with these plans: Growth yearly, Scale yearly, Enterprise yearly. The GitLab example and more: installation guide, CI/CD.
Spitfire installs on Docker or Kubernetes with one command; every testing feature and protocol is open in the free edition.