Spitfire

Version: 0.21.0This documentation is for Spitfire 0.21.0.

Runners and locations

Runners generate the load. This page covers what a runner is, the local runners that come with the installation, adding remote runners to send load from another city or cloud region, locations, labels, capacity, Kubernetes runners, the TLS pin, and why runners show up as offline.

What it is for

  • A runner is the process that sends the requests to the system under test and streams metrics to the controller every second. The controller splits a test across runners; total VU counts and arrival rates stay exactly as defined.
  • A location is where a runner sends load from (e.g. istanbul, frankfurt). You can split a run across locations by percentage and see results per location (requests, RPS, avg/p95/p99, errors, a p95-per-location chart).
  • The Runners page shows each runner's state, CPU, active VUs, clock offset and release live.

When to use it

  • After installation, to check that the runners connected.
  • When adding runners to generate more load.
  • When you want to send load from another city or cloud region where your users are.
  • When the Dashboard says "No runners connected, so tests can't start."
  • When a runner shows as Offline, older release or at its CPU limit.

What a runner is

 browser ──► controller (web UI + REST API) ──gRPC/TLS──► runners (location A)
                  │                           └─────────► runners (location B)
               Postgres                                        │
                                                   load ──► system under test
  • A runner connects out to the controller (gRPC, one long-lived stream). No inbound port has to be opened on the runner's server; only outbound TCP to the controller's gRPC port (8471 on a Docker installation) must be allowed.
  • A runner identifies itself to the controller with a registration key (sfrun_…). Keys are created on the Runners page and shown only once.
  • A runner sends a heartbeat every 2 seconds. If 3 heartbeats in a row (about 6 s) do not arrive, the controller drops the connection and the runner shows as Offline. When the connection is lost, the runner keeps trying to reconnect at growing intervals (at most 15 s).
  • Traffic to the system under test leaves from the runner; the runner needs network access to that system. Connection secrets (database, Kafka… passwords) travel from the controller to the runner over this encrypted channel when a run starts.

Local and remote runners

Local runner Remote runner
Where it comes from The installer starts it with the controller (2 by default on Docker, --runners N) You install it on another server with install.sh runner, systemd, a Windows service or docker run
Location local (change with --location NAME at installation) The location bound to the key, or the runner's --location
How it connects Over the unpublished plaintext internal port 8475 8471, with TLS + pin
Labels deploy=docker on Docker Whatever you give with --labels

To change the number of local runners, run the installer again with the new number: ~/spitfire/install.sh docker --runners 4.

The Runners page

The Runners page: the runner table with Runner, Location, Status, Labels, CPU, Active VUs, Clock offset and Last seen columns, and the Registration keys card belowThe Runners page: the runner table with Runner, Location, Status, Labels, CPU, Active VUs, Clock offset and Last seen columns, and the Registration keys card below

Table columns:

  • Runner — its name (by default the host name), release and core count. A runner on another release than the controller carries a yellow older release tag.
  • Location — the runner's location, or "none". Below it, the round trip to the controller.
  • Status — Idle, Running (with a go to run link, or "synthetic monitor check"), Draining, Offline.
  • Labels, CPU (red at its limit), Active VUs, Clock offset (yellow above 50 ms), Last seen.
  • Buttons on the right: Stop assigning new runs (drain) / Accept new runs again and Remove.

Non-admins see the table but cannot add runners.

Adding a runner

  1. Open Infrastructure → Runners (admin).
  2. On the Registration keys card at the bottom of the page:
    • Type a name you will recognise the key by into Name (e.g. eu-west runners) (required).
    • Enter the location runners connecting with this key are filed under into Location (e.g. frankfurt) (lower case, a-z, 0-9, -, _, .; at most 40 characters). Leave it empty to let the runner say its location itself with --location.
  3. Click Create key. The Runner key dialog opens.

The Runner key dialog: the key shown once, the Docker, Without Docker (systemd) and Windows tabs, ready-made install commands and the controller key pinThe Runner key dialog: the key shown once, the Docker, Without Docker (systemd) and Windows tabs, ready-made install commands and the controller key pin

  1. Mind the warning "This key won't be shown again. Copy it now.": take the key with Copy, or use the ready-made commands below it directly; they already contain the controller address, the key, the pin and the location.
  2. Pick the tab that suits the server you install the runner on: Docker, Without Docker (systemd) or Windows (sections below).
  3. Run the command on that server. Within a few seconds the runner appears in the table as Idle.
  4. Close the dialog with Close. The key stays in the list as valid; several runners can connect with the same key.
Warning

If the dialog says "The controller's external address is not configured", the address in the commands is a guess. If the remote server cannot reach it, correct it in the command; the permanent fix is --grpc-public-addr HOST:PORT at installation or SPITFIRE_GRPC_PUBLIC_ADDR on the controller.

Warning

If you see the red "The runner endpoint has no TLS (SPITFIRE_GRPC_TLS=off)" warning, connection secrets travel to runners unencrypted; do not connect from another location.

Remote runner with Docker

The installer on the Docker tab (needs Docker; checks the connection and the pin first):

bash
curl -fsSL https://spitfire.tr/install.sh | SPITFIRE_VERSION=<controller release> bash -s -- runner \
  --controller spitfire.example.com:8471 --token sfrun_... --ca-pin sha256:... --location frankfurt --runners 1
  • --runners N starts N runners on this server.
  • --tls verifies with the system CAs instead of a pin (when the controller has a company/public CA certificate).
  • The script pins the runner to the controller's release and keeps its settings in ~/spitfire/deploy/runner/.env.
  • Uninstall: ~/spitfire/uninstall.sh runner (add --purge to delete the key too; don't forget to revoke it in the UI).

The same tab also has a single docker run command:

bash
docker run -d --name spitfire-runner --restart unless-stopped \
  -e SPITFIRE_RUNNER_CONTROLLER=spitfire.example.com:8471 \
  -e SPITFIRE_RUNNER_TOKEN=sfrun_... \
  -e SPITFIRE_RUNNER_CA_PIN=sha256:... \
  -e SPITFIRE_RUNNER_EPHEMERAL=true \
  algebransoft/spitfire:<release> spitfire-runner

Linux server without Docker

The Without Docker (systemd) tab installs the runner as a systemd service on a Linux server (amd64/arm64, systemd). The controller serves its own release's runner binaries and install-runner.sh under /api/v1/runner-dist/:

bash
curl -fsSLo install-runner.sh http://CONTROLLER:8470/api/v1/runner-dist/install-runner.sh
echo "<sha256 from the page>  install-runner.sh" | sha256sum -c -
sudo bash install-runner.sh --from http://CONTROLLER:8470 \
    --controller CONTROLLER:8471 --token sfrun_... --ca-pin sha256:...
  • The script downloads the binary and checks its SHA-256 before unpacking it.
  • It creates the spitfire system user, /etc/spitfire/runner.env (0600, holds the key) and the hardened spitfire-runner@.service, then waits for the runner to connect.
  • Server without internet: download the .gz file from the page, copy it over and install with --binary ./spitfire-runner-linux-amd64.gz.
  • Running it again updates (settings are kept); --instances N runs N instances on one server; uninstall with --uninstall [--purge].
  • Logs: journalctl -u spitfire-runner@1.

The Without Docker (systemd) tab of the Runner key dialog: the install-runner.sh download, the sha256 check and the .gz files for offline installsThe Without Docker (systemd) tab of the Runner key dialog: the install-runner.sh download, the sha256 check and the .gz files for offline installs

Windows service without Docker

The Windows tab has two commands:

  • With Docker Desktop (PowerShell): — a runner in Docker through install.ps1 runner.
  • Without Docker, as a Windows service (Administrator PowerShell): — downloads install-runner.ps1 from the controller; its SHA-256 is checked in the command, and the runner binary's inside it.

The services spitfire-runner-1…N start with Windows and are restarted if they crash; settings live under %ProgramData%\Spitfire in files only SYSTEM and Administrators can read, logs in %ProgramData%\Spitfire\logs. For several runners -Count N; to uninstall -Uninstall.

The Windows tab of the Runner key dialog: the Docker Desktop command and the Windows service install commandThe Windows tab of the Runner key dialog: the Docker Desktop command and the Windows service install command

Locations

A runner's location is set in one of two ways:

  1. A location bound to the key (recommended): if you filled in Location when creating the key, every runner connecting with that key is filed under it, whatever the runner claims.
  2. What the runner says: when the key is not bound to a location, the runner says it with --location NAME (or SPITFIRE_RUNNER_LOCATION). Local runners use the installation's --location (default local).

To use them in a run: on the test page, Run → in the Run test dialog choose By location, enter each location's share (100% in total) and the number of runners to use (Spread evenly helps). A location's share is split evenly across its runners. The last split is remembered on the test.

The By location tab of the Run test dialog: each location's share and runner count, Spread evenly and StartThe By location tab of the Run test dialog: each location's share and runner count, Spread evenly and Start

Thresholds can be scoped to a location: {"metric":"req_duration","filter":{"location":"frankfurt"},"expr":"p(95)<800"}. A location filter cannot be combined with a scenario/step filter.

Note

On the free edition a run can use 1 runner and 1 location. Growth allows 10 runners and 3 locations per run, Scale 30 and 6. See License.

Labels

Labels are free key=value pairs to group runners: --labels region=eu,zone=a (or SPITFIRE_RUNNER_LABELS). By labels in the Run test dialog sends the run to runners carrying given labels. To choose runners one by one use Pick; to give just a number, By count.

Capacity: how many VUs a runner carries

The number of VUs one runner can carry is not fixed; it depends on how heavy the steps are, the response size, think time and the machine's CPU. A practical method:

  1. Run the test with a small part of the expected load.
  2. Watch the CPU column on the Runners page. A runner close to its CPU limit turns red with "Runner CPU is at its limit; results may be skewed"; latency measurements are then affected by the runner itself.
  3. If CPU is at the limit, add runners (connect new runners with the same key, or --runners N) and spread the load over more runners with By count.

--max-vus (or SPITFIRE_RUNNER_MAX_VUS) is the capacity a runner advertises and is informational only; it does not limit the load. For Kubernetes on-demand runners the suggestion is SPITFIRE_ONDEMAND_VUS_PER_RUNNER (default 1000) VUs per pod.

Runners on Kubernetes

The Kubernetes installation runs the runners as the spitfire-runner Deployment; change their number with --runners N at installation, or:

bash
kubectl -n spitfire scale deploy/spitfire-runner --replicas=N

On-demand runners: install.sh kubernetes --runners 0 --on-demand-runners 20 keeps no load generators running. When a run asks for more runners than are idle (Run test → By count; the dialog suggests a number from the test's peak VUs), the controller starts the rest as pods of the same image, the run waits in the queue until they connect (usually 20–60 s) and the pods are deleted when it ends; leftovers are removed after a controller restart. At most N pods run at once. The controller gets a Role for pods in its own namespace only, and only while the option is on (--on-demand-runners 0 removes it). Pod size is set on the controller with SPITFIRE_ONDEMAND_CPU, SPITFIRE_ONDEMAND_MEMORY, SPITFIRE_ONDEMAND_VUS_PER_RUNNER. While the feature is on, an information banner shows at the top of the Runners page.

Ephemeral runners

Runners running as containers or pods register with a new identity when they are re-created. That is why Compose and Kubernetes runners and the docker run command run with SPITFIRE_RUNNER_EPHEMERAL=true (--ephemeral): when such a runner stays offline, the controller removes it from the list after a while (5 minutes by default; SPITFIRE_RUNNER_REAP_AFTER on the controller, 0 turns it off). Old entries therefore don't pile up.

Runners installed as a systemd or Windows service are not ephemeral; they keep their identity in a file (--id-file) and show up on the same row after a restart.

TLS pin and certificates

  • Port 8471, where remote runners connect, is always TLS. The controller creates its own identity on first start and keeps it.
  • The runner verifies the controller by its key pin: --ca-pin sha256:… (several pins can be given, comma-separated). The pin is on the line "The runner verifies the controller with this key pin:" at the bottom of the key dialog.
  • If you gave the controller your own certificate (SPITFIRE_GRPC_TLS_CERT, SPITFIRE_GRPC_TLS_KEY), use --tls (system CAs) or --ca <PEM file> on the runner instead of a pin.
  • --tls-insecure does not verify the certificate; for testing only.
  • Runners on the same machine/cluster connect over the unpublished plaintext port 8475 (SPITFIRE_GRPC_INTERNAL_ADDR); never open that port to the outside.

Updating runners

Keep remote runners on the controller's release; the Runners page marks a different release with a yellow older release tag.

  • Installed from the binary (systemd, Windows service): click Update next to the older runner (or Update N runners to X above the table). The runner downloads the new release from the controller, checks Spitfire's signature itself (it accepts nothing unsigned or older) and restarts; it is disconnected for a few seconds. Runners in a run are not updated.
  • With Update runners automatically on (off by default) the controller does this by itself: as soon as a runner is idle, one at a time per location, never one in a run. A release that failed on a runner is not tried on it again.
  • Docker and Kubernetes runners update with their image: run the install command again (for a remote Docker runner, the install.sh runner command from the key dialog; updating the controller also updates the local runners).

Draining, removing and revoking a key

  • Stop assigning new runs (drain): the runner stays connected but takes no new runs (status Draining); use it before maintenance. Accept new runs again undoes it.
  • Remove: the runner is removed from the list and disconnected. It can reconnect while its key is valid. A runner in a run cannot be removed.
  • Revoke (in the key list): runners connected with a revoked key are dropped immediately unless they are in a run, and cannot connect again. A systemd runner then exits with code 3 and systemd does not restart it.

Runner environment variables

Every --flag can also be given as the SPITFIRE_RUNNER_<FLAG> environment variable:

Flag Environment variable Meaning
--controller SPITFIRE_RUNNER_CONTROLLER The controller's gRPC address (HOST:PORT)
--token SPITFIRE_RUNNER_TOKEN Registration key (sfrun_…)
--name SPITFIRE_RUNNER_NAME Name shown in the UI (default: host name)
--location SPITFIRE_RUNNER_LOCATION Location (a location bound to the key wins)
--labels SPITFIRE_RUNNER_LABELS Labels, e.g. region=eu,zone=a
--ca-pin SPITFIRE_RUNNER_CA_PIN Controller key pin; implies TLS
--tls SPITFIRE_RUNNER_TLS TLS with the system CAs
--ca SPITFIRE_RUNNER_CA PEM CA file to verify against
--ephemeral SPITFIRE_RUNNER_EPHEMERAL A runner that is replaced rather than restarted
--max-vus SPITFIRE_RUNNER_MAX_VUS Advertised capacity (informational)
--debug SPITFIRE_RUNNER_DEBUG Debug logging

Common problems

Symptom: The Dashboard says "No runners connected, so tests can't start." or a run fails with "Not enough idle runners". Cause: No runner is connected, or they are all in another run or draining. Fix: Check the states on the Runners page. If the local runners have stopped, check with cd ~/spitfire && docker compose -p spitfire -f deploy/docker/docker-compose.yml ps and run the installer again. Wait for the current run to end, or add runners.

Symptom: A runner shows Offline; Last seen keeps growing. Cause: The runner process stopped, or its network path to the controller broke (no heartbeat for over 6 s). Fix: Read the log on the runner's server: on Docker docker compose -p spitfire-runner -f deploy/runner/docker-compose.yml logs runner (in ~/spitfire) or docker logs spitfire-runner; on systemd journalctl -u spitfire-runner@1; on Windows %ProgramData%\Spitfire\logs. Test TCP access to the controller's gRPC port (8471) with nc -vz CONTROLLER 8471. If it is the old entry of an ephemeral runner, it disappears by itself after 5 minutes.

Symptom: The runner never shows up; its log says invalid enrollment token. Cause: The key was copied wrong or was revoked. Fix: Create a new key and run the command again with it.

Symptom: The installer stops at the pin check, or the runner reports a TLS error. Cause: --ca-pin belongs to another controller, the controller address is wrong, or a device in between intercepts TLS. Fix: Copy the pin again from the key dialog; make sure the address is this controller's port 8471. With your own certificate, use --tls or --ca.

Symptom: The Run test dialog says "No idle runners with a location." Cause: The runners connected without a location. Fix: Create a location-bound key and reconnect the runner with it, or start the runner with --location.

Symptom: A yellow older release tag next to a runner. Cause: The runner is on another release than the controller. Fix: Updating runners.

Symptom: The CPU column is red; "Runner CPU is at its limit; results may be skewed". Cause: The runner's machine cannot keep up with generating the load; the measured latency includes the runner's own delay. Fix: Add runners and split the load, or use a stronger machine.

Symptom: Clock offset is yellow. Cause: The runner's clock differs from the controller's by more than 50 ms. Fix: Turn on NTP on the runner's server. The controller corrects the offset, but NTP is recommended.

Symptom: Starting a run says "This run needs 2 runners; the free edition allows up to 1 per run." Cause: The license limit. Fix: Use fewer runners or upgrade the license: License.

Symptom: On Kubernetes an on-demand run waits in the queue for a long time. Cause: The pods cannot start (not enough resources, the image cannot be pulled). Fix: See why with kubectl -n spitfire get pods and kubectl -n spitfire describe pod <pod>; check the System events page.