OpenTelemetry in load testing: closer to the root cause
A load test tells you how the system looks from outside: p95 went up, the error rate rose. The real question is inside: which service, which database, which resource? Spitfire aligns your systems' own traces, metrics and logs with the run's load stages and writes the measured change with its numbers. It installs no agent.
Open standards, open tools
Spitfire speaks only open standards and open tools. An admin defines the connections on the Observability page; Spitfire talks only to the systems listed there, and tokens and passwords are stored encrypted like connection secrets.
| Source | Direction | What Spitfire uses |
|---|---|---|
| OTLP | your Collector → Spitfire | traces, metrics, error and warning logs |
| Prometheus | Spitfire → query_range | your PromQL: p95, error ratio, resources per service |
| Tempo · Jaeger | Spitfire → trace API | a sample of each stage's traces |
| Loki | Spitfire → query_range | error lines per service (LogQL) |
Prometheus-compatible backends are read through one Prometheus connection: Grafana Mimir (with its /prometheus prefix and the X-Scope-OrgID tenant), Thanos and VictoriaMetrics. The queries are your PromQL; $__interval is replaced with the query step. Every connection has a Test button.
The OTLP receiver: one exporter on the Collector
Spitfire serves OTLP/HTTP on its own HTTP port (/otlp/v1/traces, /otlp/v1/metrics, /otlp/v1/logs; protobuf or JSON). Add one exporter to your OpenTelemetry Collector, next to the ones it already has:
exporters:
otlphttp/spitfire:
endpoint: https://spitfire.example.com/otlp
headers:
Authorization: "Bearer sfotlp_…"
compression: gzip
service:
pipelines:
traces:
receivers: [otlp]
processors: [batch]
exporters: [otlphttp/spitfire] # + your existing exporters
metrics:
receivers: [otlp]
processors: [batch]
exporters: [otlphttp/spitfire]
logs:
receivers: [otlp]
processors: [batch]
exporters: [otlphttp/spitfire]The receiver keeps only what falls inside a run window (30 s before the start to 60 s after the end) and applies per-run caps; everything else is answered as rejected, so your Collector does not retry it. The token is shown once; only its hash is stored.
The Backend tab: the inside, stage by stage
Query connections are read once the run has ended (after 90 s, for late data), for the run's window only. The run page's Backend tab shows, for each load stage: p95 and error rate of every service, database and called system; a load and latency chart; the slowest operations with their trace ids; resource measurements; error log counts.
With the test's traceparent option on, every HTTP request carries a W3C traceparent header; runners keep the trace ids of each step's slowest and first failed requests, and the Backend tab finds those requests in your traces: the service's own time and the slowest span under it. It is off by default, because keeping every request's trace at load-test volume can cost storage on your tracing backend.
On the run page, Show resource metrics on the charts draws one resource series (a Prometheus query or an OTLP metric) on the run's own charts, on its own axis, with lines where each load stage begins.
Findings: only what the data shows
Backend findings follow the same rule as run findings: they write only what was measured, with its numbers, and say so when data is missing. They never say "because", only "at the same time": they report that two measured changes happened in the same stage without claiming one caused the other.
At the 450 VU stage orders-db (database) p95 rose from 18 ms to 410 ms (240 spans). At the same time the checkout step crossed its threshold: p95 520 ms, limit 300 ms. Slowest operation: SELECT … FROM order_items.
Ranking resources at the breaking point
When the run broke (a p95 or error-rate threshold was crossed, or a step of a breakpoint run failed), Spitfire takes the stage it broke in and the stage before, and ranks every resource series from the connected sources: first absolute saturation (CPU at 95% of its limit or more, memory at 90% or more, a connection pool with rising waiting requests, a queue that at least doubled), then the largest relative jump against the stage before. The top three go into one finding. The ranking is deterministic: the same data always gives the same finding.
The system broke at 560 VUs (the p95 threshold was crossed: 900 ms, limit 500 ms); at the same time the orders-db connection pool filled up (waiting connections 0 → 48) and checkout CPU went from 41% to 97%.
Frequently asked
Do I need to install an agent?
Does Spitfire tell me the cause of the slowdown?
Are values in database statements stored?
SELECT … FROM order_items); values are never stored. Raw data is pruned with the run logs; the computed result stays with the run.Spitfire installs on Docker or Kubernetes with one command; every testing feature and protocol is open in the free edition.