GraphQL load testing: per-operation latency, errors and mutations
In GraphQL every request goes to one address: POST /graphql. An HTTP load testing tool sees one endpoint, yet behind it sit a 5 ms field lookup and a 2-second query joining five tables. On top of that, GraphQL servers usually return errors with HTTP 200. So a GraphQL load test has to get two things right: split the results by operation, and see the error inside the response.
What are we measuring?
- Latency per operation: the p95 and p99 of every query and mutation on its own. A single "/graphql p95" mixes cheap and expensive queries and shows none of them right.
- Errors in the response: a response with a non-empty
errorsarray failed, even with HTTP status 200. The most common GraphQL errors under load are resolver timeouts and errors from downstream services; they do not show at the HTTP layer. - Partial results: GraphQL can return the other fields while one fails (
datatogether witherrors). Decide whether that is expected or a failure, and say so in the test. - Query cost: nested fields (product → reviews → author) can go to the database at every level (N+1). Use the depth of the queries real clients send; shallower ones make the system look faster than it is.
Shaping the load
Take the operation mix from real traffic: the queries the front end sends when a page opens, in that order, should be the journey in the test. Carry an id from one query's result into the next with a variable; querying the same id every time measures the cache. To see the rate where the system breaks, raise the load in steps and look at percentiles, not the average (the p95 and p99 guide).
Common mistakes
- Looking only at the HTTP status. A test showing a 0% error rate may carry
errorsin half of its responses. - Running mutations without noticing. In a test generated from the schema,
createOrderopens a real order in every iteration. Pick mutations on purpose and run them against a test environment. - Persisted queries and caches. If the server or a CDN caches queries, the same query with the same variables only measures the cache. Vary the variables with CSV data.
With Spitfire
In Spitfire every GraphQL operation is its own GraphQL step: the endpoint's address, the query, variables (JSON, templated: {"id": "{{userId}}"}) and the operation name. Metrics are split per step, so each operation's p95 and error rate show separately and get their own threshold. When the response carries errors, the step fails even with HTTP 200, and the first error message shows in the run page's error samples. If partial results are expected, tick Don't fail responses with errors and verify with checks.
- Query from the schema: Load schema on the step sends the introspection query with the step's address and headers. Picking a field fills in the query, sample variables and the operation name; required arguments become variables.
- Test from the schema: in Import API / HAR, give the endpoint as the address (or upload the introspection output as a file); every query and mutation you pick becomes a step (the import guide).
- Mutation guard: a step whose document has a
mutationcounts as changing data on the target; runs, trials and schedules need a write approval, recorded in the audit log.
A test where 200 VUs open the product list and go to the first product's details. The product id from the first query's result is carried into the second with a variable. The list query's p95 must stay under 250 ms, the details' under 400 ms, and the error rate (GraphQL errors included) under 1%. The test was checked with spitfire validate.
{
"name": "GraphQL: katalog ve sepet",
"variables": { "api": "https://staging.example.com/graphql" },
"scenarios": [
{ "name": "alisveris",
"executor": { "type": "ramping-vus", "startVUs": 0,
"stages": [ { "duration": "2m", "target": 200 }, { "duration": "8m", "target": 200 },
{ "duration": "30s", "target": 0 } ] },
"steps": [
{ "id": "products", "name": "Products", "protocol": "graphql",
"graphql": { "url": "{{api}}",
"query": "query Products($first: Int!) { products(first: $first) { id name price } }",
"variables": "{\"first\": 20}", "operationName": "Products" },
"checks": [ { "type": "jsonPath", "path": "$.data.products[0].id", "op": "exists" } ],
"extract": [ { "var": "productId", "from": "jsonpath", "expr": "$.data.products[0].id" } ] },
{ "id": "product", "name": "Product", "protocol": "graphql",
"graphql": { "url": "{{api}}",
"query": "query Product($id: ID!) { product(id: $id) { id name reviews(first: 5) { rating } } }",
"variables": "{\"id\": \"{{productId}}\"}", "operationName": "Product" },
"thinkTime": { "min": "1s", "max": "3s" } }
] }
],
"thresholds": [
{ "metric": "req_duration", "filter": { "step": "products" }, "expr": "p(95)<250" },
{ "metric": "req_duration", "filter": { "step": "product" }, "expr": "p(95)<400" },
{ "metric": "req_failed", "expr": "rate<0.01" }
]
}The GraphQL step produces the same metrics as HTTP (req_duration, req_failed, waiting, connect and TLS times) and shares connections with the same VU's HTTP steps. When HTTP/3 is picked as the HTTP version in the test options, GraphQL requests go over QUIC too. Subscriptions are not supported.
Spitfire installs on Docker or Kubernetes with one command; every testing feature and protocol is open in the free edition.