gRPC load testing: unary and streaming calls, HTTP/2 connections and protos
gRPC services often carry a system's internal traffic: the order, stock or payment services behind the API gateway call each other over gRPC. These calls are small and fast compared with REST, but HTTP/1.1 habits mislead in a load test: gRPC runs on long-lived HTTP/2 connections, a status code rather than an HTTP code carries the error, and messages are protobuf, not JSON. This guide covers what you need to know to set up a gRPC load test properly.
Unary and streaming
gRPC has four call types: unary (one request, one response), server streaming (one request, a sequence of responses: a price feed, a long list arriving in pieces), client streaming (a sequence of messages, one response: a bulk upload, adding cart lines one by one) and bidirectional (bidi) streaming (both sides send on the same stream whenever they like: chat, games, continuous sync). Most services consist of unary methods: every call counts as a request, and its latency and status code are measured.
A single duration is not enough for streaming methods; a stream's total time is often set by the message count and interval you chose and says little about the server's speed. The measurements to look at are:
- Time to first message: from opening the stream to the first response. Authentication, subscription setup and the first query show here; in server streaming it is the time the user waits.
- Latency per message: in bidi, the time from sending a message to its response. Measuring it means knowing which response belongs to which message; if the server does not answer messages in order, that pairing misleads.
- Time between messages: in server streaming, the gap between two consecutive responses. If it grows under load, the server cannot keep the stream fed.
- Messages sent and received, and the number of streams open at once. A long-lived stream takes one slot of the connection's stream limit for as long as it is open; the connection count below matters even more for streams.
Do not judge streaming methods in the same table as unary calls; write their thresholds on their own measurements.
The connection count: the most common mistake
HTTP/2 carries many requests on one connection at the same time (multiplexing). So a load testing tool could send every virtual user's (VU's) calls over a single connection, but that leads to two misleading results:
- The server's stream limit. Servers cap the streams open at once on one connection (often 100–128). Once the cap is reached, new calls wait in the client, and that wait is measured as latency: the server is not slow, the client is queueing.
- An L4 load balancer. A load balancer that distributes by connection (a Kubernetes Service, for example) sends all calls of one connection to one pod. A test over one connection loads only one pod of a 10-pod service.
Use a connection count close to what the clients that really call your service open. If a few client instances share many calls, a few shared connections are realistic; if there are many independent clients (mobile apps, edge devices), a connection per VU is.
Realistic messages and method mix
What loads a gRPC service is not only the number of calls; the content of the message matters too. Asking for the same id on every call comes back from the database cache and looks faster than it really is. Take ids, users and query parameters from a data file or random generators, and keep page sizes of list methods close to production. Mix the service's methods in production proportions as well: if read methods are called far more often than write methods, the test should do the same. Send authentication metadata (tokens, tenant) as the real flow does; the work an interceptor does on every call is part of the load.
Status codes and deadlines
A gRPC call's result is a status code: every code other than OK is an error, but they do not all mean the same. UNAVAILABLE usually means the connection could not be made or the server shed the load, DEADLINE_EXCEEDED that the call ran out of time, RESOURCE_EXHAUSTED that a quota or limit was hit, UNAUTHENTICATED that credentials are missing. Watching which code grows as the load grows is the quickest way to tell whether the problem is in the network, in capacity or in a limit setting.
Keep the deadline close to what production clients use. A very long deadline does not turn slowness into errors, it only stretches latency; a very short one produces errors even while the server is healthy. Write thresholds as percentiles: unary latency is low, so even a few milliseconds of drift at p99 matter (the p95 and p99 guide).
With Spitfire
Spitfire's gRPC step calls unary, server-streaming, client-streaming and bidirectional methods 0.17.0+. The method list labels each streaming method by kind (server stream, client stream, bidi stream), and the editor shows only the options that fit it.
- Connection. Add a gRPC connection under Connections: the address (
host:port), TLS if needed (CA certificate, client certificate and key, SNI), the authority and a bearer token (added to every call asauthorization: Bearer …; a step's own authorization metadata wins). Secret fields are stored encrypted. - Schema: reflection or proto. If server reflection is on, the schema is read from it. If not, upload the
.protofiles with their imports to the connection; the controller compiles them and runners receive the compiled schema. The connection's Test button checks that the services can be listed with reflection, or that the connection comes up with protos. - Step. Pick the method from the list (
package.Service/Method); Fill in the template lays out the message as JSON with every field, and you type the values. The message and metadata can use variables, CSV data and generators such as{{$uuid}}. The response goes to checks and extraction (JSONPath) as JSON; response metadata and trailers are read as headers. - Connection count. By default the VUs of a runner spread over 8 shared HTTP/2 connections by VU number; you can change the count in the connection's settings or choose a separate connection per VU. That keeps you clear of the stream-limit and load balancer traps above. Review it especially when many VUs hold long streams open.
A step calling an order service's unary method; orderId comes from a variable or a CSV file, and the response's status field is checked.
{
"id": "get_order", "name": "Siparişi getir",
"protocol": "grpc", "connection": "orders-grpc",
"grpc": {
"method": "shop.v1.OrderService/GetOrder",
"message": "{\"orderId\": \"{{orderId}}\"}",
"metadata": [ { "key": "x-tenant", "value": "acme" } ]
},
"checks": [ { "type": "jsonPath", "path": "$.status", "op": "eq", "value": "PAID" } ]
}During the run each step's requests per second, p95, p99 and error rate are watched live; error kinds are split by status code (DEADLINE_EXCEEDED timeout, UNAVAILABLE refused, UNAUTHENTICATED authorization, the rest apart). The test's request timeout (30 s by default) becomes a unary call's deadline. Thresholds are written per step (p(95)<50), and with a constant arrival rate a breakpoint test finds the rate the server can no longer keep up with. gRPC steps can also make up a user journey together with HTTP, Kafka or SQL steps in the same test.
Stream steps
One step is one stream: it opens, sends its messages, reads the responses, and the step ends when the stream does.
- Sending. Server streaming sends the request message once. Client streaming and bidi send either a list of messages (templates, in order) or the same message a given number of times;
{{__MSG}}is the message's number (0, 1, 2…). They wait the interval you set between messages, then the client closes its side. For servers that end the stream when the client closes, bidi can keep the sending side open. - Ending. The stream ends when the server closes it, when the expected number of responses arrived, or when the duration you set runs out; then the client closes it, and that is not a failure. If none of these happens it ends at the deadline: by default the test's request timeout plus the step's own sending time, and running past it is a
DEADLINE_EXCEEDEDfailure. If the server ends the stream before the expected responses arrived, the step fails asincomplete. - Checks and extraction see the last response or every response as a JSON array (
$[0].id; body contains for "in any message"). The status is the stream's gRPC status; the counts of messages sent and received are in thegrpc-messages-sent/grpc-messages-receivedheaders. - Metrics. The step's duration (
req_duration) is the whole stream.grpc_time_to_first_messageis the time from opening to the first response; in bidi each response is paired with the oldest message not yet answered, and the time between them isgrpc_message_latency; a response with nothing to pair (server streaming, extra pushes) reports the time since the previous one asgrpc_message_gap.grpc_messages_sentandgrpc_messages_receivedare counters.
A step that sends 20 messages 200 ms apart to a chat service's bidi method and waits for 20 responses; the check sees every response and confirms the last one answers the last message. The thresholds below cap the p95 of the first message and of the latency per message. The step was checked with spitfire validate in a test with a connection of that name.
{
"id": "chat", "name": "Sohbet",
"protocol": "grpc", "connection": "chat-grpc",
"grpc": {
"method": "chat.v1.ChatService/Talk",
"message": "{\"roomId\": \"{{roomId}}\", \"text\": \"mesaj {{__MSG}}\"}",
"stream": { "count": 20, "interval": "200ms", "receive": 20, "body": "all" }
},
"checks": [ { "type": "jsonPath", "path": "$[19].text", "op": "eq", "value": "mesaj 19" } ]
}"thresholds": [
{ "metric": "grpc_time_to_first_message", "filter": { "step": "chat" }, "expr": "p(95)<300" },
{ "metric": "grpc_message_latency", "filter": { "step": "chat" }, "expr": "p(95)<100" },
{ "metric": "req_failed", "expr": "rate<0.01" }
]Every stream metric accepts thresholds, and the thresholds show in the report. The run page's gRPC streams table lists, per step, the messages sent and received and the p95 of the first message, the message latency and the time between messages; they are in the CLI summary too, and comparisons put the p95 of the first message and of the message latency side by side.
Limits to know:
- The table fills when the run ends. The gRPC streams table shows once the run is stored; what you watch live during the run is the stream step's requests per second, p95 and error rate (the duration being the whole stream).
- Bidi pairing relies on order.
grpc_message_latencyis right for servers that answer each message in order. If the server answers out of order, leaves some messages unanswered or sends several responses to one, this metric misleads; look at the time to first message, the gap between messages and the stream's total time instead.
Spitfire installs on Docker or Kubernetes with one command; every testing feature and protocol is open in the free edition.