Web Vitals under load: what does the user really see?
A load test tells you the API's p95; it does not tell you what happens on the user's screen. After the API answers, the page still runs JavaScript, downloads images and lays out content. When the server goes from 200 ms to 600 ms under load, the page's main content may go from 1.8 seconds to 4 seconds, or not change at all. You only see this by opening the page in a real browser while the load runs.
What are we measuring?
Google's Core Web Vitals sum up the user experience in three numbers, and the browser measures them itself:
- LCP (Largest Contentful Paint): the moment the largest content on the page (the hero image, a heading block) shows on screen. 2.5 seconds or less is good, above 4 seconds poor.
- INP (Interaction to Next Paint): the time from a click or key press to the screen changing. 200 ms or less is good, above 500 ms poor.
- CLS (Cumulative Layout Shift): how much the content moves in front of the user. A late banner pushing the text down raises CLS. 0.1 or less is good, above 0.25 poor.
- TTFB and load time: the first byte and the page's load event. TTFB rising with LCP points at the server; LCP rising while TTFB stays flat points at the front end or another resource.
Do not generate the load with browsers
A real browser is expensive: about 350 MB of memory and a fifth of a core per browser on a real page. Imitating a thousand users with browsers takes dozens of machines, and the load hits the limit of the load generators, not of the system. The right setup is hybrid: generate the load at protocol level (HTTP, gRPC, message queues) and let a few browser users take the same journey next to it and measure. The browsers are there as a thermometer, not for load.
The thermometer must not heat up itself: if the machine running the browsers saturates, the measurements look worse than reality. Run the browsers on machines apart from the protocol load and watch the machine's busyness along with the measurements.
Common mistakes
- Not measuring before the load. Without an unloaded baseline, "LCP 3 seconds" says little. Keep the first minutes of the test at low load or raise the load in steps; the stage where Web Vitals break is the real finding.
- Looking at the average. Google rates Web Vitals by percentile too; look at p75 or p90 (the p95 and p99 guide).
- Clicking too early. User input stops the LCP measurement, and shifts within 500 ms after input do not count towards CLS. Let the page settle in the journey (wait for an element or a text), then click.
- Forgetting the cache. The same browser tab takes images and scripts from its cache on the second visit; that is a returning user. If you want first visits, keep it in mind when you read the results.
With Spitfire
In Spitfire this setup is one test. Protocol steps generate the load; a small scenario with a Browser step takes a user journey in headless Chromium at the same time. The journey is a list of actions, the first one opening a page:
- Open page (
goto): goes to the address and waits for the load event. - Click (
click) and Wait until visible (wait): by CSS selector; they wait until the element is visible. - Type (
type): types text into a field;{{variables}}and CSV data work. - Check text (
expect): the step fails if the text does not show on the page (or in an element) in time. - Sleep (
sleep): reading time. Every action has its own timeout (30 s by default).
Every VU has its own tab and cookies, like a private window; the session is kept across steps and iterations. The step produces browser_lcp, browser_cls, browser_inp, browser_ttfb, browser_load, and req_duration for the whole journey. You write thresholds on them like on any metric; a broken threshold fails the run, in CI too (the CI/CD guide).
A 14-minute test where 400 VUs load the product API while 3 browser users go from the cart to the payment page. The API's p95 must stay under 300 ms, the page's LCP p90 under 2.5 seconds and INP p90 under 200 ms. The test was checked with spitfire validate.
{
"name": "Mağaza: API yükü ve kullanıcı deneyimi",
"variables": { "base": "https://staging.example.com" },
"scenarios": [
{ "name": "api",
"executor": { "type": "ramping-vus", "startVUs": 0,
"stages": [ { "duration": "3m", "target": 400 }, { "duration": "10m", "target": 400 },
{ "duration": "1m", "target": 0 } ] },
"steps": [ { "id": "list", "name": "Ürünler", "request": { "method": "GET", "url": "{{base}}/api/products" },
"thinkTime": { "min": "1s", "max": "3s" } } ] },
{ "name": "kullanicilar",
"executor": { "type": "constant-vus", "vus": 3, "duration": "14m" },
"steps": [ { "id": "buy", "name": "Sepetten ödemeye", "protocol": "browser",
"browser": { "actions": [
{ "do": "goto", "url": "{{base}}/sepet" },
{ "do": "click", "selector": "#checkout" },
{ "do": "expect", "text": "Ödeme bilgileri", "timeout": "10s" } ] },
"thinkTime": { "min": "5s", "max": "10s" } } ] }
],
"thresholds": [
{ "metric": "req_duration", "filter": { "step": "list" }, "expr": "p(95)<300" },
{ "metric": "browser_lcp", "expr": "p(90)<2500" },
{ "metric": "browser_inp", "expr": "p(90)<200" }
]
}The Browser card on the run page colours each step's p90 values with Google's thresholds. It also measures the thermometer heating up: browser_lag is how late the runner's own timer fires; when its p95 goes above 100 ms, the card says the measurements are not reliable and suggests fewer browser VUs or a runner of their own.
The browser step needs a runner with Chromium: the algebransoft/spitfire:<version>-browser image. The For browser steps option on the Runners page prepares the command with this image. Such a runner labels itself browser, and Spitfire gives scenarios with browser steps only to these runners; the protocol load stays on the others. A runner with 4 vCPUs and 8 GB carries about 10 browser VUs.
The number of browser VUs in a run depends on the license: 1 on the free edition, 5 on Growth, 20 on Scale, unlimited on Enterprise. A thermometer rarely needs more than a few.
Spitfire installs on Docker or Kubernetes with one command; the free edition measures Web Vitals with one browser VU.