Site speed

How to measure site speed

Measuring site speed means taking one lab report, putting it next to data from real users, and picking a single task out of both. The hard part is neither taking the report nor understanding the numbers, it is resisting the pull to fix everything the tool marked red.

  • Lesson 4 of 10
  • Beginner
  • Free, no signup

One measurement cycle, five steps

Step three is where most speed projects fail, because people skip it and go straight to fixing.

  1. Take the test

    One request gives both the lab number and the field data, if it exists.

    1
  2. Read the report

    Look at the metrics separately, not only at the overall score.

    2
  3. Prioritise

    On the two axes of impact and effort, not in the order the tool printed.

    3
  4. Fix one thing

    Do several at once and you will never know which one worked.

    4
  5. Measure again

    Three runs, and compare median with median rather than best with best.

    5

This cycle shows the effect of one change, not how much money that change made you. That question has a different answer and it does not come out of this report.

Last checked: Facts and tool names in this lesson are re-checked against their sources on this date.

What to measure with, and why one call is enough

You hear three names a lot and they are not the same thing. Lighthouse is the test engine that opens the page under simulated conditions and produces a report. PageSpeed Insights is the service that runs that engine for you and, if real user data exists for that page, puts it alongside. CrUX, the Chrome User Experience Report, is that real user data, and the Core Web Vitals report in Search Console comes from the same source.

Here is the thing that is rarely said and that shortens the work: one request to the PageSpeed API returns both together. In the response, the lighthouseResult block is the lab number and the loadingExperience block is the field data. You do not need two reports from two places.

That same response also gives away a third truth: the field block may simply not be there. For most Persian language sites that is exactly what happens, and it is not a fault. The difference between these two kinds of data, and why the 75th percentile matters, is opened in the previous lesson in this track.

To take a report, the simplest route is opening the PageSpeed service in a browser. If you want to automate it, one request is enough and from there the work happens on a JSON file:

curl -s "https://www.googleapis.com/pagespeedonline/v5/runPagespeed?url=https://example.com&strategy=mobile&category=performance&key=KEY" -o psi.json

One warning that buys you time: this address no longer answers without a key. The daily quota without one is zero and the response comes back 429. The key is free from the Google Cloud console, and that single item is the difference between a real Lighthouse report and whatever a fallback tool decides to show you.

The raw file you get is far too large to read with your eyes or to paste into a chat window. This command extracts only what is useful: the score, the four lab metrics, the state of the field data, and the opportunities above a hundred milliseconds. If the field block is absent from the response it prints field: none by itself, and that is a useful answer in its own right.

python3 -c 'import json;d=json.load(open("psi.json"));l=d["lighthouseResult"];a=l["audits"];print("score",round(l["categories"]["performance"]["score"]*100));[print("lab",k,a[k]["displayValue"]) for k in ("first-contentful-paint","largest-contentful-paint","total-blocking-time","cumulative-layout-shift") if k in a];f=d.get("loadingExperience",{}).get("metrics",{});print("field: none") if not f else [print("field",k,v["percentile"],v["category"]) for k,v in f.items()];[print("fix",int(x["details"]["overallSavingsMs"]),"ms",x.get("title","")) for x in sorted(a.values(),key=lambda x:-x.get("details",{}).get("overallSavingsMs",0)) if x.get("details",{}).get("type")=="opportunity" and x["details"].get("overallSavingsMs",0)>100]'
A lab bench with a white instrument beside a window of red, green and blue night traffic, two rulers side by side

What the score is made of, and what is missing from it

The big coloured number at the top of the report is not itself a measurement. It is a weighted average of several lab metrics, and the Lighthouse documentation publishes each weight: Total Blocking Time 30%, Largest Contentful Paint 25%, Cumulative Layout Shift 25%, First Contentful Paint 10% and Speed Index 10%. The same documentation notes that the weightings have changed over time, because the Lighthouse team keeps researching what has the biggest impact on user perceived performance.

Now look at that list closely and see what is missing from it: INP. One of the three Core Web Vitals contributes nothing to the score. The reason is sensible rather than sinister: INP only means something when somebody actually clicks on the page, and in an automated run nobody clicks. Total Blocking Time is its lab stand-in, not the metric itself.

The practical consequence is one sentence: you can score a hundred and still have a page that does nothing for a full second after the user taps. If your chat widget and ad script do heavy work after the initial load, the score says nothing about it.

So what is the score good for? Comparing one page with itself, before and after a change, under the same conditions. For that it is good. For telling a client whether their site is good or bad, the individual metrics and the field data are the honest instruments.

What the score is made of, and the name missing from the list

weight of each metric in the Lighthouse performance score, per the Lighthouse documentation
  1. Total Blocking Time 30%
  2. Largest Contentful Paint 25%
  3. Cumulative Layout Shift 25%
  4. First Contentful Paint 10%
  5. Speed Index 10%

INP is not on this list, because in an automated run nobody clicks. Total Blocking Time is its lab stand-in, not the metric itself. The Lighthouse documentation also notes that these weightings have changed over time.

Why the number changes every time you test

You get 68, you run it again and get 74, the third run says 61, and nothing on the site changed. This is not a fault, and the Lighthouse documentation gives it a whole section: a lot of that variability is not due to Lighthouse at all, it is due to the conditions the test runs under. The network, momentary server load, even how idle the testing machine happened to be.

One working rule falls out of this: a single measurement is not a measurement. If you want to see the effect of a change, take three runs before and three after and look at the median, not the best number. People always line up their best before number with their best after number and lie to themselves.

Another trap fewer people know: many speed tools, ours included, cache the result for each address for a while. So if you test again a minute later, the number you see may not be a fresh run at all. A stable number can mean a stable site; sometimes it means nothing was re-run. If two consecutive tests return exactly the same figure, suspect this first.

From a long list to one task

Once the report arrives, the opportunities section hands you a long list with an estimated saving next to each item. The common mistake is starting at the top. That list is ordered by what the tool can measure, not by how much work each item is for you.

The right order comes from two axes: impact and effort. Cutting server response time usually has the highest impact and also demands the most effort, because it means going to the host. Shrinking one oversized image is high impact and low effort and always goes first. Removing unused JavaScript is medium impact and high effort, and with a purchased theme it may not even be within your reach.

One rule that takes a long time to discover on your own: do not even read any recommendation promising a saving under roughly a hundred milliseconds. The variability of the test itself is larger than that, so even if you do the work you cannot prove the effect. Our own tool leaves opportunities below that line out of its output for exactly this reason.

And finish the work with a sentence rather than a report: this month we work on this one thing, for this reason. If the outcome of your measurement is a list of ten items, you have not decided yet.

The same list, ordered by something that concerns you

  • High impact, low effort

    Today. The oversized image loading at full resolution usually lives here.

  • High impact, high effort

    Needs planning. Server response time usually lives here, because it means going to the host.

  • Low impact, low effort

    If time is left over. Never the first task of the month.

  • Low impact, high effort

    Do not do it. Any recommendation saving under roughly a hundred milliseconds usually lands here.

Work order

Where each recommendation lands in these four cells depends on your site and is not fixed. Changing hosts is a day of work for someone on a dedicated server and a project for someone on cheap shared hosting.

After the fix, what to measure again

Once the change is in, the lab number answers the same day. Take three runs and compare the median with the previous median. If the difference is smaller than the natural variability of the test, your work has no effect you can prove, and it is better to accept that than to hunt for a number that confirms you.

Field data does not answer the same day, and this is where people get frustrated. That data is collected from real visitors, and until enough of them have seen the new page the report is the old one. So the right timing is this: take the technical confirmation from the lab and the real confirmation from the field a few weeks later.

And never pull one thing out of speed data: a sales figure. No speed report says how much your revenue changes, and anyone who hands you such a number is either quoting a study of another company in another market or guessing. The relationship between speed and sales on your site comes only from your own before and after measurement. The fuller argument is in the first lesson of this track, and if you want to see what a real test method looks like across several Persian sites, our benchmark of ten Persian pages under one fixed method shows exactly that, with the table and the method open.

The fast path, with AI

The usual way is to open the report, read the opportunities from the top and copy everything red into a task list. The problem is not that it is hard, it is that the list you end up with belongs to the tool rather than to your site. The fast path hands the whole raw output plus the business role of the page to a model and gets back one priority with a reason instead of a list. The whole thing takes about ten minutes and usually deletes half of that list.

  1. Take the raw output and shrink it. The PageSpeed JSON response is far too large for a chat window, so extract only what is useful. The command in the body of this lesson does exactly that: the score, the four lab metrics, the state of the field data, and the opportunities above one hundred milliseconds.
  2. Write the business role of the page in one line and put it next to that. This page brings sales, or brings calls, or only informs. Without that line the model just sorts numbers and does exactly what the tool already did.
  3. Give it the prompt below together with both pieces. This is judgment work rather than bulk processing, so a strong model answers better; the fast cheap class is fine for summarising a report but not for picking one priority. The current pick in each class is kept in our AI reference, and no version name is written here because it would be wrong in six months.
  4. Turn the answer into one sentence. The correct output of this work is a decision: this month we work on this one thing, for this reason. If the answer is another list of ten, ask again and force it to pick one and say why not the others.

Copy-ready recipe

You are a senior web performance consultant and today your job is picking one priority, not fixing and not teaching. Judge only from the data below and add nothing from your memory.

Summarised report output:
{paste the output of the command here}

Business role of this page: {brings sales | brings calls | informs only}
What I have access to: {the theme code | only the WordPress dashboard | I can change hosting}

Write six things:
1. In one sentence, say what the main problem of this page is, based on the individual metrics rather than the overall score.
2. Put each opportunity in one of four cells: high impact low effort, high impact high effort, low impact low effort, low impact high effort. Use the "what I have access to" line to estimate effort.
3. Pick the one task to do this month, and say why that one and not the others.
4. Say which recommendations should not be done and why. Anything promising less than a hundred milliseconds belongs in this group.
5. If there was no field data in the input, write plainly that the state of real users is unknown and that your conclusion rests on a single lab run.
6. Write what has to be measured again after that one task, to tell whether it worked.

Rules: do not estimate any figure for increased sales or conversion, even if asked; no such number comes out of this data. Do not add any problem that is not in the input. If you are torn between two tasks, pick the one closer to the business role of the page and say so. If the data shows this page has no serious speed problem, write that plainly; "none of them, this page has a different problem" is an acceptable answer.

Before you trust the output: Two things the model does not know and you do. First, how much you can actually change: it estimates effort from a line you wrote yourself, so if that line is optimistic the priority comes out optimistic too. Second, one lab run is still one run; before you spend money, take the same page twice more and see whether the priority moves. And the rule you keep yourself even if the model says otherwise: accept no sales uplift figure out of this work.

AI in this kind of work

A language model is genuinely useful on this topic, but not for the question most people ask. The bad question is "how do I speed up my site"; the answer is a generic list that fits every site and decides nothing for any of them. The good question is to put your own report output in front of it and ask it to choose. In the second case the model works on your data; in the first it assembles a list from memory.

Tools that actually help

  • Claude Suited to the fast path job here: it takes the report output with the business role of the page and, instead of listing, picks one priority and says why. Iran is not on Anthropic supported-countries list, so there is no official signup and no Iranian card is accepted.
  • Gemini If you decide not to shrink the output and hand over the whole JSON file, it handles very long inputs well. Google own page says the Gemini web app runs in over 230 countries and territories, and Iran is not on that list.
  • ChatGPT Enough for summarising and ordering the report. We make no claim about access from Iran, because the OpenAI supported-countries page, like the rest of that domain, returns 403 to this server.

Where it backfires

Three risks, and the first is common enough to deserve its own sentence: the model invents audit names that are not in your report. It has seen the pattern of Lighthouse names and can produce something that looks like one but does not exist, and because the shape is right nobody questions it. The simple defence is to search every title the model wrote inside your own file; if it is not there, it is not real. Anthropic own documentation calls a model stating with confidence something it cannot back a hallucination. The second risk is mixing lab and field: the model takes the number from a simulated run and issues a verdict about the experience of real users. The third risk is about the data rather than the model: report output carries page addresses and sometimes a client internal URLs, and what happens to text you paste into a chatbot depends on that service plan and settings. If you are working on a client site, read that service data usage page once before pasting. Our position: the model may order and choose from your report data, and may not add anything to it.

Sources: Anthropic: reduce hallucinations Anthropic: supported countries Google: where Gemini Apps are available

Where this advice stops

This lesson teaches you how to measure and how to decide; it makes no page faster and produces no number about your sales. It has three other boundaries. First, the score weights and the API behaviour were read from Google documentation as it stands today, and that documentation is alive; the Lighthouse documentation itself notes the weightings have changed over time, so if a long time has passed since the verified date at the top of this page, open the source links once. Second, we publish no score as a Lighthouse score unless it came from Google own service with a key; the output of any fallback tool is a different thing and has to be called by its own name. Third, this method is written for small and medium business sites; at a scale with a dedicated performance team and continuous real user monitoring the cycle is different, and we have no first-hand experience at that scale.

From our own work

We wrote the speed test on this site ourselves against Google own API, and three things came out of that work that we have seen in no tutorial. First, one response carries both kinds of data: our code reads the lab number from lighthouseResult and five field values from loadingExperience, and we had to guard every one of those five with an existence check, because that block is very often simply absent. Second, that address without a key is dead: we called it from this very server on 8 September 2026 and got a 429 back, with a daily quota whose value is written as zero, not exhausted, zero. In that case our code falls back to an HTML level check of our own, and we do not call that output a Lighthouse report, because it is not one. Third, we cache each address result for ten minutes; good for the server, not good for someone who thinks they are taking a fresh test. Now what these decisions cost us: a free tool that depends on an API key changes shape silently and the user does not notice. That is why every number published on this site as a Lighthouse score carries its own date on the page.

Real follow-up questions

How many runs do I need before the number is reliable?

Three runs before the change and three after, and always compare medians rather than best numbers. The Lighthouse documentation itself notes that much of the variability comes from the conditions the test runs under rather than from the site. If the difference between medians is smaller than the spread, there is no provable effect.

Should I test mobile or desktop?

Both, but if you only have time for one, pick mobile. The Core Web Vitals assessment measures the two separately, and for most Persian language sites the mobile share is larger. The desktop number is almost always the better one, which is exactly why looking at it alone leaves you unjustifiably relaxed.

What is the difference between a free speed tool and a paid one?

The important difference is not which one costs money, it is which engine actually runs behind it and whether it really runs that engine. Google official service no longer answers without a key, so a tool that always answers instantly either has its own key or is running something else. The right question to ask any tool is where this number came from.