How a browser builds a page
The browser reads HTML from the first byte and builds the page tree as it goes, but it draws nothing on screen until it has understood the CSS, and if it reaches an ordinary script it stops right there. Every piece of speed advice you will hear along this path comes out of those two sentences.
- Lesson 2 of 10
- Beginner
- Free, no signup
From the HTML arriving to the first thing you see
The five steps a browser takes for every page. The first three can be slowed down; the last two are the result.
-
1
Request
Until the first byte of HTML arrives the browser has nothing to start on.
-
2
Parse and build the DOM
It proceeds gradually and stops on an ordinary script.
-
3
Read the CSS
The second map is built. Nothing is drawn until this is ready.
-
4
Render tree and layout
The two maps merge and the place of everything on screen is worked out.
-
5
Paint
The first thing the reader sees. Any delay in the three steps before shows up here.
This order is simplified. Real browsers run parts of these steps in parallel and speculatively, and that is what makes the effect of a change not always as large as this picture leads you to expect.
Last checked: Facts and tool names in this lesson are re-checked against their sources on this date.
What happens from the moment the HTML arrives
The first thing to get out of your head is that the browser waits for the whole file to arrive. It does not. It starts reading at the first byte and builds the structure of the page while the rest of it is still on the way.
What it is doing has a name: parsing. The browser turns bytes into characters, characters into tokens such as a paragraph opened here, and those tokens into nodes. The nodes connect into a tree called the DOM, the map the browser holds in memory of your page structure. Everything JavaScript changes later is changed on that tree, not on your file.
The most important property of this work is that it is incremental. The browser does not need to wait for the closing tag; if something is ready to show, it shows it. That is why you sometimes see a page whose top has arrived while its bottom is still coming. A slow site is not necessarily one where everything arrives late; more often it is one where something got in the way of that gradual start.
And that is exactly what makes the whole speed subject understandable. The right question is not how heavy is my page, it is what stopped the browser from starting sooner. The next two sections are the two main answers to that question.

Why CSS holds the display but an image does not
Having the page tree is not enough to draw. The browser knows there is a heading here, but not what size it is, what colour it is, or whether it is meant to be visible at all. It learns that from CSS and builds a second map from it, called the CSSOM. Only when both maps are ready can the browser build the render tree and paint.
So why does it wait? Because the alternative is worse. If it painted the text before the CSS was ready, you would see a moment of unstyled black on white and then everything would jump into place. Browsers treat that as a worse experience than waiting a little, which is why CSS blocks rendering by default. That is not a flaw, it is a decision.
Now the distinction that catches many people out: an image does not do this. If a picture has not arrived, the appearance of the text does not change, so the browser has no reason to wait for it; it paints the page and the image settles in later. Of course, if you reserved no space for that image, it shoves the rest of the content when it lands, and you get the jump that irritates a reader more than anything else.
A practical example you can check right now: the front page of this site has no link rel=stylesheet tag at all. The entire theme CSS is sent inside the HTML response itself, which means no network round trip is needed for the first paint. This is not the only correct route and it has a cost of its own, but it shows that optimise the CSS file is not the only option on the table.
Where JavaScript stops in this path
When the parser reaches an ordinary script tag it stops reading the HTML, fetches the file, runs it, and only then continues. While that happens no new node is added to the tree, which means the rest of the page, however ready it is, waits.
The reason is logical. A script is allowed to write into the page at that moment or to change the tree, and the browser cannot know in advance whether this one does. So it does the most cautious thing and waits. There is a point that gets made less often too: a script sitting in the head may wait for the CSS as well as blocking the parser, because it can ask how large an element is right now, and the browser cannot answer until the CSSOM is ready.
Two words change this. defer tells the browser to fetch the file now, in parallel with reading the HTML, but to leave running it until parsing is finished, and to keep the scripts in order. async says fetch it and run it wherever it lands, with no guarantee about order. For almost any script that belongs to your page, defer is the right answer; async makes sense for independent things such as an analytics tag that has nothing to do with the rest of your code.
And this is where file weight misleads you. A two megabyte image at the bottom of the page holds nothing up. A four kilobyte script sitting in the head without defer, served from a third party you do not control, pushes back the first paint of the entire page by one full round trip to that server. Size matters, but placement matters more.
One script, two different places in the path
The same file, the same size. The only difference is one word, and that word decides when the reader sees anything.
An ordinary script in the head
- the parser stops there and the page tree stops growing
- if the file comes from another server, a full round trip is spent waiting
- it may also wait for the CSS to be ready, not only block the parser
- the reader sees a blank page until all of that is over
The same script with defer
- the download starts at once, in parallel with reading the HTML
- the parser does not stop and the page tree is completed
- execution happens after parsing and the order of scripts is kept
- the reader sees the page while that file is still on its way
There is one real exception: a script that has to run before the first paint, such as deciding a light or dark theme, is deliberately left blocking so the page does not flash white and then go dark.
So CSS first and JavaScript later is not folklore
The advice is old, and in most places where it is repeated the reason given for it is wrong. The real reason is not that JavaScript is heavy. It is that rendering waits for CSS, so the CSS has to arrive early; and the parser waits for scripts, so a script should not stand in the middle of the road. It is one sentence about the order of two waits, not about the weight of files.
On the front page of this site you can count that order. In the head there are zero scripts. There are seven script tags on the page and all seven sit in the last three percent of the document, which is to say after all the content the reader is meant to see. The CSS is not a separate file at all and comes inside the response itself. The result is that for the first paint the browser waits on nothing from the network.
Now the most honest part of this lesson: the arrangement above is the old shape of the solution. Putting scripts at the end of the body works, because by the time the parser reaches them everything is already built. But the better shape today is usually to put that same script in the head with defer, because then the download starts sooner and execution still falls at the end. We have not made that change on this site yet; the gain is small, and on a page whose scripts are already this far back it has not been worth the risk of the change.
And a boundary worth knowing: a few scripts genuinely do have to run before the first paint, such as the snippet that decides light or dark theme so the page does not flash white and then go dark. That one is kept blocking on purpose, and that is the right decision. The rule push everything back has one exception, and this is it.
What holds up the first paint, and what is merely heavy
These do not hold up the display
- images, however large they are
- scripts marked with defer or async
- CSS sent inside the HTML response itself
These hold up the display
- every CSS file linked in the head
- every script in the head with neither defer nor async
- a script served from a third party domain and sitting in the head
The right column means these do not hold up the first paint, not that they are free. A large image with no reserved space shifts the content when it lands, and that is a different problem.
How to see this on your own site
You need no tool for this. Open the page source and look inside the head only. Count two things: how many link rel=stylesheet tags there are, and how many script tags with a src that have neither defer nor async. The sum of those two numbers is the list of things holding up the first paint of your page. If the second number is not zero, that is usually the first and cheapest job.
If you would rather do it from the command line, this command prints those tags so you can see for yourself which of them carry defer or async:
curl -s https://example.com/ | tr '\n' ' ' | sed 's/<\/head>.*//' | grep -oE '<link[^>]*stylesheet[^>]*>|<script[^>]*src=[^>]*>'The simpler version of this command, the one that cuts the head range line by line with sed, gives a wrong answer on minified pages, and we spent time on it ourselves. Minified HTML is usually one long line, so that range stays open to the end of the document and counts the scripts at the foot of the body as well. The command above first throws away everything after </head>, which is why it works in both cases.
One warning about reading the result: modern browsers have a separate scanner that looks ahead while the main parser is stopped on a script, and starts fetching the later files sooner. So the parser stops does not mean the network sits idle. This is one of the reasons removing a blocking script sometimes has less effect than expected, and why measuring before and after cannot be replaced with reasoning.
And that is the next step: so far you know what can hold up the display, but not which of those actually did it on your site, or by how much. That question is answered by measurement, and the lesson on measuring site speed in this path takes it on. If you want a real number from your page right now, the testing tool on the site speed page runs on the Lighthouse engine.
The fast path, with AI
When a page is slow, the ordinary move is to open the waterfall in developer tools and look for the tallest bar. The trouble is that the tallest bar is usually not the culprit; the culprit is the thing everything else waited for, and that is often a short bar near the start. A language model is good at exactly this kind of reading: it sees several dozen timing rows at once and finds the this waited for that pattern sooner than your eye does. The whole job takes ten minutes, and its first condition is that you do not put the raw file in front of the model.
- In the browser, open developer tools, go to the network tab, disable the cache and load the page once from scratch. Save the output as a HAR file.
- Do not paste the HAR file anywhere as it is. That file holds every request header, which means your login cookies and session tokens are inside it too. Pull out five columns only: address, file type, start time, duration, and whether it was in the head. The first fifty rows are enough.
- Hand the table over with the recipe below and ask explicitly for a distinction between two things: a file that took long itself, and a file that started late because it was waiting for something else. The entire value is in that separation. This needs a strong model because it is judgment; the current pick in each class is kept in our AI reference.
- Do not believe the answer, test it. Change one thing only, the one the model named, and measure again. If the number does not move, the hypothesis was wrong, and that is a result too. Two changes at once means you never learn which of them worked.
Copy-ready recipe
You are a web performance specialist. Judge only from the table below and add nothing about this site from your memory.
The table below was taken from one full load of the page {page address}. For each resource: address, type, start time in milliseconds from the beginning of the load, duration in milliseconds, and whether it was in the head.
{resource table}
Write five things:
1. Separate out the resources that could have held up the first paint: only CSS, and scripts without defer or async that were in the head. Set the rest aside and say why you set them aside.
2. For each of those, say which case it is: it took long itself, or it started late because it was waiting for another resource. For the second case, name that other resource.
3. Name the one resource that most likely added the largest delay to the first paint, and say in one sentence which columns of the table led you there.
4. Propose one specific change that can be made on its own and then measured.
5. Write down what is not in this table that could have changed your answer if it were.
Rules: invent no number that is not in the table and assume no resource that is not in it. File size alone is not a reason for blocking; if you want to appeal to size, say why it matters in this case. If the table is not enough to conclude, write that and say which column you would need instead. "There was no blocking resource in the head" is an acceptable answer.
Before you trust the output: There is an inherent limit no model removes: a waterfall says what happened when, not why. Time order looks like causation and is not, and the model sees exactly that same order. So the output of this exercise is a hypothesis rather than a diagnosis; the only thing that turns it into a diagnosis is changing that one item and measuring again. And a point that came up in this lesson too: browsers speculatively start fetching files ahead of the parser, so removing a blocking script sometimes has less effect than the chart suggests. If the number does not move after the change, that is not your failure; that is measurement doing its job.
AI in this kind of work
On this subject a language model has two completely different behaviours, and the difference between them decides whether your time is wasted. Put the real timing table of your page in front of it and it is excellent: it sees several dozen rows together and finds the waiting relationship between them sooner than you do. Put nothing in front of it and merely ask how to make your site fast, and you get the usual generic list, written for any site and accurate for none.
Tools that actually help
- Claude Suited to the job in the fast path: it takes the table and, instead of listing everything, names one resource and says which column led it there. Iran is not on Anthropic list of supported countries, so there is no official route to sign up or pay.
- Gemini If your table is large and you would rather not cut rows, it handles long inputs well. Iran is not among the regions where this service is available.
- OpenAI The most common choice, and enough for reading a fifty row table. We make no claim about access or payment from Iran, and the link goes to the maker page rather than a product page: no OpenAI consumer page opens from this server, and we do not write about something we could not see ourselves.
Where it backfires
Two risks, and the first belongs to this subject specifically. A model leans hard towards naming the largest file as the culprit, because in the text it was trained on, heavy and slow are near synonyms. But as you saw in this lesson, blocking is about placement and resource type rather than size: a large image at the bottom of the page holds nothing up, and a small script in the head holds everything up. If the answer only names the biggest row in the table, it probably answered from the pattern rather than from your data. The second risk is a security one and harder to undo: a HAR file holds every request header, meaning your login cookies and session tokens. Pasting that file raw into any online tool is in practice sending the key to your account to an outside service, and Anthropic own security guidance says anything entering the context window should be treated as untrusted input. Our position: cut the table down to five columns by hand, and before believing any resource the model names, change it once yourself and measure.
Sources: MDN: critical rendering path MDN: the script element, defer and async Anthropic: mitigate jailbreaks and prompt injections Anthropic: supported countries Google: where Gemini Apps are available
Where this advice stops
This five step picture is the classic path and it is enough to understand by, but it falls behind reality in two places. First, modern browsers do not run these steps strictly one after another: a separate scanner runs ahead of the parser and speculatively starts fetching later files, parts of layout and painting happen in parallel, and that is why the effect of a change is not always the size this diagram leads you to expect. Second and more important: if your site is a single page application and JavaScript builds the content in the browser, this picture explains only the beginning of the story. There the first paint can happen quickly with the page still empty, and the real time goes somewhere this lesson does not cover.
From our own work
The numbers in this lesson come from the front page of this site on 7 September 2026, and each can be checked by viewing the page source: zero link rel=stylesheet tags, zero scripts in the head, and seven script tags, all seven sitting in the last three percent of the document. The entire theme CSS is about 91 kilobytes and is sent inside the HTML response itself. Now the thing we actually learned while writing this lesson, which is worth more than those numbers: the first version of the command in section five cut the head range line by line with sed, and it gave a wrong answer on this very page. The HTML of this site is minified and is nearly one long line, so that range never closed and ran to the end of the document; the result was that it counted those same seven scripts at the foot of the body as scripts inside the head. Had we published that command without testing it, it would have given a wrong answer to anyone whose site has a cache, and it would have been exactly the mistake this lesson warns about: correct reasoning over data you did not check yourself.
Real follow-up questions
Why does my page show for a moment without any styling?
Because part of your CSS is being loaded in a way that does not block rendering, for example with JavaScript or with tricks that make a stylesheet non blocking. The browser does not wait, it paints the text, and then the styling arrives and everything shifts. The fix is to let the styles needed for the first screen arrive normally and blocking, and the rest afterwards.
If I put defer on every script, will something break?
It can, and the most common case is an inline snippet in the page that depends on a library which now runs later. Inline code does not follow the defer rule and runs in place, so the order breaks and you get a something is not defined error. Go one at a time, and after each change actually open the page and look at the error console.
So should I inline the CSS inside the page the way you do?
Not necessarily, and this is a trade rather than a pure improvement. Inlining removes one network round trip, but those same bytes are sent again in every response and are no longer cached across pages. It is worth it for us because the HTML of this site is cached at the edge and a visitor usually does not open several pages in a row; if your site is dynamic and a user goes through ten pages in sequence, a separate cached file is probably the better choice.