Internet and networks

The anatomy of a web address (URL)

A web address is not a string of letters but five parts with different jobs: the protocol, the domain, the path, the query and the anchor. The only part to look at when you are judging whether a link is fake is the domain, and every other part can be anything at all.

  • Lesson 5 of 12
  • Beginner
  • Free, no signup

The five parts of an address, in the order they are written

  1. 1 The protocol

    https before the colon; it says the connection is encrypted, not that the site is honest

  2. 2 The domain

    the only part that decides who owns the address, and the only part that cannot be faked

  3. 3 The path

    which page of that site; the owner writes whatever they like here

  4. 4 The query

    after the question mark; sometimes a token or a key travels here too

  5. 5 The anchor

    after the hash; never sent to the server, it is the browser's business

The order of writing is not the order of reading. To identify who owns an address, read the second link from its end and ignore the other links.

Last checked: Facts and tool names in this lesson are re-checked against their sources on this date.

What parts is a web address built from?

A URL is the address of one specific resource on the web, and RFC 3986 defines its structure. It has five parts, each doing a separate job: the protocol before the colon, the domain after the two slashes, the path after the third slash, the query after the question mark and the anchor after the hash.

On the very page you are reading: https is the protocol, rgb.ir is the domain, and /learn/network/url-anatomy/ is the path. The query and the anchor are empty here, because this page needs no parameter and has nowhere to jump to.

The order in which these parts should be read is the thing most people learned wrong. The eye reads from the left and the first things it sees are the protocol and then the first word of the domain; but what decides who the address belongs to is the end of the domain, not its beginning. The next section is exactly that.

There is also a part that is usually invisible: a port number can sit between the domain and the path, and if you leave it out the protocol's default is used. An ordinary user never needs this part, and knowing it exists is enough.

A chain of five glass segments, the second one largest and red, with a magnifier held over it

Why is reading an address a security skill?

Because the only part that cannot be faked is the domain, and everything else in the address is written by whoever wants to write it. The path can carry your bank's name, the query can contain the word secure, and neither means anything. Who owns an address is decided by the domain and by nothing else.

And the domain has to be read from the end. The last piece is the extension, the piece before it is the registered name, and everything further to the left is a subdomain, which the owner of that name can create as many of as they like. So in rgb.ir.example-site.xyz the owner of the address is example-site.xyz, and rgb.ir is only a subdomain that anybody could have created. The pattern of padding the start of an address with a familiar name is betting on exactly this: that your eye reads from the left and is convinced early.

Two other things get mentioned less. First, RFC 3986 allows something to be written before the domain and before an @ sign, and the browser goes to whatever comes after the @, not to what is written before it. Second, characters can be written in percent form, such as %2F instead of a slash, and an address full of these codes may be hiding something. In both cases the rule is the same: find the last piece of the domain, before the first slash, and ignore the rest.

Let us finish off one misconception here as well: the padlock next to the address only says the connection is encrypted, not that the site is honest. A fake site can obtain a valid certificate, and does. The difference between encryption and trust is opened up in the article on what an SSL certificate is.

One question before you click

What is the last piece of the domain, just before the first slash?

If it is a name you do not know

Do not open it

  • A familiar name at the start of the address is only a subdomain.
  • If there is an @ before the domain, the browser goes to what comes after the @.
  • Instead of clicking, type the site name into the browser yourself.
If it is the domain you expected

You may open it

  • Whatever the path and query are, the owner of the address does not change.
  • Opening it and typing a password are two separate acts; the second still needs care.
  • The padlock confirms encryption only, and says nothing about the site being honest.

This check only catches obvious fakes. A correct domain can itself be compromised, and lookalike names built from similar letters slip past the eye.

What goes on after the question mark?

Whatever the page needs in order to build its answer. The query is a set of name and value pairs separated by an & sign: a page number, a search phrase, a colour filter, a campaign identifier. That is why two addresses with the same path and a different query are two different pages.

And for exactly that reason the query affects something few people expect: caching. On this very site, when you fetch the home page with no query the header x-fp-cache: HIT-nginx comes back, meaning nginx served a prebuilt file and PHP never ran. Add ?cb= with a random number to the same address and that header does not appear at all. The same happened with ?utm_source=. So a campaign tag at the end of a link is, as far as the cache layer is concerned, a different page.

Let us be honest here: in our measurement that difference did not show up in the response time and both cases were around fifty five hundredths of a second, because most of that number is network rather than server. The cost sits elsewhere. If all your advertising links carry a tag, every click goes to PHP instead of a prebuilt file; on a busy site that difference is visible in server load, not on one visitor's stopwatch.

One warning that concerns everyone: the query is where a token sometimes sits. Password reset links and signed links carry their key right there. So forwarding a full address to somebody without looking at it can be handing over a working key.

What to put in a query and what not to

Query sheetafter the question mark

Belongs here

  • A page number, a sort order and a filter, without which the page cannot be built.
  • A search phrase used inside the site.
  • A campaign tag, knowing that the tag takes the page out of the cache.

Do not put here

  • A password, or anything that is itself a key to get in.
  • A user's personal data, which travels on with every forwarded link.
  • Information needed for routing, which belongs in the path itself.

The left column is about what you build yourself. A link somebody sends you may carry any of these, and you have no control over it.

How is the hash at the end different from the rest?

In that it is not sent to the server at all. Whatever comes after the # is the browser's business, and RFC 3986 says as much: the anchor is interpreted by the browser alone. Its ordinary job is to jump to a heading inside the same page.

One practical consequence is useful in daily work: a change of anchor means the page is not reloaded. That is why clicking a heading in an article's table of contents takes you to that section instantly.

And another consequence is useful while debugging: because this part never reaches the server, it never appears in the server log. If what you want to know is which section of a page people open most, the log will not answer.

What is a clean URL, and why does it matter for SEO?

An address a reader can understand without opening the page. Short, made of real words, with no extra parameters and no date. Its main benefit is not a rule inside a search engine but that very understandability: a readable address gets clicked more often in a search result, in a messaging app, and in a link somebody sends by hand.

One rule we follow on this site is worth stating: the page title is in Persian, its slug is in English. Left to itself, WordPress percent encodes the address of a Persian post, and that link, once copied somewhere, becomes a long meaningless string. The site's other languages also use a path prefix rather than a parameter, such as /en/; that single decision gives each language an independent address and a canonical of its own.

And the most important thing to know about an address has nothing to do with how clean it is: the address is the page's identity to a search engine. Changing it means a page that spent years earning standing loses its place, unless you put a permanent redirect in at the same moment. So do not beautify an ugly address that ranks; write the next address properly instead.

If you want to follow this from the SEO side, the on page SEO lesson puts the address next to the title and the headings, and the technical SEO article opens up its technical layer.

The fast path, with AI

Taking a long address apart by eye takes a few minutes, and one small slip in it changes the whole conclusion. With a model it takes seconds, provided you ask it for the right job: parsing, not judging.

  1. Copy the address in full, not a screenshot. The piece that gets left out is usually exactly the piece that mattered.
  2. If the address came from a suspicious message, wrap it in backticks or quotes so that it never turns into a clickable link anywhere. A fast cheap class of model is enough here; the work is text parsing.
  3. Ask for a table of the parts and, above all, ask for the registered domain separately. That single line is the one the eye gets wrong.
  4. Do not ask the model for a verdict of safe or unsafe. The model does not open the link and cannot know what is on the other side; if you ask, it guesses and says so confidently.

Copy-ready recipe

The text below is a web address. Only take it apart; do not open it.

{paste the address here}

Give a table with these rows:
protocol | full domain | registered domain (name and extension only) | subdomains | path | query parameters, one per row | anchor

Then write these three lines separately:
1. If there is an @ before the domain, say which domain the browser actually goes to.
2. Which parameters are only tracking tags, such that the address without them opens the same page.
3. Whether there is a parameter that looks like a token or a key. If there is, write only its name and do not repeat its value.

At the end, give a cleaned version of the address with the tracking tags removed.

Rule: give no opinion at all on whether this address is safe. If any part of the address is percent encoded, also write its readable form.

Before you trust the output: The output of this tells you who an address belongs to and what it is built from; it does not tell you what that site will do with you. A perfectly correct domain can itself be compromised, and lookalike names built from similar letters pass this parsing too, because textually they are healthy domains. So use it to understand, not to get permission; keep entering passwords and verification codes only on a site whose address you typed yourself.

AI in this kind of work

An address is a string of text, and a language model is good at parsing strings: it separates the parts, makes percent codes readable, and says which parameter is only a tracking tag. But there is one job never to hand it, and it is the most requested one: saying whether this link is safe. The model does not open the link and does not see what is behind it. That is our position: the model for reading an address, and for the decision about opening it the rule given on this page, which is to look at the registered domain.

Tools that actually help

  • Claude Works well for taking long, crowded addresses apart, especially when dozens of parameters run together and the eye can no longer separate them. Iran is not on Anthropic's supported countries list, so there is no official signup or payment.
  • Gemini Good for explaining in Persian what each parameter does. Google's own page says the Gemini app works in more than 230 countries and territories, and Iran is not on that list.

Where it backfires

The main risk is the thing many people do without a second thought: pasting a full address into a chat. A password reset address, a single use invitation link and signed links carry their key in the query, and copying the whole address means handing that key to another company. Google's own help page for Gemini says plainly not to enter confidential information. Our rule is simple: before pasting, delete the value of any parameter that looks like a token and leave only its name. The second risk is that confident verdict with nothing behind it: the model cannot open the link, but if you ask whether it is safe it answers in the same tone, and the answer is only a guess from the look of the text. Anthropic itself calls this confident invention hallucination in its documentation and writes about reducing it. For what paying for each of these tools looks like from Iran, see the buying guide.

Sources: Anthropic: reduce hallucinations Anthropic: supported countries Google: Gemini Apps privacy and data Google: where Gemini Apps are available

Where this advice stops

Reading an address properly catches obvious fakes and no more. A perfectly correct domain can be compromised and hand out a malware link, and names built from similar looking letters get past the eye even when you are paying attention. The cache header we showed in the third section belongs to this site's configuration: another site may ignore tracking tags and answer from cache, so look at your own site's headers before generalising.

From our own work

That a query takes a page out of the cache is something we measured on this very site the same day, and the result was cleaner than we expected. The home page with no query returns the header x-fp-cache: HIT-nginx, and the same page with ?cb= does not carry that header at all; with ?utm_source= exactly the same thing happened, so a campaign tag is no exception either. The honest part of the story is that nothing showed up in the response time: both cases were around fifty five hundredths of a second, because what we measured was mostly the network path rather than the server's work. That is also why our own testing method is this: when we want to see a page genuinely from PHP rather than from a cache file, we add a random ?cb= to the address. That small trick saves a whole day of debugging.

Real follow-up questions

What is the difference between a URL and a domain?

A domain is only the site's name, while a URL is the full address of one specific page inside that site. One domain can hold thousands of URLs, but every URL has exactly one domain. If you want to know how a name gets registered and maintained, the article on what a domain is opens that layer.

Should I use a Persian or an English slug for a page?

English, for one practical reason: a Persian address becomes percent encoded when it is copied or shared, and turns into a long unreadable string. Browsers usually display the readable form, but what lands in a messaging app or in your document is that encoded string. On this site we write the title in Persian and the slug in English.

Is there any harm in removing tracking parameters from a link?

Not for you, the same page opens. The only thing lost is the statistics the site owner collects about where the visit came from. If you are running the campaign yourself, keep those parameters; if you are only forwarding a link to somebody, removing them makes it shorter and more readable.