SEO

How to do keyword research

Keyword research means finding the phrases your customers actually type, measuring them, and turning the result into a list where each group is one page. It has five steps, and for a Persian keyword one rule differs: take the number with no country filter.

  • Lesson 4 of 15
  • Beginner
  • Free, no signup

Five steps, from a handful of seeds to a page map

The narrowing is deliberate: the real work in this process is removing, not collecting.

  1. 1

    Seed

    Ten to twenty terms from customers and Search Console, not from a tool.

  2. 2

    Expand

    Suggestion tools, related searches, related questions, competitor titles.

  3. 3

    Measure

    Search volume with no country filter. Difficulty and intent do not exist for Persian.

  4. 4

    Group

    By the job the user wants done, not by how similar the letters look.

  5. 5

    Map to pages

    Existing page, new page, or neither. The first is usually the right answer.

each group, one page

The band widths show the order of the steps only; there is no numeric ratio in them.

Last checked: Facts and tool names in this lesson are re-checked against their sources on this date.

Where do the seed terms come from?

The first common mistake is starting by opening a tool. A tool only expands what you feed it; if the input is wrong, the output is a thousand wrong keywords.

Seeds come from the business itself. What customers say on the phone, the question that keeps coming back in your messages, the sentence your salesperson answers every day. There is a detail here we have watched play out many times: the word you use inside your industry is usually not the word a customer types. They write the plain name of the thing, not the formal one on your brochure.

Three more free sources that get ignored: the queries report in Search Console, which tells you what you are already shown for and which of those you did not know about; a competitor's category list, which leaks the structure of the market; and Google's own autocomplete, which is in effect a list of things people really type.

A ten minute exercise worth more than an hour with a tool: open your last three customer conversations and write down the exact words they used, without tidying them up. Those clumsy, unprofessional phrasings are what gets typed into Google.

Ten to twenty seeds is enough. More than that, at this stage, is just noise.

How do you grow the list?

Now the tool comes in. Feed every seed to a keyword suggestion tool and collect the output without judging it; deleting is the next step's job. Our free tool on the keyword research page does this for Persian and pulls the number using the rule explained in the next section.

Collect three more sources by hand, because tools tend to miss them. The related searches at the bottom of the results page, the related questions block, and the phrasings repeated across the titles of the first ten results. The third one matters most: a phrasing several competitors put in their titles is a phrasing that has worked in this market.

At this stage your list is probably two hundred to a thousand rows, full of duplicates and junk. That is normal. Keep one rule only: never hand-write a keyword you invented yourself. The list should hold only what came from a source or from a customer's mouth.

Which number do you trust? The rule that differs for Persian

Search volume is the number of times a phrase is searched in a month, on average. Two things to take from Google itself: these numbers are rounded, and the real figure fluctuates constantly. So read two keywords with a small gap between them as equal, not as one beating the other.

Now the rule that differs for Persian. Almost every global tool asks you to pick a country, and for Persian that choice damages the number. Iran does not exist in location based products at all, and if you pick another country you have picked the exit country of a VPN, not the user. What you get back is a slice of the same Iranian audience under the wrong label.

The right move: pick no country. For a Persian phrase the number with no country filter is the real number, not a rough worldwide estimate, because the only other country with native Persian speakers is Afghanistan and its share is small. The measurements behind this are in the field-note block further down the page.

One more fact about the numbers themselves: two different tools give two different figures for the same keyword, because each makes its own estimate. Do not hunt for the correct number; stay inside one tool and compare keywords against each other. The ratios hold, the absolute figures do not.

Two columns that simply do not exist for Persian: keyword difficulty and search intent. Any number a tool shows you under those two headings did not come from Persian data. Intent has to be read off the results page, and the method is in what search intent is.

One Persian phrase: with a country filter and without one

With a country selected

  • Iran does not exist in location based products at all
  • Another country means the VPN exit, not where the user is
  • The number gets split and comes out smaller than reality
  • The same Iranian audience is counted under another country's label

With no country filter

  • The full Persian language figure comes back
  • The user behind a VPN is inside this number too
  • The only other Persian speaking country is Afghanistan and its share is small
  • This is what we print in our own tool, with the explanation beside it

This rule is for Persian phrases only. For an English or Arabic keyword the country filter does its normal job and should stay.

Where does the monthly number mislead you?

Search volume is a monthly average, and an average hides the shape of the curve. A phrase that peaks for one month and is dead for eleven can carry exactly the same average as one searched steadily all year. On paper they are equal; in practice they are two completely different jobs.

That is why our own tool keeps the last twelve months beside the number. The result reader in pack-keywords.php slices the month list down to twelve on purpose, because a seasonal decision is not made from a single figure.

Two practical consequences. If your business is seasonal, publish the page two or three months before the peak; a new page needs time to settle, and publishing in the peak week means arriving after the party. And if a phrase rose because of a one-off news event, that demand does not come back, and the page you build for it comes down with it.

Group the keywords, do not just list them

This is where keyword research turns from a useless spreadsheet into a map. Group the keywords by the job the user wants done, not by how similar the letters look.

The practical test is simple: if the results pages for two phrases look nearly the same and the same sites are on top, Google treats them as one thing and one page should answer both. If the results pages differ, you need two pages even when the keywords look alike.

Every group has one head phrase that the rest orbit. That head phrase is what goes in the page title and the heading, and the rest of the group fits naturally inside the text with no forced repetition.

Do not build two pages for one group. That is the mistake that comes back later under the name keyword cannibalization: two weak pages standing where one strong page belonged, neither of them ranking properly.

One keyword group, one page

one guide pagehead phrase in the title and heading, the rest of the group inside the text
  • The head phrase

    Not necessarily the highest volume phrase in the group; the most accurate description of the job.

    title
  • Other phrasings of the same thing

    If their results pages look nearly the same, they all belong to one page.

    in text
  • Questions on the same topic

    Each becomes a subheading, not a separate page.

    subheading
  • A phrase whose results page differs

    Similar wording is not enough; this one wants its own page.

    another group

This is a made-up example to show the shape of the work, not the keyword list of a real project. Your own group size depends on your market.

Which page does each group go to?

The last step is the one that usually does not happen: next to every group, write down which page answers it. There are only three cases. A page you already have, a page you need to build, or no page at all because the group is not worth writing.

Take the first case seriously. If you have a page written for that same intent, improve it and do not build a second one. A strong page that has been collecting links and history for years is almost always ahead of a fresh one. A new page is justified when the intent genuinely differs.

For each group, open its head phrase in Google and look at the kind of pages on top. If the first ten results are all guides and you were planning a sales page, either change the page type or drop the group. Those ten minutes prevent writing a page that goes nowhere.

Add one column nobody usually adds: the date of the decision and its reason, in one line. Six months later, when a page moves in the rankings, that note is the only thing that helps you. On a site with no record of its own changes, every drop stays a mystery.

The final output is a plain table: group, head phrase, intent, target page, and whether that page is new or existing. That table is what people call a content map, and everything else is built on it.

When should you drop a keyword?

Dropping keywords matters as much as finding them, and it is taught far less. Four cases go without hesitation.

First, when the results page is full of sites you will never replace: a government site for a government service, or a large marketplace for a product name. Even rank two on a page like that gives you nothing. Second, when the intent does not match your work, however large the number is.

Third, phrases that are simply a competitor's brand name. Fourth, phrases that bring traffic but not customers; a web design company does not need a page for every general phrase about the internet.

And one case that runs the other way and should stay: the very low volume phrase. For a service business, a phrase searched a few dozen times a month that describes exactly what you do is usually worth more than a high volume generic one. A big number is not the goal in itself.

The fast path, with AI

The slow part of this job is not collecting keywords, it is grouping a thousand rows. That single step can go to a language model and turn days into an afternoon, provided two hard constraints hold: the model may not add a keyword and may not produce a number. For this volume a fast cheap model of the Flash class is enough; our current pick is in the <a class="text-link" href="/en/ai/">AI section</a>.

  1. Export the raw list with search volumes. If the phrases are Persian, make sure the numbers were taken with no country filter.
  2. Prepare a list of your existing pages too: URL and title is enough. This is what turns the output from an academic exercise into a work plan.
  3. Fill the recipe below with both lists and run it. If the list is long, feed it in chunks and repeat the page list in every chunk.
  4. Open only the head of each group in Google and look at the page types on top. Thirty groups means thirty searches, not a thousand; that saving is the whole value of this route.

Copy-ready recipe

Role: content strategist. Your job is grouping and mapping, not research.

Hard rules:
- Add no keyword. Remove no keyword. Every input row must appear exactly once in the output.
- Produce no number. Copy the volumes exactly as I gave them.
- If you are unsure which group a keyword belongs to, put it in an "unclear" group. Guessing is not allowed.

My current site pages:
{URL | title, one per line}

The keyword list:
{keyword | search volume, one per line}

Return the output group by group, and for each group:
1. The group name, in a few words.
2. The head phrase and a one sentence reason for choosing it.
3. The rest of the keywords in that group with their volumes.
4. The page decision: "existing page: {URL}" or "new page". If the group shares its intent with one of my current pages, you must choose the existing page; a new page only when none of my pages answers that intent.
5. Up to three suggested subheadings drawn from the keywords in this same group. If the group is small, give fewer; do not write a subheading that did not come from the group's keywords.

At the end, separately from the groups, give two lists: the keywords left in the unclear group, and the groups you think overlap with more than one of my current pages.

Before you trust the output: The model has not seen the results pages, so it groups by what the words mean, not by whether Google treats them as one thing. Two phrases that mean the same but have different results pages will land in one group, and that is the only serious error on this route. This is why you open the head of every group yourself. And if the model ever prints a number next to a keyword that you did not supply, that number is fabricated; throw the whole output away and rerun the recipe with the constraint stated harder.

AI in this kind of work

Our position is simple: AI is good at sorting in this job and no good at measuring. Search volume is a measurement and has to come from a source that actually holds it; grouping is a language task and a model is fast and good at it.

Tools that actually help

  • RGB Our own keyword suggestion tool, free and with no signup. For a Persian phrase it takes the number with no country filter and explains that next to the number.
  • Google Search Console The only source holding the real queries for your own pages. The best seeds usually come from here, from phrases you got impressions on without knowing.
  • Gemini For grouping long lists on its fast model. Google's own page says the Gemini web app runs in over 230 countries and territories, and Iran is not on that list.
  • Claude Better on the existing page versus new page call, because it writes down its reasoning and you can disagree with it. Iran is not on Anthropic's supported-countries list and there is no official payment route from Iran.
  • RGB The access and payment layer for Iran, kept separate from the tools themselves.

Where it backfires

Two specific risks. First, a language model has no search volume and will invent it if you ask; the number it returns has the right shape, looks rounded, and is entirely fabricated. Never put a number in a report that did not come from a measurement source. Second, the risk one step further on: a grouped list tempts you to generate one AI page per group. Google's own spam policies describe that as scaled content abuse, and their test is not how the content was made, it is whether it is worthless to the user. A hundred generated pages that say nothing new is exactly what is written there.

Sources: Google Search Central: spam policies (scaled content abuse) Google Ads Help: search volume statistics are rounded

Where this advice stops

Keyword research tells you what people search, not whether they buy. A high volume phrase can bring no customers at all, and a phrase searched a few dozen times a month can be a month of revenue. Volumes are also rounded and they fluctuate, so a small gap between two keywords means nothing. And for a new site the length of the list does not help: one page that matches the intent exactly is ahead of fifty thin ones.

From our own work

We measured the "pick no country" rule against real data and then printed it in our own tool. A phrase tied to an Iranian judicial service, which means nothing outside Iran, shows 40,500 searches a month with no country filter, and the same phrase under a Germany filter shows 9,900. Those 9,900 are not German users; they are the same Iranian users behind a VPN. Across eight terms we also measured the share of Afghanistan, the only other Persian speaking country, at about one percent. The result is rgbr_kw_fa_worldwide_notice() in rgb-rank, which prints this explanation next to every Persian number in the tool, and the difficulty column stays empty right there because that data does not exist for Persian. That notice used to open with the words "this is not Iran" and readers took the number as second class; after this measurement we reversed the order of its sentences.

Real follow-up questions

Where do I get search volume for a Persian keyword?

From any tool connected to Google Ads data, but without selecting a country. Iran is absent from location based products, and picking another country splits the number. Our free tool on the keyword research page does exactly this.

How many keywords does one page need?

One group, not a number. A page has one head phrase plus however many other phrasings of the same intent fit naturally in the text. If you have to force a phrase in to hit a count, that phrase does not belong on this page.

How do I see keyword difficulty for Persian?

You do not, because the data does not exist for Persian and any number a tool shows you came from somewhere else. The manual substitute is this: open the results page for the phrase and look at how large the top ten sites are and how specifically their pages address the topic.