Skip to content
SEO Evaluate
Related posts
SEO5 min read

Health tourism keyword research in four languages

Collecting patient queries in TR, EN, AR and FA: the traps the languages themselves create - suffixes, letter variants, ZWNJ and transliteration.

By Roozbeh Nazari · CEO

Health tourism keyword research in four languages

The hard part of keyword research in health tourism is not which word gets searched more. The hard part is that the same patient writes the same thing in four languages in forms that bear no resemblance to one another, and that most of those forms never make it onto a standard list at all. This article is about the traps the languages themselves create when you collect queries for TR, EN, AR and FA, and how to normalise them.

We covered the cluster and page architecture side earlier in building a four-language demand map; that was at portfolio level. This is one level down: the grammatical forms of individual queries. For the applied framework on the clinic side, see our clinics page.

Let us say one thing plainly at the outset: no search volume figure is published in this article. Volume data varies by tool, by country and by measurement date; writing a number we cannot verify does nothing except make a wrong number permanent. What is described is a collection and normalisation method you can repeat with your own tool.

Turkish: suffixes and lost letters

Because Turkish is agglutinative, a single root spreads across dozens of surface forms in the search box. "Sac ekimi", "sac ekimi fiyatlari", "sac ekimi yaptiranlar" and "sac ektirmek" are different surfaces of the same intent, and the tool returns them as separate rows. Without root-based grouping the list bloats and priority disappears.

The second trap is sneakier: a significant share of users type without Turkish characters at all. Forms written in plain ASCII are real queries. At the collection stage you need to ask for both the accented and the unaccented forms separately, then reduce them to a single canonical form. Dropping the unaccented form from the list means dropping a share of the patients from the list.

English: patient language versus clinician language

On the English side the difficulty is not grammatical but lexical. For the same procedure, the medical term and the lay term live side by side: "rhinoplasty" and "nose job", "blepharoplasty" and "eyelid surgery", "bariatric surgery" and "weight loss surgery". The patient usually types the lay term into the search box while the clinic writes the page with the medical term; the two never meet.

Regional spelling differences sit on top of that. The distinctions between British and American spelling exist in the health vocabulary too. Whichever your target market is, fixing the list to it beats mixing the two.

Arabic: letter variants and dialect

There are three separate sources of variation in Arabic collection. The first is diacritics: they may or may not be present in text, and they are usually absent from queries. The second is different spellings of the same letter: alif with and without hamza, final ya versus alif maqsura, final ta marbuta versus ha. These vary freely on the user side and the tool counts them as separate queries.

The third is dialect. A patient from the Gulf and a patient from Egypt may search for the same procedure with a different word; a single list written in Modern Standard Arabic does not fully cover both. We covered this from the patient behaviour side in Arabic SEO and GCC search behaviour.

The normalisation rule is simple: collect wide, group narrow. Collect the variants separately, bind them to a single canonical form, and build the page against the canonical form. Applying that rule from the start is far cheaper than cleaning the list later.

Persian: same alphabet, different code points

What gets lost most in Persian comes from the fact that it appears to use the same alphabet as Arabic. It does not. The Persian ye and kaf are different code points from their Arabic counterparts; they are indistinguishable by eye and entirely separate characters to a machine. If the two are mixed in your list, the same word sits there as two separate rows and neither finds the other.

The second issue is the zero-width non-joiner, ZWNJ. In Persian, compound words and some suffixes are separated by this invisible character. Some users type it, some use a space, some use nothing. All three forms are real queries and all three are separate rows.

These two issues affect not only the keyword list but the slug and URL side directly; we cover the equivalent there in the technical SEO checklist for RTL sites.

Arabic and Persian written in Latin letters

This is the group most often missed in four-language lists. A significant share of Arabic- or Persian-speaking users, particularly on a mobile keyboard, type the query in Latin letters. The form that results is neither English nor a standard transliteration; it is a free spelling that varies from user to user, with digits mixed in as letters.

This has two consequences. First, these queries do not appear in the Arabic list because they are not written in Arabic script, and they do not appear in the English list because they are not English words. In a standard research flow they vanish entirely. Second, the user arriving on these queries is looking for content in the target language; routing them to the English page does not meet the intent.

What can be done at collection is to generate the Latin-letter spellings of the main target-language terms by hand, add them to the list as separate rows, and bind them to the canonical form in the target language. These rows do not require a separate page; they are added to the targeting of the existing one.

The mirror image of the same pattern exists in Turkish: an Arabic-speaking user writing in Latin letters and a Turkish user writing without Turkish characters produce the same problem for different reasons, and both are solved the same way, by binding the variant to the canonical form.

The collection and normalisation flow

  • For each language collect the exact phrases the patient writes, not the roots; root grouping is the next step.
  • Query separately for accents, diacritics, letter variants and ZWNJ; do not delete the variant from the list, bind it to the canonical form.
  • Give every row an intent label: information, comparison, price, appointment. The page type follows from that label.
  • Map the equivalent of the same intent across the four languages onto one row; a row that cannot be mapped is a need for a separate page in that language.
  • Keep all four language lists in a single table; separate files drift apart over time.

The URL structure decision comes before the list

A keyword list is useless until it is settled which language lives at which address. Google's multi-regional and multilingual sites documentation recommends using separate URLs for each language version and explicitly does not recommend the URL parameter approach. The same document also says to avoid automatically redirecting users from one language version to another; we set out why that matters in the locale auto-redirect article.

When the list is ready, the last job is to write down which page each row goes to. Leaving rows unmapped amounts to not having done the research. Keeping that mapping in the table is also the only way to answer the question "why does this page exist" six months later.

Sources

// CONTACT

Drop a brief. Send us your brief.

Our intro call is free. Once we have your brief, we'll map the market opportunity and your highest-priority growth opportunities.