Skip to content

Homoglyph domains and IDN homograph attacks

Updated October 3, 2026

A homoglyph domain replaces letters with characters that render identically or nearly so: a Cyrillic о for a Latin o, or rn for m. No typo is ever typed; the swap is built into the character encoding itself.

Encoding tell
Punycode labels start with xn--
Classic pair
Cyrillic о (U+043E) vs Latin o (U+006F)
Tracked as
MITRE ATT&CK T1583.001

Most lookalike techniques in typosquatting assume someone mistypes. Homoglyph domains assume someone misreads. The registrant swaps characters for lookalikes drawn from the same alphabet or another script entirely, producing a domain that renders the same as yours to anyone reading it.

How the substitution works

Unicode encodes thousands of characters, and many share glyphs. Latin o (U+006F) and Cyrillic о (U+043E) are different characters that most fonts draw identically. The same holds for Greek omicron, Armenian letters that mirror Latin ones, and ASCII-internal pairs like 1/l/I or rn versus m. A domain like notоlens.com, spelled with a Cyrillic о, is a different registration, owned by someone else, that looks exactly right in an email, an ad, or a browser tab.

Attackers chain this into credential theft: the domain serves a pixel-perfect copy of a login page, and the only visible anomaly is invisible. MITRE ATT&CK catalogs acquiring domains for impersonation under technique T1583.001 (Acquire Infrastructure: Domains), homograph registrations included.

Why Punycode hides it

DNS only carries ASCII, so internationalized labels are encoded as Punycode: the Cyrillic notоlens is stored and transmitted as xn--notlens-cjg. Two consequences follow:

  • The raw form betrays the trick. Any label starting xn-- is an internationalized name, and decoding it reveals which characters were swapped.
  • Almost nobody sees the raw form. Mail clients, browsers, and link previews render the decoded Unicode, which is exactly what makes the attack work.

Browsers push back with display policies: a label mixing scripts in ways the Unicode Consortium’s security report (UTR #36) flags as confusable renders in its xn-- form instead of as clean text. The rules help, but they are permissive: single-script spoofed labels and low-risk mixes still display normally, and enforcement differs across browsers and locales.

What registries do, and what is left

Some registries restrict mixed-script registrations or confusable sets at sign-up, but coverage is patchy across the TLD space. A registrant needs only one permissive combination of script and TLD. That gap is why detection has to read the encoded form, not the rendered one.

You can enumerate the homoglyph variants of your own domain with the lookalike domain generator, which produces the same confusable substitutions an attacker would try. For detection at scale, notolens decodes internationalized candidates before matching, so a swapped-script registration surfaces in your inbox like any other match, and reporting it to the sponsoring registrar follows the same abuse-desk process as any deceptive domain.

Sources

Frequently asked questions

Related tools and resources

Continuous brand monitoring

notolens checks daily registrations across 1,570 TLDs, trademark registers, and app stores. When a lookalike domain, conflicting mark, or copycat app appears, notolens checks it, explains the risk, and hands you the records and possible next steps.