IDN Homograph Checker

Two names can be spelled with different characters and still read the same way. apple.com and аpple.com — where that first letter is the Cyrillic small a — are different names as far as the network is concerned and the same word as far as a reader is concerned. This page takes a name apart and says which of those two it is: it splits the name into labels, decodes every xn-- label back to the characters it carries, counts the bytes against the DNS limits, names each look-alike character together with the script it borrows, and folds the whole name onto the shape it imitates. It reports what it finds and what it does not find, at the same level of detail, so a clean answer means the checks ran and passed rather than that nothing was looked at.
All processing happens locally in your browser — the name you type is never uploaded and never resolved.

New to this? IDN homographs explained in plain English — what punycode does and does not do, why a browser sometimes shows xn--, and why this page is a set of checks rather than a verdict.
Read the input as: Fold the name to NFC first
READS AS —
DNS FORM (what a resolver gets) —
SIZE —
Nothing checked yet. Type or paste a name above — one name at a time — or press Load an example. The page reads a bare name, a name pasted out of a URL, or a whole URL, and tells you which one it took. Every check it runs is listed under the result, including the ones that found nothing.
What is an IDN homograph?
A homograph is a word written one way and read another. In a domain name the trick usually costs one character: apple.com is seven Latin letters, while аpple.com begins with U+0430, the Cyrillic small letter a, which is drawn with the same shape in every font you have. As far as the network is concerned those are two different names. As far as your eye is concerned they are the same word, and only one of them is Apple's.

Names were ASCII-only for the first two decades of the DNS, which is why everything about them is measured in bytes and letters: labels of letters, digits and hyphens, at most 63 bytes each, at most 255 bytes in total. An internationalized domain name (IDN) is what happens when the rules are opened up to the rest of Unicode, and punycode (RFC 3492) is the spelling rule that keeps the old machinery working: a label that contains anything outside ASCII is written as a xn-- label, an ASCII string that carries the same characters in a folded form. 中文.com travels as xn--fiq228c.com, and both spellings name the same host.

The part that matters here is what punycode does not do. It converts spelling and nothing else. It has no opinion about whether аpple.com should be allowed to sit next to apple.com, and it could not be asked, because that question is about eyes and fonts rather than about bytes. The registry decides which characters it will sell, the browser decides how it will display a name it considers safe, and the reader decides whether to click. This page answers the question underneath all three: which characters is this name actually made of, and what does that make it spell?

The cases this page is checked against

Every row below is run by this page when it loads and again by the Self-test button, and the page has to report exactly what the row says it reports — the same findings, the same level, the same DNS form, the same byte count, no more and no less. That is what makes a clean answer here worth something: the code path that lets an ordinary name through is the same one that catches the tricks below.

What the case is Level Name as it arrives What the page has to report

How this page decides
  1. Take the name out of what you pasted. A bare name is used as it stands; a URL has its host taken out of it and the scheme, port, path and query are dropped, and the page says which of the two it did rather than quietly analysing a slash as if it were part of a name.
  2. Clean it the way a client would. Control characters, the zero width characters and the byte order mark are dropped, a full width or ideographic full stop counts as a dot, and a trailing dot is the root rather than a mistake. Every character removed or reinterpreted this way is reported, because a zero width space that was silently dropped is exactly the thing you need to be told about.
  3. Split it into labels and measure each one. A label is not allowed to be empty, to begin or end with a hyphen, or to carry two hyphens in positions three and four unless it is an xn-- label. Each label is measured in UTF-8 bytes against the 63 byte ceiling, and the whole name against 255.
  4. Decode every xn-- label back to its characters. The decoded label is then encoded a second time and compared with what you typed. A label that does not survive that round trip is not the label it claims to be, and one that decodes to a control character or to nothing at all is reported as such instead of being quietly dropped.
  5. Ask what each character is. Its code point, its UTF-8 bytes, the script it belongs to, whether it sits in a group of characters that are drawn alike, and what the other members of that group are. This is where a Cyrillic a stops being "a letter" and becomes "a letter that is shaped like a Latin one".
  6. Fold the name onto the shape it imitates. Every look-alike character is replaced by the character it resembles, which turns аpple.com into apple.com. That folded reading is the single most useful line on the page, because apple.com is something you can act on where "a Cyrillic a" is not.
  7. Collect the findings and report the worst one. Findings are reported once each, with the labels they occur in. A name that only looks unusual produces notes rather than findings: mixing Han and Kana, or ending in the root dot, is ordinary writing rather than a reason to worry.

The levels, and what they claim

none nothing here was worth mentioning — low the same name written a second way — medium a character that DNS will not carry, or one that folds onto a plain letter — high a name built to be read as a different name, or one that cannot be sent at all.

A level is the worst thing the findings could mean, not a probability and not a judgement about the site. A perfectly honest name can be a medium because it contains one character outside the host name alphabet, and a hostile one can be a high with nothing more sophisticated than a Cyrillic letter in it. The list under the level says why, one line per finding, and that list is the part worth reading.

The checks it runs, one by one

The table below is built from the same list the page reports from, so a row here and a finding in a result cannot disagree. A check that found nothing is listed as no: the point of the table is that you can see the question was asked and answered, rather than wondering whether it was asked at all.

Check Worst it can mean What it looks for

The same list drives the notes, which carry no level at all: one label with no dot in it, a trailing root dot, a top level domain that is not ASCII, a separator other than a dot, an upper case name, characters that were removed while cleaning, and a URL that was reduced to its host.

The limits DNS enforces

A name is a list of labels and every one of them is counted in bytes, not in characters. That distinction is the whole reason punycode exists: 中文 is two characters and six bytes, and its xn-- form is fifteen characters, so the encoded spelling is what has to fit. The page applies both ceilings and reports which one you hit.

Limit Value What happens at the edge
Why a browser sometimes shows xn--

You may have noticed that a name written in Chinese, or in Cyrillic, sometimes appears in your address bar as xn-- gibberish instead of the characters you expect. That is the browser declining to display the nice form, not the site failing to register it. A browser applies a rule of its own to a name it has never seen before, and the rule exists for the reason this whole page exists: displaying аpple.com as аpple.com is exactly the thing an attacker wants, so an unfamiliar mixed script name is shown in its encoded spelling until it has been seen often enough to be trusted.

Where the name appears What is shown Why

What this page will not do
This is a set of checks, not a verdict, and not a blocklist. It reports what a name is made of; it has no idea what the site does. A name can pass every check here and still be a phishing page, and a name can be flagged as a mixed script name and be a perfectly ordinary web site belonging to someone who writes their own language. The shape groups are a hand written subset of the confusables in UTS #39, so a character pair this page does not know about is one it will not report. And nothing here is resolved, downloaded or looked up: the name you type never leaves the browser, which also means the page cannot know whether a name exists.
Where the rules come from

Punycode is RFC 3492, and the self-test on this page runs the eleven sample labels printed in that document as its first check, so the encoder here is held to the examples in the standard rather than to itself. The processing rules an IDN goes through are RFC 5891 with the character tables of RFC 5892, and the compatibility layer that browsers actually apply is UTS #46 with the non-transitional flag; this page tells you what those rules do rather than pretending to implement them, and the folding it does apply — NFC then lower case, then RFC 3492 — is the same short sequence a client performs before a name is sent. The shape groups, the script ranges and the invisible character tables in this family of pages were written by hand for the family: the standards are named so that you can read them yourself, and no text from them is reproduced here.

Questions

Is a mixed script name always an attack?

No, and treating it that way would be wrong in a way that matters. Japanese writes Han characters and Kana in the same word, and Korean writes Han characters inside Hangul ones. What has no ordinary explanation is Latin mixed with Cyrillic or Greek in one label, because those scripts are not used together in any language, and the page draws that line rather than flagging every mixture. That is why the Japanese and Korean rows in the case table above come back clean.

Then why does my own name come back as medium?

Most likely because it contains one character that the host name alphabet does not include — an underscore, a space, an accent written as a compatibility character — and that is a medium because the name cannot be sent as it stands. Medium means "this needs a look", not "this is a forgery". The finding under the level names the exact character and its position.

Does the page tell me whether a name is dangerous, or registered, or reachable?

No, on all three. It never touches the network and never resolves anything, so it cannot know whether a name exists or who owns it, and it deliberately says nothing about what a site does. It answers a narrower question that you can check for yourself: which characters is this name made of, and if I fold the look-alikes onto the letters they imitate, what does it spell?

My name is in Chinese and the page says the top level domain is not ASCII. Is something wrong?

Nothing is wrong. A Chinese top level domain travels as an xn-- label of its own, so the last label of the name is written in its encoded form and the page notes it. Notes carry no level because they describe a fact rather than a problem.

Why does the reads-as line sometimes show the same name I typed?

Because the page only shows a folded reading when folding actually changes something. If every character in the name already is the shape it imitates, the folded reading and the name are the same string and there is nothing to tell you.

What is an A-label and a U-label?

An A-label is the ASCII spelling that travels: xn--fiq228c. A U-label is the same label written in whatever script it belongs to: 中文. They are two spellings of one label, and the page shows both for every label it decodes so that you can see what an encoded label is hiding.