Unicode Normalizer

Unicode has four normalization forms and they are not interchangeable. NFC and NFD rearrange the same character between one code point and a base plus a combining mark; NFKC and NFKD go further and rewrite characters that merely look like one another — the fi ligature, the fullwidth A, the circled ①, the Roman numeral Ⅻ. This page puts the same text through all four, prints a row per character saying what changed and why, counts graphemes against code points and UTF-8 bytes, and compares two strings under the form you choose. The second half is about the characters that never show up at all: zero width spaces, joiners, bidirectional controls, variation selectors, C0 and C1 control characters, private use code points and broken surrogate halves, each with its position and its name, and the option to strip them from a copy of the text.
All processing happens locally in your browser — your data is never uploaded to any server.

Not sure what normalization is? Unicode normalization explained in plain English — why é can be one character or two, and why that breaks a login.
Form to report and to compare with:
Compare the two boxes after:
NFC (nothing yet)
NFD (nothing yet)
NFKC (nothing yet)
NFKD (nothing yet)
Nothing compared yet. Paste the two strings and press Compare the two strings.
first string, normalized (nothing yet)
second string, normalized (nothing yet)
Nothing scanned yet. Press Scan for invisible and dangerous characters, then Strip if you want the text with those characters removed.
text with the selected classes removed (nothing yet)
Ctrl+Enter in the first box runs the active mode.
Pick a mode and click its action button. Normalize writes all four forms of your text and a row per character; Compare says whether two strings are the same after normalization; Scan lists the characters that do not show up, and can strip them. The Self-test button re-runs every documented case in this page's data file, every character the invisible and control tables name, and reports the count here.
Every character, one row each: what the character is, what the four forms do to it, and why. A shaded row is a character that at least one form changes; the NFD column shows what the character decomposes into when it decomposes at all.
How many characters is that? The same text counted four ways: grapheme clusters (what a reader counts), code points (what Unicode assigns), UTF-16 units (what JavaScript indexes) and UTF-8 bytes (what travels). Normalization moves the numbers between the second and the third column without changing the first.
What was found, in the order it appears: the position counts code points, the class is the one the character belongs to, and the note says what that character does to whatever reads the text next.
The classes this page treats as suspicious: the table is printed from the scanner itself, so the words on this page and the behaviour of the buttons cannot drift apart. Severity is about the effect on a comparison, not about how likely the character is to be innocent.
Private use characters are listed as suspicious for a different reason from the rest: they do render, they simply have no agreed meaning, so the same code point is a different picture in every font that defines it.
How large an input this page accepts, and why: normalization is a rebuild of the whole string, and this page does several of them plus a scan and a grapheme count, so the cost grows in a straight line with the length of the input. The caps below exist because the per-character table is drawn in the document, and a table with a million rows in it is what actually stops a browser.
What is Unicode normalization?

Unicode gives more than one way to write some characters. The letter é can be a single code point, U+00E9, or an ordinary e followed by a combining acute accent, U+0065 U+0301. The two look the same on screen, they print the same, and as strings they are different: they have different lengths, different UTF-8 bytes, and 'é' === 'e\u0301' is false.

That is not a curiosity, it is a class of bug. Whether you get one form or the other depends on the keyboard layout, the operating system, the input method and whether the text was ever copied through a word processor. A password check, a database unique index, a blocklist or a cache key that compares strings byte for byte will treat the two spellings as two different values — so a user can set a password on one device and be unable to log in from another, and the same word can appear twice in a list of "already taken" names.

The 30 second version: NFC and NFD are the same text written two ways — NFC composes (e + accent becomes one character), NFD decomposes (one character becomes e + accent). They are meant to be lossless and are the two forms to use for comparison and storage. NFKC and NFKD do everything the first two do and then also replace compatibility characters: the ligature fi becomes two letters, ½ becomes 1⁄2, the Roman numeral ⅲ becomes XII, a fullwidth A becomes an ordinary A. That is exactly what a search index or a keyword filter wants, and exactly what you must never apply to a password.
Why four forms and not one

The split exists because two different goals were being served. Composing and decomposing text is a canonical change: the result means the same thing and, for the first two forms, nothing is lost. Folding away ligatures, superscripts, circled letters and fullwidth forms is a compatibility change: it is done for the convenience of matching, and it does destroy the distinction between two strings that Unicode says are different. The four forms are every combination of the two decisions.

The table above is the case list from this page's data file, checked against the platform's own String.prototype.normalize, and the Self-test re-checks every row of it in front of you. Read the first two rows together: they are the same text, written two ways, and after normalization they are the same string.

Worked examples

These are the strings the Sample button walks through, normalized live by the page rather than typed in as answers. Each one was chosen because a form does something visible to it.

Compatibility characters: what NFKC costs you

NFKC and NFKD are lossy, and the loss is deliberate. They answer the question "is this the same word", not the question "is this the same string". Three consequences follow, and all three are easy to get wrong in production code:

  • Never store a password, a token or a signature in a compatibility form. Two different inputs would collapse to the same stored value, which turns a length or a character check into a weaker one. Normalize passwords with NFC (or not at all) and compare them as bytes after doing exactly the same thing on both sides.
  • A compatibility fold can turn one character into several. The ligature fi is one code point that becomes two letters, so a field that was validated for length before normalization can be longer afterwards. Normalize first, then validate, then store.
  • NFKC does not remove everything invisible. A variation selector survives NFKC and NFKD, and so does a zero width joiner. That is why this page has a scan mode: normalization answers a question about a character's shape and identity, and the scanner answers a question about whether something is there at all.
Graphemes, code points and bytes

Normalization changes the second of these counts and usually leaves the first alone. The three numbers are all "the length of the text" and all different, which is why a form field with a 20 character limit can belong to a string that is 20 code points, 30 UTF-16 units and 47 bytes long.

The number in the grapheme column is what a person counts. The gap between it and the code point count is exactly what the scanner is about: a name made of 12 code points can display as 12 letters, as 3 letters with 9 invisible marks between them, or as one emoji.

The characters you cannot see

A zero width space occupies no width, changes no glyph, and is a real character with a real code point. Two identifiers that look identical on screen can differ by one of them, which is how a name in a list can be an impostor that sorts, searches and authenticates differently from the name next to it. The same is true of the bidirectional controls, which can make a string render in an order different from the order it was written in — a trick that has been used to disguise file names.

The scanner in the third mode lists every one of them with its position: zero width and joiner characters, bidirectional controls, variation selectors, unusual spaces, C0 and C1 control characters, private use code points, and a surrogate that is not half of a valid pair, which is not a character at all. The class table further up this page says what each class does, and the stripping button removes a copy of those characters without touching your input.

The mistakes this page is built to catch
  • Comparing two strings without normalizing them. 'é' === 'e\u0301' is false. Normalize both sides with the same form before comparing, and use the same form in the index you compare against.
  • Normalizing on one side only. A lookup that normalizes the query and not the stored value finds nothing, and looks like a data problem rather than a normalization one.
  • Styling a form with a compatibility form and calling it NFC-safe. NFKC applied to a password shortens the space of possible passwords and merges inputs that were meant to be distinct.
  • Assuming a normalization form fixes a spoof. It does not remove a zero width joiner, a variation selector or a homoglyph — a Cyrillic а and a Latin a are different code points that NFC leaves alone.
  • Trusting a character count. The count a user sees, the count a database column enforces and the count a byte limit enforces are three different numbers, and the rows above show how far apart they can be.
  • Assuming NFC is what you will get. Text that arrives from an input method, a file or another system is in whatever form that system produced. This page starts on NFC because that is the form the web is expected to use, not because that is what your input is.
Quick answers

“Why are these two strings not equal when they look the same?”

Because one of them is written with a composed character and the other with a base letter plus a combining mark, or one of them contains a character that draws nothing at all. Switch to the Compare mode and put one string in each box: the page reports the first position where the normalized strings differ and what is at that position in each of them.

“Which form should I use in an application?”

NFC for anything a human will read, display, store and search, which is the form the W3C recommends for the web. NFD is the right answer when the code behind you wants base letters and separate marks, which is common in text processing and in some file systems on macOS. Use NFKC or NFKD only where you are matching rather than storing: search indexes, keyword filters, duplicate detection.

“Is a string that is already normalized still normalized after I concatenate to it?”

Not necessarily. Normalization is not preserved by concatenation: joining two NFC strings can produce a sequence that NFC would have written differently, because a combining mark at the start of the second string can attach to the character at the end of the first. Normalize after assembly, not only before it.

“Does normalization remove the invisible characters?”

No, with one partial exception: NFKC and NFKD do replace the no-break space with an ordinary space, which is a visible result of a compatibility fold. Everything else — zero width spaces, joiners, bidirectional controls, variation selectors — passes through all four forms untouched. That is what the third mode exists for.

“Why does the byte count change when the character count does not?”

Because the two forms are built from different code points that cost different numbers of bytes in UTF-8. A precomposed é is two bytes; the decomposed pair is three. The counts table above prints all four numbers side by side for the input and for each form, so the difference is visible rather than described.

“How do I know the results on this page are right?”

Press Self-test. It normalizes every case in this page's data file, re-checks the four forms against each other for the documented relationships (NFC is idempotent, NFD of an NFC string is the same as the NFD of the original, and so on), runs every character named in the invisible and control tables through the scanner, and counts the grapheme cases against Intl.Segmenter.

Standards and credit

The four forms and their definitions are from the Unicode Standard, Annex 15 (Unicode Normalization Forms). The code points used as decomposition examples are the ones the standard assigns; this page does not implement the decomposition tables itself — it calls the browser's own String.prototype.normalize, which is the same implementation the platform uses for text input, sorting and comparison, so what is reported here is what the rest of the system does with the same string. Grapheme clusters come from Intl.Segmenter where the browser has it, and from a documented approximation where it does not.

The page text was written for this site. Every worked example on this page was produced by the page's own code and then checked against the data file's recorded case list, which was in turn checked against the Unicode NormalizationTest data for the same characters. The invisible and dangerous character tables are a curated subset written for human readable warnings, not the full Unicode security profile (UTS #39), and the page says so rather than implying a completeness it does not have.