Pick a mode and click its action button. Normalize writes all four forms of your text and a row per character; Compare says whether two strings are the same after normalization; Scan lists the characters that do not show up, and can strip them. The Self-test button re-runs every documented case in this page's data file, every character the invisible and control tables name, and reports the count here.
What is Unicode normalization?
Unicode gives more than one way to write some characters. The letter é can be a single code point, U+00E9, or an ordinary e followed by a combining acute accent, U+0065 U+0301. The two look the same on screen, they print the same, and as strings they are different: they have different lengths, different UTF-8 bytes, and 'é' === 'e\u0301' is false.
That is not a curiosity, it is a class of bug. Whether you get one form or the other depends on the keyboard layout, the operating system, the input method and whether the text was ever copied through a word processor. A password check, a database unique index, a blocklist or a cache key that compares strings byte for byte will treat the two spellings as two different values — so a user can set a password on one device and be unable to log in from another, and the same word can appear twice in a list of "already taken" names.
The 30 second version: NFC and NFD are the same text written two ways — NFC composes (e + accent becomes one character), NFD decomposes (one character becomes e + accent). They are meant to be lossless and are the two forms to use for comparison and storage. NFKC and NFKD do everything the first two do and then also replace compatibility characters: the ligature fi becomes two letters, ½ becomes 1⁄2, the Roman numeral ⅲ becomes XII, a fullwidth A becomes an ordinary A. That is exactly what a search index or a keyword filter wants, and exactly what you must never apply to a password.
Why four forms and not one
The split exists because two different goals were being served. Composing and decomposing text is a canonical change: the result means the same thing and, for the first two forms, nothing is lost. Folding away ligatures, superscripts, circled letters and fullwidth forms is a compatibility change: it is done for the convenience of matching, and it does destroy the distinction between two strings that Unicode says are different. The four forms are every combination of the two decisions.
The table above is the case list from this page's data file, checked against the platform's own String.prototype.normalize, and the Self-test re-checks every row of it in front of you. Read the first two rows together: they are the same text, written two ways, and after normalization they are the same string.
Worked examples
These are the strings the Sample button walks through, normalized live by the page rather than typed in as answers. Each one was chosen because a form does something visible to it.
Compatibility characters: what NFKC costs you
NFKC and NFKD are lossy, and the loss is deliberate. They answer the question "is this the same word", not the question "is this the same string". Three consequences follow, and all three are easy to get wrong in production code:
- Never store a password, a token or a signature in a compatibility form. Two different inputs would collapse to the same stored value, which turns a length or a character check into a weaker one. Normalize passwords with NFC (or not at all) and compare them as bytes after doing exactly the same thing on both sides.
- A compatibility fold can turn one character into several. The ligature
fi is one code point that becomes two letters, so a field that was validated for length before normalization can be longer afterwards. Normalize first, then validate, then store.
- NFKC does not remove everything invisible. A variation selector survives NFKC and NFKD, and so does a zero width joiner. That is why this page has a scan mode: normalization answers a question about a character's shape and identity, and the scanner answers a question about whether something is there at all.
Graphemes, code points and bytes
Normalization changes the second of these counts and usually leaves the first alone. The three numbers are all "the length of the text" and all different, which is why a form field with a 20 character limit can belong to a string that is 20 code points, 30 UTF-16 units and 47 bytes long.
The number in the grapheme column is what a person counts. The gap between it and the code point count is exactly what the scanner is about: a name made of 12 code points can display as 12 letters, as 3 letters with 9 invisible marks between them, or as one emoji.
The characters you cannot see
A zero width space occupies no width, changes no glyph, and is a real character with a real code point. Two identifiers that look identical on screen can differ by one of them, which is how a name in a list can be an impostor that sorts, searches and authenticates differently from the name next to it. The same is true of the bidirectional controls, which can make a string render in an order different from the order it was written in — a trick that has been used to disguise file names.
The scanner in the third mode lists every one of them with its position: zero width and joiner characters, bidirectional controls, variation selectors, unusual spaces, C0 and C1 control characters, private use code points, and a surrogate that is not half of a valid pair, which is not a character at all. The class table further up this page says what each class does, and the stripping button removes a copy of those characters without touching your input.
The mistakes this page is built to catch
- Comparing two strings without normalizing them.
'é' === 'e\u0301' is false. Normalize both sides with the same form before comparing, and use the same form in the index you compare against.
- Normalizing on one side only. A lookup that normalizes the query and not the stored value finds nothing, and looks like a data problem rather than a normalization one.
- Styling a form with a compatibility form and calling it NFC-safe. NFKC applied to a password shortens the space of possible passwords and merges inputs that were meant to be distinct.
- Assuming a normalization form fixes a spoof. It does not remove a zero width joiner, a variation selector or a homoglyph — a Cyrillic
а and a Latin a are different code points that NFC leaves alone.
- Trusting a character count. The count a user sees, the count a database column enforces and the count a byte limit enforces are three different numbers, and the rows above show how far apart they can be.
- Assuming NFC is what you will get. Text that arrives from an input method, a file or another system is in whatever form that system produced. This page starts on NFC because that is the form the web is expected to use, not because that is what your input is.
Quick answers
“Why are these two strings not equal when they look the same?”
Because one of them is written with a composed character and the other with a base letter plus a combining mark, or one of them contains a character that draws nothing at all. Switch to the Compare mode and put one string in each box: the page reports the first position where the normalized strings differ and what is at that position in each of them.
“Which form should I use in an application?”
NFC for anything a human will read, display, store and search, which is the form the W3C recommends for the web. NFD is the right answer when the code behind you wants base letters and separate marks, which is common in text processing and in some file systems on macOS. Use NFKC or NFKD only where you are matching rather than storing: search indexes, keyword filters, duplicate detection.
“Is a string that is already normalized still normalized after I concatenate to it?”
Not necessarily. Normalization is not preserved by concatenation: joining two NFC strings can produce a sequence that NFC would have written differently, because a combining mark at the start of the second string can attach to the character at the end of the first. Normalize after assembly, not only before it.
“Does normalization remove the invisible characters?”
No, with one partial exception: NFKC and NFKD do replace the no-break space with an ordinary space, which is a visible result of a compatibility fold. Everything else — zero width spaces, joiners, bidirectional controls, variation selectors — passes through all four forms untouched. That is what the third mode exists for.
“Why does the byte count change when the character count does not?”
Because the two forms are built from different code points that cost different numbers of bytes in UTF-8. A precomposed é is two bytes; the decomposed pair is three. The counts table above prints all four numbers side by side for the input and for each form, so the difference is visible rather than described.
“How do I know the results on this page are right?”
Press Self-test. It normalizes every case in this page's data file, re-checks the four forms against each other for the documented relationships (NFC is idempotent, NFD of an NFC string is the same as the NFD of the original, and so on), runs every character named in the invisible and control tables through the scanner, and counts the grapheme cases against Intl.Segmenter.
Standards and credit
The four forms and their definitions are from the Unicode Standard, Annex 15 (Unicode Normalization Forms). The code points used as decomposition examples are the ones the standard assigns; this page does not implement the decomposition tables itself — it calls the browser's own String.prototype.normalize, which is the same implementation the platform uses for text input, sorting and comparison, so what is reported here is what the rest of the system does with the same string. Grapheme clusters come from Intl.Segmenter where the browser has it, and from a documented approximation where it does not.
The page text was written for this site. Every worked example on this page was produced by the page's own code and then checked against the data file's recorded case list, which was in turn checked against the Unicode NormalizationTest data for the same characters. The invisible and dangerous character tables are a curated subset written for human readable warnings, not the full Unicode security profile (UTS #39), and the page says so rather than implying a completeness it does not have.