Pick a direction and run it. All processing stays in your browser.
What is Punycode?
In one sentence: Punycode is a way of spelling characters that are not plain English letters — Chinese, accented Latin, Cyrillic, emoji — using only the 37 characters the domain name system was built for: a-z, 0-9 and -. It is defined in RFC 3492, and in a domain name you recognize it by the xn-- prefix: 中文 becomes xn--fiq228c.
So Punycode is not a language, a font or a country code — it is a translation between two spellings of the same name. The Unicode spelling is called a U-label (中文.com, bücher.de), the ASCII spelling is called an A-label (xn--fiq228c.com, xn--bcher-kva.de), and the whole idea of putting international names into DNS is called IDN (internationalized domain names) or IDNA, the set of rules built on top of it.
It exists because DNS was designed in the early 1980s around a deliberately tiny alphabet. A host name that travels through DNS, into a zone file, a resolver cache, a log line or a TLS certificate, must be plain ASCII — the wire format simply has no field for anything else. Punycode is the escape hatch: it lets a name that is not ASCII travel as if it were, and lets the receiving end reconstruct the original characters exactly. That is why a Chinese domain works in a browser at all, and why xn-- is the one string you will see in every place where the Unicode spelling cannot be sent.
Punycode in 30 seconds
| What you type (U-label) |
What DNS gets (A-label) |
What it means |
中文.com | xn--fiq228c.com | The Chinese words for “Chinese language” |
中国 | xn--fiqs8s | “China” in Chinese |
例え.jp | xn--r8jz45g.jp | Japanese examples |
bücher.de | xn--bcher-kva.de | German bücher, “books” |
münchen.de | xn--mnchen-3ya.de | Munich |
í ½í¸‚.com | xn--e28h.com | An emoji label — legal in the encoding, not in most registries |
A length comparison worth remembering: 中文 is 2 characters and 6 bytes in UTF-8, but xn--fiq228c is 11 ASCII bytes. Punycode usually makes a name longer, never shorter.
How the encoding actually works
Take bücher and watch the three moves. The engine is called a bootstring, and the whole algorithm is only a few dozen lines of code (you can read it in the script of this page).
- Copy the plain characters to the front. Every basic ASCII character is copied in its original order and kept readable:
bücher → bcher. Because something was copied and something was left over, a single - is appended as the boundary: bcher-.
- Write the rest as base-36 numbers. The remaining characters are handled in increasing code-point order, and each one is written as a variable-length number in base 36 where
a=0 … z=25 and 0=26 … 9=35. For bücher the leftover character ü becomes the digits kva.
- Stick the prefix on. Encoded part appended to the copied part gives
bcher-kva, and in domain context it is written as xn--bcher-kva.
The trick that makes this clever is what those base-36 numbers contain: not the character itself, but a delta — how far the character is from the previously encoded code point, plus where it belongs inside the string. The encoder walks code points in ascending order and each digit run describes a position and a distance. After every character a shared number called the bias is adjusted so the following number needs as few digits as possible.
Two consequences you can see for yourself above:
- Similar characters are cheap. Because only distances are stored, a run of the same character costs almost nothing: 30 copies of
中 encode to xn--fiqaaaaaaaaaaaaaaaaaaaaaaaaaaaaa, only 36 bytes. A measured, unrelated-but-nearby mix of Chinese characters packs around 41 characters into one legal 63-byte label, and a single repeated character reaches 57.
- Scattered characters are expensive. A character far from the previous one needs a longer delta number. Chinese text with many different characters typically costs roughly 2–3 extra letters per character, which is why a CJK label runs out of room somewhere around 20–40 characters even though the label limit is 63 bytes.
Decoding is the same recipe backwards: copy the ASCII prefix, read the digits, turn the deltas back into code points, and insert each new character into the string at the position the delta encodes. Only lowercase digits exist in the bootstring, which is why an uppercase letter inside an encoded part is rejected instead of guessed at — DNS is case-insensitive, so allowing both cases would make two different spellings ambiguous.
Where you will actually meet it
- In the address bar. Browsers may display
xn--fiq228c.com instead of 中文.com. It is the same site. Some browsers deliberately show the Punycode form when a name mixes scripts, precisely so that a look-alike character (Cyrillic а versus Latin a) cannot hide in the URL.
- In DNS itself. Zone files,
dig / nslookup / DNS-over-HTTPS answers and resolver logs only ever contain the xn-- form; plain Unicode is rejected or silently mangled.
- In TLS certificates. The subject alternative name for an international domain is stored as an A-label (
DNS:xn--fiq228c.com), so a certificate that does not look like the name you typed is normal, not a mismatch.
- In configuration and tooling. Web server
server_name / virtual host entries, cookie domains, hosting and CDN panels, mail domains, API keys bound to a domain, and database keys are all safer in A-label form.
- In testing and data pipelines. Some libraries pass Unicode straight to the resolver and fail; converting to Punycode first is the standard fix. Because the mapping is canonical — one Unicode label maps to exactly one A-label — the
xn-- form is also a stable identifier for logs and search.
What Punycode is not
Punycode is a reversible spelling, nothing more. It is not encryption and it is not a security feature: anyone can decode it with the tool above, so an xn-- label says “this was encoded” and nothing about who owns it. It is also not compression — it usually makes text longer, as the 6-byte to 11-byte example above shows. And it is not the same thing as the other ways Unicode travels around the internet; those solve different layers of the problem.
| Scheme |
Turns what into what |
Where it is used |
中文 becomes |
| Punycode (RFC 3492) | A Unicode domain label into ASCII letters and digits | International domain labels only | xn--fiq228c |
| UTF-8 (RFC 3629) | Code points into bytes | Every file and web request | E4 B8 AD E6 96 87 (6 bytes) |
| Percent-encoding (RFC 3986) | Bytes into %XX escapes | URL paths, query strings, form bodies | %E4%B8%AD%E6%96%87 |
| Base64 (RFC 4648) | Bytes into 64 ASCII characters | Binary inside text, JSON, e-mail, data URLs | 5Lit5paH |
| HTML entities | Code points into named or numeric references | HTML and XML source | 中文 |
| Unicode escapes | Code points into \uXXXX | Source code and JSON string literals | \u4e2d\u6587 |
The practical takeaway: Punycode applies to the domain labels of a URL, not to the whole URL. The path and query string are a different problem and use percent-encoding (see URL Encoders).
Limits and rules that actually bite
- Size limits come from DNS, not from Punycode. One label may be at most 63 bytes and a complete name at most 255 bytes (RFC 1034 / 1035). Since an A-label is pure ASCII, the byte count equals the character count, and every non-ASCII character spends part of that budget — roughly 1 character of budget per 1–4 Punycode letters.
- Punycode converts anything; registries allow only some things. IDNA rules sit on top and restrict which code points may be registered at all — uppercase letters, spaces, punctuation such as
!, and flags are not legal in a real domain. This page performs the RFC 3492 conversion, so it will happily encode text a registry would reject.
- Stay in one script. IDNA discourages mixing scripts inside a single label precisely because of look-alike attacks. A name can be encoded and still be unusable, or usable but blocklisted by browsers.
- Do not double-encode.
xn-- followed by already-encoded text is invalid, and a label that is pure ASCII never changes at all — there is nothing to encode, so no prefix is added.
- A trailing dot is fine.
例え.jp. is a fully qualified name and the empty label at the end is kept; empty labels in the middle are not valid host names.
Quick answers
“Why does my Chinese domain show up as xn--fiq228c.com in the address bar?”
Because the browser chose to render the A-label. The site is identical; use this page's Decode direction to see the Unicode form, or encode the Unicode form to check that both spellings match.
“Do I need to change my DNS records, certificate or hosting?”
No. Keep the xn-- A-label everywhere a machine reads the name (DNS, certificates, configs) and use the Unicode form only where a human reads it.
“Why is the result longer than what I typed?”
Because the encoding does not carry the characters themselves, it carries base-36 distances, and the 36-character alphabet has to spell out numbers that can be large. Growth is expected, not a bug.
“Why does decoding say ‘uppercase letter’?”
Because bootstring digits are defined as lowercase only (a-z, 0-9). An uppercase A-Z after the last hyphen is ambiguous once DNS case-insensitivity is taken into account, so it is reported instead of being silently guessed.
“Does Punycode apply to the part before the @ in an e-mail address?”
Only the domain part is a DNS name and can be Punycode. The local part is handled by the mail system, not by DNS, so this encoding does not apply to it.
“Can two different Unicode names produce the same A-label?”
No. The encoding is injective: distinct code point sequences always produce distinct bootstrings, so decoding xn--fiq228c gives back exactly 中文 and nothing else. What can collide is appearance — different characters that look alike — which is a property of the Unicode repertoire, not of the encoding.
Standards this page follows
- RFC 3492 — Punycode, the bootstring encoding implemented on this page. The known vectors quoted above are checked against the sample strings in that document.
- RFC 5890, RFC 5891, RFC 5892 and Unicode UTS #46 — IDNA2008 and its browser-side interpretation, the rules that decide which characters are legal, how labels are normalized, and how scripts may be mixed.
- RFC 1034 and RFC 1035 — the 63-byte label and 255-byte name limits inherited from DNS.
- RFC 3490 and RFC 3491 — the older IDNA2003 generation, superseded by IDNA2008. If a name behaves differently in old software, this is usually why.
Scope note: this tool performs RFC 3492 conversion and nothing else. It is not an IDN validator and does not check registration rules, pricing or availability of a domain — and since it never contacts the network, it cannot tell you whether a name exists.
Try it above: paste 中文 and press Encode Text → Punycode to see the A-label, or paste xn--fiq228c and press Decode Punycode → Text to get the characters back.