Punycode / IDN Converter

Convert Unicode text into ASCII Punycode (RFC 3492, the encoding behind xn-- internationalized domain names) or decode an xn-- label / domain back into its original characters. You can enter a single label (e.g. 中文) or a whole dotted domain (e.g. 例え.jp); each non-ASCII label is encoded independently. Also: URL Encoders | Ascii85 | Base32 / Base58.
All processing happens locally in your browser — nothing is uploaded.
Never heard of Punycode? What is Punycode — explained in plain English (with examples, myths and the rules that actually bite).
How it works: Labels that contain any non-ASCII character (CJK, accented Latin, emoji, ...) are converted to ASCII Punycode and prefixed with xn--. Pure-ASCII labels (and empty parts from a trailing dot) pass through unchanged. A label of only basic ASCII characters never changes.
Examples: 中文 → xn--fiq228c  ·  例え.jp → xn--r8jz45g.jp  ·  bücher.de → xn--bcher-kva.de  ·  😂.com → xn--e28h.com
How it works: Every label starting with xn-- (case-insensitive prefix) is decoded back to Unicode; other ASCII labels pass through unchanged. An uppercase A-Z letter in the encoded part (after the last hyphen) is ambiguous and rejected — paste lowercase A-labels.
Ctrl+Enter in the input runs the active mode.
Pick a direction and run it. All processing stays in your browser.
RFC 3492 Punycode at a glance:
Punycode is a bootstring encoding: basic ASCII code points (letters, digits, hyphens) are copied first, then every non-ASCII code point is inserted one at a time as a small integer sequence using a bias-adaptive variable-length base-36 encoding (digits a-z then 0-9). In domain names each encoded label is prefixed with xn-- so browsers and DNS can tell it apart from a normal label. The code point - acts as the delimiter between the copied ASCII part and the encoded part.

Known vectors — 中文 → xn--fiq228c, 例え → xn--r8jz45g, 中国 → xn--fiqs8s, bücher → xn--bcher-kva.
What is Punycode?
In one sentence: Punycode is a way of spelling characters that are not plain English letters — Chinese, accented Latin, Cyrillic, emoji — using only the 37 characters the domain name system was built for: a-z, 0-9 and -. It is defined in RFC 3492, and in a domain name you recognize it by the xn-- prefix: 中文 becomes xn--fiq228c.

So Punycode is not a language, a font or a country code — it is a translation between two spellings of the same name. The Unicode spelling is called a U-label (中文.com, bücher.de), the ASCII spelling is called an A-label (xn--fiq228c.com, xn--bcher-kva.de), and the whole idea of putting international names into DNS is called IDN (internationalized domain names) or IDNA, the set of rules built on top of it.

It exists because DNS was designed in the early 1980s around a deliberately tiny alphabet. A host name that travels through DNS, into a zone file, a resolver cache, a log line or a TLS certificate, must be plain ASCII — the wire format simply has no field for anything else. Punycode is the escape hatch: it lets a name that is not ASCII travel as if it were, and lets the receiving end reconstruct the original characters exactly. That is why a Chinese domain works in a browser at all, and why xn-- is the one string you will see in every place where the Unicode spelling cannot be sent.

Punycode in 30 seconds
What you type (U-label) What DNS gets (A-label) What it means
中文.comxn--fiq228c.comThe Chinese words for “Chinese language”
中国xn--fiqs8s“China” in Chinese
例え.jpxn--r8jz45g.jpJapanese examples
bücher.dexn--bcher-kva.deGerman bücher, “books”
münchen.dexn--mnchen-3ya.deMunich
😂.comxn--e28h.comAn emoji label — legal in the encoding, not in most registries

A length comparison worth remembering: 中文 is 2 characters and 6 bytes in UTF-8, but xn--fiq228c is 11 ASCII bytes. Punycode usually makes a name longer, never shorter.

How the encoding actually works

Take bücher and watch the three moves. The engine is called a bootstring, and the whole algorithm is only a few dozen lines of code (you can read it in the script of this page).

  1. Copy the plain characters to the front. Every basic ASCII character is copied in its original order and kept readable: bücher → bcher. Because something was copied and something was left over, a single - is appended as the boundary: bcher-.
  2. Write the rest as base-36 numbers. The remaining characters are handled in increasing code-point order, and each one is written as a variable-length number in base 36 where a=0 … z=25 and 0=26 … 9=35. For bücher the leftover character ü becomes the digits kva.
  3. Stick the prefix on. Encoded part appended to the copied part gives bcher-kva, and in domain context it is written as xn--bcher-kva.

The trick that makes this clever is what those base-36 numbers contain: not the character itself, but a delta — how far the character is from the previously encoded code point, plus where it belongs inside the string. The encoder walks code points in ascending order and each digit run describes a position and a distance. After every character a shared number called the bias is adjusted so the following number needs as few digits as possible.

Two consequences you can see for yourself above:

  • Similar characters are cheap. Because only distances are stored, a run of the same character costs almost nothing: 30 copies of 中 encode to xn--fiqaaaaaaaaaaaaaaaaaaaaaaaaaaaaa, only 36 bytes. A measured, unrelated-but-nearby mix of Chinese characters packs around 41 characters into one legal 63-byte label, and a single repeated character reaches 57.
  • Scattered characters are expensive. A character far from the previous one needs a longer delta number. Chinese text with many different characters typically costs roughly 2–3 extra letters per character, which is why a CJK label runs out of room somewhere around 20–40 characters even though the label limit is 63 bytes.

Decoding is the same recipe backwards: copy the ASCII prefix, read the digits, turn the deltas back into code points, and insert each new character into the string at the position the delta encodes. Only lowercase digits exist in the bootstring, which is why an uppercase letter inside an encoded part is rejected instead of guessed at — DNS is case-insensitive, so allowing both cases would make two different spellings ambiguous.

Where you will actually meet it
  • In the address bar. Browsers may display xn--fiq228c.com instead of 中文.com. It is the same site. Some browsers deliberately show the Punycode form when a name mixes scripts, precisely so that a look-alike character (Cyrillic а versus Latin a) cannot hide in the URL.
  • In DNS itself. Zone files, dig / nslookup / DNS-over-HTTPS answers and resolver logs only ever contain the xn-- form; plain Unicode is rejected or silently mangled.
  • In TLS certificates. The subject alternative name for an international domain is stored as an A-label (DNS:xn--fiq228c.com), so a certificate that does not look like the name you typed is normal, not a mismatch.
  • In configuration and tooling. Web server server_name / virtual host entries, cookie domains, hosting and CDN panels, mail domains, API keys bound to a domain, and database keys are all safer in A-label form.
  • In testing and data pipelines. Some libraries pass Unicode straight to the resolver and fail; converting to Punycode first is the standard fix. Because the mapping is canonical — one Unicode label maps to exactly one A-label — the xn-- form is also a stable identifier for logs and search.
What Punycode is not

Punycode is a reversible spelling, nothing more. It is not encryption and it is not a security feature: anyone can decode it with the tool above, so an xn-- label says “this was encoded” and nothing about who owns it. It is also not compression — it usually makes text longer, as the 6-byte to 11-byte example above shows. And it is not the same thing as the other ways Unicode travels around the internet; those solve different layers of the problem.

Scheme Turns what into what Where it is used 中文 becomes
Punycode (RFC 3492)A Unicode domain label into ASCII letters and digitsInternational domain labels onlyxn--fiq228c
UTF-8 (RFC 3629)Code points into bytesEvery file and web requestE4 B8 AD E6 96 87 (6 bytes)
Percent-encoding (RFC 3986)Bytes into %XX escapesURL paths, query strings, form bodies%E4%B8%AD%E6%96%87
Base64 (RFC 4648)Bytes into 64 ASCII charactersBinary inside text, JSON, e-mail, data URLs5Lit5paH
HTML entitiesCode points into named or numeric referencesHTML and XML source中文
Unicode escapesCode points into \uXXXXSource code and JSON string literals\u4e2d\u6587

The practical takeaway: Punycode applies to the domain labels of a URL, not to the whole URL. The path and query string are a different problem and use percent-encoding (see URL Encoders).

Limits and rules that actually bite
  • Size limits come from DNS, not from Punycode. One label may be at most 63 bytes and a complete name at most 255 bytes (RFC 1034 / 1035). Since an A-label is pure ASCII, the byte count equals the character count, and every non-ASCII character spends part of that budget — roughly 1 character of budget per 1–4 Punycode letters.
  • Punycode converts anything; registries allow only some things. IDNA rules sit on top and restrict which code points may be registered at all — uppercase letters, spaces, punctuation such as !, and flags are not legal in a real domain. This page performs the RFC 3492 conversion, so it will happily encode text a registry would reject.
  • Stay in one script. IDNA discourages mixing scripts inside a single label precisely because of look-alike attacks. A name can be encoded and still be unusable, or usable but blocklisted by browsers.
  • Do not double-encode. xn-- followed by already-encoded text is invalid, and a label that is pure ASCII never changes at all — there is nothing to encode, so no prefix is added.
  • A trailing dot is fine. 例え.jp. is a fully qualified name and the empty label at the end is kept; empty labels in the middle are not valid host names.
Quick answers

“Why does my Chinese domain show up as xn--fiq228c.com in the address bar?”

Because the browser chose to render the A-label. The site is identical; use this page's Decode direction to see the Unicode form, or encode the Unicode form to check that both spellings match.

“Do I need to change my DNS records, certificate or hosting?”

No. Keep the xn-- A-label everywhere a machine reads the name (DNS, certificates, configs) and use the Unicode form only where a human reads it.

“Why is the result longer than what I typed?”

Because the encoding does not carry the characters themselves, it carries base-36 distances, and the 36-character alphabet has to spell out numbers that can be large. Growth is expected, not a bug.

“Why does decoding say ‘uppercase letter’?”

Because bootstring digits are defined as lowercase only (a-z, 0-9). An uppercase A-Z after the last hyphen is ambiguous once DNS case-insensitivity is taken into account, so it is reported instead of being silently guessed.

“Does Punycode apply to the part before the @ in an e-mail address?”

Only the domain part is a DNS name and can be Punycode. The local part is handled by the mail system, not by DNS, so this encoding does not apply to it.

“Can two different Unicode names produce the same A-label?”

No. The encoding is injective: distinct code point sequences always produce distinct bootstrings, so decoding xn--fiq228c gives back exactly 中文 and nothing else. What can collide is appearance — different characters that look alike — which is a property of the Unicode repertoire, not of the encoding.

Standards this page follows
  • RFC 3492 — Punycode, the bootstring encoding implemented on this page. The known vectors quoted above are checked against the sample strings in that document.
  • RFC 5890, RFC 5891, RFC 5892 and Unicode UTS #46 — IDNA2008 and its browser-side interpretation, the rules that decide which characters are legal, how labels are normalized, and how scripts may be mixed.
  • RFC 1034 and RFC 1035 — the 63-byte label and 255-byte name limits inherited from DNS.
  • RFC 3490 and RFC 3491 — the older IDNA2003 generation, superseded by IDNA2008. If a name behaves differently in old software, this is usually why.

Scope note: this tool performs RFC 3492 conversion and nothing else. It is not an IDN validator and does not check registration rules, pricing or availability of a domain — and since it never contacts the network, it cannot tell you whether a name exists.

Try it above: paste 中文 and press Encode Text → Punycode to see the A-label, or paste xn--fiq228c and press Decode Punycode → Text to get the characters back.