Base62 Encode / Decode

Convert a number or a byte string with Base62: 62 characters made of digits, upper case letters and lower case letters, none of which needs escaping in a URL query or in HTML text. That is why short links and auto-increment IDs are written this way. Two modes: a decimal number (an ID becomes a short code) and bytes (text or hexadecimal input). Also: Base64 | Base32 / Base58 | Base91.
All processing happens locally in your browser — your data is never uploaded to any server.

Never heard of Base62? What is Base62 — explained in plain English (with worked examples, the three alphabets and the traps).
Alphabet:
Pad the result to a fixed width: 0 = no padding. Padding repeats the zero character of the selected alphabet, so the same number is padded differently in the letters-first alphabet.
How it works:
Input is:
How it works:
Output as:
Ctrl+Enter in the input runs the active mode.
Pick Encode or Decode and click its action button. The Self-test button runs every documented sample, every length from 0 to 64 bytes and a spread of decimal numbers through all three alphabets, and reports the count here.
The three alphabets, printed in full: Base62 has no standards body behind it, so the alphabet is not guaranteed across implementations. These are the three orderings that actually appear in the wild; each one is printed character by character with its value, so any result can be traced back to the exact alphabet that produced it. The shaded cell is the character that means zero, which is not the same character in all three.
Base62 is case sensitive in every one of the three orderings: an upper case letter and its lower case form almost always name different values, so a code that has been through a lower-casing step is a different number, not an error.
Length of the same bytes in each encoding:
How large an input this page accepts, and why: every mode here converts through one big integer, and big-integer conversion costs time proportional to the square of the length. That is a property of the format itself, not of this page: any tool that implements real Base62 has the same shape of limit. The numbers below come from a timing harness that runs the page's own routines rather than from a running tab, so they are deliberately on the pessimistic side: a desktop browser usually finishes several times faster, and that margin is what covers a slow phone. The limits are set so that even the slow case finishes instead of freezing the tab.
Above the soft limit the first click explains the cost and asks for a second click before it runs; above the hard limit the page refuses rather than freezing the tab for several seconds with no way to show progress. Byte mode is counted in bytes when it encodes and in characters when it decodes, and the character count is always the larger one: 2,048 bytes print as about 2,751 characters, which is past the character limit, so a payload of about 1,520 bytes or less is what survives a round trip on this page. For plain binary data of any real size, Base64 and Base91 are linear and have no such limit.
What is Base62?

Base62 is a way of writing a number with nothing but the characters that survive being pasted anywhere: the ten digits, the 26 upper case letters and the 26 lower case letters. No punctuation, no spaces, no padding character, nothing that has to be escaped before it goes into a URL, a file name, a cookie or a database key. Ten plus 26 plus 26 is 62, and that is where the name comes from.

The 30 second version: the number 12345 is written as 3D7, and 1234567890 is written as 1LY7VK. Going back the other way is the same operation in reverse, and it needs the exact alphabet that produced the code — a different alphabet turns the same six characters back into a different number, without any error message.

This page does two related jobs, and the difference between them matters more than any other thing on this page:

  • Integer mode treats the input as a number. A decimal ID like 1234567890 becomes 1LY7VK, and the leading zeros of the input are not part of the value, so they do not come back.
  • Byte mode treats the input as a byte string. The bytes are read as one huge big-endian integer and then divided into 62, which is the same arithmetic on a different quantity. Here the number of leading zero bytes does survive, one zero byte per zero character.
Where the 62 comes from, and what it costs

62 is not a power of two, and that single fact shapes everything else. The familiar encodings are built by slicing the input into fixed bit groups: 4 bits print as one hexadecimal character, 6 bits as one Base64 character, 5 bits as one Base32 character. Base62 cannot do that, because no whole number of bits fits a character of 62 values — 5 bits only name 32 values, 6 bits name 64. So Base62 goes through arithmetic instead: read the whole input as one integer, and print it in base 62.

One character then carries log2(62) = 5.954 bits, which is very close to Base64's 6, and in decimal mode one character replaces log62(10) = 1.792 decimal digits. A few consequences worth knowing before you pick a length:

Code lengthValues it can nameIn practice
4 characters624 = 14,776,336Enough for a counter that stays under eight figures, and it fits in a phone number's worth of space.
6 characters626 = 56,800,235,584The usual shape of an ID code. A ten digit decimal ID fits in six characters.
7 characters627 = 3,521,614,606,208Past the range of a 32 bit counter by two orders of magnitude, in the width of a short link slug.
11 characters6211 = 52,036,560,683,837,093,888More than 264 = 18,446,744,073,709,551,616, so a full 64 bit number fits in eleven characters.

Working the other way, 1,000 decimal digits need about 558 characters, and 4,096 decimal digits need 2,285. The saving is real but bounded: Base62 shortens a decimal number by roughly 44%, and it never beats the densest printable encodings, because 5.954 bits per character is simply less than 6.4 or 6.5.

Integer mode: what happens to leading zeros

An integer has one canonical written form and it starts with the alphabet's zero character, which is 0 in the two digit-first orderings and A in the letters-first ordering. Everything to the left of the first significant character is padding, and padding is not part of a value: 00012345 and 12345 both become 3D7. If a fixed width is wanted — for instance a seven character code that has to line up in a column — the page's Pad the result to a fixed width box repeats zero characters until the width is reached, so 12345 at width 7 prints 00003D7 in the digits-first alphabet and AAAADNH in the letters-first one. Padding is cosmetic: those leading characters add nothing to the value, and decoding either string gives 12345 back, because 0 × 626 and 0 × 625 are both zero no matter which character spells the zero.

Decoding a padded string therefore also works, and the page reports how many of the leading characters were zero characters so that the difference between “the code is seven characters” and “the value needed three characters” is never a mystery.

Byte mode: the same arithmetic on a byte string

Byte mode reads the input bytes as one big-endian integer — the whole byte string is treated as a single number, however long it is — and prints that number in base 62. Because a byte contains 8 bits and a character names 5.954 of them, byte mode is not compression: it needs about 134% of the input size, marginally worse than Base64's 133%. On the 92 byte sample used in the sizes table above, 92 bytes print as 124 Base62 characters, and Base64 also prints 124 (123 of them before its = padding). Base62 loses to Base64 on density and wins on character safety, which is the trade every short link and invite code makes on purpose.

Leading zero bytes are preserved, and this is easy to get wrong in both directions. Two leading zero bytes followed by 0x01 print as 001 in the digits-first alphabet and as AAB in the letters-first one, because the leading characters spell zero and are one per zero byte. A byte string of a single 0x00 prints as the single character 0, and a byte string of a single 0x01 prints as 1. This is the one place where byte mode and integer mode look different while running the same algorithm.

The cost of all this is that byte mode inherits the arithmetic's quadratic timing: 1,024 bytes are estimated at about a third of a second and 2,048 bytes at about 1.3 seconds, the same figures the limits table above gives. That is why byte mode here is capped at 2,048 bytes going out and at 2,048 characters coming back — 2,048 bytes print as about 2,751 characters, so about 1,520 bytes is the largest payload that survives a round trip through this page — and why the byte mode on this page is a reference for small values, not a substitute for Base64 or Base91 on real files.

The three alphabets, and why the alphabet is the whole story

The characters of an alphabet are a fixed sequence, and a character's value is its position in that sequence. Moving the sequence around changes what every character means. Three orderings are in wide use, and this is the same number written three ways in each of them:

Decimal0-9 then A-Z then a-z0-9 then a-z then A-ZA-Z then a-z then 0-9
000A
61zZ9
621010BA
12,3453D73d7DNH
1,234,567,8901LY7VK1ly7vkBViHfU

Read the last three columns as one alphabet and it is a single coherent scheme. Read them as interchangeable and it falls apart, because decoding is alphabet-blind in the worst possible way: any character of the input is in all three alphabets, so a wrong alphabet produces no error at all, only a different number.

A codeRead as 0-9 A-Z a-zRead as 0-9 a-z A-ZRead as A-Z a-z 0-9
3D712,34513,957211,665
1LY7VK1,234,567,8901,624,950,79248,723,527,772

Case is a second, sharper version of the same trap. Base62 is case sensitive in every ordering, so lower-casing a code silently rewrites it: 1LY7VK is 1,234,567,890 in the digits-first alphabet, while 1ly7vk is 1,624,950,792 in that same alphabet — and it happens to be 1,234,567,890 in the digits-then-lower-case alphabet, because that ordering is precisely the digits-first one with the two letter cases swapped. The zero character is not shared either: the two digit-first orderings write zero as 0, while the letters-first ordering writes it as A, so that alphabet does not begin with a digit at all.

Base62 next to the other encodings
EncodingBits per characterSize against the inputNotes
Base16 (hex)4200%Two characters per byte, trivially readable, no alphabet to look up.
Base325160%Case-insensitive and safe for human dictation, the most wasteful of the group.
Base36 (0-9 a-z)~5.17~155%Digits and lower case letters only, so it survives a lower-casing pipeline. 12345 is 9ix.
Base58~5.86~137%Chosen for addresses, not for density: four look-alike characters are dropped instead of being ordered. See Base32 / Base58.
Base62~5.95~134%Every character is a digit or a letter, so nothing is ever escaped. Density sits between Base58 and Base64.
Base646133%The default everywhere, at the cost of +, / and =, the three characters that force escaping. See Base64 converter.
Ascii85 / Base856.4125%5 characters per 4 bytes, but it divides by 85 and its alphabet needs quoting. See Ascii85.
Base91~6.5~123%The densest of these, and the one whose output is least safe to paste unescaped. See Base91.

The percentages are asymptotic costs, rounded; the measured character counts for one real 92 byte input are in the sizes table further up this page. In integer mode the comparison runs the other way round, because the input is decimal text rather than bytes: there Base62 is about 44% shorter than the digits it replaces, and Base36 about 36% shorter, which is why an ID code is normally Base62 or Base36 and not Base58.

What Base62 is not
  • Not a standard. There is no RFC and no specification. No standards body has ever defined Base62, which is why the alphabet is not guaranteed and why this page prints all three orderings instead of picking one silently.
  • Not encryption. It is a change of base, nothing more. Anyone who knows the alphabet reverses it in one pass, and there is no key, no salt and no secrecy of any kind.
  • Not compression. Byte mode grows the data by about a third. Integer mode only looks like compression because decimal text is an inefficient way to write a number in the first place.
  • Not case-insensitive. Base32 can be lower-cased and read back; Base62 cannot. A code that travels through a system that normalises case is a corrupted code.
  • Not self-checking. There is no checksum and no check character, so a single wrong character gives a valid-looking result with a different value. If that matters, the Bitcoin ordering comes with a checksum — see the Base58Check section of Base32 / Base58.
  • Not fixed length. The number of characters grows with the value, so a counter that starts at 1 and reaches 627 goes from one character to eight. Pad to a fixed width when a stable column is needed.
Traps that actually bite
  • The alphabet is the answer to nine out of ten mysteries. If another tool gives a different code for the same number, compare the two alphabets character by character before suspecting a bug. The table above shows the same value in three alphabets, and all three are correct.
  • A wrong alphabet fails silently. There is no error path for it: 3D7 decodes to 12,345, 13,957 or 211,665 depending only on which ordering is assumed. Check the first character, which is the most likely to be a zero pad, and check whether the alphabet has digits or letters at the front.
  • Leading zeros mean two different things. In integer mode 00012345 and 12345 are the same number and print identically. In byte mode a leading zero byte is data, and it prints as one extra zero character. Dropping one of those characters changes the byte string.
  • Case normalisation destroys a code. An email address, a hostname or a Windows path may be lower-cased in transit. A Base62 code cannot survive that, which is the reason Base36 still exists.
  • Long inputs are slow, not just large. The conversion is arithmetic on the entire input at once, so doubling the input length quadruples the time. This page refuses inputs past its measured limits instead of hanging; the sizes table further up gives the measured numbers.
  • A Base62 string is not proof of anything. Since almost any string of digits and letters decodes without complaint, an arbitrary URL slug will decode to some number. A code only means something if the system that issued it also stores the mapping.
Quick answers

“Why does another Base62 tool give a different code for the same number?”

Because the alphabet differs. Base62 has no specification to disagree with, so implementations pick an ordering: digits and upper case first, digits and lower case first, or letters before digits. Switch the alphabet select at the top of this page to see all three answers side by side.

“Is Base62 safe in a URL?”

Yes, in every one of the three orderings. All 62 characters are unescaped letters and digits, so none of them is expanded into a percent triplet in a query string and none of them ends a URL, a path segment or an HTML attribute. That is the entire reason short links are Base62 and not Base64.

“Which is smaller, Base62 or Base64?”

For bytes, Base64 is very slightly smaller: about 133% of the input against about 134%. On the 92 byte sample above the two tie at 124 characters. For decimal numbers, Base62 shortens the text by about 44%, which is where the practical gain is.

“Can I decode a Base64 string with this page?”

No, and it will not tell you that it failed. The two alphabets overlap in all 62 characters, so a Base64 string is usually “valid” input here too and simply decodes to a different number. Use Base64 for Base64.

“Does a Base62 code have a maximum length?”

No, in theory the algorithm has no bound. In practice this page stops at 4,096 decimal digits, or 2,048 bytes, because the arithmetic behind it is quadratic and the tab would freeze with nothing on screen to explain why. The limits table further up this page lists the measured numbers.

“How do I know the result on this page is right?”

Press Self-test. It runs every documented sample, every byte length from 0 to 64 and a spread of decimal numbers through all three alphabets, comparing the returned values with what went in, and it re-checks that each alphabet's 62 characters are all safe in a URL. The Sample button additionally compares its three results against the values recorded in this page's data file, which were verified against a second, independent big-integer implementation kept in this site's test tooling.

Standards and credit

There is no standard to implement here: Base62 is a convention, not a specification, and the three alphabets printed above are the ones that appear in real software. This page implements it with the site's own big-integer routines, written in plain arithmetic on digit arrays so that no size is ever limited by the JavaScript number type.

The page text was written for this site. Every worked example on this page was produced by the page's own code and then verified against a second, independent implementation using native big integers, which is kept in this site's test tooling; the same implementation verified the sample and size tables in the data file.