Lecture 9 min.
Base64 is a standard for encoding binary data using only 64 ASCII characters. The encoding alphabet contains the Latin alphanumeric characters A-Z, a-z and 0-9 (62 characters) and 2 additional characters that depend on the implementation. Every 3 source bytes are encoded by 4 characters (an increase of ¹⁄₃).
This system is widely used in e-mail to represent binary files within the body of a message (transport encoding).
In the MIME e-mail format, base64 is a scheme by which an arbitrary sequence of bytes is converted into a sequence of printable ASCII characters. Only Latin letters in upper and lower case (A—Z, a—z), digits (0—9), and the characters "+" and "/" are used, with the "=" character serving as a special suffix code.
The full specification of this form of base64 is contained in RFC 1421 and RFC 2045. The scheme is used to encode a sequence of octets (bytes).
To convert data to base64, the first byte is placed in the most significant eight bits of a 24-bit buffer, the next in the middle eight, and the third in the least significant eight bits. If fewer than three bytes are being encoded, the corresponding bits of the buffer are set to zero. Then each six bits of the buffer, starting from the most significant, are used as indices into the string "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/", and the characters pointed to by the indices are placed in the output string. If only one or two bytes are encoded, the result is only the first two or three characters of the string, and the output string is padded with two or one "=" characters. This prevents extra bits from being added to the decoded data. The process is repeated on the remaining input data.
For example, a quotation from Thomas Hobbes's "Leviathan":
Man is distinguished, not only by his reason, but by this singular passion from other animals, which is a lust of the mind, that by a perseverance of delight in the continued and indefatigable generation of knowledge, exceeds the short vehemence of any carnal pleasure.
when re-encoded from ASCII to base64, looks as follows:
In the example, the word Man is encoded as TWFu. The conversion process can be represented in the following table:
| Source text | M | a | n | |||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ASCII codes | 77 (0x4d) | 97 (0x61) | 110 (0x6e) | |||||||||||||||||||||
| Binary form | 0 | 1 | 0 | 0 | 1 | 1 | 0 | 1 | 0 | 1 | 1 | 0 | 0 | 0 | 0 | 1 | 0 | 1 | 1 | 0 | 1 | 1 | 1 | 0 |
| Resulting Base64 index | 19 | 22 | 5 | 46 | ||||||||||||||||||||
| Final Base64 result | T | W | F | u |
| Character | Value | Character | Value | Character | Value | Character | Value | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 10 | 2 | 8 | 16 | 10 | 2 | 8 | 16 | 10 | 2 | 8 | 16 | 10 | 2 | 8 | 16 | ||||
| A | 0 | 000000 | 00 | 00 | Q | 16 | 010000 | 20 | 10 | g | 32 | 100000 | 40 | 20 | w | 48 | 110000 | 60 | 30 |
| B | 1 | 000001 | 01 | 01 | R | 17 | 010001 | 21 | 11 | h | 33 | 100001 | 41 | 21 | x | 49 | 110001 | 61 | 31 |
| C | 2 | 000010 | 02 | 02 | S | 18 | 010010 | 22 | 12 | i | 34 | 100010 | 42 | 22 | y | 50 | 110010 | 62 | 32 |
| D | 3 | 000011 | 03 | 03 | T | 19 | 010011 | 23 | 13 | j | 35 | 100011 | 43 | 23 | z | 51 | 110011 | 63 | 33 |
| E | 4 | 000100 | 04 | 04 | U | 20 | 010100 | 24 | 14 | k | 36 | 100100 | 44 | 24 | 0 | 52 | 110100 | 64 | 34 |
| F | 5 | 000101 | 05 | 05 | V | 21 | 010101 | 25 | 15 | l | 37 | 100101 | 45 | 25 | 1 | 53 | 110101 | 65 | 35 |
| G | 6 | 000110 | 06 | 06 | W | 22 | 010110 | 26 | 16 | m | 38 | 100110 | 46 | 26 | 2 | 54 | 110110 | 66 | 36 |
| H | 7 | 000111 | 07 | 07 | X | 23 | 010111 | 27 | 17 | n | 39 | 100111 | 47 | 27 | 3 | 55 | 110111 | 67 | 37 |
| I | 8 | 001000 | 10 | 08 | Y | 24 | 011000 | 30 | 18 | o | 40 | 101000 | 50 | 28 | 4 | 56 | 111000 | 70 | 38 |
| J | 9 | 001001 | 11 | 09 | Z | 25 | 011001 | 31 | 19 | p | 41 | 101001 | 51 | 29 | 5 | 57 | 111001 | 71 | 39 |
| K | 10 | 001010 | 12 | 0A | a | 26 | 011010 | 32 | 1A | q | 42 | 101010 | 52 | 2A | 6 | 58 | 111010 | 72 | 3A |
| L | 11 | 001011 | 13 | 0B | b | 27 | 011011 | 33 | 1B | r | 43 | 101011 | 53 | 2B | 7 | 59 | 111011 | 73 | 3B |
| M | 12 | 001100 | 14 | 0C | c | 28 | 011100 | 34 | 1C | s | 44 | 101100 | 54 | 2C | 8 | 60 | 111100 | 74 | 3C |
| N | 13 | 001101 | 15 | 0D | d | 29 | 011101 | 35 | 1D | t | 45 | 101101 | 55 | 2D | 9 | 61 | 111101 | 75 | 3D |
| O | 14 | 001110 | 16 | 0E | e | 30 | 011110 | 36 | 1E | u | 46 | 101110 | 56 | 2E | + | 62 | 111110 | 76 | 3E |
| P | 15 | 001111 | 17 | 0F | f | 31 | 011111 | 37 | 1F | v | 47 | 101111 | 57 | 2F | / | 63 | 111111 | 77 | 3F |
UTF-7 is a modified variant of Base64. This encoding scheme is used for UTF-16 files as an intermediate format in MIME. UTF-7 is intended for using Unicode in e-mail without transport encoding of the content. The main difference between this variant of base64 and MIME is that the "=" character is not used for padding, since this character would require repeated escaping. Instead, the bits of the octet are padded with zeros.
The modified Base64 is standardized in RFC 2152, A Mail-Safe Transformation Format of Unicode.
In the server-to-server protocol used in IRC and compatible software, a version of base64 is used to encode client/server numerics and binary IP addresses. Client and server numerics have fixed sizes that exactly match the number of base64 characters, so no padding is needed. Binary IP addresses are extended with leading zero bits to fit. The character set differs slightly from MIME in using [] instead of +/.
Thanks to Base64, binary content can be included in HTML documents, creating a single document with no separately located pictures or other additional files. In this way, an HTML document with embedded graphics, audio, video, programs, styles and other additions becomes an excellent alternative to other formats for complexly formatted documents such as doc, docx and pdf.
Some applications encode binary data for convenient inclusion in URLs and hidden form fields.
Applying a URL encoder on top of the Base64 standard is not always convenient, because it converts the characters / and + into special hexadecimal sequences. Although this conversion is reversible, it lengthens the resulting string and somewhat complicates its subsequent parsing. In addition, the % character generated by the URL encoder may need to be escaped again when the resulting string is later passed through other systems (for example, in SQL it is a pattern element).
For this reason there is a modified Base64 for URLs, which does not use the = padding character and replaces the + and / characters with * and - respectively, so that URL encoders/decoders are no longer necessary and have no effect on the length of the encoded value, leaving the same encoded form intact for use in relational databases, web forms and object identifiers in general. The standard for Base64 encoding of URLs is the variant in which the + and / characters are replaced with - and _ respectively (RFC 3548, section 4).
Another variant is called modified Base64 for regular expressions. It uses ! and - instead of * and - to replace the standard Base64 +/, because both + and * may be reserved in regular expressions (note that the [] used above in the IRCu variant may not work in this context).
There are other variants that use _ and - or . and _ when a Base64 string must be used together with program identifiers, or . and - for use in XML name tokens (Nmtoken), or _ and : in the more restricted XML identifiers (Name). In some cases Base58, which does not use the + and / characters, is applied to URLs.
For URL encoding, some systems use Base58, which differs from Base64 in that the resulting text contains no characters that a person could read ambiguously. The characters 0 (zero), O (capital Latin o), I (capital Latin i) and l (lowercase Latin L) are excluded. The + (plus) and / (slash) characters are also excluded, since they can cause an address to be misinterpreted when encoding a URL.
Base58 is a way of encoding binary data as alphanumeric text based on the Latin alphabet. The encoding alphabet contains 58 characters. It is used to transmit data across heterogeneous networks (transport encoding). The standard is similar to Base64, but differs in that its output contains not only no service characters but also no alphanumeric characters that a person might read ambiguously. The characters 0 (zero), O (capital Latin o), I (capital Latin i) and l (lowercase Latin L) are excluded. The + (plus) and / (slash) characters are also excluded, since they can cause misinterpretation when encoding a URL.
The standard was developed to reduce visual confusion for users who enter data manually from printed text or a photograph, that is, without the ability to copy and paste by machine.
Unlike Base64, the encoding does not preserve a one-to-one byte correspondence with the source data: different combinations of the same number of bytes are encoded in Base58 as strings of different lengths.
Base58 encoding is typically used to encode addressing systems. The actual order of the letters in the alphabet depends on the application of the encoding. Therefore, giving only the term "Base58" without specifying the alphabet is not enough to describe the format completely.
| Application | Alphabet |
|---|---|
| Bitcoin addresses | 123456789ABCDEFGHJKLMNPQRSTUVWXYZabcdefghijkmnopqrstuvwxyz |
| Ripple addresses | rpshnaf39wBUDNEGHJKLM4PQRST7VWXYZ2bcdeCg65jkm8oFqi1tuvAxyz |
| Flickr short URL | 123456789abcdefghijkmnopqrstuvwxyzABCDEFGHJKLMNPQRSTUVWXYZ |
Radix-64 is a variant of Base64 encoding of binary data into a text format, used in PGP. It differs from Base64 in that a 24-bit checksum is appended at the end.
Unix-family operating systems store password hashes computed with crypt in the /etc/passwd file, using the B64 encoding.
It is similar to radix-64, but the "=" padding suffix is not used and in the alphabet the non-letter characters are placed at the beginning: ./0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz.
There are many uses of Base64. For example, Thunderbird and Mozilla Suite used Base64 to obscure passwords in POP3. Base64 can be used as a method of hiding secrets without the overhead of cryptographic key management, but this approach is completely insecure and is not recommended.
Spam scanners that do not decode Base64 messages often let them through, because they look random enough or contain no keywords in the Base64 text to be taken for spam. Spammers exploit this to bypass basic anti-spam tools.
This standard is used to encode JPEG and PNG images for embedding them in FB2 e-books .
There are applications that use the base64 technique to send small images via long SMS messages.

Fig. base64 vs base58
|
Serialization data formats
|
|
|---|---|
| Text |
|
| Internet and telecommunications |
|
| Media |
|
| Other |
|
Comments