Base64 and Base58 Encoding, Bitcoin Addresses

Lecture 9 min.



Base64 is a standard for encoding binary data using only 64 ASCII characters. The encoding alphabet contains the Latin alphanumeric characters A-Z, a-z and 0-9 (62 characters) and 2 additional characters that depend on the implementation. Every 3 source bytes are encoded by 4 characters (an increase of ¹⁄₃).

This system is widely used in e-mail to represent binary files within the body of a message (transport encoding).

MIME

In the MIME e-mail format, base64 is a scheme by which an arbitrary sequence of bytes is converted into a sequence of printable ASCII characters. Only Latin letters in upper and lower case (A—Z, a—z), digits (0—9), and the characters "+" and "/" are used, with the "=" character serving as a special suffix code.

The full specification of this form of base64 is contained in RFC 1421 and RFC 2045. The scheme is used to encode a sequence of octets (bytes).

To convert data to base64, the first byte is placed in the most significant eight bits of a 24-bit buffer, the next in the middle eight, and the third in the least significant eight bits. If fewer than three bytes are being encoded, the corresponding bits of the buffer are set to zero. Then each six bits of the buffer, starting from the most significant, are used as indices into the string "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/", and the characters pointed to by the indices are placed in the output string. If only one or two bytes are encoded, the result is only the first two or three characters of the string, and the output string is padded with two or one "=" characters. This prevents extra bits from being added to the decoded data. The process is repeated on the remaining input data.

For example, a quotation from Thomas Hobbes's "Leviathan":

Man is distinguished, not only by his reason, but by this singular passion from other animals, which is a lust of the mind, that by a perseverance of delight in the continued and indefatigable generation of knowledge, exceeds the short vehemence of any carnal pleasure.

when re-encoded from ASCII to base64, looks as follows:

 Base64 and Base58 Encoding, Bitcoin Addresses

In the example, the word Man is encoded as TWFu. The conversion process can be represented in the following table:

Source text M a n
ASCII codes 77 (0x4d) 97 (0x61) 110 (0x6e)
Binary form 0 1 0 0 1 1 0 1 0 1 1 0 0 0 0 1 0 1 1 0 1 1 1 0
Resulting Base64 index 19 22 5 46
Final Base64 result T W F u

Character-to-value correspondence in Base64

Character Value Character Value Character Value Character Value
10 2 8 16 10 2 8 16 10 2 8 16 10 2 8 16
A 0 000000 00 00 Q 16 010000 20 10 g 32 100000 40 20 w 48 110000 60 30
B 1 000001 01 01 R 17 010001 21 11 h 33 100001 41 21 x 49 110001 61 31
C 2 000010 02 02 S 18 010010 22 12 i 34 100010 42 22 y 50 110010 62 32
D 3 000011 03 03 T 19 010011 23 13 j 35 100011 43 23 z 51 110011 63 33
E 4 000100 04 04 U 20 010100 24 14 k 36 100100 44 24 0 52 110100 64 34
F 5 000101 05 05 V 21 010101 25 15 l 37 100101 45 25 1 53 110101 65 35
G 6 000110 06 06 W 22 010110 26 16 m 38 100110 46 26 2 54 110110 66 36
H 7 000111 07 07 X 23 010111 27 17 n 39 100111 47 27 3 55 110111 67 37
I 8 001000 10 08 Y 24 011000 30 18 o 40 101000 50 28 4 56 111000 70 38
J 9 001001 11 09 Z 25 011001 31 19 p 41 101001 51 29 5 57 111001 71 39
K 10 001010 12 0A a 26 011010 32 1A q 42 101010 52 2A 6 58 111010 72 3A
L 11 001011 13 0B b 27 011011 33 1B r 43 101011 53 2B 7 59 111011 73 3B
M 12 001100 14 0C c 28 011100 34 1C s 44 101100 54 2C 8 60 111100 74 3C
N 13 001101 15 0D d 29 011101 35 1D t 45 101101 55 2D 9 61 111101 75 3D
O 14 001110 16 0E e 30 011110 36 1E u 46 101110 56 2E + 62 111110 76 3E
P 15 001111 17 0F f 31 011111 37 1F v 47 101111 57 2F / 63 111111 77 3F

UTF-7

UTF-7 is a modified variant of Base64. This encoding scheme is used for UTF-16 files as an intermediate format in MIME. UTF-7 is intended for using Unicode in e-mail without transport encoding of the content. The main difference between this variant of base64 and MIME is that the "=" character is not used for padding, since this character would require repeated escaping. Instead, the bits of the octet are padded with zeros.

The modified Base64 is standardized in RFC 2152, A Mail-Safe Transformation Format of Unicode.

IRCu

In the server-to-server protocol used in IRC and compatible software, a version of base64 is used to encode client/server numerics and binary IP addresses. Client and server numerics have fixed sizes that exactly match the number of base64 characters, so no padding is needed. Binary IP addresses are extended with leading zero bits to fit. The character set differs slightly from MIME in using [] instead of +/.

Use in Web Applications

Thanks to Base64, binary content can be included in HTML documents, creating a single document with no separately located pictures or other additional files. In this way, an HTML document with embedded graphics, audio, video, programs, styles and other additions becomes an excellent alternative to other formats for complexly formatted documents such as doc, docx and pdf.

Some applications encode binary data for convenient inclusion in URLs and hidden form fields.

Applying a URL encoder on top of the Base64 standard is not always convenient, because it converts the characters / and + into special hexadecimal sequences. Although this conversion is reversible, it lengthens the resulting string and somewhat complicates its subsequent parsing. In addition, the % character generated by the URL encoder may need to be escaped again when the resulting string is later passed through other systems (for example, in SQL it is a pattern element).

For this reason there is a modified Base64 for URLs, which does not use the = padding character and replaces the + and / characters with * and - respectively, so that URL encoders/decoders are no longer necessary and have no effect on the length of the encoded value, leaving the same encoded form intact for use in relational databases, web forms and object identifiers in general. The standard for Base64 encoding of URLs is the variant in which the + and / characters are replaced with - and _ respectively (RFC 3548, section 4).

Another variant is called modified Base64 for regular expressions. It uses ! and - instead of * and - to replace the standard Base64 +/, because both + and * may be reserved in regular expressions (note that the [] used above in the IRCu variant may not work in this context).

There are other variants that use _ and - or . and _ when a Base64 string must be used together with program identifiers, or . and - for use in XML name tokens (Nmtoken), or _ and : in the more restricted XML identifiers (Name). In some cases Base58, which does not use the + and / characters, is applied to URLs.

Base58

For URL encoding, some systems use Base58, which differs from Base64 in that the resulting text contains no characters that a person could read ambiguously. The characters 0 (zero), O (capital Latin o), I (capital Latin i) and l (lowercase Latin L) are excluded. The + (plus) and / (slash) characters are also excluded, since they can cause an address to be misinterpreted when encoding a URL.

Base58 is a way of encoding binary data as alphanumeric text based on the Latin alphabet. The encoding alphabet contains 58 characters. It is used to transmit data across heterogeneous networks (transport encoding). The standard is similar to Base64, but differs in that its output contains not only no service characters but also no alphanumeric characters that a person might read ambiguously. The characters 0 (zero), O (capital Latin o), I (capital Latin i) and l (lowercase Latin L) are excluded. The + (plus) and / (slash) characters are also excluded, since they can cause misinterpretation when encoding a URL.

The standard was developed to reduce visual confusion for users who enter data manually from printed text or a photograph, that is, without the ability to copy and paste by machine.

Unlike Base64, the encoding does not preserve a one-to-one byte correspondence with the source data: different combinations of the same number of bytes are encoded in Base58 as strings of different lengths.

Base58 encoding is typically used to encode addressing systems. The actual order of the letters in the alphabet depends on the application of the encoding. Therefore, giving only the term "Base58" without specifying the alphabet is not enough to describe the format completely.

Application Alphabet
Bitcoin addresses 123456789ABCDEFGHJKLMNPQRSTUVWXYZabcdefghijkmnopqrstuvwxyz
Ripple addresses rpshnaf39wBUDNEGHJKLM4PQRST7VWXYZ2bcdeCg65jkm8oFqi1tuvAxyz
Flickr short URL 123456789abcdefghijkmnopqrstuvwxyzABCDEFGHJKLMNPQRSTUVWXYZ

Radix-64

Radix-64 is a variant of Base64 encoding of binary data into a text format, used in PGP. It differs from Base64 in that a 24-bit checksum is appended at the end.

Unix-family operating systems store password hashes computed with crypt in the /etc/passwd file, using the B64 encoding.

It is similar to radix-64, but the "=" padding suffix is not used and in the alphabet the non-letter characters are placed at the beginning: ./0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz.

Other uses

There are many uses of Base64. For example, Thunderbird and Mozilla Suite used Base64 to obscure passwords in POP3. Base64 can be used as a method of hiding secrets without the overhead of cryptographic key management, but this approach is completely insecure and is not recommended.

Spam scanners that do not decode Base64 messages often let them through, because they look random enough or contain no keywords in the Base64 text to be taken for spam. Spammers exploit this to bypass basic anti-spam tools.

This standard is used to encode JPEG and PNG images for embedding them in FB2 e-books .

There are applications that use the base64 technique to send small images via long SMS messages.

Base64 and Base58 Encoding, Bitcoin Addresses

Fig. base64 vs base58

See also

Serialization data formats
Text
  • ASCII85
  • Base58
  • Base64
  • UTF-8
    • UTF-16
    • UTF-32
  • UUE
  • yEnc
Internet and telecommunications
  • MIME
    • S/MIME
  • multipart/form-data
  • RSS
    • GeoRSS
  • Tag-length-value
Media
  • CD Video
  • Flash Video
  • MusicDNA
  • Video CD
    • Super Video CD
Other
  • CBEFF
  • RINEX
  • QTI

Comments

To leave a comment

If you have any suggestion, idea, thanks or comment, feel free to write. We really value feedback and are glad to hear your opinion.
To reply

Lectures and tutorial on "Information and Coding Theory"

Terms: Information and Coding Theory