acalculator

How do I convert binary to text?

Paste binary code to read it as text. Each group of 8 bits is one byte.

Your numbers

Text
Hello

The binary says "Hello".

Bytes
5

Text: Hello. The binary says "Hello".

How to calculate

Reads binary code (8 bits per byte) as UTF-8 text, which includes plain ASCII.

Example with the default inputs (Binary 01001000 01100101 01101100 01101100 01101111): The binary says "Hello".

Method: Split the binary into bytes of 8 bits, turn each byte into a number from 0 to 255, and read the bytes as UTF-8 text.

  • Each byte is 8 bits, most significant bit first. A group of fewer than 8 bits between spaces is padded with 0s on the left.
  • Spaces, commas, and any character other than 0 and 1 separate bytes.
  • A leading byte order mark (EF BB BF) is kept as the character U+FEFF, as RFC 3629 section 6 allows.
  • Bytes are read as UTF-8, which matches ASCII for the codes 0 to 127. A byte that is not valid UTF-8 shows as �.

Machine-readable copies: Markdown, JSON.

Worked examples

Each example is checked against the calculator on every build.

  1. Binary 01001000 01101001 gives Text Hi, Bytes 2.Source: hand calculation in content.mdx: 01001000 = 72 = H, 01101001 = 105 = i (ASCII); Python 3 bytes([72, 105]).decode()
  2. Binary 0100100001101001 gives Text Hi, Bytes 2.Source: hand calculation in content.mdx: the same 16 bits without a space split into two bytes
  3. Binary 11000011 10101001 gives Text é, Bytes 2.Source: RFC 3629 (UTF-8): C3 A9 encodes U+00E9; Python 3 bytes([0xC3, 0xA9]).decode("utf-8") = "é"

How it works

  1. Every character other than 0 and 1 (spaces, commas, new lines) separates bytes.
  2. A group of 8 bits or fewer is one byte; a shorter group is padded with 0s on the left. A longer group is split into bytes of 8 bits from the left, and it must be a multiple of 8 bits long.
  3. Each byte is read most significant bit first: the bits are worth 128, 64, 32, 16, 8, 4, 2, and 1.
  4. The bytes are read as UTF-8 text (RFC 3629). Bytes 0 to 127 are ASCII characters. A byte sequence that is not valid UTF-8 shows as the replacement character �.

Assumptions

  • Bytes are 8 bits, most significant bit first.
  • The text encoding is UTF-8, which includes ASCII.
  • A leading byte order mark (11101111 10111011 10111111) is kept as the invisible character U+FEFF, not removed (RFC 3629, section 6).

Worked examples by hand

01001000 01101001. 01001000 = 64 + 8 = 72, which is H. 01101001 = 64 + 32 + 8 + 1 = 105, which is i. The text is Hi, 2 bytes.

0100100001101001. The same 16 bits with no space split into 01001000 and 01101001, so the text is again Hi.

11000011 10101001. These are the bytes C3 and A9 in hexadecimal. In UTF-8, 110xxxxx 10yyyyyy is a two-byte character with the bits 00011 101001 = 0xE9, the code point U+00E9, which is é.

Other questions people ask

How do I convert binary to text?

Split the binary into groups of 8 bits. Turn each group into a number from 0 to 255 by adding the place values (128, 64, 32, 16, 8, 4, 2, 1) of its 1s. Then look up each number in the ASCII table, for example 72 is H and 105 is i.

What does 01001000 01101001 mean?

01001000 is 64 + 8 = 72, which is H in ASCII, and 01101001 is 64 + 32 + 8 + 1 = 105, which is i. Together they say "Hi".

Do I need spaces between the bytes?

No. Without spaces the converter splits the bits into groups of 8 from the left. With spaces, each group is one byte, and a short group like 1000001 is padded with a 0 on the left.

What is the difference between ASCII and UTF-8?

ASCII gives the codes 0 to 127 to English letters, digits, and symbols. UTF-8 uses the same one-byte codes for those characters and adds sequences of 2 to 4 bytes for every other character, such as é (11000011 10101001).

Why do I see a � in the answer?

The bytes are not valid UTF-8 at that point, for example a byte above 127 on its own. Check that no bit is missing or extra.