prerequisite

Binary and hexadecimal

How computers write numbers in base 2 and base 16, and how to read and flip individual bits.

Before this

Nothing beyond first-year college math. This is a starting page.

Why you need this

Every page in this cluster that gets close to the hardware writes numbers like 0x6000F004 and 0b0110, and every program that reads a button or a serial port tests one bit inside a bigger number. If you can convert between base 10, base 2, and base 16 and set, clear, and test a single bit, you can read register descriptions, assembly listings, and the bit tricks inside the uppercase echo.

The idea

Place value in three bases

A number written in base 10 (decimal) uses ten digits, and each position is worth ten times the one to its right. So 247 means 2×100+4×10+7×12 \times 100 + 4 \times 10 + 7 \times 1.

Base 2 (binary) uses two digits, 0 and 1, and each position is worth twice the one to its right: 1, 2, 4, 8, 16, and so on. This cluster writes binary with a 0b prefix. So 0b1011 means 1×8+0×4+1×2+1×1=111 \times 8 + 0 \times 4 + 1 \times 2 + 1 \times 1 = 11.

Base 16 (hexadecimal, or hex) uses sixteen digits: 0 to 9, then A to F for ten to fifteen. Each position is worth sixteen times the one to its right. This cluster writes hex with a 0x prefix and uppercase digits. So 0x2F means 2×16+15=472 \times 16 + 15 = 47.

The precise version: a digit string dn−1…d1d0d_{n-1} \dots d_1 d_0 in base bb has the value

v=∑k=0n−1dk bkv = \sum_{k=0}^{n-1} d_k \, b^k

where dkd_k is the digit in position kk, counting from 0 at the right.

Why hex is everywhere

One hex digit is exactly four binary digits, because 16=2416 = 2^4. That makes conversion a lookup, not arithmetic: replace each hex digit with its four bits.

Hex Binary Decimal
0x0 0b0000 0
0x4 0b0100 4
0x6 0b0110 6
0x9 0b1001 9
0xB 0b1011 11
0xF 0b1111 15

So 0xB4 is 0b1011 followed by 0b0100, which is 0b10110100, which is 180. A 32-bit address such as 0x6000F004 fits in eight hex characters instead of thirty-two bits.

Bits, bytes, and words

A bit is one binary digit. A byte is 8 bits and holds 28=2562^8 = 256 values, 0 to 255, or 0x00 to 0xFF. A word is the size the processor handles in one step; on every ESP32 chip that is 32 bits, which holds 232=4,294,967,2962^{32} = 4{,}294{,}967{,}296 values. Bits are numbered from 0 at the right (the least significant bit) up to 7 in a byte or 31 in a word. Bit kk is worth 2k2^k, so bit 5 is worth 32, which is 0x20.

Characters are numbers too

ASCII is the table that assigns a number to each English letter, digit, and punctuation mark. When you type a into a serial terminal, the wire carries the byte 0x61. Capital A is 0x41. The lowercase letters run from 0x61 (a) to 0x7A (z) in order, and the capitals from 0x41 to 0x5A.

Bitwise operations

These work on each bit position separately, so they are how a program touches one bit without disturbing the others.

Operation Symbol (C, Python, JavaScript) Result bit is 1 when Typical use
AND & both bits are 1 test a bit, clear bits
OR | either bit is 1 set bits
XOR ^ the bits differ flip bits
NOT ~ the bit was 0 build a "clear" mask
shift left << bits move toward the high end build a mask: 1 << 5 is 0x20
shift right >> bits move toward the low end pull out a field

A mask is a number with 1s only in the positions you care about. With a mask m:

  • Set those bits: x | m.
  • Clear them: x & ~m.
  • Flip them: x ^ m.
  • Test them: (x & m) != 0 is true if any of them is 1.

Worked example

Uppercase is one bit

Write a and A in binary, one above the other:

'a' = 0x61 = 0b0110 0001
'A' = 0x41 = 0b0100 0001
             bit 7 ... bit 0

They differ in exactly one position: bit 5, which is worth 0x20. Every lowercase ASCII letter is its capital plus 0x20. So there are three ways to uppercase a letter you already know is lowercase, and all three give 0x41:

Method Arithmetic Result
subtract 0x61 - 0x20 0x41
clear bit 5 0x61 & ~0x20 = 0x61 & 0xDF 0x41
flip bit 5 0x61 ^ 0x20 0x41

Flipping is only safe on a letter you know is lowercase; flip A and you get a back. That is why the echo program checks the range first. This excerpt is from the ESP32 Inspector's C6 echo payload, which has run on real hardware; it is RISC-V assembly, and t2 holds the received byte:

    li      t3, 'a'                  # 0x61
    bltu    t2, t3, tx_wait          # byte < 'a' → leave unchanged
    li      t3, 'z'+1                # 0x7B
    bgeu    t2, t3, tx_wait          # byte > 'z' → leave unchanged
    addi    t2, t2, -0x20            # subtract 0x20 to uppercase

Two comparisons keep anything outside 0x61 to 0x7A unchanged, then one subtraction does the work.

Testing a status bit

Hardware reports its state in status registers: 32-bit numbers where each bit means something. Memory maps and registers explains where they live. The C6 echo payload's status register has bit 2 set when a byte has arrived and bit 1 set when there is room to send. Bit 2 is worth 4, so the test is (status & 0x04) != 0.

Status read In binary (low 3 bits) status & 0x04 Byte waiting?
0x00000006 0b110 0x04 yes
0x00000002 0b010 0x00 no
0x00000004 0b100 0x04 yes

The same Inspector payload does this test in RISC-V with one instruction, andi t1, t1, 4, which ANDs the register with 4 and keeps the result.

A real one-bit bug

The two-bit field below comes from a RISC-V test program in the ESP32 Inspector's clone project, which passed on a real C6. It turns on the clock to the chip's UART (bit 0, called CLK_EN) and then releases the UART from reset (bit 1, called RST_EN, which holds the UART in reset while it is 1):

    ori     t1, t1, 3               # CLK_EN | RST_EN
    sw      t1, 0(t0)
    andi    t1, t1, -3              # clear RST_EN, keep CLK_EN

A computer stores negative numbers in two's complement: −n-n is written as the bits of n−1n - 1, all flipped. So −3-3 in a byte is 0b11111101, a mask with a 0 only in bit 1, and −2-2 is 0b11111110, a mask with a 0 only in bit 0.

Step Bits 1 and 0 Meaning
after ori 3 0b11 clock on, held in reset
andi -3 (correct) 0b01 clock on, running
andi -2 (the earlier bug) 0b10 clock off, held in reset

An earlier version of this program used -2. The project's notes record that a diagnostic dump of the register read 0x02, clock off and reset on, exactly the bottom row. One wrong bit, and the UART never received a byte.

In an ESP32 project

  • Reading registers. A technical reference manual describes each register as named bit fields. You turn "bit 2" into the mask 0x04 and test with AND.
  • Pulling out a field. The Inspector's classic ESP32 echo reads a count from bits 23 to 16 of a status word. In C or Python that is (status >> 16) & 0xFF: shift the field down to the bottom, then mask off everything else.
  • Serial data. Every character on a serial line is a byte sent one bit at a time.

Common mistakes

  • Mixing up bit number and bit value. Bit 5 is worth 0x20, not 0x05. Symptom: a test that is true for the wrong inputs, or a flag that never clears.
  • Clearing with the wrong mask. Writing x & 0x20 keeps only bit 5 instead of clearing it; the clear mask is ~0x20. Symptom: the register goes to 0 and the peripheral stops.
  • Off by one bit. The -2 versus -3 bug above: the program assembles, runs, and silently does the wrong thing.
  • Assignment instead of OR. Writing x = 0x04 to "set bit 2" wipes every other bit in the register. Symptom: some unrelated feature of the peripheral turns off.
  • Reading hex as decimal. 0x10 is sixteen, not ten. Symptom: a delay, size, or offset that is off by a factor near 1.6.

Cost

Bitwise operations are the cheapest things a processor does: on the ESP32's RISC-V and Xtensa cores, an AND, OR, XOR, or shift with a small constant is one instruction. Packing eight yes-or-no flags into one byte instead of eight bytes saves memory, which matters on a chip with a few hundred kilobytes of RAM. The real cost is human: a mask written wrong compiles fine and fails quietly, so write masks as 1 << n with the bit number from the datasheet, and check them against the datasheet's bit table.

Going further

  • Two's complement in more depth: why it lets the same adder handle negative numbers.
  • The full ASCII table, and UTF-8, which extends it to every alphabet.
  • Memory maps and registers, where these masks meet real addresses.
  • How a CPU runs instructions, which shows that instructions themselves are binary numbers.

Leads to

Back to ESP32 development: assembly, C, MicroPython, and CircuitPython