prerequisite

How a CPU runs instructions

The fetch-decode-execute loop, registers, the program counter, the stack, and function calls, the model every language here compiles or interprets down to.

Before this

This page assumes you are comfortable with:

Why you need this

Whether you write MicroPython, C with ESP-IDF, or assembly, the ESP32 ends up doing the same thing: running one small instruction after another. Knowing that loop explains why a busy program can still miss a byte, why assembly is fast, why a crash report shows a "PC" and a "return address", and what the cluster's RISC-V echo program is actually doing.

The idea

Instructions are numbers in memory

A CPU (central processing unit, the processor core) understands a fixed set of simple instructions: add two numbers, copy a value from memory, jump somewhere else if a value is zero. Each instruction is stored in memory as a number. On the ESP32-C6, which uses the RISC-V instruction set, most instructions are 32 bits.

For example, addi t2, t2, -0x20 ("add the constant −32-32 to register t2") encodes as the 32-bit number 0xFE038393. Split that into bit fields and the CPU can read off which operation, which registers, and which constant. The C6 also supports RISC-V's compressed extension, which lets an assembler store some common instructions in 16 bits; this one fits, as 0x1381. Either way, the program is just numbers in memory.

Registers

A register is a tiny, very fast storage slot inside the CPU, holding one 32-bit word. RISC-V has 32 of them. Assembly code calls them by short names: t0 to t6 for temporary values, a0 to a7 for function arguments, sp for the stack pointer, ra for the return address. Arithmetic happens only between registers. Memory is much larger but slower, so programs load a value from memory into a register, work on it, and store it back.

The program counter and the loop

The program counter (PC) is a special register holding the address of the next instruction. The CPU runs one loop forever:

  1. Fetch: read the instruction at the address in the PC.
  2. Decode: split its bits into the operation, the registers, and any constant.
  3. Execute: do it, which may change a register, read or write memory, or change the PC.
  4. Move the PC to the next instruction, unless step 3 already set it.

Branches and loops

A branch is an instruction that changes the PC only if a condition holds. beqz t1, echo_loop means "if t1 equals zero, set the PC to the address labeled echo_loop; otherwise carry on". A loop is just a branch that goes backward. A jump (j) changes the PC unconditionally. The if, while, and for of every higher-level language turn into branches and jumps.

The stack and function calls

A function call has to remember where to come back to. RISC-V's jal ra, msgout ("jump and link") sets the PC to msgout and saves the address of the next instruction in the return address register ra. The function ends with jalr x0, ra, which jumps back to whatever ra holds.

If that function calls another one, the inner call overwrites ra. So the function first saves ra on the stack: a region of RAM used last-in, first-out, whose current top address is held in the stack pointer sp. This excerpt is RISC-V assembly from a test program in the ESP32 Inspector's C6 clone project, which passed on real hardware. It is a function that sends a text message by calling a send-one-byte function tx for each character:

msgout:
    addi    sp, sp, -4
    sw      ra, 0(sp)
msgloop:
    lb      t2, 0(t3)
    addi    t3, t3, 1
    beq     t2, x0, msgend
    jal     ra, tx
    jal     x0, msgloop
msgend:
    lw      ra, 0(sp)
    addi    sp, sp, 4
    jalr    x0, ra

The first two lines make room for one word on the stack and store ra there. The loop loads one byte of the message (lb), advances the pointer, stops at a zero byte, and calls tx. At the end, lw gets the saved return address back, sp moves back up, and jalr returns. (x0 is a register that always reads 0, so jal x0, msgloop is a plain jump.)

Clock speed

The CPU's steps are paced by a clock: a signal that ticks a fixed number of times per second, measured in megahertz (MHz, millions of ticks per second). The author's board notes record the C6's main core at 160 MHz, the classic ESP32 and S3 at 240 MHz, and the P4 at 360 MHz. A simple instruction can finish in about one tick, but loads from slow memory, taken branches, and multiplies can take more, so clock speed gives a rough ceiling, not an exact rate.

The demo runs the uppercase echo loop, simplified from the Inspector's C6 echo payload. Queue some bytes, then press Step: the arrow is the PC, and the register line shows t0 to t3 change. Notice that beqz sends the PC back to the top while no byte is waiting, and that li with a full address is shown as one line. In the real program the assembler expands it into two instructions, lui (load the upper 20 bits) and addi (add the low 12).

Worked example

Trace the first four instructions of the echo loop, with the PC shown as a line number as the demo does. Suppose one byte, q, is waiting, so the status register at 0x6000F004 reads 0x00000006: bit 2 (a byte waiting) and bit 1 (room to send).

Step PC (line) Instruction What it does t0 t1
1 1 li t0, 0x6000F004 put the status register's address in t0 0x6000F004 0x00000000
2 2 lw t1, 0(t0) load the word at that address into t1 0x6000F004 0x00000006
3 3 andi t1, t1, 4 keep only bit 2 0x6000F004 0x00000004
4 4 beqz t1, echo_loop t1 is not zero, so do not branch 0x6000F004 0x00000004

After step 4 the PC moves to line 5, which reads the byte. Had no byte been waiting, step 2 would have loaded 0x00000002, step 3 would have produced 0, and step 4 would have set the PC back to line 1. That four-line circle is the program waiting.

How long do a million instructions take? Assume the C6's 160 MHz clock and exactly one instruction per tick, which is the best case. One tick lasts 1/160,000,0001 / 160{,}000{,}000 s =6.25= 6.25 ns (nanoseconds, billionths of a second). Then

t=1,000,000160,000,000 per second=0.00625 s=6.25 mst = \frac{1{,}000{,}000}{160{,}000{,}000 \text{ per second}} = 0.00625 \text{ s} = 6.25 \text{ ms}

Real code with loads and taken branches will take longer. At 240 MHz the same best case is about 4.2 ms.

In an ESP32 project

  • Assembly pages (RISC-V, Xtensa) write these instructions directly.
  • C is turned into instructions like these before it reaches the chip; MicroPython and CircuitPython run an interpreter made of them. Interpreters and compilers explains the difference.
  • Crash reports from ESP-IDF print the PC and a backtrace built from saved return addresses; Debugging resets and crashes reads them.
  • Polling: the echo's waiting circle is a busy-wait. Polling and interrupts counts how often it spins.

Common mistakes

  • Running assembly for the wrong instruction set. The classic ESP32 and S3 use Xtensa, the C6 and P4 use RISC-V. Symptom: the program will not assemble, or the chip crashes immediately.
  • Calling a function without saving ra. Symptom: the outer function "returns" into itself and loops forever.
  • Unbalanced stack. Moving sp down without moving it back. Symptom: a crash some time later, far from the bug.
  • Branch condition reversed. beqz where bnez was meant. Symptom: the program skips the work or spins forever.
  • Treating MHz as a speed guarantee. Symptom: timing loops calibrated by "one instruction per tick" run slow.

Cost

An instruction costs time (a few nanoseconds at these clock speeds) and space (2 or 4 bytes of flash or RAM on RISC-V). A busy-wait loop costs the whole CPU for as long as it spins, plus the power to run it. Function calls cost a few extra instructions each for saving and restoring ra and moving sp, which only matters inside very tight loops.

Going further

  • Pipelining: how a CPU overlaps fetch, decode, and execute of several instructions.
  • Caches, and why code running from flash can be slower than code in RAM.
  • The RISC-V instruction formats (R, I, S, B, U, J) and how bits map to fields.
  • Memory maps and registers, for what the addresses in the echo loop point at.
  • RISC-V assembly on the ESP32-C6 and P4, for the full echo line by line.

Leads to

Back to ESP32 development: assembly, C, MicroPython, and CircuitPython