technique

Over-the-air updates

Updating a board's code without a USB cable: manifests, hashes, safe swaps, and what happens when an update fails.

Before this

This page assumes you are comfortable with:

Why you need this

Once a board is screwed into a case, mounted on a wall, or handed to someone else, plugging in a USB cable for every fix stops being practical. An over-the-air update (OTA) lets the board fetch new code over WiFi and install it itself. It is the last part of stage 5, "Go wireless", and it is only as safe as its plan for failure: an update that goes wrong can leave a board that can no longer update.

The idea

Every OTA system answers four questions.

  1. Is there something new? The board asks a server for a small description of the current release, usually called a manifest.
  2. What do I need to download? The board compares the manifest with what it already has.
  3. Did it arrive intact? The board checks each download against a hash: a short fingerprint computed from every byte of a file. SHA-256 is the usual one. It turns any file into 32 bytes (64 hex characters), and changing even one bit of the file changes the fingerprint completely.
  4. How do I switch over without being left half-updated? The board keeps the old version usable until the new one is fully in place.

There are two common shapes on an ESP32.

ESP-IDF app OTA File-level OTA (MicroPython)
What is replaced The whole compiled app image Individual .py files in the board's filesystem
Where the new code goes The spare of two app partitions in flash A temporary file next to the old one
How the switch happens The bootloader is told to boot the other partition Files are renamed, then the board resets
Recovery if the new code is bad Automatic rollback to the previous partition (if enabled) Whatever the update code itself provides

ESP-IDF app OTA

This summary is from Espressif's ESP-IDF "Over The Air Updates" documentation. Flash must hold at least two app slots, ota_0 and ota_1, plus an otadata partition, two flash sectors (0x2000 bytes) that record which slot to boot. The app writes the new image into the slot it is not running from: esp_ota_begin() erases that slot, esp_ota_write() writes the image in pieces as it downloads, esp_ota_end() checks the image, and esp_ota_set_boot_partition() updates otadata. On the next reset the second-stage bootloader (see The boot sequence) starts the new slot.

With the option CONFIG_BOOTLOADER_APP_ROLLBACK_ENABLE turned on, a new image boots in the state ESP_OTA_IMG_PENDING_VERIFY. The new app runs its own checks and calls esp_ota_mark_app_valid_cancel_rollback() if they pass, or esp_ota_mark_app_invalid_rollback_and_reboot() if they fail. If it crashes or resets before marking itself valid, the bootloader marks it aborted and boots the previous app. The old image stays in flash the whole time, so a bad update costs one reboot, not the board.

For the common case, the esp_https_ota component wraps all of this into one call. This is the shape of the example in Espressif's ESP HTTPS OTA documentation (official documentation, not hardware-tested here):

esp_http_client_config_t config = {
    .url = CONFIG_FIRMWARE_UPGRADE_URL,
    .cert_pem = (char *)server_cert_pem_start,
};
esp_https_ota_config_t ota_config = {
    .http_config = &config,
};
esp_err_t ret = esp_https_ota(&ota_config);
if (ret == ESP_OK) {
    esp_restart();
}

Note cert_pem: the documentation requires the server's root certificate, so the board can verify it is downloading from the real server.

File-level OTA: the ESP32 Inspector's update stub

The ESP32 Inspector can provision a MicroPython board with a tiny update system. Its main.py, excerpted whole from the Inspector project, is three lines of code:

# ESP32 Inspector bootstrap: update over WiFi, then run the current payload.
# Deliberately three lines: everything lives in inspector_update so the stub
# can update itself. Ctrl-C at boot drops to the REPL (never caught).
import inspector_update

inspector_update.boot()

Everything else lives in inspector_update.py. Its boot sequence, excerpted from the Inspector project, is the whole policy in a dozen lines:

def boot():
    cfg = _read_json(CONFIG_PATH)
    if not cfg or not cfg.get("ssid"):
        ...
        config_portal()  # never returns
    if connect_wifi(cfg) is None:
        ...
        config_portal()  # never returns
    try:
        if check_and_apply(cfg):
            import machine
            ...
            machine.reset()
    except Exception as e:
        # Stale payload beats a bricked board: log and fall through.
        print("inspector: update check failed:", e)
    state = _read_json(STATE_PATH) or {}
    _run(state.get("run") or DEFAULT_RUN)

Read it as: no WiFi settings or a failed join opens a setup access point named "ESP32-Inspector" with a form for new WiFi settings; a successful update resets the board so it starts fresh on the new files; a failed update check is logged and the old program runs anyway.

Deciding what to download is one line. This function is excerpted from the same file:

def plan(manifest, state):
    """Return the manifest file entries whose sha256 differs from what the
    board last applied. state is None/{} on a fresh provision -> everything."""
    have = (state or {}).get("files", {})
    return [f for f in manifest["files"] if have.get(f["path"]) != f["sha256"]]

The board keeps inspector_state.json, a record of the hash of every file it last applied. Comparing hashes, not version numbers, means a file that did not change is not downloaded again.

Applying runs in two phases. This excerpt is from check_and_apply() in the same file:

# Phase 1: download everything to .new and verify before touching targets.
for f in todo:
    ...
    tmp = f["path"] + ".new"
    h = hashlib.sha256()
    with open(tmp, "wb") as out:
        def sink(b, _h=h, _out=out):
            _h.update(b)
            _out.write(b)
        _http_get(url, sink)
    digest = binascii.hexlify(h.digest()).decode()
    if digest != f["sha256"]:
        os.remove(tmp)
        raise ValueError("sha256 mismatch for %s (got %s)" % (f["path"], digest))
    ...
# Phase 2: swap. Not power-loss-atomic, but the window is rename-only.
for f in todo:
    ...
    os.rename(f["path"] + ".new", f["path"])

The hash is computed while the bytes stream in, so the file is never held whole in RAM. Nothing old is touched until every new file has arrived and matched. According to the author's project notes, this path has run end to end on a real board: WiFi join, TLS manifest fetch, sha256-verified downloads, and the self-reset into the new program. The setup access point had not yet been proven on hardware at the time of those notes.

Worked example

Take a one-file program. The file app.py holds print('hi') followed by a newline: 12 bytes. Its SHA-256, computed with Node's crypto module, is:

caf026f25d7140209f98072605307a438914b9ce6f3c14b23d15d9667241de52

The board is currently running the previous release, whose app.py was print('hello') plus a newline (15 bytes, SHA-256 03e693d9f2f687e0f40e36a8df7fcb4d1c22974012b7c2a55c000eb30f305824). This manifest, written for this page in the shape the Inspector's channel manifests use, announces the new release:

{
  "version": 7,
  "run": "app.py",
  "files": [
    {
      "path": "app.py",
      "url": "/payloads/files/app.py",
      "sha256": "caf026f25d7140209f98072605307a438914b9ce6f3c14b23d15d9667241de52",
      "size": 12
    }
  ]
}

Step by step:

Step What the board does Result
1 Fetch the manifest over HTTPS Version 7, one file listed
2 Read inspector_state.json app.py recorded as 03e693d9...
3 plan() compares hashes 03e693d9... differs from caf026f2..., so app.py is to do
4 Download to app.py.new, hashing as it goes 12 bytes, digest caf026f2...
5 Compare digest with the manifest Match, so keep the file
6 Rename app.py.new over app.py, write the new state State now records caf026f2... and version 7
7 machine.reset() The board boots and runs the new app.py

On the next boot, step 3 finds every hash equal and nothing is downloaded. If step 5 had found a mismatch, say because the connection dropped and the file was short, the .new file is deleted, an error is raised, boot() logs it, and the old app.py runs.

In an ESP32 project

OTA joins stage 3 and stage 5. It needs the WiFi join and HTTPS fetch from WiFi and MQTT, and in the ESP-IDF form it needs a partition table with two app slots (see Flash, RAM, and partitions) and the second-stage bootloader that chooses between them. USB remains the floor: the very first install, MicroPython itself plus main.py and the stub, goes over the cable as on Flashing and the ROM bootloader. OTA can only keep a board current once something that knows how to update is on it.

Common mistakes

  • Power lost mid-update. An ESP-IDF app OTA survives because the old slot is untouched. The Inspector stub downloads to .new files first, so only the short rename phase is exposed; its own comment says it is not power-loss-atomic. Symptom: a board that boots to a missing or mixed set of files.
  • An update that breaks the update path. The Inspector's update stub is itself listed in its manifest, so it can update itself, and a broken stub would stop all future updates. Symptom: the board runs, but never picks up the next release. Only USB fixes it, which is why rollback and careful testing of the updater matter most.
  • Trusting a hash from the same unverified source. A hash catches corruption. It does not prove who sent the file if the manifest comes over a connection whose certificate was never checked, as with the Inspector stub, which has no certificate store. ESP-IDF's cert_pem closes that gap. Symptom: none, until someone impersonates the server.
  • Wrong clock, failed TLS. Checking a certificate means checking its dates, and a board that has not set its clock does not know the date. Symptom: the handshake fails right after boot and works later.
  • Forgetting to mark the app valid. With rollback enabled, an app that never calls esp_ota_mark_app_valid_cancel_rollback() is rolled back at the next reset. Symptom: the update seems to work, then the old version returns after a reboot.

Cost

Flash is the main bill for ESP-IDF OTA. Espressif's built-in "Factory app, two OTA definitions" partition table gives factory, ota_0, and ota_1 1 MB each (1 MB = 1024 KB = 1,048,576 bytes), ending at offset 0x310000, which is 3136 KB; on a 4 MB flash that leaves 960 KB for anything else, so an app can only grow to the slot size. File-level OTA costs only the space of the changed files twice during the swap, plus the stub itself, which the Inspector's current manifest lists at 13,620 bytes (about 13.3 KB). Network cost is one small manifest fetch per boot, with one TLS handshake per request. The maker's time is spent once on the updater and on testing failure paths; after that each release is a file upload instead of a walk to every board.

Going further

  • Espressif's ESP-IDF "Over The Air Updates" and "ESP HTTPS OTA" documentation, especially the rollback and anti-rollback sections.
  • Signed updates: ESP-IDF's secure boot verifies that an image was signed by your key, which proves the sender, not just the bytes.
  • Debugging resets and crashes, for the board that reboots after an update and the reason it does.
  • The ESP32 Inspector's board detection page, which identifies a board before you provision it.

Back to ESP32 development: assembly, C, MicroPython, and CircuitPython