Skip to content

v1.1.0 - #18

Merged
diegoparrilla merged 84 commits into
mainfrom
release/v1.1.0
Sep 17, 2026
Merged

diegoparrilla merged 84 commits into
mainfrom
release/v1.1.0

Conversation

@diegoparrilla

Copy link
Copy Markdown
Contributor

v1.1.0 — the robustness release. No new features: everything here is a fix or
a recovery path. It is also the first release built with full optimisation
(Release); everything up to v1.0.1beta shipped MinSizeRel because
Release did not survive on hardware.

See CHANGELOG.md for the user-facing entry. In short:

  • Transfers of any size finish. Uploads much beyond 200 KB never used to
    arrive; a 4 MB file now round-trips byte-identical, and a slow client is no
    longer cut off.
  • A transfer that dies stops blocking the next one. A client that
    vanished mid-transfer left the lock held and FatFs handles counted until a
    reset.
  • The SD card can come and go. A reinserted card is picked up in about
    two seconds, and with no card the device says so — SD: NO CARD on the
    menu, [G]/[U] and the countdown refused, 503 no_sd_card from every
    endpoint that needs it.
  • Writes from the Atari arrive whole. A retried write chunk used to be
    written twice, losing the tail of the file.
  • Wi-Fi: the radio no longer sleeps (99 ms → 27 ms, 15% → 5% loss), the
    link rejoins by itself, a bad static IP falls back to DHCP instead of
    crashing before the menu, and the password is out of the logs.
  • Crashes are visible and survivable: panic, HardFault and an 8 s hang
    each reboot with the reason, phase and faulting PC on the menu and in
    /api/v1/system/health, behind a crash-loop guard.
  • Docs: CLAUDE.md, README, docs/api.md and programming.md describe
    v1.1 as built, with no references to untracked planning notes.

Gate

  • smoke.py 7/7 on release, 8/8 on debug; 4 MB upload and download
    byte-identical; aborted clients and a second client during a transfer all
    answered; forced panic, hardfault and hang recovered with the reason
    recorded.
  • One-hour soak: 906 rounds, 0 failures, uptime unbroken, crash_count 0,
    free heap back to its starting value, stack 3,556 / 16,384.

Ships with these known gaps

  1. Heavy tracing through the debug ABI can take the ST down. A program
    reading more than a few thousand bytes through the cartridge debug window
    during runner exec faults the Atari. The firmware side has been measured
    and cleared; the cause is not found, and the Runner has no soak evidence.
    Work in progress on epic/18-runner-execution-stability.
  2. Wi-Fi recovery is verified against injected faults, not a real AP outage.
  3. Untested on other cards: blank, exFAT, 32 KB clusters.
  4. SELECT long press not exercised, and the card has not been pulled while
    the ST had GEMDRIVE open.

Bump the three tracked version files to v1.1.0. This branch collects
the v1.1.0 epics through pull requests and merges into main only when
the version is complete.
…types (EPIC-08 STORY-01 STORY-02)

A failed step no longer gets stepped over. build.sh, rp/build.sh and
target/atarist/build.sh stop on the first error. Before this, an m68k
failure let the RP build embed the committed target_firmware.h, and a
C or link failure let build.sh copy an older UF2 from rp/dist and write
a JSON announcing the new version.

- rp/build.sh empties rp/dist before building.
- target/atarist/build.sh deletes any previous target_firmware.h and
  checks that firmware.py produced a new one, using python3 when the
  host has no python.
- build.sh falls back to md5 -q when md5sum is missing, as on macOS.

The build types are now what their names say, case-insensitively:
release is CMake Release with DEBUG_MODE=0, and debug is the same
build with DEBUG_MODE=1 for DPRINTF traces on the UART console.
RP_CMAKE_BUILD_TYPE overrides only the CMake build type, to compare
against the MinSizeRel configuration that shipped until now. A
caller-provided RELEASE_DATE is kept, so two builds of one commit can
be compared byte for byte.

FF_FS_TINY no longer follows _DEBUG, so a debug build keeps the same
FatFs objects as a release build.

With the fail-fast in place the first run stopped at a stale
.git/modules/pico-sdk/index.lock from April. The old script had been
skipping the pico-sdk pin silently ever since.
…RY-03 STORY-05)

--strip-all removed the symbol table and debug information from
rp.elf without changing a byte of the flashed image, so crash
addresses could not be decoded and a debug probe could not break at
main. The link keeps --gc-sections only.

Also removed, all without effect on the image:
- the clang-tidy block, which set CMAKE_C_CLANG_TIDY after every
  target existed and so never ran;
- the settings library: settings.c was compiled into it and a second
  time directly into rp, and only the direct copy was linked;
- pico_set_binary_type(rp), which selects nothing because the custom
  linker script takes precedence.

The stdio comment now says what the code does: USB stdio stays for
usbcdc.c, and only debug builds enable UART stdio.

Verified with RP_CMAKE_BUILD_TYPE=MinSizeRel and a fixed RELEASE_DATE:
rp.uf2 is byte-identical to the same build before these changes.
…Y-03)

CMake Tools offered its stock Debug and Release variants, which built
with DEBUG_MODE and APP_UUID_KEY unset: no traces, the default UUID,
and a CMake type the scripts never use. .vscode/cmake-variants.yaml
now defines the same two build types as ./build.sh, both CMake Release,
with DEBUG_MODE and the development UUID set at configure time.

The debug launch configuration still carried the template's name.
…-04)

Both workflows installed Ubuntu's packaged gcc-arm-none-eabi, so the
released UF2 came from a different compiler than the ARM GNU Toolchain
14.2 used locally. They now download and cache 14.2.rel1 and print the
compiler version.

- build.yml builds release and debug in a matrix, writes a size report
  (sections, code copied to RAM, heap, reserved stack, flash used) to
  the job summary, and uploads the firmware with rp.elf and its map.
- release.yml builds release in lower case, writes the same report, and
  attaches rp.elf and its map to the GitHub release. They are not
  uploaded to the S3 bucket.
- actions/checkout v4, actions/setup-python v5, and GITHUB_OUTPUT
  instead of the deprecated set-output command.
…ORY-04)

The build section still said every build is MinSizeRel and that
clang-tidy runs from CMake. It now describes the release and debug
types, the MinSizeRel override, reproducible RELEASE_DATE, symbols in
rp.elf, the VS Code variants and the CI matrix.
The stcmd that atarist-toolkit-docker's installer writes, which CI uses,
takes $VERSION as the Docker image tag. build.sh exported the firmware
version under that name, so CI asked for atarist-toolkit-docker:v1.1.0,
which does not exist, and the m68k build failed. The previous build.sh
exported it too and ignored a failed m68k build, so earlier CI and
release builds most likely embedded the committed target_firmware.h;
the logs of those runs have expired. VERSION is only used inside
build.sh.
The statistics were keyed on NDEBUG, which both build types define, so
MEM_STATS, MEMP_STATS and stats_display were compiled out everywhere.
Port Booster's block: a debug build with DEVOPS_LWIP_STATS=1 in the
environment turns them on. Release is unchanged; its lwIP objects are
byte-identical to before.
…ORY-05)

commemul_poll derived the write index from the DMA transfer count with
no lap detection, so a main loop stalled for more than 4,096 samples
silently mixed old and new data into the protocol parser. Count a lap
when that many samples arrived since the last poll, and drop the unread
ones. The channel stopped for good after 2^32 samples; re-arm it at
2^31 from its live write address.

A corrupted frame could also carry a payload size larger than the
payload buffer, and the parser wrote past it. Reject such a frame.
…(EPIC-09 STORY-01 STORY-04 STORY-06)

A panic or HardFault used to freeze the RP silently: the SDK's panic
stops at a breakpoint and fatfs-sdk's HardFault handler does the same,
which locks up an RP2040 with no debugger. A hang had no recovery at all.

health.c adds:
- panic() through PICO_PANIC_FUNCTION and a HardFault handler installed
  in the RAM vector table at boot. Both write the reason, PC, LR and SP
  to watchdog scratch registers 0-3 and reboot after 100 ms.
- An 8 s watchdog, paused under a debug probe, fed from the main loop,
  emul_pollTick, the Wi-Fi connect loop, the HTTP spin-waits and after
  the two slow boot steps. A phase ID in scratch[0] says where a hang
  happened.
- Boot decoding into power-on, reset, panic, HardFault, hang or reboot,
  shown on row 2 of the setup menu and on the debug console.
- A crash-loop guard: 3 crash reboots within 60 s stop the countdown
  until a SELECT reset or a power cycle.
- A painted core-0 stack for a high-water mark, heap free and minimum
  free from mallinfo and sbrk, and code copied to RAM.
- Debug-only fault injection for hardware verification.

sd_timeouts.c overrides fatfs-sdk's weak timeout table so a failing card
gives up well within the watchdog. The jump to Booster stops the
watchdog first, since it does not reset the chip.
A release build has no console, so nothing on the device could be read
without a debugger. GET /api/v1/system/health returns uptime, heap total,
free, minimum free and sbrk high-water, stack high-water and reservation,
code in RAM, the last reset reason with its phase or PC, the crash count
and crash-loop state, ROM3 overruns and debug-byte drops. A debug build
with DEVOPS_LWIP_STATS adds lwIP's pool counters. It is a new endpoint
rather than more ping fields, so ping stays a minimal liveness check.

`sidecart.py health` prints it, with tests. Debug builds also expose
POST /api/v1/debug/test/<fault> to trigger each recovery path on
hardware. Both are documented in docs/api.md.
…d (EPIC-09 STORY-06)

With a mounted SD card pulled, sd_read_blocks returned early from its
error path without releasing the card and SPI mutexes. Those mutexes
block without a timeout, so the next disk call waited forever: the
first GEMDRIVE or HTTP request after the pull failed, and the second
froze the RP until the watchdog rebooted it (reported as a hang in
disk_read by a diagnostic build).

Upstream fixed it in f5f31b3. 6c644cf is that fix plus a README change,
and is the commit Booster already pins. Verified on hardware: after an
idle pull, listings and downloads fail in under a second and the RP
stays up.
Flashing or resetting through a debug probe uses SYSRESETREQ, which
leaves the watchdog reason and scratch registers exactly as a real
watchdog expiry does, so every probe flash showed "Recovered: hang"
and counted as a crash. Measured on hardware: reason, watchdog control,
chip reset and the SDK's marker all match, and DHCSR reads as zero from
the firmware.

A once-a-second timer interrupt now writes a stall mark into scratch[1]
after 5 s without a watchdog feed, and clears it when feeding resumes.
A watchdog reset counts as a hang only if the mark is there. A hang with
interrupts disabled cannot set the mark and is reported as a reboot.

Verified on hardware: main-loop hang and HTTP handler hang still report
their phase; a probe reset and a probe flash now report a reboot with
crash count 0.
…d-recovery

EPIC-09: device-side visibility and recovery
…10 STORY-02 STORY-03)

The image reaches RAM through the XIP stream, whose address register has
no bits 1:0. target_firmware is a uint16_t array, only 2-byte aligned, so
at an address that is 2 mod 4 every word would be shifted and the ST would
boot without the cartridge. It has been aligned by luck. firmware.py now
emits aligned(4), and COPY_FIRMWARE_TO_RAM falls back to a word copy for
any source that is not 4-byte aligned.

The stream also moves 32-bit words, so an odd word count lost the last
16-bit word. It is now copied separately. The 64 KB window is cleared
before the copy, so nothing from a previous run (a jump from Booster, a
crash reboot) survives past the end of the image. COPY_FIRMWARE_TO_RAM_MEMCPY
took bytes while every other copy macro takes words; it takes words now.

Debug builds check the window against target_firmware after the copy and
count non-zero bytes past the image. target_firmware.h is regenerated by
the m68k build; only the declaration and the header's build date change.
…TORY-01)

__rom_in_ram_start__ was declared as `extern unsigned int`, so to the
compiler it is a 4-byte object and every read or write past those bytes
is undefined behaviour it may optimise on at -O3. Clearing the window
with memset made GCC warn that 65,536 bytes overflow a 4-byte region.
Declare it as an array of unknown size in constants.h and in the two
files that redeclare it. Every use takes its address and casts it, so no
call site changes and the code size is identical.
[G] and [U] make the m68k install GEMDRIVE at the relocation address and
memtop the RP publishes when the ST sends HELLO at cold boot. After the RP
reboots while the ST keeps running (a crash, a SELECT reset, a flash) no
HELLO arrives. Recovery only worked because the previous run's shared
variables survived in RAM; with the window now cleared, the address reads
0 and the ST halts with "Reloc/stack overlap".

The launch commands now refuse until HELLO has landed since the RP booted,
keep the full menu on screen and say "Reset the Atari ST first" on the
bottom strip. When HELLO lands the message clears, and if the refused
launch came from the countdown, the countdown starts again. Keystrokes
the ST sends while it boots are dropped at that point, or the main loop
would halt the restarted countdown before its first tick.

Verified on hardware (release): a launch after an RP reboot is refused
from the countdown and from [U] with the menu intact; after an ST reset
the countdown restarts, runs to 0 and Runner launches without the banner.
DPRINTF blocks on the UART. At the SDK's default 115,200 baud a debug
build printed its flash layout and both settings dumps before copying the
cartridge image, and a power-cycled ST booted into GEM because the
cartridge was not live yet; release, which prints nothing, showed the
menu. At 921,600 the same output is eight times faster and the
power-cycled ST shows the DevOps menu on the debug build too.

The UART clock stays at 48 MHz after the 225 MHz overclock, so the divider
gives 921,413 baud, a 0.02 % error. Only debug builds get the define;
release has no UART console. Set the serial terminal to 921600.
…RY-04)

DPRINTF waited until every byte had left the UART, so each trace stalled
the main loop, the ROM3 command ring and the ST's commands for about 1 ms
at 921,600 baud, and bursts such as GEMDRIVE's per-call traces for much
longer.

DEBUG_BUFFERED_CONSOLE in debug.h, default 1, toggles a buffered console.
DPRINTF formats with fprintf into a funopen stream that queues the text in
a 4 KB ring buffer, and the UART transmit interrupt sends it. When the
ring is full the caller waits for room, or sends bytes itself with
interrupts off, so nothing is lost. Inside an exception the queue and the
new text are sent at once, so crash reports stay complete and in order.
stderr and stdout are redirected to the same queue, so nothing else writes
to the UART behind the interrupt. Set it to 0 for the previous blocking
console, which keeps the last lines before a hang.

Everything lives in debug.h: the shared state and functions are weak
definitions merged by the linker into one copy. debug.h replaces DPRINTF
definitions made earlier, and settings.c includes it, so the settings
library's own fallback DPRINTF no longer bypasses the queue.

Debug builds only: 4 KB less heap and about 500 bytes of flash. Release is
unchanged. Verified on hardware: a power-cycled boot log is complete and in
order, and a forced panic prints its report whole, in order, before the
reboot.
usbcdc_init() called stdio_init_all() a second time in debug builds,
assuming it was idempotent. It is not: stdio_uart_init() runs uart_init(),
which resets the UART and throws away up to 32 bytes still in the console
FIFO, and stdio_usb_init() claims another IRQ and background task on every
call. With the buffered console this showed as exactly 32 bytes missing
from a boot trace.

Start USB stdio only when TinyUSB is not running yet. Release builds, where
main.c does not initialise stdio, behave as before.
__flash_binary_start and the ROM temp, Booster, config, lookup and global
config region symbols were declared as `extern unsigned int`, 4-byte
objects to the compiler. aconfig.c reads the whole 4 KB lookup table
through a pointer taken from _global_lookup_flash_start, which is undefined
behaviour the optimiser may use. Declare them as arrays of unknown size, as
__rom_in_ram_start__ already is. Every use takes the address and casts it,
so no call site changes.
…-05)

The streaming listing checked for 32 bytes of room, enough for the closing
envelope, before reading the next entry. An entry needs about 80 bytes, up
to 350 with a long name, so once a chunk had between 32 and 80 bytes left
the entry just read did not fit: it was dropped and the listing ended with
"truncated": true. Any folder with more entries than fit in one 1 KB chunk,
about 15, came back cut short.

Check for room for the largest possible entry before f_readdir, since an
entry that has been read cannot be put back; the listing then continues in
the next chunk. Found by the EPIC-10 smoke script; verified on hardware with
100 and 500 files: all entries listed, "truncated": false.
…-10 STORY-05)

FOLDER defaulted to the microfirmware template's "/test", which md-devops
never uses, so every boot created an empty /test folder on the card.
FOLDER now defaults to "/devops" and is created at boot if it does not
exist. Settings saved with the old "/test" value are changed to "/devops"
and saved once at boot, then the folder is created if needed.

Verified on hardware: the first boot logged the change and rewrote the
app settings with every other value kept; the next power-on boot read
FOLDER=/devops and mounted the card with no /test folder involved.
…iming-bugs

EPIC-10: latent layout and timing bugs
tools/dev/console.py watch reads the Debug Probe's UART at 921,600
baud, timestamps each line, appends it to tools/dev/logs/console.log and
shows it in the terminal. It finds the probe by its USB name, and waits
for it and reopens it after an unplug. since-boot, tail, grep and wait
read the log, so a boot log no longer has to be pasted by hand.

Standard library only. Verified on hardware: a probe reset produced a
complete boot log, and wait matched the GEMDRIVE ready line; unplugging
and replugging the probe resumed the capture.
…-17 STORY-02)

Every build now carries the git commit it came from: <sha7>, or
<sha7>-dirty.<diff7> with uncommitted changes, where <diff7> hashes the
diff. rp/src/build_id.cmake regenerates build_id.h on every build and
rewrites it only when the ID changes, so incremental builds never report
a stale ID and only health.c and http_server.c recompile. The same tree
gives the same ID and a byte-identical binary.

The ID is printed on the console at boot and returned by
GET /api/v1/system/health with a new "debug" flag; sidecart.py health
prints both. The health reply buffer grows by 32 bytes for them.

tools/dev/flash.sh builds out of tree and incrementally, keeps the ELF by
build ID, flashes with picotool or through the Debug Probe, and waits
until the health report shows the new ID and build type, printing the
console since boot if it does not.

Verified on hardware: debug with picotool and with the probe, release
with the probe, and the timeout path; about 18 s from command to a
confirmed boot for an unchanged build.
…C-17 STORY-02)

The developer tools must work with any microfirmware from the template
and with a hung RP, so they reach it only through picotool, the Debug
Probe and the console UART.

tools/dev/swd.py wraps OpenOCD and reads memory while the CPU runs:
running waits until VTOR holds the ELF's RAM vector table, verify
compares the whole flash with the ELF, build-id reads the new
release_build_id string from flash, read dumps a range, and program
flashes through the probe. Runs that fail on a momentary debug-port drop
are retried.

build_id.cmake now also generates build_id.c with release_build_id,
kept through --gc-sections. flash.sh checks running, verify and build-id
over SWD instead of polling GET /api/v1/system/health, which keeps
reporting the build ID for API users.

Verified on hardware: debug with picotool and release with the probe,
each running, verified byte for byte and carrying its build ID; a release
ELF against a running debug build fails verify.
swd.py screen reads the 320x200 framebuffer at the top of the cartridge
window while the RP runs, or while it is halted, and writes it as a PNG,
so the setup menu can be checked without looking at the ST. swd.py
shared prints the command sentinel, the random token and the shared
variables, named after the *_SVAR_* indexes in rp/src/include.

Neither hard-codes an address: the window comes from the ELF's
__rom_in_ram_start__ and the offsets from chandler.h, whose #defines are
evaluated by a small parser. Without --elf the tools pick the cached ELF
whose build ID the RP carries.

swd.py running now also requires core 0 not to be halted, and swd.py
resume releases both cores through DHCSR: a halted core 1 pauses the
RP2040's timer and watchdog and leaves core 0 asleep, and OpenOCD's own
resume fails in a new OpenOCD run.

Verified on hardware: the PNG matched the ST's screen before HELLO, after
an ST reset (Phystop, Screenmem and the countdown line), and after a
forced panic (Recovered: panic @10005194, the PC on the console). After
HELLO the shared variables showed reloc and memtop 0xF4000, drive C and
hook vector $70.
…Y-04)

swd.py select presses the SELECT button without firmware code: it forces
the pin's input high through the RP2040's GPIO input override for the
hold time. A long press needs --force, because md-devops then erases the
global settings.

For what the firmware must do itself, devhooks.h adds a debug-only
mailbox in RAM, self-contained like debug.h. Tools write a request over
SWD and devhooks_poll(), called from the main loop, runs it and
acknowledges it. A protocol request is queued through the new
chandler_injectProtocol() as if the ST had sent it, so swd.py key and
swd.py inject use the firmware's normal command path. An app request
goes to a handler the app registers; emul.c handles countdown stop and
restart, named by DEVHOOKS_APP_* defines that swd.py app reads. Release
builds contain none of it.

Verified on hardware with the menu: countdown stop and restart and a
key 'z' each showed the expected bottom line in swd.py screen, the
console logged the injected keystroke, and a SELECT short press was
detected after 304 ms and rebooted the RP with reason reset.
With DEVOPS_LWIP_STATS on a debug build, ran a 4 MB upload, a 4 MB
download, a listing, a 40-request burst, two concurrent clients and 20
aborted transfers. No pool reported a single allocation failure, and only
one was under real pressure.

tcp_pcb measured 4 of 4 and stayed there between requests: the server
closes every connection itself, so each request leaves a pcb in TIME_WAIT
for 2 x TCP_MSL, and lwIP's default MSL of 60 s means 2 minutes per
request. Nothing failed only because tcp_alloc evicts the oldest TIME_WAIT
pcb when the pool is empty. TCP_MSL drops to 10 s, as in Booster, and the
pool grows to 6 -- the 2 HTTP connections plus the debug stream plus
headroom. Re-measured: 25 s after the same burst the pool is back to 1 of
6, where before it sat full indefinitely.

tcp_seg peaked at 9 of 16 and pbuf_pool at 2 of 12, so both keep their
size; each now carries the measurement that justifies it. pbuf_pool is the
receive path -- the cyw43 driver allocates from it and in poll mode frees
each pbuf before taking the next -- and a pool that runs dry there drops
packets off the air, so its headroom stays.

The two extra pcbs cost 312 bytes of .bss, which comes off the heap
ceiling: total 46,224, minimum free 24,472, stack high water 2,544 of
16,384. Inside EPIC-12's budget. smoke.py passes 7/7 on release.
Audited every raw-TCP callback against the pinned lwIP (2.2.1, not the 2.1
this story assumed), reading core/tcp.c, tcp_out.c and tcp_in.c rather than
relying on the API docs.

tcp_recved() could run on a closed pcb. Both streaming branches of
srv_recv_cb called a response writer -- which closes the connection on a
FatFs error or a Runner chunk timeout -- and then accounted for the segment.
tcp_close() on a pcb whose window has not been returned sends an RST, purges
the pcb and unlinks it from the active list before returning ERR_OK, so that
tcp_recved() was touching a pcb already queued to be freed. It is now done
once at the top of srv_recv_cb, before anything can close. The adv-load
branch also continued into adv_load_finish_ok() after a failed dispatch, which
re-fired the failed chunk and wrote a second response onto a connection that
had already sent one; it now honours the drain's return value.

send_buffered treated ERR_MEM as fatal and dropped the response. lwIP queues
nothing on ERR_MEM and expects the application to wait for acks and try
again, so a client lost a reply it was still waiting for. The write is split
out as send_buffered_try: ERR_MEM leaves resp_queued false, srv_poll_cb
retries every 4 s, and the idle sweep gives up after 20 s.

stream_send_chunk wrote a chunk as three tcp_writes with no send-buffer
check on the debug-log path, so a failure after the first put a chunk-size
line on the wire with no body behind it. It now refuses to start a chunk
the send buffer cannot hold whole.

Two contract violations that are unreachable today but wrong next to correct
code: srv_recv_cb's NULL-arg path freed the pbuf and returned ERR_VAL, which
makes lwIP keep and re-deliver that pbuf (it now aborts and returns
ERR_ABRT), and conn_close's tcp_abort fallback left the callback returning
ERR_OK after an abort (it now reports through conn_takeAborted()).

An empty-body download set HC_DRAINING without stream_body_done, so
srv_sent_cb never closed it and a zero-length GET held one of the two
connection slots for 20 s. It completes in 19 ms now.

Verified on hardware, both builds: 4 MB round trips cmp identical, error
responses unchanged, 15 listings against a concurrent download, 20 aborted
transfers with no pool errors, smoke.py 8/8 debug and 7/7 release.
stream_download_drive returns when tcp_write reports ERR_MEM and waits for
a sent callback to pump the rest. That callback only arrives when the peer
acknowledges something, so a writer that stopped with nothing
unacknowledged -- ERR_MEM from an exhausted pool rather than a full window
-- waited for an event that could not come, and the transfer sat there
until the idle sweep closed it.

srv_poll_cb now drives HC_STREAM_DOWNLOAD and HC_STREAM_LISTING as well as
retrying a pending buffered response, then falls through to the idle check
so a peer that really has gone away is still swept.

The other two halves of this story landed with STORY-03: stream_send_chunk
will not start a chunk the send buffer cannot hold whole, and send_buffered
retries instead of dropping the response.

Tested with EPIC-11's heap-hold hook at 5,112 bytes free, where the heap
minimum reached 480 bytes and lwIP reported 1,595 failed allocations: error
responses, listings, and 256 KB and 4 MB downloads all completed intact, a
4 MB upload returned 201, and service recovered when the hold was released.
The same run on a458ac2 behaves the same, so that scenario does not
discriminate between the builds; these are contract fixes read out of the
lwIP source, not a reproduced failure.

What does exercise the path is a slow reader: 256 KB to a client reading at
5 KB/s now arrives complete with a matching sha256 in 53 s, with the heap
held at 5 KB free as well. In EPIC-09 that client was cut off after 8.4 s
and 39,420 bytes.
FF_FS_LOCK counts open files and open non-root directories, and it is
shared by everything on the card. The worst case this firmware can reach is
28: GEMDRIVE's 8 open files, its 16 searches (each DTA slot keeps a DIR open
between Fsfirst and Fsnext), and 2 HTTP connections holding a FIL and a DIR
each. The table was 8, so the ST on its own could exhaust it and starve any
HTTP transfer, or the reverse.

Sized to 28, with the arithmetic written next to it. Each entry is a 16-byte
FILESEM, so the heap ceiling drops by 320 bytes, from 46,224 to 45,904 --
measured, and matching the prediction. Capping searches was the alternative,
rejected because the cap would be enforced against a limit the ST cannot see.

FR_TOO_MANY_OPEN_FILES was previously reported as whatever each handler's
generic failure happened to be: Fopen said "file not found", Fcreate and
Fsfirst said "path not found", and the HTTP server called it a disk error.
It is now ENHNDL (-35) on all three GEMDOS paths and 503
too_many_open_files over HTTP, next to the existing 503
insufficient_memory. Both are retry-later conditions rather than faults.

The HTTP half is verified on hardware: a 4 MB download with 10 listings
running against it, all 200 and cmp identical, smoke.py 8/8. The ST half --
several desktop windows open during a copy and a transfer -- still needs the
Atari.
Six HTTP handlers block inside srv_recv_cb while the ST answers -- 10 s for
a Runner load, 5 s for an unload, 1 s for an advanced-load chunk and for
each of the three meminfo waits. Nothing else runs meanwhile, so for the
length of the wait the SELECT button did nothing and queued console bytes
stayed in the ring.

All six had the same three lines, so they now share one http_spinTick()
that also polls SELECT and drains USB CDC. The watchdog feed and the
HTTP_WAIT phase were already there from EPIC-09 STORY-06.

smoke.py 8/8 on debug. The check that actually enters a spin-wait -- SELECT
pressed during a 10 s runner load -- needs the ST in Runner mode and is
still open.
http_server_deinit closes every live connection, but it is not running
inside an lwIP callback, so there is nobody to return ERR_ABRT. Leaving the
flag set would make the next callback report an abort that never happened.
…TORY-08)

emul_resetRunnerSession() runs on the Runner's HELLO, which means the ST has
cold booted and owns nothing. It cleared busy, cwd and the errnos but left
runnerPendingBasepage set, so the RP stayed certain a program was still
loaded across an ST reboot: every later load answered 409
program_already_loaded, and the unload that would have cleared it answered
422, because the ST rightly refuses to Mfree a basepage from a previous life
(mfree_result=-40, EIMBA). Nothing the API offered could recover it -- one
crashed program poisoned the Runner until the RP itself was rebooted.

The session reset now clears the basepage mirror and the load errno with
everything else. Verified on hardware: after an ST reset the API reports
loaded_basepage null and loads work again.

Found while soaking the Runner for STORY-08. The soak itself is still
blocked by a separate defect that is not ours: exec takes the Atari down
about one time in three, usually cold-booting it (which the RP now recovers
from by itself) and occasionally wedging it until a physical reset. Written
up in the story.
…ssure

HTTP server and lwIP under pressure (EPIC-13)
After a link loss the device stayed off the network until someone reset it:
the link and status callbacks reset our state but nothing ever reconnected.

A non-blocking supervisor now runs in the main loop with two detectors.

The first is the ordinary one: cyw43_tcpip_link_status() reporting anything
but CYW43_LINK_UP, which is what lwIP is told when the link genuinely drops.

The second exists because the failure this story was written for is
invisible to every status call. The plan assumed the driver's join state
would show a lost association. It does not: cyw43_ctrl.c:434 collapses the
state to WIFI_JOIN_STATE_ACTIVE the moment a join completes, so a healthy
station reads 0x0001 -- exactly the value captured over SWD on a device that
had silently dropped off the network. The first implementation followed that
wrong reading and the hardware caught it: watching the join bits declared a
working link dead and rejoined it in a loop while HTTP answered 200
throughout. So the supervisor asks the network instead. Every 60 s it
flushes the interface's ARP cache and sends an ARP request for the default
gateway; no reply within 3 s is a failure, retried after 5 s, and three in a
row mean the network is gone however healthy the link claims to be.

On either signal it arms an asynchronous join after a 5 s grace period, with
backoff doubling 5 s to 60 s and holding there, BADAUTH pinned to the
ceiling and shown on the menu. network_wifiStaConnect is split into arming
and waiting so the 30 s wait loop can never run inside the main loop, and
the watchdog is fed across the rejoin.

Verified with two debug-only injections, since neither failure can be
produced on demand: a real disassociation is detected at once and recovers 9
s later, and a gateway made unanswerable while the association stays healthy
is detected in 52 s and recovers 4 s after that. The genuine silent failure
has only ever appeared spontaneously and remains unreproducible, so the
probe is built to its signature but unproven against the real thing.
The configured power mode never reached the radio. network_wifiInit applied
PARAM_WIFI_POWER, then every connect called network_resetStaInterface, whose
disable clears the only interface bit -- so the matching enable runs
cyw43_wifi_set_up with itf_state 0, and that path applies CYW43_DEFAULT_PM
unconditionally (cyw43_ctrl.c:558-561). That is CYW43_PERFORMANCE_PM, a PM2
power-save mode, so every session ran in power save whatever the setting
said.

PARAM_WIFI_POWER is now ignored outright, with the reason written where it
used to be read: it never arrived, and no value of it is wanted on a
mains-powered board answering an Atari blocked on the cartridge bus.
network_applyNoPowerSave() sets CYW43_NONE_PM and reads the mode back, and
runs after every STA bring-up -- each boot attempt and each STORY-01 rejoin
alike, since both go through the shared connect path.

The health report gained a wifi object: link_up, power_save, rejoins and
unreachable. power_save is read from the radio on every request rather than
cached at bring-up; the first version cached it and duly reported "off"
while the radio was in PM2.

Measured with a debug hook that puts PM2 back, everything else equal: ping
over 20 packets goes from 26.9 ms average and 5% loss to 99.8 ms and 15%,
and a 4 MB download from 8.82 s to 9.62 s.
The story named two prints in network.c. There were three, and a fourth leak
that mattered more than any of them: settings_print() dumps every key and
value to the console, and gconfig_init, aconfig_init and main all call it at
boot, so every boot printed WIFI_PASSWORD (STR) with the password in clear.
That is the line that ends up in pasted logs.

All three settings_print callers pass NULL, which means the debug console
and nothing else, so a value whose key contains PASSWORD now prints as <set>
or <none> there. A key-value store arguably should not know what a password
is, but a store that prints its contents into a log is the right place to
stop printing secrets.

The network.c prints -- the AP-mode password, "The password is: %s", and the
SSID/password/auth line before the async join -- report <set> or <none> too.

Verified on the device: boots before the change print the password three
times, the boot after prints <set> in all of them and the secret appears
nowhere in the log.
The static branch read the address, netmask and gateway as
settings_find_entry(...)->value with no NULL check, so a global config
missing any of the three faulted before the setup menu came up -- and the
menu is where such a setting gets fixed. The values then went through
ipaddr_addr(), which accepts "10" and "1.2" as addresses, and dhcp_stop()
ran before any of it was parsed, so a half-valid configuration left the
interface with neither DHCP nor a usable address.

Everything is now validated before the interface is touched: four decimal
octets and nothing else, no unusable address, a netmask that is a contiguous
run of ones, and a non-zero gateway inside the subnet. Any failure logs a
reason, leaves DHCP running, and records the reason for the menu, which
renders it on the IP line. dhcp_stop() and netif_set_addr() run only once
everything has passed.

Verified with a debug hook that overrides the settings in memory only
(settings_put_* does not write flash), so the real configuration was never
rewritten: an invalid address, an empty address, a non-contiguous netmask
and an off-subnet gateway each fall back to DHCP with their own message and
the API stays reachable, and a fully valid configuration is still applied
and answers on the configured address.
select_configure() ran after the network was brought up, so for the whole
boot connect -- up to three attempts of 30 s -- the SELECT button did
nothing. The one situation where a factory reset is wanted, a device that
cannot reach its access point, was exactly the situation where the button
was dead. It now runs before the network.

emul_pollTick() also polls select_checkPushReset(). That callback stands in
for the main loop while a connect runs and already serviced the ST, USB and
the terminal, but not the button. And the connect loop's wait drops from 2 s
to 10 ms per turn, matching the main loop: at 2 s a 300 ms press could land
entirely between polls, and the debounce needs several polls across 30 ms
windows to accept one at all.

STORY-01's supervisor is unaffected -- it arms an asynchronous join and
never enters this loop.

Verified with a debug hook that reproduces the boot connect (a nonexistent
SSID staged in memory, the same blocking connect, the same polling
callback): a press 8 s into the failing attempt was detected and reset the
device, and the reboot restored the real SSID from flash.
md-devops never writes the global config -- every settings_save in the tree
passes aconfig_getContext(), never gconfig_getContext() -- so these defaults
cannot be persisted over Booster's values. They are visible only for a key
missing from flash, or on a blank global config, and there they should read
the same as Booster.

Three were stale. The apps catalog URL pointed at the old host over plain
HTTP, the boot feature said CONFIGURATOR where Booster says FABRIC, and the
Wi-Fi power default differed. The last is aligned for consistency only,
since md-devops ignores that setting outright after STORY-02; a value that
matches Booster and is never read is less confusing than one that differs
and is also never read.

All 19 keys now match booster/src/gconfig.c at v2.4.2, checked by diffing
the parsed tables. On the device the stored values still win, and with
WIFI_POWER at 4 in flash the health report still reports power saving off.
Neither had a caller anywhere in rp/src, and the linker had already proved
they cost nothing: network_scan, network_getFoundNetworks and
select_coreWaitPush are all absent from the built ELF, garbage collected by
--gc-sections along with the ~6 KB scan table. The reasons for removing them
are about the source, not the image.

network_scan and friends: Wi-Fi scanning and configuration belong to
Booster, which this app only reads settings from, so there is no future in
which this is called -- and its scan callbacks were GCC nested functions,
which clang cannot parse, so the dead code degraded clang-tidy for the whole
file.

select_coreWaitPush, its disable and the orphaned blocking select_waitPush:
it launches core 1, and EPIC-07 STORY-01 already backed core 1 out because
running it froze the Wi-Fi poll loop the main loop depends on. Anyone
reaching for it would also need Booster's core-1 flash lockout. A ready-made
entry point into a known hazard is worse than no entry point. SELECT is
polled in the foreground by select_checkPushReset, which EPIC-13 STORY-07
and EPIC-14 STORY-04 build on.

A comment stands where each used to be. pico/multicore.h is gone from the
SELECT path, SELECT still resets the device, and smoke.py passes 8/8.
Stages the real SSID with a wrong password in memory, so the bad-key path
can be exercised without touching the access point. It does fail and hand
over to the supervisor, but this access point answers LINK JOIN and then
times out rather than reporting BADAUTH, so the branch that pins the backoff
to the ceiling and shows BAD AUTH on the menu is still unproven.

STORY-08 is closed as blocked: the remaining scenarios need the access point
taken down, which is not available here.
The card was mounted once at boot, so pulling it and putting it back left
GEMDRIVE and the file API dead until someone reset the device. Two separate
things kept it that way.

The driver never learns the card left. sd_card_spi_init() returns
immediately unless STA_NOINIT is set, and with card-detect disabled on this
board nothing ever sets it when a card is pulled -- so on reinsertion the
driver still believed the old card was initialised, skipped its whole init
sequence, and handed FatFs a freshly powered card it had never spoken to.
sd_card_p->deinit() sets the bit back. And while the volume stays
registered, mount_volume() never calls disk_initialize() at all, so
f_mount(NULL, "0:", 0) has to come first. Both are needed: the first version
of this fix had only the unregister, and every retry still failed with
FR_DISK_ERR until the deinit was added.

Detecting the removal cannot be left to FatFs results either. With the card
physically out, volume answered 200 from the cached free space and a listing
answered 200 with an empty directory -- FatFs serves what it cached and
reports no error, so no failure ever arrives. The main loop now asks the
card directly every 2 s by reading sector 0 through disk_read(), which the
cache cannot answer. FatFs medium errors still mark the card gone, from the
HTTP path and all 15 GEMDRIVE call sites, but the probe is what catches a
pull.

Verified with the card pulled and reinserted: detected in about 2 s,
remounted about 4 s after reinsertion, volume and listings correct again, a
256 KB upload round-tripping identical, no reset and crash count 0. A device
booted with no card also retries every 2 s and picks one up when it appears.
…15 STORY-02)

With no card the device was silently broken: the setup menu said nothing,
[G] and [U] launched a mode with no drive behind it, and the API answered
six different ways -- including two that lied. /volume answered 200 with the
cached size and free space, and a listing answered 200 with an empty
directory, because FatFs serves what it cached.

Every endpoint that needs the card now answers 503 no_sd_card: volume,
listings, downloads, uploads, folder operations, runner load and runner run,
through one guard in route() and one write_no_card(). write_fs_error()
reports a missing card as such rather than as a disk fault. The CLI needed
no change, since it maps by HTTP status and 503 already exits 6.

The menu header carries the state -- GEMDRIVE   SD: mounted / SD: NO CARD --
and emul_canLaunch() refuses [G] and [U] while no card is mounted, with
"Insert a working microSD card: GEMDRIVE needs one." on the status line.
STORY-01's retry loop is what clears it: inserting a card lifts the block
within about two seconds, with no reset.

The state went on the header line rather than a row of its own because a new
row breaks the menu: dividers are drawn at fixed pixel rows and the status
icons are right-aligned per header row, so one extra line struck Screenmem
and Hook through, collided IP address with the USB CDC heading, and ran the
status line into Select an option. The framebuffer snapshot caught it.
The device dropping a connection is routine -- a watchdog reboot, a panic or
a SELECT press all cut the socket -- and the user got a Python traceback
ending in ConnectionResetError.

urllib wraps some connection failures in URLError, which the CLI already
handled in six places, and lets the rest through as bare OSError subclasses.
Two exposures: urlopen() itself, and the reads after a successful open --
the download loop and the debug tail loop both call resp.read() in a loop
with no handler at all, which is exactly where a reboot mid-transfer lands.

CONNECTION_DROPPED names the family once, including IncompleteRead, which a
chunked body cut short raises instead of an OSError. request_json re-raises
it as URLError so the six existing handlers keep working and each still
names its own operation; the two streaming loops report through _dropped().
A dropped download leaves the partial file in place, which is what --resume
is for.

DroppedConnectionTests adds a socket server that resets with SO_LINGER 0, so
the client sees a real ECONNRESET rather than a clean close, and covers a
reset before any reply, on a JSON command, mid-download and on health. Each
asserts exit 8, no traceback, and exactly one error: line. Run against the
unfixed CLI all four error with the traceback the story was filed for, so
they fail for the right reason. Suite: 119 tests, all passing.
After a timeout the ST re-sends the same 1 KB chunk with a fresh token, and
the RP treated it as new data and appended it again: the file gained a
duplicated chunk and lost its tail. Neither side carried anything that could
tell the two apart -- a retried chunk and a new one are identical on the
wire.

An absolute offset does not work without more state than it appears: the
m68k knows only how far it has got within the current Fwrite call, not the
file position, which the RP owns and Fseek moves. A relative offset cannot
distinguish a retry of a call's first chunk from the first chunk of the next
call, since both are zero.

So the m68k carries a number that never repeats. write_seq is a longword in
the GEMDRIVE blob, bumped once per chunk and never per retry, never reset,
sent in d5 -- which send_sync_write_command_to_sidecart always transmits and
the RP was skipping. Each open file slot on the RP remembers the last
sequence it accepted and the bytes it wrote; a chunk whose sequence matches
is answered with that same count and not written again. The memo is stored
before the answer is sent, because the answer is what goes missing.

Verified on hardware with a debug hook that stalls the answer after the data
is committed. Copying a 512 KB file with two chunks stalled 10 s each, the
ST re-sent both, the RP reported "repeat of chunk N, answering 1024 again"
for each, and the copy is byte-identical to the source. A first attempt at
2 s stalls proved nothing -- no retry happened at all, because that is
inside the ST's COMMAND_TIMEOUT, and the matching checksum would have been a
false pass.

The cartridge image is now 10,144 of 10,240 bytes.
CLAUDE.md carried a memory map from before EPIC-12 -- RAM 128 K with the
cartridge window at 0x20020000, where it is now 192 K and 0x20030000 -- and
still said v1.1 "is making Release work". It now records that v1.1 ships
Release and passed the gate, the stack at the top of RAM behind an MPU
guard, FF_FS_LOCK at 28 and why, the automatic SD remount and the no-card
answers, the lwIP pools carrying their measurements, and the Wi-Fi
supervisor with the gateway probe.

The README gains the two things a user meets: what happens with no SD card
(the menu says so, [G] and [U] refuse, endpoints answer 503 no_sd_card, and
inserting one clears it without a reset) and what happens when the network
goes away (it rejoins by itself; the radio runs at full power).

The CHANGELOG entry is user-facing and leads with the upload path, which
could not complete a transfer beyond about 200 KB, then the aborted-transfer
wedge, the SD card, the duplicated write chunk, Wi-Fi, crash visibility and
the CLI. It ends with the two known limitations: heavy tracing through the
debug ABI can take the ST down (EPIC-18, cause not found), and Wi-Fi
recovery is verified against injected faults rather than a real outage.
…Y-04)

The health sample in docs/api.md no longer matched what the firmware sends:
it was missing the wifi object added in EPIC-14. Documented with the point
that matters -- power_save is read back from the radio on every request
rather than being the value we asked for, so 16 (CYW43_NONE_PM) is evidence
and not an intention.

too_many_open_files was also undocumented. It is described with the arithmetic
behind FF_FS_LOCK's 28 entries and the fact that it usually means the ST is
holding its share, so it is worth retrying rather than a fault.
The entry read like the epic notes it came from: 437 lost bytes, an idle
sweeper, a poll timer, a watchdog, MinSizeRel. None of that is what someone
downloading the firmware wants to know.

It now says what changed for them -- large files copy, a dead transfer stops
blocking the next one, a reinserted card is picked up, files written from the
Atari arrive whole, Wi-Fi is quicker and rejoins by itself -- and keeps the
measurements only where they mean something to a user, such as the response
time and packet loss from the radio no longer sleeping. The build change
moved to a short note at the end, where a developer will still find it.
Comments and docs pointed at epics, stories and files under docs/epics/ --
"(EPIC-14 STORY-02)", "see docs/epics/02-http-api.md" -- but that directory is
gitignored. Anyone reading the code was being sent to notes they do not have,
to look up an identifier that appears nowhere in the repository.

Every reference is replaced by what it was standing in for. Where the
sentence already explained the reason, the tag simply goes; where the tag was
carrying the meaning, the meaning is now written out: the scan helpers were
"removed in v1.1", the stack rule is stated rather than cited, the smoke
script's limits name the bugs they follow.

CLAUDE.md gains the rule, with the grep to run before tagging a release.
The main-loop cadence comment, the menu overlay comment and debugcap.h each
sent the reader to a "backlog entry" or a "Tier 2 backlog" that lives only in
the untracked planning notes. Each now states the thing itself: moving
chandler and the USB drain to core 1 is the step to take if 100 Hz stops
being enough, framed menu sections would need the terminal renderer to know
about them, and a consumer moved to core 1 would need explicit barriers.
The health examples in README.md and docs/api.md were taken from a build
before the RAM work: 2 KB of stack reserved, a 118 KB heap, a stack pointer
inside SCRATCH_Y. They now show what this firmware reports -- 16 KB reserved
with a 3.5 KB peak, a 49,852-byte heap, code_in_ram 33,512 -- read off the
linker symbols and the soak run.

Two field descriptions were wrong rather than stale. heap.total is capped by
the bottom of the core-0 stack, not by the cartridge window. stack.high_water
can no longer quietly exceed reserved and walk into the core-1 stack area:
the MPU guard faults and reboots instead.

Undocumented until now: the two-connection limit and the 20 s idle timeout,
which is what frees the transfer lock when a client disappears; 503 also
covers too_many_open_files; and a device powered on with no card refuses the
auto-launch countdown and waits in the menu. The README highlights say the
firmware looks after itself, since that is most of what v1.1 is.

programming.md still described the template's flash and scratch layout:
CONFIG_FLASH is 120 KB of per-app sectors with the lookup and global sectors
above it, and SCRATCH_Y is unused now that core 0's stack sits at the top of
RAM.
@diegoparrilla
diegoparrilla merged commit 1b5788d into main Sep 17, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant