v1.1.0 - #18
Merged
Merged
v1.1.0#18
Conversation
Bump the three tracked version files to v1.1.0. This branch collects the v1.1.0 epics through pull requests and merges into main only when the version is complete.
…types (EPIC-08 STORY-01 STORY-02) A failed step no longer gets stepped over. build.sh, rp/build.sh and target/atarist/build.sh stop on the first error. Before this, an m68k failure let the RP build embed the committed target_firmware.h, and a C or link failure let build.sh copy an older UF2 from rp/dist and write a JSON announcing the new version. - rp/build.sh empties rp/dist before building. - target/atarist/build.sh deletes any previous target_firmware.h and checks that firmware.py produced a new one, using python3 when the host has no python. - build.sh falls back to md5 -q when md5sum is missing, as on macOS. The build types are now what their names say, case-insensitively: release is CMake Release with DEBUG_MODE=0, and debug is the same build with DEBUG_MODE=1 for DPRINTF traces on the UART console. RP_CMAKE_BUILD_TYPE overrides only the CMake build type, to compare against the MinSizeRel configuration that shipped until now. A caller-provided RELEASE_DATE is kept, so two builds of one commit can be compared byte for byte. FF_FS_TINY no longer follows _DEBUG, so a debug build keeps the same FatFs objects as a release build. With the fail-fast in place the first run stopped at a stale .git/modules/pico-sdk/index.lock from April. The old script had been skipping the pico-sdk pin silently ever since.
…RY-03 STORY-05) --strip-all removed the symbol table and debug information from rp.elf without changing a byte of the flashed image, so crash addresses could not be decoded and a debug probe could not break at main. The link keeps --gc-sections only. Also removed, all without effect on the image: - the clang-tidy block, which set CMAKE_C_CLANG_TIDY after every target existed and so never ran; - the settings library: settings.c was compiled into it and a second time directly into rp, and only the direct copy was linked; - pico_set_binary_type(rp), which selects nothing because the custom linker script takes precedence. The stdio comment now says what the code does: USB stdio stays for usbcdc.c, and only debug builds enable UART stdio. Verified with RP_CMAKE_BUILD_TYPE=MinSizeRel and a fixed RELEASE_DATE: rp.uf2 is byte-identical to the same build before these changes.
…Y-03) CMake Tools offered its stock Debug and Release variants, which built with DEBUG_MODE and APP_UUID_KEY unset: no traces, the default UUID, and a CMake type the scripts never use. .vscode/cmake-variants.yaml now defines the same two build types as ./build.sh, both CMake Release, with DEBUG_MODE and the development UUID set at configure time. The debug launch configuration still carried the template's name.
…-04) Both workflows installed Ubuntu's packaged gcc-arm-none-eabi, so the released UF2 came from a different compiler than the ARM GNU Toolchain 14.2 used locally. They now download and cache 14.2.rel1 and print the compiler version. - build.yml builds release and debug in a matrix, writes a size report (sections, code copied to RAM, heap, reserved stack, flash used) to the job summary, and uploads the firmware with rp.elf and its map. - release.yml builds release in lower case, writes the same report, and attaches rp.elf and its map to the GitHub release. They are not uploaded to the S3 bucket. - actions/checkout v4, actions/setup-python v5, and GITHUB_OUTPUT instead of the deprecated set-output command.
…ORY-04) The build section still said every build is MinSizeRel and that clang-tidy runs from CMake. It now describes the release and debug types, the MinSizeRel override, reproducible RELEASE_DATE, symbols in rp.elf, the VS Code variants and the CI matrix.
The stcmd that atarist-toolkit-docker's installer writes, which CI uses, takes $VERSION as the Docker image tag. build.sh exported the firmware version under that name, so CI asked for atarist-toolkit-docker:v1.1.0, which does not exist, and the m68k build failed. The previous build.sh exported it too and ignored a failed m68k build, so earlier CI and release builds most likely embedded the committed target_firmware.h; the logs of those runs have expired. VERSION is only used inside build.sh.
EPIC-08: honest builds
The statistics were keyed on NDEBUG, which both build types define, so MEM_STATS, MEMP_STATS and stats_display were compiled out everywhere. Port Booster's block: a debug build with DEVOPS_LWIP_STATS=1 in the environment turns them on. Release is unchanged; its lwIP objects are byte-identical to before.
…ORY-05) commemul_poll derived the write index from the DMA transfer count with no lap detection, so a main loop stalled for more than 4,096 samples silently mixed old and new data into the protocol parser. Count a lap when that many samples arrived since the last poll, and drop the unread ones. The channel stopped for good after 2^32 samples; re-arm it at 2^31 from its live write address. A corrupted frame could also carry a payload size larger than the payload buffer, and the parser wrote past it. Reject such a frame.
…(EPIC-09 STORY-01 STORY-04 STORY-06) A panic or HardFault used to freeze the RP silently: the SDK's panic stops at a breakpoint and fatfs-sdk's HardFault handler does the same, which locks up an RP2040 with no debugger. A hang had no recovery at all. health.c adds: - panic() through PICO_PANIC_FUNCTION and a HardFault handler installed in the RAM vector table at boot. Both write the reason, PC, LR and SP to watchdog scratch registers 0-3 and reboot after 100 ms. - An 8 s watchdog, paused under a debug probe, fed from the main loop, emul_pollTick, the Wi-Fi connect loop, the HTTP spin-waits and after the two slow boot steps. A phase ID in scratch[0] says where a hang happened. - Boot decoding into power-on, reset, panic, HardFault, hang or reboot, shown on row 2 of the setup menu and on the debug console. - A crash-loop guard: 3 crash reboots within 60 s stop the countdown until a SELECT reset or a power cycle. - A painted core-0 stack for a high-water mark, heap free and minimum free from mallinfo and sbrk, and code copied to RAM. - Debug-only fault injection for hardware verification. sd_timeouts.c overrides fatfs-sdk's weak timeout table so a failing card gives up well within the watchdog. The jump to Booster stops the watchdog first, since it does not reset the chip.
A release build has no console, so nothing on the device could be read without a debugger. GET /api/v1/system/health returns uptime, heap total, free, minimum free and sbrk high-water, stack high-water and reservation, code in RAM, the last reset reason with its phase or PC, the crash count and crash-loop state, ROM3 overruns and debug-byte drops. A debug build with DEVOPS_LWIP_STATS adds lwIP's pool counters. It is a new endpoint rather than more ping fields, so ping stays a minimal liveness check. `sidecart.py health` prints it, with tests. Debug builds also expose POST /api/v1/debug/test/<fault> to trigger each recovery path on hardware. Both are documented in docs/api.md.
…d (EPIC-09 STORY-06) With a mounted SD card pulled, sd_read_blocks returned early from its error path without releasing the card and SPI mutexes. Those mutexes block without a timeout, so the next disk call waited forever: the first GEMDRIVE or HTTP request after the pull failed, and the second froze the RP until the watchdog rebooted it (reported as a hang in disk_read by a diagnostic build). Upstream fixed it in f5f31b3. 6c644cf is that fix plus a README change, and is the commit Booster already pins. Verified on hardware: after an idle pull, listings and downloads fail in under a second and the RP stays up.
Flashing or resetting through a debug probe uses SYSRESETREQ, which leaves the watchdog reason and scratch registers exactly as a real watchdog expiry does, so every probe flash showed "Recovered: hang" and counted as a crash. Measured on hardware: reason, watchdog control, chip reset and the SDK's marker all match, and DHCSR reads as zero from the firmware. A once-a-second timer interrupt now writes a stall mark into scratch[1] after 5 s without a watchdog feed, and clears it when feeding resumes. A watchdog reset counts as a hang only if the mark is there. A hang with interrupts disabled cannot set the mark and is reported as a reboot. Verified on hardware: main-loop hang and HTTP handler hang still report their phase; a probe reset and a probe flash now report a reboot with crash count 0.
…d-recovery EPIC-09: device-side visibility and recovery
…10 STORY-02 STORY-03) The image reaches RAM through the XIP stream, whose address register has no bits 1:0. target_firmware is a uint16_t array, only 2-byte aligned, so at an address that is 2 mod 4 every word would be shifted and the ST would boot without the cartridge. It has been aligned by luck. firmware.py now emits aligned(4), and COPY_FIRMWARE_TO_RAM falls back to a word copy for any source that is not 4-byte aligned. The stream also moves 32-bit words, so an odd word count lost the last 16-bit word. It is now copied separately. The 64 KB window is cleared before the copy, so nothing from a previous run (a jump from Booster, a crash reboot) survives past the end of the image. COPY_FIRMWARE_TO_RAM_MEMCPY took bytes while every other copy macro takes words; it takes words now. Debug builds check the window against target_firmware after the copy and count non-zero bytes past the image. target_firmware.h is regenerated by the m68k build; only the declaration and the header's build date change.
…TORY-01) __rom_in_ram_start__ was declared as `extern unsigned int`, so to the compiler it is a 4-byte object and every read or write past those bytes is undefined behaviour it may optimise on at -O3. Clearing the window with memset made GCC warn that 65,536 bytes overflow a 4-byte region. Declare it as an array of unknown size in constants.h and in the two files that redeclare it. Every use takes its address and casts it, so no call site changes and the code size is identical.
[G] and [U] make the m68k install GEMDRIVE at the relocation address and memtop the RP publishes when the ST sends HELLO at cold boot. After the RP reboots while the ST keeps running (a crash, a SELECT reset, a flash) no HELLO arrives. Recovery only worked because the previous run's shared variables survived in RAM; with the window now cleared, the address reads 0 and the ST halts with "Reloc/stack overlap". The launch commands now refuse until HELLO has landed since the RP booted, keep the full menu on screen and say "Reset the Atari ST first" on the bottom strip. When HELLO lands the message clears, and if the refused launch came from the countdown, the countdown starts again. Keystrokes the ST sends while it boots are dropped at that point, or the main loop would halt the restarted countdown before its first tick. Verified on hardware (release): a launch after an RP reboot is refused from the countdown and from [U] with the menu intact; after an ST reset the countdown restarts, runs to 0 and Runner launches without the banner.
DPRINTF blocks on the UART. At the SDK's default 115,200 baud a debug build printed its flash layout and both settings dumps before copying the cartridge image, and a power-cycled ST booted into GEM because the cartridge was not live yet; release, which prints nothing, showed the menu. At 921,600 the same output is eight times faster and the power-cycled ST shows the DevOps menu on the debug build too. The UART clock stays at 48 MHz after the 225 MHz overclock, so the divider gives 921,413 baud, a 0.02 % error. Only debug builds get the define; release has no UART console. Set the serial terminal to 921600.
…RY-04) DPRINTF waited until every byte had left the UART, so each trace stalled the main loop, the ROM3 command ring and the ST's commands for about 1 ms at 921,600 baud, and bursts such as GEMDRIVE's per-call traces for much longer. DEBUG_BUFFERED_CONSOLE in debug.h, default 1, toggles a buffered console. DPRINTF formats with fprintf into a funopen stream that queues the text in a 4 KB ring buffer, and the UART transmit interrupt sends it. When the ring is full the caller waits for room, or sends bytes itself with interrupts off, so nothing is lost. Inside an exception the queue and the new text are sent at once, so crash reports stay complete and in order. stderr and stdout are redirected to the same queue, so nothing else writes to the UART behind the interrupt. Set it to 0 for the previous blocking console, which keeps the last lines before a hang. Everything lives in debug.h: the shared state and functions are weak definitions merged by the linker into one copy. debug.h replaces DPRINTF definitions made earlier, and settings.c includes it, so the settings library's own fallback DPRINTF no longer bypasses the queue. Debug builds only: 4 KB less heap and about 500 bytes of flash. Release is unchanged. Verified on hardware: a power-cycled boot log is complete and in order, and a forced panic prints its report whole, in order, before the reboot.
usbcdc_init() called stdio_init_all() a second time in debug builds, assuming it was idempotent. It is not: stdio_uart_init() runs uart_init(), which resets the UART and throws away up to 32 bytes still in the console FIFO, and stdio_usb_init() claims another IRQ and background task on every call. With the buffered console this showed as exactly 32 bytes missing from a boot trace. Start USB stdio only when TinyUSB is not running yet. Release builds, where main.c does not initialise stdio, behave as before.
__flash_binary_start and the ROM temp, Booster, config, lookup and global config region symbols were declared as `extern unsigned int`, 4-byte objects to the compiler. aconfig.c reads the whole 4 KB lookup table through a pointer taken from _global_lookup_flash_start, which is undefined behaviour the optimiser may use. Declare them as arrays of unknown size, as __rom_in_ram_start__ already is. Every use takes the address and casts it, so no call site changes.
…-05) The streaming listing checked for 32 bytes of room, enough for the closing envelope, before reading the next entry. An entry needs about 80 bytes, up to 350 with a long name, so once a chunk had between 32 and 80 bytes left the entry just read did not fit: it was dropped and the listing ended with "truncated": true. Any folder with more entries than fit in one 1 KB chunk, about 15, came back cut short. Check for room for the largest possible entry before f_readdir, since an entry that has been read cannot be put back; the listing then continues in the next chunk. Found by the EPIC-10 smoke script; verified on hardware with 100 and 500 files: all entries listed, "truncated": false.
…-10 STORY-05) FOLDER defaulted to the microfirmware template's "/test", which md-devops never uses, so every boot created an empty /test folder on the card. FOLDER now defaults to "/devops" and is created at boot if it does not exist. Settings saved with the old "/test" value are changed to "/devops" and saved once at boot, then the folder is created if needed. Verified on hardware: the first boot logged the change and rewrote the app settings with every other value kept; the next power-on boot read FOLDER=/devops and mounted the card with no /test folder involved.
…iming-bugs EPIC-10: latent layout and timing bugs
tools/dev/console.py watch reads the Debug Probe's UART at 921,600 baud, timestamps each line, appends it to tools/dev/logs/console.log and shows it in the terminal. It finds the probe by its USB name, and waits for it and reopens it after an unplug. since-boot, tail, grep and wait read the log, so a boot log no longer has to be pasted by hand. Standard library only. Verified on hardware: a probe reset produced a complete boot log, and wait matched the GEMDRIVE ready line; unplugging and replugging the probe resumed the capture.
…-17 STORY-02) Every build now carries the git commit it came from: <sha7>, or <sha7>-dirty.<diff7> with uncommitted changes, where <diff7> hashes the diff. rp/src/build_id.cmake regenerates build_id.h on every build and rewrites it only when the ID changes, so incremental builds never report a stale ID and only health.c and http_server.c recompile. The same tree gives the same ID and a byte-identical binary. The ID is printed on the console at boot and returned by GET /api/v1/system/health with a new "debug" flag; sidecart.py health prints both. The health reply buffer grows by 32 bytes for them. tools/dev/flash.sh builds out of tree and incrementally, keeps the ELF by build ID, flashes with picotool or through the Debug Probe, and waits until the health report shows the new ID and build type, printing the console since boot if it does not. Verified on hardware: debug with picotool and with the probe, release with the probe, and the timeout path; about 18 s from command to a confirmed boot for an unchanged build.
…C-17 STORY-02) The developer tools must work with any microfirmware from the template and with a hung RP, so they reach it only through picotool, the Debug Probe and the console UART. tools/dev/swd.py wraps OpenOCD and reads memory while the CPU runs: running waits until VTOR holds the ELF's RAM vector table, verify compares the whole flash with the ELF, build-id reads the new release_build_id string from flash, read dumps a range, and program flashes through the probe. Runs that fail on a momentary debug-port drop are retried. build_id.cmake now also generates build_id.c with release_build_id, kept through --gc-sections. flash.sh checks running, verify and build-id over SWD instead of polling GET /api/v1/system/health, which keeps reporting the build ID for API users. Verified on hardware: debug with picotool and release with the probe, each running, verified byte for byte and carrying its build ID; a release ELF against a running debug build fails verify.
swd.py screen reads the 320x200 framebuffer at the top of the cartridge window while the RP runs, or while it is halted, and writes it as a PNG, so the setup menu can be checked without looking at the ST. swd.py shared prints the command sentinel, the random token and the shared variables, named after the *_SVAR_* indexes in rp/src/include. Neither hard-codes an address: the window comes from the ELF's __rom_in_ram_start__ and the offsets from chandler.h, whose #defines are evaluated by a small parser. Without --elf the tools pick the cached ELF whose build ID the RP carries. swd.py running now also requires core 0 not to be halted, and swd.py resume releases both cores through DHCSR: a halted core 1 pauses the RP2040's timer and watchdog and leaves core 0 asleep, and OpenOCD's own resume fails in a new OpenOCD run. Verified on hardware: the PNG matched the ST's screen before HELLO, after an ST reset (Phystop, Screenmem and the countdown line), and after a forced panic (Recovered: panic @10005194, the PC on the console). After HELLO the shared variables showed reloc and memtop 0xF4000, drive C and hook vector $70.
…Y-04) swd.py select presses the SELECT button without firmware code: it forces the pin's input high through the RP2040's GPIO input override for the hold time. A long press needs --force, because md-devops then erases the global settings. For what the firmware must do itself, devhooks.h adds a debug-only mailbox in RAM, self-contained like debug.h. Tools write a request over SWD and devhooks_poll(), called from the main loop, runs it and acknowledges it. A protocol request is queued through the new chandler_injectProtocol() as if the ST had sent it, so swd.py key and swd.py inject use the firmware's normal command path. An app request goes to a handler the app registers; emul.c handles countdown stop and restart, named by DEVHOOKS_APP_* defines that swd.py app reads. Release builds contain none of it. Verified on hardware with the menu: countdown stop and restart and a key 'z' each showed the expected bottom line in swd.py screen, the console logged the injected keystroke, and a SELECT short press was detected after 304 ms and rebooted the RP with reason reset.
With DEVOPS_LWIP_STATS on a debug build, ran a 4 MB upload, a 4 MB download, a listing, a 40-request burst, two concurrent clients and 20 aborted transfers. No pool reported a single allocation failure, and only one was under real pressure. tcp_pcb measured 4 of 4 and stayed there between requests: the server closes every connection itself, so each request leaves a pcb in TIME_WAIT for 2 x TCP_MSL, and lwIP's default MSL of 60 s means 2 minutes per request. Nothing failed only because tcp_alloc evicts the oldest TIME_WAIT pcb when the pool is empty. TCP_MSL drops to 10 s, as in Booster, and the pool grows to 6 -- the 2 HTTP connections plus the debug stream plus headroom. Re-measured: 25 s after the same burst the pool is back to 1 of 6, where before it sat full indefinitely. tcp_seg peaked at 9 of 16 and pbuf_pool at 2 of 12, so both keep their size; each now carries the measurement that justifies it. pbuf_pool is the receive path -- the cyw43 driver allocates from it and in poll mode frees each pbuf before taking the next -- and a pool that runs dry there drops packets off the air, so its headroom stays. The two extra pcbs cost 312 bytes of .bss, which comes off the heap ceiling: total 46,224, minimum free 24,472, stack high water 2,544 of 16,384. Inside EPIC-12's budget. smoke.py passes 7/7 on release.
Audited every raw-TCP callback against the pinned lwIP (2.2.1, not the 2.1 this story assumed), reading core/tcp.c, tcp_out.c and tcp_in.c rather than relying on the API docs. tcp_recved() could run on a closed pcb. Both streaming branches of srv_recv_cb called a response writer -- which closes the connection on a FatFs error or a Runner chunk timeout -- and then accounted for the segment. tcp_close() on a pcb whose window has not been returned sends an RST, purges the pcb and unlinks it from the active list before returning ERR_OK, so that tcp_recved() was touching a pcb already queued to be freed. It is now done once at the top of srv_recv_cb, before anything can close. The adv-load branch also continued into adv_load_finish_ok() after a failed dispatch, which re-fired the failed chunk and wrote a second response onto a connection that had already sent one; it now honours the drain's return value. send_buffered treated ERR_MEM as fatal and dropped the response. lwIP queues nothing on ERR_MEM and expects the application to wait for acks and try again, so a client lost a reply it was still waiting for. The write is split out as send_buffered_try: ERR_MEM leaves resp_queued false, srv_poll_cb retries every 4 s, and the idle sweep gives up after 20 s. stream_send_chunk wrote a chunk as three tcp_writes with no send-buffer check on the debug-log path, so a failure after the first put a chunk-size line on the wire with no body behind it. It now refuses to start a chunk the send buffer cannot hold whole. Two contract violations that are unreachable today but wrong next to correct code: srv_recv_cb's NULL-arg path freed the pbuf and returned ERR_VAL, which makes lwIP keep and re-deliver that pbuf (it now aborts and returns ERR_ABRT), and conn_close's tcp_abort fallback left the callback returning ERR_OK after an abort (it now reports through conn_takeAborted()). An empty-body download set HC_DRAINING without stream_body_done, so srv_sent_cb never closed it and a zero-length GET held one of the two connection slots for 20 s. It completes in 19 ms now. Verified on hardware, both builds: 4 MB round trips cmp identical, error responses unchanged, 15 listings against a concurrent download, 20 aborted transfers with no pool errors, smoke.py 8/8 debug and 7/7 release.
stream_download_drive returns when tcp_write reports ERR_MEM and waits for a sent callback to pump the rest. That callback only arrives when the peer acknowledges something, so a writer that stopped with nothing unacknowledged -- ERR_MEM from an exhausted pool rather than a full window -- waited for an event that could not come, and the transfer sat there until the idle sweep closed it. srv_poll_cb now drives HC_STREAM_DOWNLOAD and HC_STREAM_LISTING as well as retrying a pending buffered response, then falls through to the idle check so a peer that really has gone away is still swept. The other two halves of this story landed with STORY-03: stream_send_chunk will not start a chunk the send buffer cannot hold whole, and send_buffered retries instead of dropping the response. Tested with EPIC-11's heap-hold hook at 5,112 bytes free, where the heap minimum reached 480 bytes and lwIP reported 1,595 failed allocations: error responses, listings, and 256 KB and 4 MB downloads all completed intact, a 4 MB upload returned 201, and service recovered when the hold was released. The same run on a458ac2 behaves the same, so that scenario does not discriminate between the builds; these are contract fixes read out of the lwIP source, not a reproduced failure. What does exercise the path is a slow reader: 256 KB to a client reading at 5 KB/s now arrives complete with a matching sha256 in 53 s, with the heap held at 5 KB free as well. In EPIC-09 that client was cut off after 8.4 s and 39,420 bytes.
FF_FS_LOCK counts open files and open non-root directories, and it is shared by everything on the card. The worst case this firmware can reach is 28: GEMDRIVE's 8 open files, its 16 searches (each DTA slot keeps a DIR open between Fsfirst and Fsnext), and 2 HTTP connections holding a FIL and a DIR each. The table was 8, so the ST on its own could exhaust it and starve any HTTP transfer, or the reverse. Sized to 28, with the arithmetic written next to it. Each entry is a 16-byte FILESEM, so the heap ceiling drops by 320 bytes, from 46,224 to 45,904 -- measured, and matching the prediction. Capping searches was the alternative, rejected because the cap would be enforced against a limit the ST cannot see. FR_TOO_MANY_OPEN_FILES was previously reported as whatever each handler's generic failure happened to be: Fopen said "file not found", Fcreate and Fsfirst said "path not found", and the HTTP server called it a disk error. It is now ENHNDL (-35) on all three GEMDOS paths and 503 too_many_open_files over HTTP, next to the existing 503 insufficient_memory. Both are retry-later conditions rather than faults. The HTTP half is verified on hardware: a 4 MB download with 10 listings running against it, all 200 and cmp identical, smoke.py 8/8. The ST half -- several desktop windows open during a copy and a transfer -- still needs the Atari.
Six HTTP handlers block inside srv_recv_cb while the ST answers -- 10 s for a Runner load, 5 s for an unload, 1 s for an advanced-load chunk and for each of the three meminfo waits. Nothing else runs meanwhile, so for the length of the wait the SELECT button did nothing and queued console bytes stayed in the ring. All six had the same three lines, so they now share one http_spinTick() that also polls SELECT and drains USB CDC. The watchdog feed and the HTTP_WAIT phase were already there from EPIC-09 STORY-06. smoke.py 8/8 on debug. The check that actually enters a spin-wait -- SELECT pressed during a 10 s runner load -- needs the ST in Runner mode and is still open.
http_server_deinit closes every live connection, but it is not running inside an lwIP callback, so there is nobody to return ERR_ABRT. Leaving the flag set would make the next callback report an abort that never happened.
…TORY-08) emul_resetRunnerSession() runs on the Runner's HELLO, which means the ST has cold booted and owns nothing. It cleared busy, cwd and the errnos but left runnerPendingBasepage set, so the RP stayed certain a program was still loaded across an ST reboot: every later load answered 409 program_already_loaded, and the unload that would have cleared it answered 422, because the ST rightly refuses to Mfree a basepage from a previous life (mfree_result=-40, EIMBA). Nothing the API offered could recover it -- one crashed program poisoned the Runner until the RP itself was rebooted. The session reset now clears the basepage mirror and the load errno with everything else. Verified on hardware: after an ST reset the API reports loaded_basepage null and loads work again. Found while soaking the Runner for STORY-08. The soak itself is still blocked by a separate defect that is not ours: exec takes the Atari down about one time in three, usually cold-booting it (which the RP now recovers from by itself) and occasionally wedging it until a physical reset. Written up in the story.
…ssure HTTP server and lwIP under pressure (EPIC-13)
After a link loss the device stayed off the network until someone reset it: the link and status callbacks reset our state but nothing ever reconnected. A non-blocking supervisor now runs in the main loop with two detectors. The first is the ordinary one: cyw43_tcpip_link_status() reporting anything but CYW43_LINK_UP, which is what lwIP is told when the link genuinely drops. The second exists because the failure this story was written for is invisible to every status call. The plan assumed the driver's join state would show a lost association. It does not: cyw43_ctrl.c:434 collapses the state to WIFI_JOIN_STATE_ACTIVE the moment a join completes, so a healthy station reads 0x0001 -- exactly the value captured over SWD on a device that had silently dropped off the network. The first implementation followed that wrong reading and the hardware caught it: watching the join bits declared a working link dead and rejoined it in a loop while HTTP answered 200 throughout. So the supervisor asks the network instead. Every 60 s it flushes the interface's ARP cache and sends an ARP request for the default gateway; no reply within 3 s is a failure, retried after 5 s, and three in a row mean the network is gone however healthy the link claims to be. On either signal it arms an asynchronous join after a 5 s grace period, with backoff doubling 5 s to 60 s and holding there, BADAUTH pinned to the ceiling and shown on the menu. network_wifiStaConnect is split into arming and waiting so the 30 s wait loop can never run inside the main loop, and the watchdog is fed across the rejoin. Verified with two debug-only injections, since neither failure can be produced on demand: a real disassociation is detected at once and recovers 9 s later, and a gateway made unanswerable while the association stays healthy is detected in 52 s and recovers 4 s after that. The genuine silent failure has only ever appeared spontaneously and remains unreproducible, so the probe is built to its signature but unproven against the real thing.
The configured power mode never reached the radio. network_wifiInit applied PARAM_WIFI_POWER, then every connect called network_resetStaInterface, whose disable clears the only interface bit -- so the matching enable runs cyw43_wifi_set_up with itf_state 0, and that path applies CYW43_DEFAULT_PM unconditionally (cyw43_ctrl.c:558-561). That is CYW43_PERFORMANCE_PM, a PM2 power-save mode, so every session ran in power save whatever the setting said. PARAM_WIFI_POWER is now ignored outright, with the reason written where it used to be read: it never arrived, and no value of it is wanted on a mains-powered board answering an Atari blocked on the cartridge bus. network_applyNoPowerSave() sets CYW43_NONE_PM and reads the mode back, and runs after every STA bring-up -- each boot attempt and each STORY-01 rejoin alike, since both go through the shared connect path. The health report gained a wifi object: link_up, power_save, rejoins and unreachable. power_save is read from the radio on every request rather than cached at bring-up; the first version cached it and duly reported "off" while the radio was in PM2. Measured with a debug hook that puts PM2 back, everything else equal: ping over 20 packets goes from 26.9 ms average and 5% loss to 99.8 ms and 15%, and a 4 MB download from 8.82 s to 9.62 s.
The story named two prints in network.c. There were three, and a fourth leak that mattered more than any of them: settings_print() dumps every key and value to the console, and gconfig_init, aconfig_init and main all call it at boot, so every boot printed WIFI_PASSWORD (STR) with the password in clear. That is the line that ends up in pasted logs. All three settings_print callers pass NULL, which means the debug console and nothing else, so a value whose key contains PASSWORD now prints as <set> or <none> there. A key-value store arguably should not know what a password is, but a store that prints its contents into a log is the right place to stop printing secrets. The network.c prints -- the AP-mode password, "The password is: %s", and the SSID/password/auth line before the async join -- report <set> or <none> too. Verified on the device: boots before the change print the password three times, the boot after prints <set> in all of them and the secret appears nowhere in the log.
The static branch read the address, netmask and gateway as settings_find_entry(...)->value with no NULL check, so a global config missing any of the three faulted before the setup menu came up -- and the menu is where such a setting gets fixed. The values then went through ipaddr_addr(), which accepts "10" and "1.2" as addresses, and dhcp_stop() ran before any of it was parsed, so a half-valid configuration left the interface with neither DHCP nor a usable address. Everything is now validated before the interface is touched: four decimal octets and nothing else, no unusable address, a netmask that is a contiguous run of ones, and a non-zero gateway inside the subnet. Any failure logs a reason, leaves DHCP running, and records the reason for the menu, which renders it on the IP line. dhcp_stop() and netif_set_addr() run only once everything has passed. Verified with a debug hook that overrides the settings in memory only (settings_put_* does not write flash), so the real configuration was never rewritten: an invalid address, an empty address, a non-contiguous netmask and an off-subnet gateway each fall back to DHCP with their own message and the API stays reachable, and a fully valid configuration is still applied and answers on the configured address.
select_configure() ran after the network was brought up, so for the whole boot connect -- up to three attempts of 30 s -- the SELECT button did nothing. The one situation where a factory reset is wanted, a device that cannot reach its access point, was exactly the situation where the button was dead. It now runs before the network. emul_pollTick() also polls select_checkPushReset(). That callback stands in for the main loop while a connect runs and already serviced the ST, USB and the terminal, but not the button. And the connect loop's wait drops from 2 s to 10 ms per turn, matching the main loop: at 2 s a 300 ms press could land entirely between polls, and the debounce needs several polls across 30 ms windows to accept one at all. STORY-01's supervisor is unaffected -- it arms an asynchronous join and never enters this loop. Verified with a debug hook that reproduces the boot connect (a nonexistent SSID staged in memory, the same blocking connect, the same polling callback): a press 8 s into the failing attempt was detected and reset the device, and the reboot restored the real SSID from flash.
md-devops never writes the global config -- every settings_save in the tree passes aconfig_getContext(), never gconfig_getContext() -- so these defaults cannot be persisted over Booster's values. They are visible only for a key missing from flash, or on a blank global config, and there they should read the same as Booster. Three were stale. The apps catalog URL pointed at the old host over plain HTTP, the boot feature said CONFIGURATOR where Booster says FABRIC, and the Wi-Fi power default differed. The last is aligned for consistency only, since md-devops ignores that setting outright after STORY-02; a value that matches Booster and is never read is less confusing than one that differs and is also never read. All 19 keys now match booster/src/gconfig.c at v2.4.2, checked by diffing the parsed tables. On the device the stored values still win, and with WIFI_POWER at 4 in flash the health report still reports power saving off.
Neither had a caller anywhere in rp/src, and the linker had already proved they cost nothing: network_scan, network_getFoundNetworks and select_coreWaitPush are all absent from the built ELF, garbage collected by --gc-sections along with the ~6 KB scan table. The reasons for removing them are about the source, not the image. network_scan and friends: Wi-Fi scanning and configuration belong to Booster, which this app only reads settings from, so there is no future in which this is called -- and its scan callbacks were GCC nested functions, which clang cannot parse, so the dead code degraded clang-tidy for the whole file. select_coreWaitPush, its disable and the orphaned blocking select_waitPush: it launches core 1, and EPIC-07 STORY-01 already backed core 1 out because running it froze the Wi-Fi poll loop the main loop depends on. Anyone reaching for it would also need Booster's core-1 flash lockout. A ready-made entry point into a known hazard is worse than no entry point. SELECT is polled in the foreground by select_checkPushReset, which EPIC-13 STORY-07 and EPIC-14 STORY-04 build on. A comment stands where each used to be. pico/multicore.h is gone from the SELECT path, SELECT still resets the device, and smoke.py passes 8/8.
Stages the real SSID with a wrong password in memory, so the bad-key path can be exercised without touching the access point. It does fail and hand over to the supervisor, but this access point answers LINK JOIN and then times out rather than reporting BADAUTH, so the branch that pins the backoff to the ceiling and shows BAD AUTH on the menu is still unproven. STORY-08 is closed as blocked: the remaining scenarios need the access point taken down, which is not available here.
Wi-Fi lifecycle (EPIC-14)
The card was mounted once at boot, so pulling it and putting it back left GEMDRIVE and the file API dead until someone reset the device. Two separate things kept it that way. The driver never learns the card left. sd_card_spi_init() returns immediately unless STA_NOINIT is set, and with card-detect disabled on this board nothing ever sets it when a card is pulled -- so on reinsertion the driver still believed the old card was initialised, skipped its whole init sequence, and handed FatFs a freshly powered card it had never spoken to. sd_card_p->deinit() sets the bit back. And while the volume stays registered, mount_volume() never calls disk_initialize() at all, so f_mount(NULL, "0:", 0) has to come first. Both are needed: the first version of this fix had only the unregister, and every retry still failed with FR_DISK_ERR until the deinit was added. Detecting the removal cannot be left to FatFs results either. With the card physically out, volume answered 200 from the cached free space and a listing answered 200 with an empty directory -- FatFs serves what it cached and reports no error, so no failure ever arrives. The main loop now asks the card directly every 2 s by reading sector 0 through disk_read(), which the cache cannot answer. FatFs medium errors still mark the card gone, from the HTTP path and all 15 GEMDRIVE call sites, but the probe is what catches a pull. Verified with the card pulled and reinserted: detected in about 2 s, remounted about 4 s after reinsertion, volume and listings correct again, a 256 KB upload round-tripping identical, no reset and crash count 0. A device booted with no card also retries every 2 s and picks one up when it appears.
…15 STORY-02) With no card the device was silently broken: the setup menu said nothing, [G] and [U] launched a mode with no drive behind it, and the API answered six different ways -- including two that lied. /volume answered 200 with the cached size and free space, and a listing answered 200 with an empty directory, because FatFs serves what it cached. Every endpoint that needs the card now answers 503 no_sd_card: volume, listings, downloads, uploads, folder operations, runner load and runner run, through one guard in route() and one write_no_card(). write_fs_error() reports a missing card as such rather than as a disk fault. The CLI needed no change, since it maps by HTTP status and 503 already exits 6. The menu header carries the state -- GEMDRIVE SD: mounted / SD: NO CARD -- and emul_canLaunch() refuses [G] and [U] while no card is mounted, with "Insert a working microSD card: GEMDRIVE needs one." on the status line. STORY-01's retry loop is what clears it: inserting a card lifts the block within about two seconds, with no reset. The state went on the header line rather than a row of its own because a new row breaks the menu: dividers are drawn at fixed pixel rows and the status icons are right-aligned per header row, so one extra line struck Screenmem and Hook through, collided IP address with the USB CDC heading, and ran the status line into Select an option. The framebuffer snapshot caught it.
The device dropping a connection is routine -- a watchdog reboot, a panic or a SELECT press all cut the socket -- and the user got a Python traceback ending in ConnectionResetError. urllib wraps some connection failures in URLError, which the CLI already handled in six places, and lets the rest through as bare OSError subclasses. Two exposures: urlopen() itself, and the reads after a successful open -- the download loop and the debug tail loop both call resp.read() in a loop with no handler at all, which is exactly where a reboot mid-transfer lands. CONNECTION_DROPPED names the family once, including IncompleteRead, which a chunked body cut short raises instead of an OSError. request_json re-raises it as URLError so the six existing handlers keep working and each still names its own operation; the two streaming loops report through _dropped(). A dropped download leaves the partial file in place, which is what --resume is for. DroppedConnectionTests adds a socket server that resets with SO_LINGER 0, so the client sees a real ECONNRESET rather than a clean close, and covers a reset before any reply, on a JSON command, mid-download and on health. Each asserts exit 8, no traceback, and exactly one error: line. Run against the unfixed CLI all four error with the traceback the story was filed for, so they fail for the right reason. Suite: 119 tests, all passing.
After a timeout the ST re-sends the same 1 KB chunk with a fresh token, and the RP treated it as new data and appended it again: the file gained a duplicated chunk and lost its tail. Neither side carried anything that could tell the two apart -- a retried chunk and a new one are identical on the wire. An absolute offset does not work without more state than it appears: the m68k knows only how far it has got within the current Fwrite call, not the file position, which the RP owns and Fseek moves. A relative offset cannot distinguish a retry of a call's first chunk from the first chunk of the next call, since both are zero. So the m68k carries a number that never repeats. write_seq is a longword in the GEMDRIVE blob, bumped once per chunk and never per retry, never reset, sent in d5 -- which send_sync_write_command_to_sidecart always transmits and the RP was skipping. Each open file slot on the RP remembers the last sequence it accepted and the bytes it wrote; a chunk whose sequence matches is answered with that same count and not written again. The memo is stored before the answer is sent, because the answer is what goes missing. Verified on hardware with a debug hook that stalls the answer after the data is committed. Copying a 512 KB file with two chunks stalled 10 s each, the ST re-sent both, the RP reported "repeat of chunk N, answering 1024 again" for each, and the copy is byte-identical to the source. A first attempt at 2 s stalls proved nothing -- no retry happened at all, because that is inside the ST's COMMAND_TIMEOUT, and the matching checksum would have been a false pass. The cartridge image is now 10,144 of 10,240 bytes.
SD card and CLI rough edges (EPIC-15)
CLAUDE.md carried a memory map from before EPIC-12 -- RAM 128 K with the cartridge window at 0x20020000, where it is now 192 K and 0x20030000 -- and still said v1.1 "is making Release work". It now records that v1.1 ships Release and passed the gate, the stack at the top of RAM behind an MPU guard, FF_FS_LOCK at 28 and why, the automatic SD remount and the no-card answers, the lwIP pools carrying their measurements, and the Wi-Fi supervisor with the gateway probe. The README gains the two things a user meets: what happens with no SD card (the menu says so, [G] and [U] refuse, endpoints answer 503 no_sd_card, and inserting one clears it without a reset) and what happens when the network goes away (it rejoins by itself; the radio runs at full power). The CHANGELOG entry is user-facing and leads with the upload path, which could not complete a transfer beyond about 200 KB, then the aborted-transfer wedge, the SD card, the duplicated write chunk, Wi-Fi, crash visibility and the CLI. It ends with the two known limitations: heavy tracing through the debug ABI can take the ST down (EPIC-18, cause not found), and Wi-Fi recovery is verified against injected faults rather than a real outage.
…Y-04) The health sample in docs/api.md no longer matched what the firmware sends: it was missing the wifi object added in EPIC-14. Documented with the point that matters -- power_save is read back from the radio on every request rather than being the value we asked for, so 16 (CYW43_NONE_PM) is evidence and not an intention. too_many_open_files was also undocumented. It is described with the arithmetic behind FF_FS_LOCK's 28 entries and the fact that it usually means the ST is holding its share, so it is worth retrying rather than a fault.
v1.1.0 release gate (EPIC-16)
The entry read like the epic notes it came from: 437 lost bytes, an idle sweeper, a poll timer, a watchdog, MinSizeRel. None of that is what someone downloading the firmware wants to know. It now says what changed for them -- large files copy, a dead transfer stops blocking the next one, a reinserted card is picked up, files written from the Atari arrive whole, Wi-Fi is quicker and rejoins by itself -- and keeps the measurements only where they mean something to a user, such as the response time and packet loss from the radio no longer sleeping. The build change moved to a short note at the end, where a developer will still find it.
Comments and docs pointed at epics, stories and files under docs/epics/ -- "(EPIC-14 STORY-02)", "see docs/epics/02-http-api.md" -- but that directory is gitignored. Anyone reading the code was being sent to notes they do not have, to look up an identifier that appears nowhere in the repository. Every reference is replaced by what it was standing in for. Where the sentence already explained the reason, the tag simply goes; where the tag was carrying the meaning, the meaning is now written out: the scan helpers were "removed in v1.1", the stack rule is stated rather than cited, the smoke script's limits name the bugs they follow. CLAUDE.md gains the rule, with the grep to run before tagging a release.
The main-loop cadence comment, the menu overlay comment and debugcap.h each sent the reader to a "backlog entry" or a "Tier 2 backlog" that lives only in the untracked planning notes. Each now states the thing itself: moving chandler and the USB drain to core 1 is the step to take if 100 Hz stops being enough, framed menu sections would need the terminal renderer to know about them, and a consumer moved to core 1 would need explicit barriers.
The health examples in README.md and docs/api.md were taken from a build before the RAM work: 2 KB of stack reserved, a 118 KB heap, a stack pointer inside SCRATCH_Y. They now show what this firmware reports -- 16 KB reserved with a 3.5 KB peak, a 49,852-byte heap, code_in_ram 33,512 -- read off the linker symbols and the soak run. Two field descriptions were wrong rather than stale. heap.total is capped by the bottom of the core-0 stack, not by the cartridge window. stack.high_water can no longer quietly exceed reserved and walk into the core-1 stack area: the MPU guard faults and reboots instead. Undocumented until now: the two-connection limit and the 20 s idle timeout, which is what frees the transfer lock when a client disappears; 503 also covers too_many_open_files; and a device powered on with no card refuses the auto-launch countdown and waits in the menu. The README highlights say the firmware looks after itself, since that is most of what v1.1 is. programming.md still described the template's flash and scratch layout: CONFIG_FLASH is 120 KB of per-app sectors with the lookup and global sectors above it, and SCRATCH_Y is unused now that core 0's stack sits at the top of RAM.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
v1.1.0 — the robustness release. No new features: everything here is a fix or
a recovery path. It is also the first release built with full optimisation
(
Release); everything up tov1.0.1betashippedMinSizeRelbecauseReleasedid not survive on hardware.See
CHANGELOG.mdfor the user-facing entry. In short:arrive; a 4 MB file now round-trips byte-identical, and a slow client is no
longer cut off.
vanished mid-transfer left the lock held and FatFs handles counted until a
reset.
two seconds, and with no card the device says so —
SD: NO CARDon themenu,
[G]/[U]and the countdown refused,503 no_sd_cardfrom everyendpoint that needs it.
written twice, losing the tail of the file.
link rejoins by itself, a bad static IP falls back to DHCP instead of
crashing before the menu, and the password is out of the logs.
each reboot with the reason, phase and faulting PC on the menu and in
/api/v1/system/health, behind a crash-loop guard.docs/api.mdandprogramming.mddescribev1.1 as built, with no references to untracked planning notes.
Gate
smoke.py7/7 onrelease, 8/8 ondebug; 4 MB upload and downloadbyte-identical; aborted clients and a second client during a transfer all
answered; forced
panic,hardfaultandhangrecovered with the reasonrecorded.
crash_count0,free heap back to its starting value, stack 3,556 / 16,384.
Ships with these known gaps
reading more than a few thousand bytes through the cartridge debug window
during
runner execfaults the Atari. The firmware side has been measuredand cleared; the cause is not found, and the Runner has no soak evidence.
Work in progress on
epic/18-runner-execution-stability.the ST had GEMDRIVE open.