riscv: implement multikernel lifecycle and safe shutdown - #39
pro-utkarshM wants to merge 11 commits into
Conversation
The manifest code reaches into the pool's arch state for the park slot and calls mk_pool_park_regions(), which only the x86 header declares. Both are part of what CONFIG_ARCH_HAS_MK_HOST_PARK stands for, and an architecture that parks CPUs in firmware has neither, so the generic code does not build there. Declare mk_pool_park_regions() with the other host park functions and add mk_pool_park_slot() beside it, with the usual empty fallbacks. Signed-off-by: Cong Wang <cwang@multikernel.io>
Wire CONFIG_MULTIKERNEL into the 64-bit RISC-V build and add sparse hart ID translations. Reject the invalid hart sentinel before lookup so it cannot alias an unused logical CPU slot. Reserve the architecture control block for the spawn context, DTB and entry stub. Provide safe stubs for the full architecture interface so the functional SBI HSM, Image loader and doorbell work can land incrementally. [Picked from multikernel#37 with the Kconfig dependencies split out and the ARCH_HAS_MK_POOL_STATE gate on CPU removal dropped: an arch that cannot park a departing CPU should refuse in its own takedown path. Added the empty struct mk_pool_arch the core now embeds, and dropped a force-stop registration stub that nothing declares.] Signed-off-by: Nikolay Nikolaev <nicknickolaev@gmail.com> Signed-off-by: Cong Wang <cwang@multikernel.io>
The instance restore and ring handoff paths consume the live OF tree. Without CONFIG_OF, x86 multikernel configurations compile references to OF globals that cannot link and cannot restore a spawn at runtime. Reject that unusable configuration in Kconfig. Signed-off-by: Utkarsh Maurya <projects.utkarshMaurya@gmail.com>
RISC-V parks CPUs in firmware with SBI HSM, so multikernel must reuse the architecture's HART_START, HART_STOP and HART_STATUS wrappers. Make those helpers available to RISC-V architecture code while retaining the existing SBI-to-Linux errno mapping. Add only the per-instance context and entry-stub bookkeeping required by the lifecycle implementation. Hart IDs remain physical firmware identifiers and are never used as array indexes. Link: multikernel#23 Signed-off-by: Utkarsh Maurya <projects.utkarshMaurya@gmail.com>
OpenSBI does not invalidate a stopped hart's instruction cache when HART_START restarts it, and an SBI remote fence cannot target a stopped hart. Reusing an Image address can therefore execute stale instructions. Add an immutable, relocation-free entry stub whose first instruction is fence.i and whose target comes from the adjacent data page. Publish its address through the multikernel manifest and route spawn-kernel secondary starts through it as well as the primary start. The stub preserves a0 and a1, so the primary receives its DTB and secondaries receive their normal SBI boot data. Flush the local instruction cache before every HART_STOP so a later start cannot retain an older stub line. Link: multikernel#24 Signed-off-by: Utkarsh Maurya <projects.utkarshMaurya@gmail.com>
Implement the RISC-V multikernel CPU lifecycle with direct SBI HSM calls. A spawn is allowed only after HART_STATUS reaches STOPPED; STARTED, START_PENDING and STOP_PENDING are polled with a bounded timeout. Allocate the immutable entry stub and mutable target context from the instance control block, publish the stub in the manifest, and start the physical hart with the Image entry and DTB boot ABI. Use physical hart IDs for host doorbells and confirm firmware stop state before release or memory reclaim. RISC-V deliberately reports force-stop as unsupported because HSM has no remote HART_STOP operation. Link: multikernel#23 Link: multikernel#24 Signed-off-by: Utkarsh Maurya <projects.utkarshMaurya@gmail.com>
Prime every assigned hart through an immutable host-text fence.i trampoline before it fetches a rewritten instance stub or Image. Preserve distinct HSM state errors for bounded lifecycle handling. Link: multikernel#23 Link: multikernel#24 Signed-off-by: Utkarsh Maurya <projects.utkarshMaurya@gmail.com>
Reuse the kexec lock for CPU, memory and device transfers so an Image rewrite cannot race with resource mutation. Recover an active instance only after every assigned hart is confirmed stopped. Link: multikernel#23 Link: multikernel#26 Signed-off-by: Utkarsh Maurya <projects.utkarshMaurya@gmail.com>
A spawned kernel must never invoke SBI SRST or a legacy shutdown because those operations can reset the host. Detect the spawn handoff before reset registration and suppress host-wide reset handlers. Route halt, poweroff, restart, SMP stop and panic shutdown through local HART_STOP. Cache the parent endpoint for a non-allocating panic notification, set a safe spawn panic default, and prevent memory reuse until every assigned hart is confirmed stopped. Link: multikernel#25 Signed-off-by: Utkarsh Maurya <projects.utkarshMaurya@gmail.com>
Reserve instance tracking before handing a pool CPU to a running kernel. If the response is lost and firmware cannot confirm the hart stopped, conservatively keep the CPU assigned to that instance instead of exposing a possibly running hart through the free pool. Link: multikernel#23 Signed-off-by: Utkarsh Maurya <projects.utkarshMaurya@gmail.com>
Take the kexec lock while confirming parked CPUs and moving an instance back to the loaded state. This keeps asynchronous halt notifications from racing CPU ownership changes or image teardown. Link: multikernel#23 Signed-off-by: Utkarsh Maurya <projects.utkarshMaurya@gmail.com>
7fd9938 to
eef3dcb
Compare
|
@congwang-mk, before updating this draft, could you confirm the dependency ordering? We have newer local #23–#25 lifecycle commits, with #26 kept as a separate follow-up. Standalone builds passed; QEMU PLIC/AIA validation passed on an integration stack with explicit runtime dependencies, including 50-cycle same-Image and alternating-Image re-spawns. The tested stack still has the shared receive-ring reset limitation. Nickolaev’s #7 overlaps that work but is not integrated or validated in our stack. Should we depend on #7 first, or submit the bounded RISC-V series with that limitation disclosed? Which public #27/#28 commits should we use instead of our temporary attributed runtime extracts? |
|
There are no public #27/#28 commits: both are unassigned with no PR or branch, and Please don't carry extracts from Nikolay's July Don't depend on #7; disclose the ring-reset limitation. For this PR:
|
fe31d7b to
77417e1
Compare
Summary
physical hart IDs, bounded
HART_STATUSpolling, and parked-state checksfence.ion every start whilepreserving the RISC-V
a0/a1boot ABIHART_STOP, and serialize lifecycle/resource changes against image teardownuncertain, so a possibly running hart is never exposed through the free pool
Addresses #23, #24, and #25.
Base and attribution
This draft targets
multikernel-riscvatfe31d7b0ca32b0f806c88059c88d70600fda45cc, the RISC-V foundation from #37that Cong applied for follow-up work.
The previous temporary replay of Nikolay Nikolaev's foundation commit has been
dropped. His accepted foundation remains below this series with its original
authorship; every commit in this PR is above that branch point.
Scope
This PR keeps #23-#25 together because lifecycle, entry-cache coherency, and
safe shutdown share the same HSM handoff invariants.
#26 remains a separate follow-up branch. RISC-V instance-DTB device filtering
(#27) and firmware IPI receive support (#28) are out of scope.
Validation
Current head:
eef3dcbad234vmlinux Imagebuild withCONFIG_MULTIKERNEL=yvmlinux Imagebuild with multikernel disablednommu_virt_defconfigcheck provingRISCV_M_MODE=yexcludes multikernelvmlinuxbuild withCONFIG_MULTIKERNEL=y,CONFIG_OF=y, andCONFIG_KEXEC_FILE=yfence.i, noa0/a1clobber, and norelocations
git diff --checkand strict per-commitcheckpatch.pl: zero reportederrors, warnings, or checks
accounting, RISC-V stopped-state confirmation, and halted-settlement locking
Prior QEMU/OpenSBI runtime results were produced on the pre-rebase series
stacked with #26. The accepted-base head still needs runtime revalidation
before this draft is ready for merge.
Remaining draft gates
virt+ OpenSBI PLIC lifecycle tests on the accepted baserespawn, different-Image respawn, and 50-cycle stress tests
ownership/cleanup behavior
AIA runtime validation remains blocked on #27's per-hart IMSIC/APLIC and device
filtering. MKTTY/ring validation remains blocked on #28; neither is duplicated
here.