Skip to content

feat: verify externally hosted systems under test - #32

Merged
paullatzelsperger merged 6 commits into
mainfrom
feat/externalize_sut
Sep 16, 2026
Merged

paullatzelsperger merged 6 commits into
mainfrom
feat/externalize_sut

Conversation

@paullatzelsperger

@paullatzelsperger paullatzelsperger commented Sep 16, 2026 •

Copy link
Copy Markdown
Collaborator

Verify a system under test that runs outside the VE (reachable only via DSP/DCP) instead of onboarding it into the VE.

  • Hub: externallyHosted members (DID required) skip CFM provisioning; on CONFIRMED the IssuerService sends a DCP credential offer instead → CREDENTIALS_OFFERED. Lookup by DID for repeat runs.
  • Verification UI: entering a DID starts an external run that drives only the VE's side and waits for the SUT's (credentials, offer, certificate push). A DSP endpoint that 404s fails fast as a finding; no address override.
  • Reachability: Certo CCM URLs use the gateway form; setup-did-dns.sh --sut host=ip. HTTP-only is a documented constraint.
  • Platform: core platform 0.0.28 with advertisedProfile: cx-neptune (fix: make the DSP profile advertised in DID documents configurable eclipse-cfm/core-platform-distribution#26).

Verified live: managed run green; external onboarding delivered a real credential offer to a pseudo-SUT.

Open: the external run past AWAIT_CREDENTIALS is unit-tested only — needs a real vendor stack.

🤖 Generated with Claude Code

Picks up the configurable DSP profile in participant DID documents
(eclipse-cfm/core-platform-distribution#26) and sets it to `cx-neptune`,
the only dataspace profile this VE registers.

The controlplane serves one protocol context per registered profile at
/api/dsp/<participantContextId>/<profile name>. On the platform default
(`http-dsp-profile-2025-1`, the built-in profile this VE does not register)
every participant advertised a ProtocolEndpoint that 404s — harmless between
participants of this environment, which address each other by the profile they
both know they use, and fatal for an external counterparty, whose only way to
find the endpoint is the DID document.
Entering a DID on the run form verifies a system that already runs elsewhere
instead of onboarding one into this environment. The two runs are separate
sequences rather than one flow with branches (RunStep.MANAGED / EXTERNAL), and
a run's ledger holds only its own steps, so a timeline never shows a step that
was never going to be taken.

The external run drives only this environment's half. Everything on the system
under test's side is an obligation it fulfils in its own time — requesting the
credentials the issuer offers it, seeding its certificate offer, pushing a
certificate over a flow it establishes itself — so those steps are waits, not
calls, with budgets that bound a vendor's turnaround rather than a deployment's.
What the environment does drive: resolve the DID document (checkpoint 0, and the
source of the counterparty's DSP address — never an address synthesized from
this environment, which would verify the wrong system), consume the SUT's offer
as the verification participant, then retrieve and accept what arrives.

Two weakenings are deliberate. RETRIEVE_AND_VERIFY checks the delivery against
itself rather than byte-for-byte against an original, because the SUT authored
the document. And the ledger checklist is short: identity events belong to
wallets this environment provisions, and the exchange's own events carry the
verification participant's context, so onboarding plus credential delivery is
all the tracker can honestly attribute to a participant hosted elsewhere. The
credential delivery is the load-bearing one — it proves the SUT accepted this
environment's offer over DCP — and gates the exchange.

Repeat runs adopt the existing membership by DID: the Onboarding API treats an
already-registered DID as a duplicate, so re-onboarding is not a fresh start but
a guaranteed decline.

Live verification against a pseudo-SUT found the adopt lookup never matching.
RestCalls pre-encoded the DID into the path and the client encoded it again, so
a did:web carrying a port (host%3A8080) was escaped twice and matched nothing —
silent in the worst way, since an empty result reads as "not onboarded yet".
Caller-supplied values now go in as URI variables, pinned by a test asserting
the request URI.
A counterparty that answers 404/405 at the address its DID document advertises
is not slow, it is non-conformant: that document is the only way to discover
where to reach an external participant, so no amount of waiting fixes it. The
catalog wait treated it like any other failure and burned the whole budget —
15 minutes for an external run — before reporting "catalog request failed",
which reads as "the vendor has not seeded their offer yet" and sends an
operator after the wrong party.

It now fails immediately, naming the address dialled and where it came from.
Only the counterparty's OWN answer counts: a 5xx from our control plane, a SUT
still starting, or one refusing our credentials (401 — it IS serving DSP, and a
presentation can still settle) all keep retrying, since none of them is a
verdict about the address.

Detecting it means substring-matching the control plane's error text, which is
unpleasant and is what the management API leaves us: it wraps the far side's
response in a 502 whose only record of the far-side status is that message.

Deliberately no operator override of the address to go with this. Being able to
route around a wrong DID document would hide exactly the class of defect this
environment exists to find — and the one we just fixed in our own platform
(eclipse-cfm/core-platform-distribution#26). Checkpoint 0 in the SUT contract
now says so.
Certo's protocol API is COUNTERPARTY-FACING: the flow counterparty calls it
itself, carrying a Siglet flow token. Both places that named it — the CCM
transfer-type mapping installed on every member's data plane, and the baseUrl of
the CCM assets the verification runs seed — pointed at the in-cluster service,
which only a participant inside this cluster can resolve. They now use the
gateway form, which works from both sides (the CoreDNS rewrites resolve the host
in-cluster too) and passes clearglass, since the /api/certo route carries no
auth middleware and a flow token is not a jwtlet management token. Host-derived,
so both install scripts override them.

setup-did-dns.sh gains --sut <hostname>=<ip> (repeatable) and --sut-none. A SUT
lives outside the cluster, so its name cannot be rewritten to a Service the way
the VE's own hostnames are — it needs an address, via CoreDNS's hosts plugin
with fallthrough so the rest of DNS is untouched. The managed block is left
alone by runs that do not pass --sut, so re-installing does not drop it.

This matters more than plumbing: every check the VE makes about a SUT runs from
inside the cluster — the issuer resolves its did:web to deliver credential
offers, the control plane dials its ProtocolEndpoint — so a name only the
operator's machine resolves fails verification, and fails it as "the SUT is
unreachable", which reads as a finding against the vendor.

docs/sut-verification.md now states reachability in both directions (the -H
install flag being the reverse case) and records HTTP-only as a known,
deferred constraint: did:web resolution is pinned to http and the gateway
terminates HTTP only, so a SUT serving its DID document over HTTPS — the
method's own default — cannot be verified today.

What is NOT changed: siglet's EDR refresh endpoint is also in-cluster and
unrouted, but our CCM transfer-type mapping does not advertise renewal
(txRenewalSupport defaults false), so no counterparty is ever handed it. It
becomes a blocker the moment renewal is enabled.
@paullatzelsperger paullatzelsperger added the enhancement New feature or request label Sep 16, 2026
@paullatzelsperger
paullatzelsperger merged commit 3fe70b9 into main Sep 16, 2026
2 checks passed
@paullatzelsperger
paullatzelsperger deleted the feat/externalize_sut branch September 16, 2026 13:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants