feat: verify externally hosted systems under test - #32
Merged
Merged
Conversation
Picks up the configurable DSP profile in participant DID documents (eclipse-cfm/core-platform-distribution#26) and sets it to `cx-neptune`, the only dataspace profile this VE registers. The controlplane serves one protocol context per registered profile at /api/dsp/<participantContextId>/<profile name>. On the platform default (`http-dsp-profile-2025-1`, the built-in profile this VE does not register) every participant advertised a ProtocolEndpoint that 404s — harmless between participants of this environment, which address each other by the profile they both know they use, and fatal for an external counterparty, whose only way to find the endpoint is the DID document.
Entering a DID on the run form verifies a system that already runs elsewhere instead of onboarding one into this environment. The two runs are separate sequences rather than one flow with branches (RunStep.MANAGED / EXTERNAL), and a run's ledger holds only its own steps, so a timeline never shows a step that was never going to be taken. The external run drives only this environment's half. Everything on the system under test's side is an obligation it fulfils in its own time — requesting the credentials the issuer offers it, seeding its certificate offer, pushing a certificate over a flow it establishes itself — so those steps are waits, not calls, with budgets that bound a vendor's turnaround rather than a deployment's. What the environment does drive: resolve the DID document (checkpoint 0, and the source of the counterparty's DSP address — never an address synthesized from this environment, which would verify the wrong system), consume the SUT's offer as the verification participant, then retrieve and accept what arrives. Two weakenings are deliberate. RETRIEVE_AND_VERIFY checks the delivery against itself rather than byte-for-byte against an original, because the SUT authored the document. And the ledger checklist is short: identity events belong to wallets this environment provisions, and the exchange's own events carry the verification participant's context, so onboarding plus credential delivery is all the tracker can honestly attribute to a participant hosted elsewhere. The credential delivery is the load-bearing one — it proves the SUT accepted this environment's offer over DCP — and gates the exchange. Repeat runs adopt the existing membership by DID: the Onboarding API treats an already-registered DID as a duplicate, so re-onboarding is not a fresh start but a guaranteed decline. Live verification against a pseudo-SUT found the adopt lookup never matching. RestCalls pre-encoded the DID into the path and the client encoded it again, so a did:web carrying a port (host%3A8080) was escaped twice and matched nothing — silent in the worst way, since an empty result reads as "not onboarded yet". Caller-supplied values now go in as URI variables, pinned by a test asserting the request URI.
A counterparty that answers 404/405 at the address its DID document advertises is not slow, it is non-conformant: that document is the only way to discover where to reach an external participant, so no amount of waiting fixes it. The catalog wait treated it like any other failure and burned the whole budget — 15 minutes for an external run — before reporting "catalog request failed", which reads as "the vendor has not seeded their offer yet" and sends an operator after the wrong party. It now fails immediately, naming the address dialled and where it came from. Only the counterparty's OWN answer counts: a 5xx from our control plane, a SUT still starting, or one refusing our credentials (401 — it IS serving DSP, and a presentation can still settle) all keep retrying, since none of them is a verdict about the address. Detecting it means substring-matching the control plane's error text, which is unpleasant and is what the management API leaves us: it wraps the far side's response in a 502 whose only record of the far-side status is that message. Deliberately no operator override of the address to go with this. Being able to route around a wrong DID document would hide exactly the class of defect this environment exists to find — and the one we just fixed in our own platform (eclipse-cfm/core-platform-distribution#26). Checkpoint 0 in the SUT contract now says so.
Certo's protocol API is COUNTERPARTY-FACING: the flow counterparty calls it itself, carrying a Siglet flow token. Both places that named it — the CCM transfer-type mapping installed on every member's data plane, and the baseUrl of the CCM assets the verification runs seed — pointed at the in-cluster service, which only a participant inside this cluster can resolve. They now use the gateway form, which works from both sides (the CoreDNS rewrites resolve the host in-cluster too) and passes clearglass, since the /api/certo route carries no auth middleware and a flow token is not a jwtlet management token. Host-derived, so both install scripts override them. setup-did-dns.sh gains --sut <hostname>=<ip> (repeatable) and --sut-none. A SUT lives outside the cluster, so its name cannot be rewritten to a Service the way the VE's own hostnames are — it needs an address, via CoreDNS's hosts plugin with fallthrough so the rest of DNS is untouched. The managed block is left alone by runs that do not pass --sut, so re-installing does not drop it. This matters more than plumbing: every check the VE makes about a SUT runs from inside the cluster — the issuer resolves its did:web to deliver credential offers, the control plane dials its ProtocolEndpoint — so a name only the operator's machine resolves fails verification, and fails it as "the SUT is unreachable", which reads as a finding against the vendor. docs/sut-verification.md now states reachability in both directions (the -H install flag being the reverse case) and records HTTP-only as a known, deferred constraint: did:web resolution is pinned to http and the gateway terminates HTTP only, so a SUT serving its DID document over HTTPS — the method's own default — cannot be verified today. What is NOT changed: siglet's EDR refresh endpoint is also in-cluster and unrouted, but our CCM transfer-type mapping does not advertise renewal (txRenewalSupport defaults false), so no counterparty is ever handed it. It becomes a blocker the moment renewal is enabled.
jimmarino
approved these changes
Sep 16, 2026
wolf4ood
approved these changes
Sep 16, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Verify a system under test that runs outside the VE (reachable only via DSP/DCP) instead of onboarding it into the VE.
externallyHostedmembers (DID required) skip CFM provisioning; on CONFIRMED the IssuerService sends a DCP credential offer instead →CREDENTIALS_OFFERED. Lookup by DID for repeat runs.setup-did-dns.sh --sut host=ip. HTTP-only is a documented constraint.advertisedProfile: cx-neptune(fix: make the DSP profile advertised in DID documents configurable eclipse-cfm/core-platform-distribution#26).Verified live: managed run green; external onboarding delivered a real credential offer to a pseudo-SUT.
Open: the external run past
AWAIT_CREDENTIALSis unit-tested only — needs a real vendor stack.🤖 Generated with Claude Code