Context
This Issue owns one bounded reliability repair for the real ChatGPT -> Dev MCP path on macOS.
It was opened from live evidence on 2026-09-18 after repeated ChatGPT dispatch friction. The initial incident report correctly identified very slow open_workspace, but it incorrectly conflated MCP transport-session churn with durable workspace/agent loss and incorrectly classified three separate failed agent attempts as session-reset replay.
This Issue must repair the actual failing contracts without introducing transport stickiness as a new correctness authority.
Related existing work:
Live evidence
Observed through the current ChatGPT -> Dev MCP surface:
E1 — durable workspace continuity already exists
Repeated open_workspace("/Users/jameschen/Workspace/devspace", mode="checkout") in the same ChatGPT conversation returned the same workspace:
A prior incident workspace remains durably queryable:
workspaceId: ws_e405d635dc
mode: worktree
status: active
root: .../.devspace-chatgpt/worktrees/Nexus-new-2164c7ae
Therefore:
new MCP transport session != lost workspace/agent state
Do not implement IP affinity or sticky HTTP transport as the semantic fix.
E2 — three observed agent starts were distinct terminal provider failures
For task nexus-957-wave2-independent-review:
v1 agt_c08c14dc
z-ai/glm-5.3-flash
-> PROVIDER_PROTOCOL_ERROR
-> model not advertised by Cline ACP session
v2 agt_dfe6c7bb
cline-free/deepseek-v4.1-flash
-> PROVIDER_PROTOCOL_ERROR
-> model not advertised by Cline ACP session
v3 agt_2b588a1d
cline-pass/glm-5.3
-> CLINEPASS_ENTITLEMENT_REQUIRED
These were not one agent silently replayed after transport loss.
The repair target is pre-launch route/readiness admission plus replay-safe dispatch, not generic session stickiness.
E3 — open_workspace has synchronous deep-work risk
Current workspace open/reuse path reloads profile/context and calls findAvailableAgentsFiles(), which recursively walks the workspace to discover nested AGENTS.md / CLAUDE.md.
Large repositories have shown observed open_workspace latency around 40-100+ seconds.
E4 — read-only file discovery has model-facing friction
The current ChatGPT surface relies on bash for rg/find/ls/tree style inspection.
Observed behavior:
- simple
ls succeeds quickly;
- shell read-only classification is lexical and treats raw metacharacters such as
| as consequential even when they are inside a search expression;
- valid model-facing discovery commands can therefore be blocked or become expensive;
- broad
rg on Nexus-new can run long enough to time out;
- guessed paths then produce
ENOENT.
This contributes directly to path drift.
E5 — strict exact edit behavior is desirable
edit currently rejects stale/non-exact oldText.
Preserve fail-closed exactness. The missing behavior is deterministic re-read/patch recovery, not fuzzy replacement.
Architecture invariant
Correctness identity must remain:
ChatGPT conversation identity
-> durable workspaceId
-> durable agentId / operationId
-> stable attemptKey
Transport identity is separate:
HTTP connection / MCP transport session / source IP
= transport lifecycle and performance only
Sticky transport MAY be used as an optimization if evidence justifies it, but MUST NOT become the only continuity mechanism.
Required repair
G1 — make open_workspace a bounded hot path
Rework open/reuse so returning a reusable workspace does not synchronously deep-walk the repository.
Requirements:
- root validation + durable conversation binding + workspace restore/reuse stay synchronous;
- root instruction loading needed for safe operation may remain synchronous;
- nested context discovery and other expensive refreshes must be cached, indexed, incremental, lazy, or otherwise removed from the response-critical path;
- preserve correctness for nested
AGENTS.md / CLAUDE.md: do not silently stop honoring nested instructions;
- preserve durable workspace restore after server restart.
Performance acceptance on representative large local fixture:
- warm reusable checkout p95 < 1s;
- cold checkout p95 < 3s excluding worktree creation;
- no full recursive walk on warm reuse.
G2 — prove reconnect continuity instead of assuming sticky session
Add an integration test that intentionally creates a new MCP transport/session while preserving the same ChatGPT conversation identity.
Prove:
- same conversation + same checkout -> same durable workspace binding;
- old durable agentId remains queryable after transport replacement;
- no duplicate workspace is created merely because transport/session/source IP changes;
- server restart restore still works from durable state.
If a real defect is found in conversation identity propagation, fix that exact seam.
Do NOT key semantic continuity to source IP.
G3 — preflight exact provider/model/entitlement before agent side effect
Before creating/launching a durable provider attempt, use the truthful provider readiness surface available for that provider.
For Cline, integrate/reuse #59 rather than inventing another catalog.
At minimum:
- exact model not advertised -> fail before provider execution;
- entitlement missing -> fail before provider execution when discoverable;
- stale/unknown readiness stays UNKNOWN, not READY;
- admission result is structured and attributable;
- no silent provider/model substitution.
Where readiness cannot be known without a bounded probe, keep that fact explicit.
G4 — replay-safe dispatch is mandatory
Use stable attemptKey as the retry/reconciliation identity for model-facing dispatch.
Prove:
- same attemptKey + identical request -> exactly one durable agent/effect;
- same attemptKey + conflicting request -> fail closed;
- transport timeout/disconnect -> status/reconcile before any new attempt;
- no blind second
agent_start solely because the host did not receive the first response.
Do not infer retry permission from OUTCOME_UNKNOWN.
G5 — give the model a reliable read-only discovery path
Provide or restore a native read-only discovery surface suitable for ChatGPT, preferably explicit operations equivalent to:
list_directory
find_files / glob
grep_files
If existing internal pi-tools already provide these semantics, expose/reuse them rather than creating a second search engine.
Requirements:
- read-only discovery works in shared checkout without mutation authority;
- read-only discovery works in Core-bound worktrees without requiring a mutation session;
- quoted regex/glob metacharacters are not misclassified as shell effects;
- operations are bounded by result count/output size/time;
- filesystem containment remains enforced;
- no shell write capability is introduced.
A parser improvement to bash classification may be added, but a model-facing native discovery surface is preferred over relying only on shell parsing.
G6 — preserve strict edit and add deterministic stale-edit recovery
Keep exact oldText matching fail-closed.
Provide/document/test a deterministic model-facing recovery path:
read latest file
-> exact targeted edit / structured patch
-> mismatch
-> re-read exact current region
-> retry once with fresh evidence
If a structured apply_patch capability already exists internally, expose/reuse it only if it preserves workspace containment and explicit failure semantics.
Do NOT add fuzzy silent replacement.
Required tests
Add focused regression/integration coverage for all gates.
At minimum:
- warm workspace reuse does not recursive-walk full repo;
- cold and warm latency harness on a synthetic large tree;
- nested instruction correctness still holds when entering/reading nested paths;
- transport replacement preserves workspace binding;
- transport replacement preserves durable agent query;
- server restart restores workspace/agent state;
- unavailable Cline model fails pre-launch;
- missing ClinePass entitlement fails pre-launch when known;
- UNKNOWN readiness is not promoted to READY;
- exact attempt replay produces one durable agent;
- conflicting attempt replay fails closed;
- read-only list/find/grep works in shared checkout;
- read-only list/find/grep works in isolated/Core-bound worktree without mutation authority;
- regex containing
| is treated as search data, not shell pipe, on the supported native discovery path;
- discovery is output/time bounded;
- guessed/nonexistent path returns structured ENOENT without changing workspace;
- stale exact edit fails closed;
- re-read + fresh exact edit/patch succeeds;
- existing conversation-isolation, Core mutation, durable agent, provider, build/typecheck suites remain green.
Verification
Before claiming Candidate-ready:
- rebind fresh
origin/main in an isolated DevSpace worktree;
- read repository instructions and current affected contracts;
- inspect existing implementations before adding new abstractions;
- run focused tests for each gate;
- run affected regression suite;
- run typecheck/build;
- run
git diff --check;
- inspect final diff for duplicate authority/catalog/search/session mechanisms;
- perform a real local ChatGPT/Dev MCP canary where available;
- report exact base/head, changed paths, tests, residual UNKNOWNs, and any gate not physically proven.
Non-goals
- no Nexus CapabilityPlanner changes;
- no new provider/model router;
- no second Cline catalog;
- no source-IP affinity as correctness;
- no broad repository cleanup;
- no fuzzy edit semantics;
- no automatic paid/free provider substitution;
- no merge/deploy/restart authority from this Issue.
Acceptance
The Issue is complete only when all six gates are proven:
G1 open_workspace latency
G2 reconnect continuity
G3 provider readiness fail-fast
G4 retry/attempt idempotency
G5 reliable read-only discovery
G6 deterministic exact-edit recovery
Maximum claim before all six pass:
Individual Dev MCP reliability defects have been repaired, but end-to-end ChatGPT dispatch reliability is not yet fully verified.
Context
This Issue owns one bounded reliability repair for the real ChatGPT -> Dev MCP path on macOS.
It was opened from live evidence on 2026-09-18 after repeated ChatGPT dispatch friction. The initial incident report correctly identified very slow
open_workspace, but it incorrectly conflated MCP transport-session churn with durable workspace/agent loss and incorrectly classified three separate failed agent attempts as session-reset replay.This Issue must repair the actual failing contracts without introducing transport stickiness as a new correctness authority.
Related existing work:
attemptKeymechanisms remain the semantic identity layer.Live evidence
Observed through the current ChatGPT -> Dev MCP surface:
E1 — durable workspace continuity already exists
Repeated
open_workspace("/Users/jameschen/Workspace/devspace", mode="checkout")in the same ChatGPT conversation returned the same workspace:A prior incident workspace remains durably queryable:
Therefore:
Do not implement IP affinity or sticky HTTP transport as the semantic fix.
E2 — three observed agent starts were distinct terminal provider failures
For task
nexus-957-wave2-independent-review:These were not one agent silently replayed after transport loss.
The repair target is pre-launch route/readiness admission plus replay-safe dispatch, not generic session stickiness.
E3 —
open_workspacehas synchronous deep-work riskCurrent workspace open/reuse path reloads profile/context and calls
findAvailableAgentsFiles(), which recursively walks the workspace to discover nestedAGENTS.md/CLAUDE.md.Large repositories have shown observed
open_workspacelatency around 40-100+ seconds.E4 — read-only file discovery has model-facing friction
The current ChatGPT surface relies on
bashforrg/find/ls/treestyle inspection.Observed behavior:
lssucceeds quickly;|as consequential even when they are inside a search expression;rgon Nexus-new can run long enough to time out;ENOENT.This contributes directly to path drift.
E5 — strict exact edit behavior is desirable
editcurrently rejects stale/non-exactoldText.Preserve fail-closed exactness. The missing behavior is deterministic re-read/patch recovery, not fuzzy replacement.
Architecture invariant
Correctness identity must remain:
Transport identity is separate:
Sticky transport MAY be used as an optimization if evidence justifies it, but MUST NOT become the only continuity mechanism.
Required repair
G1 — make
open_workspacea bounded hot pathRework open/reuse so returning a reusable workspace does not synchronously deep-walk the repository.
Requirements:
AGENTS.md/CLAUDE.md: do not silently stop honoring nested instructions;Performance acceptance on representative large local fixture:
G2 — prove reconnect continuity instead of assuming sticky session
Add an integration test that intentionally creates a new MCP transport/session while preserving the same ChatGPT conversation identity.
Prove:
If a real defect is found in conversation identity propagation, fix that exact seam.
Do NOT key semantic continuity to source IP.
G3 — preflight exact provider/model/entitlement before agent side effect
Before creating/launching a durable provider attempt, use the truthful provider readiness surface available for that provider.
For Cline, integrate/reuse #59 rather than inventing another catalog.
At minimum:
Where readiness cannot be known without a bounded probe, keep that fact explicit.
G4 — replay-safe dispatch is mandatory
Use stable
attemptKeyas the retry/reconciliation identity for model-facing dispatch.Prove:
agent_startsolely because the host did not receive the first response.Do not infer retry permission from
OUTCOME_UNKNOWN.G5 — give the model a reliable read-only discovery path
Provide or restore a native read-only discovery surface suitable for ChatGPT, preferably explicit operations equivalent to:
If existing internal
pi-toolsalready provide these semantics, expose/reuse them rather than creating a second search engine.Requirements:
A parser improvement to
bashclassification may be added, but a model-facing native discovery surface is preferred over relying only on shell parsing.G6 — preserve strict edit and add deterministic stale-edit recovery
Keep exact
oldTextmatching fail-closed.Provide/document/test a deterministic model-facing recovery path:
If a structured
apply_patchcapability already exists internally, expose/reuse it only if it preserves workspace containment and explicit failure semantics.Do NOT add fuzzy silent replacement.
Required tests
Add focused regression/integration coverage for all gates.
At minimum:
|is treated as search data, not shell pipe, on the supported native discovery path;Verification
Before claiming Candidate-ready:
origin/mainin an isolated DevSpace worktree;git diff --check;Non-goals
Acceptance
The Issue is complete only when all six gates are proven:
Maximum claim before all six pass: