Skip to content

P0: harden ChatGPT -> Dev MCP dispatch reliability end to end #194

Description

@James3014

Context

This Issue owns one bounded reliability repair for the real ChatGPT -> Dev MCP path on macOS.

It was opened from live evidence on 2026-09-18 after repeated ChatGPT dispatch friction. The initial incident report correctly identified very slow open_workspace, but it incorrectly conflated MCP transport-session churn with durable workspace/agent loss and incorrectly classified three separate failed agent attempts as session-reset replay.

This Issue must repair the actual failing contracts without introducing transport stickiness as a new correctness authority.

Related existing work:

Live evidence

Observed through the current ChatGPT -> Dev MCP surface:

E1 — durable workspace continuity already exists

Repeated open_workspace("/Users/jameschen/Workspace/devspace", mode="checkout") in the same ChatGPT conversation returned the same workspace:

ws_c1d66f1a4c

A prior incident workspace remains durably queryable:

workspaceId: ws_e405d635dc
mode: worktree
status: active
root: .../.devspace-chatgpt/worktrees/Nexus-new-2164c7ae

Therefore:

new MCP transport session != lost workspace/agent state

Do not implement IP affinity or sticky HTTP transport as the semantic fix.

E2 — three observed agent starts were distinct terminal provider failures

For task nexus-957-wave2-independent-review:

v1 agt_c08c14dc
z-ai/glm-5.3-flash
-> PROVIDER_PROTOCOL_ERROR
-> model not advertised by Cline ACP session

v2 agt_dfe6c7bb
cline-free/deepseek-v4.1-flash
-> PROVIDER_PROTOCOL_ERROR
-> model not advertised by Cline ACP session

v3 agt_2b588a1d
cline-pass/glm-5.3
-> CLINEPASS_ENTITLEMENT_REQUIRED

These were not one agent silently replayed after transport loss.

The repair target is pre-launch route/readiness admission plus replay-safe dispatch, not generic session stickiness.

E3 — open_workspace has synchronous deep-work risk

Current workspace open/reuse path reloads profile/context and calls findAvailableAgentsFiles(), which recursively walks the workspace to discover nested AGENTS.md / CLAUDE.md.

Large repositories have shown observed open_workspace latency around 40-100+ seconds.

E4 — read-only file discovery has model-facing friction

The current ChatGPT surface relies on bash for rg/find/ls/tree style inspection.

Observed behavior:

  • simple ls succeeds quickly;
  • shell read-only classification is lexical and treats raw metacharacters such as | as consequential even when they are inside a search expression;
  • valid model-facing discovery commands can therefore be blocked or become expensive;
  • broad rg on Nexus-new can run long enough to time out;
  • guessed paths then produce ENOENT.

This contributes directly to path drift.

E5 — strict exact edit behavior is desirable

edit currently rejects stale/non-exact oldText.

Preserve fail-closed exactness. The missing behavior is deterministic re-read/patch recovery, not fuzzy replacement.

Architecture invariant

Correctness identity must remain:

ChatGPT conversation identity
  -> durable workspaceId
  -> durable agentId / operationId
  -> stable attemptKey

Transport identity is separate:

HTTP connection / MCP transport session / source IP
  = transport lifecycle and performance only

Sticky transport MAY be used as an optimization if evidence justifies it, but MUST NOT become the only continuity mechanism.

Required repair

G1 — make open_workspace a bounded hot path

Rework open/reuse so returning a reusable workspace does not synchronously deep-walk the repository.

Requirements:

  • root validation + durable conversation binding + workspace restore/reuse stay synchronous;
  • root instruction loading needed for safe operation may remain synchronous;
  • nested context discovery and other expensive refreshes must be cached, indexed, incremental, lazy, or otherwise removed from the response-critical path;
  • preserve correctness for nested AGENTS.md / CLAUDE.md: do not silently stop honoring nested instructions;
  • preserve durable workspace restore after server restart.

Performance acceptance on representative large local fixture:

  • warm reusable checkout p95 < 1s;
  • cold checkout p95 < 3s excluding worktree creation;
  • no full recursive walk on warm reuse.

G2 — prove reconnect continuity instead of assuming sticky session

Add an integration test that intentionally creates a new MCP transport/session while preserving the same ChatGPT conversation identity.

Prove:

  1. same conversation + same checkout -> same durable workspace binding;
  2. old durable agentId remains queryable after transport replacement;
  3. no duplicate workspace is created merely because transport/session/source IP changes;
  4. server restart restore still works from durable state.

If a real defect is found in conversation identity propagation, fix that exact seam.

Do NOT key semantic continuity to source IP.

G3 — preflight exact provider/model/entitlement before agent side effect

Before creating/launching a durable provider attempt, use the truthful provider readiness surface available for that provider.

For Cline, integrate/reuse #59 rather than inventing another catalog.

At minimum:

  • exact model not advertised -> fail before provider execution;
  • entitlement missing -> fail before provider execution when discoverable;
  • stale/unknown readiness stays UNKNOWN, not READY;
  • admission result is structured and attributable;
  • no silent provider/model substitution.

Where readiness cannot be known without a bounded probe, keep that fact explicit.

G4 — replay-safe dispatch is mandatory

Use stable attemptKey as the retry/reconciliation identity for model-facing dispatch.

Prove:

  • same attemptKey + identical request -> exactly one durable agent/effect;
  • same attemptKey + conflicting request -> fail closed;
  • transport timeout/disconnect -> status/reconcile before any new attempt;
  • no blind second agent_start solely because the host did not receive the first response.

Do not infer retry permission from OUTCOME_UNKNOWN.

G5 — give the model a reliable read-only discovery path

Provide or restore a native read-only discovery surface suitable for ChatGPT, preferably explicit operations equivalent to:

list_directory
find_files / glob
grep_files

If existing internal pi-tools already provide these semantics, expose/reuse them rather than creating a second search engine.

Requirements:

  • read-only discovery works in shared checkout without mutation authority;
  • read-only discovery works in Core-bound worktrees without requiring a mutation session;
  • quoted regex/glob metacharacters are not misclassified as shell effects;
  • operations are bounded by result count/output size/time;
  • filesystem containment remains enforced;
  • no shell write capability is introduced.

A parser improvement to bash classification may be added, but a model-facing native discovery surface is preferred over relying only on shell parsing.

G6 — preserve strict edit and add deterministic stale-edit recovery

Keep exact oldText matching fail-closed.

Provide/document/test a deterministic model-facing recovery path:

read latest file
-> exact targeted edit / structured patch
-> mismatch
-> re-read exact current region
-> retry once with fresh evidence

If a structured apply_patch capability already exists internally, expose/reuse it only if it preserves workspace containment and explicit failure semantics.

Do NOT add fuzzy silent replacement.

Required tests

Add focused regression/integration coverage for all gates.

At minimum:

  1. warm workspace reuse does not recursive-walk full repo;
  2. cold and warm latency harness on a synthetic large tree;
  3. nested instruction correctness still holds when entering/reading nested paths;
  4. transport replacement preserves workspace binding;
  5. transport replacement preserves durable agent query;
  6. server restart restores workspace/agent state;
  7. unavailable Cline model fails pre-launch;
  8. missing ClinePass entitlement fails pre-launch when known;
  9. UNKNOWN readiness is not promoted to READY;
  10. exact attempt replay produces one durable agent;
  11. conflicting attempt replay fails closed;
  12. read-only list/find/grep works in shared checkout;
  13. read-only list/find/grep works in isolated/Core-bound worktree without mutation authority;
  14. regex containing | is treated as search data, not shell pipe, on the supported native discovery path;
  15. discovery is output/time bounded;
  16. guessed/nonexistent path returns structured ENOENT without changing workspace;
  17. stale exact edit fails closed;
  18. re-read + fresh exact edit/patch succeeds;
  19. existing conversation-isolation, Core mutation, durable agent, provider, build/typecheck suites remain green.

Verification

Before claiming Candidate-ready:

  • rebind fresh origin/main in an isolated DevSpace worktree;
  • read repository instructions and current affected contracts;
  • inspect existing implementations before adding new abstractions;
  • run focused tests for each gate;
  • run affected regression suite;
  • run typecheck/build;
  • run git diff --check;
  • inspect final diff for duplicate authority/catalog/search/session mechanisms;
  • perform a real local ChatGPT/Dev MCP canary where available;
  • report exact base/head, changed paths, tests, residual UNKNOWNs, and any gate not physically proven.

Non-goals

  • no Nexus CapabilityPlanner changes;
  • no new provider/model router;
  • no second Cline catalog;
  • no source-IP affinity as correctness;
  • no broad repository cleanup;
  • no fuzzy edit semantics;
  • no automatic paid/free provider substitution;
  • no merge/deploy/restart authority from this Issue.

Acceptance

The Issue is complete only when all six gates are proven:

G1 open_workspace latency
G2 reconnect continuity
G3 provider readiness fail-fast
G4 retry/attempt idempotency
G5 reliable read-only discovery
G6 deterministic exact-edit recovery

Maximum claim before all six pass:

Individual Dev MCP reliability defects have been repaired, but end-to-end ChatGPT dispatch reliability is not yet fully verified.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions