fix: recover from a Letta conversation the backend no longer has - #14
Conversation
A session's stored conversation id is replayed on every later turn, so when the local backend loses that conversation (its store is replaced, the agent is re-created, or state from another backend is reused) every request for that session fails. Data's Discord session stored a pre-cutover 'con_' id that the local backend never had, so every mention failed with a 502: letta -> exit 1: Conversation con_QPvFAwS14xUBFIIb not found Detect that failure and drop the mapping, then retry once against a fresh conversation, so a stale id costs one extra turn instead of bricking the session. The detector only consults stderr when no result object was parsed, so prose in a real answer cannot trigger a retry.
|
This pull request addresses an issue where a session's stored conversation ID, when no longer recognized by the Letta backend, would cause all subsequent requests for that session to fail. The fix introduces a mechanism to detect this "missing conversation" failure, drop the stale conversation ID, and retry the request with a fresh conversation.
Reviewers, please start by examining |
1 similar comment
|
This pull request addresses an issue where a session's stored conversation ID, when no longer recognized by the Letta backend, would cause all subsequent requests for that session to fail. The fix introduces a mechanism to detect this "missing conversation" failure, drop the stale conversation ID, and retry the request with a fresh conversation.
Reviewers, please start by examining |
On the recovery path `existing` names the conversation the backend no longer has, so a fresh turn whose JSON omits `conversation_id` wrote the dead id straight back into conversations.json - recreating the exact state this recovery exists to clear. Track that recovery ran and fall back to undefined, which leaves the session unmapped so the next question starts a new conversation. Also document the behavior and its boundary: conversation loss recovers automatically, agent loss (`Agent <id> not found`) does not and needs `letta agents create` plus a `DATA_LETTA_AGENT_ID` update.
|
Added the re-brick guard as Boundary — what this does not cover. Recovery here handles conversation loss. A wiped backend also loses the agent, which surfaces differently — Verification. A stub
Typecheck clean, |
Problem
A session's stored conversation id is replayed on every later turn. When the local
Letta backend no longer has that conversation, every request for that session fails
permanently — no retry ever succeeds, because the same dead id is re-sent.
Data's Discord session held a pre-cutover
con_id that the local backend never had(it names conversations
local-conv-N). Every Discord mention returned 502:Observed live: a real mention (
1553268586628255786) failed to answer at2026-09-26T04:54:50Z for this reason.
Fix
lib/conversation-recovery.ts),which prints
Conversation <id> not foundon stderr and exits 1 with no stdout.askLetta, drop the stale mapping and retry once against a fresh conversation,so a stale id costs one extra turn instead of bricking the session.
prose in a real answer cannot trigger a spurious retry.
Verification
npm run typecheckclean;npm test26/26 pass (2 new tests).Isolated end-to-end run of
channels/http/index.tsagainst the real local Lettaagent, seeded with the stale id as a session mapping:
POST /ask {session:"stale-test"}conversation con_QPvFAwS14xUBFIIb is gone; starting a fresh one for session stale-test200with a freshlocal-conv-63The live host's stale entry was cleared and
data-httprestarted; a/askon thereal Discord session (
discord:1531147245318176848) now returns200with a freshconversation, where it previously 502'd.