You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
PR #197 (durable session identity) added a per-path, in-process mutex
(path_lock in crates/tinyagents-session/src/transcript/history.rs)
serializing FileTranscriptHistory::append_turn/append_turn_with_partial/ append/replace/clear for a given resolved path, and made Session::persist (in tinyagents-runtime) commit a compaction's successor
generation to target/self.transcript only after the append into it
succeeds (pending_generation in crates/tinyagents-runtime/src/session.rs).
Both of those are real improvements over the pre-#197 baseline, where there
was no locking of any kind and the append-vs-compaction diff was computed
purely against in-memory state — a second process could already emit a
replacement record that erased the first process's turns, with no mitigation
at all.
What remains open
Two related gaps remain, both cross-process (two separate OS processes,
not two threads/handles within one process — the in-process case above is
already closed):
TranscriptLocator::begin_generation's existence check and the
handle construction it returns are not atomic across processes. FileTranscriptHistory::new performs no I/O — it only resolves a path —
so two processes compacting the same session at the same instant can both
observe the successor generation's path as absent (via session_exists)
and both return handles bound to the same not-yet-existing file. Neither
process's own view is invalidated by the other having done the same
check.
The eventual first write into that shared successor path is not
mutually exclusive across processes either. The writer's create-fresh
branch (append_transcript_turn_with_partial's !file_exists path) uses
a plain fs::write, which has no OS-level exclusivity: whichever of the
two processes' writes lands last silently wins, discarding the other's
retained (post-compaction) message set with no error and no trace beyond
the lost data itself.
Why this needs its own design, not a quick patch
These two pull in opposite directions and have to be resolved together:
Making generation reservation atomic (e.g. an exclusive create_new on
the successor's .jsonl path, claiming the filename before any content is
known) would satisfy (1), but by itself would violate the "commit the
successor only after its opening append succeeds" invariant this PR just
established — a process that reserves the file and then fails or crashes
before completing its append would leave an empty, "existing" generation
that head_generation might select over the one that should actually be
current.
Making the content write atomic (e.g. via the same create-if-absent + fs::hard_link publish pattern write_transcript_if_absent already uses
for adoption in this PR) closes (2) but does nothing for (1): two
processes could still both build a complete, valid successor payload in
parallel and race to publish it, with the loser's compaction silently
discarded rather than retried or reported.
A correct fix needs a single cross-process protocol that reserves the
generation slot and commits its content atomically together — most likely an
OS-level advisory file lock (flock/LockFileEx, ideally via a small,
audited dependency) spanning both the existence check and the write, or a
compare-and-swap style append protocol. That is a deliberate,
separately-reviewable design decision, not something to fold into a session
identity PR.
Scope
crates/tinyagents-session/src/transcript/history.rs: TranscriptLocator::begin_generation's default and FileTranscriptLocator's override.
Context
PR #197 (durable session identity) added a per-path, in-process mutex
(
path_lockincrates/tinyagents-session/src/transcript/history.rs)serializing
FileTranscriptHistory::append_turn/append_turn_with_partial/append/replace/clearfor a given resolved path, and madeSession::persist(intinyagents-runtime) commit a compaction's successorgeneration to
target/self.transcriptonly after the append into itsucceeds (
pending_generationincrates/tinyagents-runtime/src/session.rs).Both of those are real improvements over the pre-#197 baseline, where there
was no locking of any kind and the append-vs-compaction diff was computed
purely against in-memory state — a second process could already emit a
replacement record that erased the first process's turns, with no mitigation
at all.
What remains open
Two related gaps remain, both cross-process (two separate OS processes,
not two threads/handles within one process — the in-process case above is
already closed):
TranscriptLocator::begin_generation's existence check and thehandle construction it returns are not atomic across processes.
FileTranscriptHistory::newperforms no I/O — it only resolves a path —so two processes compacting the same session at the same instant can both
observe the successor generation's path as absent (via
session_exists)and both return handles bound to the same not-yet-existing file. Neither
process's own view is invalidated by the other having done the same
check.
The eventual first write into that shared successor path is not
mutually exclusive across processes either. The writer's create-fresh
branch (
append_transcript_turn_with_partial's!file_existspath) usesa plain
fs::write, which has no OS-level exclusivity: whichever of thetwo processes' writes lands last silently wins, discarding the other's
retained (post-compaction) message set with no error and no trace beyond
the lost data itself.
Why this needs its own design, not a quick patch
These two pull in opposite directions and have to be resolved together:
create_newonthe successor's
.jsonlpath, claiming the filename before any content isknown) would satisfy (1), but by itself would violate the "commit the
successor only after its opening append succeeds" invariant this PR just
established — a process that reserves the file and then fails or crashes
before completing its append would leave an empty, "existing" generation
that
head_generationmight select over the one that should actually becurrent.
fs::hard_linkpublish patternwrite_transcript_if_absentalready usesfor adoption in this PR) closes (2) but does nothing for (1): two
processes could still both build a complete, valid successor payload in
parallel and race to publish it, with the loser's compaction silently
discarded rather than retried or reported.
A correct fix needs a single cross-process protocol that reserves the
generation slot and commits its content atomically together — most likely an
OS-level advisory file lock (
flock/LockFileEx, ideally via a small,audited dependency) spanning both the existence check and the write, or a
compare-and-swap style append protocol. That is a deliberate,
separately-reviewable design decision, not something to fold into a session
identity PR.
Scope
crates/tinyagents-session/src/transcript/history.rs:TranscriptLocator::begin_generation's default andFileTranscriptLocator's override.crates/tinyagents-session/src/transcript/writer.rs:append_transcript_turn_with_partial's create-fresh branch.timestamp-free stems; a sealed generation stays byte-identical; adoption
never touches a legacy file; a root stem never contains
__.References
path_lockand the in-process fix), see its"Known limitations" section and the review threads on
history.rs:494/history.rs:806for the original reports.