Prevent extracted-component overwrites with content-addressed filenames - #71
Merged
Volv-G merged 1 commit intoSep 29, 2026
Conversation
Dehydration named extracted component files after the component name alone while deduplicating on digest, so two different components sharing a name wrote to one path: the first spec was silently overwritten and both tasks were redirected onto the survivor. Filenames are now <stem>-<sha256-of-spec><extension>, addressed by the spec being written rather than componentRef.digest -- the hydrator digests a component's raw source text, so equal specs can carry different valid digests and an edited spec can still carry a stale one. The stem is truncated on UTF-8 bytes, not code points, and is dropped entirely when the address and extension fill the 255-byte limit. An extension too long to name at all is reported instead of surfacing as OSError from the writer.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
(AI-assisted)
What
Dehydration named extracted component files after the component name alone while deduplicating on digest. Two different components sharing a name therefore wrote to one path: the first spec was silently overwritten, and both tasks were redirected onto the survivor.
This is the parent for the upcoming YAML→Python decompiler, which reuses this same extraction path and cannot offer a distinct-component guarantee without it.
Evidence
Reproduced on
masterwith two different specs that share the nametrain model:After the change: two files, payloads
['first', 'second'], distinct URLs.A supplied digest is not proof of content. The hydrator digests a component's raw source text (
pipeline_hydrator.py:860), so two byte-different files holding the same component carry different valid digests — a leading comment is enough:The converse also occurred: an edited spec keeps the digest of what it used to be, so two different specs could share a stale digest, collapse onto one file, and bypass the overwrite check entirely.
How
<stem>-<sha256-of-spec><extension>, and the dedup key is that same address, computed with the canonicalutils.compute_spec_digest(spec)— nevercomponentRef.digest. The ref digest is a locator; the file's content is the spec._safe_filenamefolds-to_, so the separator cannot occur inside the stem._safe_filenamekeeps non-ASCII alphanumerics, and a CJK stem is three bytes per character, so'界'*62produced a 256-byte name under a code-point cap. Truncation lands on a character boundary.<address><extension>with no leading dash.ComponentFilenameTooLongErrornaming the field and byte count, instead of surfacing asOSError: [Errno 63] File name too longfrom inside the writer.Failure modes
dehydrateFILE/AUTO output is<name>-<digest>.yamlrather than<name>.yaml. Refs are rewritten consistently in the same pass and stay relative to the output file, so a dehydrate→hydrate round trip is unaffected; previously extracted files on disk are not migrated.DIGEST,NAME,URL,KEEP,AUTO's canonical-URL and library-digest preferences), URI I/O, relative-ref rewriting and subgraph extraction are unchanged.<name>_<counter>and remain traversal-order dependent. They write to a separate directory and are unique within a run, so there is no data-loss bug there; deliberately out of scope.Testing
tests/test_pipeline_dehydrator.py: same-name/different-spec (sharing a stale locator digest) keeps both payloads; identical specs with differing,"unknown"and missing digests write once and share a URL named from the canonical spec digest; and a byte-limit test covering a CJK name, the exactaddress + extension == 255boundary, and one byte over rejected without writing.Residual gaps:
tests/test_packaging.pycould not run locally (pkgs.shopify.ioreturns 400 foruv-build; identical on pristinemaster), and the byte-limit tests prove encoded length and successful creation on APFS only — native Linux was unavailable. Both are covered by CI.Review focus
components/<name>.yaml. No in-tree caller does, and the known downstream wrapper overrides only the constructor and prompt, not_save_component_to_file(whose signature lost its now-unuseddigestparameter).ComponentFilenameCollisionErroris unreachable while the address is a full digest; it is kept as a loud guard so any future shortening fails instead of silently overwriting.