Skip to content

list --sources buffers every bundle document just to read provenance #1012

Description

@jasonssdev

Problem

list --sources (application/list_service.list_provenance_sources) decodes every bundle document into a whole-bundle dict[str, str] before running any analysis, so peak memory grows with the total size of the bundle text. Only the provenance field of each document is ever used.

Why the mapping exists

bundle_provenance.provenance_source_ancestors(files: Mapping[str, str], *, object_id) takes full file texts, and its first step is _parse_provenance_by_id(files), which reduces them to an id -> provenance ids index. After that step the walk upward to a fixed point needs random access to that index, but never to the document text.

Possible direction

Build the id -> provenance index one document at a time and pass the index in (or give provenance_source_ancestors a variant that accepts it), so no document's full text outlives its own parse. Unreadable documents keep being recorded as NotRun (ADR-0022, #1011).

Other callers take the same Mapping[str, str] snapshot (for example resolve_source_raises). Keep this change scoped to the list --sources path unless the shared parse makes the others trivial.

Priority

Low. Nothing is broken, and it matters only at bundle sizes nobody has reported yet. It came up as a scoped-out item of #1002. Measure peak memory on a large bundle before and after, and keep the change only if the difference shows up.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions