Write root metadata.labels from @pipeline(labels=...) (v0.1.27) - #70
Merged
Volv-G merged 1 commit intoSep 28, 2026
Conversation
GraphBuilder carried only annotations, so the five corpus pipelines with a root metadata.labels block could not be authored in Python. Both schemas type metadata.labels as additionalProperties string, which is narrower than metadata.annotations, so the value policy is strict str -> str. Reuses check_annotations under a new PIPELINE_LABELS_POLICY, with the reserved system/ prefix rule deliberately off: that prefix is reserved for Tangle's own annotations, not for labels. labels is written before annotations, matching the corpus majority. Root only, descriptive only, and outside compile identity.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
(AI-assisted)
What
@pipeline(labels={...})writes the compiled pipeline's rootmetadata.labelsblock.GraphBuildercarried only annotations, so five corpus pipelines could not be authored in Python at all.Evidence
Fifteen corpus files carry a root
metadata.labels; ten are components, leaving five pipelines:metadatakey orderrelevance-tools/…/search_signals/pipeline.yamllabels, annotations…/join_features_from_bigquery_to_featureset_pipeline.yamllabelsonly…/storefront_searcharray_train_dsat_daily_pulse_e2e.pipeline.yamllabels, annotations…/l1_tangentable_pipeline.yamlannotations, labels…/experiments/legacy_config/real_upi_searcharray_smoke/pipeline.yamllabels, annotationsEvery key and value is a
str. The corpus is 4:1 labels-first, so that is the canonical order here;l1_tangentable_pipeline.yamlis the one shape this will not reproduce byte-for-byte.String-only, per the schemas. Both the dehydrated and the pipeline schema type
metadata.labelsasadditionalProperties: {"type": "string"}— strictly narrower thanmetadata.annotations, which the dehydrated schema also lets be a number, boolean or null. The value policy follows the schema rather than a guess.Purely descriptive — no readers. Grepping
labelsacrosstangle_clisource returns nothing: no compiler, hydrator or client code reads them. They appear only in the two schemas and in the generatedMetadataSpec.labels, and the OpenAPI spec has no label query parameters, so there is not even server-side filtering today.No task-level labels. Zero occurrences in the corpus, and neither schema gives task metadata a
labelsproperty — task metadata is annotations-only. Deliberately out of scope rather than added silently.How
PIPELINE_LABELS_POLICYinschema_validation.py, reusing the existingcheck_annotationsmachinery.AnnotationPolicy.labelnames the surface, so diagnostics readlabels value for key 'team' must be a string; got int.reject_reserved_key_prefixis deliberately off:system/is documented as reserved for Tangle's own annotations, and nothing reserves a label prefix, so enabling it would refuse a document both schemas accept. Pinned by a test, and one line to flip if that changes.InvalidPipelineLabelsErrorso a caller can tell which metadata block it got wrong.labelsthreadedPipelineFn→GraphBuilder→emit_pipeline, written beforeannotationswhichever order the keywords were passed.subpipelinechildren, consistent with root annotations, and not part of compile identity or the overrides fingerprint (that fingerprint covers config overrides only).No
pipeline_labelscompile keyword is added. Annotations got one because the varying part of that block comes from a caller's per-environment config; nothing sets labels that way,pipeline_annotationsis not even called from tangle-deploy, and the rules are already shared, so adding one later is small.Failure modes
labels=compiles byte-for-byte as before, includinglabels={}andlabels=None.annotationsfirst inmetadatacannot get it. One corpus pipeline is written that way; the block's meaning is unaffected.Review focus
system/prefix legal on labels.pipeline_labelsis the right call.Tophatting
compiles to the corpus block exactly:
Checklist
tests/test_pipeline_labels.py: each of the four distinct corpus label blocks round-tripping, the rendered YAML text, labels-before-annotations regardless of keyword order, author key order preserved, no-labels byte-identity across omitted/{}/None, two subpipeline cases, and the validation matrix.child-<hash8>token masked — child sidecar hashes fold in the output directory, so a naive cross-directory name comparison fails for an unrelated reason. With the hash masked, the only delta is the parent's metadata block.git diff --checkclean.uv.lockhand-edited on thetangle-cliself-entry only (one line);uv lockwas not run.