feat(scan): shard census records so a census stays pushable - #4
Merged
Merged
Conversation
records.jsonl reached 51.5 MiB at 31,538 servers, past the 50 MiB GitHub warns at. At 2,198 bytes per server the 100 MiB hard push limit lands near 47,700 servers, roughly two months out at the growth the last four censuses recorded. The census that crosses it cannot be pushed at all, and that would be discovered at publish time, with the scan already spent. scan now writes records/<NN>.jsonl, 64 shards keyed on a FNV-1a hash of the namespace. Re-sharding the real 2026-09-13 census gives a largest shard of 5.9 MiB against 51.5, so the limit moves from roughly 1.5 censuses away to around 17x the current registry. Keying on the namespace rather than the record means a namespace is never split, so one operator's servers are always in exactly one file. That is the per-namespace fetch a consumer actually wants, and records/manifest.json maps namespace to shard so they do not have to reimplement the hash. Nothing about a record changed, only which file it sits in. Censuses published before this keep their single records.jsonl and are still read: ReadRecords prefers shards and falls back, so every historical edition stays comparable and --compare takes a census dir, a records/ dir, or a single file. Proved by running the landing build's vendor step against the real 2026-09-13 census in both layouts: byte-identical servers.json and bundle out of both. The manifest counts distinct names rather than lines written, so the total a consumer trusts as the census size always equals what a reader of the shards sees. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019GG6QKS12HcqTh2UNRdoGZ
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why now
records.jsonlreached 51.5 MiB at 31,538 servers, past the 50 MiB GitHub warns at. At 2,198 bytes per server the 100 MiB hard push limit lands near 47,700 servers, roughly two months out at the growth the last four censuses recorded.The census that crosses 100 MiB cannot be pushed at all, and that gets discovered at publish time with the scan already spent.
What changed
scanwritesrecords/<NN>.jsonl, 64 shards keyed on a FNV-1a hash of the namespace. Re-sharding the real 2026-09-13 census:The limit moves from roughly 1.5 censuses away to around 17x the current registry.
Keying on the namespace rather than the record means a namespace is never split, so one operator's servers are always in exactly one file. That is the per-namespace fetch a consumer actually wants, and
records/manifest.jsonmaps namespace to shard so they do not have to reimplement the hash.Nothing about a record changed, only which file it sits in.
Backward compatibility, proved rather than asserted
Censuses published before this keep their single
records.jsonl.ReadRecordsprefers shards and falls back, so every historical edition stays comparable, and--comparetakes a census dir, arecords/dir, or a single file.The decisive check was running the landing build's vendor step against the real 2026-09-13 census in both layouts:
Byte-identical out of both, with a mixed root (4 legacy censuses + 1 sharded).
One defect found while testing
The manifest first counted
len(results)rather than distinct names. A test that wrote 120 records over 119 distinct names caught it. The manifest total is what a consumer trusts as the census size, so it now counts distinct names and always equals what a reader of the shards sees.Tests
records.jsonlstill reads back--compareaccepts all four path formsgo test ./...,go vet, andgolangci-lint run ./...all clean.🤖 Generated with Claude Code
https://claude.ai/code/session_019GG6QKS12HcqTh2UNRdoGZ