Skip to content

feat(scan): shard census records so a census stays pushable - #4

Merged
Solomonic merged 1 commit into
mainfrom
feat/shard-census-records
Sep 15, 2026
Merged

Solomonic merged 1 commit into
mainfrom
feat/shard-census-records

Conversation

@Solomonic

Copy link
Copy Markdown
Contributor

Why now

records.jsonl reached 51.5 MiB at 31,538 servers, past the 50 MiB GitHub warns at. At 2,198 bytes per server the 100 MiB hard push limit lands near 47,700 servers, roughly two months out at the growth the last four censuses recorded.

Census servers records.jsonl
2026-07 14,559 18.7 MiB
2026-08-03 19,804 30.1 MiB
2026-09-09 29,522 47.3 MiB
2026-09-13 31,538 51.5 MiB

The census that crosses 100 MiB cannot be pushed at all, and that gets discovered at publish time with the scan already spent.

What changed

scan writes records/<NN>.jsonl, 64 shards keyed on a FNV-1a hash of the namespace. Re-sharding the real 2026-09-13 census:

sharded 31538 records into 64 shards
largest shard 5.9 MiB   smallest 0.4 MiB

The limit moves from roughly 1.5 censuses away to around 17x the current registry.

Keying on the namespace rather than the record means a namespace is never split, so one operator's servers are always in exactly one file. That is the per-namespace fetch a consumer actually wants, and records/manifest.json maps namespace to shard so they do not have to reimplement the hash.

Nothing about a record changed, only which file it sits in.

Backward compatibility, proved rather than asserted

Censuses published before this keep their single records.jsonl. ReadRecords prefers shards and falls back, so every historical edition stays comparable, and --compare takes a census dir, a records/ dir, or a single file.

The decisive check was running the landing build's vendor step against the real 2026-09-13 census in both layouts:

IDENTICAL servers.json       8ed942f17de08f20
IDENTICAL state-of-mcp.json  4b107f47246a41c7

Byte-identical out of both, with a mixed root (4 legacy censuses + 1 sharded).

One defect found while testing

The manifest first counted len(results) rather than distinct names. A test that wrote 120 records over 119 distinct names caught it. The manifest total is what a consumer trusts as the census size, so it now counts distinct names and always equals what a reader of the shards sees.

Tests

  • a namespace is never split across shards
  • the shard mapping is pinned, since it is published and changing it breaks cached per-namespace URLs
  • the hash spreads across shards
  • a legacy records.jsonl still reads back
  • the sharded layout wins when both exist, so a re-scan does not double-count
  • manifest counts equal the shards, and every namespace it maps resolves to the shard the hash picks
  • --compare accepts all four path forms

go test ./..., go vet, and golangci-lint run ./... all clean.

🤖 Generated with Claude Code

https://claude.ai/code/session_019GG6QKS12HcqTh2UNRdoGZ

records.jsonl reached 51.5 MiB at 31,538 servers, past the 50 MiB GitHub
warns at. At 2,198 bytes per server the 100 MiB hard push limit lands near
47,700 servers, roughly two months out at the growth the last four censuses
recorded. The census that crosses it cannot be pushed at all, and that would
be discovered at publish time, with the scan already spent.

scan now writes records/<NN>.jsonl, 64 shards keyed on a FNV-1a hash of the
namespace. Re-sharding the real 2026-09-13 census gives a largest shard of
5.9 MiB against 51.5, so the limit moves from roughly 1.5 censuses away to
around 17x the current registry.

Keying on the namespace rather than the record means a namespace is never
split, so one operator's servers are always in exactly one file. That is the
per-namespace fetch a consumer actually wants, and records/manifest.json maps
namespace to shard so they do not have to reimplement the hash.

Nothing about a record changed, only which file it sits in.

Censuses published before this keep their single records.jsonl and are still
read: ReadRecords prefers shards and falls back, so every historical edition
stays comparable and --compare takes a census dir, a records/ dir, or a
single file. Proved by running the landing build's vendor step against the
real 2026-09-13 census in both layouts: byte-identical servers.json and
bundle out of both.

The manifest counts distinct names rather than lines written, so the total a
consumer trusts as the census size always equals what a reader of the shards
sees.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019GG6QKS12HcqTh2UNRdoGZ
@Solomonic
Solomonic merged commit 1c0b9f1 into main Sep 15, 2026
5 checks passed
@Solomonic
Solomonic deleted the feat/shard-census-records branch September 15, 2026 09:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant