Skip to content

Describe and check optional multi-identifier matching - #1

Draft
pid1 wants to merge 16 commits into
mainfrom
identifiers
Draft

pid1 wants to merge 16 commits into
mainfrom
identifiers

Conversation

@pid1

@pid1 pid1 commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

SPEC.md gains §5.8, which describes matching one reading position by several document identifiers, as proposed in koreader/koreader-sync-server#55. That pull request is open and not merged, and the section says so and is pinned to its branch commit.

The twenty [K-ID-…] requirements are MAY, under the optional feature identifiers added in #2. There is no flag. The verifier pushes one position naming identifiers and reads the answer:

Optional features
  no   [identifiers] multi-identifier document matching
         SPEC.md §5.8
         a push naming identifiers answered 200 {"document":"…","timestamp":…}, with no `match` field

A match field means the feature is implemented and the family is checked in full. Its absence means the family is recorded as SKIP. A server that does not implement §5.8 stays conformant; a server that does is held to all of it, and a failure there does not change the exit status.

Measured

2026-09-23, with the profile flags each server's design calls for.

Server Feature Passed MUST SHOULD Skipped Verdict
koreader/kosync:latest not found 49 0 2 20 conformant
pid1/tsundoku main not found 47 0 0 22 conformant
koreader-sync-server#55 at 49dbd38, built locally found 71 0 0 0 conformant

The first two are unchanged on every requirement that existed before this branch. The third is the proposal's own implementation, and passes all twenty.

A real client exercises [K-ID-16]: CrossPoint on an Xteink X3 sends metadata with "weak": true, and #55 at 32c1edb registers it as an alias. No client has yet sent a write that resolves only through a weak entry, so [K-ID-17] is checked by the verifier alone. On the same device, a copy run through CrossPoint's EPUB Optimizer resolves to the original's record through structure and opens at its position.

Two guards

assert() refuses an identifier that falls under a registered feature's requirement prefix without naming that feature, so one omitted feature: cannot quietly charge a server for an option it never claimed. coverage.mjs fails if a requirement defined under a heading that says proposed belongs to no optional feature at all, which is the case the first guard cannot see.

Also here

The results table gains a Skipped column, and tsundoku's SHOULD-failure count is corrected from 2 to 0. The 2 was its skipped registration assertions, written into the wrong column.

Since this was opened

The section grew two requirements and a type registry, all of them from an implementation disagreeing with the text:

  • [K-ID-8b] — the document need not be the first identifier. Pinning it there forced a filename-matching client to offer its weakest digest first and be matched on it.
  • [K-ID-12b] — an identifier ranked above the one that matched is not registered as an alias. Without it, two books a library tagged alike merged permanently: the second overwrote the first's position and correcting the tagging did not separate them, because the second book's own content digest had been glued to the other record on the way through. Measured, then fixed in #55.
  • [K-ID-5b] — a client MUST order its identifiers strongest first. [K-ID-12b] is expressed in terms of list position, so a weakest-first list defeats it entirely. Measured against the fixed server; a server cannot check this, so §12.4 says so.
  • [K-ID-14] and the type registry — type is opaque to the server but not to clients, since progress_match decides whether one device follows another's position. The registry gives content, structure, filename and metadata a recipe.
  • [K-ID-16] and [K-ID-17] — a PUT entry may be marked weak, and a write resolving only through a weak entry does not adopt that record. metadata is the only identifier that survives a spine change, so it is the only one that reaches a different edition or a re-chunked conversion; it is also the only one that can match two different works. Marking it weak makes it seed a reader without claiming a record. The flag is body-only, so the ids grammar is untouched.
  • A bug in coverage.mjs — it counted any id sharing a prefix as a refinement, so K-ID-16 looked covered by an assertion for K-ID-1. A refinement suffix is never a digit. Fixing it exposed two pre-existing gaps: K-PUT-202 was asserted and never defined, and K-FLD-1 had neither an assertion nor a reason.

The structure recipe is pinned to the point where two implementations written against it independently agree bit for bit, including on entity handling:

Text/a%20b.xhtml\nText/c&d.xhtml\n../Text/e.xhtml   47 bytes  e07ad0e2e24fbaa64b0c40a8b1ebb13f
urn:a&b:1\nch1.xhtml                                19 bytes  fb3ed76af6e07f28456616a77330b19f

🤖 Generated with Claude Code

@pid1
pid1 force-pushed the identifiers branch 2 times, most recently from 80b06e0 to aa37922 Compare September 23, 2026 01:42
pid1 and others added 15 commits September 27, 2026 13:14
SPEC.md gains §5.8: a reading position matched by several document identifiers,
as proposed in koreader/koreader-sync-server#55. The section is pinned to the
proposal's branch commit and marked as describing an unmerged pull request.

Its eighteen requirements are MAY, under the optional feature `identifiers`.
The verifier pushes one position naming identifiers and reads the answer: a
`match` field means the feature is implemented and the family is checked, its
absence means the family is recorded as SKIP. A server that does not implement
§5.8 stays conformant; one that does is held to all of it.

An assertion whose id falls under a feature's requirement prefix has to name
that feature, and coverage.mjs fails if a proposed requirement belongs to no
feature at all. Together those stop §5.8 reaching MUST and being charged to
every server.

Measured 2026-09-22:

  koreader/kosync:latest              49 passed, 0 MUST, 2 SHOULD, 18 skipped
  pid1/tsundoku main                  47 passed, 0 MUST, 0 SHOULD, 20 skipped
  koreader-sync-server#55 at 0c0f5ad  69 passed, 0 MUST, 0 SHOULD,  0 skipped

The results table also gains a Skipped column, and tsundoku's SHOULD-failure
count is corrected from 2 to 0: that 2 was its skipped registration assertions.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The list is the client's order of preference, so pinning `document` to the first
position makes preference and identity the same thing. KOReader set to match
documents by filename is addressed by the weakest digest it holds, and would
report `progress_match: "filename"` for a copy whose content is byte-identical.

K-ID-8 now asserts the list contains an entry equal to `document`. K-ID-8b
asserts that entry need not be first, that the record is created under
`document`, and that a read naming no identifiers still finds it.

Measured against koreader/koreader-sync-server#55: the new assertion fails on
the branch before the corresponding fix and passes after it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`type` is opaque to the server and not to clients: `progress_match` decides
whether one device follows another's position, so a label without a recipe is
not interoperable. The registry gives `content`, `structure` and `filename`
theirs, and K-ID-14 says an unrecognised label is never sufficient to follow a
position.

`structure` covers no file contents. CrossPoint's EPUB optimizer re-encodes
images, injects a stylesheet into every chapter and rewrites `img src`, so every
entry's bytes change while the spine does not; a digest over entry CRCs fails
against the tool this exists for. The spine href list is also what an xpointer
counts.

There is no `metadata` type. It is the only identifier that can match two
different files, and an alias is never repointed.

K-ID-12b, asserted: an identifier ranked above the one that matched is not
registered. Measured against koreader/koreader-sync-server#55 — the assertion
fails on the branch before that fix and passes after it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The K-ID family is twenty assertions, so a server without the feature skips
twenty. Re-measured against the reference image, tsundoku main, and #55 at
49dbd38.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A hand-rolled parser keeps XML entities in an href and a DOM parser decodes
them; one matches element names literally and the other on the local name.
Neither notices it made a choice.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
K-ID-12b is expressed in terms of list position, so a list that opens with its
weakest identifier matches at the first position and registers everything behind
it. Measured: two books sharing only a weak identifier, offered weakest-first,
merge and stay merged after that identifier is corrected. K-ID-5b makes the
ordering a client MUST; a server cannot check it.

The section's worked example was dated before both of today's rules and taught
the behaviour they forbid. Re-run against 49dbd38: two aliases, not four.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The href example counted seventeen characters as eighteen and thirteen as
sixteen. A reimplementer checking their work against those numbers chases a
phantom.

Also: an empty named identifier falls back, trimming is XML whitespace only, a
manifest item with no href contributes no line, and the rootfile is the first
declared package document rather than the first of any type — which keeps the
digest agreeing with whichever OPF a reader already opens.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The href rule said raw source bytes. Complying with that took expat's
XML_DefaultCurrent and a hand-rolled attribute scan in one implementation, while
the identifier and full-path arrived decoded through the same parser — so the
recipe asked for the harder thing in one place and the natural thing in two, for
no reason. Entity expansion is now uniform.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Attribute names are literal and unprefixed, because an unprefixed attribute is
in no namespace and opf:href would be a different attribute, not the same one.
An element's value is all of its character data, since a parser may split a run
at an entity reference. Only the XML-defined entities are expanded, so a lenient
parser that knows HTML names has no licence to use them. Items outside the
manifest yield no digest.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A fourth implementation cannot derive a filename at all: Readest stores books
by hash and the only name it can reconstruct is its OPF title, so a digest
labelled filename would be the metadata type this section removed, wearing a
different label. K-ID-15 says a client omits what it cannot compute honestly
rather than approximating it, and sends no list when only the document is left.

K-ID-14 gated following a position but not resolving one. A client that resolves
an untrusted position to compare it with the local one gets a number that looks
authoritative and can suppress its own conflict prompt, after which the position
is overwritten with nobody asked.

Also: one spine line is required, duplicate manifest ids take the first, and an
href is not trimmed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
metadata is the only identifier that survives a spine change, so it is the only
one that reaches a different edition or a re-chunked conversion. It is also the
only one that can match two different works. K-ID-16 lets a PUT entry carry
weak: true, and K-ID-17 says a write resolving only through a weak entry does
not adopt that record. A weak identifier seeds a reader; it does not claim a
record. progress_match told a reader not to trust a weak match while nothing
told a writer not to clobber on one.

The flag is body-only, so the ids grammar and anything parsing it are untouched.

coverage.mjs treated any id with a matching prefix as a refinement, so K-ID-16
counted as covered by an assertion for K-ID-1. A refinement suffix is never a
digit. Fixing it exposed two pre-existing gaps: K-PUT-202 was asserted and never
defined, and K-FLD-1 had neither an assertion nor a reason.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Written against the first implementation of K-ID-16 and K-ID-17, which found
all five. Two more servers are being written against the same text.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
K-ID-17 could not tell 'did not adopt' from 'never registered the weak value'.
Its second push finds the shared value only because the first push registered
it, so a server that skipped weak aliases would create the second record, leave
the first alone, and pass vacuously. K-ID-17b reads a third copy that shares
only the weak value and requires it to be seeded from the first record.

Found by the second and third implementations, which both register weak aliases
and both carry their own test for it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three clients implemented the recipe and folded case three ways: full Unicode
twice, ASCII once, because a microcontroller has no case table. An identifier
two clients fold differently is worse than one that folds less, so the fold is
ASCII only. Where the authors come from was also unstated: dc:creator element
text, never opf:file-as, contributors excluded.

A weak entry carrying the document value adopts, because writing to the record
you are addressed by is not claiming another's.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…nd K-ID-8b

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant