Skip to content

fix: boundary uses Ruby's Unicode Word property; some() repetition - #16

Merged
ronaldtse merged 1 commit into
mainfrom
fix/boundary-word-property
Oct 1, 2026
Merged

ronaldtse merged 1 commit into
mainfrom
fix/boundary-word-property

Conversation

@ronaldtse

Copy link
Copy Markdown
Contributor

Follows #15. Probed directly against Ruby this time — and the model that PR #11 reasoned from was wrong in an important way:

  • Ruby's \w is ASCII-only (س does not match /\w/).
  • But \b does not follow \w in Ruby: it is Unicode-aware over the Word property (letters + marks + digits + connectors). So ك|kasra and क|anusvara are no boundary in Ruby, while Python's \b — built on a \w that excludes combining marks — sees one.

That divergence hit both renderers: un-mar's schwa-deletion rules (क after boundary → bare "k") fired at क|ं giving kṁganā for kaṁganā, and the #11 patch (combining-marks-only word class) would itself have been wrong at Latin+mark junctions. The boundary now compiles over a word class of \w + all 290 Mark ranges (generated) + Join_Control, in both the expression layer and the subst-family renderer that was still emitting raw \b.

Also: some(X) (one-or-more) was an unsupported construct; it now renders in both paths (un-mar: from some("\U") + "0939"), max_length = inner item like Ruby's Repeat.

Direct corpus sweep: 98 failures / 27 maps → 46 / 14 — Devanagari (un-hin/mar/nep, alalc-hin/ori), Thaana (alalc-div ×2, bgnpcgn-div, mv-div), the remaining Arabic maps, and iso-mal all fully healed. Suite 44 passed / 1 xpassed.

Probed directly against Ruby: \w is ASCII-only there, but \b does
not follow \w — it is Unicode-aware over the Word property, which
counts combining MARKS as word characters. ك|kasra and क|anusvara
are no boundary in Ruby; Python's \b (built on \w, which excludes
marks) saw one, so word-final rules fired wrongly: kṁganā for
kaṁganā (un-mar schwa deletion through the subst path, which also
used raw \b), and the previous combining-marks-only patch was
itself wrong for Latin+mark junctions. The boundary now compiles
over a word class of \w plus all 290 Mark ranges plus Join_Control,
in both the expression layer and the subst-family renderer.

some(X), the one-or-more repetition, was an unsupported construct;
it now renders in both paths (un-mar uses from some("\U") +
"0939"). Its max_length is the inner item's, like Ruby's Repeat.

Direct corpus sweep: 98 failures / 27 maps -> 46 / 14 — the
Devanagari (un-hin/mar/nep, alalc-hin/ori), Thaana (alalc-div x2,
bgnpcgn-div, mv-div) and remaining Arabic families, and iso-mal,
all fully healed.
@ronaldtse
ronaldtse merged commit 121423b into main Oct 1, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant