fix: boundary uses Ruby's Unicode Word property; some() repetition - #16
Merged
Merged
Conversation
Probed directly against Ruby: \w is ASCII-only there, but \b does
not follow \w — it is Unicode-aware over the Word property, which
counts combining MARKS as word characters. ك|kasra and क|anusvara
are no boundary in Ruby; Python's \b (built on \w, which excludes
marks) saw one, so word-final rules fired wrongly: kṁganā for
kaṁganā (un-mar schwa deletion through the subst path, which also
used raw \b), and the previous combining-marks-only patch was
itself wrong for Latin+mark junctions. The boundary now compiles
over a word class of \w plus all 290 Mark ranges plus Join_Control,
in both the expression layer and the subst-family renderer.
some(X), the one-or-more repetition, was an unsupported construct;
it now renders in both paths (un-mar uses from some("\U") +
"0939"). Its max_length is the inner item's, like Ruby's Repeat.
Direct corpus sweep: 98 failures / 27 maps -> 46 / 14 — the
Devanagari (un-hin/mar/nep, alalc-hin/ori), Thaana (alalc-div x2,
bgnpcgn-div, mv-div) and remaining Arabic families, and iso-mal,
all fully healed.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follows #15. Probed directly against Ruby this time — and the model that PR #11 reasoned from was wrong in an important way:
\wis ASCII-only (سdoes not match/\w/).\bdoes not follow\win Ruby: it is Unicode-aware over the Word property (letters + marks + digits + connectors). Soك|kasraandक|anusvaraare no boundary in Ruby, while Python's\b— built on a\wthat excludes combining marks — sees one.That divergence hit both renderers: un-mar's schwa-deletion rules (
क after boundary→ bare "k") fired at क|ं givingkṁganāforkaṁganā, and the #11 patch (combining-marks-only word class) would itself have been wrong at Latin+mark junctions. The boundary now compiles over a word class of\w+ all 290 Mark ranges (generated) + Join_Control, in both the expression layer and the subst-family renderer that was still emitting raw\b.Also:
some(X)(one-or-more) was an unsupported construct; it now renders in both paths (un-mar:from some("\U") + "0939"), max_length = inner item like Ruby'sRepeat.Direct corpus sweep: 98 failures / 27 maps → 46 / 14 — Devanagari (un-hin/mar/nep, alalc-hin/ori), Thaana (alalc-div ×2, bgnpcgn-div, mv-div), the remaining Arabic maps, and iso-mal all fully healed. Suite 44 passed / 1 xpassed.