Skip to content

merge: sota-sprint-arabic — r7 canonical teacher verdicts + stable anchors (fixes dangling citations in released metadata) - #60

Merged
ronaldtse merged 113 commits into
mainfrom
merge/sota-sprint-arabic
Aug 30, 2026
Merged

ronaldtse merged 113 commits into
mainfrom
merge/sota-sprint-arabic

Conversation

@ronaldtse

Copy link
Copy Markdown
Contributor

Why

The released ara-diac-2.0 artifacts cite anchors in rababa's
docs/RESULTS.md — #r7-verdict-table-sadeeddiac-25-2026-08-28 and
#r7-ood-verdict-table-wikinews-2024-multiref-2026-08-28 (metadata
inside the release zips, models.yaml entries, ml-models RESULTS.md).
Those anchors do not exist on main — the r6/r7/r8 verdict tables
live only on sota-sprint-arabic, which advanced after PR #55 was
rebase-merged (so the lines share no commits and every co-touched
file conflicts add/add). Every citation is currently dangling.

What this merges

The Arabic campaign's canonical state: r6 morph-aux verdict
(2.5793/1.5317), r7 news-domain canonical teacher (2.2864/1.3343
ID; 17.38/11.83 WikiNews-2024 multiref OOD), r8 IPA-aux controlled
negative (2.6588), r6 beam-4 negative, Hebrew s46/s47, Urdu d1/d2
comparables, RAG homograph negative, plus the r6/r7/r8 train/eval
scripts and the resumable r6 eval generation.

Conflict resolution (5 files, all add/add)

Branch side verified as a superset in every case:

  • docs/RESULTS.md — section-level and sorted-line comm shows zero
    main-only content (s45's row lives in the branch's newer Hebrew
    table). Plus one fix: the OOD header's em dash rendered as a
    double hyphen and kept multi-ref, while the shipped metadata
    cites ...table-wikinews-2024-multiref-...; the header is now
    phrased so GitHub's rendered anchor matches the released citation
    exactly (verified against the rendered HTML on this branch).
  • train_arabic_r6.py — direct diff is only the branch's resumable
    eval generation (eval_progress.jsonl), a superset of main's plain
    loop.
  • docs/MODELS.md, docs/DISTILL-SOURCE-PROMPT.md,
    docs/SOTA_BENCHMARK.md — branch carries the r6-canonical tables
    superseding main's r5-era text.

Verification

  • Merge is clean beyond the five resolutions; git diff --cached
    audited per file.
  • Rendered anchors on this branch match both metadata citations
    (curl of the GitHub-rendered file).
  • No attribution trailers; staged set reviewed file-by-file.

ModernCharTransformer: RoPE + SDPA Flash + mHC residual + AttnRes + RMSNorm + SwiGLU. 113M params at 12L/768d. Selected via cfg.model.arch=modern. Multi-task seg head optional.

Optimizer: MuonAdamWHybrid (Muon Newton-Schulz for 2D weights, AdamW for 1D) + qk_clip_ weight-rescaling callback with anneal. Prevents attention-logit explosion during from-scratch pretrain.

Decoding: trie-constrained beam decoder with per-word exact search. scripts/build_lexicon.py builds the word-to-haraqat-sequences JSON.

Data: combined corpus (GPLv2 Tashkeela-full + Sadeed HF + QCRI EMNLP 2025) built by fetch_data on first call. Graceful fallback when HF_TOKEN unset.

Idempotency: training/resume.py auto-detects latest epoch checkpoint. train_all.py skips done stages via _status.json on checkpoints volume. scripts/status.py queries Modal volumes.

Configs: rababa_arabic_pro{,_pretrain}.yaml → arch=modern, max_len=512, optimizer=muon, with_seg_head=true, root=/datasets/arabic-combined.

References: arXiv:2606.19348 (DS V4), 2607.24653 (Kimi K3), 2512.24880 (mHC), 2507.20534 (MuonClip).
…wnload

Modal secret 'huggingface' was registered with HF_TOKEN. fetch_data now reads it via env var to authenticate the Sadeed_Tashkeela download.
Pretrain failed because MuonAdamWHybrid is not a torch.optim.Optimizer. WarmupCosine is now a duck-typed scheduler; GradScaler is skipped for Muon (bf16 autocast is enough).

Add run_sota_pipeline + sota_pipeline entrypoint: fetch -> pretrain -> train -> export ONNX/TFLite entirely on Modal via .remote() chaining. Idempotent stage skips + volume status. Survives --detach disconnect.

Fix Tashkeela-full layout discovery (tashkeela_full_train/ subdirs).
Usage: python scripts/status.py [--watch|--json] [--task rababa_arabic_pro]. Reports Modal app state, stage status JSON, pipeline log, per-epoch checkpoints, and exported artifacts.
…len 512

ModernMultiHeadCharTransformer mirrors ModernCharTransformer's encoder
(RoPE + SDPA + mHC + AttnRes + RMSNorm + SwiGLU) but with a ModuleList
of per-category linear heads for niqqud/dagesh/sin.

Encoder weights are key-compatible with ModernCharTransformer — a single
pretrain checkpoint can fine-tune into either Arabic single-head or
Hebrew multi-head.

rababa_hebrew{,_pretrain}.yaml now use:
  arch=modern_multi_head, max_len=512, optimizer=muon (MuonAdamWHybrid)

Smoke-tested: forward + backward + Muon step OK, 2.4M params at 384d/6L.
- 12L/768d/113M was too big for single-A100 pretrain on 1.7M lines (0
  checkpoints in 35min before death). 6L/512d/40M keeps the modern stack
  but is tractable.
- sota_pipeline --force now only wipes arabic-combined for Arabic tasks.
  Hebrew force-wipe was destroying Arabic corpus when both pipelines
  ran in parallel.
…on't interfere

Both Arabic and Hebrew pipelines write to the same /checkpoints/_status.json.
Without task-keying, Hebrew's 'pretrain done' marker caused Arabic's pretrain
to skip in non-force mode. Keys are now 'rababa_arabic_pro:pretrain' etc.
ByT5/seq2seq training paths alongside the char-level encoder: Muon
optimizer variants (AdaMuon, NorMuon, HTMuon, Spectral Cap), ResFormer,
MoE, ELECTRA pretraining, curriculum sampler, EMA, SAM, multi-seed and
distillation harnesses, plus eval scripts for DictaBERT, Nakdimon and
dNIKUD baselines. Arabic 0.99% DER, Hebrew 17.46% DER (beam 4).
RESULTS.md as ground truth for Arabic (0.99% DER) and Hebrew (17.46%
DER, DictaBERT 35.63% on same test), three LaTeX papers (arabic,
hebrew, umbrella), and the TODO.publish checklist.
ByT5-base run-002 (full 1.42M-line corpus, 2 epochs) to beat Claude's
1.39 DER on SadeedDiac-25. Checkpoints commit to the Modal volume at
every save so preemption no longer discards hours of training; EVAL_DONE
marker makes relaunch-after-completion a no-op.
Rejection-sampling fine-tuning on ByT5 r2 (TODO.research/12): sample K=4
per prompt, keep letter-exact-DER winners over greedy, SFT on winners.
Selection on the frozen private dev split (1,372 lines, sha256-pinned,
byte-identical to r2's held-out val); SadeedDiac-25 measured once at
the end. Self-fires when r2's EVAL_DONE marker appears; per-iter volume
commits + markers make preemptions resume cleanly.
Steers RAFT iterations: quantifies how much residual DER sits in the
word-final iʿrāb zone and which haraqat get confused, from the eval CSV.
Misraj's public corpus leaks the SadeedDiac-25 benchmark (122 exact
paragraphs + ~1k near-dup lines found via stride-1 60-char shingles).
r3 continues r2 on the decontaminated copy (1M) + MSA replay (150k) to
close the classical-Arabic domain gap behind r2's residual errors.
Preemptions every ~2h kept killing the ~4h iter-1 sampling before the
iteration marker existed, restarting from zero every time. Winners now
persist to the volume every 25 batches (with commit) and resume from
the saved prompt index.
r3 lands 2.8429/1.7589 (best non-frontier on SadeedDiac-25). Found eval
truncation: 57/1,200 preds cut at 1024B — windowed eval gives the
apples-to-apples number. RAFT now targets r3 (run-002) with mid-
sampling preemption resume.
Diacritized output is 1.4-1.6x input bytes; max_new_tokens=WINDOW
truncated windows mid-word (345-letter input, 200-letter pred).
Now WINDOW*2, plus SequenceMatcher haraqat projection onto input
letters: 759/1200 word-count mismatches -> 0, zero evaluator skips.
Nested imports made json function-local; sampling state crashed on
first access. Module-level import only.
Zero-skip windowed protocol: 2.8126/1.6877 DER, all 1,200 scored.
Papers updated with the contamination finding (122 verbatim + ~1k
near-dup benchmark paragraphs in Misraj's public corpus) and the
survivorship-bias protocol lesson (1.82 was skip-artifact).
Temperature 0, thinking disabled (reasoning mode burns minutes per
long paragraph; plain completion matches the published LLM protocol).
Checkpointed per-row; reports raw + projected zero-skip protocols.
Clean reproduction (temp 0, plain completion, 1200/1200 responses,
5 long-paragraph retries). The 2026 frontier sits at 2.51 DER, not
the published Claude-3.7 1.39; our 580M r3 trails it by ~0.3 DER and
splits metrics on the zero-skip protocol.
Distillation rejected on principle: student ceiling = teacher errors,
one systematic error poisons the chain, and gold-filtering makes the
teacher redundant. Instead: join line-split book text into ~1400-byte
paragraph units so iʿrab gets inter-sentence context, eval at the
same window with zero-skip projection.
Both GPU labelers now append per-item progress to the volume (kept
AND dropped items recorded) and commit every ~100 lines, so an app
stop costs at most one batch instead of the whole sweep — same
pattern as the eval_progress fix.
…lled attached runs; fire-and-forget removes the local dependency
…sumed run completed 13,986 news units + 400 gold-2014 lines
…al); r7 entrypoint spawn

r8 differs from r6 in exactly one variable (aux stream renders the same
r5-units as broad-phonemic IPA instead of morphology). Result on the full
1,200-paragraph windowed zero-skip harness: r8 2.6588/1.5783 vs r6
2.5793/1.5317 vs r5 2.6775/1.5965. IPA-stream probe CER 0.0230 (EM
62/200) proves the second projection was learned — the comparison is not
confounded by a failed aux task. Phonemic supervision is not the active
ingredient; lexical-morphological knowledge is. r6 stays the teacher.

r7 (news-domain) launched init from run-006-morph; entrypoint switched to
spawn() (the r8 disconnect lesson). 19,716 steps ETA ~15.7h.
-0.29pp over r6 (2.5793) on the full 1,200-paragraph windowed zero-skip
harness. The news mix improved in-domain substantially. Canonical-
teacher promotion pending the WikiNews-2024 OOD check (running via the
auto-launched actor; gate: beat r6's 19.82/12.46).
…promoted to canonical teacher

Full sweep: ID 2.2864 vs 2.5793, OOD 17.38/11.83 vs 19.82/12.46. The
r5-era paragraph-specialization trade-off is fully erased. Future
student distillations take run-007-news/best as teacher.
…resumable r6 eval

Brings the Arabic campaign's canonical state to main: r6 morph-aux
verdict (2.5793/1.5317), r7 news-domain canonical teacher (2.2864/
1.3343 ID; 17.38/11.83 WikiNews OOD), r8 IPA-aux controlled negative,
Hebrew s46/s47, Urdu d1/d2 comparables, RAG homograph negative, and
the stable verdict-table anchors that the released ara-diac-2.0
metadata cites (docs/RESULTS.md#r7-verdict-table-sadeeddiac-25-2026-
08-28 etc.) — those anchors were dangling: main never contained them.

Conflict resolution (all add/add; PR #55 was rebase-merged so the
lines diverged): the branch side is a verified superset for every
conflicted file — RESULTS.md section/line-level comm shows zero
main-only content (s45's row lives in the newer Hebrew table); the
trainer takes the branch's resumable eval generation (eval_progress.
jsonl superset of main's plain loop); MODELS/DISTILL-PROMPT/SOTA
tables carry the r6-canonical state that supersedes main's r5-era
text.
GitHub renders the em-dash header as
r7-ood-verdict-table--wikinews-2024-multi-ref-2026-08-28 (double
hyphen, hyphen kept in multi-ref), but ara-diac-2.0.metadata.yaml —
already shipped inside the release zip — cites
r7-ood-verdict-table-wikinews-2024-multiref-2026-08-28. Rephrase the
header with only proven-slug behaviors (parens/commas strip, single
spaces to single hyphens) so the document matches the released
contract; re-cutting the release to fix the citation would cost more
than the header change.
@ronaldtse
ronaldtse merged commit b9d29c2 into main Aug 30, 2026
4 of 10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant