Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
113 commits
Select commit Hold shift + click to select a range
c8b56e2
feat(arabic): Day 1 SOTA sprint — modern encoder + Muon + trie decoder
ronaldtse Aug 7, 2026
92c283b
feat(modal): wire huggingface secret into fetch_data for Sadeed HF do…
ronaldtse Aug 7, 2026
cbba8a6
fix: MuonAdamWHybrid scheduler/GradScaler + server-side sota_pipeline
ronaldtse Aug 7, 2026
ac85751
fix(modal): cap sota_pipeline timeout at Modal max 24h
ronaldtse Aug 7, 2026
ae0af29
fix(modal): commit datasets volume after fetch_data so pretrain can s…
ronaldtse Aug 7, 2026
19ac5a2
feat(scripts): status.py — pull all 5 progress signals into one report
ronaldtse Aug 7, 2026
e4c9b10
fix(scripts): parse modal app list table correctly
ronaldtse Aug 7, 2026
e7d364a
feat(hebrew): apply DS4/K3 modern stack — mHC + AttnRes + Muon + max_…
ronaldtse Aug 7, 2026
9de9272
fix(modal): --force wipes old run-001 dirs and old artifacts so resum…
ronaldtse Aug 7, 2026
6ed0be3
fix(modal): commit volume after force-wipe so next container sees cle…
ronaldtse Aug 7, 2026
e50ebd0
fix: shrink arabic_pro to 6L/512d (40M) + guard force-wipe by task
ronaldtse Aug 7, 2026
a9d1146
fix(modal): key stage status by (task, stage) so parallel pipelines d…
ronaldtse Aug 7, 2026
f5330de
feat(scripts): status.py filters stages by task prefix for parallel p…
ronaldtse Aug 7, 2026
80ecd8e
feat: modern SOTA training stack for Arabic + Hebrew diacritization
ronaldtse Aug 14, 2026
d3ca746
docs: publish results, papers, and TODO.publish execution log
ronaldtse Aug 14, 2026
45c82a2
docs: TODO.runtime-arch — training-to-usage pipeline work orders + ag…
ronaldtse Aug 15, 2026
3be1297
docs(research): RL with verifiable rewards — GLM-5.3 playbook applied
ronaldtse Aug 16, 2026
3461072
feat: Arabic ByT5-base trainer with per-save volume commits
ronaldtse Aug 16, 2026
01d78a5
feat: RAFT verifiable-reward RL + frozen private dev for Arabic
ronaldtse Aug 16, 2026
0adbc85
feat: haraqat error analyzer (word-final vs internal, confusion pairs)
ronaldtse Aug 16, 2026
14ff82e
feat: r3 domain-adaptation SFT on decontaminated Misraj corpus
ronaldtse Aug 16, 2026
b80ea76
fix: RAFT sampling survives preemption via incremental winner state
ronaldtse Aug 17, 2026
b9b1748
feat: r3 results, windowed eval, RAFT run-002 from r3
ronaldtse Aug 17, 2026
06d4815
fix: windowed eval generation cap + haraqat projection
ronaldtse Aug 17, 2026
8937c8e
fix: RAFT UnboundLocalError from nested json import
ronaldtse Aug 17, 2026
aa4b315
docs: final r3 windowed numbers in RESULTS + papers
ronaldtse Aug 17, 2026
1badb46
feat: clean GLM (z.ai) eval on SadeedDiac-25
ronaldtse Aug 17, 2026
b5f5b4b
results: GLM-5.2 verified on SadeedDiac-25 — 2.5060/1.5537
ronaldtse Aug 17, 2026
0f1979f
docs: GLM-5.2 verified frontier row in paper table
ronaldtse Aug 17, 2026
7d3fb5c
feat: r5 paragraph-context training (close the context gap)
ronaldtse Aug 17, 2026
a79e1ee
feat: GRPO trainer — gold-reward RL with negative gradients
ronaldtse Aug 17, 2026
81cbbaa
perf: cache joined paragraph units on the volume
ronaldtse Aug 17, 2026
4568950
docs: model manifest for the distillation agent
ronaldtse Aug 17, 2026
2278c9e
fix: RAFT/GRPO sized to survive the ~2h preemption cadence
ronaldtse Aug 17, 2026
e5eb9be
perf: r5 checkpoints every 1000 steps
ronaldtse Aug 17, 2026
feeec3c
fix: RAFT image missing accelerate — Trainer crash after sampling
ronaldtse Aug 18, 2026
a9d64ae
fix: r5 OOM at 40%% — batch 3 + accum 10
ronaldtse Aug 18, 2026
152c682
docs: distillation source-model usage prompt
ronaldtse Aug 18, 2026
2b84efe
fix: r5 OOM is deterministic — batch 2, unit cap 1450
ronaldtse Aug 18, 2026
aa1f248
feat: qalsadi morphological labeling for the aux-task lever
ronaldtse Aug 18, 2026
ec762d2
feat: GTPO entropy-weighted credit assignment in GRPO
ronaldtse Aug 18, 2026
78dec73
docs: arXiv sweep Aug 2026 — GTPO applied, HomoRich/DIVRIT checked
ronaldtse Aug 18, 2026
c8a9695
fix: two-tier morph labeling — exact tags + coarse fallback
ronaldtse Aug 18, 2026
fdf4e6e
fix: RAFT missing pyarabic; morph fallback on 'Not exists' analyses
ronaldtse Aug 18, 2026
f209c92
feat: multi-reference WikiNews eval (QCRI EMNLP 2025 protocol)
ronaldtse Aug 18, 2026
8e68429
fix: stem-mode scoring excluded from totals, not auto-credited
ronaldtse Aug 18, 2026
573f9d0
docs: r3 WikiNews-2024 multi-ref verdict (19.99/12.60 WER/DER)
ronaldtse Aug 18, 2026
b127779
fix: r5 OOM at 9048 — batch 1 / accum 30 (same effective batch)
ronaldtse Aug 18, 2026
7bb48ca
docs: RAFT run-002 closed — flat on benchmark (2.8515/2.8308 vs r3 2.…
ronaldtse Aug 18, 2026
3b88160
fix: expose eval args through local entrypoint
ronaldtse Aug 18, 2026
7f2ae2f
docs: r5 paragraph-context verdict — 2.6775/1.5965, beats verified GL…
ronaldtse Aug 18, 2026
2a4c213
docs: r5 SOTA in paper tables + MODELS manifest
ronaldtse Aug 18, 2026
7b37f8a
docs: umbrella Arabic headline 2.68/1.60 (r5)
ronaldtse Aug 18, 2026
b6051fb
fix: GRPO OOM — halve seqs per forward, detach entropy weights
ronaldtse Aug 18, 2026
21979c4
docs: r5 WikiNews cross-domain tradeoff (+0.53 WER)
ronaldtse Aug 18, 2026
f3bda48
fix: GRPO OOM root cause — enable gradient checkpointing
ronaldtse Aug 18, 2026
10ffe8a
fix: graded alignment-based der() — binary reward was dead
ronaldtse Aug 18, 2026
99e56af
fix: GRPO throughput — 400 steps, save 50, eval 100
ronaldtse Aug 18, 2026
2d7a5fd
fix: GRPO attention-score OOM — 1x16 micro-batch + SDPA
ronaldtse Aug 18, 2026
3aead24
perf: GRPO right-sized for A100-80GB throughput
ronaldtse Aug 18, 2026
48102da
fix: materialize best/ before final eval on flat curves
ronaldtse Aug 19, 2026
cd56105
fix: GRPO image missing pyarabic (same as RAFT bug)
ronaldtse Aug 19, 2026
211e168
docs: GTPO-GRPO closed flat — third negative RL result
ronaldtse Aug 19, 2026
493bb55
docs: full results audit, blog post, refreshed distill prompt
ronaldtse Aug 19, 2026
fb7a1ba
r6: morphological aux-task multitask (TAG: prefix format, 4x upsample…
ronaldtse Aug 19, 2026
d59cbd1
Hebrew s45: phonikud knesset 1.5M weak-pretrain then gold fine-tune (…
ronaldtse Aug 19, 2026
b2697ea
s45 stage1: batch 16 x accum 4 + expandable segments (batch-64 OOM at…
ronaldtse Aug 19, 2026
0da029c
distill prompt: in-flight successors, Persian v5 rescore verdict, Wik…
ronaldtse Aug 19, 2026
b3ef831
distill prompt: explicit start-now order; Thai gated on scaleup600k v…
ronaldtse Aug 19, 2026
7cc627c
Thai teacher verdict: scaleup600k verified 1.7260% PER (was 2.32); di…
ronaldtse Aug 19, 2026
285de08
distill order: Thai GO (verdict landed); s45 restarted note
ronaldtse Aug 19, 2026
0f6d715
MODELS: Thai 1.7260% PER (scaleup600k) — new best
ronaldtse Aug 19, 2026
e6f4ab2
Hebrew s45 phonikud curriculum verified 16.58% DER (s43 17.46): new t…
ronaldtse Aug 20, 2026
bef1538
r6: save every 300 steps (preemptions were outliving the 1000-step sa…
ronaldtse Aug 20, 2026
d3106b7
feat(improve-models): four workstreams — urdu d1, arabic r7, rag prob…
ronaldtse Aug 20, 2026
29f170c
fix(label_arabic_news): chunk articles + BBC-first fetch
ronaldtse Aug 20, 2026
2d5c627
docs: RAG homograph probe verdict — closed negative (26.07 vs 77.34)
ronaldtse Aug 20, 2026
922f8be
docs: Urdu d1 verdict — CER 6.40 vs 14.77 baseline (2.3x)
ronaldtse Aug 20, 2026
cc451d1
feat(improve-models): 05 khmer g2p + 06 urdu beam eval
ronaldtse Aug 20, 2026
d1c9800
feat(urdu): d2 second epoch; record beam-flat finding
ronaldtse Aug 20, 2026
c81a851
docs: Urdu d2 verdict — CER 5.77 / word_acc 52.47 (new best)
ronaldtse Aug 20, 2026
a29333a
fix(label_hewiki_full): pin transformers 4.38.0 for Dicta batched pre…
ronaldtse Aug 20, 2026
aa546bb
feat(eval): resumable generation — per-window predictions on the volume
ronaldtse Aug 21, 2026
2464745
feat(eval): resumable s46 sweep + shard-parallel hewiki labeling
ronaldtse Aug 21, 2026
6dd0031
docs: r6 verdict — 2.5793/1.5317 beats r5, new canonical Arabic teacher
ronaldtse Aug 21, 2026
e9b969f
docs: distill manifest — r6 is the canonical Arabic teacher
ronaldtse Aug 21, 2026
9e2841b
docs: r6 OOD verdict — 19.82/12.46, dominates r3 and r5 out-of-domain
ronaldtse Aug 21, 2026
7ce689d
feat(labelers): incrementally resumable labeling
ronaldtse Aug 21, 2026
08bf894
fix(s46): save every 500 steps (sweeps were outliving the first 2000-…
ronaldtse Aug 21, 2026
d587bca
docs: defer r7 — r6's OOD sweep absorbed its purpose; data+script pre…
ronaldtse Aug 22, 2026
da0d7dd
feat(public): remaining-work plan + s47 morph labeler
ronaldtse Aug 22, 2026
a3e6d3c
fix(morph-labeler): pin transformers 4.38.0 (Dicta batched-predict co…
ronaldtse Aug 22, 2026
0113d64
docs: Khmer v2 restoration verdict; CLE inquiry email draft
ronaldtse Aug 22, 2026
9df0bbe
feat(s47): morph aux trainer (r6 template transplant); fix labeler pr…
ronaldtse Aug 22, 2026
6b7df23
feat(r6): beam-4 probe eval (windowed zero-skip, resumable)
ronaldtse Aug 22, 2026
78f3c29
chore(labeler): wave-boundary file verification in main()
ronaldtse Aug 22, 2026
286781f
docs(public): §3 morph labeling done (200K knesset lines)
ronaldtse Aug 22, 2026
0edc088
docs: s46 verdict 16.43 DER (new Hebrew best, marginal over s45)
ronaldtse Aug 22, 2026
8bc52bf
fix(beam4): add pyarabic+prettytable to image (evaluator deps)
ronaldtse Aug 22, 2026
668bbc9
docs: r6 beam-4 probe negative (2.5588/1.5379 vs greedy 2.5793/1.5317…
ronaldtse Aug 22, 2026
a790054
docs: canonical manifest + SOTA table refresh (r6/s46/d2/khmer v2); r…
ronaldtse Aug 22, 2026
45f3ec9
docs: s47 morph-aux transplant closed negative (16.53 vs 16.43); s46 …
ronaldtse Aug 23, 2026
fe88891
feat: urdu comparable eval (urd-diac-1.0 vs d2 on one harness)
ronaldtse Aug 23, 2026
fcbd51a
docs: urd-diac-1.0 stays champion (3.74/67.51 vs d2 5.94/51.95 on one…
ronaldtse Aug 23, 2026
3e46bc9
r8: IPA auxiliary-task experiment (phonological-layer claim, controlled)
ronaldtse Aug 26, 2026
df1b4c1
r8: spawn() instead of remote() — two client-side network flaks cance…
ronaldtse Aug 26, 2026
28e03e9
label_arabic_news: spawn() entrypoint — disconnect-immune like r8; re…
ronaldtse Aug 27, 2026
e33f779
docs: r8 IPA-aux controlled verdict (2.6588 — morph aux stays canonic…
ronaldtse Aug 27, 2026
e949c20
docs: r7 news-domain ID verdict — 2.2864/1.3343, new best dedicated
ronaldtse Aug 28, 2026
c59dc3e
docs: r7 OOD verdict — WikiNews 17.3794/11.8273 beats r6 on both; r7 …
ronaldtse Aug 28, 2026
d2d3ae2
docs: stable verdict-table anchors for ara-diac-2.0 provenance
ronaldtse Aug 28, 2026
c083e03
merge: sota-sprint-arabic — r6/r7/r8 verdict tables, stable anchors, …
ronaldtse Aug 30, 2026
ee186d8
docs: OOD header must slug to the anchor the released metadata cites
ronaldtse Aug 30, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
31 changes: 31 additions & 0 deletions TODO.improve-models/01-urdu-byt5-d1.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,31 @@
# 01 — Urdu diacritization d1: ByT5-base + cross-lingual init

## Why
Urdu is our weakest shipped model (14.77% CER, urdu_diacrit/run-001,
custom char encoder trained on 635K cross-lingually machine-labeled
lines). Arabic — same script family, same task shape — sits at 2.68 DER
on ByT5-base with paragraph context. The gap is architecture + teacher
vintage, not task difficulty.

## Plan
1. Data: existing corpus on volume `urdu-diacrit-datasets`:
- `urdu-diacritized/{train,val,test}.txt` (635K machine-labeled —
WEAK labels, teacher-poison rule applies: treat as stage-1 only)
- `urdu-diacrit/*.jsonl` (HF G2P-derived pairs from WO #306)
2. `train_urdu_d1.py` (rababa):
- Init: Arabic r5 teacher `/checkpoints/rababa_arabic_byt5/
run-005-context/best` (ByT5-base) — cross-lingual init gives the
shared-abjad prior instead of starting cold.
- Stage 1: 1 epoch over the weak 635K (line units, byte tokenizer).
- Stage 2: none yet (no gold corpus found — UDD has no verifiable
public URL; do NOT fabricate one). The run's own test split is
machine-labeled too, so the eval is comparative, not absolute.
3. Eval: greedy CER via editdistance on `urdu-diacritized/test.txt`,
identical protocol to the 14.77 number + word-level accuracy.
Target: clearly under 14.77 CER (architecture+init upgrade).
4. Launch detached under supervisor app `rababa-urdu-d1`
(EVAL_DONE guard, checkpoint-resume, volume commits).

## Guards
- No LLM teacher. No RL. Weak corpus is from our Arabic model only.
- Parallel to r6 (separate Modal app + GPU).
46 changes: 46 additions & 0 deletions TODO.improve-models/02-arabic-r7-news-domain.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,46 @@
# 02 — Arabic r7: news-domain adaptation (OOD repair)

## Why
r5 paragraph-context specialized to the SadeedDiac domain and trades
~0.5 DER out-of-domain (WikiNews-2024 multi-ref: r5 20.52/12.72 vs
r3 19.99/12.60). If r6 verifies, case endings improve too — but the
OOD gap needs domain data, not morphology.

## Plan
1. `label_arabic_news.py` (rababa) — runs NOW, parallel to r6:
- Fetch unlabeled Arabic news from HF
(`khalidalt/ultimate_arabic_news`, fallback
`Abdelkareem/arabic-bbc-news`), clean (Arabic-letter fraction,
length, dedupe).
- Pseudo-label with r5 (windowed 1400B zero-skip, greedy, same
harness as eval) → volume `/datasets/arabic-news-r5/`.
- Add GOLD news: `/datasets/wikinews/WikiNews_2014.txt.diac`
(gold-diacritized, different year/documents than the 2024 probe).
- NEVER touch `WikiNews_2024*` — that is the OOD probe.
2. `train_arabic_r7.py` — launch AFTER r6 verdict + labels exist:
- INIT = r6 best if r6 verifies better on SadeedDiac, else r5.
- Mix: cached r5-units (replay, protects ID) + news units
(2014 gold upweighted + r5-pseudo modern news), news ≈ 15-20%
of steps. r5-proven batch/accum, A100-80GB.
- Gates: SadeedDiac-25 windowed zero-skip must not regress beyond
+0.1 DER of the init model; WikiNews-2024 multi-ref must improve.
3. Fold verdict into docs/RESULTS.md + DISTILL-SOURCE-PROMPT.md only
when both gates pass.

## Guards
- Domain adaptation via pseudo-labels is self-training — the replay
majority + gold 2014 news keeps ID anchored; the ID gate is the
hard stop against entrenchment.
- No RL, no LLM labels.

## DEFERRED (2026-08-22)

r6's verified OOD sweep (WikiNews-2024 full 19.82/12.46 — beats r3
AND r5) absorbed this workstream's purpose: there is no OOD deficit
left to repair. r7 as designed costs 20h of A100-80GB for marginal
OOD gains with ID-regression risk, and its concurrent A100 footprint
is exactly what triggers Modal workspace evictions. News labeling
stopped at 5,600/13,987 windows — all committed and resumable on the
volume (label_progress.jsonl), and train_arabic_r7.py is ready with
--init-run run-006-morph. Reopen ONLY if a news-heavy client use
case emerges or Arabic OOD regresses in the wild.
32 changes: 32 additions & 0 deletions TODO.improve-models/03-rag-homograph-probe.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
# 03 — RAG homograph disambiguation probe (Persian first)

## Why
Persian v1 sits at 77.34% SentenceBench homograph (ezafe-normalized) —
above published Homo-GE2PE (76.89) but flat: RL was negative, the
mapped-representation line is closed. Retrieval context is the one
untried lever, and homograph resolution is precisely a context problem.

## Plan
1. `eval_persian_rag_probe.py` (rababa-farsi) — INFERENCE ONLY, no
retraining for the probe:
- Index HomoRich TRAIN homograph sentences (TF-IDF char n-gram
retrieval over sentence contexts, pure numpy/sklearn-free).
- For each SentenceBench test sentence, retrieve top-k (k=3)
contextually nearest TRAIN sentences containing the same
homograph token, format as few-shot prefix:
`<train-sentence> => <diacritized/phonemized> ;` repeated, then
the test sentence. ByT5-small handles prefix context in bytes.
- Baseline vs RAG on the SAME harness (`eval_sentencebench.py`
protocol, ezafe-normalized homograph accuracy + exact match).
2. Decision rule:
- RAG ≥ +1.5pp homograph → invest: cache retrieval index, then a
fine-tune WITH retrieved prefixes (train-time consistency).
- +0.5..1.5pp → cheap inference-time add-on only, document.
- < +0.5pp → close the lever, record negative in docs/RESULTS.md.
3. If Persian moves, port the probe to Hebrew (Nakdimon homographs)
and Arabic (SadeedDiac residual analysis).

## Guards
- Retrieval from TRAIN splits only — zero test contamination; assert
no test sentence appears in the index.
- Teacher stays v1 (RELEASE-FROZEN); probe never modifies it.
31 changes: 31 additions & 0 deletions TODO.improve-models/04-hebrew-s46-scaled-weak.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,31 @@
# 04 — Hebrew s46: scale + diversify the weak stage

## Why
s45 proved the curriculum (+0.88 DER over gold-only: 16.58 vs 17.46)
but its weak stage was knesset only (1.5M lines, single domain —
parliamentary transcripts). hewiki (80K Hebrew-Wikipedia lines) is a
second, encyclopedic domain sitting unlabeled on the volume. Same
lever, turned up: more + more-diverse weak data before the identical
gold FT.

## Plan
1. `label_hewiki_full.py` (rababa) — runs NOW on A10G:
- Full-scale Dicta labeling of `/datasets/hewiki/train.txt`
(80K lines) via the batch recipe from `batch_distill_hewiki.py`
(DictaBERT predict batches), nikud-only targets (strip teamim),
length/Hebrew-fraction filters, 40-char window decontam vs
Nakdimon test.
- Output: `/datasets/hebrew-hewiki-dicta/{train,val}.txt`.
2. `train_hebrew_s46.py` — s45 VERBATIM with one change:
- Stage 1 weak = knesset 1.5M + hewiki-dicta (all of it,
~80K lines ≈ 5% of weak steps — a domain garnish, not a pivot).
- Stage 2 = s43 gold recipe unchanged (hebrew-v4 jsonl, 3 ep,
batch 8, LR 3e-4, warmup 500).
- Eval: beam-4 DER, identical protocol/harness as s45's 16.58.
3. Gates: must beat s45's 16.58 to replace it; otherwise record as
flat and keep s45 (weak-stage ceiling reached).

## Guards
- Dicta labels are weak only — the gold stage corrects (no-teacher-
poison, the exact s45-validated pattern).
- Zero Nakdimon-test contamination (window decontam).
25 changes: 25 additions & 0 deletions TODO.improve-models/05-khmer-g2p-byt5.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
# 05 — Khmer G2P v1: ByT5-small on the 17.9K UNGEGN word pairs

## Why
The only Khmer artifact is the legacy crystalseq transformer
(net-500-epochs.pth, loss 0.143, no held-out eval) plus its fp16 IMF
zip. Every other language in the zoo runs a modern recipe with a
measured number. Khmer is the breadth gap.

## Plan
1. Data: secryst-datasets:/data-khmer-translit/{data_kh,data_rom}.csv
— 17,911 aligned word pairs (validated: no dups, p95 word = 36B,
1 empty rom line filtered). Split 80/10/10 seeded.
2. `train_khmer_g2p.py` (secryst-train repo):
- google/byt5-small, word-level src→tgt, batch 64, LR 3e-4,
30 epochs, best-on-val-loss checkpointing (small-data recipe).
- Eval: held-out test — word accuracy (exact match) + CER via
editdistance. This is the FIRST measured Khmer number; the
crystalseq model gets the same eval for the comparison table.
3. Export to IMF v1 when the number lands (parts contract, byte+3
tokenizer) — replaces the fp16 zip of the legacy model.

## Guards
- Word-level G2P (not diacritization): no LLM-teacher concerns, but
the no-RL/no-LLM standing rules apply anyway.
- Keep the legacy crystalseq artifacts untouched (source files).
21 changes: 21 additions & 0 deletions TODO.improve-models/06-urdu-beam-and-d2.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
# 06 — Urdu d1 follow-up: beam-4 eval, then d2 decision

## Why
d1 (CER 6.40%, word_acc 47.43%) was evaluated GREEDY only. The Hebrew
s45 experience: beam-4 at inference is worth double-digit DER points
on ByT5 diacritizers. Before spending a d2 training run, collect the
free win and re-read the saturation point.

## Plan
1. `eval_urdu_d1_beam.py` (rababa): beam-4 vs greedy on the identical
test protocol (urdu-diacrit/test.jsonl, CER + word_acc, n=11,714).
2. Decision after the number:
- If beam-4 word_acc still < ~55%: d2 with a second epoch at lower
LR (1e-5) from run-001-d1 — cheap, single A100 pass.
- If beam-4 lifts word_acc >= ~55%: declare d1 final for the weak
corpus; the next real gain requires gold Urdu data (none
verifiable — do not chase it).

## Guards
- Same test set, same alignment protocol as the d1 verdict so the
numbers are comparable line-for-line.
68 changes: 68 additions & 0 deletions TODO.public
Original file line number Diff line number Diff line change
@@ -0,0 +1,68 @@
# TODO.public — remaining work after the 2026-08-22 audit

Verdicted and frozen: Arabic r6 (2.5793/1.5317 ID, 19.82/12.46 OOD),
Thai 1.73% PER, Persian v1 (77.34 homograph), Hebrew s45 16.58 (s46
verdict pending), Urdu d2 5.77 CER (within-corpus), Khmer v1 59%
word-acc. No further teacher training on Arabic/Thai/Persian — the
levers below are the only funded work.

## 1. Khmer: reframe as vowel-restoration (EXPANSION, not marginal)
The UNGEGN rule map (secryst/data-khmer-translit/ungegn-khmer-system.yaml)
handles FULL orthography; the real-world problem is Khmer's unwritten
inherent vowels and reduced homographic orthography, which rules
cannot restore. Deliverable: a model that romanizes INCOMPLETE input.
- [ ] Rules-vs-ML audit on the 895-word split (full orthography):
run the yaml map, publish both numbers honestly.
- [x] Vowel-strip augmentation: strip vowel signs/coeng from src,
keep full tgt; train Khmer v2 on (full→full) + (stripped→full).
This is a new CAPABILITY (restoration), measured by a
stripped-input benchmark where rules structurally cannot run.
DONE (run-002-restore): stripped 19.22% word_acc / 39.41 CER
(rules 0.0%); full 53.18/28.27 (−5.8pp vs v1 58.99).
- [ ] Ship decision: map for full orthography, v2 for reduced.
Recommended exactly that; optional parked lever = v1-init
two-stage curriculum to reclaim the 5.8pp.

## 2. Urdu: anchor the claim (zero GPU)
- [ ] Find the official UDD (gold-standard Urdu diacritization)
source — the paper's own release, not a guessed URL. If found:
evaluate d2 on it, publish the comparable number. If not
found: label the claim "within-corpus, no gold available"
everywhere it appears.

## 3. Hebrew s47: transplant the r6 morph template — CLOSED (negative)
Hebrew's 16.58 residual is plausibly morphology (nikud encodes
gender/number/agreement) — the same shape as Arabic's iʿrāb residual
that r6's aux-task fixed. The template is validated; this is its
first cross-language transplant.
- [x] Ground a Hebrew morph labeler (Dicta morph disambiguator /
HebPipe-class POS+morph) — must run offline-ish on Modal.
DONE: dictabert-morph (probe + labels verified).
- [x] Label knesset + hewiki lines with morph tags (A10G, ~3h).
DONE: 200K knesset lines → /datasets/hebrew-morph/train.jsonl
(dictabert-morph, segmented src + per-token tags, 100% keep
after prefix-split fix; hewiki morph not needed for s47 v1).
- [x] train_hebrew_s47.py: s45 recipe + TAG-prefixed morph stream
(init from s46-best if it beats s45, else s45), aux up x4.
- [x] Gate: beat s45/s46 DER on the identical beam-4 harness.
RESULT (2026-08-23): s47 16.53 vs s46 16.43 — FAILED.
r6's aux-task win is not template-portable; Arabic's residual
was iʿrāb-shaped, Hebrew's is not. s46 stays canonical; line
closed (no s48).

## 4. Arabic r6 beam-4 probe — CLOSED (negative)
All Arabic numbers are greedy; Hebrew gained 12 DER points from
beam. If beam helps r6, ship beam — free quality, zero training.
DONE (2026-08-23): beam-4 Total DER 2.5588 vs greedy 2.5793, Morph
1.5379 vs 1.5317 — noise both ways. Greedy stays (4x cheaper).

## 5. Delivery (agent-side per manifest; our job: keep teachers stable)
- [ ] IMF v1 export + index registration: urdu-diac d2, khmer
decision after §1, r6 teacher refresh for the agent's
ara-diac re-distillation.
- [ ] secryst.org models page refresh with the verdict table.
- [ ] Close CI: interscript-ruby #764/#763, maps #181.

## Non-goals (money discipline)
No r8, no more Thai/Persian, no LLM-teacher experiments (#313
parked), no Khmer training beyond §1's restoration framing.
Loading
Loading