Conversation
|
Update (2026-09-28): Both this PR and Dream PR #413 are now ready for review. The full local Futrell and corrected SyntaxGym evaluations have completed, with the scores and scientific boundaries in each PR description. The independent integration, plugin unit-test, and documentation checks are green. Both PRs still fail at Could a maintainer make the submission workflow fork-aware or advise an approved manual route for these two ready PRs? I will treat the scores as local benchmark results until the official scoring jobs and website publication are verified. Follow-up: I opened the minimal workflow fix in PR #414. Its |
The existing LLaDA-8B-Base plugin has public Neural scores but returns no Behavior or Engineering scores. This change adds a fixed
reading_timesoperator for Futrell2018 and SyntaxGym while keeping the published layer-24 neural readout unchanged.For each text part, the adapter independently tokenizes the available prefix and scores continuation subtokens left-to-right with one appended mask. It sums surprisal in bits, omits empty SyntaxGym regions from later context, follows the official AR adapter's punctuation spacing, and caps the input at 4095 observed tokens plus the mask. The pinned checkpoint's
last_logits_onlypath avoids computing unused all-position vocabulary logits.next_wordremains unsupported. This operator is a reading-time surrogate, not an exact diffusion joint likelihood.Validation: registration and target-hidden/empty-region tests passed. Live WeiWang smoke checks passed behavioral assembly shape, manual multi-token log-prob sum, Futrell items and SyntaxGym regions. An input-only audit found 592 completed-context tokenizations that changed preceding token IDs, motivating independent prefix tokenization. A frozen 13-item one-mask versus 16-mask check found median absolute target-surprisal change of 0.63 bits. The optimized output path preserved top-1 predictions with maximum distribution TV below 0.0001 on three input lengths; a 500-word Futrell smoke differed by about 0.000004 in normalized score.
The corrected complete local SyntaxGym run finished 31/31 suites without errors: mean over all 31 = 0.683732; mean over the 30 suites displayed on the public leaderboard = 0.673190. A diagnostic GPT-Neo-1.3B AR control, evaluated under the official adapter with empty regions handled consistently, scored 0.800313 on those 30 suites; it is scale-mismatched and not an architecture-superiority test. An output-independent 18-item late-position diagnostic found median absolute first-subtoken surprisal changes of 0.264 and 0.124 bits when the observed window was reduced from 4095 tokens to 511 or 2047 tokens, respectively. These are local results, not published website scores.
The complete local
Futrell2018-pearsonrrun finished over all 10,256 non-empty words: raw Pearson r = 0.245143, ceiling-normalized score = 0.285700. Its no-empty-parts input means the empty-region correction in this PR does not alter its behavior relative to the completed run. The local benchmark audit is complete and this PR is ready for maintainer review. Upstream's submission orchestrator currently fails to check out fork branches (see the PR comment); independent integration, plugin unit-test, and documentation checks passed. Neither benchmark score alone is a diffusion-specific brain-mechanism claim.