Skip to content

Add target-hidden behavioral scoring for LLaDA-8B Base - #412

Open
hxu129 wants to merge 1 commit into
brain-score:mainfrom
hxu129:codex/llada-brainscore-behavior-v2
Open

hxu129 wants to merge 1 commit into
brain-score:mainfrom
hxu129:codex/llada-brainscore-behavior-v2

Conversation

@hxu129

@hxu129 hxu129 commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

The existing LLaDA-8B-Base plugin has public Neural scores but returns no Behavior or Engineering scores. This change adds a fixed reading_times operator for Futrell2018 and SyntaxGym while keeping the published layer-24 neural readout unchanged.

For each text part, the adapter independently tokenizes the available prefix and scores continuation subtokens left-to-right with one appended mask. It sums surprisal in bits, omits empty SyntaxGym regions from later context, follows the official AR adapter's punctuation spacing, and caps the input at 4095 observed tokens plus the mask. The pinned checkpoint's last_logits_only path avoids computing unused all-position vocabulary logits. next_word remains unsupported. This operator is a reading-time surrogate, not an exact diffusion joint likelihood.

Validation: registration and target-hidden/empty-region tests passed. Live WeiWang smoke checks passed behavioral assembly shape, manual multi-token log-prob sum, Futrell items and SyntaxGym regions. An input-only audit found 592 completed-context tokenizations that changed preceding token IDs, motivating independent prefix tokenization. A frozen 13-item one-mask versus 16-mask check found median absolute target-surprisal change of 0.63 bits. The optimized output path preserved top-1 predictions with maximum distribution TV below 0.0001 on three input lengths; a 500-word Futrell smoke differed by about 0.000004 in normalized score.

The corrected complete local SyntaxGym run finished 31/31 suites without errors: mean over all 31 = 0.683732; mean over the 30 suites displayed on the public leaderboard = 0.673190. A diagnostic GPT-Neo-1.3B AR control, evaluated under the official adapter with empty regions handled consistently, scored 0.800313 on those 30 suites; it is scale-mismatched and not an architecture-superiority test. An output-independent 18-item late-position diagnostic found median absolute first-subtoken surprisal changes of 0.264 and 0.124 bits when the observed window was reduced from 4095 tokens to 511 or 2047 tokens, respectively. These are local results, not published website scores.

The complete local Futrell2018-pearsonr run finished over all 10,256 non-empty words: raw Pearson r = 0.245143, ceiling-normalized score = 0.285700. Its no-empty-parts input means the empty-region correction in this PR does not alter its behavior relative to the completed run. The local benchmark audit is complete and this PR is ready for maintainer review. Upstream's submission orchestrator currently fails to check out fork branches (see the PR comment); independent integration, plugin unit-test, and documentation checks passed. Neither benchmark score alone is a diffusion-specific brain-mechanism claim.

@hxu129

hxu129 commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor Author

Update (2026-09-28): Both this PR and Dream PR #413 are now ready for review. The full local Futrell and corrected SyntaxGym evaluations have completed, with the scores and scientific boundaries in each PR description. The independent integration, plugin unit-test, and documentation checks are green.

Both PRs still fail at Plugin Submission Orchestrator / 1. Detect Changes because the workflow checks out github.event.pull_request.head.ref from the upstream repository rather than from the fork that owns the branch. The workflow has additional head.ref checkout steps later, so the complete fork path may need review after the first step is corrected. I do not have push permission to brain-score/language, so an upstream-owned branch is not available from this account.

Could a maintainer make the submission workflow fork-aware or advise an approved manual route for these two ready PRs? I will treat the scores as local benchmark results until the official scoring jobs and website publication are verified.

Follow-up: I opened the minimal workflow fix in PR #414. Its Detect Changes job passed on that real fork PR, including changed-file detection and is_fork_pr=true; independent checks are green. Please review #414 when possible so #412/#413 can be retriggered against the fixed workflow.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants