Skip to content

Add Dream-v0 Base 7B Brain-Score Language subject - #413

Open
hxu129 wants to merge 1 commit into
brain-score:mainfrom
hxu129:codex/dream-v0-brainscore
Open

hxu129 wants to merge 1 commit into
brain-score:mainfrom
hxu129:codex/dream-v0-brainscore

Conversation

@hxu129

@hxu129 hxu129 commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

This adds Dream-org/Dream-v0-Base-7B (pinned revision 6572adb5535263e4d1a337b56942ba48b6dee2a9, Apache-2.0) as a Brain-Score Language ArtificialSubject.

The Neural operator uses clean passage-local text, independent text-part tokenization, and current-part mean hidden state at predeclared layer 21/28 for fMRI/ECoG. The independent boundary is necessary because Dream's custom slow tokenizer does not expose character offsets. The Behavior/Engineering operator independently tokenizes the available prefix, then scores new subtokens left-to-right with one mask (ID 151666), returning summed surprisal in bits. Empty SyntaxGym regions contribute zero and are omitted from later context. The input budget is 4095 observed tokens plus one mask; Dream evaluates with use_cache=False and num_logits_to_keep=1. next_word remains unsupported. This operator is a reading-time surrogate, not an exact diffusion joint likelihood.

WeiWang validation: the 18 checkpoint files were byte-matched against the durable source; runtime output is finite and repeatable; manual multi-token likelihood agrees with the subject; two-part neural digestion yields 3584 units. All 12 local Neural identifiers completed without execution failures. Tuckute2024-rdm returned NaN for this benchmark's one-neuroid assembly and is excluded from meaningful comparison. Key local Neural scores: Pereira-243 1.00, Pereira-384 1.00, Blank 0.233, Fedorenko 0.765.

The complete local Futrell2018-pearsonr run finished over all 10,256 words: raw Pearson r = 0.293035, ceiling-normalized score = 0.341516. The corrected SyntaxGym run finished 31/31 suites without errors; mean over all 31 suites = 0.509757, and mean over the 30 suites shown on the public leaderboard = 0.507305. These are local executions of the official benchmarks, not yet published website scores. The Futrell stimuli contain no empty parts; thus the empty-region correction in this PR does not alter its behavioral operator relative to the completed run. A 13-item one-vs-16-mask diagnostic found median absolute target-surprisal change of 0.50 bits. An output-independent 18-item late-position diagnostic found median absolute first-subtoken surprisal changes of 0.416 and 0.237 bits when the observed window was reduced from 4095 tokens to 511 or 2047 tokens, respectively.

The local benchmark audit is complete and this PR is ready for maintainer review. Upstream's submission orchestrator currently fails to check out fork branches (see the comment on PR #412); the independent integration, plugin unit-test, and documentation checks passed. Dream's neural token-boundary rule differs from the existing LLaDA plugin, so small cross-model score differences are not interpreted as architecture or brain-mechanism effects.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants