Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This adds
Dream-org/Dream-v0-Base-7B(pinned revision6572adb5535263e4d1a337b56942ba48b6dee2a9, Apache-2.0) as a Brain-Score LanguageArtificialSubject.The Neural operator uses clean passage-local text, independent text-part tokenization, and current-part mean hidden state at predeclared layer 21/28 for fMRI/ECoG. The independent boundary is necessary because Dream's custom slow tokenizer does not expose character offsets. The Behavior/Engineering operator independently tokenizes the available prefix, then scores new subtokens left-to-right with one mask (ID 151666), returning summed surprisal in bits. Empty SyntaxGym regions contribute zero and are omitted from later context. The input budget is 4095 observed tokens plus one mask; Dream evaluates with
use_cache=Falseandnum_logits_to_keep=1.next_wordremains unsupported. This operator is a reading-time surrogate, not an exact diffusion joint likelihood.WeiWang validation: the 18 checkpoint files were byte-matched against the durable source; runtime output is finite and repeatable; manual multi-token likelihood agrees with the subject; two-part neural digestion yields 3584 units. All 12 local Neural identifiers completed without execution failures.
Tuckute2024-rdmreturned NaN for this benchmark's one-neuroid assembly and is excluded from meaningful comparison. Key local Neural scores: Pereira-243 1.00, Pereira-384 1.00, Blank 0.233, Fedorenko 0.765.The complete local
Futrell2018-pearsonrrun finished over all 10,256 words: raw Pearson r = 0.293035, ceiling-normalized score = 0.341516. The corrected SyntaxGym run finished 31/31 suites without errors; mean over all 31 suites = 0.509757, and mean over the 30 suites shown on the public leaderboard = 0.507305. These are local executions of the official benchmarks, not yet published website scores. The Futrell stimuli contain no empty parts; thus the empty-region correction in this PR does not alter its behavioral operator relative to the completed run. A 13-item one-vs-16-mask diagnostic found median absolute target-surprisal change of 0.50 bits. An output-independent 18-item late-position diagnostic found median absolute first-subtoken surprisal changes of 0.416 and 0.237 bits when the observed window was reduced from 4095 tokens to 511 or 2047 tokens, respectively.The local benchmark audit is complete and this PR is ready for maintainer review. Upstream's submission orchestrator currently fails to check out fork branches (see the comment on PR #412); the independent integration, plugin unit-test, and documentation checks passed. Dream's neural token-boundary rule differs from the existing LLaDA plugin, so small cross-model score differences are not interpreted as architecture or brain-mechanism effects.