Add LLaDA-8B Base Brain-Score model plugin - #409
Conversation
Registers LLaDA-8B Base as a Brain-Score Language ArtificialSubject, with documented operating regime and plugin contract tests.
|
Submission-status update: the fresh local official-protocol neural matrix completed for all 12 registered identifiers (no failed identifiers), and the plugin contract test passed (2 passed). The remaining blocker is the submission workflow rather than model evaluation. Its first job fails before dependency installation with:
|
|
Hi Brain-Score maintainers — could you please check the status of the Jenkins scoring job triggered after merge of PR #409 for The post-merge workflow reports that the scoring trigger succeeded, but the model has not yet appeared on the public Language leaderboard or in my Language profile. The workflow also fell back to its default email/user ID because it could not resolve my GitHub username ( Could you please associate this submission and its eventual results with my Brain-Score account: For clarity, the local official-protocol neural matrix completed successfully, but I am not treating those local values as the official leaderboard result. Thank you. |
|
Hi @hxu129 thanks for submitting the model. I'll keep an eye on your model submission - it may take a few hours. If anything is acting up, please email me at kpradeep@mit.edu. |
#410) call_jenkins_language pointed at core/job/score_plugins, the bash pipeline the unified orchestrator replaced. That pipeline still submits containers with a relative `cd language`, which under the current image's `WORKDIR /language` resolves to /language/language and does not exist. Every job from PR #409 (llada-8b-base) died in seconds with exit 1 at a 60 MB peak, and the build still reported SUCCESS. Point it at core/job/gated_score_plugins, whose absolute `cd /language` was validated on builds 113 and 125. The helper also swallowed every HTTP error and returned normally, so the workflow printed "triggered successfully" whether or not Jenkins accepted the request. Raise on any non-2xx instead; the workflow step already exits non-zero when the call fails. The error carries only the status code, since the URL holds the trigger token and the query string holds the submitter's email. Move the function into submission/jenkins.py so the workflow and tests can import it without endpoints.py resolving the database secret at import time, matching how vision split call_jenkins_vision_gated in #2424. endpoints.py re-exports it for existing callers.
Brain-Score Language model submission
This PR registers
llada-8b-base(GSAI-ML/LLaDA-8B-Base) as a Brain-Score LanguageArtificialSubjectfor neural-model evaluation.What the plugin exposes
NotImplementedErroruntil a diffusion conditional-likelihood / decoding policy is separately registered.LLADA_MODEL_PATHoverride for offline or shared-GPU evaluation.Brain-Score validation
Under Brain-Score Language
0e0bb4a2d8df5d3c30fe26ec4528f27a188e1cdc:0.134269, matching the precomputed reference0.133745within the pre-registered 0.001 tolerance.Important scope boundary
Tuckute2024-rdmis undefined for both models in this Brain-Score revision: the supplied assembly has one neuroid, so its official row-wise RDM correlation has zero variance. It is documented but not averaged into a neural headline.These results establish a clean representational-alignment baseline. They do not claim a diffusion-specific mechanism, behavioral alignment, or a neural temporal-process match.
Files
Adds the model package, requirements, contract tests, and a model-specific README under
brainscore_language/models/llada8b/.