A from-scratch Rust inference stack for AMD GPUs that talks straight to the Linux kernel driver. No ROCm, HIP, or HSA runtime. Serves OLMo 2 on MI355X.
-
Updated
Aug 15, 2026 - Rust
A from-scratch Rust inference stack for AMD GPUs that talks straight to the Linux kernel driver. No ROCm, HIP, or HSA runtime. Serves OLMo 2 on MI355X.
TMLR 2026 | Mechanistic interpretability: attention-head binding (EB*) as a marker of concept emergence. 7 models, 5 architectures (Pythia 160M–2.8B, OLMo-1B, CRFM GPT-2, SmolLM3-3B, Qwen2.5-1.5B), 41 terms.
OMTR — can LLM memorization be separated from predictability? Preregistered causal probing (activation patching) in Pythia & OLMo on consumer hardware; three honest non-separations, full corrections history
PyTorch DDP and multi-GPU LLM training — from a minimal distributed example to NanoGPT speedruns, Muon, H100 profiling, and modded-nanogpt. FBA LAB https://bubblnet.com
Cross-Family Convergence of Neural Network Weight Skeletons. Companion to Zenodo paper (10.5281/zenodo.19652706).
Pre-training a ~150M parameter code-specialized language model using OLMo 3 architecture (GQA, SWA, SwiGLU, RoPE) on PHP/JS/Python/C source code.
OpenEuroLLM snapshot of the Berkeley Function Calling Leaderboard evaluation harness and OLMo evaluation orchestration.
First open-source descriptor-augmented LLM for Neglected Tropical Disease drug discovery | OLMo-7B + QLoRA + DeepChem | Bioactivity & Toxicity prediction for Leishmaniasis, Chagas, Malaria, TB
RDKit-Guided Topological State Machine (TSM) for constrained SMILES generation with OLMo-7B. Solves BPE-tokenizer mismatch via RDKit-in-the-loop decoding. Achieved 100% validity on allenai/OLMo-7B-hf.
Paired OLMo continual-training experiment on spaced review and delayed FictionalQA retention
OpenEuroLLM tau2-bench evaluation harness with OLMo and Qwen serving, user-simulator, and aggregation scripts.
AI工具类(Midjourney / Notion / ChatGPT) 订阅教程类(Netflix / Spotify / Adobe) 虚拟卡支付类(Namecheap / OpenAI / Google)
Process-level comparison of children's gaze and OLMo's self-attention on the same abstract sequence-completion task. Research Master's thesis code.
Research workspace for model diffing between pretrained and post-trained language models.
CS336A5 小显存适配:用小显卡直接进行 OLMo-2 全参数强化学习训练,已在单张 RTX 5070 Ti 16GB 上验证。提供低显存训练配置、vLLM 批量生成与复现说明。
Run LLM inference directly on AMD GPUs by bypassing ROCm for low-latency serving.
To associate your repository with the olmo topic, visit your repo's landing page and select "manage topics."