Preserving Long-Tailed Expert Information in Mixture-of-Experts Tuning (COLM 2026): ExpertCondenser: SFT for sparse MoE LLMs with gated condenser experts. Official code.
-
Updated
Sep 21, 2026 - Python
Preserving Long-Tailed Expert Information in Mixture-of-Experts Tuning (COLM 2026): ExpertCondenser: SFT for sparse MoE LLMs with gated condenser experts. Official code.
Less is MoE (EMNLP 2026 Main); Fisher-MoE: intra-expert structured pruning and compression for Mixture-of-Experts (MoE) LLMs. Prunes FFN intermediate dimensions by Fisher importance instead of dropping whole experts. Official code.
Reading the Router: plain-language, legible profiles of Mixture-of-Experts (OLMoE) expert specialization. Paper + code.
Local MoE inference in C, token-exact against transformers on any CPU: it configures itself from the machine it runs on and never swaps.
Raspberry Pi 5 16GB で 30B クラス LLM を 5.2 t/s 動作させる全層最適化スタック — 実機ベンチ + 失敗事例ログ込み
Context-length degradation ("context rot"), mechanistically localized to retrieval-head attention collapse and causally repaired — replicated across OLMoE, Granite, and Pythia
Deterministic residual Tool Expert for compressed tool schemas on OLMoE
To associate your repository with the olmoe topic, visit your repo's landing page and select "manage topics."