-
Updated
Sep 5, 2026 - Rust
activation-steering
Here are 114 public repositories matching this topic...
[ICLR 2025] General-purpose activation steering library
-
Updated
Sep 18, 2025 - Python
KV Cache Steering for Controlling Frozen LLMs
-
Updated
Aug 18, 2026 - Python
OpenAI-compatible server with a live view into any HF model's residual stream. pip install brainscope
-
Updated
Sep 4, 2026 - Python
Agents that talk through model internals — activations & J-space — instead of text. No words pass between them. The lab on top of brainscope + hidden-directions.
-
Updated
Sep 4, 2026 - Python
[ACL 2026] - Official repo for the paper: "Selective Steering: Norm-Preserving Control Through Discriminative Layer Selection"
-
Updated
May 22, 2026 - Jupyter Notebook
[Under Review] Not All Tokens Are Equally Useful for Steering: Robust Directions and Prefix Steering
-
Updated
Jul 22, 2026 - Python
Activation steering and trait monitoring for HuggingFace transformers
-
Updated
Aug 13, 2026 - Python
A bilingual awesome list for refusal suppression research: benchmarks, papers, tools, models, and ecosystem updates.
-
Updated
Jun 12, 2026
[ICLR 2026] ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
-
Updated
Sep 30, 2025 - Python
Official code for "Activation Steering for Accent Adaptation in Speech Foundation Models" (Interspeech 2026). Parameter-free accent adaptation via mean-shift steering vectors — no weight updates, consistent WER reductions across 8 accents.
-
Updated
Mar 17, 2026 - Python
Transformer interpretability research framework for circuit discovery, sparse features, and controlled residual-stream interventions.
-
Updated
Aug 16, 2026 - Python
Runtime rank-1 refusal projection for DeepSeek-V4-Flash-0731: 757KB of directions instead of a 1.54GB weight overlay, lambda as a hot-swappable dial. Full A/B measurements on 2x DGX Spark.
-
Updated
Aug 22, 2026 - Python
GEMS: Geometric Constraints Enable Multi-Semantic Superposition in LLMs
-
Updated
Jun 21, 2026 - Python
🏆[ICML 2026 Spotlight] Official implementation of "DecodeShare: Tracing the Shared Subspace of LLM Decode-Time Decisions"
-
Updated
Jun 21, 2026 - Python
Feature steering for open LLMs: find an interpretable SAE feature, steer the model, measure the causal effect
-
Updated
Jul 4, 2026 - Python
MiniMax-H3 custom nodes: steering / cache / prompt director
-
Updated
Aug 31, 2026 - Python
Does concept-injection introspection emerge with scale? A faithful, controlled reproduction charted across model-size ladders.
-
Updated
Jul 31, 2026 - Python
The paper list related to activation steering
-
Updated
May 10, 2026
Inspect, steer & monitor a real LLM on Apple Silicon — a browser-based SAE interpretability lab for Qwen Scope, powered by MLX
-
Updated
Aug 22, 2026 - Python
Add this topic to your repo
To associate your repository with the activation-steering topic, visit your repo's landing page and select "manage topics."