[ICLR 2025] General-purpose activation steering library
-
Updated
Sep 18, 2025 - Python
[ICLR 2025] General-purpose activation steering library
Benchmark evaluation code for "SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal" (ICLR 2025)
We study whether categorical refusal tokens enable controllable and interpretable safety behavior in language models.
Reproducible, evergreen benchmark for LLM refusal on biological research prompts — 19 models, 141 prompts, 13,389 adjudicated trials
🔓 Ablate — directional ablation (abliteration) toolkit for open-source LLMs. Automatic censorship/refusal removal via residual-stream direction ablation, with KL-guided search, an LLM-judge harness, and one-call push to the Hub. pip install ablate-llm
CLI for keeping Claude Fable 5 prompts, skills, and API traffic in shape — lint anti-patterns, canary silent degradation, aggregate refusal analytics. Every rule cites Anthropic docs.
中文企业公开报告 Hybrid RAG:Docling 解析 · Qdrant 稠密/稀疏检索 · 查询理解硬过滤 · 带引用生成与拒答 · 文档生命周期与评测看板
Public Driftmap harness: public-safe CSV suites + rubrics + run logs for drift detection, refusal integrity, injection resistance, and uncertainty tracking.
A probe suite that measures which conversation states an LLM cannot leave. Three arms, because two cannot tell obedience from token statistics; a null only counts when the design had the power to see the effect.
RAG with verifiable citations and measured refusal — retrieval scored separately (TF-IDF beats embeddings here), citations validated against chunks actually retrieved.
Refusal verification surface for Riverbraid fail closed policy boundaries.
RefusalScope — refusal-scope drift tracker for LLM agents. Diffs probe-pack refusal sets across snapshots so a model's newly-refused or newly-answered behavior surfaces before it ships.
Unified CLI/TUI to abliterate any (V)LLM (Heretic / OBLITERATUS / ErisForge) and benchmark the methods on one schema — G/P/S composite + Pareto front. Pure-stdlib core, no GPU for the core.
Locating and editing refusal in the J-space workspace with the Jacobian lens: refusal is legible ~10 layers before the first token, and only ~1/3 lives in the verbalizable workspace.
Training-time defense that redistributes LLM refusal via mean/covariance matching + KD, raising linear-ablation attack rank from K=1 to K≥16 (Llama-3.2-1B-Instruct)
An open reproduction of feature-level activation steering with the prompt set released, showing the capability tax that behavioural metrics miss
Public reference interfaces for proof-gated AI action, refusal, authority, and evidence boundaries.
A document copilot that cites what it says and refuses when the evidence is not there: 33 questions, baseline 12/33 to harness 32/33, control 8/8 both ways.
How does this function's cost scale? Counted, not timed — and UNDETERMINED when no complexity class settles.
This dataset compiles common phrases and statements used by AI models when they refuse to answer a query or complete a task. It's designed to help developers identify and categorize 'soft failures'—situations where a model explicitly declines a request due to safety, ethical, capability, or policy constraints, rather than generating an incorrect or
To associate your repository with the refusal topic, visit your repo's landing page and select "manage topics."