Skip to content
#

refusal

Here are 31 public repositories matching this topic...

🔓 Ablate — directional ablation (abliteration) toolkit for open-source LLMs. Automatic censorship/refusal removal via residual-stream direction ablation, with KL-guided search, an LLM-judge harness, and one-call push to the Hub. pip install ablate-llm

  • Updated Jul 18, 2026
  • Python

This dataset compiles common phrases and statements used by AI models when they refuse to answer a query or complete a task. It's designed to help developers identify and categorize 'soft failures'—situations where a model explicitly declines a request due to safety, ethical, capability, or policy constraints, rather than generating an incorrect or

  • Updated Sep 4, 2026
  • Python

Add this topic to your repo

To associate your repository with the refusal topic, visit your repo's landing page and select "manage topics."

Learn more