Dynamic Abliteration: Non-Destructive Refusal Suppression via Multi-Layer Engram Steering
- When working with open-weight LLMs like Qwen, controlling refusal behavior on security, administrative, prompts typically requires fine-tuning or permanent weight update.
- Traditional weight abliteration technique neutralizes refusal directions by projecting weight matrices orthogonal to a refusal vector.
- However, this permanently alters base model weights and can degrade performance across non-refusal tasks also.
Unverified
- When working with open-weight LLMs like Qwen, controlling refusal behavior on security, administrative, prompts typically requires fine-tuning or permanent weight update.
- Traditional weight abliteration technique neutralizes refusal directions by projecting weight matrices orthogonal to a refusal vector.
- However, this permanently alters base model weights and can degrade performance across non-refusal tasks also.
Sources: Madhukaraphatak