GRP-Obliteration: Unaligning LLMs With a Single Unlabeled Prompt
- View PDF HTML (experimental) Abstract:Safety alignment is only as robust as its weakest failure mode.
- Despite extensive work on safety post-training, it has been shown that models can be readily unaligned through post-deployment fine-tuning.
- However, these methods often require extensive data curation and degrade model utility.
Unverified
- View PDF HTML (experimental) Abstract:Safety alignment is only as robust as its weakest failure mode.
- Despite extensive work on safety post-training, it has been shown that models can be readily unaligned through post-deployment fine-tuning.
- However, these methods often require extensive data curation and degrade model utility.
Sources: Arxiv