58 words
1 minutes
reading_club_04_22

Improving Alignment and Robustness with Circuit Breakers#

arxiv

AI system can take harmful actions and could be vulnerable to adversarial attacks. This paper proposed a way called circuit breaking to improve the safety of the system especially for the adversarial attacks the model to do harmful actions.

Current alignment methods such as refusal training is easy to be fooled by attacks.

reading_club_04_22
https://lukew1999.github.io/posts/reading_club_04_22/
Author
Weiqi Wang
Published at
2025-04-21