58 words
1 minutes
reading_club_04_22
Improving Alignment and Robustness with Circuit Breakers
AI system can take harmful actions and could be vulnerable to adversarial attacks. This paper proposed a way called circuit breaking to improve the safety of the system especially for the adversarial attacks the model to do harmful actions.
Current alignment methods such as refusal training is easy to be fooled by attacks.
reading_club_04_22
https://lukew1999.github.io/posts/reading_club_04_22/