CALIBURN: Self-Calibrated LLM Unlearning Alignment

Abstract

LLM unlearning offers a practical mechanism for addressing safety and privacy concerns by removing the influence of undesirable knowledge from pretrained language models. CALIBURN calibrates unlearning updates using the target model’s confidence, enabling fine-grained forgetting while better preserving general model utility.

Publication
Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)
Zhuangdi Zhu
Zhuangdi Zhu
Assistant Professor (Tenure-Track)

My research focuses on making AI models safe and aligned.

Related