Reliability-Safety Trade-off in AI Distillation: A Renormalization-Group Approach
π arXiv:2608.08572 Β· π₯ PDF Β· 2026-08-09 Β· cond-mat.stat-mech
Authors: Y. M. Du [arXiv Β· scholar] , Miao-Miao Yi [arXiv Β· scholar] , Tan-Ji Zhou [arXiv Β· scholar] , C. P. Sun [arXiv Β· scholar]
π Abstract
Knowledge distillation transfers more than task competence: it also transmits response propensities, refusal policies, error boundaries, and latent safety biases. We formulate this behavioral inheritance as a coarse-graining model grounded in statistical mechanics, in which the student's answer and refusal decisions define two macrostates, while the teacher induces an effective field that reshapes the student's free-energy landscape. The model yields a reliability-safety trade-off relation controlled by a single parameter K, which we term the hazard discrimination capability. The predicted trade-off is consistent with refusal-token data [arXiv: 2412.06748]. In knowledge distillation, a teacher with strong hazard discrimination improves the student's attainable reliability and safety, whereas poor discrimination limits the attainable trade-off. Repeated distillation acts as an iterated renormalization-group-like transformation, under which K follows a flow across generations. The flow exhibits a tricritical structure separating regimes of K loss, stable transmission, and threshold-dependent inheritance, and yields testable scaling predictions for multigenerational distillation.