Cambridge Mofang Notes
Feb 5, 2025 · Artificial Intelligence
DeepSeek R1's Knowledge Distillation: Teacher-Student Model Compression Explained
This article explains knowledge distillation as used in DeepSeek's R1 model, detailing how a large teacher model transfers knowledge to a smaller student model via soft targets and loss functions like KL divergence, enabling efficient deployment on resource-constrained devices.
AI EfficiencyCross-Entropy LossDeepSeek R1
0 likes · 6 min read
