Community Discussion · Policy
Why Distillation Suddenly Became AI's Hottest Topic: Starting with 90% Accuracy and 10% Parameters
Last spring, I led several high school sophomores on a science innovation project focused on using AI to identify plant leaves on campus. The students' laptops couldn't handle ResNet-50, so I taught them knowledge distillation to "compress" the capabilities of large models into MobileNet. The results were intuitive: the original large model had 25 million parameters, while the student model had only 2.5 million, yet recognition accuracy dropped from 95% to 93%, with almost no loss. One student asked, "Teacher, does this count as cheating? Making the small model directly copy the big model's answers?" Before I could answer, she continued, "Then how did the big model learn? Does it have a teacher too?"
Physix Frontier