
Community Discussion · Policy
From Screen Recording to Skills: How Claude's 'Teaching' Feature Blurs RL and Supervised Learning Boundaries
I noticed an interesting detail about Claude's newly launched "Record a skill" feature. On the surface, it's a product update, but the underlying logic hits right on a classic problem in our reinforcement learning field—how to enable agents to learn efficiently from human demonstrations. As a PhD student who wrestles with reward functions and policy gradients in the lab every day, my first reaction wasn't "this feature is useful," but rather "what training scheme did Anthropic use?"
Physix Frontier