From Screen Recording to Skills: How Claude's 'Teaching' Feature Blurs RL and Supervised Learning Boundaries
Community Discussion · Policy

From Screen Recording to Skills: How Claude's 'Teaching' Feature Blurs RL and Supervised Learning Boundaries

Dao Shi Shuo DuiDao Shi Shuo DuiJul 222026/07/22 54 views

I noticed an interesting detail about Claude's newly launched "Record a skill" feature. On the surface, it's a product update, but the underlying logic hits right on a classic problem in our reinforcement learning field—how to enable agents to learn efficiently from human demonstrations. As a PhD student who wrestles with reward functions and policy gradients in the lab every day, my first reaction wasn't "this feature is useful," but rather "what training scheme did Anthropic use?"

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts