Community Discussion · Tracks

Solo Teams Aren't the End Goal; Evaluation Is

Old DengOld DengSep 62026/09/06 45 views

After the solo-team era, evaluation matters more

In this experimental design, over 80% of participants were solo teams. This is more worth noting than the 1.43 million yuan prize pool and over 30,000 submissions. It indicates that AI tools have at least reduced the collaboration cost from idea to demo. Recently, I guided students in running AI agents, comparing baselines of standard function calling versus pure prompting. Results showed that chaining retrieval, code generation, and debugging feedback into a workflow was the most effortless; the raw intelligence of the model itself wasn't as critical.

But data shouldn't just be viewed for excitement. Over 60% had no professional development background, creators with under 10k followers accounted for 88%, and about half of the top 10 had fewer than 10k followers before the competition. Barriers have indeed lowered. Dataset bias must also be considered: platform users skew young and toward content creation; being used on the Toy platform doesn't equate to reliability in production environments.

I care more about how things are evaluated afterward. Papers like ReAct remind us that tool calling moves models from answering questions to doing tasks, but we need to look at task success rates, side effects, and traceable logs. Bilibili pushing works to real users is more like an experiment than awarding prizes. Next, we'll see if popularity can translate into reusable evaluation sets.

1 replies

?
Ctrl + Enter to reply
Brother Yuan

I feel this deeply regarding evaluation sets. When I first used WorkBuddy, the model seemed pretty smart, but it crashed and burned as soon as we hit real business scenarios. The data on the Toy platform is too clean to detect long-tail issues at all.