Community Discussion · Tracks

Question Banks Dry Up; AI Starts Self-Generating Questions: Step-by-Step Guide to Data-Layer RSI

Yelin Does Not Eat Sponsored MealsYelin Does Not Eat Sponsored MealsAug 92026/08/09 234 views

Let me start with a scenario I tested over the past two days. You're training a model, feeding it a bunch of problems, and suddenly you realize performance has plateaued. It's not that the model is dumb; it's that it's exhausted the dataset. It's like solving practice problems until you've memorized all the answers—doing new ones won't help.

3 replies

?
Ctrl + Enter to reply
Zhong Jinyu

My audience feedback says that diversity in AI-generated samples is a big issue. It's hard for validation models alone to catch questions that 'look right but are logically flawed.' OP, is your validation layer purely rule-based or do you use a small model for judgment? I could make a video covering this case to help everyone avoid pitfalls.

Luguo
LuguoAug 9

Generating is easy, but filtering the questions is hard, bro. I ran into this when working with AI Agents too—AI generated questions in half an hour, but it took me two days to filter them. Bad questions were bad in all sorts of ways; I only dared to mix them into the training set after a manual review... What rules are you using for that validation layer? Pure rule matching or do you also use models to judge?

Cockpit Enthusiast

I have reservations. It sounds great, but there are plenty of issues when actually running it. I've been playing around with BYD's humanoid robot these past few weeks and tried similar data augmentation approaches. Many samples generated by AI look superficially right but are hollow inside, and the validation steps simply can't filter them out cleanly. If you use these to train models, the short-term data volume increases, but generalization capability might actually get worse. Getting the OP's workflow to run is fine, but saying "just follow this and you're good" is too light-handed.