Is it reliable to have AI impersonate users for pricing tests?
Community Discussion · Tracks

Is it reliable to have AI impersonate users for pricing tests?

Long JiLong Ji20h ago2026/10/02 47 views

Last week I took on a gig helping a client validate a pricing scheme. Three price tiers, and they wanted to talk to real users to see which tier they'd accept and why. I asked around, and the scheduling was booked out two weeks, one fee per head, and after chatting you might not even hit the point.

So I thought, can I first have AI simulate a group of people to run a round?

When searching I took a detour first. Searching "simulation" turned up robot simulation, industrial simulation, Abaqus and that kind of physical-field simulation, completely different from what I wanted. Simile does human behavior simulation—simulating how people choose and why. Founder Joon Sung Park is a Stanford PhD, and the 2023 AI town Smallville experiment was his work. He later went off to start a company, and reportedly raised $100 million in funding early this year. Stanford's Percy Liang once said "simulation is the next frontier of AI," also referring to this line.

Day one. I registered and created a project. I thought the flow was: write a description, it generates a batch of virtual users for me, then I ask whatever and it answers. It wasn't.

The interface is a project list. Clicking in, you have to choose "simulation subjects" and "scenario," where the scenario is the thing you want to ask about. But it also requires you to feed in real materials first—interview records, questionnaires, user behavior data all work. Without these, what comes out is very empty, like a general large model acting.

Don't let AI pretend to be people, it's useless.

This is what Joon Sung Park himself said. On day one I saw it and wasn't quite convinced, thinking, isn't this just letting AI pretend to be people? After getting hands-on I roughly understood—it's pretending to be "the batch of people you interviewed," not a personality conjured out of thin air. So on day one I basically got nothing running, all spent organizing the interview records on hand, formats didn't match, went back and forth revising twice.

Day three, it finally ran. I threw the three price tiers in and had it simulate a group of people's reactions to each tier—whether they'd buy, where they got stuck, what they'd compare against. The official line says it can compress months of research into half a day; mine wasn't that dramatic, but from organizing to getting results, it was indeed much faster than the two weeks I'd originally scheduled.

The pleasant surprise is that it doesn't just give a ratio, it also gives a chain of reasons. Two "won't buy" concerns were ones I hadn't thought of myself, and when I went back and checked with the client, they thought it made sense too.

The sticking points are there too. From my testing, running the same batch of input again a day later, the most popular tier didn't change, but the order of the second and third tiers wobbled a bit. So don't treat the numbers this thing produces as statistical conclusions—treat them as a list of hypotheses. Also, I ran a Chinese-language scenario this time, and I didn't fully figure out what the sample base's source was, so I'm not confident whether the conclusions are accurate—this has to be said up front.

After a week of use, I put it in the front half of the workflow: let it run a round first, cut out the obviously untenable options, and validate the rest with real people. I think this use case holds up—what it saves is screening time, not validation time.

Let me be clear about the downsides. The barrier is on data; people who don't have real interview material on hand will feel it spinning its wheels. I don't dare vouch for Chinese-language scenarios. The output is a text document, and in the end you still have to judge for yourself—it doesn't make decisions for you. I didn't ask about the price; this time I went through the trial channel.

Conclusion: depends. Suitable for product and marketing teams that already have user interview material and want to quickly screen hypotheses, doing multiple-choice things like pricing, copy, and campaign plans. Not suitable for people who have no data on hand and want a direct answer, nor for serious research that needs strict statistical significance to support decisions. As for me, I'll keep using it on my next project, but I won't let it draw conclusions on its own.

0 replies

?
Ctrl + Enter to reply
No replies yet — be the first to share your thoughts