Back to home

    A More Human-Like AI Focus Group Isn’t Necessarily a Better One

    Matching the variety of human behavior is not the same as predicting what people will do.

    Sidd Adatrao
    Written bySidd AdatraoSeptember 21, 2026

    Most evaluations of synthetic users begin by inspecting the users themselves. Do they sound distinct? Do they express believable motivations? Does the group contain a realistic range of needs and preferences?

    Those checks tell us whether the simulation represents the kinds of people we expected to see. They do not tell us whether those simulated people react correctly when price, copy, design, or context changes.

    That second question requires a different comparison. We need to measure the path from user type to observed action, not just the distribution of user types.

    A simulation has to get two things right

    Imagine dividing customers into a few broad groups: bargain hunters, convenience shoppers, loyal customers, and cautious first-time buyers.

    To predict the overall result, a simulation needs to know two things:

    1. How common each group is. What share of the real population consists of bargain hunters, loyal customers, or cautious new buyers?
    2. How each group actually behaves. When shown the same offer, how likely is each group to choose it?

    Persona diversity mainly helps with the first problem. It can make the simulated population contain a more realistic mix of people. But the final prediction also depends on the second problem: whether the agents make decisions the way their real counterparts do.

    The simple version

    Overall prediction = each group's share × how that group behaves

    Getting the group shares right is not enough. If the model gets the behavior of those groups wrong, the final answer can still be wrong.

    Simple example

    The same population mix can produce opposite predictions

    Real customers

    Population mix

    Simulated customers

    Same population mix

    What real customers do

    70% choose A

    What the simulation predicts

    30% choose A

    Where prediction error can hide

    A real-to-simulated comparison should identify which part of the prediction process is failing. There are at least three distinct sources of error.

    Population error

    The mix is wrong

    The simulation contains too many users of one kind and too few of another. Even accurate individual decisions will aggregate to the wrong result.

    Response error

    The reactions are wrong

    The population mix is correct, but simulated users respond differently from comparable real users when shown the same thing.

    Context error

    The sensitivity is wrong

    Real behavior changes with device, order, prior experience, or framing. The simulation may react too much, too little, or in the wrong direction.

    A better real-to-simulated comparison

    The cleanest evaluation is a matched study. Real and simulated users should receive the same task, the same information, and the same set of choices. The comparison should preserve the conditions that could affect the result, including presentation order, device, price, and prior context.

    1. Define the prediction first. Decide whether the simulator is predicting a choice rate, a ranking, a winning variant, retention, or something else. Score the output the product decision will actually use.
    2. Match the inputs. Show real and simulated users the same stimulus under the same randomized conditions. Otherwise a difference in outcomes may be caused by the setup rather than the simulation.
    3. Compare within groups. Measure whether bargain hunters, new customers, or other meaningful groups respond like their real counterparts. An accurate overall number can hide offsetting segment errors.
    4. Rebuild the total using real weights. Combine the simulated group-level predictions using the observed size of each real group. This isolates response error from errors in the simulated population mix.
    5. Evaluate on held-out outcomes. Data used to create, prompt, select, or tune the personas should not also be used to claim predictive performance.

    What to report

    MeasureWhat it reveals
    Outcome errorHow far the simulated choice or conversion rate is from the real rate.
    CalibrationWhether events predicted at 70% actually occur about 70% of the time.
    Segment errorWhich customer groups are predicted well and which are being averaged away.
    Decision accuracyWhether the simulation selects the real winner or produces the correct ranking.
    Decision regretThe value lost by following the simulation instead of the best real-world option.

    Every result should be shown beside a simple baseline, such as the overall historical rate, one generic agent, or the current decision process. A complex persona population is useful only when it improves the decision beyond what the simpler method already provides.

    This evaluation produces a more useful diagnosis than a single similarity score. It tells us whether the simulated population has the wrong composition, the wrong response patterns, or the wrong sensitivity to context. More importantly, it shows which failure matters for the decision being made.