Life background
Who someone is before they meet a product — life stage, household, income, geography, career arc, the priors they carry into every decision.
Without this, a persona is a costume.
As general-purpose LLMs keep scaling on complex problems, they are not improving at human behavior accuracy — we need behavior models.
SAPIENS is a purpose-built behavior model, built to model an audience's and a user's latent behavioral, psychographic, and socio-demographic traits from real user data.
As new instruction-tuned LLMs (GPT, Llama, Claude, Gemini) get better at hard math, science, PhD-level, and coding problems, they are plateauing in their ability to model behavior. The sharper these models get at reasoning, the further they drift from replicating how people actually behave.
Accuracy measures - % of Opinions where Real Opinion Matches Generated Opinion when a Product is shown to a modeled user (part of a modeled tribe)
Three standard ways to make an LLM agent act like a person. We ran every frontier model through all three. None break the ceiling and are still less accurate than SAPIENS.
Accuracy measures - % of Opinions where Real Opinion Matches Generated Opinion when a Product is shown to a modeled user (part of a modeled tribe)
It isn't only our result. The academic community is converging on the same finding: as LLMs get better at complex tasks, they get worse at modeling behavior.
CMU scored 31 models on how closely their behavior matches a real person's, on a 0–100 scale. To set the bar, they first checked humans against humans: one group of real people matched another group at 92.9. That is what "acting human" looks like. No model got close — and the newer, more capable models did worse than the older ones they replaced.
human-likeness score (0–100) · each lab's older model vs the flagship that replaced it · CMU, arXiv:2603.11245
The benchmarks that drive general-purpose LLMs — MMLU, GPQA, AIME and MATH, SWE-bench and HumanEval, ARC-AGI — all reward reasoning, knowledge retrieval, and coding tasks. Not one of them measures whether a model can behave like a person.
LLMs are instruction-tuned with RLHF to be helpful personal assistants that serve complex tasks. Real people are not that. They behave illogically, get frustrated, stay ambiguous, refuse to cooperate, and act on motivations of their own. Simulated users by general-purpose LLMs, by contrast, come out too cooperative, too uniform, and never frustrated or ambiguous — the same failure across every behavioral dimension.
As the academic community recognizes that general-purpose LLMs are not good at modeling behavior, there have been several attempts at purpose-built human behavior models — Centaur and Be.FM among them. We benchmarked against them.
behavior prediction accuracy · purpose-built behavior models
We chose the Amazon Public Dataset — the largest public record of human behavior, spanning categories: health and wellness products, video games, software products, and more.
The dataset covers millions of users across very different audiences, with extensive documented purchase history and their opinions on the products they bought. Predicting opinions across a wide variety of products is a real test of whether behavior has been captured — and it's what lets us prove SAPIENS works across categories, not just one.
The task: given a selected user and their history, predict their opinion on a new product. If the generated opinion and the real opinion match closely, that's a success. For SAPIENS, this product is genuinely unseen — it never appears in training data.
We compare the model's generated review against the person's real one on meaning — a prediction only counts when it says what they said and feels how they felt.
Did it surface what the person actually chose to write about?
Did it react to the same product details, in the same way?
Did it land on the same feeling — love it, live with it, return it?
SAPIENS is powered by dense human behavior data — which is what makes it accurate enough to model any population segment with confidence. That density is cross-pollinated across sources, because no single source covers a human.
Read: Creating the dataset of human behaviorWho someone is before they meet a product — life stage, household, income, geography, career arc, the priors they carry into every decision.
Without this, a persona is a costume.
The same human watched across hundreds of unrelated contexts — money, health, gaming, parenting, work — not just inside one category.
Behavior transfers between domains. Single-vertical data can't learn that.
The stimulus around the decision — the ads, feeds, prices, peers, reviews and rival products a person had actually seen before they chose.
Choices are reactions. Model the input or you're guessing at the output.
Revealed preference — what got clicked, bought, abandoned, cancelled, re-opened. The action that quietly contradicts the stated intention.
Everyone says they'd pay. The dataset knows who did.
// Coverage scored across the four dimensions. Directional, not benchmarked.
Identity depth. Household and segment-level demographic and psychographic structure, so a simulated user arrives with a plausible income, life stage, geography and set of values — not an invented backstory.
Healthcare depth. An intelligent platform built on the broadest ecosystem of healthcare audiences, bridging the gap between raw data and a genuine understanding of your customer in a fraction of the time you would expect.
The largest vetted user panel, wired directly into the platform. Any simulation can be put in front of a matched human cohort and scored — turning every run into a measurable claim rather than an opinion.
SAPIENS vs. general-purpose LLMs
SAPIENS — our behavior modeling approach — is roughly 30 percentage points better at predicting human behavior than general-purpose LLMs. Over 18 months we have benchmarked successive generations of frontier instruction-tuned models: despite substantial gains in reasoning and problem-solving, their accuracy on our human behavior benchmark has plateaued at 49–54%. Further academic research suggests they are becoming worse.
Behavior models differ from general-purpose LLMs on three axes — modeling objective (what they optimize for), modeling granularity (what they model), and data. Here is a detailed comparison on the same.
In order to predict what a person does, you need to understand the why behind their decisions and the behavior they exhibit. Behavior models focus on those whys, with the goal of predicting how a person will react to different stimuli.
People do not decide rationally. They decide through unstated perception gaps, biases, and the environmental factors acting on them — the reasons they would never give you in a survey, and often could not. We call these latent traits, and they are what SAPIENS learns and models.
SAPIENS learns those latent traits alongside a person's background, behavioral history, psychographics and socio-demographic traits — from vast amounts of public and proprietary data — because together they are what actually predicts what a person does.
Frontier AI labs are optimizing to replace knowledge workers on benchmarks such as SWE-bench, GPQA, AIME and MMLU. These benchmarks reward step-by-step rational decision making.
RLHF instruction-tuning optimizes for the most plausible, step-by-step logically consistent response — the machinery that wins at analytical tasks, strategic tasks, math, science and coding. Research on human behavior simulation suggests that RLHF and instruction fine-tuning is what makes these models worse at understanding human behavior.
The model is not designed to follow the trajectory of irrational decision making and behavior of different individuals; it is designed to gather and learn domain knowledge across many verticals.
Because users are modeled one at a time, the diversity and behavioral distribution of the real population is preserved rather than averaged away.
Every response is therefore grounded in an individually modeled real user — each one carrying their own experiences, preferences, attitudes and latent behavioral traits, learned from real-world data.
Asked to represent a population, it returns an averaged, rational opinion about that population — which does not represent the variance, and is always seen through the lens of rationality, giving an understanding of the domain at large.
At best, it will return you previously conducted surveys and studies from the internet.
We collect a spectrum of behavioral data on the individuals inside a specific target population: breadth across many sources, and depth on each person within them.
Breadth tells us what someone does. Depth tells us why they do it. Behavior prediction needs both.
And because behavior keeps evolving, so does the data. SAPIENS keeps learning from evolved behavior at every step, staying current with the population it models.
The data is gathered to cover domains — enough of them, well enough, to perform logical tasks in that domain. No granular, person-specific behavioral data is collected at all.
What the corpus holds is a generalized opinion across domains. Wide, shallow, unattached to any individual — and silent on the why.
And it is frozen. A general-purpose LLM is bounded by its last training date, while the behavior it is being asked to predict keeps moving.