Vectorial VECTORIAL SAPIENS

SAPIENS Benchmarking and the Rise of Behavior Models

As general-purpose LLMs keep scaling on complex problems, they are not improving at human behavior accuracy — we need behavior models.

SAPIENS is a purpose-built behavior model, built to model an audience's and a user's latent behavioral, psychographic, and socio-demographic traits from real user data.

As new instruction-tuned LLMs (GPT, Llama, Claude, Gemini) get better at hard math, science, PhD-level, and coding problems, they are plateauing in their ability to model behavior. The sharper these models get at reasoning, the further they drift from replicating how people actually behave.

Berkeley AI Research (BAIR)
Research paper submitted to EMNLP, a top peer-reviewed conference, with BAIR (Berkeley AI Research).
Read the paper →

SAPIENS vs General Purpose LLMs on Human Behavior Simulations Benchmarking

Accuracy measures - % of Opinions where Real Opinion Matches Generated Opinion when a Product is shown to a modeled user (part of a modeled tribe)

Frontier models stay near 50% across releases; SAPIENS reaches 86.1%.
01

Agent-based baselines

We measured SAPIENS against agent-based simulators built on different general-purpose LLMs.

Three standard ways to make an LLM agent act like a person. We ran every frontier model through all three. None break the ceiling and are still less accurate than SAPIENS.

Accuracy measures - % of Opinions where Real Opinion Matches Generated Opinion when a Product is shown to a modeled user (part of a modeled tribe)

AI agent with user history

Pass the user's own history to the agent and have it model their behavior.
User-history baseline vs SAPIENS 86.1%.

Tribe persona prompting

Prompt the agent with a detailed description of a niche audience segment and have it act like that tribe.
Tribe-persona baseline vs SAPIENS 86.1%.

Population persona prompting

Prompt the agent with a detailed description of a broader population.
Population-persona baseline vs SAPIENS 86.1%.
02

The academic community agrees

The smarter the model, the less human it behaves.

It isn't only our result. The academic community is converging on the same finding: as LLMs get better at complex tasks, they get worse at modeling behavior.

Snapshot of the CMU LTI paper, arXiv:2603.11245
Paper snapshot
paper-cmu.png
arXiv:2603.11245 · CMU LTI

Higher model capability does not yield more faithful user simulation

Zhou et al. — 31 models scored against a 92.9 human-to-human baseline.

"Higher general model capability does not necessarily yield more faithful user simulation." Zhou et al., CMU LTI · arXiv:2603.11245

CMU scored 31 models on how closely their behavior matches a real person's, on a 0–100 scale. To set the bar, they first checked humans against humans: one group of real people matched another group at 92.9. That is what "acting human" looks like. No model got close — and the newer, more capable models did worse than the older ones they replaced.

human-likeness score (0–100) · each lab's older model vs the flagship that replaced it · CMU, arXiv:2603.11245

Newer flagships score lower on human-likeness than the older models they replaced; the human-to-human baseline is 92.9.
92.9
Human panel vs human panel — the target
73.3
Gemini-2.0-Flash — small, cheap, a year old: the best model tested
31
Models evaluated. None came close to the human bar.
Trained to be assistants, not people.

The benchmarks that drive general-purpose LLMs — MMLU, GPQA, AIME and MATH, SWE-bench and HumanEval, ARC-AGI — all reward reasoning, knowledge retrieval, and coding tasks. Not one of them measures whether a model can behave like a person.

LLMs are instruction-tuned with RLHF to be helpful personal assistants that serve complex tasks. Real people are not that. They behave illogically, get frustrated, stay ambiguous, refuse to cooperate, and act on motivations of their own. Simulated users by general-purpose LLMs, by contrast, come out too cooperative, too uniform, and never frustrated or ambiguous — the same failure across every behavioral dimension.

Further reading
03

The rise of behavior models

SAPIENS benchmarked against the other behavior models.

As the academic community recognizes that general-purpose LLMs are not good at modeling behavior, there have been several attempts at purpose-built human behavior models — Centaur and Be.FM among them. We benchmarked against them.

behavior prediction accuracy · purpose-built behavior models

SAPIENS at 86% versus Be.FM at 71% and Centaur at 57%.
The research
04

Benchmarking dataset & study design

Benchmarking dataset

We chose the Amazon Public Dataset — the largest public record of human behavior, spanning categories: health and wellness products, video games, software products, and more.

The dataset covers millions of users across very different audiences, with extensive documented purchase history and their opinions on the products they bought. Predicting opinions across a wide variety of products is a real test of whether behavior has been captured — and it's what lets us prove SAPIENS works across categories, not just one.

Study design

The task: given a selected user and their history, predict their opinion on a new product. If the generated opinion and the real opinion match closely, that's a success. For SAPIENS, this product is genuinely unseen — it never appears in training data.

How we score a prediction

We compare the model's generated review against the person's real one on meaning — a prediction only counts when it says what they said and feels how they felt.

Same topics

Did it surface what the person actually chose to write about?

Same features

Did it react to the same product details, in the same way?

Same sentiment

Did it land on the same feeling — love it, live with it, return it?

05

The data moat

The dataset of human behavior.

SAPIENS is powered by dense human behavior data — which is what makes it accurate enough to model any population segment with confidence. That density is cross-pollinated across sources, because no single source covers a human.

Read: Creating the dataset of human behavior

What it takes to cover a human.

01

Life background

Who someone is before they meet a product — life stage, household, income, geography, career arc, the priors they carry into every decision.

Without this, a persona is a costume.

02

Cross-topic observation

The same human watched across hundreds of unrelated contexts — money, health, gaming, parenting, work — not just inside one category.

Behavior transfers between domains. Single-vertical data can't learn that.

03

Exposure

The stimulus around the decision — the ads, feeds, prices, peers, reviews and rival products a person had actually seen before they chose.

Choices are reactions. Model the input or you're guessing at the output.

04

Revealed action

Revealed preference — what got clicked, bought, abandoned, cancelled, re-opened. The action that quietly contradicts the stated intention.

Everyone says they'd pay. The dataset knows who did.

Data sources that power SAPIENS for enterprises.

Source
Life background
Cross-topic observation
Exposure to stimulus
Revealed action
Access
Public data People narrate their own lives, unprompted — the job change, the diagnosis, the move, the breakup, the new baby. The richest biography layer we have.
RedditLinkedInYouTubeApp StoreX
Life background
Cross-topic
Exposure
Revealed action
Open
Enterprise data — qualitative What customers say directly to a company when something matters enough to say it out loud. Reasons, objections, intent.
GongIntercomProductboardZendeskSurveys
Life background
Cross-topic
Exposure
Revealed action
Enterprise
Enterprise data — telemetry The behavioral record. Sessions, funnels, spend, churn — what people did when nobody was asking them a question.
AmplitudeGoogle AdsMixpanelBilling
Life background
Cross-topic
Exposure
Revealed action
Enterprise
Vectorial Pulse Holistic behavior understanding through AI-moderated interviews that ask hard, probing questions — mapping behavior across multiple dimensions in a single conversation.
Vectorial PulseAI-moderated interviews
Life background
Cross-topic
Exposure
Revealed action
Proprietary data
Kentrix.ai Household- and segment-level demographic and psychographic structure. Gives a simulated user a plausible life before the first prompt is written.
Kentrix.aiIdentity depth
Life background
Cross-topic
Exposure
Revealed action
Partnership
Konovo An intelligent platform built on the broadest ecosystem of healthcare audiences — bridging the gap between raw data and a genuine understanding of your customer in a fraction of the time you would expect.
KonovoHealthcare audiences
Life background
Cross-topic
Exposure
Revealed action
Partnership
Prolific The largest vetted user panel, integrated directly into the platform. Put a stimulus in front of a matched cohort and measure what real humans actually do.
ProlificActive panel
Life background
Cross-topic
Exposure
Revealed action
Integration

// Coverage scored across the four dimensions. Directional, not benchmarked.

The relationships that widen the moat.

Partnership
Kentrix.ai logoKentrix.ai

Identity depth. Household and segment-level demographic and psychographic structure, so a simulated user arrives with a plausible income, life stage, geography and set of values — not an invented backstory.

Fills life background
Partnership
Konovo logoKonovo

Healthcare depth. An intelligent platform built on the broadest ecosystem of healthcare audiences, bridging the gap between raw data and a genuine understanding of your customer in a fraction of the time you would expect.

Fills life background and revealed action
Integration
Prolific logoProlific

The largest vetted user panel, wired directly into the platform. Any simulation can be put in front of a matched human cohort and scored — turning every run into a measurable claim rather than an opinion.

Fills exposure and revealed action — and grades the model

Every simulation by a customer improves SAPIENS.

SIMULATE VALIDATE LABEL RETRAIN SAPIENS BEHAVIOR MODEL

SAPIENS vs. general-purpose LLMs

06

Why SAPIENS outperforms general-purpose LLMs at predicting human behavior.

SAPIENS — our behavior modeling approach — is roughly 30 percentage points better at predicting human behavior than general-purpose LLMs. Over 18 months we have benchmarked successive generations of frontier instruction-tuned models: despite substantial gains in reasoning and problem-solving, their accuracy on our human behavior benchmark has plateaued at 49–54%. Further academic research suggests they are becoming worse.

Behavior models differ from general-purpose LLMs on three axes — modeling objective (what they optimize for), modeling granularity (what they model), and data. Here is a detailed comparison on the same.

SAPIENSBehavior model
General-purpose LLMsInstruction-tuned frontier models

Trained to model human behavior — the irrational.

In order to predict what a person does, you need to understand the why behind their decisions and the behavior they exhibit. Behavior models focus on those whys, with the goal of predicting how a person will react to different stimuli.

People do not decide rationally. They decide through unstated perception gaps, biases, and the environmental factors acting on them — the reasons they would never give you in a survey, and often could not. We call these latent traits, and they are what SAPIENS learns and models.

SAPIENS learns those latent traits alongside a person's background, behavioral history, psychographics and socio-demographic traits — from vast amounts of public and proprietary data — because together they are what actually predicts what a person does.

Modeling
objectivevs.

Trained to model rational problem solving, to solve complex tasks.

Frontier AI labs are optimizing to replace knowledge workers on benchmarks such as SWE-bench, GPQA, AIME and MMLU. These benchmarks reward step-by-step rational decision making.

RLHF instruction-tuning optimizes for the most plausible, step-by-step logically consistent response — the machinery that wins at analytical tasks, strategic tasks, math, science and coding. Research on human behavior simulation suggests that RLHF and instruction fine-tuning is what makes these models worse at understanding human behavior.

The model is not designed to follow the trajectory of irrational decision making and behavior of different individuals; it is designed to gather and learn domain knowledge across many verticals.

We model each individual “real user”, then the collective target audience they belong to.

Because users are modeled one at a time, the diversity and behavioral distribution of the real population is preserved rather than averaged away.

Every response is therefore grounded in an individually modeled real user — each one carrying their own experiences, preferences, attitudes and latent behavioral traits, learned from real-world data.

Modeling
granularityvs.

One model that represents the entire world's knowledge.

Asked to represent a population, it returns an averaged, rational opinion about that population — which does not represent the variance, and is always seen through the lens of rationality, giving an understanding of the domain at large.

At best, it will return you previously conducted surveys and studies from the internet.

A spectrum of behavior — breadth and depth, per person.

We collect a spectrum of behavioral data on the individuals inside a specific target population: breadth across many sources, and depth on each person within them.

Breadth tells us what someone does. Depth tells us why they do it. Behavior prediction needs both.

And because behavior keeps evolving, so does the data. SAPIENS keeps learning from evolved behavior at every step, staying current with the population it models.

Datavs.

A corpus of world knowledge — a broad spectrum of domains ingested, with no individual-level backgrounds or trajectories explicitly captured.

The data is gathered to cover domains — enough of them, well enough, to perform logical tasks in that domain. No granular, person-specific behavioral data is collected at all.

What the corpus holds is a generalized opinion across domains. Wide, shallow, unattached to any individual — and silent on the why.

And it is frozen. A general-purpose LLM is bounded by its last training date, while the behavior it is being asked to predict keeps moving.