RESEARCH

SAPIENS learns human behavior from large-scale observation data

We introduce SAPIENS, an AI framework that directly learns rich behavioral traits from large-scale observational data. SAPIENS analyzes real user opinions to learn the latent traits and recurring experience patterns that drive how different individuals respond to stimuli, and uses these signals to instantiate an explicit synthetic population: personas with learned characteristics and backstories.

84%topic-prediction accuracy
30%higher than Be.FM, Centaur, Claude Sonnet 4 and GPT-5

Inspired by TextGrad's use of natural-language feedback based on gradients, and by matrix-factorization learning algorithms common in recommender systems, SAPIENS iteratively refines persona traits to reduce the gap between synthetic and real opinions, without fine-tuning the backbone model separately for each user. Synthetic users formed by SAPIENS significantly outperformed other behavior foundation models and standard LLMs in predicting human opinion given a product description.

WHY IT MATTERS

Synthetic users, applied

Creating synthetic users with fundamental traits across demographics, psychology, and experience, learned from real human opinions, has broad applications. Enterprises use them to run simulations at different stages of product development: understanding pain points, evaluating graphical user interfaces, and testing conversation and voice interfaces for feedback that reflects real human opinion.

Synthetic users are also important for training foundation models on reward modeling for different user populations, helping those models cater to different intents and preferences and give better personalized responses. This has significant implications for training the next generation of foundation models.

RELATED WORK

Current research on synthetic users, from Stanford and UC Berkeley

Recent work on synthetic users and behavior modeling has made important progress but leaves key gaps. Anthology conditions LLMs on rich, survey-style backstories and improves consistency and subgroup fidelity in virtual personas, yet these backstories are still largely synthetic narratives rather than traits learned from large-scale logs of how people actually act (Moon et al., 2024).

PersonaHub and Personas within Parameters move closer to data-driven personas: PersonaHub by auto-curating a billion text-based personas from web data for persona-conditioned synthetic data generation, and Personas within Parameters by fine-tuning small LMs with low-rank adapters on interaction logs to mimic specific users. Both ultimately tie behavior to particular datasets and model checkpoints; Personas within Parameters in particular requires a separate adaptation run per user, making it difficult to scale to millions of distinct personas (Ge et al., 2024; Thakur et al., 2025).

Stanford HAI's Generative Agent Simulations of 1,000 People shows that agents grounded in two-hour qualitative interviews can closely replicate survey responses and experimental outcomes for 1,052 real individuals, but this architecture depends on expensive, interview-style data collection and is optimized for reproducing survey attitudes rather than continuous product-interaction behavior at population scale (Park et al., 2024).

Behavior-oriented foundation models such as Be.FM and Centaur instead train a single large model on diverse behavioral or cognitive datasets to predict human decisions across many tasks (Xie et al., 2025; Binz et al., 2024). But because all behavior is encoded in one monolithic predictor rather than an explicit, inspectable synthetic population of individuals, these approaches make it hard to study population structure, attach persistent persona profiles, or run controlled "what-if" simulations on specific user segments.

WHERE SAPIENS DIFFERS

SAPIENS addresses these gaps by learning an explicit population of synthetic users whose traits are inferred directly from large-scale observational data, rather than from LLM-generated backstories or per-user fine-tuning. It builds high-dimensional, explainable traits from many historical opinions per user and uses these traits to predict reactions to new products. This produces more grounded, adaptable, and interpretable personas than prompt-only or monolithic behavior foundation models, while remaining practical to scale to millions of synthetic users.

BENCHMARKING

Performance, dataset, and methodology

Our synthetic user models achieve 84% accuracy in predicting which topics a user will focus on when reacting to a new product, for example price, reliability, usability, or support. This is about 30% higher than behavior foundation models Be.FM and Centaur, and general-purpose LLMs Claude Sonnet 4 and GPT-5, on the same benchmark.

Vectorial (SAPIENS 1)
84%
Be.FM
71%
Centaur
57%
Claude Sonnet 4
56%
GPT-5
54%
Success rate predicting the correct opinion topic for a persona, given a product description and its features.
DATASET FOUNDATION

Data source and prediction task

We ground our models in the Amazon product reviews dataset, a comprehensive public repository containing millions of authentic user opinions across diverse product categories. This dataset provides the real-world behavioral foundation necessary for training synthetic users in the e-commerce domain that reflect actual human responses. Access: UCSD Dataset Repository.

Given a product or feature description, our models predict the specific opinion topics that real users will discuss, not just whether feedback is positive or negative, but what aspects they'll talk about. We then measure accuracy by comparing predicted topics with actual topics.

The choice of Amazon reviews as our foundational dataset reflects a deliberate strategy: these reviews represent authentic, unstructured user feedback across an enormous variety of products and use cases. Users write about what matters to them, in their own words, without the constraints of surveys or interview scripts. This natural language data captures the true diversity of user concerns, priorities, and mental models.

WHY THIS TASK IS CHALLENGING

A single product description can trigger a wide range of possible reactions: one user might focus on privacy, another on ease of setup, and yet another on long-term durability. Even when we provide past reviews, they don't directly tell us what a user will talk about next. The model has to infer deeper traits and habits, what this person tends to care about, and then generalize them to a completely new product and context.

HOW WE EVALUATE

SAPIENS learns from each user's historical reviews to understand those underlying traits and preferences. At test time, every model, SAPIENS, Be.FM, Centaur, Claude Sonnet 4, and GPT-5, gets the same inputs: a description of an unseen product plus a sample of that user's previous reviews. Each model must predict the main topic that user will mention in their next review. We then compare the predicted topic with the ground-truth topic from the real review to compute accuracy.

DEEP DIVE

Learning user traits from real observation data

Our architecture is inspired by algorithms used in recommender systems like matrix factorization, a popular way to model a user's latent traits and an item's latent traits to predict which item a user is going to like next. We took it a step further to describe how these traits lead to certain preferences and decision making. Each time we compare a synthetic user's predictions to real user responses, we learn something about how that persona's traits should be calibrated: did we overestimate their price sensitivity, or underestimate their focus on aesthetics? These insights get encoded back into the user model, making subsequent synthetic opinions more closely matched to real ones. Over time, this creates a system that continuously learns and adapts to emerging patterns of diverse users and their decision making.

Generated versus actual opinion vector spaces, delta calculation, and user profile enrichment
Comparing generated opinion embeddings against real opinion embeddings to calculate delta and enrich persona traits, over N iterations.
1

Present product description

Synthetic users receive detailed product or feature descriptions, just as real users would encounter them in marketing materials or product launches.

2

Generate initial response

Based on their trait profiles, synthetic users produce predictions about which topics they would discuss and what concerns they would raise.

3

Compare to real feedback

We validate synthetic responses against actual user opinions from similar products, identifying matches and misalignments.

4

Update persona traits and backstory

Prediction errors trigger adjustments to user trait profiles, refining how characteristics map to opinion topics for future predictions.

PERSONA ARCHITECTURE

Layered trait modeling

The accuracy of synthetic user predictions depends fundamentally on how we model user characteristics. SAPIENS looks at users across multiple layers that capture the different dimensions of user identity and behavior that influence decision-making. This reflects a key insight: opinion topics are context-dependent, and predicting them requires understanding not just what users do, but why they do it, what constraints they face, and their past experiences.

Observational data, behavioral, psychographic and socio-demographic layers stacking into persona and industry clusters
Layers build from raw observational data up to differentiated, industry-relevant personas.
Behavioral patterns

How a user approaches decisions, whether they dive deep into technical details or prioritize ease of use.

Psychographic traits

Personality, and how it influences decision-making.

Background and experience

What a user is capable of evaluating, and what they will compare against.

Role and domain context

The broad domain to which a user's opinions will be related.

Two users might both give a product four stars, but one discusses pricing transparency while the other focuses on integration capabilities. The difference comes from their trait combinations. Consider two Amazon reviews with the same rating for the same pair of headphones: one from a daily commuter prioritizing battery life and noise cancellation, versus an audiophile focusing on subtle sound nuances and comfort for extended listening sessions.

CONTINUOUS UPDATES

Perception, knowledge, and forming intuition

User behavior evolves constantly: new product launches, technology shifts, evolving expectations, changing market conditions, and emerging technologies create new possibilities. Static user models quickly become obsolete in this dynamic environment. Vectorial AI addresses this through two foundations: a continuous data pipeline that ingests observation data from public sources such as LinkedIn, Reddit, and YouTube, and continuously learns as enterprises launch new products and features; and a three-layer continuous update system of cognition that keeps synthetic users current with evolving contexts.

Perception layer

Continuously reads new signals about how users perceive public and company events related to the product, tracking social media discussion, reviews, and shifts in behavior and key events.

Knowledge layer

Organizes incoming signals into a structured map, linking topics discussed across many different sources so scattered insight consolidates into one picture.

Intuition layer

Creates an episodic memory that tracks context, time, and reactions, recording how synthetic users responded to particular stimuli so the system can refine its predictions from past outcomes.

Public data, enterprise qualitative and telemetry sources, and real-user enrichment feeding synthetic users
Public, enterprise, and real-user data sources feed a shared observation pipeline.
Agentic sensors, knowledge graphs, episodic memory and synthetic user segments stacked as perception, knowledge and intuition layers
Perception, knowledge, and intuition layers compound into continuously updated synthetic user segments.

The continuous update system ensures our predictions remain accurate even as market conditions change, product positioning evolves, and user expectations shift. We're not predicting based on static historical patterns. We're simulating how users would respond today.

MISSION

Representing the entire world population and behavior

Vectorial AI is built on a simple but ambitious belief: to truly understand and model human behavior. Human-behavior simulation is emerging as a new branch of AI, distinct from LLMs that primarily capture world knowledge and from world models that simulate physical dynamics. While AGI aims for general cognitive capability, our north star is Behavior General Intelligence (BGI): systems that can represent how diverse people think, feel, and act across contexts. SAPIENS v1.0 is our first concrete step toward this goal, a state-of-the-art engine for creating synthetic users.