Patrick de Carvalho
EN
All articles
Patrick de Carvalho August 21, 2026 · 24 min

8.3 billion virtual humans: I read the study everyone is commenting on without having opened it

I read the study everyone is commenting on without having opened it. The decisive figure is not 8.3 billion, it is 98.3% against 27%.

By Patrick de Carvalho, CEO Apps Velocity

Contents

What MatrAIx really changes for a company director, and the figure nobody quotes

By Patrick de Carvalho, CEO and co-founder of Apps Velocity


In short: researchers from Harvard and MIT have released a method for simulating 8.3 billion synthetic personas, meaning consumer profiles generated by artificial intelligence. The figure that matters is not the size of the simulated population but the gap between models: 98.3% agreement for one, 27% for the other, on the same question. A simulation therefore does not measure a market. It measures, first of all, the model you picked.


Act 1: in 1998, finding out what my readers wanted cost me six months

In 1998 I launched planetepresse.com, one of the very first online newsstands in France. Selling newspapers and magazines over the internet, at a time when the vast majority of French people were still connecting over RTC, the switched telephone network, with a modem that crackled for forty seconds before a page appeared.

At that point I was asking myself exactly the same question any company director asks today: are people going to buy?

To answer it, I did what you did back then. I called newsagents. I stopped readers in the street. I paid for a small qualitative study on about thirty people. I spent hours on the phone with press publishers who could not understand why I wanted to put their titles "on the computer". I built an intuition by hand, one person at a time, one conversation at a time.

Six months. Maybe more. And at the end of it I had a sample of a few dozen opinions, biased by the fact that the only people willing to talk to me about an online newsstand were already, by construction, fairly curious about technology.

I remember the feeling very clearly: spending an enormous amount of energy to obtain a weak signal, blurred, and probably wrong.

Twenty-eight years later, a team of researchers from Harvard and MIT publishes work that claims to compress those six months into a single night. And to do it not on thirty people, but on a virtual population the size of the Earth's.

The project is called MatrAIx. It was filed on 4 August 2026 on arXiv, the open-access scientific publication platform where researchers deposit their work even before peer review. Since then the story has been everywhere, often badly summarised, often amplified.

I took the time to read what was actually published rather than what was said about it. Here is what I take from it, and above all what it changes for you if you run a small or mid-sized company.


Act 2: what MatrAIx actually is, with the facts

Let us set out the terms first, because half the misunderstandings start there.

An AI agent is a program that uses an artificial intelligence model to carry out a task autonomously, chaining decisions together without a human validating each step. It is not a chatbot answering a question, it is an operator pursuing an objective.

An LLM, for Large Language Model, is the artificial intelligence engine trained on enormous volumes of text and capable of producing coherent language. GPT, Claude, Gemini and Mistral are LLMs. It is the building block of everything we have been talking about for three years.

A persona, in marketing vocabulary, is a composite portrait of a typical customer. You probably have some in a folder somewhere: "Sophie, 42, procurement manager in an industrial SME, price sensitive, decides within three weeks". A Word document nobody ever reads again.

MatrAIx takes those three notions and merges them at a scale nobody had attempted before.

The Persona 8B database

The heart of the system is a database called Persona 8B. It contains 8.3 billion individual profiles, a figure calibrated to match the current world population. Each profile is described across 1,290 dimensions.

One thousand two hundred and ninety. Take a second on that number.

We are not talking about age, sex, socio-economic bracket. We are talking about the whole picture: life context, psychological traits, level of technical literacy, consumption habits, risk tolerance, decision style, right down to very fine-grained behaviours of the "how long does this person hesitate when a price goes up" variety.

Of those 8.3 billion profiles, roughly one million has been made public. And this is an important point: within that million, 600,000 are grounded in real anonymised human data, and 400,000 are synthetic, meaning artificially generated from dependency graphs that guarantee the internal coherence of the profile.

Put another way, the majority of the public sample is not invented. It is derived from real people.

The four test environments

A persona stays a dead document until a model makes it speak. That is the second component of the system: the agents are placed in situations across four types of environment.

Survey forms. Chatbot interfaces, meaning conversations with an AI. Web browsers, where the agent genuinely navigates pages. And applications, mobile or desktop.

That is not trivial. Many earlier simulations settled for asking an AI questions along the lines of "what would this person do". Here the agent acts inside an interface. It clicks, it scrolls, it abandons, it comes back.

The evaluation tasks

The third component is the tasks. The published work reports more than 1,000 standardised tasks spread across more than 25 domains: commerce, software, finance, health, cybersecurity, and others.

One of those tasks measures, for instance, a user's willingness to keep trusting an AI assistant after a first failure. That is a subject directly tied to the reliability of autonomous agents, and it is exactly the kind of question you cannot measure any other way than by observing real users over weeks.

The validation

The researchers ran a control experiment across 400 trials, measuring how closely the agents adhered to ten behavioural attributes assigned to them. The stated behaviour was correctly expressed or correctly suppressed in 366 cases, meaning 91.5%.

Human judges separately rated the quality of the personas grounded in real data at around 4.1 out of 5.

And the whole thing is published as open source, a term meaning that the source code, meaning the program's instructions, is made available to everyone, free of charge, with the right to modify and reuse it. Here under the MIT licence, one of the most permissive licences in existence. The repository is called MatrAIx-Persona-8B, it is hosted on GitHub, the reference platform for sharing code, and the one-million-profile dataset can be downloaded from Hugging Face, the reference platform for sharing AI models and datasets.

The motto displayed on the repository: Simulate Before Reality.


Act 3: the guardrails I set before going any further

I am going to do something most articles on this subject do not do. I am going to say what MatrAIx is not, because that is where the credibility of an analysis is decided.

No, nobody is simulating 8.3 billion humans in parallel. The work describes evaluations run on cohorts sampled from the database. A cohort is a subgroup selected on shared criteria in order to be observed. You draw a few thousand or a few tens of thousands of profiles, you make them act, you analyse. Nobody switches on the entire Earth on a server. The 8.3 billion figure describes the size of the reservoir, not the size of the run.

No, the 8.3 billion profiles are not public. What is available is a sample of one million, filtered on quality. That is already enormous. It is not the same thing.

No, these are not humans. They are statistical distributions to which a language model lends a voice. The nuance is fundamental and I come back to it in a moment, because that is where the real subject lies.

And on the number of researchers, caution. Part of the press announces more than 200 scientists, around forty of them from OpenAI, Anthropic, Google DeepMind and xAI. Other analyses, counting the signatures directly on the document filed on arXiv, count 93, led by Harvard and MIT with contributions from Stanford. The team is large, that much is certain. The figure of 200 deserves to be treated as a press estimate and not as an established fact, at least until you have checked it yourself.

I set these guardrails because I have seen too many LinkedIn posts turn a research publication into prophecy. When you want to be taken seriously by directors who are committing their own money, you prefer an argument that survives the first person who pushes back.


Act 4: the figure almost nobody quotes, and which is worth the whole study

Now, the passage that made me put my coffee down.

The researchers ran a price sensitivity experiment. Same web page, same product, same price increase, same cohort of personas. One single variable changes: the AI model playing the characters.

The result.

With GPT-5.5 at the controls, 98.3% of the personas hesitated to buy.

With Claude Opus 4.8 at the controls, on the same cohort, the reluctance falls to 27%.

Read that again. Same virtual population. Same stimulus. A gap of more than seventy points.

There it is. That is the AI-applied-to-business story of the year, and it is buried in the middle of articles running under a headline about the Matrix.

Why that figure matters more than the 8.3 billion

Because it destroys the illusion of neutrality.

When you launch a simulation, you believe you are questioning a market. In reality you are questioning a model that is interpreting a description of a human. And every model has been trained differently, aligned differently, calibrated differently on caution, on enthusiasm, on the propensity to say yes or no.

A bias, in artificial intelligence, is a systematic distortion of results produced by the training data or by the way the model was tuned. It is not a one-off error you correct. It is a structural inclination.

What the price experiment shows is that the bias of the underlying model can weigh more heavily on the outcome than all 1,290 dimensions of the simulated human profile put together.

Translation for a company director: if you set your pricing from a simulation, you are not making a decision about your market. You are making a decision about the AI supplier your service provider happened to pick that day. Possibly because it was cheaper per API call, the technical interface that lets two pieces of software talk to each other and that is billed by usage.

You can see the problem.

In one case you conclude that your market is highly price sensitive and you give up on the increase. In the other you conclude there is headroom and you put prices up 15%. Two opposite strategies, one single reality, and the deciding factor has nothing to do with your customers.

That is not a reason to throw the tool away

Careful, I am not saying MatrAIx is useless. I am saying exactly the opposite.

A badly calibrated thermometer is still a thermometer. What matters is knowing that it is badly calibrated, and using it for what it can do: measuring gaps, not absolute values.

If you test two versions of a page with the same model, under the same conditions, the relative comparison keeps its value. It is when you take the absolute figure literally that you put your company at risk.

That distinction between absolute value and relative gap is probably the single most useful skill a company director can develop in the face of AI in 2026.


Act 5: what this changes concretely for a French SME

Let us get down to the ground. You run a company of thirty, a hundred or three hundred people. You have no research department. You do not have 80,000 euros to put into a consumer panel. What does this news actually give you?

1. Access to something you did not have

Let us be clear on the starting point. The vast majority of French SMEs run no user testing at all. Zero. Not a bad test, not a biased test: no test.

You launch an offer because the director believes in it. You change a price because the margin is eroding. You redesign a website because the old one looked dated. Decisions get made on instinct, on experience, sometimes on the last customer comment heard on the phone on a Tuesday morning.

A biased simulation is infinitely better than no test at all, on one condition: that you know it is biased. That is the whole difference between an instrument and a belief.

2. An iteration speed that did not exist

Recruiting a panel takes weeks. Doing it again after modifying the product takes weeks more. Which is why, in real life, you never do it again. You test once, you decide, you live with it.

With a simulated cohort, you rerun the test after every modification. On Monday you test version A, on Tuesday you fix it, on Wednesday you test again.

That capacity to iterate fast is exactly what we applied at Apps Velocity on Constrok, our ERP built for house builders and project managers in the construction sector. An ERP, for Enterprise Resource Planning, is the central piece of software that runs all of a company's processes: quotes, purchasing, sites, invoicing, accounting. Historically, that is an eighteen-month project.

We developed and deployed it in one and a half months. Not because we are cleverer than anyone else. Because we compressed the feedback cycles. Every hypothesis was confronted with reality in hours, not in quarters.

MatrAIx and tools of the same kind give teams far smaller than ours access to that type of compression.

3. A pre-filter, never a decision

Here is the rule I would apply and that I recommend without reservation.

A simulation is there to eliminate bad ideas, not to validate good ones.

If a simulated cohort massively rejects your concept, whatever model you used, whatever the wording, that is a signal. Not a conviction, but a genuine signal, and it cost you an evening instead of three months.

If a simulated cohort loves your concept, that proves strictly nothing. It gives you permission to move to the next step: talking to real customers. Which remains, and will remain, the only source of truth.

That is the logic of a filter: you put a lot in at the top, you keep only what survives, and what survives goes to the real test.


Act 6: the three traps I guarantee companies will fall into

Trap 1: the mirror that looks like you

The most insidious risk is not technical, it is human.

When you configure a simulation, you describe your target customer. So you are describing your representation of your customer. You inject your assumptions into the system, and the system, obligingly, hands them back to you dressed up as data.

You think you have measured your market. You have measured your own point of view, with a layer of statistical credibility on top.

That is confirmation bias industrialised. And it is dangerous precisely because it produces tables and percentages, meaning everything that silences objections in a meeting.

There is a counter: have your assumptions tested by someone who has no interest in them being true. Have the personas configured by a person who did not write the product brief. Deliberately go looking for the cohort that ought to hate your offer.

Trap 2: the closed loop

Imagine this becomes widespread. Genuinely.

Companies design products tested on virtual populations. Those products reach the market. They meet real humans, who produce data. That data is used to improve the models. The models are used to generate new personas. On which the next products are tested.

At every turn of the loop, the share of human unpredictability shrinks. You optimise for a statistically average consumer, who exists nowhere, and who progressively becomes the only target you know how to aim at.

What I fear in that scenario is not the disappearance of humans. It is the disappearance of surprise. And innovation never comes from the centre of the distribution. It comes from the margins, from repurposed uses, from customers doing something with your product that you had not anticipated.

A system that simulates the average perfectly is a system structurally incapable of spotting what is going to break the market.

I would rather say it plainly: if you never talk to a real customer again, you will never again learn anything you did not already know.

Trap 3: personal data

A reminder of the figure: of the million public profiles, 600,000 are grounded in real anonymised human data.

Anonymised. The word is worth pausing on, because the GDPR, the General Data Protection Regulation, draws a very sharp line between anonymisation, which makes identification irreversibly impossible, and pseudonymisation, which replaces identifiers but leaves a possibility of re-identification.

The difference is not cosmetic. Genuinely anonymised data falls outside the scope of the GDPR. Pseudonymised data remains fully within it.

And academic research has shown on many occasions that a profile carrying more than a thousand attributes becomes extremely hard to anonymise properly. The more dimensions you add, the more unique every combination becomes. That is a mathematical problem, not a problem of good intentions.

I make no accusation against the MatrAIx team, who published their methodology and submitted it to the scrutiny of their peers, which is exactly the right practice.

But I say this to any director tempted to reproduce the approach in-house: if you build personas from your customer database, with hundreds of behavioural attributes, you are entering territory where legal analysis is not optional. It is not a compliance detail you settle afterwards. It is an architecture decision you take beforehand.


Act 7: how I would use it on Monday morning

Enough principles. Here is what I would concretely do if I ran an SME and wanted to get something useful out of all this, without spending a large-group budget on it.

Step 1: pick a real and costly decision. Not a demonstration exercise. A genuine pending decision: a price increase, a change to a subscription plan, a new welcome message on your site, a reorganisation of your ordering journey. Something that costs you money if you get it wrong.

Step 2: write down the hypotheses before testing. On a sheet of paper, before any simulation, note what you think is going to happen. With figures if possible. That sheet is your protection against confirmation bias. Without it, you will always find the simulation result "consistent with your intuition", because memory is an obliging narrator.

Step 3: run the same test on at least two different models. That is the direct lesson of the 98.3% against 27%. One model gives you an opinion. Two models give you a range. If the two converge, your signal is solid. If they diverge violently, you have just learned that the question cannot be settled by simulation, and that information has value.

Step 4: work on the gaps, not the values. Never take away "62% of our customers would accept this increase". Take away "version B is rejected half as often as version A, and that ratio holds across both models tested". The first statement is fiction. The second is a decision-making tool.

Step 5: confront it with five real customers. Five. Not fifty. Five twenty-minute phone conversations with real customers, chosen to be different from one another. If the real customers contradict the simulation, the real customers are right, always, no discussion.

That sequence fits into a week. It requires no team recruitment. And it turns a decision made on instinct into a decision made with an instrument, which is not the same thing at all.


Act 8: the distinction I repeat at every conference

There is one distinction I set out systematically in front of company directors, because it separates those who will extract value from AI from those who will spend money for nothing.

Operational AI automates tasks. It drafts a report, sorts emails, generates a quote, chases an unpaid invoice. The gain is real, measurable, immediate. It is also limited: you save time on what you were already doing.

Strategic AI helps you decide. It sheds light on a choice you could not have illuminated any other way, because the information did not exist, or cost too much to produce.

MatrAIx clearly belongs to the second category. It is the first time a tool of this nature has become accessible to organisations that cannot afford a research department.

And that is also why it is more dangerous. An operational AI that gets it wrong costs you an hour. A strategic AI that gets it wrong costs you a year, because you have committed the entire company to a direction founded on a hollow figure.

The closer the tool gets to the decision, the more solid the human expertise framing it has to be. That is mechanical.

I have a formula for this, and I have stuck to it from the start: AI multiplies existing expertise, it does not replace it. A director who has known their market for fifteen years and adds simulation to their arsenal becomes formidable. A director who does not know their market and replaces that knowledge with a simulation has just bought a very handsome high-speed error generator.

It is exactly the same reasoning I apply to vibe coding, the practice of having an AI produce code from instructions written in everyday language. In the hands of someone who knows what a software architecture is, it is a spectacular accelerator. In the hands of someone who does not, it is building a house without an architect or a qualified bricklayer: it stands up until the first winter.


Act 9: the real lesson, and it is not technological

I come back to 1998, one last time.

My six months of investigation on planetepresse.com did not serve to get me figures. They served to make me understand something I would never have guessed: people did not want to read the press on a screen. They wanted to stop having to go out.

That is not the same thing. And no quantitative study would have given it to me, because I did not know that was the question to ask.

That is what I fear most with the spread of simulated populations. It is not that they give bad answers. It is that they answer perfectly the questions you put to them, and that they will never tell you that you are asking the wrong one.

The role of a company director, in ten years' time as much as today, is not to collect answers. There will be plenty of those, instantly, free of charge. It is knowing which one to look for.

Every major technological rupture I have lived through from the inside, the internet bubble, mobile, cloud, subscription software, and now agentic AI, has followed the same trajectory. They destroyed neither the jobs nor the skills that came before them. They forced them to mutate.

Market research is not going to disappear. It is going to shift. The mechanical part, recruitment, fieldwork, data processing, is going to collapse in cost and in lead time. The noble part, framing the right question, interpreting what cannot be measured, deciding when the data contradicts itself, is going to take on a value it has never had.

Those who think MatrAIx replaces market research have not understood what is happening. Nor have those who think it changes nothing.

What is happening is that the cost of being wrong has just fallen enormously. And when the cost of error falls, the right strategy changes completely: you stop trying to be right first time, and you start testing a lot, fast, and correcting even faster.

That is very precisely the philosophy I have always built my companies on. The idea is worth zero, only execution counts. And execution, in 2026, is the capacity to close learn, correct, relaunch cycles faster than your competitors.

MatrAIx, for a clear-headed director, is an accelerator of those cycles. For a director in a hurry, it is a machine for manufacturing false certainties with decimal places.

The difference between the two is not decided in the tool. It is decided in the head of the person using it.


FAQ

What is MatrAIx?

MatrAIx is an evaluation infrastructure published on arXiv on 4 August 2026 by a team led by researchers from Harvard and MIT. It combines a database of 8.3 billion individual profiles described across 1,290 dimensions, four types of simulation environment, and more than 1,000 tasks spread across more than 25 domains. The code is published under the MIT licence.

What is a synthetic persona?

It is an artificially generated consumer profile, which is then played by an artificial intelligence model placed in a situation. Of the million profiles made public, around 600,000 are grounded in real anonymised human data and 400,000 are entirely synthetic, produced from dependency graphs that guarantee the internal coherence of the profile.

Can a simulation replace market research?

No, and the study itself does not claim it can. A simulation reliably measures relative gaps between two options tested under the same conditions. It does not produce an absolute value on which to base a pricing structure. The distinction between relative gap and absolute value is the decisive skill when facing these tools.

What does the 98.3% against 27% gap mean?

It is the result of a price sensitivity experiment: same page, same product, same increase, same cohort of personas, one single variable changed, the model playing the characters. With one, 98.3% of the personas hesitate to buy; with the other, 27%. More than seventy points of difference. So the simulation does not first measure a market, it measures the model that was chosen.

Is the method validated?

Partially, and that has to be said as it is. A control experiment across 400 trials measuring the agents' adherence to ten behavioural attributes gives 91.5% of behaviours correctly expressed or suppressed, and human judges rate the quality of the grounded personas at around 4.1 out of 5. This is a filing on arXiv, meaning work made public before peer review.

Can an SME use it today?

Yes, provided you fix the protocol before launching anything: one single model, one single version, identical conditions across the variants tested, and a reading in gaps and never in absolute values. Changing model mid-campaign invalidates every comparison.

Is there a personal data risk?

Yes, and it is real as soon as you ground personas in your own customer data. The GDPR, the General Data Protection Regulation, applies to any data that allows a person to be identified, directly or indirectly. A database of profiles derived from real customers remains a personal data database for as long as re-identification is possible.


Sources

  • MatrAIx: Simulating the World with 8.3 Billion Persona Agents, preprint filed on arXiv on 4 August 2026 by a team led by researchers from Harvard and MIT. arXiv:2608.04205. arxiv.org/abs/2608.04205
  • Discussion page for the work on Hugging Face, which aggregates the exchanges around the preprint. huggingface.co/papers/2608.04205
  • MatrAIx-Persona-8B code repository, published under the MIT licence, with the motto Simulate Before Reality. github.com/MatrAIx-ai/MatrAIx-Persona-8B
  • Persona 1M dataset, the million profiles made public, downloadable from Hugging Face.
  • Covered in French by Futura-Sciences, 19 August 2026.

The arXiv filing is a publication made ahead of peer review. The figures quoted here are those stated by the authors and have not been independently replicated as at the publication date of this piece.


What I am offering you

If you run an SME and you are wondering where to start with AI, without losing your budget or your team over it, I have formalised a structured methodology called RAPID: Recenser, Analyser, Piloter, Itérer, Déployer (identify, analyse, steer, iterate, deploy). Five steps, designed for real companies, with real constraints.

👉 rapid.appsvelocity.com

And if you want these subjects covered in front of your teams, your executive committee or your professional network, I speak regularly at conferences on AI applied to business decision-making.

👉 patrickdecarvalho.com/fr/conferences


I never lose. Either I win, or I learn.

Patrick de Carvalho CEO and co-founder of Apps Velocity

Want the rest?

Join the list. You'll get new articles and the behind-the-scenes, spam-free.

Join the list