Reynard has not published an accuracy number, because we have not yet run a validation we would have been willing to lose. Until we have, a figure here would be decoration. This page says what we will run instead, and how you can check it.
There is no percentage on this page. Not a rounded one, not a range, not a figure borrowed from a paper about something adjacent. Nothing has been measured that we could stand behind, so nothing is claimed.
That is an uncomfortable thing to write on a product page, and it is the whole reason the page exists. If you are evaluating this category, the sentence you should be looking for on any site is the one that describes a test that could have gone badly.
This is not about anyone in particular. It is about what makes a claim mean something, which is the same three things everywhere.
Written down here first, so the design cannot be adjusted after the fact.
Candidates: a referendum or ballot result with published outcomes, a pricing change a company made publicly and the effect it reported afterwards, or a product launch with reported adoption. The requirement is that the outcome is verifiable by anyone, and not something the simulation can simply recall as a fact.
The full prediction — the acceptance score, the segment breakdown, the direction and the size — published with a timestamp anyone can verify, before the real answer is looked up.
One run, with the panel and the prompts exactly as a customer gets them. No re-runs, no prompt tuning after seeing how it went, no quietly picking the better of two attempts.
All three, on this page, including when the gap is embarrassing. A validation you only publish when it flatters you is not a validation.
One question is an anecdote either way. The thing worth publishing is a pattern across several, including the ones that went badly.
A close match on one question proves very little. Any single prediction can land by luck, and the more freedom there is in picking the question, the more luck is available.
A pattern across several questions, predicted in advance and published whatever the outcome, is worth something: it is evidence that the simulation is tracking something real in the direction and shape of a population's reaction.
Nothing here converts a simulation into fieldwork. Even a strong track record would not give a run statistical significance, would not make it a defensible number for a board pack or a regulator, and would not mean the next question behaves like the last one. It would mean the tool is useful for deciding where to look — which is what it is for.
The experiment is designed and has not yet been run. There are no results to report, and this page carries no numbers.
When the first run has happened, the prediction, the real answer and the gap will appear on this page — in that order, and regardless of how it went.
If you want to propose a question, or check the method before we run it, get in touch. A question suggested by someone who does not work here is a better test than one we picked ourselves.
The panels, the weighting and what a run actually produces.
Read the comparisonWhat we store, where it runs, and how your input is handled.
Read the comparisonReynard against surveys, focus groups and the other simulated-audience tools.
Read the comparisonIf you can think of a question with a published, verifiable answer that a simulation should be able to get right, send it in. Proposals from outside are the ones worth running.
Get in touch