One answer is one observation
AI-answer systems can vary their retrieval, source selection, composition and presentation even when the underlying buyer need has not changed. A source may appear in one response and disappear in the next. The recommended firm or visible citation set may change too.
That makes a one-off answer useful as an example, but too weak to establish a position. The useful result is the distribution: how often an outcome appears, how the firm is classified and how precisely that rate has been estimated.
Repeat like with like
Every repeated result must stay attached to the same question family on the same engine. This is the study cell: one defined buyer need observed on one answer system.
Results from different question families or engines cannot be pooled merely to make the sample look larger. They describe different markets and different systems. Combining them would hide the very differences the measurement is meant to reveal.
Independent research has found substantial run-to-run variation in source selection and AI-visibility results. Repetition reduces the chance that an ordinary fluctuation is mistaken for a stable lead, loss or treatment effect. Preserving the responses and source evidence keeps the result auditable.
Uncertainty is part of the result
A confidence interval, or an equivalent uncertainty statement, shows how precisely a rate has been estimated. No repeat count removes uncertainty, and no universal number is right for every decision. The required depth depends on the observed rate, the difference being tested, the decision it will support and the cost of being wrong.
Delphic's complete baseline standard makes that precision-cost choice explicitly. The purpose is not to make a variable system look deterministic. It is to identify which differences are stable enough to act on.
Better evidence changes the decision
A favourable screenshot can be sold as dominance. An adverse one can provoke unnecessary intervention. Repeated evidence does something more useful: it distinguishes a persistent position from run noise and creates an honest baseline for later comparison.