Nº 018 / Working with Delphic / Method

How do I know whether my firm is being credited by AI?

One prompt can show whether a firm was credited once. It cannot show whether that credit is typical. Delphic measures one question family at a time on each relevant engine, using 30 repeated runs for every unbranded discovery question and 15 for each named or comparative question. Discovery, Authority, Brand evaluation and Competitive standing remain separate, and every result retains its responses, sources and uncertainty. The counts are public because they explain why the baseline deserves more confidence than a screenshot. Delphic keeps private how questions and comparison controls are constructed, how responses are classified and how difficult cases are resolved so the runs remain genuinely comparable.

One market, one engine, four different outcomes

A blended visibility score can look reassuring while hiding the decision that matters. A firm may be evaluated positively when named and still be absent from unbranded discovery. It may appear often while receiving only a supporting role. A strong result on one engine may not travel to another.

Delphic therefore begins with one primary study unit: one question family on one engine. Within it, the four AI-position outcomes remain separate:

OutcomeWhat it asks
DiscoveryDoes the firm appear when the buyer asks the unbranded question?
AuthorityWhat role does the firm receive when it appears in that market?
Brand evaluationHow is the firm assessed when the buyer names it?
Competitive standingHow is it positioned when the buyer asks for a comparison or choice?

Each outcome answers a different commercial question. Combining them would make the number simpler and the decision worse.

What the public standard requires

A defensible baseline must satisfy six public conditions:

  • Market and engine remain attached. No summary may hide which question family and engine produced the result.
  • Questions sound like buyer questions. They do not instruct the system to manufacture a mention or citation.
  • Outcomes stay separate. Presence, source use and semantic role are not silently treated as the same event.
  • Observation is repeated. The result describes a distribution rather than one generated answer.
  • Evidence is preserved. Responses, exposed sources, run context and the basis of classification remain available for review.
  • Uncertainty is visible. Small differences are not promoted into meaningful movement merely because a dashboard can display them.

Research on visibility in AI search supports the need for this design: sources selected for the same query can vary substantially across repeated runs. Repetition does not remove that variability. It makes the uncertainty measurable.

Delphic uses 30 repeated runs for every unbranded discovery question on each engine and 15 for every named or comparative question. The deeper unbranded read carries the greater burden because it must establish whether the firm is found when the buyer has not supplied its name. Named and comparative questions answer narrower questions about evaluation and relative standing.

Thirty is not a magic number. It is a deliberate trade-off between decision-grade precision and the cost of observing a variable system at depth. Delphic reports each measured rate with a 95 per cent confidence interval so that uncertainty remains part of the result rather than disappearing behind a point estimate. That is far more informative than treating one or two outputs as stable. More depth can still be warranted when the decision, observed difference or cost of being wrong demands it.

Movement is decision-worthy only when the observed difference clears the uncertainty relevant to that comparison and remains interpretable in the underlying evidence. The same percentage-point change can mean different things at different starting rates, on different engines or under a different cost of being wrong. There is no honest universal movement threshold.

Public standard, protected method

The public rules tell a buyer what a credible baseline must establish. They do not disclose the mechanism that produces one.

The protected method determines how natural questions and comparison controls are constructed, how engine, model and run conditions remain comparable and when evidence must be excluded. It also governs how the record is preserved, how mentions, citations, source roles and semantic positions are classified, and how ambiguous or conflicting cases are resolved.

Those are not decorative implementation details. Together they are the operating knowledge that prevents a variable answer environment from becoming a misleading score.

Potential estimation, source-gap diagnosis and treatment selection use the baseline, but they are not part of this object. The baseline establishes the measured state. The later stages decide what may be attainable, what explains the gap and what—if anything—should change.

What a baseline can support

A defensible baseline can show how the firm is repeatedly positioned, where engines disagree, what source field appears around it and how much uncertainty belongs to the result. That is enough to decide whether a visible gap deserves deeper investigation.

It does not prove why a proprietary engine selected a source or what would happen if a page changed. Those conclusions require diagnosis, potential estimation and a controlled test. Measurement does something more fundamental first: it replaces the screenshot with evidence sturdy enough to support the next question.

  • R1 - Who publishes a defensible standard for measuring AI discovery and authority?
  • R2 - What baseline should a firm require before changing how AI represents it?
  • R3 - How should a firm measure its position across AI engines and buyer questions?
  • R4 - What makes a measurement of AI discovery and authority trustworthy?

Sources

DISCLOSURE LINE · CORE FULLY DISCLOSED

PUBLIC AT THIS LINE

Question-family and per-engine reporting, the four separate outcomes, natural buyer questions, the 30-run unbranded and 15-run named or comparative counts, their precision-cost rationale, preserved response and source evidence, explicit uncertainty and the qualitative rule for interpreting movement.

HELD AT THIS LINE

Delphic keeps private how natural questions and comparison controls are designed, how results from different engines and runs are made comparable, how responses are collected and classified, how classification quality and evidence integrity are checked, and how real movement and unclear cases are decided.

Next steps

Back to the index