Methodology

A measurement contract you can trust

Instill GEO is built on a frozen, versioned methodology. The same rules govern every audit, so results are comparable over time and defensible in front of a sceptical client.

No unsupported numbers

Every asserted figure traces to a verbatim span of AI output. A finding a classifier can't locate in the raw answer is rejected, not reported.

Per-field provenance

Each value is labelled by how it was determined: observed in the text, deterministically derived, or model-inferred — and inferred values carry the model version that produced them.

Reliability, disclosed

Metrics below the minimum sample threshold, or with low agreement between repeats, are visibly flagged. We never present a thin number as if it were solid.

Correlational, never causal

We report what the AI said, not why. The methodology forbids causal claims — a cited source is an opportunity, never a proven cause of a ranking.

1. Technical Visibility Foundation

Every audit begins with a deterministic technical preflight before buyer-question research or AI visibility measurement. Its purpose is to establish whether the sampled website presents a material access, indexability, discovery, or interpretation constraint that should be understood before the AI results are read. The assessment uses fixed rules and observed website responses; it does not use a language model to decide the outcome.

The Foundation checks the submitted URL and a bounded selection of important internally linked pages, together with relevant robots and sitemap resources. The evidence retained includes:

  • Access and index eligibility — HTTP responses and redirects, robots rules for Googlebot and OAI-SearchBot, page-level robots directives, and X-Robots-Tag headers.
  • Discovery and page identity — sitemap resources, canonical annotations, and whether selected important areas can be reached through crawlable internal links.
  • Business and market clarity — titles, primary headings, descriptions, structured business data, language declarations, and geographic signals.
  • Machine-readable identity consistency (schema check) — source-HTML JSON-LD claims about the website and organisation are compared with observable evidence on the same sampled page. Missing evidence is not treated as a contradiction.
  • Content availability — whether meaningful content is present in the returned source HTML that retrieval systems can inspect.

This is deliberately not a 100-point SEO audit. It is a bounded readiness check, with the observed URL, rule, header, annotation, or resource retained behind each finding. Harmless imperfections can remain supporting evidence without changing the overall outcome.

GOOD

No material technical-readiness concern observed within the sampled scope.

ATTENTION ADVISED

A non-blocking issue was identified that may deserve review.

TECHNICAL BARRIER DETECTED

A material condition was observed that should affect how subsequent AI visibility findings are interpreted.

Interpretation rule: Technical findings may constrain discovery or interpretation, but Instill GEO does not treat them as proof of what caused an AI visibility result.

2. AI Buyer Shortlist Snapshot

The shortlist snapshot is a separate, versioned measurement instrument. It asks a narrower executive question than the broad buyer-journey panel: when an AI system is explicitly required to choose five businesses for this buying need, does the assessed business make the list, where does it place, and who is chosen instead?

  • Two frozen, brand-blind questions — one core-category shortlist and one externally grounded buyer-fit scenario.
  • Eighteen planned observations — two questions × three engines × three repeats.
  • Fail-closed validity — a primary answer counts only when it contains exactly five identifiable, unique businesses in explicit positions 1–5.
  • Separate metrics — shortlist inclusion, top-three, first-choice and median position never enter the main panel's mention or recommendation rates.
  • Evidence-linked context — exact stated selection reasons are retained; an omission follow-up runs only when the target is absent from a valid shortlist.

Interpretation rule: Omission follow-ups are model-stated diagnostic context, not observed causes. “No clear supported reason” is a valid outcome, and recurring themes are only aggregated when enough valid explanations exist.

3. What we measure, and where

We measure your brand's presence inside answers from three audited AI engines: ChatGPT / OpenAI Search, Perplexity, and Google Gemini with Search grounding. Gemini answers are model responses grounded by search; we never mislabel them as anything they are not. Each engine is measured separately — results are only combined where it is honest to do so.

Instill GEO measures defined, search-enabled AI model surfaces under recorded conditions. Results are measurements of those surfaces — they are not intended to reproduce every personalised consumer-app session a given user might see in their own account.

4. The measurement unit

The atomic unit is one buyer question, asked of one engine, with one requested model, in one market and language. The requested model is part of that identity, so results from different model releases are never silently conflated. Each unit is sampled several times to measure stability.

5. Validity and denominators

A sample only counts when the provider returns a complete, analysable answer. Every frequency uses valid samples as its denominator — never the number attempted. Failed, truncated, or refused answers are recorded and disclosed in the report's limitations, and are never billed to you as findings.

6. Evidence and provenance

Analysis is deterministic first. Brand mentions and competitor presence are located directly in the text. A language model is consulted only for judgement calls — recommendation and sentiment — and only when the brand is actually present. Any positive result must link to a verbatim quote; if the supporting quote can't be found in the raw answer, the result is rejected. Every field carries its own provenance:

  • Observed — present verbatim in a captured answer, with a pointer to it.
  • Derived — deterministically computed from observed data by a documented rule.
  • Inferred — a model judgement, carrying the model version and still backed by an observed span.
  • Unavailable — could not be determined; recorded honestly, never guessed.

7. Aggregation and reliability

Metrics are computed only after sample-level validation. Recommendation is always a strict subset of mention. Where an answer ranks options in an explicit list, we report the median position with its best-to-worst range, split per engine, so an unstable rank is visible rather than averaged away. Reliability markers are heuristic aids to reading the data — they are not presented as calibrated probabilities.

8. The causation ban

The system may never assert that a source, action, or property caused a ranking or recommendation. Recommendations are phrased as coverage gaps and observed patterns — for example, “this domain was cited in eight of the responses that recommended a competitor” — never as a causal guarantee.

9. Repeatable under recorded conditions

Generative answers vary between runs — that variability is precisely why this methodology exists. We don't claim a single answer is fixed; we claim the measurement process is repeatable under recorded conditions: the same panel, engines, models, market, and analysis method, sampled and disclosed. That is what makes two Instill GEO audits genuinely comparable rather than an apples-to-oranges snapshot.

See it in practice

The sample report shows exactly how these principles render — evidence appendix, provenance labels, per-engine position, and a plainly-stated limitations section.

  • Every quote tagged observed or inferred
  • Median list position, split by engine
  • Failures disclosed, not hidden

Rigour your clients can forward with confidence

Request an audit and see the methodology applied to your own brand.