Report method
How the AI Visibility Report 2026 was built
This page is the method behind the AI Visibility Report 2026: the panel, the queries, the platforms, the rubric as it was applied, and what the study cannot support. It is a record of the 2026 edition and it is not updated as practice changes. A later edition publishes its own method at its own address.
Scope of this page
Fieldwork ran from 30 April to 12 May 2026. The report was published on 19 May 2026. This page describes that edition, frozen as published, including the parts of the method that did not go to plan.
It is not the method of the Atlas. What qualifies as an Atlas record, how a record is verified and what "recorded" means are answered on the methodology and constitution page, which governs the Atlas and not this study. The standing commercial disclosure is stated there, permanently, under the Sandria disclosure. Defects found in this report after publication are dated and kept on the corrections page.
The 25-clinic panel
The panel was selected against four criteria: geographic diversity across the markets where premium longevity care is currently provided, business-model diversity across medical-led, resort-clinical-hybrid, membership and disruptor models, sufficient public content presence to be plausibly visible to AI search and editorial defensibility against the question a clinic CMO would reasonably ask: why these and not those.
The final list balances established category leaders, whose visibility patterns define the competitive ceiling, with newer entrants, whose visibility patterns reveal what works for a clinic without decades of accumulated authority. Its geographic distribution is United States 10, Europe 6, Middle East 4 and Asia-Pacific 4, with a disruptor tier representing newer business models within those geographies. That weighting is a methodological choice and not a definitive map.
The panel was assembled against those four criteria and was not drawn from the Atlas. It is not a sample of the Atlas record, it was not selected through the Atlas pillar taxonomy, and it is not a ranking of anything. The working roster's initial shortlist crossed several fields and narrowed to premium longevity and preventive clinics as a consequence of the criteria above, not as a starting condition. Where a panel clinic also holds an Atlas entity page, the report links to it, and that link is the only relationship between the two.
The queries and the platforms
A fixed panel of 50 questions was run across four AI platforms, Google AI Overviews, Perplexity, ChatGPT and Claude, producing 200 query-platform combinations. The queries were written in five categories of ten: brand-direct, category, service and protocol, comparison, and client-research, the last including deliberately skeptical framings. The structure surfaces both how the platforms handle a named brand and how they construct a recommendation when no brand is named.
The four platforms were chosen as the ones most likely to shape a prospective client's first impression in 2026. Gemini was excluded to avoid platform overlap inside the Google ecosystem. Tiers were not uniform: Claude was run on a paid tier that serves its largest model by default, while the other three were used on free tiers, and Google AI Overviews has no user-selectable tier at all. Findings about platform architecture are independent of tier. Findings comparing answer quality between platforms are not, and are reported with that caveat.
All queries were run by hand through the platforms' public interfaces, in private or incognito sessions, with no automated scraping, no API extraction and no circumvention of platform controls.
The capture rule
Where a query was re-run, the later capture is the published value and the earlier one is retained in the log rather than deleted. Eight query-platform combinations were re-run on 12 May 2026, all of them on ChatGPT in the service and protocol category, so the log holds 208 captures under 200 combination identifiers.
This rule decides one published figure. Tokyo Midtown Clinic appears in the 30 April capture of one re-run query and in neither the 12 May re-run of that query nor anywhere else, so it counts as a zero-observation clinic. Under the opposite rule it would not, and the count of zero-observation clinics would be two rather than three. The rule was fixed before that consequence was known.
The rubric as applied
The rubric was locked on 30 April 2026, before scoring began. It defines nine fields: a visibility score from 0 to 5, a position bonus, a sentiment penalty and a visibility total; and a brand integrity score from 0 to 5, a hallucination modifier, a stale-data modifier, a competitive-disadvantage modifier and an integrity total.
One of those nine was recorded consistently. Across the 347 scored appearances reconstructed from the query run log:
- Visibility was recorded on 347 of 347.
- Brand integrity was recorded as free text on 95, and as a number on 1.
- The position bonus was recorded on none.
- The sentiment penalty was recorded on none.
- The three integrity modifiers were recorded on none.
So the report has one axis carrying numbers and a second that was assessed in prose. That is why the report says each clinic was scored for visibility and assessed for brand integrity, rather than scored on two axes. The two remain independent, because a clinic can be highly visible and misrepresented, or accurately represented and rarely visible, and the remediation for each is different. The independence claim does not depend on both being numeric. The modifiers were defined and discussed during scoring and were never applied to a row.
Visibility was scored from the response itself, on a scale running from absent, through named in passing, described without a link, cited with a link, featured as a primary recommendation, to anchored, meaning the response was built substantially on the clinic's own content. Integrity was judged against each clinic's own published positioning, comparing the response to the clinic's homepage, services page, founder biography and most recent twelve months of press. Integrity cannot be scored by the platforms themselves without asking them to grade their own output.
Scoring was done by one author. Twenty of the two hundred rows were scored with AI assistance and the rest by hand. There was no second human scorer. Working tags kept during collection were checked against the raw captures before analysis, and where the two diverged the raw capture was treated as authoritative.
The calibration was begun and not completed
The protocol specifies an inter-rater calibration: a twenty-query sample run across both scorers, with the first ten from each cross-scored and any divergence greater than one point ruled on before bulk collection.
The calibration sheet holds ten of those twenty queries and carries no record of the cross-scoring. The calibration was therefore begun and not completed, and no inter-rater agreement figure exists or can be reconstructed. Nothing in the report rests on one.
The 11 May retest
A retest panel was run on 11 May 2026, on Google AI Overviews only, to test anomalies that had appeared in the primary panel. It is a separate panel. It is not part of the 200, and no figure describing the 200 includes it.
Its components:
- Fourteen practitioner names. Eleven were put on the bare [practitioner name] clinic template and all eleven returned no output. Three named the same practitioner on other templates and all three returned substantive answers.
- Four London phrasings. All four returned no output.
- Three Singapore phrasings. Two returned no output and the neutrally worded one was answered.
- Three service-category reframings, testing whether a tighter category framing would lift panel clinics into answer sets that had previously held none. Generally it did not.
Finding 1 rests on this panel and not on the primary 50-query panel. The suppression it reports attaches to the phrasing and not to the practitioner, which is what the three answered different-template queries establish. The retest design and its results are recorded in the study's working papers, dated 11 and 13 May 2026. The individual retest queries are not rows in the query run log, and the report's own dataset holds two practitioner-naming queries, one of which was suppressed and one of which was answered.
The 15 May correction pass
Scoring included a correction pass after collection closed. On 15 May 2026, nine corrections were entered against eight rows of the query run log, each annotated as verified against the raw capture rather than against a later summary.
Eight of the nine reassigned an organization identifier that had been recorded wrongly during scoring. The ninth added an organization that had surfaced under its operating brand and had been missed, which removed it from the report's zero-observation finding before publication. The corrected identifiers are the ones the published figures rest on.
This pass was published on the Atlas methodology page until 2026-08-15, under a heading about how Atlas records are verified. It is report method and it belongs here.
Source and custody
The primary record is a data collection workbook holding the clinic roster, the query panel, the calibration sample and the query run log, with each run's response excerpt, cited source URLs and scoring notes. Per-clinic figures exist as prose inside that log rather than as a computed sheet, so any per-clinic total is a reading of the notes and not a formula.
That workbook is not in this site's repository and was not under version control during the study. A parsed extract of the score lines, together with the rows the parse could not resolve and a written reconciliation against the published figures, is held in the repository. The counts in "The rubric as applied" above are computed from that extract. Bringing the primary record under version control is an open item for the 2027 edition.
Captures were not retained uniformly. Saved captures exist for one platform only and carry no run identifier, so no individual published result can be checked back against an archived image. The saved ChatGPT captures record the answers but not the URLs those answers cited.
Relationships with panel clinics
The 25-clinic panel was fixed and the scoring rubric locked on 30 April 2026. Fieldwork ran to 12 May 2026 and the report was published on 19 May 2026.
In August 2026, after publication, Andrea B. Maier shared the report publicly. She co-founded Chi Longevity, which is snapshot 06 in this panel, and she is named in the report. Since publication, people connected with two other panel clinics have also been in touch with the author about the Atlas.
All of these contacts began in August 2026, more than three months after the panel was fixed and after the report was published. None of them could have affected which clinics were selected, how they were scored or what the report says. No panel clinic has paid for anything, and none can.
Anyone holding a relationship with a panel clinic is excluded from reviewing this study or its successors.
This section is updated when a relationship forms, not on a schedule.
What this method cannot support
The report measures what four AI platforms surfaced in response to a fixed set of questions during a two-week window. It does not measure quality of care. A clinic scoring low here may be excellent. It does not measure organic search rankings, conversion, traffic, revenue or any clinic-side operational metric, and it does not model the investment needed to shift a result.
The 25-clinic scope is a sample. Conclusions about the most-cited clinics in it generalize within it and should not be extrapolated to clinics or markets the study did not cover without verification.
Each question was asked once per platform. A single run cannot fully separate a real pattern from ordinary platform variance, so findings that replicate across multiple queries or platforms are distinguished throughout from single-query observations. Repeat runs are a 2027 item.
The English-language weighting means non-English longevity ecosystems are underrepresented. A market with substantial domestic provision in another language may be visible in that language while appearing invisible in English-language AI search. An Arabic-language probe was run as a deliberate test of that limit and not as coverage of it.
Where a panel clinic produced no observations, the finding is that it did not surface in response to the questions clients are likely to ask. It is not a finding that the platforms are unaware the clinic exists.
AI search is a moving target. These patterns describe platform behavior during the study window. Some will hold. Others will not, as the platforms change how they retrieve, rank and cite.