A person can take three biological-age tests in one season and receive three different ages, and none of the three has made an error. That is the confusing part, and it is the part worth getting right, because the natural reading, that at least two of them are wrong, is not what is happening. What is happening is that three different instruments printed three different numbers and each called it the same word.

This piece builds on an earlier one, What counts as a biomarker, which set out what a record can verify about a measured value and what it declines to assert. The biological-age score sits one level above the biomarker. It is not a measurement. It is a model run on measurements, and the model is where the disagreement lives.

Three services, three numbers. Which one is your biological age?

A consumer biological-age test takes a sample, usually blood or saliva, and returns a single number meant to represent how old your body is functioning, as distinct from how many birthdays you have had. The number feels like a reading off a scale. It is not. It is the output of a specific algorithm trained on a specific dataset, calibrated to a specific reference population, and a different service makes different choices at every one of those steps.

Biological-age scores are not comparable across providers because each uses its own algorithm, inputs and reference population. Run the same body through three of them and the outputs can spread by years, not because the body changed between draws but because you asked three different questions that happen to share a name. The spread is not noise to be averaged away. It is the predictable result of comparing numbers that were never on the same scale.

One published self-test shows what that looks like in practice, and it is worth being exact about what it is. Cheryl McColgan, who writes the health blog Heal Nourish Grow, ordered three consumer services for herself and published the numbers: Superpower returned 45.2 in November 2025, Function Health returned 37.3 in February 2026, and Hundred Health returned 38.0 in March 2026.1 Function and Hundred landed within 0.7 years of each other. Superpower sat 7.9 years above Function, on the same body.

That is an illustration, and it is not what carries this piece. It is one person, self-reported, by an author who presents herself as a blogger rather than a clinician or a researcher; the first draw was four months before the other two, and she notes she had donated blood in between, which she says could account for some of the variance.1 A sample of one with a known confound cannot establish how far apart these services run in general, and nothing here asks it to. The load-bearing claim is structural and comes from the clock literature below: three models scored against three reference populations were never on one scale, so a spread between them is the expected output of the arrangement rather than a discovery about it. The self-test is useful only for showing what the arrangement produces when somebody buys all three, and useful in that narrow way whether the true spread is eight years or two.

The same period produced a clinician’s objection to the premise sitting underneath these panels, which is a different kind of source again. Anthony Pearson, a practising cardiologist who writes as The Skeptical Cardiologist, argued in May 2026 that the more-biomarkers-is-better assumption is unsupported, that biological-age calculators are rough directional signals at best, and that they can be gamed or misinterpreted in ways with no validated connection to living longer.2 That is an opinion piece by a clinician, not a study, and it is worth reading as exactly that: an informed argument about whether the number deserves the weight buyers put on it, which is a separate question from whether two services agree.

What the record shows:

  • Biological-age scores are not comparable across providers, because each uses its own algorithm, inputs and reference population. This is the claim the piece rests on, and it comes from the clock literature, not from any one person’s results.
  • One reviewer’s published self-test recorded 45.2, 37.3 and 38.0 from three consumer services over four months, a spread of 7.9 years between the highest and lowest, per Heal Nourish Grow, updated July 23, 2026. Recorded here as an illustration of the spread and not as evidence of its size: a sample of one, self-reported, by a non-clinician, with a four-month gap and a blood donation between draws.
  • A practising cardiologist writing as The Skeptical Cardiologist argues the premise underneath these panels is unsupported and that biological-age calculators are rough directional signals at best, per Anthony Pearson, May 6, 2026. An opinion piece by a clinician, not a study.

The most common failure: the number without the reference population. A biological age quoted as a fact about a body, with no mention of which model produced it or against whom it was scored. Stripped of its reference population the number looks portable, and it is exactly then that it stops meaning anything you can compare.

Fourteen clocks, one genome. Do they agree?

The disagreement is not a consumer-market quirk. It is visible in the science the consumer services are built on, and this is the part of the piece that is peer-reviewed rather than anecdotal. A study published in Nature Communications on December 16, 2025, by researchers at the University of Edinburgh, compared 14 epigenetic clocks against 174 incident disease outcomes in 18,859 adults from the Generation Scotland cohort, the largest unbiased comparison of its kind.3 It did not find one clock that quietly agrees with the others. It found that they diverge, and that the divergence is structured.

Two results matter for anyone holding a printed age. First, the clocks are not interchangeable in magnitude: the paper reported that first-generation clocks produced associations roughly 50 percent smaller than second-generation ones, so the same underlying risk shows up as a different-sized signal depending on which clock ran. Second, later-generation clocks predicted incident disease substantially better than first-generation clocks, and no single clock was best across all 174 outcomes. A number that is not comparable across methods, and whose usefulness depends entirely on which method produced it, is not a portable fact. It is a reading that only means something next to its own method.

The journals have registered the gap between that result and what the consumer market does with it. eBioMedicine, the Lancet group’s open-access title, devoted a February 2026 editorial to the point that epigenetic clocks still have to travel some distance before they are fit for clinical use, and that their reliability in commercial wellness settings is an open question rather than a settled one.4 An unsigned journal editorial is not a study either, and it is cited here for what it is: the field’s own statement about how far the science has and has not come.

What the record shows:

  • A 2025 Nature Communications comparison of 14 epigenetic clocks against 174 incident disease outcomes in 18,859 adults found the clocks diverge, with first-generation clocks showing associations about 50 percent smaller in magnitude than second-generation ones, per the University of Edinburgh study published December 16, 2025. Peer-reviewed, and the source the structural claim in this piece rests on.
  • No single clock performed best across all 174 disease outcomes, so there is no method-independent “true” biological age to compare against.
  • eBioMedicine, February 2026, states editorially that epigenetic clocks are not yet established for clinical use and that their reliability in commercial settings remains open. A journal editorial, not a primary study.

The most common failure: the clock mistaken for a thermometer. Treating a biological-age score as if any instrument would read the same value off the same body, the way two thermometers agree on a fever. The clocks are models with different training and different aims, and expecting them to agree is expecting three different questions to have one answer.

A number you cannot compare. What is it good for?

None of this makes the score useless. It makes it useful in one direction and misleading in another. The direction that holds is longitudinal and within a single service: take the same test, from the same provider, on the same method, at intervals, and the trend line is interpretable because the algorithm and the reference population are held constant. What moves is then a signal about you, not about the choice of clock.

The direction that fails is cross-sectional and across services: one provider’s 45 set beside another’s 38, treated as a debate about your true age. That comparison has no valid answer, and the more confident the two marketing pages sound, the more misleading the juxtaposition is. This is the same discipline the biomarker piece applied one level down, where a panel’s advertised count and a headline price turned out to be claims that needed a source and a date rather than facts that could be compared at face value. The score inherits that problem and adds one of its own, which is that even the honest number is only legible against the method that made it.

What the record shows:

  • The Atlas of the Healthspan Economy classifies epigenetic age testing as emerging, and records the organizations offering it without ranking their scores against one another.
  • The Atlas records 6 organizations offering epigenetic age testing as of August 2026.

The most common failure: the cross-service comparison. Setting one provider’s age beside another’s and asking which is right. The question assumes a shared scale that does not exist, and no amount of precision in either number supplies it.

What a record does instead

A record does not tell you your biological age, and once it is clear that the clocks behind the services diverge by design it should be clear why no honest one would try. What a record can do is hold the modality steady while the market churns: name it, classify it as emerging, list who offers it and decline to rank the scores as though they were on one scale. That is the neutral position, and it happens to match the science, which is that the number is a trend tool inside one method and not a verdict you can carry between them.

So the usable sentence is narrow and it is the one to keep. A biological-age score is legible against its own method over time, and illegible the moment it is lifted out and compared to another. Trust the trend, trust the biomarkers underneath it, and treat the single portable age for what it is, a number that lost its reference population somewhere between the lab and the landing page.


The Atlas of the Healthspan Economy is a neutral record of organizations working on healthspan. It classifies epigenetic age testing as emerging and records providers without ranking their scores. It does not recommend. See the Epigenetic age testing entry, read What counts as a biomarker and read the methodology.

A note on what this piece does not claim. It does not name a best or worst service, and it does not draw a general conclusion from one person’s results. The structural point about non-portability is what carries it, and that point rests on the peer-reviewed clock comparison, not on the self-test. The self-test is reproduced as an illustration and is explicitly not treated as evidence of how far apart these services run in general: it is a sample of one, self-reported by a non-clinician, with a four-month gap and a blood donation between draws. The cardiologist’s argument is an opinion piece and the eBioMedicine item is a journal editorial; neither is presented as a study. Nothing here is medical advice, and no service named is accused of an error, because returning a different number from a different model is not one. Study findings are as reported in the primary sources and dates cited.

Footnotes

  1. Cheryl McColgan, Heal Nourish Grow, “Superpower Health Review” and “Function Health vs Superpower,” both updated July 23, 2026 (the Superpower review first posted November 18, 2025). One author’s self-test across three consumer services: Superpower 45.2 in November 2025, Function Health 37.3 in February 2026, Hundred Health 38.0 in March 2026, against a stated calendar age of 52.8 at the Function draw. Function and Hundred fall within 0.7 years of each other; Superpower is 7.9 years above Function. The author is a health blogger and describes her background as a psychology degree with graduate clinical psychology training, a NASM personal training certification and an E-RYT yoga certification; she does not present herself as a physician or a researcher. She states that the Superpower draw was roughly four months before the other two and that she donated blood between tests, and that both facts could account for some of the variance. She also states that Function’s calculation uses eight markers (white blood cell count, ALP, glucose, creatinine, MCV, RDW, albumin, lymphocytes) while Superpower weights a different set, and advises against over-anchoring to the absolute number from either platform. Cited here as an illustration of the spread, not as evidence of its magnitude: a self-reported sample of one with a known confound. Accessed August 8, 2026. https://healnourishgrow.com/superpower-health-review/ and https://healnourishgrow.com/function-health-vs-superpower/ 2

  2. Anthony Pearson, “On the Massive BS That is Generated by Wellness Apps like Superpower and Function,” The Skeptical Cardiologist, May 6, 2026. Pearson is a practising cardiologist writing on his own Substack. He argues that testing 100-plus markers rather than the standard 10 to 15 does not translate into meaningful insight, that biological-age calculators such as PhenoAge are rough directional signals at best and “can be gamed, misinterpreted, or acted upon in ways that have no validated connection to actually living longer,” that roughly a third of Superpower’s 105 biomarkers are derived ratios of questionable clinical value, and that in his own testing every flagged abnormality proved clinically insignificant. This is a clinician’s opinion piece, not a peer-reviewed study, and it is cited as an argument about the premise rather than as evidence about any provider’s accuracy. Accessed August 8, 2026. https://theskepticalcardiologist.substack.com/p/on-the-massive-bs-that-is-generated

  3. Mavrommatis et al., “An unbiased comparison of 14 epigenetic clocks in relation to 174 incident disease outcomes,” Nature Communications, published December 16, 2025. University of Edinburgh, using 18,859 adults from the Generation Scotland cohort with 10 years of follow-up. Reports that second- and third-generation clocks significantly outperform first-generation clocks in disease settings, that first-generation clocks produce associations roughly 50 percent smaller in magnitude, and that no single clock was best across all 174 outcomes. Peer-reviewed primary research, and the source the structural claim in this piece rests on. https://www.nature.com/articles/s41467-025-66106-y

  4. eBioMedicine, “Epigenetic clocks: advancing biological age measures towards meaningful clinical use,” volume 124, February 1, 2026. DOI 10.1016/j.ebiom.2026.106175. An unsigned editorial in the journal, part of the Lancet Discovery Science group, on what still separates epigenetic clocks from established clinical use, including open questions about their reliability in commercial wellness settings. An editorial rather than a primary study, cited as the field’s own statement of where the science stands. https://www.thelancet.com/journals/ebiom/article/PIIS2352-3964(26)00056-3/fulltext