A person can take three biological-age tests in one season and receive three different ages, and none of the three has made an error. That is the confusing part, and it is the part worth getting right, because the natural reading, that at least two of them are wrong, is not what is happening. What is happening is that three different instruments printed three different numbers and each called it the same word.

This piece builds on an earlier one, What counts as a biomarker, which set out what a record can verify about a measured value and what it declines to assert. The biological-age score sits one level above the biomarker. It is not a measurement. It is a model run on measurements, and the model is where the disagreement lives.

Three services, three numbers. Which one is your biological age?

A consumer biological-age test takes a sample, usually blood or saliva, and returns a single number meant to represent how old your body is functioning, as distinct from how many birthdays you have had. The number feels like a reading off a scale. It is not. It is the output of a specific algorithm trained on a specific dataset, calibrated to a specific reference population, and a different service makes different choices at every one of those steps.

Biological-age scores are not comparable across providers because each uses its own algorithm, inputs and reference population. Run the same body through three of them and the outputs can spread by years, not because the body changed between draws but because you asked three different questions that happen to share a name. The spread is not noise to be averaged away. It is the predictable result of comparing numbers that were never on the same scale.

What the record shows:

  • Biological-age scores are not comparable across providers, because each uses its own algorithm, inputs and reference population.
  • [SOURCE PENDING: the named July 2026 hands-on review reporting one person’s three results of 45.2, 37.3 and 38.0 from three consumer services, a spread of nearly eight years. Add the outlet, author and date, then state: “Three consumer biological-age services returned results spanning nearly eight years for the same person within a few months in 2026, per [named review, date].” The exact figures are not printed here until the review is cited.]

The most common failure: the number without the reference population. A biological age quoted as a fact about a body, with no mention of which model produced it or against whom it was scored. Stripped of its reference population the number looks portable, and it is exactly then that it stops meaning anything you can compare.

Fourteen clocks, one genome. Do they agree?

The disagreement is not a consumer-market quirk. It is visible in the science the consumer services are built on. A study published in Nature Communications on December 16, 2025, by researchers at the University of Edinburgh, compared 14 epigenetic clocks against 174 incident disease outcomes, the largest unbiased comparison of its kind. It did not find one clock that quietly agrees with the others. It found that they diverge, and that the divergence is structured.

Two results matter for anyone holding a printed age. First, the clocks are not interchangeable in magnitude: the paper reported that first-generation clocks produced associations roughly 50 percent smaller than second-generation ones, so the same underlying risk shows up as a different-sized signal depending on which clock ran. Second, later-generation clocks predicted incident disease substantially better than first-generation clocks, and no single clock was best across all 174 outcomes. A number that is not comparable across methods, and whose usefulness depends entirely on which method produced it, is not a portable fact. It is a reading that only means something next to its own method.

What the record shows:

  • A 2025 Nature Communications comparison of 14 epigenetic clocks against 174 incident disease outcomes found the clocks diverge, with first-generation clocks showing associations about 50 percent smaller in magnitude than second-generation ones, per the University of Edinburgh study published December 16, 2025.
  • No single clock performed best across all 174 disease outcomes, so there is no method-independent “true” biological age to compare against.

The most common failure: the clock mistaken for a thermometer. Treating a biological-age score as if any instrument would read the same value off the same body, the way two thermometers agree on a fever. The clocks are models with different training and different aims, and expecting them to agree is expecting three different questions to have one answer.

A number you cannot compare. What is it good for?

None of this makes the score useless. It makes it useful in one direction and misleading in another. The direction that holds is longitudinal and within a single service: take the same test, from the same provider, on the same method, at intervals, and the trend line is interpretable because the algorithm and the reference population are held constant. What moves is then a signal about you, not about the choice of clock.

The direction that fails is cross-sectional and across services: one provider’s 45 set beside another’s 38, treated as a debate about your true age. That comparison has no valid answer, and the more confident the two marketing pages sound, the more misleading the juxtaposition is. This is the same discipline the biomarker piece applied one level down, where a panel’s advertised count and a headline price turned out to be claims that needed a source and a date rather than facts that could be compared at face value. The score inherits that problem and adds one of its own, which is that even the honest number is only legible against the method that made it.

What the record shows:

  • The Atlas of the Healthspan Economy classifies epigenetic age testing as emerging, and records the organizations offering it without ranking their scores against one another.
  • The Atlas records 4 organizations offering epigenetic age testing as of July 2026.

The most common failure: the cross-service comparison. Setting one provider’s age beside another’s and asking which is right. The question assumes a shared scale that does not exist, and no amount of precision in either number supplies it.

What a record does instead

A record does not tell you your biological age, and after three services disagree by the better part of a decade it should be clear why no honest one would try. What a record can do is hold the modality steady while the market churns: name it, classify it as emerging, list who offers it and decline to rank the scores as though they were on one scale. That is the neutral position, and it happens to match the science, which is that the number is a trend tool inside one method and not a verdict you can carry between them.

So the usable sentence is narrow and it is the one to keep. A biological-age score is legible against its own method over time, and illegible the moment it is lifted out and compared to another. Trust the trend, trust the biomarkers underneath it, and treat the single portable age for what it is, a number that lost its reference population somewhere between the lab and the landing page.


The Atlas of the Healthspan Economy is a neutral record of organizations working on healthspan. It classifies epigenetic age testing as emerging and records providers without ranking their scores. It does not recommend. See the Epigenetic age testing entry, read What counts as a biomarker and read the methodology.

A note on what this piece does not claim. It does not name a best or worst service, and it does not draw a general conclusion from one person’s results; the structural point about non-portability, not the anecdote, is what carries it. The specific figures from the hands-on review are withheld pending a citation to the named review. Study findings are as reported in the primary sources and dates cited.