Fitorex

How to read scientific evidence

What the evidence levels you see in each lesson really mean, and how to tell a real finding from statistical noise or marketing.

Why every lesson in this app has an 'evidence level'

Not every claim about fitness has the same backing: some ideas are supported by dozens of large, consistent studies, and others are just an expert's reasonable opinion. Assigning a level (A, B, C, D) to each claim is not bureaucracy: it tells you how much you can bet on it and how much room there is for it to change with new research. Treating a level A and a level C with the same confidence is the most common mistake in reading science.

How to apply it

  • When you read a fitness recommendation anywhere, ask what level of evidence sits behind it before following it to the letter
  • Remember that a level C or D does not mean 'false', it means 'we do not know for sure yet'

The evidence pyramid: from expert opinion to meta-analysis

At the base sit expert opinions and isolated case series (useful for generating ideas, not for confirming them); above them come observational studies (which show associations but do not prove cause); higher up, the randomised controlled trial (RCT), the design that does let you say 'this caused that'; and at the top, the meta-analysis, which combines many RCTs into a single, more reliable estimate. But the hierarchy is relative: a badly designed RCT can be worth less than a good observational study.

How to apply it

  • Give more weight to a recommendation backed by several RCTs or a meta-analysis than to a single study or an opinion, however well known the person giving it
  • Do not automatically dismiss a well-conducted observational study just because it is not an RCT: design quality matters as much as design type

The 4 types of study you will come across

The RCT (randomised controlled trial) allocates people to groups at random so you can say 'this caused the change'; the crossover design puts the same person through both conditions, useful for acute effects (like a supplement) but not for permanent adaptations (like gaining muscle); the cohort study follows people over time without manipulating anything, useful for very long-term effects that cannot be randomised; and the meta-analysis statistically combines several studies into a single figure.

How to apply it

  • If you see a supplement study with a 'crossover' design measuring long-term muscle gain, be suspicious: that design does not fit that objective
  • For questions about 'what happens over 20 years of a habit', a cohort study makes more sense than looking for an RCT, which would be prohibitively expensive and unfeasible

'Statistically significant' is not the same as 'important'

A p-value below 0.05 only says the result is unlikely if there were no real effect; it does not say whether the effect is large, useful or reproducible. With a very large sample, even a tiny, irrelevant difference (like gaining 100 grams of muscle in 12 weeks) can come out 'statistically significant'. What really matters is the effect size: how much something changes in practice, not just whether the change is mathematically improbable by chance.

How to apply it

  • When a headline says 'study proves X works', look for how big the actual effect was, not just whether it was 'significant'
  • Be suspicious of studies with thousands of participants that only report the p-value without stating the magnitude of the change

The confidence interval: what almost nobody interprets correctly

When a study says 'the effect was +3 kg (95% CI: +1 to +5 kg)', that range tells you how much real uncertainty there is around the number, something the p-value alone does not show. A very wide interval (from +0.5 to +8 kg) means the study has low precision, even if the central result looks good; a narrow interval gives you more confidence in the specific number.

How to apply it

  • If you can, look for the confidence interval as well as the average result: it tells you how much to trust the exact number
  • A very wide confidence interval is a sign that more research is needed before trusting the result too much

The most common biases in fitness studies

Survivorship bias means we only see results from those who finished the study (the most motivated people, not real-world ones); self-reported diet and exercise is systematically wrong (people underestimate what they eat by 20–40%); publication bias means studies with positive results get published more than negative ones, inflating the apparent average effect; and studies funded by a supplement's industry are 3–4 times more likely to show results favourable to whoever is paying.

How to apply it

  • If a supplement study is funded by the brand that sells it, add an extra point of scepticism
  • Remember that diets and routines from 'studies' with high dropout probably overestimate how easy they are to follow

Correlation is not causation: the example that comes up most in fitness

'People who exercise have better mental health' is a real correlation, but it does not tell you which way the causation runs: does exercise improve mood, or do people in a better mood have more energy to exercise? Only studies where exercise is assigned at random (not where you simply observe who already does it) can separate these two explanations. It is the most frequent error when reading headlines from observational studies about healthy habits.

How to apply it

  • Faced with any headline like 'people who do X live longer', ask whether the study assigned X at random or merely observed who already did it
  • Be suspicious of the 'healthy user bias': people who already train also tend to sleep and eat better, which confounds the real effect of exercise itself

Why simply taking part in a study already improves you (the Hawthorne effect)

People change their behaviour simply because they know they are being observed: in training studies, participants tend to train harder, eat better and sleep more than usual just during the study period. This can inflate the results of any intervention, because part of the improvement comes not from the treatment itself but from the attention and follow-up the participants receive.

How to apply it

  • Bear in mind that results from short, highly supervised studies can be better than what you will achieve training on your own without that follow-up
  • Logging and close tracking (like you do in this app) can also give you that extra nudge: use it to your advantage

How to read a study or a headline with healthy scepticism

Before changing your routine over a new study, ask yourself: which population was it done in (trained or sedentary, your age and sex)? How long did it last (weeks do not predict years)? Who funded it? Is the effect size large or trivial even if it is 'significant'? Are there other studies saying the same thing or is it an isolated finding? The goal is not to distrust all science, but not to change strategy over a single eye-catching headline when the body of evidence says otherwise.

How to apply it

  • Before adopting a viral trend, look for whether there is a meta-analysis or several studies supporting it, not just the one study that went viral
  • Give more weight to the accumulated body of evidence than to the latest news, however convincing the headline sounds