The Peptide CommonsEst. May 2024
Independent. We sell nothing and are affiliated with no manufacturer or pharmacy. Every moderation action is logged in public
Research Methods · Statistics

What a confidence interval means, from scratch — the long version

OV
o.vogelTL2 Moderator19 Sep 2025#1

The question in the title: What a confidence interval means, from scratch — the long version I will give what I have already checked below so nobody repeats it.

I have seen SURMOUNT-2 (Lancet, 2023) cited in support of a claim I do not think it supports, twice this month, so I would like to work through what it actually shows.

My reading is that the trial is sound for its own question and is being stretched to answer a different one. I might be wrong about that, which is why this is a topic rather than a correction.

What I would like from this discussion: someone who disagrees with me to say why, with the section of the paper they are relying on.

3 likes 10mo
TS
t.steenkampTL2Member25 Sep 2025 · edited#2

This follows the opening post rather than contradicting it.

Confidence intervals: rather than a single point estimate, a range of plausible values. A narrow interval means precise measurement; a wide interval means measurement is imprecise. Wider intervals (more uncertainty) are honest about limitation.

0 likes 10mo
AW
ai.wikstromTL2 Moderator30 Sep 2025#3

Effect sizes: the magnitude of a difference, not just whether it is statistically significant. A difference that is significant (p<0.05) might be too small to matter. A large effect might not be significant if sample size is small.

25 likes 10mo
CE
crossover_entryTL3Regular4 Oct 2025#4

Power and sample size: a study might be too small to detect a real effect (low power). Sample size calculations help determine how many participants are needed to detect an effect of a given magnitude.

12 likes 10mo
BW
br.wikstromTL2 Moderator8 Oct 2025#5
t.steenkamp, post #2: This follows the opening post rather than contradicting it. Confidence intervals: rather than a single point estimate, a range of plausible values. A narrow interval means precise measurement; a wide interval means measurement is imprecise. Wider intervals (more uncertainty) are honest about limitation. Go to post

Coming back to post #3, because the follow-up matters more than the original answer.

Two things before anyone answers the substance.

First, the context in the first post is clear and specific. Second, the question is framed so that an answer can actually address it. Both are the norm here and both matter more than they sound.

4 likes in reply to #2 10mo
GR
gradient_reviewTL2Member12 Oct 2025#6
o.vogel, post #1: The question in the title: What a confidence interval means, from scratch — the long version I will give what I have already checked below so nobody repeats it. I have seen SURMOUNT-2 ( Lancet , 2023) cited in support of a claim I do not think it supports, twice this month, so I would like to work through what it actually shows. My… Go to post

Number needed to treat: how many people need to be treated to prevent one bad outcome or achieve one good outcome. More intuitive than relative risk reduction.

0 likes in reply to #1 10mo
MM
m.marchettiTL2 Moderator15 Oct 2025#7

Absence of evidence and evidence of absence: if a study is small and finds no effect, that is absence of evidence, not evidence of absence. A larger study might find an effect that a small study missed.

0 likes 9mo
MD
methods_draftTL2Member18 Oct 2025#8

P-values and significance: p<0.05 means the data would be surprising if the null hypothesis were true, not that the null hypothesis is false. A non-significant p-value does not mean "no effect".

17 likes 9mo
NR
n.ramosTL2 Moderator22 Oct 2025#9

Relative risk and odds ratios: both compare the rate in one group to the rate in another. Relative risk is easier to understand. Odds ratios are standard in many analyses but can be misinterpreted.

0 likes 9mo
S
SHermansenTL2Member25 Oct 2025#10

Thank you for the correction. I have edited my earlier post with a note rather than silently, so the thread still makes sense to read. The error was mine and it was the kind that comes from remembering a figure instead of looking it up.

26 likes 9mo
TI
trough_indexTL3Regular28 Oct 2025#11

This follows post #8 rather than contradicting it.

Confounding: a third variable explains an apparent association. In randomised data, randomisation balances confounders. In observational data, confounders can be adjusted for but unknown ones cannot.

0 likes 9mo
EM
e.mwangiTL2 Moderator31 Oct 2025 · edited#12
crossover_entry, post #4: Power and sample size: a study might be too small to detect a real effect (low power). Sample size calculations help determine how many participants are needed to detect an effect of a given magnitude. Go to post

I read post #10 twice before replying, because I had assumed the opposite.

Regression to the mean: if you select people with extreme values (very high or very low), their next measurement is often less extreme just by chance. This can look like a treatment effect when it is just statistics.

1 like in reply to #4 9mo
NB
n.bridgewaterTL2Member3 Nov 2025#13

For anyone arriving from a search: the marked solution above is the direct answer, and the replies underneath it add the caveats that make it safe to use.

6 likes 9mo
AN
a.norgaardTL26 Nov 2025#14
TN
t.ndiayeTL2 Moderator9 Nov 2025#15

Number needed to treat: how many people need to be treated to prevent one bad outcome or achieve one good outcome. More intuitive than relative risk reduction.

5 likes 9mo
AF
a.friskTL2 Moderator11 Nov 2025#16
crossover_entry, post #4: Power and sample size: a study might be too small to detect a real effect (low power). Sample size calculations help determine how many participants are needed to detect an effect of a given magnitude. Go to post

Coming back to post #14, because the follow-up matters more than the original answer.

Two things before anyone answers the substance.

First, the context in the first post is clear and specific. Second, the question is framed so that an answer can actually address it. Both are the norm here and both matter more than they sound.

3 likes in reply to #4 9mo
K
KTurkingtonTL3Regular14 Nov 2025#17

post #16 answers the question as asked. The question underneath it is different.

Confounding: a third variable explains an apparent association. In randomised data, randomisation balances confounders. In observational data, confounders can be adjusted for but unknown ones cannot.

10 likes 8mo
TD
t.demirTL2 Moderator17 Nov 2025#18

Regression to the mean: if you select people with extreme values (very high or very low), their next measurement is often less extreme just by chance. This can look like a treatment effect when it is just statistics.

22 likes 8mo
IT
integrator_traceTL2Member20 Nov 2025#19
gradient_review, post #6: Number needed to treat: how many people need to be treated to prevent one bad outcome or achieve one good outcome. More intuitive than relative risk reduction. Go to post

For anyone arriving from a search: the marked solution above is the direct answer, and the replies underneath it add the caveats that make it safe to use.

23 likes in reply to #6 8mo
AK
ak.kravchenkoTL2 Moderator22 Nov 2025#20
trough_index, post #11: This follows post #8 rather than contradicting it. Confounding: a third variable explains an apparent association. In randomised data, randomisation balances confounders. In observational data, confounders can be adjusted for but unknown ones cannot. Go to post

Effect sizes: the magnitude of a difference, not just whether it is statistically significant. A difference that is significant (p<0.05) might be too small to matter. A large effect might not be significant if sample size is small.

0 likes in reply to #11 8mo
GD
glossary_deskTL3Regular25 Nov 2025#21

Power and sample size: a study might be too small to detect a real effect (low power). Sample size calculations help determine how many participants are needed to detect an effect of a given magnitude.

6 likes 8mo

Suggested topics

TopicParticipantsRepliesViewsActivity
Sample size intuition for a personal experiment — does this still hold?
Sample size intuition for a personal experiment — does this still hold? I have a specific reason for asking rather than idle curiosity, and the context is below. Comparing STEP 1 ( N Engl J Med , 2021) with…
AWMSNVBIFA+33 38 26k 23d
About the Statistics category
Effect sizes, intervals, multiplicity, and the difference between absent and undetected. This post is a community wiki: any member at trust level 3 or above can edit it, and every edit is recorded with its…
SLJMBVCAV+4 8 2.3k 7mo
What a confidence interval means, from scratch
What a confidence interval means, from scratch — that is the question, and I have not found it answered plainly anywhere I have looked. I have seen STEP 8 ( JAMA , 2022) cited in support of a claim I do not…
MHSNO 2 2.6k 2mo
Measurement error in home scales, with a worked standard deviation — the long version
On the subject in the title: Measurement error in home scales, with a worked standard deviation — the long version Working notes rather than a conclusion. Session topic: SURMOUNT-1 ( N Engl J Med , 2022).…
SKRRAWDYJD+96 105 2k 2mo
[2026 update] Correlation in a self-tracked dataset: what it can support
Posting this under the heading it deserves: Correlation in a self-tracked dataset: what it can support Everything below is what sits behind that. I have seen SURMOUNT-4 ( JAMA , 2024) cited in support of a…
GTARLS 2 31k 6h

Related topics — sharing the tags effect size, worked example, confounding

TopicParticipantsRepliesViewsActivity
[2026 update] Journal club: SURPASS-2 and the semaglutide 1 mg comparator
Journal club: SURPASS-2 and the semaglutide 1 mg comparator Writing it up because I had to work it out twice and would rather nobody else did. Session topic: SURPASS-4 ( Lancet , 2021). Please read it before…
EDHCITINR+126 138 31k 17mo
Coming back to: Feature request: should a dose-equivalence tool exist at all?
Feature request: should a dose-equivalence tool exist at all? I have a specific reason for asking rather than idle curiosity, and the context is below. Posting the method first, because I know what the first…
MSERSIKRMN+37 41 725 7mo
Critiquing an observational claim about a class effect
Posting this under the heading it deserves: Critiquing an observational claim about a class effect Everything below is what sits behind that. Session topic: SCALE ( N Engl J Med , 2015). Please read it before…
VRRIJV 2 6.9k 14mo
Journal club: the retatrutide phase 2 obesity paper
Journal club: the retatrutide phase 2 obesity paper — setting out what I have, and where I think it stops being reliable. I have seen FLOW ( N Engl J Med , 2024) cited in support of a claim I do not think it…
LFVDAERASV+7 11 33k 7mo
Random versus fixed effects: choosing rather than defaulting
Random versus fixed effects: choosing rather than defaulting — setting out what I have, and where I think it stops being reliable. Session topic: FLOW ( N Engl J Med , 2024). Please read it before posting;…
NPFSMLAJS+16 20 22k 13mo