Absence of evidence and evidence of absence: if a study is small and finds no effect, that is absence of evidence, not evidence of absence. A larger study might find an effect that a small study missed.
Measurement error in home scales, with a worked standard deviation — the long version posts 91–106
This is a continuation of a long topic, addressed by post number rather than by page. Start at post 1 · go to the accepted answer.
P-values and significance: p<0.05 means the data would be surprising if the null hypothesis were true, not that the null hypothesis is false. A non-significant p-value does not mean "no effect".
Coming back to post #91, because the follow-up matters more than the original answer.
Relative risk and odds ratios: both compare the rate in one group to the rate in another. Relative risk is easier to understand. Odds ratios are standard in many analyses but can be misinterpreted.
Picking up post #91: that is the part I would want checked first.
Number needed to treat: how many people need to be treated to prevent one bad outcome or achieve one good outcome. More intuitive than relative risk reduction.
Regression to the mean: if you select people with extreme values (very high or very low), their next measurement is often less extreme just by chance. This can look like a treatment effect when it is just statistics.
Confounding: a third variable explains an apparent association. In randomised data, randomisation balances confounders. In observational data, confounders can be adjusted for but unknown ones cannot.
Effect sizes: the magnitude of a difference, not just whether it is statistically significant. A difference that is significant (p<0.05) might be too small to matter. A large effect might not be significant if sample size is small.
This follows post #95 rather than contradicting it.
Thank you for the correction. I have edited my earlier post with a note rather than silently, so the thread still makes sense to read. The error was mine and it was the kind that comes from remembering a figure instead of looking it up.
On post #95 — agreed on the reasoning, with one qualification.
Power and sample size: a study might be too small to detect a real effect (low power). Sample size calculations help determine how many participants are needed to detect an effect of a given magnitude.
post #99 answers the question as asked. The question underneath it is different.
Confidence intervals: rather than a single point estimate, a range of plausible values. A narrow interval means precise measurement; a wide interval means measurement is imprecise. Wider intervals (more uncertainty) are honest about limitation.
Coming back to post #99, because the follow-up matters more than the original answer.
Relative risk and odds ratios: both compare the rate in one group to the rate in another. Relative risk is easier to understand. Odds ratios are standard in many analyses but can be misinterpreted.
Number needed to treat: how many people need to be treated to prevent one bad outcome or achieve one good outcome. More intuitive than relative risk reduction.
I read post #103 twice before replying, because I had assumed the opposite.
Multiplicity and multiple comparisons: if you test many hypotheses, the chance of finding a false positive by random chance increases. That is why pre-specifying the primary hypothesis matters.
This follows post #103 rather than contradicting it.
I disagree with the reply above, and I think the disagreement is substantive rather than terminological.
The distinction being drawn does not survive when you look at the published data for this specific question. I would be glad to be shown wrong on this, because the version I am arguing against is more convenient.
This topic was referenced in
- [2026 update] Correlation in a self-tracked dataset: what it can supportResearch Methods › Statistics · 2 replies
Suggested topics
| Topic | Participants | Replies | Views | Activity |
|---|---|---|---|---|
|
[2026 update] Correlation in a self-tracked dataset: what it can support
Posting this under the heading it deserves: Correlation in a self-tracked dataset: what it can support Everything below is what sits behind that. I have seen SURMOUNT-4 ( JAMA , 2024) cited in support of a…
|
2 | 31k | 6h | |
|
What a confidence interval means, from scratch — the long version
The question in the title: What a confidence interval means, from scratch — the long version I will give what I have already checked below so nobody repeats it. I have seen SURMOUNT-2 ( Lancet , 2023) cited…
|
+16 | 20 | 2.3k | 8mo |
|
About the Statistics category
Effect sizes, intervals, multiplicity, and the difference between absent and undetected. This post is a community wiki: any member at trust level 3 or above can edit it, and every edit is recorded with its…
|
+4 | 8 | 2.3k | 7mo |
|
Multiplicity when you track fifteen variables
Multiplicity when you track fifteen variables — setting out what I have, and where I think it stops being reliable. Comparing LEADER ( N Engl J Med , 2016) with SURMOUNT-OSA ( N Engl J Med , 2024) and finding…
|
2 | 60k | 11mo | |
|
Measurement error in home scales, with a worked standard deviation — a second dataset
Measurement error in home scales, with a worked standard deviation — a second dataset — setting out what I have, and where I think it stops being reliable. Comparing FLOW ( N Engl J Med , 2024) with SURPASS-2…
|
2 | 13k | 7mo |
Related topics — sharing the tags number needed to treat, worked example, heterogeneity
| Topic | Participants | Replies | Views | Activity |
|---|---|---|---|---|
|
Revisiting: Journal club: SURMOUNT-OSA and a hard endpoint in a soft field
Revisiting: Journal club: SURMOUNT-OSA and a hard endpoint in a soft field — setting out what I have, and where I think it stops being reliable. Comparing LEADER ( N Engl J Med , 2016) with SUSTAIN 6 ( N Engl…
|
+108 | 122 | 52k | 7d |
|
Second pass at: Absence of evidence and evidence of absence
On the subject in the title: Second pass at: Absence of evidence and evidence of absence Working notes rather than a conclusion. Session topic: SELECT ( N Engl J Med , 2023). Please read it before posting;…
|
+50 | 55 | 18k | 16h |
|
What to include when your question involves a chromatogram — one year on
Asking directly, because I could not find a straight answer: What to include when your question involves a chromatogram — one year on Question in the title. Context below, and I have tried to include the…
|
+69 | 75 | 862 | 14d |
|
Semaglutide in people without diabetes: what the evidence base looks like
Semaglutide in people without diabetes: what the evidence base looks like Writing it up because I had to work it out twice and would rather nobody else did. Session topic: SUSTAIN 6 ( N Engl J Med , 2016).…
|
+2 | 6 | 18k | 17mo |
|
Coming back to: Prediction intervals and why they are more honest than confidence intervals
Prediction intervals and why they are more honest than confidence intervals Writing it up because I had to work it out twice and would rather nobody else did. Comparing SURMOUNT-OSA ( N Engl J Med , 2024)…
|
2 | 65 | 22mo |