Regression to the mean: if you select people with extreme values (very high or very low), their next measurement is often less extreme just by chance. This can look like a treatment effect when it is just statistics.
Correlation in a self-tracked dataset: what it can support posts 151–167
This is a continuation of a long topic, addressed by post number rather than by page. Start at post 1.
post #151 answers the question as asked. The question underneath it is different.
Confounding: a third variable explains an apparent association. In randomised data, randomisation balances confounders. In observational data, confounders can be adjusted for but unknown ones cannot.
Effect sizes: the magnitude of a difference, not just whether it is statistically significant. A difference that is significant (p<0.05) might be too small to matter. A large effect might not be significant if sample size is small.
Worth separating two things that post #151 runs together.
Two things before anyone answers the substance.
First, the context in the first post is clear and specific. Second, the question is framed so that an answer can actually address it. Both are the norm here and both matter more than they sound.
post #155 is right about the mechanism and I think understates the practical bit.
Confidence intervals: rather than a single point estimate, a range of plausible values. A narrow interval means precise measurement; a wide interval means measurement is imprecise. Wider intervals (more uncertainty) are honest about limitation.
Absence of evidence and evidence of absence: if a study is small and finds no effect, that is absence of evidence, not evidence of absence. A larger study might find an effect that a small study missed.
Collapsed as off-topic by two members at trust level 3 or above
P-values and significance: p<0.05 means the data would be surprising if the null hypothesis were true, not that the null hypothesis is false. A non-significant p-value does not mean "no effect".
On post #155 — agreed on the reasoning, with one qualification.
Multiplicity and multiple comparisons: if you test many hypotheses, the chance of finding a false positive by random chance increases. That is why pre-specifying the primary hypothesis matters.
Power and sample size: a study might be too small to detect a real effect (low power). Sample size calculations help determine how many participants are needed to detect an effect of a given magnitude.
For anyone arriving from a search: the marked solution above is the direct answer, and the replies underneath it add the caveats that make it safe to use.
Effect sizes: the magnitude of a difference, not just whether it is statistically significant. A difference that is significant (p<0.05) might be too small to matter. A large effect might not be significant if sample size is small.
Collapsed as off-topic by two members at trust level 3 or above
Confidence intervals: rather than a single point estimate, a range of plausible values. A narrow interval means precise measurement; a wide interval means measurement is imprecise. Wider intervals (more uncertainty) are honest about limitation.
I read post #162 twice before replying, because I had assumed the opposite.
Multiplicity and multiple comparisons: if you test many hypotheses, the chance of finding a false positive by random chance increases. That is why pre-specifying the primary hypothesis matters.
Number needed to treat: how many people need to be treated to prevent one bad outcome or achieve one good outcome. More intuitive than relative risk reduction.
Two things before anyone answers the substance.
First, the context in the first post is clear and specific. Second, the question is framed so that an answer can actually address it. Both are the norm here and both matter more than they sound.
Picking up post #164: that is the part I would want checked first.
Confounding: a third variable explains an apparent association. In randomised data, randomisation balances confounders. In observational data, confounders can be adjusted for but unknown ones cannot.
Suggested topics
| Topic | Participants | Replies | Views | Activity |
|---|---|---|---|---|
|
About the Statistics category
Effect sizes, intervals, multiplicity, and the difference between absent and undetected. This post is a community wiki: any member at trust level 3 or above can edit it, and every edit is recorded with its…
|
+4 | 8 | 2.3k | 7mo |
|
Measurement error in home scales, with a worked standard deviation
Measurement error in home scales, with a worked standard deviation Writing it up because I had to work it out twice and would rather nobody else did. Comparing SURPASS-4 ( Lancet , 2021) with STEP 1 ( N Engl…
|
+26 | 31 | 4.3k | 4mo |
|
Regression to the mean in progress reports
On the subject in the title: Regression to the mean in progress reports Working notes rather than a conclusion. Session topic: PIONEER 6 ( N Engl J Med , 2019). Please read it before posting; the discussion…
|
+30 | 34 | 49k | 12mo |
|
What a confidence interval means, from scratch
What a confidence interval means, from scratch — that is the question, and I have not found it answered plainly anywhere I have looked. I have seen STEP 8 ( JAMA , 2022) cited in support of a claim I do not…
|
2 | 2.6k | 2mo | |
|
[2026 update] Correlation in a self-tracked dataset: what it can support
Posting this under the heading it deserves: Correlation in a self-tracked dataset: what it can support Everything below is what sits behind that. I have seen SURMOUNT-4 ( JAMA , 2024) cited in support of a…
|
2 | 31k | 6h |
Related topics — sharing the tags effect size, heterogeneity, worked example
| Topic | Participants | Replies | Views | Activity |
|---|---|---|---|---|
|
The most common factual error about semaglutide on the internet — what changed since
The most common factual error about semaglutide on the internet — what changed since — setting out what I have, and where I think it stops being reliable. Comparing SURMOUNT-2 ( Lancet , 2023) with STEP 1 ( N…
|
+42 | 46 | 1.7k | 5mo |
|
Immortal time bias in a claims-database study — one year on
Posting this under the heading it deserves: Immortal time bias in a claims-database study — one year on Everything below is what sits behind that. Comparing SURPASS-2 ( N Engl J Med , 2021) with SURMOUNT-1 (…
|
+41 | 46 | 3.7k | 22h |
|
Retatrutide's phase 2 heart-rate signal and how to think about it
On the subject in the title: Retatrutide's phase 2 heart-rate signal and how to think about it Working notes rather than a conclusion. I have seen PIONEER 6 ( N Engl J Med , 2019) cited in support of a claim…
|
+19 | 23 | 20k | 7mo |
|
Second pass at: Journal club: LEADER as the historical anchor
Posting this under the heading it deserves: Second pass at: Journal club: LEADER as the historical anchor Everything below is what sits behind that. Comparing SCALE ( N Engl J Med , 2015) with SURPASS-2 ( N…
|
+75 | 81 | 2k | 2y |
|
Peak-to-trough ratio at steady state for a weekly agent
On the subject in the title: Peak-to-trough ratio at steady state for a weekly agent Working notes rather than a conclusion. I would like to understand what this number means before I repeat it anywhere. A…
|
+58 | 64 | 9.7k | 2mo |