The Peptide CommonsEst. May 2024
Independent. We sell nothing and are affiliated with no manufacturer or pharmacy. Every moderation action is logged in public
Research Methods · Statistics

Measurement error in home scales, with a worked standard deviation

Solved Wiki
Solved by an.zamora in post #7
Effect sizes: the magnitude of a difference, not just whether it is statistically significant. A difference that is significant (p

Jump to the accepted answer →

TK
t.karlsenTL2 Moderator30 Jan 2026#1

Measurement error in home scales, with a worked standard deviation Writing it up because I had to work it out twice and would rather nobody else did.

Comparing SURPASS-4 (Lancet, 2021) with STEP 1 (N Engl J Med, 2021) and finding the comparison harder than it looks.

Different populations, different durations, different endpoints defined slightly differently, and in one case a different estimand. People compare the headline percentages anyway, including me until recently.

Is there a defensible way to put these side by side, or is the honest answer that there is not and we should stop?

11 likes 6mo
S
SHermansenTL2Member2 Feb 2026#2
Community wiki post. Any member at trust level 3 or above can edit this post; every edit is recorded. Last edited by j.delacroix on 21 May 2026.
  • 27 Apr 2026 — customs_ledger: Replaced an unsourced figure with the published one and cited it.
  • 10 Jun 2026 — ppm_error: Removed a claim that the cited source did not support.
  • 21 May 2026 — j.delacroix: Clarified the distinction that was causing repeat questions below.
Editors: customs_ledger, ppm_error, j.delacroix

the opening post answers the question as asked. The question underneath it is different.

For anyone arriving from a search: the marked solution above is the direct answer, and the replies underneath it add the caveats that make it safe to use.

1 like 6mo
TV
t.vargaTL2 Moderator5 Feb 2026#3
SHermansen, post #2: the opening post answers the question as asked. The question underneath it is different. For anyone arriving from a search: the marked solution above is the direct answer, and the replies underneath it add the caveats that make it safe to use. Go to post

Relative risk and odds ratios: both compare the rate in one group to the rate in another. Relative risk is easier to understand. Odds ratios are standard in many analyses but can be misinterpreted.

31 likes in reply to #2 6mo
MD
methods_draftTL2Member8 Feb 2026#4

Number needed to treat: how many people need to be treated to prevent one bad outcome or achieve one good outcome. More intuitive than relative risk reduction.

16 likes 6mo
EB
e.bakkenTL2 Moderator10 Feb 2026#5

Worth separating two things that the opening post runs together.

Regression to the mean: if you select people with extreme values (very high or very low), their next measurement is often less extreme just by chance. This can look like a treatment effect when it is just statistics.

10 likes 6mo
F
FConsidineTL1Member12 Feb 2026#6

Confounding: a third variable explains an apparent association. In randomised data, randomisation balances confounders. In observational data, confounders can be adjusted for but unknown ones cannot.

3 likes 5mo
AZ
an.zamoraTL2 Moderator Solution14 Feb 2026#7
methods_draft, post #4: Number needed to treat: how many people need to be treated to prevent one bad outcome or achieve one good outcome. More intuitive than relative risk reduction. Go to post

Effect sizes: the magnitude of a difference, not just whether it is statistically significant. A difference that is significant (p<0.05) might be too small to matter. A large effect might not be significant if sample size is small.

7 likes in reply to #4 5mo
R
RodriguesTL3Regular16 Feb 2026#8

This follows post #5 rather than contradicting it.

Power and sample size: a study might be too small to detect a real effect (low power). Sample size calculations help determine how many participants are needed to detect an effect of a given magnitude.

22 likes 5mo
YA
y.adeyemiTL2 Moderator18 Feb 2026#9

Confidence intervals: rather than a single point estimate, a range of plausible values. A narrow interval means precise measurement; a wide interval means measurement is imprecise. Wider intervals (more uncertainty) are honest about limitation.

15 likes 5mo
JR
j.rasmussenTL2Regular20 Feb 2026 · edited#10
t.karlsen, post #1: Measurement error in home scales, with a worked standard deviation Writing it up because I had to work it out twice and would rather nobody else did. Comparing SURPASS-4 ( Lancet , 2021) with STEP 1 ( N Engl J Med , 2021) and finding the comparison harder than it looks. Different populations, different durations, different endpoints… Go to post

Multiplicity and multiple comparisons: if you test many hypotheses, the chance of finding a false positive by random chance increases. That is why pre-specifying the primary hypothesis matters.

5 likes in reply to #1 5mo
SS
s.silvaTL2 Moderator21 Feb 2026#11

P-values and significance: p<0.05 means the data would be surprising if the null hypothesis were true, not that the null hypothesis is false. A non-significant p-value does not mean "no effect".

0 likes 5mo
EF
e.ferrariTL2 Moderator23 Feb 2026#12
SHermansen, post #2: the opening post answers the question as asked. The question underneath it is different. For anyone arriving from a search: the marked solution above is the direct answer, and the replies underneath it add the caveats that make it safe to use. Go to post

Practical note that does not fit anywhere else. Whatever you conclude from this topic, write down what you did and when. The single most useful thing in your own records is not any individual result; it is that they are dated and consecutive.

0 likes in reply to #2 5mo
AW
a.weissTL2 Moderator25 Feb 2026#13
methods_draft, post #4: Number needed to treat: how many people need to be treated to prevent one bad outcome or achieve one good outcome. More intuitive than relative risk reduction. Go to post

This follows post #10 rather than contradicting it.

Number needed to treat: how many people need to be treated to prevent one bad outcome or achieve one good outcome. More intuitive than relative risk reduction.

7 likes in reply to #4 5mo
KR
k.roosTL2 Moderator27 Feb 2026#14

Relative risk and odds ratios: both compare the rate in one group to the rate in another. Relative risk is easier to understand. Odds ratios are standard in many analyses but can be misinterpreted.

18 likes 5mo
V
VThorvaldsenTL3Regular28 Feb 2026#15

I disagree with the reply above, and I think the disagreement is substantive rather than terminological.

The distinction being drawn does not survive when you look at the published data for this specific question. I would be glad to be shown wrong on this, because the version I am arguing against is more convenient.

0 likes 5mo
IB
i.brobergTL2 Moderator2 Mar 2026 · edited#16
VThorvaldsen, post #15: I disagree with the reply above, and I think the disagreement is substantive rather than terminological. The distinction being drawn does not survive when you look at the published data for this specific question. I would be glad to be shown wrong on this, because the version I am arguing against is more convenient. Go to post

Multiplicity and multiple comparisons: if you test many hypotheses, the chance of finding a false positive by random chance increases. That is why pre-specifying the primary hypothesis matters.

1 like in reply to #15 5mo
K
KnowltonTL3Regular3 Mar 2026#17

Picking up post #14: that is the part I would want checked first.

Power and sample size: a study might be too small to detect a real effect (low power). Sample size calculations help determine how many participants are needed to detect an effect of a given magnitude.

12 likes 5mo
EK
e.kimaniTL2 Moderator5 Mar 2026#18

Coming back to post #16, because the follow-up matters more than the original answer.

Effect sizes: the magnitude of a difference, not just whether it is statistically significant. A difference that is significant (p<0.05) might be too small to matter. A large effect might not be significant if sample size is small.

25 likes 5mo
KL
k.laurentTL27 Mar 2026#19
DO
d.oyelaranTL3Pharmacist8 Mar 2026#20
Knowlton, post #17: Picking up post #14: that is the part I would want checked first. Power and sample size: a study might be too small to detect a real effect (low power). Sample size calculations help determine how many participants are needed to detect an effect of a given magnitude. Go to post

P-values and significance: p<0.05 means the data would be surprising if the null hypothesis were true, not that the null hypothesis is false. A non-significant p-value does not mean "no effect".

0 likes in reply to #17 5mo
AS
a.silvaTL2 Moderator10 Mar 2026#21

Confounding: a third variable explains an apparent association. In randomised data, randomisation balances confounders. In observational data, confounders can be adjusted for but unknown ones cannot.

10 likes 5mo
P
PSkarbekTL3Regular11 Mar 2026#22

Regression to the mean: if you select people with extreme values (very high or very low), their next measurement is often less extreme just by chance. This can look like a treatment effect when it is just statistics.

3 likes 5mo
MA
m.achebeTL2 Moderator13 Mar 2026#23

Absence of evidence and evidence of absence: if a study is small and finds no effect, that is absence of evidence, not evidence of absence. A larger study might find an effect that a small study missed.

0 likes 5mo
VK
v.krastevTL2 Moderator14 Mar 2026#24
Knowlton, post #17: Picking up post #14: that is the part I would want checked first. Power and sample size: a study might be too small to detect a real effect (low power). Sample size calculations help determine how many participants are needed to detect an effect of a given magnitude. Go to post

Picking up post #21: that is the part I would want checked first.

For anyone arriving from a search: the marked solution above is the direct answer, and the replies underneath it add the caveats that make it safe to use.

30 likes in reply to #17 4mo
AH
a.hartmannTL2 Moderator16 Mar 2026#25
VThorvaldsen, post #15: I disagree with the reply above, and I think the disagreement is substantive rather than terminological. The distinction being drawn does not survive when you look at the published data for this specific question. I would be glad to be shown wrong on this, because the version I am arguing against is more convenient. Go to post

Effect sizes: the magnitude of a difference, not just whether it is statistically significant. A difference that is significant (p<0.05) might be too small to matter. A large effect might not be significant if sample size is small.

15 likes in reply to #15 4mo
EC
excursion_checkTL3Regular17 Mar 2026#26

Power and sample size: a study might be too small to detect a real effect (low power). Sample size calculations help determine how many participants are needed to detect an effect of a given magnitude.

6 likes 4mo
KH
k.haddadTL2 Moderator19 Mar 2026#27

I read post #25 twice before replying, because I had assumed the opposite.

Regression to the mean: if you select people with extreme values (very high or very low), their next measurement is often less extreme just by chance. This can look like a treatment effect when it is just statistics.

1 like 4mo
BM
buffer_marginTL3Regular20 Mar 2026 · edited#28
m.achebe, post #23: Absence of evidence and evidence of absence: if a study is small and finds no effect, that is absence of evidence, not evidence of absence. A larger study might find an effect that a small study missed. Go to post

This follows post #25 rather than contradicting it.

Confounding: a third variable explains an apparent association. In randomised data, randomisation balances confounders. In observational data, confounders can be adjusted for but unknown ones cannot.

0 likes in reply to #23 4mo
TV
t.verhoevenTL2 Moderator21 Mar 2026#29

On post #25 — agreed on the reasoning, with one qualification.

Absence of evidence and evidence of absence: if a study is small and finds no effect, that is absence of evidence, not evidence of absence. A larger study might find an effect that a small study missed.

21 likes 4mo
LC
l.chevalierTL3Regular23 Mar 2026#30

post #29 answers the question as asked. The question underneath it is different.

I disagree with the reply above, and I think the disagreement is substantive rather than terminological.

The distinction being drawn does not survive when you look at the published data for this specific question. I would be glad to be shown wrong on this, because the version I am arguing against is more convenient.

9 likes 4mo