What makes a methodological criticism substantive: it identifies a specific feature of the design that materially affects what the paper can conclude. "Small sample size" alone is weak. "Small sample size for a rare outcome, so the confidence interval is wide" is stronger.
[2026 update] Confounding by indication, explained with a concrete example posts 91–117
This is a continuation of a long topic, addressed by post number rather than by page. Start at post 1.
Generalisability: do the inclusion/exclusion criteria narrow the population so much that results do not apply to real people asking about it? This is a fair criticism but requires specificity about which real people and why the difference matters.
This follows post #90 rather than contradicting it.
Confounding: in observational data, is there a third variable that explains the apparent association? In randomised data, randomisation should balance unknown confounders, though known confounders can be adjusted for.
Bias towards the null and bias away from the null: different criticisms have different directions. Differential dropout might bias away from null; conservative statistical analysis might bias toward null.
Publication bias: a single published positive trial is weaker evidence than multiple published trials with consistent results. Asking whether there are unpublished negative trials is a fair critical question.
Practical note that does not fit anywhere else. Whatever you conclude from this topic, write down what you did and when. The single most useful thing in your own records is not any individual result; it is that they are dated and consecutive.
Picking up post #94: that is the part I would want checked first.
Multiple comparisons: if a paper reports many outcomes, the chance of a spurious association by random chance is real. Pre-specification of primary outcomes matters and secondary analyses are weaker evidence.
Coming back to post #96, because the follow-up matters more than the original answer.
Choosing the worst interpretation: "The confidence interval includes a harmful effect" is true if the CI goes from -1 to +5. But assuming the worst-case scenario is not how you use the evidence. The point estimate and the precision both matter.
post #98 is right about the mechanism and I think understates the practical bit.
Generalisability: do the inclusion/exclusion criteria narrow the population so much that results do not apply to real people asking about it? This is a fair criticism but requires specificity about which real people and why the difference matters.
Collapsed as off-topic by two members at trust level 3 or above
Worth separating two things that post #96 runs together.
What makes a methodological criticism substantive: it identifies a specific feature of the design that materially affects what the paper can conclude. "Small sample size" alone is weak. "Small sample size for a rare outcome, so the confidence interval is wide" is stronger.
Collapsed as off-topic by two members at trust level 3 or above
Defending a paper against criticism: if the authors respond, they might clarify something the paper explained poorly. Their response might also miss your point. Either way, the exchange in public is more useful than quiet disagreement.
When you change your mind: if a reply convinces you that your criticism was not well-founded, say so plainly. The critique might still be real but smaller than you originally thought. That is not a failure — it is how discussion works.
Picking up post #100: that is the part I would want checked first.
I disagree with the reply above, and I think the disagreement is substantive rather than terminological.
The distinction being drawn does not survive when you look at the published data for this specific question. I would be glad to be shown wrong on this, because the version I am arguing against is more convenient.
Coming back to post #102, because the follow-up matters more than the original answer.
Publication bias: a single published positive trial is weaker evidence than multiple published trials with consistent results. Asking whether there are unpublished negative trials is a fair critical question.
Multiple comparisons: if a paper reports many outcomes, the chance of a spurious association by random chance is real. Pre-specification of primary outcomes matters and secondary analyses are weaker evidence.
Practical note that does not fit anywhere else. Whatever you conclude from this topic, write down what you did and when. The single most useful thing in your own records is not any individual result; it is that they are dated and consecutive.
Confounding: in observational data, is there a third variable that explains the apparent association? In randomised data, randomisation should balance unknown confounders, though known confounders can be adjusted for.
I read post #106 twice before replying, because I had assumed the opposite.
Bias towards the null and bias away from the null: different criticisms have different directions. Differential dropout might bias away from null; conservative statistical analysis might bias toward null.
post #108 answers the question as asked. The question underneath it is different.
I disagree with the reply above, and I think the disagreement is substantive rather than terminological.
The distinction being drawn does not survive when you look at the published data for this specific question. I would be glad to be shown wrong on this, because the version I am arguing against is more convenient.
Collapsed as off-topic by two members at trust level 3 or above
On post #106 — agreed on the reasoning, with one qualification.
Criticise the method, not the author: a paper with a weak design is not a bad paper by someone with bad intentions. It is a paper that answers a limited question. Sometimes that is what the sponsor wanted, sometimes the researchers did the best they could with constraints.
Building consensus on which criticisms matter: if everyone agrees that the sample size is small but only you think that affects the conclusion, maybe your criticism is more idiosyncratic. That does not make it wrong but it is worth noticing.
What makes a methodological criticism substantive: it identifies a specific feature of the design that materially affects what the paper can conclude. "Small sample size" alone is weak. "Small sample size for a rare outcome, so the confidence interval is wide" is stronger.
Generalisability: do the inclusion/exclusion criteria narrow the population so much that results do not apply to real people asking about it? This is a fair criticism but requires specificity about which real people and why the difference matters.
When you change your mind: if a reply convinces you that your criticism was not well-founded, say so plainly. The critique might still be real but smaller than you originally thought. That is not a failure — it is how discussion works.
On post #113 — agreed on the reasoning, with one qualification.
Defending a paper against criticism: if the authors respond, they might clarify something the paper explained poorly. Their response might also miss your point. Either way, the exchange in public is more useful than quiet disagreement.
This topic was referenced in
- Criticising the method without criticising the authors — one year onEvidence › Study critique · 20 replies
Suggested topics
| Topic | Participants | Replies | Views | Activity |
|---|---|---|---|---|
|
A well-designed study with a badly written abstract
Posting this under the heading it deserves: A well-designed study with a badly written abstract Everything below is what sits behind that. Comparing SCALE ( N Engl J Med , 2015) with LEADER ( N Engl J Med ,…
|
+117 | 128 | 7.2k | 14mo |
|
Coming back to: Measurement error in a self-reported exposure
Measurement error in a self-reported exposure — setting out what I have, and where I think it stops being reliable. I have seen FLOW ( N Engl J Med , 2024) cited in support of a claim I do not think it…
|
+23 | 27 | 49k | 3mo |
|
A structured critique template this community uses — a second dataset
On the subject in the title: A structured critique template this community uses — a second dataset Working notes rather than a conclusion. Session topic: SELECT ( N Engl J Med , 2023). Please read it before…
|
+47 | 51 | 16k | 15h |
|
Coming back to: A critique that turned out to be unfair, retracted by its author
A critique that turned out to be unfair, retracted by its author — setting out what I have, and where I think it stops being reliable. I have seen STEP 4 ( JAMA , 2021) cited in support of a claim I do not…
|
2 | 1.3k | 14mo | |
|
Immortal time bias in a claims-database study — one year on
Posting this under the heading it deserves: Immortal time bias in a claims-database study — one year on Everything below is what sits behind that. Comparing SURPASS-2 ( N Engl J Med , 2021) with SURMOUNT-1 (…
|
+41 | 46 | 3.7k | 22h |
Related topics — sharing the tags observational data, disputed, risk of bias
| Topic | Participants | Replies | Views | Activity |
|---|---|---|---|---|
|
Journal club: SUSTAIN 6 and its retinopathy signal
On the subject in the title: Journal club: SUSTAIN 6 and its retinopathy signal Working notes rather than a conclusion. I have seen SELECT ( N Engl J Med , 2023) cited in support of a claim I do not think it…
|
+110 | 122 | 46k | 1d |
|
The "under review" status pill and when a moderator applies it — a second dataset
The "under review" status pill and when a moderator applies it — a second dataset Writing it up because I had to work it out twice and would rather nobody else did. I have seen SURPASS-4 ( Lancet , 2021)…
|
+50 | 54 | 1.4k | 11mo |
|
Secretagogues and glucose tolerance: the mechanistic concern — what changed since
Secretagogues and glucose tolerance: the mechanistic concern — what changed since Writing it up because I had to work it out twice and would rather nobody else did. I have seen STEP 1 ( N Engl J Med , 2021)…
|
+2 | 6 | 31k | 2mo |
|
Journal club: SUSTAIN 6 and its retinopathy signal — does this still hold?
The question in the title: Journal club: SUSTAIN 6 and its retinopathy signal — does this still hold? I will give what I have already checked below so nobody repeats it. I have seen SURMOUNT-2 ( Lancet ,…
|
+60 | 66 | 12k | 16mo |
|
Revisiting: Filing a first vendor report: a worked example
On the subject in the title: Revisiting: Filing a first vendor report: a worked example Working notes rather than a conclusion. Structured report rather than an opinion, following the format the maintainers…
|
2 | 60k | 11mo |