Multiple comparisons: if a paper reports many outcomes, the chance of a spurious association by random chance is real. Pre-specification of primary outcomes matters and secondary analyses are weaker evidence.
A structured critique template this community uses — a second dataset posts 31–52
This is a continuation of a long topic, addressed by post number rather than by page. Start at post 1 · go to the accepted answer.
Publication bias: a single published positive trial is weaker evidence than multiple published trials with consistent results. Asking whether there are unpublished negative trials is a fair critical question.
Worth separating two things that post #29 runs together.
Choosing the worst interpretation: "The confidence interval includes a harmful effect" is true if the CI goes from -1 to +5. But assuming the worst-case scenario is not how you use the evidence. The point estimate and the precision both matter.
When you change your mind: if a reply convinces you that your criticism was not well-founded, say so plainly. The critique might still be real but smaller than you originally thought. That is not a failure — it is how discussion works.
Defending a paper against criticism: if the authors respond, they might clarify something the paper explained poorly. Their response might also miss your point. Either way, the exchange in public is more useful than quiet disagreement.
I disagree with the reply above, and I think the disagreement is substantive rather than terminological.
The distinction being drawn does not survive when you look at the published data for this specific question. I would be glad to be shown wrong on this, because the version I am arguing against is more convenient.
Collapsed as off-topic by two members at trust level 3 or above
On post #33 — agreed on the reasoning, with one qualification.
Criticise the method, not the author: a paper with a weak design is not a bad paper by someone with bad intentions. It is a paper that answers a limited question. Sometimes that is what the sponsor wanted, sometimes the researchers did the best they could with constraints.
post #37 answers the question as asked. The question underneath it is different.
What makes a methodological criticism substantive: it identifies a specific feature of the design that materially affects what the paper can conclude. "Small sample size" alone is weak. "Small sample size for a rare outcome, so the confidence interval is wide" is stronger.
I read post #37 twice before replying, because I had assumed the opposite.
Two things before anyone answers the substance.
First, the context in the first post is clear and specific. Second, the question is framed so that an answer can actually address it. Both are the norm here and both matter more than they sound.
Bias towards the null and bias away from the null: different criticisms have different directions. Differential dropout might bias away from null; conservative statistical analysis might bias toward null.
Coming back to post #40, because the follow-up matters more than the original answer.
Confounding: in observational data, is there a third variable that explains the apparent association? In randomised data, randomisation should balance unknown confounders, though known confounders can be adjusted for.
post #42 answers the question as asked. The question underneath it is different.
Publication bias: a single published positive trial is weaker evidence than multiple published trials with consistent results. Asking whether there are unpublished negative trials is a fair critical question.
Practical note that does not fit anywhere else. Whatever you conclude from this topic, write down what you did and when. The single most useful thing in your own records is not any individual result; it is that they are dated and consecutive.
This follows post #42 rather than contradicting it.
When you change your mind: if a reply convinces you that your criticism was not well-founded, say so plainly. The critique might still be real but smaller than you originally thought. That is not a failure — it is how discussion works.
I read post #44 twice before replying, because I had assumed the opposite.
Choosing the worst interpretation: "The confidence interval includes a harmful effect" is true if the CI goes from -1 to +5. But assuming the worst-case scenario is not how you use the evidence. The point estimate and the precision both matter.
Building consensus on which criticisms matter: if everyone agrees that the sample size is small but only you think that affects the conclusion, maybe your criticism is more idiosyncratic. That does not make it wrong but it is worth noticing.
Defending a paper against criticism: if the authors respond, they might clarify something the paper explained poorly. Their response might also miss your point. Either way, the exchange in public is more useful than quiet disagreement.
Picking up post #46: that is the part I would want checked first.
Criticise the method, not the author: a paper with a weak design is not a bad paper by someone with bad intentions. It is a paper that answers a limited question. Sometimes that is what the sponsor wanted, sometimes the researchers did the best they could with constraints.
On post #47 — agreed on the reasoning, with one qualification.
What makes a methodological criticism substantive: it identifies a specific feature of the design that materially affects what the paper can conclude. "Small sample size" alone is weak. "Small sample size for a rare outcome, so the confidence interval is wide" is stronger.
Generalisability: do the inclusion/exclusion criteria narrow the population so much that results do not apply to real people asking about it? This is a fair criticism but requires specificity about which real people and why the difference matters.
This topic was referenced in
- Criticising the method without criticising the authors — one year onEvidence › Study critique · 20 replies
Suggested topics
| Topic | Participants | Replies | Views | Activity |
|---|---|---|---|---|
|
Reverse causation in a cohort study of weight and outcome — a second dataset
Reverse causation in a cohort study of weight and outcome — a second dataset Writing it up because I had to work it out twice and would rather nobody else did. Session topic: SURMOUNT-OSA ( N Engl J Med ,…
|
+112 | 131 | 1.5k | 10mo |
|
Measurement error in a self-reported exposure
Posting this under the heading it deserves: Measurement error in a self-reported exposure Everything below is what sits behind that. Session topic: SURPASS-2 ( N Engl J Med , 2021). Please read it before…
|
+117 | 129 | 45k | 5mo |
|
Immortal time bias in a claims-database study — one year on
Posting this under the heading it deserves: Immortal time bias in a claims-database study — one year on Everything below is what sits behind that. Comparing SURPASS-2 ( N Engl J Med , 2021) with SURMOUNT-1 (…
|
+41 | 46 | 3.7k | 22h |
|
Immortal time bias in a claims-database study
Immortal time bias in a claims-database study — setting out what I have, and where I think it stops being reliable. I have seen FLOW ( N Engl J Med , 2024) cited in support of a claim I do not think it…
|
+29 | 34 | 28k | 11mo |
|
Critiquing an observational claim about a class effect
Posting this under the heading it deserves: Critiquing an observational claim about a class effect Everything below is what sits behind that. Session topic: SCALE ( N Engl J Med , 2015). Please read it before…
|
2 | 6.9k | 14mo |
Related topics — sharing the tags observational data, effect size, confounding
| Topic | Participants | Replies | Views | Activity |
|---|---|---|---|---|
|
TB-500 and thymosin beta-4: the fragment versus the protein
TB-500 and thymosin beta-4: the fragment versus the protein Writing it up because I had to work it out twice and would rather nobody else did. I have seen SURPASS-4 ( Lancet , 2021) cited in support of a…
|
+107 | 118 | 22k | 11mo |
|
Contributing data without breaching anyone's privacy
Contributing data without breaching anyone's privacy Writing it up because I had to work it out twice and would rather nobody else did. I would like to understand what this number means before I repeat it…
|
+65 | 71 | 839 | 5mo |
|
Revisiting: Sample size in a voluntary survey: the selection problem
Revisiting: Sample size in a voluntary survey: the selection problem — setting out what I have, and where I think it stops being reliable. Posting the method first, because I know what the first three replies…
|
2 | 56k | 9mo | |
|
How this community labels and discusses unrefereed work
How this community labels and discusses unrefereed work — that is the question, and I have not found it answered plainly anywhere I have looked. I have seen SCALE ( N Engl J Med , 2015) cited in support of a…
|
+112 | 119 | 1.3k | 2y |
|
Journal club: STEP-HFpEF and symptom endpoints
Journal club: STEP-HFpEF and symptom endpoints — setting out what I have, and where I think it stops being reliable. Comparing SURMOUNT-2 ( Lancet , 2023) with SURMOUNT-1 ( N Engl J Med , 2022) and finding…
|
+18 | 22 | 2.5k | 11mo |