The Peptide CommonsEst. May 2024
Independent. We sell nothing and are affiliated with no manufacturer or pharmacy. Every moderation action is logged in public
Evidence · Trials

Revisiting: Primary endpoint hierarchies and why order matters

AH
a.hartmannTL2 Moderator3 May 2026#1

Posting this under the heading it deserves: Revisiting: Primary endpoint hierarchies and why order matters Everything below is what sits behind that.

Comparing SURMOUNT-4 (JAMA, 2024) with SURPASS-2 (N Engl J Med, 2021) and finding the comparison harder than it looks.

Different populations, different durations, different endpoints defined slightly differently, and in one case a different estimand. People compare the headline percentages anyway, including me until recently.

Is there a defensible way to put these side by side, or is the honest answer that there is not and we should stop?

43 likes 3mo
G
GDashwoodTL3Regular4 May 2026 · edited#2

Practical note that does not fit anywhere else. Whatever you conclude from this topic, write down what you did and when. The single most useful thing in your own records is not any individual result; it is that they are dated and consecutive.

0 likes 3mo
JT
j.teixeiraTL2 Moderator4 May 2026#3

Confounding in observational data: a third variable can explain an apparent association. In a randomised trial, randomisation balances unknown confounders. In observational data, observed confounders can be adjusted for but unknown ones cannot.

3 likes 3mo
CV
c.vermeulenTL2 Moderator5 May 2026#4

Surrogate endpoints: an endpoint that is not the outcome that matters but is measured as a stand-in. HbA1c is a surrogate for long-term glucose control and the short-term complications it prevents. Weight loss is a surrogate for metabolic health and long-term outcomes. Surrogates are useful but not identical to the endpoint that matters.

11 likes 3mo
MS
m.silvaTL2 Moderator5 May 2026#5

post #4 is right about the mechanism and I think understates the practical bit.

Thank you for the correction. I have edited my earlier post with a note rather than silently, so the thread still makes sense to read. The error was mine and it was the kind that comes from remembering a figure instead of looking it up.

17 likes 3mo
VS
v.salgadoTL2 Moderator6 May 2026 · edited#6
a.hartmann, post #1: Posting this under the heading it deserves: Revisiting: Primary endpoint hierarchies and why order matters Everything below is what sits behind that. Comparing SURMOUNT-4 ( JAMA , 2024) with SURPASS-2 ( N Engl J Med , 2021) and finding the comparison harder than it looks. Different populations, different durations, different endpoints… Go to post

Worth separating two things that post #2 runs together.

Risk of bias: structured appraisal of internal validity. Key things to assess: randomisation method (was it truly random or could someone predict the next assignment), concealment (could randomisation be subverted), blinding (who was blinded and why or why not), completeness of outcome reporting.

32 likes in reply to #1 3mo
LC
lu.cabreraTL2 Moderator6 May 2026#7

The estimand: what the trial set out to estimate. Two trials can be identical in structure but estimate different things by using different handling rules for people who stop taking the drug. Treatment-policy and hypothetical approaches are both legitimate but answer different questions.

1 like 3mo
NS
n.stanescuTL2 Moderator7 May 2026#8

Multiplicity and multiple comparisons: if a trial tests many hypotheses, the chance of a false positive on at least one by random chance increases. This is why pre-specification of the primary endpoint matters and why secondary endpoints are weaker evidence.

7 likes 3mo
LC
l.cabreraTL2 Moderator7 May 2026#9

Intent-to-treat versus per-protocol: ITT includes everyone assigned regardless of whether they took the drug. Per-protocol includes only those who completed it as intended. The two can give substantially different results.

11 likes 3mo
AS
a.stephanopoulosTL3Regular7 May 2026#10

On post #6 — agreed on the reasoning, with one qualification.

Absolute numbers, not just relative: a 30% relative reduction tells you the ratio but not the practical magnitude. The event rate in each arm and the difference between them tells you how many people benefit.

24 likes 3mo
EF
endo_fellow_rkTL3Endocrinology fellow8 May 2026#11

Practical note that does not fit anywhere else. Whatever you conclude from this topic, write down what you did and when. The single most useful thing in your own records is not any individual result; it is that they are dated and consecutive.

0 likes 3mo
YA
y.adebayoTL2 Moderator8 May 2026#12

Population narrowness: most trials in this class enrolled fairly specific groups. Baseline body mass index ranges, exclusion of renal disease, exclusion of certain comorbidities, all narrow the population. Applying point estimates to someone well outside the range is an extrapolation.

32 likes 3mo
MH
ms_hollowayTL4Mass spectrometrist8 May 2026#13
lu.cabrera, post #7: The estimand: what the trial set out to estimate. Two trials can be identical in structure but estimate different things by using different handling rules for people who stop taking the drug. Treatment-policy and hypothetical approaches are both legitimate but answer different questions. Go to post

On post #9 — agreed on the reasoning, with one qualification.

Dropout is information: high dropout rates can indicate tolerability problems or lower efficacy than the summary suggests. Where the analysis handled dropouts matters. An intention-to-treat analysis with many dropouts can give a smaller apparent effect than per-protocol analysis.

16 likes in reply to #7 3mo
MI
m.ibarraTL2 Moderator9 May 2026#14

Generalisability: the enrolled population was selected in ways that matter. Entry criteria, run-in periods, and the simple fact that people who agree to a multi-year trial differ from people who do not, all narrow the population. That is how internal validity is bought, at the cost of external validity.

6 likes 3mo
FN
formulary_notesTL3Regular9 May 2026#15

Surrogate endpoints: an endpoint that is not the outcome that matters but is measured as a stand-in. HbA1c is a surrogate for long-term glucose control and the short-term complications it prevents. Weight loss is a surrogate for metabolic health and long-term outcomes. Surrogates are useful but not identical to the endpoint that matters.

0 likes 3mo
CA
c.amankwahTL2 Moderator9 May 2026 · edited#16

Confounding in observational data: a third variable can explain an apparent association. In a randomised trial, randomisation balances unknown confounders. In observational data, observed confounders can be adjusted for but unknown ones cannot.

24 likes 3mo
TH
TL4_HalvorsenTL4Leader · Journal club10 May 2026#17
n.stanescu, post #8: Multiplicity and multiple comparisons: if a trial tests many hypotheses, the chance of a false positive on at least one by random chance increases. This is why pre-specification of the primary endpoint matters and why secondary endpoints are weaker evidence. Go to post

Worth separating two things that post #13 runs together.

Risk of bias: structured appraisal of internal validity. Key things to assess: randomisation method (was it truly random or could someone predict the next assignment), concealment (could randomisation be subverted), blinding (who was blinded and why or why not), completeness of outcome reporting.

11 likes in reply to #8 3mo
CC
ch.correiaTL2 Moderator10 May 2026#18
m.silva, post #5: post #4 is right about the mechanism and I think understates the practical bit. Thank you for the correction. I have edited my earlier post with a note rather than silently, so the thread still makes sense to read. The error was mine and it was the kind that comes from remembering a figure instead of looking it up. Go to post

post #17 is right about the mechanism and I think understates the practical bit.

For anyone arriving from a search: the marked solution above is the direct answer, and the replies underneath it add the caveats that make it safe to use.

3 likes in reply to #5 3mo
NM
n.moreauTL2 Moderator10 May 2026#19

Coming back to post #17, because the follow-up matters more than the original answer.

Open-label design: unblinded trials admit expectation effects. For weight-loss trials where one arm loses substantial weight and the other does not, complete blinding is impossible anyway. The unblinded nature is a limitation worth noting.

3 likes 3mo
PM
p.mbekiTL2 Moderator10 May 2026#20
lu.cabrera, post #7: The estimand: what the trial set out to estimate. Two trials can be identical in structure but estimate different things by using different handling rules for people who stop taking the drug. Treatment-policy and hypothetical approaches are both legitimate but answer different questions. Go to post

Picking up post #17: that is the part I would want checked first.

Intent-to-treat versus per-protocol: ITT includes everyone assigned regardless of whether they took the drug. Per-protocol includes only those who completed it as intended. The two can give substantially different results.

0 likes in reply to #7 3mo
S
SHermansenTL2Member11 May 2026#21

The estimand: what the trial set out to estimate. Two trials can be identical in structure but estimate different things by using different handling rules for people who stop taking the drug. Treatment-policy and hypothetical approaches are both legitimate but answer different questions.

29 likes 3mo
NR
n.ramosTL2 Moderator11 May 2026#22
lu.cabrera, post #7: The estimand: what the trial set out to estimate. Two trials can be identical in structure but estimate different things by using different handling rules for people who stop taking the drug. Treatment-policy and hypothetical approaches are both legitimate but answer different questions. Go to post

Coming back to post #20, because the follow-up matters more than the original answer.

Multiplicity and multiple comparisons: if a trial tests many hypotheses, the chance of a false positive on at least one by random chance increases. This is why pre-specification of the primary endpoint matters and why secondary endpoints are weaker evidence.

0 likes in reply to #7 3mo
R
RodriguesTL3Regular11 May 2026#23

post #22 answers the question as asked. The question underneath it is different.

Generalisability: the enrolled population was selected in ways that matter. Entry criteria, run-in periods, and the simple fact that people who agree to a multi-year trial differ from people who do not, all narrow the population. That is how internal validity is bought, at the cost of external validity.

20 likes 3mo
AZ
an.zamoraTL212 May 2026#24
IS
isotonic_sheetTL3Regular12 May 2026#25
an.zamora, post #24: Absolute numbers, not just relative: a 30% relative reduction tells you the ratio but not the practical magnitude. The event rate in each arm and the difference between them tells you how many people benefit. Go to post

Dropout is information: high dropout rates can indicate tolerability problems or lower efficacy than the summary suggests. Where the analysis handled dropouts matters. An intention-to-treat analysis with many dropouts can give a smaller apparent effect than per-protocol analysis.

21 likes in reply to #24 3mo
PT
p.trevinoTL2 Moderator12 May 2026#26
v.salgado, post #6: Worth separating two things that post #2 runs together. Risk of bias: structured appraisal of internal validity. Key things to assess: randomisation method (was it truly random or could someone predict the next assignment), concealment (could randomisation be subverted), blinding (who was blinded and why or why not), completeness of… Go to post

Population narrowness: most trials in this class enrolled fairly specific groups. Baseline body mass index ranges, exclusion of renal disease, exclusion of certain comorbidities, all narrow the population. Applying point estimates to someone well outside the range is an extrapolation.

0 likes in reply to #6 3mo
Promoted into the documentation commons. The content of this topic is maintained at PIONEER 6 — trial digest, with named maintainers and a review date. The promotion was discussed in doc review. Corrections are best raised against the document, which is the version that gets kept current.

Suggested topics

TopicParticipantsRepliesViewsActivity
Run-in periods and the population they select — the long version
Run-in periods and the population they select — the long version Writing it up because I had to work it out twice and would rather nobody else did. Session topic: SUSTAIN 6 ( N Engl J Med , 2016). Please read…
ALFREK 2 8.9k 4mo
About the Trials category
Individual randomised trials: design, population, endpoints, results, and limitations. This post is a community wiki: any member at trust level 3 or above can edit it, and every edit is recorded with its…
JMBNSSZOEM+6 10 27k 12mo
Non-inferiority margins: how they are chosen and how they are abused
Posting this under the heading it deserves: Non-inferiority margins: how they are chosen and how they are abused Everything below is what sits behind that. Comparing SURPASS-2 ( N Engl J Med , 2021) with…
NTKCTDOPA+42 46 10k 5mo
Interim analyses and stopping rules
Interim analyses and stopping rules Writing it up because I had to work it out twice and would rather nobody else did. I have seen SUSTAIN 6 ( N Engl J Med , 2016) cited in support of a claim I do not think…
TKHJRIGEL+25 29 35k 16d
Coming back to: Sample size calculations, read backwards from the published number
On the subject in the title: Sample size calculations, read backwards from the published number Working notes rather than a conclusion. Session topic: PIONEER 6 ( N Engl J Med , 2019). Please read it before…
PRIEBMMTV+82 88 1.8k 11mo

Related topics — sharing the tags estimand, surrogate endpoints, discontinuation & dropout

TopicParticipantsRepliesViewsActivity
Adjudicated events and why the definition matters — the long version
Adjudicated events and why the definition matters — the long version Writing it up because I had to work it out twice and would rather nobody else did. Comparing SURPASS-4 ( Lancet , 2021) with STEP 4 ( JAMA…
JSEFIBSSRE+125 146 13k 15mo
Follow-up: Open-label extensions: what survives and what does not
Open-label extensions: what survives and what does not — setting out what I have, and where I think it stops being reliable. Comparing SURMOUNT-2 ( Lancet , 2023) with SURMOUNT-1 ( N Engl J Med , 2022) and…
MATPBLCMA+22 26 397 16mo
Interim analyses and stopping rules — does this still hold?
Interim analyses and stopping rules — does this still hold? — that is the question, and I have not found it answered plainly anywhere I have looked. Comparing SCALE ( N Engl J Med , 2015) with SURMOUNT-4 (…
JDMHKMFTHL+30 34 29k 13mo
Journal club: the CagriSema phase 2 combination paper
On the subject in the title: Journal club: the CagriSema phase 2 combination paper Working notes rather than a conclusion. I have seen SURMOUNT-1 ( N Engl J Med , 2022) cited in support of a claim I do not…
PFKMM 2 329 21h
About the Journal club category
Our recurring session. One named paper per topic, methods first, conclusions last. Everyone reads before posting. This post is a community wiki: any member at trust level 3 or above can edit it, and every…
KHSDEMSCA+2 6 12k 16mo