Exhibit two

A large effect, published as absent.

Relapse fell from 83% on placebo to 47% on methotrexate. The p-value came in at 0.06, and the trial entered the literature as negative. The same paper, in the same table, also reports the analysis at p = 0.019.

Nothing about the treatment changed between those two numbers.

The trial

Methotrexate in giant-cell arteritis

A randomised, double-blind, placebo-controlled trial asked whether methotrexate prevents relapse in giant-cell arteritis. This is the completion-of-treatment analysis, as published.

Jover JA, et al. Ann Intern Med. 2001 · Relapse, completion-of-treatment analysis · N = 33
ArmRelapseNo relapseRate
Methotrexate 7 a 8 b 46.7%
Placebo 15 c 3 d 83.3%

Fisher’s exact, two-sided: above 0.05, so the classification is not significant. A 36.7-point difference in relapse rate is reported as no difference.

Fisher’s exact, two-sidedp = 0.061

Fragility — fr

One reclassified patient moves this trial across the threshold.

Fragility measures the stability of the statistical classification: what proportion of patients would have to be reclassified before the classification changes. Here that proportion is 1 in 33 — GFQ = 0.0303, well below the 0.05 cutpoint, so the classification is fragile.

Take one placebo patient who did not relapse. Count that patient instead as a methotrexate patient who did. That is a single reallocation out of 33 — the toggle d → a — and the classification changes.

Methotrexate

7 relapses of 15

Placebo

15 relapses of 18
After one global movep = 0.061

The same count, two different move-sets

The FI* may only toggle outcomes inside one arm, holding that arm’s total fixed. The Global Fragility Index searches every cell-to-cell reallocation in the table, with neither margin fixed. Both find a single move here — but not the same move, and not the same resulting p-value.

FI* = 1

a → b

Confined to one arm, selected by the rule that takes the arm with fewest events — methotrexate, 7 against 15. That arm’s total stays at 15. One patient who relapsed instead does not.

(6, 9, 15, 3)

p = 0.014274

GFI = 1

d → a

The whole table, neither margin fixed. One placebo patient without relapse is instead a methotrexate patient with relapse.

(8, 8, 15, 2)

p = 0.025510

A within-arm toggle is itself a cell-to-cell reallocation, so every move available to the FI is also available to the global search. The FI is a constrained special case, which makes GFI ≤ FI a mathematical proof, confirmed empirically. When the two counts agree, as they do here, they still name different toggles reaching different tables. When they disagree, the FI reports a trial as more stable than it is using an unconstrained toggle space.

Robustness — nb

The effect is robust, with strong separation from neutrality.

Robustness measures how far the observed result sits from therapeutic neutrality, on a bounded 0–1 scale. It is a property of the point estimate, not of its precision — which is exactly why it can disagree with a p-value.

nb (RQ) 0.363636 Strong. The published bands are weak below 0.075, moderate from 0.075 to 0.227, strong at 0.227 and above.
NDI 3 The same distance stated as a patient count: 3 paired moves bring this table to the point closest to no effect. NDI = round(|ad − bc|/N) = round(99/33).

NDI — the Neutrality Distance Index — is a count, so it grows with trial size and is not comparable between trials. CREST’s NDI is 7 against this trial’s 3, yet CREST is weak and this is strong. Use RQ to compare trials; use NDI to state one trial’s distance in patients.

Neutrality Distance Index: walking the table to therapeutic neutrality

A paired move relocates two patients, one in each arm: (a, b, c, d) → (a+1, b−1, c−1, d+1). Because it takes one from each arm and gives one back to each, both margins stay fixed — the arms remain 15 and 18, and the outcome totals remain 22 and 11 at every step. Only the risk ratio moves.

Paired move(a, b, c, d)Risk, MethotrexateRisk, PlaceboRRad − bc
as published (7, 8, 15, 3) 46.67% 83.33% 0.5600 -99
move 1 (8, 7, 14, 4) 53.33% 77.78% 0.6857 -66
move 2 (9, 6, 13, 5) 60.00% 72.22% 0.8308 -33
move 3 (10, 5, 12, 6) 66.67% 66.67% 1.0000 0

Each paired move shifts ad − bc by exactly 33 — that is, by N. Reachable values are therefore spaced N apart, which is why NDI = round(|ad − bc| / N). Here 99 divides evenly by 33, so the table lands on exact neutrality: RR = 1.0000. That is not typical — most tables stop at the reachable point nearest RR = 1 without hitting it.

Why this is not the GFI

The two indices count moves, but they use different moves, hold different things fixed, and aim at different targets.

NDI — robustnessGFI — fragility
MovePaired: two patients, one per armSingle transfer: one patient, any cell to any cell
MarginsBoth fixed — 15/18 and 22/11 throughoutNeither fixed — here they become 16/17 and 23/10
TargetRR = 1, therapeutic neutralityp crossing α = 0.05
Test usedNone. It depends only on ad − bcFisher’s exact, two-sided
Answer here3 paired moves, RR 0.56 → 1.001 move, d → a, p 0.061 → 0.026

Follow Fisher’s exact along the ladder above and it reads 0.061, 0.163, 0.488, 1.000 — walking away from significance while the NDI walks toward neutrality. NDI never consults it. The two indices point in opposite directions here, which is the clearest possible sign that fragility and robustness are measuring different things.

The triplet

Complete evidence for this trial

pattern (not significant, fragile, strong)

Fragile nonsignificance, yet far from neutrality. The published interpretation of this cell is unambiguous: likely false negative — the effect exists but was not detected. The trial did not show that methotrexate fails. It showed that this trial was too small to say.

The demonstration

The same trial, analysed twice, in the same table.

Table 2 of the paper reports two analysis populations. They straddle the significance threshold. Watch which number moves and which number does not.

Completion of treatmentCompletion of follow-up
2×2 7 / 8 vs 15 / 3 9 / 11 vs 16 / 3
N3339
p 0.0613 — not significant 0.0187 — significant
GFI12
fr (GFQ) 0.0303 — fragile 0.0513 — stable
nb (RQ) 0.3636 — strong 0.3918 — strong
Pattern not significant, fragile, strong significant, stable, strong
Impression Likely false negative — increase power Concordant (strong) evidence

The p-value moved from 0.061 to 0.019 and carried the verdict with it. The robustness moved from 0.364 to 0.392. Robustness was strong and never in question. Only the classification was — but without complete statistical evidence, only the significance classification (the p-value) was what the literature reported.

This is what the triplet is for. Reported alone, the first analysis says “negative trial”. Reported as p–fr–nb, it says “large effect, unstable classification, underpowered” — which is both true and actionable.

Against exhibit one

Identical fragility. Opposite failures.

CREST and this trial each sit one reallocation from the opposite classification. GFI = 1 for both. What separates them is robustness.

CREST 2011 — myocardial infarctionJover 2001 — relapse
p 0.0290 — significant 0.0613 — not significant
GFI11
fr (GFQ) 0.0004 — fragile 0.0303 — fragile
nb (RQ) 0.0115 — weak 0.3636 — strong
Pattern significant, fragile, weak not significant, fragile, strong
Impression Thin evidence — remain skeptical Likely false negative — increase power

One trial reported a negligible effect as established. The other reported a substantial effect as absent. Both passed peer review at a major journal. Both are a single reclassified patient from the opposite conclusion, and neither p-value carried that information.

Check these numbers.

Every figure on this page is reproducible from the two 2×2 tables above. The calculators and the full specification of GFI, GFQ, RQ and NDI are free.

Read the specification →

Trial data from Jover JA, et al. Ann Intern Med. 2001, Table 2 (doi:10.7326/0003-4819-134-2-200101160-00010). Both analysis populations are as published; the p-values here are recomputed by Fisher’s exact test and agree with the paper’s reported 0.06 and 0.018.

I endorse the Complete Evidence Standard.

Endorsement is a public statement of practice, not a membership fee. There is nothing to pay and nothing to renew.

Your email is used to confirm your endorsement and is never published or shared.

Signatories

1

researcher or clinician has endorsed the Standard.

Thomas F. Heston, MD, MSc