The Internet Medical Association

“Statistically significant” is not a verdict. It is one of three numbers.

Every trial you read reports the first number. Almost none report the other two — and without them you cannot tell a finding that will replicate from one that is already indistinguishable from nothing.

Below is a published trial. Judge it the way you normally would, then keep scrolling.

Exhibit one

A trial that changed a conversation

In a randomised comparison of carotid endarterectomy against stenting, surgical patients had roughly twice the rate of myocardial infarction. The difference was statistically significant.

Blackshear JL, et al. Circulation. 2011 · Myocardial infarction · N = 2,502
ArmMINo MIRate
Endarterectomy 28 1,212 2.26%
Stenting 14 1,248 1.11%
Fisher’s exact, two-sidedp = 0.029

Dimension two — fragility

One patient.

Move a single myocardial infarction from the surgical arm to the stenting arm — one reallocation, out of 2,502 patients — and the finding is no longer significant.

Endarterectomy

28 MI of 1,240

Stenting

14 MI of 1,262
After one global movep = 0.029

This is the Global Fragility Index

The GFI is the smallest number of cell-to-cell reallocations, anywhere in the table, that flips the significance classification. Neither row nor column totals are held fixed — the search is global.

FI*
2
Confined to a single arm
GFI
1
The true global minimum
GFQ = GFI / N
0.0004
0.04% of patients reclassified. Below 5% → fragile

The FI* reported 2 here because it may only toggle outcomes within the arm that had fewer events. Allowed to search the whole table, the answer is 1. Restricting the search to one arm overstates how stable this classification is — and it does so systematically. We come back to that.

Dimension three — robustness

It was never far from nothing.

Fragility tells you the statistical classification is unstable. It does not tell you whether there was much of an effect to begin with. That is a separate question, and it has a separate answer.

RQ — distance from therapeutic neutrality, where neutrality means the two arms are indistinguishable.

0.011CREST
p — significance
0.029
Significant
fr — fragility (GFQ)
0.0004
0.04% of patients — fragile
nb — robustness (RQ)
0.011
At the neutrality boundary

The two arms differ by 1.1 percentage points — 2.26% against 1.11%. On the neutrality scale that is RQ = 0.011: the result sits almost exactly on the boundary where treatment and control become indistinguishable.

pattern (significant, fragile, weak)
This is the SFW pattern: significant, fragile, weak. A strong indicator that a “statistically significant” result should be viewed with skepticism — the classification flips on a single patient, and the effect it describes is barely distinguishable from zero. Nothing in the published report tells you either of those things. The p-value is identical to that of a trial you would be right to act on.

SFW applies to a claim that an effect exists. A claim that an effect is absent reads the same three dimensions differently: there, a high p-value with a stable classification and a result close to neutrality is what supports the null.

Exhibit two

Now the same p-value, from a different trial.

Tocilizumab for giant cell arteritis. Complete remission at twelve weeks. p = 0.030 — statistically indistinguishable from the trial you just dismantled.

Villiger PM, et al. The Lancet. 2016 · Complete remission at 12 weeks · N = 30
ArmRemissionNo remission
Tocilizumab173
Placebo46
p
0.030
Significant
fr (GFQ)
0.033
3.3% of patients — below 5%, fragile
nb (RQ)
0.400
Clearly separated

Both trials on the same neutrality scale:

0.011CREST 0.400Tocilizumab
pattern (significant, fragile, strong)
Not SFW — a real effect that needs replication. Also fragile: one move flips the classification. But 36 times farther from neutrality than the first trial. Same p-value, same fragility verdict — and the third dimension is the only thing that tells them apart.

Fragility flagged both classifications as unstable. Only nb distinguished a finding worth confirming from one worth discarding. No standard reporting requirement asks for it. No other framework computes it.

Exhibit two, continued

They followed the same patients for a year.

Same trial, same thirty people, same drug — complete remission measured at fifty-two weeks instead of twelve. Watch all three numbers move.

Villiger PM, et al. The Lancet. 2016 · Complete remission at 52 weeks · N = 30
ArmRemissionNo remission
Tocilizumab173
Placebo28
p
0.00097
31× smaller
fr (GFQ)
0.100
10% of patients — above 5%, stable classification
nb (RQ)
0.578
Strongly separated
pattern (significant, stable, strong)
Complete statistical evidence. Significant, with a stable classification, and far from neutrality. This is what a finding looks like when it has earned its conclusion.

Standard reporting called all three of these results “significant” and stopped. The triplet separated them into discard, replicate, and act.

The framework

Complete statistical evidence is a triplet.

Reporting a p-value alone is partial evidence. It answers one question and is silent on two others that determine whether a finding is decision-ready.

p — significance
Compatible with no effect?
The only number most trials report.
fr — fragility
What proportion of patients would have to be reclassified to flip statistical significance?
A quotient, not a count: fr = 0.05 means 5% of patients. Higher fr = more stable classification. Below 5% may be considered fragile.
nb — robustness
How far is this from no effect at all?
Higher nb = farther from neutrality.

Both fr and nb are computed from published summary statistics alone — no raw data, no simulation, no distributional assumptions, no access to the trial team. Any reader can compute them from a table already printed in the paper. Full specification →

Why the global index

Restricting the search to one arm overstates classification stability.

The FI* may only toggle outcomes within one arm, holding that arm’s total fixed. The Global Fragility Index searches every cell-to-cell reallocation in the table, with neither margin fixed. Because a within-arm toggle is itself a reallocation, the FI is a constrained special case of that search — so GFI ≤ FI always. That is a mathematical result, not a tendency.

We recomputed both indices from the published 2×2 of 130 trials by exhaustive Fisher’s-exact search. The inequality held every time.

Trials recomputed
130
GFI never once exceeded FI
FI overstated
83
64% of trials — median 42% too high
Verdict changed
9
Called stable by FQ, fragile by GFQ

Nine trials where the answer changes, not just the number

In these nine, the conventional quotient lands at or above the 0.05 cutpoint — stable — while the global quotient lands below it — fragile. Same table, same threshold, opposite verdict. A reader trusting the arm-restricted index would file all nine as secure.

TrialNFIFQGFIGFQ
Taylor 2025Lancet Infect Dis 50 5 0.1000stable 2 0.0400fragile
Wang 2021J Hematol Oncol 42 4 0.0952stable 2 0.0476fragile
Lee 2025PLOS ONE 115 8 0.0696stable 4 0.0348fragile
Guo 2025Phytomedicine 212 15 0.0708stable 10 0.0472fragile
Guido 2019Anesth Analg 158 10 0.0633stable 6 0.0380fragile
Zucchelli 1992Kidney Int 69 4 0.0580stable 2 0.0290fragile
Chevalet 2000J Rheumatol 97 5 0.0515stable 3 0.0309fragile
Brede 2026Crit Care 179 10 0.0559stable 8 0.0447fragile
Norgaard-Pedersen 2025BMJ Open 80 4 0.0500stable 3 0.0375fragile

Both quotients use the same 0.05 cutpoint. FQ = FI / N. GFQ = GFI / N.

How wide the gap runs

FI GFI
Solomon 2006N = 8,912
FI 15
GFI 8  ·  FI overstates by 88%
Marx 2025N = 2,596
FI 18
GFI 10  ·  FI overstates by 80%
Altinbas 2014N = 1,585
FI 27
GFI 16  ·  FI overstates by 69%
Bonati 2009N = 413
FI 19
GFI 12  ·  FI overstates by 58%
Reid 2018N = 2,000
FI 34
GFI 22  ·  FI overstates by 55%
Ng 2025N = 980
FI 27
GFI 19  ·  FI overstates by 42%
Bonati 2019N = 1,530
FI 40
GFI 29  ·  FI overstates by 38%
Blackshear 2011N = 2,502
FI 2
GFI 1  ·  FI overstates by 100%

Across the 83 trials where the two disagreed, the FI ran a median of 42% too high, and as much as 150% too high. Every published fragility analysis that confines its search to a single arm is reporting a number more generous than the truth.

Indices recomputed from each published 2×2 by exhaustive search, not taken from any spreadsheet. Search depth capped at 40 reallocations, which excluded 2 trials of 132. Check any row yourself at fragilitymetrics.org.

Issued by the Internet Medical Association

The Complete Evidence Standard

Statistical evidence is reported completely when all three dimensions are reported together.

  1. Report p, fr, and nb for every primary outcome.
  2. Report loss to follow-up alongside fragility. When attrition exceeds the fragility index, say so plainly.
  3. Do not call a finding significant without stating whether that classification is stable and whether the effect is distinguishable from no effect.
  4. Report these values whether or not they favour the conclusion.

The metrics are model-free, computed from published summary statistics, and free to calculate. Compliance costs nothing but candour.

I endorse the Complete Evidence Standard.

Endorsement is a public statement of practice, not a membership fee. There is nothing to pay and nothing to renew.

Your email is used to confirm your endorsement and is never published or shared.

Signatories

2

researchers and clinicians have endorsed the Standard.

Jack
Thomas F. Heston, MD, MSc