The Internet Medical Association

“Statistically significant” is not a verdict. It is one of three numbers.

Every trial you read reports the first number. Almost none report the other two — and without them you cannot tell a finding that will replicate from one that is already indistinguishable from nothing.

Below is a published trial. Judge it the way you normally would, then keep scrolling.

Exhibit one

A trial that changed a conversation

In a randomised comparison of carotid endarterectomy against stenting, surgical patients had roughly twice the rate of myocardial infarction. The difference was statistically significant.

Blackshear JL, et al. Circulation. 2011 · Myocardial infarction · N = 2,502
ArmMINo MIRate
Endarterectomy 28 1,212 2.26%
Stenting 14 1,248 1.11%
Fisher’s exact, two-sidedp = 0.029

Dimension two — fragility

One patient.

Move a single myocardial infarction from the surgical arm to the stenting arm — one reallocation, out of 2,502 patients — and the finding is no longer significant.

Endarterectomy

28 MI of 1,240

Stenting

14 MI of 1,262
After one global movep = 0.029

This is the Global Fragility Index

The GFI is the smallest number of cell-to-cell reallocations, anywhere in the table, that flips the significance classification. Neither row nor column totals are held fixed — the search is global.

Classic FI
2
Confined to one arm, one direction
GFI
1
The true global minimum
GFQ = GFI / N
0.0004
0.04% of patients reclassified. Below 5% → fragile

The classic fragility index reported 2 here because it may only toggle outcomes within the arm that had fewer events. Allowed to search the whole table, the answer is 1. The classic index was overstating how stable this classification is — and it does so systematically. We come back to that.

Dimension three — robustness

It was never far from nothing.

Fragility tells you the statistical classification is unstable. It does not tell you whether there was much of an effect to begin with. That is a separate question, and it has a separate answer.

RQ — distance from therapeutic neutrality, where neutrality means the two arms are indistinguishable.

0.011CREST
p — significance
0.029
Significant
fr — fragility (GFQ)
0.0004
0.04% of patients — fragile
nb — robustness (RQ)
0.011
At the neutrality boundary

The two arms differ by 1.1 percentage points — 2.26% against 1.11%. On the neutrality scale that is RQ = 0.011: the result sits almost exactly on the boundary where treatment and control become indistinguishable.

pattern (significant, fragile, weak)
This is the SFW pattern: significant, fragile, weak. A strong indicator that a “statistically significant” result should be viewed with skepticism — the classification flips on a single patient, and the effect it describes is barely distinguishable from zero. Nothing in the published report tells you either of those things. The p-value is identical to that of a trial you would be right to act on.

SFW applies to a claim that an effect exists. A claim that an effect is absent reads the same three dimensions differently: there, a high p-value with a stable classification and a result close to neutrality is what supports the null.

Exhibit two

Now the same p-value, from a different trial.

Tocilizumab for giant cell arteritis. Complete remission at twelve weeks. p = 0.030 — statistically indistinguishable from the trial you just dismantled.

Villiger PM, et al. The Lancet. 2016 · Complete remission at 12 weeks · N = 30
ArmRemissionNo remission
Tocilizumab173
Placebo46
p
0.030
Significant
fr (GFQ)
0.033
3.3% of patients — below 5%, fragile
nb (RQ)
0.400
Clearly separated

Both trials on the same neutrality scale:

0.011CREST 0.400Tocilizumab
pattern (significant, fragile, strong)
Not SFW — a real effect that needs replication. Also fragile: one move flips the classification. But 36 times farther from neutrality than the first trial. Same p-value, same fragility verdict — and the third dimension is the only thing that tells them apart.

Fragility flagged both classifications as unstable. Only nb distinguished a finding worth confirming from one worth discarding. No standard reporting requirement asks for it. No other framework computes it.

Exhibit two, continued

They followed the same patients for a year.

Same trial, same thirty people, same drug — complete remission measured at fifty-two weeks instead of twelve. Watch all three numbers move.

Villiger PM, et al. The Lancet. 2016 · Complete remission at 52 weeks · N = 30
ArmRemissionNo remission
Tocilizumab173
Placebo28
p
0.00097
31× smaller
fr (GFQ)
0.100
10% of patients — above 5%, stable classification
nb (RQ)
0.578
Strongly separated
pattern (significant, stable, strong)
Complete statistical evidence. Significant, with a stable classification, and far from neutrality. This is what a finding looks like when it has earned its conclusion.

Standard reporting called all three of these results “significant” and stopped. The triplet separated them into discard, replicate, and act.

The framework

Complete statistical evidence is a triplet.

Reporting a p-value alone is partial evidence. It answers one question and is silent on two others that determine whether a finding is decision-ready.

p — significance
Compatible with no effect?
The only number most trials report.
fr — fragility
What proportion of patients would have to be reclassified to flip statistical significance?
A quotient, not a count: fr = 0.05 means 5% of patients. Higher fr = more stable classification. Below 5% may be considered fragile.
nb — robustness
How far is this from no effect at all?
Higher nb = farther from neutrality.

Both fr and nb are computed from published summary statistics alone — no raw data, no simulation, no distributional assumptions, no access to the trial team. Any reader can compute them from a table already printed in the paper. Full specification →

Why the global index

The classic fragility index overstates classification stability.

The conventional index may only toggle outcomes within one arm, in one direction. The Global Fragility Index searches every cell-to-cell reallocation in the table. Because the classic procedure is a constrained special case, GFI ≤ FI always — and the gap is routinely large.

Classic FI GFI
Desai 2016Eur Heart J
FI 599
GFI 419  ·  FI overstates by 43%
Garvey 2025NEJM
FI 409
GFI 383  ·  FI overstates by 7%
Hastings 2025Ann Intern Med
FI 239
GFI 142  ·  FI overstates by 68%
Black 2007NEJM
FI 199
GFI 107  ·  FI overstates by 86%
Cummings 2009NEJM
FI 140
GFI 76  ·  FI overstates by 84%
Saag 2017NEJM
FI 76
GFI 44  ·  FI overstates by 73%
Solomon 2006Circulation
FI 15
GFI 9  ·  FI overstates by 67%
Blackshear 2011Circulation
FI 2
GFI 1  ·  FI overstates by 100%

Every published fragility analysis using the conventional index is reporting a number more generous than the truth.

Issued by the Internet Medical Association

The Complete Evidence Standard

Statistical evidence is reported completely when all three dimensions are reported together.

  1. Report p, fr, and nb for every primary outcome.
  2. Report loss to follow-up alongside fragility. When attrition exceeds the fragility index, say so plainly.
  3. Do not call a finding significant without stating whether that classification is stable and whether the effect is distinguishable from no effect.
  4. Report these values whether or not they favour the conclusion.

The metrics are model-free, computed from published summary statistics, and free to calculate. Compliance costs nothing but candour.

I endorse the Complete Evidence Standard.

Endorsement is a public statement of practice, not a membership fee. There is nothing to pay and nothing to renew.

Your email is used to confirm your endorsement and is never published or shared.

Signatories

1

researchers and clinicians have endorsed the Standard.

Thomas F. Heston, MD, MSc