Reference
Two dimensions, and the indices that measure them. Every worked value below comes from the two trials on this site: CREST and Jover.
Work in progress. This page is being built out. The canonical, versioned definitions live in the specification at fragilitymetrics.org.
fr · a dimension, not an index
The instability of the p-value to cross the alpha threshold of 0.05 with small perturbations in the data.
Fragility asks how securely a result holds its side of the threshold. It is direction-agnostic: the same measure applies whether the result starts significant and would be lost, or starts non-significant and would be gained. A fragile classification is not a wrong one — it is one that a handful of reclassified patients would reverse.
Measured by an index (a count of patients) or a quotient (that count as a proportion of N). The quotient is what gets compared across trials.
nb · a dimension, not an index
The geometric distance from therapeutic neutrality.
Robustness asks how far the observed result sits from no effect. It is a property of the point estimate alone — it does not depend on sample size, precision, or any significance test. That separation is deliberate: it is what lets the triplet tell a large-but-imprecise effect apart from a genuinely null one.
Published bands: weak below 0.075, moderate from 0.075 to 0.227, strong at 0.227 and above.
Heston TF. The Neutrality Boundary Framework: Quantifying Statistical Robustness Geometrically. arXiv:2511.00982. 2025.
FI · count · fragility
The smallest number of patients whose outcome must be reclassified, within a single arm, to move the p-value across alpha.
The original fragility measure. One arm is selected and outcomes are toggled inside it until the classification flips; the arm total stays fixed. Because the search is confined to one arm, the FI is a constrained special case of the global search below.
This site uses the Heston FI: the arm with fewer events is selected, ties broken to the smaller arm, and toggling then runs in whichever direction crosses alpha. Full rule.
Walsh M, Srinathan SK, McAuley DF, et al. The statistical significance of randomized controlled trial results is frequently fragile: a case for a Fragility Index. J Clin Epidemiol. 2014;67(6):622–628.
FQ · quotient · fragility
The Fragility Index expressed as a proportion of the trial’s total sample.
FQ = FI / N
A raw count grows with trial size, so two FIs from trials of different size cannot be compared. Dividing by N removes that dependence and turns the count into a proportion of patients — which is comparable.
Ahmed W, Fowler RA, McCredie VA. Does Sample Size Matter When Interpreting the Fragility Index? Crit Care Med. 2016;44(11):e1142–e1143.
GFI · count · fragility
The smallest number of cell-to-cell reallocations, anywhere in the table, that moves the p-value across alpha.
The search is global: a patient may be moved from any cell to any other, and neither the row nor the column margins are held fixed. Only N stays constant. Because a within-arm toggle is itself a reallocation, every move available to the FI is also available here — so GFI ≤ FI always. That is a mathematical result, confirmed empirically, not a tendency.
Both indices count moves to the same boundary, but they search different move-sets, so an equal count can still name a different toggle and a different resulting p-value.
Heston TF. Fragility Metrics Toolkit. Zenodo. doi:10.5281/zenodo.17254763
GFQ · quotient · fragility
The Global Fragility Index expressed as a proportion of the trial’s total sample.
GFQ = GFI / N
GFQ is the value the framework reports as fr. Read it as a proportion, not a count: it answers what proportion of patients would have to be reclassified before the classification changes. Below 0.05 the classification is called fragile — the same cutpoint as the p-value it mirrors.
Heston TF. Fragility Metrics Toolkit. Zenodo. doi:10.5281/zenodo.17254763
NDI · count · robustness
The smallest number of paired, fixed-margin moves that brings the table to the reachable point closest to therapeutic neutrality.
NDI = round(|ad − bc| / N) = round(N · RQ / 4)
A paired move relocates two patients, one in each arm, so both margins stay fixed. Each such move shifts the cross-product difference by exactly N, which is why reachable values are spaced N apart. NDI invokes no significance test at all — it depends only on ad − bc. Its target is RR = 1, not alpha.
NDI is a count, so it grows with trial size and is not comparable between trials. CREST’s NDI of 7 exceeds Jover’s 3, yet CREST is weak and Jover is strong. Use RQ to compare trials; use NDI to state one trial’s distance in patients. Worked step by step.
Heston TF. Fragility Metrics Toolkit. Zenodo. doi:10.5281/zenodo.17254763
RQ · quotient · robustness
The normalised geometric distance of a 2×2 table from therapeutic neutrality, on a bounded 0–1 scale.
RQ = |ad − bc| / (N² / 4)
RQ is the value the framework reports as nb for a 2×2 table, and it is the measure to use when comparing trials. It is scale-invariant: multiplying every cell by a constant leaves it unchanged. Neutrality is ad = bc, which is RR = 1, so RQ = 0 means the table sits exactly on the neutrality boundary.
RQ and NDI express the same distance in different units: RQ as a bounded proportion, NDI as an integer count of paired moves. RQ is not NDI / N.
Heston TF. The Neutrality Boundary Framework: Quantifying Statistical Robustness Geometrically. arXiv:2511.00982. 2025. · Heston TF. Fragility Metrics Toolkit. Zenodo. doi:10.5281/zenodo.17254763
Compute any of these free at fragilitymetrics.org. Spotted an error? Tell us.
Endorsement is a public statement of practice, not a membership fee. There is nothing to pay and nothing to renew.
Signatories
1
researcher or clinician has endorsed the Standard.