Splicing the tail: the severity modeling choices that actually matter
The body of the severity distribution is easy; the tail is the entire game. A walk through the modeling choices that actually move the answer: splicing, thresholds, and the copula default nobody questions.
Operational loss severity is where quantification programs go to die. The body of the distribution is easy: thousands of small losses, any reasonable parametric family fits, nobody’s capital number moves. The tail is the entire game: three or four observations, a choice of distributional family, and a capital estimate that swings by multiples depending on decisions most documentation never surfaces. This note walks through the choices that actually move the answer, in the order they bite.
1. Why moments lie to you
Heavy-tailed severity breaks the intuitions that carry analysts through credit and market risk. For a distribution with tail index near or below one, the sample mean is dominated by the largest observation you happen to have; the sample variance is close to meaningless; and “average loss given event” is a number that will be revised violently the next time a large event lands. Any methodology that fits tail parameters by matching moments is fitting noise. Quantile-based and likelihood-based estimation on the tail region is not a stylistic preference; it is the difference between estimating the distribution and memorializing your sample.
2. The spliced architecture
The standard resolution, and the right one, is to stop asking one distribution to describe two regimes. Below a threshold u, fit the body empirically or with a lognormal/Weibull, from the institution’s own loss history, where data is plentiful and idiosyncratic experience is genuinely informative. Above u, switch to the generalized Pareto distribution, whose role is not convenience but theorem: the Pickands–Balkema–de Haan result establishes the GPD as the limiting distribution of threshold exceedances for essentially every distribution of practical interest. The spliced density is stitched at u with continuity constraints, and the tail is parameterized by shape (ξ) and scale (β), with ξ carrying almost all of the capital sensitivity.
Two practical consequences follow. First, the institution’s own data is usually sufficient for the body and almost never sufficient for the tail, which is the principled argument for informing tail parameters with industry or consortium data: pooling where the theory says shape is shared, staying local where experience is genuinely yours. Second, every downstream number should be reported with its ξ sensitivity, because a ±0.1 movement in shape is both plausible estimation error and a material capital delta.
3. Threshold selection: the decision hiding in plain sight
The threshold u is a modeling decision dressed as a technicality. Set it too low and the GPD is contaminated by body behavior it was never meant to describe; too high and you are estimating two parameters from five points. The classical diagnostics (mean-excess plots that should go linear above a valid threshold, and parameter-stability plots across candidate thresholds) are genuinely useful and genuinely subjective, which is why the honest practice is not “we chose the optimal threshold” but “here is the capital number across the defensible threshold range.” If that range moves the answer materially, that fact is the finding, and it belongs in front of the risk committee, not buried in an appendix. Hill-plot instability on operational loss data is normal; anyone showing you a perfectly stable tail estimator is showing you a smoothing choice.
4. Dependency: where the aggregate distribution is really decided
Severity choices set the marginal distributions; the aggregate tail is set by how event categories move together. The industry default (model categories independently and add capital) is not conservative, whatever its reputation: independence assumptions understate exactly the joint-stress states capital exists for.
Copulas make the dependency choice explicit, and the choice that matters is tail behavior, not correlation level. A Gaussian copula, whatever correlation you feed it, has zero tail dependence: in the limit, it asserts that extreme outcomes never coincide, which is a strong empirical claim smuggled in as a software default. A t-copula with the same correlation matrix and finite degrees of freedom asserts the opposite: joint extremes co-occur with positive probability, with the degrees-of-freedom parameter governing how strongly. On identical marginals, that single choice routinely moves high quantiles of the aggregate distribution more than any plausible re-estimation of the correlation matrix. The uncomfortable truth is that dependency parameters are the least identifiable objects in the whole architecture, which is an argument for scenario-based stress of the copula assumptions, never for defaulting to the assumption that happens to minimize capital.
5. What good looks like
A defensible severity stack, compressed to a checklist: spliced body/tail with the splice point diagnosed and sensitivity-tested, not asserted. Tail shape estimated by likelihood on exceedances, informed by pooled data, reported with confidence bounds. Every aggregate number accompanied by its ξ-, threshold- and copula-sensitivity. Parameter provenance flagged: estimated, pooled, or elicited. Seeded, reproducible simulation with convergence evidence at the quantiles that matter, which for operational risk means far past the 99th percentile, where Monte Carlo error is largest exactly where attention is highest.
None of this requires vendor tooling. It requires taking seriously that in heavy-tailed territory, the model is the assumptions, and assumptions govern best when they are visible.
The reference implementation of this architecture (spliced severity, explicit copulas, full known-answer test suite) is published as TEMPEST under a PolyForm Noncommercial license, demonstrated on the fictional Meridian Bancorp. For the roadmap, get in touch.