Skip to content

Why I'm building in public: a research agenda for AI-native operational risk tooling

GRC systems became glorified registers while regulatory expectations leapt ahead. Twelve modules, two flagships, one platform: the gap closes only when the builders understand the domain.

I have spent eighteen years inside the risk functions of North American banks, most of it in the rooms where operational risk gets quantified, challenged, and defended: stress-testing pipelines, board risk committees, regulatory examinations. For the last several of those years I have watched a gap widen between what supervisors expect operational risk functions to demonstrate and what the tooling inside banks can actually do.

This essay is the founding document for what I intend to do about it: a twelve-module research agenda for AI-native operational risk tooling, built in public as working prototypes and blueprints. Two are already live. Here is the reasoning, because the reasoning matters more than the tools.

The stagnation nobody disputes

Corner any operational risk practitioner at a conference and ask what their GRC platform actually does for them. The honest answer is: it stores things. Assessments, once completed elsewhere, get recorded there. Incidents, once understood elsewhere, get logged there. Controls, once designed elsewhere, get listed there. The register was digitized twenty years ago and then the innovation stopped.

Meanwhile the discipline’s center of gravity moved. The UK’s operational resilience regime (PRA SS1/21 and its FCA counterpart, in force with the transition period complete since March 2025) reframed the question from “what risks have you recorded” to “which important business services would fail, how fast, and within what tolerance.” DORA, Regulation (EU) 2022/2554, has applied to EU financial entities since January 2025 and asks a version of the same question about digital resilience, in binding detail. OSFI’s revised Guideline E-21 renamed the discipline itself to operational risk and resilience, with full adherence expected by September 2026 and scenario testing of critical operations by 2027. The U.S. agencies said the quiet part in their 2020 interagency paper on strengthening operational resilience. And the Basel Committee’s revised Principles for the Sound Management of Operational Risk sit under all of it.

The common demand across those texts is not documentation. It is demonstration: map your services, set tolerances, test against severe but plausible scenarios, and produce evidence on request. A register cannot demonstrate anything. Most banks’ tooling is a register.

Why the gap persists

The economics are unglamorous but decisive. The work that would close the gap is deeply domain-specific: scenario generation with real challenge, dependency mapping that stays current, classification of loss events against taxonomies, regulatory change mapped to obligations. Vendors avoid it because the addressable market for “tooling that can survive an OCC examination of your scenario process” is too narrow and too demanding. Internal change budgets avoid it because it is second-line plumbing that never demos as well as a customer-facing feature.

The result: the largest banks build fragments in-house, the smallest take whatever their core provider bundles, and the mid-tier institutions, carrying nearly the full weight of the expectations above with a fraction of the build capacity, improvise with spreadsheets and heroics. I have watched skilled teams spend a quarter of their year assembling board reporting by hand. That is not a talent problem. It is a tooling vacuum.

What changed

Large language models are, bluntly, well matched to the substance of operational risk work in a way no previous technology has been. The daily material of the discipline is structured judgment expressed in language: reading regulatory text and extracting obligations, drafting scenario narratives and attacking their assumptions, classifying incident descriptions against event taxonomies, spotting the inconsistencies across three hundred RCSAs, converting a quarter’s data into prose a board can act on. These are exactly the tasks the models handle well, and exactly the tasks that never got tooling investment.

But I want to be precise about the claim, because our industry has been burned by technology evangelism before. The model is not the tool. A language model applied naively to risk management produces fluent artifacts that collapse under examination: scenario narratives with no coherent loss logic, classifications with no auditable rationale, summaries that a challenger would shred. What makes a tool is the scaffolding around the model: the workflow that enforces challenge, the audit trail that makes every output reconstructable, the access control that mirrors a real control environment, the domain judgment about what an examiner will actually probe. That scaffolding can only be designed by someone who has operated inside it. This is the entire thesis: AI closes the tooling gap when the builder understands the domain, and not otherwise.

The agenda

The research agenda is twelve modules: my working answer to the question “what would a complete, AI-native operational risk stack for a mid-tier bank look like.” Not twelve disconnected tools. Twelve components of one integrated platform, joined on a shared taxonomy and entity spine, so that a loss event, a control, a vendor, and an important business service mean the same thing everywhere they appear. I call the platform AEGIS. Two flagships are live as working prototypes.

DELPHI rebuilds scenario analysis as a continuous, challenged, auditable process: severe-but-plausible narratives with quantified loss ranges, a two-pass challenger workflow that institutionalizes effective challenge, and governance that mirrors real bank control expectations. ORBIT is an operational resilience workbench built to the shape the regimes converge on: important business services, impact tolerances, dependency mapping, and the scenario testing and evidence a supervisor can inspect. It maps what must not fail and proves it stays within tolerance under test.

Behind them, in research: KRI intelligence, tail-event modeling for stress loss, third-party dependency mapping, loss-event classification, an RCSA copilot, regulatory-change tracking, impact-tolerance calibration, control rationalization, external loss intelligence, and a comprehensive issues-management module that captures self-identified, audit, and regulatory issues in one place, runs them through remediation workflows, and aggregates exposure across every taxonomy node. Binding all twelve is a reporting and analytics layer that turns the underlying data into the board narrative, with every quantitative claim traceable to its source record, so the report cannot drift from the evidence. Each module corresponds to a specific gap I have personally hit. The sequencing follows the pain.

Why in public

Three reasons, in candor.

First, credibility flows from artifacts. The industry does not need another commentator asserting that AI will transform risk management. It needs working demonstrations that can be examined, challenged, and cited, along with a documented record of the design decisions behind them. Publishing the builds is the difference between an opinion and a contribution.

Second, challenge improves the work. I spent a career in functions whose entire value was effective challenge; I would be a poor student of my own discipline if I exempted my prototypes from it. Building in public invites the criticism that makes blueprints better, from the practitioners and supervisors best positioned to give it.

Third, honesty about what this is. These are prototypes and blueprints, not products. Nothing here is for sale. An independent research platform can say true things about vendors, regulators, and banks precisely because it is not selling to any of them. That independence is worth more than any revenue line it forecloses, and I intend to keep it.

What to expect

The essays here will document the rebuild of scenario analysis, of resilience testing, and of the second line’s tooling, with the specificity of someone who has to make the arguments hold. The Lab will grow as research builds mature into working prototypes. The OpRisk Signal will carry the curation and one essay per issue.

If you run an operational risk function and any of this matches or contradicts your experience, I want to hear it; that is what the contact page is for. The discipline is being rebuilt whether or not practitioners lead the rebuild. I think we should.