The second line doesn't need fewer people. It needs instruments.
Every AI conversation in risk starts with the wrong question: headcount. The second line's binding constraint has never been intelligence. It's attention, and attention is exactly what the new instruments return.
Every conversation about AI in risk management eventually arrives at the same question, usually asked with a careful glance around the room: how many people is this going to replace?
I understand why the question gets asked. I think it is the wrong question, and I think it reveals a misunderstanding of what actually constrains a second line, one worth correcting precisely, because the institutions that correct it first will spend the next five years setting the standard everyone else is examined against.
The real constraint
Walk the floor of any operational risk function and inventory what the skilled people are actually doing. A senior analyst, someone who can smell a bad control from across the room, is reconciling three spreadsheets of RCSA results because the business units used inconsistent scales. A scenario lead who has defended numbers in front of the Fed is reformatting last year’s workshop output for this year’s template. A regulatory affairs manager is reading a 200-page guideline to determine which of its paragraphs touch which of four hundred controls. A reporting team is assembling the quarterly board deck by hand, from nine sources, for the fortieth consecutive quarter.
None of these people lacks intelligence or judgment. What they lack is time to use it. The judgment-bearing hours of the second line are a small fraction of hours worked: the hours spent actually challenging a scenario, probing a control, or arguing with the first line about a risk that matters. The rest disappears into collation, reconciliation, formatting, and search. That is the binding constraint: not headcount, not intelligence. Attention.
So the replacement question is malformed. You cannot meaningfully “replace” a function that is already operating at a tenth of its judgment capacity. You can, however, equip it, and equipment is the right frame for what large language models actually offer this discipline.
Instruments, precisely
Risk management has always advanced by instrumentation. Loss distributions were not computable by hand; Monte Carlo made them routine. Nobody described the arrival of simulation as replacing quants; it turned quant judgment loose on questions that were previously unanswerable. The current generation of AI deserves the same framing, because what it does well maps with almost suspicious precision onto the second line’s attention sinks.
It reads regulatory text and maps it to obligations and controls: the 200-page guideline becomes a structured delta against your framework, produced in an afternoon and reviewed, rather than produced over a quarter and skimmed. It classifies incident narratives against event taxonomies with a written rationale for each call, which converts loss-data quality from a perpetual cleanup project into a reviewable pipeline. It reads three hundred RCSAs and surfaces the inconsistencies: the identical risk scored high in one unit and low in another with no difference in controls, the assessment untouched for six quarters, the control mapped to no risk anywhere. It drafts the first version of the scenario narrative from the bank’s own profile and, more valuably, drafts the attack on that narrative: what it assumes, what it ignores, where the loss logic is soft. It turns a quarter’s structured data into the first draft of board prose.
Notice what every item on that list has in common: each one produces an artifact for a human to challenge, not a decision for a human to obey. That is not a limitation to apologize for. That is the design.
The guardrails that make it defensible
I hold the enthusiasm with discipline, because a second line that adopts undisciplined tooling has traded one credibility problem for a worse one. Three requirements are non-negotiable, and they are the same requirements we have always imposed on anything that touches a regulated process.
Every output must be reconstructable. If a model classified an event or drafted a scenario, the record must show what went in, what came out, what a human changed, and who approved it: the audit trail as exhaust of the process, not a report written after the fact. Challenge must be structural, not aspirational: the workflow separates producers from challengers from approvers, with role-based access control doing the separating, exactly as a bank separates them everywhere else it is serious. And the models themselves belong inside the model risk framework: inventoried, validated to the standard their materiality warrants, monitored in use. U.S. readers have had the blueprint since SR 11-7; the only new work is applying it honestly to a class of models that produces prose instead of numbers.
None of this is exotic. It is the same scaffolding the second line already knows how to build. Which is why the practitioners best positioned to deploy AI in risk are not the ones who understand the models best; they are the ones who understand what an examiner will ask.
Where to start
Not every use case deserves the same risk appetite, and the sequencing matters more than the ambition. The right early candidates share three properties: the output is reviewed by a human before anything depends on it, the blast radius of an error is contained, and the payoff is asymmetric because the task is currently done badly or not at all.
By that test, start with classification and consistency: loss-event classification with rationale, RCSA coherence review, regulatory-change triage. Each is drowning in backlog, each produces reviewable artifacts, and each failure mode is visible on inspection. Scenario drafting and challenge come next: high value, and safe precisely because the entire process exists to challenge its own outputs. What comes last, deliberately: anything that feeds capital numbers without human mediation, and anything customer-facing. A second line that cannot show this kind of sequencing judgment about its own tooling has no business opining on the first line’s.
The multiplier math
Now run the arithmetic that the replacement question obscures. Take a ten-person operational risk team operating, honestly assessed, at ten or fifteen percent judgment capacity. Equip it with instruments that absorb the collation, the reconciliation, the first drafts, and the searching. If judgment hours merely triple, a conservative outcome in my estimation, you have not replaced anyone. You have conjured the equivalent of twenty additional senior practitioners out of the hours your existing ones were already being paid for. And they are your practitioners: people who know the bank, hold the context, and can defend the outputs, which no vendor and no model can do for them.
For mid-tier banks, this arithmetic is not an optimization; it is the only available answer to a structural squeeze. The regulatory expectations described across this site fall on them with nearly full weight; the headcount to meet those expectations by brute force does not exist and is not coming. The choice is not between the current team and a smaller one. It is between an equipped second line and an exhausted one.
The replacement question will keep being asked, and it will keep being boring. The equipment question is the urgent one: whether the judgment your institution already employs gets to spend its hours on the work that requires it. The instruments now exist. The banks that pick them up first, with the audit trails, the challenge workflows, and the model-risk discipline that make them defensible, will define what “adequately resourced second line” means for everyone else.