A worker made of software: what an AI agent actually is
The word “agent” has moved from AI conferences into board papers, usually without a definition attached. The plain version: an agent is software you hand a task and a written procedure. It reads the evidence, operates the systems it is given access to, and finishes the casework, recording what it did as it goes. The right comparison is a new hire rather than a new system. And that worker suits non-financial risk work, because the work is recurring, evidence-based, and judged on whether it can be shown.
This piece assumes you know what a control test is and have no particular reason to know what a “large language model” is. That is the right starting point. The vocabulary around AI agents is new, but the management questions they raise are ones you have handled before, and by the end I hope to convince you that you already know most of what you need to govern one.
RPA repeats steps, a model computes, an agent reads the case
An agent is not robotic process automation, and it is not a statistical or rules-based model.
Robotic process automation repeats recorded steps. If the report always arrives in the same format and the decision is always the same lookup, RPA will perform that task a thousand times without complaint, then stop dead the moment the format changes. Models compute: a score, a threshold breach. Both are useful inside those limits. Neither, on its own, reads a contract and judges it against the approval behind it.
An agent does that, because it is software built around a large language model, the same technology behind the public chat assistants. The difference is that it acts on what it reads. Given a task and a written procedure, it works through documents the way a person does, reading them for what they say: a contract, a policy, an approval email. It logs into and operates whatever systems it is given credentials for, a case-management tool, a GRC platform, a document archive, using them as its tools. Those systems stay where they are. It carries a piece of casework from instruction to conclusion without a person steering each step. And it records what it did and why, step by step, as it goes.
It works through a queue of cases on its own. Anyone whose only encounter with this technology is a chat window should set that picture aside and think instead of a worker.

The same test, run on six thousand contracts instead of twenty-five
Take a control of the kind second-line teams test routinely: new customer contracts must be activated on the terms that were approved.
Suppose the procedure reads the way most testers would recognise. Pull a sample, say twenty-five contracts out of six thousand. Retrieve each approval, compare it with the activated record, apply the materiality rules to any differences, write the workpaper. An agent is onboarded onto that procedure, whatever yours actually says, and follows it: retrieves the approval, compares it with the activated record, applies the same materiality rules, decides pass or fail, and hands back a workpaper recording every retrieval, comparison and decision it made.
Two things differ. The agent can work all six thousand. And every case is worked against the same versioned procedure, with the reasoning recorded case by case, so variation between cases is visible rather than assumed. Whoever signs the conclusion is then relying on an executed test for every contract in the population, instead of an extrapolation from twenty-five to the five thousand nine hundred and seventy-five nobody looked at.

The most cautious function in the bank is where this fits best
Risk work is an unusually good fit for this worker, which sounds odd for the most cautious function in the building. Three things about it explain why.
- A lot of it is recurring, evidence-based casework. Control testing, due-diligence file reviews, incident write-ups, monitoring follow-ups: gather the evidence, apply a defined standard, record a conclusion, repeat next quarter. That is the kind of work an agent does well, and it tends to arrive in populations of near-identical cases.
- Coverage in this work is bounded by reading time. Sample sizes are set for several reasons, and how much one person can read in the time available is usually one of them. The same goes for how often a control gets tested. An agent takes that particular constraint out of the arithmetic. What your methodology should then say is your call.
- The work is judged on whether it can be shown. An agent documents as it works, so the workpaper is a by-product of the test rather than a task after it.
Across European banking supervision, breaches of internal governance were the largest single category of sanctioning proceedings in both 2023 and 2024, at 57 and 55 per cent.¹ Those are findings about how a bank runs itself and about the record behind it.
In Sweden, under its own national regulator, Finansinspektionen issued Klarna a remark and a SEK 500 million sanction in December 2024. The regulator did not claim that any laundering had actually occurred. The deficiencies were gaps in the framework itself: a general risk assessment that had not assessed how the bank’s own products could be misused, and due-diligence procedures that did not cover every situation requiring them.²
Be careful about what a case like that does and does not argue. No volume of control testing writes a missing risk assessment; that document is the bank’s to produce. What testing supplies is the other half of the same supervisory demand, the record showing that the controls you do have operate as designed, case after case, quarter after quarter. That half is still largely made by hand.
Accountability moves onto the process you govern and the supplier you manage
Paying someone else to do work you remain answerable for is nothing new. Banks already pay employees, consultants and service providers to do work that someone else signs for. The difference that matters with an agent is that it cannot bear accountability. It has no mandate and nothing to lose. So the accountability relocates onto the humans and the process that govern it.
One way to handle that is to have a person re-check every case the agent touched. That quietly rebuilds the capacity ceiling you were trying to escape, and adds the cost of an agent on top. The alternative is that a named person answers for a governed process: a written, versioned procedure; automated checks built independently of the agent; an adjudicated set of cases kept as the answer key; clear escalation for anything ambiguous. That is how anyone stands behind a team they do not shadow, by governing how the work is done.
If you can govern a testing team, much of that machinery is already in the building. The management questions are the familiar ones: whose procedure, whose review, whose signature. But an agent working inside your systems is also a supplier arrangement, so a second set of questions applies, and those are the ones you already ask of any outsourced service: what it has access to, what happens on exit, what your third-party register says about who is doing the work. Neither set of questions is new. Assembling both around the same worker is.

You buy completed work, priced like a person and governed like one
How you buy an agent affects how much control you keep.
We offer a service called Agent Workforce, by Digital Workforce. You buy completed work: control tests with the full work shown, delivered by an agent with an operating team behind it, priced like a person’s time rather than a licence seat, and governed the way you would govern the person doing that job. You keep the written procedures and the records, and you can walk away. The first job is usually control testing, because it is bounded, recurring and easy to judge against the human team doing it today. In time the same worker can take on the neighbouring casework: due-diligence file reviews, incident write-ups, monitoring follow-ups.
It cannot sign, and it cannot be its own check
A primer that only lists strengths is a brochure, so here is what an agent cannot do.
An agent cannot sign. Relocated accountability is a permanent condition of using one, and the technology maturing will not change it. There will always be a named human answering for the mandate.
Some cases have to go to a person, and always should: the ambiguous ones, and the first of a kind. The aim is that human effort tracks those cases and not the size of the book. It never falls to zero.
An agent cannot be its own check, and a second model is a smaller fix than it sounds: different models fail on the same items far more often than chance, and the overlap gets worse as the models get better. A check is worth something when it holds what the agent did not, a source record to reconcile against, a rule that computes instead of judging, a set of cases already adjudicated and kept as the answer key. That is close to how automated controls have long been assured. Once the logic is tested and shown to work, assurance moves off the output and onto the environment: what can change the logic, who can change it, how you would know. What is left for a person is the criteria and the exceptions, not a sample of the cases. Building that layer is real engineering work, and any provider who waves it away should worry you.

No agent is ready on day one, and no engagement should assume the bank is either. The sensible start is an agent replicating a process you already run, alongside the people who run it, until its output has been benchmarked and the acceptance criteria agreed. You build the readiness during the engagement.
You govern this the way you govern a new hire
Take away the vocabulary and this is a familiar management problem. You would not install a new employee. You hire one: give them written procedures, check their early work against people you trust, widen their responsibilities as the evidence comes in, and hold a named manager accountable for how they perform. That is the right frame for an agent, and it is a discipline your institution runs every time it hires.
The questions that follow this one each deserve their own treatment: why the enforcement climate makes this urgent, what gets better beyond cost, and how to source an agent without losing control of the work. I have written about those separately. What I wanted to do here was make one word mean something specific. If it left you with questions, I am glad to compare notes.
Sources
- 1. ECB Banking Supervision, SSM sanctioning reports for 2023 and 2024: breaches of internal governance as a share of sanctioning proceedings across European banking supervision (57% of 371 proceedings in 2023; 55% of 441 in 2024). https://www.bankingsupervision.europa.eu/press/other-publications/publications/sanctioning-report/html/ssm.sr2024~fa110879af.en.html https://www.bankingsupervision.europa.eu/press/other-publications/publications/sanctioning-report/html/ssm.sr2025.en.html
- 2. Finansinspektionen, “Klarna får en anmärkning och sanktionsavgift,” decision of 11 December 2024, FI dnr 22-11505 (fi.se decision page and decision PDF). https://www.fi.se/sv/publicerat/sanktioner/finansiella-foretag/2024/klarna-far-en-anmarkning-och-sanktionsavgift/ https://www.fi.se/contentassets/9dc40cb5565c478f9bd5775a02449682/klarna-bank-ab-beslut.pdf
About the author
Jonatan Larsen is a risk practitioner turned builder at Digital Workforce, where he leads Agent Workforce, the firm’s agentic risk and compliance capacity for financial institutions.

Share this post