You Cannot Depose a Model

Sean BaenenThe agentic coworker will be right most of the time. The trouble is the day it is wrong, and the year after that, when someone asks it to explain itself.

Advisor Perspectives welcomes guest contributions. The views presented here do not necessarily represent those of Advisor Perspectives.

A year after the fact, an examiner conducting your audit asks a straightforward question. On the first of the month, when your system approved that exception, what exactly did it see? Which client record, which version of the fee schedule, which permission, and which sentence of which policy did it rely on to make that judgment?

In the world we grew up in, a person answered because a person could be found, seated, and asked to walk through the file. Assuming competence, that person kept notes or at least had a memory that someone else could jog.

What if your agent kept neither? The basis for the decision cannot be found, cannot be asked about, and the agent remembers nothing past the moment it acted. Even in retrospect, what it did was probably right. However, you cannot prove it and, in this industry, the phrase “I cannot prove it” is a sentence with consequences.

The Scorecard Measures the Wrong Day

Many firms in the wealth management industry are piloting agentic AI right now, and most use similar scorecards. Is it accurate? How often does it get the answer right? What is the error rate against a person doing the same task?

Those are reasonable questions, but they are the wrong ones to lead with because they measure the machine on the day it performs and say nothing about the day it has to account for its actions. The performance is graded that afternoon in your own office, while the accounting is graded a year later by an examiner holding the file with no reason to take your word for it.

The better question, which I rarely see on the scorecard, is whether you can prove what the agent saw and why it acted, and whether you can reconstruct that a year later for someone who is not inclined to take your word for it.

Capability is what the vendor demos. Accountability is what the examiner opens the file to find. In a regulated business, those are not the same test, and your firm can ace the first and fail the second badly enough to wish it had never turned the thing on.

When you are in the business of advice, you make judgment calls. By definition, a person who makes a judgment call leaves a trail almost by accident. There is an email, a meeting note, a colleague who remembers the conversation, or a calendar that puts everyone in a room. The trail is thin and human, but it has saved a thousand firms in a thousand exams. An agent leaves nothing like it unless you decided, on purpose, measured against your standard and before you deployed, that it would.