The agentic coworker will be right most of the time. The trouble is the day it is wrong, and the year after that, when someone asks it to explain itself.
Advisor Perspectives welcomes guest contributions. The views presented here do not necessarily represent those of Advisor Perspectives.
A year after the fact, an examiner conducting your audit asks a straightforward question. On the first of the month, when your system approved that exception, what exactly did it see? Which client record, which version of the fee schedule, which permission, and which sentence of which policy did it rely on to make that judgment?
In the world we grew up in, a person answered because a person could be found, seated, and asked to walk through the file. Assuming competence, that person kept notes or at least had a memory that someone else could jog.
What if your agent kept neither? The basis for the decision cannot be found, cannot be asked about, and the agent remembers nothing past the moment it acted. Even in retrospect, what it did was probably right. However, you cannot prove it and, in this industry, the phrase “I cannot prove it” is a sentence with consequences.
The Scorecard Measures the Wrong Day
Many firms in the wealth management industry are piloting agentic AI right now, and most use similar scorecards. Is it accurate? How often does it get the answer right? What is the error rate against a person doing the same task?
Those are reasonable questions, but they are the wrong ones to lead with because they measure the machine on the day it performs and say nothing about the day it has to account for its actions. The performance is graded that afternoon in your own office, while the accounting is graded a year later by an examiner holding the file with no reason to take your word for it.
The better question, which I rarely see on the scorecard, is whether you can prove what the agent saw and why it acted, and whether you can reconstruct that a year later for someone who is not inclined to take your word for it.
Capability is what the vendor demos. Accountability is what the examiner opens the file to find. In a regulated business, those are not the same test, and your firm can ace the first and fail the second badly enough to wish it had never turned the thing on.
When you are in the business of advice, you make judgment calls. By definition, a person who makes a judgment call leaves a trail almost by accident. There is an email, a meeting note, a colleague who remembers the conversation, or a calendar that puts everyone in a room. The trail is thin and human, but it has saved a thousand firms in a thousand exams. An agent leaves nothing like it unless you decided, on purpose, measured against your standard and before you deployed, that it would.
The Machine Forgets by Default
An agent is a witness who cannot be cross-examined. A person remembers whether they want to or not, which is why we have built an entire industry of substances and services to help us forget stuff.
A machine is the reverse. It remembers nothing unless you tell it to. It takes no notes, it cannot be recalled to the stand, and it has zero innate memory of the file it worked on an hour ago, much less a year ago.
It acted on a bundle of context, some client data, some policy, some permission, some definition of what a given field meant that particular afternoon, and then it let all of it go. Ask the agent later what it saw, and it will happily assemble a plausible answer that may or may not be what actually governed the decision. If you have ever been in an SEC, FINRA, or IRS audit, you know that is worse than saying nothing at all.
This is the odd asymmetry of AI. The entire appeal of the agent is that it carries no baggage, holds no grudge, is not looking at its watch, and starts every task clean. That same blankness is a liability the moment anyone downstream needs to know what happened. The human layer you are replacing was doing two jobs: the work itself and remembering the work itself.
A technologist will object that agentic systems log constantly, and they do. A log, though, is a quality control (QC) function built for debugging that is usually held by the vendor. It captures what the system did rather than what governed the decision: which version of the fee schedule, which sentence of which policy, or which permission on that particular afternoon. A log is not a record.
The Record Was Always the Product
All leaders of wealth management firms know that, under the Advisers Act, a registered firm must preserve its books and records for at least five years measured from the end of the fiscal year in which the last entry was made. Firm records must be kept in an easily accessible place for that whole span, in the firm’s own office for the first two years, and produced as a true and complete copy on demand.
For decades, that meant trade tickets, advertisements, and the communications that led to a recommendation. The rule never cared whether a person or a machine made the call, only whether you could show what the decision rested on.
Point an agent at a client recommendation and the record expands. What you now have to reproduce is not just the output, but the context the machine stood on when it acted: the data, the permissions, and the rule as written that day, preserved for five years in an easily accessible place.
That is the oldest recordkeeping standard in the business meeting the newest technology in the building. Whether or not your AI pilot meets it has everything to do with whether or not anyone wrote it into the requirements.
If you think regulators aren't paying attention and that everyone will get a pass early on, I would caution you to think again. In March 2024, the SEC settled its first cases charging advisors over their AI claims, and the tell was the method. Examiners pressed the firms to substantiate what their AI actually did, and the SEC found they did not have the AI and machine-learning capabilities they had advertised.
Two advisors paid $400,000 between them, and a priceless reminder followed. When the SEC asks what the machine did, the answer comes from your records or it comes from your confession. There is no door three.
The Vendors Are Remarkably Quiet About This
This is not a surprise when you consider who is across the table selling you the agent. The platform has every reason to demo capability, and its deck is full of accuracy, throughput, and hours saved.
The demo is where every vendor is all hat and no cattle, because capability on a stage costs nothing. That pitch probably isn't as headline-driven on point-in-time reconstruction, on who holds the log, or on whether the record it generates would survive contact with an auditor who wasn't in the room and isn't feeling generous.
Many vendors selling you an agent, particularly those not native to wealth management, run a version of the same play: They show you a capable machine and let you assume capable and deployable are the same word. In a supervised business, they are not remotely the same. The deployable unit is not a capable agent. It is an auditable one.
Auditability is not something you buy bundled, because the record must reflect your data, your permissions, your rules, and your unique means of production, all composed and kept by you. The vendor can sell you the agent. Whether that agent can be deposed is entirely up to, and on, you.
The Folder They Closed Without a Question
Early in my career, I helped launch a financial services firm that was truly novel for its moment, regulated by both FINRA and the SEC, and moving faster than either was used to seeing. Someone recommended an old-school chief compliance officer to us. He was, candidly, an outlier in a Silicon Valley start-up machine that measured everything in velocity. He kept meticulous, seemingly redundant memos and generally bugged everyone. Our investors insisted he be part of our executive committee, and he documented all of our management conversations. It was never invasive, and we understood it was the job, but I will be honest: At the time it felt like a bureaucratic tax on people in a hurry.
Then, about 18 months after launch, FINRA came in. The examiners had real questions about the nature of our disclosures and how we were marketing our securities because we were pushing the envelope, running television ads in a corner of the industry that did not usually run television ads.
Those are the exams that end companies. The examiners sat in our offices, and months of decisions that had felt obvious in the moment were suddenly being read back to us cold by a group of people with no particular reason to be generous.
The FINRA team opened the file, read the compliance officer’s memos, and closed the folder. No follow-up. Not one. The stickler we had quietly resented for harshing our mellow was the reason there was nothing to chase, because he had written down what we saw, what we did, and why we did it while we were doing it, not a year later from memory.
He was our court reporter, and he saved us inordinate time, money, and maybe more. I did not fully understand what he had done for us until I watched that folder close. Replacing a person like that without replacing the recording layer they provide is operational suicide, and it is exactly the kind of trade a firm makes the day it puts an agent in the seat and forgets to build the reporter back in.
Build the Court Reporter First
When leading a firm through a change like this, the instinct is to treat governance as the compliance tax you bolt on after the smart part works. In many AI projects underway right now, I can assure you governance is viewed as the brake that slows the pilot down.
Those deploying AI effectively started in part from a place that recognizes that, in a regulated firm, the record is not what slows the agent down. In reality, it's the only thing that lets you deploy the agent at all. Those who see it this way will get a gift the laggards do not.
Build the court reporter in from the first day, make the agent keep a real, reconstructable account of what it saw and why it acted, and you will have done something the human era never quite managed. You will have put the judgment on the record.
The person you are replacing was never this accountable, because that person's reasoning lived in someone’s head and died with the meeting. If we do this right, the machine goes from the biggest risk in the building to the best-documented decision-maker your firm ever hired.
I am in and around a lot of different businesses, and the narrative is shifting away from pure speed and efficiency. In what is a good sign, many discussions around automation generally and AI specifically are increasingly moving to questions about compliance and liability. Basically, people are asking whether the agentic coworker they just hired is smart enough to trust.
In my view, those who will win safely should ask instead whether that agent's job description is documented enough to hire against, build for that first, and then find out those two questions were asking the same thing. The deployable agent is the one that leaves a record. However capable the rest may be, in a regulated business, you cannot deploy what you cannot prove.
Sean Baenen is co-founder of SMART Growth Partners. He was previously President of NorthRock, which grew both assets and revenue at a 30% CAGR during his nearly six-year tenure. He writes Built to Compound, a newsletter on growth and capital strategy for advisory firms, as well as House on Fire which offers leaders of growing firms principles for holding the line through the chaos.
A message from Advisor Perspectives and VettaFi: Discover something new! Click here to register for our upcoming webcasts.
More Tax Planning Topics >