A supervision rule built to catch promises of investment returns will flag the word "guarantee." It will also flag an adviser telling a colleague "I guarantee this will be on your desk by Friday." The word that makes the rule effective is the same word that makes it noisy, and every noisy flag makes it a little easier to stop reading closely.
This is the core tension in communications monitoring: rules broad enough to catch real risk are broad enough to catch much more. Firms have lived with that trade-off for years, but it's becoming harder to ignore.
The SEC's fiscal year 2026 exam priorities are explicit that AI-driven supervision tools will be judged on more than whether they exist. Examiners want to understand the logic behind a flag, not just confirm that flagging happens. A system that can't explain why it raised an alert, or why it didn't, is a harder system to defend under that scrutiny.
This guide breaks down why false positives happen, what they cost a compliance program, and how to bring the queue back under control.
Most false positives trace back to one design choice: the system detects individual words, but has no way to understand how those words relate to each other.
Keyword and lexicon-based detection
Traditional monitoring tools work by matching a list of terms against every message that passes through. "Guarantee." "Can't lose." "Off the books." The list catches genuine red flags, but it has no way to tell the difference between an adviser promising returns and an adviser promising a document by Friday. The word is the same though the meaning is not, and the tool only sees the word.
Rules built defensively, not precisely
When a compliance team misses something a lexicon should have caught, the instinct is to add more terms and widen the net. Over time, rule sets grow to cover every conceivable phrasing of every conceivable risk, because narrowing a rule feels like taking on liability, while broadening it feels safe. The result is a system tuned to avoid missing anything, at the cost of flagging almost everything.
No context awareness
A word means something different depending on who sent it, who received it, and what came before it in the conversation. "Guarantee" from a compliance officer confirming a filing deadline is not the same risk as "guarantee" from a registered rep talking to a retail client about returns. Lexicon-based tools don't see that difference because they aren't built to look at the relationship between words, only the presence of one.
The same rigidity that flags "guarantee" in a harmless sentence also lets genuine risk through undetected. Someone who avoids the exact trigger word, phrasing a promise differently or splitting it across two messages, passes a lexicon-based check without issue. A tool built to catch words rather than meaning fails in both directions.
The clearest cost of a false positive is the time it takes to rule out. Someone must read the flagged message, understand its context, and confirm there's no real issue, then move to the next one. Multiplied across a compliance program, that adds up to a lot of hours spent on messages that were never a problem.
MirrorWeb's 2025 survey of 200 senior compliance decision-makers puts a number on part of this picture. Looking specifically at mobile communications, 27% of teams reported false positives at least once a day, and a further 51% at least once a week. Teams spent an average of 308 hours a year on mobile communications supervision alone, and for 16% of firms that figure passed 500 hours. Firms estimated an average annual cost of $232,457 tied to this inefficiency, with 13% putting the figure above half a million dollars.
Those numbers only cover mobile. Most firms are also monitoring email, Teams, Slack, and other channels, so the true scale of time and cost lost to false positives across a full communications program is certainly higher.
The less visible cost
The bigger risk isn't the time spent, it's what that time does to attention. A reviewer working through a constant stream of irrelevant flags has less capacity left for the ones that matter. A system tuned to avoid missing anything, at the cost of flagging everything, doesn't just waste hours. It works against the reason the system exists.
Use these questions to check whether a system is built on keyword matching or genuine context awareness:
If most of these point the wrong way, the system is built on keyword matching rather than genuine context awareness. That's the gap context-aware, explainable supervision is built to close.
Rather than matching against a fixed list of terms, an AI model can interpret a message the way a person would, weighing the relationship between words, the roles of the people involved, and the surrounding conversation. A word like "guarantee" stops being an automatic trigger and becomes one input among several the system considers before deciding whether something needs a closer look.
MirrorWeb's Mira works this way. It's built to reduce the volume of irrelevant alerts by understanding context rather than just detecting keywords, and to show its reasoning behind a flag rather than producing a bare alert with no explanation. That combination of fewer false positives and a visible rationale for the ones that remain - effectively an audit trail for every decision - is what lets a compliance team spend its time on real risk instead of clearing noise.
This isn't a wholesale replacement for human judgment - a reviewer still makes the final call. What changes is what reaches them: a shorter, better-prioritized queue, with the reasoning behind each item already attached.
False positives are not a minor inconvenience. They're a design flaw with a real cost - hours spent clearing alerts that were never a risk, and, per MirrorWeb's 2025 survey, an average of $232,457 a year in mobile supervision inefficiency alone. That's before accounting for the bigger risk: attention worn down to the point where a genuine issue can slip through unnoticed.
Fixing that starts with understanding why keyword-based systems generate so much noise in the first place, and it ends with a system that can explain its own reasoning, something regulators are now asking for directly. Firms that make this shift aren't just cutting down on wasted time and cost. They're building a supervision program that can truly stand behind its own decisions.
Book a demo of Mira to see a system that shows its work.