
The email landed on a Tuesday afternoon, routed to a department head who had been using an AI tool to stay on top of regulatory updates in her industry.
She had asked the tool a straightforward question: what were the current reporting requirements for a specific type of client transaction her company handled regularly? The response came back detailed, organized, and authoritative. It cited what appeared to be the relevant framework, laid out the key thresholds, and even included a timeline. She incorporated the guidance into a summary she shared with her team. Decisions were made based on it.
The requirement the AI described had changed. The information in its response reflected an earlier version of the rule. Nothing about the answer looked wrong. The formatting was clean, the logic was coherent, and the tone projected the same confidence as every other response she had gotten from the tool. There was no disclaimer, no hedge, no flag that any of it needed verification. The problem only surfaced later, when an external review caught the discrepancy.
This is not a story about carelessness. It is a story about a tool communicating in such a way that makes it very difficult to know when to push back.
Why AI Sounds Right Even When It Isn’t
AI tools are built on a straightforward principle: given a body of text → predict what a useful, coherent response looks like. They are extraordinarily good at this. But that capability is not the same as accuracy, and the gap between the two is where business risk lives.
When an AI model encounters a question it cannot fully answer (because its training data is outdated, incomplete, or simply doesn’t contain the right information), it often doesn’t stop and say so. It does what it was built to do: generate a response that fits the pattern of a correct answer. The result is what the industry calls a hallucination.
The word “hallucination” is worth understanding once and then setting aside. What matters is the practical reality it describes: AI generates content that looks like accurate output even when it isn’t, and it does so without signaling uncertainty. This is a materially different risk than using a search engine, reading an article, or asking a colleague → all of which come with cues that help you calibrate how much of the information is worth trusting.
Where the Stakes Are Highest
Not all AI-assisted work carries the same risk. The relevant question is not whether you use AI, it’s whether the work being done with AI has meaningful consequences if it turns out to be wrong.
For most organizations, the highest-risk categories include:
- Client-facing communications where accuracy reflects directly on your credibility
- Compliance and regulatory research where incorrect guidance can create legal or financial exposure
- Financial summaries and projections where numbers inform decisions
- HR and employment guidance where errors can become policy
- Any situation where AI output is used to draft something that gets sent, signed, or acted on without a second set of eyes
Low-stakes use such as brainstorming, formatting, or a first-draft prose on internal documents, carries far less risk due to the limited consequences of an error and the output being naturally reviewed before it matters.
The practical principle is proportionality: the higher the business consequence of being wrong, the more human verification is required before acting on what AI produces.
The Problem Gets Worse Over Time
There is a compounding dynamic worth naming directly. When AI tools consistently perform well (and for a wide range of tasks, they do) employees begin to naturally trust them more. Review steps that felt important at the beginning start to feel unnecessary. “The tool has been right every other time!”
This is where individual errors become organizational habits. When an AI-assisted workflow loses its verification step, what’s left is an unreviewed information source embedded in business decisions. The AI isn’t more or less reliable than it was before. The scrutiny has simply eroded.
Air Canada experienced a version of this at scale. Its customer service chatbot provided a passenger with incorrect information about bereavement fare policies. The passenger acted on it, the company did not honor what the chatbot had communicated, and Air Canada was successfully sued. The court rejected the argument that the chatbot’s output was somehow separate from the company’s responsibility. The AI spoke for the organization, and the organization was held accountable for what it said.
The lesson here is less about AI chatbots and more about human oversight. Once AI output is integrated into customer-facing or decision-facing workflows, the business owns the consequences regardless of what produced the answer.
Questions Leaders Should Be Asking
Business leaders don’t need to understand the technical reasons AI produces incorrect output. They do, however, need to ask whether their organization has thought carefully about where AI answers are being trusted without verification.
Start here:
Which AI-assisted work in our organization gets reviewed before it’s used and by whom? If the answer is “it depends on the employee,” then it’s an unmanaged variable, not a policy. Organizations who’ve thought this through can name the categories of work that require review and the very person responsible for it.
Do employees understand that confident AI output is not the same as accurate AI output? Most people who use AI tools regularly develop a feel for when they’re working well. But that intuition doesn’t reliably signal when the tool is wrong on something factual. Employees need to understand that the format and tone of an AI response carry no information about its accuracy.
Have we defined which categories of work require human verification before acting? This doesn’t need to be a complex policy. It can be a short list: client-facing content, compliance research, financial figures, and anything regulatory. That gives employees a clear standard to apply rather than leaving the judgment call to individual discretion.
For IT Leaders: Building the Verification Layer
Misinformation risk is manageable but managing it requires moving beyond “train employees to be careful.” A few specific steps:
- Identify the high-stakes AI-assisted workflows in your organization. These are the places where AI output influences a communication, document, or decision with real consequences. These are the workflows that need a built-in review step, not a reminder to use good judgment.
- Document a short list of content categories that require human verification before use. Push this list to the relevant departments and make it part of onboarding for any team adopting AI tools.
- If AI tools are integrated with client-facing systems, communications platforms, or customer service workflows, treat their output with the same scrutiny as any other externally visible content. The Air Canada case is the clearest available example of why this matters.
- Log and periodically review AI-assisted outputs in high-stakes categories. Not to police employees, but to catch patterns of error before they become patterns of exposure.
Confidence Is Not the Same as Accuracy
AI tools are genuinely useful. They help people work faster, communicate more clearly, and handle tasks that would otherwise take significantly more time. None of that is in dispute.
What is also true is that they produce confident-sounding output regardless of whether they are correct. The situations where organizations are most likely to rely on AI answers are often the situations where being wrong carries real consequences.
The goal is not to stop or slow AI use. It is to guide its use with enough structure that the risk of acting on a wrong answer is proportional to the stakes involved.
If your team is using AI for client communications, compliance research, or financial decisions, this article is worth passing along. The next piece in this series covers where sensitive business information can unintentionally end up when employees use AI tools along with why it matters more than most organizations realize.
Source: This article draws on guidance from the OWASP Top 10 for LLM Applications, 2025 (LLM09: Misinformation).
