The most expensive AI mistake I have made was not a wrong answer. It was a confident answer I did not check. I installed a fuse based on a specification an assistant invented, and I only caught it because I happened to open the trailer manual one more time. Nothing burned. The lesson cost me an evening, and it changed how I work.
This is the tenth entry in The AI Test Log and the fifth in Workbench Skills. If you are learning how to use AI as a beginner, verification is the skill that separates a useful tool from a liability. It is also the skill nobody teaches, because it is not a feature and you cannot buy it. What follows is the system I now use on every task that matters.

Why Verification Is a Skill, Not Doubt
The Wrong Way to Think About It
Some people treat verification as paranoia, and some treat it as an insult to the tool. Both framings miss the point. I do not verify because I distrust AI. I verify because I know what the tool is doing: producing the most plausible next words based on patterns.
Plausible and correct overlap often. They are not the same thing, and the gap between them is where mistakes live.
What Verification Actually Costs
On a typical task, verification adds two to fifteen minutes. On a task involving money, safety, or a legal requirement, it can add an hour. That number sounds high until you compare it to the cost of being wrong.
This is why I decide the verification budget before I start, not after. If checking the answer will take longer than doing the task myself, I do the task myself.
The Four Risk Tiers
Tier One — Money and Commitments
Anything that spends money, sets a price, or creates an obligation. Pricing a product. Choosing a shipping carrier. Estimating a tax category. Signing off on a quote.
These get verified outside the chat, always, no exceptions. A pricing model does not know my cost per jar or what my market will bear, and it will not tell me it does not know.
Tier Two — Safety, Rules, and Compliance
Electrical specifications. Allergen statements. Shelf-life claims. Permit requirements. Labeling rules. Anything a regulator, an inspector, or a physical hazard touches.
This tier is where I once got an incomplete answer about cottage food labeling in Colorado. The response was organized, plausible, and missing a requirement. I caught it because I checked the state guidance directly.
For this tier, the model can help me build the question. It does not get to be the answer.
Tier Three — Anything That Changes
Hours, availability, prices, schedules, rules that update, and anything tied to a live system. A campground with a rolling reservation window. A store that closed last month. A trail rule that is seasonal.
Language models do not have live data unless a tool gives it to them, and even then I confirm. This tier is easy to spot once you ask one question: could this have changed since the model was trained?
Tier Four — Anything That Sounds Precisely Right
Specific numbers, part numbers, citations, statistics, and quotes. This is the tier people miss, because specificity feels like evidence. An invented part number reads exactly like a real one.
My rule: if a number, name, or quote did not come from a document I supplied, I treat it as unverified until I find it somewhere I can point at.
The Verification Moves I Actually Use
The Sourcing Question
One sentence, every time: which parts of this came from what I gave you, and which parts did you infer?
This separates reading from guessing, and it changes the answer. When I asked it about the trailer fuse block, the model admitted it was inferring from similar part numbers. That admission is what stopped me.
The Reverse Test
I ask the same question from the opposite direction. If I asked which tool is best for a task, I ask which situations that tool handles badly. If I asked whether a plan works, I ask what would make it fail.
Consistent answers across both directions are a weak signal of reliability. Contradictions are a strong signal to slow down.
The Second Tool
I keep a second assistant for one purpose. Same question, different model, compare the answers. When they agree, I still verify facts. When they disagree, I know the topic is contested or thin, and I go to an outside source.
The Two-Minute Rule
If I can verify a claim in under two minutes, I verify it immediately instead of trusting it. Most claims fall into this bucket. A quick search, a page in a manual, a phone call, a look at the actual object in front of me.
The claims that take longer are the ones that carry more weight, and those are exactly the ones worth the time.

Three Times Verification Saved Me
The Fuse Specification
Covered above. An invented amperage rating for a circuit that also feeds a water pump. Verified by checking the manual's parts list and calling a supplier who could document the modern equivalent.
The Cottage Food Label
An incomplete answer about required label elements. Verified against the state guidance page. The fix took ten minutes. The consequence of shipping mislabeled product could have been far worse than a fine.
The Campsite That Was Full
An assistant recommended a campground without noting that reservations open months ahead. Verified with one phone call. We camped somewhere better, and I avoided a four-hour drive to a full lot.
None of these were dramatic. That is the point. Verification is boring, and boring is what keeps small errors from becoming expensive ones.
How I Mark Unverified Claims
I keep a simple habit on longer documents. Anything in a draft that did not come from my own material gets a small mark in the margin. Then I work through the marks before the document goes anywhere.
Three categories. From my input, verified. Inferred, needs checking. Unknown source, delete or replace.
That last category is important. If I cannot find where a claim came from, it does not belong in something with my name on it.
What Does Not Need Heavy Verification
Brainstorming. Structural options. Rewriting something I already wrote. Explaining a concept I can test by using it. Summarizing a document I supplied and can re-read.
For these, a quick read is usually enough, because I am the source and the judge. The risk is low and the time saved is real. Knowing which tasks fall into this group is what makes the whole system practical instead of exhausting.
What Still Needs Human Judgment
The decision about what tier a claim belongs to. That call is mine, and it changes by context. A part number for a decorative bracket is tier four and low stakes. The same kind of number for a fuse is tier two and high stakes.
I also keep the final read on anything customer-facing. If a sentence claims something I cannot back up, it gets cut, no matter how good it sounds.
Cost and Time
One subscription plus a second for cross-checking. Verification time on an average work task: five to ten minutes. On a task involving money, safety, or rules: thirty to sixty minutes. Time saved by catching one wrong specification before it becomes a repair, a refund, or a compliance problem: considerably more than that.
Final Verdict
Keep it.
Verification is not the tax you pay on using AI. It is the part of the work that makes the rest of it usable. Four risk tiers, one sourcing question, a reverse test, a second tool, and a two-minute rule. That is the whole system, and it fits on a single index card.
If you remember one thing from this log, make it this: ask where the answer came from before you act on it.
Take it apart first. Then ask AI.
No comments yet — be the first to share a thought.