AI-generated business content fails on integrity, not only on quality: a generated document can state a rule, assert that every tactic in it was checked against that rule, and violate the rule in the same file while reading as competent throughout, because fluency is what the generator optimizes and truthfulness is not a property any document can establish about itself.
Somebody on your team is going to hand you a document an AI wrote, and it is going to read well. The structure will be clean, the tone will be right, the numbers will sit in tables. You will skim it and nothing will look wrong, and that is the part worth being careful about, because reading well and being true are different properties and only one of them is what the machine was optimizing.
I had one written for my own business a few days ago: a full go-to-market plan, down to templates that were ready to send to strangers. One adversarial review pass returned eight findings graded blocking. The worst of them for my credibility was a sentence claiming a working relationship that does not exist, addressed to people who had never heard of me.
Nothing reached anyone. It was caught in the template, which is the only reason this is an example and not an apology. But nobody in the normal review path would have stopped it, because that path is someone reading the document and judging whether it is any good, and by that standard it was fine.
Two of the eight are worth your time. Neither is the kind of error a careful read catches, and the person who owns catching them is whoever signs off on outbound content before it ships.
The plan included a ready-to-send template. One line read:
Your {{colleague_or_office}} and I traded notes about the intake step where the same client details get entered twice.
Later, in its own list of standing constraints, the same document said this:
No warm intros, no referral asks, no "who else should I talk to", no alumni networks, no past-client reactivation. Dan has no network in this market and every motion here is cold by design. This is the single most common way a plan of this shape fails, and every tactic in it was checked against this rule.
Read those in order, because the sequence is the point. The document states the fact that makes the template false. It states the rule forbidding it. It certifies that every tactic in it was checked against that rule. And it contains the tactic.
The reviewer's verdict on that certification was four words long: This tactic was not.

The falsehood is not a phrase that slipped through. It is a placeholder, named for a colleague or an office, waiting to be filled in at send time. And the plan's own mitigation does not save it: sending only to a firm that had already replied positively to something else does not turn "a firm replied to a cold email" into "your colleague and I traded notes."
The same plan projected two to three sales in sixty days, from a stated close rate of twenty-five to thirty-five percent, labeled as an inference.
The only rate the underlying research had actually grounded was an outbound-meeting-to-close rate of about nine percent, with a note that consultants without case studies sit at or below that. The plan applied a three to four times uplift to the one number its own corpus had measured, and never reconciled the two.
The contact base was overstated too: three hundred claimed, roughly two hundred to two hundred and twenty with enough runway to convert inside the window. Across all three channels, email, dials and LinkedIn, the reviewer's own recomputation put seven to eight qualifying calls in total. Five to seven of them land early enough to close inside the window, and at nine percent that is between 0.45 and 0.63 expected sales. The reviewer's verdict on the sixty-day target was a qualified yes for one, closer to a coin flip than a forecast.
Be fair to the machine here, because the reviewer was. There is a legitimate argument that a small offer with a capped downside closes better than the engagement the nine percent came from. It was never made and no number was attached to it. The defect is not dishonesty. It is optimism never reconciled against the one figure in the room that had a source.
Both are invisible to a careful read.
The template line is invisible because judging it requires a fact that is not in the sentence: whether that conversation happened. Fluency cannot encode that. The sentence is grammatical, plausible and in the right register, and those are exactly the properties being optimized.
The close rate is invisible because judging it requires holding two numbers from two different files at once and noticing that one quietly replaced the other. Nothing flags the substitution. Both are formatted identically, and one carries a label saying it is an inference, which reads like disclosure and functions like permission.
This is the difference between reviewing for quality and reviewing for integrity. Quality asks whether the artifact is good. Integrity asks whether its claims are true.
Internal consistency is checkable from inside the document. Does the artifact contradict itself? Does a number in one section survive a number in another? Does it obey the rules it states? Both of the findings above are this kind. They are internal to the generated corpus, though not both to a single file, which is why one adversarial reader with no outside access could catch them. That is the cheap layer, and almost nobody runs it.
Correspondence to the world is not checkable from inside at any price. Did that conversation happen. Is that person a customer. Does that regulation say what the paragraph claims. It requires leaving the artifact and asking someone or something outside it, and no amount of rereading substitutes.
Be precise about which layer your pass actually performs, because the flattering reading always overstates it. A reader checking a document against itself catches contradiction. It does not catch a false statement the document never contradicts. Both findings above were caught by the cheap layer only: the false relationship surfaced because the same file elsewhere said no such relationships exist, not because anyone checked whether that conversation happened. Omit the rule and the template survives the review.
It helps to sort claims by two things at once: what it would cost to check them, and what it costs if they are wrong. Most attention goes to the expensive checks. Some of the damage is there. The rest is somewhere nobody is looking.
Claims about the artifact itself are safe. "This sequence has six steps." Countable, verifiable in seconds, without leaving the page.
Claims about your product are mostly safe and wrong in boring ways. "The report includes a failure-mode map." Someone on your team knows, the correction is cheap, nobody's credibility moves.
Claims about the world are where the danger sits. A relationship, a prior conversation, a credential, an industry benchmark, what a competitor charges. The generator has no access to any of it. It has patterns describing what sentences of that shape usually look like, which is a different thing, and the result is indistinguishable in tone from the ones that are true.
Claims about whether it was checked are the worst category, and they are worst for an awkward reason: they are among the cheapest claims to disprove and almost nobody tries. "Every tactic here was checked against this rule." "All figures sourced." These read as the output of a process and are the output of a sentence generator that has learned what documents say near the end. Disproving one costs a few minutes, because a single tactic that breaks the rule is enough and it is sitting in the same file. Confirming one is the expensive direction, because no amount of reading shows you that a process actually ran. The consequence of not verifying it is unbounded, because the sentence actively suppresses the scrutiny that would catch everything else: a reader told every tactic was checked is less likely to check the tactics.
The rule that falls out of it: the more a sentence sounds like it is closing an inquiry, the more it deserves one.
Four questions, none needing a specialist, all answerable in an afternoon.
Which sentences assert a fact about the world? Not about your product, about the world: a relationship, a prior conversation, a credential, a result someone else got. Mark each one and name who can confirm it. Anything nobody can confirm is cut or rewritten as the weaker true version.
Does the document make a claim about itself? Phrases like "every item here was checked against" or "this has been validated for" are self-attestation, and a generated document is the least reliable possible source for them. Treat any self-issued certificate as unverified until something outside the document confirms it.
For every load-bearing number, where is the sourced one it replaced? Find the earlier figure, put them side by side, and require the change to be argued rather than labeled. An inference tag is not an argument. And if you cannot find an ancestor at all, stop there: a load-bearing number with no traceable source is not a weaker finding than a substituted one, it is the same finding, earlier.
Which of those will anyone actually check? Take the list from question one and mark each sentence with who is going to confirm it and by when. This is the triage step, and it is where the correspondence layer gets paid for or does not. Anything nobody is going to verify does not get stated as fact; it gets stated as an assumption or cut.
If you only run one, run question two. It is the cheapest, almost nobody has a habit for it, and it is the one that caught the thing in mine I would least have wanted to send.

The check cannot belong to whoever produced the artifact, and when the artifact is AI-written that includes the person who ran the generator. Not because they are careless, but because they read the output already knowing what it was supposed to say, which is the condition under which a plausible sentence stops being checkable.
It also should not sit with the most senior person available. This is a mechanical pass requiring time and a list, not a judgment call requiring authority, and routing it upward is how it becomes the thing that gets skipped in a busy week. The person who signs off still owns the outcome. What they should own is the slot in the process, not the task itself.
What it does need is that slot, an adversarial brief rather than an editorial one, and standing to stop a send. That last one matters more here than it would elsewhere: a document that reads well generates no urgency, so a finding against it that arrives as a suggestion loses to a deadline almost every time.
The flattering version of this story is that I spotted it. I did not, and I read the plan.
It was caught by a scheduled adversarial pass, a step I had put in that pipeline for exactly this. It was not caught by the pipeline that wrote the plan, which reported a clean run.
That is the operational point. The check has to run whether or not anyone is worried, because a review that happens when someone is already suspicious will never run on the document that reads well, and that is the one that needs it.
Apply all of this to what you are reading. This essay was drafted with AI assistance, about an AI-generated plan, and it went through the four questions above. They did not catch the worst thing in this essay.
An earlier version took the rule from the reviewer's shortened rendering, which carried an ellipsis, and replaced that ellipsis with a period while capitalizing the next word, which invented a sentence break that was not there. What the ellipsis had marked, and the period did not, was a cut: the clause in which the plan states that I have no network in this market. That clause is the spine of the whole argument, and I had removed it while quoting the document I was accusing of misrepresenting itself.
What caught it was an adversarial pass diffing the quotation character by character against the source, which is a check the four questions do not contain. It does not belong on a list, it belongs in my publishing pipeline, and that pipeline does not have it yet. Until it does, it is a standing rule that every quotation gets diffed before the first review round, which is the weaker of the two, a rule rather than a gate, and I would rather tell you that than describe the version I intend to build. A checklist that misses the thing it was written for is worth more to you than one that reports a clean run.
What makes this account checkable is narrow and I will not overstate it: both quotations are word-for-word against the source, once the list marker and bold markup are stripped, the four-word verdict came from a reviewer that was not the author, and the expected-sales figure carries the intermediate step it was derived from.
The eight were not equally serious and lumping them together would be its own small dishonesty. One was a false claim of a relationship. One was unreconciled arithmetic. The rest were structural in other ways, among them a failure gate that could not trip on the plan's own schedule, no scheduled way to take payment, a list that could not be built in the hours budgeted.
The arithmetic class is the one that keeps coming back. A rebuilt version failed review again the next day, on a class of defect it had already been corrected for twice: a threshold the schedule underneath it could not reach. That is when you stop adjusting the number and start asking whether you are checking the right thing. Note what it is not: it is not the self-attestation defect this edition is about. The one I would have bet on recurring is not the one that did.
If you are putting AI-generated content in front of customers, the question is not whether the model is good enough to write it. It clearly is, and that is the difficulty. The question is who checks whether it is true, on what schedule, and against what.
You do not need me for the first pass. Take the artifact you would least like an AI to speak for you in and search it for one thing: any sentence saying the document has been checked, reviewed, validated or sourced. Then look for one item it does not cover. Finding one takes minutes and tells you the certificate is worthless. Not finding one is not a clean bill, because nothing you can read shows you the check was ever run.
If it turns out the check belongs somewhere in the workflow rather than at the end of it, that is what the Readiness Scan my agency runs is for: one workflow, its failure modes mapped, and a plan for who owns it after launch. Work with OIA on one workflow.