All guides

AI newsletters and hallucinations: how fact-checking actually works

Why AI drafts contain confident errors, the specific mistakes that show up in newsletters, how a claim-by-claim checking pass catches them, and what still needs a human.

The short version

  • A language model predicts plausible text; it does not look anything up unless a tool makes it.
  • The common newsletter errors are rounded numbers, proposals reported as decisions, misattributed quotes and links that do not say what the sentence claims.
  • Checking means opening every link and comparing each claim to the source, sentence by sentence.
  • Brieva’s checker marks each claim verified, softened or cut, and shows you the scorecard before you approve.
  • A checker catches factual drift; only a person catches the claim that is accurate but unwise to send.

Why does AI make things up?

Because that is what it is built to do. A language model produces the most plausible next words given everything before them. Plausible is not the same as true. When the model has read the source and the fact is in it, plausible and true coincide. When the source is ambiguous, or the model is asked for a detail the source did not contain, it fills the gap with something that sounds right, in the same confident tone as everything else.

The word hallucination makes it sound like a malfunction. It is closer to the ordinary behaviour of a very fluent writer who has not been told that saying "I do not know" is allowed. Left alone, the model will always prefer a specific, confident sentence to a vague, honest one, because specific confident sentences are what most of its training text looks like.

This matters more for a newsletter than for a chat. In a chat you read the answer and judge it. In a newsletter the answer goes to two hundred people who trust the sender, and the sender is you.

What do the errors actually look like?

They are rarely dramatic. Nobody’s newsletter announces a war that did not happen. The errors are small, plausible and fluent, which is what makes them dangerous. From checking thousands of drafts, these are the shapes that recur.

Rounded numbers. The source says a 3.7% increase; the draft says "nearly 4%", then a later sentence says "4%", then the subject line says "4% rise". Each step is defensible; the result is wrong.

Proposals reported as decisions. A regulator publishes a consultation paper; the draft says the rule "will change". A council debates a bylaw; the draft says it "has introduced" one. The tense drifts toward the more newsworthy version.

Misattributed quotes. A trade body’s press release quotes its chief executive; the draft attributes the line to the minister who was also mentioned. Or the quote is real but the person’s title is last year’s.

Links that do not say what the sentence says. The link is real, the page exists, and it is about the topic, but the specific claim in the sentence is not on it. This is the most common error and the hardest to spot by eye, because the link "looks right".

Invented specificity. The source says "several regions"; the draft names three. The source says "in recent months"; the draft says "since March". The detail was not in the source; it was supplied because the sentence read better with it.

Stale facts from training. The model knows what the threshold was two years ago and writes that, even though the source it was given says something else. This shows up most with tax rates, thresholds and dates.

Why does "use a better model" not fix it?

Better models make fewer errors and make them more fluently. The rate goes down; the detectability goes down with it. A weak model’s mistakes are often obvious. A strong model’s mistakes read exactly like its correct sentences.

Model quality also does nothing about the structural problem: the model is asked to write from sources, and the sources may not contain the answer to the question the draft is implicitly asking. When a strong model does not have the fact, it still produces a fluent sentence. The only fix for that is a step that goes back to the source and checks.

This is why the question to ask any AI newsletter tool is not "which model do you use" but "what checks the draft".

What does a checking pass actually do?

It treats the draft as a set of claims and tests each one against the source it cites. In practice that means, for every sentence that asserts a fact: identify the claim, open the linked source, find the passage the claim rests on, and decide whether the source supports the claim as written, supports a weaker version of it, or does not support it at all.

Brieva does this with a separate pass that runs after the draft is written and before you see it. It is a different job from writing, with different instructions: the writer is asked to be useful and readable; the checker is asked to be sceptical and literal. Splitting them matters. A model asked to write and check in one go will grade its own work generously.

Every claim ends up in one of three states. Verified: the source says this, as written. Softened: the source says something weaker, so the sentence has been adjusted to match (a "will" becomes "is proposed to"; a "4%" becomes "3.7%"). Cut: no source supports it, so it has been removed, and the scorecard says what was removed and why.

Links are checked separately for the boring failure: does the URL resolve, and is the page it resolves to actually about this? A link to a 404, a login wall or a homepage instead of the article is flagged even when the claim itself is fine.

What the scorecard shows you

With every draft you get a report card: an overall verdict (looks good, passed with warnings, or needs review), how many links were checked and how many resolved, and a list of what was softened or cut, each with the original sentence, what changed, and the line from the source that decided it.

The point of showing the softened and cut items rather than silently fixing them is that you learn what your sources are like. If the same trade publication keeps producing claims that get softened, its headlines overstate its articles, and you may want it out of the source list. If a section keeps producing cuts, the section is asking for content the sources do not provide.

It also means you can disagree. A checker is literal, and sometimes the softened version is too cautious for a reader who knows the industry. You can send the draft back with a note ("the WorkSafe item is confirmed, see their media release from Tuesday") and the revised draft will reflect it.

What still needs a person?

Three things a checker cannot see.

Whether a true claim is wise to send. Your supplier’s price rise is accurately reported; do you want to be the one telling your customers about it? A regulator’s enforcement action against a competitor is real; is it your place to circulate it? The checker verifies; it does not advise.

Whether the emphasis is right. Four verified items in the wrong order tell a different story from the same four in the right one. The lead is a judgement about what your readers care about, and it is yours.

What happened inside the business. The job that illustrates the point, the question three customers asked this week, the thing you changed: no source publishes these, so no checker can verify them, and they are the most valuable sentences in the issue. You type them, and they go out on your say-so.

This is why Brieva does not send without approval and does not offer a way to turn that off. The checking pass makes the five minutes of approval a real five minutes instead of an hour of anxious rereading. It does not make them zero.

How to read a draft in five minutes

Read the scorecard first, not the draft. If it says looks good and every link resolved, read the draft for emphasis and tone. If it says needs review, go to the flagged items before anything else.

For each softened item, ask whether the softer version is still worth including. Often it is; sometimes an item that was interesting as a decision is not interesting as a proposal, and you cut it.

For each cut item, ask whether you know the fact from somewhere else. If you do, and you have a source, send the draft back with the source. If you do not, the cut was right.

Then read the draft top to bottom as a reader would, on your phone. Check the lead is the lead. Add your sentence. Approve.

If you find yourself doing more than this every issue, something upstream is wrong: too many sources, sources that overstate, or a section asking for things the sources do not have. Fix the configuration rather than editing every draft.

A worked example: one draft, four flags

A fortnightly issue for an accounting practice, four items, sources: Inland Revenue, MBIE, the Beehive, and a trade publication. The draft reads well. Here is what the checker did with it.

Item one said provisional tax dates "have been brought forward for the 2027 year". The Inland Revenue page it linked describes a consultation on that change. Softened: the sentence now reads "Inland Revenue is consulting on bringing forward…", with the submission date added because it was on the page and it is the thing a reader can act on.

Item two quoted a 6.2% figure for a wage-growth statistic. The linked Stats NZ release says 6.2% for one measure and 4.9% for the one the sentence was actually describing. Softened: figure corrected to 4.9% and the measure named.

Item three said a bill "passed its third reading on Tuesday". The Beehive release was from Tuesday and announced the second reading. Softened: "passed its second reading; the third reading is expected before the end of the month" was cut back further, because the "expected" part had no source; the final sentence says only what the release says.

Item four was a nicely written paragraph about a change to trust reporting requirements. No linked source supported it; the model had written it from what it remembered about a previous year’s change. Cut entirely, with a note in the scorecard saying so. The practice manager knew of the change, found the current guidance, sent the draft back with the link, and the revised draft included the item with the correct year and threshold.

Overall verdict on the first draft: needs review. Time for the practice manager to deal with all four: about six minutes, most of it finding the trust guidance. Without the checker, three of the four errors would have gone to four hundred clients, and the fourth would have gone with last year’s numbers.

How the checker avoids grading its own homework

The obvious way to check an AI draft is to ask the same model whether it is correct. That does not work well: a model asked to review text it just produced tends to agree with itself, for the same reason it produced the text in the first place. Three things make a checking pass actually independent.

It is a separate step with the source in front of it. The checker is not asked "is this right?" in the abstract; it is given the sentence and the page and asked whether the page supports the sentence. That is a reading task, not a knowledge task, and models are good at reading tasks.

It is asked to be literal. The writing pass is told to be helpful and fluent. The checking pass is told to be pedantic: a proposal is not a decision, 3.7 is not 4, "several" is not "three". The instructions pull in opposite directions on purpose.

It has to show its working. Every softened or cut claim comes with the sentence from the source that decided it. A checker that says "this is wrong" without a quotation can be wrong itself; one that has to quote the source can be checked by you in ten seconds.

The link check is separate again and does not involve a model at all: it fetches the page and looks at the response. A dead link is a dead link regardless of what any model thinks.

Configuring for fewer flags

Flags are information about your sources and your sections, and after a few issues they tell you how to configure for fewer of them.

If one source keeps producing softened claims, its headlines overstate its articles, which is common with trade publications that rewrite press releases. Either remove it or add a voice rule telling the writer to treat that source as a lead to the primary source rather than as the source itself.

If a section keeps producing cuts, it is asking for something the sources do not provide. A "what this means for you" section will generate opinion with no source unless you tell the writer that this section is yours and should be left as a placeholder for you to fill. A "by the numbers" section will invent a number when the sources have none unless it is allowed to be empty.

If the same kind of fact keeps being stale (thresholds, rates, dates), add the current value to the newsletter’s standing notes so the writer has it in front of it every issue rather than reaching for what it remembers.

And tell the writer what it is allowed not to know. A voice rule as simple as "if the sources do not say, do not say" removes most invented specificity on its own.

What this costs, and why it is worth it

A checking pass roughly doubles the AI work per issue compared with generating a draft and sending it. On a flat plan that cost is absorbed; it is one of the reasons Brieva prices per list size rather than per generation. It also adds a few minutes to the time between research and draft, which nobody notices because the draft arrives overnight.

Against that, consider what one wrong claim costs. Not the correction, which is a sentence in the next issue. The reader who knew it was wrong, who now discounts everything else you send. Trade newsletters go to the people best placed to catch errors in your trade. The margin for being confidently wrong in front of them is small, and the checking pass is what keeps you inside it.

Want this written for you every issue?