← Journal

Three inputs and an adversary

A mind map settles the argument, a tone profile is data rather than an instruction, and the model that checks the work comes from a different vendor than the one that wrote it. Why those three constraints produce something worth publishing.

Three inputs and an adversary

A mind map settles the argument, a tone profile is data rather than an instruction, and the model that checks the work comes from a different vendor than the one that wrote it. Why those three constraints produce something worth publishing.

Plausible is the failure mode

Give a model a prompt and it hands back prose that reads fine. The reading-fine is what causes the trouble. Bad prose gets read carefully, because the reader has to work to follow it; good prose gets skimmed. That is my own observation, from watching it happen, and it is why fluency should not be read as a sign that the specifics hold. Good prose is still worth having; it is just not evidence of anything. A draft that sounds wrong gets fixed. A draft that sounds right gets published.

The register arrives borrowed. A model writes in whatever voice it read most of, and what comes out reads like a press release nobody signed — confident, evenly warm, faintly corporate, attached to no one. That is opinion rather than finding, but I would defend it.

The specifics are where it actually goes wrong. An invented figure, date or tool name sits on the page in the same typeface as a real one, with the same air of having been looked up. Nothing about the sentence signals which it is. And nothing in the exchange has been told what the piece is meant to argue, so nothing is holding a supporting point apart from a decorative one; the model will supply either on request, and both will scan.

So the interesting question about a writing pipeline is not how good the writing is. It is which decisions the model has been denied. What follows is that list, roughly in the order the decisions matter.

The argument is settled before a word is written

A flat illustration: a pencil, a microphone and a video camera above the words should you blog, vlog or podcast

The structure comes in as a mind map, and the mind map is the author’s. The top level is the spine of the argument, the children are the supporting points, the leaves are the specifics. That shape is written before the drafting starts, and the stages that follow answer to it rather than propose one. Denying the model the structure does not stop a draft drifting off it; drift still has to be caught downstream. What it does stop is the model inventing a shape of its own without anyone noticing.

One convention does more work than the rest. A node whose text ends in a question mark is treated as an open question. The ideator must either answer it from the author’s notes or drop it, and say which. It is never asserted as a claim by the post. Half-formed thoughts are the ones most likely to arrive in the draft dressed as conclusions, and the convention stops that.

The map is read where it lies, in whichever of three formats is present — a nested markdown list, OPML or FreeMind XML, taken in that order. Converting it to a single canonical format would produce a copy, and the copy would go stale the first time the author edited the original. Freeplane’s note pane is read too, and passed to the ideator with the node it belongs to, which is usually where the text actually worth having lives.

The claims the post will rest on are extracted here, at ideation, before the draft exists. Assembling a claim list from finished prose tells you what the model wrote. Extracting it beforehand tells you what the piece is for.

A rule nobody checks is a preference

“Warm but professional” tells a model nothing that can be verified after the fact. It cannot fail a check, so it is an aspiration.

A tone of voice profile therefore carries two halves doing different jobs. One is free-text prose describing how the voice should sound, which reaches the writer as prose. The other is hard rules and banned phrases. The prose half is the one no linter can check, and it is not redundant — it is how a voice gets communicated at all. The rules half is the part that can fail a build. Treating the profile as only one of these throws away whichever half you dropped.

The deterministic checks run first, on every iteration of the revision loop, before anything goes near a model. They are cheap and repeatable, and they clear the mechanical faults so the expensive pass spends its attention on the arguable ones. Sample texts, where an author supplies them, are used to propose rules by inference, and a proposal is not applied on its own say-so. What ends up enforced is still something a person chose.

The sources you chose beat the sources it finds

Reference material the author ingests is trusted above anything live search returns. This is an ordering, not a ban on search — the author’s own sources are the ones with a person’s judgement already attached.

A retrieved page is stored with the date it was retrieved. Pages change. The date says when the page was seen, so a citation made today is not quietly reinterpreted by whatever the page says next month. The stored copy does the other job: it keeps the text the claim was checked against available to look at.

A supplied document can be marked as informing the writing without being citable. Background reading that shaped the argument does not belong in the reference list, and having somewhere to put it is what stops it drifting in.

All of it — mind maps, notes, profiles, drafts, claim ledgers, citation records — is plain text in a repository, with git as the history. The history is git’s to keep rather than mine.

A model checking its own work agrees with itself

The adversary runs on a different vendor from the writer. Not a different prompt, not a different temperature, not a sibling model from the same firm. Two models from the same firm agree with each other more readily than either agrees with a source, and a model asked to audit its own output is being asked to have a different opinion than it just had.

Vendor is the unit, not the access channel. The claude CLI and the Anthropic API reach the same models through different billing; splitting a writer and an adversary across those two satisfies nothing at all while looking like a control. A local open-weights model is counted as its own vendor for the purposes of the rule. That is a routing convention; whether a given checkpoint really shares no training with a hosted provider is a separate question, and running it on your own machine does not settle it. The rule is enforced in the router when configuration loads, and it fails loudly. A check that can be switched off by a config file that nobody reads is not a check.

The adversary returns findings, not a rewrite. A rewrite hides what was wrong; a finding makes somebody decide.

Every claim extracted from the draft gets one of three statuses, which is not the same as every factual sentence in it — the last section comes back to that gap. Supported. Contradicted. Unsupported. A contradicted claim blocks the post, and there is no flag that overrides the block — the value of the block is entirely in its being unconditional. An unsupported claim halts the pipeline for a human decision, and if the human decides to keep it, the unsupported flag travels through into the output rather than being resolved by the act of publishing.

No two roles may share a vendor. That leaves room for a third: an arbiter, to settle claims the writer and the adversary cannot agree on between them. It is specified in the architecture and designed for. It is not yet wired.

The bibliography is the giveaway

Ask a model for a reference list and it will produce one that looks correct. Plausible formatting, plausible authors, plausible years, plausible DOIs. That is the worst property a reference list can have, because a list that looks right invites less scrutiny than one that looks wrong.

So no model writes it. The reference list is generated from the evidence collected during checking, and nothing else goes in. Quotes are verified against the stored page rather than taken from what a model remembers of it.

The unexamined share

What comes out alongside the post is a ledger: every claim that was extracted, and the evidence that settled it. Attached to the post, not remembered by whoever happened to run it. There is also a provenance record of what the post drew on and what a later post may cite it for.

And there is the number I actually care about. The tool reports which verification passes ran, and which parts of the finished text no pass ever questioned. That is where a confident invention sits undisturbed: in the sentences nobody thought to extract a claim from. A pass that ran and got it wrong will let one through as well, so the checked claims are not a guarantee either; the unquestioned share is the part nothing even looked at. Reporting that share is what lets a reader see where the checking stopped.

No accuracy figure appears here, and none should. I have not measured one, and a percentage would be doing rhetorical work rather than evidentiary work. The argument is structural: here is what the tool refuses to let a model decide, and the refusing is the whole of the guarantee.

As of 12 September 2026 the pipeline had produced a cited draft end to end through live models, and all three front doors — the CLI, the MCP server and the local web view — were working.