← Journal

The Model Is Not the System

On 18 September the US military came within minutes of boarding a Chinese ship on the strength of an AI-generated intelligence report that was, in the words of one source, entirely false. Everyone will call it a hallucination. It was an architecture, and the same one is being bought by the licence in ordinary companies every week. The essay, a one-page brief for the board, and a company that built its product the right way round.

A man in a dark suit stands before an enormous speech bubble on a stone plinth, striking through its lines of text with a red pen.

AI is only as safe as the decisions, checks and people built around it. The US military nearly boarded a Chinese ship on an AI-generated report. Most of us will not start a war. We can still lose.

This one is in three parts. The first is the argument, worked through a near-miss at sea. The second is the whole thing on one page for the board, and the designed brief comes with the PDF download at the end of the essay. The third is a company that built its product the right way round, which turns out to be rarer than it should be.

Who this is for · jump to your part
Everyone

What happened, and why the word hallucination is doing all the wrong work. Start at the incident and read to which model was it.Part one · about twelve minutes · one table, one pipeline

Boards and executives

The risk, the gap, six commitments and the business case, on one page. The brief on one page.Part two · four minutes · the designed brief comes with the PDF download

Anyone choosing a tool

Why picking the tool you already have is picking an architecture without knowing it, and the three ways the human check fails. The same failure, everywhere else.Part one · about eight minutes

Bid, legal and risk teams

A tendering company that built its AI to do the groundwork and refuse the exam question. Guided, not spoon-fed.Part three · five minutes

In a hurry

Governance is whether it is safe to say the report is wrong. That section stands on its own.Three minutes

Part one · The argument

The incident

A container ship in a harbour, its cargo made of stacked documents and speech bubbles, with a red targeting reticle hanging in the sky above it. A man on the quay studies it through a magnifying glass.

On 18 September 2026, CNN reported that the US military came within minutes of boarding a Chinese vessel in the Middle East on the strength of an intelligence report that was, in the words of one source, “entirely false”.

The sequence, as far as it can be reconstructed from CNN and the outlets that followed it, runs like this.

  1. Intelligence reporting on the ship’s manifest originated with US Special Operations Command Pacific in Hawaii.
  2. An analyst at Special Operations Command put that reporting to an AI chatbot.
  3. The chatbot fused open-source information with classified signals intelligence held in government systems and concluded the ship was carrying components linked to a nuclear weapons programme.
  4. The analyst then used AI again to package the conclusion into a standard-format intelligence report.
  5. The report circulated across the US military during the spring, in the middle of the war with Iran, and set off alarm bells.
  6. Preparations for an interception began. Armed personnel were preparing to board. Aircraft were already airborne.
  7. Shortly before the operation, officials examined the underlying intelligence more closely, found it had been AI-generated, and found the cargo had been misidentified.
  8. The operation was halted. CNN could not establish what the cargo actually was.

CNN could not determine whether the chatbot was a commercial product or a government-built one. Neither the Pentagon nor SOCPAC responded to requests for comment. Sources told CNN that similar AI-related errors have occurred elsewhere in the intelligence community as the technology has spread, which is the sentence in the story that should worry you most, and the one that got the least attention.

Everything below is analysis of a fast-moving story. The facts may sharpen or shift as the reviews are done. The argument, I suspect, will not.

This is not a hallucination story

The tempting headline is “AI hallucinated and nearly started a war”. It is true, and it is almost useless. It puts the fault in the model, implies the fix is a better model, and leaves untouched the thing that actually failed.

What failed was the architecture around the model. Read properly, the incident is close to a textbook illustration of the wrong tool being handed the wrong job inside a system with no gates, and I say that with the weary affection of someone who has drawn the diagram of how not to do this more times than he would like.

It is not one failure. It is a stack of them, and each layer on its own would have been survivable.

LayerWhat appears to have gone wrong
Retrieval and contextOpen-source reporting and classified signals intelligence were combined without provenance or verification. The problem was not too little context. It was too much, badly governed.
Model selectionA probabilistic generative model was allowed to do factual identification and evidence reconciliation, a job that needs far tighter controls than free-form synthesis.
Output formatThe conclusion was then rendered, again by AI, into the visual and linguistic conventions of an intelligence product. It acquired the appearance of institutional knowledge.
Human judgementThe output gained credibility faster than it gained verification. It was believed because it looked like the things that are normally believed.
OrchestrationThere was no mandatory, independent validation stage between “the chatbot produced an answer” and “the answer entered an operational decision chain”.

Together, with no gate between them, they route a probabilistic guess straight into a boarding party.

What a defensible architecture looks like

A factory conveyor belt carries a document from a chatbot through a series of stations: a magnifying glass, a set of scales, two small robots comparing notes, and finally a man in a shirt and tie who reads it under a lamp and stamps it before it passes through a door into a boardroom. Beneath the belt, a container ship sits in grey water.

My design rule for this has not changed in two years, and I am starting to think it will not need to:

Code where certainty matters. Micro-models where specialisation matters. Frontier models where intelligence matters. Context everywhere. Human judgement over consequence.

The incident collapsed all five into “ask the chatbot”.

For a task like cargo identification, a safer pipeline separates the jobs, and the separation is the point:

  1. Raw intelligence
  2. Deterministic extraction manifest fields, entities, identifiers
  3. Source and provenance reconciliation observed or inferred; who reported it; how reliable
  4. Specialist classification model cargo category, dual-use flags
  5. Confidence and contradiction checks what conflicts? what is uncorroborated?
  6. Independent human analyst
  7. Only then, synthesis by a frontier model
  8. Human decision gate before anything operational

The frontier model is still useful in this pipeline. Its proper questions are the interesting ones. What are the plausible interpretations of these observations? What evidence contradicts the current hypothesis? What additional intelligence would tell the hypotheses apart?

What it should not be doing on its own is converting ambiguous source material into a flat factual assertion, this ship is carrying nuclear weapons components, and letting that assertion become operational truth. A model that is superb at “what might this mean?” is a liability at “what is this?”, and the two questions look almost identical on a screen.

Too much context, badly governed

The reflexive diagnosis for an AI error is “it didn’t have enough information”. Here, everything was present. Open-source information. Classified signals intelligence. Cargo and manifest data. A generative model. Formatting tools. Military reporting conventions. The larder was full.

What was missing was the machinery that decides:

  • Which source is authoritative?
  • What conflicts with what?
  • What must be independently corroborated before it can be asserted?
  • What is inference rather than observation?
  • How uncertain is the conclusion, and is that uncertainty carried forward or quietly erased?
  • What must never be promoted into a decision without a human looking at it?

This is the context-compiler argument in its sharpest form. The value is not in the model’s intelligence. It is in the machinery that decides what the model knows when it starts thinking, and, just as importantly, what status its output has when it stops.

Confidence is not context

A giant chat bubble wearing a general's peaked cap sits on a plinth between two flags while a line of staff with clipboards wait to salute it.

The most dangerous moment in the sequence was not when the chatbot sounded confident. Chatbots always sound confident. It is their one reliable feature. The dangerous moment was when the organisation became confident.

Once the conclusion was wrapped in the format of an intelligence report, everything downstream treated it as institutional knowledge rather than machine output with an unknown error bar. The format did the persuading. Nobody had to be fooled by the model. They only had to trust a document that looked like every other document they trust, which is a thing people are trained for years to do.

I wrote about this failure at boardroom scale in The Room Where No One Can Tell You You’re Wrong. There, confident-sounding output leads to a bad technology investment. Here it leads to aircraft in the air. The mechanism is identical. A probabilistic output was laundered into certainty by presentation, and presentation is cheap.

Which model was it? It doesn’t matter

Speculation about whether the chatbot was Grok, Gemini, Claude or something built in-house is understandable, and it is beside the point. The reporting does not identify the tool. The Pentagon has access to several.

The stronger observation is that the architecture failed regardless of vendor. If the workflow lets a general-purpose chatbot turn mixed intelligence sources into an operationally consequential factual claim without hard provenance, contradiction checks, independent verification and a human validation gate, then swapping one frontier model for another solves nothing. It only changes which model gets blamed.

The model is not the system. The decision environment around it is.

The micro-model extension

The incident also strengthens the case for progressive specialisation, which regular readers will recognise as the thing I cannot stop talking about.

Instead of pointing a frontier chatbot at every intelligence task, a system can accumulate the recurring specialist sub-tasks and compile them down into tested, deterministic or narrowly trained components:

  • manifest parsing
  • entity resolution
  • cargo classification
  • contradiction detection
  • source reliability scoring
  • observation-versus-inference tagging

Each of these is auditable, testable and versioned in a way a chatbot prompt is not. The frontier model then receives only the genuinely ambiguous residue, and its output is labelled as such.

The design principle: reduce uncertainty as far as possible with deterministic and specialist systems, then spend frontier intelligence on what remains uncertain. Frontier intelligence is expensive and occasionally inventive. You want it doing the crossword, not the accounts.

The same failure, everywhere else

It would be comforting to treat this as a military story. It isn’t. The pattern trickles straight down into ordinary industry, and it repeats in almost every organisation now deciding where to put AI. Nobody is boarding a ship. The mechanism is the same.

Choosing by availability, not by fit

Most technology selection is not selection at all. Organisations reach for what they already have: the chatbot bundled into the productivity suite, the assistant that came with the cloud contract, the tool someone on the team already pays for. It is cheaper. It is already approved. It requires no architecture conversation, and nobody enjoys an architecture conversation.

What is rarely asked is why one system is better than another for a particular task, or how those systems are built inside. A general-purpose chat portal, a retrieval pipeline, a specialist classifier and a deterministic rules engine are not interchangeable products with different price tags. They are different machines that handle uncertainty, provenance and validation in fundamentally different ways. Picking one because it is already on the desktop is picking an architecture without knowing it.

The analyst almost certainly did not select the chatbot after a careful evaluation of alternatives. It was there. That is the industry pattern in miniature.

The risk you cannot see

The cheaper, already-available option carries risk the buyer is usually unaware of, and not because the risk is hidden. It is because nobody in the selection process is equipped to see it. The person choosing the tool understands the price and the interface. They do not understand that the tool has no concept of source authority, no contradiction checking, no distinction between what it observed and what it inferred, and no gate before its output becomes a document.

The result is a complete mismatch of context caused by portal choice alone. The task needed a system that knew which source to trust. The organisation supplied a system that treats every token in its window as equally true. No amount of good intent from the user repairs that mismatch, because it was decided before the user ever typed a prompt.

Three ways the human gate fails

A man stands on top of a towering, leaning stack of paper, writing on the top sheet under a huge angled lamp, while a row of five board members sit at a long table far below and look up at him.

The last layer of the design rule is human judgement over consequence. In practice that gate fails in three distinct ways, and organisations usually blur them into one.

  1. Unwilling. The user has the best of intentions but does not judge the output, because the tool is fast, the output looks finished, and checking it feels like redundant work. This is the laziness case, and it is structural rather than moral. The tool has been designed to make its output look done.
  2. Unable. The user wants to validate but cannot. The analyst may have had no way to check which of the fused sources drove the conclusion, or whether the classified intelligence actually said what the model implied it said. You cannot validate what the tool has already flattened into a summary.
  3. Unequipped. The user validates, but the tooling has no validation routines. No provenance trail, no confidence surfaced, no contradiction flagged, no “this is inference” marker. The human is asked to be the safety system for a machine that gives them nothing to be safe with.

In the CNN case it is not yet known which of the three applied. It may well have been all three. The point for industry is that the third case is the one organisations can actually fix, and it is the one they most consistently ignore, because fixing it means building architecture rather than buying a licence.

What this means for technology selection

The selection question is not “which AI should we use?”. It is:

  • What kind of task is this: synthesis, classification, extraction, decision?
  • What kind of system handles that task with the right treatment of certainty?
  • Does the tool we already have expose provenance, confidence and contradiction, or hide them?
  • If a human is the gate, does the tool give that human anything to gate with?
  • What is the cost of the output being wrong, and does the tool’s architecture match that cost?

A cheap tool with no validation routines is not cheap. It has moved the cost from the licence to the consequence, and the consequence does not come with a free trial.

Governance is whether it is safe to say the report is wrong

There is one more failure that sits above all the others, and it is the one organisations are most tempted to commit after an incident like this.

CNN’s reporting is itself a form of governance. It is the independent check that surfaced the error, forced the question, and made a review unavoidable. The reflex response to that kind of scrutiny, restrict access, shut the reporter out of the room, treat the disclosure as the problem, does not remove the failure. It removes the last mechanism that would have found it. The organisation stops learning about its own mistakes precisely at the moment it most needs to.

Inside a company the same function wears a different name: internal audit, the risk committee, the whistleblowing channel, the post-incident review, the analyst who says “I’m not sure this report is right”. The mechanism differs. The failure mode is identical. Punish or sideline the person who surfaced the error and you have deleted the validation gate you were relying on. The next bad AI output goes unchallenged, not because nobody noticed, but because noticing became career-limiting.

Every gate in this essay, the provenance checks, the contradiction detection, the independent validation, the competent human at the end, depends on someone being willing to say that the confident, well-formatted output is wrong. Governance is not a department. It is whether that sentence is safe to say.

Conclusion

The CNN story is not evidence that AI doesn’t work in intelligence. It is evidence that AI systems must be engineered as systems, not deployed as chatbots with authority.

A generative model can be entirely appropriate for synthesis and entirely inappropriate as the final authority on a factual claim. The engineering question was never whether AI is used. It is which intelligence is permitted to do which job, with what evidence, at what confidence, and what must happen before its output becomes consequential. And whether the people who catch the errors are protected when they do.

Part two · For the board

The brief on one page

Infographic titled The model is not the system, laid out in four columns: the risk of getting it wrong, the gap and why it closes, six commitments for doing it right, and better outcomes for the business, with a row of board-level business outcomes along the bottom.

Boards do not read twenty-five-minute essays, and nor should they have to. This is the same argument reduced to what a board actually needs: the risk, the gap, what we commit to, and why it is worth it. The designed four-page version comes with the PDF download at the end of this essay.

The risk of getting it wrong. One unchecked AI output, dressed up as a real report, escalates fast. Bad decisions and lost business. Reputation damage. Loss of face with clients and regulators. In the extreme, company failure. The cheap tool is not cheap. It moves the cost from the licence to the consequence.

The gap, and why it closes. These tools are becoming critical to how we work. The knowledge gap is wide. The ignorance gap is wider: most people cannot say why one system suits a task and another does not. That gap is closable. It is a small amount of reading and a modest amount of re-architecting to take the next step, not a re-skilling programme. Picking the tool we already have is picking an architecture without knowing it.

Doing it right. Six commitments.

  1. Right architecture. The right kind of system for each task.
  2. Right context. Governed sources, known provenance.
  3. Right judgement. Human intervention wherever it is critical.
  4. Independent validation. Multiple models, different vendors, checking each other.
  5. A competent human checks it. Someone who knows what they are doing.
  6. Protect the people who catch errors. Governance is whether it is safe to say the report is wrong.

What the business gets. Higher quality decisions and fewer costly errors. Faster, better-informed decisions. Higher win rates on bids and deals. Stronger trust with clients, regulators and partners. Productivity without added risk. An AI capability that turns into a durable advantage rather than a liability with a logo.

This is not a headcount strategy. The vision is not to save money by running lean or freezing hiring. It is to make sure every person we have is up to the job of judging what these tools produce. If they are not, that is a different problem, and it deserves its own conversation rather than a quiet assumption.

Part three · A company that got it right

Guided, not spoon-fed

Split illustration. On the left, a man in a suit sits in a high chair with a bib, mouth open, as a robotic arm spoon-feeds him a document; crumpled papers and a teddy bear lie at his feet. On the right, the same man stands at a drawing board working under a lamp while a small robot holds up a clipboard with a tick on it.

The argument above is that the risk in AI is rarely the model. It is the architecture around it, and above all whether a competent human sits between the machine’s output and the consequence. Here is a company that has bet its product on it. This is a first-hand case: I know the firm, I have seen the platform work, and what follows is what I saw rather than what the brochure says.

The company

The firm is a UK tendering specialist, founded in the early 2000s by an engineer. For more than twenty years it has run bids as communications campaigns rather than form-filling, on several hundred tenders across engineering, construction, rail and infrastructure, for contracts from £100m to £10bn, and has helped win well over £100bn. I have not named it.

The team is not a software company’s. Bid writers, designers, strategists, researchers, journalists, behavioural coaches and sector specialists: people whose expertise is knowing how an evaluation panel thinks at four o’clock on a Thursday afternoon with eleven submissions still to read.

The platform came out of the consultancy, not the other way round, and the order matters. The software encodes two decades of watching what wins and what gets binned.

The approach

The firm sees the AI bidding market as split into two philosophies. Most tools are built to write the answer: paste in the RFP, receive a draft. This one was built to do the groundwork.

The platform reads the tender and pulls out the requirements, finds evidence in the bidder’s own past submissions, builds an answer plan and storyboard for each response, and scores drafts against the criteria the evaluator will actually apply. It runs inside the client’s own cloud tenancy, so the evidence base is the organisation’s own history rather than a scraped corpus.

What it will not do is answer the exam question. It clears the reading, the retrieval, the structure and the scoring so that the humans can spend their time on the one thing the machine cannot supply: why this team, on this bid, for this client.

On one recent bid, a joint venture of three large contractors chasing a rail programme worth well over a billion pounds, the platform is reported to have saved more than a thousand hours over four months. Nobody used those hours to shrink the team. They used them to sharpen the submission.

Why this is the right design

Consider what a tender actually is. On a large infrastructure or government programme it can be worth billions. It commits the bidder to prices, timelines, liabilities and reputational exposure for years. Every clause is a promise. Every number is a commitment someone will be held to.

Now put to it the question this essay puts to the intelligence analyst: why would you let a probabilistic text generator have the final say on that?

These models, remember, are machines for producing the most likely next word. Not the truest, or the wisest. The likeliest. A wonderful property when you want a plausible first draft; a hair-raising one when the draft is a binding promise to deliver a railway.

Bid for a multi-billion-pound project and you have an enormous amount to lose. A tool that hands you all the answers is then offering to remove the one safeguard that matters: an expert who knows the domain, reads the output sceptically, and can say “this is wrong” before it goes out of the door.

Being spoon-fed feels like help. On a job of this size it is the opposite. The value of a good tendering process is that the people running it know why each answer is what it is. Hide that reasoning behind a generated draft and the bid is no better; the team is simply less able to defend it.

The second reason is the more interesting one. Procurers are steadily homogenising response formats to make evaluation easier and cheaper. The effect is to push every bidder towards the same answer, so the decision can be made on price. A tool that writes the answer for you produces the average response, faster. A tool that does the groundwork and leaves the differentiation to humans is the only sort that helps a bidder stand out inside a format designed to stop anyone standing out.

The other side of the ledger

I have watched the other half of this story too, and it is the half nobody writes up because nobody likes admitting to it.

Bid teams elsewhere are pasting the same tenders into whichever general-purpose chatbot the company already pays for. They get back something fluent, well-structured and finished-looking, in minutes. It reads like a bid. It has the headings a bid has. And it is, almost without exception, the average answer: the response the model has seen most often, stripped of anything that only this team could have said, with the occasional confident claim about a past project that did not quite happen that way. The team reads it, thinks “that’s rather good”, tidies the formatting and submits.

Then they lose, and the debrief tells them they scored well on compliance and poorly on differentiation, and they nod, because that is what the debrief always says. What it does not say is that they handed the exam question to a machine built to produce the likeliest text, and that a panel of people whose entire job is telling submissions apart could tell.

It is the container ship again, in a different harbour. A general model was given a factual, consequential job with no gate. Its output was formatted into something that looked like the real thing. Everyone downstream believed it because it looked like the things they normally believe. The only difference is scale of consequence. Nobody scrambles aircraft over a lost tender. You simply do not win the billion-pound contract, and you rarely find out that the tool is why, because the tool was already approved and the team believed it was doing the right thing.

That is the whole argument in one industry. Two bidders, the same tender, the same underlying models. One built its system to do the groundwork and hand the judgement to experts. The other pointed a chatbot at the question. One of them is winning contracts the size of a small country’s rail budget. The other is wondering why the debrief keeps using the word generic.

Mapped to the pipeline

The architecture lines up with the layered pipeline in Part one:

This essayThe bidding platform
Deterministic extractionTender requirements pulled out of hundreds of pages
Provenance and context governanceEvidence retrieved from the bidder’s own submissions, inside its own tenancy
Specialist modelsTwo decades of sector-specific bid intelligence built into the workflow
Confidence and contradiction checksDraft scoring against the evaluator’s actual criteria
Human judgement over consequenceThe experts write the part that wins; the machine does not answer the exam question

What to ask at the demo

Guidance scales expertise. Spoon-feeding replaces it, and on anything that matters you cannot afford the replacement. The right tool is the one built to need the right people, an old idea we keep having to relearn.

Next time someone demonstrates a tool that promises to write the whole bid, ask who in the room will be able to explain the answer when the client does.

And if you are running bids of this size and want to see the guided approach for yourself, ask me and I will point you in the right direction. I have kept the firm’s name out of print, not out of reach.

Thank you for reading.

Further reading

The earlier essays this one leans on.

Sources: CNN, 18 September 2026, on the AI-generated intelligence report that nearly triggered a US interception of a Chinese vessel, and the outlets that followed it. Part three draws on the unnamed company’s public materials, September 2026.

That was the summary

Read the rest with your email

The full essay is about 25 minutes. Leave your email and the page unlocks here. Nothing is sent: the PDF edition, with contents, page numbers and a glossary, is a separate download at the end of the essay, yours when you ask for it.