---
title: The CEO Gap
source: https://steadman.ai/newsletters/david/the-ceo-gap.html
published: 2026-06-20
summary: Every AI answer sits some distance from the correct one. This is a picture of that gap: how big it is, what it costs to close, and why the checking it needs keeps shrinking.
---

# The CEO Gap

*Check, Edit, Own. And the day you might not have to.*

Published 20th June 2026.

An AI has just handed you a draft. Do you trust it? Send it on, or check every line? You make that judgement dozens of times a day, and it decides how good your work is. The **CEO principle (Check, Edit, Own)** is how we make it today: by hand, on instinct. This page is the picture behind the instinct: how far an AI's answer sits from the right one, what it costs to close that gap, and why the checking keeps shrinking. The interesting question is whether it ever reaches zero, and for which tasks.

## The four bands

Every answer lands in one of four quality bands.

- **Not good enough** — A wrong or off-target answer. Rarer now on a good tool: usually a vague brief, sometimes a real model mistake. Mostly avoidable.
- **Good first draft** — In the right area, but it absolutely needs the full CEO. Cheap to get, expensive to trust as-is.
- **Good enough to send** — Might carry the odd human-level mistake, the kind you'd make yourself. Light CEO, or just press send.
- **Verified** — Known to be right, not just probably right: a calculation the machine re-runs and confirms, a quote matched word-for-word to its source. No reason left to check.

## Chart 1: Closer and closer to the answer

The AI's response climbs toward the correct answer as you put in more effort and spend (the horizontal axis is effort and cost: better prep, better prompt, better process). The space left above the line is **the CEO gap**: how much checking, editing and owning the work still needs.

The four bands sit left to right along the effort-and-cost axis, each narrower than the last: "not good enough" is the widest, then "good first draft", then "good enough to send", then a narrow "verified" band at the right. The curve rises steeply, then flattens, and meets the correct-answer line inside the verified band.

- The not-good-enough band (widest, on the left) shows a **wide gap** between the AI's response and the correct answer.
- The good-first-draft band shows only a **small gap**.
- By the verified band the curve has met the line: **no gap**.

A better AI lifts the whole curve: the same effort carries you further along the bands. Where you choose to stop is set by what the task is worth, not by where the curve could reach. The curve flattens fast, and only in the verified band, where the answer can be mechanically checked, does it actually meet the line. Short of that, a last sliver of doubt remains: the hardest, and often the most expensive, to remove.

## Aim for the right band, on purpose

Which band you're aiming for is a decision you make before you start, and it changes from task to task.

- **Good first draft.** A client proposal, a strategy paper, anything carrying your name and your thinking. For this work a draft is all you want. Take it, then stop and pour human judgement into the editing and the rewrite. That pause is where the value is.
- **Good enough to send.** A meeting summary, an internal note, a routine reply. Aim here, do a quick check, and press go.
- **Verified.** A calculation, a data pull, a quote against its source. For this narrow but growing set, aim high enough that the answer can go out on its own, automatically, without even a light check.
- **Not good enough.** Never the aim. It's wasted effort, and with a clear brief and the right context up front, it's the one band you can almost always avoid.

## Chart 2: How much checking does this task earn?

Two things set how much CEO a task needs: how close the answer is likely to be, and **how much a mistake would cost**. Cross them and the rule for any task falls out. An answer can be probably right and still earn a full check, because the one time it isn't might be the time that matters.

A two-by-two grid. The horizontal axis runs from wide gap to narrow gap (how close is it likely to be?). The vertical axis runs from low stakes to high stakes (what would a mistake cost?).

- **Wide gap, high stakes: THE FULL CEO.** Check every line, edit hard, own it. Example: a client proposal drafted from cold.
- **Narrow gap, high stakes: CHECK IT ANYWAY.** Probably right is not enough here. Example: a board paper, a market-sizing estimate.
- **Wide gap, low stakes: A QUICK LOOK.** Fix the obvious, then let it go. Example: a brainstorm for your eyes only.
- **Narrow gap, low stakes: PRESS SEND.** The odd human-level slip is fine. Example: a meeting summary, a routine reply.

The warm top row is where the checking lives; the cool bottom row is where you let go. Most everyday tasks live in the bottom row. The top row is where your judgement earns its keep.

## Chart 3: The cost of getting there

Take one real task: read four transcripts and six documents (a morning's reading on its own), then draft a proposal with two revision passes. Here is what it costs, two ways. The AI's price is the easy part. **The expensive input is rework: the redo, the second look, the "not quite, try again".**

- **Best model, right first time:** $4.56 all-in, no rework.
- **Cheapest model, then one fix:** $1.05 of model plus ten minutes of you (about $40 of a senior person's time), for about $41 total, almost all of it the rework.
- **Doing it by hand:** 4 to 8 hours, roughly $1,000 to $2,000. Off the top of the chart: 25 to 50 times the bar at right.

If you pay a flat monthly subscription you never see this bill, and the rule is simpler still: give every task the strongest model on offer. For those who do pay per task, the whole spread from the cheapest model to the dearest here is $3.51, less than a minute of a senior manager's time. The moment the cheap model costs you even a minute of redo, the dear one was the cheaper choice. Spend up, and buy most of the rework out of existence. (Numbers from a verified cost model on June 2026 list prices and May 2026 firm usage data; the two models compared are Claude's Fable 5, the dearest, and its Sonnet 4.6, the cheapest in this comparison.)

## Chart 4: Where the work lands, over time

The same picture, three moments apart. As tools improve, the mass of everyday tasks moves out of the left-hand bands and into **good enough to send** and beyond. The **verified** band on the right, all but absent at the start, slowly opens up.

- **Late 2022 (when ChatGPT launched):** not good enough 35%, good first draft 45%, good enough to send 18%, verified 2%.
- **Today (a good model, used well):** not good enough 5%, good first draft 25%, good enough to send 50%, verified 20%.
- **Six months on? (if the trend holds):** not good enough 2%, good first draft 12%, good enough to send 46%, verified 40%.

Illustrative, not measured. Each bar reads left to right in the same band order as the chart above. The shape is the point: the bands don't change, but the work shifts rightwards through them over time, and the checking each task needs falls away with it.

## Two things to hold onto

**The gap is shrinking, but you still own the result.** More work landing in the higher bands doesn't move accountability onto the machine. You direct it, you check where it matters, you own what goes out, exactly as you would for anything a colleague handed you. The further the AI reaches on your behalf, the more it's your judgement, not your absence, that makes the work good.

**Does the top band really exist? Sometimes, and it's growing.** The honest answer is yes, for a narrow but growing class of tasks. A sum the machine re-runs and confirms, a quote matched to its source: there the answer is mechanically known to be right and there's nothing left to check. The open questions are how big that band gets, how fast, and where we're willing to draw the line that says "this part is verified, trust it."

## Try this: make every task tell you what to check

Each piece of work an AI hands back should arrive with its own label: roughly which band it landed in, and what kind of checking it needs. "Here's your answer, it's a good first draft, check the figures." One day the tools will carry that label themselves. You don't have to wait. End your next prompt with one line: **"Which band did this land in: not good enough, good first draft, good enough to send, or verified? Tell me what to check before I put my name on it."** Treat the answer as a pointer, not a verdict: it knows where it already has doubts, and it cannot see its own confident mistakes. Never let its self-grade lower the checking a high-stakes task earns. But it starts your checking in the right places, and it turns the CEO principle from a rule you have to remember into part of the task itself.

---

Part of a connected set. See where tasks sit on the [AI Usage Spectrum](https://steadman.ai/newsletters/david/ai-usage-spectrum.html), and what you're managing at each level in [Three Generations of AI](https://steadman.ai/newsletters/david/three-generations.html).

*Saturday AI Thoughts, by David Boyle. This is an artefact created for David's weekly email: [see the others here](https://steadman.ai/newsletters/david/).*
