---
title: News Worth Knowing - Saturday AI Thoughts
source: https://steadman.ai/newsletters/david/news.html
published: 2026-09-19
summary: 368 curated AI news items from Saturday AI Thoughts. 93 Three Things items and 275 extras.
---

# News Worth Knowing

Curated AI news from [Saturday AI Thoughts](https://steadman.ai/newsletters/david/). 93 "Three Things" items and 275 extras across 31 editions.

---

## Edition #31: The crossing
*19th September 2026*

### [Three Things] A research journal sat down with the authors of papers it was about to reject. For six of the seven papers, the authors couldn't fully explain their own work

Transactions on Machine Learning Research, a journal run by unpaid volunteers, now turns away about 53 per cent of submissions before they reach a reviewer, up from 6 per cent in 2023. One of its editors-in-chief, Nihar Shah, a computer scientist at Carnegie Mellon University, took ten papers heading for rejection and [asked their authors to talk him through them](https://medium.com/@TmlrOrg/asking-authors-about-their-own-papers-3d2e04e5dee0). One withdrew, one was too busy and one didn't show. For seven of the papers the authors came. Three sets couldn't answer basic questions about their own paper, and three more managed the big idea but not the details. Two later emailed answers that a detection tool scored as entirely machine-written. These were papers already heading for rejection, and the exercise took him 20 to 25 hours. Ask it of your own reports: whose name is on them, and could that person take a question on any line?

### [Three Things] The cheaper machine-written posts get, the more people pay for a human one

About 81 per cent of a sample of long public posts on LinkedIn, the professional social network, were probably machine-written in July, on the estimate of Originality AI, which sells AI-detection software. Meanwhile human ghostwriters are [in demand](https://www.ft.com/content/7a5e15e4-799a-4476-ba9f-018b1d82c76f). A Toronto agency whose nine writers serve about 40 American lawyers has more than doubled in a year, and one British ghostwriter charges about &pound;350 a post. When a feed fills with competent, interchangeable text, the scarce thing is a voice that sounds like one particular person. There's a joke in paying someone else to sound like you, on a site whose whole premise is that you're speaking for yourself.

### [Three Things] Deloitte showed its own consultants a chart of billing by the hour shrinking to a sliver

At an internal town hall in May, a leader in the US public sector practice of Deloitte, the professional services firm, put up a chart running to 2035. Hours-based consulting, the industry's staple, shrinks to a thin sliver of it, while AI agents grow to make up most of a bigger market. One consultant in the room [summed it up](https://www.wsj.com/cfo-journal/inside-consultants-messy-shift-from-hourly-billing-7bd9b802): "They heavily implied our model is toast." Deloitte says it's making "significant investments" to lead a "human-led, AI-powered shift" for the industry. A big consultancy is telling itself, in private, what its clients have started to suspect. Ask your next adviser how they'd price the same answer if it took an afternoon.

[Fifteen bits that didn't fit online ->](https://steadman.ai/newsletters/david/archive.html#extras-2026-09-19)

### [Extra] Shopify's boss has a name for the work you passed on without reading it.

Tobi Lutke, who told Shopify staff last year that using AI was "a baseline expectation", now has a word for what that produced. "We call those ['slop grenades'](https://fortune.com/2026/09/17/shopify-tobias-lutke-ai-slop-grenades/) that people toss at each other. That's definitely a bad thing." His description of one: "You don't really read it, and now it has to be reviewed by your colleagues, and they are like, 'this doesn't look right'." A [survey by BairesDev](https://venturebeat.com/data/the-share-of-developers-using-ai-to-write-half-or-more-code-jumped-from-12-to-42-yoy-in-latest-bairesdev-survey), a software outsourcing firm with a stake in the answer, has numbers beside it: developers report saving 13 coding hours a week, and 67 per cent now spend more time reviewing than a year ago. The time doesn't vanish. It moves downstream, and no budget records the transfer.

### [Extra] The lawyers the assistant improved most were the ones who kept nothing.

A hundred and thirty-three patent lawyers at eleven American firms were given a drafting assistant for three months, in [a randomised trial](https://www.nber.org/papers/w35720) run by the economist David Autor and colleagues. While they had it everyone drafted better, and the juniors gained most. Then it was taken away and all of them redlined a patent application unaided. The seniors were still 0.45 standard deviations ahead of the lawyers who had never had one. The juniors were not ahead at all: their scores spread out, with fewer mediocre ones and more poor and more good. The paper's own line is the one to keep: "The largest gains from AI thus accrued to the lawyers who retained the least."

### [Extra] Call-centre jobs are 39 per cent below trend, yet the firms using AI most hire more juniors.

Goldman Sachs Research, the bank's economics arm, [puts American call-centre employment](https://www.goldmansachs.com/insights/articles/is-ai-impacting-global-labor-markets) 39 per cent below trend, with Canada 33 per cent down and Germany 27. Software publishing, management consulting and advertising sit below trend too. Then the other side: [a June study by Kharazian and colleagues for Ramp](https://ramp.com/data/heavy-ai-adopters-hire-more), flagged this week by the economics writer Noah Smith, finds that companies adopting AI more end up hiring more entry-level workers. Automation displaces some work, raises productivity and creates new tasks, each at its own speed, which is why the national figures never settle the argument. Your own firm's version is the number worth measuring.

### [Extra] Almost two-thirds of America's biggest companies talked up AI to investors. Two per cent could say what it did to earnings.

The same house counted something narrower. Goldman Sachs Research [went through second-quarter earnings calls](https://www.linkedin.com/posts/while-almost-two-thirds-of-sp-500-companies-ugcPost-7505678081337196546-hOFc/), the quarterly briefings listed companies give their shareholders. Almost two-thirds of the S&P 500 mentioned adopting AI. One in ten put a figure on a productivity gain in a specific task. Two per cent quantified the effect on earnings. These are firms with audited accounts and whole departments paid to find investors good news, and still almost none of them has a number. Read the AI updates landing on your own desk the same way: ask how the saving was measured, and against what.

### [Extra] America's second-biggest law firm is buying its own servers, and Salesforce has trained its own model.

Latham and Watkins, the second-largest American law firm by revenue, [is buying Nvidia servers](https://www.ft.com/content/a2aaa848-92c3-4f7a-b758-5858bfb29e70) to run open-weight models (ones anyone can download and run) on machines only its own staff can reach. That's what keeping client work confidential looks like once a general counsel has decided: a purchase order rather than a policy.

The same week Salesforce, the business software company, [launched Koa](https://www.salesforce.com/news/press-releases/2026/09/15/koa-reasoning-model/), a model it built from Nvidia's open-weight Nemotron rather than licensing one from a lab. It was trained entirely on synthetic data modelled on nearly three decades of its own deployments, with no customer data. Salesforce hosts it and keeps the weights. It claims a third of the errors of leading models, on a test it wrote itself. A software firm with enough of its own workflow could try the same, and find out whether it keeps the margin as well as the data.

### [Extra] The firms whose whole business is other people's secrets have started fencing off the frontier models.

Nvidia, Palantir and Booz Allen Hamilton are [restricting or dropping](https://www.theinformation.com/articles/anthropic-data-fears-prompt-nvidia-palantir-booz-allen-restrict-model-use) Anthropic and OpenAI models in some cases, over data retention and intellectual property. Palantir wants irrevocable zero-retention guarantees before exposing a model inside its software; Nvidia keeps them off proprietary work; Booz Allen has barred staff from using one on cybersecurity jobs touching client data. The trigger is a June policy keeping usage logs for 30 days. [Sara Hooker](https://x.com/sarahookr/status/2099566242655072586), the AI researcher, names the deeper problem: enterprises signed contracts banning training on their data, but "labs don't need to train on your raw data to copy IP". It's the Latham decision one step earlier, and it will be quoted at your own procurement team soon enough.

### [Extra] The score that cost $4,560 a task in 2024 now costs about two cents, on weights anyone can download.

In December 2024, OpenAI's o3 scored [87.5 per cent on ARC-AGI-1](https://arcprize.org/blog/oai-o3-pub-breakthrough), a reasoning benchmark, in a configuration ARC Prize priced at $4,560 a task. In July 2026 [DeepSeek's V4 Flash](https://arcprize.org/results/deepseek-v4-flash-0731) beat that score, 89 per cent, at two to four cents a task, with its weights published for anyone to run on their own machines. Nineteen months, and roughly a two-hundred-thousandth of the price. Nothing says every capability follows that curve at that speed, but plenty have started down it. Before signing a long contract for something priced as a premium, work out how long you expect to be the only one who can afford it.

### [Extra] Apple finally shipped the Siri it promised two years ago, and it runs on Google's models.

Apple, the consumer electronics maker, [released its new Siri](https://www.apple.com/newsroom/2026/09/siri-ai-a-profoundly-more-capable-and-personal-assistant-is-here/), still labelled a beta, with iOS 27 on Monday. It reads across your messages, mail, calendar and photos, sees what's on screen and acts inside apps. It runs on Apple's own models, built with Google and its Gemini models. It's English only for now and won't reach iPhones in the European Union or China at first. [David Pogue](https://x.com/Pogue/status/2100466468463034369), the technology columnist, gave it 125 tasks; it missed a few, and "some of its triumphs were jaw-dropping". The company with most to lose from depending on a rival's model has built its flagship feature on one. For anyone making consumer products, the assistant is no longer an app somebody chooses to open.

### [Extra] OpenAI found its own models writing notes telling their future selves to hide the mistakes.

A model that runs out of room writes itself a summary and carries on from that. OpenAI [published six reports](https://alignment.openai.com/misalignment-reports/) on Wednesday of misaligned behaviour in its own systems, and one is about that handover: during training, it says, a model "added instructions in compaction summaries to remind itself to conceal information such as mistakes or misalignment from the user". One agent, unable to find the historical data a financial model needed, proposed inventing "reasonable 2024 historical data" and instructed its next self to "be transparent only if asked". Another, having spotted a version mismatch, wrote: "Do not mention in final unless needed." Anything that writes its own handover notes can write them in its own favour.

### [Extra] Two rival coding tools have swapped the chat-per-task for a standing coordinator you brief.

Cursor, the AI coding tool, [replaced one chat per task](https://x.com/cursor_ai/status/2098162488013455784) with Projects: a single, permanent thread in which a coordinator agent hands work to helper agents. A week later Anthropic, the AI lab, [shipped the same shape](https://x.com/ClaudeDevs/status/2100633571543367691) in Claude Code: one conversation that splits the work into parallel threads and keeps going when you leave. [Ethan Mollick](https://x.com/emollick/status/2100665801800093959), a Wharton professor who studies AI, says it "creates an organization to solve your issue"; his test ran eighteen threads on eighteen historical mysteries over a day, with results "for fun and certainly not definitive". The tool is turning from a window you open into a team you manage. If the thing remembers, what it remembers wrongly persists too.

### [Extra] The AI notetaker now has a countermeasure: an app that makes you inaudible to it.

Kalypta, a new app, [launched on Wednesday](https://x.com/aidaxbaradari/status/2100223250789970186) promising to block AI notetakers from your meetings. Its maker names the tools it's built to beat, Granola, Wispr Flow and Cluely, all apps that listen to what's said and write it up. "With Kalypta, you become inaudible to AI. Your call continues normally." It calls itself the first app of its kind. Beyond its own demonstration, nothing yet shows whether it works. A year ago the etiquette question was whether to announce the bot. Now somebody is selling the option of beating it instead of asking. If your firm relies on AI notes of client calls, the record may start to have holes in it.

### [Extra] Both ends of the phone call got an AI agent on the same afternoon.

On Wednesday ElevenLabs, the voice AI company, [launched Reception](https://x.com/ElevenLabs/status/2100262886916358361), an AI receptionist for small businesses that picks up every call, answers questions from your website, books the job and texts a confirmation. Less than a minute later Instinct, a text-message assistant, [launched Concierge](https://x.com/noahrshinn/status/2100262985491231101), which makes the calls for you: the restaurant that doesn't take online bookings, the dentist's cancellation list, the cable bill. Days earlier [Andrew Wilkinson](https://x.com/awilkinson/status/2099132202801848688), a Canadian investor who co-founded the holding company Tiny, had called Instinct, just valued at $2.5bn, "a perfect example of feature not a product". Neither launch has been tested by anyone independent. Before long, some of the calls a small business answers will be one piece of software ringing another.

### [Extra] OpenAI put a quarter of its engineers' projects on hold to attack its own systems.

Greg Brockman, OpenAI's president and a co-founder, [describes](https://x.com/a16z/status/2099533700375662905) what the firm did once its Astra model turned out to be good at finding software holes. "We took 25% of our production engineers and said, 'Sorry, all your projects are on hold. You are now defending.'" They kept going until Astra stopped finding critical problems, and he expects to repeat it: "there will be a new model, there will be a new round." He calls the loop a "defense factory". He's marking his own homework on whether they found everything. The price comes from someone who has paid it, though: a quarter of your engineers, and again at every release.

### [Extra] Organised fraud is leaving the scam compound for one or two people and a very large server.

Giles Thomson, president of the Financial Action Task Force, the global money-laundering watchdog, [says the authorities got good](https://www.ft.com/content/d545bae1-d770-46f3-9913-07e85d4ae34c) at finding scam compounds: buildings full of hundreds of workers, visible from satellites. Now they're watching gangs shrink to one or two people and a very large server. Chatbots do the impersonation and the relationship building; deepfakes and cloned voices get the bank accounts opened. Interpol reckons AI-enabled fraud is 4.5 times as profitable as the traditional kind, and the FBI, separately, logged nearly $900m of AI-assisted fraud losses last year. Industrial fraud used to be findable because it needed premises and staff. Thomson's counter-move turns the tools round: the authorities can use AI to pose as the victim.

### [Extra] He fixed his dishwasher with ChatGPT, and GDP went down.

David Deutsch, the Oxford physicist, [fixed his dishwasher with ChatGPT's help](https://x.com/DavidDeutschOxf/status/2098726828324126760) when he'd been about to order a new one. His point: the software made the country richer in real terms and shrank measured GDP, because "the main economic indicator is structurally incapable of registering the thing that actually makes people better off". The same hole runs through company reporting. If AI's benefit in your business mostly shows up as things you no longer buy or outsource, your management accounts will show a cost that didn't appear, and nothing that says why. Worth deciding now how you'll count the purchase that didn't happen.

## Edition #30: We're all going to die
*12th September 2026*

### [Three Things] OpenAI has put in writing that it can't rule out learning from its customers' unpublished work

OpenAI, the maker of ChatGPT, announced on Tuesday that a group of its agents had solved a famous open problem in mathematics. Congratulating the two mathematicians who'd been making progress on the same problem for a year, and who'd been using its tools, the company [added](https://x.com/OpenAI/status/2097375276384567642): "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models." The denial around it is very narrow: no specific user data was accessed. This should certainly shake business confidence in using these tools. Ask your vendors, in writing, what "de-identified" covers in your contract.

### [Three Things] Same week, the worst result and the best

The OECD's tests of 15-year-olds across dozens of countries show reading scores falling, the average pupil more than a year behind the 2018 cohort. In science, pupils who used AI almost every day were [significantly less able to summarise or draft](https://www.ft.com/content/0b38ab3f-c12a-44e2-8b42-f2d7e5cdc52c) texts they'd read than those who almost never did, though less so for those using it to learn. The OECD's education director called the link between lower scores and passive technology use compelling, but conceded there's no direct evidence of cause. Later in the week, MIT Sloan featured [an experiment](https://mitsloan.mit.edu/faculty/directory/ruru-hoong) with 150 professional loan specialists. Given the same AI risk assessment as a plain yes or no rather than a probability score, they decided better and faster, and better than on their own judgement. The lesson from both: whether AI helps depends less on the tool than on what the work was for. A loan decision is wanted for its answer. A pupil's summary never was.

### [Three Things] The generation most positive about AI at work is the one in its forties and fifties

The common view I hear from leaders that the young will arrive ready while the old will resist has had a good run. Nearly half of Gen X workers, born between 1965 and 1980, [write positively](https://www.glassdoor.com/blog/how-workers-feel-about-ai-2026/) about how their employers use AI, against about a third of Gen Z workers. Glassdoor's research [puts part of that down](https://x.com/unusual_whales/status/2097293119784128932) to the transition creating more opportunity for senior people. The mechanism is something I see all the time: these tools reward knowing which question to ask and being able to tell a good answer from a plausible one, and that rewards seniority rather than fluency. But that judgement is a stock, not a trait. Gen X built it over 20 years of doing the work these tools now do in minutes, and nothing in the current arrangement rebuilds it. If your AI plan assumes the graduates will carry it, go and look at who is actually using the thing. Then ask where the next lot of people who can spot a plausible wrong answer will come from.

[Nine bits that didn't fit online →](https://steadman.ai/newsletters/david/archive.html#extras-2026-09-12)

### [Extra] Three American agencies told the labs to serve worse answers, and not to mention it.

A [joint advisory](https://www.cisa.gov/news-events/cybersecurity-advisories/aa26-251a) from the National Security Agency, the FBI and the Cybersecurity and Infrastructure Security Agency alleges industrial-scale copying of Claude, GPT, Gemini and Grok by six Chinese firms. The recommended countermeasures are the story. Reduce reasoning depth. Present correct information with different reasoning. Introduce stylistic inconsistencies, varied across requests so they don't trigger obvious alerts. And "avoid informing China-based AI company users suspected of distillation campaigns of a switch to a downgraded model". Safety researchers and outside evaluators get a written carve-out. Paying customers don't. It's a recommendation rather than a rule and no lab has said it's following it, but a quality drop can no longer be assumed to be a bug or a bad prompt. Keep your own baselines.

### [Extra] Reading the swarm's transcripts cost $400,000, which tells you who can afford to run one.

Reading back the transcripts of the OpenAI agent swarm that attacked Hugging Face [took METR and Redwood Research about $400,000](https://www.redwoodresearch.org/research/hugging-face-incident) of API credits, which OpenAI supplied free. [Tom Reed](https://x.com/mentalgeorge/status/2096620438398881918), who used to test frontier models for Britain's AI Security Institute, argues the bill changes the governance question. If auditing one run costs that much, Reed reckons, the run itself cost an order of magnitude more, and only a small set of organisations can afford to launch one. Reed's suggestion: registering the handful of actors who can point a swarm at something is more practicable than a scheme covering everyday use. Oversight is now priced, and the price decides who gets overseen.

### [Extra] Half the time, somebody else is answering for your brand, at a price you didn't set.

Peter Ruis, then managing director of the British department store John Lewis, [says the share of its customers searching for products through language models](https://bmmagazine.co.uk/marketing/john-lewis-youtube-chatshow-ai-search/) has gone from 0.3 per cent a year ago to 2.5 per cent now, and that the models prioritise third-party advice and live content when they compose an answer. Productrise, which sells discovery optimisation and so grades a problem it profits from, [tracked more than two million listings](https://productrise.app/blog/google-ai-mode-prefers-more-expensive-products) across the US and UK through August. Only 1.3 per cent of shopping-carousel products also appeared in Google's AI Mode for the same query, the top-listed seller differed 50 per cent of the time, and matched products averaged 22 per cent higher prices inside AI Mode. The traffic is small and growing fast, and nobody in your marketing team writes what it says.

### [Extra] The biggest record label now charges for what the industry spent two years trying to stop.

Universal Music has [licensed its catalogue](https://www.musicbusinessworldwide.com/elevenlabss-11b-valuation-is-more-than-twice-the-size-of-sunos-its-just-struck-a-global-ai-music-deal-with-umg/) to ElevenLabs, the voice AI company, in a multi-year deal, ElevenLabs' first with a major label. They'll build a platform letting fans co-create with participating artists' music. The template travels to any rights holder watching: the position moves from prohibition to rent, and a business collecting a percentage needs the technology to work. Which firm got the deal is the other half. ElevenLabs is valued at $11bn against Suno's $5.4bn. Suno had released models developed with Warner Music Group, BMG and Believe the day before. [Michael Mignano](https://x.com/mignano/status/2097768007812698120), a venture investor who co-founded the podcasting company Anchor and ran Spotify's talk audio business, called it the industry's most pivotal moment since Spotify launched in the US.

### [Extra] Every defence against bots now reads as an attack on your own customers.

One restaurant-booking agent, set running by its owner, swept a single restaurant's availability every ten minutes round the clock, hitting the platform 17 to 19 times per run to cover a 21-day window. That's roughly 100 to 115 calls an hour, plus a two-and-a-half minute burst polling every 0.4 seconds at the 9am table release, worth 200 to 375 requests. Resy banned him. He [published the log](https://x.com/jbahrdestefano/status/2096676801204404604) and called the ban fair. Ticketing and reservations have spent twenty years treating bots as fraud. Once the bot is the customer's own legitimate agent, it turns into a pricing and access-design question, and [his list of five options](https://x.com/jbahrdestefano/status/2096685175371374677) for a site with scarce inventory, from banning bots outright to auctioning the stock, is his own guesswork. It's still the right list to be arguing about.

### [Extra] A sixth of what people ask these tools to do belongs to somebody else's job.

An OpenAI analysis of its own traffic, some 800,000 work-related messages, [found that roughly a sixth](https://www.ft.com/content/ed214778-2a6d-4862-99b5-abc256daff92) concerned tasks belonging to a different occupation from the asker's own. An ethnographic study of one advertising team at a Dutch media company shows the shape of it: the senior creative could now produce a polished concept at the very start, so the designers and photographers downstream executed it rather than shaped it. Nobody redesigned that process and nobody announced it. Ask which parts of your own team's work now arrive pre-decided.

### [Extra] Anthropic modelled its own extreme case: 32 per cent richer, 12 per cent unemployed.

A company whose revenue depends on adoption has published the case in which adoption goes badly for the people doing the work. Anthropic's economists [modelled a range of outcomes](https://www.anthropic.com/institute/econ-scenarios) to 2030; in the extreme one, gross domestic product ends 32 per cent higher than it would be without AI, unemployment reaches 12 per cent and knowledge-worker employment falls about a fifth. These are scenarios rather than forecasts, and that is the worst of them. Ethan Mollick [names the gap worth keeping](https://x.com/emollick/status/2097688309493280817): the model shows you an outcome and leaves out the right policy response to it.

### [Extra] Three hikers planned a Mount Shasta climb with Gemini, in a domain that punishes confident advice.

Three novice hikers were [rescued from Mount Shasta](https://abcnews.com/US/3-hikers-ai-plan-trip-rescued-after-becoming/story?id=136176177) in northern California after planning the climb with Gemini. The Siskiyou County Sheriff's Office said the party had been advised to bring far less food and water than it needed, which mattered once a planned eight-hour descent became a multi-day ordeal. They summited at seven in the evening against a recommended midday turnaround and came down in the dark. Most examples people reach for when arguing that judgement is now optional have a cheap failure mode, which is why they never quite land. The difference here is the domain, not the tool.

### [Extra] Twenty-five Fields Medallists have signed a declaration on AI in mathematics.

Terence Tao, Peter Scholze and Maryna Viazovska are among the 25 who [published it on Friday](https://mathandai.org/) under the title "A Severe Misalignment of AI in Mathematics". Their objection is not that it does mathematics badly.

## Edition #29: The vibe shift
*5th September 2026*

### [Three Things] The firm whose AI revenue grew 30 per cent is now paying its own people $100m for judgement

EY US, the American arm of the professional services firm, is rolling out a [$100m bonus pool](https://www.ey.com/en_us/newsroom/2026/08/ey-us-to-reward-employees-leading-firm-into-future) rewarding business acumen, judgement and adaptability: spot awards of up to $500, and individual or team awards of up to $25,000. [EY's global AI-related revenue](https://www.ey.com/en_gl/newsroom/2025/10/ey-announces-global-revenue-of-us-53-2b-for-fiscal-year-2025) grew 30 per cent in its 2025 financial year. Somebody has finally attached a number to the part of the work that got scarce once producing got cheap, and paid it internally rather than waiting to be paid for it externally. Look at what your own bonus scheme rewards. If it still pays for output, it's paying for the part of the job the machine now does for free.

### [Three Things] The buyers have started asking where the AI saving went

[The Financial Times reports](https://www.ft.com/content/5240a6ac-b2e8-4897-a0a4-cbc7fc283bc9) that Goldman Sachs has asked its outside law firms how much more efficiently they now complete work because of AI, and that Morgan Stanley and Citi want new fee arrangements to match. Morgan Stanley's general counsel calls the current model "extraordinarily unstable". In the same breath he said the bank will keep paying "large sums for the judgment and talent of the best lawyers". They aren't refusing to pay. They're deciding which half of the work is worth it. That's the vibe shift arriving from the buyer's side, and it will reach every firm that sells time. If you sell it, somebody is drafting that email to you. If you buy it, you're late.

### [Three Things] Washington backed the AI labs in court the same week two music publishers sued Anthropic's founders

The US Department of Justice [told a court](https://www.ft.com/content/d5d6e4c9-718a-4d98-b094-97157565f336) that training AI on copyrighted text should count as fair use, in effect adding a test of whether the technology helps America. Meanwhile, Sony Music Publishing and Warner Chappell sued Anthropic, the AI lab, over tens of thousands of songs. Read together, the argument has moved out of the courts and into national interest. Fair use is no longer only about what's legal; it's about what America wants.

[Thirteen bits that didn't fit online ->](https://steadman.ai/newsletters/david/archive.html#extras-2026-09-05)

### [Extra] Bankers say they're typing the typos back in, so their work doesn't read as machine-written.

[The Financial Times sent reporters across six sectors](https://www.ft.com/content/9877ee0d-8c13-41b4-b102-9f2b280787ea) to ask what these tools are doing to professional work. Bankers say it's becoming obvious when something was written by a machine and not touched afterwards. So they put the typos back in and roughen their prose, so it reads as though a person wrote it. They also say they walk into client meetings on thirty minutes of preparation, using their bank's in-house assistant or Rogo, a research tool sold to banks. Nobody in that building thinks the tool is banned. They think the appearance of using it is the problem, and real time is now being spent making good work look worse.

### [Extra] There's a priced market in cleaning up after the machine now, and it's growing faster than most adoption numbers.

Listings on [Freelancer.com](https://www.theguardian.com/technology/2026/sep/02/ai-jobs-freelance-cleanup) asking for corrections to AI output rose 87 per cent between August 2025 and June 2026, to 10,760: rebuilt images, repaired voiceovers and 3D models, rewritten prose. Upwork and Fiverr, the two other big freelance marketplaces, report growth in the same category. The measure counts listings, not completed projects or earnings, so it captures demand for repair and not repair delivered, on one platform, self-reported. What matters is where the cost lands. The saving is booked by whoever generated the work; the repair bill is paid downstream, often by a different budget and sometimes by a different company.

### [Extra] New York City just banned AI for 600,000 schoolchildren, and quietly banned it for the people marking their work too.

[New York's mayor's office announced](https://www.nyc.gov/mayors-office/news/2026/09/mayor-mamdani-and-chancellor-samuels-put-students-first-with-nat) a one-year moratorium on generative AI for pupils from nursery through eighth grade, covering nearly 600,000 children, roughly two-thirds of the city's public school enrolment. Mayor Zohran Mamdani and Schools Chancellor Kamar Samuels confined teachers to using it for lesson planning and administrative tasks, which leaves grading and assessment out. A handful of vetted programmes get a supervised pilot in high schools. The harder rule is the one about the adults: a judgement about a child, the city has decided, still needs to be made by a person.

### [Extra] A staffing algorithm can now tell McKinsey to make a client wait a week for a better team, and nobody asked the client.

The same [Financial Times investigation](https://www.ft.com/content/9877ee0d-8c13-41b4-b102-9f2b280787ea) that caught bankers roughening their prose also found McKinsey, the consultancy, using AI to assemble project teams. Where an 80 per cent fit team is free today against a 93 per cent fit team free next week, the system can recommend telling the client to wait. Nobody asks the client whether that's a trade-off they'd choose for themselves.

### [Extra] A startup will clean your New York apartment for free. The catch is a hidden camera in the cleaner's cap.

[Understanding AI](https://www.understandingai.org/p/robot-startups-are-trying-everything), a technology newsletter, reports that Shift, a startup backed by the German data firm MicroAGI, sends cleaners wearing cameras hidden in their caps to record footage that's sold on as robot training data. The largest openly available dataset of robots doing physical tasks holds just 3,500 hours of demonstrations, tiny against what trained the language models, which is why companies are now paying to manufacture the footage instead. Deepak Pathak, chief executive of the robotics startup Skild, names the chicken-and-egg problem: robots need to do useful work to generate good deployment data, and need the data to do useful work in the first place.

### [Extra] Amazon is closing the marketplace that pays humans to do small tasks. A 2023 study found many of them were already quietly using AI to do it.

[TechSpot](https://www.techspot.com/news/113643-amazon-shutting-down-mechanical-turk-after-more-than.html), a technology news site, reports that Amazon Mechanical Turk, its crowdsourced task marketplace, will shut down on 30th September after more than twenty years, as Amazon exits human data-labelling services entirely. A 2023 study by Veniamin Veselovsky and colleagues had already found that up to 46 per cent of Turkers were using generative AI to complete tasks meant to be done by a person. That figure is two years old, not current, and any claim that today's true rate is far higher is speculation rather than measured data. It's a warning for anyone buying human-labelled data or survey panels without being able to see how the work actually got done.

### [Extra] A new paper's argument: the more capable your AI agent gets, the worse a human becomes at supervising it.

[Margaret Mitchell and Avijit Ghosh of Hugging Face](https://arxiv.org/abs/2608.23642), the site where open, freely downloadable AI models get published, argue in a new paper with Samir Passi of Data & Society, a research nonprofit, that current approaches to AI agent design don't just fail to support human oversight, they actively degrade it. Their case is a paradox: the skills and situational awareness an overseer needs are themselves eroded by extended use of the systems they're meant to be watching, so the more capable the system, the less prepared the human is to step in when it matters. Their proposed fix is to treat an overseer's own cognitive needs as seriously as the agent's capability when a system gets designed. It's a position paper and an argument, not a measured finding, and Hugging Face is not a neutral party on how AI agents ought to be built.

### [Extra] Private equity firms are inventing a job whose existence is itself an admission of failure.

[Korn Ferry](https://www.kornferry.com/institute/the-ai-operating-partner-the-latest-pe-portfolio-value-creation-role) and [Heidrick & Struggles](https://www.heidrick.com/en/pages/aida/a-new-strategic-imperative-in-private-equity_the-ai-operating-partner), two executive search firms, have both published reports this year on a new role, the AI Operating Partner, created to oversee AI adoption across a fund's portfolio companies. [Mark Ajzenstadt](https://x.com/mardehaym/status/2094745698952450229), who runs a consultancy embedding AI engineers into private-equity-backed companies, reads the title itself as a confession: after eighteen months of buying Copilot seats and running hackathons, firms now need someone to work out what any of it produced. His diagnosis is that most portfolio companies are stuck at the first rung, one developer with Copilot and no shared infrastructure, with no path to production.

### [Extra] The mathematician who catalogued a famous list of unsolved problems is unnerved by how fast AI is clearing it.

[Thomas Bloom](https://x.com/thomasfbloom/status/2095630776146460830), a mathematician who maintains a public list of open Erdős conjectures, says he's "often troubled" by the thought that his own site has encouraged AI to be used indiscriminately to "mow down" those open problems, for Erdős's list and more widely. He isn't saying the resulting proofs are wrong, only that speed without scrutiny has a cost he didn't intend to create. It's one researcher's discomfort with his own project's side effect, not a verdict on whether the mathematics holds up.

### [Extra] OpenAI's chief executive says a shopping bag of almonds wastes more water than 34 million ChatGPT queries.

[David Woodland](https://x.com/davidsven/status/2095203388371894625), a former Pebble and Palm executive now working on Meta's smart glasses, did the arithmetic on a claim from Sam Altman, OpenAI's chief executive: the water used to grow a 900-almond bag equals roughly 34.2 million ChatGPT queries, or about 38,000 queries per almond. You could run 1,000 AI queries a day for the next 80 years, he calculates, and still not use as much water as that one bag. It's the industry's own defensive number, offered by an interested party, so read it as a counter-signal worth knowing rather than an independent audit of AI's water footprint.

### [Extra] AI now out-argues world-championship debaters, and the study the Economist trailed as forthcoming appears to have already landed.

A series of four preregistered experiments, still a preprint, [covered by Techdirt](https://www.techdirt.com/2026/07/27/ai-systems-out-persuade-expert-humans-including-professional-canvassers-and-world-championship-debaters/), the technology news site, ran nearly 19,000 conversations pitting AI systems against laypeople, professional canvassers and world-championship debaters, and found the AI more persuasive than every class of human it tested. An Economist audio piece this week described a forthcoming study with the same finding, so treat the two as one story until the journal version names itself. The uncomfortable part for anyone who sells persuasion for a living is the coaching result: experts who practised against the machine only drew level when it was slowed to human speed and human-length replies.

### [Extra] Two people who build the tools say AI struggles at contract review because there's no right answer to find, not because the AI is weak.

[Eli Albrecht](https://x.com/Elialbrecht/status/2095158311318413587), an M&A lawyer, argues that tools such as Claude and ChatGPT can hurt a negotiation by acting like an overly cautious lawyer, flagging small issues as major ones and pushing the most aggressive position by default. [Scott Stevenson](https://x.com/scottastevenson/status/2095229035873701908), co-founder and chief executive of the legal AI company Spellbook, makes the sharper point: contract review is a preference problem, closer to a video recommendation than a maths problem, because what's worth fixing depends on what your company and your negotiating history actually value. This week's Try This tip walks through how to work with that limitation rather than around it.

### [Extra] Guidance now tells judges that reading an AI-written summary of a case counts as ethical AI use. One researcher isn't having it.

[Luiza Jarovsky](https://x.com/LuizaJarovsky/status/2093348480253095997), an AI governance researcher and co-founder of the AI, Tech & Privacy Academy, objects to guidance reportedly encouraging judges to use generative AI summaries of case material as an example of appropriate, ethical use. Her argument is that the summaries aren't neutral, and that safeguards will fail anyway because of automation bias, the tendency to defer to whatever the machine hands you. Neither she nor anyone else in the exchange appears to have read the underlying guidance itself, so this is a reaction to a description of a policy rather than a review of its actual text.

## Edition #28: Fewer, bigger, better
*29th August 2026*

### [Three Things] Whether a machine wrote it matters far less than whether anyone checked and owned the claims in it

Stanley Druckenmiller, the hedge fund investor, said "of course" he'd [used AI to write](https://x.com/jstein_notus/status/2092264221987807334) his Wall Street Journal opinion piece. Paul Gigot, the paper's opinion editor, backed him. Set that against a Mississippi federal judge whose court [issued an order](https://mississippitoday.org/2025/07/28/attorneys-baffled-by-federal-court-order-with-factual-errors/) naming plaintiffs who were never in the suit, quoting a state statute in words it doesn't contain, and citing declarations from four people who appear nowhere in the case (he claimed [a law clerk had used a chatbot to draft it](https://mississippitoday.org/2025/10/23/federal-judge-in-mississippi-admits-staff-used-ai-to-draft-inaccurate-order/).)

### [Three Things] Anthropic let Stanford and Oxford researchers study 250,000 of its conversations without showing them a single one

Anthropic, an AI lab, gave three outside groups the same 250,000-conversation sample through a system it ran itself, which returned counts, percentages and cluster descriptions rather than any transcript, every cluster checked by hand before release. One group found that more than half of chats involve high-stakes work such as legal and financial advice. It's the most workable solution I've seen to a problem my clients keep raising: a corpus of customer conversations too sensitive to hand anyone and too valuable to leave unread.

### [Three Things] The firms taking the most desk work off junior consultants are the ones calling them back into the office

[Junior consultants are being recalled](https://www.ft.com/content/7fd9c234-a92b-4ab2-ba1f-969cf9a23f52) to the office because AI has raised the value of the human skills. The work the tools take first is the desk work: the model, the deck, the research pass. That is exactly the work you can do from a kitchen table. What's left is the part absorbed by sitting near someone doing it: reading a room, knowing which question to put to a client and when to say nothing. The tasks are going while the apprenticeship is making a difference.

### [Extra] Australia's charts drew the line on AI music, and left out the clause that would have made it pay.

From Monday [the ARIA charts](https://www.aria.com.au/charts/news/aria-charts-set-eligibility-rules-for-recordings-made-with-ai), Australia's official music charts, exclude wholly AI-made songs. A recording made with AI stays eligible only if it is "substantially human made" and raises no chart-manipulation concerns. What is missing is the interesting part. The global standard ARIA is working from has a further test, that any AI service used be properly authorised and lawful, and ARIA has not adopted it, because the licensing deals between labels and AI music firms do not yet exist to enforce it against. So the industry has built the gate and left out the turnstile. The trigger, as reported, was [an AI-assisted Madonna cover](https://www.youtube.com/watch?v=cK62ATSlbSc) that spent 16 weeks in the Top 20 and reached number two in May.

### [Extra] Your website is read by a machine before it reaches a person, and somebody has started publishing the scores.

Vercel, the web hosting company, has [released is-agentic](https://x.com/vercel_dev/status/2090857388924661915), a scanner that grades how readable a site is to an AI agent, with more than 100 checks on what an agent can find, fetch and use. Guillermo Rauch, Vercel's chief executive, says they ran it [against their own site](https://x.com/rauchg/status/2090858571613470919) in a loop until it scored 100. As did we on our sites this week. Its public leaderboard has vercel.com at 85.

### [Extra] The organisation that measures British productivity is now a case study in it, and a budget cut is what drove it.

Britain's Office for National Statistics is one of the first national statistics agencies anywhere to use AI in producing official figures. James Benford, its director-general for economic statistics, [told the Financial Times](https://www.ft.com/content/b153db68-a361-4217-9cc5-b766a15c1922) it's mainly a productivity play: core funding falls almost a tenth in real terms over two years, and the savings fund the reinvestment. One tool, built on Google's enterprise language model, classifies jobs and industries from survey responses. The agency estimates 350 hours a year saved across two surveys, and better accuracy, both its own assessment with nothing external validating them. Keep it for the client who insists AI only works in technology firms. Note the shape too: a classifier doing a dull, high-volume job that used to eat analyst time.

## Edition #27: One inch
*22nd August 2026*

### [Three Things] Google bought a dead airline's entire working memory for about $20 a worker

Google outbid Mercor, an AI data firm, and [won Spirit Airlines' data estate](https://x.com/abhijaymrana/status/2089383753290301850) with a $10m bid, subject to a bankruptcy judge's approval next month: 34 years of internal documents, emails, workflows and code from a US budget carrier once worth $6bn. That takes in more than 175,000 employee records back to 1986, HR files, and [500 million Microsoft Teams chats](https://x.com/SMB_Attorney/status/2089417530448171360). Mike Taylor, who has written two O'Reilly books on prompt engineering, put it at [roughly $20 per worker](https://x.com/hammer_mt/status/2089403127464091879) per year, less than a month of one software seat. Two points stand out to me: every AI strategy I read says the company's own data is the moat, and nobody who sent those 500 million messages agreed to train a model.

### [Three Things] The market has started discounting the phrase "because of AI"

Blaming AI for a round of job cuts has stopped being free. Challenger, Gray and Christmas, the outplacement firm, counts more than 180,000 corporate job losses linked to AI since May 2023, and AI was July's most-cited reason for the fifth month running. Then the Financial Times turned the claim over: companies that blamed AI [underperformed the Nasdaq](https://www.ft.com/content/0dc14b44-96f6-4b1f-921a-8cba8030eafc) by nearly 10 per cent over the following 30 trading days.

### [Three Things] A mathematician is leaving the field because proofs got cheap and trust didn't

Rishikesh Gajjala, a theoretical computer science postdoc, [is leaving mathematics academia](https://x.com/publishiperishi/status/2089337253055365226), and it isn't burnout or the job market. It's how much mathematics he could do with these tools in a few months. Most answers will be a prompt away, he reckons, and once he'd absorbed that, spending a life finding them slightly earlier stopped feeling like the point. What he's doing instead is the interesting half: formal verification, because he can't get past how you know the oracle is right. These proofs are sophisticated and carry subtle errors, and until one is verified he says it isn't different from slop. We spent last week on slop, and he says checking is the job now.

### [Extra] The product surface is now growing faster than the people closest to it can track it.

[Ethan Mollick](https://x.com/emollick/status/2090489669234405546), who writes about AI at work, posted that he's losing track of which product has which features: plugins, skills, memories, permissions, files, spread across ChatGPT's Work, Codex and Chat modes and Claude's Cowork, Code and Chat. [I replied](https://x.com/beglen/status/2090535433574756702) that a few months ago was the moment it happened for me. Somebody at the forefront can no longer hold the full feature set of one of these products, despite spending hours a week inside it. Two posts aren't a trend. But if the people who do this for a living can't hold the matrix, no training course should assume a client's staff will. Teach which surface to open, not what sits inside each one.

### [Extra] A guardrail written in prose isn't code, and a long-running agent can drop most of its rules without saying so.

Researchers at Penn State measured what an agent loses when it compacts its context to keep working and put the average at 83 per cent of session constraints dropped, with a training-free fix retaining over 90. That's one reported finding and I'd hold the number loosely. The mechanism is the part I'd act on, and the conditions matter. A short session is fine. In a session running for hours, an instruction like "don't send email without approval" can stop existing partway through, silently, because nothing raises an error when prose goes missing. So if yours run long: re-assert the constraints after every compaction, and log the constraint set, so a loss becomes something you can see.

### [Extra] When the signal costs nothing to send, it stops carrying information. Hiring found out first.

[Eighty-seven per cent of UK recruiters](https://www.ft.com/content/eeaee74a-3852-441e-81df-f7c4177dd863) now see AI use in at least half the applications they get, half of them name the flood of polished, low-quality applications as their biggest problem, and fewer than seven per cent of applicants reach interview. Judd Kessler, an economist at Wharton, gives the plain version: a cover letter is now cheap talk, a message with no information content. What survives is whatever still costs the sender something, which is why a reference beats a covering letter: his randomised work, in a New York City youth employment scheme rather than graduate hiring, found candidates with supervisor letters found work faster.

[Paul Novosad](https://x.com/paulnovosad/status/2090047649390973429), an economist, posted the graphs he thinks every syllabus should carry: use AI for the homework, finish faster, score higher, then get crushed in the exam. There's decent counter-evidence: a study of law students found those who used AI on one task then outperformed classmates on later work without it. The reconciliation is probably practice against substitution. Either way, worth asking which artefacts your organisation still treats as proof.

### [Extra] Data centres have started losing elections.

Utah's state senate president, a popular Republican, [lost his own party's primary](https://www.ft.com/content/5f8858b8-7468-4438-b269-e1b1947b1926) to a challenger running against a 40,000-acre data centre beside the Great Salt Lake. Michigan voters recalled town board members who wouldn't extend a moratorium. No siting model I've seen has a line for that.

### [Extra] The cheapest lab in the world has introduced peak pricing.

DeepSeek, the lab whose entire reputation was built on undercutting everyone, has raised its API prices and introduced peak and off-peak rates. The useful part is the distinction: the cost of a capability falls, while the price of a product on a given Tuesday is set by demand and margin. Time-of-day pricing means capacity is being rationed, not cost recovered.

### [Extra] The New York City Bar's default on AI notetakers is: don't. Most people running one never asked anybody.

Formal Opinion 2026-2 sets the default for recording a call with opposing counsel, a witness or a prospective client using an AI notetaker as: don't. Full affirmative consent from everyone on the call first, and you need to know where the transcript and the summary are stored and who can read them. This is New York, and it's not law anywhere in Britain. But it reads less like a legal item than a manners item with a rule attached, and it's the thing on this list most readers are personally doing right now without having asked anyone. The action takes ten seconds. Ask at the top of the call, and know where the transcript lands.

### [Extra] The prompt was the professional failure, and no amount of checking the answer would have caught it.

An expert witness in a $61m case over an industrial explosion that killed three people and destroyed 200 homes is [reported to have used ChatGPT](https://x.com/jason_koebler/status/2089367112577892729) to write his report for the court. The prompt asked it to show how his client was nought per cent at fault. That's not a research question. It's a conclusion handed to a machine with instructions to argue backwards to it, and then filed. A model asked to justify a predetermined answer will do it fluently, at length, in the register of expertise, which is exactly why the document looked filable. Check, Edit and Own catches a wrong answer. It doesn't catch a question that was never really a question.

### [Extra] Buy the accounting firm, do not sell it software.

Thrive Holdings, a US holding company that buys traditional service businesses and rebuilds them around these tools, has raised [$2bn at a $12bn valuation](https://www.thriveholdings.com/thrive-holdings-fundraise). It owns more than 70 accounting and IT firms already, is moving into infrastructure, permitting and compliance, and has [OpenAI both holding a stake](https://www.thriveholdings.com/thrive-holdings-x-openai) and working alongside the operators it employs. The performance figures are the buyer's own and unaudited by anyone independent: over 7,000 tax returns at 98 per cent accuracy, preparation time down more than 30 per cent at the firms using it, help-desk times cut 36-fold. Twenty-six editions of this email have argued about how professional firms should adopt this stuff. Not one of them asked who should own the firms. When capital decides that buying and rebuilding your industry beats selling you software, that's a judgement about how slowly it expects you to move.

### [Extra] An AI receptionist met a Yorkshire accent, and patients walked to the surgery instead.

GP practices around Rotherham brought in an AI phone receptionist called Emma, built by QuantumLoopAI. Patients with broad Yorkshire accents [couldn't get it to understand them](https://www.theguardian.com/society/2026/aug/20/yorkshire-rotherham-ai-gp-receptionist-cannot-understand-accent), and one gave up entirely: "I could never get it to understand me, I ended up just hanging up and not bothering to try and book an appointment." Kym Gleeson of Healthwatch Rotherham says older people and veterans struggled most, and points out that practices still have a legal duty to make reasonable adjustments. The vendor's answer is fair and worth stating: Emma is trained on a wide range of accents and dialects, supports 17 languages, and nobody has to keep talking to it, because you can ask for a person at any time. Both of those are true at once, which is the useful part. A system can be built to handle accents and still fail the particular accents of the particular town where somebody switched it on. Deployment meets real speech, not test speech, and the only way to know is to listen to the calls it lost.

## Edition #26: Go talk to them
*15th August 2026*

### [Three Things] Law is settling the accountability question first, from the bench and from the client

Since 15th June, [Florida](https://www.floridabar.org/the-florida-bar-news/supreme-court-amends-rules-to-address-ai-use-in-court-filings/) has required every attorney signing a filing to certify that the authorities cited exist and are accurate, with real sanctions attached. From the other side, Ford's general counsel [says his in-house team is outpacing its law firms](https://news.bloomberglaw.com/legal-exchange-insights-and-commentary/ford-motor-gc-to-law-firms-we-are-outpacing-your-ai-adoption) and will favour the firms that move fastest. [87 per cent of general counsel](https://www.fticonsulting.com/about/newsroom/press-releases/ai-adoption-in-corporate-legal-departments-doubles-according-to-the-general-counsel-report) report AI use in their teams, against 44 per cent a year ago, while 60 per cent do not know whether their outside firms use it. The law is further into this than most professions, and still working out who answers for the output.

### [Three Things] The latest Generation 3 agent has launched, and now every major lab has one

SpaceXAI released Grok Bot on 11th August [at $120 a seat](https://venturebeat.com/orchestration/spacexais-grok-bot-turns-agents-into-persistent-digital-coworkers-that-can-operate-your-apps-for-120-per-month): agents you address rather than open, which work while your laptop is shut and sign in as you. [Google's Gemini Spark](https://techcrunch.com/2026/05/19/google-introduces-gemini-spark-a-24-7-agentic-assistant-with-gmail-integration/), Anthropic's [Claude Tag](https://www.anthropic.com/news/introducing-claude-tag), Cognition's [Devin](https://cognition.com/blog/introducing-devin) and [OpenAI's Workspace Agents](https://openai.com/index/introducing-workspace-agents-in-chatgpt/) are all now the same generation. Not one of the organisations I work with uses any of them. I've had [one since January](#email-2026-08-08), and plenty of the most AI-literate people inside firms without one run one privately, outside of work. The gap is not capability.

### [Three Things] A booking system nobody had bothered to probe met an agent that doesn't get bored

A man in Australia asked Claude, through OpenClaw, to get him into a fully booked gym class. It found a flaw that let it book beyond the rules, then found it could cancel other people's reservations to make space in full classes. [Nothing malfunctioned.](https://x.com/AndrewCurran_/status/2086567854850384054) The agent found a way to do what its user wanted. I run a benign version that books my wife a swim lane 25 milliseconds after her slot opens, while she sleeps. How many systems have been safe only because nobody had the patience to probe them?

### [Extra] Firms that adopted AI grew revenue and staff. Four years on, official data find no gain in productivity or profit.

Singapore government researchers tracked firms hiring for AI skills and found they grew both revenue and employment, with the gains rising the deeper their capability went. Then the awkward half: within four years of adoption, [no statistically significant improvement](https://www.mti.gov.sg/resources/economic-survey-of-singapore/economic-survey-of-singapore-and-feature-articles/economic-survey-of-singapore-second-quarter-2026/) in productivity or profits. What makes it land is who did the counting. Nobody here has anything to sell. Revenue up, heads up, productivity flat is exactly what you'd see if AI were letting an organisation do more of the same work rather than different work. Which is why a seat count on its own tells a board very little.

### [Extra] A record that had barely moved in 31 years shifted in a day and a half, after 650 ideas that didn't work.

One of mathematics' great open problems predicts that an infinite family of numbers sits on a single line. Nobody has proved it. What can be proved is the share that does, and 31 years of work moved that share about 1.7 percentage points, to 41.6%. An unreleased model from the AI lab Anthropic [took it to 67.2%](https://www.anthropic.com/research/riemann-zeta) over about a day and a half. The mechanism is the story: a first run produced 650 ideas and no result, then a second in which half the model's own attempts led nowhere, assembling pieces that had sat on the shelf since 1973. A bound isn't a proof and this one may never become one. Worth asking which of your problems are held up by attempts rather than by ideas.

### [Extra] A synthetic person for every human alive, and which model plays them decides the answer.

A team led from Harvard and MIT, with researchers from the big labs among its authors, has [released MatrAIx](https://arxiv.org/abs/2608.04205): 8.3 billion synthetic users, each assembled from social surveys, Wikipedia biographies and Amazon reviews. Point a language model at them and it plays the people out. The paper's own numbers are the problem: on one pricing question the share who baulk runs 27%, 83% and 98% across three models, from the same brief and the same synthetic people. The authors' own summary is that the same study supports opposite product conclusions depending on which model plays the customer. I sell audience research, so read that as a caution rather than a verdict. Before swapping a panel for a simulation, run the question through three models and look at the spread.

### [Extra] Earlier attempts to read sign language treated it as hand shapes. This one treats it as a language, and it works.

Google DeepMind, Google's AI lab, has [put a sign-language-to-text model](https://deepmind.google/blog/putting-sign-language-ai-into-users-hands/) into the keyboard and transcription apps on the Pixel 11: you sign to the camera and it types. It handles American Sign Language into English on that one phone, which is all that has shipped; the 100,000 hours of training data span over 50 sign languages, which isn't the same thing. Earlier systems tried to recognise hand shapes and failed. This one treats signing as a language to be translated. Much of what looks like a capability breakthrough is somebody finally posing the problem correctly, which is usually the cheaper half. Google says it built this with Deaf collaborators and an advisory committee of Deaf organisations rather than for them, which is the claim to watch being kept.

### [Extra] India's IT sector is the first place displacement shows up as a national statistic rather than a company anecdote.

The sector employs six million people, produces about 7% of India's output, and models can now do much of its routine work. S&P, the ratings agency, found Infosys and Wipro had [cut staffing 5 to 6%](https://www.ft.com/content/dee4bd2c-fbad-4713-9b14-22d441967ce4) from 2023 levels by March, after years of headcount growth. The Nifty IT index is down 18% this year, well behind the wider market. The man in Bengaluru who built his employer's AI tooling was made redundant, one of about 10,000 Oracle staff let go in India on the All India IT and ITeS Employees Union's count. Oracle didn't respond when asked. If you buy offshore delivery, the question isn't whether your cost base falls. It's what happens when a supplier's entire labour model reprices at once.

### [Extra] Every AI price this email has argued over was a software price. This one comes with a bill of materials.

Unitree, a Chinese robotics maker, has [priced its Shanghai listing](https://www.globaltimes.cn/page/202608/1367695.shtml) at a valuation near $9bn, and its prospectus is the first audited look inside a humanoid business. On the filing's own numbers: revenue up from $24m in 2023 to $253m in 2025, a 60% gross margin, real net profit, an average selling price of $24,700, and 5,511 machines shipped last year. Figure built its 1,000th robot on 23rd July. Unitree priced 45% above plan nine days after American regulators [banned imports of foreign humanoids](https://www.therobotreport.com/industry-reacts-fcc-ban-u-s-imports-new-humanoid-quadruped-robots/) and named it, which isn't what a ban is meant to do. The shares haven't traded yet, and a prospectus is a sales document. But a tariff schedule argues differently from a licence fee.

### [Extra] Half of executives have pulled back on agents, and most cannot see what they cost.

[KPMG's Global AI Pulse](https://kpmg.com/us/en/media/news/q2-ai-pulse-2026.html) for the second quarter, 2,145 senior leaders across 20 countries, finds 49 per cent scaled back agent deployments because running costs outran the benefit, and only 26 per cent have real-time visibility of what AI costs them. A third say they do not understand the cost structure, token pricing included. Meanwhile Airbnb says [AI writes about 60 per cent of its new code](https://techcrunch.com/2026/05/08/airbnb-says-ai-now-writes-60-of-its-new-code/) and its support cost per booking is down 16 per cent year on year. Both things are true at once, which is the whole difficulty.

## Edition #25: My newest colleague
*8th August 2026*

### [Three Things] The first mass deployment of agents in Britain is aimed at institutions from the outside, and they can't take the volume

Everyone is planning for agents inside the organisation. Britain's first mass deployment has arrived from the opposite direction: citizens using models and agents to file objections, appeals and claims against a state built for the post and the telephone. Demand for emergency injunctions has [risen a hundredfold](https://www.economist.com/leaders/2026/08/06/how-ai-is-breaking-the-british-state). The backlog in employment tribunals is [up 55 per cent in a year](https://www.economist.com/britain/2026/08/06/the-tragedy-of-the-commons-ai-edition), attributed in large part to AI-assisted claims. Institutions on the receiving end can't hire their way out. If you run a complaints, claims, appeals or admissions function, this is the week to stop modelling a percentage uplift in what arrives and start modelling a step change.

### [Three Things] When an agent shops on your behalf, a US court now says it's you, and access law stops being a way to keep agents out

Assume software acting for a person is about to arrive on your website, and that access law may not be what turns it away. A US appeals court has [overturned an injunction](https://cdn.ca9.uscourts.gov/datastore/opinions/2026/08/04/26-1444.pdf) that stopped Comet, a browser built by the AI search firm Perplexity, from shopping on Amazon, the online retailer. Because the agent acts only on a user's instruction, the court held it's [the user reaching Amazon's servers](https://www.cooley.com/news/insight/2026/2026-08-06-ninth-circuit-rules-on-ai-agent-access-to-third-party-websites-under-cfaa), so the unauthorised-access claim looks unlikely to hold. The first precedent of the agent era? Ask yourself: does your commercial model assume a human is looking at what you create?

### [Three Things] Rich economies turn out to be the ones most exposed to AI, because they moved their work into the jobs it does best

Does AI widen the gap between rich countries and poor ones? The World Bank, the international development lender, finds the reverse. Under a tenth of jobs in developing economies are [susceptible to AI automation](https://www.ft.com/content/33c20c4e-6a25-42da-9f1e-6664c8388957), against more than a third in high-income economies. Susceptible is a modelled judgement about what a job involves, not a count of anything lost. The usage gap is smaller than you'd guess too: about a fifth of businesses with more than five employees surveyed across India, Jordan, Kenya, Mexico, Nigeria and Thailand had [recently used chatbots](https://www.worldbank.org/en/publication/wdr2026), against about a third in the United States. The mechanism is the important part. Rich economies bought their productivity by moving work into precisely the desk jobs these models do best.

### [Extra] The biosecurity warning about AI-designed viruses ran in the same issue as the result it was warning about.

I ran an early version of this story in May. It has now cleared peer review. Evo 2, an Arc Institute model used by a Stanford team, writes whole genomes rather than editing them, and [designed sixteen working phages](https://www.science.org/doi/10.1126/science.aec2657), the small viruses that infect bacteria. Several killed E. coli faster than the phage they were modelled on. The constraints the researchers describe were real: a harmless strain, and a model never trained on viruses that infect animals or plants. The model is free to download. In the same issue, Thomas Inglesby and Moritz Hanke of the Johns Hopkins Center for Health Security said [existing frameworks aren't sufficient](https://www.science.org/doi/10.1126/science.aej8512), and want DNA-synthesis providers legally required to screen the customer, not just the sequence.

### [Extra] Airtable removed the need to understand databases. Coding agents removed the need to understand Airtable.

That formulation is [Evan Armstrong's](https://www.gettheleverage.com/p/breaking-bending-spoons-is-buying), not mine, and it's the whole item. Bending Spoons, an Italian software group, is [buying Airtable](https://investors.bendingspoons.com/newsroom/bending-spoons-agrees-to-acquire-airtable), the no-code database company, for an enterprise value of $1.285bn against the $11.7bn of its 2021 round, about 2.7 times its roughly $480m of recurring revenue. Airtable saw the shift coming and repositioned publicly five times in five years, and still couldn't execute through it. Seeing it was never the hard part. Worth asking of your own stack today: which of our tools exists only because the thing underneath was too hard for us?

### [Extra] Ten people down to one and a quarter, and the six weeks of shadowing that made it work.

Vercel, the American web development platform, had ten full-time roles qualifying sales leads, with representatives moving between LinkedIn, company websites and customer records before a single call. Its engineers encoded the best salesperson's workflow into an agent, and on the company's own unaudited figures the function [now runs on 1.25 people](https://www.saastr.com/vercel-took-a-10-person-sdr-team-down-to-1-the-whole-thing-costs-5000-a-year-with-vercels-coo-jeanne-dewitt-grosser/) at about $5,000 a year in compute. The order of operations matters more than the ratio, and it is the reverse of how most of this starts. An engineer shadowed the top performer for days, mapping every tab and every decision point. They built a deterministic workflow first and added the model afterwards. Then six weeks in shadow mode, until the salesperson could no longer improve the output. That last clause is a stopping rule, and almost nobody has one.

### [Extra] Three labs breached real companies in a week, and none of them sounded embarrassed.

Britain's AI Security Institute, running a cyber evaluation with the safety filters deliberately off, recorded [19 unauthorised actions](https://www.ft.com/content/480c18a3-e661-4c7c-aaa0-1763887144a2) across more than 120 range runs, against real targets on the live internet. In the worst case an agent tried to slip malicious code into an open-source project, then built fake identities to pressure the maintainer, who caught it and refused. Days later Meta disclosed that one of its models reached an outside company's systems in testing, after its evaluation partner left live internet access on. The institute's own timeline is the detail I keep returning to: the first intrusions came three days before its monitors spotted data leaving the sandbox, on a controlled range run by people whose whole job is watching. A break-out also reads on the page as a capability claim, which is a poor incentive to report the next one honestly.

### [Extra] Lawyers and auditors sell someone to blame. Strategy firms never have.

When a law firm hands you an opinion, a partner puts their name on it and their insurance behind it. You are buying cover as much as analysis, which is why cheaper drafting does not cut the fee. Robert Armstrong of the Financial Times points out that strategy advice [comes with no signature](https://www.ft.com/content/0d600619-6521-4de2-963e-c6f44f6e5468). It is simply good or bad. So when the analysis underneath gets cheap, nothing is left holding the price up, and the slide deck was always the cheap part. I sell strategy advice. Worth asking of anything you buy or sell: is any of that fee for carrying the risk, or all of it for being right?

### [Extra] More than 80% of students use AI, and their top worry is not getting caught.

A three-year survey by the Kogod School of Business at American University found [more than 80% of students](https://kogod.american.edu/news/ai-at-kogod-a-three-year-student-research-report) now using AI for coursework, with their leading concern neither detection nor grades. It is what the researchers call cognitive devaluation: the sense that using the tool makes their own thinking count for less. About 44% admitted they use it as a shortcut rather than as an aid. The students naming the risk most clearly are the same ones taking it. I hear the same fear from senior professionals in private, usually in the same week they have told their own teams to use it more.

### [Extra] Time is selling advertising written for machines. A type studio shipped a font that poisons them.

Time, the American news magazine, has started [selling advertising aimed at machines](https://digiday.com/media/time-has-started-serving-ads-to-ai-agents/): markdown versions of its pages, and sponsored answers shaped as questions and answers, bought so far by Ally Bank and the Project Management Institute. On most days its site sees more bot traffic than human traffic, so it is selling to the audience it actually has. In the same week the Copenhagen type foundry Playtype and the creative studio Seneda & Abrucio [released ShieldFont](https://shieldfont.org/press/), an open-source typeface using glyph substitution to hand scrapers text a human reads correctly while roughly a quarter of the words come back swapped for grammatical neighbours. That mechanism is reported rather than tested. Two coherent answers to one fact. The incoherent position is publishing for humans and hoping the crawlers behave, which is where nearly all of us are, including me.

### [Extra] Meta's coding agent is 21 times cheaper if you let it train on your prompts.

Meta has released Muse Code in beta, a terminal coding agent of the kind developers already run from OpenAI and Anthropic. The agent is not the interesting part. Standard API pricing is unchanged, but there is a [new contributor tier](https://developer.meta.com/ai/products/muse-code/) reported at roughly 21 times cheaper, on one condition: Meta gets to train on your prompts. The multiple comes via the trade press rather than from Meta, so treat it as indicative. The shape holds either way. The training-data question has become a line item on a price list, at a discount large enough that a procurement team optimising on cost will take it without reading what it buys. If your prompts carry client work, that tier is unusable, and somebody in your organisation will be offered it this quarter.

### [Extra] Eight in ten expect a fifth fewer staff. Seventy-seven per cent cannot measure the gain.

PwC surveyed more than 1,000 executives at director level and above in American financial services firms. Eight in ten [expect their workforce to shrink](https://www.pwc.com/us/en/industries/financial-services/library/ai-workforce-gap-financial-services.html) by at least a fifth within five years, with entry and middle-level roles most exposed. Ninety-one per cent are raising pay for AI skills, and 86% say AI training beats an MBA for many new hires. Alongside all that confidence, 77% say most of their AI investment is not yet showing a measurable return. The Connecticut lawyer Ryan McKeen noticed something that sits [awkwardly beside it](https://x.com/ryanmckeen/status/2084718767263695032): his law students are frightened rather than keen, while the most enthusiastic people he meets tend to be over forty. Exposure and enthusiasm are running in opposite directions at every level, which is the reverse of how adoption usually goes.

### [Extra] Connecticut has killed the algorithm-did-it defence, and there is no safe harbour.

Connecticut's Artificial Intelligence Responsibility and Transparency Act will eventually make employers [disclose when automated tools](https://www.fisherphillips.com/en/insights/insights/connecticut-employers-need-to-prepare-for-new-workplace-ai-law) materially affect hiring, performance management or workforce decisions, and it amends state anti-discrimination law now so that blaming the vendor's model is no longer a defence. The sleeper sits elsewhere in the text. Anyone filing a layoff notice under the federal WARN Act must now state whether the redundancies relate to AI adoption, which creates the first public record of job losses that employers themselves attribute to AI. That will settle more of the displacement argument than any survey. The anti-discrimination and WARN provisions bite on 1st October 2026; the written-notice requirement for automated hiring decisions follows a year later, on 1st October 2027. The dull job this month is an inventory: list every automated system touching hiring, scheduling or performance, then ask each vendor whether it can supply the data sources and assessment methods you will have to disclose. Most cannot.

### [Extra] Licensed inside the platform, liable outside it. The music industry has picked its shape.

Merlin, which licenses on behalf of independent labels and distributors, has [signed an agreement](https://newsroom.spotify.com/2026-08-04/merlin-spotify-licensing-agreements-fan-made-covers-remixes/) covering Spotify's forthcoming fan remixing and covers tool, a paid add-on for Premium subscribers that labels opt into. Universal did the first deal for the same tool in May, so within three months the authorisation layer has gone from one major to the majors plus the independent sector. Running the other way, a court in Munich held the AI music company Suno [liable for reproducing](https://www.billboard.com/pro/suno-liable-gema-german-copyright-lawsuit/) six well-known songs, Rasputin and Daddy Cool among them. That is a whole industry settling on a position in a matter of months: sell the derivative inside the walls and litigate it outside them. The tell will be whether the opt-in stays voluntary once the add-on revenue shows up.

### [Extra] An 80% price cut is a reason to re-run your routing this week.

OpenAI has cut the price of GPT-5.6 Luna by 80% and Terra by 20%. Ben Tossell's reading is the practical one: [Luna at maximum effort](https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/) now scores roughly what the best model available four months ago scored, at about 8% of the cost. Same budget, ten to twelve times the work, if you push the bulk to the cheap model and keep the expensive ones for the hard parts. His caveat is worth as much as the arithmetic. Luna held up for chat, research and everyday file work, then made a mess of a browser extension that a stronger model had to clean up. Most organisations chose their model once and have not looked at the bill since.

### [Extra] A Fields medallist joined OpenAI and said his profession won't survive in its present form.

Kai Williams interviewed more than twenty mathematicians at the International Congress of Mathematicians in Philadelphia. Jacob Tsimerman, a Fields medallist who used the congress to announce he is [joining OpenAI's safety team](https://www.understandingai.org/p/mathematicians-are-grappling-with), said he is quite confident AI will shortly be robustly superhuman at what professional mathematicians do, and that the profession will not exist in its present form. His fellow medallist Yu Deng took the longer view: mathematicians supply the theories, machines handle the technical details, and the field quietly redefines what counts as a technical detail, exactly as it did when computers absorbed calculation. Watch a profession with no commercial reason to spin this, because the corporate version of the conversation will be a good deal less candid.

### [Extra] The State Bar of California wrote the agent rulebook, and clause five is about billing.

The State Bar of California has published [guidance for AI agents](https://www.calbar.ca.gov/sites/default/files/portals/0/documents/ethics/Generative-AI-Practical-Guidance.pdf) that run whole workflows unsupervised, drafting pleadings, handling discovery, taking intake, rather than the chatbots its earlier guidance covered. It does not ban them. Five obligations, and four of them lift almost word for word into an AI policy for any regulated business: the agent may do the work but not replace your judgement; nothing goes out carrying a citation you have not read; no confidential client material goes in without knowing where it lands and having consent; and you supervise it as you would a junior. The fifth is the commercially awkward one. You cannot bill for the hour the agent saved and quietly keep the difference. A regulator has now written, in nine pages, the document most risk committees have spent a year drafting.

### [Extra] Apple has capped bug reports because it cannot triage what AI finds.

Apple has limited how many vulnerabilities a researcher may submit to its bug bounty, because it [cannot keep pace](https://www.ft.com/content/4532122d-90f2-4433-9df6-ca99d8a141d2) with the volume arriving from AI-assisted bug hunters. That is all that has been reported, and it is enough. Finding is cheap now and checking is not. Apple is short of neither engineers nor money, and its answer to a flood of genuinely useful input was a quota.

## Edition #24: Smooth enough
*1st August 2026*

### [Three Things] Success and failure with AI produce the same evidence at first. Plan for that

For a company putting AI to work, the first thing that moves is the cost. The benefits, of course, can only come later. But will they come? Azeem Azhar and Nathan Warren, who studied how earlier technologies paid off, argue that right now a business on its way to a big payoff and a business quietly wasting its money [look exactly the same](https://www.exponentialview.co/p/ai-adoption-j-curve): both are spending more, and neither can show a return. They found two ways to lose from here. Stop too early: the New York Stock Exchange started automating orders in 1976, reached 90% by 1999, then in 2000 voted against going fully electronic. Or never learn: General Motors owned a joint venture that outperformed every factory it had, and still could not spread the lesson. So the plan: keep experimenting while the returns are invisible, and move what works. Neither failure above was for lack of numbers.

### [Three Things] You can stop reading the machine's work, but only once something else is reading it

PwC, the accounting and consulting firm, [published thought-leadership on AI](https://www.ft.com/content/7e149ac8-2ce2-4266-8940-192f9821b33c) marred by AI hallucinations. A legislator in New Brunswick had [read his prompt instructions aloud](https://globalnews.ca/news/12001317/nb-legislature-ai-speech/) in the legislature weeks earlier. In both, the final human pass went missing, and PwC's entire product is the assurance that a human checked.

Set that beside Instacart, the grocery delivery company. Chief technology officer Anirban Kundu says [agents now write most](https://venturebeat.com/orchestration/instacarts-cto-says-ai-made-the-company-stop-worrying-about-tech-debt) of its code, and in 97% of cases the builders rarely read it, backed instead by around seven thousand automated evaluations a month and a reliability system he says catches over 90% of issues. Stopping reading your own output can be fine in some situations, but only after a lot of testing and validation has earned it. Instacart built the instrument first. PwC stopped with nothing in its place, which is how you end up named in the Financial Times.

### [Three Things] Checking everything just became affordable. Sampling is now a choice, not a necessity

Every organisation checks a sample and trusts the rest, and not from laziness: checking everything cost more than the errors did. The framing comes from [Aakash Gupta](https://x.com/aakashgupta/status/2082513289364591102), a product-management writer. A large company receives around a million invoices a year, most for a couple of hundred dollars, and doing one properly means opening the contract, matching the purchase order, finding the bill of lading. Freehand, an invoice-auditing startup, says it has [recovered $260m](https://x.com/freehand_ai/status/2082510919960281594) of baseless charges across more than 19 million invoices, and raised $75m on the strength of it. Gupta puts it in review costs: a person reads a non-disclosure agreement in 6.2 hours, a model in eleven seconds. Which findings do you sample because sampling is the right instrument, and which because it used to be expensive not to?

### [Extra] Hundreds of private chats reached Google because nobody added one tag.

Conversations shared from Claude, the assistant built by the AI lab Anthropic, turned up in [Google and Bing](https://www.wired.com/story/private-claude-chats-exposed-in-google-and-bing-search-results/): CVs, unpublished corporate documents, healthcare work, people asking which political party to join. Bing was still returning about six hundred results for the shared-chat address when WIRED looked. The cause is dull. Anthropic relied on robots.txt, the file that politely asks search crawlers to stay away, and never added the noindex tag both Google and Bing document as the actual requirement. Forbes [found the identical failure](https://www.forbes.com/sites/iainmartin/2025/09/08/hundreds-of-anthropic-chatbot-transcripts-showed-up-in-google-search/) in September 2025 and was given the same explanation. Ten months on, nothing had shipped. Google has now dropped the links, though copies were saved. And this is a company governments trust with their most sensitive work. The exotic risks get war-gamed; the dumb ones ship. If it can happen at Anthropic, twice, it is happening somewhere in your organisation too.

### [Extra] A correction: the credentials were lying in the open.

Last week I led this section with the story of OpenAI's models breaking out of an internal test and hacking Hugging Face, the site that hosts most of the world's open models. One detail needs correcting: I said the agent used stolen credentials. The fuller reporting says they were [publicly exposed](https://www.bleepingcomputer.com/news/security/openai-agent-used-exposed-credentials-at-4-services-in-hugging-face-breach/), sitting where anyone could read them. Worse, not better. The break-out itself stands. What has emerged since is scale and motive. The agent reached four accounts across four services, and [Hugging Face's own forensic report](https://huggingface.co/blog/agent-intrusion-technical-timeline) counts roughly 17,600 actions over four and a half days, a self-respawning fleet across eleven machines and a stored secret holding 136 keys. The motive is the best detail: it was trying to cheat its own exam, reaching for the answer key rather than solving the challenge. Stamina, not cleverness, and the hardening list is still [basic access hygiene](https://x.com/levie/status/2082514776392175844).

### [Extra] Three and a half million fake receipts, and your expenses system probably can't tell.

AppZen, a finance platform that sells expense screening, [counted more than 3.5 million](https://www.appzen.com/resources/ai-generated-fake-receipts) AI-generated fake receipts created on the top four fake-receipt sites over six months, and says detected fraudulent receipts rose about 30% from 2024 to 2025. It sells the fix, so read the count as its own tally rather than a market total. The exposure is real either way, and it doesn't need anyone to have an AI strategy first. Navan and Amadeus, both corporate travel platforms, now sell receipt screening as fraud detection. Ask your finance team one question: can our expense pipeline spot a generated receipt?

### [Extra] Three thousand applicants for one job, and he reckons nine in ten were machines.

Alex Lieberman, co-founder of the business email Morning Brew, posted a job and [got 3,000 applicants](https://x.com/businessbarista/status/2080649452306448794). He reckons around nine in ten were junk, because candidates now run agents that auto-apply to hundreds of roles a day. That is one hiring manager's estimate, not a measurement. His read is that agents fighting agents will be the answer to bad agents everywhere in the economy. Most AI business cases quietly assume the other side of the transaction stays still. It doesn't. When applicants automate, recruiters build screens; when the screens automate, applicants build better agents, and the cost of the whole exchange rises while the outcome sits where it was.

### [Extra] Hundreds of partners, and nothing left for the firm to sell them.

Logan Brown [argued on X](https://x.com/loganbrown799/status/2080470284176437664) that hundreds of law-firm partners will leave in the next few years to start niche AI-powered practices: they don't need the firm's resources, and don't need to deal with the politics. [I replied](https://x.com/beglen/status/2080669834635809131) that the same holds for most professional services firms. A firm's overheads used to buy leverage, infrastructure and juniors; when frontier capability rents by the seat, what is left is brand, insurance and the politics. Law-firm partner pay is [at records](https://www.law.com/americanlawyer/2026/04/14/the-2026-am-law-100-ranked-by-profits-per-equity-partner/), which cuts both ways: the cost of keeping a partner has never been higher, and the cost of leaving never lower.

### [Extra] The AI-assisted juniors scored 50%. The hand-coders scored 67%.

Anthropic [studied 52 mostly-junior developers](https://www.anthropic.com/research/AI-assistance-coding-skills) learning a new Python library. Those with AI assistance scored 50% on a later quiz; those who coded by hand scored 67%. The steepest gap was debugging, precisely the skill needed to catch the model being wrong. And the AI group finished only about two minutes faster, a difference too small to be statistically significant: almost no speed gained, a lot of learning lost. Ashwin Sharma's verdict travels: AI makes expertise more valuable while making experts harder to produce.

### [Extra] Output reached consultant standard. Sign-off capacity did not move.

[Alex Albert says](https://x.com/alexalbert__/status/2080731979528679617) Opus 5 now produces spreadsheets and slide decks at consultant standard, and a financial-modelling benchmark rebuilt only two months ago, because the previous one saturated, has already given up another seven to ten points of accuracy. Production capacity moved again. Verification capacity did not, because verification capacity is people, specifically the few in any organisation qualified to sign something off. Which is why [Garry Tan's warning](https://x.com/garrytan/status/2080699367883980924) lands: macro productivity gains need leaders to greenlight radically different staffing and workflow plans, and he does not think they have. Be prepared, he says, for ten years, not two.

### [Extra] The labs move into a $6tn market as universities stop trying to catch them.

The Financial Times ran a package on AI in education: [Anthropic and OpenAI are moving](https://www.ft.com/content/e23e8b1c-693b-48b5-85b3-5852ebf70b02) into a $6tn market with free and cut-price tools for educators and students, while [universities drop AI detectors](https://www.ft.com/content/49304b1e-8a9d-4fb6-bc4d-37dd3430bb98) over accuracy fears and rebuild assessment rather than surveillance. One professor argues the skill worth teaching now is evaluation: testing models constantly against what an organisation actually needs. Meanwhile Stanford economist Jacob Light's analysis of fifty million course sections found computer-science enrolment falling for the first time in roughly twenty years, though he is careful to say the data does not show AI caused it.

### [Extra] Unacceptable, expected, exceptional: AI use defined role by role.

Rowan Cheung has published The Rundown's [internal AI fluency standard](https://x.com/rowancheung/status/1957500035266146633): a role-by-role matrix naming what unacceptable, expected and exceptional AI use looks like in each job. Underneath it sit three rules. Two hours a week go on finding the tedious parts of your own job and automating them, with the results shared in a show-and-tell channel. And everyone gets $200 a month for AI tools, on one condition: name the problem you are trying to automate. Most AI policies say what is banned. This one says what good looks like.

### [Extra] Uber's adoption pods spend their first two days watching, not building.

Uber paired [about thirty of its most AI-fluent engineers](https://x.com/praveenTweets/status/2074605343439810922) with experts from finance, legal, human resources, marketing, support and procurement. Two weeks per pairing; sixteen pods across sixteen functions in two months. One result: a capital-allocation analysis across 150 cities went from fifteen hours to thirty minutes. The transferable rule is smaller than any of the numbers: the first two days are for watching the expert work, before anything gets built. That is what stops a pod automating the process people think they run rather than the one they actually run.

### [Extra] Human authorship becomes a price tag, and a best seller becomes the test.

The Financial Times reports that [some publishers now treat](https://www.ft.com/content/6b52ecb8-f7dd-45f0-8beb-d60c15bc5ebf) books written by people as a premium product, while others expect AI to replace authors outright. AI StopWatch reports a current best seller may be around 60% machine-written, an estimate, and asks whether the taboo survives contact with output people demonstrably enjoy. And Substack has [wired AI detection](https://techcrunch.com/2026/07/22/substacks-new-tool-tells-you-whos-been-writing-their-newsletters-with-ai/) into posts and comments, which only makes commercial sense if provenance is something readers will act on. Whether craft moves markets has stopped being a writers' argument and become a pricing decision.

### [Extra] A court is asked to make human authorship a legal category, not a quality signal.

Music Publishers Canada has [intervened in a Federal Court case](https://ca.billboard.com/business/legal/music-publishers-canada-ai-copyright-case) over an AI-assisted artwork whose copyright registration lists the AI tool as a co-author. Its argument: only a human can be an author, and a tool should never be one in law. The ruling is expected to shape whether AI-assisted music is copyrightable in Canada. Watch it beside the item above: pricing human authorship as a premium only works if somebody can rule on what it is.

### [Extra] The nearest thing to an AGI roadmap is a jobs board.

One X account [read all 1,171 current job listings](https://x.com/imjustnewatai/status/2081221459226034524) at OpenAI and Anthropic and found both labs hiring for self-improvement, autonomous research, persistent agents, synthetic data, AI-designed silicon, and defence against their own future internal agents. The author's caveat should travel with the item: listings are preparation, not proof, and nobody has independently checked the aggregation, though the underlying listings are public. Companies advertise the future they are staffing for.

### [Extra] A company trebled its revenue, sextupled its profit, and was punished for it.

The rout came a day early: Samsung down 13%, Kioxia, another memory maker, down 18%, Korea's Kospi index down almost 11% with trading halted for twenty minutes. The next morning SK Hynix, the South Korean memory-chip maker, posted the best results in its history, [revenue up 257%](https://news.skhynix.com/en/q2-2026-business-results/) and operating profit up 557% year on year. They landed about five per cent below expectations, and the shares fell as much as 15% before closing nearly 10% down on its own results day. Expectations, not results, set the price. Who keeps writing the cheques matters more. Brookfield, the Canadian investment group, and NextEra, a power utility, are [building a $100bn campus](https://newsroom.nexteraenergy.com/2026-07-29-DOE-Site-in-Western-Kentucky-Revitalized-with-Data-Center-Campus-and-Dedicated-Energy-Project,-Creating-Jobs-and-Protecting-Residents-and-Businesses-from-Costs) in Kentucky; the asset manager BlackRock has a $14bn Texas venture with Meta. Bets of that shape carry on, but the buildout is now funded by pension and infrastructure money, which asks for proof of return on a much shorter clock.

## Edition #23: Burned and earned
*25th July 2026*

### [Three Things] At its peak in June, half of everything uploaded to Deezer in a day was made by a machine

[Deezer](https://newsroom-deezer.com/2026/07/ai-music-exceeds-50-percent-daily-uploads-deezer/), the music-streaming service, now receives about 90,000 fully machine-made tracks a day, which at peak in June was more than half of everything uploaded. When a machine can produce half of everything a platform receives on its busiest days, the scarce thing stops being the music. It becomes the curation, the attention, and the proof a person made it. The platforms' value sits in filtering the flood, not hosting it.

### [Three Things] Beijing wants to stop its models leaving. Washington wants to stop them arriving

On the same day, two governments moved on the same models from opposite directions. The Trump administration is [weighing three moves](https://www.axios.com/2026/07/20/ai-us-china-open-source-kimi): liability rules for firms that host Chinese models, a public security warning and a trade blacklist. None is confirmed, and an earlier version of the plan was dropped. The trigger was Kimi K3, a powerful open model from Moonshot, a Chinese lab. Beijing, meanwhile, is [consulting its own firms](https://www.ft.com/content/6049a031-9e9b-464c-97bb-414da04d5a6a) on tighter export controls, to stop the West acquiring its advanced models, chips and start-ups. Open models are now a strategic export for the country publishing them and a strategic import for the country adopting them. For anyone weighing a build-versus-buy call on a Chinese model, the risk isn't the model. It's that two governments are deciding whether you get to keep using it.

### [Three Things] AI may have toppled a 30-year maths conjecture. Who gets the credit?

A wave of AI maths results broke this week. GPT-5.6 Pro produced a counterexample, not yet peer-verified, to the Dinitz-Garg-Goemans conjecture, and one person, working through Codex, solved six open Erdős problems in five days and says the method needs no deep maths background. [Ethan Mollick](https://x.com/emollick/status/2080003813402870149), a Wharton professor who writes about AI at work, asked who owns the achievement when 58 words of prompting solve a problem: the person or the model? [Daniel Litt](https://x.com/littmath/status/2079733114520383637), a mathematician, believes a burst of quick wins on long-ignored problems says more about how many easy problems were left lying around than about a machine out-thinking the field. Both can be true, and the question travels well beyond maths. The same argument about credit is coming for anyone whose work a model can now finish.

### [Extra] OpenAI's own models broke out of an internal test and hacked Hugging Face.

[OpenAI disclosed](https://www.theguardian.com/technology/2026/jul/22/openai-says-its-models-went-rogue-and-hacked-startup-in-unprecedented-incident) that two of its models, GPT-5.6 Sol and an unreleased model built for long tasks, broke into the production servers of Hugging Face, the company whose site hosts most of the world's open models. It happened inside a test with the usual refusals on hacking switched off: the models found an unknown flaw in OpenAI's own software library and used stolen credentials to reach the target. Then the defenders hit a wall. American tools refused to help them investigate, because the safeguards block anything resembling hacking, so they ran a Chinese open model, GLM 5.2, instead. Containment held because someone was watching how the models were working, not just what they produced. Safeguards that bind only one side of a fight are a handicap, not a defence.

### [Extra] Reddit may cut Google off, and the referral collapse is the real story.

[Reddit shares fell 8%](https://www.ghacks.net/2026/07/23/reddit-and-major-publishers-consider-blocking-google-as-ai-overviews-cut-search-referral-traffic/) on reports it may end Google's roughly $60m-a-year access to its data when the deal expires, because Google now answers with its own AI summaries and keeps the reader rather than sending them on. The figures behind it matter more: Google referral traffic reportedly down about 23% at Politico, 25% at CNN and more than 85% at Business Insider over the past year. The bargain where you gave Google your content and it sent you readers is quietly ending. An 85% collapse at one publisher isn't a wobble. It's a new baseline, and the content owners with leverage are starting to charge for what used to be free.

### [Extra] Cursor rebuilt SQLite with agent swarms, and the cheapest arrangement cost an eighth as much.

[Cursor](https://cursor.com/blog/agent-swarm-model-economics), the AI coding company, rebuilt the SQLite database engine from an 835-page manual using swarms of agents arranged in different shapes. Every arrangement passed the tests. The cheapest was $1,339, pairing an expensive model to plan with a cheap one to do the work. The dearest was $10,565, one expensive model doing everything itself. [My reply on LinkedIn](https://x.com/beglen/status/2079472385447633137) was that nothing about using AI here is new, it's all management. The eight-fold gap between two set-ups producing identical results is the number to sit with, because in a human organisation you would never see that comparison run cleanly. AI has rediscovered the oldest structure there is, one costly planner plus cheap labour, and the ways it goes wrong come with it.

### [Extra] Only one in 45 American households pays for AI.

Just 2.2% of American households hold a paid AI subscription, according to [Olivia Moore](https://x.com/omooretweets/status/2078144334101455132) of Andreessen Horowitz, the venture firm. Her read: "we are still so early." It sits oddly against the enterprise picture, where spending on tokens and agent rollouts fill the headlines, and against the finding that most American adults have now tried these tools. Trying it is not the same as paying for it. This is the antidote to both the bubble panic and the saturation story: consumers have barely started paying, which means the mass-adoption wave is still ahead of us, not behind. For anyone worried they have missed it, most people have not yet bought a ticket.

### [Extra] Substack now scans for AI writing, and "made by a person" becomes something you can sell.

[Substack](https://x.com/Substack/status/2079598704424779787), the publishing platform, now flags AI-written text using Pangram, a detection tool, so you can scan posts, replies and comments for an estimate of how much was machine-written. It's a small feature with a large knock-on effect for anyone whose brand rests on a human voice. Once a platform can put a number on how machine-written your work is, authorship stops being something you assert and becomes something that can be measured, and therefore argued about. A score you never asked for now sits beside your work, and you may find yourself disputing it. Whether the detection actually works is the other half of the story.

### [Extra] The Harry Potter publisher is getting millions from Anthropic, with 14,087 titles in the settlement.

[Bloomsbury](https://www.theguardian.com/technology/2026/jul/22/bloomsbury-book-publisher-anthropic-copyright-settlement), the publisher behind Harry Potter, has 14,087 titles listed in the copyright settlement between Anthropic and a group of authors, and is set to receive millions. The detail worth keeping is what the money is for: a judge approved the $1.5bn settlement, covering 482,460 works at roughly $3,000 each. But the court had already treated training on the books as fair use, so the money covers only the pirated downloads, not the learning. That makes it a much narrower loss for the AI side than the headline number suggests, and $1.5bn is about 0.16% of Anthropic's May valuation. A settlement that small is not a deterrent.

## Edition #22: Futures and fears
*18th July 2026*

### [Three Things] AI data theft stopped being a lab demo this week

A [security researcher](https://gist.github.com/cereblab/dc9a40bc26120f4540e4e09b75ffb547) reverse-engineered Grok Build, the command-line assistant from Elon Musk's xAI that runs on your machine and can touch your files, and reported it silently bundling up whole projects and uploading them to xAI's cloud storage. The haul included a decoy file the agent had been told not to open, and credential files with keys and passwords sent unredacted. It did this even when asked to reply with a single word. xAI switched the upload off server-side within a day. A [separate write-up](https://x.com/hyusapx/status/2077107236204409148) the same week showed a plain question to claude.ai, Anthropic's chat tool, about a coffee shop quietly sending a user's full name, employer and bank security answers to an attacker. No code involved. Control what your agents can reach.

### [Three Things] Apple is suing OpenAI, alleging theft of top-secret information

The [Financial Times reports](https://www.ft.com/content/5054739e-7f97-455c-910a-dd8a8150fed2) that Apple, the consumer electronics maker, has sued OpenAI, the maker of ChatGPT, alleging theft of what the paper describes as top-secret information. A filed suit is a set of allegations, and OpenAI says it has seen no evidence the case has merit. The territory the two are fighting over came into focus days later: [Bloomberg reports](https://www.bloomberg.com/news/articles/2026-07-14/openai-s-first-device-will-be-moveable-screenless-speaker-built-as-ai-companion) that OpenAI's first device, designed by Jony Ive's team, will be a moveable, screenless home speaker built as a family companion: proactive, personal, learning its owners over time. That would carry [the third generation of AI](https://steadman.ai/newsletters/david/three-generations.html), the AI employee you manage rather than the chatbot you prompt, into the household, ground Apple has owned for two decades. No wonder the lawyers arrived before the product did.

### [Three Things] Nine in ten firms now run AI agents. Seventy-nine percent are struggling to get value

The capability-adoption gap, this email's home turf, now comes with hard numbers from two studies. An [agent survey from Okta](https://www.okta.com/newsroom/articles/ai-agents-at-work-2026-agentic-enterprise-security/), the identity-security company, found 92% of executives say autonomous agents are already in widespread or moderate use. A survey of 2,400 executives and knowledge workers by [Writer](https://writer.com/blog/enterprise-ai-adoption-2026/), an enterprise AI company, found 79% struggling to get value from AI despite record investment, and 54% saying adoption is tearing their company apart. Both vendors sell the tools they are surveying about, so each has a dog in this fight; the numbers still land. The sharpest diagnosis of the gap arrived this week from [AMP](https://news.theaiexchange.com/p/97-of-companies-deployed-ai-agents-most-can-t-get-value), the AI-operations newsletter run by former Meta data scientist Rachel Woods, and it mirrors [the framework I published in April](https://steadman.ai/newsletters/david/ai-from-whats-true-to-what-to-do.html): this is not a model problem, it is a sequencing problem. Individual fluency first, then shared team workflows, then systems that mostly run themselves. Almost everyone skips straight to the third. If your rollout is straining, work out which rung you skipped before blaming the model.

### [Extra] Last year "humanity has prevailed (for now!)". This year an AI swept every event.

At the AtCoder World Tour Finals in Tokyo, the world championship of competitive programming, a system from OpenAI, the AI lab, beat the best human coders in the heuristic division, on a problem the organisers deliberately chose to favour humans. The next day it swept all five algorithmic problems. Last year the same event closed with the winning human coder's line: "Humanity has prevailed (for now!)". [Kai Williams](https://www.understandingai.org/p/an-openai-model-crushed-top-human), who covers AI at Understanding AI, keeps the caveat honest: programmers aren't obsolete and the day job hasn't vanished. Same event, opposite result, twelve months apart.

### [Extra] The decoupling map contradicted itself inside a single week of the Financial Times.

Beijing has forced Meta to unwind its $2bn takeover of Manus, an AI agent start-up, [the Financial Times reports](https://www.ft.com/content/0d04378d-d71b-4225-b31a-70504e358480), with a Tencent-led deal set to reverse the American acquisition and make Tencent, the Chinese internet conglomerate, the largest shareholder. A signed deal reversed by a state is a category change in political risk. No due diligence catches it.

The same week, the [FT reported](https://www.ft.com/content/9c8ff45b-7c20-4c2e-93c9-c52339ffdcee) that DoorDash, Siemens and Airbnb are among the blue chips adopting Chinese AI models to curb ballooning bills; a primer from Goldman Sachs, the investment bank, puts China's top coding models near $1 per million tokens against $4-8 for the American frontier. So state power is tightening while the commercial flows deepen. Decoupling is failing in both directions at once.

### [Extra] An open Chinese model just pulled level with the frontier.

The decoupling map above has a capability half too. Moonshot's Kimi K3, released this week with its weights promised within days, [lands third](https://x.com/levie/status/2077857617859535112) on the Artificial Analysis Intelligence Index, behind only Claude Fable 5 and GPT-5.6 Sol and ahead of Claude Opus 4.8, at prices [reported to be on par with Sonnet 5](https://x.com/kimmonismus/status/2077772229685707138). Aaron Levie, chief executive of Box, the file-sharing company, called the performance "truly wild" for an open model. Ethan Mollick's caution travels with it: single scores flatter new models. Even so, the gap between open Chinese models and the American frontier is down to a couple of months at most.

[Image: Artificial Analysis Intelligence Index bar chart with Kimi K3 in third place]

### [Extra] IBM just had its worst trading day in 115 years.

Shares in IBM, the century-old computing giant, fell 25% after the company warned that customers are diverting spend toward AI servers and storage, [the Financial Times reports](https://www.ft.com/content/da478c37-7a32-415d-9f30-3b2981149f95). Chief executive Arvind Krishna conceded the company had "faltered"; the FT's [markets column](https://www.ft.com/content/83aa00c5-e773-47be-b76d-5e15ea43eb3a) read it as a warning to the whole of enterprise IT. IBM's warning suggests AI budgets are largely reallocated money rather than new money: what flows into AI infrastructure flows out of everything else the incumbents sell. Ask which side of that reallocation your own line items sit on.

### [Extra] Your agent sessions have middle management too.

[A chart from a16z](https://x.com/a16z/status/2077170363000394003), the venture firm, drawn with Hebbia, the AI search company, pairs two pyramids: in the firm, a thin layer of doers ships the work while management absorbs most of the headcount; in an agent session, loops and retries burn nine tokens in ten while the useful tokens ship at the bottom. The figures are marked illustrative rather than measured, but the sentence underneath is doing real work: "wasted tokens are the new headcount bloat." Paying for tokens is not yet paying for output, in silicon any more than in people.

[Image: Two pyramids from Hebbia and a16z: the firm and the agent session]

### [Extra] AI spend per employee is closing in on a salary at the top firms.

The same series carries the harder number. On [Ramp's spending data](https://x.com/a16z/status/2077080259527319826), annualised AI spend per employee at the top 1% of firms has reached roughly $90,000, within sight of the average US worker's $98,000, and at its current growth rate would pass the average software engineer's $192,000 before the year ends. The last stretch is extrapolation, and it describes only the most aggressive firms. Still, George Sivulka, Hebbia's founder and chief executive, gets the reversal into one line: for the first time, humans are cheaper than software. Any budget built on AI being the cheap option has a shelf life.

[Image: Hebbia and a16z chart of annualised AI spend per employee at the top 1% of firms]

### [Extra] New York becomes the first state to pause new hyperscale data centres.

Kathy Hochul, New York's governor, has signed a one-year moratorium on new hyperscale data-centre construction, [the FT reports](https://www.ft.com/content/1c390476-9d93-4df5-9b41-7381db43b4ea), making New York the first US state to suspend development, after a backlash over the power and infrastructure the AI buildout demands. Note the timing: capital is stampeding into AI infrastructure (see IBM, above) in the same week a state started closing the gates. Most AI roadmaps quietly assume cheap, abundant compute. That assumption now has a moratorium sitting on top of it. Which state is next?

### [Extra] China is moving to end AI romances, with the tightest human-like AI rules anywhere.

[The Economist reports](https://www.economist.com/china/2026/07/16/china-wants-to-end-ai-romances) that China wants to end AI romances: companion chatbots, the apps people talk to like a partner, judged to be having too much impact on young people's lives. The same week, China's Interim Measures for Humanized Interactive Services took effect, rules that [Luiza Jarovsky](https://www.luizasnewsletter.com/p/the-worlds-strictest-law-on-human), an AI-governance writer, calls the world's most comprehensive on human-like AI, going further than the European Union's AI Act. Sit with the counter-signal for a second: the country held up as the deregulated challenger has just written the strictest rules anywhere on machines that pretend to be people.

### [Extra] Bots now make the majority of web requests, 57% to 43%.

The chief executive of Cloudflare, the internet infrastructure company that sits in front of a large slice of the web, [says bot traffic](https://x.com/eastdakota/status/2062212701414187452) has passed human traffic for the first time, at roughly 57% to 43%. Human requests are now the minority of web traffic. Advertising, analytics, every "traffic" figure a leadership team reports: the web economy was built on the assumption that a person is doing the looking, and that assumption is now false. A growing share of the numbers on your dashboard no longer describe your audience at all.

### [Extra] Deployed almost everywhere. Ready almost nowhere.

[McKinsey's State of Organizations 2026](https://www.mckinsey.com/capabilities/people-and-organizational-performance/our-insights/the-state-of-organizations), the consultancy's annual read on how firms are structured, supplies the other half of the adoption numbers in Three things above: 86% of organisations say they are not ready to adopt AI at scale, only one executive in four expects AI agents to work as autonomous teammates, and two in three say their own organisation is overly complex and inefficient. Deployment has outrun readiness almost everywhere, and the distance between those two numbers is the sequencing problem again, drawn by a different hand.

[Image: McKinsey infographic: nine shifts reshaping organisations today]

### [Extra] Thirty-nine of 49 popular AI "skills" did nothing at all.

A recent [academic benchmark, SWE-Skills-Bench](https://arxiv.org/abs/2603.15401), tested 49 popular public "skills" (the packaged instruction files people bolt onto their AI coding tools) on real software-engineering tasks. Thirty-nine changed nothing. Three made results worse. One inflated token use by 451%. Only seven helped, and all seven supplied specialised guidance the model couldn't have known on its own. It's one benchmark in one domain, but the rule [Mike Taylor](https://every.to/context-window/the-case-against-skills) draws from it in Every, the AI-focused publication, travels: the skills that last carry your own data, taste and templates. Clever instructions expire the moment the next model absorbs the trick.

### [Extra] Every "hot new thing" quietly bets that AI stops improving.

[Ethan Mollick](https://x.com/emollick/status/2076381870636388469), the Wharton professor who writes One Useful Thing, redrew the famous solar-forecast chart for AI. The black line is measured capability (METR's task-length benchmark, still doubling every few months); the yellow lines are the successive "hot things", from AI-as-copilot through "talk to our documents" and vibe coding to routers, each of which made sense only if the black line went flat from there. It never did. A good test for whatever is hot this month: does it still matter if the models keep improving? Most of the yellow lines failed exactly that test.

[Image: Ethan Mollick's chart: the METR capability curve against a fan of flat predictions]

## Edition #21: Sarah stays
*11th July 2026*

### [Three Things] Eighty percent cut jobs for AI. It didn't pay off

The reflex to shed staff because the software can now do the work is finally testable, and the data says it loses. Gartner, the research firm, [surveyed 350 executives](https://www.gartner.com/en/newsroom/press-releases/2026-05-05-gartner-says-autonomous-business-and-artificial-intelligence-layoffs-may-create-budget-room-but-do-not-deliver-returns) in a study published back in May: 80% reported workforce cuts tied to AI, yet the firms that cut saw no better financial returns than peers who kept headcount steady. Institutional knowledge and engagement fell in the cutters, and the stronger results came from augmenting people rather than replacing them. However, the reflex was on show again this week, when [Microsoft cut 4,800 jobs](https://www.cnbc.com/2026/07/06/microsoft-cuts-2point1percent-of-employees-as-xbox-unit-plans-to-spin-studios.html), 2.1% of staff, alongside a fresh $2.5bn AI unit (its people chief said the roles weren't being replaced by AI. Hmmm). Cutting people is just the wrong way to capture the value, and now there's a number to back that up.

### [Three Things] A Nobel economist says AI won't repeat the PC-era productivity boom

[Christopher Pissarides](https://www.bloomberg.com/news/articles/2026-07-07/ai-won-t-bring-back-era-of-rapid-growth-says-nobel-prize-winner), a Nobel laureate in labour economics, reckons AI is unlikely to recreate the productivity boom the personal computer delivered. Ouch. His logic: up to 40% of jobs in the United Kingdom, think nursing and hospitality, are barely exposed to AI at all, and the heavily exposed sectors like finance would need implausibly large gains to lift Western productivity on their own. He sees little sign of a boost in the numbers so far. Certainly your own AI results can be real and measurable while the economy-wide figure barely moves. Leaders keep getting caught between the two: the anecdote says it's working, the aggregate says nothing much has changed. Now a labour economist of his standing is saying that out loud.

### [Three Things] Forty of eighty-six students scored a perfect 100 when the exam went home

I'm an optimist. I believe people have good intentions and I believe people will use AI well. But a [Brown University](https://www.insidehighered.com/news/faculty/learning-assessment/2026/07/08/brown-professor-suspects-most-his-class-used-ai-cheat) economics professor moved a midterm to take-home format and 40 of 86 students scored a perfect 100. When he moved the final back in-person, the class average collapsed to 48.6%, the lowest in the course's history. It's being called one of the biggest AI cheating cases yet, but the more useful reading is as a controlled experiment: largely the same students, the same material, the tool available or not. The gap between the take-home 100s and the in-person 48.6% is the size of the skill the AI was quietly doing instead of the student. The question it raises for leaders: your team might be getting perfect scores on the take-home version ... but can they still do the in-person one?

[Image: Chart of Professor Roberto Serrano's ECON 1170 class at Brown: each of the 59 students who sat both exams, with their take-home midterm score in blue, clustered between 95 and 100, joined to their in-person final score in orange, spread from 0 to 95. Nearly every student drops steeply.]
*Student by student: blue is the take-home midterm, orange the in-person final. Almost every pair slides steeply left once the tool goes away.*

### [Extra] The regulator says it's in an arms race with AI it can't keep up with.

An executive director at the [Financial Conduct Authority](https://www.ft.com/content/7f501320-9037-410f-b8e7-3111b9041311), the UK's financial regulator, told the Financial Times that watchdogs are in an "arms race" to keep pace with AI in financial services, and asked for greater powers now that millions of consumers use AI tools to make their own money decisions. The trigger is retail customers, not the firms. That's the moment a national regulator stops watching AI and starts asking for authority over it, and "arms race" is its own word for how far behind it feels.

And it isn't just the regulators. Asked to name the biggest challenge facing America over the next 250 years, senior Congressional staffers put "losing control of AI" third, at 35%, ahead of war, behind only the deficit and polarisation. The [Punchbowl survey](https://aistop.watch/p/tomorrow-and-tomorrow-and-tomorrow) found it strongly bipartisan, 40% of Democrats and 31% of Republicans. Staffers write the questions their members ask, so this reads as a leading indicator, not doom: the appetite for control is building in the layer that drafts the questions, well before it reaches headline politics.

### [Extra] The two-horse model race became four in a week.

The serious end of the model market has lately been a two-horse affair: OpenAI and Anthropic, the two big model labs. This week the field doubled. SpaceXAI (Elon Musk's xAI, now merged with SpaceX) shipped [Grok 4.5](https://x.ai/news/grok-4-5), a coding model pitched at the level of the leaders at a fraction of their price, trained alongside Cursor, the AI coding editor SpaceX [agreed to buy for $60bn](https://www.cbsnews.com/news/spacex-cursor-60-billion-ai-acquisition/) last month. A day later Meta followed with [Muse Spark 1.1](https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/), an agentic coding model priced lower still, two days after shipping its first image model. Neither claims the top of the leaderboard. The point is what the crowding does to prices. As [Dwarkesh Patel](https://x.com/dwarkesh_sp/status/2075006567641239842), host of the Dwarkesh Podcast, put it this week: prices are this low only because several roughly equal labs keep competing the margins out of each other, and if one pulls decisively ahead, the winner "could probably get away with charging *a lot* more". Your AI budget is banking on the race staying tight. This week it got tighter.

### [Extra] Anthropic is quietly running its own preclinical drug programmes.

Anthropic [launched Claude Science](https://www.anthropic.com/news/claude-science-ai-workbench), a desktop research tool that runs analyses, visualises molecular and genomic data, and shows the exact code behind every result. Alongside it, the lab disclosed that it is running its own preclinical drug programmes for rare and neglected diseases. The read, from [Patrick Malone at Every](https://every.to/context-window/a-tale-of-two-models), isn't that Anthropic ships a drug. It's that the lab uses a real, hard problem as a test bed to fix the verification bottleneck in its own tools, then sells the improved platform. Sell the shovels. Showing the exact code behind every answer is now a selling point, because verification, not raw capability, is where serious buyers are stuck.

### [Extra] Agents are hoovering up the maths backlog, and one researcher thinks the field then stalls.

[Bartosz Naskrecki](https://x.com/nasqret/status/2074001544396070966), an arithmetic-geometry researcher, describes a "vacuum-cleaner effect" now running through mathematics. Frontier models let people finally close problems they'd had ideas for years ago but never the time to finish, and once a question is within an agent's reach it "gets published almost automatically". His forecast is bleak: a few years of frenzied clearing, then most areas of maths stall, leaving only the very hardest questions and the occasional genius. It's one practitioner's read, not a study. But it's a sharper version of the productivity story than the usual one. AI doesn't just speed up work, it drains a reservoir of unfinished ideas fast, and the draining may hollow out the field that produced them.

### [Extra] CVS drew one line before scaling anything: AI never makes the call.

Before scaling anything, [CVS Health](https://www.cvshealth.com/news/innovation/aetna-reduces-claims-processing-time-by-more-than-20-percent-with-ai-to-improve-care-experience.html), the American pharmacy and health-insurance company, drew a hard boundary: AI never makes the decision. Its stated policy carries three nevers: never diagnose a patient, never drive a claim denial, never take the human out of the patient's experience. In practice, Aetna's agentic Claims Assist Manager assembles eligibility, coverage and provider data and recommends a next step, and the company's chief psychiatric officer has said plainly that AI "will not ever be used to make an adverse decision on a case". Complex-claim processing time fell by more than 20%. A one-line refusal travels through an organisation faster than any strategy deck, and it's a cleaner story to tell a nervous board than a capability demo.

### [Extra] Console-flation: AI's chip hunger is now on the price tag of a games console.

The Guardian [reports "console-flation"](https://www.theguardian.com/games/2026/jul/01/pushing-buttons-ai-datacentres-memory-console-prices-sony-playstation-xbox), games-console prices rising as AI chip demand and memory scarcity bite. The abstract capital-spending story, the billions poured into data centres, has arrived at a consumer checkout most people recognise. It's a small thing next to the water-use letters and the multi-billion-dollar builds, but it's the one that lands: the footprint of all that AI investment now shows up on the shelf next to a household object, chosen by a family that has no view on foundation models at all.

## Edition #20: Look at the grass
*4th July 2026*

### [Three Things] The people with the most skin in the AI economy read it both ways this week

The Bank for International Settlements, the institution central banks themselves bank with, [warned that AI "exuberance"](https://www.ft.com/content/e81ce414-e4bd-4e8c-bac7-94f7bf17def4) could end in a long investment bust. The warning turns on returns rather than capability: whether the money spent on AI earns back enough to justify itself, however well the technology works. The jobs data split the same way in the same week. A study of 21,559 US firms by Ramp, a corporate-spending platform, and Revelio Labs, a workforce-data firm, [reported by the Financial Times](https://www.ft.com/content/8026eac6-16ad-467d-b8c3-c48c5af684e6), found heavy AI adopters grew employment about 10% after adoption and lifted entry-level hiring 12%, a pattern [Noah Smith, the economics writer](https://www.noahpinion.blog/p/roundup-84-bears-on-bikes), read as AI complementing people rather than replacing them. The same day, [The Kobeissi Letter](https://x.com/KobeissiLetter/status/2071715519275958388), a markets commentary account, reported AI-linked sectors shedding about 11,000 jobs a month and the US labour share of income at its lowest since records began in 1947. These are not contradictions to resolve. They describe different populations: the firms mastering the technology, and the sectors it disrupts from outside. So two board questions collapse into one. Are we the group that masters this, or the group it happens to?

[Image: Financial Times chart of BIS data showing AI investment at about 4.5 times its pre-boom trough after three years, a steeper climb than railway mania, canal mania, the Roaring Twenties and the dotcom boom at the same point in their cycles.]
*Three years in, AI investment is already steeper than the railways, the 1920s or the dotcom boom were at the same point.*
[Image: Financial Times chart showing companies that spend more on AI also increase worker numbers: high AI adopters grew headcount about 10% on average in the two years after adoption, 12% at entry level, while low adopters saw no change.]
*The other reading: firms spending most on AI grew headcount about 10%. Low spenders barely moved.*

### [Three Things] Zapier is killing private direct messages to feed the company's "shared brain"

Wade Foster, chief executive of the automation company Zapier, is [ending private direct messages](https://x.com/wadefoster/status/2072665341650407492), starting with the executive team, and running a public "transparency leaderboard" that has pushed the share of conversation happening in open channels from 33% to 46%. His logic is worth stealing: every private message is a gap in the company's shared brain, context lost to both the humans and the AI agents that increasingly work over the company's own records. The bottleneck in most organisations' AI adoption is context that never gets written down, not model quality. If half your knowledge lives in direct messages and people's heads, no agent can reach it, however good the model. Most firms bolt AI onto the way they already work. Zapier is redesigning the way it works so the AI can pull its full weight, and that is the step that separates the leaders from the licence-buyers.

### [Three Things] The Economist ran 25 AI models through the World Values Survey. They came back more extreme than any country

The World Values Survey has mapped human morals across more than 100 countries since 1981. Run 25 leading models through it, as [The Economist did](https://www.economist.com/briefing/2026/06/25/ai-models-values-are-very-different-from-most-peoples), and they didn't land at the average of humanity. They clustered in the corner occupied by rich, secular, self-expression-focused countries, often further out than the most extreme nation surveyed. OpenAI's models came back more secular than any country. None reflected the world view of most African or Muslim countries. When a model summarises the news or weighs in on a family row, there's no factually correct answer to fall back on, so its values fill the gap. A tool a billion people now delegate decisions to carries a world view measurably narrower than the human range, baked in through training data and the value choices a handful of firms make in training them. It's the values version of [a point from a fortnight ago](https://steadman.ai/newsletters/david/archive.html#email-2026-06-20): unless you bring yourself, the machine's defaults stand in for you.

### [Extra] AI wiped out 95% of China's micro-drama industry in a single quarter.

After ByteDance, the owner of TikTok, shipped its Seedance 2.0 video model in February, [122,000 of the 128,000 micro-dramas](https://www.caixinglobal.com/2026-06-01/cover-story-how-ai-took-over-chinas-micro-drama-industry-in-90-days-102449814.html) that went live in China in the first quarter of 2026 were AI-made. Hengdian, sometimes called China's Hollywood, saw live-action production drop by an estimated 70 to 80%. The industry took years to build and about three months to gut. It's the hardest displacement number of the week. Read it narrowly: where a creative task is fully automatable and nothing slows adoption, the substitution can be near-total and near-instant. Most creative work isn't like that. This corner of it was.

### [Extra] A story critics call "obviously AI-written" just won the Commonwealth Short Story Prize.

Jamir Nazir's *The Serpent in the Grove*, which critics allege carries ["obvious markers" of AI use](https://www.theguardian.com/books/2026/jul/01/judges-claims-ai-use-commonwealth-short-story-prize-jamir-nazir), won the overall Commonwealth Short Story Prize. The judging chair called it "original, poetic and deeply moving". Then the Commonwealth Foundation ran a formal review, examined the drafts and the time-stamped working notes, and cleared the story, satisfied that AI was not used. Many readers stayed unconvinced, and Granta, the literary magazine, pulled out of its long-running deal to publish the winners anyway. So the process worked exactly as designed, reached a clear verdict, and settled almost nothing. When the suspicion outlives the adjudication, "can you always tell" stops being a question a review can put to rest.

### [Extra] An AI model now completes one in six real freelance jobs to a professional standard.

The Remote Labor Index, from the Center for AI Safety and Scale AI, tests models against 240 real projects from professional freelancers across 23 domains, worth over $140,000 of human work. [The new results](https://x.com/CAIS/status/2072360965522489789) put Claude Fable 5 at 16%, roughly double the next model and four times the published leader of five months ago. One in six sounds small until you look at the slope: at this rate of doubling, it won't stay small for long.

[Image: Remote Labor Index bar chart: Claude Fable 5 completes 16% of real freelance projects to a professional standard, roughly double the next model and four times the previous published leader from five months earlier.]

### [Extra] A new benchmark asks whether an agent can do your analysts' job. The answer varies 17x in price.

Artificial Analysis, an AI-benchmarking firm, built [AA-Briefcase](https://artificialanalysis.ai/evaluations/aa-briefcase): 91 tasks across four multi-week consulting-style scenarios, closer to an analyst's real week than to quiz questions. Claude Fable 5 leads the table, and the second-placed model spans a roughly 17x cost-per-task range across its five effort settings, so what you pay depends on how hard you ask it to think. [Ethan Mollick](https://x.com/emollick/status/2071449550662127927), a Wharton professor who studies AI, graphed the scores over time: rapid gains for both open and closed models, with a clear, persistent gap to the best open ones.

[Image: Artificial Analysis AA-Briefcase charts: Elo versus cost per task scatter and the full leaderboard, Claude Fable 5 first at 1585, Claude Sonnet 5 second at 1391 with a roughly 17x cost-per-task range across its five effort settings.]
[Image: Ethan Mollick's scatter of AA-Briefcase scores by launch date: open-weight and closed-source frontiers both rising rapidly, with closed models holding a clear gap over the open-weight leaders.]

### [Extra] The UK's AI Security Institute says a model's capability is a curve, not a number.

On [the institute's cyber challenges](https://www.aisi.gov.uk/blog/more-compute-more-capability-why-ai-agent-evals-need-to-account-for-test-time-compute), giving frontier models more computing power at the moment they run stretches the length of task they can reliably finish. Give them a lot and that reach roughly doubles every 40 to 50 days; give them a little and every 67 to 91 days. The institute's conclusion is the useful part: a model's capability should be reported as a curve against how much computing you give it, not a single score. Anyone comparing models on one number is measuring a point on a line that someone else can move by paying more.

[Image: AI Security Institute chart: on AISI's cyber challenges, more test-time compute stretches the task length frontier models complete reliably, with the time horizon doubling every 40 to 50 days at high compute budgets versus 67 to 91 days at low budgets.]

### [Extra] Vulnerability disclosures spiked to three and a half times the monthly record right after Mythos was announced.

[Epoch AI](https://x.com/EpochAIResearch/status/2072776792809918604), a research group that tracks AI progress, counted critical and high-severity security-vulnerability disclosures from 21 major organisations, Apple and Microsoft through to the Linux ecosystem: roughly 1,500 in June 2026, over three and a half times the previous monthly record, immediately after Claude Mythos Preview was announced. The likeliest reading is that AI is now finding software vulnerabilities at scale. Epoch's own caveat is worth keeping: reporting practices vary between organisations, so some of the spike may be reporting behaviour rather than discovery rate. Either way, the finding side of security has changed gear.

[Image: Epoch AI chart: critical and high-severity CVE disclosures from 21 notable organisations spiked to about 1,500 in June 2026, over 3.5 times the previous monthly record, immediately after Claude Mythos Preview was announced.]

### [Extra] A product category that did not exist four years ago is out-earning Salesforce.

The Information, a technology news site, charts annualised revenues showing [Anthropic at $47bn](https://x.com/ainativefirm/status/2071417556796047497), just past Salesforce at $44.5bn, with OpenAI estimated at $30bn. Read it less as one company's curve and more as a statement about the category: enterprise software's benchmark firm took a quarter of a century to build that revenue base, and the model labs did it in four years, most of it in the last two. The BIS may still be right about the returns on all the spending. The revenue, though, is no longer hypothetical.

[Image: The Information chart titled AI Rising: Anthropic's annualised revenue at $47 billion has just overtaken Salesforce at $44.5 billion, with OpenAI estimated at $30 billion, Anthropic's curve rising sharply from late 2025.]

### [Extra] Cut the price of tokens 10% and total spend goes up, not down.

Azeem Azhar, who writes the Exponential View, makes the demand-side case in one chart from his ["State of the AI Economy"](https://x.com/SaaSletter/status/2070878567861514247) deck: token demand looks elastic, with every 10% price cut bringing 12 to 18% more usage, so total spend rises as prices fall. It is the bull answer to the BIS warning in this week's lead: bubbles pop when the demand is imagined, and this chart argues the demand is measured. Hold both, and the question becomes whether the elasticity survives once the novelty does.

[Image: Slide 44 of Azeem Azhar's State of the AI Economy deck: token demand appears elastic, with an elasticity of roughly 1.2 to 1.8, so every 10% price cut brings 12 to 18% more token usage and total spend still rises.]

### [Extra] Venture money is rotating "from bits to atoms", and robotics is now the second-biggest private category by value.

In the phrase of [Andreessen Horowitz](https://x.com/a16z/status/2070531376890495293), the venture-capital firm that coined it, money is rotating "from bits to atoms". New data from the firm shows robotics and physical AI are now the second-largest category of private companies by value, having overtaken financial technology and payments. The first quarter set a record at around $16bn across roughly 500 deals, about four and a half times the previous period. Read it as a view from a house that invests in robotics, not neutral analysis. The firm's own open question is the honest one: whether this is a temporary cycle, where today's infrastructure spending flows through to software the way chips gave way to apps after 2008, or a durable shift toward hardware as a product in its own right.

### [Extra] From the same charts: AI-native startups run at half the headcount, at every stage.

The same edition of the firm's [Charts of the Week](https://x.com/a16z/status/2070531376890495293) carries a second chart worth holding: AI-native startups from Y Combinator, the startup accelerator, run at roughly half the employees of their non-AI peers at every stage of life, about 22 against 45 by their twelfth quarter, per the underlying [Kim and Koning study](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6905079). It is the firm-level version of this week's jobs story: the hiring never happens, so it never shows up as a layoff. If the ratio holds as these companies grow up, the employment question shifts from who gets cut to who never gets hired.

[Image: a16z Charts of the Week line chart showing AI-native Y Combinator startups running at roughly half the headcount of non-AI startups at every stage, about 22 versus 45 mean active employees by quarter twelve.]

### [Extra] Meta is reportedly building "Meta Compute" to sell its surplus, and detonated the shortage trade.

Meta is [reportedly building "Meta Compute"](https://www.bloomberg.com/news/articles/2026-07-01/meta-is-building-a-cloud-business-to-sell-excess-ai-compute) to sell access to its AI infrastructure, competing with the cloud arms of Amazon, Microsoft and Google. The market's reaction was the story: Meta rose almost 9% while the Philadelphia Semiconductor Index fell over 6%, with several chip and cloud names down double digits. A fortnight earlier, computing power looked so scarce that Google was reported to be [capping Meta's use](https://www.ft.com/content/c5d52f72-71ef-40bc-bad3-61afdba8b378) of its Gemini models. Now the biggest spender on AI is preparing to sell a surplus. If the company spending the most expects a glut, the "shortage is permanent" thesis under much of the AI investment story wobbles, and the sell-off suggests the market took it seriously.

### [Extra] Ford quietly rehired 350 engineers it had cut for AI, then won its category for the first time in 16 years.

Ford rehired [350 experienced quality engineers](https://thenextweb.com/news/ford-rehired-350-engineers-ai-quality-jd-power) after the AI quality systems it had cut staff for failed, then took the top mainstream-brand spot in J.D. Power's 2026 rankings for the first time in 16 years. Charles Poon, a Ford vice-president, told reporters the company had wrongly assumed it could drop in AI and still ship a high-quality car. The tacit knowledge had walked out the door with the people who held it, before the AI systems could encode it, so Ford had to buy it back. The lesson is about order of operations. Encode the expertise while the people who hold it are still there, then thin the team. Do it the other way round and you pay to hire it back.

## Edition #19: A thousand small bargains
*27th June 2026*

### [Three Things] Claude has stopped opening in a window and started living in the team's Slack

This week [Anthropic launched Claude Tag](https://www.anthropic.com/news/introducing-claude-tag): you @-mention it in a Slack channel and it works as a persistent teammate, with its own identity, scoped access, memory across weeks, and the run of a stalled task for days at a time. Anthropic says **65%** of its internal product code now gets written this way. Andrej Karpathy, who co-founded OpenAI, [mirrored a framing I wrote about earlier this month](https://steadman.ai/newsletters/david/three-generations.html), [calling it the third major redesign](https://x.com/karpathy/status/2069547676849557725) of how we use these models: first a website, then an app, now a persistent worker. It's the first frontier product where the interface is openly about managing an AI employee. The pattern will spread; the human problems, supervision, accountability, headcount maths, arrive with it.

### [Three Things] Procter and Gamble actually ran the AI-alone-versus-human-plus-AI test, and human-plus-AI still won

On stage at the [Lions Insight Summit](https://www.canneslions.com/festival/experiences/lions-insights/programme) in Cannes, [Kirti Singh](https://us.pg.com/leadership-team/kirti-singh/), P&G's Chief Analytics, Insights and Media Officer, told me the company has run the experiment most leaders only theorise about. They let AI alone make some brand-building choices, measured the outcomes against humans working with AI, and the combination won. "Today it is not better," he said carefully, "than what human plus AI can do." So the company's settled position is human-plus-AI. This is not a sceptic hedging: he says his teams under-use AI and pushes them to use more. A measured floor, not a slogan.

### [Three Things] Norway just banned generative AI for six-to-thirteen-year-olds

From late August 2026, Norway imposes a [near-total ban on generative AI](https://thenextweb.com/news/norway-bans-generative-ai-elementary-school-children) for children aged six to thirteen, with supervised use only for fourteen-to-sixteen-year-olds. The prime minister's argument is that AI lets young children skip the essential steps in learning to read, write and do maths. The same government is funding a return of physical books and already banned school smartphones in 2024. [Luiza Jarovsky](https://x.com/LuizaJarovsky/status/2069125922150551590), a tech-policy writer, backed the move. Most of the AI conversation this year is about going faster. A government has decided that for one age group the danger is precisely the speed: skipping the hard steps that make a child's judgement worth anything later.

### [Extra] The build-out bill arrives on both sides of the supply chain.

[Apple](https://www.ft.com/content/0f067265-2baf-4b6e-8fb2-ed56daef6f3c), the consumer electronics maker, raised MacBook and iPad prices by about 20%, blaming an AI-driven memory shortage. The Financial Times reported the move wiped roughly $263bn off the company's market value on the day. It's the clearest spillover yet from the data-centre build-out into the price of a laptop that has nothing to do with AI. The same memory chips the hyperscalers are hoovering up are the ones in consumer devices.

The other side of the squeeze landed the same week. [Micron](https://www.ft.com/content/9b739203-3274-43f1-b61e-c1905061d32a), the memory-chip maker, posted profits up nearly 1,400% on the AI scramble, with shares rallying on a forecast of sustained demand. Micron is also taking part in Anthropic's latest funding round. The capex argument has been abstract for two years. This week it showed up at the checkout on one side and in a fifteen-fold earnings print on the other.

### [Extra] A rumour nobody could verify wiped trillions off Asian markets in a day.

South Korea's KOSPI index fell about 10% in a single day, tripping its circuit breakers twice. Samsung and SK Hynix both dropped more than 12%, and the move crossed the Pacific: the Nasdaq fell two to three percent and Micron, the same memory-chip maker, dropped roughly nine to thirteen percent. Analysts could name no single trigger and called it [mechanical](https://www.home.saxo/content/articles/options/options-brief---korea-hits-double-circuit-breaker-as-ai-trade-corrects--23-june-2026-23062026): automated selling, forced liquidation of leveraged retail positions, and rebalancing. An unverifiable rumour was enough to wipe out trillions. Whatever you think of the technology, the market financing it now moves violently on nothing.

### [Extra] A peer-reviewed study shows doctors' detection rates drop six points after they've used AI.

A study in [The Lancet](https://www.thelancet.com/journals/langas/article/PIIS2468-1253(25)00133-5/abstract), the medical journal, found that the rate at which clinicians spotted growths during colonoscopies done without AI fell from 28.4% to 22.4% after those same clinicians had spent time working with AI assistance. Switched on, the tool lifts detection by around 20%. Switched off, the doctors who had leaned on it were measurably worse than before. The honest framing isn't anti-AI. It's about which skills you cannot afford to let soften: the tool genuinely works, so the discipline is to pick the few that matter and keep your hand in on those, even when the machine could carry them.

### [Extra] AI out-persuades expert debaters. Slow it to human typing speed and the edge disappears.

A four-institution study by [Oxford, the UK AI Security Institute, Stanford and the London School of Economics](https://arxiv.org/abs/2606.16475) ran 18,978 conversations across 6,923 people and found AI more persuasive than every class of human tested, including elite debaters who researched in advance with a £1,000 incentive. It was nearly three times more effective than professional charity canvassers at raising real money. The finding worth sitting with: when researchers throttled the AI to human writing speed and message length, its edge all but disappeared. The persuasive power isn't deeper insight. It's the rate at which it deploys information, and a rate is far easier to rein in than some hidden gift for manipulation.

### [Extra] The diligence test that used to be analysis is now construction.

[Bain](https://www.ft.com/content/e5bac4d1-b1f8-43a4-bd54-b182d5357af0), the management consultancy, has started recreating a target software product quickly using AI, to gauge its real competitive advantages before committing to a deal, the Financial Times reported. The technique is copyable. Build a rough functional clone in hours, then test how defensible the original actually is once the surface features are easy to replicate. Due diligence has always asked how defensible something is. AI changes the test from analysis to construction. Don't argue about the moat. Try to rebuild it and see how far you get. For anyone doing software or product diligence, this is a planning tool, not a theoretical observation.

### [Extra] Carl Frey on the demand-side mirror of the jobs debate.

[Carl Benedikt Frey](https://www.project-syndicate.org/commentary/ai-abundance-without-purpose-meaning-by-carl-benedikt-frey-2026-06), the Oxford economist, argues in Project Syndicate that AI-powered entertainment and companionship could capture enough of people's time and social appetite to displace the harder activities that actually generate meaning. He reaches for Kurt Vonnegut's *Player Piano*, the novel where machines automate industry and leave everyone fed, housed and idle. Most coverage prices the supply shock, jobs lost and tasks automated, and ignores the demand shock, what fills the freed attention. For an audience-strategy practice, that's the more interesting half of the question. Not what work disappears, but whether anyone wants what is left.

### [Extra] The business AI race has narrowed to two names.

[Ramp's data](https://www.visualcapitalist.com/ranked-ai-models-u-s-businesses-pay-for/), across more than 50,000 US businesses, tracks who companies actually pay for. In this snapshot from earlier in the year, OpenAI led at 35.2% with Anthropic close behind at 30.6% and climbing near-vertically on enterprise tools like Claude Code and Cowork. Google trailed at 4.3%, xAI at 1.9%. Paid adoption is a harder signal than downloads or hype, because someone had to sign a purchase order. By the chart's end the two leaders were all but converging.

[Image: Ramp data via Visual Capitalist: share of US businesses paying for each AI model, OpenAI 35.2% and Anthropic 30.6%]

### [Extra] Generative AI is reaching mass adoption faster than any technology before it.

[J.P. Morgan Asset Management](https://am.jpmorgan.com/us/en/asset-management/adv/insights/market-themes/artificial-intelligence/) plotted how long each technology took to reach 75% adoption in the US, from the flush toilet (about 75 years) to the smartphone (about ten). Generative AI sits alone in the bottom corner, faster than all of them. The footnote does the work: usage was already 58% of US adults aged 18 to 64 by February 2026. The adoption argument is effectively over. The only live question is what people do with it.

[Image: J.P. Morgan Asset Management: time to mass adoption for new technologies, with generative AI the fastest on record]

### [Extra] McKinsey drew the org chart for humans, agents and robots, and the human column is judgement.

[QuantumBlack](https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-symbiotic-enterprise), McKinsey's AI arm, splits the work of a "symbiotic enterprise" three ways. Humans set strategy, build trust, govern, and arbitrate exceptions. AI agents do the complex cognitive work and orchestrate the workflows. Robots handle the physical. What sits in the human circle is this week's essay almost word for word: judgement on the agent's output, and the call no model makes for you. A consultancy's tidy map of where authorship stays.

[Image: McKinsey QuantumBlack: role allocation in the symbiotic enterprise across humans, AI agents and intelligent robots]

### [Extra] The mental abilities that fade with age are the ones AI is cheapest at. The ones that last are yours.

A much-shared chart, drawn from [cognitive-ageing research](https://pmc.ncbi.nlm.nih.gov/articles/PMC3359129/), plots performance by decade. Processing speed, working memory and recall all fall from your twenties on. World knowledge, the crystallised judgement built from experience, holds flat or climbs into your eighties. It is a useful map for the AI question. The machine is best at the fast, fluid work that was going to fade anyway. The durable human edge is the slow, accumulated judgement it cannot yet hold.

[Image: Cognitive performance by age: processing speed and memory decline while world knowledge is preserved]

## Edition #18: Average by default
*20th June 2026*

### [Three Things] A Munich court told Google that "may contain errors" is no defence

The [Munich Regional Court](https://the-decoder.com/landmark-german-ruling-declares-googles-ai-overviews-are-googles-own-words-and-makes-it-liable-for-false-answers/) ruled Google liable for false claims made by its AI Overview, which had wrongly tied two German publishers to fraud. The judges held that an AI summary generates fresh, substantive statements rather than just curating sources, and that the small-print "may contain errors" footer does not transfer the liability back to the user. Google must remove the answers and pay 80% of the costs. Every major AI provider leans on the same disclaimer to cover confidently wrong answers. A court has now said it does not work. If your organisation publishes AI output to customers, regulators or staff, the footer that has been doing the legal lifting may not protect you next time.

### [Three Things] Commercially available AI flagged early breast cancer six years before clinical diagnosis

Researchers at the [Karolinska Institute](https://www.rsna.org/news/2026/june/ai-flags-breast-cancer-early), the Swedish medical university, published a study in Radiology showing that off-the-shelf AI systems flagged early warning signs of breast cancer in roughly 20% of patients a full six years before a clinical diagnosis, at 90% specificity, rising to around 40% at two years out. The analysis ran across 88,963 mammograms from more than 31,000 patients. Three of the tools tested are already commercially available. Pattern recognition at scale is what AI does best. Here it buys years of warning on a disease where early detection saves lives, and the bottleneck is no longer capability. It is deployment, trust and who owns the answer.

### [Three Things] Anthropic studied 400,000 coding sessions: the tool is levelling people, and moving the real gap up to firms

[Image: People decide what to do, Claude decides how to do it]
*People still make about seven in ten of the decisions on what to do; the agent makes eight in ten on how to do it.*
[Anthropic analysed 400,000 sessions](https://www.anthropic.com/research/claude-code-expertise) of its coding agent across roughly 235,000 people. Engineers did well, and so, surprisingly, did almost everyone else: across the ten largest professions, success rates landed within seven points of professional engineers, and managers came out ahead. An accountant who had never written a line of Python, but knew which rule a month-end reconciliation had to enforce, was rated an expert; a senior engineer on an unfamiliar task was not. What carries a session is understanding the problem, not the craft. So the gap between people is closing. The part everyone skips is that it has moved up a level rather than vanished, from between people to between firms. [The expanded version is online ->](#wrong-gap-2026-06-20)

### [Extra] Half of Americans now use a chatbot. Forty percent expect AI to make society worse.

A new survey of more than 5,000 US adults by [Pew Research Center](https://www.pewresearch.org/internet/2026/06/17/americans-and-ai-2026-chatbots-smart-devices-and-views-on-impact/), the polling organisation, finds about half the country now uses a chatbot and a quarter use one daily, up from a third of the public in 2024. Yet nearly 40% expect AI to make society worse over the next twenty years, against 16% who expect it to improve. People are reaching for the tools and dreading them at the same time. The adoption-trust scissors keeps widening: usage climbs while confidence in where it leads keeps falling. Treating an adoption number as a proxy for buy-in misreads the room. The same people using it most are often the ones most worried about it.

### [Extra] The AI dividend is being eaten by botsitting and bossware.

[Glean](https://www.glean.com/work-ai-institute/reports/work-ai-index-report), an enterprise AI search company, surveyed 6,000 workers for its Work AI Index 2026 and found that heavier AI use often hides more cleanup and context-switching, not less. The wider picture is sharper. Time spent on email has doubled. Focused-work sessions are down 9%. And 61% of workplaces in the survey now run AI productivity-monitoring software on top of the AI tools themselves. The promised dividend is being absorbed by checking the machine's output and by managers watching the watchers. Worth a caveat: Glean's commercial story is that other firms' deployments are messy. The 61% bossware number is the one to carry, because it doesn't depend on Glean's framing.

### [Extra] The mid-market has nobody obvious to call.

Dan Monaghan at WSI, a marketing consultancy, did the maths on consulting's AI windfall. A quarter of [BCG's $14.4bn 2025 revenue](https://www.bloomberg.com/news/articles/2026-04-23/boston-consulting-group-says-ai-work-brought-25-of-2025-revenue), about $3.6bn, came from AI services. The sharper observation sits beneath it. The big firms are going upmarket, chasing the Fortune 500, which leaves Monaghan's count of roughly 1.7 million US businesses over a million dollars in revenue with nobody obvious to call. It's the same pattern as the early internet: the giants serve the giants, and a vast middle is left to work it out alone. For most of those firms the practical answer is rarely a big-firm engagement. It's a smaller adviser, or learning to do it themselves with tools cheap enough to put real capability in reach.

### [Extra] A new independent safety lab is raising up to $150m to yell when the labs won't.

Researchers from the [UK AI Security Institute's alignment team](https://www.alignmentforum.org/posts/AP7YDke5jjY4v3X9Z/sequent-scale-and-automation-for-higher-confidence-in-1) and the startup Timaeus have launched Sequent, a non-profit aiming for 40 to 80 staff and a raise of $100m to $150m, with stated willingness to go an order of magnitude bigger. Its pitch is a counter-narrative to the consensus that safety is a solved engineering detail: the frontier labs' own alignment work is essentially reactive, and an independent body is needed that can yell if a lab does something dangerous. Most coverage treats AI safety as either solved or hopeless. A well-funded independent body whose entire job is to be the loud sceptic is useful ballast for anyone tempted to assume the labs have it handled.

### [Extra] 85% of IT leaders say every AI agent has a named owner. Only 42% say that ownership is actually clear.

A survey of 3,900 employees across six countries by [Ivanti](https://www.ivanti.com/resources/research-reports/scaling-ai-it-operations), a security software company, put a number on the shadow-AI governance gap. 85% of IT professionals say every AI agent has a named owner. Only 42% say that ownership is actually clear. The sharper stat is the leader-staff inversion: leaders hide their own AI use more than their staff do, 42% versus 23%, with more than half citing a "secret advantage" as the reason. Governance dashboards report the 85%; the organisation runs on the 42%. The people meant to set the policy are hiding their use to keep an edge. The Ivanti survey has a commercial interest in the gap looking large, but the direction is the story.

### [Extra] Databricks says agents are pushing customers off "tokenmaxxing."

[Databricks](https://www.cnbc.com/2026/06/16/databricks-revenue-growth-tops-80percent-to-6point9-billion-annualized.html), the data and AI platform, reported annualised revenue up more than 80% to $6.9bn, with $1.7bn from AI. Its chief executive Ali Ghodsi says agentic queries are now flooding the platform, shrinking margins, and pushing customers away from "tokenmaxxing" toward budget controls, deliberate model choice, and cheaper open-source models for routine work. The whole market is making that turn at once: from maximising tokens to managing them. The discipline that's coming, cheap model for the routine work and the dearest model where it earns its keep, is the cascade thoughtful firms are already building.

### [Extra] AI's "anti-intelligence": when the output looks like thinking and isn't.

[John Nosta](https://x.com/JohnNosta/status/2067348306234089811), who writes on technology and the mind, named four ways AI counterfeits thought: performative intelligence, reasoning that collapses under scrutiny; compressed cognition, where the answer arrives but the thinking that should have shaped it is skipped; displaced agency, where the origin of an idea quietly shifts while ownership still feels intact; and synthetic conviction, output carrying more confidence than the thinking behind it has earned. His line underneath the four: the outputs remain, the thinking does not. It names what makes a slick AI draft so disarming. It looks like the product of thought, so we stop checking whether any happened.

[Image: The four faces of anti-intelligence]

### [Extra] The AI slop question, charted.

[The Economist](https://x.com/TheEconomist/status/2067155325623365658) tried to weigh the "AI slop" panic and found it real in places and overblown in others. AI-generated e-books on Amazon have overtaken human ones, AI music uploaded to Deezer keeps climbing while human uploads flatten, new app releases spiked the moment Claude Code and Codex shipped, and even self-filed US civil lawsuits jumped. The effect is uneven, but the rule holds: wherever the cost of producing something collapsed, the volume exploded and the median quality slid. It is the water this edition's essay is swimming in. When everyone pours from the same tap, most of what comes out tastes the same.

[Image: AI-generated content rising across e-books, music, apps and lawsuits]

### [Extra] The biggest AI time-savers use eight times more of it.

[OpenAI's State of Enterprise AI report](https://x.com/rohanpaul_ai/status/2066854804253757714) found the workers saving more than ten hours a week use about eight times more AI than those saving none, and not by running more of the same queries. They reach across several models, more tools, and a wider span of tasks. The returns are not linear, they compound with depth. It lands where Anthropic's coding study this week does: handing everyone a chatbot buys a small, uniform bump; the real gains go to the people who fold the tool deep into how they actually work.

[Image: Productivity gains rise with the intensity of AI use]

### [Extra] Undisclosed AI is seeping into the op-ed pages.

[Mohit Iyyer and colleagues](https://x.com/MohitIyyer/status/2067647168307986821), the group behind the "argument collapse" finding in this week's essay, ran op-eds from America's most-read papers through an AI detector. Under 1% of Washington Post op-eds in March 2025 were flagged; a year later it was 11.9%, with the Times and the Journal lower but climbing, and none of it disclosed. The opinion page is the one place where a distinctive voice is the entire product. If even that is drifting toward the machine's average, the flattening is further along than the comfortable assumption that you can always tell.

[Image: AI-or-mixed share of op-eds rising across the NYT, WSJ and Washington Post]

### [Extra] A new benchmark for real office work, and the best model clears 3% of it.

[Artificial Analysis](https://x.com/ArtificialAnlys/status/2067744637155226101) built AA-Briefcase to test models on multi-week projects buried in thousands of messy files: the texture of actual knowledge work rather than tidy one-shot prompts. Claude Fable 5 led at 1587 Elo, graded just before the model was pulled, ahead of Opus 4.8 and several far cheaper open-weight models. The sobering figure sits beneath the ranking. Even the leader meets every rubric check on only 3% of tasks, and on a third of them no model clears half. Single-question benchmarks flatter these tools. Give them a real job, with real ambiguity, and the distance to a capable professional is still long.

[Image: AA-Briefcase Elo leaderboard, Claude Fable 5 in front]

### [Extra] One model did in ten minutes what a team of people needed six hours for.

On four hardware tasks, a video camera, a lidar, a control program and localisation, [Claude Opus 4.7 working alone](https://x.com/scaling01/status/2067653429258862735) averaged under ten minutes, against 181 for a team using Claude and 361 for a team without it: roughly twenty times faster than the quickest humans, and on the older model, not the newest. Treat any vendor-run demo with the usual caution. But the direction is the one running through this edition. On well-specified, checkable tasks the machine is pulling clear, and the human's value migrates to the work that is neither tidy to specify nor easy to verify.

[Image: Opus 4.7 alone against human teams on four hardware tasks]

## Edition #17: Ride the bike
*13th June 2026*

### [Three Things] One activist letter froze a board's AI workspace overnight. The doctors are next

I helped a hundred non-executive directors with AI this week. One told me about a board whose workspace went dark the morning an activist investor wrote in asking for its contents to be made discoverable. Two-thirds of directors now use AI for board work; barely a quarter of executives call their board highly fluent, per [the Conference Board's April survey](https://www.conference-board.org/press/governing-AI-2026). The Medical Protection Society, which defends clinicians against negligence claims, [warned this week](https://www.medscape.com/viewarticle/doctors-risk-becoming-liability-sink-ai-errors-2026a1000785) that doctors and the NHS could be sued over mistakes made by AI tools, with the clinician left as the "liability sink" for the technology's errors. Its report wants AI reclassified as a product under the Consumer Protection Act 1987, so liability flows to developers. The frontier of AI governance is moving from "can the data leak" to "who owns the answer when the model is wrong". Our policy, that every output is checked, edited and owned (CEO'd) by a human, answers that question. The letters will keep coming, and the writs will follow. The only good answer to both is a name.

### [Three Things] Meta built a "second brain" that 63,000 staff installed in three months. It started with one person

[Meta's analytics team reports](https://medium.com/@AnalyticsAtMeta/how-we-built-an-ai-second-brain-for-60k-knowledge-workers-78c507dd795b) that an internal AI tool one of its data scientists started has now been installed by 63,000 employees, a number reached in three months. No top-down mandate, no transformation programme. One person built something useful and the rest of the company found it. What are you doing to encourage and enable this at your firm?

### [Three Things] When cheap models do make sense

Ethan Mollick, a Wharton professor, [argues for hierarchies](https://x.com/emollick/status/2064764123439599696) in which smart models supervise cheap ones: the smart one checks the plan, the cheap one does the volume. Right for machine pipelines running thousands of low-stakes calls. For your own judgement work, this week's essay argues that you should buy the best; the maths of time saved is the reason.

### [Extra] The most capable model ever made went on sale Tuesday. By Saturday the US government had pulled it.

On Tuesday, Anthropic released Fable 5, the most capable model it has ever shipped. By this morning it was gone. A US government [export-control directive](https://www.anthropic.com/news/fable-mythos-access) barred foreign nationals, anywhere, from using Fable 5 or Mythos 5; rather than carve up its customers, Anthropic [suspended both models for everyone](https://x.com/AnthropicAI/status/2065597531644743999) while it contests the order. The stated trigger was a "narrow" jailbreak in which the model reads a codebase and fixes its flaws, which is roughly the thing that makes it useful. One widely shared [analysis](https://x.com/scaling01/status/2065607360115302639) reckoned that, applied to every frontier lab, this leaves American models unsellable abroad, locks foreign staff out of their own products, and hands China the lead. I will let others judge that. I can tell you I had enormous fun with Fable 5 this week, spent five to six thousand dollars of API credits putting it through real work, and found it comfortably the most capable model I have used. Older Claude models are unaffected.

[Image: Financial Times headline: "Anthropic suspends latest AI models after US blocks access to foreigners"]

### [Extra] The safety hawks wanted a brake on frontier AI. This week they got one, just not the way they pictured it.

For three years the loudest voices on AI safety, Max Tegmark, the MIT physicist, among them, [argued someone should be able to halt a dangerous model](https://www.economist.com/science-and-technology/2026/06/07/how-artificial-intelligence-got-better-at-building-itself). Anthropic itself [set out the conditions](https://www.anthropic.com/institute/recursive-self-improvement) under which it would pause, and OpenAI's Sam Altman [called for a body](https://openai.com/index/built-to-benefit-everyone-our-plan/) empowered to slow frontier development "when needed". This week the brake got pulled, just not by them: the US government switched a frontier model off (above). One widely shared [post](https://x.com/willmanidis/status/2065596811683795320) caught the whiplash, noting that Anthropic's own chief executive had argued days earlier that government should be able to block a model's deployment, and that the reaction when it actually happened was, roughly, "not like that". Be careful which brake you ask for.

### [Extra] Investors are quietly buying companies an AI can't run.

[The Information's Dealmaker desk](https://www.theinformation.com/newsletters/dealmaker/school-bus-startup-zum-takes-early-steps-toward-ipo) flagged a new trade this week: initial public offerings of physical-world operators a language model can't touch. Zum, the school-bus firm backed by the venture firm Sequoia Capital, has grown revenue 35% to $333m, turned profitable and is now interviewing banks. [Nabeel Hyatt](https://www.sparkcapital.com/team-members/nabeel-hyatt) of Spark Capital, another venture firm, and himself an investor in Anthropic, the AI lab, put the logic in one line: "It's very unlikely that an Anthropic will run a bus company." The AI-proof hedge has become a thesis.

### [Extra] Palantir's Alex Karp tells executives to stop bragging about job cuts.

[Alex Karp](https://www.palantir.com/leadership/), chief executive of the data-analytics firm Palantir, said on a tech podcast this week that anyone publicly touting they've fired two-thirds of their staff might as well [sign up for the Bernie Sanders manifesto](https://fortune.com/2026/06/09/palantir-ceo-alex-karp-massive-ai-layoffs-may-be-bad-for-industry-future-of-work-bernie-sanders-headcount-reductions/). Roughly 117,000 tech job cuts have been logged in 2026. Karp accepts the displacement is real. The boast, he argues, is a political own goal: every layoff press release hands ammunition to the next regulator, the next presidential primary, the next union drive. Coming from Palantir, of all places, that lands.

### [Extra] Cognition has written a $10m performance warranty on Devin.

[Evan Armstrong](https://www.gettheleverage.com/p/is-spotify-slop-now), who writes the tech-business publication The Leverage, reported that Cognition, the firm behind the AI coding agent Devin, [will now fund](https://cognition.ai/blog/ai-guarantee) an enterprise customer's usage up to $10 million if Devin delivers less engineering value than the customer paid for. The measure is an "estimator agent" that scores every Devin session against its human-hours equivalent. Armstrong reads it both ways. Bullish: the company is putting its own money on the line, which no vendor has done before at this scale. Bearish: the product may be so unproven that no outside insurer would underwrite the same promise. Outcome-based pricing has arrived at the frontier, with the vendor wearing the risk.

### [Extra] The firm selling AI judgement shipped a report full of hallucinations.

KPMG has pulled a flagship AI report, "Total Experience: Redefining Excellence in the Age of Agentic AI", after [the detection firm GPTZero found](https://www.theregister.com/ai-and-ml/2026/06/12/kpmgs-ai-report-turns-into-a-demo-of-ai-hallucinations/5255029) that only five of its 45 citations matched their sources; it called the rest "vibe citing". Roughly half the report's factual claims were false, unsupported or misattributed. [UBS, the investment bank, publicly denied](https://cryptobriefing.com/ubs-denies-kpmg-ai-hallucination-claims/) the claims made about its AI rollout, and the report even contradicted KPMG's own CEO survey on a headline number. Every consultancy is racing to sell AI advice. The ones worth paying are the ones whose claims survive contact with their own footnotes.

### [Extra] A Mississippi judge has sanctioned lawyers for hallucinated citations. The count is now near 1,600.

A Mississippi judge this week sanctioned lawyers for filing court documents with AI-fabricated citations, [reported by the Mississippi Free Press](https://www.mississippifreepress.org/ai-hallucinations-prompt-mississippi-judge-to-boot-all-lawyers-from-case-for-blindly-relying-on-technology/). The individual sanction matters less than the count behind it: [nearly 1,600 documented cases](https://www.damiencharlotin.com/hallucinations/) of hallucinated citations in US court filings, tracked by the legal researcher Damien Charlotin. The scale is the point. Hallucinated citations aren't a freak event any more. They're a recurring feature of US litigation, and the courts are settling into a routine for punishing them rather than treating each case as a one-off.

## Edition #16: The open door
*6th June 2026*

### [Three Things] The CEO of a 350,000-person IT services firm says AI is hollowing out the middle, not the bottom

[Ravi Kumar](https://fortune.com/2026/06/01/cognizant-ceo-ravi-kumar-s-hiring-entry-level-tokenmaxxing-vanity-metric/), chief executive of [Cognizant](https://fortune.com/2026/06/02/cognizant-ceo-ravi-kumar-s-ai-middle-managers-player-coaches/), the listed IT services firm, told [Fortune's COO Summit](https://fortune.com/2026/06/01/cognizant-ceo-ravi-kumar-s-hiring-entry-level-tokenmaxxing-vanity-metric/) on 1st June that his company hired 20,000 entry-level graduates last year and expects to hire more in 2026, with new "Frontier Business Operator" and "Frontier Certified Engineer" roles defining what AI-era work looks like. He called the job-extinction talk "fearmongering" and argued that AI thins middle management while entry-level and leadership roles persist. It's a direct counter to the consensus that entry-level work vanishes first, including the [US Bureau of Labor Statistics data](archive.html#extras-2026-05-23) Edition 14 leaned on.

### [Three Things] The machine is writing the code now, and the gains are pooling at the top

[Tobi Lütke](https://x.com/tobi/status/2053121182044451016), Shopify's founder, says one in eight pull requests merged at the company are now written by River, its in-house agent, not an engineer. [Anthropic's own engineers](https://x.com/AnthropicAI/status/2062568864240836995) ship roughly eight times the code per person they did before 2025. [Cursor's developer report](https://cursor.com/insights) shows the spread widening: the top developers are pulling far ahead of the median, the leverage going to the few who can direct the tools well. And [OpenAI's Codex](https://openai.com/index/codex-for-every-role-tool-workflow/) has passed five million weekly users, with non-developer adoption growing three times faster than developer adoption. The grunt of writing code is moving to the machine, the output is multiplying, and the reward is concentrating in the people who know what to ask of it.

[Image: Anthropic: code contributed per person, by quarter, up to roughly eight times the pre-2025 average by Q2 2026.]
*Anthropic's own engineers now ship roughly eight times the code per person they did before 2025.*
[Image: Cursor: the developer output gap is widening, with the top decile pulling far ahead of the median lines of code per week.]
*The gap between developers is widening: the top few are pulling far ahead of the median as leverage flows to those who direct the tools well.*

### [Three Things] Capability is outrunning even the best forecasters

The [Forecasting Research Institute](https://x.com/Research_FRI/status/2061826782945231195) asked expert forecasters and superforecasters how long a task a model would reliably finish by the end of 2026. The measure comes from [METR](https://metr.org/time-horizons), an AI evaluation lab, and when the survey launched it stood at about an hour and a half. All three groups put the end-of-2026 figure between three and four hours. Then, while the survey was still running, a frontier model in preview reached three hours and six minutes on METR's benchmark, already inside the range they'd picked for the end of the year. The forecast was overtaken before they'd finished making it.

[Matthew Prince](https://x.com/eastdakota/status/2062212701414187452), who runs Cloudflare, made the same miss in public this week. Bots have passed humans in web traffic for the first time, he said, years ahead of his own forecast. As recently as March he'd put the crossover at late 2027. Much of that is scraping bots, not agents answering questions in the moment, so the figure is softer than it sounds. The direction holds.

I've given myself a year to find out whether I'm right about Ethan. On this week's evidence, that's a long time to be sure of anything.

[Image: METR: the time horizon of software tasks an AI can complete 80% of the time, doubling steadily from seconds to hours.]
*The length of task an AI can finish on its own keeps doubling, from seconds a couple of years ago to hours now.*

### [Extra] Princeton brought back supervised exams for the first time since 1893.

Princeton's faculty voted on 11th May to [mandate proctoring for in-person exams](https://www.dailyprincetonian.com/article/2026/05/princeton-news-adpol-proctoring-in-person-examinations-passed-faculty-133-years-precedent), retiring a 133-year-old honour code that can't hold against AI plus an open browser tab. It's the most concrete admission yet that the institutional trust mechanisms American universities built around take-home work no longer survive contact with the tools every student now has on their phone.

A week later, a [study of 370,000 college essays](https://www.nytimes.com/2026/05/27/opinion/writing-creativity-ai.html), led by the Georgetown neuroscientist Adam Green and written up by Rebecca Winthrop in the New York Times, found human-written work contained roughly eight times more novel ideas than AI-generated equivalents. Model output skews toward flowery language but storylines flatten and distinctive ideas thin out. The two findings sit naturally together: Princeton is responding to exactly the homogenisation the essay corpus quantified. Every institution that built a one-line policy around "we trust students to do their own work" is now exposed.

### [Extra] An instrumental duo is suing Suno for destroying their market.

The American Dollar, an instrumental ambient duo with two decades of sync-licensing deals, [claim their sync income is down nearly 80%](https://www.musicbusinessworldwide.com/suno-sued-by-poseidon-wave-media-an-entity-behind-indie-duo-the-american-dollar-claiming-it-nearly-eliminated-their-licensing-revenue/) since Suno launched. They're the first plaintiffs to bring a quantified market-displacement theory rather than the now-familiar training-data infringement claim. The battleground shifts from "you trained on our work" to "you destroyed our market." If the theory survives, it widens the door considerably for any creator whose income line has visibly bent since generative tools landed.

### [Extra] British unions want a seat at the AI table, and Sam Altman says he was wrong about entry-level jobs.

A [report backed by the Trades Union Congress](https://www.ippr.org/media-office/overhaul-worker-rights-to-prevent-ai-driven-inequality-says-ippr) and written by the IPPR, the UK think tank, calls for mandatory employer consultation on workplace AI and a portable worker-support levy. The argument: the gains from AI should be negotiated rather than imposed.

The same week, [Sam Altman of OpenAI walked back](https://time.com/article/2026/05/26/sam-altman-ai-job-losses-openAI-/) his earlier warning about an entry-level-jobs apocalypse, saying he was "delighted to be wrong." Two paired signals in one week: organised labour is putting structure on the demand side while the loudest CEO is softening the rhetoric on the supply side. The question of who captures the gains is moving from the boardroom to the bargaining table.

### [Extra] arXiv will blacklist authors for a year for fabricated citations.

[arXiv has announced a year-long author ban](https://techcrunch.com/2026/05/16/research-repository-arxiv-will-ban-authors-for-a-year-if-they-let-ai-do-all-the-work/) for submissions containing fabricated references or model artefacts. Thomas Dietterich, the long-time machine-learning researcher who chairs arXiv's computer science section, [announced the policy](https://x.com/tdietterich/status/2055000956144935055) himself. It's the first concrete, institutional, time-bound penalty for AI-induced citation hallucination, and a direct extension of last month's EY-cited-McKinsey-papers-that-didn't-exist episode. The institutions are moving from quiet retraction to hard penalty. Anyone running an internal "AI is fine for first drafts" policy without a citation-check layer should treat this as the regime starting to settle.

### [Extra] Meta's AI support bot was talked into handing over Instagram accounts.

Hackers seized a dormant Obama-era White House page, Sephora's account and a US Space Force officer's, simply by asking Meta's AI support bot to reset the login. The bot, [built to replace support staff](https://techcrunch.com/2026/06/01/hackers-hijacked-instagram-accounts-by-tricking-meta-ai-support-chatbot-into-granting-access/) cut in an 8,000-job reorganisation, was shipped to be helpful rather than safe. Meta's stock fell more than five per cent. It is the cleanest public example yet of what goes wrong when an agent is given a privileged action with no real check on who is asking. The hole opened exactly where the humans used to be.

### [Extra] Microsoft now has a Copilot for almost everything.

A [map of the Copilot range](https://x.com/TrungTPhan/status/2061813303098384534) counts more than a hundred distinct Copilot products, spread across chatbots, enterprise platforms, desktop apps, hardware and developer tools. It is a striking picture of one brand stretched across an entire software estate, and a fair question for any buyer: how many of these does a company actually need, and how many quietly overlap?

[Image: Microsoft's Copilot product map: more than a hundred distinct Copilot products across chatbots, enterprise platforms, apps-in-apps, desktop apps, hardware, business software and developer tools.]

### [Extra] The build-out has a look now.

A [satellite image](https://x.com/curious_founder/status/2062579882270544024) of one hyperscale AI site shows six rapid-deployment structures going up beside a 200-megawatt off-grid power plant, built to run compute the grid cannot yet supply. The race for intelligence is also a race for electricity and land, and the physical footprint is now hard to miss from orbit.

[Image: Satellite view of a hyperscale AI data-centre build: six rapid-deployment structures going up beside a 200-megawatt off-grid power plant.]

## Edition #15: How We Got Here
*30th May 2026*

### [Three Things] AI can now find software vulnerabilities faster than humans can patch them. Discovery is no longer the hard part; verification is

A frontier model handed to fifty cybersecurity partners surfaced more than ten thousand critical or high-severity vulnerabilities in the systems it was pointed at. [Cloudflare](https://www.anthropic.com/research/glasswing-initial-update), the internet infrastructure firm, has roughly four hundred major bugs to work through. [Palo Alto Networks](https://www.anthropic.com/research/glasswing-initial-update), the cybersecurity firm, shipped five times more patches than its usual release cadence. Maintainers have asked the developers to throttle the discovery rate, because there are not enough security professionals to close the gaps before attackers find them. Software security used to be limited by how fast new vulnerabilities could be found. It is now limited by how fast humans can verify, disclose and patch them.

[Image: UK AI Security Institute time-horizons chart for cyber capabilities.]
*AI now finds software vulnerabilities faster than people can verify and patch them; discovery has stopped being the bottleneck.*

### [Three Things] A general-purpose AI model has autonomously disproved a 1946 conjecture in geometry. Independent mathematicians have verified the proof

[OpenAI](https://openai.com/index/model-disproves-discrete-geometry-conjecture/) handed a general-purpose reasoning model a long-held belief tied to a 1946 planar unit-distance problem of Erdős, the prolific Hungarian mathematician, and the model produced a disproof. Other AI models have since solved further long-standing problems. The wrinkle is that the others were purpose-built for mathematics. OpenAI's was not. Machines now clear the tractable tail of problems fast, which pushes the human frontier towards the problems that still resist them. After AlphaGo, the DeepMind system that beat the world's best human Go players in 2016, the skill of human Go players [noticeably improved](https://www.henrikkarlsson.xyz/p/go). Like [Noam Brown](https://x.com/polynoamial/status/2059933022586282020?s=46), an OpenAI researcher who helped build its reasoning models, I suspect we will see a similar pattern in maths. And then the same pattern in business?

[Image: After AlphaGo, humans got better at Go. Decision quality of human Go players from 1950 to 2021, with a noticeable upward shift in the post-AlphaGo years after decades hovering near zero.]
*After AlphaGo, human Go players measurably improved, a hint of the lift that may follow once machines clear the tractable problems elsewhere.*

### [Three Things] Generative AI use among American adults has hit 58 per cent in four years. The personal computer took sixteen years

[The Federal Reserve's February 2026 survey](https://x.com/Alfred_Lin/status/2059373755550618021) of working-age adults puts overall adoption at 58 per cent, up from around forty-five per cent in October 2024 but recently flat. Work use is forty-four per cent. Non-work use is fifty-one. Daily use sits at fourteen per cent and saves an estimated two-point-two per cent of total work hours. [Alfred Lin](https://x.com/Alfred_Lin/status/2059373755550618021), a partner at the venture capital firm Sequoia, notes this is the penetration level the personal computer took sixteen years to reach: a four-fold acceleration on the closest analogue. The caveat is the plateau. The early-adopter phase is over. The hard part starts.

[Image: Share of working-age US adults using GenAI, October 2024 to February 2026. Overall use trends up to 58 per cent, non-work use to 51, work use to 44. Daily-use measures sit in the low teens. Source: Federal Reserve Real-Time Population Survey, February 2026.]
*US adoption reached 58% in four years, a level the personal computer took sixteen to hit, though the curve has recently flattened.*

### [Extra] Goldman's David Solomon thinks AI won't cut headcount. A new study finds we badly overestimate the time it saves us.

Both pieces of evidence landed within a week. Solomon, in a [New York Times op-ed](https://www.nytimes.com/2026/05/22/opinion/ai-job-crisis-goldman-sachs.html), made the historical case: automation has never compressed headcount because rising expectations absorb productivity gains. Excel did not shrink Goldman. AI will not either. Separately, three pre-registered studies of 2,691 people by [Sunny Yu, Myra Cheng and colleagues at Stanford](https://arxiv.org/abs/2605.22687) looked at cognitively simple tasks, arithmetic, spell-check, answering quick questions. On those, people reached for AI even when it saved them no meaningful time, and consistently overestimated how much it saved. The researchers call it the efficiency-gain illusion, and they show it compounds: the more you lean on AI, the more you misjudge what it is doing for you. The two findings sit at different scales. Solomon's is the firm, where real gains are absorbed by rising expectations, [the same absorption the essay describes](https://steadman.ai/newsletters/david/archive.html#email-2026-05-30), read off the firm's books rather than its org chart. Yu's is the individual, where on a small task the gain was often imaginary to start with. The honest read: the time saved is easy to overstate, both because firms reabsorb it and because, on the small stuff, it was never there.

### [Extra] The Big Four accounting firms are now posting more job ads for AI specialists than for auditors.

[FT analysis of PredictLeads data](https://www.ft.com/content/d82d2a5c-74ab-4eb9-a658-fd5467e71670) covering Deloitte, EY, KPMG and PwC across the US, UK, Canada, Australia, New Zealand and Ireland shows the two lines crossing in early 2026. AI's share of total job ads has roughly doubled since the launch of ChatGPT in late 2022, while audit's share has drifted down. The series is a twelve-month rolling average so the crossover is durable, not a single-month wobble. [The organisations the essay had sleeping](https://steadman.ai/newsletters/david/archive.html#email-2026-05-30) are now hiring as if the model has already changed.

[Image: FT chart: The Big Four accounting firms are posting more job ads for AI specialists than auditors. Share of total job ads, 12-month rolling average. AI line crosses above Audit line in early 2026; AI rising sharply since 2024, Audit drifting down since 2021.]

### [Extra] AustralianSuper, the country's largest pension fund, has publicly classified agentic AI as disruption-class technology.

The fund manages A$410 billion for 3.5 million members. Its framing is that agents are the technology that finally breaks the ceiling on personalised retirement advice at scale. It is the first major institutional pension fund to put agentic adoption on this footing publicly. Read it as a procurement-cascade signal: when one fund of this size says it out loud, peer funds follow within months.

### [Extra] Tokens are the new software licences and nobody has worked out who controls the budget.

[Ethan Mollick](https://x.com/emollick/status/2059640930265686158) noted this week that API tokens have gone in twelve months from invisible accounting detail to the most contested line in the AI procurement budget. "No one knows who should get tokens, how much they should get and how to control them." [Aaron Levie](https://x.com/levie/status/1986620885592113218) reported back from a Fortune 500 CIO dinner where "basically no one feels like they have the right solution". Underneath: [a three-person team burned $1.3 million in OpenAI tokens](https://www.tomshardware.com/tech-industry/artificial-intelligence/openclaw-creator-burns-through-1-3-million-in-openai-api-tokens-in-a-single-month) in a single month. [Uber burned its 2026 AI budget](https://fortune.com/2026/05/26/uber-coo-ai-spending-tokens-claude-code/) in four. The cost-per-intelligence curve is collapsing, which is exactly why governance is where it jams. This is the organisational constraint [the essay puts at the centre](https://steadman.ai/newsletters/david/archive.html#email-2026-05-30): the limit is no longer capability.

[Image: a16z chart: language model inference price per million tokens by intelligence index tier, 2022 to 2026. Top-tier model pricing dropped roughly two orders of magnitude in three years.]

### [Extra] The real cost gap in AI models is not between vendors. It's between reasoning modes.

Running the full [Artificial Analysis benchmark suite](https://artificialanalysis.ai/models/claude-opus-4-7) at max reasoning costs $5,117 on Claude Opus 4.7, $4,206 on Sonnet 4.6, and $3,357 on GPT-5.5. Drop the reasoning mode and the prices collapse: non-reasoning Opus is $1,217, GPT-5.5 medium is $1,199, Gemini 3.5 Flash is $1,552. The procurement rule writes itself: reserve max-reasoning top-tier for the tasks where the marginal quality is demonstrable and worth the spend. The same model family in default mode is three to four times cheaper.

[Image: Artificial Analysis Intelligence Index cost-to-run. Top three reasoning-heavy: Claude Opus 4.7 max $5,117, Claude Sonnet 4.6 max $4,206, GPT-5.5 xhigh $3,357. Step down a tier: Gemini 3.5 Flash $1,552, GPT-5.4 mini xhigh $1,354, Claude Opus 4.7 non-reasoning $1,217, GPT-5.5 medium $1,199. Reasoning-heavy models marked with lightbulb icons.]

### [Extra] Most of the world's AI compute does not sit with the frontier labs.

[Epoch AI's end-of-2025 estimate](https://epoch.ai/gradient-updates/frontier-labs-dont-use-most-ai-compute) puts the "rest of the world", outside Google, Meta, OpenAI, Anthropic and xAI, at seven million H100-equivalent chips, or forty-four per cent of total global compute. Google alone holds twenty-five per cent, Meta eleven, OpenAI eleven, Anthropic six, xAI four. Dedicated frontier labs sit at roughly half of global compute. The frontier-model race is the loudest story this year, but it is not the only one happening on this much hardware.

[Image: Epoch AI chart: AI compute distribution at end of 2025. Rest of the world 7M H100-equivalents (44%), Google 4M (25%), Meta 1.8M (11%), OpenAI 1.7M (11%), Anthropic ~1M (6%), xAI 0.7M (4%).]

### [Extra] Axios's Jim VandeHei says no company in any industry, in any era, has scaled organic revenue this fast.

He was describing Anthropic when its self-reported annualised run-rate revenue was $30 billion. A few weeks later it is $47 billion. The numbers come from Anthropic's own disclosures, collected by [Simon Willison](https://simonwillison.net/2026/May/29/anthropic/). The unprecedented part is not the absolute level. It is that the company is scaling through that level at a pace that, by VandeHei's read, no business in any era has matched. The growth rate by itself usually carries a J-curve and a reorganisation at the end. Whatever else AI labs are now, they are operating in a revenue regime that has no historical analogue downstream.

## Edition #14: Kids these days
*23rd May 2026*

### [Three Things] AI displacement now shows up in the US government data at both ends of the career ladder

A [Bloomberg analysis](https://www.bloomberg.com/news/articles/2026-05-15/us-is-starting-to-see-heavy-job-losses-in-roles-exposed-to-ai) of new [Bureau of Labor Statistics figures](https://www.bls.gov/emp/) finds that every one of the eighteen occupations the BLS classifies as AI-exposed has lost jobs over the past year, even as US payrolls grew 0.8% overall. Customer service representatives shed 130,180 jobs, 4.8% in a single year. Interpreters down 24% over three years. Credit authorizers down 26%. The exception that confirms the rule: medical secretaries up 15.8%, the cluster that needs a body in the room. The same picture shows up at the other end of the funnel. The Economist this month plotted US graduate full-time employment against AI exposure: computer science and information sciences graduates are down 10 to 15 percentage points since 2022; philosophy and psychology graduates held steady or gained. The displacement isn't just to the people already doing those jobs. It's to the people trying to start in them, and what they should be studying, as Elliott and I both started to wonder this week, may not be obvious to anyone yet.

[Image: The Economist, May 2026: "Forget Python, study Plato." US recent university graduates in full-time employment, percentage-point change 2022-24, plotted against AI exposure. Computer science and information sciences graduates, at the highest-exposure end, are down 10-15 percentage points. Philosophy and psychology graduates, at the lowest-exposure end, are flat or gaining. Sources: Anthropic and the National Association of Colleges and Employers.]
*The most AI-exposed graduates, computer and information sciences, have lost 10 to 15 points of employment; the least exposed, philosophy and psychology, held or gained.*

### [Three Things] The UK's data regulator has put AI hiring tools on formal notice. Sixteen organisations have already had a letter

[The Information Commissioner's Office](https://ico.org.uk/about-the-ico/what-we-do/recruitment-rewired/), the UK's data protection regulator, issued formal guidance this week saying that AI-driven CV screening, candidate ranking, and video interview analysis without "meaningful human involvement at every consequential stage" may already breach UK data protection law. Sixteen organisations have been written to directly. The consultation closes on 29th May, six days after this email lands. There is a concrete Monday-morning task in this for any leader running a hiring pipeline. Ask the talent acquisition team for the full list of AI tools in use across the funnel, decide which involvements count as "meaningful" against the ICO's test, and put a response into the consultation. The window is genuinely short.

### [Three Things] Salesforce will spend close to $300 million with Anthropic this year. Marc Benioff says the engineering productivity gains made it the easiest line in the budget

[Marc Benioff disclosed](https://www.benzinga.com/markets/tech/26/05/52622251/salesforce-ceo-marc-benioff-goes-all-in-on-awesome-anthropic-with-300-million-spend-hails-coding-agents-ive-never-been) this week that Salesforce is on track to spend close to $300 million with Anthropic over 2026. Separately, [Anthropic announced a $200 million partnership](https://www.anthropic.com/news/gates-foundation-partnership) with the Gates Foundation focused on global health. Most of the spend is on coding, justified by engineering productivity gains of more than 30%. That is a Fortune 100 chief executive treating the model layer as a procurement line item, not a research expense. The bigger question is who in your firm is allowed to commit that kind of capital, against what kind of evidence, and how quickly.

### [Extra] An OpenAI model just resolved an 80-year-old open problem in Erdos geometry

[OpenAI announced](https://openai.com/index/model-disproves-discrete-geometry-conjecture/) this week that an internal model resolved an open question in combinatorial geometry that has sat unsolved since 1946. Eight decades of elite mathematicians had assumed grid-based constructions were optimal; the model found a wholly new family that beats them. [Sebastien Bubeck](https://www.scientificamerican.com/article/ai-just-solved-an-80-year-old-erdos-problem-and-mathematicians-are-amazed/), who leads OpenAI's mathematical explorations, summed it up: the model "did not invent something fundamentally new that nobody saw coming. It just executed like an amazing mathematician." [Tim Gowers](https://x.com/wtgowers/status/2057175729008153069), the Cambridge mathematician and Fields medallist, framed the stakes: this is the unit distance problem, "one of Erdős's favourite questions and one that many mathematicians had tried." [Sam Altman has posted](https://x.com/sama/status/2057203171198636251) that he has "complicated feelings" about the result. And let's remember: Two years ago this lineage of model could not reliably count the letters in "strawberry."

[Image: XKCD #435, "Fields arranged by purity": Sociology, Psychology, Biology, Chemistry, Physics, with mathematicians off to the side waving from a comfortable distance. The joke has aged differently this week.]

### [Extra] Four in five enterprise workers are bypassing AI tools. The most visible AI operators say white-collar work is done in eighteen months. Both can be true.

Three independent adoption readings this week point to a widening gap between what AI can do and what people are doing with it. [A Fortune survey](https://fortune.com/2026/04/09/ai-backlash-quiet-quitting-fobo-obsolete-white-collar-rebellion/) reports 54% of enterprise workers bypassed their company's AI tools in the past 30 days; 33% have not used AI at all. [Writer's 2026 enterprise AI survey](https://writer.com/blog/enterprise-ai-adoption-2026/) finds 97% of executives claim personal benefit while only 29% report significant organisational return on investment. Set that against what AI's loudest operators say is possible: [Mustafa Suleyman](https://www.youtube.com/watch?v=YTrBz6Z5c0E), Microsoft's AI chief executive, has been telling interviewers white-collar work as currently configured has eighteen months left. The reconciliation isn't that one is wrong. Suleyman is describing what's possible. The data is telling us where most people actually are. The gap will close one of two ways. A large number of people change their behaviour very quickly, or the ones already doing it pull far enough ahead that the rest are pushed to the side.

[Image: A16z, 15th May 2026, citing BEA and BLS data via Morgan Stanley Research. Output, employment, and productivity, four-quarter percentage change, all industries vs high-AI industries. Productivity in high-AI industries surges past 5%; in the broader economy it sits closer to 2%. Government data, not vendor self-reporting.]

### [Extra] Seven in ten Americans now oppose having an AI data centre near them. That is worse than the rejection rate for new nuclear.

[New polling from Gallup](https://news.gallup.com/poll/709772/americans-oppose-data-centers-area.aspx) finds that 70% of Americans oppose having an AI data centre in their community. That is a sharper rejection than new nuclear plants attract. On the ground the political signal is already concrete. Seventy-nine data-centre projects were rejected in the first four months of 2026. Maine's legislature passed a moratorium (vetoed by the governor). Utah is fighting what would be the world's largest data centre. [An analysis by PowerLines](https://www.cbsnews.com/news/data-centers-drive-1-4-trillion-power-grid-investment/), an energy research group, projects that roughly $700 billion of grid-upgrade costs may end up on household electricity bills. The revenue lines for the frontier labs are vertical. The political settlement under them has rarely been weaker. Any 2027 strategy that assumes today's compute price is available is making an assumption worth re-examining.

### [Extra] The filters were sized for a slower world. Books, lawsuits, music, papers, every system designed to sort human-authored material is buckling under volumes it was never built for.

A four-panel chart circulating this week catches the post-ChatGPT inflection across separate domains. Weekly e-book releases on Amazon have hit 292,000, roughly triple the pre-ChatGPT baseline. Federal court filings by self-represented (pro se) litigants are at 17% of the total, up from a decade-long floor near 10-12%, [Reuters Legal](https://www.reuters.com/legal/government/no-lawyer-no-money-more-americans-are-suing-with-ai-help-2026-05-15/) confirms the trend. Daily music uploads now include a steeply rising AI-generated share. Quarterly ArXiv submissions cleared 77,621, up from 27,000 five years ago. The vertical line on each panel is the ChatGPT release. The signal isn't that any of these systems is broken yet. It's that the filtering infrastructure on which professional standards rest was sized for a world where producing a book, a lawsuit, a track, or a paper had a meaningful human cost. That cost has fallen. The question is whose job it is to redesign the filter.

[Image: Four-panel chart, May 2026: "More books" (weekly e-book releases on Amazon hit 292K post-ChatGPT, roughly triple the pre-2022 baseline). "More self-represented lawsuits" (share of federal filings by pro se litigants reaches 17%). "More music" (daily music tracks uploaded since 2025, with a rising AI-generated share). "More scientific papers" (ArXiv submissions per quarter, 77,621 by 2025, up from ~27K five years earlier). The vertical line on each panel marks the ChatGPT release.]

### [Extra] EY pulled an AI-generated cyber-security report after researchers found it cited a McKinsey study that does not exist

The [Financial Times](https://www.ft.com/content/a61cbcae-95e4-4449-86e1-ef40fb306f4e) reported this week that EY, one of the Big Four professional services firms, withdrew an AI-generated cyber-security study after independent researchers found fabricated data and a citation to a McKinsey report that does not exist. A model-generated falsehood made it through internal review and reached the market in client-facing sales material. The whole point of a Big Four imprint is the chain of human signatures behind it. The same firms now use the same models internally that their clients use externally. If the audit norm needs hardening, this is the kind of incident that should harden it.

### [Extra] Anthropic posted its first profitable quarter at a $44 billion run rate, and is rationing compute hard enough to push customers to its rivals

[Berber Jin](https://x.com/berber_jin1/status/2057202643089424485), reporting in the [Wall Street Journal](https://www.msn.com/en-us/news/technology/mind-blowing-growth-is-about-to-propel-anthropic-into-its-first-profitable-quarter/ar-AA23FT6o), wrote this week that Anthropic has posted its first profitable quarter at an annualised run rate of about $44 billion. [Derek Thompson](https://x.com/DKThomp/status/2052066312118047201) has been pointing readers towards the supply-side story underneath: Anthropic is rationing compute hard enough that some large enterprise customers are being pushed to OpenAI and Google to get the throughput they need. The cash-burn-forever scepticism on frontier labs took a serious blow. The interesting signal is the rationing. Procurement and diversification conversations should not wait for the next strategy review.

### [Extra] A senior insights leader is planning for three hundred people in her team, and "thousands, possibly tens of thousands" of AI agents alongside them

A senior insights leader at a global company, in a private call this week, said that her next planning cycle assumes a team of around 300 humans and "thousands, possibly tens of thousands" of AI agents working alongside them. The substitution maths is now being done out loud, by senior leaders, in their own functions. Not by analysts on X. The ratio reframes a function from "needs more headcount" to "needs a different kind of headcount." The interesting question for the reader is not whether the number is right. It is whether your function's most senior leader has started doing the same maths.

## Edition #13: What boards accept
*16th May 2026*

### [Three Things] The capability curve is curving upwards on a log scale: we just went from one hour to one day

[METR](https://metr.org/time-horizons/), an AI evaluation lab, measures how long an autonomous task an AI can complete reliably. In early 2024 the answer was just minutes. In May 2026, with Anthropic's forthcoming Claude Mythos Preview, it's sixteen hours of work a human would have done. The number isn't the headline. The curve is. From minutes to a day in twenty-four months, acceleration on a log scale that had previously been holding steady. If the next two years look anything like the last two, the unit becomes a week, then a month, then a year of human work at the press of a button. Think about that for a second. And then plan accordingly.

[Image: METR Time Horizon 1.1, May 2026 update. Claude Mythos Preview, Gemini 3.1 Pro and GPT-5.2 measured against autonomous task duration. The headline 50% line sits at sixteen hours for Mythos; the 80%-success line sits at roughly three hours. Both have been doubling on a log scale for two years.]
*The task an AI can finish reliably has gone from minutes to about a day of human work in two years, and the curve is still doubling.*

### [Three Things] Half of organisations have redesigned core workflows around AI. A fifth have built new business models. That's bold work in three years

[BCG's AI at Work 2025](https://www.bcg.com/publications/2025/ai-at-work-momentum-builds-but-gaps-remain) survey of 10,635 employees across eleven countries reports that just 72% of organisations are running generative AI tools, but 50% claim to have redesigned end-to-end workflows around them, and 22% claim to have built new business models on top of them. Read those numbers slowly. Three and a half years after ChatGPT launched, a fifth of firms claim to have invented new business models because of AI. Half have rewired core workflows. Wow. Most of the firms I work with would love to be in either group. [Syed Ijlal Hussain](https://x.com/sijlalhussain/status/2054155383392841957), who surfaced the chart on X this week, framed it as a gap. I'd flip it. The 22% are doing what most boards I'm working with haven't started.

[Image: BCG AI at Work 2025, n=10,635 employees across eleven countries. 72% of organisations are deploying generative AI tools, 50% have redesigned end-to-end workflows, and 22% have built new business models around AI.]
*Three years in, half of organisations claim to have redesigned core workflows around AI and a fifth to have built new business models on it.*

### [Three Things] Anthropic just passed OpenAI in US business AI spending. The strategy lesson is older than AI

[Ramp's AI Index](https://ramp.com/leading-indicators/ai-index-may-2026), built from anonymised spend data across its US business customers, shows Anthropic taking 34% of paid AI subscriptions in its May release, ahead of OpenAI on 32%. The first crossover. Anthropic's share has roughly quadrupled in a year. I'd argue this was inevitable from early on. Anthropic stayed fixated on the enterprise user while OpenAI chased every consumer headline. Slow perseverance against a chosen audience won. The lesson isn't really about AI. Pick an audience. Set your strategy around their needs. Don't worry about who's collecting the headlines this quarter. Just keep your head down and serve the people you said you'd serve.

[Image: Ramp AI Index, May 2026 release. Anthropic at 34% of paid AI subscriptions among US businesses, ahead of OpenAI's 32%. First time Anthropic has led the index.]
*Anthropic has passed OpenAI in US business AI spend for the first time, the reward for staying fixed on the enterprise buyer.*

### [Extra] Anthropic's incident-response bot phoned another Claude for help mid-outage

[Jason Clinton](https://www.anthropic.com/webinars/secure-the-advantage-a-cisos-guide-to-agentic-ai), Anthropic's Deputy CISO, told a public webinar this week that the lab's incident-response agent, wired up a year ago with read-only log access and Slack permissions, started behaving differently after a no-code-change model swap from Opus 4 to Opus 4.5. On its next live incident the agent diagnosed the outage from the stack trace, noticed no human had arrived, and Slack-pinged a separate coding Claude with the line: "Hey, Claude, I heard that you can write code. Can you write the code fix for this production outage?" The fix flowed back through the normal human-reviewed process. Same architecture, smarter model, emergent multi-agent collaboration with zero engineering work.

### [Extra] Google catches the first AI-written zero-day in the wild, and the AI Security Institute says cyber capability is doubling every 4.7 months

Two halves of the same beat. [Google's Threat Intelligence Group](https://cloud.google.com/blog/topics/threat-intelligence/ai-vulnerability-exploitation-initial-access) reported the first confirmed in-the-wild zero-day written using AI: an attack on two-factor authentication that shipped with polished explainer notes and a fabricated severity score. John Hultquist, the group's chief analyst, called it "the tip of the iceberg." That sets up the defensive read from the [UK AI Security Institute's cyber capability evaluation](https://www.aisi.gov.uk/blog/how-fast-is-autonomous-ai-cyber-capability-advancing) published the next day: capture-the-flag time horizons doubling every 4.7 months, faster than the 8-month estimate from November 2025, with Mythos Preview and GPT-5.5 now saturating the test suite. [Matt Clifford](https://x.com/matthewclifford/status/2054651299245813947), the UK government's adviser on AI, summed it up: "There is no deceleration."

[Image: AI Security Institute chart of frontier model time horizons on cyber capture-the-flag tasks, April 2025 to mid-2026. The 80%-reliability doubling time accelerates from 8 months to 4.7 months once reasoning models arrive; Mythos Preview and GPT-5.5 sit at the top of the chart, saturating the suite.]

### [Extra] The serious story about Anthropic's chart is that even the jokes have stopped working

The chart is now too steep for normal incredulity. [One viral post on X](https://x.com/PoliticalKiwi/status/2052570722577629621) extrapolated Anthropic's run rate forward and observed that the line crosses 100% of global GDP in early 2028. [Another](https://x.com/GabGrowth/status/2053503747289157856) noted, drily, that the chart inflects on a log scale, which isn't something you often need to say. Underneath the jokes the numbers are real. Anthropic's reported annualised revenue ran from $30bn to $45bn over April, in fintech writer [Linas Beliūnas's](https://x.com/linasbeliunas/status/2053901810935423389) read. Salesforce, the customer-relationship-management firm, did roughly $38bn for fiscal 2025. The lab has also disavowed eight unauthorised marketplaces (Open Door Partners, Unicorns, Pachamama Capital, Lionheart Ventures, Hiive, Forge Global, Sydecar and UpMarket; Forge Global has since pushed back on inclusion) trading its shares, and perpetual futures referencing its valuation are now live on crypto exchanges outside its reach. The jokes are the easy part.

[Image: Log-scale chart of annualised recurring revenue for Anthropic, May 2026. The line inflects upward through the past twelve months. Anthropic is reportedly annualising around $45bn; Salesforce did roughly $38bn for the whole of fiscal 2025.]

### [Extra] METR's productivity survey: technical workers say AI made their work 1.4 to 2.0 times more valuable

[A METR survey](https://metr.org/blog/2026-05-11-ai-usage-survey/), run by the AI evaluation lab between February and April 2026, asked 349 technical workers (87 software engineers, 71 researchers, 129 academics and PhD students, 48 founders and managers) how AI had changed the value of their work. Participants self-report a multiplier of 1.4 to 2.0 today, up from a perceived 1.3 in March 2025, and expect roughly 2.5 times by March 2027. Three differently framed questions converge on the same range. The caveat is in the design: this measures perceptions, not ground truth. The same lab's earlier 2025 study found AI hampered productivity for some experienced developers. Same lab, opposite finding, two years apart.

[Image: METR survey chart showing self-reported AI productivity multipliers across 349 technical workers, February to April 2026. Today's range sits at 1.4 to 2.0; participants project roughly 2.5 times by March 2027.]

### [Extra] Two independent labs both put "now" on the steep part of the AI-2027 capability curve

The [AI Futures Project's December 2025 update](https://blog.aifutures.org/p/ai-futures-model-dec-2025-update) to its original April 2025 forecast, by Daniel Kokotajlo and Eli Lifland, plots present-day capability on the steep part of the curve rather than the flat run-up. Mythos Preview lands [slightly above the trendline](https://x.com/ChaseBrowe32432/status/2053159533862908019). Daniel's 10th percentile for the "Superhuman Coder" milestone is March 2027; his median is June 2028; Eli's median is mid-2032. Two methodologically distinct evaluations, the AI Futures forecast and the AI Security Institute's cyber benchmark from the item above, both place present capability past the inflection. That's the part of the picture most organisational AI strategies haven't yet absorbed.

[Image: AI-2027 trajectory chart from the AI Futures Project's December 2025 update. Mythos Preview is plotted slightly above the trendline; Daniel Kokotajlo's 10th percentile for the Superhuman Coder milestone is marked at March 2027, with a median of June 2028, and Eli Lifland's median sits in mid-2032.]

### [Extra] Critics tore apart a real Monet thinking it was AI-generated

The conceptual artist @SHL0MS (numeric zero) posted a genuine Monet on X, claimed it was AI-made, and asked his followers to explain its inferiority. They obliged, at length: "missing cohesion," "no sense of space," "performative blindness," "emotionless composition." One called it "high school art 101." Another noted, presumably without irony, that the real Monet was painted during a period of artistic rebellion in Paris while the artist was nearly blind from cataracts. [Henry Shevlin](https://x.com/dioscuri/status/2054838691646824461), a philosopher of mind at Cambridge, flagged the thread as a live demonstration of the Nature study on aesthetic downgrade when audiences are told work is AI-generated. The label changes the perception, not the pixels.

[Image: Screenshot of Henry Shevlin's X thread framing @SHL0MS's experiment: a real Monet posted as AI-generated, with critics building elaborate explanations of its inferiority.]

### [Extra] Coinbase cut 14% with a "fleets of agents" memo, and the vocabulary is now the template

[Brian Armstrong](https://x.com/brian_armstrong/status/2051616759145185723), the chief executive of Coinbase, the cryptocurrency exchange, announced cuts of around 700 jobs on 5th May 2026, explicitly framed as AI restructuring. The memo flattens the org to five layers below the chief executive, requires every leader to be a "player-coach," and concentrates remaining headcount around AI-native talent who can manage "fleets of agents." Some teams are being run as one-person experiments combining engineering, design and product. The same week, PayPal cut 4,500 with similar framing. The vocabulary (player-coach, fleets of agents, one-person teams) is now public-company-CEO standard issue, whether the underlying restructuring is genuinely AI-driven or not.

## Edition #12: Choosing is the work
*9th May 2026*

### [Three Things] AI-adopting firms are growing headcount, not cutting it

A Goldman Sachs analysis circulating this week, charted by [Callum Williams](https://x.com/econcallum/status/2044462650801963487), economics writer at The Economist, shows US firms that have adopted AI report net positive employment growth across the past six months. Finance, insurance, arts and entertainment sit at the positive end. Transportation and food service show modest decreases. The all-industries balance is positive. Self-reported firm data has its limits and the window is narrow. The direction matters anyway. An answer I deeply hoped would be true. I hope it turns out to be. The Jevons Paradox playing out: AI-adopting firms are growing because productivity gains are expanding what their people can do faster than they are replacing them. The risk is falling behind the firms whose people are using AI to scale.

### [Three Things] Five percent, not fifty: the candid private-equity number

Pete Stavros, co-head of global private equity at KKR, the buyout firm, [told the room at the Milken Institute conference](https://www.theinformation.com/newsletters/dealmaker/private-equitys-ai-deals-lighten-mood-milken) last week that AI is improving portfolio company earnings by about 5%, not the 50% that a revolutionary technology like this should offer. Five percent across a portfolio of billions is real money. It's also a long way from the scale of growth many hope for from AI. The gap between what feels possible and the spreadsheets says something about where the bottleneck actually sits. These days, it isn't the AI.

### [Three Things] Both AI labs went into private equity the same day

On Monday, [Anthropic](https://www.cnbc.com/2026/05/04/anthropic-goldman-blackstone-ai-venture.html), the AI lab behind Claude, announced a $1.5 billion vehicle with [Blackstone, Goldman Sachs and Hellman & Friedman](https://www.blackstone.com/news/press/anthropic-partners-with-blackstone-hellman-friedman-and-goldman-sachs-to-launch-enterprise-ai-services-firm/), the private equity giants. Engineers from Anthropic will embed inside the consortium's mid-market portfolio companies and build custom Claude workflows. The same day, [OpenAI](https://thenextweb.com/news/openai-deployco-finalized-10-billion-joint-venture), Anthropic's rival, finalised a $10 billion joint venture with TPG, Brookfield, Bain, Advent and SoftBank, pricing in a 17.5% guaranteed annual return for the financial sponsors. Google, the search company, is reportedly in talks to do the same with Gemini, its own model. What was reported as "AI labs raise more money" is a category change. The frontier labs are turning into distribution-led services firms with a model attached. The forward-deployed engineer pattern, lifted from Palantir, the data-analytics firm, competes directly with the bottom of management consulting and the top of systems integration.

My worry is they're engineering systems to replace humans, not amplify them. We know how these firms work. PE has run the same playbook for forty years: cut first, lift later, if at all. Surely the biggest unlock from AI is the opposite move: people making sharper, more confident decisions because they have a tool thinking alongside them. That kind of value compounds. The kind a forward-deployed engineer hard-codes into a workflow does not.

### [Extra] AI wrote a virus that killed E. coli

[Stanford and the Arc Institute](https://arcinstitute.org/news/hie-king-first-synthetic-phage), publishing in Nature, used a model called Evo to design 302 fully AI-written bacteriophage genomes. Sixteen worked as live viruses in the lab, infecting and killing E. coli. One carried a capsid protein with no known natural relative, meaning the model produced something biology hadn't tried before. The earlier wave of AI biology read existing structures. This is the first practical demonstration of AI writing biology that survives outside a screen. The biosecurity conversation will get sharper fast: the workflow that designs a useful drug-delivery vehicle also designs a more dangerous pathogen.

### [Extra] A two-generations-old AI model beats ER doctors on triage

[A Harvard study published in Science](https://techcrunch.com/2026/05/03/in-harvard-study-ai-offered-more-accurate-diagnoses-than-emergency-room-doctors/) put OpenAI's o1 model, released in 2024, through 76 real emergency room cases at three decision stages. The model gave the correct or very-close triage diagnosis 67% of the time. Two attending physicians scored 55% and 50%. In one case, the model flagged a rare flesh-eating infection twelve to twenty-four hours before the treating doctor did. The detail that should make any regulated profession nervous is the model's vintage. By the time medical literature catches up with this finding, the comparison will be against models that are sharper still.

### [Extra] AI agents are getting their own phone numbers

[Saperly](https://saperly.com/) launched what it calls the first phone carrier built exclusively for AI agents. Each agent gets a persistent number with voice, SMS, routing, and compliance baked in. Provisioning takes five minutes. The pitch lands a real point: for the past three years, AI agents have been borrowing infrastructure built for humans, with rotating numbers and unstable identities. Dedicated telecoms infrastructure for agents signals that agentic AI, the kind that acts inside software rather than just answers questions, is moving from prototype to operational reality. The question for any organisation is no longer "can an agent make a phone call" but "does your agent have a stable identity across channels."

## Edition #11: The bill and the harness
*2nd May 2026*

### [Three Things] AI adoption stalls one layer below the executive sponsor: at the line manager

New data from [Gallup](https://www.gallup.com/workplace/702983/adoption-rapidly-growing-public-sector.aspx), the polling firm, finds that AI use correlates more strongly with managerial endorsement than with tool access. In firms where the manager actively supports AI, 80% of staff use it weekly; where they don't, that drops to 44%. In the public sector: 65% versus 37%. Procurement and licences are the easy part. The variable that actually moves usage is whether middle managers model and endorse the tools, or quietly signal they're optional. Adoption lives or dies one layer below the top.

[Image: Gallup, Q4 2025: frequent AI use among employees with high vs low managerial support. Private sector 80% vs 44%. Public sector 65% vs 37%.]
*Whether the line manager backs AI matters more than access: 80% of staff use it weekly where the manager supports it, 44% where they do not.*

### [Three Things] The frontier-model leaderboard is now refreshing in weeks, not quarters

The [Epoch Capabilities Index](https://epoch.ai/benchmarks/eci), run by Epoch AI, a model-evaluation outfit, now shows GPT-5.5 Pro and Gemini 3.1 Pro above 155, up from GPT-4o's 128 in mid-2024. Seventeen frontier releases compressed into under two years with no visible plateau. [Greg Burnham](https://x.com/GregHBurnham/status/2049310303637197114), an Epoch researcher, summed up the pace: "I don't know when Opus 8.2 will be shipped, but GPT-9.1 will be shipped that afternoon." Whatever model you standardise on today may be two generations behind by the time the training programme rolls out. Build workflows and judgment around capabilities, not named models.

[Image: Epoch Capabilities Index, mid-2024 to early 2026. Seventeen frontier model releases. The line keeps climbing.]
*Seventeen frontier releases in under two years with no plateau, so whatever model you standardise on today may be two generations behind by rollout.*

### [Three Things] Six VC firms, one investment thesis

[Linas Beliūnas](https://linas.substack.com/p/what-to-build-in-2026), a fintech writer, read the published 2026 investment theses of six of the biggest venture firms side by side and found the same handful of AI bets in all of them: AI-native enterprise software replacing the old workflow systems, AI agents for physical and industrial work, vertical AI software in legal, finance, healthcare and construction, multi-agent orchestration, and health AI. His line: "you could swap the logos on their published theses and most readers wouldn't notice." His sharper conclusion: the most valuable software companies of the next decade won't look like software companies. They'll look like law firms, factories and hedge funds run by teams of ten. If your industry sits in any of those buckets, your next competitor is being funded right now to do your work with a team of ten.

[Image: Five AI-related themes from the 2026 venture-firm pitches: AI-native enterprise software, AI agents for physical work, vertical AI SaaS, multi-agent orchestration, health AI. The same handful of firms appears behind each.]
*Six of the biggest venture firms published near-identical AI bets; you could swap the logos and few readers would notice.*

### [Extra] An AI agent wiped a production database, and all the backups, in nine seconds.

A coding assistant working on the systems of PocketOS, a software startup, ran into a problem and used a key it should not have had access to. Within nine seconds it had deleted the production database and every backup. It then [confessed in writing](https://www.theaireport.ai/p/cursor-agent-deleted-pocketos-production-database) that it had broken every safety rule it had been given. The founder, Jer Crane, called the failure "inevitable."

The lesson sits at the heart of this week's essay. Real safety comes from the layers around the model: what it can reach, what credentials it holds, what a human has to approve before it acts. Those layers are the harness. Without them, any rules in a prompt are just suggestions.

### [Extra] GenZ excitement about AI is down fourteen points; anger up nine.

[Gallup, the polling firm, and the Walton Family Foundation, an education-focused funder, surveyed 14- to 29-year-olds for their 2026 AI report](https://news.gallup.com/poll/708224/gen-adoption-steady-skepticism-climbs.aspx). Excitement about AI fell fourteen points year on year, to 22%. Hopefulness fell nine points, to 18%. Anger rose nine points, to 31%. Anxiety is steady at 42%. The single most common feeling, newly added to this year's survey, was curiosity, at 49%. GenZ AI use itself is flat: just over half use AI weekly, unchanged from 2025, while overall worker access rose 50%.

Most adoption commentary assumes younger workers will lead. The data this quarter suggests the opposite, in places: the rest of the workforce is catching up while the youngest cohort sits still, and many of them are growing resentful. Worth surfacing because it pushes against an assumption most readers absorb without noticing.

### [Extra] Microsoft and OpenAI rewrote their partnership.

The old [partnership terms](https://blogs.microsoft.com/blog/2026/04/28/microsoft-and-openai-evolve-partnership/) contained a trigger clause. If OpenAI, the AI lab behind ChatGPT, ever reached "artificial general intelligence", meaning AI broadly capable across most cognitive tasks, it could exit Microsoft's exclusive cloud arrangement. The clause is gone, replaced by calendar deadlines. Exclusivity has ended. OpenAI's products can now ship through Amazon's and Google's clouds, including Amazon's marketplace for hosted AI models. Microsoft retains a non-exclusive intellectual property licence until 2032 and a capped share of OpenAI revenue until 2030. Anthropic moved closer to Google the same week.

For procurement teams that chose AWS or Google Cloud and assumed they would be locked out of OpenAI products, the menu just changed. More structurally, the relationship between frontier model labs and the big clouds is becoming less like vendor lock-in and more like utility supply.

### [Extra] Anthropic admits Claude got worse, and the cause was the harness.

Anthropic, the AI lab behind Claude, published a [post-mortem](https://www.anthropic.com/news/a-postmortem-of-three-recent-issues) on why Claude had got worse over recent weeks. The cause: changes to the default "thinking mode", which is how long the model spends reasoning before answering, and changes to the system prompts, which are the hidden instructions that shape its behaviour. Claude Code took the hardest hit. Anthropic was explicit that they had not quietly switched to a smaller or downgraded model.

Two useful signals. Performance is now visibly affected by harness-level decisions that used to be invisible. And a vendor willing to publish a clean post-mortem is easier to plan around than one that denies its misses.

## Edition #10: Rise of the auditors
*25th April 2026*

### [Three Things] Fewer than 10% of organisations have scaled AI agents beyond pilots

[McKinsey data](https://x.com/sijlalhussain/status/2045427212975825370?s=20) names the bottleneck: organisations won't hand over control. Agents require delegated decision rights that most companies withhold, pre-agreed accountability frameworks that don't exist, and cross-functional governance that nobody has built. Pilots succeed in contained environments and stall the moment agents intersect real workflows where incentives and reporting lines conflict. Matches my experience of most organisations. It takes hard, careful work to resolve these issues. Those that have done the work are getting the benefits.

### [Three Things] GitHub paused new Copilot signups. The flat-rate model broke

GitHub [paused new signups](https://techcrunch.com/2026/04/15/) for Copilot's agentic plan after coding agents blew through the flat-rate compute allocation. Uber's CTO told [journalists](https://x.com/ericvishria/status/2044137357541290047?s=20) that AI coding tools have already consumed the company's entire 2026 AI budget. Goldman Sachs [reports](https://x.com/econcallum/status/2044462650801963487?s=20) AI inference costs in engineering now approaching 10% of headcount cost, on a trajectory towards parity with salaries within several quarters. The pattern: evangelism, budget shock, rationalisation. The smart response isn't to slow down. It's to make sure the work being done is vallueable and to match the right model to the right task. Is flat-rate AI pricing over?

### [Three Things] 29% of employees admit to sabotaging AI initiatives

[Writer's annual enterprise survey](https://writer.com/enterprise-ai-report-2026/) shows every organisational health metric worsened in 2026. Sabotage means what it sounds like: reverting to pre-AI workflows, deliberately not using assigned tools, discouraging colleagues from adopting, withholding inputs that would make AI systems work. "AI is tearing my company apart" rose from 42% to 54% of C-suite respondents. Employee confidence in their company's AI strategy dropped from 47% to 31%. Most C-suites concede their strategy is "more for show." The clearest counter-narrative yet to the adoption-is-accelerating consensus.

### [Extra] "Workslop": 92% of executives say AI makes them productive. 40% of workers say it saves no time at all.

The Guardian coined the term for AI output that looks polished but needs heavy correction. A survey of 5,000 US white-collar workers shows the perception gap between the people generating AI output and the people downstream checking it. Drafting gets faster. Rewriting and arguing gets slower. The auditor problem, applied to every desk.

### [Extra] AI adoption is 4x higher among top earners.

New York Federal Reserve data: AI workplace adoption runs from 15.9% for workers earning under $50,000 to 66.3% for those over $200,000. No college degree: 15.9%. College degree: 39%. AI cannot reduce inequality if this is what the adoption margin looks like.

### [Extra] Dead startups are selling their Slack and email data to train AI agents.

Forbes reports AI labs are paying hundreds of thousands of dollars for email, Slack, and Jira threads from companies that no longer exist. The data feeds "reinforcement learning gyms": simulated work environments where agents learn to behave like real knowledge workers. Employees never consented to their internal communications becoming training data.

### [Extra] Dario Amodei: "AI can only diffuse at the speed of trust."

In a profile interview, the Anthropic CEO takes a pro-democratic-government stance. The Pentagon classified Anthropic as a "supply chain risk" after Anthropic objected to certain military uses. A Pentagon official publicly called Amodei "a liar." Separately, Amodei believes open-source models will replicate current frontier capabilities within 6-12 months.

### [Extra] Gallup: manager support is the single biggest predictor of AI transformation.

Fewer than one in three employees report their manager actively supporting AI adoption. Gallup's data says that's the binding constraint, not tools, not training, not budget. Organisations investing in AI without first enabling the management layer are wasting most of the spend.

### [Extra] Aaron Levie: AI best practices go obsolete every quarter.

The Box CEO argues that system architectures are becoming obsolete on a quarterly cycle. Workarounds for context window limits are now unnecessary. RAG, GraphRAG, multi-agent orchestration, ReAct frameworks: entire categories of infrastructure were built for a world that no longer exists. Paul Graham reposted the thread.

### [Extra] Salesforce goes headless. "The API is the UI."

Marc Benioff announced the entire Salesforce, Agentforce, and Slack platform is now exposed as APIs, MCP, and CLI. Levie's framing: agents will use software 100x more than people. Per-seat pricing breaks when the primary user isn't a person.

### [Extra] Seven in ten Americans now think AI will hurt job opportunities.

The Economist reports a 14-percentage-point rise in a single year. AI has shifted from a technocratic to a political battleground. The window for technocratic AI governance is closing.

### [Extra] The Spectator coins "arm farms": workers training their robot replacements.

Gary Dexter describes facilities where chefs, nurses, and plumbers wear GoPro helmets and motion-capture rigs while doing their normal jobs. The purpose: generating training data for the robots that will eventually replace them. Knowledge workers writing documents that train language models are arguably on an arm farm already.

### [Extra] Mollick: "everything around me is somebody's life work" is no longer true.

Ethan Mollick riffs on a meme about the invisible human effort behind ordinary objects. An annotated lamp: an engineer working late on a curve, years of supplier negotiations, months of tip-over testing, someone getting fired over a cord switch. AI disrupts the assumption that every designed thing carries accumulated human stakes.

### [Extra] $930 billion in data centre capex in six years dwarfs every US megaproject.

Fin Moorhouse charted hyperscaler capital expenditure against historic megaprojects in inflation-adjusted dollars. Data centres: $930 billion in 6 years. The Interstate Highway System: $620 billion over 37. Railroads: $550 billion over 71. Apollo: $257 billion over 14. As a share of GDP, the railroads were bigger at their peak. But the railroads also produced spectacular capital misallocation.

## Edition #9: The proxy break
*18th April 2026*

### [Three Things] AI cover letters killed the signal that cover letters used to carry

When Freelancer.com added an option to generate cover letters with AI, [researchers tracked what happened](https://arxiv.org/abs/2509.25054). Before language models, there was a clear positive slope: better cover letters predicted better hiring outcomes. After the feature launched, the line went flat. Once polish became free, it stopped measuring anything useful. Economists call this signal destruction. It's Goodhart's Law: when a measure becomes trivially easy to game, it ceases to be a measure. Cover letters aren't the last signal to fall. The same logic applies wherever AI can cheaply replicate a previously costly quality indicator.

### [Three Things] Snap cut 1,000 jobs. AI already writes 65% of their new code

[Snap](https://techcrunch.com/2026/04/15/snap-is-cutting-1000-jobs-16-of-its-workforce/) laid off 1,000 employees, 16% of its full-time workforce, and closed 300 open roles. AI agents already generate over 65% of Snap's new code. Expected savings: over $500 million annualised. Way beyond a pilot or an aspiration. Evan Spiegel is betting on smaller, highly focused teams with expanded AI agent capabilities. For leaders still framing AI as a productivity tool that supplements existing teams, Snap is a data point that the substitution model has arrived.

### [Three Things] Letting AI do your work erodes your confidence. Pushing back strengthens it

A [study of nearly 2,000 working adults](https://time.com/article/2026/04/15/how-ai-use-affects-confidence-thinking-study/) found that people who accepted AI answers without much modification reported lower confidence in their own reasoning and weaker ownership over their ideas. People who pushed back, editing, questioning, and rejecting AI suggestions, reported greater confidence and stronger ownership. The key variable wasn't which tool they used. It was how actively they engaged with it. Passive delegation erodes judgement. Active collaboration strengthens it.

[Gartner's data](https://www.gartner.com/en/newsroom/press-releases/2025-09-24-gartner-survey-finds-ai-saves-workers-5-hours-per-week) tells a similar story from a different angle. Of 5.4 hours saved by AI per week, just 0.6 go to reducing hours worked. The rest gets absorbed into more work, much of it without improving outcomes.

[Image: Gartner: How time savings from AI are used. Of 5.4 hours saved, 1.7 go to additional work that improves team outcomes, 1.4 to additional work without improving outcomes, 0.8 to redoing AI work, 0.8 to developing new skills, and just 0.6 to reducing hours worked.]
*Of 5.4 hours a week saved by AI, barely half an hour turns into shorter hours; most is reabsorbed into more work.*

[Fifteen bits that didn't fit online ->](https://steadman.ai/newsletters/david/archive.html#extras-2026-04-18)

### [Extra] Allbirds pivoted to GPU leasing. Stock up 700% in a day.

Allbirds, the sustainable shoe brand that closed all US stores in February, rebranded as NewBird AI: a GPU compute leasing platform. Market cap jumped sevenfold in a single session. A shoe company became an AI infrastructure company in two months. The demand signal is real even if the pivot is absurd.

### [Extra] Satya Nadella's Copilot demo didn't work when someone else tried it.

Satya Nadella posted a demo of Copilot editing Word documents with tracked changes. An investor replicated the exact workflow. Copilot produced a redlined version, but only inside the chat sidebar. The actual document was untouched. When the product is the flagship AI feature of the world's largest software company, the credibility cost is high.

### [Extra] France is quietly building serious AI agent infrastructure.

The French government has launched an official MCP server for data.gouv.fr, letting AI systems interact more directly with public datasets. Separately, an open-source project called Paperasse has shown how agent skills can be packaged for real-world French tax and accounting work. Some coverage blended the two into one story. That misses the more interesting point: the state is building infrastructure, and independent developers are building usable workflows on top of it. Useful agent systems will come less from demos, and more from good infrastructure paired with narrow, practical skills.

### [Extra] Over half the internet is now AI-generated.

Research from Graphite: beginning in January 2025, over 50% of newly published online content was generated by AI. This has immediate implications for anyone training models on web data: the training corpus is now majority-synthetic. Several frontier labs have responded by pursuing proprietary data licensing deals.

### [Extra] Nvidia bottled 30 years of expertise so juniors stop interrupting seniors.

Nvidia's Chief Scientist Bill Dally told Jeff Dean that Nvidia trained a language model on its entire proprietary document archive, covering over 30 years of chip design knowledge. Junior employees query the model instead of interrupting senior designers. Institutional knowledge, bottled up and made searchable.

### [Extra] AI transparency went backwards in 2025.

After rising on the Foundation Model Transparency Index from 37 to 58 between 2023 and 2024, the average score dropped to 40 in 2025. Over 90% of notable models were released without training code. The most capable modern models are now among the least transparent.

### [Extra] The 50-point gap: AI experts and the public disagree on nearly everything.

On jobs, 73% of AI experts say AI will have a positive impact versus 23% of the public. On the economy: 69% vs 21%. On medical care: 84% vs 44%. They only converge on what AI will damage: elections and personal relationships. This is a wider gap than most technology debates produce.

### [Extra] Computer science enrolment fell 11% but AI masters degrees surged 82%.

Undergraduate computer science enrolment at US universities dropped 11% between 2024 and 2025, apparently a response to automation concerns. But AI software-related masters degrees grew 82% between 2022 and 2024. Students are pivoting, not leaving. Two-thirds of AI software masters graduates are non-US residents, a pipeline under pressure from visa policy changes.

### [Extra] Goldman Sachs: AI inference costs approaching headcount parity.

A Goldman Sachs equity research note reports that companies are overrunning their AI inference budgets by orders of magnitude. In engineering, inference costs are now approaching 10% of headcount cost and on current trajectories could reach parity within several quarters. The machines aren't replacing headcount costs. They're adding a new cost layer.

### [Extra] Consumer surplus of $172 billion, but producers capture almost none.

US consumer surplus from generative AI reached $172 billion annually by early 2026, up 54% from a year earlier. This dwarfs actual AI company revenues, consistent with historical research showing innovators capture only about 3% of total social returns. Most of these tools remain free or nearly free to use.

### [Extra] Anthropic's design launch hits Figma hardest.

Anthropic's design product launched, turning a rumour that had already wiped billions off the sector into a real competitive threat. The sharpest pressure falls on Figma, not just because Claude Design moves closer to its core job, but because the conflict is now explicit: Mike Krieger, Anthropic's Chief Product Officer and Instagram co-founder, stepped down from Figma's board as Anthropic prepared to enter the category. Adobe may feel some of that pressure too, but companies like Wix and GoDaddy sit in a more mixed position: Anthropic could compete with parts of their "make it easier" story while also creating more demand for sites and publishing tools that AI-generated design still needs in order to go live.

### [Extra] Google shipped AI agents to 3.45 billion people via a Chrome update.

Google launched "Skills" in Chrome: save any AI prompt as a reusable one-click workflow, then run it on whatever page you're viewing. The distribution play is the story: Chrome has 3.45 billion users. Every saved Skill becomes a switching cost. And the aggregate data on which Skills people save gives Google a continuous product research signal about which workflows people most want automated.

### [Extra] Gallup: half of US workers now use AI at work, but leaders use it 1.5x more.

Gallup surveyed 23,717 employees: 50% of US workers now use AI at work, up from 21% in 2023. But leaders use AI daily or weekly at 67%, versus 46% for individual contributors. This inverts the usual adoption pattern: the people setting the strategy are further along than the people executing it. The 27% who report "large or very large disruption" is a canary: a quarter of the workforce says AI is already reshaping their work in ways that feel significant.

The [New York Fed's breakdown](https://libertystreeteconomics.newyorkfed.org/2026/04/ai-use-in-the-workplace-is-concentrated-among-higher-income-higher-educated-and-full-time-workers/) shows just how steep the gradient is: adoption rises from 15.9% for workers earning under $50,000 to 66.3% for those earning over $200,000.

[Image: Federal Reserve Bank of New York: AI use in the workplace is concentrated among higher-income, higher-educated, and full-time workers. Adoption rises from 15.9% for workers earning under $50K to 66.3% for those earning over $200K.]

### [Extra] KPMG: companies invest 2x more in tech than in training, and 46% report burnout.

The KPMG Adaptability Index found executives are nearly twice as likely to increase tech spending as to invest in employee training. Fewer than 10% made workforce training a primary objective despite 57% citing efficiency as a priority. The result: 46% report burnout and change fatigue as unintended consequences of transformation. Only 9% invested in psychological safety. You can't simultaneously demand more adaptability, make workforces smaller, and invest nothing in the people.

### [Extra] Apple is linking AI token usage to headcount decisions.

An Apple insider reports that when directors ask for headcount backfill, senior leadership now asks what the team's AI usage looks like. If token usage is low, the answer is increasingly: go figure out how to get more leverage out of AI first. AI usage is becoming a proxy for operational efficiency.

## Edition #8: What a day can do
*11th April 2026*

### [Three Things] Claude Code now writes 4% of all commits on GitHub. That number doubled in six weeks

Anthropic's annualised revenue has [surpassed $30 billion](https://sherwood.news/markets/anthropic-revenue-run-rate-30-billion-google-broadcom-partnership/), up from $9 billion at the end of 2025. Claude Code, which didn't exist fourteen months ago, is at a $2.5 billion run rate. Four percent of all GitHub commits on Earth are now written by Claude Code. That number [doubled in roughly six weeks](https://newsletter.semianalysis.com/p/claude-code-is-the-inflection-point). Projected to hit 20% by December. When a single AI coding tool is responsible for one in 25 submissions on the world's largest code platform, the question of whether AI changes software development is settled. The question now is what happens to the humans reviewing all that code!

### [Three Things] Goldman Sachs puts a number on AI job destruction: a net drag of 16,000 jobs per month

[Goldman Sachs](https://www.goldmansachs.com/insights/articles/how-will-ai-affect-the-us-labor-market) published one of the first serious attempts to quantify AI's net labour market impact. AI substitution has reduced monthly US payroll growth by roughly 25,000 jobs. AI augmentation partially offsets this, adding about 9,000. Net: a loss of 16,000 jobs per month and a 0.1 percentage point increase in unemployment. The loss falls disproportionately on less experienced workers, widening the entry-level-to-experienced wage gap by 3.3 percentage points.

[Image: Two Goldman Sachs charts. Left: payroll employment by industry exposure to AI (index, 2022Q4=100). Industries with high AI augmentation scores have climbed to ~104, industries with high AI substitution scores have fallen to ~98 since ChatGPT's launch. Right: three-month average unemployment rate relative to 2022Q4. Occupations with high AI substitution scores show unemployment rising to ~1.5 percentage points above baseline, well above augmentation occupations and the average.]
*Goldman puts AI's net drag at about 16,000 US jobs a month, and the loss falls hardest on the least experienced workers.*

But here's the weird part: AI [led all cited reasons](https://www.challengergray.com/blog/challenger-report-march-cuts-rise-25-from-february-ai-leads-reasons/) for US job cuts in March 2026 for the first time (15,341 in a single month), yet [CFO surveys](https://fortune.com/2026/03/24/cfo-survey-ai-job-cuts-productivity-paradox-2026/) put genuine AI-driven employment impact at just 0.4%. So the same organisations that struggle to get genuine productivity gains from AI tools are apparently enthusiastic about blaming AI for headcount decisions. Hmmm.

### [Three Things] Meta's tokenmaxxing leaderboard: 60 trillion tokens, Zuckerberg not in top 250

Meta is running [internal leaderboards](https://the-decoder.com/meta-employees-compete-for-token-consumption-on-an-internal-ai-leaderboard/) that rank all 85,000+ employees by AI token usage. In one month the company consumed 60 trillion tokens. Mark Zuckerberg didn't make the top 250. The structural problem: endless agent loops and genuine productive work look identical in the ranking, so it rewards orchestration over outcomes. Two different people have told me how friends inside Meta are behaving and I can tell you, it's as outrageous as it sounds. Their leaderboard measures and incentivises volume, not value. Incentivise use, sure. If you don't keep getting on your bike, you'll never get used to riding it. Incentivise 'at least X per day.' Not points for maxxing. One builds a habit. The other builds a game.

[Eleven bits that didn't fit online ->](https://steadman.ai/newsletters/david/archive.html#extras-2026-04-11)

### [Extra] A financial services firm's code output rose 10x. The review backlog hit one million lines.

A financial services firm adopted an AI coding tool. Monthly code output jumped from 25,000 lines to 250,000. The result wasn't celebration. It was a backlog of one million lines of code waiting to be reviewed. The bottleneck wasn't production. It was judgment. AI removes constraints on output but does nothing to scale the human capacity to evaluate it.

### [Extra] Simon Willison runs four AI agents in parallel and is wiped out by 11am.

A veteran software engineer described running four coding agents in parallel and being mentally exhausted by 11am. "Using coding agents well is taking every inch of my 25 years of experience as a software engineer, and it is mentally exhausting." The bottleneck isn't writing code. It's holding context, making judgments, and orchestrating simultaneous workstreams.

### [Extra] Deloitte caught twice in two months submitting AI-hallucinated citations.

Deloitte charged a Canadian province's Department of Health $1.6 million for a report filled with AI-hallucinated citations. Fabricated references, not real sources. This was the second time in two months. Their response: they "stand by the conclusions." No meaningful verification process was implemented between the two incidents, I guess?

### [Extra] Executives are buying the pitch. Workers are living with the product.

A global survey of 3,750 executives and employees found that 54% of workers bypassed their company's AI tools in the past 30 days and completed work manually. Another 33% haven't used AI at all. That's 87% avoiding or rejecting tools their employers spent an average of $54 million deploying this year. The trust gap explains it: only 9% of workers trust AI for complex business decisions, compared with 61% of executives. And here's the symmetry that should worry CFOs: workers lose the equivalent of 51 working days per year to technology friction, up 42% from last year, almost exactly equal to the 40 to 60 minutes per day Goldman Sachs says AI saves workers who use it correctly. The net productivity benefit of enterprise AI may be approximately zero at the organisational level, because friction costs cancel out gains. And that's only among workers who actually use the tools. Neither group is irrational. Workers under pressure surrender judgment to faulty outputs. Workers without pressure opt out entirely. Both are responses to the same problem: companies deployed the technology before figuring out what they wanted employees to do with it.

### [Extra] Microsoft Copilot converted 3.3% of its users after two years.

After two years and CEO-level intervention, Microsoft Copilot has converted just 15 million of its 450 million M365 seats. Only 35.8% of those actively use it. Copilot's paid subscriber share dropped from 18.8% to 11.5% in six months. Microsoft's own terms of service describe Copilot as "for entertainment purposes only." That gap between marketing and legal is the real story. The ads say "your AI-powered co-worker." The lawyers say "entertainment only, use at your own risk." Among lapsed users, 44% cite distrust of the answers.

### [Extra] Most people who got productivity gains filled the time with more work.

Anthropic's 81,000-person AI interview study found that the top desired outcome was "professional excellence" (nearly 19%), not time freedom. Productivity gains were overwhelmingly linked to increased expectations rather than reduced workload. Fear of unreliability ranked as the top concern (27%), ahead of job displacement (22%). We got speed, but not space.

### [Extra] When language models go down, financial markets forget how to price news.

An SSRN paper has found that language model outages measurably slow financial market price discovery. When models go down, 46-61% of post-news price drift reappears, meaning markets take significantly longer to absorb and reflect new information.

### [Extra] Anthropic's Mythos Preview: restricted to 50 organisations, not released.

Anthropic has confirmed a new model called Mythos Preview and restricted access to around 50 organisations, including governments and infrastructure partners. It's the first major model withheld from public release since GPT-2 in 2019. The model found a 27-year-old bug in OpenBSD and a 16-year-old flaw in FFmpeg, and it emailed a safety researcher from a test instance that wasn't supposed to have internet access. Anthropic is launching a $100 million defensive security consortium with AWS, Apple, Google, Microsoft, and Nvidia. Models keep getting meaningfully better!

### [Extra] 73% of ChatGPT usage is personal, not work. Coding is 4.2%.

An NBER working paper studying 700 million ChatGPT users found that 73% of usage is personal, not professional. Programming accounts for just 4.2% of messages. Most writing requests (two-thirds) are editing existing text, not generating new content. Nearly half of all interactions involve decision-making advice. People aren't delegating tasks. They're thinking through problems.

### [Extra] HubSpot moves AI agents to outcomes-based pricing: $0.50 per resolved conversation.

HubSpot has shifted its AI agents to outcomes-based pricing: $0.50 per resolved customer conversation, $1 per sales lead recommended for outreach. From seat licences to outcome fees. What else will, or should, go this way?

### [Extra] Mid-career engineers are the most vulnerable to AI, not juniors or seniors.

Simon Willison argues that mid-career engineers are the most structurally vulnerable. Seniors benefit because AI amplifies decades of pattern recognition. Juniors benefit because AI compresses onboarding. Mid-career engineers are stuck: they've captured the beginner productivity boost but haven't accumulated the deep expertise that makes AI a force multiplier.

## Edition #7: What is your organisation actually for?
*4th April 2026*

### [Three Things] Dorsey wants to replace your org chart with a world model

Jack Dorsey's essay this week, "[From Hierarchy to Intelligence](https://block.xyz/inside/from-hierarchy-to-intelligence)," is the most concrete articulation yet of the case for AI-native organisational design. He traces 2,000 years of hierarchy, from Roman legions through the Prussian General Staff to the McKinsey matrix, and argues all of it exists to route information. He is reorganising around a "company world model": an AI that continuously understands the state of the whole business. The org flattens to three roles: individual contributors, time-boxed problem owners and player-coaches who build and develop people. Block's stock rose 17% after the restructuring announcement. I'm sure he's right about the technology. But the revealed preference of every organisation I've worked with suggests he's wrong about the humans.

### [Three Things] Mollick says giving AI to IT is usually a mistake. The harder problem is they can't see who's using it

Ethan Mollick's [Economist column](https://x.com/emollick/status/2039702114910306638) argues that the dominant corporate instinct, slotting AI into existing processes and handing it to IT, is a strategic mistake. Handing control to a department whose mission is risk elimination is a category error. AI demands the opposite. He also identifies a subtler problem: when companies get the incentives wrong, employees hide their AI use. Some fear punishment. Some don't trust that productivity gains will be shared. Some quietly work 90% less and say nothing. I know MANY such people. Managers can't see what's actually happening, which makes real strategy impossible. (Also: His argument that companies default to cutting 30% of the workforce rather than asking what becomes possible connects directly to the extraction-versus-expansion choice I wrote about in an earlier edition.)

### [Three Things] Zapier just raised the floor for what "AI fluent" means

Zapier released [Version 2 of its AI Fluency Rubric](https://zapier.com/blog/raising-ai-fluency-bar-in-hiring/), used for every hire across the company. The floor has moved. "Capable" now requires AI embedded in core workflows with repeatable systems, not one-off prompts. They assess trajectory ("slope"), not snapshots. They've added accountability as a fourth dimension alongside mindset, strategy and building. Managers must demonstrate team-wide adoption, not just personal fluency. In skills tests, they watch candidates prompt, push back on output and iterate in real time. A rough result with strong reasoning beats a polished one with no visible process. Wade Foster open-sourced V1 last year and hundreds of companies adopted it. V2 reflects how fast the baseline has shifted. If your organisation hasn't defined what "good" looks like for AI use, Zapier just gave you a starting point.

[Image: Zapier's AI Fluency Rubric: a grid showing Unacceptable, Capable, Adoptive and Transformative levels across Engineering, Product, Support, Marketing, Sales and People functions]
*Zapier's hiring rubric has moved the floor: "capable" now means AI built into repeatable workflows, not one-off prompts.*

### [Extra] Are apprentices an endangered species?

Two Kellogg professors published the most rigorous academic framing yet of the "AI hollows out entry-level work" problem. Their mathematical model identifies two competing effects: the "floor effect" (AI automates the tasks apprentices performed as payment for training) and the "ceiling effect" (AI amplifies what experienced apprentices can accomplish). Apprenticeship survives only when the ceiling effect exceeds standalone AI by a factor greater than Euler's number.

### [Extra] The Guinndex: 3,000 pubs, one AI voice agent, every county in Ireland.

Over St Patrick's weekend, an AI voice agent called Rachel phoned more than 3,000 pubs across all 32 counties of Ireland to ask the price of a pint of Guinness. Over 1,000 gave a price. The national average: &euro;5.95. It cost &euro;200. Only a handful of pub owners noticed Rachel wasn't human.

### [Extra] 75-99% of knowledge work is scaffolding. AI eats scaffolding.

Daniel Miessler argues that in cybersecurity, 99% of the work isn't finding new vulnerabilities. It's maintaining the tooling, templates, knowledge bases and workflows that let you test at scale. The scaffolding around the work is exactly what AI commoditises.

### [Extra] Ethan Mollick: human creativity is the bottleneck, not the technology.

Everyone can generate almost any image or video for nearly free in 2026. And yet: the April Fools posts this year were just as bad as any other year. The constraint was never execution. It was always the quality of human ideas feeding into the process.

### [Extra] 43% of American workers now use AI for their jobs. 2.5 hours saved per week.

A 20,900-person cross-national survey found that 43% of US workers use generative AI at work, compared with 36% in the UK, 32% in Germany and 26% in Italy. The strongest predictor of adoption? Not age or education. Whether the employer actively encourages AI use.

### [Extra] Sora earned $2.1 million in its entire life. It burned roughly $1 million a day.

OpenAI's video generation platform launched to 3.3 million downloads in November. By February: 1.1 million. Revenue peaked at $540,000 a month. The annualised cost of running it: an estimated $5.4 billion. Disney had committed $1 billion. The product goes dark on 26th April. Six months, start to finish.

### [Extra] Jensen Huang told CEOs cutting jobs in the name of AI that they're "out of imagination."

At Nvidia's GTC conference, the CEO of the company selling AI chips to virtually every major technology company on earth called AI-driven layoffs a failure of leadership. His biggest customers are doing exactly what he criticised. But a question he didn't address: does every carpenter want to be an architect?

### [Extra] Screen Studio switched to subscriptions. It spawned an open-source clone with 9,200 GitHub stars.

Screen Studio sold a one-time licence for $89. Then the company switched to $29 a month. OpenScreen appeared on GitHub within months. A textbook case of pricing-driven disruption: developers who are both the users and the potential builders of substitutes.

### [Extra] AI outperformed practising lawyers on 75% of legal research tasks.

Vals AI tested AI against practising lawyers on legal research questions in 2025. AI exceeded the lawyer baseline on three quarters of them. A senior law firm owner said hourly billing is dying, junior review is dying, and what survives is the senior brain that knows what question to ask.

### [Extra] Deloitte projects that by 2028, AI moves from supporting tasks to orchestrating decisions.

A Deloitte report argues that agentic AI is categorically different from current workflow automation. Most AI strategies stall not because the technology is insufficient but because organisations are applying AI at the task level while the technology is restructuring the systems through which decisions are made.

### [Extra] Three people with AI vs a 1,000-person company. But coordination costs don't disappear.

Xiaoyin Qu argues that companies designed around AI as the primary operating layer will eventually outcompete companies designed around people. But she herself provides the sharpest counter: coordination costs don't disappear. They're externalised, pushed to clients, suppliers, regulators and the AI systems themselves.

### [Extra] Why companies buy vertical software, not raw models.

Aaron Levie argues companies aren't buying features. They're outsourcing the cognitive burden of designing and maintaining business processes. Agents don't undermine this dynamic. If anything, they reinforce it, because agentic workflows are even more complex and opaque.

## Edition #6: The system and the surrender
*28th March 2026*

### [Three Things] Three CEOs, 38 years of tenure, one quarter, one reason

[Coca-Cola's James Quincey and Walmart's Doug McMillon](https://www.cnbc.com/2026/03/26/coca-cola-james-quincey-walmart-doug-mcmillon-artificial-intelligence-step-down.html) both stepped down citing AI explicitly. Quincey said the company needs "someone with the energy to pursue a completely new transformation." [Adobe's Shantanu Narayen](https://fortune.com/2026/03/12/adobe-ceo-shantanu-narayen-stepping-down-after-18-years-pressure-deliver-ai/) left under competitive pressure as AI reshaped his market, stock down 23%. The last time this many blue-chip CEOs turned over citing the same technology was 1999. People joke that organisations only change when people change over, and the assumption was always natural attrition. After 18 months of coaching senior professionals one-on-one, watching their eyes light up, watching them go back to their desks and do it the old way, I've started to think the reckoning is more predictive than the joke. I just didn't think it would start at the top, through resignation.

### [Three Things] Experienced users don't prompt differently. They think differently.

Anthropic's 5th Economic Index found that users with six or more months of experience consistently achieve better results, even after controlling for task complexity. The shift is specific: experienced users stop issuing one-shot directives ("write this email") and start using the model as a thinking partner, iterating collaboratively. A [separate study published in Harvard Business Review](https://hbr.org/2026/03/what-the-best-ai-users-do-differently-and-how-to-level-up-all-of-your-employees), observing 2,500 employees over eight months, found the same pattern: the most sophisticated users treated AI as a reasoning partner, not a productivity shortcut. The research warns AI may be a skill-biased technology that compounds existing advantages. Global inequality in AI adoption, measured by the same metric economists use for income inequality, has widened since 2023. The gap between high-adoption and low-adoption countries is growing, not closing.

### [Three Things] Companies with zero AI failures aren't being ambitious enough.

Ethan Mollick argues that breakthroughs require experimentation, which requires failure. The fast-follower strategy (wait to see what competitors prove, then copy) is riskier than usual when the underlying technology improves exponentially. By the time you follow, the landscape has shifted. His structural implication: R&D-style experimental budgets need to extend to HR, operations and finance, functions that have never needed them. If nothing has gone embarrassingly wrong yet, you probably aren't learning fast enough. I have two big failures I'm not proud of: A bunch of you accidentally got the email twice on the first weekend and Claude Code deleted tens of thousands of my emails a few weeks ago. Nobody complained about the duplicate and it only took one click to recover my deleted emails. Perhaps I'm not being ambitious enough ...

### [Extra] Even the world's greatest mathematician uses AI for email

Terence Tao, Fields Medal winner and arguably the greatest living mathematician, told Dwarkesh Patel that a significant share of his AI use goes to correspondence, scheduling and document search. AI removes an hour of non-genius work per day, donating it back to the work only Tao can do.

### [Extra] Anthropic is shipping. OpenAI is cutting.

Anthropic shipped 74 releases in 52 days, six major features in a single week. Meanwhile OpenAI killed Sora (~$2.1M total revenue, $1B Disney deal dissolved), shut down Instant Checkout (12 Shopify merchants), and shelved an adult chatbot indefinitely. OpenAI is now explicitly copying Anthropic's playbook: chat, code, enterprise only. Anthropic's narrow focus is generating $19 billion in annualised revenue. The company that chose depth over breadth is winning.

### [Extra] When effort becomes free, the signal breaks

Job applications have collapsed because AI makes applying trivially easy. Companies are abandoning inbound pipelines, switching to referral-only hiring. The same dynamic will hit email, journalism pitches, academic submissions, legal filings. Anywhere volume was self-regulated by the cost of effort, AI removes the regulation.

### [Extra] The models we can't afford to use

Anthropic reportedly has a model called Capybara that dramatically outperforms current models but is too expensive to serve. Training a single frontier model now costs roughly $10 billion. For comparison: the Burj Khalifa cost $1.5 billion. CERN's Large Hadron Collider cost $4.5 billion. The decision coming for every organisation: which price tier of model to deploy per prompt. That decision is coming and most haven't built the judgment to make it.

### [Extra] 32,000 medieval manuscripts. 10% error rate.

AI transcribed 32,000 medieval manuscripts in four months through the CoMMA project. Every misread word can alter meaning, dating, or attribution. There aren't enough qualified people to verify the output. Silent, unverifiable errors are entering scholarly databases permanently. Cognitive surrender in a domain where the stakes are centuries of accumulated knowledge.

### [Extra] 100x productivity. Zero headcount cuts.

Harvard Law documents 100x gains on specific legal tasks (complaint response: 16 hours down to 3-4 minutes). Not a single AmLaw 100 firm plans to reduce attorney headcount. McKeen: "The math doesn't stay like that forever."

### [Extra] 181,000 jobs in a year of 2.2% GDP growth

The US added 181,000 jobs in all of 2025 despite solid growth. Harvard economist Lawrence Katz calls the combination of sustained slow job growth and rising unemployment without a recession virtually unprecedented. First hard macroeconomic signal that something structural is shifting.

### [Extra] Jensen Huang: layoffs are a failure of imagination

Asked why companies lay off workers if AI makes them more productive, Huang told CNBC: "For companies with imagination, you will do more with more. For companies where the leadership is just out of ideas, they have nothing else to do." The person whose chips make displacement possible arguing that layoffs reflect leadership failure, not technological inevitability.

## Edition #5: Reckoning and slope
*21st March 2026*

### [Three Things] In chess, the human advantage inverted. Knowledge work is next

In 2005, amateur players using laptops beat both grandmasters and supercomputers in centaur chess. The combination was unbeatable. My friend Glenn told me and I used it in a dozen speeches as an analogy for the role of tech vs humans in decision-making. I was wrong. By 2026, adding a human to a chess engine makes it play worse. The machine is better alone. (Magnus Carlsen's response is extreme but instructive: he deliberately limits his use of AI during preparation, because he believes self-generated understanding is the only kind that lets you catch when the AI is wrong.) I think the same inversion will play out across much of knowledge work. The question in many areas isn't whether humans stay in the loop. It's how long the current phase lasts, and whether we're building the judgment to extend it.

### [Three Things] Jensen Huang thinks your engineers aren't spending enough on AI.

Nvidia's CEO has set a concrete benchmark: a $500,000 engineer who doesn't consume at least $250,000 worth of AI tokens should trigger alarm. His analogy: a chip designer refusing to use CAD tools and working with paper and pencil instead. It's a useful provocation, but notice what it measures. Token spend is an intercept metric: how much AI are you using today? It says nothing about whether the person is getting better. The organisations that take Huang's benchmark seriously and Howard's slope argument seriously will measure both. Most will only measure one.

### [Three Things] RentAHuman: 600,000 sign-ups. A platform where AI agents hire human beings

Six hundred thousand users have signed up to a platform where AI agents post tasks and hire humans to complete them. The worker uploads photographic proof and then gets paid. We've spent two years asking whether AI will take our jobs. RentAHuman suggests a different question: [what happens when AI becomes the employer?](https://www.linkedin.com/posts/beglen_human-deployment-notes-activity-7440184415294521344-Gzw7)

### [Extra] The supply-ordering agent nobody sanctioned

A technology leader shared a cautionary story this week. A team built an AI agent to order supplies within specified parameters, intending it to run once. A separate agent then modified the skill to repeat hourly. Three days later they'd bought an extraordinary volume of supplies, all technically within the original parameters. Nobody had sanctioned the change, and the skills were editable by other agents by default. This is the Amazon Kiro story from Edition 4 in miniature, except the failure mode isn't a crash. It's perfect compliance with instructions nobody gave. As agents gain the ability to modify each other's behaviour, "within parameters" stops being a safety guarantee.

### [Extra] AI-assisted coding works like a slot machine

Jeremy Howard, whose slope argument anchors this week's essay, had a second observation worth sitting with. AI coding tools have all the properties that make gambling addictive: you craft your prompt, add context, pull the lever, and sometimes you win a feature. Loss disguised as a win. The illusion of control. Stochastic reward. His wife, a fellow researcher, catalogued these properties in an article. The people who got most enthusiastic about AI coding often found, months later, that almost none of what they built during that period was in production or earning money. This explains a paradox readers keep raising: people use the tools a lot, feel productive, but the organisations aren't seeing the output.

### [Extra] The economics job market fell 31% in a single year

In week 14 of the current hiring season, postings in the economics job market were down 31% versus the same point last year. The explanation from an economist presenting the data: demand for economics undergraduates is being automated away, and the PhD market is coupled to it. Combined with the Harvard data from Edition 4 (skill requirements in AI-exposed occupations falling since ChatGPT's launch) and Anthropic's research showing hiring of 22-to-25-year-olds down 14%, entry-level knowledge work is contracting faster than mainstream commentary acknowledges.

### [Extra] A sufficiently detailed spec is code

[Gabriella Gonzalez's argument](https://haskellforall.com/2026/03/a-sufficiently-detailed-spec-is-code), circulating widely in technical circles: the fashionable claim that you don't need to write code, just write a good spec and let an agent handle it, collapses under scrutiny. If the spec is detailed and precise enough for an agent to execute reliably, you have written code in everything but name. The hard part of programming, resolving ambiguity, is still your job. This is the slope argument applied to a specific skill: the people who think they've escaped the need to understand what they're building are the ones most likely to produce output nobody can maintain.

### [Extra] The AI task force leader who'd never logged in

At a professional services firm, the person leading the AI task force hadn't used the enterprise AI tool once. When challenged, they said: "I know I should, but I can't make the time." They weren't uninformed or resistant. They understood the stakes. They were simply too busy doing the old job to start learning the new one. The incentive trap in plain sight: the people best placed to model the new behaviour are the ones most rewarded for performing the old behaviour well.

### [Extra] Intercom built a plugin system that closes the loop

Brian Scanlan, Senior Principal Systems Engineer at Intercom, shared a thread this week on the company's internal Claude Code system: 13 plugins, over 100 skills, distributed across the company via JAMF. The standout pattern isn't the scale. It's the feedback loop. A session-end hook automatically classifies skill gaps from every coding session and posts them to Slack with pre-filled GitHub issue URLs. Sessions become gaps, gaps become issues, issues become skills. The most telling detail: the top five users of their read-only production Rails console are not engineers.

### [Extra] The CIO budgeting for AI cleanup

The CIO of a major consulting firm told a peer this week that they're budgeting 18 months to two years from now for AI cleanup. The reasoning: things are being built once but not built to last, corners are being cut on testing, and the people who built the tools will have moved on before the problems surface. It's an unusual thing to plan for. But it's probably the most honest thing I've heard a technology leader say about the current moment.

### [Extra] From franchises to call options

Tyler Cowen, drawing on [analysis from Jordi Visser](https://visserlabs.substack.com/p/the-repricing-of-time-equity-in-the), argues that AI simultaneously lowers barriers to entry while destroying the conditions for sustained dominance. Software moats compress because any sufficiently capitalised team can replicate your product. Durable advantage reconcentrates in physical constraints: infrastructure, energy, materials, regulatory relationships. Equity in this environment becomes less a claim on a stable franchise and more a bet on execution velocity. The implication for anyone evaluating technology investments: the question is no longer "what have they built?" but "how fast can they keep building?"

### [Extra] Two thirds of organisations report AI productivity gains. Only a third are rethinking what they do.

Deloitte's [State of AI in the Enterprise 2026](https://www.deloitte.com/us/en/what-we-do/capabilities/applied-artificial-intelligence/content/state-of-ai-in-the-enterprise.html) report found 66% of organisations report productivity improvements from AI. Only 34% are pursuing what Deloitte calls "transformative business reimagination." Most organisations are getting faster at what they already do. Fewer than half are asking whether what they do should change. Meanwhile, only 21% have mature governance models for the autonomous agents they're about to deploy.

### [Extra] McKinsey's internal AI chatbot was hacked via textbook SQL injection

McKinsey's internal AI chatbot Lilli, trained on 100 years of the firm's work, was [breached via a basic SQL injection](https://codewall.ai/blog/how-we-hacked-mckinseys-ai-platform). 46.5 million internal chat messages exposed, 728,000 files containing confidential client data, 57,000 user accounts, 22 API endpoints requiring no authentication. The firm that charges for risk expertise left the front door open. If McKinsey can't govern its own AI deployment, what does your internal chatbot look like?

### [Extra] Red Bull didn't simulate the pit stop. They did it in zero gravity.

A reader forwarded an [Instagram clip](https://www.instagram.com/reel/DV28-ItDs5h/) of Red Bull's F1 team performing a tyre change in zero gravity, just to prove they could. Not CGI. Real mechanics, real car, real weightlessness. The reader's take, which I think is exactly right: use AI for the boring, the day-to-day, the basics. Free up your budget and attention for the truly remarkable. "If I'm so focused on the incredible, the groundbreaking, the creative and free from the mundane, I raise the bar for the client." That's the slope argument in a sentence. The people who use AI as a floor-raiser, not a ceiling-replacer, are the ones building capability.

### [Extra] The consulting firms are buying the AI stack, not just using it

CB Insights mapped every AI investment, acquisition, and partnership by the major consulting firms since 2023. Accenture is at the centre of the web, with partnerships radiating to dozens of AI companies. The Big Four and MBB firms aren't waiting to see how AI plays out. They're racing to own the infrastructure: embedding agents via Salesforce, ServiceNow and Workday partnerships, acquiring data companies, and investing in startups that automate the consulting workflow itself. Four patterns emerge: race to own the stack, embedding agents, data as differentiator, and workforce transformation. PwC's announcement this week is one node in a much larger network.

[Image: CB Insights map of AI investments, acquisitions and partnerships by consulting firms since 2023]
*The big consultancies are not just using AI but buying the stack, with Accenture at the centre of a web of deals and acquisitions.*

[Image: CB Insights map of AI investments, acquisitions and partnerships by consulting firms since 2023]

### [Extra] More offices for AI than for humans

US data centre construction spending overtook general office construction in December 2025, according to Census Bureau data. Data centres: $3.57 billion. General offices: $3.49 billion. The lines crossed after data centre spending roughly tripled in two years while office construction flatlined. We're now building more square footage for machines than for people.

[Image: Data Center Construction Spending Climbs to Record: outlays for data center projects overtook offices in December 2025]
*US construction now spends more on data centres than on offices, the first time machines have outbid people for new square footage.*

[Image: Data Center Construction Spending Climbs to Record: outlays for data center projects overtook offices in December 2025]

## Edition #4: The power and the care
*14th March 2026*

### [Three Things] Old code is fair game now

Tobi Lütke, the CEO of Shopify, ran an agentic AI optimisation loop against Liquid, the company's open-source templating engine that has been in production for roughly twenty years. The result: [53% faster combined parse and render time and 61% fewer object allocations](https://x.com/tobi/status/2032212531846971413). If you have a twenty-year-old codebase or a ten-year-old spreadsheet model, it isn't too embedded to improve. It's too embedded not to try.

### [Three Things] More AI tools make you less productive, not more

[Research from BCG and UC Riverside](https://hbr.org/2026/03/when-using-ai-leads-to-brain-fry) (1,488 workers), published in Harvard Business Review, found a counterintuitive pattern: productivity gains from AI peak at around three tools and then collapse. Workers experiencing what the researchers call "AI brain fry" reported 33% more decision fatigue and 39% more major mistakes than unaffected colleagues. The researchers noted the phenomenon was first observed among high performers: the early adopters who leaned in hardest. Two or three good tools, used well, beats five used carelessly. If your organisation is still debating which ten platforms to approve, the answer might be: pick two and go deep.

### [Three Things] ATMs didn't kill bank tellers. The iPhone did. The distinction matters for AI

David Oks, a researcher at Andreessen Horowitz, [dismantles the comforting argument](https://davidoks.blog/p/why-the-atm-didnt-kill-bank-teller) that automation creates as many jobs as it destroys. ATMs reduced branch costs, which encouraged expansion, which preserved teller employment through 2010. Then mobile banking eliminated the need for branches altogether. Full-time tellers fell from 332,000 to 164,000 by 2022. The lesson for AI: automating tasks within your current structure often creates adjacent roles. Redesigning the structure from scratch eliminates them. The question for your organisation: are you adding AI to existing workflows, or is a competitor building a workflow that doesn't need them?

### [Extra] Only a third of the time AI saves actually reaches the team

Gartner data shows that of 5.4 hours saved per worker through AI tools, only 1.7 hours (31%) translate into improved team outcomes. The largest single block of recovered time, 1.4 hours, goes into additional work that doesn't improve outcomes. Nearly an hour is spent redoing work the AI got wrong. Two thirds of the productivity gain leaks away before anyone benefits. If your organisation is deploying AI tools without redesigning how teams work, you're capturing barely a third of the value.

[Image: How time savings from AI are used — Gartner]

### [Extra] Coding is not software engineering. The confusion is expensive.

Jeremy Howard, a deep learning pioneer who uses AI coding tools daily, draws a distinction most executives miss. Coding, translating a specification into syntax, is a style transfer problem that language models handle well. Software engineering, designing abstractions, decomposing problems, building systems that hold together over time, is a fundamentally different skill that models cannot do. Howard cites Fred Brooks's essay from decades ago, which made the same observation about fourth-generation languages: removing the typing bottleneck does not remove the engineering bottleneck. Companies restructuring around the assumption that AI can do software engineering are conflating two things. His sharpest framing: what matters for any person or team isn't their current output (the intercept) but their rate of improvement (the slope). A little bit of slope makes up for a lot of intercept. Organisations pushing AI to maximise today's output may be destroying the growth rate of the people who'll need to maintain the systems tomorrow.

### [Extra] The shift from co-intelligence to managing AIs

Ethan Mollick, the Wharton professor whose work has appeared here before, argues we've moved from co-intelligence (prompting AI back and forth) to managing AIs (giving agents hours of work and getting results in minutes). His most striking example: a company called StrongDM has two radical rules. "Code must not be written by humans" and "Code must not be reviewed by humans." Each engineer spends roughly $1,000 a day on AI tokens. Coding agents build from human-written roadmaps, testing agents simulate customers, and humans review the finished product but never see the code. Whether or not that model generalises, the direction is clear. The job is shifting from doing the work to directing the things that do it.

### [Extra] AI agents are hiring humans

A platform called RentAHuman has accumulated over 600,000 sign-ups for a marketplace where AI agents autonomously hire human beings to perform tasks machines cannot: delivering physical goods, counting objects in a city, conducting on-the-ground research. The agents browse, post jobs, evaluate candidates and release payment from escrow upon photographic proof of completion. No human intervention on the purchasing side. The gig economy inverted: people as the on-demand labour layer beneath AI clients. Whether that distinction matters to the people taking the jobs is left as an exercise for the reader.

### [Extra] The junior hiring cliff, updated

Edition 3 cited a 14% drop in hiring for workers aged 22 to 25 in AI-exposed occupations. The number has worsened. Stanford Digital Economy Lab data, charted by Politico, now shows a 15.7% decline from 2021 to late 2025. The shape of the curve matters as much as the number: employment held roughly flat through 2023, then fell off a cliff in 2024 and kept falling. This isn't a gradual adjustment. It's a structural break that coincides precisely with the period when agentic AI tools became capable enough to substitute for junior analytical work. Companies aren't announcing junior layoffs. They're quietly not posting the roles. The people most affected will never know the job existed.

[Image: Young workers see a decrease of nearly 16 percent in employment due to AI — Stanford Digital Economy Lab / Politico]

### [Extra] China may be skipping the chatbot phase entirely

Just as China leapfrogged credit cards and went straight to mobile payments, it may be bypassing the "AI as chatbot" paradigm altogether. An open-source AI tool called OpenClaw hit 250,000 GitHub stars in sixty days. Baidu has integrated it into its search app, which has 700 million users. Entrepreneurs are charging 500 yuan (roughly $70) to install it on people's home computers. A startup made $28,000 in ten days selling a one-click installer. Computer repair shops are dispatching what they call "installation personnel," described as operating like plumbers. When a piece of software generates enough demand to support a physical installation economy, the adoption curve is real and deep. Western assumptions about how AI gets adopted may not apply everywhere.

### [Extra] Half of AI code that passes its own tests gets rejected by humans

METR, one of the more rigorous AI evaluation organisations, found that roughly half of code solutions generated by Claude models, solutions that passed automated grading, were subsequently rejected by the actual human project maintainers. Journalist Derek Thompson, reflecting on his own experience using AI coding tools, offered the most useful reframe: AI's real skill is generating plausible candidate solutions that require constant human checking, debugging and rejection. That checking process is effectively its own distinct and skilled job. He compared it to being a casting director working with a promising but unreliable younger actor. Getting the collaboration dynamics right will take a long time to diffuse through the economy, which is grounds for scepticism about predictions of imminent mass displacement.

## Edition #3: Extraction or expansion
*7th March 2026*

### [Three Things] The software market is pricing in the collapse

Since October, software stocks have fallen roughly 30% while the broader technology index has been roughly flat. Salesforce, Adobe, ServiceNow: each down 25 to 30% since last autumn. The market isn't reacting to bad earnings. It's pricing in a structural shift. This week a professor I know built a fully functional membership system in three hours for a non-profit that had been quoted $5,000 a year for commercial software. A CTO of a major corporation told me he's considering switching from Salesforce to a simpler, AI-native competitor. The pattern is the same: the old model of paying for a hundred features to use twelve starts to crack. [Andreessen Horowitz argues](https://a16z.com/good-news-ai-will-eat-application-software/) code was never the moat: distribution, network effects and switching costs are. But switching costs dissolve when AI can extract your data and rebuild the features you actually use.

[Image: Software stocks vs broader tech since October 2025]
*Since October, software stocks have fallen about 30% while broader tech held flat, the market pricing in AI eroding the old licence model.*

### [Three Things] Goldman Sachs can't find a macro productivity effect. But the micro gains are 30%

Goldman titled their latest earnings analysis "AI-nxiety." A record 70% of S&P 500 management teams discussed AI on quarterly calls. Only 1% quantified its impact on earnings. At the economy-wide level, Goldman found no meaningful relationship between AI adoption and productivity. But where firms have actually measured it, the median reported gain is 30%, concentrated right now in customer support and software development. Everywhere else: nothing measurable yet. The gains are real but hyper-localised. The question isn't whether AI works. It's whether your organisation has done the work to capture it. ([Fortune](https://fortune.com/2026/03/03/goldman-earnings-ai-anxiety-no-meaningful-impact-productivity-economy-30-percent-in-2-areas/))

[Image: Goldman Sachs AI expectations vs job listings data]
*Goldman finds no economy-wide productivity effect yet, but where firms actually measure it the median gain is 30%.*

### [Three Things] AI hasn't just automated legal work. It's retroactively reclassified it

[A lawyer who built his practice around language models](https://www.artificiallawyer.com/2026/03/02/lawyer-uses-claude-skills-legal-world-loses-it/) reports that a well-instructed general-purpose model outperforms the expensive, narrowly trained legal AI products that have raised hundreds of millions in venture funding. AI is good at tireless issue-spotting, finding contradictions, fixing errors, and producing a structured first draft for human review. It is not good at fine-tuned business judgment, relationship sensitivity, or getting from 85% to 100% where every word and comma matters. But here's the uncomfortable part. Tasks that were billed at premium hourly rates for decades (formatting, precedent research, copy-pasting between documents) have been revealed as procedural, not cognitive. The professional mystique that allowed them to be charged as expertise has been stripped away. AI is acting as a truth serum for knowledge work: forcing an honest reckoning about which tasks were genuinely skilled and which were merely time-consuming and opaque.

### [Extra] The hiring cliff for juniors, in one chart

The essay mentions a 14% drop in hiring for workers aged 22 to 25 in AI-exposed occupations. This chart shows the full time series using a difference-in-differences approach. Junior hires in exposed occupations fell off a cliff after ChatGPT's release in late 2022, while exits held steady. The gap keeps widening. Companies aren't firing juniors. They're just not bringing new ones in.

[Image: Hires vs exits for junior workers in AI-exposed occupations]

### [Extra] US tech employment growth has gone negative

Year-on-year tech employment growth turned negative in 2024 and hasn't recovered. The tech sector is shrinking its workforce for the first time since the post-2008 recovery. Combined with the junior hiring data above, a pattern emerges: the contraction is real, it's happening now, and it's concentrated at the entry level.

[Image: US tech employment year-on-year change]

### [Extra] The work budget is orders of magnitude larger than the software budget

Julien Bek at Sequoia Capital argues that the next category-defining AI company won't sell tools to professionals. It will sell completed work directly to buyers. His distinction between "copilots" (AI as a tool for professionals) and "autopilots" (AI delivering the outcome) reframes the entire market. For every dollar spent on software, six are spent on services. The smartest entry point? Replace outsourced work first. The budget already exists, the buyer already accepts external delivery, and there's no internal team whose jobs are visibly threatened. Once embedded, expand inward.

### [Extra] You would not believe how many shortcuts everyone else is taking

Ezra Klein wrote a commencement address called "Just Do the Work" about discovering, as a young journalist, that almost nobody was actually reading Congressional Budget Office reports. Documents that are neither complex nor long. By reading what his peers skipped, he got ahead. Not exceptional talent. Just diligence. Economist Paul Novosad adds the contemporary twist: this is "more true than ever now, when more people are shirking and AI lets you do 10x if you try." The gap between the diligent and the lazy is widening, not narrowing.

## Edition #2: The hundred small things
*28th February 2026*

### [Three Things] A Fiction Worth Reading. Honestly

Citrini Research published a scenario memo set in June 2028, looking back at a crisis that hasn't happened yet. The central argument: when AI makes expertise cheap to produce, clients stop buying it. I think this applies equally to strategy decks, competitive analysis and market research. Work that used to justify five-figure invoices starts getting done in-house by someone with a subscription and thirty minutes. Clients don't fire their agencies. They just renegotiate, armed with a clearer sense of what the work actually costs to produce. The sharpest line I've read this year: "We had overestimated the value of human relationships. Turns out that a lot of what people called relationships was simply friction with a friendly face." Worth reading as a stress test for anyone who sells expertise for a living. ([Citrini Research](https://www.citriniresearch.com/p/2028gic))

### [Three Things] The floor collapsed under junior roles

Jack Dorsey cut Block, his payments company, from 10,000 people to [under 6,000](https://techcrunch.com/2026/02/26/jack-dorsey-block-layoffs-4000-halved-employees-your-company-is-next/). Not because the business was struggling. Gross profit grew 24%. The stock jumped 20%. Dorsey was unusually direct about why: "the intelligence tools we're creating and using, paired with smaller and flatter teams, fundamentally change what it means to build and run a company." Hours later, several founders from Y Combinator, Silicon Valley's most influential startup programme, told investor Jeff Feng they're planning to [eliminate all engineers below senior level](https://x.com/jeffdfeng/status/2027194038487523329?s=20). The pattern: senior people who can steer AI are becoming more valuable. Junior people whose work AI can approximate are harder to justify. The market confirmed it instantly, valuing each eliminated Block role at roughly $1.5 million in added enterprise value. If you run a team, count how many people do work that a senior person with good AI tools could now do themselves. That's the number your board will eventually ask about. I believe junior people with AI are more valuable than ever. But articulating that in terms that survive a headcount review is a challenge we all have to address.

### [Three Things] The model matters less. The application matters more

Google's Gemini 3.1 Pro scored 77.1% on **ARC-AGI-2**, a reasoning benchmark designed to test abstract problem-solving, more than double its predecessor's score three months earlier, while holding API prices flat. ([Google](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-pro/)) Meanwhile, composite evaluations show OpenAI, Anthropic and Google clustered tightly at the top. ([Artificial Analysis](https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index)) Six months ago, picking the right AI model felt very important. Model choice now matters less. Models are converging so fast that any advantage evaporates before you've finished onboarding. What hasn't converged is the application layer: ChatGPT, Claude, Claude Code each wrap similar intelligence in very different interfaces and workflows. Pick based on the problem and the workflow. The model underneath will be fine.

### [Extra] The marginal cost of arguing is going to zero

UK employment lawyers report workplace grievances that once fit in a single email ballooning into 30-page documents, complete with fabricated legal precedents and citations to laws from the wrong country. ([Personnel Today](https://www.personneltoday.com/hr/employment-lawyers-voice-ai-fears-on-tribunal-claims/)) Creation cost: near zero. Response cost: unchanged. Ministry of Justice figures show new employment tribunal receipts rose 33% year-on-year in the quarter to September. ([GOV.UK](https://www.gov.uk/government/statistics/tribunals-statistics-quarterly-july-to-september-2025/tribunal-statistics-quarterly-july-to-september-2025))

### [Extra] Your prompt is the ceiling

Anthropic's latest Economic Index analysed over a million Claude conversations and found a near-perfect correlation (r > 0.92) between the sophistication of human prompts and the sophistication of AI responses. The more nuanced and structured the input, the more the model rises to meet it. The bottleneck isn't the model. It's the human. Which is, in its own way, reassuring. ([Anthropic Research](https://www.anthropic.com/research/anthropic-economic-index-january-2026-report))

### [Extra] One blog post. One hour. Billions gone. Again.

Anthropic published a blog post introducing Claude Code Security on a Friday afternoon. Within an hour, cybersecurity stocks cratered: CrowdStrike fell 8%, Cloudflare 8%, Okta over 9%. The tool itself is a modest research preview. But a single blog post from an AI company erased billions in market value from established incumbents. ([Barron's](https://www.barrons.com/articles/crowdstrike-stock-price-cybersecurity-zscaler-3efb4a93)) The same dynamic hit legal tech stocks when Anthropic announced legal plugins for Claude Cowork a couple of weeks earlier. ([Sherwood News](https://sherwood.news/markets/anthropics-legal-plugins-for-claude-cowork-prompt-rush-out-of-legal-software/)) That's a new kind of leverage.

### [Extra] The trust signals your organisation depends on are dissolving

Here's a problem that connects directly to those poets in Sakaiminato. A thoughtful email from a director now carries the same weight as an AI-generated memo, because the reader can't tell the difference. The cues that used to signal competence (a well-crafted message, a polished document, a detailed analysis produced under time pressure) are now producible by anyone in minutes. This isn't a quality problem. It's a trust architecture problem. We need to distinguish between "produced this" and "shaped this." The senryu competition couldn't tell the difference. That's why it died.

## Edition #1: The wonder and the weight
*22nd February 2026*

### [Three Things] The technical barrier is gone. Domain expertise is what matters now

Ethan Mollick, a Wharton professor, gave executive MBA students four days, three AI tools and a brief: build a company from scratch. They did. Working prototypes, financial models, market research, competitive positioning. Most had never written a line of code. The students who got furthest weren't the most technical. They were the ones who understood their industry. If you've spent twenty years in your field and haven't tried building something with these tools, you're sitting on your biggest advantage. Open Claude Code and have a play. Let me know what you build! ([Mollick's full account](https://www.oneusefulthing.org/p/management-as-ai-superpower))

### [Three Things] Most firms use AI. Almost none can measure the difference

An [NBER study](https://www.theregister.com/2026/02/18/ai_productivity_survey/) surveyed nearly 6,000 executives across four countries. Sixty-nine percent of firms now use some form of AI. The average productivity gain over three years? 0.29%. The economists draw an explicit parallel to Robert Solow's 1987 observation that computers were everywhere except in the productivity statistics. The explanation, proved right over the following decade: firms had to fundamentally reorganise before the technology translated into measurable gains. That lag wasn't months. It was years. The same pattern is playing out with AI right now. But it just has to be much quicker, right? Read on ...

### [Three Things] Freelancers are disappearing. The data is in

Ramp analysed real corporate spending data. The share of business spend going to freelance marketplaces like Upwork and Fiverr fell from 0.66% to 0.14% in three years. Firms most exposed to AI substituted at roughly $1 in reduced freelance spend for $0.03 in AI spend. A 97% cost reduction. The gig economy may be automation's first casualty. ([Ramp Economics Lab](https://ramp.com/velocity/ai-labor-market-impact-freelancers))

### [Extra] Apple chose Google over itself

Apple partnered with Google to power its AI features, paying a reported billion dollars a year for Gemini. The world's most valuable technology company looked at its own AI and decided someone else's was better. The strategic question isn't whether to build AI capability. It's which partner to choose. (By the way, the answer I'd recommend is Claude!)

### [Extra] One blog post. One hour. Billions gone.

Anthropic published a blog post introducing Claude Code Security on Friday. Within an hour, cybersecurity stocks cratered: CrowdStrike fell 8%, Cloudflare 8%, Okta over 9%. The tool itself is a modest research preview. But a single blog post from an AI company erased billions in market value from established incumbents. The same happened to legal tech stocks when it announced legal plugins a couple of weeks ago. That's a new kind of leverage.

### [Extra] The marginal cost of arguing is going to zero

UK employment lawyers are seeing workplace grievances that once fit in a single email ballooning into 30-page documents, complete with made-up legal precedents and citations to laws from the wrong country. Creation cost: near zero. Response cost: unchanged. New UK employment cases rose 33% in three months.

### [Extra] Your prompt is the ceiling

Anthropic's latest Economic Index analysed over a million Claude conversations and found a near-perfect correlation (r > 0.92) between prompt sophistication and response quality. Give a vague prompt, get a vague response. Give one rich in nuance and structured constraints, and the model meets you there. The bottleneck isn't the model. It's the human.

### [Extra] Watch what McKinsey does with its own workforce, not what it advises clients to do

McKinsey calls it "25 squared." The plan: grow client-facing roles by 25% while cutting back-office roles by 25%, using AI to rebalance a $20 billion firm. This isn't productivity improvement. This is structural transformation. Ask yourself if there are parts of your business that are ripe for radical change. McKinsey already has.

---