Reliability, oversight and lock-in
Agent reliability is improving far slower than raw capability, which makes human oversight a permanent cost rather than a temporary one. That cost doesn't stay in engineering — it reshapes the business case, the operating model and, eventually, the contract. This theme follows the chain from the research bench to the terms you sign
Reports
Brief: Enterprise AI vendor Advisory (S)
This brief is about the evidence underneath enterprise AI investment cases: how the adoption, revenue and reliability numbers in front of a CIO are produced, and what they hold up to once you know. I should be clear about who I'm picturing, because it shapes the whole thing. You're a CIO past the "should we try this" stage — several agent pilots run, some already in production and wired into live workflows. The board pack in front of you quotes figures on how many peers run agents, how much revenue vendors book from them, how much they save. This brief is about whether those figures mean what the slide implies. Three claims follow, each resting on the last. That the headline adoption and revenue figures are definitional artefacts, not measurements: change the population or the survey question and the same data moves by a factor of three. That a separate research literature, one your vendors don't cite, says the automation business case was wrong on the day it was approved, because the arithmetic assumed a reliability the technology doesn't deliver. And that both conclusions change the terms of the contract you're about to sign, in a direction nobody is currently negotiating. Every figure carries three things: where it came from, how much weight it bears, and who benefits if you believe it. Where a widely-quoted number couldn't be traced to a source, it's flagged as unverified rather than used. There's a fourth finding, and it surprised me enough to earn its own section: why the fix for the first problem quietly becomes the largest switching cost in your stack
£199
Brief: Enterprise AI vendor Advisory (T)
This brief is about the evidence underneath enterprise AI investment cases: how the adoption, revenue and reliability numbers in front of a CIO are produced, and what they hold up to once you know. I should be clear about who I'm picturing, because it shapes the whole thing. You're a CIO past the "should we try this" stage — several agent pilots run, some already in production and wired into live workflows. The board pack in front of you quotes figures on how many peers run agents, how much revenue vendors book from them, how much they save. This brief is about whether those figures mean what the slide implies. Three claims follow, each resting on the last. That the headline adoption and revenue figures are definitional artefacts, not measurements: change the population or the survey question and the same data moves by a factor of three. That a separate research literature, one your vendors don't cite, says the automation business case was wrong on the day it was approved, because the arithmetic assumed a reliability the technology doesn't deliver. And that both conclusions change the terms of the contract you're about to sign, in a direction nobody is currently negotiating. Every figure carries three things: where it came from, how much weight it bears, and who benefits if you believe it. Where a widely-quoted number couldn't be traced to a source, it's flagged as unverified rather than used. There's a fourth finding, and it surprised me enough to earn its own section: why the fix for the first problem quietly becomes the largest switching cost in your stack
£499
