LAB
Khelion Lab

What the evidence actually says.

Before we publish results of our own, we publish our reading of everyone else's — sourced, dated, and with the weak sources called weak. Each note below states its claim, links its primary source, and says what we would not conclude from it.

01 — The gap between belief and measurement

Around 90% of chief executives expect AI agents to deliver measurable return in 2026 (BCG, 2,360 executives). In the same period, 39% of organisations report any EBIT impact from AI at all, and most of those below 5% (McKinsey, 1,993 respondents). Both numbers are probably honest. They are measuring different things: one measures intention, the other measures the P&L.

What we would not conclude: that AI does not work. Only that the distance between an executive's expectation and a finance department's ledger is currently very large, and that anyone selling into that gap should expect to be asked for proof.

StrengthBoth are executive surveys. Declarative, large samples, primary URLs.

02 — Why agent projects get cancelled

Gartner predicts more than 40% of agentic AI projects will be cancelled by the end of 2027, naming three causes: escalating costs, unclear business value, and inadequate risk controls. Note the order. Two of the three are not technical problems, and the third is a governance problem.

What we would not conclude: that this is measured. It is an analyst prediction, and it should always be quoted as one. We cite it because it names the failure modes we build against, not because a forecast is evidence.

StrengthPrediction, not measurement. Primary press release.

03 — The one number in this field that is properly measured

A controlled study of 5,179 customer-support agents found productivity up 14% on average, and 34% for the least experienced workers — the strongest people gained almost nothing. It is a quasi-experiment with a real control, published in a peer-reviewed journal. It is also about an assistant, not an autonomous agent, and about one function.

What we would not conclude: that agents raise productivity 14% anywhere. The distribution is the finding: this technology lifts the floor much more than the ceiling. That has consequences for where you deploy it first.

SourceBrynjolfsson, Li & Raymond, NBER w31161 · published in QJE, April 2025
StrengthThe best evidence in this note set. Peer-reviewed, controlled.

04 — A widely cited number we refuse to use

You will see “95% of AI pilots fail” everywhere. It comes from an MIT Project NANDA report built on 52 interviews and 153 survey responses, describing itself as “directionally accurate” and not peer-reviewed. It is also about generative AI in general, not about agents.

Why this matters to you: any vendor quoting it as an agent failure rate either has not read it or is counting on you not having. We would rather lose the dramatic slide than cite a number that does not support the sentence it is placed in.

StrengthWeak methodology, honestly disclosed by its own authors.

05 — What a regulator just made mandatory

Commission Regulation (EU) 2026/247 of 2 February 2026 replaces Annex II of Regulation 300/2008. From 1 January 2028, every Member State must run a process for reporting, classifying, processing, storing, protecting, analysing and aggregating aviation security occurrences — with a confidential channel and a common classification. Separately, the 2026 digital omnibus moved the AI Act's Annex III high-risk obligations to 2 December 2027.

What we would not conclude: that either date is a sales argument on its own. Both create documentation work that has to exist whether or not anyone buys anything from us.

StrengthPrimary legal texts. Verified at source.

06 — The paperwork burden, measured by the industry itself

Surveying 138 device manufacturers and 73 diagnostics makers, MedTech Europe found that 70% need up to four months a year simply to update their post-market surveillance reports, and that 86% of large manufacturers and 91% of SMEs cannot recruit the qualified regulatory staff they need.

Caveat we will not hide: this is a survey by an industry association that is lobbying to soften the same regulation. The direction is credible, the framing is interested. We use it because the labour shortage it describes is corroborated by every job posting in the field.

StrengthLarge sample, interested party. Both true at once.

Coming from our own work

These are ours to produce, and they will appear here when the data is real — not before.

Forthcoming

The warrant format, published in full

Schema, blank template and the runtime checks that enforce it. Free to copy, including by people who will never buy anything from us.

Open template
Forthcoming

What 72 hours of unsupervised agent actually produces

Autonomous production is not autonomous revenue. We measure the gap on our own operation, and publish the cost side too.

Method & results

Notes are published under an open licence at khelion.com/lab. If a source here is wrong, tell us and we will correct it in public.