The argument in brief· 3–5 min read· full paper 26 pp

Speed Is the Moat

A slogan, and what happens when somebody finally tests it.

Speed is not a moat. The dispersion in inference speed is real and larger than most people assume — but a lead inside it has a measured half-life of one to two months, which is shorter than the procurement cycle meant to capture it.

01The claim, and its half-life

The slogan has a source, and the source has an interest:

“Speed is the moat.” Anush Elangovan, VP GPU Software, AMD

None of which makes it wrong. Interested parties are often correct, and a claim’s origin is not an argument against it. So the paper tests it instead of dismissing it — and the test is what a lead is worth once you measure how long it survives.

“The half-life of a software-derived speed lead is roughly one to two months. Most procurement cycles are longer than that.” The central finding

That is the number that decides the question. A moat is a structural advantage that persists while a competitor tries to close it. An advantage that halves every six to eight weeks, on unchanged silicon, from software alone, is a lead — a real one, worth having, worth competing for. It is not a moat, and the distinction is not pedantry: it determines whether you build a strategy on it or a quarter.

There is a second reason the metric is softer than it looks. Much of what gets reported as a hardware or engineering result is a choice somebody made about where to sit on a cost curve:

“Where a provider sets their batch size is a business decision as much as a technical one.” On the economics of inference serving

Which reframes the whole exercise. A speed programme run on providers who can reprice their own latency at will is not measuring a durable capability. It is sampling a business decision, on the day you sampled it.

02The interactive case runs backwards

The intuitive defence of speed is the user: people want answers now, so faster wins. The one controlled experiment on interactive latency found close to the opposite.

“Participants who waited two seconds rated the model’s output less thoughtful and less useful than participants who waited nine to twenty.” Controlled experiment, n=240, CHI 2026

Delay, in that setting, read as deliberation. That is one experiment on one band of latency with one family of tasks, and the paper says so rather than generalising it into a law — but it is the only controlled evidence pointing either way, and it does not point where the slogan assumes.

What survives is narrower and more useful than the claim it replaces. Latency is a priced contract term, not an engineering win: something you specify, pay for and hold a supplier to, in the same register as uptime. Speed pays where it is contractual or where it changes what the system can attempt at all. It does not pay as a standing advantage, because nothing that reprices on a two-month cadence can be one.

What this argument is not

The half-life is measured over one window, on public benchmark data whose method the paper reproduces and criticises in the same breath. Vendor-published throughput figures are used where nothing independent exists, and are labelled as such.