All field notes

Independent analysis

When AI Gets Cheaper, Verification Becomes the Factory Floor

Falling model costs do not automatically lower the cost of completed work. In many AI workflows, scarce capacity to test, reconcile, review and correct outputs becomes the binding production constraint.

MP
Max PerfiljevFounder & CEO, AES · Architect of Autonomous Organizations
Read in Russian

The most expensive part of an AI workflow may soon be the person—or system—waiting at the end of it.

Inference is getting cheaper and generation is getting faster. That changes the economics of work, but not in the simple way token-price comparisons suggest. If an AI system produces drafts, analyses, code changes or recommendations faster than an organization can establish that they are acceptable, the scarce resource is no longer generation. It is verification.

This is not an argument that every AI output needs a human reviewer. Some work can be checked deterministically, reconciled against a source of record, or sampled statistically where the cost of an error is bounded. The argument is narrower and more useful: in workflows where acceptance capacity cannot expand as quickly as generation, verification becomes the binding production constraint. More model output then creates queues, retries, rework and delay rather than proportionate business value.

A cheap draft is not a cheap completed task

OpenAI made this point explicitly in its July 17 scorecard for the AI age. Its proposed measure is the cost per successful task, including model costs, employee time, human review, retries and rework—and counting only work that meets the required quality bar. This is a meaningful correction to AI accounting. A token is an input. A generated artifact is intermediate work. Neither is an outcome unless it is accepted for its intended use.

Anthropic’s January 2026 Economic Index reaches a related conclusion from a different direction. It treats task success as economically relevant both to whether a task can be automated and to how many attempts automation requires. Its modeled productivity effects fall sharply when automated occupational tasks depend on slower, complementary tasks that remain unautomated. Under one complementarity assumption, its illustrative success-adjusted estimates decline to 0.8 percentage points for Claude.ai use and 0.6 points for API use. These are model-dependent estimates, not observed economy-wide productivity results. But the mechanism is important: accelerating one step does not remove the work around it.

METR supplies a practical warning against substituting confidence for measurement. In a randomized study of 246 tasks performed by 16 experienced open-source developers in an early-2025 setting, AI-tool access increased measured completion time by 19 percent. Participants nevertheless believed the tools had made them faster. The finding does not show that AI generally slows software development. It concerns a small, specific group working in mature repositories with the tools available at that time. It does show that apparent local speed can diverge from end-to-end completion time.

Generation raises the arrival rate; review has a fixed service rate

The operational mechanism is ordinary queueing. Little’s result relates average work in process, arrival rate and average time in a stable system. If an AI system increases the arrival rate of candidate work into a review stage while reviewer capacity remains fixed, something must give: the review queue grows, cycle time rises, or incoming work is constrained. Often all three happen. A team may celebrate a tenfold increase in generated pull requests or customer-response drafts while quietly creating a backlog of work no one has capacity to trust.

The same constraint appears in the logic behind Amdahl’s 1967 result: speeding one portion of a process leaves overall improvement bounded by the portion that remains unaccelerated. Suppose generation had taken half of a workflow’s elapsed time and verification, correction and release took the other half. Making generation nearly instantaneous cannot make the whole workflow nearly instantaneous. And if cheaper generation causes more failed attempts or more candidate artifacts to inspect, the unaccelerated portion can grow.

The production function should be designed around the scarce validator, not the abundant generator.

Measure verification-adjusted throughput

The practical metric is verification-adjusted throughput: accepted outcomes divided by the full cost of delivering them. The denominator should include model use, reviewer time, automated test and reconciliation capacity, retries, rework, waiting time and any delay that materially affects the workflow. The exact financial treatment will differ by business. The discipline should not. Count completed, accepted work; then account for the resources required to get there.

This measure changes the questions operators ask. Instead of asking how many cases an agent processed, ask how many cases reached the required standard without creating downstream repair work. Instead of asking whether a model reduced drafting time, ask whether the total time from request to accepted output fell. Instead of buying more inference to increase parallelism, locate the acceptance stage and measure its capacity, queue and failure modes.

Three verification regimes

Workflows should be classified by how their outputs can be accepted. The classification is an economic design choice before it is a tooling choice.

  1. Machine-verifiable work has executable acceptance tests. Examples include constrained transformations, calculations with defined inputs, schema validation, deterministic business rules and reconciliations against authoritative records. Here, investment should focus on improving tests, constraints and reference data so acceptance can scale with generation.
  2. Statistically reviewable work permits sampling because error costs are bounded and sampling provides enough confidence for the use case. The essential work is to define the sampling method, the tolerated error rate and the response when sampled work fails—not simply to call sampling “automation.”
  3. Expert-verifiable work depends on scarce professional judgment. Legal interpretation, complex code review, high-context customer decisions and novel analysis may fit this regime. AI can still help prepare and narrow the work, but throughput remains capped until the demand for expert verification falls or expert capacity rises.

The point is not to force every process into the first category. It is to avoid pretending that expert-verifiable work has machine-verifiable economics. An organization that sends five times as many ambiguous cases to the same expert team has not increased productive capacity. It has increased the arrival rate at its bottleneck.

Reduce verification demand before scaling generation

The strongest response is usually not hiring reviewers to chase a larger output stream. It is redesigning the workflow so fewer outputs need costly scrutiny. Add deterministic tests where the acceptance condition can be made explicit. Use structured outputs to reduce ambiguity. Reconcile proposed actions against reference data and systems of record. Narrow task definitions so the model is not asked to infer broad intent from incomplete context. Separate retrieval, transformation and decision steps when each can be checked differently. These investments turn some judgment-heavy review into lower-cost validation; they also make failure easier to locate and repair.

Only then should generation be scaled. A useful operating rule is simple: increase output supply only when acceptance capacity increases with it, or when the workflow has a deliberate way to limit arrivals. This is not an anti-AI position. It is a production-design position. Cheaper intelligence changes what is abundant. It does not repeal the economics of bottlenecks.

For founders and operators, the next AI budget review should therefore begin at the far end of the workflow. Identify who or what validates the result, what makes an outcome acceptable, how often work must be retried, and how long accepted work waits behind unfinished review. The best AI system will not be the one that produces the most. It will be the one that increases accepted outcomes without overwhelming the capacity to know they are right.

BUILD WITH AES

Turn architecture into an operating company.

AES connects strategy, tasks, organizational memory, knowledge, agents, people and approvals in one execution environment.