All field notes

Research operations

OpenAI’s Research Team Now Runs More AI Time Than Human Time

OpenAI reports 3.1 eight-hour AI-runtime workdays for each human workday in its research organization. The figure is not an FTE or productivity measure, but it makes inference capacity a visible input to research operations.

MP
Max PerfiljevFounder & CEO, AES · Architect of Autonomous Organizations
Read in Russian

OpenAI has published an unusually concrete view of how AI is being used inside its own research organization. The headline number is not a benchmark score: by mid-August, the organization logged 3.1 eight-hour AI-runtime workdays for every human workday.

That ratio does not mean OpenAI has 3.1 AI employees for every researcher. It does not establish a comparable gain in productivity, completed research, or labor substitution. OpenAI is explicit that the measure is runtime, not output quality or full-time-equivalent capacity. But it does mark a material operating change. In frontier-model research, inference capacity is becoming a measurable production input alongside human time.

The September 6 disclosure calls its automated system a “research intern”: a system that completes well-defined research tasks under human direction, including tasks that might take a skilled researcher several days. The label is deliberately narrower than autonomous scientist. The useful lesson is not the label, however. It is the measurement discipline behind it.

What changed: research work is being counted in two clocks

Most organizations still estimate research capacity primarily through headcount, hiring plans and utilization. OpenAI’s data suggests that this view is no longer sufficient where models can run substantial task sequences. A team’s effective operating envelope also depends on how much model runtime it can allocate, how many jobs can run concurrently, and which tasks are structured well enough to delegate.

OpenAI reports that its median research-organization user consumed more than $600 per day of inference when priced at API rates by mid-August. This is not OpenAI’s internal marginal cost; it is a valuation proxy. Even so, it provides a useful operational signal. Model use has moved beyond occasional assistance and into a daily resource allocation that can be observed, budgeted and constrained.

The organization also reports that experiments per active experimenter reached their highest level since it began tracking the measure in January 2025. OpenAI says the rise correlated with Codex adoption. It does not claim Codex caused it: available compute increased substantially over the same period, so the contribution of each factor cannot be cleanly separated.

That caveat matters. More experiments are not automatically more scientific progress. Code counts and experiment counts are difficult to translate into research outcomes, OpenAI says. Still, the operational pattern is clear: when a model can construct code, support technical work and monitor processes, more of a researcher’s day can be spent initiating, comparing and selecting experiments rather than building every component manually.

What did not change: people retain the research judgment

The disclosure does not describe an independent research system. High-level planning remains a minimal share of agent output tokens. OpenAI says people still set research priorities, evaluate which results are worth pursuing, and decide whether systems should be scaled, paused or deployed.

The task data points in the same direction. Measured task success increased from January through July across several estimated-duration categories. Yet, during the preceding six months, more than half of successful tasks estimated at four to eight hours involved at least one human intervention. Long-running work is advancing, but successful completion often remains a joint process rather than unattended execution.

The evidence also has clear methodological limits. It is first-party and preliminary. Some tool use is excluded from coverage. Tasks with uncertain outcomes were omitted from the success analysis. And OpenAI cautions against treating code and experiment counts as a direct measure of research progress. These limits should narrow the conclusion, not erase the signal.

The operating consequence: plan capacity as a portfolio of people, runtime and task shape

For research leaders and technical operators, the practical implication is straightforward: headcount alone is becoming an incomplete capacity plan. A serious plan for AI-assisted research needs at least three connected budgets.

  • Human attention: the time needed to set priorities, formulate work, inspect results, intervene in failures and make consequential choices.
  • Inference and concurrency: the runtime available for delegated work, its allocation across teams, and the number of tasks that can progress at once.
  • Task shape: the share of work that is well-defined enough to hand to a model, with usable inputs, evaluable outputs and an appropriate intervention path.

These are not interchangeable resources. Adding model runtime will not resolve an ambiguous research question. Adding researchers will not necessarily increase throughput if technical construction is already the bottleneck. And a large inference allocation can be wasted if tasks have no clear completion condition or if their outputs cannot be evaluated in time.

This is why cost accounting for model use should mature beyond a monthly platform bill. Teams need to see where runtime goes: exploratory coding, experiment construction, technical support, monitoring, evaluation and rework. They should also distinguish API-price valuation from actual internal cost, just as OpenAI does. The objective is not to make every token defensible in isolation. It is to understand whether the allocation is expanding the organization’s useful research options.

Measure the joint system without inventing an equivalence

The tempting simplification is to convert AI runtime into employee equivalents. OpenAI’s disclosure is a reason not to do that. An eight-hour model runtime is a quantity of computation. A human workday contains problem framing, domain judgment, coordination, accountability and decisions about what should happen next. The two clocks interact; they do not represent the same unit of output.

A better operating review asks four questions. Which task classes are actually being delegated? How often do people intervene before a task succeeds? What resource limits prevent more concurrent work? And, after outputs are evaluated, which experiments or technical results changed the team’s next decision?

Those questions preserve the important distinction in OpenAI’s report. The system can increase the volume and duration of supervised technical work without taking over research direction. That is already enough to change how a research organization allocates capacity.

A bounded but important disclosure

OpenAI has not shown autonomous research, recursive self-improvement, or a general productivity multiplier for enterprises. It has shown intensive internal use of models under human direction, rising capability on measured tasks, and a growing volume of AI runtime in its own research operations.

That is the news. Research capacity at the frontier is starting to be expressed in two measurable inputs: human labor time and model runtime. Organizations that use AI for sustained technical work should begin planning for both—while keeping the difference between computation, successful work and scientific progress intact.

Source: OpenAI, “Research acceleration: a view inside OpenAI,” September 6, 2026. https://openai.com/index/research-acceleration-view-inside-openai/

BUILD WITH AES

Turn architecture into an operating company.

AES connects strategy, tasks, organizational memory, knowledge, agents, people and approvals in one execution environment.