The completed task is the right economic unit

In The Work Now Within Reach , published September 8, OpenAI connects model capability, product distribution and computing infrastructure to one business proposition: better systems can make previously uneconomic work practical. The essay is by Sarah Friar , OpenAI's chief financial officer.

The most useful part of that proposition is its unit of measurement. A model request is not the same thing as a completed task. The request may be quick and inexpensive while the job remains costly because someone must retry the prompt, check the result, reconcile it with source material or repair it before use. The accepted result is where capability, reliability, time and price finally meet.

That distinction turns a broad corporate strategy into a practical evaluation method. Define what completion means, count every attempt required to reach it and include the human work around the model. The result is not a universal ranking of providers. It is a clearer measure of whether a particular system makes a particular workflow more economical.

Define acceptance before counting savings

Suppose a team asks an AI system to prepare a short document from an approved set of sources. A weak definition of completion is that a file exists. A useful definition might require accurate citations, the requested structure, no unsupported claims and acceptance by the person responsible for publication.

Those definitions create different success counts. If ten files are generated but only six meet the agreed criteria, the denominator is six completed jobs, not ten. The four rejected outputs consumed model usage, elapsed time and review effort. Excluding them would make the system look cheaper without making the workflow more productive.

A fair comparison therefore begins with a fixed acceptance rule and a representative set of tasks. Record attempts, accepted outputs, human review time, repairs and cases handed to a fallback process. The method can start in a spreadsheet. Its rigor comes from keeping the definition stable while alternatives are compared.

The broader idea of total cost of ownership supplies useful context: the visible purchase price is only one part of the overall cost. For AI work, the categories should match the actual job. A casual first draft, a customer-facing answer and a document supporting a consequential decision require different levels of review. Treating them as one generic task would hide the very costs the analysis is meant to reveal.

Three cost boundaries need separate ledgers

OpenAI's argument touches three related but distinct economic boundaries. The first is the provider's cost to serve an interaction. The second is the customer's price for access or usage. The third is the organization's total cost of obtaining an acceptable result.

An improvement at one boundary does not automatically prove an improvement at the others. Lower serving cost may support lower prices, higher capacity, better margins or a different service tier. A lower customer price may still accompany more retries. A more expensive model may reduce total workflow cost if it reaches acceptance more reliably and demands less correction.

The organizational ledger includes work outside the AI invoice. For the document example, that can include preparing the source set, explaining requirements, checking citations, correcting language and escalating exceptions. Some of those activities may also exist in the previous workflow, so the comparison must apply the same accounting rules to both paths.

Time needs two columns as well. Elapsed time measures how long the process takes. Human effort measures how much active attention it consumes. An unattended hour can be operationally cheaper than ten minutes of continuous supervision, unless that hour blocks a deadline or another dependent process. Combining those measures too early can produce an apparently precise number that answers the wrong question.

Hardware efficiency is an input, not the outcome

OpenAI's essay points to its earlier Jalapeño inference-chip results . The August 25 report describes InferenceX tests across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T. OpenAI reports 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency, normalized against published accelerator power. It presents these as first-party results and says deployment is planned to begin by year-end.

Those measurements address the tested systems and workloads. They do not measure every customer's completed task, guarantee a subscription-price change or establish a reduction in an organization's electricity use. Each claim belongs to a different boundary and would require different evidence.

For an agent workflow, the bridge from hardware to economics is concrete but not automatic. Faster inference can shorten response time. Greater efficiency can increase available capacity. The business result still depends on whether the system produces acceptable work, how often it must retry and how much supervision it needs.

That is why a throughput or latency figure cannot finish the comparison. It explains one component of the delivery stack. The completed-task measure tests whether improvements across the stack combine into something useful for the buyer. The same discipline applies to vendor benchmarks, competitor results and independent tests: preserve the scope of the measurement instead of turning one strong number into a claim about the whole workflow.

A small comparison exposes the trade-off

Return to the hypothetical document task. Imagine two systems judged against the same acceptance criteria. System A costs 20 pence per attempt and averages three attempts for an accepted result. System B costs 45 pence per attempt and averages one and a quarter attempts. Before human work, the expected model charge per completed task is 60 pence for A and about 56 pence for B.

That simple calculation reverses the apparent price ranking, but it is still incomplete. If A's corrections take two minutes and B's take five, the labor result may reverse again. If one system fails on a common document type and needs manual routing, that exception rate belongs in the ledger. If quality varies sharply between easy and difficult tasks, averages should be split by task class instead of allowed to blur the difference.

The comparison should retain failures rather than silently removing them. A workflow that performs well on straightforward cases but regularly stalls on an important exception needs a defined fallback. The time spent detecting, routing and resolving those cases is part of the cost of completion.

Our analysis of Nvidia's Personal AI Router examines a related distinction between routing a request and the infrastructure that serves the model. Here, the corresponding lesson is that an improved component should not stand in for the performance of the complete workflow.

OpenAI's essay sets out a coherent full-stack strategy. Its completed-task framing is also the right place for customers to test that strategy. Fix the acceptance criteria, compare representative work, retain failures and keep provider cost, purchase price and organizational effort separate. That method shows when lower latency, better models or cheaper inference actually become a less expensive accepted result.