Microsoft proposes “useful yield” as the real measure of AI infrastructure
Microsoft is arguing that the next phase of AI infrastructure should be measured by “useful yield”: how effectively silicon, memory, networking, power and software combine to produce useful work.
Rani Borkar, president of Azure Hardware Systems and Infrastructure, presented the idea at SEMICON Taiwan 2026. The company points to measures such as tokens per dollar, tokens per watt, throughput and latency instead of relying only on a processor’s theoretical peak performance.
Why system-wide efficiency matters
An accelerator can look impressive in isolation while spending time waiting for memory, network transfers or software scheduling. Cooling, power delivery and failure rates also affect the cost and environmental footprint of an AI service. Microsoft’s approach treats the data centre as one co-designed system, from chips and racks to compilers and workloads.
The concept is particularly relevant as model use shifts from training runs to continuous inference. A small gain repeated across billions of queries can change both operating cost and electricity demand.
What buyers should ask for
“Useful” depends on the workload. A system optimised for short chatbot replies may not lead on scientific inference, image generation or long-context analysis. Comparisons should therefore state model quality, response length, latency target, power boundary, hardware utilisation and error rate.
Useful yield is Microsoft’s framing rather than a universal industry standard. It can improve discussion only if vendors publish enough methodology for customers and researchers to compare equivalent tasks.
Why one efficiency number is not enough
Tokens per watt or tokens per dollar can be useful operational measures, but neither is meaningful without a defined workload. Tokenisation differs between model families, and two systems can produce the same number of tokens with different accuracy, latency and usefulness. A credible comparison must hold the task and quality target reasonably constant and explain whether energy use covers only an accelerator, an entire server or the complete data-centre facility.
Throughput and latency can also pull in different directions. Batching many requests may improve total output while making an individual wait longer. Reserving spare capacity can make a service more reliable during demand spikes while lowering average utilisation. “Useful yield” is consequently better treated as a family of workload-specific measurements than as one universal score.
Where cross-layer design can help
Microsoft points to memory management, compiler placement, network topology, power controls and fleet operations as parts of the same system. The argument is that a bottleneck sometimes has to be solved outside the component where it appears. Faster chips provide little value if they frequently wait for model weights, network traffic or unavailable electrical capacity.
The company cites its Maia accelerator and Cobalt CPU work as examples of hardware-software co-design. Those are vendor descriptions, and buyers still need independently comparable measurements, failure-rate assumptions and workload details. The larger point is sound: procurement based only on a headline chip specification can overlook cooling, networking, utilisation and software costs that determine the service delivered to users.
What customers should demand
A useful disclosure would identify the model and precision, task mix, quality threshold, input and output lengths, request latency, batch size, hardware utilisation, energy-measurement boundary and total cost assumptions. Without that context, an efficiency claim may be technically correct but impossible to compare. Transparent methodology is what would turn Microsoft’s framing into a dependable industry measure rather than another marketing label.
Source: Microsoft official announcement and keynote summary.
Screenshot of Microsoft’s official announcement. Source: Microsoft.



