decryptingtech

Technology. Business models. Market debates.

Browse this section

Compute

An accelerator can be available, powered and technically healthy while producing little economic value. It may be waiting for data, synchronising with other machines or reserved for demand that has not arrived. Compute is the business of turning expensive processing hardware into useful work. The most important number is often how much valuable output the system delivers under real operating conditions, rather than the performance printed on its specification sheet.

The investment question

The compute investment question is who earns attractive returns from growing demand for processing: the chip designer, server integrator, cloud operator or customer using the resulting service. Each can report AI-related growth while facing a different margin structure and capital burden. This page examines deployed computing systems; the chip-design deep dive addresses the upstream semiconductor business.

Our view is that utilisation, software and fleet management increasingly determine the economics. Better hardware matters, but customers need reliable throughput at an acceptable cost and response time. Suppliers with strong integration and operational capabilities can create value beyond the component bill, while undifferentiated capacity is more exposed to price competition and generation turnover.

How AI compute works

CPUs coordinate general-purpose tasks and system operations. GPUs and other accelerators perform suitable workloads using specialised or parallel execution. Memory holds the data those processors need, while interconnects move information within and between servers. A usable platform also requires drivers, compilers, libraries, scheduling, storage access and monitoring.

Training adjusts model parameters using data and optimisation. Inference runs an already trained model to produce outputs. For language models, inference commonly includes processing the prompt and generating subsequent tokens. These phases can stress the system differently. NVIDIA’s Dynamo documentation describes separating prefill and decode into different worker pools so each can be managed independently.

Disaggregation is a trade-off, not a free improvement. It introduces data transfers and scheduling requirements, and its benefit depends on workload shape and hardware configuration. Similarly, batching requests can improve throughput while affecting latency. A system optimised for the maximum number of tokens per second may provide an unsuitable experience for an interactive application.

Market structure and competitive advantage

Layer Illustrative participants Customer value
Accelerator platforms NVIDIA, AMD, hyperscaler designs Hardware capability and software ecosystem
Servers and integrated racks Dell, HPE, Supermicro and manufacturing partners Deployment, integration and support
General cloud platforms AWS, Microsoft Azure, Google Cloud, Oracle Cloud Provisioning and broader services
Specialist compute operators CoreWeave and other providers Capacity access and workload-focused operations
Revenue at one layer can include components purchased from another; adding these revenues overstates final customer spending.

A server integrator earns value through engineering, supply-chain execution, delivery and support. Much of a system’s selling price may reflect expensive purchased components, so revenue growth alone is a poor guide to margin potential. A cloud operator earns from providing capacity and services over time, but carries financing, utilisation and operational risk.

Specialist providers can compete through availability, performance tuning and focused support. Hyperscalers can combine compute with databases, security, distribution and customer commitments across a broader platform. Neither approach guarantees superior economics for every workload. The relevant comparison includes service quality, actual pricing and the resources needed to move or operate the application.

Economics: useful output per unit of cost

The cost of compute includes depreciation or lease expense, electricity, networking, facilities, software operations and support. Some costs scale with usage, while others remain even when the fleet is idle. Realised revenue depends on contracted pricing, customer utilisation, service credits and the mix of commitments versus more flexible demand.

Benchmark comparisons need matching conditions. Model quality, numerical precision, prompt length, output length, concurrency and latency requirements can all change results. MLCommons’ inference benchmark guidance uses defined workloads and rules to make comparisons more meaningful. A vendor’s peak arithmetic figure or power-supply rating is not a measured cost-per-answer result.

Illustratively, if a server’s annual ownership cost is 100 and productive utilisation rises from 40% to 60%, fixed cost per productive hour falls by one-third. The calculation assumes unchanged capability and fixed cost; energy and other variable costs still need to be added. It shows why scheduling, reliability and customer demand can be financially as important as a new processor generation.

Asset life is another major sensitivity. Faster hardware can reduce the competitive price of older capacity before the older equipment is physically worn out. Accounting depreciation is not a guarantee of economic life. Older systems may remain useful for less demanding workloads, but that requires customers, compatible software and pricing that covers continued operating costs.

AI and hyperscalers: internal workloads and external demand

Hyperscalers can allocate compute between internal applications and external customers. Their custom accelerators may improve the economics of workloads they understand well, while merchant platforms provide flexibility and broad software support. This creates a portfolio decision rather than a simple choice of one chip supplier for all AI.

The same infrastructure can support model developers, enterprises and applications owned by the platform. Demand signals should therefore be interpreted carefully. A growing internal model programme may create upstream equipment orders without producing immediate external cloud revenue. Conversely, AI can improve advertising or software economics without appearing as a separately reported compute product.

Network, power and cooling constraints can prevent ordered hardware from entering service. Compute demand should be evaluated alongside commissioning capacity, not only accelerator availability. A provider promising future capacity must coordinate the complete deployment and fund the gap between purchasing equipment and collecting customer revenue.

Current market debates — September 2026

Recent results illustrate the scale of orders and the need to distinguish financial measures. Dell’s fiscal second-quarter 2027 release, published in September 2026, reported $60.9 billion of AI server orders, $16.4 billion of AI server revenue and a $95 billion backlog. Orders, recognised revenue and backlog are different quantities; none by itself measures the customer’s deployed utilisation or return.

The immediate debate is whether strong commitments translate into sustained demand through successive hardware generations. The constructive case is that model use expands and customers value reliable access to capacity. The countercase is that rapid replacement, deployment delays and pricing pressure reduce returns on the installed fleet. Monitor actual service availability and customer consumption alongside upstream sales.

Another debate concerns how inference is organised. The development of specialised serving software and separated processing stages can improve efficiency but change the optimal hardware mix. MLCommons’ endpoint benchmark framework highlights the relationship between throughput, concurrency and response characteristics. This is a better basis for evaluating customer value than declaring a universal winner from one performance number.

Structural debates: commoditisation and financing

Compute becomes more commodity-like when customers can move equivalent workloads easily between providers and compare delivered performance transparently. Differentiation persists where software, reliability, data location or integration make switching costly. Open frameworks can reduce friction while leaving significant operational differences between platforms.

Financing is equally structural. An operator can grow quickly by committing to equipment and facilities, but debt service and replacement spending remain real obligations. Long-term contracts may reduce demand risk, yet customer concentration and counterparty quality still matter. A large revenue backlog should be assessed against capital requirements, service commitments and the timing of cash receipts.

The final question is whether efficiency expands the market enough to offset falling resource use per task. Lower costs can enable new applications, while competition passes some gains to customers. A successful compute thesis must explain both growth in useful workloads and the provider’s ability to retain part of the resulting economic benefit.

What to watch

Track productive utilisation, realised pricing, service reliability, workload-specific performance and the economics of older hardware. Compare orders with recognised revenue and commissioned capacity. Examine financing costs, customer concentration and free cash flow after growth investment.

The strongest provider is the one that keeps delivering useful computation at competitive economics as hardware and models change. Owning the newest machines is an input to that outcome, not the outcome itself.

Explore this sector

AI infrastructure & data centres — sector overview

Related sectors: Semiconductors & chipmaking · Data platforms & analytics