decryptingtech

Technology. Business models. Market debates.

Browse this section

Networking

Thousands of accelerators can behave like one powerful machine only if they communicate effectively. When information arrives late, expensive processors wait. When a network becomes congested or unreliable, a large job can lose time far beyond the cost of the failed component. AI networking is therefore a productivity system for the compute fleet, with economics that extend well beyond the price of switches and cables.

The investment question

The investment question is whether networking suppliers retain value as AI systems become larger and more integrated. Demand for bandwidth is attractive, but the profit pool can move between switch silicon, complete systems, network interfaces, software and optics. Growing traffic does not guarantee rising margins for every participant.

Our view is that the most durable advantage lies in predictable performance at scale, supported by software and operational trust. Customers want networks that maintain useful compute utilisation under demanding traffic patterns. Open standards broaden competition, while integrated platforms can reduce deployment risk. The market will be shaped by that tension rather than a simple contest between two protocol labels.

How AI networks work

Scale-up connects accelerators within a tightly coupled computing domain. Scale-out links servers or groups of accelerators across a larger cluster. Connections between data centres add another set of distance and latency constraints. These boundaries vary by architecture; a rack is a physical enclosure, not a universal definition of the network domain.

AI workloads can require coordinated exchanges involving many devices. Congestion, uneven paths and a slow participant can reduce the progress of the wider job. Network design therefore considers topology, bandwidth between groups, traffic scheduling, congestion control and failure handling. Remote direct memory access, or RDMA, can move data with less host-CPU involvement, but the surrounding implementation determines the delivered benefit.

Ethernet and InfiniBand are important scale-out approaches. NVIDIA also supplies NVLink for tightly coupled accelerator communication. Its networking overview distinguishes these roles. Ethernet’s broad ecosystem is valuable, but an ordinary enterprise network is not automatically an efficient AI fabric. Likewise, a specialised fabric still needs operational integration, software support and suitable economics.

Market structure and competitive advantage

Layer Illustrative participants Basis of competition
Switch silicon Broadcom, NVIDIA, Cisco Bandwidth, latency, features and power
Network systems Arista, Cisco, NVIDIA and others Software, integration and operating reliability
Network interfaces Platform and semiconductor suppliers Data movement, offload and ecosystem support
Cloud-designed networks Hyperscalers and manufacturing partners Fleet-wide optimisation and purchasing scale
Silicon, systems and operational software are distinct layers, even when one vendor supplies several of them.

Merchant switch silicon allows system vendors to focus on software, integration and customer requirements. Proprietary silicon can enable differentiated features but requires substantial development scale. Hyperscalers can choose complete systems, design parts of their own network or combine components from several suppliers. The commercial opportunity depends on which functions they buy externally.

Software creates staying power through configuration, telemetry, automation and troubleshooting. Customers value consistent behaviour across a large fleet and the ability to diagnose problems without interrupting workloads. A low-priced switch that increases operational complexity can be more expensive over its life than a better-supported alternative.

Economics: protect compute utilisation

Networking revenue depends on port volumes, speeds, topology and product mix. A transition to higher speeds can raise selling prices while reducing cost per unit of bandwidth. The number of required switches can also change as switch capacity and architecture evolve. Revenue should therefore be modelled from the actual network design rather than by multiplying accelerator growth by a fixed ratio indefinitely.

The customer’s economic comparison includes switches, interfaces, optics, cabling, electricity and operations. It also includes the compute time lost to communication delays. Illustratively, if a fleet costs 100 per hour and improved networking recovers five hours of productive time, it creates 500 of additional available compute value before considering whether that capacity can be monetised. The example shows why performance consistency can justify a networking premium.

Margins depend on the supplier’s contribution. A semiconductor vendor, a system integrator and a software-rich networking platform have different cost structures. Large cloud buyers can also negotiate aggressively and change sourcing decisions between deployments. Strong revenue growth should be read alongside customer concentration and gross-margin trends.

Backlog requires interpretation. Customers may order ahead of deployment to secure supply, and component availability can affect shipment timing. A backlog reduction can reflect successful delivery rather than weak demand; an increase can reflect constraints rather than immediately stronger end use. Compare orders, shipments and operational deployments together.

AI and hyperscalers: architecture owners matter

Hyperscalers shape networking because they operate at sufficient scale to optimise the entire fleet. They can tune software to their own workloads, influence standards and qualify several suppliers. Their internal accelerators introduce additional interconnect choices and can change the balance between proprietary and open technologies.

Merchant accelerator platforms pursue integration for similar reasons. NVIDIA’s Spectrum-X description combines Ethernet switches, network interfaces and software as a coordinated platform. The strategic benefit is reduced integration burden; the customer’s countervailing concern is dependence on one supplier across more of the system.

Inference does not remove the networking opportunity, but it changes its shape. Model distribution, retrieval, cache movement and separated serving stages create traffic that differs from large training collectives. Some workloads need tightly coupled communication; others can be distributed more independently. The investment case should follow workload architecture rather than assuming training and inference require identical network spending.

Current market debates — September 2026

Recent results point to strong activity. Arista’s August 2026 release reported its first quarter above $3 billion of revenue and highlighted new 1.6-terabit AI fabric platforms. Cisco’s fiscal 2026 results described strong networking orders and hyperscaler AI momentum. Company-wide figures are not directly comparable measures of AI-networking market share.

The immediate debate is how much Ethernet gains in demanding AI environments and which vendors benefit. The Ultra Ethernet Consortium’s first public specification established an industry effort to address AI and high-performance-computing communication requirements. Publication of a specification is a milestone, but product interoperability, deployment experience and customer adoption remain separate tests.

The constructive case for independent system vendors is that buyers want open ecosystems and supplier choice. The integrated-platform case is that customers prioritise time to deployment and proven performance. Evaluate repeat production wins, realised margins and the customer’s chosen operating model instead of assuming that protocol adoption maps directly to one company’s success.

Structural debates: openness and system boundaries

Open standards can reduce dependence on proprietary interfaces while leaving significant differentiation in implementation. Two products supporting the same standard can differ in congestion behaviour, observability, software maturity and support. Openness expands the potential supplier set; it does not automatically make networks interchangeable.

Another structural question is how far tightly coupled computing domains expand. Larger domains can improve some workloads while increasing design and reliability demands. At the same time, distributed inference and geographically separated capacity introduce different requirements. These trends can coexist, supporting multiple architectures rather than one universal network design.

Optics increasingly affects the boundary between semiconductor and system suppliers. Moving optical functions closer to switching silicon can reduce electrical-path constraints but change qualification, servicing and the allocation of value. The optical-connections deep dive examines this separately; the networking implication is that future system choices may alter today’s component relationships.

What to watch

Track production AI deployments, customer concentration, speed transitions, software adoption and realised margins. Distinguish switch-silicon sales from complete-system revenue. Look for evidence that new fabrics improve job completion time and reliability under representative workloads.

The strongest networking franchise makes the compute fleet more productive and easier to operate across successive generations. That outcome is more durable than winning an isolated port-speed comparison.

Explore this sector

AI infrastructure & data centres — sector overview

Related sectors: Semiconductors & chipmaking · Data platforms & analytics