
Three Networks Inside the AI Factory
Understanding the optical challenge begins with recognizing that an AI cluster contains several distinct networking environments.
Gartner divided AI infrastructure into three broad tiers: scale-up, scale-out and scale-across. Each operates over a different distance, carries a different level of traffic and creates a different set of requirements for the interconnect.
Scale-up describes the connections within a rack, where operators place as many GPUs as possible inside servers and then pack those servers into the available rack footprint. Gartner estimated that the bandwidth within this environment can be approximately 500 times that of a traditional wide-area network application.
Scale-up connections still rely heavily on electrical interfaces because electrical interconnects remain relatively inexpensive and power efficient over short distances.
Once the compute capacity of a rack has been exhausted, the cluster must expand into additional racks. This is the scale-out network, where 400G and 800G pluggable optics connect large numbers of GPU systems operating in parallel.
Gartner characterized scale-out bandwidth as roughly 50 times the capacity associated with a conventional WAN environment.
The third tier, scale-across, emerges when a data center reaches its practical power limit and the AI infrastructure must extend into another facility. Those data centers may be separated by tens or hundreds of kilometers, requiring coherent optical technology capable of carrying extremely high-capacity signals over longer distances.
Scale-across networks can represent approximately 14 times traditional WAN bandwidth, according to Gartner.
Taken together, the three tiers illustrate why optics has become inseparable from the AI infrastructure discussion. Network capacity must expand inside the rack, across rows of racks and increasingly between separate data centers—all without consuming an untenable share of the power budget or introducing failures that leave GPUs idle.
When One Link Slows the Whole Cluster
The reliability requirement for AI networks differs sharply from that of conventional enterprise or internet infrastructure.
In a traditional TCP/IP network, an error may trigger a packet retransmission. The application or user may never realize that the error occurred.
AI training clusters operate differently. GPUs work on large problems in parallel and must remain synchronized. When an error occurs on one link, the impact can extend across the cluster.
“All of the GPUs need to stop and back up to a checkpoint and restart,” Gartner said. “Because everything is running in parallel in a GPU infrastructure, a link error on one link actually impacts neighboring links.”
Those failures can be caused by a bad connection, a dirty optical connector, noise on the link or other physical and electronic problems. The resulting instability is often described as a link flap—a connection repeatedly dropping and recovering.
Gartner said information Cisco has received from hyperscale customers indicates that link flaps can reduce the efficiency of GPU infrastructure by as much as 40%.
That figure changes the practical meaning of optical reliability. A transceiver is no longer simply one component among thousands in the network. An unstable link can undermine the utilization of the far more expensive accelerators attached to it.
“We have every incentive to make sure that the optics are super reliable,” Gartner said.
The economics are stark. Operators are investing billions of dollars in AI campuses designed to keep accelerators running at the highest possible utilization. A network issue that forces thousands of GPUs to stop, return to a checkpoint and restart can erode the productivity of the entire deployment.
In that environment, optical reliability becomes a compute-efficiency metric.
1.6T Moves Toward Deployment
Most large AI training networks are currently being built around 400G and 800G connectivity. Gartner described those technologies as ubiquitous within hyperscale training infrastructure, with each serving different portions of the network.
The next step is already approaching.
Cisco is sampling 1.6-terabit optics with customers, and Gartner said the first volume deployments are expected to begin later this year. As with previous transitions from 10G to 100G, 100G to 400G and 400G to 800G, the appeal lies in improving cost and power efficiency per bit.
“There’s no question that there’s a demand for 1.6T,” he said. “We’re looking beyond 1.6T as well, as to what happens next.”
Adoption will remain uneven across customer segments.
Hyperscalers are driving most of the demand and volume for 800G and 1.6T technology. Traditional service providers remain earlier in their 400G and 800G deployments, while 400G is still in the early stages within many enterprise networks.
That distinction matters because hyperscale purchasing volume will help determine the technology’s cost curve. Gartner advised service providers and enterprises to closely watch the architectures being adopted by the largest cloud and AI infrastructure operators.
“If you want to be on the low-cost curve, you should be on the hyperscaler cost curve,” he said.
The Coming Electrical-to-Optical Transition
The progression from 400G to 800G and then 1.6T is significant, but Gartner views it as a familiar generational evolution. The more disruptive transition will occur when electrical connections can no longer carry the bandwidth required within the scale-up network.
The constraint comes down to physics.
As the bitrate of a signal increases, the distance it can travel generally decreases. Electrical interconnects remain attractive inside the rack because they are inexpensive and efficient, but the industry will eventually reach a point where they cannot support the required capacity over even short distances.
“My bet is ultimately on an optical solution,” Gartner said.
Several potential architectures are under development, including co-packaged optics, near-packaged optics, wide electrical buses using slower-rate signals and radio-frequency technologies.
Co-packaged optics and near-packaged optics move optical components away from the front panel of the switch and closer to the switching silicon. Instead of transmitting a high-speed electrical signal several inches across a line card to a pluggable transceiver, the system converts the signal into light nearer to the chip.
That shorter electrical path can reduce the need to repeatedly retime the signal, lowering power consumption. Removing pluggable modules from the faceplate could also increase port density because optical connectors require less space than complete transceivers.
Cisco expects co-packaged solutions to enter the market this year, giving operators an opportunity to gain practical experience with the architecture.
But the transition introduces operational tradeoffs.
Pluggable optics are replaceable, interoperable and available from multiple vendors. When one fails, a technician can replace the affected module. In a co-packaged design, the optical components may be mounted directly to a line card or integrated with the switch silicon.
That raises a more difficult question: What happens when a single optical channel fails?
Operators may need to rely on software to route around the failed link, replace a larger portion of the system or reconsider how they stock and service network equipment. Questions involving multi-vendor sourcing and component interoperability will also have to be resolved.
Gartner expects those operational issues to be worked through over the next several years before co-packaged or near-packaged optics become commonplace in scale-up networks.
Routed Optics Reduce the Scale-Across Penalty
Optical architecture is already changing how operators connect data centers.
Traditional long-distance optical networks use transponder line cards housed in dedicated chassis. The transponder receives an optical signal from a switch or router, converts it into a form capable of traveling over a long distance and combines it with other wavelengths on the same fiber.
Through its acquisition of Acacia, Cisco developed an architecture known as routed optical networking. The approach replaces the standalone transponder line card with a coherent pluggable optic inserted directly into the router.
Gartner said the architecture can consume approximately 90% less power than the conventional transponder-based approach while requiring less space and reducing cost.
Hyperscalers were among the earliest adopters, using coherent pluggables to connect AI infrastructure across data centers. The technology is now moving into service-provider networks and enterprise campus environments.
For AI operators, routed optical networking makes scale-across architectures more practical. When the power ceiling at one location prevents a cluster from expanding further, coherent optical links can extend the training environment into another facility.
Rather than acting as a bottleneck, Gartner said, optics is helping AI networks continue to scale beyond the physical and electrical limits of a single building.
“We’re going to see increasing use of multiple data centers involved in a single training infrastructure,” he said.
Coherent Optics Moves Inside the Data Center
The boundaries between scale-out and scale-across may also become less distinct.
Conventional data center optics commonly use intensity-modulation direct-detection technology, or IMDD, for connections rated over distances of approximately two kilometers. As bitrates continue increasing, maintaining that reach becomes more difficult.
The physical distance between switches does not become shorter simply because the network has moved to a faster generation. Operators expect each new technology to support the cabling and layouts already in place.
At some point, Gartner said, conventional data center optics may no longer be able to maintain the required reach. The industry may then need to adopt techniques from coherent optics, which have historically been used for metro, regional and long-haul networks.
“The likelihood is that we’ll have to borrow some of the coherent techniques that are used outside the data center and bring them in the data center,” he said.
That evolution could further blur the distinction between a data center campus, a metro cluster and geographically distributed AI infrastructure. Optical systems capable of supporting connections beyond 1,000 kilometers are already extending the practical reach of data center interconnection.
Bringing coherent techniques inside the facility would create new power and cost considerations, but it could allow operators to continue increasing bitrates without redesigning every physical network around shorter reaches.
Inference Changes the Optimization Target
Training clusters currently command much of the industry’s attention because of their enormous scale and capital requirements. Gartner expects the next major deployment wave to arrive as trained models move into inference.
Inference environments generally require smaller systems than frontier training clusters, but they will be deployed much more broadly across hyperscalers, service providers and enterprises.
The optimization target will therefore change.
Training places a premium on scale and the ability to connect enormous numbers of GPUs. Inference will place greater pressure on cost efficiency, power efficiency and the ability to support specialized models closer to users and applications.
That will keep optics at the center of the infrastructure equation, even as the type and size of the deployment changes.
For data center operators, the underlying lesson is that networking can no longer be treated as a secondary layer added after decisions about compute, power and facilities have been made. Optical cost, power consumption, reach and reliability are becoming design inputs for the AI factory itself.
As Gartner put it, optics is not limiting the growth of AI networks today. It is one of the technologies allowing them to move beyond the power and geographic boundaries of a single data center.



















