Main content

Scale-beyond: when AI systems become more than their clusters

map of the world with focus on India

Artificial intelligence (AI) is reshaping the infrastructure of cloud computing. Over the past several years, I have seen the industry response be primarily focused on building bigger pools of compute. Accelerators became more powerful, clusters became larger, and networks evolved to connect unprecedented numbers of resources together. Whether described as scale-up, scale-out, or more recently scale-across, the underlying objective remained the same: create larger compute domains capable of supporting increasingly demanding AI workloads. 

That strategy has delivered extraordinary results. Training clusters now contain hundreds of thousands of accelerators. New networking technologies have extended cohesive infrastructure domains beyond the boundaries of individual buildings. What once required a supercomputer can now be achieved by infrastructure operating at cloud scale. Yet as AI continues to expand, a different set of constraints is beginning to shape architectural decisions. 

For years, the industry's scaling terms have all described the same basic ambition: how to build a larger compute domain: 

  • Scale-up makes a single node or rack larger 
  • Scale-out connects nodes and racks together  
  • Scale-across extends scale-out across facilities 

These three together work to create one larger compute domain that behaves as a whole. Scaling up and out are foundational concepts evolved from high performance computing. More recently, as the limits of a single facility became real, scale-across allowed us to stretch that computing domain within a campus or region, for a single operator. But, these facilities alone don’t complete the picture for scaling AI connectivity globally. 

As I have met with customers over the past year, two things came into focus:  

  1. AI is entering a phase of massive global adoption and distribution,   
  2. Our shared language of how networks support that distribution was lacking.  

Established service providers want to build products that enable them to participate in the AI supercycle; end-user cloud providers and enterprises want to consume these products. Both lack a shared industry understanding of exactly what is unique about networks with AI workloads, and it acts to stifle innovation and growth. In the same way, suppliers want to innovate and deliver new products capable of meeting the coming challenges. How should the industry approach this problem – more than just modernizing today’s global Internet, and distinct from the existing networks which focus on compute scale rather than distribution? 

The answer is informed by networking history. As the Internet was built over the last 30 years, we learned how to scale both size and operations. Federation, domains, interworking and other key architecture concepts allowed for the robust connectivity we enjoy today, while maintaining control and sovereignty. That same pattern will repeat as AI becomes the dominant workload across the globe. Now is the right time to define the discovery space for those innovations, so that operators and enterprises can collectively grow and scale in cadence with the growth of AI technology itself. 

This new AI network category is called scale-beyond

Global AI infrastructure at scale 

At Nokia, we have a long heritage deploying carrier-grade networks around the globe. These networks are used by billions of people everyday, across mobile, fixed networks, enterprises and more. Designed for dependability and scale, they enable telcos, carriers, and enterprises to deliver rich connectivity services worldwide. 

Our experience with these networks gave us a key insight: the same capabilities we’ve deployed for backbones, mobile, fixed networks, and enterprises worldwide are now required to realize the full vision of AI infrastructure.  

A modern AI system may include compute, storage, data, KV cache, prefill, CPU racks, and more. These pieces increasingly live across multiple locations, multiple operators, and multiple administrative domains. When that's the reality, the central challenge stops being a larger single compute infrastructure, but instead coordination across independent, geographically dispersed systems. 

As I talk to network operators and AI/cloud providers, they all seem to agree that a new type of connectivity solution is required to enable these geographically distributed compute resources to operate as a single unified intelligent platform. This solution would bring rich, operationally robust networking capabilities to the massive scale and demands of modern AI systems. 

Building on the existing terms we have for solutions that connect AI accelerators within a single cluster: scale-up, -out, and -across; the term scale-beyond seemed appropriate. Scalable connectivity that extends beyond individual clusters. Of course, the existing solutions remain necessary building blocks; you still need scalable, fast and predictable fabrics to pool compute. Scale-beyond recognizes that the next frontier isn't just connecting more accelerators. It's coordinating increasingly distributed AI resources and operational domains into a single federated system.

What is scale-beyond?

Scale-beyond describes the interconnection of disparate compute resources and operational domains in AI infrastructure. And while that typically will involve connections across longer distances, scale-beyond addresses the underlying requirements of the network traffic rather than being defined by geography. The majority of scale-beyond networks will echo today’s service provider networks in footprint, but they will be powered by capabilities and features that support the uniqueness of AI workloads. 

While scale-beyond might simply describe a purpose-built AI native data center interconnect between operators, it also could be private peering connections between enterprises. Or, the aggregation of secure connections to AI-RAN sites across the globe. Consumer and small business fiber broadband can have the intelligence to handle high performance AI traffic. Scale-beyond allows the industry to drive AI aware connectivity deep into our global infrastructure, unlocking new opportunities and business models. 

A key attribute of scale-beyond is the ability to organize communication between participants in the AI infrastructure ecosystem. When we all align on terms and categories, it makes the business of growth easier. Carriers and telcos can more easily create products that serve the needs of the market. Cloud operators and enterprises sourcing connectivity can do so in a common language. 

Another benefit is that suppliers and industry analysts can more accurately quantify the set of products and offerings for each use case, and measure their growth. This helps cloud operators source products that have common alignment, through more robust product strategy, roadmaps and supply chain sourcing. 

Scale-beyond also recognizes learnings from the last five decades of building the Internet. We use different products to build regional and global networks than we use inside a data center, as each is optimized for a purpose. Switches are designed as high-radix, lower cost devices for massive connectivity. Routers and coherent optics are used to manage both the propagation delay introduced by distance, as well as the impedance mismatch of bandwidth between domains. Scale-beyond gives the entire AI infrastructure ecosystem a natural point to drive these carrier-like requirements, freeing the building and campus networks from being overloaded with meaning. 

Conclusion

Distributed AI infrastructure is now becoming a quilt of connectivity, leveraging existing capacity with massive growth, provided by a variety of global and regional operators. Cloud operators are deploying private infrastructure, which will both use and interconnect with these operators in every region. Scale-beyond is the framework that describes this phenomenon, and helps point our attention at the right problem. Just as scale-up and scale-out evolved from high performance computing, scale-beyond builds on the rich heritage of service provider connectivity and applies it to a new era of AI infrastructure. 

The industry spent a decade learning how to connect more accelerators. The next decade is about coordinating everything those accelerators depend on. That's the work of scale-beyond. 

Clayton Wagar

About Clayton Wagar

Clayton leads AI and High Performance Computing field activities for the Network Infrastructure Group at Nokia. He previously has served in both the development and deployment of many major technology shifts, including broadband, metro Ethernet, and over-the-top video streaming. Additionally, he is a part of a non-profit foundation serving North America’s largest motion picture studio, where he consults on the intersection between technology and storytelling.

Article tags