NVIDIA is positioning AI factories as a new class of infrastructure designed to produce intelligence at scale, combining accelerated computing, networking, memory, storage, software, power and cooling to support always-on AI workloads.

AI factories represent a shift from traditional data centres, which primarily store and process information, toward infrastructure designed to continuously generate AI outputs. NVIDIA describes the primary product of these systems as intelligence, measured through metrics such as tokens per second, tokens per watt, cost per token, utilisation and uptime.

The shift is being driven by the rapid growth of agentic AI, where AI systems do more than respond to individual prompts. Autonomous agents can reason, plan, search, retrieve information, use tools, write code and take actions, creating longer and more compute-intensive workloads.

According to NVIDIA, these workloads require infrastructure capable of keeping the entire AI workflow moving efficiently. Accelerated computing works alongside high-speed memory, storage, networking and CPUs, while software orchestrates the different components to maintain throughput, responsiveness and utilisation.

NVIDIA also highlights the importance of full-stack codesign in AI factories. Hardware, networking, memory, storage and software are designed and continuously optimised together to increase utilisation, improve performance per watt and reduce the cost of producing AI tokens.

As AI workloads become more interactive, inference is also becoming a real-time orchestration challenge. AI factories must route requests, manage memory, coordinate services and balance latency with throughput while keeping computing resources highly utilised.

NVIDIA says AI factories can range from smaller systems supporting individual business units to massive facilities designed for high-performance AI training and inference. Its DSX reference designs are intended to help organisations design and optimise large-scale AI factories, including gigawatt-scale infrastructure.

The company is also using digital twins to address the complexity of building and operating large AI facilities. The NVIDIA Omniverse DSX Blueprint connects facility design with hardware and software through digital representations, allowing organisations to model, validate and optimise infrastructure before construction and throughout the operational lifecycle.

Energy efficiency is another critical consideration. NVIDIA emphasises performance per watt as an important measure of AI factory competitiveness because computing efficiency directly influences the economics of producing intelligence at scale.

The growing adoption of agentic AI is therefore changing the requirements for AI infrastructure. Instead of treating compute, networking, storage, power and cooling as separate components, AI factories bring these elements together into an integrated system designed for continuous AI production.

NVIDIA’s approach reflects a broader transition in computing infrastructure as enterprises and organisations prepare for AI systems that increasingly reason, act and operate continuously. The company expects full-stack AI factories to become an important foundation for scaling the next generation of AI applications and services.

Source: This article is based on an official press release issued by NVIDIA
NVIDIA AI Factory Compute Is Becoming an Investable Asset Class | NVIDIA Blog