How AI Is Redefining Data Center Design and Infrastructure
Ahmed Alsunbel, Regional Partnerships Lead, MEA, Turkey, and Africa at Submer, explains how AI is redefining data center design, the challenges of scaling AI infrastructure, the growing role of digital twins, and why speed, resilience and security have become the defining priorities for next-generation AI facilities.
How is the rise of AI workloads changing the design and architecture of modern data centers?
AI, and generative AI in particular, has broken the assumptions that traditional data center design was built on. Enterprise IT workloads were largely CPU-based, modestly power-dense, and tolerant of latency. AI training and inference workloads are the opposite: they run on dense clusters of GPUs or accelerators, they generate heat and draw power at a scale legacy facilities were never designed for, and they depend on extremely low-latency, high-bandwidth interconnects between nodes.
This is pushing architecture in a few clear directions. Rack densities that used to average 5–10 kW are now routinely 40–100 kW, with some AI training racks exceeding 100 kW. Facilities are being designed around GPU clusters as the core unit of design rather than the server rack. Data center footprints are shifting toward purpose-built AI factories rather than general-purpose colocation space, with power availability not just land , now the primary siting constraint. In short, AI hasn’t just added a new workload type; it has forced a redesign of the electrical, thermal, and network backbone of the facility itself.
What are the biggest challenges in scaling infrastructure to support high-density AI compute environments?
Three challenges dominate right now. The first is power availability. Utility interconnection queues in many markets now stretch years, and AI campuses are being planned around where megawatts can be secured, not where customers or fiber happen to be. This is reshaping site selection strategy across the industry.
The second is thermal management at density levels air cooling simply cannot handle. Beyond roughly 30–40 kW per rack, air cooling loses efficiency and floor space rapidly, forcing a transition to liquid cooling a shift that touches facility plumbing, rack design, and maintenance practices simultaneously.
The third is the pace mismatch between construction timelines and compute demand. Traditional data center builds take 18–36 months; AI demand cycles move in months. Operators are responding with modular and prefabricated designs, phased power delivery, and retrofit strategies for existing facilities, but the fundamental tension between “build for scale” and “build fast” remains unresolved for most of the industry.
How are power, cooling, and networking requirements evolving with AI-driven infrastructure?
Power is evolving from a facilities afterthought to a strategic constraint. Operators are pursuing on-site generation, direct power purchase agreements, grid-interactive designs, and in some cases nuclear or gas-based dedicated generation to guarantee capacity independent of utility timelines.
Cooling is moving decisively toward liquid direct-to-chip cooling and immersion cooling are becoming standard for high-density AI racks rather than niche solutions. This requires new plumbing infrastructure, coolant distribution units, and a shift in facility design from air-handling-centric to fluid-handling-centric engineering.
Networking is being redesigned around GPU-to-GPU communication rather than server-to-storage traffic. High-bandwidth, low-latency fabrics (such as InfiniBand and emerging Ultra Ethernet approaches) are essential because AI training performance is often bottlenecked by interconnect speed, not raw compute. This is driving investment in optical interconnects and rethinking traditional network topologies to minimize hop count between accelerators.
Where do you see the convergence between physical infrastructure management and cybersecurity today?
The line between physical infrastructure and cybersecurity has effectively dissolved. Building management systems, power distribution units, cooling controllers, and rack-level sensors are increasingly IP-connected and software-controlled, which means they now sit inside the same threat surface as IT systems rather than outside it.
This convergence shows up in a few concrete ways: security operations centers are beginning to ingest telemetry from both IT and OT (operational technology) systems; access control, surveillance, and environmental monitoring are being integrated into unified platforms rather than siloed systems; and physical security incidents (unauthorized rack access, tampering with cooling or power systems) are being treated as potential precursors to or components of cyberattacks. The practical implication for operators is that a vulnerability in a cooling controller or a power management interface is no longer just a facilities issue ,it’s a security issue with the same severity classification as a network intrusion.
What new security risks emerge as data centers become more software-defined and AI-orchestrated?
Software-defined, AI-orchestrated infrastructure introduces risks that didn’t meaningfully exist in traditional facilities. Expanded attack surface is the most obvious one every sensor, controller, and management interface that becomes IP-addressable and software-controlled is a potential entry point, including systems that were historically air-gapped or isolated by design.
AI orchestration systems themselves become high-value targets, since compromising the system that allocates compute, manages workload placement, or controls thermal and power balancing could allow an attacker to degrade performance, cause physical damage, or exfiltrate data at scale without ever touching a traditional application layer.
Supply chain risk grows as well, given the number of specialized hardware and firmware components (GPUs, accelerators, cooling controllers, power electronics) sourced from a concentrated set of vendors, each a potential vector for compromised firmware or hardware.
Finally, there’s a newer category: adversarial risk to the AI systems themselves model poisoning, data exfiltration through inference APIs, or manipulation of training pipelines which sits at the intersection of traditional cybersecurity and AI-specific security practices that many organizations are still building expertise in.
How important is real-time simulation or digital twin technology in planning and managing AI infrastructure?
It’s becoming close to essential rather than optional. Given the cost and lead time of AI infrastructure a single high-density AI campus can represent capital investment in the billions the ability to model power distribution, thermal behavior, and network topology before physical construction, and then continuously during operation, materially reduces risk.
Digital twins are proving valuable in several ways: pre-construction, they allow operators to simulate rack layouts, airflow or coolant flow, and power distribution under various load scenarios; in operation, they enable predictive maintenance and early detection of thermal or electrical anomalies before they cause downtime; and for capacity planning, they let operators model the impact of adding new AI workloads or hardware generations without disrupting live environments.
As AI hardware generations turn over faster than traditional server refresh cycles, the ability to simulate infrastructure changes before committing capital is shifting from a competitive advantage to a baseline operational requirement.
How are vendors adapting to the need for faster deployment cycles in AI infrastructure environments?
Vendors across the ecosystem are restructuring around speed as a primary design constraint, not just a cost or efficiency one. On the hardware side, this means modular and prefabricated data center components power skids, cooling units, and even entire modular data halls that can be manufactured off-site and assembled on-site in a fraction of traditional construction time.
Liquid cooling vendors are shipping standardized, pre-engineered CDU (coolant distribution unit) systems rather than fully custom installations, reducing design and integration time. On the power side, vendors are offering pre-permitted, containerized generation and battery storage systems that can be deployed ahead of full utility interconnection, letting operators bring compute online in phases.
And on the software side, infrastructure management platforms are increasingly offering AI-driven automation for commissioning, workload placement, and capacity optimization , compressing timelines that used to require extensive manual configuration.
The broader trend is a shift from bespoke, fully customized data center builds toward standardized, repeatable “AI factory” designs that can be deployed and scaled faster, even if that means somewhat less site-specific optimization. Speed to power-on has become as important a competitive metric as efficiency or cost per megawatt.



