|
AI Collab Score: 8 / 2 For years, enterprise infrastructure solved growth with a simple formula: More applications meant more servers. Most of those servers ran one workload, consumed a full set of power and cooling resources, and spent much of their life underutilized. We built larger server rooms, added more racks, and expanded the facilities' footprint because the architecture gave us little choice. Then virtualization changed the question. VMware did not repeal the laws of physics. It did not make computing free, and it did not eliminate heat—it concentrated more useful work into smaller, highly utilized physical footprints by eliminating so much empty capacity. Virtualization made compute a shared resource. It introduced abstraction, pooling, scheduling, mobility, and policy. It allowed us to ask a much better question: Do we actually need another physical server, or are we simply failing to use the infrastructure we already own intelligently? AI infrastructure is reaching a similar inflection point. We are building larger GPU clusters, denser racks, bigger liquid-cooling systems, and increasingly power-hungry campuses. Much of that investment is necessary. Accelerated computing is not conventional enterprise IT. A modern AI factory requires a radically different power, network, storage, and cooling architecture. But we should still ask the harder question: Are we building the most intelligent systems in history with a pre-virtualization infrastructure mindset? The Water Debate Is Really a Design DebateWhen the public talks about AI data centers, the conversation often reduces to water. How much water do they use? That is the right concern, but it can lead to the wrong level of analysis. The liquid circulating through GPU cold plates and CDUs is often part of a tightly controlled closed loop. The more consequential water issue is frequently the facility heat-rejection design. In an evaporative system, water carries heat away through evaporation. Minerals remain behind, concentrations rise, some water is discharged as blowdown, and makeup water is added to keep the system operating within chemistry limits. While that is a conventional and proven thermal-management model rather than an engineering failure, it deserves severe scrutiny as AI campuses move toward hundreds of megawatts and eventually gigawatt-class ambitions. Not because AI should stop growing. Because water availability should no longer be treated as a secondary facilities detail. It should be treated as a first-order architecture constraint. We need to ask questions that are harder than, “Can the site support another cooling tower?”
Those are not anti-AI questions. They are the questions responsible infrastructure leaders should be asking before the concrete is poured and the GPUs are ordered. The New Version of GPU SprawlBefore we solve for more cooling, we need to confront why so much expensive accelerated infrastructure sits idle or gets used inefficiently in the first place. The early signs of GPU sprawl are already familiar:
Enterprise IT has seen this movie before. The old version was physical server sprawl. Then it became VM sprawl. AI can create GPU sprawl at a far more expensive and resource-intensive scale. Every wasted GPU-hour becomes more power drawn, more heat generated, more cooling required, and potentially more water consumed. That is not primarily a cooling problem. It is a utilization, orchestration, and governance problem. What Is the VMware of AI?The answer is not one product. AI’s VMware moment will likely be a stack of capabilities that makes accelerated compute shareable, schedulable, measurable, and accountable. At the hardware layer, GPU partitioning and virtual-GPU capabilities allow supported accelerators to be divided into isolated slices. A smaller inference service, development environment, or specialized model should not automatically consume a full GPU simply because that is the only allocation model available. At the orchestration layer, the scheduler needs to understand more than whether a GPU is technically free. It must account for GPU-to-GPU connectivity, NUMA alignment, network topology, storage locality, queue position, data gravity, and the difference between a latency-sensitive inference request and a checkpointable training job. At the control-plane layer, enterprises need policy:
A GPU environment without policy is simply a more expensive version of VM sprawl. And at the model layer, the system must continually ask whether the workload needs that much infrastructure at all:
That is the real AI infrastructure conversation. Not just “How many GPUs do we need?” But, “How do we ensure that every GPU-hour produces meaningful value?” The most sustainable AI data center may not be the one with the best cooling system. It may be the one that needs less cooling because it uses the infrastructure more intelligently. Better Cooling Still Matters—But It Cannot Be the Entire StrategyHigh-density AI requires better thermal engineering. Direct-to-chip liquid cooling, CDUs, warm-water loops, rear-door heat exchangers, immersion designs, and improved heat-rejection systems all matter. But a perfectly engineered cooling system can still support a poorly utilized GPU estate. The next major breakthrough may not be a cooling system that enables a one-gigawatt AI campus. It may be the control plane that prevents us from needing half of that campus in the first place. That is the opportunity in front of the industry: treat the platform layer as a sustainability layer. GPU pooling, fractional allocation, intelligent scheduling, topology-aware placement, model routing, inference caching, and precision optimization are not just operational tools. They are conservation technologies. Conservation Is Becoming an Engineering DisciplineThere are already companies approaching the problem from different angles. Google’s Hamina data center in Finland demonstrates that coastal heat-rejection architectures are not theoretical. The facility uses seawater from the Bay of Finland and is paired with offsite heat recovery, showing what becomes possible when a site is designed around the thermal characteristics of its location rather than assuming every facility must use the same cooling model. Nautilus Data Technologies offers another important example. Its liquid-cooling infrastructure can use seawater, river water, or lake water as a heat-rejection source while keeping the technology cooling loop separate. In its natural-water configuration, source water is filtered, passes through heat exchange, and returns only slightly warmer. The company describes this approach as having virtually zero consumptive water use compared with cooling-tower evaporation. That does not make the ocean, river, or lake a free cooling resource. It still requires careful intake design, filtration, thermal-discharge analysis, environmental permitting, ecosystem safeguards, and long-term maintenance planning. But it does prove that the industry has alternatives to treating freshwater evaporation as the default answer. Ecolab is attacking the challenge closer to the operations layer: cooling-loop efficiency, water chemistry, coolant health, corrosion prevention, fouling control, monitoring, and optimization. That may sound less exciting than floating data centers, but it is where operational resilience is won or lost. A liquid-cooling strategy that ignores chemistry, contamination, flow, leakage, and asset health does not eliminate risk. It simply moves it. These are not competing ideas. They are components of a more mature design philosophy: Treat water as a constrained resource, not an unlimited utility. Treat cooling as part of workload architecture, not just facilities engineering. Treat utilization as a sustainability metric, not merely a finance metric. Coastal Compute Is Worth Exploring—But It Is Not a Free PassThe ocean is an obvious place to look for large-scale heat rejection. It offers an enormous thermal resource and could reduce dependence on freshwater evaporation for appropriately located facilities. But “build it near the ocean” is not a complete strategy. A coastal data-center architecture still has to solve for corrosion, biofouling, filtration, heat-exchanger design, marine ecosystem impact, thermal discharge, storm resilience, permitting, maintenance access, power delivery, and network redundancy. The right model is not pumping seawater through servers. It is a carefully engineered, layered heat-exchange architecture where seawater remains physically separated from the controlled cooling loop that serves the facility and IT equipment. The goal is not to shift an environmental problem from municipal water systems into marine ecosystems. The goal is to design a heat-rejection system that is transparent, resilient, and less dependent on freshwater consumption. Frontier Ideas Matter—Even When They Do Not Become the AnswerMicrosoft’s Project Natick explored underwater data-center deployment. Other companies are exploring floating compute platforms and the idea that energy-intensive infrastructure may need to follow energy resources rather than forcing energy to follow compute demand. Those concepts are valuable because they challenge assumptions we tend to accept too easily. What if compute followed abundant energy? What if thermal design was determined by the environment rather than bolted on after the fact? What if location strategy started with power, water, and heat rejection—not just tax incentives, fiber routes, and available land? Even orbital data centers deserve a place in the thought-experiment column. But they should remain there for now. Space does not make heat disappear. It turns thermal management into a radiative engineering challenge. And for tightly coupled distributed training or latency-sensitive real-time inference, an orbital hop is an awfully expensive way to add distance. The point is not that every unusual idea will win. The point is that conventional thinking should not win by default. The Jevons Problem: Efficiency Can Increase ConsumptionThere is another uncomfortable truth. Efficiency does not automatically reduce total resource use. Economists call this the Jevons paradox: when a technology becomes cheaper or more efficient, demand can increase so much that total consumption rises rather than falls. Virtualization gave us a version of this lesson. It improved server utilization. It also made it dramatically easier to create new virtual machines. Many organizations solved physical-server sprawl only to create VM sprawl. AI can follow the same pattern. A more efficient model may make AI cheap enough to deploy across a thousand more workflows. Better GPU scheduling may unlock capacity that encourages more demand. Lower inference costs may increase agentic activity, automated decisioning, and machine-to-machine workloads. That does not mean efficiency is pointless. It means efficiency alone is not sustainability. The goal should not be to promise that smarter infrastructure will make the aggregate AI footprint smaller. It may not. The goal should be to maximize intelligence yield per watt, per liter of water, per GPU-hour, and per dollar of infrastructure deployed. Efficiency must be paired with governance, transparent measurement, and intentional choices about where and when demand should grow. A Better AI Infrastructure ScorecardPUE changed the conversation because it made data-center overhead measurable. AI needs a broader scorecard. We should measure:
A data center should not be judged only by how much compute it contains. It should be judged by how intelligently it converts constrained resources into outcomes. The Real Infrastructure InnovationThe next breakthrough may be a more efficient cold plate. It may be a new heat-exchanger material, a better coolant, a coastal thermal architecture, or a smarter way to reuse waste heat. But history suggests the larger transformation may happen one layer above the hardware. Virtualization taught us that more physical servers did not automatically mean more useful computing. AI must teach us that more GPUs, more cooling, and more water do not automatically mean more intelligence. The sustainable AI factory will be built by facilities engineers, power specialists, network architects, platform teams, data scientists, and business leaders working from the same design principles. It will not begin with the question: “How do we build a larger AI data center?” It will begin with a better one: “How do we produce more useful intelligence while demanding less from the world around us?” Additional Resources
0 Comments
Your comment will be posted after it is approved.
Leave a Reply. |
RSS Feed