Source: Unite.AI
Over the past few years, and counting, the AI infrastructure market has spent a great amount of time and money (close to $7 trillion of data-centre investment by 2030, by McKinsey’s estimate) learning how to build large data centres wherever power is available. “Wherever” being the key word because the location wasn’t very important. A training cluster simply needs access to sufficient power. But location is starting to become increasingly relevant nowadays as AI’s usage grows beyond chat interfaces to machines, voice systems, video, and applications interacting with the physical world. Such operations require computing power to be closer to the urban centres where those interactions happen to keep latency down. While tools like chatbots can tolerate a half-second delay, automated service applications like remotely supervised delivery robots, live voice assistants or automated fulfillment systems may need to react within milliseconds.
This is the new challenge on the block for AI companies. How to find smaller blocks of reliable power in and around large metropolitan areas? And even after locations have been scouted, how to deal with the backlash that even remote data centres are facing?
The Demand for AI in Urban Centres Is Rising
There is an increasing demand within the urban sphere for AI that goes beyond personal use of chatbots or any hyperfixation on robots. An IMF working paper estimates that electricity consumption from data centres and AI had already reached 400–500 TWh globally in 2023, and AI-driven consumption could potentially reach 1,500 TWh by 2030. The automation required for making simple operations more efficient — be it in processing paperwork or archiving videos — requires a lot of power at scale.
Cities are at the forefront of this, as seen in an OECD review of 250 cities across 78 countries. It found that 56% were already actively using AI, while another 35% were piloting or planning deployments. To illustrate one simple use case, let’s consider surveillance cameras in the city. IHS Markit estimated several years ago that the number of surveillance cameras worldwide would pass one billion. Whilst they are mainly used for recording purposes, they are also now used for other, more computationally demanding jobs, e.g., tracking objects across a network, or searching large video archives using natural language. Such processes cannot take place in every camera per the current AI processing we have. Should they be sent out of the city then to a far-off data centre to process? That is impractical due to data regulations and also the cost of moving such large volumes of traffic.
A recent Gallup survey shows that around 30% of private and public sector employees use AI in their role and the numbers are rising. These sectors — be it banking, hospitals or public-sector organisations — run into the same issues. They want AI infrastructure within their own jurisdiction and physically close enough to control. Their interactive voice systems and video applications degrade as network latency rises. Even if the problem can be initially bypassed by retaining a human operator for the first generation of autonomous systems, that supervision in itself is not very forgiving towards latency. For instance, delivery robots and home robots supervised over a video connection require a full communication loop that has to remain below approximately 80 milliseconds as measured by a remote robotic teleoperation study. Autonomous cars, counterintuitively, are the exception: driving stays on board for safety reasons, and Waymo’s remote operators work at a quarter of a second of latency, some of them from another continent. What the city inherits from autonomous fleets is their data, not their steering wheel.
The data-centre market is underestimating this surge in demand. The World Bank published a paper showing how municipalities in countries like Argentina, India, and Ethiopia, among others, have adopted AI in areas of justice, anti-corruption, customs, public administration, and service improvement. This isn’t a local occurrence. It’s a global transformation. At Teravolt, our estimate is that a metropolitan area of 10 to 15 million people could support roughly 6 to 13 MW of continuous latency-sensitive AI load by 2030, before general-purpose robots are mass deployed. Arguably, it’s not that big of a number by hyperscale standards, but finding it at the right spot is quite harder than it looks. By 2036, the same metropolitan area moves to 22 to 50 MW of latency-sensitive load — and to 60 to 100 MW if the car industry starts producing robots at automotive scale. The demand economics are unforgiving in the best way: a few hundred dollars a year of server cost on a machine that replaces $30,000 to $60,000 of labour is demand that does not haggle over price.
Adding a Fourth Tier to Computing
Currently, computational power for AI exists in three levels. The top layer consists of gigafactories that are used to train models. Their location matters to developers only insofar as it secures large amounts of power at an acceptable price. Then come regional AI facilities, typically around 20 to 200 MW, where much of the recurring inference demand in major cloud and enterprise operations occurs. Lastly, we have the computing power embedded directly in the device or the robot.
There appears to be a gap between the last two layers that I identify as the urban layer—the gap responsible for latency issues. These would need to be located inside or very close to large cities. They would need to be measured in single-digit or low double-digit megawatts while being positioned within metropolitan areas and connected to high-quality fibre. The starting point could be around 2 MW if the site has guaranteed power. To illustrate the practicality of it, consider a delivery robot. It handles different requests and not all can be processed on the device at hand. While it is trained at the gigafactory, more advanced assistance in real-time would be needed to be processed by a nearby urban facility instead of sending it to a regional facility every time.
The Two-Way Traffic Issue
Urban data centres will not simply be intermediary points of receiving data from training centres, but they will also produce large amounts of data to be sent back up the pipeline.
Every autonomous machine and sensor-rich system will be accumulating information that will need to be processed and sent back to be integrated into training. Fleets of robots returning to their charging points after a shift will be sending enormous useful information about edge cases and daily operations that data centres need in order to improve functions. Imagine all this traffic being sent at once in the same direction. It will create a bottleneck. This is also where the urban data centre comes in handy as it can process some of the load before it overwhelms the regional computing facility. It can identify and discard some erroneous data while compressing and moving forth the useful parts.
Essentially, the urban data centre provides another layer of filtering to the two layers that precede it making the flow of data all the more organised.
Is the Future a Server on Every Block?
Although I present urban servers as a solution, it is also simply too expensive to solve latency issues with unbridled proliferation. No matter the size, a data centre still requires a grid connection, transformers, cooling, security, networking and staff. These costs do not shrink proportionally with IT load. To put it into perspective, a 0.5 MW facility still carries many of the same fixed costs as a much larger site.
In our projects, we’ve seen that a standalone urban data centre becomes economically viable at around 5 to 8 MW. Even when assuming that several sites could share a network operations centre, the threshold only falls to 1.5 to 3 MW. Therefore, the bottom line is still the same. Splitting 20 MW across twenty locations instead of two of three saves almost no latency inside a metropolitan area — the city fibre ring is only a few milliseconds across — while multiplying capital and operating cost.
This is why I believe the industry has to move down in size and build what I call “substations for intelligence”. Currently, developers are learning how to build regional facilities in the 20 to 100 MW range. The future goal is to apply that to even smaller, urban sites in order to reach the 2 MW level suitable for major cities.
The Political Impediment
Moving past the economic viability, the political challenge still remains. Data centres have become increasingly vilified in the media for their adverse impact on the environment. According to a Pew survey, more than half of Americans believe data centres to be harmful to the environment and hold negative views towards it. The argument of job creation pales in comparison to the amount of water, land, and electricity consumed; as well as noise and emissions generated. Getting approvals even for smaller urban facilities would prove to be difficult.
We’re already seeing examples of data centre developers and city authorities at loggerheads. In the US, xAI’s use of gas turbines to accelerate its Memphis build-out has led to permitting disputes, community opposition and litigation. Europe has gone further: Amsterdam imposed a moratorium on facilities above 70 MW, Dublin now admits new data centres only if they bring their own generation or storage, and the key substation of the West London corridor will not be reinforced before the early 2030s. Even outside of such high profile cases, with this being a new area of contention, many city authorities don’t have a clear approval process for approving an additional 10 MW of AI infrastructure inside the urban area. As noted by a National League of Cities report, one of the most common challenges for cities is that older local zoning codes do not explicitly address data centres leading to case-by-case reviewing. Such legal uncertainties can delay even technically viable projects.
Infrastructure Decisions Today Shape Our Capacity for the Future
All this is to say that we need to preemptively plan and build the infrastructure required for future progressions in AI. Because while models develop rapidly, the power required to run them depends on a type of infrastructure that takes time to build. Grid connections can take years, and in some major data-centre markets queues now stretch to seven to ten years. So if cities are ramping up AI computing power by 2030, that electricity infrastructure needs to be put into place now. At Teravolt, our working model estimates global urban AI demand at 9 to 18 GW by 2036. That implies around 1000 to 1500 significant urban and metropolitan sites worldwide, depending on the average facility size, and in sum, could represent a market worth approximately $500 billion to $1 trillion.
If the necessary planning is not carried out, we will be stuck with advanced AI capacity that we cannot harness. The mobile era needed to have the infrastructure of towers for radio coverage to follow people otherwise, the mobile phone would have been moot. Similarly, we must recognise the mismatch between the evolution of AI with the installment of electrical infrastructure. Where the mobile era measured coverage in signal bars, the AI era will measure it in milliseconds and watts per resident — and it will need not millions of towers, but roughly a thousand facilities the size of a substation. Fibre, milliseconds and megawatts are the currency in this field and those that recognise this scarcity ahead of time are better suited to shape the future.
