1GW Compute Park: The Real Challenge Is Delivery
This week I ran a batch of inference stress tests in the terminal. Cloud GPUs had been used for about a month, with samples coming from the re-inspection queue of industrial visual inspection. The model wasn't large, but once concurrency ramped up, latency jitter exploded before VRAM did. While pondering whether to move checkpoints to a parallel file system, I scrolled past news that Tata Consultancy Services (TCS) plans to build a 1GW AI data center in Hyderabad, with an estimated investment of 700 billion rupees.
There are other numbers in the news: HyperVault acquired 264 acres of land, OpenAI is the first customer initially using 100 megawatts, with potential expansion to 1 gigawatt in the future. My first reaction landed on delivery. Putting GPU racks into a building is just the surface for an AI data center. The difficulties lie in power, cooling, networking, storage, and scheduling. It's more like a high-power, high-noise, highly-coupled factory.
In the short term, 1GW tends to distract attention. Land is secured, capacity is written down, and the investment amount is scary, but what makes compute run is power, cooling, networking, storage, and scheduling. This relationship is critical. 100MW can already support high-intensity deployment; 1GW is more like a long-term roadmap. Customers buying AI compute won't just look at park area; they'll look at whether it can stably provide GPU hours, network throughput, fault recovery, and data residency.
In my tests, training and inference bottlenecks have different emphases. In training clusters, frequent collective communications between GPUs mean if the network jitters, the entire card group waits. Inference is more fragmented; KV cache, batching, cold starts, and long-tail requests all affect first-token and last-token latency. Using the same AI data center to serve both types of business results in completely different operational complexities.
Power, heat dissipation, networking, storage, and software are all hard thresholds. Substations, transformers, UPS, switching times, redundant paths—missing any one slows down rack-up. High-density racks struggle with air cooling; liquid cooling isn't just about connecting pipes. Training looks at RDMA, multi-rail, bandwidth, and latency; inference looks at egress, load balancing, and cross-region scheduling. Model weights, datasets, and checkpoints all require read/write access. If the parallel file system can't handle metadata, GPUs still starve. Scheduling, fault tolerance, monitoring, billing, and multi-tenant isolation—these determine whether it can be sold to enterprises.
So the short-term judgment is: don't rush to applaud it as globally leading. First see if the 100MW can be accepted in batches based on power, SLA, and customer load. HCLTech is investing 35 billion rupees with a long-term goal of 50MW; Google is also pouring $15 billion into India for data centers and submarine cables. Indian IT giants entering en masse indicates everyone is looking at the foundation for enterprise AI deployment in the coming years.
In the long term, this is more interesting than the server room itself. TCS used to excel at outsourcing delivery: engineers, processes, SLAs, client onsite support. Now it's acquiring land, building AI data centers, and preparing to let thousands of employees use enterprise ChatGPT and deploy Codex to improve software engineering outcomes. The route is clear: it wants to stuff AI into its delivery system, then sell that system to clients.
I've said before that AI-generated code is just part of delivery; the key remains acceptance criteria, test records, and rollback mechanisms. Compute is similar now. 1GW sounds like scale, but customers buy certainty: Can the model run on time? Can data stay local? Can faults be recovered? Can costs be calculated clearly?
India has engineers, an English-speaking market, data localization demands, and a tradition of enterprise software outsourcing. If TCS can make HyperVault verifiable infrastructure, its role will change. It needs to shift from taking projects to providing compliant inference, industry fine-tuning, private deployment, and compute scheduling. This is attractive to Chinese overseas expansion, European enterprises, local finance, and manufacturing clients.
However, the shortcomings are blunt. High-end GPU supply chains, liquid cooling equipment, power engineering, cross-region networking, and O&M talent can't be filled by announcing plans. The battery industry competes on parameters; the compute industry competes on wattage. In the end, it comes down to failure validation. Model training failures can roll back checkpoints; data center power outages are accidents. Customers won't forgive 100MW instability just because you have a 1GW plan.
My long-term judgment is that success depends on whether it can break "AI-readiness" into testable items, including availability per megawatt, network packet loss, storage write-back, GPU failure rate, customer data isolation, and disaster recovery time. If these metrics pass, 1GW is an asset. If not, it's a pile of expensive cabinets.
If you're currently doing model deployment or selecting AI infrastructure for your company, my advice is: don't be scared by gigawatt numbers. Break the problem down small. Test the business link first, then discuss model selection. Customer service Q&A, document extraction, and visual quality inspection have completely different compute requirements. Real-time interaction looks at latency; batch processing looks at throughput; privatization looks at data residency; training looks at fault recovery. If you can't get business metrics, you can't judge how 1GW relates to you.
Require suppliers to provide phased acceptance. Don't just listen to "AI data center," "high-performance GPU," "green compute." Look at power caps, PUE or cooling efficiency standards, network bandwidth, SLAs, fault drills, data isolation, audit logs, and rollback mechanisms—can these be written into contracts? Without these, procurement is just buying concepts.
Acceptance units should fall to whether each megawatt can run stably at full capacity. When I ran inference stress tests, benchmarks showed bottlenecks often came from token output, network jitter, and storage write-back. Large AI data centers will be the same. The bigger the number, the more it tests engineering details.
So regarding the TCS news, focus on the 100MW landing in the short term, and in the long term, focus on whether it can convert outsourcing capabilities into compute delivery capabilities. 1GW is impressive, but AI infrastructure ultimately competes on stable power supply, stable cooling, stable scheduling, and stable rollback.
Physix Frontier