Cloud repatriation is the deliberate movement of specific workloads from public cloud back to private cloud, colocation, on-premises or edge infrastructure. It is selective, not a cloud exit: 69% of IT leaders planned some repatriation in 2025 (Broadcom), yet only 8–9% plan full-scale workload repatriation (IDC). The triggers are workload economics, data sovereignty and always-on AI inference.
For most of the last decade, the infrastructure decision inside large enterprises was effectively pre-made. Cloud-first was on the policy. Migration was the mandate, and the only serious question was sequencing.
That assumption is now being audited. Enterprises are working through their estates application by application, asking which workloads genuinely earn their place in a rented environment, and which have been paying a premium for elasticity they stopped using years ago. The word repatriation suggests retreat. The data suggests something more disciplined.
In Broadcom’s 2025 survey of 1,800 IT leaders, 69% were considering moving workloads from public to private cloud and about a third had already done so, yet only 10% wanted to run private cloud alone.
Flexera’s 2025 State of the Cloud report puts roughly 21% of workloads and data as already moved back to private or on-premises environments.
At the same time, Secondary reports citing IDC indicates only about 8-9 % of organizations are moving everything off cloud, and Gartner forecasts India’s public cloud spending will rise 28.1% to $17.5 billion in 2026, while worldwide sovereign cloud IaaS spending grows 35.6% to $80.4 billion.
Read together, those figures describe a correction that enterprises are not leaving the cloud. They are narrowing what they send there. For infrastructure teams this distinction is important because it shifts the decision from being based on platform to individual workloads.
None of the forces behind this shift amount to “cloud is bad.” Each one describes a category of workload that stopped fitting the rented model.
The first is cost gravity. Cloud pricing is best suited for fluctuating workloads and it rewards the variability. Whereas a steady state application might just have to pay for the elasticity it never needs. As FinOps practices mature and organizations gain visibility into per-workload unit economics, those line items become difficult to defend.
The second is regulation. The EU AI Act, India’s Digital Personal Data Protection Act and sector-specific rules across banking, insurance and healthcare have transitioned data residency from just an engineering dilemma to a board level priority. The physical presence of data now has a legal factor attached and the cleanest way to ride that is a sovereign or in-country infrastructure. This same pressure is driving the broader move toward sovereign AI strategies.
The third is latency and control. Checkout systems, pricing engines, industrial telemetry and real-time operations perform better when compute sits close to where the work happens. Predictable billing belongs to the same argument, and finance teams value a known monthly cost more than most architecture reviews acknowledge.
The fourth force is reshaping the arithmetic faster than the other three and deserves its own section.
Public cloud remains an excellent environment for AI training. Training is bursty by nature: intense demand for a period, then nothing. That profile is exactly what rented infrastructure was built to serve.
Inference behaves differently. Once a model reaches production it runs continuously, and continuous GPU consumption produces a very different cost curve. Deloitte’s Tech Trends 2026 projects that roughly two-thirds of all AI compute will be inference by 2026, up from about a third in 2023.
That matters because of a threshold that most cloud vs on-premise cost comparisons tend to converge on: somewhere around 60 to 70% sustained utilization, owned GPUs typically become cheaper than rented ones. Below it, cloud wins on flexibility. Above it, the enterprise is renting an asset it uses like an owned one.
The practical question is no longer whether to use cloud for AI. It is how many hours a day a given GPU will actually run, a question that sits at the centre of modern AI cloud architectures.
This is not a theoretical debate. Walmart operates what it describes as a triplet model, blending public cloud, private cloud and roughly 10,000 edge nodes in stores, reportedly cutting 10 to 18% of annual cloud spend while strengthening its business continuity strategy during outages.
Airtel turned its own infrastructure into a sovereign, India-built cloud, positioning it for regulated sectors that need full data residency. KDDI built MShip3, an OpenStack-based private cloud across six sites, serving sovereignty-sensitive industries. At a very different scale, 37signals left AWS for owned hardware and projects around $10 million in savings over five years with the same team size.
What links these examples is not a rejection of public cloud. Each organization kept it for the workloads where it performs best and moved the ones where it did not.
The wrong lesson to draw is that cloud is finished and everything should come home. Public cloud still wins decisively on elasticity, global reach, managed services and speed of experimentation. Rebuilding those capabilities internally is expensive, and doing it badly is more expensive still.
Repatriation also carries a hidden requirement. A distributed estate spanning public, private, sovereign, edge and on-premises environments demands real platform engineering maturity: consistent Kubernetes operations, infrastructure as code, unified observability and a multi-cloud and hybrid cloud security posture that holds across all of it. Organizations that move workloads without building that operating model tend to trade a predictable cloud bill for unpredictable operational risk.
The strategic shift underway is smaller than the headlines suggest and more demanding than it sounds. Cloud-first was a destination. Workload-first is a strategic discipline that weighs economics, sovereignty, latency and control for each application rather than just having default answer to all.
The organizations handling this well are building that framework now, before AI inference volumes make the cost of guessing wrong considerably higher. Through our cloud modernization services and cloud-native engineering work with enterprises and operators, the harder problem is rarely the migration. It is running a multi-environment estate with one operating model, one security model and one view of cost.
That is the capability worth building next, and it is where the next round of infrastructure advantage will be decided. Hughes Systique helps enterprises decide where each workload belongs and run public, private and edge environments as one estate.
Cloud repatriation is the deliberate movement of specific workloads from public cloud back to private cloud, colocation, on-premises or edge infrastructure. It is a workload-level decision rather than a platform-level one, which is why most programmes move a subset of applications rather than an entire estate. Broadcom’ 2025 survey found 69% of IT leaders planned some repatriation during the year. The most common candidates are steady-state applications, regulated data and sustained AI inference.
No, and the distinction matters commercially. IDC data indicates only around 8% of organizations are moving everything off cloud, while Gartner forecast global public cloud spending at $723 billion for 2025, growing about 21% year over year. The same period saw roughly 21% of workloads move back to private or on-premises environments, per Flexera. Both trends are real at once: the cloud market keeps expanding while enterprises become more selective about what they place there.
Three profiles come up repeatedly. Steady-state applications with flat, predictable demand pay for elasticity they never use, so their unit economics usually improve on owned or private infrastructure. Workloads bound by data residency rules under the EU AI Act, India’s DPDP Act or sector regulation in banking and healthcare often need sovereign or in-country hosting. Sustained AI inference is the third, once GPU utilization stays consistently high.
It depends on utilization rather than on the workload type. Industry TCO analyses generally place the crossover somewhere around 60 to 70% sustained utilization, above which owned GPUs tend to beat rented ones for continuous inference. This is becoming a larger question because Deloitte projects roughly two-thirds of AI compute will be inference by 2026, up from about a third in 2023. Training, which is bursty, usually remains better suited to public cloud.
Platform engineering maturity, primarily. Running a distributed estate across public, private, sovereign, edge and on-premises environments requires consistent Kubernetes operations, infrastructure as code, unified observability and a security model that holds identically everywhere. Without those, organizations trade a predictable cloud bill for unpredictable operational risk. A structured workload placement framework and a clear data migration strategy applied before migration begins, is what separates a cost-driven repatriation from an expensive reversal.