Vivold Consulting
Business & Enterprise

Why Enterprises are Moving Critical AI Workloads On-Premise

Cost, latency and data-sovereignty rules are pulling enterprise AI back on-premise

Key Insights

After a decade of cloud-first migration, enterprises are bringing mission-critical AI workloads back into their own data centers, driven by soaring cloud costs, latency, and data-sovereignty rules like the EU AI Act and GDPR. New liquid-cooled, Blackwell-class hardware from Dell, HPE, and Lenovo now puts hyperscale-equivalent compute within reach of any well-capitalized firm. The result is a hybrid model, with Goldman Sachs, Siemens, and NTT DATA among the most advanced adopters.

Stay Updated

Get the latest insights delivered to your inbox

The cloud-first consensus is quietly reversing

For most of the past decade, enterprise IT had one default answer: put everything in the public cloud. AI is rewriting that. As models move from pilots to mission-critical infrastructure, the limits of a cloud-only approach - latency, data sovereignty, regulatory compliance, and cost - are pushing companies to bring AI workloads back behind their own walls. Purpose-built private infrastructure for training and inference is becoming a central pillar of enterprise strategy rather than a niche concern.

The numbers behind the shift

The spending signals are hard to ignore:

  • IDC reported enterprise compute and storage hardware for AI grew 166% year-on-year in Q2 2025, while Gartner pegged 2025 AI spending at US$1.5tn, with data-center systems up nearly 47%.

  • The GPU server market, worth US$171bn in 2025, is forecast to hit US$730bn by 2030.

  • For firms in regulated industries or under data-residency laws, the cloud isn't just costly - it can be a legal risk, with confidentiality obligations sometimes requiring on-premise deployment outright.
What changed on the supply side is that the hardware caught up: liquid-cooled GPU servers built on NVIDIA's Blackwell architecture, available through Dell, HPE, and Lenovo, now deliver petaflop-scale inference in racks a company can own and secure itself. Most organizations are landing on a hybrid model - public cloud for elastic, non-sensitive work; private data centers for inference and fine-tuning; edge for latency-critical tasks.

What it means for the data center

Bringing AI in-house is not just racking more servers. Densities can reach 100 kilowatts per rack, which makes traditional air cooling inadequate and turns power resilience, grid connectivity, and thermal management into strategic concerns - the data center becomes, in effect, an AI factory.

Who's furthest ahead

The piece profiles three very different adopters. Goldman Sachs has built a private agentic stack and became the first major bank to roll out Cognition's autonomous engineer Devin across its 12,000 developers, reporting three-to-four-times productivity gains in software lifecycle work - funded partly by capital redirected from its retreat from consumer banking. Siemens pushes AI onto the factory floor via its Industrial Edge platform and is building modular, lower-carbon data-center units. And NTT DATA runs agentic AI inside its Cyber Defense Centers to protect private infrastructure, cutting alert volumes by up to 90%. The throughline: on-premise AI is now as much an engineering and security discipline as a software one.

More in Business & Enterprise

All Business & Enterprise stories

Nadella's warning: you're paying for AI twice - once in tokens, once in your own IP

In a blog post, Satya Nadella warned that enterprises using proprietary AI models are paying twice - once in money for tokens, and again in the proprietary knowledge they must reveal to make those models useful, since models learn from the 'exhaust' of prompts, tool use, and especially corrections. He argued it is inconsistent for labs to claim fair-use rights to train on the world's public data while restricting others from distilling their models in return. Nadella - whose company invests in both OpenAI and Anthropic - later doubled down on CNN, saying firms without their own models or an AI gateway layer separating prompts, memory, and harness from the model won't survive as firms, having 'outsourced your thinking.'

Amazon retires Mechanical Turk: the platform that secretly powered 'AI' for 21 years is done

Amazon will close Mechanical Turk to new customers on July 30, 2026, moving the 21-year-old crowdsourcing marketplace into maintenance mode with no new features - and reporting indicates SageMaker Ground Truth and Amazon Augmented AI close to new customers the same day. Launched in 2005 as 'artificial artificial intelligence,' MTurk annotated the data that trained a generation of models; by 2023 a study found 33-46% of its workers were using LLMs to do the tasks, dissolving the platform's reason to exist. If your research, labeling, or human-review pipeline touches MTurk, you now have a migration deadline.

Zuckerberg's candid admission: AI agents 'haven't accelerated the way we expected'

At an internal town hall on July 2, Mark Zuckerberg told employees that AI agent development over the last four months has not accelerated as expected, that Meta's sweeping reorganisation was not as clean as it could have been, and that its bets on the new structure have not yet paid off - remarks first reported by Reuters from a recording. The admission stings because Meta laid off about 10% of its workforce and reassigned ~7,000 people to AI teams in May, with executives who planned the reorg reportedly optimistic about tools like Claude Code. Zuckerberg still expects significant AI benefits within three to six months - but the gap between agent hype and agent reality just got named by its biggest spender.