Your legal team says the data can't leave the EU. Your CFO says the API bill can't keep doubling. There are now six credible ways to run enterprise AI — from a €4,000 box under a desk to a sovereign cloud — and most companies only know about one of them.
Two years ago, "how will we run this?" was not an interesting question. You got an API key from OpenAI or Anthropic, you put it in an environment variable, and the architecture discussion was over.
That's no longer true, for three unrelated reasons that happened to arrive at the same time.
Open weights got good enough. The gap between the best model you can rent and the best model you can download is now measured in months, not generations. Alibaba's Qwen family and DeepSeek dominate open-model downloads and lead most open benchmarks; OpenAI's `gpt-oss` and Mistral are the main Western options. That's a real choice you didn't have in 2024.
The regulatory picture stopped being theoretical. The EU AI Act's general-purpose AI obligations have been in force since August 2025. The high-risk obligations were pushed back by the Digital Omnibus — Annex III systems to 2 December 2027, regulated products to 2 August 2028 — but deferred is not cancelled, and procurement teams are already writing the questions into RFPs.
Somebody looked at the bill. A successful pilot has a nasty property: the better it works, the more it costs. Token spend scales with adoption, and unlike a licence it never plateaus — a pattern we also mapped in Where Business AI Actually Works in 2026.
So the question is live again. Here is the actual landscape — and the honest catch on each option.
First: the two questions that decide everything
Before comparing hardware, answer these. Everything else follows.
1. What data actually touches the model? Not "what's in our system" — what goes into a prompt. Public marketing copy and internal CVs with names, salaries and health disclosures are not the same risk, and shouldn't get the same architecture.
2. What is your sustained token volume? Not your peak demo, not your ambition. Tokens per day, today. This number decides the economics, and almost nobody measures it before the debate starts.
With those two numbers, the six options sort themselves.
Option 1 — Hosted frontier API (the baseline)
Anthropic, OpenAI, Google, direct. Best models, zero operations, per-token pricing, enterprise terms available: signed DPAs, no training on your data, zero-retention modes on request.
Who it's for: anyone who hasn't yet proven the use case, and most companies who have. It is genuinely hard to beat on cost below a few thousand euros a month.
The catch: the vendor is a US entity, which means the CLOUD Act sits behind whatever the contract says about storage location. Your cost scales linearly with success, forever. And you are one pricing change or deprecation notice away from re-architecting.
Option 2 — Frontier models inside your own cloud tenancy
The most under-used option in the list. The same frontier models, served through your existing AWS, Azure or Google contract — Amazon Bedrock, Microsoft Foundry, Google Vertex AI. Your VPC, your IAM, your logging, your existing cloud DPA.
Who it's for: companies who already have a serious cloud footprint and a procurement process that has already cleared that provider. This is usually the shortest path from "legal said no" to "legal said yes."
The catch, and it is a sharp one: an EU endpoint does not mean EU processing. As of now, AWS Bedrock's EU inference profile is the one that contractually pins execution to European regions (Frankfurt, Ireland, Paris, Stockholm, Milan, Spain). Equivalent Claude deployments on Foundry and Vertex have shipped as "global standard," meaning inference can be routed outside the EU regardless of which endpoint you call. If data residency is the reason you're here, read the inference-routing documentation, not the marketing page. And the CLOUD Act still applies to all three.
Option 3 — Open-weight models on managed European inference
Per-token APIs, EU-owned companies, EU data centres, no GPUs to babysit. Scaleway (France, ISO 27001 and HDS-certified), Nebius (Netherlands, Finnish and French capacity), OVHcloud, IONOS, T-Systems (Germany, BSI C5 and TISAX), Exoscale (Switzerland), and 3DS OUTSCALE, which holds French SecNumCloud certification — the closest thing Europe has to a government-grade sovereignty stamp.
Who it's for: teams whose blocker is jurisdiction rather than raw capability. You get the operational simplicity of Option 1 with an EU legal entity on the other end of the contract.
The catch: you're limited to open-weight models, which still trail the frontier on the hardest reasoning and long-horizon agentic work. Smaller providers mean thinner rate limits, fewer regions, and less mature tooling.
Option 4 — Self-hosted open weights in your own cloud
You rent GPUs — from a hyperscaler, or from Hetzner, CoreWeave, Nebius, Lambda — and you run the serving stack yourself: vLLM or SGLang, a gateway, autoscaling, evaluations, observability. The model file is yours. Nothing leaves your account.
Who it's for: high, predictable volume on a task where a mid-size open model is good enough. Document classification, extraction, enrichment, embedding — the boring high-throughput work.
The catch, which is where most business cases quietly die: the GPU bills by the hour whether it is saturated or idle. On-demand H100 capacity runs around $2.89/hour, H200 around $3.59. At full saturation the per-token cost is excellent — one recent analysis put a `gpt-oss-120b` deployment near $0.10 per million tokens against $0.17–$0.60 from hosted providers. At 10% utilisation, it's ten times worse. Add engineering time, failover and idle headroom and realistic all-in cost lands at three to five times the raw GPU rental. The rough thresholds circulating among people who have actually done this: under ~$50K/year of API spend, don't; $50K–$500K, hybrid; above $500K, a well-utilised cluster can win. You need roughly 50–80% sustained utilisation just to tie a managed open-weight provider.
Option 5 — Hardware you buy once
The option most people mean when they ask this question, and it's really two very different products.
Desk-class appliances. NVIDIA's DGX Spark packs 128 GB of coherent unified memory and about 1 petaFLOP of FP4 compute into a small box at $4,699 MSRP, and will hold models up to ~200B parameters in memory. A Mac Studio M3 Ultra with 256 GB of unified memory (~$6,000, 819 GB/s) is the same idea with a nicer OS. An AMD Strix Halo mini-PC with 128 GB lands near $2,000. Dell, Lenovo and others now ship equivalents on the same NVIDIA silicon.
These are transformative for development — a full model, air-gapped, on a desk, for less than three months of a mid-size API bill. They are not production servers. Memory bandwidth, not capacity, sets token speed: a 70B dense model on a 128 GB Mac Studio runs at roughly 8–15 tokens per second for a single user. Put ten concurrent users on it and the experience collapses. If you need one workstation that serves a small team, an NVIDIA RTX PRO 6000 with 96 GB of VRAM at ~$22,000 is the honest entry point.
Rack-class systems. This is the real "buy it once" tier: an NVIDIA DGX B200 (eight Blackwell GPUs, 1,440 GB of GPU memory) at roughly $300K–$500K, or an integrated platform — HPE Private Cloud AI, Dell AI Factory, Lenovo Hybrid AI Advantage, Cisco Secure AI Factory, Nutanix Enterprise AI — which bundle the GPUs with an inference stack, RBAC, monitoring and air-gapped operation so your team isn't assembling it from parts.
Who it's for: genuinely air-gapped requirements (defence, classified, some clinical and legal work), data that legally cannot traverse a public network, or organisations with enough sustained load that three years of amortisation beats three years of tokens.
The catch: you are buying a depreciating asset in a market where a hardware generation lasts about 18 months. The capex is the easy part; the hard part is power, cooling, rack space, and the two engineers who now maintain an inference platform instead of shipping product. And the box does not come with the model evaluation, retrieval layer or guardrails that make it useful.
Option 6 — European vendors with self-deployment rights
The middle path that gets forgotten: buy a commercial model from a European company, with the contractual right to run it inside your own perimeter.
Mistral (France) licenses its models for on-premise and private-VPC deployment with full EU data residency. Aleph Alpha (Germany) built its Pharia suite specifically for regulated industries and offers on-premise, private VPC, air-gapped and hybrid deployment. LightOn (France) does the same for document intelligence. Silo AI (Finland, now part of AMD) covers Nordic and low-resource European languages that the US labs handle poorly.
Who it's for: regulated buyers who want a vendor to call, a support contract, and an EU legal entity — but also want the weights inside their own building.
The catch: you pay commercial licence fees and run the infrastructure, so it is rarely the cheapest option. Capability still trails the frontier on the hardest tasks.
The comparison, on one screen
1. Hosted frontier API — Data leaves your control: Yes (contractual limits) — Ops burden: None — Cost shape: Per token, scales forever — Best model available: Frontier — Sovereignty: Weak — CLOUD Act
2. Frontier in your tenancy — Data leaves your control: Partly — check routing — Ops burden: Low — Cost shape: Per token + cloud commit — Best model available: Frontier — Sovereignty: Medium (Bedrock EU profile strongest)
3. Managed EU inference — Data leaves your control: No — EU entity, EU region — Ops burden: None — Cost shape: Per token — Best model available: Open-weight tier — Sovereignty: Strong
4. Self-hosted in your VPC — Data leaves your control: No — Ops burden: High — Cost shape: Fixed GPU-hour, utilisation-driven — Best model available: Open-weight tier — Sovereignty: Strong
5. Owned hardware — Data leaves your control: No — Ops burden: Highest — Cost shape: Capex + power + staff — Best model available: Open-weight tier — Sovereignty: Total
6. EU vendor, self-deployed — Data leaves your control: No — Ops burden: High — Cost shape: Licence + infrastructure — Best model available: Near-frontier (EU) — Sovereignty: Total
What most companies should actually do
Almost nobody should pick one. The pattern that works looks like this:

Route the high-volume, low-judgement work — classification, extraction, tagging, embedding, summarising a known format — to a small open model on Option 3 or 4. This is usually 80–90% of your tokens and about 10% of the difficulty.
Route the low-volume, high-judgement work — the hard reasoning, the multi-step agent, the thing a customer actually sees — to a frontier model via Option 1 or 2. This is where capability is worth paying for. For agent workloads that forget, contradict, or re-ask, see The Four Kinds of Memory Every Enterprise AI Agent Needs.
Put both behind one internal gateway so the routing decision is a config change, not a rewrite. That gateway is the single most valuable thing you can build here, because it means the six options above stop being an irreversible bet and become a dial you can turn as prices, models and regulation move.
And be honest about what problem you're solving. "We need it on-prem" is sometimes a compliance requirement and sometimes an instinct dressed as one. The two lead to very different budgets.
Coming next
This piece maps the options. The follow-up puts numbers on them: what each path actually costs at 10 million, 100 million and 1 billion tokens a month — capex, GPU hours, licence fees, and the engineering time nobody puts in the business case.
If you're weighing one of these decisions and want a second opinion on the architecture before you commit budget, that's most of what we do at Future Proof Technology.
