Where to Host LLM Inference with EU Data Residency (2026)
Blog

Where to Host LLM Inference with EU Data Residency (2026)

Where to run LLM inference with EU data residency in 2026: EU-native providers, hyperscaler regions, and dedicated GPU options for GDPR and the AI Act.

Tessera 8 min read OpenAIApertusRegoloGoogle Vertex AIAzure OpenAI

Where to Host LLM Inference with EU Data Residency (2026)

You can host LLM inference with EU data residency on a managed EU-native service like Tessera AI, which serves open-source models on dedicated EU and LATAM infrastructure behind an OpenAI-compatible API, or on hyperscaler European regions from AWS, Google Cloud, and Azure. These options keep personal data inside the EEA and align with GDPR Articles 44-49 and the EU AI Act; EU-native providers add the jurisdictional control that a region label alone does not give.

EU Data Residency Requirements for LLM Inference

You must keep personal data inside the EEA or rely on valid transfer mechanisms like the Data Privacy Framework. GDPR Articles 44-49 block transfers outside the EEA unless you use Standard Contractual Clauses or adequacy decisions. Following Schrems II, assess destination-country laws individually to confirm protection levels match EU standards.

The EU AI Act entered into force on 1 August 2024 and becomes fully applicable on 2 August 2026. It adds deployer duties for high-risk systems, including mandatory logging and human oversight. If you use a third-party API, those duties still apply to your use case; our AI Act compliance guide for LLM inference walks through them in detail.

EU data residency is not the same as GDPR compliance. A US-hosted stack may still be lawful if the transfer mechanism is valid, but Schrems II requires assessing whether US law undermines protections promised by SCCs. The AI Act adds governance and traceability obligations on top of GDPR but does not replace transfer rules.

Mapping GDPR Transfer Mechanisms to LLM Workflows

True residency depends on more than server location. Prompts, responses, system instructions, and embedding vectors count as personal data if they identify individuals. Logs and debugging traces often capture raw prompts and must be regionally isolated to prevent accidental cross-border transfers.

Treat a provider’s processing location, remote support access, logging infrastructure, and backups as part of the transfer analysis. Do not rely solely on a visible “EU region” label. Verify that the provider’s legal entity, infrastructure, and inference engine are anchored in Europe.

Remote support staff outside the EEA constitute a transfer requiring its own SCCs or adequacy assessment.

Managed Inference Providers in the EU

Managed inference services meet EU residency requirements by running dedicated regional endpoints and enforcing strict retention policies, without requiring you to manage GPU clusters. EU-native providers operate under European law and outside US jurisdictional reach, which makes them the strongest option for true data sovereignty.

Tessera AI serves open-source models like Qwen3.6-35B-A3B on dedicated EU and LATAM infrastructure through an OpenAI-compatible API. Dedicated tenancy keeps inference workloads on isolated hardware, flat monthly pricing replaces variable token billing, and a standard Data Processing Agreement covers GDPR Article 28 requirements.

Other EU-native options include Apertus, which hosts open-source models exclusively in the EU behind an OpenAI-compatible API, and Regolo, which runs inference in Italy on renewable-powered data centers without retaining user data for training. Whichever provider you shortlist, verify its subprocessors, operational disclosures, and service maturity.

Evaluating Subprocessor Chains and Logging Controls

Document the exact location of every subprocessor, including vector databases, monitoring tools, and error-tracking services. If any component processes data outside the EEA, implement supplementary technical measures like encryption with keys held solely in-region.

The EU AI Act requires high-risk systems to include risk assessment and mitigation, logging, detailed documentation, human oversight, and robustness measures. When using a third-party API, you remain responsible for deployer-facing duties. Verify that the provider’s logging system retains audit trails within the EU and that you can export them for regulatory review; our LLM DPA and subprocessors guide covers the clause-by-clause checklist.

Hyperscaler EU Regions and Limitations

Hyperscalers offer EU endpoints for LLM inference, but EU-region processing is not the same as full sovereignty. CLOUD Act exposure and provider-controlled subprocessors remain relevant even when data stays in-region. Model availability is often the real constraint, as not every model variant is available in every EU region and regional parity may lag the US market.

OpenAI supports data residency for ChatGPT Enterprise and Edu customers in Europe, running GPU inference in-region for eligible content when the setting is enabled. Google Vertex AI lists EU regions in Belgium, the Netherlands, and Finland, though Gemini availability depends on the specific region. Azure OpenAI deploys to Sweden Central and France Central; AWS Bedrock places EU model endpoints in Ireland.

Check subprocessor locations and logging controls before relying on any hyperscaler for compliance. If you require that no data leaves the EU, choose an endpoint with explicit EU-region inference rather than a global API with residency settings alone.

Model Availability and Regional Parity

Providers often roll out new architectures to US regions first, leaving European teams waiting for regional parity. This lag can force architectural compromises, such as routing non-sensitive workloads to US endpoints while keeping sensitive data in-region, or sticking with older model versions that lack critical performance improvements.

Check the provider’s documentation for region-specific release schedules. Review the available models list to verify which architectures support EU-region deployment. Plan your deployment strategy to handle regional gaps, whether that means maintaining multiple model versions or building a routing layer that directs traffic based on model availability and data sensitivity.

EU vs US Latency and Cost Considerations

EU-hosted inference removes a transatlantic round trip from every request, which lowers time to first token for European users compared to calling US endpoints. For interactive workloads like chat and support automation, that overhead recurs on every call.

Routing between European regions costs far less than crossing the Atlantic. AWS’s cross-region inference guidance describes EU geography profiles that keep requests inside European regions, with pricing calculated on the source request region and no extra charge for cross-region routing. The bulk of user-visible delay comes from model inference time, so in-region hosting removes the avoidable network share without touching model quality.

Direct EU-vs-US same-model price benchmarks for open-source models like Qwen, Llama, or Mistral are not well-documented in public sources; a comparable price table requires direct vendor quotes or independent testing.

Cost Modeling for High-Volume Inference

Predicting inference costs requires looking beyond per-token pricing. Self-hosted open-source models shift costs from variable token fees to fixed GPU infrastructure, deployment, logging, access control, and rollback. Account for GPU utilization rates, idle time during low-traffic periods, and the engineering overhead of maintaining serving infrastructure.

Managed inference gateways offer a middle ground by enforcing jurisdiction-aware routing so EU traffic goes to EU-region backends without leaving the continent. The gateway itself adds another vendor to the chain, so exposure depends on the gateway design and the providers behind it. Calculate total cost of ownership by combining gateway fees, backend API costs, and the operational burden of managing routing policies, observability, and provider failover.

Latin America Data Transfer Considerations

EU hosting gives Latin American controllers a recognized adequacy baseline, reducing compliance friction when sending data to European infrastructure. EU data protection law is widely treated in the region as a benchmark for adequate safeguards, while US-hosted inference raises questions about onward access.

Brazil’s LGPD requires ANPD-approved standard contractual clauses for international transfers under Resolution CD/ANPD No. 19/2024. EU infrastructure simplifies documentation by providing established safeguards, making it easier to present as a controlled transfer destination than a US environment.

Argentina holds an adequacy decision from the European Commission, so EU-hosted inference removes transfer risk for Argentine entities. Colombia permits transfers to adequate countries or under specific exceptions, meaning EU hosting fits that model directly. Mexico’s LFPDPPP enforces strict consent rules and cross-border restrictions that apply regardless of hosting location, but EU hosting aligns with a stronger privacy regime and simplifies destination assessments.

Aligning LATAM Privacy Frameworks with EU Hosting

Documenting cross-border transfers for LATAM operations requires matching each country’s specific mechanism to your hosting architecture. For Brazil, prepare ANPD-approved standard contractual clauses and verify that your EU provider can supply the required technical and organizational safeguards. For Colombia, ensure your transfer falls under express consent or contract performance exceptions if you cannot rely on adequacy.

Mexico’s law imposes consent requirements, privacy notices, and cross-border transfer restrictions for controllers. Align your privacy notices to explicitly disclose the EU hosting location and the legal basis for the transfer. Argentina’s framework is recognized as providing adequate protection, so EU hosting satisfies its baseline requirements without additional contractual layers.

Choosing a Residency-Compliant Inference Stack

A residency-compliant stack needs dedicated EU hardware, clear subprocessor controls, and pricing that handles high-volume workloads without penalties. The core trade-off is a three-way balance between residency guarantees, subprocessor jurisdiction exposure, and operational burden.

Tessera AI covers all three for teams that want managed infrastructure: open-source models like Qwen3.6-35B-A3B run on dedicated EU and LATAM infrastructure, the OpenAI-compatible API lets teams follow standard migrating from OpenAI patterns without rewriting code, and flat monthly pricing keeps high-volume budgeting predictable. For the full compliance picture, including Article 28 contracts and audit logging, see our guide to GDPR-compliant LLM hosting.

Managed inference gateways can enforce jurisdiction-aware routing so EU traffic goes to EU-region backends and never leaves the continent, if the backend selection is correctly configured. The gateway itself adds another vendor subprocessor to the chain, so exposure depends on the gateway design and the providers behind it. Self-hosting open-source models on your own hardware gives the lowest external subprocessor exposure but carries the highest burden: GPU infrastructure, deployment, logging, access control, model serving, and rollback are all customer responsibilities.

FAQ

Is EU data residency the same as GDPR compliance?

No. GDPR permits transfers outside the EEA if you satisfy Articles 44-49. Residency means keeping processing inside the EU; compliance covers the whole legal framework. You must still implement valid transfer mechanisms and supplementary measures where required.

Does hosting in an EU region eliminate US CLOUD Act risks?

Not entirely. Public-cloud EU regions can still expose data to US jurisdiction through provider subprocessors. EU-native providers or dedicated infrastructure reduce that exposure by operating strictly under European law. Verify each provider’s subprocessors and operational disclosures regardless of the hosting model.

How does EU hosting affect latency for European users?

It removes a transatlantic round trip from every request, cutting network overhead before the model starts generating. Routing between European regions adds only a small fraction of that cost. Most user-visible delay comes from model inference time, so in-region hosting eliminates the avoidable network share.

Can I use open-source models with EU data residency?

Yes. Self-hosting open-source models on EU infrastructure gives you full control over the inference path. Managed EU-native providers like Tessera AI serve open-source models such as Qwen3.6-35B-A3B on EU infrastructure, and managed inference gateways can route open-source workloads to EU-region backends while maintaining jurisdiction-aware policies.