Key Takeaways
- Data residency is a requirement about where data physically sits. Most vendors answer it by opening a regional instance, which satisfies the contract question and leaves the architectural one untouched.
- Residency and sovereignty are different constraints. Data can sit in Frankfurt and still be subject to a foreign jurisdiction's legal reach, depending on who operates the infrastructure and who holds the keys.
- The rules differ in kind, not just in strictness. France requires EEA-only storage under HDS certification, Germany allows a list of permitted areas but demands a German establishment, Australia prohibits offshore processing outright for its national record system, and HIPAA imposes no location requirement at all.
- The requirement shapes at least four decisions: whether you run one deployment or one per region, where identifiers and derived data live, where model inference executes, and what happens to cross-region analytics.
- Cross-region aggregation is where most architectures quietly break. Cohort analysis, population baselines, and score calibration all want data pooled in one place, which is the thing residency prevents.
- An AI layer can undo the whole arrangement in a single call. Health data kept in-region and then sent to a hosted model endpoint elsewhere has left the region, whatever the database diagram says.
- Retrofitting residency after launch means changing data topology, not configuration. Teams that treat it as a deployment flag discover the cost when the first enterprise buyer asks a precise question.
Is Your HealthTech Product Built for Success in Digital Health?
.avif)
Why Data Residency Is an Architecture Problem
"We have an EU region now, so data residency is handled." That sentence ends the conversation in a US sales meeting and starts a much longer one in a European procurement review. Standing up a second instance in Frankfurt answers where the rows are stored. It says nothing about who can compel access to them, where the derived metrics are computed, or what happens when the analytics pipeline needs data from both regions at once.
Data residency requirements arrive as a compliance line item and land as an architecture problem. For teams handling health data, the gap between those two things is wider than usual, because health data carries stricter processing conditions under GDPR as a special category, sits under HIPAA's own safeguards in the US, and increasingly passes through model inference somewhere in the stack. Our GDPR for HealthTech and HIPAA vs GDPR pieces cover the regulatory mechanics; this one covers what the requirement does to the system you are building.
What follows is the set of decisions residency actually forces, in the order they usually surface: what the term means next to sovereignty, where the requirements come from, the four architectural choices that follow, and the two places (cross-region analytics and the AI layer) where teams most often find out they got it wrong.
What Data Residency Actually Means
Data residency is a requirement that specified data be stored, and often processed, within a defined geographic boundary. The boundary is usually a country, sometimes an economic area like the EEA, occasionally a single named data center for a single named customer.
The requirement can come from three different places, and they behave differently. Law imposes it directly in some jurisdictions, most often for health records and government workloads. Contracts impose it more often than law does, through a customer's own data processing agreement. Internal policy imposes it in the remaining cases, usually because a security team decided it was the cleanest answer to a question they did not want to litigate per vendor.
None of those three sources tells you how to build for it. A law says the data stays in-country. A DPA says the processing happens in the EEA. Neither says whether that means one deployment per region, one deployment with partitioned storage, or one deployment that happens to run in the right place. That choice is yours, and it has consequences that outlast the contract that prompted it.
Data Residency vs. Data Sovereignty
These two get used as synonyms in vendor marketing and they are not the same constraint. Residency is about location. Sovereignty is about jurisdiction, meaning which legal system can compel access to the data and under what process.
The distinction matters because they can diverge. Data stored in Frankfurt on infrastructure operated by a US-headquartered cloud provider is resident in Germany. Whether it is sovereign to the EU is a separate question, one that depends on the provider's corporate structure, the applicable extraterritorial statutes, and who holds the encryption keys. A European hospital's security review will ask the second question, not just the first, and answering it with a region name does not land.
For health data this shows up as a practical test rather than a theoretical one. If a foreign authority served a lawful order on your infrastructure provider, could that provider produce readable patient data without your involvement? If the answer is yes, you have residency without sovereignty. Key management, operator identity, and deployment control are what separate the two, which is why the question tends to push teams toward infrastructure they run themselves rather than infrastructure they rent access to. Momentum's own self-hosted infrastructure and health data compliance piece covers what that shift does and does not settle.
Where Data Residency Requirements Come From
Four sources produce most of the residency requirements a health product will meet.
GDPR, indirectly. GDPR does not mandate that personal data stay in the EU. It restricts transfers to countries without an adequacy decision, and it treats health data as a special category with stricter processing conditions on top. In practice, many organizations find keeping data in the EEA cheaper than maintaining transfer safeguards and documenting them for every processor in the chain, so a legal restriction on transfers turns into an operational rule about storage. That is why "GDPR requires data residency" is technically wrong and practically close enough to how procurement behaves.
National and sectoral health data rules. Several jurisdictions impose location requirements on health data specifically, separate from general data protection law. These vary enough that a product selling into multiple European markets can meet different rules in each, which is precisely the situation that breaks a single-region design. The section below sets out what four of them actually say.
Customer contracts. The most common source by volume. An enterprise customer's DPA specifies where processing occurs, names permitted subprocessors, and requires notification before that list changes. This is where a wearable data aggregator sitting in your pipeline becomes a problem: it is a subprocessor with its own hosting decisions, and you inherit its answer to the residency question rather than making your own. Teams hit this when they outgrow SaaS wearable aggregators for reasons that started out unrelated to geography.
Sector and public-sector procurement. Hospital groups, national health services, and insurers frequently carry their own hosting requirements, applied uniformly to every vendor regardless of size. A small vendor cannot negotiate these the way it might negotiate an SLA, and passing a security review as a small supplier usually means meeting them rather than arguing about them.
What Specific Jurisdictions Actually Require
Generic advice about residency is less useful than four concrete examples, because the requirements differ in kind rather than in degree. One is a certification regime with a storage boundary attached, one is a permitted-area list with an establishment condition, one is an outright prohibition, and one is not in force yet.
France: certification plus an EEA storage boundary. Hosting personal health data collected in France requires HDS certification under article L1111-8 of the Code de la santé publique. The current certification referential, version 2.0, approved by the arrêté of 26 April 2024, added a storage-location requirement that the earlier version did not have. Requirement 28 states that where a hosting activity involves storing personal health data, the host or its subcontractors must store that data exclusively within the European Economic Area, and must document the storage location and disclose it to the client. Remote access from outside the EEA is permitted only on the basis of an adequacy decision under article 45 of GDPR or, failing that, the safeguards under article 46, and hosts must publish a mapping of transfers outside the EEA. Anyone working from the older referential will describe HDS as certification of the host regardless of location, which is no longer accurate.
Germany: a permitted-area list plus a German establishment. Section 393 of the Sozialgesetzbuch V governs cloud use in healthcare. Paragraph 2 permits processing of social and health data via a cloud service in Germany, in an EU member state, in a state treated as equivalent, or in a third country covered by an adequacy decision under article 45 of GDPR, and adds a condition the other regimes do not have: the processing entity must have an establishment in Germany. Paragraph 4 also requires a C5 attestation, tightened from Type 1 to Type 2 with effect from 1 July 2025. The location rule here is a list of permitted areas rather than a strict in-country mandate, which makes it more permissive than France on geography and stricter on corporate presence.
Australia: an outright prohibition, for a defined system. Section 77 of the My Health Records Act 2012 prohibits the System Operator and registered repository, portal and contracted service providers from holding or taking My Health Record records outside Australia, or processing or handling the related information outside Australia, or causing another person to do either. The exception is narrow and applies only to the System Operator, only for records containing no personal or identifying information about a healthcare recipient. Breach carries imprisonment for five years or 300 penalty units under the fault-based offence, with a civil penalty of 1,500 penalty units. The scope matters: this binds participants in the My Health Record system specifically, and it is not a general rule that all Australian health data must stay onshore.
The European Health Data Space: an EU-level requirement, from 2029. Regulation (EU) 2025/327 introduces the first EU-level location requirement for health data, and it applies to secondary use rather than to running a product. Article 87 requires health data access bodies, trusted health data holders and the Union health data access service to store and process personal electronic health data in the Union when carrying out pseudonymisation, anonymisation and the other processing operations listed in the regulation, through secure processing environments or HealthData@EU, and extends that requirement to any entity performing those tasks on their behalf. Storage in a third country covered by an adequacy decision is permitted. For primary use, article 86 takes the opposite approach: it preserves the ability of member states to require storage in the Union rather than imposing it. The timing is the part worth planning around, because the regulation applies from 26 March 2027 and the secondary-use chapter containing article 87 applies from 26 March 2029.
Two patterns are worth taking from this. The requirements are not interchangeable, so a system designed to satisfy the strictest one does not automatically satisfy the others: Germany's establishment condition is not a geography problem at all, and France's disclosure obligations survive a correct storage decision. And the rules move, which is why a design that treats the permitted region as configuration rather than as an architectural assumption ages better than one that hardcodes today's answer.
The Four Architecture Choices Residency Forces
Once a residency requirement is real, four decisions follow. Teams usually make the first one deliberately and the other three by accident.
One Deployment per Region, or One Deployment That Runs in a Region
A single deployment hosted in the required region satisfies a single-market requirement completely, and it is the right answer for a product selling into one jurisdiction. It stops working the moment a second market arrives with a different rule, because the choices at that point are to migrate everyone into the stricter regime or to split.
Separate deployments per region satisfy any combination of rules and cost you a multiple of the operational work: every release ships more than once, every schema migration runs more than once, and every incident has a scope question attached before anyone can start debugging. Teams underestimate this in proportion to how automated their deployment already is, which is why the cost lands later than expected rather than never.
The middle option, one logical system with storage partitioned by region, reads well on a diagram and is the hardest to hold. Every new query path is an opportunity to route data across a boundary that was supposed to be enforced, and enforcement lives in application logic rather than in topology. It works when the team treats the boundary as a hard invariant with tests behind it, and it degrades quietly when the boundary is a convention.
Where the Identifiers Live
Residency conversations focus on measurements and skip identity, which is where the actual regulatory exposure sits. A per-region deployment still needs to know that a given account exists, and a central account directory holding names, emails, and region assignments is a store of personal data in whatever jurisdiction it happens to run.
The workable pattern separates identity from measurement, keeps the identifiable record in-region, and lets the central layer hold only opaque references that mean nothing without the regional store. That is more work than a shared users table, and it is the difference between a residency claim that survives a data protection audit and one that survives a sales call. Our piece on what counts as PHI under HIPAA is a useful check on which fields carry the exposure.
Where Derived Data Is Computed and Stored
Raw wearable data is only the input. A health product generates derived data on top of it: daily aggregates, baselines, health scores, trend lines, model features. Each of those is computed somewhere, and the computation reads the raw data.
If a score is calculated in-region and stored in-region, the derived data inherits the residency of its input and nothing changes. If the calculation runs in a central job that pulls raw data from every region into one worker, the raw data has left its region during processing, whatever the storage diagram says. Processing location is part of most residency requirements, not an exception to them, and this is the most common way a compliant-looking system fails a precise question.
Where the Boundary Is Enforced
The last choice is where enforcement actually lives, and it determines whether the other three hold up. Enforcement in infrastructure means a service in one region has no network path and no credentials to reach another region's data, so a mistake in application code cannot cross the boundary. Enforcement in application logic means the code is expected to check a region field before it queries, and the boundary holds as long as every path remembers.
Infrastructure-level enforcement is the one that answers an auditor's question with a configuration rather than a promise. It is also the one that constrains what the product can do later, which is the tradeoff worth making consciously rather than discovering.
A self-hosted data layer changes the shape of these four decisions rather than removing them. Running the deployment yourself means the region, the topology, the key management, and the enforcement boundary are all yours to choose, and the constraints come from your architecture instead of from a vendor's regional roadmap. Momentum built Open Wearables as a self-hosted wearable data layer for exactly that reason, and the tradeoff is honest: the decisions above become yours to make and yours to operate.
Cross-Region Analytics Is Where This Usually Breaks
Every residency architecture works until someone needs a number that spans regions. The requests are ordinary and they arrive from every direction: how many active users across all markets, what the population baseline looks like, whether the sleep model performs differently in one country than another, what the cohort retention curve says. Each of those wants data pooled in one place, and pooling is the thing residency prevents.
There are two honest answers and one that gets teams in trouble.
Aggregate in-region and combine only the aggregates. Each region computes its own counts, distributions and model metrics locally, and a central layer receives only numbers that no longer describe individuals. This preserves the boundary and costs you granularity: you can compare regional distributions, and you cannot re-slice the combined population by an attribute nobody thought to aggregate on. Most analytical questions survive that constraint. Some do not.
Move a properly de-identified extract. If the data leaving a region is anonymized to the standard the applicable regulation actually requires, it is no longer personal data and the transfer restriction stops applying. The catch is that this standard is much higher than dropping a name column, and health data with high-resolution timestamps is notoriously easy to re-identify. Treating pseudonymized data as anonymized is a common and consequential mistake, since pseudonymized data remains personal data under GDPR.
The answer that causes problems is a central analytics warehouse that quietly ingests raw records from every region because the pipeline was built before the residency requirement existed. It usually predates the compliance conversation, it rarely appears on the architecture diagram anyone shows the auditor, and it is the single most likely thing to surface during a serious review.
Model training deserves its own note. Training on pooled multi-region data is the same transfer as any other, and a model trained on personal data can retain traces of it. Our piece on creating synthetic training data for healthcare AI covers what to do when the dataset you are allowed to use is smaller than the one you want.
The AI Layer Undoes This Quietly
A team can get all four architectural decisions right and hand the whole arrangement back with one API call.
The pattern is common enough to be predictable. The data layer is self-hosted in-region, identity is separated from measurement, scores are computed locally, and the boundary is enforced in infrastructure rather than in application code. Then a coaching feature, a natural-language query interface, or a triage summary is wired to a hosted model endpoint, and health data from a regulated region is now in a POST body headed somewhere else. Nothing in the database diagram changed. The residency claim is no longer true.
Inference location is processing location. If the model runs outside the region, the data was processed outside the region, and the fact that the provider does not retain it changes the retention analysis rather than the transfer analysis. This is the newest version of the problem and the one where diligence questions are getting sharper fastest, because AI features ship faster than the compliance review around them.
The options are the ordinary ones: run the model in-region, whether self-hosted or through a provider offering genuine regional inference with contractual commitments behind it; minimize what reaches the model so nothing identifiable leaves; or keep the feature in-region only and accept it launches per market. Our pieces on building secure AI models for HealthTech and AI agent compliance in healthcare go into what defensible looks like here. What does not work is treating the AI layer as an application concern once the data layer has been declared sovereign.
What Changes When You Enter a Second Market
Most teams meet residency properly at their second market, not their first, and the sequence is consistent.
The first market sets a topology, usually one deployment in one region, and nobody calls it a residency decision because nothing forced the question. Then a customer in another jurisdiction arrives with a DPA that names a different location. At that point the honest options are to run a second deployment, restructure into a partitioned single system, or decline the requirement.
What makes this expensive is that the decision is about data topology, not configuration. Splitting a system that assumed one shared database means separating identity from measurement, deciding where derived data is computed, rebuilding analytics to aggregate before it combines, and doing all of it while the first market's product keeps running. None of that is a deployment flag. Teams that considered the second market before writing the first one's schema pay a fraction of what teams that did not end up paying.
The question worth answering early is narrow: if a customer in another jurisdiction demanded their data stay there, what in the current system would have to change? A clear answer means the design already accounts for it. An unclear answer is the finding.
Where to Start
If a residency requirement is already on the table, the useful first step is establishing where data currently sits, where it is processed, and which of the four decisions above the current architecture has made implicitly. That map is usually short and frequently surprising, particularly around analytics pipelines and model calls.
Talk to us if you want a second pair of eyes on what your architecture currently commits you to, before a customer's security review finds it first.




