Private AI on Azure OpenAI and Amazon Bedrock in Canadian Regions
A private AI deployment runs language models inside your own Azure or AWS account, in Canadian regions where the model allows, with network controls and logging you manage. Remolda deploys one workload in six weeks for $9,800 CAD + HST.
In short
- Price: $9,800 CAD + HST for the six-week AI Pilot Sprint: one AI workload deployed in your cloud tenant.
- Azure OpenAI: Standard deployments in Canada East keep prompts and responses in the Canadian geography for listed models (gpt-4o, gpt-4.1-mini, embeddings); Global deployments may process anywhere.
- Amazon Bedrock: Claude models are called from Canada (Central) with logs and data at rest in Canada; inference runs through cross-region profiles in other regions.
- Setup includes private networking, managed identities, key vault, logging and budget alerts.
- You receive a data-flow record that states, per component, where data is stored and where it is processed.
Your situation
A private deployment comes up when AI must run inside your own cloud controls. Typical cases:
- Security asks where prompts go. The answer needs to name regions, services and retention.
- A Quebec privacy impact assessment is due. It needs a factual map of data flows outside Quebec.
- An application needs a model endpoint. Developers want an API inside the company's Azure or AWS account.
- Sensitive workloads. Health, financial or government data with contractual residency terms.
What gets deployed
| Layer | Azure | AWS |
|---|---|---|
| Model endpoint | Azure OpenAI in Foundry, Canada East or Central | Amazon Bedrock from Canada (Central) |
| Network | Private endpoints, VNet integration | VPC endpoints (PrivateLink) |
| Identity | Managed identities, Entra ID roles | IAM roles, least privilege |
| Secrets | Key Vault | Secrets Manager / KMS |
| Retrieval data | Azure AI Search in-region | Bedrock Knowledge Bases in-region |
| Monitoring | Azure Monitor, budget alerts | CloudWatch, AWS Budgets |
Model-to-region availability is checked against Microsoft's and AWS's tables at build time, because it changes every few months. The feature that uses the endpoint is built through LLM API integration or integration with existing systems; ongoing monitoring can move to AIOps.
Canadian data residency, stated plainly
- Azure OpenAI. Data at rest stays in the chosen Azure geography. Standard and Regional Provisioned deployments process prompts in that geography; Data Zone deployments cover US, EU or APAC only; Global deployments may process in any region.
- Amazon Bedrock with Claude. Called from Canada (Central) with data at rest in Canada; inference runs in other regions through cross-region profiles.
- Quebec Law 25, s. 17. A privacy impact assessment and a written agreement are required before personal information is communicated outside Quebec.
What the AI Pilot Sprint includes
Architecture and data-flow record. Regions, services and retention for each component.
One workload deployed. Endpoint, networking, identity, secrets and monitoring in your tenant.
Model and region test. Candidate models compared for quality, cost and processing location.
Infrastructure as code. Terraform or Bicep, so the setup is repeatable.
Handover. Runbook and a session for your cloud team.
How long it takes
Six weeks from kickoff. In our experience the usual split is two weeks for access, quotas and security review, two weeks of build, and two weeks of testing with the first application. Quota approvals for specific models can take longer.
What it costs
The AI Pilot Sprint is $9,800 CAD + HST, fixed, for one workload. Cloud and model usage are billed under your Microsoft or AWS agreement. Choosing between providers first? See AI vendor selection. All packages are on the pricing page.
Why Remolda for private AI in Canada
- Facts from the vendor tables. Residency statements cite Microsoft and AWS documentation checked at build time.
- Your tenant. Everything runs under your cloud agreement and controls.
- Privacy assessment ready. The data-flow record supports Law 25 and PIPEDA work.
- Repeatable setup. Infrastructure as code.
- Fixed price. One workload, six weeks.
How the work runs
The Sprint covers Audit, Strategy and Implement in the Remolda Cycle (Audit → Strategy → Implement → Empower → Evolve).
- Scope call. 30 minutes: workload, cloud, data.
- Architecture. Regions and services chosen, data-flow record drafted.
- Build. Infrastructure as code, security controls.
- First application. Connected and tested.
- Review. Cost, latency and residency confirmed, then go / adjust / stop.
Frequently asked questions
Can we run Azure OpenAI in Canada?
Yes, with limits that change often. As of September 2026, Microsoft lists Standard (regional) deployments in Canada East for gpt-4o, gpt-4.1-mini and embedding models, which process prompts within the Canadian geography. Newer models such as gpt-5.x are offered in Canada through Regional Provisioned deployments (reserved capacity, processed in the geography) or Global deployments, which may process outside Canada.
Does Amazon Bedrock process Claude requests in Canada?
AWS lists in-region inference in Canada (Central) as not supported for current Claude models. Requests from ca-central-1 use cross-region inference; AWS states that data at rest, logs and knowledge bases stay in Canada (Central) while processing may occur in another region.
How much does a private AI deployment cost?
Remolda's six-week AI Pilot Sprint is $9,800 CAD + HST to deploy one workload with networking, security and monitoring. Cloud and model usage are billed by Microsoft or AWS under your agreement.
Does Quebec Law 25 allow processing outside Quebec?
It requires a privacy impact assessment before communicating personal information outside Quebec, including when a provider outside Quebec processes it on your behalf, and a written agreement. The data-flow record we produce feeds that assessment.
When are ChatGPT Enterprise or Copilot enough?
For staff productivity, often. A private deployment fits when you build your own application, need control over networking and logs, or must pin processing to a region for specific models.
Can we run open-weight models instead?
Yes. Azure AI Foundry and Bedrock host models such as Llama and Mistral, and some teams run them on their own GPUs. We compare quality, cost and hosting effort on your use case.
How long does it take?
Six weeks from kickoff. In our experience, cloud subscription access, quota requests and security review take the first two weeks.
Sources
- Microsoft Learn — Deployment types for Foundry Models (data processing location)
- Microsoft Learn — Model region availability (Azure OpenAI in Foundry)
- AWS — Amazon Bedrock model support by Region
- AWS Machine Learning Blog — Amazon Bedrock cross-Region inference in Canada (Nov 24, 2025)
- LégisQuébec — Act respecting the protection of personal information in the private sector (P-39.1), s. 17
Facts checked:
Related services
Approach phases
Related insights
LLM Integration into Existing CRM and ERP Software in Canada: Architecture, Data Residency and Cost
Mitacs AI Advantage: Ottawa Commits $162M to 10,000 AI Work Placements
AI for Canadian Municipalities: Where It Works in 2026
Talk to an AI transformation consultant
A 30-minute call: you describe the situation, we tell you what to do first and what it would cost.
Book a 30-min call30 minutes. English or French.