LLM integration patterns for enterprise architecture fall into six types: direct API calls, an LLM gateway, retrieval-augmented generation (RAG), tool calling, fine-tuning and self-hosted models. Most production systems combine a gateway, RAG and tool calling.
This guide covers architecture and governance: how the pieces fit, where data goes and who controls it. For the hands-on build (SDKs, current models, API costs and step-by-step implementation), see our practical LLM integration guide for 2026.
What are the main LLM integration patterns in enterprise architecture?
Six patterns cover almost every enterprise case. They stack. A typical internal assistant uses a gateway, RAG over policies and tool calls into the ticketing system.
| Pattern | What it is | Use it when | Main risk |
|---|---|---|---|
| Direct API | Application calls a provider endpoint | Prototypes, low-sensitivity data | Keys spread across apps, no central logging |
| LLM gateway | One internal service in front of all models | Two or more apps or providers | A new component to run and secure |
| RAG | Retrieve documents, pass them as context | Knowledge changes often, citations needed | Stale or over-permissive index |
| Tool calling | Model calls your APIs through defined functions | The AI must read or update CRM, ERP, tickets | Actions taken on bad instructions |
| Fine-tuning | Adjust model weights on your examples | Narrow format or vocabulary, high volume | Retraining when the base model retires |
| Self-hosted open-weight | Model runs in your cloud or data centre | Data may not leave your environment | Capability gap, operations cost |
Two of these are architecture decisions that are hard to reverse: the gateway and the choice of where inference runs. Settle them first.
How does an LLM gateway fit into enterprise architecture?
An LLM gateway sits between your applications and every model provider. Applications call one internal endpoint; the gateway decides which model answers and records what happened.
A reference flow looks like this:
- Application (portal, CRM plug-in, internal chat) sends a request with the user's identity.
- Gateway checks the caller, strips or masks personal information, applies the prompt template and a cost limit.
- Router picks the model: a small model for classification, a larger one for reasoning, a Canadian-region deployment for regulated data.
- Provider returns the response, which passes through output checks (format, blocked content, citation present).
- Log store keeps timestamp, model version, user, tools called and a hash or redacted copy of the prompt.
Commercial and open-source gateways exist, and cloud platforms offer their own. The pattern matters more than the product. What counts is one place for keys, logs, budgets and policy.
When should you use RAG, tool calling or fine-tuning?
Use RAG for knowledge, tool calling for actions and fine-tuning for form. Most teams reach for fine-tuning too early.
- RAG fits policies, contracts, product documentation and case files. Answers can cite their source, and access rules from the source system can be enforced at retrieval time. The index must respect permissions. A retrieval layer that ignores them will show HR files to the whole company.
- Tool calling fits "look up the order", "create the ticket", "update the CRM record". Define each tool with a narrow schema, give the service account the smallest possible permissions and require human approval for writes that move money or change records of consequence. The Model Context Protocol (MCP) is one common standard for exposing tools to models.
- Fine-tuning fits a fixed output format or specialist vocabulary at high volume, after prompting and RAG have been tried. The comparison is covered in RAG vs fine-tuning for enterprise.
Which external LLM provider options keep data in Canada?
As of September 2026, Canadian residency is available for storage in several places and for processing in a few. Check each model: residency is set per deployment, and the tables change often.
| Provider path | Data at rest in Canada | Prompts processed in Canada | Notes |
|---|---|---|---|
| Azure OpenAI (Microsoft Foundry) | Yes, in the chosen Azure geography | Canada East Standard: gpt-4o, gpt-4.1-mini, text embeddings. Regional provisioned (PTU) in Canada Central and Canada East for gpt-4o, gpt-4.1, gpt-5.x and others | Global deployments may process anywhere; there is no Canada data zone |
| OpenAI API and ChatGPT Enterprise | Yes, for eligible customers (ca.api.openai.com, via sales) | No; in-region inference is offered for Europe, US and UAE only | Requires Modified Abuse Monitoring or Zero Data Retention |
| Anthropic API (first party) | No | No; inference_geo is "global" or "us" | "us" costs 1.1x |
| Claude on Amazon Bedrock (ca-central-1) | Yes: logs, knowledge bases, configuration | No; the Canada geo profile routes inference to US Regions | In-Region inference "not-supported" for current Claude models |
| Gemini on Google Vertex AI (northamerica-northeast1, Montréal) | Yes | Yes for Gemini 3.5 Flash, Gemini 2.5 Flash and Gemini 2.5 Pro; newer 3.6–3.8 Flash models only in US and EU multi-regions | Claude on Vertex has no Canada processing column |
| Self-hosted open-weight model | Yes | Yes | You run capacity, patching and evaluation |
For many Canadian organizations the workable split is this: regulated data goes to an in-Canada deployment, and low-sensitivity work goes to the global endpoint of the strongest model. The gateway enforces that split by data class.
How do you secure an LLM API integration for financial or personal data?
Secure it in the contract, in the data path and in the logs. Each layer covers what the others miss.
Contract. OpenAI states that it does not use data from ChatGPT Business, Enterprise or the API platform for training by default. Anthropic says the same for Claude for Work and the API. Google says Workspace and Vertex AI customer data is not used to train its models without prior permission or instruction. Get the data processing addendum signed, set retention and ask for zero data retention where it is offered.
Data path.
- Mask names, account numbers and health identifiers before the prompt leaves your network; re-insert them after the response if the user needs them.
- Use private networking to the cloud provider where available.
- Keep API keys in your secrets manager, rotate them, and give each application its own key through the gateway.
- Treat retrieved documents and web content as untrusted input. Prompt injection arrives through them.
Logs. Record model version, user, tools called and outcome for every call. Store prompts redacted or hashed when they contain personal information. The log is the first thing a security review or a privacy complaint will ask for.
What legal considerations apply when integrating an LLM into existing workflows?
In Canada, privacy law applies to every prompt that contains personal information, and there is no stand-alone federal AI act. Bill C-27 and its Artificial Intelligence and Data Act died in January 2025.
- PIPEDA governs private-sector personal information. The OPC's 2023 principles for generative AI stress legal authority and consent, openness, limiting collection and safeguards.
- Quebec Law 25. Section 3.3 requires a privacy impact assessment for any project to acquire, develop or overhaul an information system involving personal information. Section 17 requires another assessment and a written agreement before personal information goes outside Quebec, including to a cloud or AI vendor. Section 12.1 covers decisions made exclusively by automated processing.
- Federal institutions follow the Treasury Board Directive on Automated Decision-Making, including an Algorithmic Impact Assessment before production.
- Pending: Bill C-36, the Protecting Privacy and Consumer Data Act, has been at first reading since June 15, 2026. It would add transparency duties for automated decision-making.
An AI compliance review for PIPEDA and Law 25 maps these duties onto a specific architecture before build.
How do you govern LLMs once they are in production?
Pin versions, test every change and plan for deprecation. Model behaviour shifts between versions without a code change on your side.
- Pin model versions in the gateway configuration. Move through a planned migration.
- Keep an evaluation set of real inputs with expected outputs. Run it before each model or prompt change.
- Track deprecation dates for every model in use and put migrations in the annual plan.
- Name owners: one business owner per use case, one technical owner for the gateway, one privacy contact.
- Review monthly: cost per use case, error reports, blocked requests, retrieval misses.
Where should an enterprise start?
Start with an inventory and one pattern decision. List the applications that already call LLMs, the data classes they touch and the providers in use. Decide where inference may run for each data class. Then put a gateway in front of the first production use case.
Remolda designs and builds LLM integration with existing systems such as CRM, ERP and document stores, and helps with choosing between Copilot, ChatGPT Enterprise and Claude. Work starts with fixed-price packages from $490 CAD, with scope agreed in writing.
Sources
- Microsoft Learn — Deployment types for Foundry models
- Microsoft Learn — Model region availability
- OpenAI — Business data privacy
- OpenAI — Your data (API data residency)
- Anthropic — Data residency
- Anthropic Privacy Center — Is my data used for model training?
- AWS — Amazon Bedrock model Region compatibility
- AWS blog — Amazon Bedrock cross-Region inference in Canada
- Google Cloud — Generative AI data residency
- Google Cloud — Generative AI data governance (training restriction)
- Google Workspace — Generative AI in Google Workspace Privacy Hub
- OPC — Principles for responsible, trustworthy and privacy-protective generative AI
- LégisQuébec — CQLR c. P-39.1
- Treasury Board — Directive on Automated Decision-Making
- Parliament of Canada — Bill C-36