To integrate an LLM into existing business software in 2026, pick one measurable task, call a current model through its official API, validate and log every response, and test on real examples before release. Seven steps cover it.
This is the hands-on guide: models, prices, Python patterns and steps. Architecture questions (gateways, RAG at scale, Canadian data residency, governance) are in our enterprise LLM integration patterns guide.
How do you integrate an LLM into existing business software?
Follow seven steps, in order. Steps 1 and 5 carry the most weight.
- Pick one task with a number attached. "Draft first replies to support emails" beats "add AI to the CRM". Write down today's time per item and error rate.
- Map the data. What goes into the prompt, where it comes from, whether it contains personal information and who may see it.
- Choose model and access path. Vendor API, or a cloud platform (Azure, Amazon Bedrock, Google Cloud) where that model is offered and your contracts and residency needs point there.
- Build a thin integration layer. One module owns prompts, model calls, validation, retries and logging. Application code calls that module only.
- Build a test set. 50 to 200 real, anonymized examples with the expected answer. Score every prompt or model change against it.
- Ship behind a flag. Start with a pilot group and a human review step. Watch cost, latency and error reports daily for two weeks.
- Hand over. Document prompts, model versions, owners and the rollback plan.
Which LLM models can you integrate directly in 2026?
Three vendor families cover most business integrations, each with a large, a mid and a small tier. Prices below are standard API list prices per million tokens, as of September 2026.
| Vendor | Model | Input (US$/1M) | Output (US$/1M) | Typical use |
|---|---|---|---|---|
| OpenAI | GPT-6 Astra | 10.00 | 50.00 | Hardest reasoning, low volume |
| OpenAI | GPT-6 Sol | 2.00 | 10.00 | Coding, agents, general work |
| OpenAI | GPT-6 Luna | 0.10 | 0.50 | Classification, extraction at volume |
| Anthropic | Claude Opus 5.5 | 4.00 | 20.00 | Long documents, complex analysis, agents |
| Anthropic | Claude Sonnet 5 | 2.00 | 10.00 | General business workloads |
| Anthropic | Claude Haiku 4.5 | 1.00 | 5.00 | Fast, high-volume tasks |
| Gemini 3.1 Pro (preview) | 2.00 | 12.00 | Long context, multimodal (up to 200k tokens at this rate) | |
| Gemini 3.8 Flash | 0.75 | 3.75 | Mid-tier, through Dec 31, 2026 (then 1.50 / 7.50) | |
| Gemini 3.5 Flash-Lite | 0.30 | 2.50 | Low-cost volume |
The OpenAI and Anthropic batch APIs cost 50% less for work that can wait. Open-weight models (Llama, Mistral, Qwen and others) can run in your own cloud when data may not leave it; you then pay for compute and operations instead of tokens.
Pick by task, then test. Benchmarks rarely match your documents. Run your test set on two or three candidates and keep the cheapest one that passes.
How much does LLM integration cost?
Token cost is usually the smaller part. A worked example at list prices: 10,000 requests a month, 2,000 input tokens and 500 output tokens each (20M input, 5M output).
| Model | Monthly token cost (US$) |
|---|---|
| GPT-6 Luna | about 4.50 |
| Gemini 3.8 Flash | about 34 |
| Claude Haiku 4.5 | about 45 |
| Claude Sonnet 5 / GPT-6 Sol | about 90 |
| Claude Opus 5.5 | about 180 |
| GPT-6 Astra | about 450 |
Three levers change the bill:
- Prompt caching. A repeated system prompt or document is billed at a fraction of the input price on a cache hit (0.05x for Claude Opus 5.5; US$0.20 against US$2.00 for GPT-6 Sol).
- Batch processing. Half price on OpenAI and Anthropic for overnight or background jobs.
- Model routing. Send simple classification to a small model and only hard cases to a large one.
The larger costs are engineering, test sets, integration with existing systems and maintenance when models are retired.
What are the Python LLM integration patterns for 2026?
Six patterns keep a Python integration stable. They apply equally to the OpenAI, Anthropic and Google SDKs.
- Config from the environment. Model name, API key and limits come from environment variables (a
.envfile locally, a secrets manager in production). Changing models becomes a config change. - Structured output. Ask for JSON that matches a schema and validate it (for example with Pydantic). Reject and retry on invalid output.
- Retries with backoff and a fallback. Retry rate-limit and server errors with exponential backoff. After a set number of failures, switch to a second model or send the item to a human queue.
- Streaming for chat interfaces. Stream tokens to the user so a long answer starts appearing at once.
- Tool calling. Expose narrow functions ("get_order", "create_ticket") with typed parameters. The model proposes the call; your code checks permissions and executes.
- Batch for background work. Nightly document runs go through the batch API at half price.
A minimal shape of the integration layer, with the vendor call left abstract:
import os, time
from pydantic import BaseModel, ValidationError
MODEL = os.environ["LLM_MODEL"] # set per environment
MAX_RETRIES = int(os.environ.get("LLM_MAX_RETRIES", "3"))
class TicketSummary(BaseModel):
category: str
urgency: int
summary: str
def summarize(ticket_text: str) -> TicketSummary | None:
for attempt in range(MAX_RETRIES):
try:
raw = call_model(MODEL, prompt_for(ticket_text)) # vendor SDK call
result = TicketSummary.model_validate_json(raw)
log_call(MODEL, attempt, ok=True)
return result
except (ValidationError, TimeoutError) as err:
log_call(MODEL, attempt, ok=False, error=str(err))
time.sleep(2 ** attempt)
return None # caller routes the ticket to a human queue
Keep this module small and owned by one person. Every prompt lives there, with a version number.
How do you handle API keys and authentication?
Keep keys out of code, give each application its own key and route user identity through your own login. That covers most audit questions.
- Local development:
.envfile excluded from the repository. - Production: the cloud secrets manager (Azure Key Vault, AWS Secrets Manager, Google Secret Manager). Rotate keys on a schedule.
- Cloud platforms: on Azure, Bedrock or Vertex AI, use managed identities or IAM roles instead of static keys where possible.
- Users: people sign in to your application through your SSO. The application calls the model with a service identity and logs which user asked.
- Spend limits: set monthly caps per key in the vendor console.
How do you connect an LLM to email, CRM or a chat interface?
Connect through the systems' own APIs and keep a person in the loop for anything sent outside. The common first integrations:
| Integration | What the LLM does | Control to add |
|---|---|---|
| Email (Microsoft 365, Gmail) | Classifies incoming mail, drafts replies, extracts fields | Drafts only; a person sends |
| CRM (Salesforce, HubSpot, Dynamics) | Summarizes account history, fills notes after calls | Writes to a draft field or with approval |
| Help desk | Suggests category, priority and a reply from the knowledge base | Agent accepts or edits |
| Internal chat UI | Answers questions over policies and documents | Citations to source documents |
| Document intake | Extracts data from invoices, forms, contracts | Confidence threshold, human check below it |
For email and document intake at volume, see AI workflow automation. For connecting several systems through well-defined endpoints, see AI API design and integration.
How do you keep an LLM integration stable in production?
Treat the model as an external dependency that changes. Monitor it the way you monitor a payment provider.
- Pin the model version and upgrade on purpose, after the test set passes.
- Log model, prompt version, latency, token count, cost and outcome for every call.
- Alert on error rate, latency and daily spend.
- Collect feedback from users with a one-click "wrong answer" button, and add those cases to the test set.
- Track deprecation dates for each model you use.
- Check privacy duties before personal information enters prompts: PIPEDA federally, and in Quebec a privacy impact assessment under Law 25 for new systems and for transfers outside the province.
When does an LLM integration need an outside team?
A small feature inside one app is a normal job for an in-house developer with this checklist. Work that spans several systems, touches personal information or needs an audit trail benefits from people who have built the test sets, logging and privacy reviews before.
Remolda builds custom AI features into existing software, starting with a scoped pilot. The six-week AI Pilot Sprint is a fixed-price package at $9,800 CAD, and the one-week readiness review starts at $490 CAD.