An enterprise AI system is much more than a model behind an API. A sound enterprise AI architecture separates the system into layers so that each layer has one job, a clear contract with its neighbours, and can be swapped without rewriting the rest. This article is that reference design: each layer's responsibilities, decisions and typical technologies, and how it changes across cloud, hybrid and on-prem deployments.
This is an architecture reference, not a process guide. For how the system gets built in stages with sign-off gates, read how FDEs take AI from POC to production. For what goes wrong when layers are missing, see why AI demos fail in enterprise production. Here we stay with components and boundaries.
The layered reference architecture
The diagram below is a generative AI reference architecture for an organisation that runs several assistants and agents, not one chatbot. The top half is the request path; the bottom half is the shared foundation every use case reuses.
CHANNELS Teams | Web | Mobile | Partner API
|
GATEWAY authN/Z, rate limits, routing
|
ORCHESTRATION agents, workflows, state, memory
|
+----------+--------------+-----------------+
|KNOWLEDGE | MODEL LAYER | TOOLS / INTEGR. |
|ingest, | model gateway| MCP servers, |
|hybrid | hosted + | REST/SOAP APIs, |
|search, | open-weight | ServiceNow,Jira |
|ACLs | models | |
+----------+--------------+-----------------+
SAFETY input/output guardrails, PII
DATA app DB, event and feature stores
PLATFORM K8s/serverless, IaC, CI/CD
OBSERVABILITY traces, metrics, cost, evals
GOVERNANCE registry, policy, audit, risk
Safety, observability and governance are drawn as horizontal bands because they are not steps in the flow. They wrap every call that crosses a layer boundary. LLM application architecture fails at the boundaries, so that is where the controls belong.
The layers, one by one
1. Channels
Responsibilities: meet users where they already work. In most enterprises that means Microsoft Teams or Slack, an existing intranet or portal, a mobile app, and an API for other systems.
Design decisions: keep channels thin. A Teams bot should translate a message into a request and render the response, nothing more. Business logic in a channel adapter gets duplicated the moment a second channel appears. Decide early how citations and approvals render in each channel.
Examples: Teams bot via the Bot Framework, a React web client, a FastAPI endpoint for system-to-system calls.
2. Gateway
Responsibilities: authenticate the caller, establish who the end user is, enforce rate limits and quotas per tenant or team, and route to the right use case or version.
Design decisions: the gateway is where identity enters the system, so it must produce a token or context object that carries the user's identity downstream (more on identity propagation below). It is also the natural place for canary routing to a new prompt or model version.
Examples: Amazon API Gateway, Azure API Management, an ingress controller on Kubernetes, Microsoft Entra ID for OAuth/OIDC.
3. Orchestration
Responsibilities: decide what happens for a request: a single retrieval-and-answer, a deterministic workflow, or an agent that plans, calls tools and loops. Holds conversation state, short-term memory and human-in-the-loop checkpoints.
Design decisions: prefer explicit workflows (a graph with defined nodes and edges) over open-ended agent loops wherever the task is known. Use agents only where the path genuinely varies. Cap iterations, persist state so a long-running task can resume, and put approval steps before any action that changes a system of record.
Examples: LangGraph for stateful graphs and multi-agent patterns, LangChain for components, Amazon Bedrock Agents for a managed option, a Python/FastAPI service hosting the orchestration code.
4. Knowledge
Responsibilities: ingest documents and records, chunk and embed them, keep them fresh, and retrieve the right passages at query time, filtered by what the user is allowed to see. This is the core of RAG architecture in the enterprise.
Design decisions:
- Hybrid search. Vector search finds paraphrases; keyword search (BM25) finds product codes, clause numbers and names. Combine them and add a reranker.
- Permissions at retrieval time. Store the source system's access control list with every chunk and filter on it in the query. Never retrieve first and filter in the prompt.
- Incremental ingestion. Pipelines must handle updates and deletions from the source, not only additions, and record which document version each chunk came from.
Examples: PostgreSQL with pgvector (vectors, metadata and full-text in one database), Amazon OpenSearch or Azure AI Search for larger corpora, Bedrock Knowledge Bases as a managed pipeline. For the concept itself, see what RAG is.
5. Model layer
Responsibilities: provide access to language, embedding and reranking models through one internal interface.
Design decisions: put a model gateway (or router) between the application and providers. It holds credentials, applies per-team budgets, logs tokens, fails over, and routes simple requests to smaller, cheaper models. Without it, every team hardcodes a provider. Decide which classes of data may go to which models, and write it down.
Examples: Amazon Bedrock, Azure OpenAI and Google Gemini on Vertex AI for hosted models; open-weight models served on GPU nodes for workloads that must stay inside your network; a gateway service you build or adopt in front of all of them.
6. Tools and integration
Responsibilities: let agents read from and act on enterprise systems: ticketing, CRM, ERP, code repositories, core banking.
Design decisions: wrap each system once, behind a well-described interface, rather than letting every agent call raw APIs. The Model Context Protocol (MCP), an open protocol introduced by Anthropic in late 2024, standardises how AI applications discover and call tools and read resources, so one ServiceNow MCP server can serve many agents. Underneath, the MCP server still calls the system's normal API, with the user's delegated identity, and with read and write tools separated. The trade-offs are covered in MCP vs API.
Examples: MCP servers for ServiceNow, Jira and GitHub; REST or SOAP clients for legacy systems; message queues for asynchronous actions.
7. Safety
Responsibilities: check inputs and outputs at every boundary: prompt injection detection, PII detection and redaction, topic and content filters, output schema validation, and policy checks before tool calls.
Design decisions: treat retrieved documents and tool outputs as untrusted input, not only user messages. Guardrails are a layer of defence, not a substitute for least-privilege permissions on tools. The full threat model is in AI security for enterprises.
Examples: Amazon Bedrock Guardrails, Azure AI Content Safety, custom validators in the orchestration code.
8. Data
Responsibilities: hold the application's own state (conversations, feedback, approvals), plus events and features that agents use for context, such as a customer's recent transactions or a ticket's history.
Design decisions: separate operational data from logs and traces, set retention per data class, and decide where conversation history lives and who can read it.
Examples: PostgreSQL for application state, an event stream such as Kafka or Amazon EventBridge, a feature store where ML features already exist.
9. Platform
Responsibilities: run everything reliably, repeatably and securely.
Design decisions: Most enterprise AI platform architecture mixes Kubernetes (control, GPU scheduling) with serverless (spiky, stateless pieces). Everything is defined as code, deployed through pipelines, and promoted through environments, with prompts, model configuration and index versions treated as deployable artefacts.
Examples: Docker, Kubernetes (EKS, AKS, GKE), AWS Lambda, Terraform, GitHub Actions, Argo CD for GitOps. See AWS for AI engineers for how these map onto AWS services.
10. Observability and evaluation
Responsibilities: trace every request across layers (retrieved chunks, prompts, model calls, tool calls, latency, tokens, cost), collect user feedback, and run evaluation sets on every change.
Design decisions: instrument with OpenTelemetry so traces flow into the enterprise's existing monitoring, and use an LLM-aware tool on top. Evaluation is part of the CI pipeline, not an occasional exercise. Details are in AI observability.
Examples: OpenTelemetry, LangSmith, Langfuse, Ragas for RAG metrics, CloudWatch or Azure Monitor.
11. Governance
Responsibilities: know which models, prompts, data sources and agents are in production, who owns them, what risk review they passed, and keep an audit trail of what the system did and why.
Design decisions: keep a registry of use cases and the models and data each one uses; require a review before a new use case or data class goes live; keep audit logs immutable.
Examples: a model and prompt registry, policy-as-code in pipelines, centralised audit logging, Microsoft Entra ID groups for role-based access to the platform itself.
Cross-cutting concerns
Identity propagation. The end user's identity must travel from the channel through the gateway, orchestration and into every retrieval and tool call. If an agent calls ServiceNow with a shared service account, it can see and do everything that account can, regardless of who asked. Use token exchange or on-behalf-of flows so downstream systems enforce their own permissions, and log the user identity on every trace.
Multi-tenancy. Whether tenants are business units, countries or external customers, isolate them in the index (separate indexes or a mandatory tenant filter), in conversation storage, in caches and in quotas. A shared cache that ignores tenant is a data leak waiting to happen.
Data residency. Know which region each model endpoint processes data in, where embeddings and logs are stored, and whether any cross-region inference is enabled. For regulated Indian workloads this often means pinning to Indian cloud regions and checking that every managed service in the path honours that.
Cost control. Token usage is attributed per use case and per team at the model gateway, with budgets and alerts. Levers: smaller models, trimmed context, capped agent loops.
Caching. Three distinct caches: provider prompt caching for long, stable system prompts; a response cache for repeated questions (keyed on tenant, user permissions and index version, never on the question text alone); and an embedding cache in ingestion. Invalidate on reindex.
Designing these concerns across a real customer's estate is the heart of Forward Deployed Engineering. If you want to practise it on a full simulated engagement rather than read about it, Cloudsoft's AI Forward Deployed Engineer course takes you through it across five enterprise projects and the GlobalBank capstone.
Build vs buy, layer by layer
No enterprise builds or buys every layer. The table gives a typical starting position.
| Layer | Usually buy / use managed | Usually build or configure |
|---|---|---|
| Channels | Teams, Slack, existing portal | Thin bot or UI adapter |
| Gateway | API Gateway / APIM, Entra ID | Routing rules, quotas, canary config |
| Orchestration | Frameworks (LangGraph), managed agents for simple cases | Workflows, agent graphs, approval steps |
| Knowledge | Vector DB or search service, managed ingestion for simple corpora | Connectors, chunking, ACL sync, freshness |
| Model layer | Hosted models via cloud providers | Model gateway policies, routing logic |
| Tools | Vendor-supplied MCP servers where trustworthy | MCP servers for internal and legacy systems |
| Safety | Managed guardrail and content services | Domain policies, output validators |
| Data | Managed PostgreSQL, event streaming | Schemas, retention rules |
| Platform | Managed Kubernetes, serverless | Terraform modules, pipelines, GitOps |
| Observability / eval | LangSmith or Langfuse, cloud monitoring | Eval sets, dashboards, alerts |
| Governance | Audit logging, policy tooling | Registry, review process, ownership |
A useful test: build where the layer encodes something specific to the business, and buy undifferentiated plumbing.
Three deployment patterns
Pattern A: single-cloud managed
Everything runs in one cloud using its managed AI service: Amazon Bedrock on AWS, Azure OpenAI on Azure, or Gemini on Vertex AI on Google Cloud. Models, guardrails, knowledge bases and agents sit inside the same network, identity and logging boundary as the rest of the estate, which makes this the fastest route for an enterprise standardised on one cloud. The risk is coupling: keep the model gateway and orchestration code provider-agnostic so a later move is a configuration change, not a rewrite. For AWS, Cloudsoft's AWS training and AWS Solutions Architect course cover the networking, IAM and service design this pattern depends on.
Pattern B: hybrid with open-weight models
Hosted models handle general reasoning, while open-weight models run on the enterprise's own Kubernetes GPU nodes for sensitive data, high-volume narrow tasks such as classification or extraction, or cases where a fine-tuned model is needed. The model gateway is what makes this workable: applications call one interface and routing policy decides, by data class and task, which model serves the request. The cost is operational: GPU capacity planning, model serving, patching and evaluating every model you host. Kubernetes training is the foundation for running that serving tier.
Pattern C: on-prem or air-gapped
Some regulated or defence-adjacent workloads cannot send data to any public cloud endpoint. Here every layer runs inside the data centre: open-weight models on local GPUs, a self-hosted vector store, self-hosted observability, and an internal package and model mirror. Plan for limited GPU capacity, slower model upgrades, self-hosted guardrails and longer change approvals. The layered design still holds; only the technology inside each layer changes.
An illustrative bank architecture
Consider a mid-sized bank whose technology team, run from a GCC in Hyderabad, wants three things on one platform: a policy assistant for branch staff, an agent that drafts responses to customer complaints, and an IT-operations agent that triages incidents. Mapped to the layers, an architecture might look like this:
- Channels: Teams for staff; the internal service desk portal for IT operations.
- Gateway: Azure API Management with Microsoft Entra ID; quotas per business unit.
- Orchestration: LangGraph services on EKS; the complaints flow is a fixed workflow with a mandatory human approval node before anything reaches a customer.
- Knowledge: PostgreSQL with pgvector holding policy circulars and product terms, with branch and role ACLs synced nightly from the document management system and an effective-date filter so superseded circulars are never cited.
- Model layer: a model gateway routing to Amazon Bedrock models in an Indian AWS region; complaint text with account details is masked before it leaves the bank's VPC.
- Tools: MCP servers for ServiceNow (incidents) and the complaints system, read tools separated from write tools, with write tools requiring an approval token.
- Safety: Bedrock Guardrails plus a bank-specific validator that blocks any response containing an account number.
- Data: conversation and approval records in PostgreSQL with retention aligned to the bank's records policy.
- Platform: Terraform, GitHub Actions, Argo CD, separate dev, UAT and production accounts.
- Observability: OpenTelemetry traces into Langfuse and CloudWatch; Ragas evaluation in CI against a versioned test set.
- Governance: a use-case registry reviewed by the risk team, with immutable audit logs of every tool action and the user it was performed for.
The point is reuse: the second and third use cases plug into layers that already exist.
Architecture anti-patterns
- The monolith chatbot. Channel, prompt, retrieval and tool calls in one service. Every change is risky and nothing is reusable.
- Provider SDKs everywhere. Each team calls a model provider directly, so there is no central cost view, no failover and no easy model change.
- Service-account agents. Agents act with a powerful shared identity instead of the user's delegated one.
- Filter-in-the-prompt permissions. Retrieving restricted documents and asking the model not to show them.
- Agents for everything. Open-ended agent loops for tasks a three-step workflow would handle more cheaply and predictably.
- Vector-only retrieval. Missing exact-match search for identifiers and codes.
- Observability as an afterthought. No traces, so a bad answer cannot be explained.
Frequently asked questions
What is enterprise AI architecture?
Enterprise AI architecture is the set of layers, components and boundaries that let an organisation run AI assistants and agents securely and reliably at scale. It covers channels, gateway, orchestration, knowledge retrieval, models, tool integration, safety, data, platform, observability and governance, plus cross-cutting concerns such as identity, residency and cost.
What is a model gateway and do I need one?
A model gateway is an internal service between your applications and model providers. It handles credentials, routing, failover, token logging and budgets. Once more than one team or model is involved, you need one to control cost and change models without code changes.
Should I use agents or workflows?
Use a deterministic workflow whenever the steps are known in advance, and an agent only where the path genuinely depends on what the model discovers. Many production systems are workflows with one or two agentic steps inside.
Where should document permissions be enforced in a RAG architecture?
At retrieval time, in the search query itself. Each chunk carries the access control information from its source system, and the query filters on the current user's identity and groups. Restricted content should never reach the model's context window.
Where does MCP fit in an enterprise AI architecture?
MCP sits in the tools and integration layer. An MCP server wraps an enterprise system such as ServiceNow or Jira and exposes its capabilities as tools that any MCP-compatible agent can discover and call. It does not replace the system's API; it standardises how AI applications use it.
Kubernetes or serverless for AI applications?
Most enterprises use both. Kubernetes suits long-running orchestration services, self-hosted models on GPUs and workloads that need fine control. Serverless suits event-driven ingestion and spiky, stateless functions. Choose per component.
How is a hybrid AI architecture different from a single-cloud one?
A hybrid architecture combines hosted models from a cloud provider with open-weight models you run yourself, usually to keep sensitive data in your network or reduce cost on high-volume tasks. A model gateway routes each request by data class and task.
What skills does an engineer need to design this architecture?
Cloud architecture and networking, identity and access management, Python service development, RAG and agent design, Kubernetes and infrastructure as code, CI/CD, observability and evaluation, and the ability to work with security and risk teams.
Enterprise architecture is core Forward Deployed Engineering work: discovering a customer's constraints, designing the layers, and integrating them with real systems. If you want to learn it hands-on, the Cloudsoft FDE PRO program runs 12 weeks with 60+ labs, five enterprise projects including a ServiceNow AI agent via MCP and a Secure Banking AI Assistant, and the GlobalBank capstone. Join in the classroom at Ameerpet or live online, and call +91 96660 19191 for a free demo.



