AWS for AI engineers comes down to a small set of services used well: Amazon Bedrock for models, retrieval, agents and guardrails; Lambda, ECS on Fargate or EKS for your application; S3, OpenSearch, Aurora PostgreSQL with pgvector or DynamoDB for data; and IAM, VPC endpoints, KMS and CloudWatch to make it safe and observable. You do not need every AWS service to ship an enterprise AI assistant. You need to know which dozen matter, how they connect, and where the security and cost traps are. This guide is that map.
If you are a cloud engineer weighing a move into AI work, the career side is covered in moving from cloud engineer to AI Forward Deployed Engineer. This article stays technical. Service names and model line-ups change often, so treat it as a map of building blocks and confirm details against current AWS documentation.
The AWS services that matter for AI applications
An enterprise LLM application on AWS has the layers of any production system, plus a model layer. This service map is the short list; most other choices are variations on it.
| Need | AWS service | Notes |
|---|---|---|
| Foundation models | Amazon Bedrock | Models from several providers behind one API; access enabled per account and region |
| Managed RAG | Knowledge Bases for Amazon Bedrock | Ingest, chunk, embed and store in a vector store you choose |
| Agents | Bedrock Agents, Amazon Bedrock AgentCore | Managed orchestration, or infrastructure for your own agent code |
| Safety controls | Amazon Bedrock Guardrails | Topic, content, PII and grounding policies |
| Custom models | Amazon SageMaker AI | Fine-tuning, training and hosting models you control |
| Application compute | Lambda, ECS on Fargate, EKS | Pick by workload shape and team skills |
| Documents | Amazon S3 | Source documents, evaluation datasets, logs |
| Vector and keyword search | Amazon OpenSearch Service | Hybrid search at scale; managed or Serverless |
| Vectors with business data | Aurora PostgreSQL with pgvector | Similarity search plus SQL filters in one query |
| Session state | Amazon DynamoDB | Conversation history and feedback, with TTL |
| Private connectivity | VPC, PrivateLink endpoints | Keep model and service calls off the public internet |
| Identity | IAM, IAM Identity Center, Cognito | Roles for workloads, SSO for people, IdP federation |
| Secrets and keys | Secrets Manager, AWS KMS | API keys, credentials, customer managed keys |
| Observability | CloudWatch, X-Ray, OpenTelemetry | Metrics, logs, traces, token usage |
| Cost control | Budgets, Cost Explorer, tags | Spend per project and team, early alerts |
| Infrastructure as code | Terraform, AWS CDK | Reproducible dev, test and production stacks |
Where these pieces sit in a wider design (gateways, orchestration, evaluation, governance) is the subject of enterprise AI architecture. Here we go service by service.
Amazon Bedrock: the model layer
Amazon Bedrock is AWS's managed service for foundation models. You call models from several providers, including Amazon's own, through AWS APIs authenticated with IAM, billed to your account and reachable privately from your VPC. That is the main reason enterprises build GenAI on AWS: the model call becomes another AWS API with the same identity, network and audit controls as everything else.
Calling models
Alongside model-specific invocation APIs, Bedrock offers the unified Converse API, one request shape for chat, tool use and streaming across supported models. Use it so switching models is a configuration change, not a rewrite. Four mechanics catch people out:
- Model access is per account and region. A model enabled in your sandbox may not be enabled, or offered, in the customer's production region.
- Quotas apply. Requests and tokens per minute are limited. Load test and request increases early.
- Capacity options differ. On-demand is billed per token; provisioned throughput reserves capacity. Cross-region inference routes requests across regions for capacity, so check which regions a profile uses against data residency rules.
- Invocation logging is your choice. Logging prompts and responses to CloudWatch Logs or S3 must be configured, and those logs may hold sensitive data.
Knowledge Bases for RAG
Knowledge Bases for Amazon Bedrock is managed retrieval-augmented generation (primer: what RAG is). You point it at a source such as an S3 bucket, choose an embedding model, a chunking strategy and a vector store such as OpenSearch Serverless or Aurora PostgreSQL, and Bedrock handles ingestion and sync. Your application calls a retrieve API (you build the prompt) or retrieve-and-generate (Bedrock returns an answer with citations). The enterprise-critical feature is metadata filtering: tag documents with department or access group and filter at query time, so a branch employee never retrieves the risk team's documents.
Agents and AgentCore
Bedrock Agents is managed orchestration: you write instructions, attach action groups (tools backed by Lambda and described with an API schema) and knowledge bases, and Bedrock runs the reason-act loop. Amazon Bedrock AgentCore is AWS's agent runtime and infrastructure offering: you bring agent code written in a framework of your choice, such as LangGraph, and AgentCore supplies production plumbing such as a managed runtime, memory, identity, tool gateways and observability. Agents means "configure an agent"; AgentCore means "run the agent you built".
Guardrails
Amazon Bedrock Guardrails applies policies to inputs and outputs: denied topics, content and word filters, sensitive information filters that block or mask PII, and contextual grounding checks that flag answers unsupported by retrieved sources. A standalone API applies a guardrail to any text, which helps when part of your pipeline runs outside Bedrock. Guardrails are one layer, not a substitute for access control; see AI security for enterprises.
SageMaker: when you need your own models
Most enterprise assistants never train a model. Amazon SageMaker AI is for when you must: fine-tuning on your data, hosting open-weight models on endpoints you control, batch inference, or serving classic ML models (fraud scores, churn predictions) that an agent calls as a tool. The trade-off is that you own instance sizing, scaling and idle cost. Start with Bedrock; move to SageMaker only for a concrete reason.
Compute: Lambda vs ECS/Fargate vs EKS
| Option | Good fit | Watch out for |
|---|---|---|
| AWS Lambda | S3-triggered ingestion, agent action groups, spiky low-volume APIs | Execution time limits, cold starts, front-door timeouts on long generations |
| ECS on Fargate | A containerised FastAPI chat service with streaming and steady traffic | Scaling policies need tuning for long-lived requests |
| Amazon EKS | Customers standardised on Kubernetes, many services, GitOps, GPU-hosted models | Highest operational load; worth it only if the platform team runs it |
A sensible default for one assistant is ECS on Fargate behind an Application Load Balancer, plus Lambda for ingestion. Choose EKS when the customer's platform team expects every workload there, and use EKS Pod Identity or IAM roles for service accounts so pods get scoped permissions. Deployments, autoscaling and GPU scheduling are covered in Kubernetes for AI applications.
Data: S3, OpenSearch, Aurora pgvector, DynamoDB
- Amazon S3 holds source documents and evaluation data. Block public access, enable versioning, encrypt with KMS, and restrict bucket policies to specific roles and your VPC endpoint.
- Amazon OpenSearch Service combines vector and keyword search, so hybrid retrieval catches exact strings such as policy numbers that semantic search misses. Good for large corpora or customers already running OpenSearch.
- Aurora PostgreSQL with pgvector keeps embeddings beside business data, letting one query combine similarity with tenant, entitlement and date filters, on a database your team already operates. For moderate corpora it is often the simplest good answer.
- Amazon DynamoDB is for conversation history, feedback and agent session state, with TTL to expire what you should not keep, not for vectors in this design.
Networking: keeping traffic private
A bank or hospital security team will first ask whether any prompt or document leaves the private network. The answer is VPC design:
- Application compute in private subnets with no direct internet route.
- Interface VPC endpoints (AWS PrivateLink) for the Bedrock runtime and agent runtime, Secrets Manager, KMS and CloudWatch Logs; gateway endpoints for S3 and DynamoDB.
- Endpoint policies limiting allowed actions, plus IAM conditions that deny calls not arriving through your endpoint.
- Users reach the app through an internal load balancer over Direct Connect or VPN, or a public one with AWS WAF and SSO if policy allows.
- Tool calls to third-party SaaS go out through controlled, logged egress.
IAM: least privilege for AI workloads
Roles for services, never keys in code
Each component gets its own role: the API task role, the ingestion Lambda role, the action group role. No long-lived access keys. Scope bedrock:InvokeModel and its streaming variant to specific model or inference profile ARNs, and grant only the knowledge base, bucket prefix and KMS key each role needs.
Separate humans from workloads
Engineers use IAM Identity Center with short-lived sessions. Keep dev, test and production in separate AWS accounts governed by AWS Organizations and service control policies, so a developer role cannot touch production data.
Per-user access to data
The model never decides who sees what. Users sign in through the customer's identity provider, often Microsoft Entra ID, via OIDC or SAML directly or through Amazon Cognito. Your API validates the token, reads the user's groups and applies them as metadata filters or SQL predicates at retrieval time. Tools that act for the user should pass the user's identity downstream and be authorised by the target system, not by a broad service role. Log the user, documents retrieved and tools called on every request.
Secrets Manager and KMS
Keep database credentials and third-party keys (a ServiceNow integration user, say) in AWS Secrets Manager with rotation where supported, fetched at runtime through the IAM role. Use AWS KMS customer managed keys for S3, Aurora, OpenSearch, DynamoDB and log groups when the customer wants to control key policies and revoke access. Key policies are least privilege too.
Building this end to end, with IAM, private endpoints, evaluation and a simulated customer engagement, is what the Cloudsoft FDE PRO program trains, with AWS as its primary platform alongside Azure and Google Cloud.
Observability: CloudWatch, X-Ray and OpenTelemetry
Bedrock publishes CloudWatch metrics including invocations, latency, throttles and input and output token counts. Alarm on throttles and errors, and graph tokens daily, because tokens drive cost. Metrics will not explain a wrong answer, though. Instrument the app with OpenTelemetry (AWS Distro for OpenTelemetry is AWS's supported distribution) and send traces to X-Ray, CloudWatch or tools such as Langfuse or LangSmith. A useful trace shows the request, retrieval query, chunks with scores, guardrail result, model call with tokens and each tool call. More in AI observability.
Cost controls
- Tag everything (project, environment, cost centre, owner) and activate cost allocation tags. Tagged Bedrock application inference profiles attribute model spend per application.
- AWS Budgets alerts per account and tag, plus Cost Anomaly Detection for runaway loops.
- Monitor tokens per user and feature. Cost per resolved question means more to a business owner than monthly spend.
- Design for cost: smaller models for routing, capped context and output, capped agent iterations, caching where policy allows, batch inference for offline jobs.
Infrastructure as code: Terraform or CDK
An assistant that exists only in one console is a demo. Define VPC, endpoints, roles, buckets, databases, knowledge base, guardrail, compute and alarms in code, so the stack deploys identically to every environment and security can review a pull request instead of screenshots. Terraform is common in multi-cloud enterprises and services firms; AWS CDK suits AWS-only teams who prefer TypeScript or Python. Version prompts, guardrail settings and evaluation datasets in the same repository. Terraform training and AWS DevOps training cover the pipeline side.
Reference deployment: a private RAG assistant on AWS
User (corporate network, SSO via IdP) | Internal ALB + WAF (private subnets) | FastAPI service on ECS Fargate | task role: least privilege |-- validate token, read user groups |-- DynamoDB: session history (TTL) |-- Retrieve: Aurora pgvector or | Bedrock Knowledge Base | (filter by user groups) |-- Guardrail check (input/output) |-- Bedrock model call (streaming) |-- OpenTelemetry trace -> X-Ray | VPC endpoints: bedrock-runtime, bedrock-agent-runtime, secrets, kms, logs; S3 gateway endpoint Ingestion: S3 (KMS) -> event -> Lambda -> parse, chunk, tag ACL metadata -> embed (Bedrock) -> vector store Ops: CloudWatch alarms, Budgets, tags, Terraform/CDK, CI/CD
The properties that matter: no public model calls, no long-lived keys, every retrieval filtered by the signed-in user's entitlements, every request traced with token counts, and the whole stack reproducible from code.
Illustrative walkthrough: an insurer's policy assistant
Consider an insurer whose IT is run from a GCC in Hyderabad. Claims handlers want answers from policy wordings and internal circulars, with citations. The path to production follows the chain from customer problem to business outcome:
- Discovery. Real handler questions reveal two hard requirements: some circulars are restricted to senior adjusters, and every answer must cite the clause.
- Data. Documents land in a KMS-encrypted S3 bucket. A Lambda chunks them by section and tags each chunk with product line and access group.
- Architecture. The team already runs Aurora PostgreSQL, so embeddings go into pgvector. A FastAPI service on Fargate handles chat; history lives in DynamoDB.
- Security. Handlers sign in with Entra ID; groups become SQL filters at retrieval. Bedrock is reached only through interface endpoints, the task role can invoke only approved models, and a guardrail masks policyholder PII and flags ungrounded answers.
- Deployment. Terraform defines the stack; the pipeline deploys to a test account, runs the evaluation suite and promotes on approval.
- Evaluation and optimization. A golden set of handler questions is scored for faithfulness and citation accuracy on every change. Token dashboards show long answers driving cost, so output caps and a smaller model for simple lookups cut spend without lowering scores.
Nothing here is exotic. It works because every AWS decision maps to a customer requirement, which is the difference between an AI demo and an enterprise outcome.
Security checklist for AI workloads on AWS
- Separate accounts for dev, test and production, with service control policies.
- Private subnets; Bedrock and other services reached via VPC endpoints with endpoint policies.
- One IAM role per component, scoped to specific models, knowledge bases, buckets and keys; no access keys.
- Entitlements enforced at retrieval and in every tool call, never by the prompt.
- Secrets in Secrets Manager; customer managed KMS keys where required.
- S3 public access blocked and bucket policies restricted.
- Guardrails for PII and grounding, plus defences against prompt injection in retrieved content and tool output.
- Invocation logging decision documented; log groups encrypted and access-restricted.
- Regions and cross-region inference reviewed against data residency rules.
- CloudTrail on; alarms for throttles, errors and token spikes.
- Human approval for agent actions that change records or move money.
How to learn this stack
If IAM, VPC and compute are new, start with AWS training and add design depth through AWS Solutions Architect training. Then build the reference deployment yourself, first in the console, then from code. Build retrieval twice, with Knowledge Bases and with your own pgvector pipeline, and practise defending each design choice to a security reviewer.
Frequently asked questions
What AWS services should an AI engineer learn first?
Start with IAM, VPC, S3 and one compute option, then Amazon Bedrock with Knowledge Bases and Guardrails. Add CloudWatch and Terraform or CDK. That set covers most enterprise LLM applications.
Is Amazon Bedrock enough, or do I need SageMaker?
Bedrock is enough for most assistants because they use pretrained models through an API. Use SageMaker to fine-tune or train models, host open-weight models you control, or serve classic ML models.
Should I use Bedrock Knowledge Bases or build RAG myself?
Use Knowledge Bases to move fast when its options fit. Build your own pipeline, for example with Aurora PostgreSQL and pgvector, when you need full control over chunking, hybrid search or ranking. Knowing both is valuable.
What is the difference between Bedrock Agents and AgentCore?
Bedrock Agents is managed orchestration: you configure instructions, tools and knowledge bases and Bedrock runs the loop. AgentCore provides infrastructure such as runtime, memory, identity and observability for agent code you write yourself.
How do I keep Bedrock traffic private?
Run the application in private subnets and use interface VPC endpoints (AWS PrivateLink) for Bedrock and related services, with endpoint policies and IAM conditions that only allow calls through them.
Lambda, ECS or EKS for an LLM application?
Lambda for event-driven ingestion and agent tools, ECS on Fargate for a streaming chat API, and EKS when the customer already runs Kubernetes or you host models on GPU nodes.
How do I control the cost of a generative AI app on AWS?
Tag resources, set AWS Budgets and anomaly alerts, and monitor token counts per model and feature. Cap context, output and agent iterations, and route simple tasks to smaller models.
How do I enforce per-user document access in a RAG assistant?
Authenticate users through the customer's identity provider, read their groups in your API and apply them as retrieval filters. Never rely on the prompt to hide documents, and re-check permissions in every tool.
Your next step
Knowing the service list is the easy part; shipping a private, observable, evaluated assistant that a security team signs off is the skill employers want. To build exactly that on AWS with Bedrock, pgvector, EKS and Terraform, ending in the GlobalBank simulated customer engagement, look at FDE PRO, Cloudsoft's Forward Deployed Engineer course in Hyderabad: 12 weeks, 60+ labs and five enterprise projects, in Ameerpet beside the Metro or live online, with placement support until you're placed. For a Bedrock-only focus, the AWS Bedrock GenAI course is the shorter route. Call +91 96660 19191 for a free demo.



