AWS Bedrock interview questions in 2026 test whether you can turn a managed model API into a governed enterprise system: choosing models and inference options for cost and latency, building Knowledge Bases that retrieve the right chunk, enforcing Guardrails, evaluating quality, and running agents on Bedrock AgentCore. This guide works through 80 Amazon Bedrock interview questions with answers, from fundamentals to production architecture and troubleshooting scenarios. Every answer explains why you would pick a particular Bedrock capability, because that reasoning is what interviewers score.
How to use this guide. What interviewers usually probe at each level:
- Freshers and early-career engineers: what Bedrock is, how the Converse API works, what a Knowledge Base does, what Guardrails filter. Clear definitions, plus a small project you can explain end to end.
- Developers and cloud/DevOps engineers moving into GenAI: IAM, VPC endpoints, KMS, logging, inference options, cost control, Infrastructure as Code. Your existing AWS depth counts for a lot here.
- Senior, architect and Forward Deployed Engineer loops: trade-offs, failure modes, evaluation strategy, multi-account governance and the scenario questions at the end. Expect follow-ups like "why not just use X?"
Bedrock changes quickly. Treat the product details here as a snapshot checked against AWS documentation, and confirm Region and model availability in the current docs before an interview.
- Bedrock fundamentals (Q1โQ8)
- Model selection (Q9โQ14)
- Inference options and cost (Q15โQ24)
- Knowledge Bases and RAG (Q25โQ36)
- Agents and AgentCore (Q37โQ43)
- Guardrails (Q44โQ49)
- Evaluation (Q50โQ54)
- Model customization (Q55โQ59)
- Security and networking (Q60โQ65)
- Observability (Q66โQ69)
- Production architecture (Q70โQ74)
- Troubleshooting scenarios (Q75โQ80)
- Key takeaways
- Interview preparation checklist
- FAQ
Bedrock fundamentals
1. What is Amazon Bedrock, and why would an enterprise choose it over calling a model provider's API directly?
Answer: Amazon Bedrock is a fully managed AWS service that gives you API access to foundation models from Amazon and third-party providers, plus the building blocks around them: Knowledge Bases for RAG, Guardrails, evaluation, prompt management, model customization, and AgentCore for running agents. The model is only part of the reason to choose it. Enterprises pick Bedrock because it runs inside the AWS controls they already have. Access goes through IAM, traffic can stay on AWS PrivateLink, keys live in KMS, API calls show up in CloudTrail and metrics in CloudWatch, and billing lands on the same AWS account. A bank that has already passed an AWS security review can add GenAI without signing and assessing a new vendor for every model it tries.
Interview tip: Lead with governance and integration, not "it has many models". Interviewers want to hear that you know why a security team says yes.
2. What are the main building blocks of the Bedrock platform?
Answer: Think of it in layers:
- Models and inference: the model catalogue, the Converse and InvokeModel APIs, OpenAI-compatible Responses and Chat Completions APIs, batch inference, Provisioned Throughput, service tiers and cross-Region inference profiles.
- Data: Knowledge Bases (managed or customer-managed) for retrieval-augmented generation.
- Safety and quality: Guardrails, and Bedrock Evaluations for models and RAG.
- Development: Prompt management, Flows, intelligent prompt routing, prompt caching.
- Customization: fine-tuning, reinforcement fine-tuning, distillation and Custom Model Import.
- Agents: Amazon Bedrock AgentCore, the current platform for building, deploying and operating agents.
Being able to draw this map in an interview shows you understand Bedrock as a platform and not just a chat endpoint.
3. Explain the Converse API. Why is it usually preferred over InvokeModel?
Answer: Converse (and ConverseStream) gives you one request and response format for all supported models: messages with roles, a system prompt, inference configuration, tool definitions and guardrail configuration. InvokeModel takes each provider's native JSON body. Converse is usually the right default because you can switch models without rewriting payload code, and tool use, documents and images all follow one shape. InvokeModel still has a place when you need a provider-specific parameter that Converse doesn't expose, or when you use Custom Model Import models that Converse doesn't support.
import boto3
brt = boto3.client("bedrock-runtime")
resp = brt.converse(
modelId=MODEL_OR_PROFILE_ID,
system=[{"text": "Answer from policy only."}],
messages=[{"role": "user",
"content": [{"text": question}]}],
inferenceConfig={"maxTokens": 512,
"temperature": 0.2},
)
print(resp["output"]["message"]["content"][0]["text"])
print(resp["usage"])
Interview tip: Mention that you log usage from every response. It is the basis for cost attribution later.
4. What is the difference between the bedrock, bedrock-runtime, bedrock-agent, bedrock-agent-runtime and bedrock-mantle endpoints?
Answer: bedrock is the control plane: managing models, customization jobs, guardrails, Provisioned Throughput and logging configuration. bedrock-runtime is the inference data plane: Converse, InvokeModel, ApplyGuardrail, and the OpenAI-compatible Responses and Chat Completions APIs. bedrock-agent is the build-time API for Knowledge Bases, data sources, prompts, flows and Agents Classic. bedrock-agent-runtime serves Retrieve, RetrieveAndGenerate, InvokeFlow and InvokeAgent. bedrock-mantle is a newer endpoint that serves OpenAI-compatible and Anthropic Messages-style APIs and has its own IAM action prefix. This matters in practice because each one needs its own VPC interface endpoint and IAM permissions. "Inference works but retrieval times out in our private subnet" usually means someone created only the runtime endpoint.
5. How do tool use (function calling) and structured output work on Bedrock?
Answer: With Converse you pass a toolConfig containing JSON-schema tool specifications. When the model decides to call a tool, it stops with a tool_use stop reason and returns the tool name and arguments. Your code runs the tool and sends back a toolResult message, and the loop continues until the model produces a final answer. Bedrock also documents a way to get validated JSON from supported models, which is cleaner than parsing free text. Never trust tool arguments blindly: validate them against your own schema, check authorisation in your code, and make side-effecting tools idempotent. For more depth, see function calling and structured outputs.
6. Does Bedrock store my prompts or share them with model providers?
Answer: Don't answer from memory; answer from the current data-protection and data-retention pages. AWS documents that each model runs in a Bedrock-owned model deployment account that model providers can't access, so providers don't see your logs, prompts or completions. Retention is now an explicit, per-Region setting with modes: none (zero retention), default (the model's own policy, which may include retention for abuse detection), aws_review (retained within AWS and possibly reviewed by AWS, required by some models), a legacy provider_data_share mode, and inherit. If you set zero retention and call a model that requires retention, Bedrock blocks the request. You can enforce the allowed modes with SCPs. Your own responsibilities still apply: invocation logs (which you enable) hold sensitive prompts, so encrypt them, restrict access and set retention. Read each service's data terms, too. AgentCore's documentation, for example, notes it may use and store your content to improve your own service experience.
Interview tip: A regulated customer will ask "is anything retained?" The right answer is "it depends on the retention mode and the model, and here is how we lock it down", not a blanket "no".
7. What is an inference profile, and why do you see IDs with prefixes like "us." or "global."?
Answer: An inference profile is a Bedrock resource that defines a model and the Regions requests can be routed to. System-defined cross-Region profiles carry a geography prefix (for example us., eu., apac.) or global., and you pass the profile ID or ARN as the modelId. You can also create application inference profiles to tag usage for a team or application, which helps with cost allocation. Many newer models are offered primarily through inference profiles, so "model not supported with on-demand throughput" errors often just mean you should call the profile instead of the bare model ID.
8. What are Bedrock Flows, and when would you use them instead of code?
Answer: Flows lets you link prompts, models, Knowledge Bases, Lambda functions, conditions and other nodes into a versioned workflow that you call through InvokeFlow. It suits deterministic pipelines, such as classify, retrieve, draft and format, where the steps are known in advance and you want versioning and aliases without writing orchestration code. Use code (or an agent framework on AgentCore) when the path depends on model reasoning, you need complex retries and state, or you want ordinary unit tests and code review. Many teams use Flows for prototypes and move hot paths into code once the logic settles.
Model selection
9. How do you choose a foundation model on Bedrock for an enterprise use case?
Answer: Start from requirements, not leaderboards: task type (extraction, summarisation, reasoning, coding, multimodal), quality bar, latency target, context length, languages, cost per request at expected volume, Region and data-residency constraints, and the features you need (tool use, prompt caching, service tiers, batch, fine-tuning). Shortlist two or three candidates, run them against a golden dataset from the customer's real documents, and compare quality, latency and cost together. Pick the cheapest model that clears the quality bar, and keep the model ID in configuration so you can swap it later. Our guide to choosing an LLM for enterprise covers the selection process in more detail.
Real-world example: Consider an insurer classifying incoming claim emails. A small, fast model reaches the needed accuracy on a few hundred labelled emails. The larger model is held back for the drafting step, where nuance matters.
10. Which model families are available on Bedrock, and why does breadth matter?
Answer: The catalogue includes Amazon's own Nova models and Titan embeddings, along with models from providers such as Anthropic, Meta, Mistral AI, Cohere, OpenAI and others, plus Bedrock Marketplace models. The exact list and Regions change often, so check the "models at a glance" pages rather than memorising versions. Breadth matters for three practical reasons: you can match model size to the task, avoid lock-in to a single provider, and keep a fallback model ready when one model is throttled or deprecated. Before you start, check that the shortlisted models have model access enabled in your account.
11. How do you handle model lifecycle and deprecation?
Answer: Bedrock models move through lifecycle states and are eventually marked legacy and retired on published dates. Plan for this from the first sprint. Keep model IDs in configuration, not code. Keep a regression evaluation suite you can rerun against a replacement model in an afternoon. Subscribe to AWS Health and deprecation notices. Never let prompts depend on one model's quirks without a test that would catch the difference. A model upgrade should be a pull request with evaluation results attached, not an emergency.
12. How do you choose an embeddings model for a Knowledge Base?
Answer: Compare models on retrieval quality for your language and domain, the vector dimensions they produce, and whether your vector store supports them. Bedrock documents Amazon Titan Text Embeddings and Cohere Embed (English and multilingual), and multimodal embeddings for image, audio and video content. Fewer dimensions mean less storage and faster search but can lose some precision, and Titan V2 lets you choose the size. OpenSearch Serverless and OpenSearch Managed Clusters are the only Knowledge Base stores that support binary vectors. Changing the embeddings model means re-embedding everything, so test on a real query set before you commit. Background reading: embeddings explained.
13. What prompt engineering practices matter most on Bedrock?
Answer: Be explicit about role, task, constraints and output format. Put stable instructions in the system prompt. Delimit retrieved context clearly (Anthropic models respond well to XML tags). Ask for citations, and tell the model what to do when the context doesn't contain the answer. Add a few examples for format-heavy tasks. For cost, put static content first so prompt caching can reuse it. Most importantly, treat prompts as versioned artefacts with tests. Prompt management (Q14) and an evaluation set turn prompt changes from guesswork into engineering.
14. What does Bedrock Prompt management give you, and why use it instead of prompts in code?
Answer: Prompt management stores prompts with variables, lets you create variants that differ in message, model or inference settings, compare them side by side in the prompt builder, optimise a prompt, and save immutable versions. You run a managed prompt by passing its ARN to model inference, or by using it as a prompt node in a Flow. You can also enable prompt caching for supported models. The point is to separate prompt changes from code deploys: a product owner can iterate on wording in a controlled place, every change is versioned, and CloudTrail can record RenderPrompt as a data event. Keep prompts in code when they are tightly coupled to logic and change only with releases. Many teams use both.
Inference options and cost
15. Compare on-demand, batch, Provisioned Throughput and the service tiers. When do you pick each?
| Option | How it works | Choose it when |
|---|---|---|
| On-demand (Standard tier) | Pay per token, shared capacity, account quotas | Default for most interactive apps and variable traffic |
| Priority tier | Per-request service_tier flag; served ahead of Standard and Flex at a price premium | Customer-facing paths where latency matters but a 24x7 reservation isn't justified |
| Flex tier | Per-request flag; discounted, longer processing times acceptable | Evaluations, summarisation, background agent work |
| Reserved tier | Reserved tokens-per-minute capacity for one or three months, overflow to Standard | Mission-critical, steady, high-volume workloads (via your AWS account team) |
| Batch inference | JSONL input in S3, asynchronous job, output to S3 | Large offline jobs: backfills, bulk classification, document enrichment |
| Provisioned Throughput | Model units billed hourly; no-commitment, one-month or six-month terms | Dedicated, predictable throughput for a model, and required to serve fine-tuned custom models |
Answer: Choose based on the latency requirement and how predictable the traffic is, not on habit. Interactive and spiky traffic goes on-demand, with Priority for the few paths that need it. Offline volume goes to batch or Flex. Only steady, high, measured volume justifies reserved capacity.
16. What are the constraints of batch inference that catch teams out?
Answer: Batch inference processes each JSONL record independently. AWS documents that it does not support tool calling or structured output (response_format), it isn't supported for provisioned models, and prompt caching isn't supported with the batch API. You submit a job, wait, and read results from S3. EventBridge can notify you when the job changes state, so you don't need to poll. Design for partial failures: match outputs back to inputs by record ID and retry failed records. Use batch when the work is large and not time-sensitive, and check the pricing page for the batch rate for your model.
17. Explain cross-Region inference. Geographic or global, and why?
Answer: Cross-Region inference sends your request through an inference profile so Bedrock can serve it from another Region's capacity. That gives you higher throughput and resilience during demand spikes. Geographic profiles keep processing inside a geography such as the US, EU or APAC. Global profiles can use any supported commercial Region and AWS lists them at a lower price than geographic profiles. AWS documents that there is no extra routing charge, that pricing follows the source Region, that data stays on the AWS network and is encrypted in transit, and that CloudTrail records the request in the source Region with an additionalEventData.inferenceRegion field. Choose geographic when you have data-residency obligations. Choose global when cost and availability matter more than where processing happens.
Interview tip: Mention SCPs. Geographic profiles need every destination Region allowed. Global profiles need aws:RequestedRegion "unspecified" allowed. Many "access denied" errors on profiles are really SCP problems.
18. How does prompt caching reduce cost and latency, and how do you design for it?
Answer: Prompt caching reuses a processed prompt prefix across requests. Bedrock supports implicit caching (automatic, no markers, with no promise of a hit) and explicit caching, where you place cachePoint checkpoints in tools, system or messages. Cached reads are billed at a reduced rate. Depending on the model, cache writes can cost more than normal input tokens. Cache entries have a TTL that resets on each hit, and some models offer a longer TTL option. Design rules: put the stable content (tool definitions, system prompt, a long policy document) first and the variable content last, because changing an earlier section invalidates the cache for every section after it. Confirm hits from cacheReadInputTokens in the response instead of assuming them.
19. What is intelligent prompt routing?
Answer: A prompt router is a single serverless endpoint that predicts response quality for each request across models in the same family and routes the request to the model offering the right quality for the cost. A configured router lets you choose two models, a fallback model and a response-quality-difference criterion. AWS notes that it is optimised for English prompts and can't learn from your application-specific performance data. Use it for mixed traffic where many requests are simple. Evaluate it on your own dataset before trusting it, and log which model actually answered each request (the response includes it).
20. How do you estimate the monthly Bedrock cost of a RAG assistant before building it?
Answer: Build a simple model from measured or assumed inputs, clearly labelled as assumptions: requests per day ร (input tokens per request ร input price + output tokens per request ร output price). Then add embedding costs for ingestion and re-syncs, vector store costs (OpenSearch Serverless capacity, Aurora instance, or S3 Vectors storage and queries), reranking calls, guardrail evaluations, semantic chunking (which itself uses a model), logging storage and Lambda. Input tokens in RAG are dominated by retrieved context, so the number of chunks and their size drive cost more than the question length does. Present a low, expected and high range and the levers for each: fewer or smaller chunks, caching, a smaller model, batch for offline work. For more cost levers, see cloud cost optimization for AI.
21. Scenario: your Bedrock bill tripled in a month and nobody knows why. How do you investigate?
Answer: Treat it like any cost incident. Attribute the spend first, then fix the cause.
What I would check:
- Cost Explorer by usage type and model, and by application inference profile tags if they exist.
- CloudWatch
InputTokenCountandOutputTokenCountper model over time, to see whether input or output grew. - Invocation logs grouped by
identity.arnandrequestMetadatato find which caller or team changed. - Recent deploys: a larger retrieval
numberOfResults, a new agent loop, a retry storm, or a model change. - Agent traces for runaway tool loops, which multiply calls per user request.
Production consideration: Prevent the next one. Add per-application inference profiles, request metadata tags, AWS Budgets alerts, max-iteration limits on agents, and token budgets per request.
22. Scenario: during a festive-season sale, a retailer's assistant starts returning ThrottlingException errors. What do you do?
Answer: Bedrock enforces requests-per-minute and tokens-per-minute quotas for each model and Region. Stabilise the service first, then add capacity.
What I would check:
InvocationThrottlesand token metrics: is it the request quota or the token quota?- Retry behaviour: exponential backoff with jitter and capped attempts, so retries don't amplify the load.
- Whether a cross-Region inference profile can absorb the peak.
- Whether non-interactive work (summaries, tagging) can move to a queue, batch or the Flex tier.
- Whether input tokens can drop through fewer chunks or prompt caching, which also eases token quotas.
Production consideration: Request quota increases with usage evidence before the next peak. Consider the Reserved tier or Provisioned Throughput for steady critical load. Fail gracefully with a clear message rather than queuing users forever.
23. How do you reduce latency for a chat assistant on Bedrock?
Answer: Stream responses with ConverseStream so users see the first tokens quickly. Use the smallest model that meets the quality bar. Cut input tokens (fewer, better chunks, trimmed history). Use prompt caching for long static prefixes. Run retrieval and other independent calls in parallel. Use the Priority tier or latency-optimized inference where supported. Keep compute in the same Region as the model endpoint. Measure time-to-first-token and total latency separately, because they have different causes. AWS documents output tokens per second as a way to tell whether higher InvocationLatency comes from longer outputs or slower generation. Our LLM latency optimization guide goes deeper.
24. How do you attribute Bedrock cost to teams in a shared platform account?
Answer: Create application inference profiles per team or application and tag them, so cost allocation tags flow into billing. Add per-request metadata tags, which appear in invocation logs as requestMetadata. Because identity.arn is captured automatically in invocation logs, you can also group token usage by IAM role with CloudWatch Logs Insights. Combine these into a chargeback dashboard. An LLM gateway in front of Bedrock can add per-team budgets and rate limits. See LLM gateway explained for that pattern.
Knowledge Bases and RAG
25. What is a Bedrock Knowledge Base, and why use it instead of building RAG yourself?
Answer: A Knowledge Base is managed retrieval-augmented generation. It connects to data sources, parses and chunks documents, creates embeddings, stores them in a vector store, keeps them in sync, and serves retrieval through the Retrieve, RetrieveAndGenerate and RetrieveAndGenerateStream APIs with citations. You choose it to avoid building and maintaining ingestion pipelines, sync jobs and citation plumbing. You build RAG yourself when you need retrieval logic it can't express, full control of the index, or a vector store it doesn't support. If RAG itself is new to you, read what RAG is first.
26. What is the difference between a Bedrock Managed Knowledge Base and a customer-managed Knowledge Base?
Answer: With a Bedrock Managed Knowledge Base, Bedrock runs the ingestion, the datastore, indexing and retrieval. It includes a managed embedding model and a managed semantic reranker, agentic and hybrid retrieval, native connectors (S3, SharePoint, Confluence, Web Crawler, Google Drive, OneDrive and custom), a built-in multimodal parser, access-control-list awareness, and AgentCore Gateway integration. AWS's documentation recommends it for the combination of ease of use, accuracy and cost. With a customer-managed Knowledge Base, you choose, provision and scale your own vector store and choose your parser and chunking strategy, which you want when you need a specific vector database configuration or direct access to the index.
Interview tip: Mention the connector change. AWS states that from 30 September 2026, new Confluence, SharePoint, Salesforce and Web Crawler connectors can no longer be created on customer-managed Knowledge Bases. Existing ones keep working, and AWS points new needs to the managed option.
27. Which data sources can a Knowledge Base use?
Answer: Customer-managed Knowledge Bases list Amazon S3, Confluence, Microsoft SharePoint, Salesforce, a web crawler and custom data sources (you push documents directly through the API). The connector restriction in Q26 applies to new connectors. Multimodal files (images, audio, video) are supported only with S3 and custom sources. Knowledge Bases can also connect to structured data stores through a query engine that turns natural language into SQL. That fits questions like "total claims by branch last quarter", where vector search is the wrong tool. Supported document formats include text, Markdown, HTML, Word, CSV, Excel and PDF, with per-file size quotas.
28. Which vector stores can a customer-managed Knowledge Base use, and how do you choose?
| Vector store | Why you would pick it |
|---|---|
| Amazon OpenSearch Serverless | Quick-create default, hybrid search, rich filtering, binary vectors; pay for serverless capacity |
| Amazon OpenSearch Service Managed Clusters | You already run OpenSearch and want control over capacity; note the domain must use public access for Knowledge Bases |
| Amazon S3 Vectors | Low-cost, durable storage for large, infrequently queried corpora; float vectors only, metadata size limits |
| Amazon Aurora PostgreSQL (pgvector) | Teams already on PostgreSQL; hybrid search; same-account cluster required |
| Amazon Neptune Analytics | GraphRAG, where relationships between entities matter |
| Pinecone, Redis Enterprise Cloud, MongoDB Atlas | An existing third-party platform standard; credentials through Secrets Manager |
Answer: Choose based on query volume, filtering and hybrid-search needs, and cost profile, and on what the operations team already runs. A rarely queried archive suits S3 Vectors. A busy assistant that needs hybrid search suits OpenSearch Serverless or Aurora. The vector databases explainer covers the general trade-offs.
29. Explain the chunking strategies Knowledge Bases offer and when you use each.
Answer:
- Default: chunks of roughly 300 tokens that respect sentence boundaries. A sensible baseline.
- Fixed-size: you set maximum tokens per chunk and an overlap percentage. Predictable, good for uniform prose.
- Hierarchical: small child chunks are retrieved for precision, then replaced by their larger parent chunk for context. Good for long manuals and policies. You may get fewer results than requested, and AWS advises against it with S3 Vectors because of metadata size limits.
- Semantic: splits where meaning changes, controlled by max tokens, buffer size and breakpoint percentile threshold. It uses a model, so it adds ingestion cost.
- No chunking: each file is one chunk. Use it when you have already split documents yourself. You lose page-number citations.
- Custom transformation: a Lambda function in the ingestion pipeline for your own chunking or metadata logic.
Pick by testing retrieval metrics on real questions, not by preference. See RAG chunking strategies.
30. What parsing options exist, and why do they matter for PDFs with tables?
Answer: Customer-managed Knowledge Bases offer the default text parser, a foundation-model parser, and Amazon Bedrock Data Automation. The managed option has a built-in multimodal parser. The default parser extracts text, which is fine for clean prose but can flatten tables and lose figures. A foundation-model parser or Data Automation can interpret tables, charts and scanned pages into structured text before chunking, at extra cost. For a bank's rate sheets or an insurer's benefit tables, poor parsing is the most common root cause of "the assistant gives the wrong number". More on this in parsing PDFs, tables and scans for RAG.
31. Retrieve versus RetrieveAndGenerate: when do you use each?
Answer: Retrieve returns ranked chunks with scores and metadata, and you build the prompt and call the model yourself. RetrieveAndGenerate retrieves, generates the answer with citations and manages session context in one call. Use RetrieveAndGenerate for fast delivery and standard Q&A. Use Retrieve when you need control: custom prompts, combining several retrieval sources, permission filtering in your own code, your own reranking, or an agent that treats retrieval as one tool among many. In production, most serious systems end up on Retrieve plus Converse.
kb = boto3.client("bedrock-agent-runtime")
r = kb.retrieve(
knowledgeBaseId=KB_ID,
retrievalQuery={"text": question},
retrievalConfiguration={"vectorSearchConfiguration": {
"numberOfResults": 8,
"overrideSearchType": "HYBRID",
"filter": {"equals": {"key": "dept",
"value": user_dept}}}})
chunks = [x["content"]["text"]
for x in r["retrievalResults"]]
32. How do hybrid search, reranking and query decomposition improve retrieval?
Answer: Hybrid search combines vector similarity with keyword search over raw text, which helps with product codes, policy numbers and exact terms that embeddings handle poorly. AWS documents hybrid support for Amazon RDS (Aurora), OpenSearch Serverless and MongoDB stores that have a filterable text field. Other stores fall back to semantic search. Reranking runs a reranker model over the retrieved candidates and reorders them by relevance, so you can retrieve broadly and pass only the top few to the model. Query decomposition breaks a compound question into sub-queries before retrieval. Each adds latency and cost, so turn them on when evaluation shows they help. See hybrid search and reranking.
33. How do metadata filtering and implicit filtering work?
Answer: You attach metadata to documents (for S3, a .metadata.json sidecar file) and filter at query time with operators such as equals, in, greaterThan and listContains, combined with andAll and orAll. Operator support varies by vector store. Implicit filtering asks a supported model to generate the filter from the user's question and a metadata schema you describe, for example inferring year or region. Use explicit filters for anything security-related, and implicit filters only for relevance. The model should never decide which documents a user is allowed to see.
34. Scenario: an HR assistant must only show documents a user is entitled to see. How do you implement permission-aware retrieval?
Answer: Enforce access in retrieval, never in the prompt.
What I would check:
- Where entitlements come from (Entra ID groups, HRMS roles), and whether each document can carry an ACL or classification label in metadata at ingestion.
- Whether the managed Knowledge Base's ACL awareness covers the connector in use, or whether I need explicit metadata filters built server-side from the user's verified token.
- That filters are applied in the backend, using claims from the authenticated session, never values the client sends.
- Tests with users from different groups proving that restricted chunks never appear in
Retrieveresults, not just in final answers. - How quickly permission changes propagate (sync schedule), and whether the lag is acceptable to the customer's security team.
Production consideration: Log the filter applied with every request so an auditor can reconstruct why a user saw a document. For highly sensitive tiers, a separate Knowledge Base per classification can be simpler to defend than complex filters.
35. How do you keep a Knowledge Base fresh, and what does a sync actually do?
Answer: An ingestion job (sync) crawls the data source and processes new, modified and deleted documents incrementally, then re-embeds only what changed. Schedule syncs with EventBridge Scheduler or trigger them from source events, and monitor job statistics for failed documents. For custom data sources you can ingest or delete documents directly through the API for near-real-time updates. Set the data deletion policy deliberately. And put a "last indexed" timestamp in metadata so answers can say how current they are.
36. Scenario: users say the assistant answers confidently from an outdated policy document. What do you do?
Answer: This is a data problem that shows up as a model problem.
What I would check:
- Whether the old and new versions both sit in the source, and both get indexed.
- Whether the last sync succeeded and the new document was actually ingested (check the job's failure list).
- Whether metadata includes effective dates or a status field to filter on.
- Whether retrieval ranks the old document higher because of wording, and whether reranking fixes it.
- Whether the prompt tells the model to prefer the most recent effective policy and to cite its date.
Production consideration: Fix the root cause with document lifecycle ownership: archive superseded versions out of the indexed location, add a status filter, and add regression questions about recently changed policies to the evaluation set.
Agents and AgentCore
This section stays brief on purpose. For component-level depth, see the Bedrock AgentCore interview questions guide.
37. What is the current status of Amazon Bedrock Agents?
Answer: AWS's documentation states: "Amazon Bedrock Agents (now Amazon Bedrock Agents Classic) is no longer open to new customers. For capabilities similar to Bedrock Agents Classic, explore Amazon Bedrock AgentCore. Existing customers can continue to use the service as normal." The maintenance-mode page dates this to 30 July 2026. Accounts with Bedrock Agents activity in the previous 12 months are allowlisted. New accounts get an AccessDeniedException on CreateAgent and InvokeInlineAgent. No new features are planned, the Agents Classic model catalogue is frozen at that date, and there is no announced end-of-life or migration deadline. Bedrock models, Knowledge Bases and Guardrails are not affected. For new work, AWS recommends AgentCore.
Interview tip: Say "Agents Classic is in maintenance mode, closed to new customers" precisely. Saying "Bedrock Agents was shut down" is wrong, and saying "use Bedrock Agents for new builds" is out of date.
38. What is Amazon Bedrock AgentCore, and what are its components?
Answer: AgentCore is AWS's platform for building, deploying and operating agents with any framework and any model, inside or outside Bedrock. Its services can be used together or independently:
- Harness: a managed agent loop you configure with a model, system prompt and tools.
- Runtime: serverless hosting for agents and tools with session isolation, and support for MCP and A2A.
- Memory: short-term (within a session) and long-term (across sessions) memory.
- Gateway: turns APIs, Lambda functions and services into MCP tools, fronts existing MCP servers and other agents, and handles inbound and outbound authentication.
- Identity: workload identities for agents and credential management for acting on behalf of users.
- Code Interpreter and Browser: sandboxed code execution and a managed browser.
- Observability: OpenTelemetry-compatible tracing through CloudWatch.
- Evaluations, Optimization, Policy, Registry and Payments: agent quality scoring, recommendation-driven tuning with A/B tests, deterministic rules on tool calls (Cedar-compatible), a governed catalogue of agents and tools, and agent micropayments.
Check Region availability for each service before you design around it.
39. Why would you choose AgentCore over running LangGraph or Strands on your own EKS cluster?
Answer: AgentCore removes the undifferentiated work every production agent needs: isolated sessions, long-running execution, identity and OAuth token handling for third-party tools, memory storage, tool gateways and tracing. Your framework code can stay the same, because Runtime hosts LangGraph, Strands, CrewAI, LlamaIndex, OpenAI Agents SDK and custom code. Self-hosting on EKS still makes sense when the customer requires everything in their own cluster, needs a network topology AgentCore can't meet, or already has a mature platform team and observability stack. The honest answer weighs build-and-run cost against control.
40. What is the Strands Agents SDK?
Answer: Strands Agents is an open-source SDK from AWS for building agents in Python and TypeScript using a model-driven approach. You give the agent a model, a system prompt and tools (plain functions with a @tool decorator, or tools from MCP servers), and the model plans and calls tools in a loop. It supports Bedrock and other model providers, multi-agent patterns (agents as tools, swarm, graph, workflow), and OpenTelemetry-based observability. It deploys to Lambda, Fargate, EKS, EC2 or AgentCore Runtime. Choose it when you want a lightweight, AWS-friendly framework. Choose LangGraph when you need explicit state machines with checkpoints, as covered in our LangGraph interview questions.
41. How does AgentCore Gateway relate to MCP?
Answer: The Model Context Protocol (MCP) is an open protocol, introduced by Anthropic in late 2024, for connecting AI applications to tools and data. AgentCore Gateway exposes OpenAPI, Smithy and Lambda targets as MCP-compatible tools, connects to existing MCP servers, offers integrations with tools such as Salesforce, Slack and Jira, and provides semantic tool search so agents can find the right tool among many without loading every definition into the prompt. It also handles inbound authentication (who may call the gateway) and outbound credential exchange (how the gateway authenticates to each tool). Agents talk one protocol, and the gateway handles the translation and credentials. Background: what is MCP.
42. Scenario: a customer runs Bedrock Agents Classic and wants to move to AgentCore. How do you plan the migration?
Answer: AWS documents two paths. The AgentCore harness is the closest equivalent: you declare the model, tools and instructions. Code-defined agents on Runtime suit custom orchestration and multi-agent supervisors. There is no migration deadline, so plan it deliberately.
What I would check:
- The inventory: action groups (these become Gateway tools), Knowledge Bases (Gateway or a retrieval tool), return-of-control (inline function tools) and session memory (AgentCore Memory).
- Gaps AWS lists: stage-specific prompt overrides aren't replicated directly, and routing-mode multi-agent collaboration needs custom framework code.
- Whether the AgentCore CLI import or the agent-toolkit migration skill covers this agent.
- AgentCore Region availability for the customer's Region.
Production consideration: Gate the cutover on evaluation. Replay the same test conversations against the old and new agents, compare outcomes, then shift traffic gradually.
43. How do you keep an agent from taking a harmful action?
Answer: Use layers, with the model as the least trusted one. Apply least-privilege IAM to each tool. Scope OAuth tokens per user through AgentCore Identity. Use deterministic rules on tool calls through AgentCore Policy at the Gateway. Require human approval for irreversible or high-value actions (inline function tools pause the loop and return control to your code). Set input and output Guardrails. Cap iterations and token budgets. Keep a full trace of every action. Then red-team it. See AI agent identity and access and the agentic AI interview questions for broader patterns.
If you want to practise this end to end, building an agent on Bedrock, wiring tools through MCP and defending the design to a sceptical "customer", that is the core of Cloudsoft's AI Forward Deployed Engineer course. Its Secure Banking AI Assistant and ServiceNow AI Agent via MCP projects use exactly these patterns.
Guardrails
44. What policy types does Amazon Bedrock Guardrails provide?
Answer: Six safeguards, each configurable on inputs, outputs or both:
- Content filters: Hate, Insults, Sexual, Violence, Misconduct and Prompt Attack, with adjustable strength, for text and images.
- Denied topics: natural-language definitions of subjects to block, such as investment advice in a banking app.
- Word filters: exact words and phrases, plus a managed profanity list.
- Sensitive information filters: detect PII entities and custom regex patterns, then block or mask them.
- Contextual grounding checks: flag responses that aren't grounded in the provided source or aren't relevant to the query.
- Automated Reasoning checks: validate responses against logical rules derived from your policies, and highlight unstated assumptions.
Guardrails also has Classic and Standard safeguard tiers. Standard extends detection into code elements such as comments and string literals. More context: AI guardrails explained.
45. What is the ApplyGuardrail API, and why does it matter?
Answer: ApplyGuardrail evaluates any text against a guardrail without invoking a model. That decouples safety from the model provider. You can screen user input before it reaches a self-hosted or third-party model, check tool outputs before an agent uses them, or scan retrieved chunks. One guardrail definition can then protect every model in the estate, which is easier to govern than separate filters for each provider.
46. How do you apply guardrails only to the user's input in a RAG prompt?
Answer: Use input tagging. Wrap the user's text in guardrail tags through the SDK so the guardrail evaluates only that section and ignores system instructions, retrieved results and few-shot examples. Without tagging, a content filter can fire on a legitimate retrieved policy paragraph, or a prompt-attack filter can flag your own system prompt. AWS notes that selective evaluation is available through the SDK, not the console playground.
47. Scenario: after enabling Guardrails, legitimate questions in a hospital assistant are being blocked. How do you tune it?
Answer: Over-blocking is as much a production failure as under-blocking, so tune with data.
What I would check:
- The guardrail trace on blocked requests, to see which policy intervened.
- Whether clinical vocabulary is tripping the Violence or Sexual content filters at high strength, and whether lowering the strength for inputs only is acceptable.
- Whether denied-topic definitions are too broad (for example "medical advice" when the app exists to give clinical guidance to staff).
- Whether input tagging is missing, so the guardrail is scanning retrieved clinical text.
- A labelled set of allowed and disallowed prompts, run through ApplyGuardrail before and after each change.
Production consideration: Version the guardrail, record the false-positive and false-negative counts for each version, and get the clinical safety owner's sign-off. Guardrail settings are a policy decision, not just an engineering one.
48. How do contextual grounding checks differ from Automated Reasoning checks?
Answer: Contextual grounding checks score whether a response is supported by the source you supply and relevant to the query. They catch RAG answers that add facts not in the retrieved chunks. Automated Reasoning checks validate a response against formal rules built from a policy document, for example eligibility rules in an HR leave policy, and can suggest corrections. Use grounding broadly on RAG answers. Use Automated Reasoning where rules are crisp and errors are costly. Neither replaces offline evaluation.
49. How do you enforce guardrails consistently across many teams and accounts?
Answer: Publish versioned guardrails from a platform team and require them through IAM conditions on inference calls where supported. Bedrock also documents Guardrails enforcements that apply safeguards across accounts in an organisation, so an enforced guardrail evaluates calls even when the application didn't specify one. Log ApplyGuardrail data events in CloudTrail for audit. AgentCore Gateway can enforce guardrail policy on agent tool traffic. Consistency comes from central ownership plus enforcement, not from asking every team to remember a parameter.
Evaluation
50. What evaluation types does Amazon Bedrock offer?
Answer: Bedrock Evaluations supports:
- Automatic programmatic model evaluations on built-in or custom prompt datasets.
- Human-based evaluations, using your own team or subject-matter experts.
- Model evaluations with an LLM as a judge, where a second model scores responses and explains its scores.
- RAG evaluations that score retrieval and generation for Bedrock Knowledge Bases or for RAG sources outside Bedrock, using a dataset with ground truth.
It can evaluate foundation models, custom and imported models, prompt routers and provisioned models, and it can also score responses you generated elsewhere. For agents, AgentCore Evaluations scores sessions, traces and spans. Concepts in depth: LLM evaluation.
51. How do you build an evaluation dataset for a Bedrock RAG assistant?
Answer: Start from real user questions (from logs or discovery workshops), not invented ones. Cover common questions, edge cases, questions the corpus can't answer, and adversarial prompts. For each, record the expected answer and the expected source document or passage, reviewed by a domain expert. Keep a frozen regression set and a rotating set refreshed from production failures. A smaller, expert-labelled set is worth more than a large synthetic one, though synthetic generation helps fill coverage gaps.
52. Which metrics matter for evaluating RAG on Bedrock?
Answer: Separate retrieval from generation. For retrieval: context relevance and context coverage (did the right passages come back?). For generation: correctness, completeness, faithfulness or grounding to the retrieved context, citation precision, and responsible-AI metrics such as harmfulness and refusal behaviour. Bedrock's RAG evaluations compute LLM-judged versions of several of these, and frameworks such as Ragas compute similar ones. Track latency and cost alongside quality, because a correct answer that takes too long still fails. See RAG evaluation metrics.
53. How do you make LLM-as-a-judge evaluation trustworthy?
Answer: Calibrate the judge against human labels on a sample and measure agreement before relying on it. Use clear rubrics with defined score levels. Use a judge model that differs from the model under test where possible. Check for position and verbosity bias. Re-check calibration when you change the judge model. Report judge scores with confidence ranges and spot-check samples by hand. A judge is a measuring instrument, and it needs calibrating like one.
54. Where does evaluation sit in a CI/CD pipeline for a Bedrock application?
Answer: Run a fast subset on every pull request that touches prompts, retrieval settings or model IDs, and run the full suite before release. Gate promotion on thresholds agreed with the business owner, not on gut feel. Post results as pipeline artefacts. In production, sample live traffic for online evaluation and feed failures back into the dataset. Our guide to CI/CD for AI applications shows the pipeline shape.
Model customization
55. Which customization methods does Bedrock support today?
Answer: AWS's documentation lists three:
- Supervised fine-tuning: train on labelled input and output pairs.
- Reinforcement fine-tuning: you define reward functions (for example in Lambda) that score responses, and you can supply prompt datasets or existing invocation logs.
- Distillation: a larger teacher model generates responses to your prompts, and Bedrock fine-tunes a smaller, cheaper student model on them.
Separately, Custom Model Import brings in open-weight models you customised elsewhere. Training is billed on tokens processed times epochs, plus monthly model storage. Check the customization docs for which base models support which method.
56. When do you fine-tune instead of using RAG or better prompts?
Answer: Fine-tune for behaviour, RAG for knowledge. Fine-tuning helps with consistent format, tone, a specialised task (classification, extraction into a fixed schema) or a smaller model matching a larger one. RAG is the right tool for facts that change, need citations, or must respect permissions. Fine-tuning won't keep facts current and makes them hard to audit. The usual order is prompt engineering, then RAG, then distillation or fine-tuning, with each step justified by evaluation results. See RAG vs fine-tuning and the fine-tuning guide.
57. What is model distillation, and why is it attractive?
Answer: Distillation transfers a large model's task performance to a smaller model. You supply prompts for your use case, Bedrock generates teacher responses and fine-tunes the student on them. You choose it when the large model meets the quality bar but its cost or latency doesn't work at volume, for example a retailer classifying millions of product queries. Validate the student on the same evaluation set as the teacher, and decide in advance what quality gap is acceptable.
58. What is Custom Model Import, and what are its limits?
Answer: Custom Model Import brings open-weight models you trained or fine-tuned elsewhere (for example in SageMaker AI) into Bedrock, so you can call them through Bedrock APIs with on-demand billing. You supply Hugging Face format files (safetensors weights, config and tokenizer files) from S3, for supported architectures that include Llama, Mistral, Mixtral, Flan, Qwen, GPTBigCode and GPT-OSS. AWS documents limits: it can't be used with batch inference or CloudFormation, it doesn't support embedding models, it is offered in a subset of Regions, and there are weight-size and context-length ceilings. Choose it when you need a specific open-weight model but want Bedrock's security, logging and API surface.
59. How do you serve a fine-tuned model, and what does it cost structurally?
Answer: The Provisioned Throughput documentation states that if you customized a model, you must purchase Provisioned Throughput to use it. That means hourly billing for model units, with no-commitment, one-month or six-month terms, priced like the base model. That fixed cost is why the business case matters: a fine-tuned model serving low traffic can cost more than a larger on-demand model. Check the current docs for any on-demand deployment option for your specific custom model type, since this area changes. Imported models (Q58) use on-demand billing.
Security and networking
60. How do you design IAM for a Bedrock application?
Answer: Give each workload its own role, scoped to specific actions and resources: bedrock:InvokeModel and InvokeModelWithResponseStream on approved model and inference-profile ARNs only, bedrock:Retrieve on specific Knowledge Base ARNs, bedrock:ApplyGuardrail on approved guardrails. Avoid AmazonBedrockFullAccess outside sandboxes. Use SCPs at the organisation level to deny unapproved models or Regions. Knowledge Bases and customization jobs use service roles that need their own least-privilege access to S3, the vector store and KMS. For the newer bedrock-mantle endpoint, remember the separate bedrock-mantle: action prefix.
61. How do you keep Bedrock traffic private?
Answer: Create interface VPC endpoints (AWS PrivateLink) for each Bedrock endpoint you use: com.amazonaws.region.bedrock-runtime for inference, bedrock-agent-runtime for retrieval, bedrock and bedrock-agent for control-plane and build-time calls, bedrock-mantle if you use it, and FIPS variants in supported Regions. Enable private DNS so SDK calls resolve to the endpoints without code changes. Attach endpoint policies that restrict actions and principals. Run customization and batch jobs with VPC configuration so training data in S3 stays on private paths, and use OpenSearch Serverless with VPC network access for the vector store.
App (private subnet)
|
v
VPC interface endpoints (bedrock-runtime,
| bedrock-agent-runtime)
v
Amazon Bedrock --> Knowledge Base --> vector store
| (VPC access)
v
KMS keys, CloudWatch Logs, CloudTrail
Interview tip: Mention the gotcha. AWS documents that OpenSearch Managed Cluster domains behind a VPC are not supported for Knowledge Bases, so a "fully private" requirement points you to OpenSearch Serverless with VPC access, Aurora, or another supported store.
62. Where does KMS fit in a Bedrock design?
Answer: Use customer-managed KMS keys where the service supports them: custom model artefacts and customization jobs, Knowledge Base resources and transient ingestion data, vector stores (OpenSearch, S3 Vectors with SSE-KMS, Aurora), guardrails, agent resources, invocation log destinations (S3 bucket and CloudWatch Logs group), and Secrets Manager secrets for third-party vector stores. Write key policies that grant the Bedrock service role only the operations it needs. Customer-managed keys give you revocation, rotation and an audit trail of key use, which is what regulated customers usually ask for.
63. Scenario: a bank's security team asks you to prove Bedrock usage meets its controls. What do you present?
Answer: Give evidence, not assurances.
What I would check:
- Network: VPC endpoint IDs, endpoint policies and VPC Flow Logs showing no internet path.
- Identity: IAM roles and SCPs restricting models and Regions, with access reviews.
- Data: KMS key policies, encryption settings, invocation-log retention, and the account data-retention mode, enforced by SCP.
- Residency: geographic inference profiles only, with SCPs blocking global profiles if required.
- Safety: guardrail versions, their test results and enforcement configuration.
- Audit: CloudTrail trails including data events for Retrieve, ApplyGuardrail and agent calls, plus GuardDuty findings coverage.
Production consideration: Map each control to the bank's own framework and any applicable regulation (in India, the DPDP Act for personal data). Hand over a control matrix the security team can file. See generative AI in banking.
64. How do you defend a Bedrock RAG application against prompt injection?
Answer: Assume the documents you retrieve may contain instructions. Turn on the Guardrails Prompt Attack filter on user input (with tagging). Delimit retrieved content clearly and instruct the model to treat it as data. Never give the model credentials or unscoped tools. Validate tool arguments in code. Use ApplyGuardrail on tool and retrieval outputs in agent flows. Sanitise ingested content where possible. Red-team regularly with indirect injection examples planted in test documents. No single control is enough, so layer them.
65. How do you handle multi-account and data-residency requirements for Indian enterprises?
Answer: Use AWS Organizations with separate accounts per environment and workload, SCPs that allow only approved Regions and models, and geographic cross-Region inference profiles that keep processing within the permitted geography. Check whether the models you need are available in the Mumbai or Hyderabad Regions or in the relevant geographic profile, and document what is processed where. For GCC teams in Hyderabad or Bengaluru serving European clients, the client's residency rules usually decide the design. Our DPDP Act for AI applications article covers the Indian side.
Observability
66. What should you monitor for a Bedrock workload?
Answer: From CloudWatch runtime metrics: Invocations, InvocationLatency, InvocationClientErrors, InvocationServerErrors, InvocationThrottles, InputTokenCount and OutputTokenCount per model, plus service-tier dimensions (ResolvedServiceTier shows which tier actually served the request). Add application metrics: time-to-first-token, retrieval latency and result counts, guardrail intervention rate, cache hit tokens, cost per request, and quality signals such as user feedback and online evaluation scores. Alarm on throttles, error rate and cost anomalies. More in AI observability.
67. What is model invocation logging, and what should you be careful about?
Answer: It records full request and response bodies and metadata for Converse, ConverseStream, InvokeModel and InvokeModelWithResponseStream calls on the bedrock-runtime endpoint. Delivery goes to CloudWatch Logs, S3 or both, in the same account and Region. Large bodies and binary data go to S3. It is disabled by default. It is invaluable for debugging and cost analysis because each record includes identity.arn, token counts and optional requestMetadata. It is also a store of potentially sensitive prompts, so encrypt it with KMS, restrict access, set retention, and decide with the customer whether PII masking happens before prompts reach the model. AWS notes that calls through bedrock-mantle are not currently captured.
68. What does CloudTrail capture for Bedrock, and what needs extra configuration?
Answer: CloudTrail records InvokeModel, InvokeModelWithResponseStream, Converse and ConverseStream as management events, logged by default, with who, when, from where and which model. Prompt content is not included. Several important calls are data events that you must enable with advanced event selectors (and pay for): Retrieve and RetrieveAndGenerate (Knowledge Base resource type), InvokeAgent, InvokeFlow, RenderPrompt and ApplyGuardrail. Teams often discover this gap during an audit, so enable data events for Knowledge Bases and Guardrails from day one in regulated environments. GuardDuty analyses CloudTrail for suspicious Bedrock activity, such as someone deleting guardrails.
69. How do you trace an agent or RAG request end to end?
Answer: Propagate one trace ID from the API layer through retrieval, model calls, guardrail checks and tool calls. Use OpenTelemetry instrumentation (Strands and AgentCore emit OTEL-compatible telemetry, and AgentCore Observability shows it in CloudWatch). Record spans for each step with token counts, latency, the retrieved document IDs and guardrail results, and put the same ID into requestMetadata so invocation logs link to traces. Tools such as LangSmith or Langfuse can sit alongside for prompt-level debugging. When a user reports a bad answer, you should be able to replay exactly what was retrieved and what the model saw.
Production architecture
70. Design a production Bedrock RAG assistant for an enterprise. Walk through the architecture.
Answer: A defensible reference design:
Users --> SSO (Entra ID / Cognito)
|
v
API (API Gateway + Lambda/ECS) -- auth, rate limit
|
+--> Guardrail (input, tagged)
+--> KB Retrieve (filters from user claims)
| |-- vector store + rerank
+--> Converse (stream) + prompt cache
+--> Guardrail (output + grounding)
|
v
Answer + citations --> feedback capture
Logs/metrics/traces --> CloudWatch, CloudTrail
Explain the choices. Retrieve plus Converse keeps control of the prompt and permissions. Filters come from verified claims. Guardrails sit on both sides. Streaming keeps the experience responsive. Everything runs through VPC endpoints with KMS encryption. Infrastructure is defined in Terraform or CDK. Evaluation runs in CI before every change. State the non-functional targets (latency, availability, cost per answer) agreed with the customer, and how you will measure each.
71. Scenario: the primary model is throttled or degraded for an hour during business hours. How should the system behave?
Answer: It should degrade in a planned way, not fail.
What I would check:
- Whether calls already go through a cross-Region inference profile for extra capacity.
- Whether a fallback model, already tested against the evaluation set, can be switched in through configuration or a circuit breaker.
- What the degraded experience is, for example returning retrieved sources without a generated summary.
- Whether non-urgent work queues instead of competing for capacity.
Production consideration: Rehearse failover in a game day. An untested fallback model is a guess. For critical steady paths, consider reserved capacity.
72. How would you deploy Bedrock resources with Infrastructure as Code?
Answer: Define Knowledge Bases, data sources, vector stores, guardrails and their versions, IAM roles, KMS keys, VPC endpoints, logging configuration and alarms in Terraform or CDK/CloudFormation, with separate dev, test and prod accounts. Promote guardrail and prompt versions through environments just as you promote code. Keep data ingestion (syncs) as a pipeline step, not a manual console action. Note that some features, such as Custom Model Import, aren't supported in CloudFormation, so script those through the SDK. See Terraform for AI infrastructure.
73. Scenario: a retailer wants a customer-support assistant that answers from FAQs and can check order status. Design it on Bedrock and explain your service choices.
Answer: Split the knowledge path from the action path.
What I would check:
- Where the FAQs and policies live (a managed Knowledge Base with a native connector, or S3 sync from the CMS), and how often they change.
- The order system's API and authentication. Expose it as a Gateway tool with outbound OAuth through AgentCore Identity, scoped to the logged-in customer.
- Whether order lookup is read-only (safe to automate) and whether refunds need human approval.
- Languages customers write in, which drives the choice of model and embeddings.
- Peak traffic during sales, which drives cross-Region profiles, quotas and the cost model.
Production consideration: Run the agent on AgentCore Runtime with Memory for the conversation, guardrails for PII masking, and a small fast model for routing with a stronger model only for complex replies. Measure containment rate and escalation quality with the support team, not just answer accuracy.
74. Scenario: you are building a multi-tenant SaaS product on Bedrock for many corporate clients. How do you isolate tenants?
Answer: Choose an isolation model per tenant tier, and prove it with tests.
What I would check:
- Pooled tenants: one shared Knowledge Base with mandatory tenant-ID metadata filters applied server-side.
- Premium or regulated tenants: a dedicated Knowledge Base or a dedicated account.
- Cost attribution through application inference profiles or per-tenant request metadata.
- Per-tenant rate limits at a gateway, and per-tenant guardrails where policies differ.
Production consideration: Run cross-tenant leakage tests in CI on every release. See multi-tenant AI SaaS architecture.
Troubleshooting scenarios
75. Scenario: Converse returns AccessDeniedException for a model that works in another account.
Answer: Work through the permission chain from the outside in.
What I would check:
- Whether model access is enabled for that model in this account and Region, including any provider agreement.
- The IAM policy resource ARNs: foundation-model ARN versus inference-profile ARN, and whether a profile call also needs permission on the destination Regions' model ARNs.
- SCPs denying the Region, the model, or
aws:RequestedRegion"unspecified" for global profiles. - VPC endpoint policies that allow only certain actions.
- Whether it is actually the Agents Classic maintenance-mode error, if the call was CreateAgent or InvokeInlineAgent in a new account.
Production consideration: Use the IAM policy simulator and CloudTrail error details rather than trial and error, and capture the working policy as code.
76. Scenario: a Knowledge Base sync finishes, but retrieval returns nothing relevant for obvious questions.
Answer: Confirm the data is in the index before tuning the search.
What I would check:
- Ingestion job statistics: documents scanned, indexed and failed, and the failure reasons (unsupported format, size limits, parsing errors).
- Whether scanned PDFs went through a parser that extracts text at all.
- Whether the embeddings model and vector dimensions match the index configuration.
- Whether a metadata filter in the query is silently excluding everything (test without it).
- Chunk content: retrieve with a high
numberOfResultsand read what is actually stored. - Whether hybrid search is available for this store and needed for exact codes.
Production consideration: Add a smoke test after every sync that runs known questions and checks the expected document IDs come back.
77. Scenario: answers are correct but slow during business hours, with no errors.
Answer: Break the latency down by step before changing anything.
What I would check:
- Traces split into retrieval, reranking, guardrails and generation time.
- Time-to-first-token versus total time, and output tokens per second, to see whether outputs got longer or generation slowed.
- Input token growth from more or larger chunks, or from long conversation history.
- Throttle-driven retries hidden inside the SDK (check
InvocationThrottles). - Vector store capacity during peaks (OpenSearch Serverless capacity or Aurora CPU).
Production consideration: Fix the dominant step first. Common wins are streaming, trimming context, prompt caching, a cross-Region profile, and the Priority tier for the customer-facing path.
78. Scenario: the model ignores retrieved context and answers from general knowledge.
Answer: Check whether the right context reached the model before blaming the model.
What I would check:
- Whether the right chunks were actually retrieved and included (inspect the final prompt in invocation logs).
- Whether a custom RetrieveAndGenerate prompt template dropped
$search_results$or$output_format_instructions$, which also removes citations. - Prompt wording: is there an explicit instruction to answer only from the context and to say when the answer isn't there?
- Whether the context is so long that the relevant passage is buried. Reranking and fewer chunks help.
- Whether contextual grounding checks are enabled to catch ungrounded answers.
Production consideration: Add faithfulness to the evaluation suite so this regression is caught before release next time. LLM hallucinations explained covers the underlying causes.
79. Scenario: an agent on AgentCore loops, calling the same tool repeatedly, and costs spike.
Answer: Look at the tool's responses before looking at the model's reasoning.
What I would check:
- The trace: is the tool returning an error or an empty result the model keeps retrying?
- Whether the tool description and schema are ambiguous, so the model can't tell success from failure.
- Whether the tool returns a clear, structured error with a next step ("order not found; ask the user for the order ID").
- Maximum iteration and token budget limits on the loop.
- Whether memory is feeding stale state back into each turn.
Production consideration: Enforce hard caps, alarm on tool calls per session, and add the failing conversation to the agent evaluation set. See AI agent evaluation.
80. Scenario: a fine-tuned model scored well in testing but performs worse than the base model in production.
Answer: Suspect the data before the model.
What I would check:
- Distribution shift: do production inputs look like the training and test data (length, language, format, channel)?
- Leakage: did test examples overlap with training data, inflating scores?
- Prompt mismatch: is production using a different system prompt or template from training?
- Overfitting: was it trained for too many epochs on a small dataset, losing general ability?
- Whether the failures cluster in tasks the base model handled through general knowledge.
Production consideration: Keep the base model behind a feature flag for quick rollback, compare both models on a held-out production sample, and consider whether RAG or distillation would have been the better tool.
Key takeaways
- Bedrock interviews reward reasoning about why: governance, cost, latency and data residency drive service choices more than model benchmarks do.
- Know the inference menu (on-demand, service tiers, batch, Provisioned Throughput, cross-Region profiles, caching, routing) and when each one pays off.
- Knowledge Bases are a data-engineering problem first: parsing, chunking, metadata, permissions and sync freshness decide answer quality.
- Agents Classic is in maintenance mode and closed to new customers. New agent work goes on AgentCore, using any framework, including Strands and LangGraph.
- Guardrails are policy. Tune them with labelled data, version them, and enforce them centrally.
- Evaluation gates every change to a prompt, model or retrieval setting, in CI and on live traffic.
- Security answers need evidence: IAM, PrivateLink endpoints, KMS, CloudTrail data events and invocation-log handling.
Interview preparation checklist
- Call a model with Converse and ConverseStream from Python, and log
usagefrom every response. - Build a Knowledge Base on your own documents. Try two chunking strategies and compare retrieval on ten real questions.
- Implement Retrieve plus Converse with a metadata filter derived from a user's role.
- Create a guardrail with denied topics, PII masking and grounding checks, and test it with ApplyGuardrail.
- Run a Bedrock RAG evaluation or a Ragas evaluation, and explain each metric.
- Deploy a small Strands or LangGraph agent with one tool, ideally on AgentCore Runtime, and read its trace.
- Write a least-privilege IAM policy and create a bedrock-runtime VPC endpoint in a test account.
- Enable invocation logging and CloudTrail data events, then query token usage by role in Logs Insights.
- Prepare a cost estimate for one design with explicit assumptions.
- Rehearse three scenario answers out loud using the "what I would check" structure.
- Read the current Bedrock and AgentCore "what's new" pages the week before the interview.
FAQ
What skills are required for an AWS Bedrock role?
Python, core AWS (IAM, VPC, S3, Lambda, CloudWatch), RAG fundamentals, prompt engineering, evaluation methods and basic agent design. Strong candidates also understand security, cost and Infrastructure as Code.
Is AWS Bedrock good for freshers to learn?
Yes. It lets freshers build real GenAI applications with managed services and learn the cloud security and operations skills employers expect, without running GPU infrastructure.
How should I prepare for an Amazon Bedrock interview?
Build one small RAG project and one agent on your own AWS account, practise explaining the design choices, and rehearse scenario answers on cost, latency, permissions and failures.
Do I need to know machine learning to work with Bedrock?
Not deeply for most roles. You need to understand embeddings, context windows, evaluation and when fine-tuning helps. Training models from scratch is rarely part of Bedrock application work.
Is Bedrock Agents still worth learning in 2026?
Learn its concepts, because existing customers still run it, but focus new effort on Bedrock AgentCore. AWS has put Agents Classic into maintenance mode and closed it to new customers.
What is the difference between Bedrock and SageMaker AI?
Bedrock gives managed API access to foundation models and GenAI building blocks. SageMaker AI is for building, training and hosting your own machine learning models with more infrastructure control.
Which AWS certification helps for Bedrock roles?
AWS offers AI and machine learning certifications alongside its associate-level cloud certifications. Check AWS Training and Certification for the current list, and pair any certification with a deployed project.
Is an AWS Bedrock career a good choice?
For engineers who like combining cloud, data and AI, it is a practical path. Enterprises on AWS increasingly need people who can take GenAI from prototype to governed production.
How long does it take to prepare for a Bedrock interview?
If you already know AWS, a few focused weeks of hands-on building and scenario practice is typical. Freshers usually need longer, starting with Python and cloud basics.
Interviewers can tell when a candidate has only read about Bedrock and when they have shipped something on it. Cloudsoft's FDE PRO program is a 12-week, project-led course (120+ hours live, 60+ labs, five enterprise projects and the GlobalBank capstone) built on Amazon Bedrock, LangGraph, MCP, Ragas and AWS, with placement support until you're placed. If you want a broader AI, ML, cloud and cyber security foundation first, look at APEX. For a focused Bedrock track, see the Amazon Bedrock and GenAI course. Classes run in Ameerpet (beside Ameerpet Metro) or live online. For a free demo, call +91 96660 19191.



