These Bedrock AgentCore interview questions are aimed at engineers who will build, secure and run AI agents on AWS. Amazon Bedrock AgentCore is AWS's agentic platform for deploying and operating agents built with any framework and any foundation model, so interviewers care less about one SDK call and more about whether you understand Runtime, Gateway, Memory, Identity, Policy, Observability and Evaluations as one production system. The 50 questions below run from fundamentals to architecture and real-world scenarios. Each answer is written the way a senior engineer would explain it in an interview.
How to use this guide
- Freshers and career switchers: make sure you can answer questions 1β20 clearly. Interviewers check that you know what each AgentCore service is for and how a session works.
- Developers and AI engineers: spend your time on Runtime, Gateway, Memory, frameworks and MCP/A2A (questions 7β18 and 30β34). Expect follow-ups on session handling, tool design and state.
- Cloud, DevOps and platform engineers: security, identity, observability, AgentOps and cost (questions 21β29 and 39β40) come up most at your level.
- Senior and architect roles: architecture, migration and scenario questions (35β50) decide the outcome. Talk about trade-offs and failure modes, not only features.
AgentCore changes quickly. Before an interview, skim the current AgentCore Developer Guide so you can say which features you have actually used and which you have only read about. Being honest about that boundary comes across well.
Contents
- AgentCore fundamentals (Q1βQ6)
- AgentCore Runtime (Q7βQ11)
- AgentCore Gateway (Q12βQ15)
- AgentCore Memory (Q16βQ18)
- Code Interpreter and Browser (Q19βQ20)
- Security, Identity and Policy (Q21βQ25)
- Observability, Evaluations and AgentOps (Q26βQ29)
- Frameworks on AgentCore (Q30βQ32)
- MCP and A2A (Q33βQ34)
- Architecture (Q35βQ36)
- Migration from Bedrock Agents (Q37βQ38)
- Cost (Q39βQ40)
- Real-world scenarios (Q41βQ50)
- Key takeaways
- Interview preparation checklist
- FAQ
AgentCore fundamentals
1. What is Amazon Bedrock AgentCore, and what problem does it solve?
Answer: AgentCore is a set of managed AWS services for building, deploying and operating AI agents securely at scale. It works with any open-source framework (Strands Agents, LangGraph, CrewAI, LlamaIndex and others) and any foundation model, inside or outside Amazon Bedrock. The problem it solves is the gap between a working agent on a laptop and one that can serve many users in an enterprise. That gap includes per-user isolation, long-running sessions, secure tool access, credential handling, memory, policy enforcement, tracing and quality evaluation. Teams used to build all of that themselves. AgentCore offers each piece as a managed service, and you can use the pieces together or separately.
Interview tip: Call it "agent infrastructure" rather than "an agent framework". AgentCore hosts and governs agents. It does not replace the framework that writes the reasoning loop, except in the managed harness (see Q5).
2. What does "framework-agnostic and model-agnostic" mean in practice?
Answer: AgentCore Runtime defines a service contract (HTTP endpoints and ports) rather than a programming model. Anything that implements the contract can be hosted: a Strands agent, a LangGraph graph, a CrewAI crew, a Google ADK or OpenAI Agents SDK agent, or plain Python with no framework. The model is called from your code, so it can be a Bedrock model, an OpenAI or Gemini model, or any other endpoint your code can reach. In practice, the framework choice belongs to the team writing the agent and the platform choice belongs to the team running it. That separation is useful in large organisations where different teams have already standardised on different frameworks.
Real-world example: Consider a GCC in Hyderabad where one team built a LangGraph claims assistant and another built a Strands IT-ops agent. Both can run on the same Runtime with the same identity, gateway and observability setup, without either team rewriting its agent.
3. Name the AgentCore services and what each one does.
Answer: The core services, by their current names, are:
| Service | What it does |
|---|---|
| Runtime | Serverless hosting for agents and tools, with an isolated microVM per session, long-running sessions, streaming and HTTP/MCP/A2A/AG-UI protocols |
| Harness | A managed agent loop: you declare the model, tools and instructions as configuration, and AgentCore runs it |
| Gateway | Turns APIs, Lambda functions and existing services into MCP tools, connects to existing MCP servers, and fronts agents and model providers behind one secured endpoint |
| Memory | Managed short-term (session events) and long-term (extracted records) memory |
| Identity | Agent workload identities, inbound authentication, and outbound credentials (OAuth, API keys) kept in a token vault |
| Code Interpreter | Sandboxed code execution (Python, JavaScript, TypeScript) |
| Browser | Managed, isolated cloud browser for agents that need to use web applications |
| Observability | OpenTelemetry-compatible traces, metrics and logs stored in Amazon CloudWatch |
| Evaluations | LLM-as-a-judge quality scoring of sessions, traces and spans, with built-in and custom evaluators |
| Policy | Deterministic, Cedar-based authorization of every tool call that passes through Gateway |
The AgentCore documentation also lists newer capabilities: Optimization (recommendations and A/B tests built on Evaluations), Registry (a governed catalog of agents, MCP servers, tools and skills) and Payments (agent micropayments using protocols such as x402). Check which of these are available in your Region before you claim hands-on experience with them.
4. Can you use AgentCore services independently? Give an example.
Answer: Yes. AWS describes the services as modular, usable together or on their own. One common pattern is to keep an existing agent where it already runs, for example on Amazon EKS, and adopt only Gateway (tools as MCP), Identity (credentials) or Memory. AgentCore Identity explicitly supports agents running on AgentCore Runtime, in self-hosted environments or in hybrid deployments. Observability uses OpenTelemetry, so telemetry fits the tools you already have. Adopting incrementally is often the realistic path, because a customer will not re-platform a working agent just to get a token vault.
Interview tip: Saying that an incremental path exists, instead of assuming everything moves to Runtime on day one, shows you have worked with real customers.
5. What is the AgentCore harness, and when would you choose it over a code-defined agent on Runtime?
Answer: The managed harness is a configuration-based agent. You declare the model, tools, skills and instructions, and AgentCore provides the orchestration loop, compute, memory, identity, networking and observability. Each harness session is stateful and runs in its own isolated microVM backed by Runtime, with a filesystem and shell. AWS documents that the harness is built on Strands Agents, that it can call models from Bedrock, OpenAI, Gemini and other compatible providers, and that it has no separate charge beyond the AgentCore capabilities it uses. AWS's guidance is to start with the harness unless you need to own the loop. Code-defined agents on Runtime are the right choice for custom orchestration, multi-agent supervisor or routing patterns, an existing framework codebase, or stage-specific prompt control. A harness can also be exported to Strands code when configuration is no longer enough.
6. How does AgentCore relate to the rest of Amazon Bedrock?
Answer: Bedrock still provides model inference, Knowledge Bases and Guardrails. AgentCore is the agent layer that sits on top of them and runs alongside them. An agent on AgentCore can call Bedrock models, reach a Knowledge Base through a Gateway-fronted retrieval tool, and keep Bedrock Guardrails on the model while Gateway enforces agent-level policy. The original Bedrock Agents service is now called Amazon Bedrock Agents Classic and is in maintenance mode (see Q37). For a broader view of where these services fit in an AWS AI platform, see AWS for AI and FDE engineers. For Bedrock-specific question practice, use the AWS Bedrock interview questions.
AgentCore Runtime
7. How does AgentCore Runtime host an agent? Walk through the service contract.
Answer: Runtime hosts a containerised application, either a container image in Amazon ECR or code that the AgentCore CLI packages and uploads. The application must implement a protocol contract. For the HTTP protocol, the container listens on port 8080 and serves /invocations (with /ws for WebSocket) and a /ping health endpoint. MCP servers listen on port 8000 at /mcp. A2A servers listen on port 9000 at the root path. AG-UI uses port 8080. The AgentCore Python SDK's BedrockAgentCoreApp with an @app.entrypoint decorator implements the HTTP contract for you. Clients call the InvokeAgentRuntime API (or the WebSocket variant) with the runtime ARN, a session ID and a payload. Runtime routes each session to its own microVM.
Interview tip: Interviewers like hearing that the contract is what makes Runtime framework-agnostic. If you can serve /invocations and /ping, Runtime can host you.
8. Explain session isolation in AgentCore Runtime. Why does it matter for agents?
Answer: Each user session runs in a dedicated microVM with its own CPU, memory and filesystem. When the session ends, the microVM is terminated and its memory is sanitized. Agents need this more than ordinary APIs do because they hold stateful reasoning context, run privileged tool calls on a user's behalf, and behave non-deterministically. A hardware-level boundary per session means one user's context, files or credentials cannot leak into another session, even if the agent code has a bug. There is an important catch: AgentCore does not enforce which user owns which session ID. Your backend has to map users to session IDs, enforce limits such as the maximum number of sessions per user, and never let a client choose another user's session ID.
9. How does Runtime support long-running and asynchronous agents?
Answer: Unlike a request-scoped function, a Runtime session stays alive between invocations. At the time of writing, a session lifecycle can last up to 8 hours on microVMs (up to 14 days on the Instances compute type). The documented default idle timeout is 15 minutes, and both values are configurable through lifecycle settings (idleRuntimeSessionTimeout, maxLifetime). A session moves between Active, Idle and Stopped states. For background work, the agent reports HealthyBusy from /ping so Runtime knows the session is still working even with no open request. If a stopped session is invoked again, Runtime provisions new compute. In-memory state is lost unless you configured persistent session storage or saved state to AgentCore Memory.
Real-world example: An insurer's document-review agent accepts a batch of claim files, returns "accepted" right away, keeps processing in the background while reporting HealthyBusy, and writes results somewhere durable for the client to poll.
10. How do Runtime versions and endpoints work, and how would you roll back a bad release?
Answer: Every configuration change, such as a new container image or a protocol or network setting, creates a new immutable version. Version 1 is created with the runtime. A DEFAULT endpoint is created automatically and follows the latest version. You can create your own endpoints (for example dev, staging, prod) that each point to a specific version, and update them without downtime. A controlled release keeps production clients on a named endpoint, deploys a new version, tests it through a staging endpoint, then moves the production endpoint to the new version. Rolling back means pointing the endpoint at the previous version again. Do not point production clients at DEFAULT, because every update would go straight to them.
11. What compute and platform choices does Runtime offer, and how would you pick?
Answer: There are two compute types. microVMs are fully managed and serverless, scale on demand and are billed for what you use. Instances run on AWS-managed Amazon EC2 capacity in your account and support persistent multi-day sessions, GPU workloads and several cooperating agents on one instance. For microVMs there is also a platform version setting. V1 is the default. V2 starts agents from a snapshot so cold starts stay consistent regardless of image size or concurrency. V2 has trade-offs: create and update take minutes, the container must report healthy quickly, the documented environment-variable size limits are smaller, and Region coverage is narrower. Choose microVMs for interactive, bursty, per-user agents. Choose Instances for long-lived, GPU-bound or co-located workloads. Evaluate V2 when cold-start latency on large images is a real user complaint.
AgentCore Gateway
12. How does AgentCore Gateway turn APIs and Lambda functions into MCP tools?
Answer: You create a gateway, which is an MCP endpoint, and attach targets to it. For a Lambda target you supply a tool schema describing the function's inputs. For an OpenAPI or Smithy target, Gateway reads the specification and exposes the operations as tools. When an agent sends an MCP tools/call, Gateway translates it into the matching Lambda invocation or HTTP API request, injects the right outbound credentials, and returns the result in MCP format. MCP targets work in aggregation mode, so the agent sees a single tools/list covering every attached target. Teams expose existing enterprise APIs as agent tools without writing or hosting a separate MCP server for each one.
Interview tip: If asked "why not just call the REST API from the agent?", explain that MCP gives you one discoverable tool interface across many backends, and that centralising tools in Gateway centralises authentication, policy and audit as well. What is MCP covers the protocol basics.
13. What target types does Gateway support?
Answer: AWS documents three categories:
- MCP targets (aggregated into one virtual MCP server): AWS Lambda functions, Amazon API Gateway REST API stages, OpenAPI schemas, Smithy models, existing MCP servers, built-in templates from integration providers, and built-in connectors. AWS lists one-click integrations for tools such as Salesforce, Slack, Jira, Asana and Zendesk.
- HTTP targets: traffic is passed straight through without translation. This is how Gateway fronts other agents, including AgentCore Runtime agents and A2A traffic.
- Inference targets: model-based routing of LLM requests across providers through one endpoint.
You can attach a different credential provider to each target, so a Salesforce target and an internal Lambda target can authenticate in completely different ways behind the same gateway.
14. What is semantic tool search in Gateway, and why does it matter?
Answer: As the tool catalogue grows, putting every tool definition into the prompt costs tokens and latency and makes the model more likely to pick the wrong tool. Gateway's semantic tool selection lets an agent search the available tools by task context and load only the relevant ones. AWS positions this as the way to let agents work with very large tool collections while keeping prompts small. Search does not fix bad tool design, though. Names and descriptions still have to be specific and distinct, because the search and the model both rely on them.
15. How does Gateway handle authentication on both sides?
Answer: Gateway handles inbound and outbound authentication separately. Inbound, it verifies who is calling the gateway, typically with OAuth/JWT from your identity provider or with IAM. Outbound, it fetches and injects the credential each target needs, such as an IAM role for Lambda, an API key, or an OAuth token (including three-legged OAuth at target level for user-delegated access). Credentials are stored through AgentCore Identity rather than in agent code. Gateway is also where Policy is enforced (Q24), where Bedrock Guardrails can be applied to agent traffic, and where request and response interceptors can be added. That makes it the natural control point in an enterprise design.
AgentCore Memory
16. Explain short-term and long-term memory in AgentCore.
Answer: Short-term memory stores raw interactions as events, written with CreateEvent and read back with ListEvents, GetEvent and ListSessions. It is scoped by actor and session, so an agent can rebuild a conversation exactly, even after the Runtime session has ended. Long-term memory stores memory records extracted from those events, such as facts, preferences and summaries, and keeps them across sessions. Extraction runs asynchronously in the background, and consolidation merges new insights with existing records. Agents usually query long-term memory with RetrieveMemoryRecords, which does a semantic search for the records most relevant to the current request. If you configure no strategies, no long-term records are extracted.
For the broader design space (working memory, episodic memory and forgetting), see AI agent memory explained.
17. What are memory strategies, and how would you choose between them?
Answer: A strategy defines what gets extracted from raw events into long-term memory. The built-in strategies are semantic (facts and knowledge), user preference, summary (session summaries) and episodic (episodes with reflection across them). There are three ways to run them. Built-in strategies are fully managed with little configuration. Built-in with overrides let you change the extraction prompts and use Bedrock models invoked in your account. Self-managed strategies give you the whole pipeline, including schemas and namespaces. A single memory resource can combine several strategies. Pick based on the use case: a support agent usually needs user preference plus summary, a research agent benefits from semantic memory, and a regulated domain with strict rules about what may be remembered may need overrides or self-managed extraction.
18. How do you keep one user's memory from reaching another user or tenant?
Answer: Scope every write and read with a trustworthy actorId and sessionId, and organise long-term records into namespaces by actor, session or strategy. AWS's AgentOps guidance names namespaces as the memory-isolation control. Derive the actor ID from the authenticated identity on the server side, never from a value the client sends. Add IAM permissions that limit which roles can read which memory resources. For multi-tenant SaaS, decide between a separate memory resource per tenant and shared resources with tenant-prefixed namespaces, based on your isolation and blast-radius requirements (see multi-tenant AI SaaS architecture). Also remember that event metadata is not encrypted with a customer-managed key, so keep sensitive data out of it.
Code Interpreter and Browser
19. When would an agent use AgentCore Code Interpreter, and how do you run it safely?
Answer: Use it when the task needs exact computation rather than reasoning: data analysis on CSV or Excel files, calculations, chart generation or file transformation. Code runs in an isolated sandbox with pre-built Python, JavaScript and TypeScript runtimes and common libraries. Large datasets can be referenced from Amazon S3, and execution is logged in CloudTrail. The documented default execution time is 15 minutes, extendable to 8 hours. To run it safely, choose the most restrictive network mode the task allows, give it an execution role scoped to the specific S3 prefixes it needs, close sessions when you are done, and treat any code the model generates as untrusted, because it is.
20. What is AgentCore Browser used for, and what operational features does it have?
Answer: Browser gives an agent a managed, isolated cloud browser for web applications that have no API: navigating, filling forms, clicking and extracting data. Sessions are ephemeral, with a configurable timeout (documented default 15 minutes, maximum 8 hours). Agents drive the browser through WebSocket-based automation with libraries such as Playwright, Strands or Nova Act. A Live View endpoint lets a human watch the session and take over, which is useful for human-in-the-loop steps. Custom browsers can record sessions, including DOM changes, actions, console logs and network events, to your S3 bucket for replay. Session recording is also strong audit evidence when an agent acts on a regulated web portal.
Security, Identity and Policy
21. Explain inbound and outbound authentication in AgentCore.
Answer: Inbound auth decides who may invoke an agent or gateway. Runtime supports IAM SigV4 (the default) or JWT bearer tokens from an OpenID Connect provider such as Amazon Cognito, Okta or Microsoft Entra ID. JWT validation uses a discovery URL, allowed audiences, clients and scopes, and optional required custom claims. A given runtime accepts SigV4 or JWT, not both at once, although different versions can be configured differently. Outbound auth decides how the agent reaches downstream systems. AgentCore Identity supplies OAuth tokens or API keys, either user-delegated (acting as the end user) or autonomous (acting with its own service credentials). Keeping the two concerns separate makes it possible to answer two different questions: "is this user allowed to talk to this agent?" and "is this agent allowed to touch this system, as whom?"
For identity patterns beyond AWS, read AI agent identity and access management.
22. What are 2LO and 3LO flows in AgentCore Identity, and what is the token vault?
Answer: 2LO is the OAuth 2.0 client credentials grant: machine-to-machine, with no user involved, used when the agent acts for itself. 3LO is the authorization code grant: the user is sent to a consent screen and the agent receives a token scoped to that user's data. The token vault stores OAuth tokens, client credentials and API keys, encrypted with AWS KMS (customer-managed or service-managed keys). A credential is released only to a workload that presents verifiable proof of identity. In code, SDK decorators such as @requires_access_token and @requires_api_key fetch and inject credentials. With a 3LO flow, the SDK returns an authorization URL when consent is needed. Identity includes pre-configured credential providers for services such as Google, GitHub, Slack, Salesforce and Atlassian, plus configurable providers for any OAuth 2.0 server.
23. What is a workload identity, and how does a user's token become a downstream token?
Answer: A workload identity is the agent's own identity in the AgentCore Identity directory, with its own ARN. Runtime creates one automatically for each agent runtime. In the documented flow, Runtime first validates the user's inbound JWT. It then exchanges that JWT for a workload access token (GetWorkloadAccessTokenForJWT) and passes it to the agent code. When the agent needs, for example, the user's Google Drive, it uses the workload access token to request an OAuth token from the vault. If the user has not yet consented, the agent receives a 3LO consent URL. The resulting token is cached in the vault, keyed by the workload identity together with the user ID, so the user is not asked to consent again until the token expires. This keeps raw downstream credentials out of prompts, logs and agent memory.
24. What is Policy in AgentCore, and how does it differ from IAM and Bedrock Guardrails?
Answer: Policy lets you create a policy engine, store deterministic authorization policies in it, and attach it to a gateway. Every tool call through that gateway is intercepted and checked before it executes. Policies are written in Cedar, AWS's open-source authorization language, or drafted in natural language, which AgentCore converts to Cedar, validates against the tool schema, and checks with automated reasoning for rules that are too permissive, too restrictive or impossible to satisfy. Rules can use the principal's identity, the tool and the tool's input parameters. Temporal policies can also take into account what already happened earlier in the session. The differences from the other two controls:
- IAM controls which AWS principals can call which AWS APIs. It knows nothing about an agent's tool arguments.
- Guardrails filter model inputs and outputs, such as harmful content, denied topics and sensitive information.
- Policy checks "may this agent, for this user, call this tool with these parameters right now?" It runs outside the agent's code, so a prompt injection cannot talk its way past it.
Interview tip: The phrase interviewers want to hear is "deterministic controls outside the model". Prompts are guidance. Policy is enforcement.
25. How do you stop clients from bypassing your gateway and calling the agent runtime directly?
Answer: If you put Gateway in front of a runtime for policy, Guardrails and interceptors, then direct access to the runtime has to be closed. Otherwise the controls are optional. For a runtime using SigV4, attach a resource-based policy that allows only the gateway's execution role to invoke the runtime and explicitly denies every other principal. Then lock down the gateway role's trust policy with aws:SourceArn and aws:SourceAccount conditions so only your gateway can assume it. For a runtime using JWT, set allowedWorkloadConfiguration on the JWT authorizer so requests are accepted only when the identity chain includes your gateway. A related trap is the X-Amzn-Bedrock-AgentCore-Runtime-User-Id header. AgentCore does not verify that header, so grant bedrock-agentcore:InvokeAgentRuntimeForUser only to trusted principals, or deny it where you don't need it.
Interviewers in this area want someone who can connect identity providers, OAuth, policy, cloud networking and observability across a whole customer deployment. That end-to-end integration work is the focus of the AI Forward Deployed Engineer course at Cloudsoft, where the Secure Banking AI Assistant and ServiceNow AI Agent via MCP projects are built against these concerns.
Observability, Evaluations and AgentOps
26. How does AgentCore Observability work?
Answer: AgentCore emits telemetry in an OpenTelemetry-compatible format and stores metrics, spans and logs in Amazon CloudWatch. Built-in metrics cover agents, gateways and memory resources, including session count, latency, duration, token usage and error rates. Memory spans and logs are available when you enable them. For agents on Runtime, the CloudWatch console provides an agent observability dashboard with trace visualisations and error breakdowns. You can instrument your own code with OpenTelemetry to add spans and custom metrics, and because the format is standard you can also send data to an observability stack you already run. A good trace shows the full path of a request: the inbound request, each model call with its token counts, each tool call through Gateway with its latency and result, memory reads and writes, and policy decisions.
For the general discipline, see AI observability.
27. What is AgentCore Evaluations, and how do online and on-demand evaluation differ?
Answer: Evaluations scores agent quality using LLM-as-a-judge techniques on sessions, traces and spans. It works with traces from Strands and LangGraph agents instrumented with OpenTelemetry or OpenInference. Built-in evaluators have public ARNs such as Builtin.Helpfulness. Custom evaluators are private resources controlled by IAM. On-demand evaluation scores specific traces or spans, which suits development and pre-release gates. Online evaluation continuously samples and scores live production traffic, which suits detecting drift. The documentation also covers batch evaluation, dataset evaluation and simulation. Results appear in AgentCore Observability in CloudWatch, so a drop in a quality score can trigger an alarm just as a latency spike does.
Evaluation design itself, including test sets, rubrics and judge calibration, is covered in AI agent evaluation.
28. What is AgentOps, and what does AWS recommend for it?
Answer: AgentOps is the operational discipline for deploying, managing and continuously improving agents in production. AWS's AgentOps guidance for AgentCore is organised into four pillars:
- Governance and security: a multi-account strategy, deterministic controls (Policy with Cedar at Gateway), reasoning controls and human-in-the-loop, Identity for authentication and authorization, and Memory namespaces for isolation.
- Build and operations: agents treated as versioned artefacts, separate repositories for infrastructure, agents, tools and applications, CI/CD that builds container images to ECR and deploys to Runtime, and memory configuration treated as versioned infrastructure.
- Evaluation: evaluation at the tool, conversation-turn, session-outcome and system levels, on-demand evaluation in development, online evaluation in production, and a rule that a build does not reach production until evaluation passes.
- Observability and monitoring: framework, service, infrastructure and application telemetry layers, so you can see what the agent decided, what AgentCore did, how the compute behaved and what the business outcome was.
29. Design a CI/CD pipeline for an agent on AgentCore with quality gates.
Answer: On every commit: run unit tests for tools and prompt templates, build the image, push it to ECR, and create a new Runtime version. Point a staging endpoint at the new version. Replay a fixed evaluation dataset against staging and score it with on-demand Evaluations. Block promotion if task success, tool-selection accuracy or safety scores fall below the agreed thresholds. On pass, move the production endpoint to the new version, ideally after a canary period. Keep the previous version so rollback is a single endpoint update. Keep Gateway targets, Policy and Memory configuration in infrastructure as code with their own pipelines, since a policy change can break an agent as easily as a code change. Where Optimization is available, its traffic splitting through Gateway can A/B test prompt or tool-description changes with significance reporting before full rollout.
commit -> tests -> image to ECR -> new Runtime version
-> staging endpoint -> eval dataset + Evaluations
-> pass? -> move prod endpoint (canary) -> online eval
-> fail? -> stop; prod stays on previous version
Frameworks on AgentCore
30. How do you deploy a Strands Agents agent to AgentCore Runtime?
Answer: Wrap the agent in BedrockAgentCoreApp, expose an entrypoint, and deploy with the AgentCore CLI (agentcore create, agentcore dev to test locally, then agentcore deploy). The SDK handles the HTTP contract and the /ping health check.
from strands import Agent
from bedrock_agentcore.runtime import BedrockAgentCoreApp
app = BedrockAgentCoreApp()
agent = Agent(tools=[lookup_order])
@app.entrypoint
def invoke(payload, context):
result = agent(payload.get("prompt", ""))
return {"result": result.message}
if __name__ == "__main__":
app.run()
Strands is AWS's open-source agent framework, and the harness is built on it, so it has the closest integration with AgentCore Memory, Gateway tools and A2A helpers. In an interview, also mention input validation in the entrypoint and passing context.session_id through to memory calls.
31. You have a LangGraph agent. What changes when you move it to AgentCore?
Answer: Very little of the graph changes. You add a BedrockAgentCoreApp entrypoint that invokes the compiled graph. The real decisions are about state and tools. LangGraph checkpoints held in process memory disappear when a Runtime session stops, so durable state has to live elsewhere: AgentCore Memory (which documents LangGraph and LangChain integration), session storage, or your own checkpoint store. Tools that were Python functions inside the graph can stay there, or move behind Gateway as MCP tools so identity, policy and audit are handled centrally. Human-in-the-loop interrupts need a resume path that sends the same session ID. Evaluations supports LangGraph traces instrumented with OpenTelemetry or OpenInference, so add that instrumentation as part of the move. The LangGraph interview questions cover the graph side in more depth.
32. Can CrewAI or other frameworks run on AgentCore? Are there any caveats?
Answer: Yes. AWS lists CrewAI, LlamaIndex, Google ADK and the OpenAI Agents SDK alongside Strands and LangGraph, and any custom code that implements the service contract can be hosted. The caveats are about depth of integration rather than whether it works at all. Some capabilities integrate more deeply with particular frameworks. For example, Evaluations documents Strands and LangGraph trace support, and Memory lists LangGraph, LangChain, Strands and LlamaIndex integrations. With other frameworks, check how you will instrument traces and connect memory. Multi-agent frameworks also need you to decide whether the agents run inside one Runtime session or as separate runtimes talking over A2A. For a framework-by-framework comparison, see AI agent frameworks compared.
MCP and A2A
33. When would you host an MCP server on Runtime instead of exposing tools through Gateway?
Answer: Gateway fits when the tool logic already exists as an API or Lambda function and you want MCP in front of it without writing a server. It adds aggregation, semantic search, credential injection and Policy. An MCP server on Runtime (port 8000, path /mcp) fits when the tool itself needs custom code, per-session state, long-running work or isolation, such as a tool that holds a working directory or a stateful MCP session. The two combine well: you can host the MCP server on Runtime and register it as an MCP server target behind Gateway, so agents still use one governed endpoint. The protocol is covered in our MCP explainer, and there is more practice in the MCP interview questions.
34. How does AgentCore support the A2A protocol?
Answer: Runtime can host A2A servers. The container runs a streamable HTTP server on port 9000 at the root path. Runtime acts as a transparent proxy, passing JSON-RPC payloads from InvokeAgentRuntime through unchanged while adding SigV4 or OAuth 2.0 authentication and per-session isolation. Standard A2A discovery still works: the agent card is served at /.well-known/agent-card.json, describing the agent's skills, endpoint and auth requirements. The AgentCore SDK provides a serve_a2a helper, and the CLI can scaffold A2A projects with Strands, LangChain/LangGraph or Google ADK. Gateway can also front A2A traffic through HTTP passthrough targets. Use A2A when separately owned agents need to delegate work to each other. Use MCP when an agent needs a tool. The A2A protocol explainer covers the protocol itself.
Architecture
35. Sketch a production architecture for an enterprise agent on AgentCore.
Answer: Users authenticate with the enterprise identity provider. The client calls Gateway with a JWT. Gateway applies Policy and Guardrails, then routes to the agent on Runtime, and direct access to the runtime is blocked (Q25). The agent reads and writes Memory with a server-derived actor ID, calls models, and reaches business systems only through Gateway tools, with credentials supplied by Identity. All telemetry goes to CloudWatch, and Evaluations scores sampled sessions.
User (web / Teams)
| JWT (Entra ID / Okta / Cognito)
v
AgentCore Gateway --- Policy (Cedar) + Guardrails
|
v
AgentCore Runtime (microVM per session)
|-- Memory (actor/session namespaces)
|-- Model (Bedrock or other provider)
|-- Identity (token vault, 2LO/3LO)
v
Gateway tools -> Lambda / OpenAPI / MCP servers
|
v
Observability (OTEL -> CloudWatch) -> Evaluations
Production consideration: Use separate AWS accounts for development, staging and production, as AWS's AgentOps guidance recommends. Use VPC connectivity for private backends. Write a runbook covering what happens when the model provider, a tool backend or the identity provider is degraded. See also enterprise AI architecture.
36. How would you design a multi-agent platform on AgentCore for many internal teams?
Answer: Treat AgentCore as a paved road. A platform team owns shared gateways with approved tool targets, shared Policy engines, Identity credential providers, observability standards and a CI/CD template. Product teams own their agents as separate runtimes. Where Registry is available, it gives you a governed catalog for publishing, reviewing and finding agents, MCP servers, tools and skills, which helps prevent duplicate tools and unknown "shadow" agents. Agents that work together communicate over A2A or use each other as tools through Gateway. Plan isolation from the start: each team gets its own execution roles, its own memory namespaces or resources, and its own cost tags. Without that, one team's misbehaving agent becomes everyone's incident.
Migration from Bedrock Agents
37. Is Amazon Bedrock Agents still available to new customers?
Answer: No, not to new customers. AWS renamed the original service Amazon Bedrock Agents Classic, and it closed to new customers on July 30, 2026. It is in maintenance mode. Accounts with Bedrock Agents activity in the previous 12 months are allowlisted and keep working. For accounts without that history, CreateAgent and InvokeInlineAgent return an AccessDeniedException. No new features are planned. The model catalogue inside the Classic orchestration layer is frozen at the maintenance-mode date, while Bedrock itself keeps adding models. AWS has announced no end-of-life date and no migration deadline, and it recommends AgentCore for new agent development. Bedrock Knowledge Bases and Guardrails are not affected.
Interview tip: State this precisely. "Bedrock Agents is deprecated" is not quite accurate. "Bedrock Agents Classic is in maintenance mode and closed to new accounts, and AgentCore is the recommended path" is.
38. How would you migrate a Bedrock Agents Classic agent to AgentCore?
Answer: AWS documents two target paths: the managed harness (closest to the Classic experience) and code-defined agents on Runtime. It also provides an amazon-bedrock migration skill in the agent toolkit for AWS, and an AgentCore CLI import, both of which inventory the existing agent and map it across. The component mapping:
| Bedrock Agents Classic | AgentCore equivalent |
|---|---|
| Action groups (OpenAPI/function schema + Lambda) | Gateway MCP tools wrapping the same APIs and Lambda functions |
| Knowledge Base attached to the agent | Gateway-fronted knowledge base or a retrieval tool |
| Session and memory settings | AgentCore Memory with strategies |
| Return of control / user input | Inline function tools in the harness |
| Code interpreter action group | AgentCore Code Interpreter |
| Guardrails on the agent | Bedrock Guardrails plus Gateway policy enforcement |
| Custom orchestrator, multi-agent collaboration | Code-defined agents on Runtime |
Production consideration: Make the gaps clear to stakeholders. Stage-specific prompt overrides (pre-processing, KB response generation, post-processing) are not directly replicated in the harness, and routing-mode multi-agent collaboration needs custom framework code. Run the old and new agents side by side on the same evaluation set before switching traffic.
Cost
39. How is AgentCore priced?
Answer: AgentCore is consumption-based, with no upfront commitment or minimum fee, and each service bills on its own usage. In general terms: Runtime microVMs, Code Interpreter and Browser bill per second for CPU and memory actually consumed. Runtime aims to avoid CPU charges during I/O wait, such as while waiting for a model response. Runtime Instances bill EC2 capacity plus a management fee. Gateway bills per API invocation, per search query and per tool indexed. Memory bills on events, storage and retrieval for short-term memory and on records and retrievals for long-term memory. Identity bills per token or API-key request for non-AWS resources. AWS documents no extra Identity charge when it is used through Runtime or Gateway. Policy bills per authorization request. Evaluations bills on tokens or evaluations performed. Observability follows CloudWatch pricing. The harness adds no separate charge. Model inference is billed separately by the model provider. Always check the current AgentCore pricing page instead of quoting rates from memory.
40. What levers do you have to control AgentCore cost?
Answer: The levers, roughly in order of impact:
- Model tokens: usually the largest line. Use semantic tool search so only relevant tools go into the prompt, put summaries in memory instead of replaying whole transcripts, and route easy steps to cheaper models.
- Session lifecycle: set idle timeouts and maximum lifetimes to match real usage, and call
StopRuntimeSessionwhen a conversation clearly ends. - Memory design: extract only the strategies you need. The documentation notes that built-in strategies cost more for storage than overrides or self-managed strategies.
- Telemetry volume: set CloudWatch log retention and sample verbose spans.
- Evaluation sampling: score a representative sample online rather than every session.
- Attribution: tag resources per team or tenant so cost conversations rely on data rather than guesswork.
Wider tactics are covered in cloud cost optimization for AI.
Real-world scenarios
41. A bank wants a customer-support agent that reads a logged-in customer's account data from core banking APIs. How do you design identity?
Answer: Every downstream call should carry the customer's own authority, never a broad service account. The agent should only be able to see what that customer could see in the mobile app.
What I would check:
- Inbound: Runtime or Gateway configured with a JWT authorizer pointing at the bank's identity provider, with allowed audiences, clients and scopes set.
- The actor ID for Memory taken from the validated token's subject, not from the request body.
- Core banking APIs exposed as Gateway tools, using a user-delegated flow or a token exchange the bank's API gateway supports, with credentials held in the Identity token vault.
- Policy rules restricting tools by customer segment and input parameters, for example read-only balance tools for retail customers.
- Direct runtime access blocked so every request passes through Gateway.
Production consideration: Expect the bank's security team to ask for CloudTrail and policy-decision logs, data-residency confirmation for the chosen Region, and a documented threat model covering prompt injection through account notes and transaction descriptions. The generative AI in banking article covers the domain controls.
42. An agent works locally with agentcore dev but fails after deployment to Runtime. How do you debug it?
Answer: Most of these failures come down to the contract, permissions or network. Work from the outside in.
What I would check:
- Runtime status and health. Does the container listen on the expected port and path for its protocol, and does
/pingreturn healthy promptly? On V2, health must be reported within the documented startup window. - CloudWatch logs for the runtime: import errors, missing dependencies, or an
424 RuntimeClientErrorpointing at the agent code. - The execution role: model invocation permissions, Memory and Gateway access, and ECR pull rights.
- Network: whether a VPC-configured runtime can reach the model endpoint and private backends.
- Configuration that differs from local: environment variables (watch the size limits), Region and model IDs.
- Authentication: whether the client uses the inbound auth type the runtime was configured with (SigV4 or JWT).
Production consideration: Add a deployment smoke test that invokes a staging endpoint after every deploy, so this is caught by the pipeline and not by users.
43. A user reports seeing details from someone else's conversation. What happened, and how do you fix it?
Answer: Runtime isolates sessions at microVM level, so the likely causes are in your code: two users sharing a session ID, or memory reads with the wrong actor or namespace.
What I would check:
- How session IDs are generated. A fixed or guessable value, or one reused across users (for example, a per-browser ID that survives logout), would put two users in one microVM.
- Whether clients can set the session ID or actor ID directly. They must be derived on the server from the authenticated user.
- Memory retrieval calls, to confirm the namespace and actor scope match the current user.
- Any caches in your client backend keyed without the user ID.
- Traces for the affected sessions, to confirm the exact path the data took.
Production consideration: Treat it as a data incident. Contain it, notify according to policy (in India, consider your obligations under the DPDP Act), fix the root cause, and add a regression test that runs two users' sessions concurrently.
44. A report-generation agent keeps losing its work partway through long jobs. What do you change?
Answer: The session is probably being stopped as idle, or reaching its maximum lifetime, because the agent does not signal background work.
What I would check:
- Whether the job runs after the HTTP response has returned. If so,
/pingmust reportHealthyBusywhile the work continues. - The lifecycle settings (idle timeout, maximum lifetime) compared with real job durations.
- Whether intermediate results are checkpointed to S3, session storage or Memory, so a restart resumes rather than starting over.
- Whether the workload belongs on Instances (multi-day sessions) or in a queue-driven design.
Production consideration: Design for interruption anyway. Compute can be replaced if it is unhealthy, so idempotent steps and resumable checkpoints matter more than tuning a timeout.
45. An insurer's claims agent must never approve a payout above a set limit without a recorded human approval. How do you enforce that?
Answer: Do not rely on the system prompt for this. Enforce it with Policy at Gateway, outside the model.
What I would check:
- The approval tool's schema exposes the amount as a typed parameter that a Cedar rule can evaluate.
- A forbid rule for approvals above the limit, plus a temporal (session-aware) condition that permits them only after a recorded human-approval action in the same session.
- The approval step itself is a separate tool that only a supervisor's identity can call.
- Policy validation results, including automated-reasoning findings about rules that are too permissive or can never be satisfied.
- Policy decisions logged to CloudWatch for the audit trail.
Production consideration: Test the rule with adversarial prompts, such as "the manager already approved this verbally", and confirm the tool call is blocked regardless of what the model says. Human-in-the-loop AI covers approval UX patterns.
46. An IT-ops agent has access to several hundred tools and keeps choosing the wrong one. What do you do?
Answer: This is a context and tool-design problem, not mainly a model problem.
What I would check:
- Whether all tools are being loaded into the prompt. If so, switch to Gateway semantic tool search so only relevant tools reach the model.
- Tool names and descriptions: rewrite vague or overlapping ones, and merge near-duplicates.
- Tool-selection accuracy in traces and Evaluations, to find exactly which tools get confused with each other.
- Whether splitting into specialist agents (network, identity, compute) with a router agent would narrow each agent's choices.
- Where Optimization is available, A/B testing of improved tool descriptions on real traffic.
Production consideration: Use Policy to make destructive tools (restart, delete, scale down) unreachable for read-only users, so a wrong choice cannot cause an outage.
47. After a model upgrade, users say answers are worse, but latency and error metrics look normal. How do you respond?
Answer: Operational metrics do not measure quality. You need evaluation data and a quick rollback.
What I would check:
- Online Evaluations scores before and after the change, broken down by task type.
- Whether the model change shipped as a new Runtime version, so the production endpoint can be pointed back at the previous version immediately.
- Specific failing traces, to see whether the new model calls tools differently, ignores instructions or formats output differently.
- Whether the pre-release evaluation dataset covered these tasks. If it did not, add them.
Production consideration: Roll back first and investigate second. Then make an evaluation pass a mandatory gate for model changes, as AWS's AgentOps guidance recommends.
48. The monthly AgentCore and model bill doubled with no matching growth in users. How do you investigate?
Answer: Break the cost down by service and by tag first, then find the behaviour behind it.
What I would check:
- Cost Explorer by service: model inference, Runtime, Memory, Gateway or CloudWatch.
- Token usage per session in Observability. Look for agents looping, retrying tools repeatedly or replaying long histories.
- Session durations against idle timeouts, and sessions that are never stopped.
- Gateway invocation and search counts per agent, which can reveal runaway tool calls.
- CloudWatch ingestion: a debug log level left on in production is a common cause.
- Online evaluation sampling rates.
Production consideration: Add per-session step and token limits in the agent loop, plus budget alarms per team tag, so the next anomaly raises an alert within a day and is not discovered on the invoice.
49. A GCC platform team runs LangGraph agents on EKS and asks whether it should move to AgentCore. What do you recommend?
Answer: Adopt by capability, not as one migration. Start where AgentCore removes the most custom work and risk.
What I would check:
- What they have built themselves today: credential storage, tool wrappers, memory, tracing. Those are candidates for Identity, Gateway, Memory and Observability, all of which can be used from agents still running on EKS.
- Pain points with per-user isolation and long-running sessions on EKS. Those are the strongest reasons to move hosting to Runtime.
- Requirements that need full cluster control (GPU scheduling, custom sidecars, on-premises connectivity), which may keep some workloads on EKS or point to Runtime Instances.
- Regional availability of the AgentCore features they need.
Production consideration: Pilot one agent end to end, compare operational effort and cost against the EKS baseline using real traces, and present the result as a decision with data behind it rather than a mandate.
50. A retailer wants an agent to update stock in a supplier portal that has no API. How would you build it safely?
Answer: Use AgentCore Browser to drive the portal, but treat it as the riskiest tool in the system and wrap it in controls.
What I would check:
- That the supplier's terms permit automated access, and whether an API, EDI feed or file upload exists that would be a more reliable choice.
- Portal credentials stored through Identity, never in prompts, with a dedicated low-privilege portal account.
- A custom browser with session recording to S3 so every change can be replayed for audit.
- Live View with a human approval step before submitting changes above an agreed threshold.
- Verification after each action: read the value back and compare it with the intended change.
Production consideration: Web UIs change without notice. Add a daily synthetic run that alerts when the expected page structure changes, and a fallback queue for manual processing.
Key takeaways
- AgentCore is agent infrastructure that works with any framework and any model. Know every service by its current name and the job it does.
- Runtime gives each session its own microVM and supports long-running sessions, but mapping users to session IDs is your responsibility.
- Gateway is the control point. Tools become MCP, credentials are injected per target, and Policy checks every tool call deterministically.
- Identity separates inbound authentication (who can call the agent) from outbound credentials (how the agent reaches systems, and as whom).
- Memory, Evaluations and Observability turn a demo into an operated system. Evaluation gates belong in CI/CD.
- Bedrock Agents Classic closed to new customers on July 30, 2026. Know the harness and code-defined migration paths and their gaps.
- Scenario answers carry the most weight. Give a diagnosis order, the controls you would add, and a production consideration.
Interview preparation checklist
- Deploy one agent to AgentCore Runtime with the AgentCore CLI and invoke it using two different session IDs.
- Expose one Lambda function and one OpenAPI spec through Gateway and call them as MCP tools.
- Configure JWT inbound auth with Cognito or Entra ID, then complete one 3LO outbound flow.
- Write one Cedar policy that blocks a tool call based on an input parameter, and test it with an adversarial prompt.
- Add short-term and long-term memory, and explain which strategy you chose and why.
- Open the traces in CloudWatch and walk through one full request, covering model, tool and memory spans.
- Run an on-demand evaluation and set a pass threshold you can defend.
- Practise drawing the Q35 architecture from memory in under five minutes.
- Prepare one story about debugging an agent and one about a security trade-off.
- Read the current Bedrock Agents Classic maintenance-mode page and the AgentCore pricing page the week before your interview.
FAQ
How should I prepare for a Bedrock AgentCore interview?
Build and deploy a small agent that uses Runtime, Gateway, Identity and Memory, then practise explaining its traces, security model and failure modes. Interviewers value hands-on detail, such as session handling and inbound versus outbound auth, more than memorised feature lists.
Do I still need to learn Bedrock Agents Classic?
Learn it at the level needed for migration. Bedrock Agents Classic is closed to new customers and in maintenance mode, but many organisations still run it, so knowing how action groups, knowledge bases and return of control map to AgentCore is a practical interview topic.
Is AgentCore only for models hosted on Amazon Bedrock?
No. AgentCore works with any foundation model your agent code can call, including models outside Bedrock. Bedrock remains the natural choice on AWS for model access, Knowledge Bases and Guardrails.
Which agent framework should I learn for AgentCore roles?
Strands Agents and LangGraph are the most useful to know, because AWS documents the deepest integrations with them, including Evaluations trace support. AgentCore hosts other frameworks too, so the concepts carry over.
What programming skills do AgentCore roles need?
Python is the main language in the AgentCore SDK examples, and the AgentCore CLI is installed through npm. You also need working knowledge of IAM, OAuth 2.0 and OpenID Connect, containers, infrastructure as code and OpenTelemetry.
Can freshers apply for roles that involve AgentCore?
Yes, if they can show real work. A fresher who has deployed an agent with Gateway tools, scoped identity and an evaluation report, and can explain every design decision, is more convincing than one with only course certificates.
What projects demonstrate AgentCore skills well?
Good options include an enterprise knowledge assistant with permission-aware retrieval, an IT helpdesk agent that calls ServiceNow or Jira through Gateway, or a banking assistant with user-delegated OAuth and Cedar policies. Each should include traces, an evaluation set and a short architecture write-up.
How is AgentCore different from LangGraph or CrewAI?
LangGraph and CrewAI are frameworks for writing an agent's reasoning and orchestration. AgentCore is the managed infrastructure that hosts, secures, connects and monitors agents built with those frameworks. They complement each other rather than compete.
Is AgentCore knowledge useful outside AWS-only roles?
Yes. The underlying ideas, such as session isolation, MCP tools, A2A, OAuth delegation, policy as code, OpenTelemetry tracing and evaluation gates, apply on any cloud. AgentCore is a concrete way to learn and demonstrate them.
If you want to go beyond interview answers and build these systems against realistic customer requirements, look at Cloudsoft's FDE PRO program: 12 weeks of live sessions, 60+ labs, five enterprise projects and the GlobalBank capstone, run as a simulated customer engagement. Placement support continues until you're placed, with resume, portfolio review and mock interviews. Classes run in Ameerpet (beside Ameerpet Metro) or live online. For a broader track that covers AI/ML, cloud and cyber security together, see the APEX AI, ML, Cloud and Security program. If you want Bedrock skills on their own, there is the Amazon Bedrock and GenAI course. Call +91 96660 19191 for a free demo.



