These AWS serverless interview questions cover what cloud, DevOps and backend interviews in 2026 actually test: AWS Lambda internals, API Gateway, Step Functions, SQS, SNS, EventBridge and DynamoDB, and how they fit together into event-driven systems that stay correct under retries and load. Interviewers rarely stop at "what is Lambda"; they ask why a function was throttled, why an SQS consumer charged a customer twice, or how you would run a job that needs more than 15 minutes, and they listen for idempotency, concurrency and failure handling in your answer. The 55 questions below run from fundamentals to architecture, serverless for AI workloads on Amazon Bedrock, and eleven production scenarios.
How to use this guide:
- Freshers and juniors are usually tested on the fundamentals: invocation models, limits, execution roles, SQS queue types and the difference between SNS, SQS and EventBridge. Be able to sketch an API Gateway, Lambda and DynamoDB request path on a whiteboard.
- Mid-level cloud and DevOps engineers get the cold start, concurrency, API Gateway and Step Functions questions, plus at least one debugging scenario.
- Senior and architect roles are pushed on idempotency, retry and DLQ design, cost modelling, security boundaries and when serverless is the wrong choice.
- Answer in three beats: the one-line answer, the mechanism, then the trade-off.
Contents
- Serverless and Lambda fundamentals (Q1โQ10)
- Lambda cold starts, concurrency and new compute options (Q11โQ18)
- API Gateway interview questions (Q19โQ24)
- Step Functions, EventBridge and DynamoDB (Q25โQ31)
- Idempotency, retries and dead-letter queues (Q32โQ35)
- Observability, security, cost and tooling (Q36โQ40)
- AI in serverless: Bedrock from Lambda (Q41โQ44)
- Real-world scenario questions (Q45โQ55)
- Key takeaways
- Interview preparation checklist
- FAQ
Serverless and Lambda fundamentals
1. What does "serverless" mean on AWS, and which services count as serverless?
Answer: Serverless means you do not provision, patch or capacity-plan servers; the service scales with demand (often to zero) and you pay for usage rather than for idle capacity. Servers still exist, but AWS owns them. On AWS the core serverless services are Lambda (compute), API Gateway and Lambda function URLs (HTTP entry points), SQS, SNS and EventBridge (messaging and events), Step Functions (orchestration), DynamoDB (database), S3 (storage), AppSync (GraphQL and real-time APIs), Cognito (user identity) and Fargate (containers without managing EC2 hosts).
Interview tip: Say that serverless is a spectrum. Fargate removes host management but still bills for running tasks; Lambda bills per invocation and duration. Showing you know the difference is a quick credibility win.
2. How does AWS Lambda work under the hood?
Answer: Lambda runs your code in an isolated execution environment (a Firecracker microVM on the default compute type) that has a runtime, your code and the memory you configured. Each environment goes through three phases:
- Init: Lambda starts extensions, bootstraps the runtime and runs your static code outside the handler (imports, SDK clients, connection setup). This phase is limited to 10 seconds; if it does not finish, Lambda retries it at the first invocation under the function timeout.
- Invoke: Lambda calls your handler with the event. One environment handles one request at a time on the default compute type.
- Shutdown: after a period of inactivity, Lambda freezes and eventually discards the environment.
When the next request arrives and a warm environment exists, Lambda reuses it and skips Init. When none is free, it creates a new one, which is a cold start.
3. What are the ways a Lambda function can be invoked?
Answer: There are three invocation models, and the retry behaviour differs for each:
| Model | Examples | Who retries |
|---|---|---|
| Synchronous | API Gateway, function URLs, SDK Invoke, Cognito triggers | The caller. Lambda returns the error and does not retry. |
| Asynchronous | S3 events, SNS, EventBridge rules | Lambda's internal queue retries (two more attempts by default), then sends to a destination or DLQ. |
| Event source mapping (polling) | SQS, Kinesis, DynamoDB Streams, Amazon MQ, Kafka | Lambda's poller. Behaviour depends on the source: SQS redelivers after visibility timeout; streams retry the batch until success or expiry. |
4. How do memory, CPU and timeout work in Lambda, and what are the key limits?
Answer: You configure memory from 128 MB to 10,240 MB in 1 MB steps, and Lambda allocates CPU in proportion to memory; at 1,769 MB a function gets the equivalent of one vCPU. That is why raising memory often makes a CPU-bound function faster and sometimes cheaper, because duration falls. The timeout is up to 900 seconds (15 minutes).
| Limit | Value (check the Lambda quotas page for changes) |
|---|---|
| Timeout | 15 minutes (Lambda Managed Instances allow up to 90 minutes for asynchronous and most event-source invocations) |
| Synchronous payload | 6 MB request and 6 MB response; up to 200 MB with response streaming |
| Asynchronous payload | 1 MB |
| Deployment package | 50 MB zipped upload, 250 MB unzipped including layers; container images up to 10 GB |
| Layers | 5 per function |
| Environment variables | 4 KB total |
/tmp storage | 512 MB to 10,240 MB |
5. What is the difference between a Lambda execution role and a resource-based policy?
Answer: They control opposite directions of access. The execution role is an IAM role the function assumes; it decides what the function can call (read a DynamoDB table, write to S3, invoke a Bedrock model, write CloudWatch logs). The resource-based policy is attached to the function itself and decides who can invoke it (API Gateway, S3, SNS, EventBridge, or another account). When you add an S3 trigger in the console, Lambda adds a resource-policy statement allowing s3.amazonaws.com to invoke it, ideally scoped with SourceArn and SourceAccount conditions.
Event source mappings are the exception that confuses people: for SQS, Kinesis and DynamoDB Streams, Lambda polls the source, so the execution role needs permission to read from the queue or stream, and no resource policy is needed on the function.
6. When would you use Lambda layers versus a container image?
Answer: Layers are zip archives of shared code or dependencies (an internal logging library, a database driver) mounted under /opt. They suit small, shared, slowly changing pieces, but they count toward the 250 MB unzipped limit and create version-coupling problems if many teams depend on one layer. Container images (up to 10 GB) suit large dependencies such as ML libraries, teams already building OCI images, and organisations that want the same scanning and signing pipeline they use for containers.
Default: zip for small functions, a layer only for genuinely shared code, container images when dependencies outgrow zip limits. For container build habits, see the Docker interview questions guide.
7. Explain Lambda versions and aliases.
Answer: A version is an immutable snapshot of code and configuration with its own ARN; $LATEST is the mutable working copy. An alias is a named pointer (for example prod) to a version, and callers invoke the alias ARN, so promoting a release means moving the pointer, not changing every trigger. Aliases support weighted routing between two versions, which is how canary and linear deployments work: AWS CodeDeploy (used by SAM's DeploymentPreference) shifts traffic gradually and rolls back automatically if CloudWatch alarms fire.
Versions also matter for features: provisioned concurrency and SnapStart apply to published versions or aliases, not to $LATEST.
8. How do you manage configuration and secrets in Lambda securely?
Answer: Use environment variables for non-sensitive configuration (table names, feature flags, log level). They are encrypted at rest with KMS, but anyone with GetFunctionConfiguration can read them in plain text, so secrets do not belong there. Store secrets in AWS Secrets Manager (rotation, auditing) or SSM Parameter Store SecureString, grant the execution role read access to only those ARNs, and fetch them during Init so they are cached for the life of the environment. The AWS Parameters and Secrets Lambda Extension, or the Powertools parameters utility, adds a local cache with a TTL so you are not calling Secrets Manager on every request.
9. What are the key SQS concepts: Standard vs FIFO, visibility timeout, long polling and message size?
Answer:
- Standard queues offer very high throughput, at-least-once delivery and no strict ordering, so duplicates and reordering are possible. FIFO queues preserve order within a message group and deduplicate within a five-minute window using a deduplication ID or content hash, at lower throughput (higher with high-throughput mode).
- Visibility timeout hides a received message from other consumers while one consumer works on it. If the consumer does not delete it before the timeout, the message reappears and is processed again.
- Long polling (
ReceiveMessageWaitTimeSecondsup to 20 seconds) waits for messages instead of returning empty responses immediately, cutting empty receives and cost. - Message size is up to 1 MiB. For larger payloads, store the body in S3 and send a pointer (the claim-check pattern); the SQS Extended Client Libraries for Java and Python do this for you.
- Delay queues and message timers postpone delivery by up to 15 minutes. Retention is configurable, with a default of four days and a maximum of 14 days.
10. When do you use SNS, SQS or EventBridge?
Answer: They solve different problems and are often combined.
| Service | Model | Use it for |
|---|---|---|
| SQS | Queue, pull, one consumer group per message | Buffering, load levelling, retries, decoupling a producer from a slower worker |
| SNS | Pub/sub topic, push to many subscribers | Fan-out of the same message to several queues, functions, HTTP endpoints, email or SMS |
| EventBridge | Event bus with content-based rules | Routing domain events by pattern, SaaS and AWS service events, cross-account buses, archive and replay |
A common production pattern is SNS (or EventBridge) to several SQS queues: each consumer gets its own queue, retries and DLQ, so one slow consumer cannot hold back the others.
Lambda cold starts, concurrency and new compute options
11. What is a cold start and what makes it worse?
Answer: A cold start is the extra latency when Lambda must create a new execution environment: download code, start the runtime and run your Init code before the handler executes. It happens on the first request, when traffic scales beyond the current warm environments, after a deployment, and after idle environments are reclaimed.
What makes it worse: large packages and many dependencies, heavyweight frameworks (dependency injection containers, ORMs) initialised eagerly, runtimes with JIT warm-up such as Java and .NET without SnapStart, expensive network calls during Init, and a large container image.
Interview tip: Cold starts usually matter only for synchronous, latency-sensitive paths. For an SQS worker, a cold start is noise. Saying this shows you prioritise.
12. How do you reduce cold start impact, and where does SnapStart fit?
Answer: Work from cheapest to most expensive:
- Shrink Init: trim dependencies, use tree-shaking or bundling for Node.js, lazy-load rarely used modules, and create SDK clients once outside the handler.
- Tune memory: more memory means more CPU, which speeds Init.
- Use Graviton (arm64) where your dependencies support it, usually for better price-performance.
- SnapStart: Lambda runs Init when you publish a version, snapshots the initialised microVM memory and disk, and resumes new environments from the snapshot. It supports Java 11 and later, Python 3.12 and later, and .NET 8 and later. It does not work with provisioned concurrency, EFS, or ephemeral storage above 512 MB, and only applies to published versions. Code that generates unique values or opens connections during Init must be made snapshot-safe using runtime hooks.
- Provisioned concurrency: keeps a set number of environments initialised so they respond without cold starts. It is the strongest option for strict latency targets, and it is billed while configured.
Real-world example: Consider a bank's card-status API in a Hyderabad GCC built with Java and Spring. Turning on SnapStart for the published alias removes most Init latency at no provisioned cost; the team adds provisioned concurrency only for the business-hours peak, scheduled with Application Auto Scaling.
13. How does Lambda scale under load?
Answer: Concurrency is the number of in-flight requests, and on the default compute type each execution environment serves one request at a time, so concurrency equals active environments. For synchronous calls, throughput is roughly concurrency multiplied by requests per second per environment (Lambda caps each environment at 10 requests per second for synchronous invocation). Each function can add up to 1,000 execution environments every 10 seconds, until it reaches the account's concurrency limit for the Region (1,000 by default for established accounts, lower for new accounts until AWS raises it automatically, and increasable to tens of thousands through Service Quotas).
When a function cannot get concurrency, synchronous callers receive a 429 throttling error, asynchronous events are retried from Lambda's internal queue for up to six hours, and event source mappings back off and retry.
14. What is the difference between reserved and provisioned concurrency?
Answer: They solve different problems and can be used together.
| Reserved concurrency | Provisioned concurrency | |
|---|---|---|
| Purpose | Carve out a slice of account concurrency; acts as both a floor and a ceiling | Keep environments pre-initialised to remove cold starts |
| Cost | No extra charge | Billed for the configured amount while enabled, plus usage |
| Typical use | Protect critical functions; cap a function so it cannot overwhelm a database or a downstream quota | Latency-sensitive APIs with predictable peaks |
| Set to zero | Stops all invocations (an emergency off switch) | Not applicable |
You can reserve up to the account's unreserved concurrency minus 100; the remaining 100 units stay available for functions without reservations.
15. Is Lambda stateless? How should you handle state between invocations?
Answer: Treat it as stateless. Warm environments do keep global variables, open connections and /tmp contents between invocations, and you should exploit that for caching SDK clients, configuration and connection pools. But you cannot control which environment receives a request, environments are discarded without notice, and many run in parallel. So anything that must be correct (sessions, counters, workflow progress, uploaded files) belongs in DynamoDB, S3, ElastiCache or a workflow engine such as Step Functions or durable functions.
16. How do you connect Lambda to a database in a VPC safely?
Answer: Attach the function to private subnets in at least two Availability Zones with a security group that the database's security group allows. For outbound internet access (third-party APIs), route through a NAT gateway, or prefer VPC endpoints for AWS services like S3, DynamoDB, Secrets Manager and Bedrock so traffic stays private and avoids NAT data charges.
The main risk with relational databases is connection exhaustion: each concurrent environment opens its own connections, and a traffic spike can open hundreds. Use RDS Proxy, which maintains a connection pool, queues or throttles excess application connections and integrates with Secrets Manager and IAM authentication. Combine it with reserved concurrency to cap how many functions can talk to the database at once. If the access pattern is key-value, DynamoDB avoids the problem entirely.
17. What are Lambda function URLs and response streaming?
Answer: A function URL is a dedicated HTTPS endpoint for one function, secured with IAM auth or left public (NONE), with CORS support. It is simpler and cheaper than API Gateway when you do not need API keys, usage plans, request validation, WAF directly on the API or custom authorizers; teams often put CloudFront in front of it.
Response streaming lets a function send partial output as it is produced, through function URLs, the InvokeWithResponseStream API or an API Gateway REST proxy integration. It improves time to first byte and lifts the response size to 200 MB (the first 6 MB is uncapped, the rest is bandwidth-limited). Native streaming is supported on Node.js managed runtimes; Python and other languages need a custom runtime or the Lambda Web Adapter. One trap: if the client disconnects, the function keeps running and you pay for the full duration, so keep timeouts sensible.
18. What are Lambda Managed Instances and Lambda durable functions?
Answer: Both are recent additions that change classic Lambda answers, so check current documentation before an interview.
- Lambda Managed Instances run your functions on current-generation EC2 instances in your account (including Graviton and network-optimised types) through a capacity provider, while AWS still handles provisioning, patching, routing and scaling. One execution environment can serve multiple invocations concurrently, which suits I/O-heavy code. Pricing is EC2-based plus a management fee, so EC2 Savings Plans and Reserved Instances apply.
- Lambda durable functions let you write multi-step workflows in ordinary code (Node.js, Python, Java and .NET are supported) that can run for up to a year. The runtime checkpoints each step; after an interruption the function replays from the start and skips completed steps. Waits suspend execution without compute charges, which suits human approvals and polling.
Interview tip: Frame durable functions versus Step Functions as code-first versus graph-first orchestration. AWS's own guidance: durable functions when workflow logic is tightly coupled to application code; Step Functions for visual, cross-service orchestration with native integrations.
API Gateway interview questions
19. What is the difference between REST APIs, HTTP APIs and WebSocket APIs in API Gateway?
Answer: REST and HTTP APIs are both request/response products; WebSocket APIs keep a two-way connection open.
| Capability | REST API | HTTP API | WebSocket API |
|---|---|---|---|
| Positioning | Full feature set | Minimal features, lower price | Persistent, bidirectional connections |
| Endpoint types | Edge-optimised, Regional, private | Regional only | Regional |
| Auth | IAM, Cognito, Lambda authorizers, resource policies | IAM, JWT authorizers (including Cognito), Lambda authorizers | IAM, Lambda authorizer on $connect |
| API keys, usage plans, per-client throttling | Yes | No | No |
| Request validation, caching, AWS WAF, X-Ray | Yes | No | No |
| Response streaming | Yes (proxy integrations) | No | Not applicable |
| Integration timeout | 29 seconds default; can be raised for Regional and private APIs | 30 seconds maximum | 29 seconds per message integration |
WebSocket connections last at most two hours and close after ten minutes idle, so clients need reconnect logic. Choose HTTP APIs for straightforward JWT-protected Lambda backends, REST APIs when you need API management, WAF, private endpoints or streaming, and WebSocket APIs (or AppSync Events, a managed pub/sub WebSocket service) for real-time push.
20. How do you secure an API Gateway endpoint?
Answer: Layer the controls:
- Authentication: Cognito user pools or another OIDC provider such as Microsoft Entra ID via JWT authorizers (HTTP API) or Cognito authorizers (REST); IAM SigV4 for service-to-service calls; Lambda authorizers for custom tokens or header logic, with result caching to keep latency down.
- Network: private REST APIs reachable only through interface VPC endpoints, plus resource policies that restrict by VPC endpoint, source IP or account.
- Edge protection: AWS WAF on REST APIs (or on CloudFront in front of HTTP APIs) for rate-based rules and common exploit patterns; mutual TLS where partners present client certificates.
- Abuse control: stage and method throttling, usage plans per client.
- Least privilege behind the API: the integration's role or Lambda's execution role should reach only what that route needs.
Interview tip: State clearly that API keys identify clients for usage plans; they are not an authentication mechanism on their own.
21. What is the difference between Lambda proxy and non-proxy integration?
Answer: With proxy integration (AWS_PROXY), API Gateway passes the whole HTTP request (method, path, headers, query string, body) to Lambda as a structured event, and Lambda must return status code, headers and body in a defined format. It is simple and keeps HTTP logic in code. With non-proxy (custom) integration on REST APIs, you use VTL mapping templates to transform the request before it reaches the backend and the response before it reaches the client. That lets API Gateway call AWS services directly, for example writing straight to SQS or DynamoDB with no Lambda in the path, and it lets you reshape legacy payloads.
Request validation (JSON Schema models and required parameters) on REST APIs rejects malformed requests before they cost a Lambda invocation.
22. Explain edge-optimised, Regional and private endpoints, plus custom domains and CORS.
Answer: An edge-optimised REST API routes clients through the CloudFront network to the API's Region, useful for geographically spread clients. A Regional endpoint serves clients directly and lets you put your own CloudFront distribution in front. A private endpoint is reachable only from VPCs through interface endpoints, for internal APIs.
Custom domains map your hostname to one or more APIs and stages through base path mappings, with certificates from ACM. CORS must be configured when a browser app on another origin calls the API: HTTP APIs have built-in CORS configuration; for REST APIs you configure the OPTIONS preflight and return the CORS headers, and with proxy integrations your Lambda must also add the headers to its responses, which is the step people forget.
23. How do throttling, usage plans and API keys work in API Gateway?
Answer: API Gateway throttles with a token-bucket model: a steady-state rate and a burst capacity. Limits apply at several levels: an account-level Regional quota, stage and method throttling you configure, and, on REST APIs, per-client limits through usage plans that bind API keys to rate, burst and quota (requests per day, week or month). Requests over the limit receive a 429.
24. How do stages, canary releases, caching and timeouts work in API Gateway?
Answer: A stage (dev, test, prod) is a named reference to a deployment, with its own settings, stage variables, logging and throttling. On REST APIs, a canary setting sends a percentage of stage traffic to a new deployment so you can compare metrics before promoting.
Caching (REST only) stores responses per stage with a TTL (default 300 seconds, up to 3,600), keyed on chosen parameters; it cuts latency and backend load for read-heavy endpoints but must be keyed carefully so one user's data is never served to another. Timeouts: REST integrations default to a maximum of 29 seconds; you can raise this for Regional and private REST APIs, possibly at the cost of a lower Region-level throttle quota. HTTP APIs stop at 30 seconds. Anything slower should go asynchronous (accept the job, return 202 with a job ID, deliver the result later) or use response streaming.
Step Functions, EventBridge and DynamoDB
25. What is AWS Step Functions and what are its main state types?
Answer: Step Functions is a managed workflow service: you define a state machine in Amazon States Language (JSON, or visually in Workflow Studio), and the service runs each step, tracks state, handles retries and records history. Main state types:
- Task: do work (invoke Lambda, call an AWS SDK API, run an ECS task, call an HTTPS endpoint).
- Choice: branch on input data.
- Parallel: run fixed branches concurrently.
- Map: iterate over items (inline, or Distributed Map for very large datasets such as millions of S3 objects).
- Wait: pause for a duration or until a timestamp.
- Pass, Succeed, Fail: shape data and end executions.
Data moves between states as JSON. You can shape it with the classic JSONPath fields (InputPath, Parameters, ResultSelector, ResultPath, OutputPath) or with JSONata expressions, which Step Functions now supports and which are simpler to read for new workflows.
26. What is the difference between Standard and Express workflows?
Answer:
| Standard | Express | |
|---|---|---|
| Maximum duration | Up to one year | Five minutes |
| Execution semantics | Exactly-once (unless you add retries) | Asynchronous: at-least-once; synchronous: at-most-once |
| Pricing | Per state transition | Per execution, duration and memory |
| History | Full history via API and console for 90 days after completion | CloudWatch Logs when logging is enabled |
| Patterns | All, including .sync job runs, callbacks with task tokens, Distributed Map, activities | No .sync, no .waitForTaskToken, no Distributed Map, no activities |
Use Standard for long-running, auditable, non-idempotent business processes (payments, onboarding, approvals). Use Express for high-volume, short, idempotent event processing (IoT ingestion, stream transformation, synchronous microservice orchestration behind an API). The type cannot be changed after creation, and a common pattern nests Express workflows inside a Standard parent.
27. How do you handle errors, retries and timeouts in Step Functions?
Answer: Each Task can declare Retry rules matched on error names (for example Lambda.TooManyRequestsException, States.Timeout or your own error types) with IntervalSeconds, BackoffRate, MaxAttempts, an optional MaxDelaySeconds and jitter. After retries are exhausted, Catch routes the error to a fallback state: a compensation step, a notification, or a Fail state with a clear cause. TimeoutSeconds bounds how long a task can run, and HeartbeatSeconds detects a worker or callback that has silently died.
Two habits interviewers listen for: retry only transient errors (throttling, timeouts, 5xx), never validation errors; and design compensation for partial failure (the saga pattern), because a workflow that charged a card and then failed to reserve stock must refund or retry, not just stop.
28. What service integration patterns does Step Functions support?
Answer: Three patterns:
- Request-response: call the service and move on as soon as it responds (for example, put an item in DynamoDB or publish to SNS).
- Run a job (
.sync): start a job (ECS task, Glue job, nested state machine, Batch job) and wait until it completes. - Wait for callback (
.waitForTaskToken): pass a task token to an external system (an SQS message, a human approval email, a partner webhook) and pause until something callsSendTaskSuccessorSendTaskFailurewith that token.
Through AWS SDK integrations, Step Functions can call a very wide range of AWS APIs directly, so many "glue" Lambda functions are unnecessary.
29. What are EventBridge event buses, rules, Scheduler and Pipes?
Answer: An event bus receives events (from AWS services, your applications via PutEvents, or SaaS partners). Rules match events by content patterns and route them to targets such as Lambda, SQS, Step Functions or another bus, with input transformation, retry policies and DLQs per target. Archives and replay let you re-run past events after fixing a bug.
- EventBridge Scheduler is a separate, highly scalable scheduler for one-time and recurring schedules (cron, rate, time zones, flexible windows) that can invoke a large number of AWS service APIs directly. It is the modern replacement for scheduled rules, and handles millions of schedules, for example one reminder per customer.
- EventBridge Pipes connects one source (SQS, Kinesis, DynamoDB Streams, Kafka, MQ) to one target with optional filtering, enrichment (Lambda, Step Functions, API destination) and transformation. It replaces small "poll, filter, forward" Lambda functions in point-to-point integrations.
30. Why is DynamoDB the default database for serverless, and how do on-demand mode and Streams work?
Answer: DynamoDB is HTTP-based (no connection pools to exhaust), scales horizontally, has single-digit-millisecond reads at the key level and charges by usage in on-demand mode, which suits Lambda's bursty concurrency. On-demand capacity bills per read and write request and adapts to traffic without capacity planning; provisioned capacity with auto scaling is usually cheaper for steady, predictable traffic.
DynamoDB Streams capture a time-ordered log of item-level changes (keys only, new image, old image, or both), retained for 24 hours. A Lambda event source mapping on the stream enables change data capture: update a search index, publish domain events, maintain aggregates or replicate to another store. Records are ordered per item (per partition key), and the consumer must keep up or it loses data after 24 hours.
31. Orchestration or choreography: when do you use Step Functions, durable functions or plain events?
Answer: Choreography means services react to events on a bus with no central controller. It keeps teams decoupled and suits "notify everyone who cares" flows (order placed: email, analytics, loyalty points). Its weakness is visibility: nobody owns the end-to-end flow, and debugging a stalled process means tracing events across services.
Orchestration puts the sequence, retries, timeouts and compensation in one place. Choose Step Functions when the flow spans many AWS services, benefits from a visual graph and execution history, or needs callbacks and Distributed Map. Choose durable functions when the workflow is tightly bound to application code and the team prefers writing it in their language with normal testing tools.
Most real systems mix both: orchestrate inside a bounded context (order fulfilment), choreograph between contexts (fulfilment emits OrderShipped and other domains react). This is also how we explain long-running AI agents in durable, long-running AI agent workflows.
Idempotency, retries and dead-letter queues
32. What is idempotency and how do you implement it in a serverless system?
Answer: An operation is idempotent if running it more than once has the same effect as running it once. It is mandatory in serverless because almost every path can deliver duplicates: SQS Standard and Lambda event source mappings are at-least-once, asynchronous invocations can run more than once, clients retry on timeouts, and Express workflows are at-least-once.
Implementation:
- Pick an idempotency key that identifies the business operation: a client-supplied
Idempotency-Keyheader, an order ID plus action, or a hash of the relevant payload fields (not the whole event, which may contain timestamps). - Record it atomically before side effects with a DynamoDB conditional write (
attribute_not_exists), statusIN_PROGRESS, and a TTL. - On duplicate, return the stored result if completed, or reject or retry later if still in progress.
- Make downstream calls idempotent too: pass the same key to payment providers that support it, and use conditional updates rather than blind increments.
The Powertools for AWS Lambda idempotency utility implements this with a DynamoDB persistence layer, payload-subset keys, expiry windows and handling for Lambda timeouts.
33. How do retries, destinations and DLQs work for asynchronous Lambda invocations?
Answer: For asynchronous calls, Lambda queues the event internally. If the function returns an error, Lambda retries twice by default (waiting about one minute, then two minutes). For throttling and service errors it keeps retrying with exponential backoff for up to six hours by default. You can lower MaximumRetryAttempts (0 to 2) and MaximumEventAgeInSeconds.
When attempts are exhausted, the event goes to an on-failure destination (SQS, SNS, S3, another Lambda function or an EventBridge bus) or to a legacy dead-letter queue (SQS or SNS). Prefer destinations: the record includes the request, the error and the response context, and you can also configure an on-success destination to chain processing.
34. How does error handling work when Lambda processes SQS messages?
Answer: Lambda polls the queue, invokes the function with a batch, and deletes the batch if the function succeeds. If the function throws, the whole batch becomes visible again after the visibility timeout and is retried, including messages that had already succeeded. Key settings:
- Partial batch response: enable
ReportBatchItemFailuresand return the IDs of failed messages inbatchItemFailures, so only those are retried. The Powertools batch processor handles this logic. - Visibility timeout: set the queue's visibility timeout to at least six times the function timeout, so throttled or retried batches are not redelivered while still in progress.
- DLQ on the queue (a redrive policy with
maxReceiveCount), not a Lambda async DLQ, because this is an event source mapping. SQS also supports redrive from the DLQ back to the source. - Scaling control: maximum concurrency on the event source mapping limits how many concurrent invocations one queue drives, or provisioned mode reserves pollers for spiky, latency-sensitive queues (they are mutually exclusive).
- FIFO queues: ordering holds per message group, and a failing message blocks its group until it succeeds or moves to the DLQ.
35. How do you handle poison records in Kinesis or DynamoDB Streams consumers?
Answer: Stream sources are ordered per shard, and by default Lambda retries a failed batch until it succeeds or the records expire, so one bad record can block the shard; this shows up as rising IteratorAge. Mitigations on the event source mapping: set MaximumRetryAttempts and MaximumRecordAgeInSeconds, enable BisectBatchOnFunctionError to split failing batches and isolate the bad record, return partial batch responses, and configure an on-failure destination (SQS, SNS or S3) that receives metadata about the skipped records so you can inspect and replay them.
Observability, security, cost and tooling
36. How do you build observability for a serverless application?
Answer: Cover the three signals with correlation across services:
- Logs: structured JSON logs in CloudWatch Logs with a correlation ID, function request ID and business keys (order ID), with log levels and retention set explicitly.
- Metrics: built-in Lambda metrics (
Invocations,Errors,Throttles,Duration,ConcurrentExecutions,IteratorAge, async event age and dropped events), SQSApproximateAgeOfOldestMessage, API Gateway latency and 4xx/5xx, Step Functions failed executions, plus business metrics published cheaply with CloudWatch Embedded Metric Format. - Traces: AWS X-Ray active tracing across API Gateway, Lambda and Step Functions to see where latency sits. AWS has put the X-Ray SDKs and daemon into maintenance mode (from February 2026) and points new instrumentation to OpenTelemetry, for example the AWS Distro for OpenTelemetry, so mention OTel in new designs.
Powertools for AWS Lambda (Python, TypeScript, Java, .NET) packages a structured logger, tracer, metrics, idempotency, batch processing and parameters, and is the usual answer for "how do you standardise this across fifty functions?" For the AI side, see AI observability, and for metrics stacks outside CloudWatch, the Prometheus and Grafana interview questions.
37. How do you apply least privilege in a serverless application?
Answer: One execution role per function, scoped to the exact actions and resource ARNs it needs: dynamodb:GetItem on one table, not dynamodb:* on *. In SAM, policy templates such as DynamoDBReadPolicy keep this concise; in CDK, grant methods (table.grantReadData(fn)) generate scoped policies. Further controls:
- Resource policies on functions, queues, topics, buckets and buses with
SourceArnandSourceAccountconditions to prevent confused-deputy access. - KMS key policies for encrypted queues, topics and tables, so a role needs both service and key permission.
- Permission boundaries and SCPs so developers can create roles without escalating privilege.
- IAM Access Analyzer to find externally shared resources and generate policies from CloudTrail activity.
- Code signing for Lambda and dependency scanning in CI, since a function's supply chain is part of its attack surface.
38. How is serverless priced, and where do the surprise costs come from?
Answer: Each service has its own meter. Lambda charges per request plus duration in GB-seconds (memory multiplied by billed time, at millisecond granularity), with Arm usually cheaper than x86 per GB-second; provisioned concurrency adds a charge for the configured amount. API Gateway charges per request (HTTP APIs cost less than REST), plus caching and data transfer. Step Functions Standard charges per state transition; Express charges per execution, duration and memory. DynamoDB on-demand charges per request plus storage. SQS, SNS and EventBridge charge per request or event.
Surprise costs interviewers want you to name: CloudWatch Logs ingestion from verbose logging, NAT gateway data processing for Lambda in VPCs, chatty Standard workflows with many tiny states, recursive invocation loops, provisioned concurrency left on at night, and over-sized memory on I/O-bound functions. Serverless is cheap at low and spiky volume; at steady high volume a container or instance platform, or Lambda Managed Instances, can cost less. For deeper FinOps practice, see the FinOps interview questions.
39. SAM, CDK or Terraform: how do you deploy serverless applications?
Answer: AWS SAM is an open-source framework whose templates extend CloudFormation with shorthand serverless resources (AWS::Serverless::Function, Api, StateMachine), and the SAM CLI builds, tests locally (sam local invoke, sam local start-api), syncs changes to the cloud quickly during development and deploys with gradual traffic shifting. AWS CDK defines infrastructure in TypeScript, Python, Java, C# or Go and synthesises CloudFormation; its constructs and grant methods suit larger systems and platform teams that want reusable patterns. Terraform suits organisations standardising on one tool across clouds and non-AWS resources.
The choice matters less than discipline: everything in code, one stack per service, and CI/CD that runs integration tests against real services before promoting. See the Terraform interview questions for the IaC side.
40. When should you NOT use serverless, and when is Fargate the better fit?
Answer: Serverless is a poor default when:
- Work runs longer than 15 minutes per unit and cannot be split or checkpointed.
- Traffic is steady and high, so per-invocation pricing exceeds always-on compute.
- You need GPUs, special hardware, or very large memory or local disk.
- Latency must be consistently in single-digit milliseconds with no tolerance for cold starts.
- The workload needs long-lived connections or in-memory state (game servers, some trading systems).
Fargate (on ECS or EKS) runs containers without managing EC2 hosts and has no 15-minute limit, so it fits long-running workers, services that already exist as containers, and steady APIs. Many systems use both: Lambda for event glue and spiky endpoints, Fargate for long batch jobs and steady services. For container orchestration questions, see the Kubernetes interview questions.
If you want to practise these designs hands-on, with Lambda, API Gateway, Step Functions and DynamoDB built and broken in labs, Cloudsoft's AWS training in Hyderabad runs in Ameerpet and live online. Call +91 96660 19191 for a free demo.
AI in serverless: Bedrock from Lambda
41. How do you call Amazon Bedrock from a Lambda function, and what limits matter?
Answer: Use the AWS SDK's bedrock-runtime client, preferably the Converse API (Converse, or ConverseStream for streaming), which gives one request shape across Bedrock models that support messages; InvokeModel and InvokeModelWithResponseStream remain for model-specific payloads. Grant the execution role bedrock:InvokeModel (and the streaming action) on only the model or inference-profile ARNs you use, and reach Bedrock through a VPC endpoint if the function is in a VPC.
The constraints are mostly about time and quotas. Model calls can take many seconds, so set the function timeout and the SDK read timeout deliberately; API Gateway's default 29-second integration timeout is often the first thing to break. Bedrock enforces per-model request and token quotas, so handle throttling with SDK retries plus backoff and jitter, and cap Lambda concurrency so a traffic spike cannot burn the whole quota. The AWS Bedrock interview questions guide goes deeper on models, knowledge bases and guardrails.
42. How do you stream LLM responses to users in a serverless architecture?
Answer: Streaming lowers perceived latency because users see the first tokens quickly. Options:
- Lambda response streaming through a function URL (often behind CloudFront), relaying
ConverseStreamchunks as server-sent events. Native on Node.js managed runtimes; Python needs the Lambda Web Adapter or a custom runtime. - API Gateway REST API with response transfer mode set to
STREAMon a Lambda proxy integration. It streams for up to 15 minutes, gets past the 29-second and 10 MB buffered limits, and is subject to idle timeouts (five minutes for Regional and private endpoints, 30 seconds for edge-optimised). Caching and VTL response transformation are not available in streaming mode. - WebSocket APIs or AppSync when the conversation is two-way or results arrive from a background worker: the worker posts tokens or progress to the open connection.
Browser
| HTTPS (SSE)
v
CloudFront / API Gateway (STREAM)
|
v
Lambda --ConverseStream--> Amazon Bedrock
| relays chunks as they arrive
v
Partial tokens back to the browser
Production consideration: a streaming function keeps running and billing if the user closes the tab, so enforce a maximum output length and a sensible timeout. For more on latency, see LLM latency optimisation.
43. Which asynchronous patterns suit AI workloads on serverless?
Answer: Anything slower than a comfortable HTTP request should become a job:
- Queue-and-worker: API Gateway writes the request to SQS and returns 202 with a job ID; a Lambda worker with capped concurrency calls Bedrock, stores the result in DynamoDB or S3, and the client polls or receives a push over WebSocket. The queue absorbs bursts and shields model quotas.
- Step Functions: for document pipelines (extract, chunk, embed, index), with retries per step, Map for parallelism and a task-token callback for human review. Step Functions can call Bedrock through its service integration without a Lambda in between.
- Bedrock batch inference: for large offline jobs (classify a backlog of tickets overnight), submit prompts in bulk and read results from S3 instead of invoking per item.
- Durable functions: for agent-style loops that chain model calls, wait for tools or approvals, and must survive interruptions without paying for idle waiting.
The design principles for multi-hour and multi-day agents are covered in durable agents and long-running workflows.
44. How do you control cost and reliability for an AI feature built on Lambda and Bedrock?
Answer: Treat the model as an expensive, rate-limited dependency:
- Cap concurrency with reserved concurrency or SQS maximum concurrency so the function cannot exceed model quotas or budget.
- Cache repeated answers (DynamoDB or ElastiCache keyed on normalised prompt and context version) and use Bedrock prompt caching where the model supports it.
- Route simple tasks to smaller, cheaper models and reserve larger models for hard cases.
- Set maximum output tokens, log token usage per request as a metric, and alarm on cost anomalies with AWS Budgets.
- Make retries idempotent so a timeout does not trigger duplicate expensive calls or duplicate actions taken by an agent.
For a fuller cost view see cloud cost optimisation for AI; for central quota and routing control many teams add an LLM gateway.
Real-world scenario questions
45. Scenario: during a sale, your order API starts returning 429 errors from Lambda. What do you do?
Answer: A 429 from Lambda means it could not get concurrency: the function hit its reserved concurrency, the account hit its Regional limit, or traffic grew faster than Lambda's scaling rate. API Gateway throttling also returns 429, so first confirm which layer is throttling.
What I would check:
- Lambda
Throttlesversus API Gateway 429 counts, to locate the throttle. ConcurrentExecutionsfor the function and account against reserved concurrency and the account quota.- Whether another function (a batch job or an SQS consumer) consumed the shared unreserved pool at the same time.
Duration: if a slow downstream (database, payment provider) doubled duration, the same request rate needs double the concurrency.
Production consideration: Short term, raise or rebalance reserved concurrency and request a quota increase. Longer term, reserve concurrency for the order path, cap background consumers, add provisioned concurrency for known peaks, fix the slow dependency, and move non-critical work (emails, analytics) off the synchronous path into SQS.
46. Scenario: customers report being charged twice; the payment worker reads from SQS. How do you investigate and fix it?
Answer: Duplicate processing is expected behaviour with SQS Standard and Lambda event source mappings, so the fix is idempotency, but first find which duplicate source is firing.
What I would check:
- Function duration versus the queue's visibility timeout: if processing sometimes exceeds the timeout, the message reappears and a second worker processes it.
- Whether the handler throws after the charge succeeded (for example, a failure writing the receipt), causing the whole batch, including the successful charge, to be retried.
- Whether partial batch responses are enabled, or one bad message is forcing retries of good ones.
- Logs grouped by order ID and message ID to see whether duplicates share a message ID (redelivery) or not (producer duplicates).
Production consideration: Add an idempotency key per payment (order ID plus attempt intent), recorded with a DynamoDB conditional write before calling the provider, and pass the same key to the provider's idempotency feature. Set visibility timeout to at least six times the function timeout, enable ReportBatchItemFailures, and configure a DLQ with alarms. FIFO with deduplication IDs reduces producer duplicates but does not replace consumer idempotency.
47. Scenario: p99 latency on an API Gateway and Lambda API spikes every morning and after deployments. How do you diagnose it?
Answer: Morning and post-deployment spikes point to cold starts as concurrency ramps from a low overnight base, or after new versions replace warm environments, but verify before tuning.
What I would check:
- API Gateway
LatencyversusIntegrationLatency: a large gap means time in the gateway (authorizer, mapping); a small gap means the backend. - Lambda
Init Durationin REPORT log lines and traces: how often and how long cold starts are. - Lambda authorizer latency and whether its result caching is enabled.
- Downstream spans in traces: database connection setup, Secrets Manager calls on each request, cold caches.
- Memory size: a CPU-starved function has slow Init and slow handlers.
Production consideration: Fix Init first (trim dependencies, initialise clients once, cache secrets), raise memory if CPU-bound, enable SnapStart on supported runtimes, and schedule provisioned concurrency for the morning ramp. Use gradual deployments so new versions warm up under a slice of traffic.
48. Scenario: design an order-processing workflow with Step Functions.
Answer: Use a Standard workflow, because payment is non-idempotent, the process can wait on external events, and auditors want execution history.
API GW -> Lambda: validate, save PENDING order
|
v StartExecution (name = orderId)
[Reserve stock] --fail--> [Fail: out of stock]
|
[Charge payment] --fail--> [Release stock] -> Fail
|
[Create shipment: task token, wait for warehouse]
| timeout -> [Refund] -> [Release]
v
[Update order SHIPPED] -> EventBridge OrderShipped
|
email / analytics / loyalty consumers
Design points to say out loud:
- Idempotent start: use the order ID as the execution name, so a retried request cannot start two workflows for the same order.
- Direct integrations: DynamoDB updates and the EventBridge publish use SDK integrations, with no Lambda needed.
- Retries with backoff on transient errors for each task; Catch blocks route to compensating steps (release stock, refund), which is the saga pattern.
- Callback with
.waitForTaskTokenand a heartbeat timeout for the warehouse system. - Payment idempotency key passed to the provider, so a retried charge step cannot double-charge.
- Choreography after completion: downstream domains react to
OrderShippedon EventBridge rather than being steps in the workflow.
Production consideration: Alarm on ExecutionsFailed and ExecutionsTimedOut, log execution history to CloudWatch, version state machines with aliases for safe rollout, and keep payloads small by passing IDs and storing documents in S3 or DynamoDB.
49. Scenario: a nightly Lambda job that processes a large file now times out at 15 minutes. What are your options?
Answer: Do not just set the timeout to the maximum and hope. The job has outgrown a single invocation, so split it, checkpoint it, or move it.
What I would check:
- Where the time goes: CPU-bound processing (raise memory), slow downstream writes (batch them), or pure volume.
- Whether the work is divisible by record, row range, S3 object or customer.
- Whether it needs ordering or can run in parallel.
- How it currently handles a crash halfway: does it restart from zero?
Production consideration: Options in rough order: (a) fan out, with a coordinator splitting the file into chunks onto SQS for parallel workers, or Step Functions Distributed Map over S3 objects or lines; (b) checkpoint progress in DynamoDB so each invocation continues where the last stopped, or use durable functions, which checkpoint steps for you; (c) run it as an ECS task on Fargate (or AWS Batch) started by EventBridge Scheduler or a Step Functions .sync step, with no 15-minute ceiling; (d) for asynchronous or event-source workloads, Lambda Managed Instances allow longer timeouts. Whatever you choose, make chunks idempotent so a retried chunk does not duplicate output.
50. Scenario: an internal GenAI assistant behind API Gateway fails with 504 errors on long answers. How do you fix it?
Answer: A 504 means the integration exceeded API Gateway's integration timeout (29 seconds by default on REST, 30 seconds on HTTP APIs) while the model was still generating.
What I would check:
- Lambda duration distribution versus the gateway timeout.
- Time split between retrieval (vector search), the model call and post-processing in traces.
- Prompt size and maximum output tokens: oversized context and unlimited outputs drive latency.
- Bedrock throttling retries hidden inside the SDK, which add delay.
Production consideration: For chat, switch to streaming (REST API STREAM mode or a function URL with response streaming) so tokens flow immediately. For long, non-interactive tasks (summarise a 200-page policy), move to an asynchronous job with SQS or Step Functions and notify on completion. Then reduce latency at the source: smaller retrieved context, a smaller model where quality allows, output limits and caching.
51. Scenario: the Lambda bill jumped sharply overnight with no traffic increase. What happened?
Answer: Classic causes are a recursive loop, a retry storm or runaway logging, so look at invocation counts per function before anything else.
What I would check:
- Which function's
InvocationsandDurationrose, using Cost Explorer by resource and CloudWatch metrics. - Recursion: a function triggered by an S3 prefix that writes back to the same prefix, or an SQS consumer that re-enqueues to its own queue. Lambda's recursive loop detection stops loops between functions, SQS, S3, SNS and EventBridge custom buses (when a supported SDK is used) and notifies you, but it cannot see loops that pass through other services such as DynamoDB.
- Retry storms: a broken downstream causing every message to fail and be redelivered until
maxReceiveCount. - CloudWatch Logs ingestion from debug logging left on, and provisioned concurrency left enabled.
Production consideration: Stop the bleeding by setting reserved concurrency to zero on the looping function. Then fix the trigger (separate input and output prefixes or buckets), keep recursive loop detection on, add DLQs with sensible receive counts, and set AWS Budgets alerts plus anomaly detection so the next spike pages someone within hours.
52. Scenario: a DynamoDB Streams consumer's iterator age keeps growing and the search index is hours behind. What do you do?
Answer: Rising IteratorAge means the consumer is falling behind; because stream records expire after 24 hours, this becomes data loss if ignored.
What I would check:
- Function
Errors: a poison record retried endlessly blocks its shard. Durationand downstream latency: the search cluster may be slow or throttling bulk writes.- Batch size, batching window and parallelisation factor on the event source mapping.
Production consideration: Configure maximum retry attempts, maximum record age, bisect on error and an on-failure destination so a bad record cannot block the shard. Increase parallelisation factor and batch size, and use bulk indexing. Make index updates idempotent (upsert by key with a version check) so replays are safe. Alarm on iterator age well before the 24-hour window.
53. Scenario: during peak traffic, your Lambda functions cause "too many connections" errors on Aurora PostgreSQL. How do you fix it?
Answer: Each concurrent Lambda environment opens its own database connection, so a concurrency spike multiplies connections faster than the database can accept them.
What I would check:
ConcurrentExecutionsagainst the database's connection count and limit.- Whether connections are created inside the handler (one per invocation) instead of reused per environment.
- Whether connections leak on error paths.
Production consideration: Put RDS Proxy in front of Aurora to pool and multiplex connections, open the connection once per environment, cap the function with reserved concurrency sized to what the database can handle, and push write bursts through SQS so the database sees a steady rate. For read-heavy, key-based access, consider DynamoDB or a cache for that path.
54. Scenario: a Friday release broke a critical Lambda function in production. How should deployments work so this is recoverable in minutes?
Answer: Rollback should be a pointer move, not a rebuild. Callers invoke an alias, new code ships as a new version, and traffic shifts gradually.
What I would check:
- Whether triggers point at the alias or at
$LATEST. - Whether the release included an incompatible change in shared state (a DynamoDB attribute renamed, an event schema changed) that a rollback alone will not undo.
Production consideration: Use SAM deployment preferences (CodeDeploy canary or linear traffic shifting) with CloudWatch alarms on errors and latency for automatic rollback, and pre-traffic hooks that run smoke tests. Make schema and event changes backward compatible (expand, migrate, then contract), version Step Functions with aliases, and keep a feature flag for risky logic. The same release discipline applies across the DevOps engineer interview questions.
55. Scenario: a team wants to move a steady, high-traffic REST service from ECS to Lambda "to save money". How do you advise them?
Answer: Model it before agreeing. Serverless saves money when traffic is spiky or low and idle capacity dominates the bill; a service that runs hot around the clock may cost more on per-request Lambda and API Gateway pricing than on well-utilised containers.
What I would check:
- Request rate profile across the day and week, average duration and memory: compute Lambda plus API Gateway cost against current Fargate or EC2 cost, including Savings Plans.
- Latency requirements and the service's start-up time (cold start impact).
- Long-lived connections, in-memory caches or background threads the code relies on.
Production consideration: Often the right answer is hybrid: keep the steady core on containers (or evaluate Lambda Managed Instances for steady Lambda-shaped workloads), move spiky, event-driven edges (file processing, webhooks, scheduled jobs) to Lambda, and decide with numbers from a pilot rather than a slogan.
Key takeaways
- Know Lambda's execution environment lifecycle; it explains cold starts, concurrency, state reuse and most debugging answers.
- Always answer failure questions per invocation model: synchronous, asynchronous or event source mapping.
- Design for at-least-once delivery: idempotency keys, conditional writes, partial batch responses and monitored DLQs or destinations.
- Use reserved concurrency as a safety valve for downstream systems and provisioned concurrency or SnapStart for latency.
- Pick API Gateway REST, HTTP or WebSocket by features needed, and know the integration timeouts.
- Use Step Functions Standard for auditable, non-idempotent processes and Express for high-volume, short, idempotent ones; durable functions are the code-first alternative.
- For AI workloads, stream for interactive chat, go asynchronous for long jobs, and cap concurrency to respect model quotas and budgets.
Interview preparation checklist
- Build an API Gateway (HTTP API), Lambda and DynamoDB CRUD service with SAM or CDK, with one execution role per function.
- Add an SQS worker with partial batch responses, a DLQ and a redrive, then deliberately send a poison message.
- Implement idempotency with the Powertools utility and prove a duplicate message has no second effect.
- Measure cold starts at three memory sizes, then compare with SnapStart or provisioned concurrency.
- Write a Standard Step Functions workflow with Retry, Catch, a compensation step and a task-token callback.
- Instrument a function with structured logs, metrics and traces, and build a dashboard with throttles, errors, duration and queue age.
- Stream a Bedrock response through a function URL or a REST API in STREAM mode.
- Estimate the monthly cost of one of your projects at two traffic levels and explain where containers become cheaper.
FAQ
Which AWS services should I study for a serverless interview?
Focus on Lambda, API Gateway, SQS, SNS, EventBridge, Step Functions and DynamoDB first, then add S3 events, Cognito, AppSync, Fargate and CloudWatch with X-Ray or OpenTelemetry.
Are AWS Lambda interview questions mostly theory or hands-on?
Mid-level and senior interviews lean heavily on scenarios: throttling, duplicates, timeouts and latency spikes. Freshers get more definitions, but even they are often asked to sketch an architecture or explain a retry flow.
How should I prepare for a serverless architecture interview with no production experience?
Build two or three small but complete projects with infrastructure as code, break them on purpose (poison messages, throttling, timeouts), and fix them. Being able to describe what you observed in metrics and logs is far more convincing than memorised answers.
Do I need to know SAM, CDK or Terraform for serverless roles?
You should be comfortable with at least one. SAM is the quickest to learn for Lambda-centric work, CDK is common in AWS-heavy product teams, and Terraform is common in organisations that manage several clouds. Knowing why infrastructure as code matters is more important than which tool you pick.
Which programming language is most useful for AWS Lambda work?
Python and Node.js (TypeScript) are the most common for serverless and event-driven work, and Java and .NET are frequent in enterprises. Pick one, learn its cold start behaviour and Powertools library, and be able to write a clean handler from memory.
Is serverless knowledge useful for DevOps and cloud engineer roles?
Yes. DevOps and cloud engineers build the pipelines, IAM boundaries, monitoring and cost controls for serverless applications, and many internal automation tasks run on Lambda and EventBridge. It complements container and Kubernetes skills rather than replacing them.
How do Step Functions interview questions differ from Lambda questions?
Step Functions questions focus on workflow design: Standard versus Express, Retry and Catch, callbacks with task tokens, compensation and data passing between states. Expect to be asked to design a workflow such as order processing on a whiteboard.
Do AWS certifications help in serverless interviews?
Certifications such as AWS Certified Developer or Solutions Architect can help you get shortlisted and structure your study, but interviews still test hands-on reasoning. Pair any certification with projects you can explain in depth.
How is AI changing serverless interviews?
More interviews now include a question on calling a foundation model from Lambda: timeouts, streaming responses, asynchronous job patterns, quota handling and cost control. The core serverless skills stay the same; AI adds a slow, rate-limited dependency to design around.
If you want to turn these answers into working systems, Cloudsoft's AWS course in Hyderabad covers Lambda, API Gateway, Step Functions, DynamoDB and event-driven design through hands-on labs, in our Ameerpet classroom or live online. For a broader path that combines cloud with AI, ML and security, explore the APEX AI, ML, Cloud and Cyber Security program. Call +91 96660 19191 to book a free demo.

