New batches starting this week ยท Limited seats

Python Interview Questions for Cloud and DevOps Engineers 2026 (60 Questions)

60 Python interview questions for cloud and DevOps engineers with tested code, from core Python and safe subprocess calls to boto3 paginators, uv packaging, concurrency, Lambda, moto testing and twelve real automation scenarios.

Python for cloud and DevOps interview questions and answers 2026 - Cloud Soft Solutions
Last updated ยท 44 min read ยท 9,754 words

These Python interview questions for DevOps, cloud and platform roles focus on what interviewers in 2026 actually probe: whether you can write a small, safe, testable automation script that talks to AWS, Azure or Kubernetes without leaking credentials, hanging forever or deleting the wrong thing. Nobody is hired for reciting what a decorator is; you are hired for the script that cleans up a thousand volumes in dry-run mode first, retries throttled API calls with backoff and still works when cron runs it at 2 a.m. This guide gives 60 high-value questions with tested code, from core Python to boto3, concurrency, packaging with uv and twelve real coding and operations scenarios.

How to use this guide:

  • Freshers and support engineers are usually tested on core Python: data types, mutability, comprehensions, generators, exceptions, files and JSON. Type every snippet yourself and break it on purpose.
  • Mid-level DevOps and cloud engineers get a Python scripting interview round on subprocess, logging, CLIs, HTTP retries and boto3 interview questions such as paginators, waiters and the credentials chain.
  • Senior and SRE-leaning roles add concurrency and the GIL, packaging, testing with moto, typing, secrets and Lambda design, plus live Python automation interview questions where you write code against a realistic task.
  • This page is about Python for infrastructure and operations automation. For Python in LLM applications (FastAPI, Pydantic, embeddings, streaming), use the companion Python for AI interview questions guide instead of repeating it here.

Contents

Core Python fundamentals

1. Why is Python so widely used for cloud and DevOps automation, and which version should you target in 2026?

Answer: Python reads like pseudo-code, ships with a strong standard library (subprocess, pathlib, json, logging, argparse, concurrent.futures), and every major cloud has a first-class Python SDK: boto3 for AWS, the Azure SDK for Python, the Google Cloud client libraries and the official Kubernetes client. Ansible, many CLIs and AWS Lambda all speak Python natively, so one language covers glue scripts, serverless functions and internal tools.

On versions: at the time of writing, Python 3.14 is the current stable line, 3.15 is scheduled for October 2026, and 3.11 to 3.13 receive security fixes only. Python 3.9 reached end of life in October 2025 and 3.10 in October 2026, and boto3 ended support for 3.9 in April 2026. Target a supported version that your runtime offers; Lambda currently provides python3.12, python3.13 and python3.14 on Amazon Linux 2023. Check the Python developer guide's version page before you pin anything.

Interview tip: Saying "3.9 is end of life, so our scripts moved to 3.12 or later and we pinned that in pyproject.toml" sounds like someone who maintains real automation.

2. Explain list, tuple, set and dict. Which do you reach for in automation scripts?

Answer: A list is an ordered, mutable sequence (instance IDs to process). A tuple is ordered and immutable, so it can be a dict key or set member, for example (region, account_id). A set holds unique hashable items with fast membership tests, ideal for "which volumes are not in the allow-list?" using set difference. A dict maps keys to values with fast lookup and preserves insertion order; boto3 responses are nested dicts and lists.

Real-world example: To find security groups not attached to any network interface, build all_sgs and used_sgs as sets and compute all_sgs - used_sgs. Doing that with nested list loops is slower and harder to read.

3. What is the difference between mutable and immutable objects, and what is the mutable default argument bug?

Answer: Immutable objects (int, str, tuple, frozenset) cannot change in place; operations create new objects. Mutable objects (list, dict, set) can. Default argument values are evaluated once, when the function is defined, so a mutable default is shared across calls.

def add_tag(tag, tags=[]):      # bug: shared list
    tags.append(tag)
    return tags

def add_tag_fixed(tag, tags=None):
    tags = [] if tags is None else list(tags)
    tags.append(tag)
    return tags

Calling the buggy version twice returns ["a", "b"] on the second call. In a long-running worker or a warm Lambda container, that leak survives between invocations, which is exactly where it hurts.

4. Shallow copy versus deep copy: when does it matter?

Answer: A shallow copy (list(x), dict.copy(), copy.copy()) creates a new outer container but shares the inner objects. A deep copy (copy.deepcopy()) recursively copies everything. It matters with nested configuration: if you shallow-copy a base config dict and change cfg["tags"]["env"] = "prod", you also changed the base config, because tags is the same inner dict.

Interview tip: Mention that deep copies are slower and can fail on objects such as open sockets or SDK clients, so copy plain data (dicts loaded from YAML), not clients.

5. How does Python manage memory?

Answer: CPython uses reference counting as the primary mechanism: when an object's reference count drops to zero, it is freed immediately. A cyclic garbage collector (the gc module) periodically finds reference cycles that counting cannot free. Small objects come from CPython's own allocator pools, so freed memory is often reused by the process rather than returned to the operating system, which is why RSS may not shrink after a big list is deleted.

For automation the practical lessons are: stream large data instead of loading it (generators, line-by-line reads, paginators), avoid unbounded caches and global lists in long-running workers, and use tracemalloc to find where memory grows. Q60 is a scenario on exactly this.

6. When do you use a list comprehension versus a generator expression?

Answer: A list comprehension [x for x in items if cond] builds the whole list in memory; use it when the result is small or you need to index or iterate it more than once. A generator expression (x for x in items if cond) produces items lazily; use it for large or streaming inputs and for feeding functions like sum(), any() or max().

Example: total = sum(v["Size"] for v in volumes) never builds an intermediate list. Dict and set comprehensions follow the same pattern: {i["InstanceId"]: i["State"]["Name"] for i in instances}.

7. What is a generator, and why is it useful for log processing?

Answer: A generator is a function that uses yield to produce values one at a time, pausing between them. Memory stays flat no matter how large the input, and generators chain into pipelines.

def read_lines(path):
    with open(path, encoding="utf-8",
              errors="replace") as fh:
        for line in fh:
            yield line.rstrip("\n")

def only_errors(lines):
    return (ln for ln in lines if " ERROR " in ln)

errors = only_errors(read_lines("app.log"))

Nothing is read until you iterate errors. A generator can be consumed once; if you need two passes, re-create it or materialise a bounded result.

8. What is a decorator? Write one that times a function.

Answer: A decorator is a callable that takes a function and returns a wrapped function, adding behaviour (timing, retries, logging, permission checks) without changing the original code. Use functools.wraps so the wrapper keeps the original name and docstring, which matters for logs and debugging.

import functools, logging, time

log = logging.getLogger(__name__)

def timed(func):
    @functools.wraps(func)
    def wrapper(*args, **kwargs):
        start = time.perf_counter()
        try:
            return func(*args, **kwargs)
        finally:
            ms = (time.perf_counter() - start) * 1000
            log.info("%s took %.1f ms",
                     func.__name__, ms)
    return wrapper

The finally block logs the duration even if the function raises. A retry decorator with backoff is Q53.

9. What is a context manager, and where do you use one in DevOps code?

Answer: A context manager wraps setup and teardown around a block via with: files close, locks release, temporary directories are deleted, even when an exception occurs. Write your own with a class (__enter__/__exit__) or with contextlib.contextmanager.

import contextlib, os

@contextlib.contextmanager
def working_dir(path):
    old = os.getcwd()
    os.chdir(path)
    try:
        yield
    finally:
        os.chdir(old)

Typical uses: tempfile.TemporaryDirectory() for build artefacts, threading.Lock(), SSH or database connections, and contextlib.suppress(FileNotFoundError) for cleanup where a missing file is fine.

10. How should exception handling look in an automation script?

Answer: Catch specific exceptions you can handle, let the rest propagate, and keep context. Use try/except/else/finally: else runs only on success, finally always runs. Re-raise with raise NewError(...) from err to keep the original traceback. Never write a bare except:, because it also swallows KeyboardInterrupt and SystemExit, and never except Exception: pass.

At the top level of a CLI, catch expected failures, log them once with context, and exit with a non-zero code so the pipeline or scheduler sees the failure. Since Python 3.11, ExceptionGroup and except* handle several errors from concurrent tasks (for example asyncio.TaskGroup).

Interview tip: Mention EAFP ("easier to ask forgiveness than permission"): try to open the file and handle FileNotFoundError, rather than checking exists() first and racing another process.

11. What does if __name__ == "__main__": do, and how do you structure a reusable script?

Answer: When a file is run directly, its __name__ is "__main__"; when imported, it is the module name. Guarding the entry point means importing the module (from tests or another tool) does not trigger side effects such as deleting resources. A clean structure is: pure functions that take inputs and return results, one main(argv=None) that parses arguments and wires things together, and raise SystemExit(main()) under the guard. That layout makes every function unit-testable.

Files, data formats, regex and the OS

12. os.path or pathlib: which do you use and why?

Answer: Prefer pathlib.Path for new code. Paths become objects with readable operations: Path("/var/log") / "app" / "app.log", .exists(), .suffix, .stat().st_size, .glob("*.log"), .rglob("*.gz"), .read_text(), .mkdir(parents=True, exist_ok=True). It handles separators across Linux and Windows. os.path still works and you will read plenty of it in older scripts; the two interoperate because most APIs accept path-like objects.

13. Which os and shutil functions do you use most in ops scripts?

Answer: From os: os.environ.get("NAME", default) for configuration, os.scandir() and os.walk() for fast directory traversal, os.replace() for atomic renames, os.getpid(), and os.cpu_count(). From shutil: copy2() (keeps metadata), copytree(), move(), rmtree(), disk_usage() for free-space checks, which() to confirm a binary exists before calling it, and make_archive().

Interview tip: Point out that shutil.rmtree() on a path built from user input is dangerous; validate that the resolved path sits under an expected base directory before deleting.

14. How do you read and write JSON, YAML and TOML safely?

Answer: JSON: json.load(fh) / json.dump(obj, fh, indent=2); use default=str when dumping datetimes from boto3 responses. YAML: always yaml.safe_load() from PyYAML, never plain yaml.load() on untrusted input, because the full loader can construct arbitrary Python objects. TOML: the standard library's tomllib (Python 3.11+) reads TOML such as pyproject.toml; it is read-only, so use a third-party library to write.

Validate after parsing. A Kubernetes manifest or pipeline config that parses is not necessarily correct; check required keys and types, or validate with a schema or Pydantic model, before acting on it.

15. How do you process CSV reports correctly?

Answer: Use the csv module, not line.split(","), because quoted fields can contain commas. Open files with newline="" as the documentation requires, and use csv.DictReader / csv.DictWriter to work with column names. For a cost or inventory report, stream rows and aggregate with a dict or collections.Counter; reach for pandas only when the analysis really needs it, because it is a heavy dependency for a Lambda or a container image.

16. How do you use regular expressions to parse log lines?

Answer: Compile the pattern once, use raw strings, and use named groups so the code reads like the log format. re.match anchors at the start of the string; re.search scans anywhere; re.fullmatch must match the whole string.

import re

ACCESS = re.compile(
    r'(?P<ip>\d{1,3}(?:\.\d{1,3}){3}) \S+ \S+ '
    r'\[(?P<ts>[^\]]+)\] "(?P<method>[A-Z]+) '
    r'(?P<path>\S+) [^"]*" (?P<status>\d{3})'
)

m = ACCESS.search(line)
if m and m["status"].startswith("5"):
    print(m["ip"], m["path"])

For structured (JSON) logs, skip regex and parse with json.loads. Q49 turns this into a full "top error IPs" task.

17. How do you run shell commands from Python safely?

Answer: Use subprocess.run() with a list of arguments, check=True, a timeout, and captured output. Avoid shell=True, and never combine it with user input: a value like "x; rm -rf ~" becomes a second command. With a list, each element is passed directly to the program, so there is no shell to inject into.

import subprocess

def disk_report(mount: str) -> str:
    result = subprocess.run(
        ["df", "-h", "--", mount],
        capture_output=True, text=True,
        check=True, timeout=30,
    )
    return result.stdout

If you truly need shell features such as pipes, build the pipeline in Python or quote every untrusted value with shlex.quote(). Avoid os.system(): it always uses a shell and gives you no output or useful error handling. Catch subprocess.CalledProcessError and subprocess.TimeoutExpired separately so the log says which one happened.

18. How do you write a config or state file so a crash never leaves it half-written?

Answer: Write to a temporary file in the same directory, flush and fsync it, then os.replace() it over the target. The rename is atomic on the same filesystem, so readers see either the old file or the new one, never a partial one.

import json, os, tempfile

def atomic_write_json(path, data):
    folder = os.path.dirname(path) or "."
    fd, tmp = tempfile.mkstemp(dir=folder)
    try:
        with os.fdopen(fd, "w") as fh:
            json.dump(data, fh, indent=2)
            fh.flush()
            os.fsync(fh.fileno())
        os.replace(tmp, path)
    except BaseException:
        os.unlink(tmp)
        raise

19. How do you handle dates and time zones in logs and automation?

Answer: Work in timezone-aware UTC internally: datetime.now(timezone.utc), not the naive datetime.now() or the deprecated datetime.utcnow(). boto3 returns aware datetimes for fields such as LaunchTime and CreateTime, and comparing an aware datetime with a naive one raises TypeError. Convert to local time only for display, with zoneinfo.ZoneInfo("Asia/Kolkata"). Use datetime.fromisoformat() for ISO 8601 strings and log timestamps in ISO format so they sort correctly.

CLIs, logging, environments and packaging

20. argparse or click: how do you build a command-line tool?

Answer: argparse is in the standard library, so it suits scripts that must run anywhere with no dependencies. click (and Typer, built on it) gives decorators, nested command groups, prompts, environment-variable defaults and easy testing with CliRunner; it suits larger internal tools. Either way: destructive tools default to dry-run, need an explicit flag to act, and print what they would do.

import argparse

def build_parser():
    p = argparse.ArgumentParser(
        description="Delete unattached volumes")
    p.add_argument("--region", required=True)
    p.add_argument("--execute",
                   action="store_true",
                   help="really delete")
    p.add_argument("-v", "--verbose",
                   action="count", default=0)
    return p

args = build_parser().parse_args(
    ["--region", "ap-south-1"])
assert args.execute is False

21. How should an automation script log, and why not just use print?

Answer: Use the logging module: it gives levels, timestamps, module names and pluggable handlers, and it can be silenced or redirected without code changes. Create a logger per module with logging.getLogger(__name__), configure handlers once in main(), and use lazy formatting (log.info("deleted %s", vol_id)). Send logs to stderr and real output (a report, JSON) to stdout so the tool composes in pipelines.

In containers and Lambda, log to stdout/stderr and let the platform ship logs; structured JSON lines make CloudWatch Logs Insights or Loki queries far easier. Never log secrets, tokens or full request bodies, and use log.exception() inside except blocks to keep the traceback.

22. What do exit codes, stdout/stderr and signals mean for a script that runs in CI or cron?

Answer: Exit code 0 means success; any non-zero code means failure, and pipelines, cron wrappers, systemd and Kubernetes Jobs all depend on it. Use sys.exit(main()) or raise SystemExit(code); an uncaught exception exits with code 1. Distinguish codes when useful, for example 2 for bad arguments (argparse does this) and 3 for "found drift". Handle SIGTERM (what Kubernetes and systemd send before killing you) with the signal module to finish the current item and exit cleanly. Our Linux interview questions guide covers signals and systemd from the OS side.

23. Why do you need virtual environments, and how do pip and requirements files fit in?

Answer: A virtual environment (python -m venv .venv) isolates a project's packages from the system Python and from other projects, so upgrading one tool does not break another or the OS's own Python tools. Many Linux distributions now mark the system Python as externally managed and refuse pip install outside a venv. Pin dependencies for reproducible deployments: a lockfile or a fully pinned requirements.txt (ideally with hashes), separate from the loose ranges you declare in pyproject.toml. Never sudo pip install.

24. What is uv, and how does it change a Python DevOps workflow?

Answer: uv is a Python package and project manager from Astral (the makers of Ruff), written in Rust. It aims to replace several tools in one: pip, pip-tools, pipx, virtualenv, pyenv and much of Poetry's role. Key commands:

  • uv init creates a project with pyproject.toml; uv add boto3 adds a dependency.
  • uv lock writes a cross-platform uv.lock; uv sync makes the environment match it.
  • uv run script.py runs inside the project environment, and supports inline script metadata (PEP 723), which is handy for single-file ops scripts.
  • uv python install 3.13 installs interpreters; uvx ruff check . runs a tool without installing it globally.
  • uv pip install ... offers a pip-compatible interface for existing workflows.

Interview tip: In CI, uv sync --locked (install from the lockfile and fail if it no longer matches pyproject.toml) gives fast, reproducible installs. Commit uv.lock for applications and tools.

25. What goes into pyproject.toml, and how do you ship an internal CLI?

Answer: pyproject.toml is the standard project file. [build-system] names the build backend (for example hatchling or setuptools), [project] holds name, version, requires-python and dependencies, [project.scripts] creates console commands, [dependency-groups] holds development-only groups, and [tool.*] tables configure Ruff, mypy and pytest.

[project]
name = "opskit"
version = "0.4.0"
requires-python = ">=3.12"
dependencies = ["boto3>=1.40", "click>=8.1"]

[project.scripts]
opskit = "opskit.cli:main"

[dependency-groups]
dev = ["pytest", "moto[ec2]", "ruff", "mypy"]

[build-system]
requires = ["hatchling"]
build-backend = "hatchling.build"

Build a wheel (uv build or python -m build), publish it to an internal index such as AWS CodeArtifact or Azure Artifacts, and install it with uv tool install opskit or pipx. Or package the tool into a container image so CI runners get a fixed version.

26. When do you use paramiko, Fabric, the Docker SDK or Ansible instead of subprocess?

Answer: paramiko is a low-level SSH library (run a command, transfer a file); Fabric wraps it for running tasks across hosts. Use them for small, ad hoc remote actions, and load known host keys rather than AutoAddPolicy, which silently trusts any server. The Docker SDK for Python (docker package) talks to the Docker API directly, so you get structured objects instead of parsing CLI text. For repeatable configuration of many servers, use Ansible, which is written in Python and gives idempotent modules and inventory. In cloud environments, prefer the provider's run-command service (for example AWS Systems Manager) over opening SSH at all.

HTTP clients and concurrency

27. How do you call a REST API with requests so it never hangs and retries safely?

Answer: requests has no default timeout, so a call can wait forever on a dead server. Always pass timeout=(connect, read). Use a Session for connection reuse, and mount urllib3's Retry to retry transient failures with backoff.

import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

def make_session():
    retry = Retry(
        total=4, backoff_factor=0.5,
        status_forcelist=(429, 500, 502, 503, 504),
        allowed_methods=frozenset({"GET", "PUT",
                                   "DELETE"}),
    )
    s = requests.Session()
    s.mount("https://", HTTPAdapter(max_retries=retry))
    return s

resp = make_session().get(
    "https://jenkins.example.internal/api/json",
    timeout=(3.05, 15))
resp.raise_for_status()

POST is not in the retry list on purpose: retrying a non-idempotent call (such as triggering a Jenkins build) can run it twice unless the API supports an idempotency key. Retry respects Retry-After headers by default.

28. What does httpx offer over requests?

Answer: httpx has a requests-like API plus a native async client (httpx.AsyncClient), optional HTTP/2, and a default timeout (five seconds) instead of none. It also ships httpx.MockTransport, which makes testing clients without a network easy. Use requests for simple synchronous scripts that already depend on it; use httpx when you need async fan-out to many endpoints or want the stricter defaults. Retries in httpx are limited to connection failures (transport=httpx.HTTPTransport(retries=3)), so status-code retries still need your own logic or a library such as tenacity.

29. What is the GIL, and what is the free-threaded build?

Answer: The Global Interpreter Lock lets only one thread execute Python bytecode at a time in standard CPython. It is released during blocking I/O, so threads still speed up network-bound work (API calls, SSH, downloads), but they do not speed up pure-Python CPU work; for that you use processes.

CPython also has a free-threaded build (PEP 703) without the GIL, typically installed as python3.14t. It was experimental in 3.13; under PEP 779 it became officially supported but still optional in 3.14, and it is not the default. Single-threaded code runs somewhat slower on it, memory use is higher, and C extensions that are not marked compatible can re-enable the GIL. You can check with sys._is_gil_enabled(). For typical DevOps automation, which is I/O-bound, the standard build is still the right choice.

30. Threads, asyncio or multiprocessing: how do you choose?

ApproachSuitsWatch out for
ThreadPoolExecutorBlocking I/O with sync libraries: boto3, requests, paramikoShared mutable state; create one boto3 client per thread or share a client, never a resource
asyncioVery many concurrent network calls with async libraries (httpx, aiohttp)One blocking call stalls everything; boto3 is not async
ProcessPoolExecutorCPU-heavy work: compressing, hashing, parsing huge filesPickling overhead, higher memory, start-up cost

Answer: Decide by where the time goes. Waiting on the network with sync libraries: threads. Thousands of concurrent requests with async libraries: asyncio. Burning CPU: processes. Whatever you choose, bound the concurrency, because cloud APIs throttle and a hundred threads usually just collect throttling errors faster.

31. Write code that inventories EC2 instances across regions concurrently.

Answer: Use a thread pool with a bounded number of workers, give each region its own client, and use a paginator inside each task. Collect results and failures per region rather than letting one failed region kill the whole report.

from concurrent.futures import (
    ThreadPoolExecutor, as_completed)
import boto3

def count_running(session, region):
    ec2 = session.client("ec2", region_name=region)
    pages = ec2.get_paginator(
        "describe_instances").paginate(
        Filters=[{"Name": "instance-state-name",
                  "Values": ["running"]}])
    return sum(len(r["Instances"])
               for p in pages
               for r in p["Reservations"])

def inventory(regions, workers=8):
    session = boto3.Session()
    out, errors = {}, {}
    with ThreadPoolExecutor(workers) as pool:
        futs = {pool.submit(count_running,
                            session, r): r
                for r in regions}
        for f in as_completed(futs):
            region = futs[f]
            try:
                out[region] = f.result()
            except Exception as exc:
                errors[region] = str(exc)
    return out, errors

boto3 clients are generally thread-safe; sessions and resources are not, so create clients up front or per thread rather than sharing a resource object across threads.

32. How do you limit concurrency and add timeouts with asyncio?

Answer: Use one shared AsyncClient, an asyncio.Semaphore to cap in-flight requests, and a timeout per call. asyncio.timeout() (Python 3.11+) or the client's own timeout bounds each operation; asyncio.gather(..., return_exceptions=True) keeps one failure from cancelling the rest.

import asyncio, httpx

async def check(client, sem, url):
    async with sem:
        try:
            async with asyncio.timeout(5):
                r = await client.get(url)
            return url, r.status_code
        except (httpx.HTTPError, TimeoutError) as e:
            return url, type(e).__name__

async def check_all(urls, limit=20):
    sem = asyncio.Semaphore(limit)
    async with httpx.AsyncClient() as client:
        return await asyncio.gather(
            *(check(client, sem, u) for u in urls))

Never call blocking code (boto3, time.sleep, requests) inside a coroutine; wrap it in asyncio.to_thread() if you must.

boto3, Lambda, Azure SDK and Kubernetes client

33. What is the difference between a boto3 session, client and resource?

Answer: A session holds configuration and credentials (profile, region). A client is a low-level interface that maps one-to-one to service API operations and returns dicts; it covers every operation and is generally thread-safe. A resource is a higher-level object interface (s3.Bucket("x").objects.all()) that is convenient but covers fewer services and is not thread-safe. The boto3 documentation states that the SDK team does not intend to add new features to the resources interface and that newer features are available through clients.

Interview tip: "I default to clients, especially for new services and in threaded code; I only use resources in existing code that already relies on them" is a strong, current answer.

34. How does boto3 find credentials, and how should production scripts authenticate?

Answer: boto3 walks a provider chain and stops at the first hit. In order: credentials passed to client(), then to Session(), environment variables, assume-role and web-identity providers configured in the profile, IAM Identity Center (SSO), the shared credentials file, console-login credentials (aws login), the config file, the legacy boto2 config, the container credential provider (ECS, EKS Pod Identity) and finally the EC2 instance metadata service.

In production, use roles: an instance profile on EC2, a task role on ECS, IRSA or EKS Pod Identity on Kubernetes, an execution role for Lambda, and OIDC federation from GitHub Actions or GitLab to assume a role in CI. Locally, use SSO profiles. Long-lived access keys in code or environment files are the classic finding in every security review.

35. Why do you need paginators? Show one.

Answer: Most list and describe APIs return results in pages with a NextToken or Marker. Code that reads only the first response silently misses resources, which is how a cleanup script "passes" in a small dev account and misbehaves in production. Paginators handle tokens for you, and JMESPath search() can flatten results.

import boto3

def list_bucket_keys(bucket, prefix=""):
    s3 = boto3.client("s3")
    paginator = s3.get_paginator("list_objects_v2")
    for page in paginator.paginate(
            Bucket=bucket, Prefix=prefix):
        for obj in page.get("Contents", []):
            yield obj["Key"], obj["Size"]

The function is a generator, so even buckets with millions of objects stream through with flat memory. PaginationConfig={"PageSize": ...} tunes page size where a service supports it.

36. What are waiters, and when do you use them?

Answer: A waiter polls a describe call until a resource reaches a state, for example ec2.get_waiter("instance_running"), "volume_available", "snapshot_completed" or CloudFormation's "stack_create_complete". It replaces hand-written while True: sleep() loops and raises WaiterError on failure or timeout.

waiter = ec2.get_waiter("instance_running")
waiter.wait(InstanceIds=[iid],
            WaiterConfig={"Delay": 10,
                          "MaxAttempts": 30})

Set WaiterConfig to match the real operation, so a Lambda does not wait longer than its own timeout. For long workflows (stack updates, AMI builds), prefer Step Functions or event-driven designs over a script blocking for many minutes.

37. How do you handle boto3 errors and configure retries?

Answer: Service errors arrive as botocore.exceptions.ClientError; inspect err.response["Error"]["Code"] to branch on "NoSuchKey", "AccessDenied", "ThrottlingException" and so on. Some clients also expose modelled exceptions such as s3.exceptions.NoSuchBucket. Network-level problems raise other botocore exceptions such as EndpointConnectionError.

Retries are configured with botocore.config.Config. The boto3 documentation lists three modes: legacy (documented as the default), standard and adaptive (adds client-side rate limiting, documented as experimental). Set the mode explicitly and add timeouts:

from botocore.config import Config

cfg = Config(
    retries={"mode": "standard",
             "total_max_attempts": 6},
    connect_timeout=5, read_timeout=30,
)
ec2 = boto3.client("ec2", config=cfg)

Note the subtlety: in a Config object max_attempts excludes the first call, while total_max_attempts (and the AWS_MAX_ATTEMPTS environment variable) includes it.

38. How do you run the same automation across many AWS accounts and regions?

Answer: From a central tooling account, call STS assume_role into a role that exists in every member account (often deployed with CloudFormation StackSets or Terraform), build a boto3.Session from the temporary credentials, then loop or fan out over regions. Get the account list from AWS Organizations and the region list from ec2.describe_regions(), skipping regions that are not enabled.

def session_for(account_id, role="OpsReadOnly"):
    sts = boto3.client("sts")
    creds = sts.assume_role(
        RoleArn=f"arn:aws:iam::{account_id}"
                f":role/{role}",
        RoleSessionName="inventory",
    )["Credentials"]
    return boto3.Session(
        aws_access_key_id=creds["AccessKeyId"],
        aws_secret_access_key=creds[
            "SecretAccessKey"],
        aws_session_token=creds["SessionToken"],
    )

Use a read-only role for reporting and a separate, tightly scoped role for changes, and record the account and region in every log line. Assumed-role credentials expire, so long jobs must re-assume.

39. Write a Python Lambda function triggered by S3 uploads. What matters?

Answer: The handler receives (event, context). Create clients outside the handler so warm invocations reuse them, decode object keys (they arrive URL-encoded, so a space appears as +), log in a structured way, and keep the function idempotent because events can be delivered more than once.

import json, logging, urllib.parse
import boto3

log = logging.getLogger()
log.setLevel(logging.INFO)
s3 = boto3.client("s3")  # reused when warm

def handler(event, context):
    done = []
    for rec in event.get("Records", []):
        bucket = rec["s3"]["bucket"]["name"]
        key = urllib.parse.unquote_plus(
            rec["s3"]["object"]["key"])
        head = s3.head_object(Bucket=bucket,
                              Key=key)
        log.info(json.dumps({
            "bucket": bucket, "key": key,
            "size": head["ContentLength"]}))
        done.append(key)
    return {"processed": done}

Give the execution role only s3:GetObject on that bucket, and never write output back to the same prefix that triggers the function, or it will invoke itself in a loop. The AWS serverless interview questions guide goes deeper on triggers, concurrency and Step Functions.

40. How do you package, version and deploy Python Lambda functions?

Answer: Small functions ship as a zip with dependencies vendored in; heavier ones (native libraries, large ML packages) ship as container images. Lambda layers share common dependencies across functions. Build dependencies for the right architecture and Python version (x86_64 or arm64, manylinux wheels), ideally inside a matching container, because a wheel built on a Mac will not import on Lambda.

Deploy with SAM, CDK, Terraform or the Serverless Framework, never by hand-editing code in the console. Publish immutable versions and point an alias (such as live) at them; aliases support weighted routing for canary releases and quick rollback. Bundle your own boto3 version rather than relying on the one in the runtime, as AWS recommends for dependency control. Upgrade runtimes before deprecation dates; python3.9 is already deprecated.

41. How do you authenticate with the Azure SDK for Python?

Answer: Use azure-identity. DefaultAzureCredential tries a chain in order: environment variables (service principal), workload identity, managed identity, then developer tools (Visual Studio Code, Azure CLI, Azure PowerShell, Azure Developer CLI) and a broker; interactive browser login is off by default. The same code therefore works on a laptop after az login and on an Azure VM or AKS pod with a managed or workload identity.

from azure.identity import DefaultAzureCredential
from azure.mgmt.resource import (
    ResourceManagementClient)

cred = DefaultAzureCredential()
rm = ResourceManagementClient(cred, sub_id)
for rg in rm.resource_groups.list():
    print(rg.name, rg.location)

Microsoft's own guidance is to replace DefaultAzureCredential with a specific credential such as ManagedIdentityCredential once deployed, because chains are harder to debug and can change behaviour if someone sets environment variables. In recent azure-identity versions the AZURE_TOKEN_CREDENTIALS environment variable can restrict the chain to prod or dev credentials or to one named credential. See the Azure interview questions guide for the platform side.

42. How do you use the Kubernetes Python client?

Answer: Install kubernetes, load configuration with config.load_kube_config() on a laptop or config.load_incluster_config() inside a pod (it uses the pod's service account token), then call typed APIs such as CoreV1Api and AppsV1Api.

from kubernetes import client, config

try:
    config.load_incluster_config()
except config.ConfigException:
    config.load_kube_config()

v1 = client.CoreV1Api()
pods = v1.list_namespaced_pod(
    "payments",
    field_selector="status.phase=Failed")
for pod in pods.items:
    print(pod.metadata.name,
          pod.status.reason)

Give the service account a namespaced Role with only the verbs it needs, use label and field selectors instead of listing everything, and use watch.Watch().stream() for event-driven tools. For anything that continuously reconciles state, a proper operator framework (Kopf in Python, or a Go controller) beats a cron script. The Kubernetes interview questions guide covers RBAC and controllers in detail.

Testing, typing, linting and secrets

43. How do you test automation code with pytest?

Answer: Keep logic in pure functions, then test them with plain assert. Use fixtures for setup (tmp_path for a scratch directory, monkeypatch for environment variables), parametrize to cover many inputs, and pytest.raises for error paths.

import pytest

@pytest.mark.parametrize("line,status", [
    ('1.2.3.4 - - [x] "GET / HTTP/1.1" 503 9',
     "503"),
    ('5.6.7.8 - - [x] "POST /a HTTP/1.1" 200 1',
     "200"),
])
def test_status(line, status):
    assert ACCESS.search(line)["status"] == status

def test_find_large(tmp_path):
    (tmp_path / "big.bin").write_bytes(b"x" * 2048)
    (tmp_path / "small.txt").write_text("hi")
    top = largest_files(tmp_path, n=1)
    assert top[0][1].endswith("big.bin")

Run tests in CI on every pull request. The GitHub Actions interview questions guide shows how to wire pytest, Ruff and mypy into a workflow.

44. How do you test code that calls AWS without touching a real account?

Answer: Three options, from most realistic to most targeted:

  • moto fakes AWS services in memory. Since moto 5, one decorator or context manager, mock_aws, covers all services. Set dummy credentials and a region in a fixture so nothing can reach a real account.
  • botocore Stubber queues exact expected requests and canned responses, good for asserting you sent the right parameters or for simulating a specific error such as throttling.
  • unittest.mock patches your own wrapper functions; keep it at your boundaries rather than mocking boto3 internals.
import os, pytest
from moto import mock_aws

@pytest.fixture
def aws():
    os.environ.update(
        AWS_ACCESS_KEY_ID="testing",
        AWS_SECRET_ACCESS_KEY="testing",
        AWS_DEFAULT_REGION="ap-south-1")
    with mock_aws():
        yield boto3.client("ec2")

moto does not emulate every behaviour or IAM policy evaluation exactly, so keep a small set of integration tests that run against a sandbox account too.

45. Why use type hints, mypy and Ruff in ops scripts?

Answer: Type hints (def tag(vol_id: str, tags: dict[str, str]) -> None) document intent and let a checker catch mistakes before a script runs in production. Python does not enforce them at runtime. mypy (or Pyright) checks them statically; the boto3-stubs or types-boto3 packages add types for boto3 clients so typos in parameter names are caught. Ruff, from Astral, is a fast linter and formatter that can replace flake8 and many of its plugins plus isort, and ruff format covers Black-style formatting. It also catches bugs such as mutable defaults and bare excepts, and has security rules modelled on Bandit, including shell=True checks.

Configure both in pyproject.toml, run them in pre-commit and CI, and fail the build on errors. Gradual typing works: start with new modules and public functions.

46. How do you handle passwords, API keys and tokens in Python scripts?

Answer: Never hard-code them or commit them, including in .env files, notebooks or test fixtures. Fetch them at runtime from a secret store: AWS Secrets Manager or SSM Parameter Store (SecureString), Azure Key Vault (SecretClient with DefaultAzureCredential), Google Secret Manager or HashiCorp Vault. Access is granted to the workload's identity (role or managed identity), not to a shared key.

In code: cache the secret in memory with a short TTL rather than calling the store per request (the AWS Parameters and Secrets Lambda Extension does this for Lambda), never log it, avoid passing it on a command line where ps can see it, and keep it out of exception messages. Add secret scanning (for example gitleaks or GitHub secret scanning) to CI so leaks are caught before merge. Rotation is Q51.

AI in Python automation

47. How do you call an LLM API safely from an automation script?

Answer: Treat it like any other external dependency with extra risks. Keep the API key in a secret store or use the cloud's identity-based access (Amazon Bedrock with an IAM role, Azure OpenAI in Microsoft Foundry with Entra ID). Set timeouts and retry only on throttling and server errors, with backoff. Put a token and cost budget on each run. Redact secrets, customer data and personal data from logs and tickets before they go into a prompt. Log the prompt version and model identifier with each call so a behaviour change can be traced. And treat everything the model returns as untrusted input, never as a command.

Real-world example: Consider a GCC operations team in Hyderabad that summarises overnight alarms into a morning digest. The script sends alarm names and metric summaries, not raw log lines with customer IDs, and posts the summary to Teams with a link back to the source alarms so humans can verify it.

Python patterns specific to LLM applications (streaming, Pydantic models, FastAPI services) are covered in the Python for AI engineers guide.

48. An LLM suggests remediation actions for alerts. How do you use its output without letting it break production?

Answer: Ask for structured output (a JSON schema or the provider's structured-output or tool-calling feature), validate it in code, and only let validated, allow-listed, low-risk actions run automatically. Everything else becomes a suggestion for a human. The model never chooses arbitrary commands; it picks from actions your code already implements.

from typing import Literal
from pydantic import BaseModel, ValidationError

class Action(BaseModel):
    action: Literal["restart_service",
                    "scale_out", "open_ticket"]
    target: str
    reason: str

AUTO_OK = {"open_ticket"}

def decide(raw_json: str):
    try:
        act = Action.model_validate_json(raw_json)
    except ValidationError:
        return "rejected", None
    if act.action in AUTO_OK:
        return "auto", act
    return "needs_approval", act

Add guardrails around it: validate target against your inventory, run in dry-run first, require approval for anything that changes production, and log every decision. Our explainers on function calling and structured outputs and AI guardrails go deeper.

If you want to practise these scripts on real AWS, Azure and Linux environments with feedback from trainers, Cloudsoft's Multi-cloud DevOps with Linux & Python course covers Python automation alongside Linux, CI/CD, containers and the major clouds, in the Ameerpet classroom or live online.

Python for DevOps scenarios and coding tasks

49. Parse a web server access log and print the top ten IP addresses causing 5xx errors.

Answer: Stream the file line by line, extract IP and status with a compiled regex, count with collections.Counter, and print most_common(). Accept the path as an argument and handle gzipped rotated logs too.

import collections, gzip, re, sys

LINE = re.compile(
    r'^(?P<ip>\S+) .*?" (?P<status>\d{3}) ')

def open_log(path):
    if path.endswith(".gz"):
        return gzip.open(path, "rt",
                         errors="replace")
    return open(path, errors="replace")

def top_error_ips(paths, n=10):
    counts = collections.Counter()
    for path in paths:
        with open_log(path) as fh:
            for line in fh:
                m = LINE.match(line)
                if m and m["status"][0] == "5":
                    counts[m["ip"]] += 1
    return counts.most_common(n)

if __name__ == "__main__":
    for ip, hits in top_error_ips(sys.argv[1:]):
        print(f"{hits:8d}  {ip}")

What I would check:

  1. The real log format: behind a load balancer, the first field may be the balancer's IP, and the client IP is in X-Forwarded-For or the load balancer's own logs.
  2. Malformed lines: they are skipped, but count them so a format change is noticed.
  3. Whether 5xx errors come from a few clients (a bad integration or bot) or from everyone (a backend problem).

Production consideration: For a one-off on a server, this script or awk is fine. For recurring analysis, query centralised logs (CloudWatch Logs Insights, OpenSearch, Loki) instead of copying files around.

50. Write a script that deletes unattached, untagged EBS volumes safely.

Answer: Find volumes in the available state (not attached), skip any with an owner or keep tag, skip recently created ones, default to dry-run, and require an explicit flag to delete. Report everything it would do.

from datetime import datetime, timedelta, timezone
import boto3

KEEP_TAGS = {"Owner", "keep"}

def candidates(ec2, min_age_days=7):
    cutoff = (datetime.now(timezone.utc)
              - timedelta(days=min_age_days))
    pages = ec2.get_paginator(
        "describe_volumes").paginate(
        Filters=[{"Name": "status",
                  "Values": ["available"]}])
    for page in pages:
        for v in page["Volumes"]:
            tags = {t["Key"] for t in
                    v.get("Tags", [])}
            if tags & KEEP_TAGS:
                continue
            if v["CreateTime"] > cutoff:
                continue
            yield v["VolumeId"], v["Size"]

def cleanup(ec2, execute=False, min_age_days=7):
    report = []
    for vol_id, size in candidates(
            ec2, min_age_days):
        if execute:
            ec2.delete_volume(VolumeId=vol_id)
        report.append((vol_id, size,
                       "deleted" if execute
                       else "would delete"))
    return report

What I would check:

  1. That the account's tagging policy actually uses the tags I am treating as "keep", and with the right spelling and case.
  2. Whether a final snapshot is required before deletion for audit or recovery.
  3. The IAM role: it should allow ec2:DeleteVolume only in the intended accounts, ideally conditioned on tags.

Production consideration: Run dry-run, send the report to the owning teams, wait a grace period (or tag volumes as pending-delete first), then delete. EC2's own DryRun=True parameter only checks permissions; it is not a substitute for your script's dry-run report. Tested here with moto.

51. How would you automate rotation of a database password stored in AWS Secrets Manager?

Answer: First, check whether managed rotation or an AWS-provided rotation template already covers the database (RDS and others have them). For a custom target, write a rotation Lambda that Secrets Manager calls four times with a Step: createSecret (generate a new value and store it as AWSPENDING), setSecret (change the password on the database), testSecret (log in with the pending value) and finishSecret (move AWSCURRENT to the new version; the old one becomes AWSPREVIOUS).

import json

def create_secret(sm, arn, token):
    try:
        sm.get_secret_value(SecretId=arn,
                            VersionId=token,
                            VersionStage="AWSPENDING")
        return  # idempotent: already created
    except sm.exceptions.ResourceNotFoundException:
        pass
    cur = json.loads(sm.get_secret_value(
        SecretId=arn,
        VersionStage="AWSCURRENT")["SecretString"])
    cur["password"] = sm.get_random_password(
        PasswordLength=32,
        ExcludeCharacters="/@\"'\\")["RandomPassword"]
    sm.put_secret_value(
        SecretId=arn, ClientRequestToken=token,
        SecretString=json.dumps(cur),
        VersionStages=["AWSPENDING"])

What I would check:

  1. That applications fetch the secret at runtime and refresh on authentication failure, otherwise rotation breaks them.
  2. Single-user versus alternating-users strategy: alternating users avoids a window where the old password no longer works.
  3. Network path: the rotation Lambda needs to reach both the database and the Secrets Manager endpoint (VPC endpoint or NAT).

Production consideration: Each step must be idempotent because Secrets Manager may retry it, and the function must never log secret values.

52. A nightly script sometimes hangs for hours and never finishes. How do you find and fix it?

Answer: A hang almost always means something is waiting without a timeout: an HTTP call with no timeout, a subprocess waiting for input, an SSH session, a lock or a database query. Find where it is stuck, then put a bound on every external wait.

import subprocess, requests

def fetch_status(url):
    return requests.get(url,
                        timeout=(3, 20)).json()

def run_backup(cmd, limit=900):
    try:
        subprocess.run(cmd, check=True,
                       timeout=limit,
                       stdin=subprocess.DEVNULL)
    except subprocess.TimeoutExpired:
        log.error("backup exceeded %ss", limit)
        raise

What I would check:

  1. Where it is stuck, using py-spy dump --pid on the live process, or faulthandler.dump_traceback_later() built into the script.
  2. Every network call and subprocess for a missing timeout, including SDK clients (connect_timeout/read_timeout in botocore Config).
  3. Subprocesses that prompt for input (a password prompt, a host-key question), which wait forever; stdin=DEVNULL makes them fail fast instead.

Production consideration: Add an overall deadline as well, such as systemd's RuntimeMaxSec, a Kubernetes Job's activeDeadlineSeconds or timeout 1h in cron, plus an alert when a run does not finish.

53. Write a retry decorator with exponential backoff and jitter.

Answer: Retry only errors that are transient, cap the number of attempts and the delay, add random jitter so many clients do not retry in lockstep, and re-raise the last error when attempts run out.

import functools, random, time

def retry(exceptions, attempts=5, base=0.5,
          cap=20.0, sleep=time.sleep):
    def deco(func):
        @functools.wraps(func)
        def wrapper(*args, **kwargs):
            for n in range(1, attempts + 1):
                try:
                    return func(*args, **kwargs)
                except exceptions:
                    if n == attempts:
                        raise
                    delay = min(cap, base * 2 ** n)
                    sleep(random.uniform(0, delay))
        return wrapper
    return deco

@retry((ConnectionError, TimeoutError))
def call_api():
    ...

This is "full jitter": a random delay between zero and the exponential ceiling. The injectable sleep makes the decorator testable without waiting.

What I would check:

  1. That the operation is idempotent; retrying "create order" or "trigger build" can duplicate it.
  2. That the SDK is not already retrying underneath (boto3 does), which multiplies attempts.
  3. That permanent errors such as 400 or AccessDenied are not in the retry list.

Production consideration: In real code, the tenacity library or the SDK's built-in retry configuration is usually better than a hand-rolled decorator; interviewers ask you to write it to see that you understand the trade-offs.

54. A server's disk is almost full. Write a Python tool to find the largest files.

Answer: Walk the tree with os.walk (or os.scandir for speed), skip symlinks, ignore files that vanish or cannot be read, and keep only the top N with heapq.nlargest so memory stays bounded.

import heapq, os

def iter_files(root):
    for folder, dirs, files in os.walk(root):
        for name in files:
            path = os.path.join(folder, name)
            try:
                st = os.lstat(path)
            except OSError:
                continue  # vanished/no access
            if os.path.isfile(path) and \
               not os.path.islink(path):
                yield st.st_size, path

def largest_files(root, n=20):
    return heapq.nlargest(n, iter_files(str(root)))

What I would check:

  1. Whether df and the file sizes disagree: a deleted file still held open by a process uses space but is invisible to this walk; lsof +L1 finds it.
  2. Whether to stay on one filesystem (compare st_dev) so the scan does not wander into network mounts.
  3. Which files are safe to remove: rotated logs and caches, not database files.

Production consideration: Fix the cause (log rotation, retention policies, a container writing to the root filesystem), and add a disk-usage alert with shutil.disk_usage() or your monitoring agent. The Linux guide's disk-full scenario covers the shell side.

55. Write a script that checks TLS certificate expiry across a list of domains.

Answer: Open a TLS connection with verification on, read the peer certificate's notAfter, compute the days remaining, and check domains concurrently with timeouts. A failed handshake (already expired, wrong name, untrusted chain) is itself a finding.

import socket, ssl, time
from concurrent.futures import ThreadPoolExecutor

def days_left(host, port=443, timeout=5):
    ctx = ssl.create_default_context()
    with socket.create_connection(
            (host, port), timeout=timeout) as sock:
        with ctx.wrap_socket(
                sock, server_hostname=host) as tls:
            cert = tls.getpeercert()
    expires = ssl.cert_time_to_seconds(
        cert["notAfter"])
    return int((expires - time.time()) // 86400)

def check(host, warn=21):
    try:
        d = days_left(host)
        return host, d, "WARN" if d < warn else "OK"
    except (OSError, ssl.SSLError) as exc:
        return host, None, f"ERROR {exc}"

def check_all(hosts):
    with ThreadPoolExecutor(10) as pool:
        return list(pool.map(check, hosts))

What I would check:

  1. Every endpoint behind a hostname: different load balancer nodes or CDN edges can serve different certificates.
  2. Internal domains that use a private CA, which need that CA bundle in the context (cafile=), not verification turned off.
  3. Whether ACM, Key Vault or cert-manager already renews these, so the script becomes a safety net, not the process.

Production consideration: Run it on a schedule, emit results as metrics or alerts (for example to CloudWatch or Prometheus), and alert well before expiry so there is time for change approval.

56. A script works when you run it by hand but fails when cron runs it. Why?

Answer: cron runs with a minimal environment: a short PATH, a different working directory (usually the home directory), no shell profile, no activated virtualenv, often no HOME-based config you expect, no terminal, and possibly a different user. Each of these breaks something that worked interactively.

# crontab: absolute paths, venv python,
# log output, overall time limit
15 2 * * * cd /opt/opskit && timeout 1h \
  /opt/opskit/.venv/bin/python -m opskit.report \
  >> /var/log/opskit/report.log 2>&1

What I would check:

  1. Where output goes: without redirection, errors are mailed locally or lost. Capture stderr first; the traceback usually names the problem.
  2. Relative paths in the script: resolve them from Path(__file__).resolve().parent, not the current directory.
  3. Credentials: an AWS profile or SSO session that existed in your shell does not exist for cron; on EC2 rely on the instance role, elsewhere use a dedicated profile with AWS_PROFILE set in the crontab.
  4. Binaries called through subprocess: use shutil.which() at start-up and fail with a clear message, or call them by absolute path.
  5. Locale and encoding: open files with an explicit encoding="utf-8".

Production consideration: systemd timers give logging in the journal, dependency ordering and resource limits for free, and on Kubernetes a CronJob with a fixed image removes the "works on my machine" problem entirely.

57. Your multi-region compliance script gets throttling errors and takes too long. What do you change?

Answer: Reduce calls first, then add bounded concurrency with proper retries. Use filters server-side (Filters=, tag filters) instead of fetching everything and filtering in Python. Use paginators with larger page sizes where allowed. Fan out across regions with a small thread pool, as in Q31, but keep concurrency per API low. Configure botocore retries explicitly (standard or adaptive mode) instead of wrapping calls in your own retry loop on top of the SDK's.

What I would check:

  1. Which API is throttled, from ClientError codes in the logs; limits are often per account and per region, and other tools share them.
  2. Whether the same data is available from a cheaper source: AWS Config, Resource Explorer or a Config aggregator query can replace thousands of describe calls.
  3. Whether the script makes per-resource calls in a loop (an N+1 pattern) that a batch API could replace.

Production consideration: For continuous compliance, managed rules plus event-driven checks scale better than a script that sweeps every account nightly. Reliability patterns like these are also covered in our SRE interview questions.

58. An SQS-triggered Python Lambda processes some messages twice and sends others to the DLQ unnecessarily. Why?

Answer: By default, if the function raises for one message, the whole batch becomes visible again, so successful messages are processed again. Enable ReportBatchItemFailures on the event source mapping, catch errors per message and return only the failed message IDs. Also make processing idempotent, because SQS standard queues deliver at least once.

def handler(event, context):
    failures = []
    for rec in event["Records"]:
        try:
            process(json.loads(rec["body"]))
        except Exception:
            log.exception("failed %s",
                          rec["messageId"])
            failures.append(
                {"itemIdentifier": rec["messageId"]})
    return {"batchItemFailures": failures}

What I would check:

  1. The queue's visibility timeout against the function timeout: if processing takes longer than the visibility timeout, the message reappears while still being processed.
  2. An idempotency key (message ID or a business key) recorded in DynamoDB with a conditional write, so a duplicate is detected and skipped.
  3. The redrive policy's maximum receive count, which decides how quickly poison messages reach the DLQ.

Production consideration: Powertools for AWS Lambda (Python) has batch-processing and idempotency utilities that implement these patterns, which is usually better than writing them yourself.

59. Code review: what is wrong with this function, and how would you fix it?

def backup(db, user, password, dest, opts=[]):
    try:
        os.system(f"mysqldump -u {user} "
                  f"-p{password} {db} > {dest}")
        opts.append("done")
    except:
        pass

Answer: At least six problems: shell injection through every argument; the password appears on the command line, visible in ps and shell history; os.system ignores failures (it returns an exit status that is thrown away, and it raises nothing); the bare except: pass hides everything; the mutable default opts=[] leaks state between calls; and there is no timeout. A fixed version:

def backup(db, user, dest, timeout=3600):
    env = {**os.environ,
           "MYSQL_PWD": get_secret("db/backup")}
    with open(dest, "wb") as out:
        subprocess.run(
            ["mysqldump", "-u", user, "--", db],
            stdout=out, env=env, check=True,
            timeout=timeout)

The password comes from a secret store and is passed through the environment of the child process only (an option file with tight permissions is even better), arguments are a list, failures raise, and the call is bounded. Interviewers use this question to see how many issues you spot and in what order; lead with security.

60. A script that analyses a large log export gets OOM-killed in its container. How do you fix it?

Answer: It is almost certainly loading everything into memory: fh.read(), readlines(), json.load() on a huge file, building a big list of parsed records, or calling pd.read_csv() on the whole file. Switch to streaming: iterate lines, parse each record, aggregate into small structures and discard the rest.

import collections, gzip, json

def errors_per_service(path):
    counts = collections.Counter()
    with gzip.open(path, "rt",
                   encoding="utf-8") as fh:
        for line in fh:  # one record at a time
            try:
                rec = json.loads(line)
            except json.JSONDecodeError:
                counts["_bad_lines"] += 1
                continue
            if rec.get("level") == "ERROR":
                counts[rec.get("service", "?")] += 1
    return counts

What I would check:

  1. The file format: JSON Lines streams naturally; a single giant JSON array needs a streaming parser such as ijson.
  2. Memory growth with tracemalloc snapshots, to confirm which structure grows.
  3. The container's memory limit versus what the job actually needs, and whether the data should be queried where it lives (Athena, Logs Insights) instead of downloaded.

Production consideration: If pandas is genuinely needed, read in chunks (chunksize=) with explicit column types, or use a columnar format such as Parquet and read only the needed columns.

Key takeaways

  • Python for DevOps interviews reward safe automation: dry-run by default, timeouts on every external call, bounded retries and clear exit codes.
  • Know the 2026 baseline: Python 3.14 is the current stable line, 3.9 and 3.10 are end of life, and the free-threaded build is supported but optional.
  • In boto3, default to clients, always paginate, use waiters instead of sleep loops, and set retry mode and timeouts explicitly in Config.
  • Authenticate with roles and identities (instance profiles, IRSA or Pod Identity, managed identity, OIDC in CI), never long-lived keys in code.
  • Pick concurrency by bottleneck: threads for blocking I/O, asyncio for many async calls, processes for CPU, and always bound it.
  • Package and test like software: pyproject.toml, uv with a lockfile, pytest with moto, Ruff and mypy in CI.
  • LLM output in automation is untrusted input: validate structure, allow-list actions and keep humans approving risky changes.

Interview preparation checklist

  • Write, from memory, a generator-based log parser with a compiled regex and Counter.
  • Explain mutable defaults, shallow versus deep copy and the GIL in under a minute each.
  • Rewrite an os.system call as a safe subprocess.run list with check and timeout.
  • Build a small CLI with argparse or click that defaults to dry-run and logs with the logging module.
  • Create a project with uv init, add boto3 and pytest, commit uv.lock and run tests with uv run pytest.
  • Write a boto3 script with a paginator, a waiter, a ClientError branch and an explicit retry Config, and test it with moto's mock_aws.
  • Describe the boto3 credentials chain and DefaultAzureCredential's chain, and how each one picks up a role or managed identity in production.
  • Deploy one small Python Lambda with an alias, and be ready to discuss idempotency and partial batch responses.
  • Prepare two stories: a script you made safer (dry-run, timeouts, tests) and an automation failure you diagnosed.
  • Run Ruff and mypy on your own scripts and fix what they report before the interview.

FAQ

How much Python do DevOps engineers need to know?

Enough to write and maintain reliable automation: core data types, functions, exceptions, files and JSON, subprocess, logging, a CLI library, HTTP calls with timeouts, one cloud SDK such as boto3, and pytest. Deep object-oriented design and advanced metaprogramming are rarely tested for DevOps roles.

Is Python or Bash better for DevOps automation?

Both have a place. Bash is fine for short glue around existing commands. Once a script needs error handling, data structures, API calls, tests or more than a page of logic, Python is easier to make safe and maintainable.

Which Python version should I learn for DevOps in 2026?

Learn on a currently supported version such as Python 3.12, 3.13 or 3.14. Python 3.9 and 3.10 are end of life, and cloud runtimes like AWS Lambda are dropping older versions on their own schedules.

What are the most common boto3 interview questions?

Clients versus resources, the credentials chain, paginators, waiters, error handling with ClientError, retry configuration, assuming roles across accounts and writing a small cleanup or inventory script that runs in dry-run mode first.

Do I need to know uv for a Python DevOps interview?

It helps. Many teams have moved to uv for fast installs and lockfiles, so knowing uv init, uv add, uv lock, uv sync and uv run shows current practice. Understanding virtual environments and pyproject.toml matters more than any single tool.

How should I prepare for a live Python scripting interview?

Practise small tasks end to end in a plain editor: parse a log, call an API with retries, list cloud resources with pagination, and write a test. Talk through edge cases, timeouts and failure handling as you code, because interviewers grade the reasoning as well as the result.

Is Python for DevOps a good career skill?

Yes. Python automation is useful across DevOps, cloud, SRE, platform engineering and security roles, and it is the base for newer work such as AI-assisted operations. It is a skill that keeps paying off as you move between roles.

Do DevOps interviews include AI questions now?

Some do, usually practical ones: how you would call an LLM API safely from a script, validate its output and keep a human in the loop for risky actions. Deep machine learning theory is not expected for DevOps roles.

Ready to turn these answers into working skills? Cloudsoft's multi-cloud DevOps training with Linux and Python takes you through Python automation, Linux, AWS, Azure, CI/CD and Kubernetes with hands-on labs, in our Ameerpet classroom or live online. If you also want AI, ML and cloud security in one track, look at the APEX AI, ML, Cloud & Cyber Security program. Book a free demo on +91 96660 19191.

New ยท AI Career Guide

Meet Aanya โ€” ask anything about courses, fees & placement

Instant answers from verified Cloudsoft info โ€” courses, fees, formats, placement support and free demos. Available 24/7, right here on the site.

How Aanya works โ†’
Share๐•infโœ‰
EnrollWhatsAppCall us