These Docker interview questions cover what DevOps, cloud and platform interviews in 2026 actually probe: how images and containers work under the hood, how to write a Dockerfile that is small, fast to build and safe to run, and how to debug a container that misbehaves at 2 a.m. Interviewers rarely stop at "what is a container"; they ask you to explain layers, cgroups, BuildKit and image signing, then hand you a broken container and watch how you reason. The 60 questions below go from fundamentals to Compose, security, GPU containers for AI workloads and twelve real-world scenarios.
How to use this guide:
- Freshers and juniors are usually tested on the fundamentals and Dockerfile sections: image vs container, layers, CMD vs ENTRYPOINT, volumes, basic networking. Be able to type a working Dockerfile from memory.
- Mid-level DevOps engineers get the BuildKit, Compose, resource limits, registries and CI/CD questions, plus at least one debugging scenario.
- Senior and platform roles are pushed on security (rootless, capabilities, signing, supply chain), multi-platform builds, the Docker/containerd/Kubernetes relationship and GPU containers.
- For every answer, practise saying the one-line answer first, then the mechanism, then a trade-off. That structure is what separates a confident Docker for DevOps interview answer from a recited definition.
Contents
- Docker fundamentals (Q1βQ10)
- Dockerfile interview questions and BuildKit (Q11βQ21)
- Networking, volumes and storage (Q22βQ26)
- Docker Compose, limits, healthchecks and logging (Q27βQ32)
- Security, registries and tagging (Q33βQ39)
- Containers for AI and GPUs (Q40βQ43)
- Ecosystem, CI/CD and debugging (Q44βQ48)
- Real-world scenario questions (Q49βQ60)
- Key takeaways
- Interview preparation checklist
- FAQ
Docker fundamentals
1. What is Docker and why do teams use it?
Answer: Docker is a platform for building, shipping and running applications as containers: an application plus its runtime, libraries and configuration packaged into an image that runs the same way on a laptop, a CI runner and a production server. Teams use it for three reasons. Consistency: the image that passed tests is byte-for-byte the image that is deployed, which removes "works on my machine" drift. Isolation: each service gets its own filesystem, process tree and network stack without the overhead of a full virtual machine. Speed: containers start in seconds and images are layered, so rebuilding and shipping a change moves only what changed.
Interview tip: Separate the parts. "Docker" can mean the CLI, the Docker Engine daemon, Docker Desktop, Docker Hub or the image format. Naming which one you mean signals real experience.
2. What is the difference between an image and a container?
Answer: An image is an immutable, read-only template made of stacked filesystem layers plus metadata (default command, environment, exposed ports, user). A container is a running (or stopped) instance of an image: the runtime adds a thin writable layer on top, plus namespaces, cgroups and a network configuration. You can start many containers from one image; each gets its own writable layer, and anything written there disappears when the container is removed unless it went to a volume.
A useful analogy is class and object, but the more precise statement is: image = layers + config, container = image + writable layer + process + isolation settings.
3. What are image layers and why do they matter?
Answer: Each filesystem-changing Dockerfile instruction (RUN, COPY, ADD) produces a layer: a tarball of the files added, changed or deleted relative to the layer below. Layers are content-addressed by digest, so identical layers are stored once and shared between images and pulls. They matter for three things:
- Build cache: if an instruction and its inputs have not changed, the cached layer is reused. Order instructions from least to most frequently changed.
- Size: deleting a file in a later layer does not shrink the image; the bytes still live in the earlier layer. Clean up in the same
RUNthat created the mess. - Security: anything ever written to a layer, including a secret you "removed" later, can be extracted with
docker historyor by unpacking the image.
4. How are containers different from virtual machines?
Answer: A VM virtualises hardware and runs a full guest kernel on a hypervisor; a container is an isolated process sharing the host kernel. Containers are therefore lighter (megabytes, not gigabytes; seconds, not minutes) but the isolation boundary is weaker: a kernel vulnerability affects every container on the host. You also cannot run a different kernel in a container, which is why Docker Desktop on macOS and Windows runs a lightweight Linux VM behind the scenes.
| Aspect | Container | Virtual machine |
|---|---|---|
| Kernel | Shared host kernel | Own guest kernel |
| Isolation | Namespaces, cgroups, seccomp, capabilities | Hypervisor boundary |
| Start time | Seconds or less | Tens of seconds to minutes |
| Typical use | Microservices, CI jobs, packaging | Multi-tenant isolation, different OSes |
Interview tip: Mention that sandboxed runtimes such as gVisor or Kata Containers exist for when container isolation is not strong enough, for example running untrusted code.
5. What are namespaces and cgroups, and what does each do for a container?
Answer: They are the two Linux kernel features containers are built from. Namespaces control what a process can see: PID (its own process tree, with the app as PID 1), network (its own interfaces, routes and ports), mount (its own filesystem view), UTS (hostname), IPC, user (UID mapping) and cgroup namespaces. Control groups (cgroups) control what a process can use: memory, CPU shares and quotas, block I/O and the number of processes. Modern distributions use cgroup v2, a single unified hierarchy that Docker supports.
Seccomp profiles, Linux capabilities and AppArmor/SELinux add a third layer: what a process is allowed to do even inside its namespaces.
Real-world example: A memory leak in one container hits its cgroup memory limit and that container is OOM-killed, while neighbouring containers on the same host keep running. Without the limit, the host's OOM killer could pick any process.
6. Walk through Docker's architecture. What happens between the CLI and a running process?
Answer: The docker CLI is a client that calls the Docker Engine API (over a Unix socket by default). The daemon, dockerd, handles images, networks, volumes and builds, and delegates container lifecycle to containerd. containerd pulls and unpacks images, then uses a shim and an OCI runtime, normally runc, to create the namespaces and cgroups and start the process.
docker CLI
| REST over /var/run/docker.sock
dockerd (images, networks, volumes, builds)
| gRPC
containerd (image store, container lifecycle)
|
containerd-shim --> runc --> your process
The shim keeps the container running even if containerd or dockerd restarts. The Open Container Initiative (OCI) standardises both the image format and the runtime spec, which is why an image built by Docker runs under containerd, CRI-O or Podman.
7. What are registries, repositories, tags and digests?
Answer: A registry is a server that stores and serves images (Docker Hub, Amazon ECR, Azure Container Registry, Google Artifact Registry, GitHub Container Registry, Harbor). A repository is a named collection of related images inside it, such as myorg/payments-api. A tag is a human-readable, mutable pointer, such as :1.4.2 or :latest. A digest (@sha256:...) is the content hash of the image manifest and is immutable: the same digest always means the same bytes. A multi-platform image has a manifest list (image index) whose digest points to per-architecture manifests.
Interview tip: Say that latest is just a tag name with no special meaning beyond being the default when none is given; it is not "the newest image".
8. What is the difference between CMD and ENTRYPOINT?
Answer: ENTRYPOINT defines the executable the container runs; CMD provides default arguments (or a default command if no entrypoint is set). Arguments after the image name in docker run replace CMD, while replacing ENTRYPOINT needs the explicit --entrypoint flag. A common pattern is ENTRYPOINT ["python", "-m", "app"] with CMD ["--port", "8000"], so users can override flags without retyping the program. Another is an entrypoint script that runs migrations or templating and ends with exec "$@" so the real process replaces the shell and receives signals.
9. COPY vs ADD: when would you use each?
Answer: Use COPY by default; it copies files from the build context into the image and nothing else. ADD also auto-extracts local tar archives and can fetch remote URLs (including Git repositories in newer Dockerfile syntax). That extra behaviour is implicit and easy to misuse, so good practice is to use ADD only when you specifically want local tar extraction or a checksummed remote fetch, and COPY otherwise. COPY --from=<stage> is how multi-stage builds move artefacts between stages, and COPY --chown sets ownership without an extra chown layer.
10. What does a restart policy do, and which would you choose?
Answer: A restart policy tells the daemon what to do when a container's main process exits: no (default), on-failure[:max-retries] (restart only on non-zero exit), always (restart regardless, and on daemon start), and unless-stopped (like always, but respects a manual stop). For a long-running service on a single host, unless-stopped is the usual choice; for batch jobs, on-failure with a retry cap. Note that restart policies react to process exit, not to a failing healthcheck: plain Docker marks an unhealthy container but does not restart it. Orchestrators (Kubernetes, Swarm, ECS) act on health.
Dockerfile interview questions and BuildKit
11. What is a multi-stage build and why use it?
Answer: A multi-stage build uses several FROM stages in one Dockerfile: early stages carry compilers, package managers and test tooling; the final stage starts from a slim runtime base and copies in only the built artefacts with COPY --from. The result is a much smaller image with a smaller attack surface, because build tools, source code and caches never reach production. BuildKit also skips stages the target does not depend on and builds independent stages in parallel.
# syntax=docker/dockerfile:1
FROM python:3-slim AS build
WORKDIR /src
COPY requirements.txt .
RUN pip wheel -r requirements.txt -w /wheels
FROM python:3-slim
RUN useradd -r -u 10001 app
COPY --from=build /wheels /wheels
RUN pip install --no-cache-dir /wheels/* \
&& rm -rf /wheels
COPY --chown=app:app src/ /app/
USER app
CMD ["python", "-m", "app"]
Interview tip: Mention --target to build a specific stage, for example a test stage in CI that is never shipped.
12. How does the build cache work, and how do you order instructions to use it well?
Answer: For each instruction the builder checks whether it has a cached result with the same parent layer and the same inputs: the instruction text for RUN, and file contents (checksums) for COPY/ADD. The first miss invalidates every instruction after it. So put rarely changing steps first: base image, OS packages, then the dependency manifest (package.json/lockfile, requirements.txt, go.mod), then the dependency install, and only then COPY . . for the source. A one-line code change then reuses the expensive dependency layer.
Two gotchas: RUN apt-get update on its own line gets cached forever and goes stale, so combine it with the install; and a changing ARG used early invalidates everything below its first use.
13. What does .dockerignore do and what belongs in it?
Answer: .dockerignore excludes paths from the build context sent to the builder. It speeds builds (less to transfer), stops unnecessary cache invalidation (an edited README no longer busts COPY . .), and prevents leaks. Typical entries: .git, node_modules, virtualenvs, __pycache__, build output, test reports, local .env files, credentials, IDE folders and large datasets or model files. Some teams invert it: ignore everything with * and re-include only what the build needs with ! patterns.
14. How do you reduce image size?
Answer: Work from the biggest wins down:
- Choose a smaller base:
-slimvariants, Alpine, distroless or a vendor-hardened minimal image instead of a full OS image. - Use multi-stage builds so compilers, headers and source stay out of the final stage.
- Install only runtime dependencies (
--no-install-recommends, production-only npm installs) and clean package caches in the sameRUN, or use BuildKit cache mounts so caches never enter a layer. - Add a tight
.dockerignore. - Keep large data and model weights out of the image entirely (Q42).
Then measure: docker image ls for total size, docker history for per-layer size, and a layer explorer such as dive to find the culprit file. Scenario Q49 walks through this end to end.
15. Alpine, slim, distroless or scratch: how do you choose a base image?
Answer: It is a trade-off between size, compatibility and debuggability.
| Base | Strength | Watch out for |
|---|---|---|
| Debian/Ubuntu slim | glibc, broad compatibility, familiar tools | Larger; more packages to patch |
| Alpine | Very small, has a shell and apk | musl libc: some Python wheels and native libraries behave differently or must compile |
| Distroless | Runtime only, no shell or package manager, few CVEs | Harder to debug; needs ephemeral debug containers |
| scratch | Empty; ideal for static Go or Rust binaries | You must add CA certificates, timezone data and a non-root user yourself |
Interview tip: For Python ML stacks, say you would avoid Alpine because many prebuilt wheels target glibc, which turns a two-minute build into a long compile. That one detail shows hands-on experience.
16. Why should containers run as a non-root user, and how do you do it?
Answer: By default the process in a container runs as UID 0. It is namespaced and capability-reduced, but it is still root on files it can reach: a mounted host directory, a container escape bug or a writable Docker socket turns that into host root. Running as non-root limits the blast radius. In the Dockerfile, create a user with a fixed UID (RUN useradd -r -u 10001 app), give it ownership of only what it must write (COPY --chown), and switch with USER 10001. Using a numeric UID lets Kubernetes' runAsNonRoot check verify it. Bind to ports above 1024, or grant only NET_BIND_SERVICE. At runtime, docker run --user can override it.
17. What is the difference between shell form and exec form, and why does PID 1 matter?
Answer: Exec form (CMD ["node", "server.js"]) runs the program directly as PID 1. Shell form (CMD node server.js) wraps it in /bin/sh -c, so the shell becomes PID 1. That matters because docker stop sends SIGTERM to PID 1, waits a grace period (10 seconds by default) and then sends SIGKILL. Many shells do not forward SIGTERM to the child, so the app never gets a chance to drain connections and is killed hard. PID 1 also has special signal semantics (no default handlers) and must reap zombie child processes. Fixes: use exec form, end wrapper scripts with exec, and add a tiny init with docker run --init (or tini in the image) for apps that spawn children.
18. ARG vs ENV: what is the difference?
Answer: ARG is a build-time variable, set with --build-arg, visible only during the build (and scoped to the stage where it is declared; an ARG before FROM is only usable in FROM lines unless redeclared). ENV sets an environment variable that persists in the image and in every container started from it. Neither is a safe place for secrets: build args are recorded in image metadata and history, and ENV values are visible to anyone who can inspect the image. Use BuildKit secrets at build time and runtime secret injection instead (Q20, Q37).
19. What are BuildKit cache mounts and when do they help?
Answer: A cache mount gives a RUN step a persistent directory that survives between builds but is never committed into the image layer. It is ideal for package manager caches:
RUN --mount=type=cache,target=/root/.cache/pip \
pip install -r requirements.txt
When the requirements change, the layer cache misses, but pip still finds previously downloaded wheels in the cache mount, so the reinstall is much faster. The same works for apt, npm, Go modules and Maven. BuildKit has been the default builder for Docker Engine on Linux since version 23.0 and in Docker Desktop. One caveat: cache mounts live on the builder, so ephemeral CI runners lose them unless you persist the builder or export cache (Q47).
20. How do you use secrets during a build without leaking them into the image?
Answer: Use BuildKit secret mounts. The secret is mounted as a file (or exposed as an environment variable) only for the duration of that RUN step and is not written to any layer or the build history:
RUN --mount=type=secret,id=npmrc,target=/root/.npmrc \
npm ci
# docker build --secret id=npmrc,src=$HOME/.npmrc .
For private Git dependencies, RUN --mount=type=ssh forwards your SSH agent instead of copying a key. The anti-patterns interviewers want you to name are COPY id_rsa followed by rm (still in the earlier layer), passing tokens via --build-arg (visible in history) and baking .env files into the image.
21. How do multi-platform builds work?
Answer: docker buildx build --platform linux/amd64,linux/arm64 -t repo/app:1.2.0 --push . builds one image per architecture and pushes them under a single tag backed by a manifest list; each client pulls the variant matching its CPU. Buildx can produce the non-native variants three ways: QEMU emulation (simplest, slow for heavy compiles), native builder nodes per architecture (fast, more infrastructure), or cross-compilation inside the Dockerfile using the automatic BUILDPLATFORM and TARGETPLATFORM/TARGETARCH args, which works especially well for Go and Rust.
Real-world example: A team moving workloads to Arm-based cloud instances, or shipping models to Arm edge devices, needs arm64 images; the same skills show up in edge AI interview questions. The classic failure is an "exec format error", meaning an amd64-only image was pulled onto an arm64 host.
Networking, volumes and storage
22. How do containers communicate, and what network drivers does Docker offer?
Answer: Containers on the same user-defined network reach each other by container or service name, because Docker runs an embedded DNS server for those networks. The default bridge network does not provide name resolution, which is why you should always create your own network (Compose does this automatically per project). Drivers:
- bridge: private network on one host with NAT to the outside; the default for single-host apps.
- host: the container shares the host's network namespace; no port mapping, no isolation, minimal overhead.
- none: loopback only, for jobs that should have no network.
- overlay: spans multiple hosts (Swarm).
- macvlan/ipvlan: containers get addresses on the physical network, used for legacy integrations.
23. What does EXPOSE do compared with -p?
Answer: EXPOSE is documentation and metadata: it declares which port the app listens on, but publishes nothing. -p 8080:80 (or ports: in Compose) actually publishes container port 80 on host port 8080 by setting up forwarding rules; -P publishes all exposed ports to random high ports. By default -p 8080:80 binds on all host interfaces, which on a cloud VM can expose a database to the internet; bind to loopback with -p 127.0.0.1:5432:5432 when only local access is needed. Also remember that Docker's published-port rules can take precedence over host firewall front-ends such as ufw, a common surprise.
24. Volumes vs bind mounts vs tmpfs: when do you use each?
Answer: Volumes are managed by Docker, live under Docker's data directory, survive container removal, can use volume drivers for remote storage, and are the preferred choice for persistent data such as databases. Bind mounts map an exact host path into the container; they are ideal in development (live code reload) and for injecting host config, but they couple the container to the host's directory layout and permissions. tmpfs mounts are in-memory only and vanish when the container stops, which suits scratch space and sensitive temporary files. Prefer the explicit --mount type=...,source=...,target=... syntax in scripts because it fails loudly on typos instead of silently creating an empty directory.
25. How does the container filesystem work (storage drivers, copy-on-write)?
Answer: Docker uses a union filesystem, normally overlay2 on Linux, to present the read-only image layers plus the container's writable layer as a single directory tree. When a process modifies a file from a lower layer, the file is first copied up into the writable layer (copy-on-write). Consequences: writing large or frequently changing files inside the container filesystem is slow and grows the container; deleting a lower-layer file only hides it with a whiteout. So databases, logs and caches belong on volumes, and you should treat the container filesystem as disposable. Newer Docker releases can also use the containerd image store, which changes how images are stored and enables features such as multi-platform images locally.
26. How does a container reach a service on the host, and when would you use host networking?
Answer: Inside a container, localhost is the container itself. To reach the host, Docker Desktop provides the name host.docker.internal; on Linux Engine you can add it with --add-host=host.docker.internal:host-gateway, or use the bridge gateway address. --network host removes the network namespace so the container uses the host's interfaces directly: useful for very high packet rates, some monitoring agents or tools that need to see host interfaces, but you lose port isolation and can hit port clashes. Host networking behaves differently on Docker Desktop because the "host" is the Desktop VM, so check current documentation before relying on it there.
Docker Compose, limits, healthchecks and logging
27. What is Docker Compose and what changed with Compose v2?
Answer: Compose defines a multi-container application (services, networks, volumes, secrets, configs) in a YAML file and runs it with one command. Compose v2 is a Docker CLI plugin invoked as docker compose (with a space); the old Python-based docker-compose v1 stopped receiving updates in July 2023. v2 follows the Compose Specification, so the top-level version: field is obsolete and only produces a warning. Everyday commands: docker compose up -d, ps, logs -f, exec, down (add -v to remove named volumes), config to render the merged file, and watch for syncing code changes in development.
Interview tip: Saying "Compose is for local development and simple single-host deployments; for multi-host production I would use Kubernetes or a managed container service" shows you know where the tool's boundary is.
28. Does depends_on make a service wait for its dependency to be ready?
Answer: Plain depends_on controls only start order; it does not wait for the database to accept connections. To wait for readiness, give the dependency a healthcheck and use the long syntax:
services:
api:
build: .
depends_on:
db:
condition: service_healthy
db:
image: postgres:16
healthcheck:
test: ["CMD-SHELL", "pg_isready -U app"]
interval: 5s
retries: 10
Other conditions are service_started and service_completed_successfully (useful for a migration job). Even so, the application should still retry connections with backoff, because dependencies restart in production too.
29. How do you manage environments with Compose: overrides, profiles and variables?
Answer: Keep a base compose.yaml and layer environment-specific files: compose.override.yaml is merged automatically (handy for dev bind mounts and debug ports), and others are combined explicitly with -f compose.yaml -f compose.ci.yaml. profiles: mark optional services (an admin UI, a local Kafka broker, a mock SMTP server) that start only with --profile. Variable interpolation (${TAG:-dev}) reads from the shell and a .env file in the project directory; env_file: passes variables into the container. Run docker compose config to see exactly what will run. If your stack includes streaming components, the Kafka interview questions guide covers the broker side.
30. How do you set resource limits on containers, and what happens when they are exceeded?
Answer: Use cgroup-backed flags: --memory (hard limit), --memory-reservation (soft), --cpus (CPU quota, for example 1.5), --cpu-shares (relative weight under contention), --pids-limit (stops fork bombs), and in Compose the equivalents under deploy.resources or service-level keys. Exceeding memory triggers the kernel OOM killer inside the cgroup: the container dies with exit code 137 and docker inspect shows OOMKilled: true. Exceeding CPU only throttles, so the symptom is latency, not a crash. Runtimes such as the JVM, Node and Python worker pools should be configured to respect the container limit rather than the host's total memory and CPU count.
31. What does HEALTHCHECK do and how do you write a good one?
Answer: HEALTHCHECK runs a command inside the container on an interval; after a configured number of consecutive failures the container's status becomes unhealthy. Options: --interval, --timeout, --retries, --start-period (grace time during boot) and --start-interval. A good check is cheap, tests the app's own readiness endpoint, and does not depend on downstream systems, otherwise a slow database marks every API container unhealthy at once. Remember the probe tool must exist in the image; distroless images have no curl, so use a small built-in health subcommand or a static probe binary. Kubernetes ignores Docker's HEALTHCHECK and uses its own liveness, readiness and startup probes.
32. How does container logging work, and how do you stop logs filling the disk?
Answer: Docker captures the main process's stdout and stderr and hands them to a logging driver. The default json-file driver writes to files on the host with no rotation unless configured, which is a classic cause of full disks. Set rotation per container (--log-opt max-size=10m --log-opt max-file=3) or globally in daemon.json, or switch to the local driver, which rotates and compresses by default. Other drivers ship logs to journald, syslog, Fluentd, Splunk, Amazon CloudWatch Logs or Google Cloud Logging. Good practice for the app: log to stdout/stderr in structured JSON, never to files inside the container, and let the platform collect.
If you want to practise these topics on real pipelines rather than only reading about them, Cloudsoft's DevOps training in Hyderabad builds Dockerfiles, Compose stacks and CI pipelines in hands-on labs, in the Ameerpet classroom or live online.
Security, registries and tagging
33. How do you scan images for vulnerabilities, and what is an SBOM?
Answer: A scanner (Trivy, Grype, Docker Scout, or the scanner built into your registry such as Amazon ECR or Harbor) inventories the OS packages and language libraries in an image and matches them against vulnerability databases. A Software Bill of Materials (SBOM), in SPDX or CycloneDX format, is that inventory as a document you can store and re-scan later when new CVEs are published. BuildKit can attach SBOM and provenance attestations at build time (docker buildx build --sbom=true --provenance=true). Scan in CI as a gate (fail on fixable critical and high findings, with an exceptions process), and re-scan images already in the registry continuously, because an image that was clean on Monday can be vulnerable by Friday.
Interview tip: Emphasise triage over volume: base image updates fix most findings, and "not reachable" or "no fix available" findings need documented risk acceptance, not silent suppression. The DevSecOps interview questions guide goes deeper on pipeline gates.
34. How do image signing and provenance work?
Answer: Signing proves who produced an image and that it has not changed since. The common modern tools are Sigstore cosign (key-based or keyless signing using short-lived certificates tied to a CI identity, such as a GitHub Actions OIDC token, recorded in a transparency log) and Notation from the Notary Project. Signatures and attestations are stored in the registry next to the image, referenced by digest. Provenance attestations, often in the SLSA format, record how the image was built: source repository, commit, builder and parameters. The enforcement half is what matters: an admission controller in Kubernetes (for example Kyverno or a Sigstore policy controller) or a deploy-time check refuses images that are unsigned or not built by your pipeline. Docker Content Trust, based on the original Notary, is legacy; check current Docker documentation for its status.
35. What is rootless Docker and how is it different from userns-remap?
Answer: In rootless mode both the Docker daemon and the containers run as an unprivileged user, using user namespaces and subordinate UID/GID ranges (/etc/subuid, /etc/subgid). A daemon or runtime escape then lands as that ordinary user, not host root. With userns-remap, the daemon still runs as root but container root is mapped to an unprivileged host UID range. Rootless has trade-offs: binding to ports below 1024 needs extra configuration, some storage and networking features differ or are slower, and certain cgroup limits need cgroup v2 with delegation. Podman runs rootless by default, which is one reason security-focused teams like it (Q44).
36. How do Linux capabilities, seccomp and read-only filesystems harden a container?
Answer: Root's powers are split into capabilities; Docker grants a reduced default set. Hardening means dropping all and adding back only what is needed (--cap-drop=ALL --cap-add=NET_BIND_SERVICE). Seccomp filters system calls; Docker applies a default profile that blocks dangerous syscalls, and you should never run with seccomp=unconfined casually. --security-opt no-new-privileges stops setuid binaries from gaining privileges. --read-only makes the root filesystem immutable, with tmpfs for the paths that must be writable. And --privileged disables almost all of this, granting all capabilities and device access, so treat it as a red flag in any review.
Real-world example: Consider a bank's platform team in a Hyderabad GCC. Its baseline policy might be non-root, read-only root filesystem, all capabilities dropped, default seccomp, no privileged containers and no host namespaces, enforced by admission policy in Kubernetes, with exceptions reviewed by security.
37. How should applications in containers receive secrets at runtime?
Answer: Inject secrets at runtime from a secrets manager, never bake them into images. Options, from weakest to strongest: environment variables (simple, but visible in docker inspect, crash dumps and child processes); files mounted into the container (Compose and Swarm secrets: appear under /run/secrets/; Kubernetes Secrets as volumes); and fetching from a manager such as AWS Secrets Manager, Azure Key Vault, Google Secret Manager or HashiCorp Vault using the workload's identity (an IAM role or managed identity) so no long-lived credential exists at all. Rotate secrets, scope them per service, and make sure they never appear in logs.
38. What is a good tagging strategy, and why do immutable tags and digests matter?
Answer: Tag every build with something traceable and unique, such as the Git commit SHA, plus a semantic version for releases (1.4.2, optionally 1.4 and 1 as moving convenience tags). Avoid deploying latest: you cannot tell what is running, rollbacks are guesswork, and two nodes may pull different images under the same tag. Turn on tag immutability in the registry (Amazon ECR, Azure Container Registry and others support it) so a release tag cannot be overwritten, and pin deployments to digests (app@sha256:...) for full reproducibility. Promote the same digest from dev to staging to production rather than rebuilding per environment. Pin base images by digest too, and let a dependency bot propose updates.
39. Why is mounting the Docker socket into a container dangerous?
Answer: -v /var/run/docker.sock:/var/run/docker.sock gives the container full control of the Docker daemon. Anyone with that access can start a privileged container that mounts the host's root filesystem, which is effectively root on the host. Being in the docker group carries the same power. CI systems and monitoring tools often ask for the socket; safer alternatives are rootless BuildKit or daemonless builders for CI, a socket proxy that allows only read-only API calls for monitoring, or a dedicated, isolated build VM. If the socket must be mounted, treat that container as part of the host's trusted computing base.
Containers for AI and GPUs
40. How do GPU containers work with Docker?
Answer: The container does not contain the GPU driver; the host does. The host needs the NVIDIA kernel driver plus the NVIDIA Container Toolkit, which configures the container runtime to inject the GPU device nodes and the matching user-space driver libraries into the container at start. Setup is roughly: install the toolkit, run sudo nvidia-ctk runtime configure --runtime=docker, restart Docker, then docker run --gpus all ... (or --gpus '"device=0"' for a specific GPU). The image supplies the CUDA runtime and libraries (for example a CUDA base or a framework image). In Compose you reserve GPUs under deploy.resources.reservations.devices with driver: nvidia and capabilities: [gpu]. The toolkit also supports the Container Device Interface (CDI), which NVIDIA recommends for Podman.
Interview tip: The key compatibility rule: the CUDA version in the image must be supported by the host driver. Newer drivers run older CUDA images, not the reverse. For more on how GPUs work, see GPU basics for AI engineers.
41. AI images are often many gigabytes. How do you keep them manageable?
Answer: Start from the smallest image that has what you need: a CUDA runtime image rather than the devel image (compilers and headers belong in a build stage), and only the framework you use. Install exactly pinned wheels without caches, and avoid pulling in unrelated extras. Split CPU and GPU variants so a CPU-only API service does not carry GPU libraries. Order layers so the large, stable framework layer is cached and shared between images and only the small application layer changes per release. Pre-pull or cache large base layers on GPU nodes to avoid slow cold starts, and keep the registry in the same region as the cluster. For a deeper walkthrough, see Docker for AI applications.
42. Should model weights go inside the image?
Answer: Usually not. Baking multi-gigabyte weights into the image couples model releases to code releases, multiplies registry storage for every tag, makes pulls and scale-out slow, and often violates licence or data-handling controls. Common patterns instead: download weights at start from object storage or a model registry into a mounted volume, with a checksum check; mount a shared read-only volume or network filesystem on the node; use an init container in Kubernetes to fetch them before the server starts; or store weights as a separate OCI artifact. Baking them in is reasonable for small models, air-gapped environments or when you need one fully immutable artefact for audit. Whatever you choose, version the model separately and record which model version each container served.
Real-world example: Consider a hospital running a self-hosted summarisation model. Keeping weights in a controlled bucket with access logging lets security audit model access, and swapping the model version becomes a config change rather than a rebuild. See self-hosting LLMs for the wider design.
43. How would you structure a local Compose stack for an AI application?
Answer: One service per concern: the API (for example FastAPI), a worker for ingestion and embedding jobs, PostgreSQL with pgvector or another vector store, a cache or queue, and optionally a local model server with a GPU reservation. Healthchecks plus condition: service_healthy control ordering, named volumes persist the database and model cache, a profile keeps the GPU model server optional so developers without GPUs can point at a hosted model API instead, and secrets come from a local, git-ignored file. The same images then move to Kubernetes in production; the Kubernetes for AI interview questions cover that next step, and LLM inference and serving interview questions cover model servers.
Ecosystem, CI/CD and debugging
44. Docker vs Podman vs containerd: how do they compare?
Answer: All three run OCI containers from OCI images, so images are portable between them.
| Tool | Model | Typical use |
|---|---|---|
| Docker Engine | Client plus daemon (dockerd), uses containerd and runc | Developer workflow, builds, single-host services, Compose |
| Podman | Daemonless, rootless by default, Docker-compatible CLI, pods | Security-sensitive hosts, RHEL-family systems, systemd integration |
| containerd | Lower-level runtime daemon; CLIs such as ctr, nerdctl and crictl | Kubernetes nodes and platforms embedding a runtime |
| CRI-O | Runtime built only for the Kubernetes CRI | Kubernetes nodes, including OpenShift |
Daemonless builders such as Buildah, and BuildKit running rootless, are common in CI where you do not want a privileged Docker daemon.
45. Kubernetes removed dockershim. Does that mean Docker images no longer work on Kubernetes?
Answer: No. Kubernetes talks to runtimes through the Container Runtime Interface (CRI). Docker Engine never implemented CRI, so the kubelet carried an adapter called dockershim. It was deprecated in v1.20 and removed in Kubernetes v1.24. Nodes now run a CRI runtime directly, typically containerd or CRI-O. Images built with Docker are standard OCI images and run unchanged. What broke was anything that depended on the Docker daemon on the node: mounting docker.sock into pods, Docker-in-Docker builds on nodes, or monitoring agents that queried dockerd. Teams that need Docker Engine as the Kubernetes runtime can use the separately maintained cri-dockerd adapter. For more cluster questions, see the Kubernetes interview questions.
46. What should an engineer know about Docker Desktop licensing?
Answer: Docker Desktop (the desktop application for macOS, Windows and Linux) is free for personal use, education, non-commercial open-source projects and small businesses below Docker's published employee and revenue thresholds; larger organisations and government use need a paid subscription (Pro, Team or Business). The open-source Docker Engine and Moby projects are licensed separately and remain free to use, so a Linux server running Docker Engine is not affected. Terms change over time, so in an interview say you would check the current Docker subscription terms and your company's licence position, and mention alternatives some teams use on laptops (Podman Desktop, Rancher Desktop, Colima, or Docker Engine in a Linux VM). Do not quote figures you have not checked.
47. How does Docker fit into a CI/CD pipeline?
Answer: A typical pipeline on a pull request and on merge:
commit -> lint Dockerfile (hadolint)
-> buildx build (cache from registry)
-> unit tests in --target test stage
-> scan + SBOM -> fail on policy
-> push :<git-sha> -> sign by digest
-> deploy digest to dev -> promote
Key practices: build once and promote the same digest through environments; tag with the commit SHA; persist build cache across ephemeral runners with --cache-from/--cache-to (registry cache, or the GitHub Actions cache backend type=gha); authenticate to the registry with short-lived OIDC credentials instead of stored passwords; and run integration tests with docker compose up --wait against real dependencies. For AI-specific stages such as evaluation gates, see CI/CD for AI applications; for a Jenkins-centred angle, see these Docker interview questions with a CI/CD focus.
48. What is your toolkit for debugging a misbehaving container?
Answer: Work from outside in:
docker ps -afor status and exit code;docker logs --tail 100 -ffor output.docker inspectfor the effective command, env, mounts, networks,State.ExitCode,State.OOMKilledand health log.docker exec -it <c> shto look inside a running container (if it has a shell).docker statsfor live CPU and memory;docker eventsfor restarts and kills;docker topfor processes.- For slim or distroless images, attach a toolbox container to the same namespaces:
docker run --rm -it --network container:<c> --pid container:<c> <toolbox-image>. - For a container that dies at start, run the image with
--entrypoint shand execute the command by hand;docker cppulls files out of a stopped container.
Host-side skills matter just as much; Linux for AI engineers covers the processes, permissions and networking commands you will lean on.
Real-world scenario questions
49. Your image is 3 GB and slow to pull. How do you shrink it?
Answer: Measure first, then remove the biggest contributors: usually the base image, build tooling left in the final stage, package caches, and data or model files that should never have been copied in.
What I would check:
docker history --no-truncand a layer explorer to find which layers and files are large.- The build context and
.dockerignore: isCOPY . .pulling in.git, datasets, virtualenvs or model checkpoints? - The base: full OS or CUDA devel image where slim or runtime would do?
- Whether compilers and dev headers are in the final stage; move them into a build stage.
- Cleanups done in a separate
RUN(they do not reclaim space); merge them or use cache mounts. - Large weights or reference data that can move to object storage or a volume.
Production consideration: Add an image-size budget check in CI so regressions are caught at pull request time, and make sure the slimmer image still passes integration tests: removing a "useless" system library is a classic way to break TLS or timezone handling.
50. A container exits immediately after docker run. How do you troubleshoot it?
Answer: A container lives exactly as long as its main process. It exits immediately either because that process crashed or because it finished: it was a short command, or the service daemonised itself into the background so PID 1 returned.
What I would check:
docker ps -afor the exit code: 0 means it completed; 1 is an app error; 125 means docker run itself failed; 126 means not executable; 127 means command not found; 137 means SIGKILL, often OOM.docker logs <c>for the stack trace or missing-config message.docker inspectfor the effective entrypoint and command, especially if CMD was overridden.- Whether the service runs in the foreground (for example nginx with
daemon off, not a startup script that launches a background process and exits). - Missing environment variables, secrets or mounted files the app needs at start.
- Line endings or architecture: a script saved with Windows CRLF line endings, or an "exec format error" from the wrong platform.
- Run interactively with
--entrypoint shand start the command manually.
Production consideration: Do not "fix" it with tail -f /dev/null; that hides the real failure. Make the app fail fast with a clear log line and let the restart policy or orchestrator handle it.
51. docker run fails with "port is already allocated". What do you do?
Answer: Another process or container already owns that host port on that interface. Find it, then either stop it or publish on a different host port; the container port can stay the same.
What I would check:
docker ps --filter publish=8080to see if another container holds it, including one from a different Compose project.sudo ss -ltnp | grep 8080(orlsof -i :8080) for host processes.- Stopped-but-not-removed containers and leftover Compose stacks:
docker compose ls, thendocker compose downin the right project. - Whether you actually need a host port: services talking on a shared Docker network do not need
ports:at all. - Then remap, for example
-p 8081:8080, or parameterise it in Compose as"${API_PORT:-8080}:8080".
Production consideration: On shared CI runners, hardcoded host ports make parallel jobs collide. Let Docker assign random host ports (-p 8080 or -P) and discover them with docker port, or run tests on the Docker network with no published ports.
52. The build or tests work on your laptop but fail in CI. How do you approach it?
Answer: Treat it as an environment-difference hunt. The usual causes are architecture, cache, context and credentials.
What I would check:
- Architecture: an Apple silicon laptop builds arm64 by default; CI is often amd64. Build with an explicit
--platformand compare. - Cached state: the laptop has a warm cache with an old, working layer; CI builds fresh. Reproduce with
docker build --no-cache --pull. - Unpinned inputs: a floating base tag or unpinned dependency that changed. Pin versions and base digests.
- Build context: a file present locally but git-ignored (a generated config, a local
.env), or a.dockerignoredifference. - Credentials and network: private registry or package mirror access, proxies, rate limits on anonymous base image pulls.
- Runtime differences: CI runs as a different UID, has less memory, or uses rootless/daemonless builders; timing-sensitive tests that assume a dependency is instantly ready.
Production consideration: Make local and CI builds the same command, for example a Makefile or script calling docker buildx bake or docker compose, and run the same build in a clean container locally before blaming CI.
53. Your container cannot see the GPU. How do you debug it?
Answer: Check the stack bottom-up: hardware and driver on the host, the NVIDIA Container Toolkit and runtime config, the docker run flags, then the CUDA and framework versions inside the image.
What I would check:
nvidia-smion the host. If it fails, it is a driver or hardware problem, not Docker.- That the NVIDIA Container Toolkit is installed and configured (
nvidia-ctk runtime configure --runtime=docker) and that Docker was restarted afterwards. - The run command actually requests GPUs:
--gpus all, or the Compose device reservation. Running without it is the most common cause. docker run --rm --gpus all <cuda-base-image> nvidia-smias a known-good test, separating platform issues from image issues.- CUDA compatibility: the image's CUDA version needs a host driver new enough to support it.
- Inside the app: is it a CPU-only build of the framework? Is
CUDA_VISIBLE_DEVICESset to an empty or wrong value? - Rootless Docker, Podman or Kubernetes setups may need CDI or the GPU device plugin rather than the default path.
Production consideration: Bake the known-good test into node provisioning and alert when GPU nodes come up without visible devices, so a bad driver update does not surface as "the model is mysteriously slow" because it fell back to CPU.
54. A container keeps restarting with exit code 137. What is happening?
Answer: Exit code 137 is 128 + 9: the process received SIGKILL. The most common cause is the out-of-memory killer, followed by docker stop timing out or a manual kill.
What I would check:
docker inspect --format '{{.State.OOMKilled}}'and host kernel logs (dmesg) for OOM events.docker statsto watch memory climb before death: a leak (steady growth) or a spike (a large request, a batch load)?- Whether the limit is realistic for the workload, and whether the runtime respects it (JVM heap settings, worker counts per container).
- Whether the host itself ran out of memory, killing containers that had no limit.
Production consideration: Set memory limits from measured peaks plus headroom, alert on memory nearing the limit, and size worker concurrency to memory. Raising the limit without understanding the growth just postpones the crash.
55. In Compose, the API cannot connect to the database at localhost:5432. Why?
Answer: Inside the API container, localhost is the API container itself. The database is a different container, reachable by its service name on the Compose network, so the connection string should use db:5432.
What I would check:
- The connection host is the service name, not localhost or the host's IP.
- Both services are on the same network (
docker network inspect). - The database is actually ready: add a healthcheck and
condition: service_healthy, plus retries in the app. - The database listens on all interfaces inside its container, not just its own loopback.
- Use the container port (5432) from other containers, not a remapped host port.
Production consideration: Make the database host configurable through an environment variable so the same image connects to a container in development and to a managed database in production.
56. After switching the image to a non-root user, the app gets "permission denied" writing to a volume. How do you fix it?
Answer: The mounted directory is owned by root (or a different UID) on the host or in the volume, and the new UID cannot write to it. Fix ownership deliberately rather than going back to root.
What I would check:
- The UID the process runs as (
docker exec <c> id) versus the directory owner (ls -ln). - For named volumes: create the directory in the image with the right owner (
RUN mkdir /data && chown 10001 /data); a new, empty volume copies that ownership on first mount. - For bind mounts: align the host directory's owner or group with the container UID, or run with a matching
--user. - SELinux hosts: the mount may need a relabel option (
:zor:Z). - Whether the app writes somewhere it should not, such as its own code directory; point it at a dedicated data or tmpfs path.
Production consideration: In Kubernetes the equivalent tool is fsGroup in the pod security context. Standardise one UID per image so storage ownership is predictable.
57. A security review finds an API key in your image's history. What do you do?
Answer: Treat the key as compromised: anyone who pulled the image can extract it. Rotate first, then fix the build, then clean up.
What I would check:
- Revoke and rotate the key immediately, and review its usage logs for abuse.
- Find how it got in:
ARG/ENV, a copied.envor config file, or a file deleted in a later layer. - Fix the Dockerfile with BuildKit secret mounts for build-time needs and runtime injection for the app; add
.envand credential files to.dockerignore. - Delete or overwrite affected tags in every registry, including mirrors and caches, and rebuild.
- Add secret scanning of images and repositories to CI so it is caught before push.
Production consideration: Rewriting the image is not enough on its own, because copies may exist on nodes, in CI caches and in pulled developer environments. Rotation is the actual fix; cleanup reduces exposure.
58. Builds take 20 minutes because the dependency layer rebuilds on every commit. How do you fix it?
Answer: Something early in the Dockerfile changes on every commit and busts the cache, or the CI runner starts with no cache at all.
What I would check:
- Instruction order: is
COPY . .before the dependency install? Copy only the lockfile first, install, then copy source. - A changing build arg (commit SHA, build date) declared and used early; move it to the end or into a label.
- Missing
.dockerignore, so changes to.gitor test output invalidate the copy. - Ephemeral CI runners: export and import cache with
--cache-to/--cache-from(registry ortype=gha), usingmode=maxto include intermediate stages. - Add BuildKit cache mounts for package managers so even a cache miss downloads little.
Production consideration: Track build duration as a pipeline metric. Slow builds tempt teams to skip scanning or tests, which costs more than the minutes saved.
59. A build server's disk is full of Docker data. How do you recover and prevent it?
Answer: Docker accumulates images, stopped containers, unused volumes, build cache and logs. Find which category is large, clean it safely, then automate.
What I would check:
docker system df -vto see images, containers, volumes and build cache sizes.- Build cache:
docker builder prune(with a filter or a keep-storage budget). - Dangling and unused images:
docker image prune(add-acarefully on CI runners, not on hosts that run services). - Container logs under the Docker data directory: configure rotation (Q32).
- Volumes last and with care:
docker volume prunedeletes data that may matter.
Production consideration: Schedule pruning with retention filters on build hosts, set a BuildKit garbage-collection storage limit, monitor disk usage, and consider putting Docker's data directory on its own volume so a full disk cannot take down the OS.
60. Every deployment drops in-flight requests and docker stop takes 10 seconds. What is wrong?
Answer: The app is not receiving or handling SIGTERM, so it never drains; Docker waits the default grace period and then SIGKILLs it, cutting active requests.
What I would check:
- Shell-form
CMDor an entrypoint script withoutexec, leaving a shell as PID 1 that swallows the signal (Q17). - Whether the app has a SIGTERM handler that stops accepting new work, finishes in-flight requests and exits.
- A
STOPSIGNALmismatch: some servers expect a different signal for graceful shutdown. - Whether the grace period (
docker stop -t, Composestop_grace_period) is long enough for the real drain time. - The load balancer: it should stop sending traffic before the container stops (deregistration delay, readiness going false).
Production consideration: Test graceful shutdown explicitly by sending load while stopping a container. In Kubernetes the same issue appears as errors during rolling updates, solved with signal handling, a short preStop delay and readiness probes.
Key takeaways
- Explain containers from the kernel up: namespaces isolate what a process sees, cgroups limit what it uses, and images are content-addressed layers.
- Good Dockerfiles are ordered for caching, multi-stage, use minimal bases, run as a fixed non-root UID and use exec form so signals reach the app.
- BuildKit cache mounts, secret mounts and multi-platform builds are now everyday interview material, not advanced trivia.
- Security is layered: scanning and SBOMs, signing enforced at deploy, dropped capabilities, no privileged containers, no socket mounts, and runtime secrets.
- Deploy immutable digests, not
latest; build once and promote the same artefact. - For AI workloads, the host provides the GPU driver, the image provides CUDA, and model weights usually live outside the image.
- In scenarios, show a method: exit code, logs, inspect, reproduce, then fix and prevent.
Interview preparation checklist
- Write a multi-stage Dockerfile for a Python or Node service from memory, with a non-root user and a healthcheck.
- Shrink a real image and be ready to describe the before and after, layer by layer.
- Build a Compose stack with an API, a database with a healthcheck, a named volume and a profile for an optional service.
- Use a BuildKit secret mount and a cache mount once, and run a multi-platform build with buildx.
- Scan an image, generate an SBOM and sign the image by digest with cosign in a CI pipeline.
- Reproduce each failure mode once: exit 137, a port conflict, a localhost-in-Compose error and a permission-denied volume.
- If you can access a GPU machine, run a CUDA test container with
--gpus alland explain each layer of the stack. - Be able to explain the CLI, dockerd, containerd and runc chain, and what the dockershim removal changed.
- Prepare one story from your own work about a container problem you diagnosed, with the commands you used.
FAQ
How should I prepare for a Docker interview?
Combine concepts with hands-on practice. Build and shrink your own images, run a multi-service Compose stack, break things on purpose and debug them with logs, inspect and exec. Interviewers value a clear method and real commands over memorised definitions.
Is Docker still worth learning in 2026 if Kubernetes no longer uses it as a runtime?
Yes. Kubernetes stopped using Docker Engine as its node runtime, but Docker remains a common tool for building images, local development and CI. The images it builds are standard OCI images that Kubernetes runs directly, so Docker skills carry straight into Kubernetes work.
What Docker topics do freshers need for interviews?
Freshers should know images versus containers, layers and caching, basic Dockerfile instructions, CMD versus ENTRYPOINT, volumes, port publishing, container networking by name and the everyday CLI commands. Being able to write and run a simple Dockerfile live is often tested.
Which Docker topics matter most for DevOps interviews?
For DevOps roles, expect multi-stage builds, BuildKit caching, registries and tagging strategy, Docker in CI/CD pipelines, resource limits, logging, image scanning and signing, and debugging scenarios such as containers that exit immediately or fail only in CI.
Do I need Linux skills to do well in Docker interviews?
Yes. Containers are Linux processes, so understanding processes, signals, file permissions, users, networking and disk usage makes Docker debugging far easier. Many Docker scenario questions are really Linux questions in disguise.
Should I learn Podman as well as Docker?
Learn Docker first because it is the most common developer workflow, then understand how Podman differs: daemonless, rootless by default and Docker-compatible on the command line. Knowing the trade-offs is usually enough for interviews unless the role is on a Podman-based platform.
How important are Docker skills for AI and ML engineering roles?
Increasingly important. AI services are shipped as containers, and interviews for AI platform roles ask about GPU containers, image size with ML libraries, keeping model weights out of images and reproducible environments. The fundamentals are the same as for any service.
Is Docker Compose used in production?
Compose is mainly used for local development, testing and simple single-host deployments. Multi-host production systems usually run on Kubernetes or a managed container service, but Compose files remain a common way to describe and test the same services.
What projects should I show to prove Docker skills?
Show a containerised multi-service application with a lean multi-stage Dockerfile, a Compose file with healthchecks and volumes, and a CI pipeline that builds, scans, tags by commit and pushes the image. A short write-up of the size and build-time improvements you made adds credibility.
Docker questions rarely appear alone; they come bundled with CI/CD, Kubernetes, cloud and security. If you want guided, lab-based preparation across that whole stack, look at the Cloudsoft DevOps course, available in our Ameerpet classroom or live online. For a broader path that combines cloud, DevOps, AI/ML and security engineering, explore the APEX AI, ML, Cloud and Cyber Security program. Book a free demo on +91 96660 19191.



