These Ansible interview questions cover what DevOps, cloud and platform engineering panels actually probe in 2026: agentless architecture, inventories, playbooks, roles, variable precedence, idempotency, Vault, rolling updates, performance at scale, testing, Event-Driven Ansible and how AI fits into automation work. Ansible interviews are rarely about reciting module names; they test whether you can write automation that is safe to run twice, safe to run on a thousand hosts and safe to hand to the next engineer. Every answer below starts with the direct response and then explains the reasoning an experienced engineer would give.
How to use this guide
- Freshers and juniors are usually tested on fundamentals: control node versus managed node, inventory, playbook structure, modules, handlers and what idempotency means.
- Mid-level engineers get variable precedence, roles, Jinja2 templates, Vault, error handling, check mode and dynamic inventory.
- Senior and lead engineers are pushed on rolling updates, performance at scale, testing pipelines, execution environments, Ansible Automation Platform, Event-Driven Ansible and design trade-offs against Terraform.
- Scenario rounds (Q45 onwards) are where most offers are won or lost. Practise answering them out loud: symptom, hypothesis, checks, fix, prevention.
Contents
- Fundamentals (Q1โQ11)
- Intermediate: variables, roles, templates, Vault (Q12โQ27)
- Advanced and architecture (Q28โQ40)
- AI in Ansible automation (Q41โQ44)
- Real-world scenarios (Q45โQ55)
- Key takeaways
- Interview preparation checklist
- FAQ
Fundamentals
1. What is Ansible, and how does it differ from Puppet or Chef?
Answer: Ansible is an open-source automation tool for configuration management, application deployment, orchestration and provisioning. It is agentless and push-based: you run it from a control node, and it connects to managed nodes over SSH (or PowerShell remoting for Windows), runs small units of work, and disconnects. Puppet and Chef traditionally install an agent on every node that pulls its desired state from a central server on a schedule.
The practical differences: Ansible needs no agent lifecycle to manage, no certificate authority for agents and no always-on master; you describe work in YAML playbooks rather than a Ruby-based DSL; and execution is ordered, which makes it a natural fit for orchestration (drain this node, patch it, check health, move on). The trade-off is that without an agent there is no continuous enforcement unless you schedule runs yourself, for example from automation controller or a CI job.
Interview tip: Don't say "Ansible is better". Say which model fits which problem: pull agents suit continuous drift correction across huge estates; push suits orchestrated change and teams that want low setup overhead.
2. Explain the Ansible architecture: control node, managed nodes and the pieces in between.
Answer: The control node is any Linux/Unix-like machine with Ansible installed; it holds the playbooks, inventory and configuration. Managed nodes are the targets; Linux targets need SSH access and Python for most modules, Windows targets are reached over WinRM or, increasingly, SSH, using PowerShell-based modules. Network devices often use connection plugins such as network_cli or httpapi where the module logic runs on the control node.
The moving parts are: inventory (which hosts and groups exist), modules (units of work, such as installing a package), plugins (connection, lookup, filter, callback, inventory, become, strategy), playbooks (ordered plays mapping hosts to tasks) and ansible.cfg (behaviour settings). There is no database and no daemon in plain ansible-core.
control node
ansible.cfg + inventory + playbooks
|
| SSH / WinRM / API
v
+-----------+ +-----------+ +-----------+
| web01 | | db01 | | switch01 |
| python | | python | | network |
+-----------+ +-----------+ +-----------+
3. What is the difference between ansible-core and the Ansible community package?
Answer: The Ansible documentation describes two community packages. ansible-core is the minimal language and runtime: the engine, the CLI tools and the ansible.builtin modules and plugins. ansible is the larger "batteries included" package that bundles ansible-core with a community-curated set of collections (for example cloud, network and POSIX collections). Each version of the ansible package pins a specific ansible-core version.
In production, many teams install ansible-core and declare only the collections they need in a collections/requirements.yml, which keeps environments smaller and upgrades deliberate. Others package everything into an execution environment image (see Q34).
Interview tip: If asked "which version of Ansible do you use", answer with both: the ansible-core version and how you pin collections.
4. What are modules, and what happens when a task runs?
Answer: A module is a reusable unit of work with declared arguments, such as ansible.builtin.copy, ansible.builtin.service or amazon.aws.ec2_instance. When a task runs against a Linux host, Ansible packages the module code with its arguments, transfers it over the connection (or streams it with pipelining), executes it with the remote Python, collects a JSON result containing fields like changed, failed and module-specific return values, and cleans up the temporary files.
Well-written modules check current state before acting and only change what is needed, which is where idempotency comes from. The command and shell modules cannot know what your command does, so they report "changed" every time unless you tell them otherwise.
5. What are collections and fully qualified collection names (FQCNs)?
Answer: A collection is the distribution format for Ansible content: modules, plugins, roles and playbooks packaged together under a namespace.collection name, for example community.general or amazon.aws. An FQCN is the full reference to a piece of content: namespace.collection.module, such as ansible.builtin.template or community.postgresql.postgresql_db.
Use FQCNs everywhere. The documentation recommends it because the short-name lookup depends on what is installed and on the collections keyword, and roles do not inherit the collections keyword from the playbook that calls them. FQCNs remove ambiguity when two collections ship a module with the same short name, and ansible-lint flags short names by default.
- name: Render nginx config
ansible.builtin.template:
src: nginx.conf.j2
dest: /etc/nginx/nginx.conf
6. What is an inventory? Compare static and dynamic inventories.
Answer: The inventory defines the hosts Ansible manages, the groups they belong to and variables attached to them. A static inventory is an INI or YAML file you maintain by hand; it suits stable, small estates and lab environments. A dynamic inventory is generated at run time by an inventory plugin that queries a source of truth: a cloud API (amazon.aws.aws_ec2, azure.azcollection.azure_rm, google.cloud.gcp_compute), a CMDB, or a virtualisation platform.
Groups matter because they are how you target work (hosts: web), how you attach variables (group_vars/web.yml) and how you express structure, with nested groups such as prod containing prod_web and prod_db. In the cloud, static files go stale within days, so dynamic inventory with tag-based grouping is the norm.
7. Describe the structure of a playbook.
Answer: A playbook is a YAML list of plays. Each play maps a host pattern to an ordered list of tasks and sets play-level keywords such as become, vars, gather_facts and serial. Tasks call modules; handlers are tasks that run only when notified; roles bring in packaged tasks, handlers, templates and defaults. Plays can also have pre_tasks and post_tasks, which run before and after roles.
- name: Configure web tier
hosts: web
become: true
vars:
nginx_port: 8080
tasks:
- name: Install nginx
ansible.builtin.package:
name: nginx
state: present
- name: Deploy config
ansible.builtin.template:
src: nginx.conf.j2
dest: /etc/nginx/nginx.conf
notify: Restart nginx
handlers:
- name: Restart nginx
ansible.builtin.service:
name: nginx
state: restarted
8. When would you use an ad hoc command instead of a playbook?
Answer: Ad hoc commands (ansible web -m ansible.builtin.ping, ansible db -m ansible.builtin.shell -a "uptime") are for one-off, read-mostly actions: checking connectivity, gathering a fact across the fleet, or an emergency restart. Anything you will run again, anything that changes state in a way someone needs to review, and anything that needs ordering or error handling belongs in a playbook in version control. A good habit is: if you ran the same ad hoc command twice, write the playbook.
9. How do handlers and notify work?
Answer: A task with notify: Restart nginx queues that handler only if the task reports changed. Handlers run once at the end of the play (or the end of each serial batch), in the order they are defined in the handlers section, not the order they were notified, and only once even if notified many times. That is how you avoid restarting a service five times when five config files change.
Two details interviewers like: ansible.builtin.meta: flush_handlers forces queued handlers to run immediately (useful when a later task needs the restarted service), and if a play fails after a handler was notified, the handler is skipped unless you use --force-handlers or force_handlers: true. Handlers can also listen to a topic so several handlers react to one notification.
10. What are facts, and how do they differ from variables?
Answer: Facts are variables Ansible discovers about a host, collected by the setup module when gather_facts is on: OS family, distribution, IP addresses, memory, mounts and so on, available under ansible_facts. Variables are values you define (inventory, playbooks, roles, extra vars). Facts describe reality; variables describe intent.
You can also create host-scoped values at run time with set_fact, and drop static custom facts in /etc/ansible/facts.d on the managed node. For large fleets, fact gathering is expensive, so you limit it with gather_subset, disable it where unneeded, or enable fact caching (for example the jsonfile or redis cache plugins) so later plays reuse earlier results.
11. What is idempotency, and how does Ansible achieve it?
Answer: An idempotent task produces the same end state no matter how many times you run it, and reports changed only when it actually changed something. Ansible achieves this through declarative modules: ansible.builtin.package with state: present checks whether the package is installed before acting; ansible.builtin.lineinfile checks whether the line exists.
Idempotency breaks with command/shell, with tasks that append rather than set, and with templates that embed timestamps. Fix it by preferring a real module, using creates/removes arguments, adding changed_when based on output, or checking state first and running the command conditionally. The practical test is simple: run the playbook twice; the second run should report zero changes.
Interview tip: Mention that you test idempotency automatically in Molecule (Q32). That single sentence tells the panel you have done this in a team.
Intermediate: variables, roles, templates, Vault
12. Explain Ansible variable precedence.
Answer: When the same variable is defined in several places, Ansible applies a documented order. From lowest to highest, the main points are: role defaults at the bottom; then inventory group variables (with all below specific groups); then host variables; then host facts; then play vars, vars_prompt and vars_files; then role vars; then block and task vars; then include_vars; then registered vars and set_fact; then role and include parameters; and extra vars (-e) always win.
The design rule that falls out of it: put tunable values in role defaults/main.yml so inventory can override them, put values that must not be overridden casually in role vars/, keep environment differences in group_vars, and reserve -e for pipeline inputs like a release version. Also note that command-line options such as -u are not variables; a connection variable like ansible_user in inventory overrides them.
Interview tip: Nobody expects you to recite all levels. Get the bottom (role defaults), the top (extra vars) and the defaults-versus-vars distinction right, and explain how you avoid needing the full list.
13. How do you manage configuration for dev, staging and production?
Answer: Keep one set of roles and playbooks and separate the data. The common layout is one inventory per environment, each with its own group_vars and host_vars, and secrets encrypted per environment with Vault IDs:
inventories/
dev/
hosts.yml
group_vars/all.yml
group_vars/web.yml
prod/
aws_ec2.yml
group_vars/all.yml
group_vars/all/vault.yml
roles/
playbooks/site.yml
The pipeline selects the environment with -i inventories/prod. Promotion means the same commit of the roles moves from dev to prod; only inventory data differs. Avoid if env == 'prod' logic sprinkled through tasks, because it hides behaviour that should be visible in data.
14. How do Jinja2 templates work in Ansible?
Answer: Ansible uses Jinja2 for all variable interpolation ({{ var }}) and for template files rendered on the control node by ansible.builtin.template, then copied to the target. Templates support loops, conditionals and filters, for example {{ users | map(attribute='name') | join(',') }}, default(), to_nice_yaml and ipaddr-style filters from collections.
Good practice: keep logic light in templates, compute complex data structures in variables, add an {{ ansible_managed }} header comment, and use validate on the template task (for example nginx -t -c %s or visudo -cf %s) so a broken file is never put in place. Using backup: true keeps the previous version for quick rollback.
15. How do you control task flow with when, loop and until?
Answer: when evaluates a raw Jinja2 expression (no braces) to skip or run a task, for example when: ansible_facts['os_family'] == 'Debian'. loop iterates over a list, with loop_control for the loop variable name, labels that keep output readable, and pauses. until with retries and delay re-runs a task until a condition is true, which is the standard pattern for waiting on a health endpoint.
- name: Wait for app health
ansible.builtin.uri:
url: "http://{{ inventory_hostname }}:8080/health"
status_code: 200
register: health
until: health.status == 200
retries: 20
delay: 6
Watch for performance: many package modules accept a list directly, so passing the list once is faster than looping one package at a time.
16. What is a role, and how is it structured?
Answer: A role is a self-contained, reusable unit of automation with a fixed directory convention, so Ansible loads its parts automatically:
roles/nginx/
defaults/main.yml # overridable inputs
vars/main.yml # internal constants
tasks/main.yml
handlers/main.yml
templates/
files/
meta/main.yml # dependencies, platforms
meta/argument_specs.yml
A well-designed role does one thing (configure nginx, harden SSH), exposes a clear input contract through defaults, prefixes its variables with the role name to avoid collisions, and validates inputs with meta/argument_specs.yml, which ansible-core checks before the role runs. Roles promote reuse because a team can share one hardened, tested role instead of every project copying tasks.
17. Roles versus playbooks, and import versus include: when do you use each?
Answer: A playbook decides what runs where and in which order; a role packages how to do one capability. Playbooks should be thin: target hosts, set a few play keywords, call roles.
import_role/import_tasks are static: processed at parse time, so tags and --list-tasks see their contents and conditions are applied to every imported task. include_role/include_tasks are dynamic: processed at run time, so you can loop over them or choose the file from a variable, but tags and listing only see the include itself unless you use apply. Rule of thumb: import by default for predictability, include when you need runtime decisions.
18. How do you use Ansible Galaxy, and what are its pros and cons?
Answer: Galaxy is the community hub for collections and roles; ansible-galaxy collection install -r requirements.yml installs pinned dependencies. Enterprises using Ansible Automation Platform often use automation hub (Red Hat certified and validated content) or a private automation hub to curate what is allowed.
Pros: you avoid rewriting common automation, and popular collections are maintained by vendors or active communities. Cons: quality varies, abandoned content exists, and pulling unpinned latest versions makes runs non-reproducible. Pin versions, review content before adopting it, mirror approved content internally, and treat third-party roles like any other dependency with a supply-chain risk.
19. What do check mode and diff mode do?
Answer: --check runs a dry run: modules that support check mode report what they would change without changing it. --diff shows the before-and-after for file-type changes such as templates and lineinfile. Together, --check --diff is the standard pre-change review, and many teams run it in a pull request against staging.
The limits are important: modules that do not support check mode are skipped; command/shell tasks are skipped; and tasks that depend on the result of a previous change (install a package, then template its config) can fail or mislead because the first change never happened. You can force a read-only task to run in check mode with check_mode: false, and use diff: false on tasks that handle secrets so their content is not printed.
20. How do you secure sensitive data with Ansible Vault?
Answer: Ansible Vault encrypts files or individual values with AES-256 so secrets can live in version control. Use ansible-vault encrypt for whole files or ansible-vault encrypt_string for single variables. A common pattern is group_vars/all/vault.yml holding vault_db_password, referenced from a plain vars.yml as db_password: "{{ vault_db_password }}", so you can still grep for variable names.
Use vault IDs (--vault-id dev@prompt, --vault-id prod@/path/to/client-script) to separate keys per environment. Never commit the vault password; supply it from a CI secret, a password client script, or the credential store in automation controller. Vault protects secrets at rest in Git; it does not stop them leaking into logs, which is the job of no_log (Q47).
21. How does block/rescue/always error handling work?
Answer: A block groups tasks; if any task in it fails, execution jumps to rescue; always runs regardless. It is Ansible's try/catch/finally, and it is how you build safe changes with automatic rollback.
- block:
- name: Deploy new release
ansible.builtin.unarchive:
src: "app-{{ version }}.tgz"
dest: /opt/app/releases/
- name: Switch symlink
ansible.builtin.file:
src: "/opt/app/releases/{{ version }}"
dest: /opt/app/current
state: link
rescue:
- name: Roll back symlink
ansible.builtin.file:
src: "/opt/app/releases/{{ previous }}"
dest: /opt/app/current
state: link
always:
- name: Re-enable monitoring
ansible.builtin.include_tasks: monitoring_on.yml
Inside rescue, ansible_failed_task and ansible_failed_result tell you what broke. A rescued host is not counted as failed, so if you want the run to still fail after rollback, end the rescue with ansible.builtin.fail.
22. Explain ignore_errors, failed_when and changed_when.
Answer: ignore_errors: true lets the play continue when a task fails; it is a blunt instrument and often hides real problems. failed_when defines failure precisely, for example failing only when output contains "ERROR" or when the return code is not 0 or 2. changed_when defines what counts as a change, which is how you make command tasks honest: changed_when: false for read-only commands, or a condition on stdout for commands that report whether they did something.
Interview tip: Say you avoid ignore_errors in reviewed code and prefer failed_when with a specific condition. Panels read ignore_errors everywhere as a red flag.
23. How do you use tags effectively?
Answer: Tags let you run or skip subsets of a playbook: --tags config, --skip-tags packages. They are useful for fast iteration and targeted operations, such as rotating certificates without re-running a full build. Keep a small, documented tag vocabulary (packages, config, service, certs), apply tags at role or block level rather than on every task, and remember the special tags always and never. Over-tagged code becomes a maze where nobody knows which partial run is safe, so test the common tag combinations too.
24. How do delegate_to, local_action and run_once work?
Answer: delegate_to runs a task on a different host than the one being iterated, while keeping the current host's variables. Classic uses: removing a web node from a load balancer (delegate_to: lb01), calling an API from the control node (delegate_to: localhost), or updating DNS. local_action is older shorthand for delegating to localhost; delegate_to: localhost is clearer. run_once: true runs a task on one host of the batch and applies the result to all, useful for a database migration that must run exactly once.
Gotchas: facts gathered by a delegated task are assigned to the original host unless you set delegate_facts: true, and run_once combined with serial runs once per batch, not once per play.
25. How do you configure dynamic inventory for a cloud provider?
Answer: Use the provider's inventory plugin with a YAML config file whose name matches what the plugin expects (for AWS, a file ending in aws_ec2.yml). Filter to the hosts you want, build groups from tags with keyed_groups, and choose the address to connect to with compose or hostnames:
plugin: amazon.aws.aws_ec2
regions: [ap-south-1]
filters:
tag:Environment: prod
instance-state-name: running
keyed_groups:
- key: tags.Role
prefix: role
compose:
ansible_host: private_ip_address
Authentication should come from an instance role or a short-lived credential, not keys in the file. Check the result with ansible-inventory -i aws_ec2.yml --graph. Tag hygiene becomes part of your automation contract: an untagged instance is an unmanaged instance.
26. How do you debug a failing playbook?
Answer: Work from cheap to deep. Increase verbosity (-v to -vvvv; four levels show connection details). Print values with ansible.builtin.debug and register. Use --start-at-task and --step to avoid re-running everything, --limit to one host, and --syntax-check and --list-tasks before running. Enable the playbook debugger (debugger: on_failed) to inspect and modify variables at the point of failure. For connection problems, test with ansible host -m ansible.builtin.ping -vvvv and a plain SSH command using the same user and key.
Real-world example: A role worked on Ubuntu but failed on RHEL. -vvv showed the template task succeeded but the handler failed; ansible_facts['os_family'] revealed the service name differed, and the role needed an OS-specific vars file loaded with include_vars.
27. When do you use async and poll?
Answer: async sets the maximum run time for a task and runs it in the background on the managed node; poll sets how often Ansible checks it. Use it for long operations that would otherwise hit SSH timeouts (large package updates, database dumps) or to start many long tasks in parallel. With poll: 0 the task is fire-and-forget; you register the job and later check it with ansible.builtin.async_status in an until loop. Do not use poll: 0 for operations that need exclusive locks, such as package managers, if other tasks in the play touch the same lock.
Advanced and architecture
28. How do serial and max_fail_percentage enable rolling updates?
Answer: By default Ansible runs each task across all targeted hosts before moving on. serial makes the whole play run on batches: a number, a percentage, or a list that ramps up, such as serial: [1, 5, 20] (one canary host, then five, then twenty at a time). Each batch completes the entire play, including handlers, before the next batch starts.
max_fail_percentage stops the run if failures in a batch exceed the threshold; with serial, failure scope becomes the batch. any_errors_fatal: true aborts everything on the first failure, which is what you usually want for a cluster where a half-applied change is dangerous. Combined with load balancer delegation and health checks, this gives you a controlled rolling deployment (see Q48).
29. What are execution strategies, and how does throttle differ from forks and serial?
Answer: The default linear strategy runs each task on all hosts in the batch (up to the fork limit) before starting the next task. free lets each host run through the play as fast as it can, which helps when hosts vary widely in speed and tasks are independent. The debug strategy drops into the debugger on failure.
forks is the global cap on parallel workers (default 5). serial limits how many hosts are in a play batch. throttle limits workers for a specific task or block, for example calling a rate-limited API with throttle: 1. Throttle can only reduce concurrency below forks or serial, never raise it.
30. How do you make Ansible runs faster?
Answer: Measure first, then apply the levers that matter:
- Forks: raise from the default of 5 to a value your control node CPU, memory and network can sustain.
- Pipelining:
pipelining = Truereduces SSH operations per task by streaming modules instead of copying files. It is off by default because it conflicts withrequirettyin sudoers, so confirm that setting on targets. - SSH multiplexing: keep
ControlMaster/ControlPersistenabled and set a sensible persist time so connections are reused. - Facts: disable gathering where not needed, use
gather_subset, and enable fact caching. - Task design: pass lists to package modules instead of looping, avoid unnecessary dynamic includes, use
asyncfor long independent work. - Strategy:
freewhen tasks are independent across hosts.
Profile with the ansible.posix.profile_tasks callback (enabled via callbacks_enabled) to see which tasks dominate, rather than guessing.
31. When would you write a custom module or plugin?
Answer: When existing modules cannot express the operation idempotently and you would otherwise chain shell commands with fragile parsing, for example managing objects in an internal API. A custom module is usually Python using AnsibleModule for argument parsing, must support check mode where possible, returns changed accurately and documents itself with DOCUMENTATION, EXAMPLES and RETURN blocks. Ship it inside a collection (plugins/modules/) with unit tests and versioning.
Choose the plugin type by job: a filter for data transformation, a lookup for fetching data on the control node, an inventory plugin for a custom source of truth, a callback for reporting to a logging or chat system. Before writing anything, search existing collections; maintaining your own module is a long-term commitment.
32. How do you test Ansible code with ansible-lint and Molecule?
Answer: Testing has layers. ansible-playbook --syntax-check catches YAML and structure errors. ansible-lint enforces style and catches anti-patterns: missing FQCNs, unnamed tasks, command where a module exists, risky file permissions, and tasks that are not idempotent-looking. Molecule tests roles and collections end to end: it creates disposable instances (containers or VMs), runs a converge playbook, runs it again to check idempotence (the second run must report no changes), then runs a verify step that asserts the outcome, and finally destroys the instances.
Both tools are part of the Ansible development tools package (ansible-dev-tools), alongside ansible-navigator, ansible-builder and ansible-creator. In CI, run lint on every pull request and Molecule on every role change, ideally across each OS the role claims to support.
33. How do you integrate Ansible with a CI/CD pipeline?
Answer: Treat automation as code with the same lifecycle as applications. A typical pipeline in Jenkins, GitHub Actions or GitLab CI:
- Pull request: ansible-lint, syntax check, Molecule for changed roles.
- Merge: run
--check --diffagainst staging and publish the diff as an artifact. - Apply to staging, run smoke tests.
- Approval gate, then apply to production with the same commit and a pinned execution environment.
Secrets come from the CI secret store or a vault, never from the repository. For larger organisations, the pipeline often calls automation controller's API to launch a job template rather than running ansible-playbook on a CI runner, which centralises credentials, RBAC and audit logs. If you want the pipeline-tool side covered in depth, see these DevOps engineer interview questions.
34. What are execution environments, and why do they matter?
Answer: An execution environment (EE) is a container image that bundles ansible-core, the collections, Python libraries and system packages your automation needs. You build one with ansible-builder from a definition file and run playbooks inside it with ansible-navigator or automation controller. EEs solve the "works on my laptop" problem: the CI runner, a colleague's machine and the production controller all run the identical dependency set.
Version EE images like any artifact, rebuild them on a schedule to pick up security fixes, and keep them lean, since each extra collection is extra attack surface and pull time. Container skills transfer directly here; the Docker interview questions guide covers image hygiene.
35. What is Red Hat Ansible Automation Platform, and what are its main components?
Answer: Ansible Automation Platform (AAP) is Red Hat's supported enterprise product built around Ansible. Per Red Hat's documentation for recent releases, its main components are:
- Automation controller (previously Ansible Tower): web UI and API for job templates, workflows, schedules, credentials, RBAC and job history. AWX is its upstream open-source project.
- Automation hub: certified and validated content, plus private hub for internal collections and EE images.
- Event-Driven Ansible controller: runs rulebooks that trigger automation from events (Q37).
- Platform gateway: a single entry point for authentication and access across the components.
- Execution environments and automation mesh: containerised runtimes and a network of execution nodes that distribute work across sites.
Red Hat also bundles generative AI assistance (Q41). Component names have shifted across releases, so check the current documentation for the version a company runs.
36. How do you handle access control and credentials in automation controller or AWX?
Answer: Organise by organisations and teams, map them to your identity provider (LDAP, SAML or OIDC, for example Microsoft Entra ID), and grant roles on specific objects: who can use a credential, execute a job template, admin an inventory. Engineers launch approved job templates without ever seeing the SSH key or cloud secret, because credentials are stored encrypted and injected at run time; external credential plugins can fetch them from HashiCorp Vault, CyberArk, Azure Key Vault or AWS Secrets Manager.
Add surveys for controlled inputs, workflow approval nodes for production changes, and ship job logs and activity streams to a central log platform for audit. The principle is that people get access to outcomes, not to raw credentials.
37. What is Event-Driven Ansible, and how does a rulebook work?
Answer: Event-Driven Ansible (EDA) reacts to events instead of waiting for someone to launch a job. The engine, ansible-rulebook, runs rulebooks: YAML files containing rulesets, each with sources (event inputs such as a webhook, Kafka or a monitoring alert source from the ansible.eda collection), rules with conditions evaluated against each event, and actions such as run_job_template, run_workflow_template, run_playbook or debug. In AAP the Event-Driven Ansible controller manages rulebook activations.
- name: Disk alerts
hosts: all
sources:
- ansible.eda.webhook:
host: 0.0.0.0
port: 5000
rules:
- name: Clean logs on disk-full alert
condition: event.payload.alert == "disk_full"
action:
run_job_template:
name: clean-app-logs
organization: ops
Keep the rulebook thin: it decides whether to act, while the job template, with its RBAC and audit trail, decides how. Start with low-risk, well-understood remediations and add approval gates for anything destructive (Q43).
38. Ansible vs Terraform: what is the difference, and when do you use each?
Answer: Terraform is declarative infrastructure as code with a state file: it builds a dependency graph, plans the difference between desired and actual resources, and creates, updates or destroys cloud resources (VPCs, clusters, databases). Ansible is procedural-with-declarative-modules and largely stateless: it is strongest at configuring what runs inside and on top of infrastructure (OS settings, packages, application deployment) and at orchestrating ordered operational change.
| Aspect | Terraform | Ansible |
|---|---|---|
| Primary job | Provision and lifecycle cloud resources | Configure, deploy, orchestrate |
| State | State file tracks managed resources | No state; checks live system each run |
| Deletion | Removing code plans a destroy | Removing a task does nothing; you write state: absent |
| Ordering | Dependency graph | Top-to-bottom task order |
| Preview | terraform plan | --check --diff (approximate) |
The common pattern is Terraform to provision, Ansible to configure (Q52). Ansible can create cloud resources, but without state it will not know about resources it created and later stopped describing. For the provisioning side, see the Terraform interview questions.
39. How does Ansible compare with Puppet, Chef and Salt, and with image-based approaches?
Answer: Puppet and Chef are agent-based, pull-model tools with strong continuous enforcement and mature reporting; they suit very large, long-lived server estates where drift correction every few minutes matters. Salt offers both agent and agentless modes with a fast event bus. Ansible wins on low setup cost, readability and orchestration.
The bigger shift in 2026 is that many workloads no longer need in-place configuration management at all: immutable images built with Packer, containers on Kubernetes, and GitOps controllers that continuously reconcile cluster state. In those estates Ansible's role moves to building images, bootstrapping clusters, managing the remaining VMs and network devices, and orchestrating operational runbooks. Saying this clearly shows the panel you understand where the tool fits, rather than defending it everywhere.
40. How do you manage secrets beyond Ansible Vault?
Answer: Vault-encrypted files work, but they share one key across everyone who can decrypt them and make rotation manual. In larger environments, pull secrets at run time from a central secret manager: lookup plugins such as community.hashi_vault.hashi_vault or amazon.aws.aws_secret, or credential plugins in automation controller. That gives short-lived credentials, per-role access policies, central rotation and audit.
Regardless of storage, protect secrets in use: no_log: true on tasks that handle them, diff: false on templates that contain them, restrictive file modes (mode: "0600") on rendered files, and no secrets in -e arguments that land in shell history or CI logs. For how this fits wider pipeline security, the DevSecOps interview questions go further.
If you want to practise these patterns hands-on, with roles, Molecule, dynamic inventory and CI pipelines on real cloud labs, Cloudsoft's DevOps training in Hyderabad covers Ansible alongside Terraform, Docker, Kubernetes and CI/CD, in the Ameerpet classroom or live online.
AI in Ansible automation
41. What AI assistance does the Ansible ecosystem offer today?
Answer: Red Hat introduced generative AI for Ansible under the name Red Hat Ansible Lightspeed. Red Hat now presents it as two capabilities: an automation coding assistant (formerly Ansible Lightspeed) that generates task, playbook and role YAML from natural-language prompts inside the Ansible VS Code extension, and an automation intelligent assistant, a chat interface inside Ansible Automation Platform that helps administrators and operators with questions and troubleshooting. Red Hat's product page lists IBM watsonx Code Assistant, Google Gemini on Vertex and Red Hat AI as model options for the coding assistant. Naming and availability have changed several times, so check current Red Hat documentation before quoting specifics in an interview.
General-purpose coding assistants also write Ansible well because the YAML is readable and the modules are well documented. The engineering question is not which tool, but how you review and test what it produces.
42. How would you use AI-generated playbooks safely in a team?
Answer: Treat AI output exactly like code from a new team member: useful, often right, never trusted without review. The workflow I would put in place:
- Prompt with context: target OS, collections in use, team conventions (FQCNs, naming, variable prefixes).
- Read every task. Check that each module exists in the pinned collection version, that arguments are real, and that
shellhas not been used where a module exists. - Run ansible-lint and Molecule, including the idempotence check, in CI. AI-generated code fails idempotence surprisingly often.
- Run
--check --diffagainst staging and have a human approve the pull request. - Never paste secrets, hostnames or internal data into an external assistant; follow the company's approved-tool policy.
The value is speed on boilerplate (role skeletons, Molecule scenarios, converting shell scripts to modules); the risk is plausible-looking YAML with invented parameters. For the wider governance picture, see AI coding assistants in the enterprise.
43. Design AIOps-triggered remediation with Ansible. Where do approval gates go?
Answer: An AIOps platform or monitoring stack detects and correlates an anomaly; Event-Driven Ansible receives the event; a rulebook decides which job applies; automation controller runs it with audited credentials. Approval gates depend on risk, not on how confident the AI sounds.
monitoring / AIOps alert
|
v
EDA rulebook: match + enrich
|
low risk? ---- yes ----> job template runs
| (restart, clear tmp)
no
v
workflow approval node --> on-call approves
|
v
remediation job --> verify --> ticket update
Low-risk, reversible actions on a pre-approved list (restart a stuck service, clear a temp directory, scale out) can run automatically with rate limits. Anything destructive or wide-impact (failovers, data changes, firewall rules, anything touching many hosts) goes through a workflow approval node with the evidence attached. Every run updates the incident ticket, and a circuit breaker stops automation if the same remediation fires repeatedly, because that means the root cause is unresolved. The AIOps interview guide and human-in-the-loop AI cover the decision design in more depth.
44. Where should AI not be trusted in automation?
Answer: AI should not be the final authority on anything irreversible: deleting data, changing access controls, modifying production network policy, or deciding blast radius. It should not hold credentials directly; an agent that triggers automation should call a scoped job template, not get SSH keys. It should not write compliance controls without a human who understands the control objective. And it should never bypass the pipeline: AI-written code goes through the same lint, test, review and approval steps. A good interview answer frames AI as a drafter and a triage assistant, while ownership and accountability stay with the engineer.
Real-world scenarios
45. A playbook reports "changed" on every run and restarts services each time. How do you fix it?
Answer: Non-idempotent tasks cause unnecessary restarts and make "changed" meaningless, so nobody notices real drift. Find which tasks report changed on a second run, then fix each at the cause.
What I would check:
- Run the playbook twice on a test host and list tasks that are changed on the second run.
command/shelltasks: replace with a module, or addcreates/removesor achanged_whenbased on output.- Templates: look for timestamps, random values or unordered dictionaries that render differently each time.
lineinfileregex that does not match the line it inserts, so it appends again.get_urlor download tasks without checksums, file permission or ownership fights between two roles managing the same file.- Handlers wired to tasks that should never notify.
Production consideration: Add Molecule's idempotence step to CI so the regression cannot return, and treat any unexpected "changed" in scheduled production runs as a drift alert worth investigating.
46. A playbook takes hours against 1,000 hosts. How do you speed it up?
Answer: At that scale the defaults (5 forks, full fact gathering, file transfer per task) dominate. Profile, then fix the biggest costs first.
What I would check:
- Enable
ansible.posix.profile_tasksto find slow tasks; often a few loops or waits account for most of the time. - Forks: raise them and check control node CPU, memory and file descriptors; split across execution nodes or automation mesh if one node is saturated.
- Enable pipelining (after confirming
requirettyis off) and SSH connection reuse. - Facts:
gather_facts: falsewhere unused,gather_subsetelsewhere, and fact caching between plays. - Loops over packages or users: pass lists to the module in a single call.
- Strategy
freeif hosts are independent;asyncfor long tasks. - Unreachable or slow hosts holding up batches: tune timeouts and handle unreachable hosts explicitly.
Production consideration: Speed must not undermine safety. For changes with impact, keep serial batches and accept a longer run; optimise the per-host time, not the blast radius.
47. A database password appeared in a CI job log. What do you do?
Answer: Treat it as a security incident first and a code fix second: the secret is compromised the moment it is in a log that others can read.
What I would check:
- Rotate the password immediately and confirm applications pick up the new value.
- Restrict and purge the log: CI job logs, artifacts, any log shipping destination and controller job output.
- Find the source: a task without
no_log: true, adebugof a variable,-vvvverbosity in CI, a template rendered with--diff, or a secret passed via-eon the command line. - Fix: add
no_log: trueon secret-handling tasks (including loops, which print items),diff: falseon secret templates, and fetch secrets from a lookup instead of passing them as arguments. - Add a check in CI: ansible-lint rules plus a secret scanner on logs and the repository.
Production consideration: no_log also hides useful error output, so apply it to the specific tasks that handle secrets, not to whole plays, and record the incident with a blameless review.
48. Design a rolling OS patch for a web tier behind a load balancer with zero downtime.
Answer: Patch a few hosts at a time, take each out of rotation first, verify health before returning it, and stop at the first sign of trouble.
- name: Rolling patch web tier
hosts: web
become: true
serial: [1, 2, 5]
max_fail_percentage: 0
pre_tasks:
- name: Drain from load balancer
ansible.builtin.include_tasks: lb_drain.yml
tasks:
- name: Apply updates
ansible.builtin.dnf:
name: "*"
state: latest
register: patch
- name: Reboot if needed
ansible.builtin.reboot:
when: patch is changed
- name: Wait for health
ansible.builtin.uri:
url: "http://{{ ansible_host }}/health"
register: h
until: h.status == 200
retries: 30
delay: 10
post_tasks:
- name: Return to load balancer
ansible.builtin.include_tasks: lb_enable.yml
What I would check:
- Capacity: can the remaining hosts carry peak traffic with one batch out? That decides the maximum batch size.
- Connection draining time on the load balancer before reboot.
- The health check tests the application, not just the port.
- Patch the canary first and pause or watch metrics before ramping up.
- A tested rollback path, or at least snapshots for stateful hosts.
Production consideration: Run it in a change window with monitoring visible, use the drain and enable tasks delegated to the load balancer or its API, and keep patch versions pinned to a tested repository snapshot so staging and production get the same packages. On Debian-family hosts, swap dnf for apt; on Kubernetes nodes, cordon and drain instead.
49. A playbook caused unexpected changes and an outage. How do you respond and prevent it?
Answer: Stabilise first, then learn. Stop further runs (disable the schedule or EDA activation), roll back using the previous known-good commit or a rescue path, and restore service. Then find out why the change was a surprise.
What I would check:
- What changed between this run and the last good one: role version, collection version, inventory data, extra vars.
- Whether a variable override landed with unexpected precedence (for example a group_vars file overriding a role default for more hosts than intended).
- Whether the host pattern or dynamic inventory matched more hosts than expected.
- Whether
--check --diffran before production and whether anyone read it.
Production consideration: Prevention is process: mandatory check-and-diff review, serial with a canary, pinned collections and EEs, --limit safeguards, and an approval step for production job templates.
50. You inherit a large, messy Ansible repository. How do you refactor it?
Answer: Refactor incrementally behind tests; a big-bang rewrite of automation that runs production is how outages happen.
What I would check:
- Inventory what runs: which playbooks are scheduled, which are dead, which hosts each touches.
- Run ansible-lint to get a baseline and fix mechanical issues first (FQCNs, task names, deprecated syntax).
- Pick one high-value area, extract a role with a clear defaults contract and argument specs, add Molecule tests, and switch callers over.
- Move secrets to Vault or an external manager; remove plaintext.
- Package roles into internal collections with versions once the structure stabilises.
Production consideration: Before and after each change, run --check --diff on representative hosts; an empty diff proves the refactor kept behaviour identical.
51. Your estate grows from 50 to several thousand hosts across regions. How do you scale Ansible?
Answer: Move from engineers running playbooks on laptops to a platform: a central controller (automation controller or AWX), execution environments, dynamic inventory per region, and distributed execution close to the targets.
What I would check:
- Execution capacity: execution nodes or automation mesh hops in each region so SSH traffic stays local.
- Inventory sync schedules and caching so runs do not hammer cloud APIs.
- Job slicing for very large inventories, which splits one job into parallel slices.
- Team ownership: collections per domain, RBAC per team, standard job templates.
- Logging and metrics on job duration and failure rate.
Production consideration: Bake more into images so runs only apply the delta; fewer tasks per host scales better than more forks.
52. Terraform creates new EC2 instances. How do you get them configured by Ansible automatically?
Answer: Keep the tools in their lanes and connect them through tags and inventory, not by having Terraform shell out to Ansible.
What I would check:
- Terraform tags each instance with role and environment.
- Ansible's
aws_ec2inventory plugin groups by those tags. - The pipeline runs Terraform apply, waits for instances to be reachable (
wait_for_connection), then runs the matching playbook with--limitto new hosts; or an EDA rulebook or a launch callback triggers the configuration job. - For autoscaling groups, prefer a pre-baked image plus a small first-boot configuration, because scale-out must not depend on a pipeline run.
Production consideration: Decide ownership clearly. If Terraform manages a security group, Ansible must not edit it, or the two tools fight and each run reverts the other's change.
53. A bank's audit team wants proof that servers meet a hardening baseline. How would you use Ansible?
Answer: Consider a bank whose GCC operations team in Hyderabad manages Linux servers that must follow an internal baseline derived from CIS benchmarks. Ansible can both enforce and report: a hardening role applies the controls, and a scheduled --check --diff run reports drift without changing anything.
What I would check:
- Each control maps to a tagged task with a control ID in its name, so reports map to audit items.
- Controls that would break applications are parameterised with documented exceptions, approved by the risk team.
- The scheduled check-mode job output is shipped to a central store and summarised per host.
- Remediation runs only through an approved change, with
serialbatches.
Production consideration: Ansible reports what its tasks check, nothing more. For full coverage, pair it with a dedicated compliance scanner and use Ansible for remediation.
54. A playbook works on your laptop but fails in CI and in automation controller. What do you investigate?
Answer: Almost always an environment difference: different ansible-core or collection versions, missing Python libraries, different credentials or network paths.
What I would check:
ansible --versionandansible-galaxy collection listin both places.- Python dependencies the modules need (for example boto3 for AWS modules) inside the execution environment.
- Which
ansible.cfgis picked up; Ansible uses the first one it finds, and settings likeroles_pathdiffer. - Credentials and SSH keys injected in CI versus your agent forwarding locally.
- Network: can the runner reach the hosts, bastion and proxy?
Production consideration: Develop inside the same execution environment with ansible-navigator so local runs match CI and the controller.
55. Some hosts randomly show UNREACHABLE during large runs. How do you troubleshoot?
Answer: Intermittent unreachability at scale is usually resource or network limits, not broken hosts.
What I would check:
- Run
-vvvvagainst an affected host to see whether it is a timeout, authentication or host key failure. - Bastion or jump host limits:
MaxStartupsandMaxSessionsin sshd, connection rate limits, and forks set higher than the bastion can accept. - Control node limits: file descriptors, CPU, ephemeral ports.
- Dynamic inventory returning terminated or stopped instances.
- Timeouts: raise the connection
timeoutfor slow networks; useignore_unreachableonly deliberately.
Production consideration: Track unreachable hosts as a metric. A host that cannot be reached by automation cannot be patched, which makes it a security concern, not just an annoyance. The Linux commands interview questions cover the SSH and system-level debugging behind this.
Key takeaways
- Idempotency is the core skill: prefer modules over shell, make
changedhonest, and prove it with a second run in Molecule. - Know the shape of variable precedence (role defaults at the bottom, extra vars at the top) and design so you rarely depend on the middle.
- Use FQCNs, pinned collections and execution environments so every run is reproducible.
- Safe change at scale means
serialwith a canary, health checks,max_fail_percentage, delegation to load balancers, and block/rescue rollback. - Vault protects secrets at rest;
no_log,diff: falseand external secret managers protect them in use. - Ansible configures and orchestrates; Terraform provisions and tracks state. Use both with clear ownership.
- AI can draft playbooks and trigger remediation through EDA, but review, testing and approval gates stay with engineers.
Interview preparation checklist
- Build a lab with one control node and three managed nodes (two Linux distributions if possible).
- Write a role with defaults, handlers, a validated template and argument specs; publish it in a Git repository.
- Add ansible-lint and a Molecule scenario with the idempotence check to a CI pipeline.
- Configure a dynamic inventory from a cloud account using tags and
keyed_groups. - Encrypt secrets with Vault IDs for two environments, and test a task with
no_log. - Implement a rolling update with
serial, load balancer drain viadelegate_to, health checks and a rescue rollback. - Profile a slow run with
profile_tasksand tune forks, pipelining and fact gathering. - Write a small EDA rulebook that triggers a playbook from a webhook.
- Prepare two stories from your own work: one automation failure you fixed and one process you automated end to end.
- Revise Linux fundamentals (SSH, sudo, systemd, package managers) and Git workflows; see these Git interview questions.
FAQ
Is Ansible still worth learning in 2026?
Yes. Even in container-heavy organisations, teams use Ansible to configure virtual machines, network devices and Windows servers, to build images and to orchestrate operational runbooks, and it appears regularly in DevOps and cloud job descriptions alongside Terraform and Kubernetes.
What skills are required for an Ansible-focused DevOps role?
Solid Linux administration, SSH and networking basics, YAML and Jinja2, Git, one cloud platform, and the ability to write tested, idempotent roles. Python helps for custom modules and filters, and CI/CD knowledge ties it all together.
How should a fresher prepare for Ansible interview questions?
Build a small lab, write playbooks and one role from scratch, run them twice to understand idempotency, and be ready to explain inventory, modules, handlers, variables and Vault in your own words with examples from your lab.
Do I need to know Python to use Ansible?
Not to write playbooks and roles, which are YAML with Jinja2. Python becomes useful when you write custom modules or filter plugins, debug module failures, or manage the Python dependencies that cloud modules need.
Should I learn Ansible or Terraform first?
If your work is mostly server configuration and operations, start with Ansible; if it is mostly cloud provisioning, start with Terraform. Most DevOps roles expect both, so learn the second soon after and practise using them together.
What is the difference between AWX and automation controller?
AWX is the open-source upstream project. Automation controller, previously called Ansible Tower, is the supported component of Red Hat Ansible Automation Platform built from that work, with Red Hat support and lifecycle commitments.
Is there an Ansible certification?
Red Hat offers Ansible-related certification exams as part of its certification programme. Check Red Hat's official certification pages for the current exam names and objectives, because they change across product versions.
How many Ansible questions are asked in a DevOps interview?
It varies by company and role. Ansible-heavy roles may spend a full round on it, while general DevOps interviews often mix a few Ansible questions with Linux, CI/CD, containers, Terraform and cloud, usually ending with a scenario.
Can Ansible manage Windows servers?
Yes. Ansible connects to Windows over WinRM or SSH and uses PowerShell-based modules from the ansible.windows and community.windows collections to manage features, services, registry settings and updates.
Ansible rarely gets hired on its own; panels want engineers who can connect it to Linux, cloud, CI/CD and containers in one working pipeline. Cloudsoft's DevOps course builds exactly that stack through hands-on labs, in our Ameerpet classroom beside the Metro or live online. If you want to combine DevOps with AI, ML and cyber security, look at the APEX AI, ML, Cloud and Cyber Security program. Call +91 96660 19191 for a free demo.



