New batches starting this week ยท Limited seats

Terraform Interview Questions and Answers 2026 (60 Questions)

60 Terraform interview questions with practical answers for DevOps and cloud roles, from state, locking and modules to moved and removed blocks, ephemeral values, policy as code, OpenTofu vs Terraform and twelve real production scenarios.

Terraform interview questions and answers 2026 - Cloud Soft Solutions
Last updated ยท 45 min read ยท 9,894 words

These Terraform interview questions cover what DevOps, cloud and platform interviews in 2026 actually probe: how state works and fails, how to structure modules and environments, how to refactor live infrastructure safely, and how to run Terraform through a reviewed pipeline. Interviewers rarely stop at "what is Terraform"; they ask how you would recover a stuck state lock, reconcile drift after a console change, or rename a resource without destroying it, and they watch how carefully you reason about blast radius. The 60 questions below go from fundamentals to state, modules, the modern language features (import, moved, removed, check blocks, tests, ephemeral values), policy as code, CI/CD, OpenTofu vs Terraform, AI in IaC and twelve real-world scenarios.

How to use this guide:

  • Freshers and juniors are usually tested on the fundamentals: the init/plan/apply workflow, providers, resources vs data sources, variables and the purpose of state. Be able to write a small configuration from memory.
  • Mid-level DevOps and cloud engineers get the Terraform state interview questions (remote backends, locking, drift), Terraform modules interview questions (structure, versioning), count vs for_each, lifecycle rules and at least one scenario.
  • Senior and platform roles are pushed on multi-environment layouts, refactoring with moved and removed blocks, policy as code, plan review in CI/CD, secrets handling and the OpenTofu vs Terraform decision.
  • For every answer, practise the shape: one-line answer, the mechanism, then the trade-off or failure mode. That is what separates an engineer who has run terraform apply against production from one who has read about it.

Contents

Terraform fundamentals

1. What is Terraform and what is Infrastructure as Code?

Answer: Infrastructure as Code (IaC) means defining infrastructure (networks, compute, databases, IAM, DNS) in version-controlled files instead of clicking through consoles. Terraform is a declarative IaC tool from HashiCorp: you describe the desired end state in HCL (HashiCorp Configuration Language), and Terraform computes a plan of create, update and delete actions to make reality match, then executes it through provider plugins that call each platform's API. Because it is provider-based, the same workflow manages AWS, Azure, Google Cloud, Kubernetes, GitHub, Datadog and hundreds of other APIs.

Interview tip: Mention that IBM completed its acquisition of HashiCorp in February 2025, so Terraform is now an IBM product line. It shows you follow the ecosystem without turning the answer into trivia.

2. What does "declarative" mean in Terraform, and how is it different from a script?

Answer: Declarative means you state what should exist, not the steps to get there. If you declare three subnets and two already exist, Terraform creates one. Run the same configuration again and nothing happens, because reality already matches. A Bash or Python script using the cloud CLI is imperative: run it twice and you may get six subnets or an error. Terraform's idempotency comes from comparing configuration, state and the real world on every plan. The trade-off is less control over ordering, which Terraform derives from references (Q8).

3. Walk through the core workflow: init, plan, apply and destroy.

Answer: terraform init prepares the working directory: it configures the backend, downloads providers (recording exact versions and checksums in .terraform.lock.hcl) and fetches modules. terraform plan refreshes the known state of real resources, compares it with configuration and prints the proposed changes. terraform apply executes a plan; in pipelines you should apply a saved plan file (plan -out=tfplan, then apply tfplan) so what runs is exactly what was reviewed. terraform destroy removes everything the configuration manages.

Interview tip: The phrase "plan shows the diff, apply executes it, always review the plan" is the minimum. The senior answer adds saved plans and the fact that a plan can go stale if someone else applies in between.

4. What is a provider, and how do you pin provider versions?

Answer: A provider is a plugin that translates Terraform resources into API calls for one platform, for example hashicorp/aws or hashicorp/azurerm. You declare it in a required_providers block with a source and a version constraint such as ~> 6.0 (allow minor and patch upgrades, block the next major). The .terraform.lock.hcl file records the exact version selected and its checksums; commit it so every engineer and every CI run uses the same provider build. Upgrade deliberately with terraform init -upgrade in a dedicated pull request, read the provider's changelog for the major version, and run a plan in each environment before merging. Use alias for several instances, such as two regions.

5. What is the difference between a resource and a data source?

Answer: A resource block declares an object Terraform owns: it creates, updates and destroys it, and records it in state. A data block reads an existing object that something else owns, for example the latest approved AMI, an existing VPC from the network team, or the current caller identity, and exposes its attributes for use. Data sources never modify infrastructure. Rule of thumb: if your team does not own an object's lifecycle, read it rather than import it.

6. Explain input variables, locals and outputs.

Answer: Input variables are the parameters of a module or root configuration, typed (string, number, bool, list, map, object), optionally with defaults, descriptions, validation rules and sensitive = true. Values come from .tfvars files, -var flags, TF_VAR_ environment variables or workspace variables in HCP Terraform. Locals are named expressions computed inside the module, useful for naming conventions and merged tag maps; they are not settable from outside. Outputs expose values to the caller, to the CLI, or to other configurations.

7. How does Terraform compare with CloudFormation, Bicep, Pulumi and Ansible?

Answer: CloudFormation (AWS) and Bicep (Azure) are first-party, single-cloud tools where the cloud service manages state for you; they are attractive when an organisation is firmly on one cloud. Terraform is multi-provider and has the largest module and provider ecosystem, at the cost of managing state yourself. Pulumi uses general-purpose languages (TypeScript, Python, Go) instead of HCL, which suits teams that want loops, tests and abstractions in a familiar language. Ansible is primarily configuration management: it is procedural, agentless and excellent at configuring operating systems and applications on servers that already exist. A common pattern is Terraform to provision, Ansible or cloud-init to configure, and Kubernetes manifests or Argo CD for workloads.

ToolStyleStateTypical sweet spot
Terraform / OpenTofuDeclarative HCLYou manage (backend)Multi-cloud and SaaS provisioning
CloudFormation / BicepDeclarativeManaged by the cloudSingle-cloud estates
PulumiDeclarative via codeService or self-managedTeams wanting real languages
AnsibleProcedural tasksMostly statelessOS and app configuration

If the role also involves configuration management, prepare with the Ansible interview questions as well.

8. What is the dependency graph, and how does Terraform decide the order of operations?

Answer: Terraform builds a directed acyclic graph from references between blocks. If a subnet references aws_vpc.main.id, the subnet depends on the VPC, so the VPC is created first and destroyed last. Independent resources are created in parallel (ten at a time by default, adjustable with -parallelism). You rarely need to state order explicitly; the reference itself is the dependency. Cycles usually come from two resources referencing each other (for example inline security group rules); split them into separate rule resources.

9. What is the .terraform.lock.hcl file and should it be committed?

Answer: Yes, commit it. It is the dependency lock file for providers: it records the exact provider version chosen during init and the hashes of the packages for each platform. Without it, two engineers or a CI runner could silently get different provider builds within the same version constraint, which produces different plans. A practical gotcha: if developers use macOS on Apple silicon and CI runs on Linux, the lock file may contain hashes for only one platform; run terraform providers lock with the needed -platform flags so CI does not fail on a checksum mismatch. Modules are not recorded in this file, which is why module sources should pin a version or a Git tag.

Terraform state interview questions

10. What is Terraform state and why does it matter?

Answer: State is a JSON document that maps each resource address in your configuration (for example aws_s3_bucket.logs) to the real object's ID and its last known attributes. Terraform needs it for three reasons: to know which real objects it owns (cloud APIs cannot tell it "these were created by this configuration"), to compute accurate diffs, and to track metadata such as dependencies for correct destroy order. Treat state as critical, sensitive data: it often contains secrets in plain text, it must be backed up and versioned, and only one process should write it at a time.

Interview tip: Say plainly that you never edit the state JSON by hand; you use terraform state commands or, better, declarative import, moved and removed blocks.

11. What is a remote backend, and which ones have you used?

Answer: A backend decides where state is stored and how operations lock it. The default local backend writes terraform.tfstate to disk, which breaks the moment two people work on the same infrastructure. Remote backends store state centrally: S3 on AWS, the azurerm backend (Blob Storage) on Azure, gcs on Google Cloud, or HCP Terraform and Terraform Enterprise, which add remote runs, a UI, variable management and policy checks. A good remote backend setup has encryption at rest, versioning (so you can recover a previous state), tight access control (state readers can see secrets), locking, and one state file per environment per component.

Real-world example: Consider a GCC platform team in Hyderabad running AWS for a retail client. They keep one S3 bucket per account for state, with versioning, KMS encryption, a bucket policy restricted to the CI role and two break-glass admins, and a key path like prod/network/terraform.tfstate.

12. What is state locking, and how does it work on S3 today?

Answer: Locking prevents two operations from writing the same state at once, which would corrupt it or lose changes. When a plan or apply starts, Terraform acquires a lock; others wait or fail with a "state locked" error showing the lock ID, who holds it and when it was taken. Historically the S3 backend used a DynamoDB table for locks. Current Terraform supports S3-native locking with use_lockfile = true, which writes a lock object next to the state file, and HashiCorp's S3 backend documentation marks DynamoDB-based locking as deprecated and due for removal in a future minor version. During migration you can configure both simultaneously. Azure Blob uses blob leases, GCS uses a lock object, and HCP Terraform locks workspaces itself.

terraform {
  backend "s3" {
    bucket       = "acme-tfstate-prod"
    key          = "network/terraform.tfstate"
    region       = "ap-south-1"
    encrypt      = true
    use_lockfile = true
  }
}

Interview tip: Many candidates still answer "DynamoDB" from older tutorials. Mentioning S3-native locking and the deprecation shows your knowledge is current; add "check the backend docs for your version".

13. What is drift, and how do you detect and handle it?

Answer: Drift is when real infrastructure diverges from what configuration and state describe, usually because someone changed something in the console, a script or another tool modified it, or the cloud provider changed a default. terraform plan detects drift on managed attributes because it refreshes real values first; terraform plan -refresh-only shows only what changed outside Terraform without proposing configuration changes, and apply -refresh-only accepts those changes into state. Then you decide per change: revert it by applying the configuration, adopt it by updating the code, or ignore it deliberately with ignore_changes when another system legitimately owns that attribute. Continuous detection means a scheduled plan in CI, or HCP Terraform's health assessments, alerting when a plan is not empty. Drift on objects Terraform does not manage at all is invisible to plan; that needs import (Q32) or a cloud inventory tool.

14. Why is state sensitive, and how do you protect it?

Answer: State stores every attribute of every managed resource, including generated database passwords, private keys created by the TLS provider and connection strings. Marking a variable sensitive only hides it in CLI output; it is still written to state in plain text. Protection is layered: encrypted backend storage with a customer-managed key, strict IAM on the bucket or container, no state in Git, no local state on laptops, audit logging on state access, and short-lived credentials for pipelines. The structural fix is to keep secrets out of state entirely: let the cloud generate and store credentials (for example RDS-managed master passwords in Secrets Manager), or use ephemeral resources and write-only arguments (Q37). OpenTofu additionally offers client-side state and plan encryption.

15. When would you use terraform state mv, state rm and state pull/push?

Answer: These are imperative surgery tools. state mv renames an address in state (for example when moving a resource into a module) so Terraform does not destroy and recreate it. state rm makes Terraform forget an object without deleting it, for example when handing ownership to another team. state pull downloads the current state for inspection; state push overwrites remote state and is dangerous, used only in recovery. state list and state show are safe read commands. In 2026 the preferred approach for mv and rm is declarative: moved and removed blocks go through code review, show up in the plan, and run in every environment identically, while CLI surgery happens once on one machine with no review trail. Before any state command, take a backup (state pull > backup.tfstate) or confirm bucket versioning.

16. How do configurations share data: terraform_remote_state or data sources?

Answer: terraform_remote_state reads the outputs of another configuration's state. It works but couples you tightly: the consumer needs read access to the entire producer state (including its secrets), and renaming an output breaks consumers silently. Preferred alternatives: look the object up with a provider data source by tag or name (for example aws_vpc filtered by a tag), or have the producer publish values to a parameter store such as SSM Parameter Store or Azure App Configuration, which consumers read with a data source. HCP Terraform also supports sharing outputs between workspaces with explicit access control.

17. How do you decide how to split state, and what is blast radius?

Answer: Blast radius is how much infrastructure one bad apply can damage. A single state holding the network, databases, Kubernetes cluster and every application means a typo in an app's tag can produce a plan that touches the VPC, and every plan is slow because it refreshes everything. Split state along lines of ownership and change frequency: foundation (accounts, network, DNS) changes rarely and is owned by the platform team; shared data services change occasionally; application stacks change daily. Each environment gets its own state too. The cost of splitting is more cross-state references and ordering between stacks, so do not go to one resource per state either. A useful rule: a state should be owned by one team and be safe to plan in a minute or two.

live/
  prod/
    network/      (platform team)
    data/         (platform team)
    app-orders/   (orders team)
  dev/
    network/
    app-orders/
modules/
  vpc/  rds/  service/

Terraform modules interview questions

18. What are modules and why do they matter?

Answer: A module is a directory of Terraform files with inputs (variables) and outputs; every configuration is a module, and the one you run commands in is the root module. Child modules are reusable, parameterised groups of resources, for example "a VPC with public and private subnets and flow logs" or "a service with its load balancer target, IAM role and alarms". Modules encode standards once (encryption on, tags applied, logging enabled) so every team gets them by default, keep root configurations short, and let you test infrastructure patterns in isolation. The goal is consistency and safe defaults, not just less code.

19. How do you structure a reusable module?

Answer: A conventional layout is main.tf, variables.tf, outputs.tf, versions.tf (required Terraform and provider versions), a README.md with usage, an examples/ directory and tests/ with .tftest.hcl files. Design principles: expose a small typed interface, choose secure defaults, output the IDs and ARNs consumers need, and do not configure providers inside a reusable module (Q23). Keep a module focused on one concept; a "platform" module that creates network, cluster and databases together is hard to version and impossible to reuse partially.

20. How do you version modules and roll out a new version safely?

Answer: Publish modules with semantic versions: a registry (public Terraform Registry, the HCP Terraform private registry, or another private registry) with a version constraint, or a Git source pinned to a tag such as ?ref=v2.3.0. Never point production at a branch. Breaking changes (renamed inputs, resources that will be replaced) go in a major version with release notes and, ideally, moved blocks inside the module so consumers' resources are renamed instead of recreated. Roll out environment by environment: bump the version in dev, review the plan, then staging, then prod. Tools such as Renovate or Dependabot can open those version-bump pull requests automatically, but a human still reads each plan.

Interview tip: Say "a module upgrade is a code change whose impact is only visible in each consumer's plan". That framing shows you understand why pinning matters.

21. Public registry modules or in-house modules?

Answer: Well-maintained public modules (for example the widely used community AWS VPC and EKS modules) save weeks and encode lessons you have not learned yet, but they expose many options, change on their own schedule and may not match your security baseline. Large organisations often wrap them: an internal module calls the public module with company defaults locked in and exposes only the inputs teams should change. Review the source before adoption, pin versions, and mirror it privately if you need supply-chain control.

22. What is module composition, and when should you not create a module?

Answer: Composition means building systems from small modules wired together in the root, passing outputs of one (the VPC's subnet IDs) as inputs to another (the cluster), instead of nesting modules many levels deep. Flat composition keeps dependencies visible and each module independently testable. Do not create a module for a single resource with a thin wrapper of variables that mirror the resource's arguments one to one; it adds indirection without encoding any standard.

23. Why should reusable modules not contain provider blocks?

Answer: If a module configures its own provider, the caller cannot choose the region, account or credentials, cannot use count or for_each on the module, and removing the module later becomes painful because Terraform needs the provider configuration to destroy its resources. Instead, the module declares required_providers (with configuration_aliases if it needs more than one instance, for example a primary and a disaster-recovery region), and the root module passes providers in with the providers argument. This keeps credentials and regions as decisions of the root configuration and the pipeline.

Language: for_each, count, dynamic blocks and lifecycle

24. count vs for_each: what is the difference and when do you use each?

Answer: Both create multiple instances of a resource or module. count takes a number and addresses instances by index (aws_instance.web[0]). for_each takes a map or a set of strings and addresses instances by key (aws_s3_bucket.this["logs"]). The practical difference is identity: with count over a list, removing the second item shifts every later index, so Terraform plans to destroy and recreate resources that did not change. With for_each, each instance is tied to a stable key, so removing one key removes exactly one resource. Use for_each for collections of distinct things; use count for identical replicas or the conditional pattern count = var.enabled ? 1 : 0. Keys in for_each must be known at plan time, so do not derive them from attributes of resources that do not exist yet.

resource "aws_s3_bucket" "this" {
  for_each = toset(["logs", "artifacts", "models"])
  bucket   = "acme-${each.key}-prod"
}

25. What are dynamic blocks and when are they appropriate?

Answer: A dynamic block generates repeated nested blocks inside a resource, for example multiple ingress blocks in a security group or multiple setting blocks, from a collection. It has a for_each, an optional iterator name and a content block. Use it when a module must accept a variable number of nested settings from callers. Do not use it to hide simple, static configuration; three literal blocks are easier to read and review than a dynamic block over a local list.

dynamic "ingress" {
  for_each = var.ingress_rules
  content {
    from_port   = ingress.value.port
    to_port     = ingress.value.port
    protocol    = "tcp"
    cidr_blocks = ingress.value.cidrs
  }
}

26. Explain the lifecycle meta-argument and its options.

Answer: lifecycle changes how Terraform handles create, update and delete for a resource:

  • create_before_destroy: when a change forces replacement, build the new object first, then delete the old one. Essential for things referenced by live traffic (certificates, launch templates). Names must not collide, so use name prefixes.
  • prevent_destroy: any plan that would destroy the resource fails. A seatbelt for production databases and state buckets, though it does not help if someone deletes the resource block itself.
  • ignore_changes: ignore differences on listed attributes, for when another system legitimately manages them (an autoscaler changing desired count, a tag added by a compliance tool).
  • replace_triggered_by: force replacement when a referenced resource or attribute changes.
  • precondition and postcondition: assertions evaluated during plan or apply (Q29).

Interview tip: Warn that ignore_changes = all turns Terraform into a create-only tool for that resource and hides drift; use it with a comment explaining why.

27. What causes a forced replacement, and how do you spot it in a plan?

Answer: Some attributes cannot be changed in place by the cloud API: an RDS instance identifier, an EC2 instance's subnet, a storage account's name, a Kubernetes cluster's network settings. Changing them makes Terraform plan a destroy and create, shown as -/+ (or +/- with create_before_destroy) and annotated "forces replacement" next to the attribute. In review, search the plan for "must be replaced" and the summary's destroy count before anything else. For stateful resources, a replacement means data loss unless you have planned a migration, so pipelines often fail automatically if a plan contains deletes of tagged critical resources, using policy as code (Q42) or a script over terraform show -json.

28. When do you need depends_on, and why is it a code smell?

Answer: depends_on declares a dependency Terraform cannot infer from references, for example an application that needs an IAM policy attachment to exist before it starts, although it only references the role. It is legitimate for such hidden dependencies, but overuse is a smell: it makes plans more conservative (Terraform may defer reading data sources until apply, showing more "known after apply" values) and often hides a missing reference. Prefer passing the actual attribute you depend on.

29. How do variable validation, preconditions, postconditions and check blocks differ?

Answer: All four express assumptions, at different points. validation inside a variable rejects bad input early (for example the environment must be dev, staging or prod; a CIDR must be a /16). A precondition in a resource's lifecycle checks an assumption before the resource is planned, for example that a looked-up AMI is from the approved owner. A postcondition checks the result after apply, such as that an instance got a private IP. Failing any of these stops the run. A check block (Q35) is different: it validates the overall infrastructure and only produces warnings, so it suits health checks that should not block deployment. Choose the earliest point that catches the problem with the clearest message.

30. Why are provisioners considered a last resort?

Answer: local-exec and remote-exec provisioners run scripts during create or destroy, but Terraform cannot model what they do: they are not in the plan, they do not re-run when the script changes, failures leave a resource "tainted", and remote-exec needs network access and credentials to the machine. Better options: cloud-init or user data for first boot, pre-baked images built with Packer, configuration management with Ansible, a provider resource that does the job natively, or terraform_data with triggers_replace if you genuinely need to run a command when certain inputs change. Mention that terraform_data is the built-in replacement for the old null_resource pattern.

31. Which expressions and functions do you use most in real modules?

Answer: for expressions to transform collections ({ for k, v in var.subnets : k => v.cidr if v.public }), merge() for tag maps, lookup() and try() for optional values, coalesce() for fallbacks, cidrsubnet() to carve subnets from a VPC range, templatefile() for user data and policies, jsonencode() for IAM policy documents, flatten() with nested for to turn a map of lists into a flat set for for_each, and the splat operator [*]. Writing the flatten pattern on a whiteboard is a common intermediate test.

Modern Terraform features

32. How do import blocks work, and how are they better than terraform import?

Answer: An import block declares, in code, that an existing real object should be brought under management at a given address: it takes to (the resource address) and id (the provider's identifier; newer provider resources also support an identity attribute set). The import happens during a normal plan and apply, so it is reviewed in a pull request and shows in the plan, whereas the old terraform import CLI command wrote straight to state from one engineer's machine. Running terraform plan -generate-config-out=generated.tf drafts the matching resource blocks for you, which you then clean up into proper modules and variables. Import blocks also accept for_each, so you can import a whole set of buckets or DNS records from a map. After a successful apply, the import blocks can be deleted.

import {
  to = aws_s3_bucket.this["logs"]
  id = "acme-logs-prod"
}

33. What is a moved block and why does it matter for refactoring?

Answer: A moved block tells Terraform that the object previously at one address now lives at another, for example aws_instance.web to module.web.aws_instance.this, or aws_s3_bucket.logs to aws_s3_bucket.this["logs"] when switching from separate resources to for_each. Without it, Terraform sees one address disappear and a new one appear, and plans destroy plus create. With it, the plan shows the move and no infrastructure change. Because it is code, the same refactor applies consistently in dev, staging and prod, and module authors can ship moved blocks inside a new major version so consumers upgrade without replacements. Keep moved blocks for a while (until every state using the module has applied), then remove them.

34. What does a removed block do?

Answer: A removed block declares that a resource or module is leaving the configuration and says what should happen to the real object. With lifecycle { destroy = false }, Terraform forgets the object (removes it from state) without deleting it, which is the declarative, reviewable version of terraform state rm. It is used when handing a resource to another team's state, or when decommissioning Terraform management of something that must keep running. With destroy left at its default, Terraform destroys the object, and you can attach destroy-time provisioners for clean-up. The key point for interviews: deleting a resource block alone means "destroy it"; a removed block with destroy = false means "stop managing it".

35. What are check blocks used for?

Answer: A check block contains one or more assert conditions, optionally with a scoped data source, that Terraform evaluates at the end of every plan and apply. Failures are reported as warnings, not errors, so they never block a deployment. Typical uses: confirm a public endpoint returns HTTP 200 using the http data source, verify a certificate is not close to expiry, or check that a DNS record resolves. If a condition must stop the run, use a precondition or postcondition instead.

36. How do you test Terraform code with terraform test?

Answer: The native test framework (Terraform v1.6 and later) discovers .tftest.hcl files, usually in a tests/ directory. Each file has one or more run blocks that execute in sequence, set variables, and make assert statements against the result. command = plan gives fast unit-style tests that create nothing; the default command = apply creates real, short-lived infrastructure, asserts against it and destroys it afterwards, which makes it an integration test. Mock providers (from v1.7) let you fake provider responses so unit tests run without cloud credentials. A sensible pipeline runs fmt -check, validate, a linter such as TFLint, a security scanner such as Checkov or Trivy, plan-mode tests on every pull request, and apply-mode tests for modules in a sandbox account on release.

run "bucket_is_encrypted" {
  command = plan
  variables { name = "test-logs" }
  assert {
    condition     = aws_s3_bucket.this.bucket == "test-logs"
    error_message = "Bucket name not passed through"
  }
}

37. What are ephemeral resources, ephemeral values and write-only arguments?

Answer: They solve the long-standing problem of secrets landing in state. Terraform 1.10 introduced ephemeral values: ephemeral input variables and outputs, and ephemeral resource blocks that open or fetch something (a secret from AWS Secrets Manager or Azure Key Vault, a short-lived Kubernetes token) only for the duration of a run. Terraform does not store ephemeral resources in state or plan files. Terraform 1.11 added write-only arguments on managed resources, so an ephemeral value can be passed into, for example, a database password argument that is sent to the API but never persisted; you change a version argument alongside it to trigger an update. Ephemeral values can only be referenced in specific places (write-only arguments, provider configuration, other ephemeral blocks, locals), and availability depends on each provider implementing them.

Interview tip: Contrast with sensitive = true, which only redacts output. Ephemeral means the value is never written to state at all. Check the provider docs for which resources support write-only arguments.

Environments, workspaces and Stacks

38. How do you handle multiple environments?

Answer: The goal is separate state per environment, shared modules, and per-environment values. The most common production layout is a directory per environment (and per component) that calls versioned modules with its own backend configuration and .tfvars, ideally with each environment in its own cloud account or subscription. Differences between environments are explicit inputs (instance sizes, replica counts, deletion protection), never copy-pasted resource code. Promotion is a version bump of the module reference moving from dev to staging to prod, each with its own reviewed plan. Terragrunt reduces the boilerplate; HCP Terraform workspaces or Stacks are managed equivalents.

PR merged
   |
   v
dev plan --> auto apply --> tests
   |
   v
staging plan --> review --> apply
   |
   v
prod plan --> approval --> apply

39. Terraform CLI workspaces vs separate state: which do you prefer?

Answer: CLI workspaces (terraform workspace new staging) give one configuration several state files in the same backend, selected by terraform.workspace. They are handy for short-lived copies of identical infrastructure, such as a per-feature-branch test environment. For long-lived dev, staging and prod they are risky: all environments share the same backend and credentials, the active workspace is invisible in the code so applying to the wrong one is easy, and real environment differences end up as terraform.workspace == "prod" ? ... : ... conditionals scattered through the code. Separate directories (or separate HCP Terraform workspaces, which are a different concept from CLI workspaces) with separate backends and credentials give stronger isolation and clearer review.

AspectCLI workspacesSeparate state per env
IsolationSame backend and credentialsSeparate backend, account, role
Visibility in codeHidden (selected at runtime)Explicit directory or config
Env differencesConditionals on workspace nameExplicit tfvars inputs
Good forEphemeral, identical copiesLong-lived dev/staging/prod

40. What is HCP Terraform, and what are Terraform Stacks?

Answer: HCP Terraform is HashiCorp's managed service for running Terraform; it was called Terraform Cloud until April 2024. It provides remote state with locking, remote plan and apply runs, VCS-driven workflows, a private module registry, variable sets, policy enforcement with Sentinel or OPA, run tasks for third-party scanners, drift detection and team-based access. Terraform Enterprise is the self-hosted edition. Terraform Stacks are a newer HCP Terraform feature, now generally available, for managing infrastructure that is repeated across many deployments (environments, regions, accounts): you define component blocks that reference modules in .tfcomponent.hcl files and declare deployments in .tfdeploy.hcl files, and HCP Terraform orchestrates them, including ordering between components. Stacks are not available on legacy HCP Terraform team plans; check the current documentation for plan availability and for Terraform Enterprise support.

Secrets, policy as code, CI/CD and OpenTofu

41. How do you handle secrets in Terraform?

Answer: Rules in order of importance. Never commit secrets to .tf or .tfvars files. Do not give pipelines long-lived cloud keys; use OIDC federation so GitHub Actions, GitLab CI or Azure DevOps exchange a short-lived token for a scoped cloud role, or HCP Terraform's dynamic provider credentials. Let the platform generate and store secrets where possible (managed database passwords in Secrets Manager or Key Vault), and pass references, not values. Read secrets at runtime with ephemeral resources and feed them through write-only arguments so they never reach state. Mark any remaining secret variables and outputs sensitive, protect state as if it were a secret store, and run a secret scanner in CI. For the wider pipeline security picture, the DevSecOps interview questions guide covers scanning, signing and supply-chain controls.

42. What is policy as code, and how do Sentinel and OPA compare?

Answer: Policy as code means writing organisational rules (no public S3 buckets, only approved regions and instance families, mandatory cost-centre tags, no deletes of production databases without an exception) as code that is evaluated automatically against every plan, instead of relying on reviewers to remember. Sentinel is HashiCorp's policy language, built into HCP Terraform and Terraform Enterprise, with enforcement levels (advisory, soft-mandatory with override, hard-mandatory). OPA (Open Policy Agent) uses the Rego language; HCP Terraform supports OPA policy sets too, and outside HCP Terraform teams run OPA through Conftest against terraform show -json output in CI. Static scanners like Checkov, Trivy or TFLint check code before a plan exists; policy on the plan JSON sees the real computed values. Mature teams use both: fast static checks on every commit, plan-time policy as the gate.

Real-world example: Consider an insurer whose security team requires that nothing in production gets a public IP. A hard-mandatory policy on the plan blocks it in every workspace, and a documented exception process (a soft-mandatory override with an approver) handles the rare legitimate case.

43. Design a CI/CD pipeline for Terraform. What does a good plan review look like?

Answer: On every pull request: fmt -check, init, validate, linting, security scanning, terraform test in plan mode, then plan -out for each affected environment, with the plan summary posted as a PR comment and policy checks run on the plan JSON. On merge: apply the saved plan (or re-plan and compare) to dev automatically, and to production only after a required approval in a protected environment. Credentials come from OIDC, scoped per environment: the PR plan role is read-only where possible, and only the apply job on the main branch can assume the write role. Serialise applies per state (concurrency groups) and keep every plan and apply log. Tools in this space include GitHub Actions or GitLab CI, Atlantis (plan and apply driven by PR comments), HCP Terraform, and commercial platforms like Spacelift or env0.

A good plan review reads the summary line first (add, change, destroy counts), then every destroy and replacement, then IAM and network changes, then everything else. Reviewers should ask "is this change what the PR description says, and nothing more?" Unexpected changes usually mean drift or an unpinned provider.

PR opened
  |-- fmt, validate, lint, scan
  |-- terraform test (plan mode)
  |-- plan -out=tfplan  --> PR comment
  |-- policy check on plan JSON
merge to main
  |-- apply dev (auto)
  |-- apply prod (approval + OIDC role)

If you want hands-on practice building exactly this kind of pipeline with remote state, modules and policy checks, Cloudsoft's Terraform training in Hyderabad works through it in labs, in the Ameerpet classroom or live online.

44. What changed with Terraform's licence, and what is OpenTofu?

Answer: In August 2023 HashiCorp moved Terraform (and its other products) from the Mozilla Public License 2.0 to the Business Source License (BSL/BUSL 1.1) for future releases. The BSL permits most internal use but restricts building competing commercial products on the code, and its ambiguity worried vendors and some enterprises. In response, a group including Gruntwork, Spacelift, Harness, env0 and Scalr forked the last MPL-licensed version as OpenTofu, which was accepted into the Linux Foundation in September 2023, remains under MPL 2.0, and joined the CNCF as a sandbox project in 2025. IBM then completed its acquisition of HashiCorp in February 2025. Terraform itself is still actively developed by HashiCorp; the licence affects what vendors may build, not whether you can use Terraform to manage your own infrastructure.

45. OpenTofu vs Terraform: how would you choose for a new project?

Answer: They share HCL, the provider protocol and most workflow, and OpenTofu reads state from Terraform 1.5.x and earlier, so migration from that point is usually straightforward. They have diverged since: OpenTofu has features such as client-side state and plan encryption, early evaluation of variables in backend and module source configuration, and provider iteration with for_each, while Terraform has HCP Terraform integration, Stacks, Sentinel and its own newer language features. Decision factors: licence and open-governance requirements (some organisations mandate OSI-approved licences); whether you rely on HCP Terraform or Terraform Enterprise; which features you need; and your vendor's support model. Avoid flip-flopping on a live estate: state files written by newer versions of either tool may not be readable by the other, so pick one per estate, pin the version, and test any migration on a copy of state first.

Interview tip: Interviewers asking OpenTofu vs Terraform want a balanced, factual answer, not a side. Say what each offers, what would drive the decision in their organisation, and that you would check both projects' current docs for feature parity.

46. How do you manage Terraform and provider versions across many repositories?

Answer: Pin required_version for Terraform (for example ~> 1.16) and constraints for every provider in each root module, commit lock files, and use a version manager such as tfenv, tenv or asdf locally so the CLI matches the pipeline. Run the CLI in CI from a pinned container image. Upgrade on a schedule rather than when something breaks: a bot opens provider and module bump pull requests, the pipeline plans every affected environment, and an engineer reviews unexpected diffs. Upgrade the CLI one minor version at a time, reading the upgrade notes.

AI in IaC

47. How should AI-assisted Terraform code generation be used, and how do you review it?

Answer: AI coding assistants are useful for drafting modules, converting -generate-config-out output into clean code, writing .tftest.hcl tests, explaining an unfamiliar plan and writing policies. The risk is that generated HCL looks plausible while using deprecated arguments, outdated provider versions, overly broad IAM ("Action": "*"), public network exposure or missing encryption. Treat AI output exactly like a junior engineer's pull request: it goes through the same pipeline (validate, lint, security scan, policy on the plan, tests) and a human reads the plan before apply. Never give an assistant credentials that can apply to production. Grounding helps: HashiCorp publishes a Terraform MCP server that gives AI assistants current provider documentation, registry module details and Sentinel policies instead of stale training data, and your private module registry is a better source than generic examples. For the broader enterprise picture, see AI coding assistants in the enterprise.

Interview tip: The phrase interviewers want to hear is "the plan is the contract". Whoever or whatever wrote the code, the reviewed plan and the policy checks decide whether it ships.

48. How do you use Terraform to provision infrastructure for AI workloads?

Answer: AI platforms are still infrastructure, with a few extra concerns. Typical Terraform scope: GPU node groups or managed endpoints (with quotas checked per region before you plan), vector databases or PostgreSQL with pgvector, object storage for documents and model artefacts, private endpoints so model API traffic stays off the public internet, IAM roles scoped to specific model and knowledge-base resources, secrets for API keys handled with ephemeral values, and cost tags on every GPU and model resource so spend is attributable. Modules per layer (network, data, model access, application) keep the blast radius small, and new provider resources for fast-moving AI services sometimes lag behind the console, so you may temporarily need a generic provider such as AWS Cloud Control (awscc) or AzAPI (azapi). The Terraform for AI infrastructure guide walks through module structure, private endpoints and cost tagging in detail, and cloud cost optimisation for AI covers the spend side.

Real-world scenario questions

49. Scenario: an apply fails with "Error acquiring the state lock". What do you do?

Answer: First find out whether the lock is legitimate. The error shows the lock ID, the operation, who holds it and when. If a colleague or pipeline run is genuinely applying, wait. A lock is "stuck" only when the process that took it is gone, for example a CI runner killed mid-apply or a laptop that lost network. Only then run terraform force-unlock <LOCK_ID>, and then run a plan to see whether the interrupted apply left anything half-done.

What I would check:

  1. The lock info: who, which operation (plan or apply) and the timestamp.
  2. CI history: is a job for this state still running, queued or recently cancelled?
  3. Ask the person named in the lock before touching it.
  4. If the holder is confirmed dead, force-unlock with the exact lock ID (on S3-native locking, the lock object; on legacy DynamoDB, the table item).
  5. Run a plan immediately; if the interrupted run was an apply, expect partial changes and review them.
  6. Check the state file version history in case the interrupted write left it inconsistent.

Production consideration: Prevent repeats with pipeline concurrency controls (one run per state), timeouts that cancel gracefully, and restricting force-unlock to a break-glass role. Never force-unlock just because a plan is slow.

50. Scenario: someone changed a security group and an instance size in the console. The next plan shows unexpected changes. How do you handle it?

Answer: The plan is showing drift: Terraform wants to put things back the way the code says. Before applying, find out why the changes were made. A manual change during an incident might have been the fix, and reverting it blindly could bring the incident back.

What I would check:

  1. Run terraform plan -refresh-only to see exactly what changed outside Terraform.
  2. Use CloudTrail, the Azure Activity Log or the GCP audit logs to find who changed it and when, and link it to an incident or ticket.
  3. Decide per change: keep it (update the code to match, then the plan becomes empty) or revert it (apply the code).
  4. If another system legitimately owns an attribute, add a commented ignore_changes.
  5. Apply through the normal pull-request pipeline, not from a laptop.

Production consideration: Fix the cause: schedule drift-detection plans, restrict console write access in production to break-glass roles, and agree a rule that any emergency console change gets a follow-up PR within a day.

51. Scenario: you need to refactor a flat configuration into modules and switch from count to for_each without destroying anything. How?

Answer: Use moved blocks so every old address maps to its new address, and prove with the plan that there are zero adds and zero destroys before merging.

What I would check:

  1. List current addresses with terraform state list in each environment.
  2. Write the module and update the root to call it.
  3. Add a moved block per resource, for example from aws_subnet.private[0] to module.network.aws_subnet.private["ap-south-1a"]; mapping indexes to keys needs care.
  4. Run a plan in dev: the only lines should be "has moved to". Any create or destroy means a missing or wrong mapping, or a changed argument.
  5. Repeat in staging and prod; keep moved blocks until every state has applied.
moved {
  from = aws_subnet.private[0]
  to   = module.network.aws_subnet.private["a"]
}

Production consideration: Make refactoring PRs pure: no functional changes mixed in, so a clean "0 to add, 0 to change, 0 to destroy" plan is the acceptance test. Add prevent_destroy on stateful resources during the refactor as a seatbelt.

52. Scenario: a pipeline destroyed a production database by accident. Walk through your response.

Answer: Restore service first, then find out how a destroy got past review, then make it structurally impossible to repeat.

What I would check:

  1. Stop further runs on that state: pause the pipeline or lock the workspace.
  2. Restore from the latest automated backup, final snapshot or point-in-time recovery; this is why deletion protection, final snapshots and backup retention belong in the module defaults.
  3. Bring the restored database back under management: an import block to the same address, then a plan that should be empty or close to it.
  4. Read the plan and apply logs: was the destroy visible? Typical causes are a renamed resource without a moved block, a replace-forcing attribute change, a wrong workspace or tfvars file, or a deleted module call.
  5. Run a blameless post-incident review.

Production consideration: Layered controls: prevent_destroy and the cloud's own deletion protection on stateful resources; a policy that fails any production plan containing deletes of critical resources without an approved exception; mandatory human approval with the destroy count highlighted; separate state for data layers; and least-privilege apply roles that cannot delete backups.

53. Scenario: removing one item from a list makes the plan destroy and recreate five unrelated instances. Why, and how do you fix it?

Answer: The resources use count over a list. Removing an item from the middle shifts the index of every later element, so Terraform thinks instance [2] is now a different thing from what it was. The fix is to migrate to for_each with stable keys and add moved blocks from each old index to its new key, so the migration itself is a no-op plan. After that, removing an item removes exactly one instance. In the meantime, do not apply the plan; revert the list change.

54. Scenario: you join a team whose AWS estate was built by hand in the console. How do you bring it under Terraform?

Answer: Incrementally, by layer, with import blocks, and without trying to codify everything in one go.

What I would check:

  1. Inventory resources (tags, AWS Config or Resource Explorer) and agree ownership boundaries and state layout first.
  2. Start with low-risk, foundational layers such as VPCs, subnets and DNS zones, then data stores, then applications.
  3. Write import blocks (with for_each for repeated resources) and use -generate-config-out to draft configuration.
  4. Refactor the generated code into modules and variables, then iterate until the plan shows only imports and no changes.
  5. Only after a clean plan, tighten console permissions for those resources.

Production consideration: Generated configuration mirrors whatever is there, including bad settings. Import first with a no-change plan, then fix security issues in separate, reviewed PRs so each change has a clear plan. The AWS interview questions guide covers the underlying services you will meet in such an estate.

55. Scenario: a security review finds plain-text database passwords in your state files. What do you do?

Answer: Treat it as a credential exposure: rotate the secrets, restrict who can read state, and change the design so the secret is no longer persisted.

What I would check:

  1. Who and what has read access to the state bucket or container, including old versions, CI logs and laptops with local copies.
  2. Rotate the exposed credentials.
  3. Redesign: platform-managed master passwords, or ephemeral resources plus write-only arguments where the provider supports them, so new state no longer contains the value.
  4. Tighten bucket policies and enable access logging; consider whether old state versions that contain the secret need to be purged.

Production consideration: Add a check in CI that scans terraform show -json output for known secret attributes, and make "no secrets in state" a module review criterion.

56. Scenario: after a major provider upgrade, the plan wants to change dozens of resources. What do you do?

Answer: Do not apply. Major provider versions rename arguments, change defaults and split resources (for example S3 bucket settings moving into separate resources in an earlier AWS provider major version), and the plan reflects that.

What I would check:

  1. Read the provider's upgrade guide for that major version.
  2. Classify each diff: cosmetic (defaults now shown explicitly), argument renames, or real changes.
  3. Update code to the new schema, using moved blocks or import blocks where resources were split.
  4. Re-plan until the diff is empty or contains only understood, harmless changes.
  5. Roll out per environment; revert the lock file if you need to back out.

Production consideration: Upgrade one major version at a time, in its own PR, never combined with feature changes.

57. Scenario: plans on your main configuration have become very slow and often hit API rate limits. How do you fix it?

Answer: The state is too big. Every plan refreshes every resource, which means thousands of API calls. Short-term relief: -refresh=false for quick local checks (never for the plan that will be applied), lower -parallelism if throttled, and -target only for emergencies, since targeted applies leave the rest unreconciled. The real fix is splitting the state along ownership and change-frequency lines (Q17): move resources into new configurations with removed blocks (destroy = false) in the old one and import blocks in the new one, validated by no-change plans on both sides, and replace direct references with data sources or published outputs.

58. Scenario: a reviewed plan was approved, but by the time it applied, another merge had changed the same state. What can go wrong and how do you prevent it?

Answer: If the pipeline re-plans at apply time, it may apply changes nobody reviewed. If it applies the saved plan, Terraform rejects it as stale when the state has changed since the plan was created, which is the safe outcome. Prevention: always apply saved plans, serialise runs per state with concurrency groups or workspace queues, require branches to be up to date before merge, and re-plan and re-review if the saved plan is rejected.

59. Scenario: an apply fails halfway through, after some resources were created. What state are you in, and what next?

Answer: Terraform writes state as it goes, so resources that were created successfully are recorded and those that failed are not, or are marked tainted if creation partly succeeded. Terraform does not roll back. Read the error (often a quota, a permissions gap, a name collision or an API timeout), fix the cause, and run a new plan: it will show only the remaining work plus replacement of tainted resources. If the state write itself failed, Terraform saves an errored.tfstate locally; push it carefully after checking it against the remote version.

60. Scenario: your company wants to evaluate moving from Terraform to OpenTofu. How would you run the evaluation?

Answer: As an engineering decision with a test plan, not a debate.

What I would check:

  1. The drivers: licence policy, governance preference, cost of HCP Terraform, or need for OpenTofu-only features such as state encryption.
  2. Dependencies on HCP Terraform, Terraform Enterprise, Sentinel or Stacks that would need replacing (for example with OPA and another runner).
  3. The Terraform versions in use and which language features each configuration relies on, against current OpenTofu compatibility notes.
  4. A pilot on a non-production configuration using a copy of state: init, plan, expect an empty plan.
  5. Pipeline, tooling (linters, scanners, IDE plugins) and provider registry behaviour.

Production consideration: Migrate one state at a time with a rollback plan, back up state before the first OpenTofu write, and freeze tool versions during the transition. The same evaluation approach applies in Azure-heavy estates; the Azure interview questions guide covers that platform side.

Key takeaways

  • State is the heart of Terraform: remote, encrypted, versioned, locked, split by ownership, and never edited by hand.
  • Know the current details: S3-native locking with use_lockfile, HCP Terraform (previously Terraform Cloud), ephemeral values and write-only arguments, Stacks now GA.
  • Prefer for_each with stable keys over count for distinct objects, and use declarative import, moved and removed blocks instead of state surgery.
  • Separate state and credentials per environment beat CLI workspaces for long-lived environments.
  • The reviewed, saved plan plus policy as code is the deployment contract, whether a human or an AI assistant wrote the code.
  • For OpenTofu vs Terraform, give a balanced answer grounded in licence, governance, platform dependencies and features.
  • In scenarios, slow down: confirm the facts, protect production first, then fix the cause.

Interview preparation checklist

  • Write a small configuration from memory: provider block with version constraints, a resource, a data source, variables with validation, outputs.
  • Set up an S3 (or Azure Blob) backend with encryption, versioning and locking, and deliberately trigger a lock conflict to see the error.
  • Build one module with typed inputs, a README, examples and a .tftest.hcl file; publish a tagged version and consume it.
  • Convert a count resource to for_each using moved blocks and get a no-change plan.
  • Import existing resources with an import block and -generate-config-out.
  • Make a console change, then detect it with plan -refresh-only and decide how to reconcile it.
  • Build a GitHub Actions pipeline with OIDC, a saved plan, a PR comment and an approval gate for prod.
  • Write one policy (Sentinel or OPA/Conftest) that blocks public buckets or unapproved regions.
  • Prepare a two-minute, balanced explanation of the BSL licence change and OpenTofu.
  • Rehearse the four classic scenarios: stuck lock, drift, refactor without destroy, accidental destroy.

FAQ

What skills are required for a Terraform interview in 2026?

You need solid HCL, a clear understanding of state and remote backends, module design and versioning, for_each and lifecycle rules, multi-environment layouts, CI/CD with plan review, secrets handling and policy as code, plus working knowledge of at least one cloud such as AWS or Azure.

How should I prepare for Terraform interview questions?

Build and break real infrastructure in a free-tier or sandbox account: set up a remote backend, write and version a module, import resources, refactor with moved blocks, and run everything through a pipeline. Then practise explaining each scenario out loud in a clear order.

Is Terraform still worth learning after the licence change?

Yes. Terraform remains widely used for managing cloud infrastructure, the licence change mainly affects vendors building competing products, and almost everything you learn transfers directly to OpenTofu because they share HCL, providers and workflow.

Should I learn Terraform or OpenTofu?

Learn the shared core first, since the language and workflow are largely the same. Then learn the differences that matter for the employers you target, such as HCP Terraform and Sentinel on one side and state encryption on the other.

Do I need a Terraform certification to get a DevOps job?

A certification can help you structure study and pass resume filters, but interviewers mostly test hands-on reasoning about state, modules and incidents. A public repository with a tested module and a working pipeline is strong evidence either way.

Which cloud should I learn alongside Terraform?

Pick the cloud most common in the roles you are applying for, often AWS or Azure in Indian GCCs and services firms, and learn its networking, IAM and storage well, because most Terraform interview questions are really about those resources.

How much coding do Terraform interviews involve?

Expect to read and write HCL rather than general programming: for_each and for expressions, a dynamic block, variable validation and a module call. Some roles add a little Bash or Python for pipeline scripting.

Are freshers asked Terraform interview questions?

Yes, for DevOps and cloud trainee roles. Freshers are usually asked fundamentals: what IaC is, the init-plan-apply workflow, providers, state and modules, often with a small hands-on task.

How is Terraform used with Kubernetes?

Terraform commonly provisions the cluster, node groups, networking and IAM, while workloads inside the cluster are usually deployed with Helm, Kubernetes manifests or GitOps tools such as Argo CD, keeping each layer's ownership clear.

If you want to practise these Terraform interview questions on real labs, from remote state and module versioning to OIDC pipelines and policy checks, explore Cloudsoft's Terraform course in Ameerpet and online. For a broader path that combines cloud, DevOps, AI and security engineering, look at the APEX AI, ML, Cloud and Cyber Security program. Call +91 96660 19191 to book a free demo.

New ยท AI Career Guide

Meet Aanya โ€” ask anything about courses, fees & placement

Instant answers from verified Cloudsoft info โ€” courses, fees, formats, placement support and free demos. Available 24/7, right here on the site.

How Aanya works โ†’
Share๐•infโœ‰
EnrollWhatsAppCall us