New batches starting this week ยท Limited seats

Linux Interview Questions and Answers 2026 (70 Questions)

70 Linux interview questions with accurate commands and practical answers for admin, DevOps, cloud and SRE roles, from permissions and systemd to LVM, networking, Bash, SELinux, cgroups and thirteen real troubleshooting scenarios.

Linux interview questions and answers 2026 - Cloud Soft Solutions
Last updated ยท 47 min read ยท 10,381 words

These Linux interview questions cover what system administrator, DevOps, cloud and SRE interviews in 2026 actually test: the commands you type every day, the concepts behind them, and how you reason through a broken server. Interviewers rarely accept a list of commands; they hand you a symptom such as "disk full", "load average of 40" or "permission denied" and watch how you narrow it down. This guide gives 70 high-value questions with accurate commands, short explanations and the follow-ups that usually come next.

How to use this guide:

  • Freshers and support engineers are usually tested on the fundamentals: filesystem layout, permissions, users, processes, package managers and everyday commands. Practise typing them on a real VM, not reading them.
  • Linux admin interview rounds for mid-level roles focus on systemd, journalctl, disks, LVM, fstab, networking, SSH and text processing pipelines.
  • Linux for DevOps interview and SRE rounds go further: shell scripting discipline, performance analysis, SELinux/AppArmor, sysctl, namespaces and cgroups, plus several Linux troubleshooting interview scenarios.
  • For every answer, say the one-line answer first, then the command, then what could go wrong. That order sounds like experience.

Contents

Fundamentals: filesystem, permissions, users, packages

1. Explain the Linux filesystem hierarchy. Where do configuration, logs and binaries live?

Answer: Linux has a single tree rooted at /, laid out by the Filesystem Hierarchy Standard. The directories you must know: /etc (system configuration), /var (variable data: /var/log, /var/lib for service state such as databases and container images, /var/cache), /usr (installed software: /usr/bin, /usr/lib; on modern distributions /bin and /sbin are symlinks into /usr), /home and /root (user homes), /tmp (temporary, often cleared at boot), /opt (third-party packages), /boot (kernel and initramfs), /dev (device files), and the virtual filesystems /proc and /sys, which expose kernel and process state as files.

2. What is an inode, and what is the difference between a hard link and a symbolic link?

Answer: An inode stores a file's metadata (owner, permissions, size, timestamps, pointers to data blocks) but not its name. A directory entry maps a name to an inode number. A hard link (ln file link) is another name for the same inode: deleting one name leaves the data intact until the link count reaches zero and no process holds it open. Hard links cannot cross filesystems and normally cannot point to directories. A symbolic link (ln -s target link) is a separate small file containing a path; it can cross filesystems and point to directories, but it breaks if the target moves. Use ls -li to see inode numbers and link counts, and stat file for full metadata.

3. How do Linux file permissions work? Explain chmod in symbolic and octal form.

Answer: Every file has an owner, a group and three permission sets (user, group, others), each with read (4), write (2) and execute (1). ls -l shows them as -rwxr-x---. Octal adds the bits per set: chmod 750 deploy.sh gives the owner rwx, the group r-x, others nothing. Symbolic form changes bits relative to the current state: chmod u+x,g-w,o= file. On directories the meanings differ: r lists names, w creates, deletes and renames entries, and x lets you enter the directory and access files inside it by name. chmod -R applies recursively; chmod -R u=rwX,go=rX dir (capital X) adds execute only to directories and files that are already executable.

Interview tip: Say why chmod 777 is the wrong fix: it hides the real problem (ownership, a parent directory, SELinux) and makes the file writable by any process on the box.

4. What do chown and chgrp do, and who is allowed to run them?

Answer: chown user:group file changes the owner and group; chown -R app:app /srv/app does it recursively; chgrp devs file changes only the group. Only root (or a process with CAP_CHOWN) can give a file away to another user.

5. What is umask, and what permissions does a new file get with umask 022?

Answer: umask is a mask of permission bits removed from newly created files and directories. With 022, new files get 644 (rw-r--r--) and directories 755. With 027, files get 640 and directories 750, so "others" get nothing, which is common on hardened servers. Check it with umask (or umask -S for symbolic). Set it per shell in profile files, per service with UMask= in a systemd unit, or system-wide via the distribution's login defaults.

6. Explain SUID, SGID and the sticky bit with examples.

Answer: They are special permission bits with octal values 4, 2 and 1 in a fourth digit placed first.

  • SUID (4) on an executable runs it with the file owner's privileges. /usr/bin/passwd is SUID root so users can update /etc/shadow. Shown as s in the user execute slot (-rwsr-xr-x).
  • SGID (2) on an executable runs it with the file's group; on a directory, new files inherit the directory's group. chmod 2775 /srv/shared is the classic shared team folder.
  • Sticky bit (1) on a directory means only the file owner, directory owner or root can delete or rename files inside. /tmp is 1777 (drwxrwxrwt).

A capital S or T means the special bit is set but the underlying execute bit is not. Security teams audit SUID binaries with find / -xdev -perm -4000 -type f because an unexpected SUID root binary is a privilege-escalation path.

7. What are ACLs, and when do you need them instead of standard permissions?

Answer: POSIX ACLs give permissions to additional named users and groups beyond the single owner and group. Use them when, for example, the app user owns a log directory but the monitoring user also needs read access and you don't want to change group ownership. Commands: setfacl -m u:monitoring:rx /var/log/app, setfacl -d -m u:monitoring:rx /var/log/app (default ACL so new files inherit it), getfacl /var/log/app to view, setfacl -b to remove all. A + at the end of the ls -l permission string shows an ACL exists. The ACL mask caps the effective permissions of named entries and the group, which is a frequent cause of "I added the ACL but it still doesn't work".

8. Where are users and groups stored, and how do you create, modify and lock a user?

Answer: /etc/passwd holds accounts (name, UID, GID, home, shell), /etc/shadow holds password hashes and ageing (root-readable only), /etc/group holds groups and members. Query them through NSS with getent passwd alice so LDAP or SSSD users also show up. Commands: useradd -m -s /bin/bash alice (or the friendlier adduser on Debian/Ubuntu), passwd alice, usermod -aG docker alice, usermod -L alice or passwd -l alice to lock, chage -l alice for password ageing, userdel -r alice to remove with home directory. Check identity with id alice, whoami and groups.

Interview tip: Point out that usermod -G docker alice without -a replaces all supplementary groups, and that new group membership only applies to new login sessions.

9. How does sudo work, and how do you grant limited sudo access safely?

Answer: sudo runs a command as another user (root by default) according to rules in /etc/sudoers and /etc/sudoers.d/, and logs each use. Always edit with visudo (or visudo -f /etc/sudoers.d/deploy) because it syntax-checks before saving; a broken sudoers file can lock everyone out of root. Grant the narrowest rule: deploy ALL=(root) NOPASSWD: /usr/bin/systemctl restart myapp rather than ALL. Members of the sudo group (Debian/Ubuntu) or wheel group (RHEL family) usually get full access. sudo -l shows what the current user may run; sudo -i gives a root login shell. Avoid allowing editors, shells or tar/find with arbitrary arguments, because they can spawn a root shell.

10. Compare apt and dnf. How do you install, find which package owns a file, hold a version and roll back?

Answer: apt (with dpkg underneath) manages .deb packages on Debian and Ubuntu; dnf (with rpm underneath) manages .rpm packages on RHEL, Rocky, AlmaLinux and Fedora, replacing yum.

Taskapt / dpkgdnf / rpm
Refresh metadata and upgradeapt update && apt upgradednf upgrade
Install / removeapt install nginx / apt remove nginxdnf install nginx / dnf remove nginx
Which package owns a filedpkg -S /usr/sbin/nginxrpm -qf /usr/sbin/nginx
Which package provides a missing commandapt-file search bin/digdnf provides '*/dig'
List files in a packagedpkg -L nginxrpm -ql nginx
Show versions and source repoapt policy nginxdnf info nginx
Pin a versionapt-mark hold nginxdnf versionlock add nginx (plugin)
Undo a transactionreinstall a specific versiondnf history then dnf history undo N

11. Which everyday commands should you be able to use without thinking?

Answer: Interviewers often run a quick-fire round. Be fluent with these and their common flags:

AreaCommands
Files and directoriesls -lah, cd, pwd, mkdir -p, cp -a, mv, rm -r, touch, cat, less, head, tail -f, wc -l, diff -u, file
System informationuname -r, hostnamectl, uptime, lscpu, nproc, whoami, id, cat /etc/os-release, dmesg -T
Archivestar -czf app.tgz dir/, tar -xzf app.tgz -C /opt, tar -tzf to list, gzip/gunzip, zip/unzip, zstd
Copying between hostsscp, rsync -avz --delete src/ host:/dst/ (the trailing slash matters), wget, curl -O
Shell convenienceshistory, Ctrl-R, alias, which/type, man, echo $?

Interview tip: Mention that ifconfig, netstat and route come from the deprecated net-tools package and are often not installed on minimal images; the modern equivalents are ip and ss.

Processes, signals, systemd and journalctl

12. How do you list processes and find the ones using the most CPU or memory?

Answer: ps aux (BSD style) or ps -ef (System V style) show everything. For targeted output: ps -eo pid,ppid,user,stat,%cpu,%mem,etime,cmd --sort=-%mem | head. pgrep -a nginx finds by name, pstree -p shows parent-child relationships, and /proc/PID/ holds details (cmdline, status, fd/, limits). The STAT column matters: R running, S interruptible sleep, D uninterruptible sleep (usually waiting on I/O), Z zombie, T stopped.

13. What does top show, and which fields do you read first?

Answer: The header shows uptime and load average, task counts by state, CPU breakdown and memory. In the CPU line read us (user), sy (kernel), wa (waiting on I/O), st (steal: time the hypervisor gave your vCPU to someone else, important on cloud VMs) and id. Interactive keys: P sort by CPU, M by memory, 1 per-CPU view, c full command, k kill. In the process list, RES is resident memory and VIRT is virtual address space, which is usually much larger and rarely the number to worry about.

14. Explain the common signals. What is the difference between kill -15 and kill -9?

Answer: A signal is an asynchronous notification to a process. SIGTERM (15, the default for kill) asks the process to exit; it can catch it, finish requests, flush data and clean up. SIGKILL (9) is handled by the kernel and cannot be caught, so there is no cleanup: open transactions, temp files and lock files may be left behind. Others to know: SIGHUP (1, many daemons reload configuration), SIGINT (2, Ctrl-C), SIGSTOP/SIGCONT (pause and resume), SIGCHLD (child exited). Commands: kill -TERM 1234, pkill -HUP nginx, killall name, kill -l to list. The professional sequence is TERM, wait, then KILL only if needed, which is also what systemctl stop and container runtimes do.

15. What are nice and renice, and how do you lower the impact of a heavy job?

Answer: Niceness ranges from -20 (highest priority) to 19 (lowest) and influences the CPU scheduler's share for a process. nice -n 10 ./backup.sh starts a job at lower priority; renice -n 15 -p 4321 changes a running one. Ordinary users can only make processes nicer; raising priority needs root. For disk-heavy jobs CPU niceness is not enough, so also use ionice -c3 (idle class) or ionice -c2 -n7. For firmer control, run the job in its own cgroup: systemd-run --scope -p CPUWeight=20 -p MemoryMax=2G ./job.sh (a CPUQuota= property sets a hard CPU cap).

16. What are zombie and orphan processes, and how do you deal with them?

Answer: A zombie (Z) has exited but its parent hasn't read its exit status with wait(). It uses no CPU or memory beyond a process-table entry, and you can't kill it because it is already dead. The fix is the parent: signal it to reap children, or restart it so the zombies get re-parented to PID 1, which reaps them. Find the parent with ps -o ppid= -p ZOMBIE_PID. An orphan is a running process whose parent died; it is adopted by PID 1 (or a subreaper) and is normally harmless.

17. How do jobs, fg, bg, nohup and disown work, and what do you use instead for long-running work?

Answer: Append & to run a command in the background; jobs lists shell jobs; Ctrl-Z suspends the foreground job; bg %1 resumes it in the background and fg %1 brings it back. When the terminal closes, the shell sends SIGHUP to its jobs; nohup cmd & ignores it and writes output to nohup.out, and disown removes a job from the shell's table. For anything you need to reconnect to, use tmux or screen. For anything that must survive reboots and restart on failure, write a systemd service rather than relying on a backgrounded shell process.

18. What is systemd, and what is the difference between systemctl start, enable, restart, reload and mask?

Answer: systemd is PID 1 on most current distributions: it boots the system, supervises services, manages dependencies, mounts, timers, sockets and cgroups, and collects logs through journald. Its objects are units (.service, .socket, .timer, .mount, .target).

  • systemctl start nginx starts it now; enable makes it start at boot (creates symlinks under a target's .wants); enable --now does both.
  • restart stops and starts; reload asks the service to re-read config without dropping connections (only if the unit defines ExecReload); reload-or-restart picks whichever is supported.
  • mask links the unit to /dev/null so nothing can start it, even as a dependency.
  • systemctl status, is-active, is-enabled, list-units --failed, cat (shows the unit and its overrides) and daemon-reload after editing unit files.

19. Write a systemd service unit for a Python API and explain each directive.

Answer: Put custom units in /etc/systemd/system/; never edit vendor files in /usr/lib/systemd/system/ (use systemctl edit unit to create a drop-in override).

[Unit]
Description=Orders API
After=network-online.target
Wants=network-online.target

[Service]
User=orders
WorkingDirectory=/srv/orders
EnvironmentFile=/etc/orders/env
ExecStart=/srv/orders/venv/bin/uvicorn app:app \
  --host 0.0.0.0 --port 8000
Restart=on-failure
RestartSec=5
LimitNOFILE=65536
NoNewPrivileges=true

[Install]
WantedBy=multi-user.target

After/Wants order the start after the network is up; User drops root; EnvironmentFile keeps secrets out of the unit (restrict its permissions); Restart=on-failure with RestartSec gives supervision without a tight crash loop; LimitNOFILE raises the open-file limit (shell ulimit settings do not apply to services); WantedBy ties enable to normal multi-user boot. Then run systemctl daemon-reload && systemctl enable --now orders.

Interview tip: Mention hardening directives such as ProtectSystem=strict, PrivateTmp=true and ReadWritePaths=, and systemd-analyze security orders to score a unit.

20. How do you read logs with journalctl?

Answer: journald stores structured logs from the kernel, services and syslog. Commands you should know:

  • journalctl -u nginx -f: follow one unit.
  • journalctl -u nginx --since "1 hour ago" --until "10 min ago": time window.
  • journalctl -p err -b: errors and worse since this boot; -b -1 is the previous boot, essential after a crash or unexpected reboot.
  • journalctl -k: kernel messages (OOM kills, disk errors, NIC resets).
  • journalctl -u app -o json-pretty: structured fields; _PID=1234 filters by field.
  • journalctl --disk-usage and journalctl --vacuum-size=500M to control size.

If -b -1 shows nothing, the journal is probably volatile (kept in /run); set Storage=persistent in /etc/systemd/journald.conf or create /var/log/journal.

Disks, filesystems, LVM, memory and swap

21. What is the difference between df and du, and why can they disagree?

Answer: df -h asks the filesystem how many blocks are used and free; du -sh dir walks the directory tree and adds up file sizes it can see. They disagree when space is used by something du cannot see: deleted files still held open by a process (Q58), files hidden underneath a mount point, or filesystem metadata and reserved blocks (ext4 reserves a percentage for root by default). Useful forms: du -xh --max-depth=1 / | sort -h (stay on one filesystem), du -sh /var/* 2>/dev/null | sort -h | tail, and df -hT to show filesystem types. ncdu is handy interactively if installed.

22. How do you see block devices, filesystems and mounts?

Answer: lsblk -f shows disks, partitions, LVM volumes, filesystem types, UUIDs and mount points as a tree. blkid prints UUIDs and types. findmnt shows the mount tree with options; findmnt /data checks one path. To mount: mount /dev/nvme1n1 /data; to unmount: umount /data (if "target is busy", find the culprit with lsof +D /data or fuser -vm /data). mount -o remount,ro /data changes options in place.

23. Explain /etc/fstab. How do you add a disk safely so a mistake doesn't stop the server booting?

Answer: Each line has six fields: device, mount point, type, options, dump, fsck order. Example: UUID=3f1c... /data xfs defaults,nofail 0 2. Safe procedure:

  1. Create the filesystem (mkfs.xfs /dev/nvme1n1) and get its UUID with blkid.
  2. Use the UUID, not /dev/sdX or /dev/nvmeXn1.
  3. Add nofail for non-root data disks so boot continues if the disk is missing (important for detachable cloud volumes).
  4. Run systemctl daemon-reload, then mount -a and findmnt --verify before rebooting. If mount -a errors, the next boot would probably have dropped into emergency mode (see Q70).

24. What is LVM, and how do you extend a logical volume online?

Answer: LVM adds a layer between disks and filesystems: physical volumes (PVs) are pooled into a volume group (VG), from which you carve logical volumes (LVs) that can be resized, snapshotted and moved without repartitioning. Inspect with pvs, vgs, lvs. To grow /var with a new disk:

pvcreate /dev/nvme2n1
vgextend vg_data /dev/nvme2n1
lvextend -r -L +20G /dev/vg_data/lv_var

-r resizes the filesystem as well (it calls xfs_growfs or resize2fs). Both XFS and ext4 can grow while mounted; XFS cannot shrink at all, and ext4 shrinks only offline. If the VG already has free space, skip the first two steps.

Real-world example: On a cloud VM without LVM, the equivalent is: enlarge the volume in the console, then growpart /dev/nvme0n1 1 and xfs_growfs / or resize2fs /dev/nvme0n1p1.

25. What does "No space left on device" mean when df -h shows free space?

Answer: The filesystem has run out of inodes, not blocks. Every file needs an inode, so millions of tiny files (session files, mail queues, cache shards, build artefacts) can exhaust them while gigabytes remain free. Check with df -i. Find the directory with the most files: du --inodes -x / 2>/dev/null | sort -n | tail (GNU du), or find /var -xdev -type f | cut -d/ -f2-3 | sort | uniq -c | sort -n | tail. Fix by deleting or archiving the small files and fixing whatever creates them. On ext4 the inode count is fixed at creation; XFS allocates inodes dynamically, which is one reason it handles this better.

26. How do you read free -h? Is a server with little "free" memory in trouble?

Answer: Not necessarily. Linux uses spare RAM as page cache (buff/cache) and gives it back when applications need it. The column that matters is available: the kernel's estimate of memory that can be allocated without swapping. A server with low free but healthy available is normal. Real memory pressure shows as low available, swap usage that keeps growing, non-zero si/so in vmstat, and OOM kills in the kernel log. For per-process detail, look at RSS in ps or top, or smem if installed; /proc/meminfo has the raw numbers.

27. What is swap, how do you add a swap file, and what does vm.swappiness control?

Answer: Swap is disk space the kernel uses to page out memory that isn't actively used, freeing RAM for active work. It buys time under pressure but heavy swapping makes everything slow. Add a swap file:

fallocate -l 4G /swapfile   # or dd on some filesystems
chmod 600 /swapfile
mkswap /swapfile
swapon /swapfile
echo '/swapfile none swap sw 0 0' >> /etc/fstab

Check with swapon --show. vm.swappiness (0โ€“200 on recent kernels, default 60) tunes how readily the kernel swaps anonymous memory versus dropping page cache; database hosts often lower it.

28. What does vmstat tell you, and how do you read its columns?

Answer: vmstat 1 5 prints one line per second (the first line is an average since boot, so ignore it). Columns: r runnable processes (compare with CPU count), b processes blocked in uninterruptible I/O, swpd/free/buff/cache memory, si/so swap in and out per second (sustained non-zero values mean memory pressure), bi/bo block I/O, in/cs interrupts and context switches, and CPU us sy id wa st.

29. How does the OOM killer work, and how do you find out it killed your process?

Answer: When the kernel cannot reclaim enough memory, the OOM killer picks a victim, mainly by memory usage adjusted by oom_score_adj (-1000 to 1000), and sends SIGKILL. Evidence: journalctl -k | grep -i -E 'out of memory|killed process' or dmesg -T | grep -i oom, which lists the victim, its memory and the cgroup. Inspect a process's score at /proc/PID/oom_score. There are two flavours: global OOM (the whole host ran out) and cgroup OOM (a container or systemd service hit its own memory.max while the host still had RAM). A container killed this way exits with code 137 (128 + 9). Protect critical daemons with OOMScoreAdjust=-900 in their unit, but the real fix is right-sizing limits and finding leaks. For the AI-workload version of this problem, the Linux for AI engineers guide walks through an OOM-killed inference service.

Networking, firewalls and SSH

30. Which ip commands do you use to check addresses, routes and links?

Answer: ip -br addr (brief view of interfaces and addresses), ip addr show eth0, ip link set eth0 up, ip route (routing table and default gateway), ip route get 10.20.0.15 (which interface and gateway the kernel will actually use for a destination, very useful with multiple NICs or VPNs), ip neigh (ARP/neighbour table) and ip -s link (packet and error counters).

31. How do you find which process is listening on a port?

Answer: ss -tulpn lists TCP and UDP listening sockets with the owning process (run as root to see other users' processes). ss -tlnp 'sport = :8080' filters one port; lsof -i :8080 also works. Check the local address column: 127.0.0.1:8080 accepts only local connections, 0.0.0.0:8080 or [::]:8080 accepts on all interfaces. ss replaces netstat.

32. How do you troubleshoot DNS with dig?

Answer: dig example.com shows the answer, TTL and which server responded; dig +short example.com A just the records; dig @1.1.1.1 example.com queries a specific resolver to compare with your default; dig +trace example.com follows delegation from the root, useful after a nameserver change; dig -x 10.0.0.5 does a reverse lookup; dig example.com MX, TXT or CNAME for other types. Remember applications resolve through NSS (/etc/nsswitch.conf, /etc/hosts, then DNS), so test what the app sees with getent hosts example.com. On systemd-resolved systems, resolvectl status shows the real upstream servers, because /etc/resolv.conf points at the local stub 127.0.0.53.

33. How do you use curl to debug an HTTP service?

Answer: curl -v https://api.example.com/health shows DNS resolution, TCP connect, TLS handshake, request and response headers. curl -I sends a HEAD request. Timing breakdown: curl -s -o /dev/null -w '%{http_code} dns=%{time_namelookup} connect=%{time_connect} tls=%{time_appconnect} total=%{time_total}\n' URL. curl --resolve api.example.com:443:10.0.1.20 https://api.example.com/ tests a specific backend while keeping the correct hostname for TLS and virtual hosting. Avoid -k outside debugging, because it disables certificate checks. For raw port reachability without HTTP, nc -zv host 5432 replaces the old telnet host port trick.

34. Compare iptables, nftables, firewalld and ufw.

Answer: All of them configure the kernel's netfilter packet filter. iptables is the older rule syntax with tables (filter, nat, mangle) and chains (INPUT, OUTPUT, FORWARD); on current distributions the iptables command is often iptables-nft, a compatibility layer writing nftables rules. nftables (nft) is the modern framework with one syntax for IPv4, IPv6 and more, and atomic rule-set loading (nft list ruleset). firewalld is the RHEL-family manager with zones and services: firewall-cmd --permanent --add-service=https then firewall-cmd --reload; without --permanent, changes are lost on reload. ufw is Ubuntu's simple front-end: ufw allow 22/tcp, ufw status verbose. Pick one manager per host; mixing them produces rules that overwrite each other.

Interview tip: In the cloud, also mention the network layer outside the VM (security groups, network security groups, NACLs). A packet must pass both.

35. How do you set up SSH key authentication and harden sshd?

Answer: Generate a key with ssh-keygen -t ed25519 -C "alice@laptop" (protect it with a passphrase and use ssh-agent), then install the public key with ssh-copy-id user@host or by appending it to ~/.ssh/authorized_keys. Permissions matter: ~/.ssh 700, authorized_keys 600, and the home directory must not be group- or world-writable, or sshd ignores the keys (with StrictModes, the default). Harden /etc/ssh/sshd_config or a file in sshd_config.d/: PasswordAuthentication no, PermitRootLogin no (or prohibit-password), AllowGroups ssh-users. Always validate with sshd -t and keep an existing session open while you reload, so a mistake doesn't lock you out.

36. What can you do with ~/.ssh/config and ProxyJump?

Answer: The client config turns long commands into short aliases and encodes how to reach private hosts through a bastion:

Host bastion
  HostName 203.0.113.10
  User ec2-user
  IdentityFile ~/.ssh/prod_ed25519

Host app-*
  User ubuntu
  ProxyJump bastion
  ServerAliveInterval 30

Now ssh app-1 hops through the bastion automatically (one-off equivalent: ssh -J bastion ubuntu@10.0.2.15). ProxyJump is safer than agent forwarding (-A) because your private key and agent never become usable from the bastion.

37. Explain SSH local, remote and dynamic port forwarding.

Answer:

  • Local (-L): ssh -N -L 5433:db.internal:5432 bastion makes localhost:5433 on your laptop reach the private database through the bastion. Typical use: connecting a SQL client to a database in a private subnet.
  • Remote (-R): ssh -N -R 9000:localhost:3000 server exposes your local port 3000 as port 9000 on the server. Used for webhooks during development; usually restricted by policy.
  • Dynamic (-D): ssh -N -D 1080 bastion creates a SOCKS proxy so a browser can reach internal web consoles.

-N means no remote command; -f backgrounds it. Tunnels bypass network controls, so many enterprises disable forwarding in sshd (AllowTcpForwarding no) except on approved bastions.

If you want to practise these commands on real cloud servers with a mentor reviewing your troubleshooting, Cloudsoft's Multi-cloud DevOps with Linux & Python course covers Linux administration, shell and Python automation across AWS, Azure and Google Cloud, with classroom batches in Ameerpet and live online sessions.

Text processing, shell scripting and scheduling

38. Show grep options you use daily.

Answer: grep -i error app.log (case-insensitive), grep -rn "DB_HOST" /etc/app/ (recursive with line numbers), grep -v DEBUG (invert), grep -c 500 (count lines), grep -E 'timeout|refused' (extended regex), grep -o 'user_id=[0-9]*' (only the match), grep -A 3 -B 2 Exception (context lines), grep -l (file names only), grep -w (whole word), grep -F (fixed string, faster and no regex surprises).

39. How do you use sed for in-place edits and line selection?

Answer: sed 's/old/new/g' file substitutes and prints; sed -i.bak 's/^Port 22$/Port 2222/' /etc/ssh/sshd_config edits in place keeping a backup. sed -n '100,120p' file prints a range; sed '/^#/d;/^$/d' file strips comments and blank lines; sed -i '/^PasswordAuthentication/c\PasswordAuthentication no' file replaces a whole matching line. Use a different delimiter for paths: sed 's|/opt/old|/opt/new|g'. -E enables extended regex. For repeatable configuration changes across many servers, a configuration management tool beats ad hoc sed; see the Ansible interview questions.

40. What can awk do that cut can't? Give examples.

Answer: cut slices fixed delimiters (cut -d: -f1,7 /etc/passwd, cut -c1-10) but cannot handle runs of spaces or apply logic. awk splits on whitespace by default, supports conditions, variables and arithmetic:

  • awk -F: '$3 >= 1000 {print $1}' /etc/passwd: regular users.
  • awk '$9 >= 500 {print $7}' access.log: paths returning 5xx in a common log format.
  • awk '{bytes += $10} END {print bytes/1024/1024 " MB"}' access.log: total bytes.
  • awk '{c[$1]++} END {for (ip in c) print c[ip], ip}' access.log | sort -rn | head: requests per IP.
  • df -P | awk 'NR>1 && $5+0 > 85 {print $6, $5}': filesystems above a threshold.

41. Find the top 10 client IPs in an access log. Explain each part of the pipeline.

Answer: awk '{print $1}' access.log | sort | uniq -c | sort -rn | head -10. awk extracts the first field; sort groups identical IPs together, which is required because uniq only collapses adjacent duplicates; uniq -c counts each group; sort -rn sorts numerically in reverse; head trims. For huge logs, LC_ALL=C sort is noticeably faster.

42. Show useful find commands and how to combine find with xargs safely.

Answer:

  • find /var/log -name '*.log' -mtime +7: logs older than seven days.
  • find / -xdev -type f -size +500M -exec ls -lh {} +: large files on the root filesystem only.
  • find /srv -user olduser, -perm -o+w, -newer marker, -empty.
  • find /tmp/build -type f -name '*.tmp' -delete: delete matches (put -delete last; test without it first).

When piping file names, spaces and newlines break naive xargs. Use find ... -print0 | xargs -0 cmd. xargs -n 1 passes one argument per command, -P 4 runs four in parallel, -I {} places the argument where you want: cat hosts.txt | xargs -P 8 -I {} ssh {} uptime. -exec cmd {} + batches arguments like xargs and avoids the pipe entirely.

43. What does set -euo pipefail do, and what are its gotchas?

Answer: It makes Bash scripts fail fast instead of carrying on after errors:

  • -e: exit when a command returns non-zero.
  • -u: treat unset variables as errors, so rm -rf "$DIR/" with an empty DIR fails instead of targeting /.
  • -o pipefail: a pipeline returns the failure of any command in it, not just the last. Without it, curl ... | tar xz can "succeed" after a failed download.

Gotchas: -e is ignored inside if conditions and on the left of &&/||; local var=$(cmd) hides cmd's failure because local succeeds; grep returning 1 for "no match" will stop the script unless you write grep ... || true; with -u, use ${VAR:-} for optional variables. Pair it with trap cleanup EXIT and run shellcheck on every script.

44. What are exit codes, and what do 0, 1, 2, 126, 127, 130, 137 and 143 mean?

Answer: Every command returns 0โ€“255; echo $? shows the last one. 0 is success, 1 a general error, 2 usually misuse of a command or shell builtin. 126 means found but not executable (permissions or a noexec mount); 127 means command not found (wrong PATH, missing package, or a Windows line ending in the shebang). Codes above 128 mean "killed by signal (code minus 128)": 130 is SIGINT (Ctrl-C), 137 is SIGKILL (often the OOM killer or a forced container stop), 143 is SIGTERM. In pipelines, ${PIPESTATUS[@]} holds each command's status.

45. Explain variables, quoting, loops and functions in Bash with a short example.

Answer: Assign without spaces (name=web), read with "$name", and almost always quote expansions to avoid word splitting and glob expansion. $(cmd) captures output; ${VAR:-default} supplies a default; "$@" passes all arguments preserving spaces ($* joins them). Example that checks services listed in a file:

#!/usr/bin/env bash
set -euo pipefail

check() {
  local svc="$1"
  if systemctl is-active --quiet "$svc"; then
    echo "OK   $svc"
  else
    echo "FAIL $svc"; return 1
  fi
}

rc=0
while IFS= read -r svc; do
  [[ -z "$svc" || "$svc" == \#* ]] && continue
  check "$svc" || rc=1
done < "${1:-services.txt}"
exit "$rc"

while IFS= read -r reads lines safely (no backslash or whitespace mangling); for f in /var/log/*.log iterates over globs; never iterate over $(ls).

46. Explain cron syntax and the most common reasons a cron job doesn't run.

Answer: Five time fields then the command: minute, hour, day of month, month, day of week. 30 2 * * 1-5 /opt/scripts/backup.sh runs at 02:30 on weekdays; */15 * * * * runs every 15 minutes. User crontabs are edited with crontab -e and listed with crontab -l; files in /etc/cron.d/ and /etc/crontab have an extra user field. Common failures: cron's minimal environment (short PATH, no profile, different shell), relative paths, % characters which cron treats as newlines unless escaped, missing execute permission, a file in /etc/cron.d with a dot in its name (ignored by run-parts-style rules on Debian), and the server timezone differing from your assumption. Always redirect output: ... >> /var/log/backup.log 2>&1. See Q64 for the full scenario.

47. What are systemd timers, and when would you choose them over cron?

Answer: A timer unit activates a service unit on a schedule. For backup.service, create backup.timer:

[Timer]
OnCalendar=*-*-* 02:30:00
Persistent=true
RandomizedDelaySec=10min

[Install]
WantedBy=timers.target

Enable with systemctl enable --now backup.timer; list with systemctl list-timers; test expressions with systemd-analyze calendar "Mon..Fri 02:30". Advantages over cron: logs land in the journal (journalctl -u backup), Persistent=true runs missed jobs after downtime, the job gets cgroup resource limits and sandboxing, dependencies are explicit, and a job can't overlap itself because the service is already active.

Performance, boot, security, kernel and containers

48. What is load average, and how do you interpret it?

Answer: The three numbers from uptime or /proc/loadavg are exponentially damped averages over 1, 5 and 15 minutes of tasks that are running, waiting for a CPU, or in uninterruptible sleep (D state, typically waiting on disk or NFS). On Linux that last part is key: load can be high while CPUs are idle, because processes are stuck on I/O. Interpret it relative to CPU count (nproc): a load of 8 on 16 vCPUs is comfortable, on 2 vCPUs it means a queue. Compare the three values to see the trend (1-minute much higher than 15-minute means it is rising). Then use vmstat, top and iostat to tell CPU saturation from I/O blocking (Q59).

49. How do you use iostat, sar, mpstat and pidstat for performance analysis?

Answer: They come from the sysstat package.

  • iostat -xz 1: per-device extended stats. Read r/s/w/s, rkB/s/wkB/s, r_await/w_await (average latency in ms, including queueing) and aqu-sz (queue depth). %util near saturation is meaningful for single spinning disks but misleading for SSDs, NVMe and cloud volumes, which serve requests in parallel; latency and queue depth tell you more.
  • sar: historical data collected every few minutes when the sysstat collector is enabled. sar -u CPU, sar -r memory, sar -q load and run queue, sar -d disks, sar -n DEV network; sar -f reads an earlier day's file. This answers "what happened at 3 a.m.?"
  • mpstat -P ALL 1: per-CPU usage, revealing one core pinned by a single-threaded process.
  • pidstat -d 1 or -u, -r: per-process I/O, CPU and memory, so you can name the culprit. iotop is an interactive alternative.

50. Describe the Linux boot process from power-on to login.

Answer:

Firmware (UEFI or BIOS): POST, find boot device
  -> Bootloader (GRUB2): menu, load kernel + initramfs
  -> Kernel: init hardware, mount initramfs
  -> initramfs: load drivers, find and mount real root
  -> systemd (PID 1): mounts, services, targets
  -> default.target (multi-user or graphical)
  -> getty / sshd / display manager: login

Know the tools: systemctl get-default and set-default, systemd-analyze blame and critical-chain for slow boots, journalctl -b for boot logs. Recovery: at the GRUB menu, edit the kernel line and add systemd.unit=rescue.target (single-user with basic services) or emergency.target (minimal, root mounted read-only).

51. What is SELinux, and how do you troubleshoot an SELinux denial without disabling it?

Answer: SELinux is a mandatory access control system, enabled by default on the RHEL family. Every process and file has a security context (type such as httpd_t or httpd_sys_content_t), and policy decides which types may interact, on top of normal permissions. Modes: getenforce shows Enforcing, Permissive (logs only) or Disabled; setenforce 0 switches to permissive temporarily for diagnosis. Troubleshooting flow:

  1. ls -Z /srv/web and ps -eZ | grep nginx to see contexts.
  2. ausearch -m avc -ts recent (or journalctl -t setroubleshoot) to see the denial; audit2why explains it.
  3. Fix the label persistently: semanage fcontext -a -t httpd_sys_content_t "/srv/web(/.*)?" then restorecon -Rv /srv/web.
  4. Or enable a boolean: setsebool -P httpd_can_network_connect on for a web server proxying to a backend; non-standard ports need semanage port -a -t http_port_t -p tcp 8081.

Interview tip: "I set it to permissive to confirm, fixed the context, and set it back to enforcing" is the answer interviewers want. "I disabled SELinux" is a red flag.

52. What is AppArmor, and how does it differ from SELinux?

Answer: AppArmor is the mandatory access control used by default on Ubuntu and SUSE. Instead of labelling every file, it attaches path-based profiles to programs (in /etc/apparmor.d/) listing which files, capabilities and network access they may use. aa-status shows loaded profiles and their modes; aa-complain /etc/apparmor.d/usr.sbin.mysqld switches a profile to log-only, aa-enforce back (both from the apparmor-utils package); denials appear in the kernel log as apparmor="DENIED". Classic case: moving MySQL's data directory works with permissions set correctly but fails until the profile (or its local override file) allows the new path.

53. What is sysctl, and which kernel parameters do DevOps engineers commonly tune?

Answer: sysctl reads and writes kernel parameters exposed under /proc/sys. sysctl net.ipv4.ip_forward reads, sysctl -w net.ipv4.ip_forward=1 sets it until reboot, and a file such as /etc/sysctl.d/90-app.conf followed by sysctl --system makes it persistent. Commonly touched: net.ipv4.ip_forward (routers, NAT gateways, Kubernetes nodes), net.core.somaxconn and net.ipv4.tcp_max_syn_backlog (listen queues for busy servers), net.ipv4.ip_local_port_range (outbound connection-heavy proxies), vm.swappiness, vm.max_map_count (search engines such as Elasticsearch/OpenSearch require a higher value), and fs.file-max / fs.inotify.max_user_watches. Distinguish the system-wide fs.file-max from per-process limits (ulimit -n, /etc/security/limits.conf, or LimitNOFILE for services).

54. Which Linux features make containers possible? Explain namespaces.

Answer: A container is an ordinary Linux process isolated by namespaces, limited by cgroups, and usually further restricted by capabilities, seccomp and an LSM (SELinux/AppArmor), running on a layered root filesystem (overlayfs). Namespaces give the process its own view of: pid (process IDs, so the app can be PID 1), net (interfaces, routes, ports), mnt (mount table), uts (hostname), ipc, user (UID mapping, so root in the container can be unprivileged on the host), cgroup and time. Explore them with lsns, create them with unshare --pid --fork --mount-proc bash, and enter a container's network namespace with nsenter -t PID -n ip addr, which is a powerful debugging trick when the image has no tools. The Docker interview questions guide goes deeper on images, layers and runtimes.

55. What are cgroups, and how do v1 and v2 differ?

Answer: Control groups account for and limit CPU, memory, I/O and process counts for groups of processes. Most current distributions use cgroup v2: a single unified hierarchy mounted at /sys/fs/cgroup (check with stat -fc %T /sys/fs/cgroup, which prints cgroup2fs), with files such as memory.max, memory.current, cpu.max, io.max and pids.max, and better memory pressure accounting (memory.pressure, PSI). systemd organises everything into slices and services: systemd-cgls shows the tree, systemd-cgtop live usage, and systemctl set-property app.service MemoryMax=2G applies limits (CPUQuota= and IOWeight= work the same way). docker run --memory 1g --cpus 1.5 and Kubernetes resource limits are written into these same files.

AI in Linux operations

56. How should a Linux engineer use AI assistants for commands and troubleshooting?

Answer: AI assistants are useful for explaining an unfamiliar flag, drafting an awk one-liner, summarising a long journal excerpt or suggesting hypotheses for a symptom. The risk is that generated commands can be plausible and wrong: wrong distribution syntax, a destructive flag, or a command that targets the wrong path. A safe practice: read every flag before running it, try destructive commands with a dry-run or on a non-production host first (rsync -n, find without -delete, sed without -i), never paste secrets, hostnames or customer data from logs into external tools unless your organisation permits it, and never pipe generated scripts straight into sudo bash. AIOps platforms apply the same idea at scale, correlating metrics and logs to propose likely causes; the AIOps interview guide covers that side.

57. What Linux knowledge matters when a server runs AI or GPU workloads?

Answer: The same fundamentals, with larger numbers. Model weights and container images are big, so disk (df, du, separate volumes for model caches) and image cleanup matter. Inference servers hold models in RAM or GPU memory, so you need to distinguish a host OOM kill (kernel log) from a GPU out-of-memory error (application log). GPU hosts need the right kernel driver, and nvidia-smi shows GPU utilisation, memory and processes; containers need the NVIDIA container toolkit. For the full picture, read Linux skills for AI engineers and GPU basics for AI engineers.

Linux troubleshooting interview scenarios

58. The disk is full, but du finds no large files. What is happening?

Answer: Most often a process still holds a deleted file open. Someone ran rm on a large log while the application was writing to it: the directory entry is gone (so du can't see it) but the inode and its blocks stay allocated until the last file descriptor closes (so df still counts it).

What I would check:

  1. df -h versus du -xsh / to confirm the gap, and df -i to rule out inode exhaustion.
  2. lsof +L1 (open files with link count zero) or lsof -nP | grep '(deleted)' to find the process and file descriptor.
  3. Restart or reload the process cleanly so it reopens its log; if you can't restart it, truncate through the descriptor: : > /proc/PID/fd/FD.
  4. If no deleted files are found, check for data hidden under a mount point (files written to /data before the volume was mounted): bind-mount / elsewhere (mount --bind / /mnt/rootview) and run du there.

Production consideration: The lasting fix is log rotation that signals the process (postrotate reload) or uses copytruncate, log size limits, and disk-usage alerts at a warning threshold, plus a team rule: truncate live logs, don't rm them.

59. Load average is very high but CPU usage looks low. How do you investigate?

Answer: On Linux, load includes tasks in uninterruptible sleep, so high load with idle CPUs usually means processes are stuck waiting on I/O: a slow or saturated disk, a hung NFS mount, or a cloud volume hitting its throughput or IOPS limit.

What I would check:

  1. uptime and nproc to put the number in context; vmstat 1 for the b column and wa.
  2. ps -eo state,pid,cmd | awk '$1=="D"' to list D-state processes, and cat /proc/PID/stack (as root) to see where they're blocked.
  3. iostat -xz 1 for rising await and queue depth on a device; pidstat -d 1 or iotop to name the writer.
  4. findmnt -t nfs,nfs4 and dmesg -T for "server not responding" or disk errors.
  5. top for st (steal), which points to a noisy neighbour or an exhausted burstable-CPU credit balance on cloud instances.
  6. sar -q and sar -d to correlate with when it started (a backup, a batch job, log rotation).

Production consideration: If load is high with high us, it really is CPU saturation, and the next step is top/pidstat -u, then profiling or scaling out. Alert on latency and saturation per resource, not on load average alone.

60. A service fails to start after a configuration change. Walk through your approach.

Answer: Read what systemd and the service already told you before changing anything.

What I would check:

  1. systemctl status myapp: the active state, exit code or signal, and the last log lines.
  2. journalctl -u myapp -b --no-pager -n 100 for the full error; -xe adds explanations.
  3. Validate config with the program's own checker: nginx -t, sshd -t, apachectl configtest, named-checkconf.
  4. systemctl cat myapp to see the effective unit plus drop-ins; if you edited it, run systemctl daemon-reload.
  5. Exit code clues: 203/EXEC means the binary path is wrong or not executable; 217/USER means the User= doesn't exist; "Address already in use" means a port conflict (ss -tlnp).
  6. Run the ExecStart command manually as the service user (sudo -u myapp ...) to reproduce with the same environment, then check permissions on config, data and log paths, and SELinux/AppArmor denials.
  7. If it hit "start request repeated too quickly", fix the cause, then systemctl reset-failed myapp.

Production consideration: Keep config under version control, validate it in the pipeline before deployment, and keep the previous version ready for a fast rollback.

61. ssh to a server returns "Connection refused". What do you check?

Answer: "Connection refused" means a TCP RST came back: the host is reachable but nothing is listening on that port, or a firewall is actively rejecting. A timeout points to a silent drop (security group, NACL, firewall DROP or routing), and "Permission denied (publickey)" means the network is fine and authentication failed.

What I would check:

  1. From the client: ssh -vvv user@host and nc -zv host 22 to confirm the exact failure and port.
  2. Via console or another access path: systemctl status ssh (Ubuntu/Debian) or sshd (RHEL), and ss -tlnp | grep ssh for the port and address it listens on.
  3. sshd -t for config errors after a recent edit; a bad config stops the restart.
  4. Port changes: on recent Ubuntu releases sshd is socket-activated through ssh.socket, so after changing Port you need systemctl daemon-reload and systemctl restart ssh.socket; on SELinux systems the new port needs semanage port -a -t ssh_port_t -p tcp 2222.
  5. Host firewall: nft list ruleset, iptables -L -n -v, firewall-cmd --list-all or ufw status; also fail2ban bans.

Production consideration: Always keep a second access path (cloud serial console, session manager, out-of-band management) and change SSH config with an existing session open.

62. "Permission denied" persists even after chmod 777 on the file. Why?

Answer: File mode is only one of several checks. The usual causes, in the order I check them:

What I would check:

  1. Parent directories: every directory in the path needs execute (x) for that user. namei -l /srv/app/data/file shows permissions of each component in one view.
  2. Which user is really running it: the service may run as www-data or a dedicated user, not you. ps -o user= -p PID.
  3. Mount options: findmnt -T /path shows noexec (scripts on /tmp fail with 126), ro, or nosuid.
  4. Immutable or append-only attributes: lsattr file; remove with chattr -i if appropriate.
  5. SELinux or AppArmor: wrong context or a profile denying the path (Q51, Q52); check ausearch -m avc or the kernel log for DENIED.
  6. ACL mask limiting effective rights (getfacl), or a network filesystem (NFS root squash, SMB) applying its own rules.
  7. For scripts: a missing interpreter or a CRLF shebang produces confusing errors; file script.sh reveals CRLF line endings.

Production consideration: Revert the 777 once you've found the real cause, and fix ownership and least-privilege permissions in the deployment automation rather than by hand.

63. An application on a VM works with curl localhost but can't be reached from other machines. How do you debug it?

Answer: Work outward from the process to the network edge.

What I would check:

  1. ss -tlnp | grep 8000: if it shows 127.0.0.1:8000, the app is bound to loopback only; change its bind address to 0.0.0.0 or the private IP.
  2. Host firewall rules (nftables, firewalld zone, ufw).
  3. From another host in the same subnet: nc -zv 10.0.1.20 8000 and curl -v.
  4. Cloud controls: security group or NSG inbound rule, subnet NACLs, route tables, and the load balancer's target health check.
  5. tcpdump -ni eth0 port 8000 on the server to see whether SYN packets arrive at all. If they arrive and get no reply, it's local; if they never arrive, it's upstream.

Production consideration: Keep the app on a private address behind a load balancer or reverse proxy rather than opening it to the internet, and document the expected listening ports per service.

64. A script runs fine by hand but fails from cron. What is different?

Answer: The environment. Cron runs with a minimal PATH, no login profile, often /bin/sh instead of Bash, a different working directory, and no terminal or SSH agent.

What I would check:

  1. Capture output: ... >> /tmp/job.log 2>&1, and check the cron log (grep CRON /var/log/syslog on Ubuntu, /var/log/cron on RHEL, or journalctl -u cron / -u crond).
  2. Use absolute paths for commands and files, or set PATH= at the top of the crontab or script.
  3. Add a proper shebang (#!/usr/bin/env bash) and execute permission; escape % in the crontab line.
  4. Check which user's crontab it is in and whether that user can read the files and credentials it needs.

Production consideration: For important jobs, move to a systemd timer (journal logging, missed-run catch-up) and add a "last successful run" check to monitoring, so a silently failing job is noticed.

65. Name resolution fails on a server, but ping to an IP works. How do you fix it?

Answer: Network connectivity is fine; the resolver path is broken.

What I would check:

  1. getent hosts example.com (what applications see) versus dig example.com and dig @<known resolver> example.com.
  2. cat /etc/resolv.conf and, on systemd-resolved systems, resolvectl status to see the real upstream servers per interface.
  3. Whether outbound UDP and TCP port 53 is allowed by the host firewall and cloud rules.
  4. /etc/nsswitch.conf and /etc/hosts for overrides.
  5. Split-horizon or VPN DNS: a private zone may only be resolvable through the corporate or VPC resolver.

Production consideration: Don't hand-edit /etc/resolv.conf on hosts where Netplan, NetworkManager or DHCP manage it, because your change will be overwritten; fix it at the source.

66. A Java or Python process keeps disappearing with no error in its own log. What do you suspect?

Answer: A process killed by SIGKILL can't log anything, so suspect the OOM killer first.

What I would check:

  1. journalctl -k -b | grep -i -E 'oom|killed process', including the previous boot if the host restarted.
  2. systemctl status app for "status=9/KILL" or "oom-kill" as the result; for containers, exit code 137 and the runtime's OOMKilled flag.
  3. Whether it was a cgroup limit (MemoryMax=, container limit) or a host-wide shortage (free -h, sar -r history).
  4. Memory growth over time (a leak) versus a sudden spike (a large request or batch). For JVMs, compare heap settings with the container limit, because heap is only part of total memory.

Production consideration: Set explicit memory limits with headroom, alert on memory approaching the limit, and capture heap or memory profiles before the next kill rather than just raising the limit.

67. After a deployment the server becomes slow and iowait is high. How do you find the cause?

Answer: Identify the device, then the process, then the reason for the I/O.

What I would check:

  1. iostat -xz 1 to find the busy device and its latency.
  2. pidstat -d 1 or iotop -o to name the process doing the reads or writes.
  3. lsof -p PID to see which files it is hitting; common culprits are debug-level logging left on, a cache directory on the root disk, or a swap storm (check vmstat si/so).
  4. On cloud volumes, compare throughput and IOPS with the volume's provisioned limits in the provider's metrics.

Real-world example: Consider a bank's internal reporting service whose new release logged every SQL query at debug level. Disk writes jumped, request latency rose, and the database client timed out. The fix was a config change, plus a pipeline check rejecting debug logging in production profiles.

Production consideration: Put logs and high-write data on separate volumes, rate-limit logging, and include disk latency in deployment health checks so a bad release is rolled back automatically.

68. "Too many open files" errors appear under load. How do you fix it properly?

Answer: The process hit its file descriptor limit (sockets count as files).

What I would check:

  1. cat /proc/PID/limits | grep 'open files' for the actual limit, and ls /proc/PID/fd | wc -l for current usage.
  2. Whether usage grows steadily (a descriptor leak: connections or files not closed) or simply peaks with traffic.
  3. For systemd services, raise LimitNOFILE= in the unit or a drop-in; /etc/security/limits.conf only affects PAM login sessions, which is why "I changed limits.conf and nothing happened" is so common.

Production consideration: Raise the limit to a sensible value, but also fix leaks and add descriptor usage to monitoring; an unlimited ceiling just moves the failure.

69. A deployment script reported success, but the application was half-updated. How do you make shell automation safer?

Answer: The script almost certainly continued after a failed step: no set -e, a failure hidden inside a pipeline, or a command whose exit code was ignored.

What I would check:

  1. Add set -euo pipefail and a trap that logs the failing line (trap 'echo "failed at line $LINENO" >&2' ERR).
  2. Run shellcheck and fix quoting so paths with spaces or empty variables can't misbehave.
  3. Make the deploy atomic: unpack into a new release directory, verify, then switch a symlink (ln -sfn releases/v42 current) and reload, so a failure leaves the old version running.
  4. Make steps idempotent so a re-run is safe, and add a post-deploy health check (curl --fail) that exits non-zero.

Production consideration: Beyond a handful of hosts, move this logic into configuration management or a CI/CD pipeline with proper rollbacks; the DevOps engineer interview questions cover that workflow.

70. After editing /etc/fstab and rebooting, the server drops into emergency mode. How do you recover?

Answer: A mount listed in fstab failed (typo, wrong UUID, missing disk without nofail), and systemd treats required local mounts as essential for boot.

What I would check:

  1. Get a console (cloud serial console, hypervisor console, or attach the root volume to a rescue instance).
  2. At the emergency shell, read the error: journalctl -xb | grep -i -E 'mount|fstab' or systemctl --failed.
  3. If root is read-only: mount -o remount,rw /.
  4. Fix or comment out the bad line, compare UUIDs with blkid, add nofail for non-essential disks, then systemctl daemon-reload and mount -a to test.
  5. Reboot and confirm with findmnt.

Production consideration: Always run mount -a and findmnt --verify before rebooting after fstab changes, manage fstab through automation, and keep tested console access for every production host.

Key takeaways

  • Strong Linux interview answers state the command, the concept behind it and one failure mode, not just a definition.
  • Know the modern tooling: ip and ss instead of net-tools, systemd and journalctl, nftables or firewalld, cgroup v2.
  • Most scenario questions reduce to a few patterns: df versus du, load versus CPU versus I/O, OOM evidence in the kernel log, and permissions beyond the file mode.
  • Read before you change: systemctl status, journalctl, the program's config test, then the fix.
  • Treat shell scripts as production code: set -euo pipefail, quoting, meaningful exit codes, shellcheck.
  • Fix security controls instead of disabling them: SELinux contexts and booleans, AppArmor profiles, narrow sudo rules.
  • Containers are Linux processes with namespaces and cgroups, so Linux skill is the foundation for Docker and Kubernetes work.

Interview preparation checklist

  • Spin up one Ubuntu and one RHEL-family VM and practise every command in this guide on both.
  • Write a systemd service and timer from memory, break them deliberately, and debug them with journalctl.
  • Add a disk, create an LVM volume, extend it online and mount it through fstab with nofail.
  • Reproduce the deleted-open-file disk problem and fix it with lsof +L1.
  • Fill a cgroup's memory limit to trigger an OOM kill, then find the evidence in the kernel log.
  • Set up SSH keys, a ~/.ssh/config with ProxyJump, and a local port forward to a private service.
  • Write three scripts with set -euo pipefail: a log analyser using awk/sort/uniq, a health checker, and a cleanup job using find safely.
  • On the RHEL VM, serve files from a non-default directory with SELinux enforcing and fix the context properly.
  • Practise explaining load average, free output and iostat columns aloud in under a minute each.
  • Prepare two real troubleshooting stories in the format symptom, evidence, root cause, fix, prevention.

FAQ

What Linux skills are required for DevOps and cloud roles?

You need confident command-line use, permissions and users, systemd services and logs, disks and LVM, networking and SSH, text processing, Bash scripting and performance troubleshooting. For container work, add namespaces, cgroups and SELinux or AppArmor basics.

How should I prepare for a Linux interview?

Practise on real virtual machines rather than only reading. Break things on purpose, such as filling a disk, misconfiguring a service or blocking a port, then fix them while explaining each command aloud. Interviewers value a clear troubleshooting process over memorised flags.

Which Linux distribution should I practise on for interviews?

Practise on one Debian-family distribution such as Ubuntu and one RHEL-family distribution such as Rocky Linux or AlmaLinux. Most commands are identical, but package management, firewall tools, log file names and the default security module differ, and interviewers often ask about those differences.

Are Linux commands interview questions enough, or do I need concepts too?

Commands alone are not enough. Interviewers follow up with why: why load average is high with idle CPUs, why df and du disagree, why a chmod did not fix access. Learn the concept behind each command so you can reason about new problems.

How much shell scripting do I need for a Linux admin interview?

You should be able to write a short script with variables, loops, conditionals, functions and correct exit codes, use set -euo pipefail sensibly, quote variables correctly and process files line by line. Many interviews include a small live scripting task.

Do I need Linux if I work mostly with Kubernetes or managed cloud services?

Yes. Containers are Linux processes, nodes are Linux machines, and most debugging still ends at a shell inside a pod or on a node. Managed services reduce administration work but not the need to read logs, check networking and understand resource limits.

Are Linux certifications useful for interviews?

A recognised Linux certification can help freshers get shortlisted because it signals structured preparation. In interviews, hands-on troubleshooting ability matters more, so combine any certification with practical labs and real examples you can talk through.

Is a Linux administrator career still relevant in 2026?

Yes, though the role has broadened. Pure server administration roles increasingly overlap with DevOps, SRE, cloud and platform engineering, where Linux remains the operating foundation. Engineers who combine Linux depth with automation, cloud and containers have the widest set of options.

How long does it take to become interview-ready in Linux?

It depends on your starting point and daily practice time. Someone already comfortable with a terminal can cover the fundamentals in a few weeks of regular lab work; scenario-level confidence comes from repeated hands-on troubleshooting rather than a fixed number of days.

Linux is the base layer for every DevOps, cloud and SRE role, and interviewers can tell quickly whether you have practised on real servers. Cloudsoft's multi-cloud DevOps training with Linux and Python builds those skills through hands-on labs on AWS, Azure and Google Cloud, in classroom batches beside Ameerpet Metro or live online. If you also want AI, machine learning and cyber security on top of cloud foundations, look at the APEX AI, ML, Cloud & Cyber Security program. Call +91 96660 19191 to book a free demo.

New ยท AI Career Guide

Meet Aanya โ€” ask anything about courses, fees & placement

Instant answers from verified Cloudsoft info โ€” courses, fees, formats, placement support and free demos. Available 24/7, right here on the site.

How Aanya works โ†’
Share๐•infโœ‰
EnrollWhatsAppCall us