DEV Community

Hugo | DevOps | Cybersecurity
Hugo | DevOps | Cybersecurity

Posted on Originally published at valtersit.com

iSCSI Storage for Hypervisors: TrueNAS, LIO, Multipath

Two paths, one dead for three weeks, nobody noticed. That's the story I keep running into. A cluster of a dozen VMs ran fine on a single surviving iSCSI path until a switch firmware reboot took the second one down at 02:00. Every VM froze for five minutes, then errored out. The monitoring was green the entire time because nobody was watching path state — only "is the LUN up." This guide is about making sure that doesn't happen to you.

This is for sysadmins and platform engineers running KVM-based hypervisors (Proxmox VE, plain libvirt, XCP-ng) who need block storage that behaves under failure. After reading, you'll be able to design an iSCSI fabric, build a target on TrueNAS SCALE or raw LIO, write a multipath.conf that isn't cargo-culted, and verify failover before you trust it with production VMs.

:::note[TL;DR]

  • Block storage wins for VM workloads because there's no double-caching and no file-locking pathologies — but only if you build it right.
  • TrueNAS SCALE wraps the same kernel LIO target you'd get with targetcli, just with a GUI and guardrails. Pick based on who's on call at 3am.
  • A single path is a SPOF with a nice name. Multipath is not optional.
  • Tune no_path_retry and dev_loss_tmo, or your VMs will hang for minutes on path failure.
  • Monitor path state, not LUN state. Save your targetcli config every time. :::

Prerequisites

  • Two servers: one target (TrueNAS SCALE or Debian/Ubuntu/RHEL with targetcli), one or more initiators.
  • Dedicated NICs or VLANs for iSCSI traffic. Do not share with VM traffic.
  • open-iscsi and multipath-tools installed on every initiator.
  • targetcli-fb (or targetcli) on raw LIO targets.
  • Root/sudo on both sides and a maintenance window for the first failover test.

Why Block Storage Still Wins (And When It Doesn't)

The reflex is "just use NFS, it's simpler." Sometimes that's correct — shared filesystems, template stores, and anything where multiple hosts need the same file are NFS territory. But for VM disk images, block wins, and it's not close.

The Case For iSCSI Over NFS for Hypervisors

With iSCSI, the client is stateless. The hypervisor sees a LUN, puts XFS or LVM-thin or ZFS on it, and the target has no idea what filesystem lives inside. That means no client-side cache coherency to reason about, no NFSv4.1 session trunking to debug, and no file-locking pathologies when a VM does something stupid with flock().

NFSv4.1 sessions are not multipathing. They give you trunking across multiple connections to the same server, but if that server dies, everything dies. Real multipath gives you two independent paths to two independent controllers, and the kernel picks one.

I moved a cluster off NFS after a coherency incident: two KVM hosts mounted the same NFS export, one host's page cache went stale after a network blip, and a VM wrote a corrupted qcow2 over a file the other host had just modified. iSCSI with a properly sized LUN per VM eliminated the class of problem entirely.

When You Should Absolutely Not Do This

If you have one initiator, one target, one switch, and one cable — you've built a SPOF with extra steps. iSCSI without multipath is worse than NFS because at least NFS will give you a clean I/O error. A dead iSCSI path with default timeouts will hang your VM for minutes before anything surfaces.

The "I'll add multipath later" crowd always adds it after the first outage. Do it now.


⚠️ TRUNCATED VERSION
This is an abbreviated cross-post. Full article (all config files, architecture diagrams, images): valtersit.com


🛠️ Partner Tools for Developers

ValtersIT curates partner deals for developers and sysadmins — VPS hosting, security tools, monitoring platforms and dev productivity gear. No filler, no affiliate spam.

➜ valtersit.com/deals/

Top comments (0)