DEV Community

Aditya Raj
Aditya Raj

Posted on

Why is Linux Lying to you ? Inodes explained

No space left on device, but df says the disk is 1% full

I hit this on a box with 56 MB free and a filesystem that refused to create an empty file.

$ df -h /mnt/demo
Filesystem  Size  Used Avail Use%
/dev/loop0   60M   24K   56M   1%

$ touch index.js
touch: cannot touch 'index.js': No space left on device
Enter fullscreen mode Exit fullscreen mode

The disk is almost empty. The error says it is full. Both are true, because they are talking about different things.

Reproduce it in thirty seconds

You don't need a broken server. Make a tiny filesystem with a tiny inode table:

dd if=/dev/zero of=/tmp/tiny.img bs=1M count=64
mkfs.ext4 -N 128 -F /tmp/tiny.img
mkdir -p /mnt/demo
mount -o loop /tmp/tiny.img /mnt/demo

for i in $(seq 1 128); do touch /mnt/demo/$i; done

touch /mnt/demo/newfile     # No space left on device
df -h /mnt/demo             # 56M free
df -i /mnt/demo             # IUse 100%
Enter fullscreen mode Exit fullscreen mode

-N 128 is the whole trick. It tells mkfs.ext4 to build an inode table with room for 128 inodes and nothing more, ever.

What an inode actually is

A file on Linux is three separate things.

The inode is a record holding the metadata: permissions, owner, timestamps, size, and pointers to where the data lives. The data blocks hold the contents. And the directory entry holds the name.

That last one surprises people. The filename is not in the inode. A directory is a list that maps names to inode numbers, and that is the only place a name exists.

$ ls -li
12 -rw-r--r-- 1 eradon eradon 4232 notes.txt
Enter fullscreen mode Exit fullscreen mode

That leading 12 is the inode number. The name notes.txt is just an entry in the directory pointing at inode 12.

Why moving a huge file is instant

Once names and inodes are separate, mv inside one filesystem makes sense. It doesn't move data. It removes one directory entry and creates another pointing at the same inode. A 40 GB file "moves" in a millisecond because nothing moves.

Across filesystems it's a real copy, because the inode number means nothing over there.

Why ext4 runs out

On ext4 the inode table is built when the filesystem is created, sized from a bytes-per-inode ratio, and it never grows. You can't add inodes later. There is no command for it.

Every file costs exactly one inode, regardless of size. A 4 GB video costs one. An empty file costs one. So a filesystem full of tiny files runs out of inodes long before it runs out of space — logs, cache entries, session files, mail queues, node_modules. None of them are big. All of them cost an inode.

That's the whole bug. The disk had 56 MB free and zero inodes, and touch needs an inode.

Recovering from it

Your options on ext4 are thin:

  • delete files you don't need, which frees inodes immediately
  • move a directory tree to another filesystem
  • create a new filesystem with a bigger inode count and mount it

resize2fs grows the filesystem, not the inode table. tune2fs -l will show you the count but won't change it.

XFS behaves differently — it allocates inodes dynamically, so it doesn't have a fixed ceiling in the same way. If you know a filesystem will hold millions of small files, that's a real argument for XFS, or for passing a much lower bytes-per-inode value at mkfs time.

Hard links: many names, one inode

If names and inodes are separate, nothing stops two names pointing at the same inode. That's a hard link.

ln notes.txt backup.txt
ls -li notes.txt backup.txt
Enter fullscreen mode Exit fullscreen mode

Both show the same inode number, and the link count goes to 2. There's no original and no copy — the two names are equal, and the data exists once.

Two consequences follow:

Hard links can't cross filesystems. An inode number only means something within its own filesystem, so ln across a mount point gives you Invalid cross-device link.

A hard link needs no new inode. On our full filesystem, touch fails but this works:

ln /mnt/demo/1 /mnt/demo/another-name    # fine
ln -s /mnt/demo/1 /mnt/demo/softlink     # No space left on device
Enter fullscreen mode Exit fullscreen mode

A hard link adds a name to a directory. A symlink is a new file, so it needs an inode, so it fails. That one line is the clearest demonstration of the difference I know.

What rm actually does

rm doesn't delete a file. It removes a directory entry and decrements the link count. The syscall is even called unlink.

The inode and its data survive while the link count is above zero. Only when the count hits zero and no process holds the file open does the kernel reclaim the inode and the blocks.

That second condition is the other classic "disk is full" mystery: delete a 40 GB log that nginx still has open, and du says the space is gone while df says it's still used. The space comes back when the process closes the descriptor, or you restart it. lsof +L1 lists these.

Symlinks

A symlink is a file whose contents are a path. It gets its own inode, and that inode holds text like /home/eradon/notes.txt.

Because it stores a path rather than an inode number, it can point anywhere, including another filesystem. Resolution costs an extra lookup: you hit the symlink's inode, the kernel reads the path out of it, and then resolves that path normally.

Permissions work differently between the two. Hard links all share one inode, so they share its permission bits — chmod through any name affects all of them. A symlink has its own inode and its own bits, but they're ignored: the kernel checks the target's permissions. A symlink never grants access you don't already have.

"Everything is a file", so did we break everything?

If files need inodes and inodes ran out, can the machine still fork a process or read /proc?

Yes, because there are two different things called an inode.

On-disk inodes belong to a filesystem like ext4. They live in that fixed table. They're what we exhausted.

VFS inodes are struct inode objects in kernel memory. The VFS is the layer that makes everything look like a file, and not everything behind it is on a disk. /proc and /sys are synthesised by the kernel when you look at them:

$ stat -c '%i %n' /proc/self/status
4026532042 /proc/self/status

$ df -i /proc
Filesystem  Inodes IUsed IFree IUse%
proc             0     0     0     -
Enter fullscreen mode Exit fullscreen mode

It has an inode number, but procfs has no table and nothing to run out of. Exhausting ext4's inodes exhausts that filesystem, not some global pool.

A rule of thumb that gets you most of the way: if you can hold a file descriptor to it, it has an inode — sockets, pipes, epoll, timerfd, /proc entries. There are exceptions, but not many.

What to actually monitor

Most monitoring watches disk space and stops there. Add inodes:

df -i                          # inode usage per mount
du --inodes -x -d1 /var        # which directory is eating them
find /var/spool -xdev | wc -l  # count in a suspect tree
Enter fullscreen mode Exit fullscreen mode

Alert on df -i the same way you alert on df -h. In production the things that generate millions of tiny files are exactly the things nobody watches: log directories without rotation, cache trees, session stores, CI workspaces.

And in containers, remember overlayfs draws on the host filesystem's inodes. One runaway container can starve the host.

The takeaway

Free space and a usable filesystem are not the same thing. When writes fail and df -h looks fine, check df -i before anything else.


Top comments (0)