Difficulty: Intermediate
How does a Unix file system store files? What is an inode, how do direct and indirect pointers work, and what is the difference between hard and soft links?
A file system is the part of the OS that turns a raw array of disk blocks into named, organized, persistent files. The user wants names, folders and permissions; the disk offers only numbered blocks. Think of a library. The books are the data blocks, the catalogue card for each book records where it lives and who can borrow it (the inode), and the index of titles that points to the cards is the directory.
In Unix-style file systems (ext4, ext2, UFS), every file is described by an inode (index node), a fixed-size on-disk structure. It stores the file's metadata: type (regular, directory, symlink, device), owner and group IDs, permission bits, size, timestamps (access, modify, change), link count, and pointers to the data blocks. What it does not store is the file name. That surprises many people. Names live in directories. A directory is itself a special file whose contents are a list of (name, inode number) pairs. To resolve /home/alice/notes.txt, the kernel starts at the root inode, reads the root directory's data, finds the entry for home, gets its inode number, reads that directory, and so on down the path.
How does the inode locate the data? The classic design has 12 direct pointers straight to data blocks, one single indirect pointer to a block full of pointers, one double indirect pointer (a block of pointers to blocks of pointers) and one triple indirect pointer. This is efficient because small files, which are the majority, need no extra reads, while the structure scales to huge files. Take a 4 KB block and 4-byte pointers, so a block holds 1,024 pointers. Direct: 12 x 4 KB = 48 KB. Single indirect: 1,024 x 4 KB = 4 MB. Double indirect: 1,024^2 x 4 KB = 4 GB. Triple indirect: 1,024^3 x 4 KB = 4 TB. The maximum file size is roughly 4 TB plus small terms. Modern ext4 replaces block lists with extents (a start block and a length), which are more compact for large contiguous files.
The on-disk layout has a superblock (file system parameters), an inode bitmap and data block bitmap tracking free space, the inode table, and the data blocks. The number of inodes is fixed when the file system is created, so you can run out of inodes ("No space left on device") even with free blocks if you create millions of tiny files.
Links are a favourite topic. A hard link is another directory entry pointing to the same inode; the inode's link count increases. The file's data is deleted only when the link count reaches zero (and no process has it open). Hard links cannot cross file systems, since inode numbers are only unique within one, and normally cannot link directories. A soft (symbolic) link is a separate small file whose content is a path name; it can cross file systems and point to directories, but it dangles if the target is removed.
Also cover allocation methods briefly: contiguous (fast, but external fragmentation), linked (FAT uses a table of links, no random access in the pure form), and indexed (inode approach, supports random access). Journaling file systems such as ext4 and NTFS log metadata changes before applying them, so a crash can be recovered quickly by replaying the journal instead of scanning the whole disk with fsck.
$ echo "hello" > a.txt
$ ln a.txt hard.txt # hard link: same inode
$ ln -s a.txt soft.txt # symbolic link: new inode holding a path
$ ls -li a.txt hard.txt soft.txt
131074 -rw-r--r-- 2 user user 6 Sep 29 10:00 a.txt
131074 -rw-r--r-- 2 user user 6 Sep 29 10:00 hard.txt
131075 lrwxrwxrwx 1 user user 5 Sep 29 10:00 soft.txt -> a.txt
$ rm a.txt
$ cat hard.txt
hello
$ cat soft.txt
cat: soft.txt: No such file or directory
a.txt and hard.txt share inode 131074 with link count 2, so the data survives deleting a.txt. soft.txt has its own inode and dangles.
inode, file system, block allocation, hard link, symbolic link