failfs¶
failfs is a kernel-internal filesystem that fails every operation
reaching it with EOPNOTSUPP. It is the counterpart to nullfs. Where
nullfs is permanently empty, failfs means “nothing is supported here”.
It cannot be mounted from userspace, nothing can be mounted on top of
it. It cannot be cloned.
The only way into it is the FD_FAILFS_ROOT file descriptor sentinel which
is understood by fchdir(2) and fchroot(2).
Semantics¶
Every path walk of a component through failfs fails with
EOPNOTSUPP before that component is parsed, including ..
No path lookup can open the root, not even with O_PATH.
A process with its working directory in failfs fails every
AT_FDCWD-relative lookup. As with any working directory that is
unreachable from the process root, the getcwd(2) system call returns
a path prefixed with (unreachable).
A process with its root directory in failfs fails every absolute path lookup including absolute symlinks and the interpreter of dynamically linked binaries. In other words, this fails exec.
Lookups anchored at explicit directory file descriptors keep working. It
is the fs_struct equivalent of RESOLVE_BENEATH. The process must
anchor every lookup at a file descriptor it explicitly holds.
Entering¶
fchroot(FD_FAILFS_ROOT, 0) requires CAP_SYS_CHROOT in the
caller’s user namespace, mirroring chroot(2). Unprivileged callers
may enter if all of the following hold:
no_new_privsis set: setuid binaries on regular mounts remain reachable via inherited directory file descriptors and executing them with an unusable root directory is the classic confused deputy.The caller is not already chrooted: the root directory is what confines
..resolution and the failfs root can never be reached by walking up a real mount tree, so moving the root of a chrooted task to failfs would allow it to escape its chroot viaopenat(fd, "..").The caller does not share its
fs_struct:no_new_privsis checked on the calling thread, but the root lives in thefs_struct. ACLONE_FSsibling withoutno_new_privscould otherwise execute a setuid binary with the failfs root, so entry requiresfs->users == 1, the same restrictionsetns(2)applies for the mount and user namespaces.
Leaving¶
Backing out is currently hard, but this is a property of the current
implementation, not a guaranteed interface, and may be loosened later.
For now a process that entered failfs counts as chrooted, so it cannot
create user namespaces to regain CAP_SYS_CHROOT, and chroot(2)
or fchroot(2) back out require CAP_SYS_CHROOT. The remaining way
out today is setns(2) with a mount namespace file descriptor, which
requires CAP_SYS_ADMIN over the target mount namespace as well as
CAP_SYS_CHROOT and CAP_SYS_ADMIN in the caller’s user namespace
and resets both root and working directory. A process that holds no such
file descriptor and restricts *chdir()/*chroot()/setns() via
seccomp cannot currently get back out.