summaryrefslogtreecommitdiff
path: root/kernel
diff options
context:
space:
mode:
authorChristian Brauner <brauner@kernel.org>2026-06-04 11:14:28 +0200
committerChristian Brauner <brauner@kernel.org>2026-06-29 10:54:43 +0200
commit1b8a585da1b50b3db0e7bbd073ca274875539ea8 (patch)
treeecc916b89a305d4fe9765e2a873877e63c8b7861 /kernel
parentdc59e4fea9d83f03bad6bddf3fa2e52491777482 (diff)
parent272fa19991cd6c40602b0d27d4f07117d25792c0 (diff)
Merge patch series "fs,kthread: start all kthreads in nullfs"
Christian Brauner <brauner@kernel.org> says: Summary: * all kthreads are isolated in a separate SB_KERNMOUNT of nullfs. -> no lookup of anything else, no mounting on top of it, completely isolated. * init has a separate fs_struct from all kthreads * scoped_with_init_fs() allows a kthread to temporarily assume init's fs_struct for filesystem operations. So this is a bit of a crazy series. When the kernel is started it roughly goes like this: init_task ==> create pid 1 (systemd etc.) ==> pid 2 (kthreadd) After this point all kthreads and PID 1 share the same filesystem state. That obviously already came up when we discussed pivot_root() as this allows pivot_root() to rewrite the fs_struct of all kthreads. This rewriting is really weird and mostly done so kthread can use init's filesystem state when they would like to. But this really should be discouraged. The rewriting should also stop completely. I worked a bit to get rid of it in a more fundamental way. Is it crazy? Yes. Is it likely broken? Yes. Does it at least boot? Yes. Instead of sharing fs_struct between kernel threads and pid 1, pid 1 get's a completely separate fs_struct. All kthreads continue sharing init_fs as before and pid 1's fs_struct is isolated from kthread's filesystem state. IOW, userspace init cannot affect kthreads filesystem state anymore and kthreads cannot affect userspace's filesystem state anymore - without explicit opt-in. All kthreads are anchored in a kernel internal mount of nullfs that cannot be mounted on and that cannot be used to follow other mounts. It's a completely private mount that insulates kthreads. This series makes performing mountains of filesystem work such as path lookup and file opening and so on from kthreads hard - painfully so. I think this is a benefit because it takes the idea of just offloading _security sensitive_ operations in init's filesystem state and running random binaries or opening and creating files to kthreads difficult behind the shed... And imho it should. The only remaining kernel tasks that actually share init's filesystem state are usermodhelpers - as they execute random binaries in the root filesystem. Another concept we should really show the back of the shed. This gives a lot stronger guarantees than what we have now. This also makes path lookup from kthreads fail by default. IOW, it won't be possible anymore to just lookup random stuff in init's filesytem state without explicitly opting in to that. The places that need to perform lookup in init's filesystem state may use scoped_with_init_fs() which will temporarily override the caller's fs_struct with init's fs_struct. We now also warn and notice when pid 1 simply stops sharing filesystem state with us, i.e., abandons it's userspace_init_fs. On older kernels if PID 1 unshared its filesystem state with us the kernel simply used the stale fs_struct state implicitly pinning anything that PID 1 had last used. Even if PID 1 might've moved on to some completely different fs_struct state and might've even unmounted the old root. This has hilarious consequences: Think continuing to dump coredump state into an implicitly pinned directory somewhere. Calling random binaries in the old rootfs via usermodehelpers. Be aggressive about this: We simply reject operating on stale fs_struct state by reverting userspace_init_fs to nullfs. Every kworker that does lookups after this point will fail. Every usermodehelper call will fail. This is a lot stronger but I wouldn't know what it means for pid 1 to simply stop sharing its fs state with the kernel. Clearly it wanted to separate so cut all ties. I've went through the kernel and looked at hopefully everything that does path lookup from kthreads (workqueues, ...). TL;DR: ==== PID 1 (systemd) ==== root@localhost:~# stat --file-system /proc/1/root File: "/proc/1/root" ID: e3cb00dd533cd3d7 Namelen: 255 Type: ext2/ext3 root@localhost:~# cat /proc/1/mountinfo | wc -l 30 ==== PID 2 (kthreadd) ==== root@localhost:~# stat --file-system /proc/2/root File: "/proc/2/root" ID: 200000000 Namelen: 255 Type: nullfs root@localhost:~# cat /proc/2/mountinfo | wc -l 0 * patches from https://patch.msgid.link/20260601-work-kthread-nullfs-v4-0-77ee053060e0@kernel.org: (25 commits) fs: stop rewriting paths for PF_EXITING | PF_DUMPCORE fs: stop rewriting kthread fs structs fs: start all kthreads in nullfs nullfs: make nullfs multi-instance devtmpfs: create private mount namespace fs: add umh argument to struct kernel_clone_args fs: stop sharing fs_struct between init_task and pid 1 af_unix: use scoped_with_init_fs() for coredump socket lookup initramfs: use scoped_with_init_fs() for rootfs unpacking pnfs/blocklayout: use scoped_with_init_fs() for SCSI device lookup ksmbd: use scoped_with_init_fs() for VFS path operations ksmbd: use scoped_with_init_fs() for filesystem info path lookup ksmbd: use scoped_with_init_fs() for share path resolution fs: use scoped_with_init_fs() for kernel_read_file_from_path_initns() coredump: use scoped_with_init_fs() for coredump path resolution btrfs: use scoped_with_init_fs() for update_dev_time() scsi: target: use scoped_with_init_fs() for APTPL metadata scsi: target: use scoped_with_init_fs() for ALUA metadata crypto: ccp: use scoped_with_init_fs() for SEV file access rnbd: use scoped_with_init_fs() for block device open ... Link: https://patch.msgid.link/20260601-work-kthread-nullfs-v4-0-77ee053060e0@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
Diffstat (limited to 'kernel')
-rw-r--r--kernel/fork.c53
-rw-r--r--kernel/kcmp.c2
-rw-r--r--kernel/umh.c6
3 files changed, 36 insertions, 25 deletions
diff --git a/kernel/fork.c b/kernel/fork.c
index 13e38e89a1f3..b85b649c710d 100644
--- a/kernel/fork.c
+++ b/kernel/fork.c
@@ -1613,9 +1613,27 @@ static int copy_exec_state(u64 clone_flags, struct task_struct *tsk)
return task_exec_state_copy(tsk);
}
-static int copy_fs(u64 clone_flags, struct task_struct *tsk)
+static int copy_fs(u64 clone_flags, struct task_struct *tsk, bool umh)
{
- struct fs_struct *fs = current->fs;
+ struct fs_struct *fs;
+
+ /*
+ * Usermodehelper may copy userspace_init_fs filesystem state but
+ * they don't get to create mount namespaces, share the
+ * filesystem state, or be started from a non-initial mount
+ * namespace.
+ */
+ if (umh) {
+ if (clone_flags & (CLONE_NEWNS | CLONE_FS))
+ return -EINVAL;
+ if (current->nsproxy->mnt_ns != &init_mnt_ns)
+ return -EINVAL;
+ fs = userspace_init_fs;
+ } else {
+ fs = current->fs;
+ VFS_WARN_ON_ONCE(current->fs != current->real_fs);
+ }
+
if (clone_flags & CLONE_FS) {
/* tsk->fs is already what we want */
read_seqlock_excl(&fs->seq);
@@ -1628,7 +1646,7 @@ static int copy_fs(u64 clone_flags, struct task_struct *tsk)
read_sequnlock_excl(&fs->seq);
return 0;
}
- tsk->fs = copy_fs_struct(fs);
+ tsk->real_fs = tsk->fs = copy_fs_struct(fs);
if (!tsk->fs)
return -ENOMEM;
return 0;
@@ -2276,7 +2294,7 @@ __latent_entropy struct task_struct *copy_process(
retval = copy_files(clone_flags, p, args->no_files);
if (retval)
goto bad_fork_cleanup_semundo;
- retval = copy_fs(clone_flags, p);
+ retval = copy_fs(clone_flags, p, args->umh);
if (retval)
goto bad_fork_cleanup_files;
retval = copy_sighand(clone_flags, p);
@@ -2818,6 +2836,7 @@ pid_t user_mode_thread(int (*fn)(void *), void *arg, unsigned long flags)
.exit_signal = (flags & CSIGNAL),
.fn = fn,
.fn_arg = arg,
+ .umh = 1,
};
return kernel_clone(&args);
@@ -3215,7 +3234,7 @@ static int unshare_fd(unsigned long unshare_flags, struct files_struct **new_fdp
*/
int ksys_unshare(unsigned long unshare_flags)
{
- struct fs_struct *fs, *new_fs = NULL;
+ struct fs_struct *new_fs = NULL;
struct files_struct *new_fd = NULL;
struct cred *new_cred = NULL;
struct nsproxy *new_nsproxy = NULL;
@@ -3246,6 +3265,10 @@ int ksys_unshare(unsigned long unshare_flags)
if (unshare_flags & CLONE_NEWNS)
unshare_flags |= CLONE_FS;
+ /* No unsharing with overriden fs state */
+ VFS_WARN_ON_ONCE(unshare_flags & (CLONE_NEWNS | CLONE_FS) &&
+ current->fs != current->real_fs);
+
err = check_unshare_flags(unshare_flags);
if (err)
goto bad_unshare_out;
@@ -3293,23 +3316,13 @@ int ksys_unshare(unsigned long unshare_flags)
new_nsproxy = NULL;
}
- task_lock(current);
+ if (new_fs)
+ new_fs = switch_fs_struct(new_fs);
- if (new_fs) {
- fs = current->fs;
- read_seqlock_excl(&fs->seq);
- current->fs = new_fs;
- if (--fs->users)
- new_fs = NULL;
- else
- new_fs = fs;
- read_sequnlock_excl(&fs->seq);
- }
-
- if (new_fd)
+ if (new_fd) {
+ guard(task_lock)(current);
swap(current->files, new_fd);
-
- task_unlock(current);
+ }
if (new_cred) {
/* Install the new user namespace */
diff --git a/kernel/kcmp.c b/kernel/kcmp.c
index 7c1a65bd5f8d..76476aeee067 100644
--- a/kernel/kcmp.c
+++ b/kernel/kcmp.c
@@ -186,7 +186,7 @@ SYSCALL_DEFINE5(kcmp, pid_t, pid1, pid_t, pid2, int, type,
ret = kcmp_ptr(task1->files, task2->files, KCMP_FILES);
break;
case KCMP_FS:
- ret = kcmp_ptr(task1->fs, task2->fs, KCMP_FS);
+ ret = kcmp_ptr(task1->real_fs, task2->real_fs, KCMP_FS);
break;
case KCMP_SIGHAND:
ret = kcmp_ptr(task1->sighand, task2->sighand, KCMP_SIGHAND);
diff --git a/kernel/umh.c b/kernel/umh.c
index 48117c569e1a..6e2c7bb315c6 100644
--- a/kernel/umh.c
+++ b/kernel/umh.c
@@ -71,10 +71,8 @@ static int call_usermodehelper_exec_async(void *data)
spin_unlock_irq(&current->sighand->siglock);
/*
- * Initial kernel threads share ther FS with init, in order to
- * get the init root directory. But we've now created a new
- * thread that is going to execve a user process and has its own
- * 'struct fs_struct'. Reset umask to the default.
+ * Usermodehelper threads get a copy of userspace init's
+ * fs_struct. Reset umask to the default.
*/
current->fs->umask = 0022;