* [PATCH 01/10] file: let dup_fd() leave the punched hole behind
2026-09-21 14:15 [PATCH 00/10] files,close_range: add CLOSE_RANGE_{CLOEXEC_ONLY,EXCEPT} Christian Brauner
@ 2026-09-21 14:15 ` Christian Brauner
2026-09-21 14:15 ` [PATCH 02/10] selftests/core: test the hole CLOSE_RANGE_UNSHARE leaves behind Christian Brauner
` (9 subsequent siblings)
10 siblings, 0 replies; 14+ messages in thread
From: Christian Brauner @ 2026-09-21 14:15 UTC (permalink / raw)
To: Jann Horn, linux-fsdevel, Oleg Nesterov
Cc: Alexander Viro, Jan Kara, Neil Brown, Jeff Layton,
Christian Brauner (Amutable)
close_range(CLOSE_RANGE_UNSHARE) passed the range it is about to close
to dup_fd() so the clone is done without that range. This only works
when the last open descriptor falls into the range. A range in the
middle of the table is copied like everything else. That's wasteful.
Don't copy them. Descriptors in a skipped range stay open in the source
fdtable. They are never copied into the new table and no reference is
taken on them. Figuring out the size of the table follows the same rule.
That drops the special-case it has now.
This also means we stop calling ->flush() on fds from a table they
were never part of. With the range at the top of the table that was
already mostly the case.
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
fs/file.c | 60 ++++++++++++++++++++++++++++++++++++++++++++++--------------
1 file changed, 46 insertions(+), 14 deletions(-)
diff --git a/fs/file.c b/fs/file.c
index 628ca07dc4b1..95dbdfee555a 100644
--- a/fs/file.c
+++ b/fs/file.c
@@ -352,27 +352,51 @@ static inline bool fd_is_open(unsigned int fd, const struct fdtable *fdt)
return test_bit(fd, fdt->open_fds);
}
+/* Bits of [range->from, range->to] that fall into word @i of a bitmap. */
+static unsigned long fd_range_word(struct fd_range *range, unsigned int i)
+{
+ unsigned int first = i * BITS_PER_LONG;
+ unsigned int last = first + BITS_PER_LONG - 1;
+
+ if (range->to < first || range->from > last)
+ return 0;
+ return GENMASK(min(range->to, last) - first,
+ max(range->from, first) - first);
+}
+
+/* Bits of word @i that dup_fd() leaves behind. */
+static unsigned long dup_fd_dropped_word(unsigned int i, struct fd_range *punch_hole)
+{
+ if (!punch_hole)
+ return 0;
+ return fd_range_word(punch_hole, i);
+}
+
/*
* Note that a sane fdtable size always has to be a multiple of
* BITS_PER_LONG, since we have bitmaps that are sized by this.
*
* punch_hole is optional - when close_range() is asked to unshare
- * and close, we don't need to copy descriptors in that range, so
- * a smaller cloned descriptor table might suffice if the last
- * currently opened descriptor falls into that range.
+ * and close, dup_fd() leaves the descriptors in that range behind,
+ * so the cloned table only has to reach the last open descriptor
+ * outside of it.
*/
static unsigned int sane_fdtable_size(struct fdtable *fdt, struct fd_range *punch_hole)
{
unsigned int last = find_last_bit(fdt->open_fds, fdt->max_fds);
+ unsigned int i;
if (last == fdt->max_fds)
return NR_OPEN_DEFAULT;
- if (punch_hole && punch_hole->to >= last && punch_hole->from <= last) {
- last = find_last_bit(fdt->open_fds, punch_hole->from);
- if (last == punch_hole->from)
- return NR_OPEN_DEFAULT;
+ /* Only words up to the last open descriptor can hold a kept one. */
+ i = last / BITS_PER_LONG + 1;
+ while (i--) {
+ unsigned long dropped = dup_fd_dropped_word(i, punch_hole);
+
+ if (fdt->open_fds[i] & ~dropped)
+ return (i + 1) * BITS_PER_LONG;
}
- return ALIGN(last + 1, BITS_PER_LONG);
+ return NR_OPEN_DEFAULT;
}
/*
@@ -384,7 +408,8 @@ struct files_struct *dup_fd(struct files_struct *oldf, struct fd_range *punch_ho
{
struct files_struct *newf;
struct file **old_fds, **new_fds;
- unsigned int open_files, i;
+ unsigned int open_files, fd;
+ unsigned long dropped = 0;
struct fdtable *old_fdt, *new_fdt;
newf = kmem_cache_alloc(files_cachep, GFP_KERNEL);
@@ -451,13 +476,18 @@ struct files_struct *dup_fd(struct files_struct *oldf, struct fd_range *punch_ho
*
* Instead of trying to placate userspace racing with itself, we
* ref the file if we see it and mark the fd slot as unused otherwise.
+ * Descriptors dup_fd() is asked to leave behind get the same treatment.
*/
- for (i = open_files; i != 0; i--) {
+ for (fd = 0; fd < open_files; fd++) {
struct file *f = rcu_dereference_raw(*old_fds++);
- if (f) {
+
+ if (!(fd % BITS_PER_LONG))
+ dropped = dup_fd_dropped_word(fd / BITS_PER_LONG, punch_hole);
+ if (f && !(dropped & BIT_MASK(fd))) {
get_file(f);
} else {
- __clear_open_fd(open_files - i, new_fdt);
+ f = NULL;
+ __clear_open_fd(fd, new_fdt);
}
rcu_assign_pointer(*new_fds++, f);
}
@@ -848,10 +878,12 @@ SYSCALL_DEFINE3(close_range, unsigned int, fd, unsigned int, max_fd,
swap(cur_fds, fds);
}
- if (flags & CLOSE_RANGE_CLOEXEC)
+ if (flags & CLOSE_RANGE_CLOEXEC) {
__range_cloexec(cur_fds, fd, max_fd);
- else
+ } else if (!fds) {
+ /* If we unshared, dup_fd() left the range behind already. */
__range_close(cur_fds, fd, max_fd);
+ }
if (fds) {
/*
--
2.53.0
^ permalink raw reply related [flat|nested] 14+ messages in thread* [PATCH 02/10] selftests/core: test the hole CLOSE_RANGE_UNSHARE leaves behind
2026-09-21 14:15 [PATCH 00/10] files,close_range: add CLOSE_RANGE_{CLOEXEC_ONLY,EXCEPT} Christian Brauner
2026-09-21 14:15 ` [PATCH 01/10] file: let dup_fd() leave the punched hole behind Christian Brauner
@ 2026-09-21 14:15 ` Christian Brauner
2026-09-21 14:15 ` [PATCH 03/10] file: rename dup_fd()'s punch_hole to range Christian Brauner
` (8 subsequent siblings)
10 siblings, 0 replies; 14+ messages in thread
From: Christian Brauner @ 2026-09-21 14:15 UTC (permalink / raw)
To: Jann Horn, linux-fsdevel, Oleg Nesterov
Cc: Alexander Viro, Jan Kara, Neil Brown, Jeff Layton,
Christian Brauner (Amutable)
Cover a range in the middle of a full word:
- the descriptors in it are gone from the clone and the ones around it
are still there
- the slots are handed out again, the word is not left marked full
- the table the child cloned from is untouched
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
tools/testing/selftests/core/close_range_test.c | 44 +++++++++++++++++++++++++
1 file changed, 44 insertions(+)
diff --git a/tools/testing/selftests/core/close_range_test.c b/tools/testing/selftests/core/close_range_test.c
index f14eca63f20c..afae3462502d 100644
--- a/tools/testing/selftests/core/close_range_test.c
+++ b/tools/testing/selftests/core/close_range_test.c
@@ -236,6 +236,50 @@ TEST(close_range_unshare_capped)
EXPECT_EQ(0, WEXITSTATUS(status));
}
+TEST(close_range_unshare_hole)
+{
+ int i, status;
+ pid_t pid;
+ struct __clone_args args = {
+ .flags = CLONE_FILES,
+ .exit_signal = SIGCHLD,
+ };
+
+ /* Fill the first two words of the table. */
+ for (i = 3; i < 128; i++)
+ ASSERT_GE(dup2(0, i), 0);
+
+ pid = sys_clone3(&args, sizeof(args));
+ ASSERT_GE(pid, 0);
+
+ if (pid == 0) {
+ /* Punch a hole into the second word, behind a full first one. */
+ if (sys_close_range(70, 80, CLOSE_RANGE_UNSHARE))
+ exit(EXIT_FAILURE);
+
+ for (i = 3; i < 128; i++) {
+ bool closed = i >= 70 && i <= 80;
+
+ if (closed == (fcntl(i, F_GETFD) != -1))
+ exit(EXIT_FAILURE);
+ }
+
+ /* A stale full bit on word 1 would hand out 128, not 70. */
+ if (dup(0) != 70)
+ exit(EXIT_FAILURE);
+
+ exit(EXIT_SUCCESS);
+ }
+
+ EXPECT_EQ(waitpid(pid, &status, 0), pid);
+ EXPECT_EQ(true, WIFEXITED(status));
+ EXPECT_EQ(0, WEXITSTATUS(status));
+
+ /* The shared table the child unshared from is untouched. */
+ for (i = 3; i < 128; i++)
+ EXPECT_NE(-1, fcntl(i, F_GETFD));
+}
+
TEST(close_range_cloexec)
{
int i, ret;
--
2.53.0
^ permalink raw reply related [flat|nested] 14+ messages in thread* [PATCH 03/10] file: rename dup_fd()'s punch_hole to range
2026-09-21 14:15 [PATCH 00/10] files,close_range: add CLOSE_RANGE_{CLOEXEC_ONLY,EXCEPT} Christian Brauner
2026-09-21 14:15 ` [PATCH 01/10] file: let dup_fd() leave the punched hole behind Christian Brauner
2026-09-21 14:15 ` [PATCH 02/10] selftests/core: test the hole CLOSE_RANGE_UNSHARE leaves behind Christian Brauner
@ 2026-09-21 14:15 ` Christian Brauner
2026-09-21 14:15 ` [PATCH 04/10] file: let dup_fd() drop everything outside of the range Christian Brauner
` (7 subsequent siblings)
10 siblings, 0 replies; 14+ messages in thread
From: Christian Brauner @ 2026-09-21 14:15 UTC (permalink / raw)
To: Jann Horn, linux-fsdevel, Oleg Nesterov
Cc: Alexander Viro, Jan Kara, Neil Brown, Jeff Layton,
Christian Brauner (Amutable)
dup_fd() will be taught what to do with the range that was passed to it.
It won't always be dropped. Rename it to a plain "range" from
"punch_hole". close_range() keeps a pointer named drop for the one case
it has today.
No functional changes.
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
fs/file.c | 28 ++++++++++++++--------------
1 file changed, 14 insertions(+), 14 deletions(-)
diff --git a/fs/file.c b/fs/file.c
index 95dbdfee555a..6234548be88b 100644
--- a/fs/file.c
+++ b/fs/file.c
@@ -365,23 +365,23 @@ static unsigned long fd_range_word(struct fd_range *range, unsigned int i)
}
/* Bits of word @i that dup_fd() leaves behind. */
-static unsigned long dup_fd_dropped_word(unsigned int i, struct fd_range *punch_hole)
+static unsigned long dup_fd_dropped_word(unsigned int i, struct fd_range *range)
{
- if (!punch_hole)
+ if (!range)
return 0;
- return fd_range_word(punch_hole, i);
+ return fd_range_word(range, i);
}
/*
* Note that a sane fdtable size always has to be a multiple of
* BITS_PER_LONG, since we have bitmaps that are sized by this.
*
- * punch_hole is optional - when close_range() is asked to unshare
+ * range is optional - when close_range() is asked to unshare
* and close, dup_fd() leaves the descriptors in that range behind,
* so the cloned table only has to reach the last open descriptor
* outside of it.
*/
-static unsigned int sane_fdtable_size(struct fdtable *fdt, struct fd_range *punch_hole)
+static unsigned int sane_fdtable_size(struct fdtable *fdt, struct fd_range *range)
{
unsigned int last = find_last_bit(fdt->open_fds, fdt->max_fds);
unsigned int i;
@@ -391,7 +391,7 @@ static unsigned int sane_fdtable_size(struct fdtable *fdt, struct fd_range *punc
/* Only words up to the last open descriptor can hold a kept one. */
i = last / BITS_PER_LONG + 1;
while (i--) {
- unsigned long dropped = dup_fd_dropped_word(i, punch_hole);
+ unsigned long dropped = dup_fd_dropped_word(i, range);
if (fdt->open_fds[i] & ~dropped)
return (i + 1) * BITS_PER_LONG;
@@ -402,9 +402,9 @@ static unsigned int sane_fdtable_size(struct fdtable *fdt, struct fd_range *punc
/*
* Allocate a new descriptor table and copy contents from the passed in
* instance. Returns a pointer to cloned table on success, ERR_PTR()
- * on failure. For 'punch_hole' see sane_fdtable_size().
+ * on failure. For 'range' see sane_fdtable_size().
*/
-struct files_struct *dup_fd(struct files_struct *oldf, struct fd_range *punch_hole)
+struct files_struct *dup_fd(struct files_struct *oldf, struct fd_range *range)
{
struct files_struct *newf;
struct file **old_fds, **new_fds;
@@ -431,7 +431,7 @@ struct files_struct *dup_fd(struct files_struct *oldf, struct fd_range *punch_ho
spin_lock(&oldf->file_lock);
old_fdt = files_fdtable(oldf);
- open_files = sane_fdtable_size(old_fdt, punch_hole);
+ open_files = sane_fdtable_size(old_fdt, range);
/*
* Check whether we need to allocate a larger fd array and fd set.
@@ -455,7 +455,7 @@ struct files_struct *dup_fd(struct files_struct *oldf, struct fd_range *punch_ho
*/
spin_lock(&oldf->file_lock);
old_fdt = files_fdtable(oldf);
- open_files = sane_fdtable_size(old_fdt, punch_hole);
+ open_files = sane_fdtable_size(old_fdt, range);
}
copy_fd_bitmaps(new_fdt, old_fdt, open_files / BITS_PER_LONG);
@@ -482,7 +482,7 @@ struct files_struct *dup_fd(struct files_struct *oldf, struct fd_range *punch_ho
struct file *f = rcu_dereference_raw(*old_fds++);
if (!(fd % BITS_PER_LONG))
- dropped = dup_fd_dropped_word(fd / BITS_PER_LONG, punch_hole);
+ dropped = dup_fd_dropped_word(fd / BITS_PER_LONG, range);
if (f && !(dropped & BIT_MASK(fd))) {
get_file(f);
} else {
@@ -858,7 +858,7 @@ SYSCALL_DEFINE3(close_range, unsigned int, fd, unsigned int, max_fd,
return -EINVAL;
if ((flags & CLOSE_RANGE_UNSHARE) && atomic_read(&cur_fds->count) > 1) {
- struct fd_range range = {fd, max_fd}, *punch_hole = ⦥
+ struct fd_range range = {fd, max_fd}, *drop = ⦥
/*
* If the caller requested all fds to be made cloexec we always
@@ -866,9 +866,9 @@ SYSCALL_DEFINE3(close_range, unsigned int, fd, unsigned int, max_fd,
* use them.
*/
if (flags & CLOSE_RANGE_CLOEXEC)
- punch_hole = NULL;
+ drop = NULL;
- fds = dup_fd(cur_fds, punch_hole);
+ fds = dup_fd(cur_fds, drop);
if (IS_ERR(fds))
return PTR_ERR(fds);
/*
--
2.53.0
^ permalink raw reply related [flat|nested] 14+ messages in thread* [PATCH 04/10] file: let dup_fd() drop everything outside of the range
2026-09-21 14:15 [PATCH 00/10] files,close_range: add CLOSE_RANGE_{CLOEXEC_ONLY,EXCEPT} Christian Brauner
` (2 preceding siblings ...)
2026-09-21 14:15 ` [PATCH 03/10] file: rename dup_fd()'s punch_hole to range Christian Brauner
@ 2026-09-21 14:15 ` Christian Brauner
2026-09-21 14:15 ` [PATCH 05/10] close_range: turn the flags into an enum Christian Brauner
` (6 subsequent siblings)
10 siblings, 0 replies; 14+ messages in thread
From: Christian Brauner @ 2026-09-21 14:15 UTC (permalink / raw)
To: Jann Horn, linux-fsdevel, Oleg Nesterov
Cc: Alexander Viro, Jan Kara, Neil Brown, Jeff Layton,
Christian Brauner (Amutable)
Add FD_RANGE_EXCEPT. It turns the meaning of range around. Instead of
indicating that the descriptors in the range are the ones that are left
out of the copy they indicate the range that makes it into the copy. All
other files are left behind. The clone only has to reach the last open
descriptor inside the range.
Nothing passes the flag yet.
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
fs/file.c | 15 ++++++++++-----
include/linux/fdtable.h | 6 ++++++
2 files changed, 16 insertions(+), 5 deletions(-)
diff --git a/fs/file.c b/fs/file.c
index 6234548be88b..ae7d021398dd 100644
--- a/fs/file.c
+++ b/fs/file.c
@@ -367,19 +367,24 @@ static unsigned long fd_range_word(struct fd_range *range, unsigned int i)
/* Bits of word @i that dup_fd() leaves behind. */
static unsigned long dup_fd_dropped_word(unsigned int i, struct fd_range *range)
{
+ unsigned long dropped;
+
if (!range)
return 0;
- return fd_range_word(range, i);
+ dropped = fd_range_word(range, i);
+ if (range->flags & FD_RANGE_EXCEPT)
+ dropped = ~dropped;
+ return dropped;
}
/*
* Note that a sane fdtable size always has to be a multiple of
* BITS_PER_LONG, since we have bitmaps that are sized by this.
*
- * range is optional - when close_range() is asked to unshare
- * and close, dup_fd() leaves the descriptors in that range behind,
- * so the cloned table only has to reach the last open descriptor
- * outside of it.
+ * range is optional. When close_range() is asked to unshare dup_fd()
+ * will leave any files behind according to the range and its flags. The
+ * cloned table only has to reach the last open descriptor that is
+ * carried over.
*/
static unsigned int sane_fdtable_size(struct fdtable *fdt, struct fd_range *range)
{
diff --git a/include/linux/fdtable.h b/include/linux/fdtable.h
index c45306a9f007..d6c6c7a3400d 100644
--- a/include/linux/fdtable.h
+++ b/include/linux/fdtable.h
@@ -101,8 +101,14 @@ struct task_struct;
void put_files_struct(struct files_struct *fs);
int unshare_files(void);
+enum fd_range_flags {
+ /* Leave behind all descriptors outside of the specified range. */
+ FD_RANGE_EXCEPT = (1U << 0),
+};
+
struct fd_range {
unsigned int from, to;
+ enum fd_range_flags flags;
};
struct files_struct *dup_fd(struct files_struct *, struct fd_range *) __latent_entropy;
void do_close_on_exec(struct files_struct *);
--
2.53.0
^ permalink raw reply related [flat|nested] 14+ messages in thread* [PATCH 05/10] close_range: turn the flags into an enum
2026-09-21 14:15 [PATCH 00/10] files,close_range: add CLOSE_RANGE_{CLOEXEC_ONLY,EXCEPT} Christian Brauner
` (3 preceding siblings ...)
2026-09-21 14:15 ` [PATCH 04/10] file: let dup_fd() drop everything outside of the range Christian Brauner
@ 2026-09-21 14:15 ` Christian Brauner
2026-09-21 14:15 ` [PATCH 06/10] file: add CLOSE_RANGE_EXCEPT Christian Brauner
` (5 subsequent siblings)
10 siblings, 0 replies; 14+ messages in thread
From: Christian Brauner @ 2026-09-21 14:15 UTC (permalink / raw)
To: Jann Horn, linux-fsdevel, Oleg Nesterov
Cc: Alexander Viro, Jan Kara, Neil Brown, Jeff Layton,
Christian Brauner (Amutable)
Turn the close_range() flags into an enum. Makes debugging a lot easier.
No functional changes.
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
include/uapi/linux/close_range.h | 21 +++++++++++++++++----
1 file changed, 17 insertions(+), 4 deletions(-)
diff --git a/include/uapi/linux/close_range.h b/include/uapi/linux/close_range.h
index 2d804281554c..b1b637d74deb 100644
--- a/include/uapi/linux/close_range.h
+++ b/include/uapi/linux/close_range.h
@@ -2,11 +2,24 @@
#ifndef _UAPI_LINUX_CLOSE_RANGE_H
#define _UAPI_LINUX_CLOSE_RANGE_H
-/* Unshare the file descriptor table before closing file descriptors. */
-#define CLOSE_RANGE_UNSHARE (1U << 1)
+/*
+ * A macro of one of these names defined before this header is parsed, by
+ * a libc or by a program's own fallback, would replace the enumerator.
+ */
+#undef CLOSE_RANGE_UNSHARE
+#undef CLOSE_RANGE_CLOEXEC
-/* Set the FD_CLOEXEC bit instead of closing the file descriptor. */
-#define CLOSE_RANGE_CLOEXEC (1U << 2)
+enum close_range_flags {
+ /* Unshare the file descriptor table before closing file descriptors. */
+ CLOSE_RANGE_UNSHARE = (1U << 1),
+
+ /* Set the FD_CLOEXEC bit instead of closing the file descriptor. */
+ CLOSE_RANGE_CLOEXEC = (1U << 2),
+};
+
+/* Keep #ifdef working and let glibc skip its own definitions. */
+#define CLOSE_RANGE_UNSHARE CLOSE_RANGE_UNSHARE
+#define CLOSE_RANGE_CLOEXEC CLOSE_RANGE_CLOEXEC
#endif /* _UAPI_LINUX_CLOSE_RANGE_H */
--
2.53.0
^ permalink raw reply related [flat|nested] 14+ messages in thread* [PATCH 06/10] file: add CLOSE_RANGE_EXCEPT
2026-09-21 14:15 [PATCH 00/10] files,close_range: add CLOSE_RANGE_{CLOEXEC_ONLY,EXCEPT} Christian Brauner
` (4 preceding siblings ...)
2026-09-21 14:15 ` [PATCH 05/10] close_range: turn the flags into an enum Christian Brauner
@ 2026-09-21 14:15 ` Christian Brauner
2026-09-21 14:15 ` [PATCH 07/10] selftests/core: test CLOSE_RANGE_EXCEPT Christian Brauner
` (4 subsequent siblings)
10 siblings, 0 replies; 14+ messages in thread
From: Christian Brauner @ 2026-09-21 14:15 UTC (permalink / raw)
To: Jann Horn, linux-fsdevel, Oleg Nesterov
Cc: Alexander Viro, Jan Kara, Neil Brown, Jeff Layton,
Christian Brauner (Amutable)
close_range() operates on [fd, max_fd]. Add CLOSE_RANGE_EXCEPT. This new
flag instructs it to operate on all file descriptors outside of the
range instead.
Without any other flags that means it closes everything except the file
descriptors in the specified range. Together with CLOSE_RANGE_CLOEXEC it
marks every file descriptor except the ones in the range as
close-on-exec. With CLOSE_RANGE_UNSHARE dup_fd() never takes a reference
on what is dropped.
A task between clone(CLONE_FILES | CLONE_VM | CLONE_VFORK) and execve()
that has the descriptors for the child in one window can shed the rest
of the shared table in one call:
close_range(lo, hi, CLOSE_RANGE_UNSHARE | CLOSE_RANGE_EXCEPT)
Nothing outside of the window is referenced by the child at any point.
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
fs/file.c | 76 +++++++++++++++++++++++++++++++---------
include/uapi/linux/close_range.h | 5 +++
2 files changed, 64 insertions(+), 17 deletions(-)
diff --git a/fs/file.c b/fs/file.c
index ae7d021398dd..39651c7a992e 100644
--- a/fs/file.c
+++ b/fs/file.c
@@ -794,34 +794,67 @@ static inline unsigned last_fd(struct fdtable *fdt)
}
static inline void __range_cloexec(struct files_struct *cur_fds,
- unsigned int fd, unsigned int max_fd)
+ struct fd_range *range)
{
struct fdtable *fdt;
+ unsigned int last;
- /* make sure we're using the correct maximum value */
spin_lock(&cur_fds->file_lock);
fdt = files_fdtable(cur_fds);
- max_fd = min(last_fd(fdt), max_fd);
- if (fd <= max_fd)
- bitmap_set(fdt->close_on_exec, fd, max_fd - fd + 1);
+ /* make sure we're using the correct maximum value */
+ last = last_fd(fdt);
+ if (!(range->flags & FD_RANGE_EXCEPT)) {
+ if (range->from <= last)
+ bitmap_set(fdt->close_on_exec, range->from,
+ min(range->to, last) - range->from + 1);
+ } else {
+ if (range->from > 0)
+ bitmap_set(fdt->close_on_exec, 0,
+ min(range->from - 1, last) + 1);
+ if (range->to < last)
+ bitmap_set(fdt->close_on_exec, range->to + 1,
+ last - range->to);
+ }
spin_unlock(&cur_fds->file_lock);
}
-static inline void __range_close(struct files_struct *files, unsigned int fd,
- unsigned int max_fd)
+/* Next open descriptor in [fd, max_fd] that @range selects. */
+static inline unsigned int next_fd_to_close(struct fdtable *fdt,
+ unsigned int fd, unsigned int max_fd,
+ struct fd_range *range)
+{
+ fd = find_next_bit(fdt->open_fds, max_fd + 1, fd);
+ /* Hop over the window the range keeps. */
+ if ((range->flags & FD_RANGE_EXCEPT) &&
+ fd >= range->from && fd <= range->to) {
+ if (range->to >= max_fd)
+ return max_fd + 1;
+ fd = find_next_bit(fdt->open_fds, max_fd + 1, range->to + 1);
+ }
+ return fd;
+}
+
+static inline void __range_close(struct files_struct *files,
+ struct fd_range *range)
{
struct file *file;
struct fdtable *fdt;
- unsigned n;
+ unsigned int fd, max_fd;
spin_lock(&files->file_lock);
fdt = files_fdtable(files);
- n = last_fd(fdt);
- max_fd = min(max_fd, n);
+ if (range->flags & FD_RANGE_EXCEPT) {
+ /* Outside of the range means the whole table. */
+ fd = 0;
+ max_fd = last_fd(fdt);
+ } else {
+ fd = range->from;
+ max_fd = min(range->to, last_fd(fdt));
+ }
- for (fd = find_next_bit(fdt->open_fds, max_fd + 1, fd);
+ for (fd = next_fd_to_close(fdt, fd, max_fd, range);
fd <= max_fd;
- fd = find_next_bit(fdt->open_fds, max_fd + 1, fd + 1)) {
+ fd = next_fd_to_close(fdt, fd + 1, max_fd, range)) {
file = file_close_fd_locked(files, fd);
if (file) {
spin_unlock(&files->file_lock);
@@ -849,21 +882,30 @@ static inline void __range_close(struct files_struct *files, unsigned int fd,
* This closes a range of file descriptors. All file descriptors
* from @fd up to and including @max_fd are closed.
* Currently, errors to close a given file descriptor are ignored.
+ *
+ * With CLOSE_RANGE_EXCEPT the range names what to leave alone instead:
+ * every open file descriptor outside of [@fd, @max_fd] is closed, or
+ * marked close-on-exec with CLOSE_RANGE_CLOEXEC.
*/
SYSCALL_DEFINE3(close_range, unsigned int, fd, unsigned int, max_fd,
unsigned int, flags)
{
struct task_struct *me = current;
struct files_struct *cur_fds = me->files, *fds = NULL;
+ struct fd_range range = {fd, max_fd};
- if (flags & ~(CLOSE_RANGE_UNSHARE | CLOSE_RANGE_CLOEXEC))
+ if (flags & ~(CLOSE_RANGE_UNSHARE | CLOSE_RANGE_CLOEXEC |
+ CLOSE_RANGE_EXCEPT))
return -EINVAL;
if (fd > max_fd)
return -EINVAL;
+ if (flags & CLOSE_RANGE_EXCEPT)
+ range.flags |= FD_RANGE_EXCEPT;
+
if ((flags & CLOSE_RANGE_UNSHARE) && atomic_read(&cur_fds->count) > 1) {
- struct fd_range range = {fd, max_fd}, *drop = ⦥
+ struct fd_range *drop = ⦥
/*
* If the caller requested all fds to be made cloexec we always
@@ -884,10 +926,10 @@ SYSCALL_DEFINE3(close_range, unsigned int, fd, unsigned int, max_fd,
}
if (flags & CLOSE_RANGE_CLOEXEC) {
- __range_cloexec(cur_fds, fd, max_fd);
+ __range_cloexec(cur_fds, &range);
} else if (!fds) {
- /* If we unshared, dup_fd() left the range behind already. */
- __range_close(cur_fds, fd, max_fd);
+ /* If we unshared, dup_fd() already left behind what we'd close. */
+ __range_close(cur_fds, &range);
}
if (fds) {
diff --git a/include/uapi/linux/close_range.h b/include/uapi/linux/close_range.h
index b1b637d74deb..39eddb3ab613 100644
--- a/include/uapi/linux/close_range.h
+++ b/include/uapi/linux/close_range.h
@@ -8,6 +8,7 @@
*/
#undef CLOSE_RANGE_UNSHARE
#undef CLOSE_RANGE_CLOEXEC
+#undef CLOSE_RANGE_EXCEPT
enum close_range_flags {
/* Unshare the file descriptor table before closing file descriptors. */
@@ -15,11 +16,15 @@ enum close_range_flags {
/* Set the FD_CLOEXEC bit instead of closing the file descriptor. */
CLOSE_RANGE_CLOEXEC = (1U << 2),
+
+ /* Act on every file descriptor outside of the given range instead. */
+ CLOSE_RANGE_EXCEPT = (1U << 3),
};
/* Keep #ifdef working and let glibc skip its own definitions. */
#define CLOSE_RANGE_UNSHARE CLOSE_RANGE_UNSHARE
#define CLOSE_RANGE_CLOEXEC CLOSE_RANGE_CLOEXEC
+#define CLOSE_RANGE_EXCEPT CLOSE_RANGE_EXCEPT
#endif /* _UAPI_LINUX_CLOSE_RANGE_H */
--
2.53.0
^ permalink raw reply related [flat|nested] 14+ messages in thread* [PATCH 07/10] selftests/core: test CLOSE_RANGE_EXCEPT
2026-09-21 14:15 [PATCH 00/10] files,close_range: add CLOSE_RANGE_{CLOEXEC_ONLY,EXCEPT} Christian Brauner
` (5 preceding siblings ...)
2026-09-21 14:15 ` [PATCH 06/10] file: add CLOSE_RANGE_EXCEPT Christian Brauner
@ 2026-09-21 14:15 ` Christian Brauner
2026-09-21 14:15 ` [PATCH 08/10] file: let dup_fd() drop only close-on-exec descriptors Christian Brauner
` (3 subsequent siblings)
10 siblings, 0 replies; 14+ messages in thread
From: Christian Brauner @ 2026-09-21 14:15 UTC (permalink / raw)
To: Jann Horn, linux-fsdevel, Oleg Nesterov
Cc: Alexander Viro, Jan Kara, Neil Brown, Jeff Layton,
Christian Brauner (Amutable)
Cover the inverted range:
- with plain close the window stays and everything outside of it goes,
stdio included
- a window at the top, one that cannot hold a descriptor and one of a
single descriptor keep just what they name, and so does the unshare
form on a table that is not shared
- with CLOSE_RANGE_CLOEXEC everything outside of the window is marked
and the window is not, in place and in a clone, for a
window in the middle, at the bottom, at the top and above the table
- with CLOSE_RANGE_UNSHARE the table the child cloned from is untouched,
a window near the top of the table comes back whole, one at the bottom
keeps stdio, one that cannot hold a descriptor keeps nothing, and the
slots left behind are handed out again from the bottom
- the bounds are checked before the range is turned around
The extra cases came out of walking the window positions that
__range_close(), __range_cloexec() and dup_fd() tell apart: at the
bottom, in the middle, at the top, above the table, a single descriptor
and none, on a shared and on a private table.
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
tools/testing/selftests/core/close_range_test.c | 426 ++++++++++++++++++++++++
1 file changed, 426 insertions(+)
diff --git a/tools/testing/selftests/core/close_range_test.c b/tools/testing/selftests/core/close_range_test.c
index afae3462502d..caeb2f1ea800 100644
--- a/tools/testing/selftests/core/close_range_test.c
+++ b/tools/testing/selftests/core/close_range_test.c
@@ -36,6 +36,14 @@ static inline int sys_close_range(unsigned int fd, unsigned int max_fd,
return syscall(__NR_close_range, fd, max_fd, flags);
}
+static void clear_cloexec(const int *fds, size_t n)
+{
+ size_t i;
+
+ for (i = 0; i < n; i++)
+ fcntl(fds[i], F_SETFD, 0);
+}
+
TEST(core_close_range)
{
int i, ret;
@@ -637,6 +645,424 @@ TEST(close_range_cloexec_unshare_syzbot)
EXPECT_EQ(close(fd3), 0);
}
+TEST(close_range_except)
+{
+ int i, ret, status;
+ pid_t pid;
+ int open_fds[101];
+ struct __clone_args args = {
+ .exit_signal = SIGCHLD,
+ };
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++) {
+ int fd;
+
+ fd = open("/dev/null", O_RDONLY);
+ ASSERT_GE(fd, 0) {
+ if (errno == ENOENT)
+ SKIP(return, "Skipping test since /dev/null does not exist");
+ }
+
+ open_fds[i] = fd;
+ }
+
+ /* A range covering everything keeps everything. */
+ ret = sys_close_range(0, UINT_MAX, CLOSE_RANGE_EXCEPT);
+ if (ret < 0) {
+ if (errno == ENOSYS)
+ SKIP(return, "close_range() syscall not supported");
+ if (errno == EINVAL)
+ SKIP(return, "close_range() doesn't support CLOSE_RANGE_EXCEPT");
+ }
+ ASSERT_EQ(0, ret);
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++)
+ EXPECT_NE(-1, fcntl(open_fds[i], F_GETFD));
+
+ /* The bounds are checked before the range is turned around. */
+ EXPECT_EQ(-1, sys_close_range(open_fds[20], open_fds[10],
+ CLOSE_RANGE_EXCEPT));
+ EXPECT_EQ(EINVAL, errno);
+
+ /* Everything above open_fds[50] goes. */
+ ASSERT_EQ(0, sys_close_range(0, open_fds[50], CLOSE_RANGE_EXCEPT));
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++)
+ EXPECT_EQ(i <= 50, fcntl(open_fds[i], F_GETFD) != -1);
+
+ /* A window in the middle takes stdio with it, so do that in a fork. */
+ pid = sys_clone3(&args, sizeof(args));
+ ASSERT_GE(pid, 0);
+
+ if (pid == 0) {
+ ret = sys_close_range(open_fds[10], open_fds[20],
+ CLOSE_RANGE_EXCEPT);
+ if (ret)
+ exit(EXIT_FAILURE);
+
+ for (i = 0; i <= 50; i++) {
+ bool kept = i >= 10 && i <= 20;
+
+ if (kept != (fcntl(open_fds[i], F_GETFD) != -1))
+ exit(EXIT_FAILURE);
+ }
+
+ if (fcntl(STDERR_FILENO, F_GETFD) != -1)
+ exit(EXIT_FAILURE);
+
+ exit(EXIT_SUCCESS);
+ }
+
+ EXPECT_EQ(waitpid(pid, &status, 0), pid);
+ EXPECT_EQ(true, WIFEXITED(status));
+ EXPECT_EQ(0, WEXITSTATUS(status));
+
+ /* The fork had a table of its own. */
+ for (i = 0; i <= 50; i++)
+ EXPECT_NE(-1, fcntl(open_fds[i], F_GETFD));
+}
+
+TEST(close_range_except_cloexec)
+{
+ int i, ret;
+ int open_fds[101];
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++) {
+ int fd;
+
+ fd = open("/dev/null", O_RDONLY);
+ ASSERT_GE(fd, 0) {
+ if (errno == ENOENT)
+ SKIP(return, "Skipping test since /dev/null does not exist");
+ }
+
+ open_fds[i] = fd;
+ }
+
+ ret = sys_close_range(open_fds[10], open_fds[20],
+ CLOSE_RANGE_CLOEXEC | CLOSE_RANGE_EXCEPT);
+ if (ret < 0) {
+ if (errno == ENOSYS)
+ SKIP(return, "close_range() syscall not supported");
+ if (errno == EINVAL)
+ SKIP(return, "close_range() doesn't support CLOSE_RANGE_EXCEPT");
+ }
+ ASSERT_EQ(0, ret);
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++) {
+ bool inside = i >= 10 && i <= 20;
+ int flags = fcntl(open_fds[i], F_GETFD);
+
+ EXPECT_NE(-1, flags);
+ EXPECT_EQ(inside ? 0 : FD_CLOEXEC, flags & FD_CLOEXEC);
+ }
+
+ /* stdio sits outside of the window too. */
+ EXPECT_EQ(FD_CLOEXEC, fcntl(STDERR_FILENO, F_GETFD) & FD_CLOEXEC);
+
+ /* A window that starts at 0 marks only what lies above it. */
+ clear_cloexec(open_fds, ARRAY_SIZE(open_fds));
+ ASSERT_EQ(0, fcntl(STDERR_FILENO, F_SETFD, 0));
+ ASSERT_EQ(0, sys_close_range(0, open_fds[20],
+ CLOSE_RANGE_CLOEXEC | CLOSE_RANGE_EXCEPT));
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++)
+ EXPECT_EQ(i <= 20 ? 0 : FD_CLOEXEC,
+ fcntl(open_fds[i], F_GETFD) & FD_CLOEXEC);
+ EXPECT_EQ(0, fcntl(STDERR_FILENO, F_GETFD) & FD_CLOEXEC);
+
+ /* One at the top marks only what lies below it, stdio included. */
+ clear_cloexec(open_fds, ARRAY_SIZE(open_fds));
+ ASSERT_EQ(0, sys_close_range(open_fds[80], UINT_MAX,
+ CLOSE_RANGE_CLOEXEC | CLOSE_RANGE_EXCEPT));
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++)
+ EXPECT_EQ(i < 80 ? FD_CLOEXEC : 0,
+ fcntl(open_fds[i], F_GETFD) & FD_CLOEXEC);
+ EXPECT_EQ(FD_CLOEXEC, fcntl(STDERR_FILENO, F_GETFD) & FD_CLOEXEC);
+
+ /* One that cannot hold a descriptor marks everything. */
+ clear_cloexec(open_fds, ARRAY_SIZE(open_fds));
+ ASSERT_EQ(0, fcntl(STDERR_FILENO, F_SETFD, 0));
+ ASSERT_EQ(0, sys_close_range(UINT_MAX, UINT_MAX,
+ CLOSE_RANGE_CLOEXEC | CLOSE_RANGE_EXCEPT));
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++)
+ EXPECT_EQ(FD_CLOEXEC, fcntl(open_fds[i], F_GETFD) & FD_CLOEXEC);
+ EXPECT_EQ(FD_CLOEXEC, fcntl(STDERR_FILENO, F_GETFD) & FD_CLOEXEC);
+}
+
+TEST(close_range_except_cloexec_unshare)
+{
+ int i, ret, status;
+ pid_t pid;
+ int open_fds[101];
+ struct __clone_args args = {
+ .flags = CLONE_FILES,
+ .exit_signal = SIGCHLD,
+ };
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++) {
+ int fd;
+
+ fd = open("/dev/null", O_RDONLY);
+ ASSERT_GE(fd, 0) {
+ if (errno == ENOENT)
+ SKIP(return, "Skipping test since /dev/null does not exist");
+ }
+
+ open_fds[i] = fd;
+ }
+
+ /* A range covering everything marks nothing. */
+ ret = sys_close_range(0, UINT_MAX,
+ CLOSE_RANGE_CLOEXEC | CLOSE_RANGE_EXCEPT);
+ if (ret < 0) {
+ if (errno == ENOSYS)
+ SKIP(return, "close_range() syscall not supported");
+ if (errno == EINVAL)
+ SKIP(return, "close_range() doesn't support CLOSE_RANGE_EXCEPT");
+ }
+ ASSERT_EQ(0, ret);
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++)
+ EXPECT_EQ(0, fcntl(open_fds[i], F_GETFD) & FD_CLOEXEC);
+
+ pid = sys_clone3(&args, sizeof(args));
+ ASSERT_GE(pid, 0);
+
+ if (pid == 0) {
+ ret = sys_close_range(open_fds[10], open_fds[20],
+ CLOSE_RANGE_UNSHARE | CLOSE_RANGE_CLOEXEC |
+ CLOSE_RANGE_EXCEPT);
+ if (ret)
+ exit(EXIT_FAILURE);
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++) {
+ bool inside = i >= 10 && i <= 20;
+ int flags = fcntl(open_fds[i], F_GETFD);
+
+ if (flags == -1)
+ exit(EXIT_FAILURE);
+ if ((flags & FD_CLOEXEC) != (inside ? 0 : FD_CLOEXEC))
+ exit(EXIT_FAILURE);
+ }
+
+ exit(EXIT_SUCCESS);
+ }
+
+ EXPECT_EQ(waitpid(pid, &status, 0), pid);
+ EXPECT_EQ(true, WIFEXITED(status));
+ EXPECT_EQ(0, WEXITSTATUS(status));
+
+ /* The shared table the child unshared from is untouched. */
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++)
+ EXPECT_EQ(0, fcntl(open_fds[i], F_GETFD) & FD_CLOEXEC);
+}
+
+TEST(close_range_except_bounds)
+{
+ int i, c, ret, status;
+ pid_t pid;
+ int open_fds[101];
+ struct __clone_args args = {
+ .exit_signal = SIGCHLD,
+ };
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++) {
+ int fd;
+
+ fd = open("/dev/null", O_RDONLY);
+ ASSERT_GE(fd, 0) {
+ if (errno == ENOENT)
+ SKIP(return, "Skipping test since /dev/null does not exist");
+ }
+
+ open_fds[i] = fd;
+ }
+
+ /* A range covering everything keeps everything. */
+ ret = sys_close_range(0, UINT_MAX, CLOSE_RANGE_EXCEPT);
+ if (ret < 0) {
+ if (errno == ENOSYS)
+ SKIP(return, "close_range() syscall not supported");
+ if (errno == EINVAL)
+ SKIP(return, "close_range() doesn't support CLOSE_RANGE_EXCEPT");
+ }
+ ASSERT_EQ(0, ret);
+
+ struct {
+ unsigned int fd, max_fd, flags;
+ } cases[] = {
+ /* A window at the top drops everything below it. */
+ { open_fds[50], UINT_MAX, CLOSE_RANGE_EXCEPT },
+ /* One that cannot hold a descriptor keeps nothing. */
+ { UINT_MAX, UINT_MAX, CLOSE_RANGE_EXCEPT },
+ /* One of a single descriptor keeps just that. */
+ { open_fds[30], open_fds[30], CLOSE_RANGE_EXCEPT },
+ /* The unshare form on a table that is not shared acts in place. */
+ { open_fds[10], open_fds[20],
+ CLOSE_RANGE_UNSHARE | CLOSE_RANGE_EXCEPT },
+ };
+
+ /* Each of them takes stdio with it, so do that in a fork. */
+ for (c = 0; c < ARRAY_SIZE(cases); c++) {
+ pid = sys_clone3(&args, sizeof(args));
+ ASSERT_GE(pid, 0);
+
+ if (pid == 0) {
+ ret = sys_close_range(cases[c].fd, cases[c].max_fd,
+ cases[c].flags);
+ if (ret)
+ exit(EXIT_FAILURE);
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++) {
+ unsigned int fd = open_fds[i];
+ bool kept = fd >= cases[c].fd &&
+ fd <= cases[c].max_fd;
+
+ if (kept != (fcntl(fd, F_GETFD) != -1))
+ exit(EXIT_FAILURE);
+ }
+
+ if (fcntl(STDERR_FILENO, F_GETFD) != -1)
+ exit(EXIT_FAILURE);
+
+ exit(EXIT_SUCCESS);
+ }
+
+ EXPECT_EQ(waitpid(pid, &status, 0), pid);
+ EXPECT_EQ(true, WIFEXITED(status));
+ EXPECT_EQ(0, WEXITSTATUS(status));
+ }
+
+ /* Each fork had a table of its own. */
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++)
+ EXPECT_NE(-1, fcntl(open_fds[i], F_GETFD));
+}
+
+TEST(close_range_except_unshare)
+{
+ int i, ret, status;
+ pid_t pid;
+ int open_fds[200];
+ struct __clone_args args = {
+ .flags = CLONE_FILES,
+ .exit_signal = SIGCHLD,
+ };
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++) {
+ int fd;
+
+ /* Odd slots are close-on-exec, which makes no difference here. */
+ fd = open("/dev/null", O_RDONLY | (i % 2 ? O_CLOEXEC : 0));
+ ASSERT_GE(fd, 0) {
+ if (errno == ENOENT)
+ SKIP(return, "Skipping test since /dev/null does not exist");
+ }
+
+ open_fds[i] = fd;
+ }
+
+ /* A range covering everything keeps everything. */
+ ret = sys_close_range(0, UINT_MAX, CLOSE_RANGE_EXCEPT);
+ if (ret < 0) {
+ if (errno == ENOSYS)
+ SKIP(return, "close_range() syscall not supported");
+ if (errno == EINVAL)
+ SKIP(return, "close_range() doesn't support CLOSE_RANGE_EXCEPT");
+ }
+ ASSERT_EQ(0, ret);
+
+ pid = sys_clone3(&args, sizeof(args));
+ ASSERT_GE(pid, 0);
+
+ if (pid == 0) {
+ /* The window sits near the top, so the clone is sized off its end. */
+ ret = sys_close_range(open_fds[150], open_fds[160],
+ CLOSE_RANGE_UNSHARE | CLOSE_RANGE_EXCEPT);
+ if (ret)
+ exit(EXIT_FAILURE);
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++) {
+ bool kept = i >= 150 && i <= 160;
+
+ if (kept != (fcntl(open_fds[i], F_GETFD) != -1))
+ exit(EXIT_FAILURE);
+ }
+
+ if (fcntl(STDERR_FILENO, F_GETFD) != -1)
+ exit(EXIT_FAILURE);
+
+ /* What was left behind is handed out again, from the bottom. */
+ if (dup(open_fds[150]) != 0)
+ exit(EXIT_FAILURE);
+
+ exit(EXIT_SUCCESS);
+ }
+
+ EXPECT_EQ(waitpid(pid, &status, 0), pid);
+ EXPECT_EQ(true, WIFEXITED(status));
+ EXPECT_EQ(0, WEXITSTATUS(status));
+
+ /* A window at the bottom keeps just that, stdio included. */
+ pid = sys_clone3(&args, sizeof(args));
+ ASSERT_GE(pid, 0);
+
+ if (pid == 0) {
+ ret = sys_close_range(0, open_fds[10],
+ CLOSE_RANGE_UNSHARE | CLOSE_RANGE_EXCEPT);
+ if (ret)
+ exit(EXIT_FAILURE);
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++) {
+ if ((i <= 10) != (fcntl(open_fds[i], F_GETFD) != -1))
+ exit(EXIT_FAILURE);
+ }
+
+ if (fcntl(STDERR_FILENO, F_GETFD) == -1)
+ exit(EXIT_FAILURE);
+
+ /* The first slot left behind is the next one handed out. */
+ if (dup(0) != open_fds[10] + 1)
+ exit(EXIT_FAILURE);
+
+ exit(EXIT_SUCCESS);
+ }
+
+ EXPECT_EQ(waitpid(pid, &status, 0), pid);
+ EXPECT_EQ(true, WIFEXITED(status));
+ EXPECT_EQ(0, WEXITSTATUS(status));
+
+ /* A window that cannot hold a descriptor keeps nothing. */
+ pid = sys_clone3(&args, sizeof(args));
+ ASSERT_GE(pid, 0);
+
+ if (pid == 0) {
+ ret = sys_close_range(UINT_MAX, UINT_MAX,
+ CLOSE_RANGE_UNSHARE | CLOSE_RANGE_EXCEPT);
+ if (ret)
+ exit(EXIT_FAILURE);
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++)
+ if (fcntl(open_fds[i], F_GETFD) != -1)
+ exit(EXIT_FAILURE);
+
+ if (fcntl(STDERR_FILENO, F_GETFD) != -1)
+ exit(EXIT_FAILURE);
+
+ exit(EXIT_SUCCESS);
+ }
+
+ EXPECT_EQ(waitpid(pid, &status, 0), pid);
+ EXPECT_EQ(true, WIFEXITED(status));
+ EXPECT_EQ(0, WEXITSTATUS(status));
+
+ /* The shared table the child unshared from is untouched. */
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++)
+ EXPECT_NE(-1, fcntl(open_fds[i], F_GETFD));
+}
+
TEST(close_range_bitmap_corruption)
{
pid_t pid;
--
2.53.0
^ permalink raw reply related [flat|nested] 14+ messages in thread* [PATCH 08/10] file: let dup_fd() drop only close-on-exec descriptors
2026-09-21 14:15 [PATCH 00/10] files,close_range: add CLOSE_RANGE_{CLOEXEC_ONLY,EXCEPT} Christian Brauner
` (6 preceding siblings ...)
2026-09-21 14:15 ` [PATCH 07/10] selftests/core: test CLOSE_RANGE_EXCEPT Christian Brauner
@ 2026-09-21 14:15 ` Christian Brauner
2026-09-21 16:24 ` Jann Horn
2026-09-21 14:15 ` [PATCH 09/10] file: add CLOSE_RANGE_CLOEXEC_ONLY Christian Brauner
` (2 subsequent siblings)
10 siblings, 1 reply; 14+ messages in thread
From: Christian Brauner @ 2026-09-21 14:15 UTC (permalink / raw)
To: Jann Horn, linux-fsdevel, Oleg Nesterov
Cc: Alexander Viro, Jan Kara, Neil Brown, Jeff Layton,
Christian Brauner (Amutable)
Add FD_RANGE_CLOEXEC_ONLY. When set dup_fd() leaves everything behind
except for fds that are close-on-exec. Without FD_RANGE_EXCEPT the
clone loses the close-on-exec descriptors in the range. With it the
clone keeps the range and loses the close-on-exec descriptors everywhere
else. Descriptors without the flag are carried over either way.
Nothing passes the flag yet.
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
fs/file.c | 43 ++++++++++++++++++++++++++++++++++---------
include/linux/fdtable.h | 3 +++
2 files changed, 37 insertions(+), 9 deletions(-)
diff --git a/fs/file.c b/fs/file.c
index 39651c7a992e..8952efd2cf3a 100644
--- a/fs/file.c
+++ b/fs/file.c
@@ -365,7 +365,8 @@ static unsigned long fd_range_word(struct fd_range *range, unsigned int i)
}
/* Bits of word @i that dup_fd() leaves behind. */
-static unsigned long dup_fd_dropped_word(unsigned int i, struct fd_range *range)
+static unsigned long dup_fd_dropped_word(struct fdtable *fdt, unsigned int i,
+ struct fd_range *range)
{
unsigned long dropped;
@@ -374,6 +375,8 @@ static unsigned long dup_fd_dropped_word(unsigned int i, struct fd_range *range)
dropped = fd_range_word(range, i);
if (range->flags & FD_RANGE_EXCEPT)
dropped = ~dropped;
+ if (range->flags & FD_RANGE_CLOEXEC_ONLY)
+ dropped &= fdt->close_on_exec[i];
return dropped;
}
@@ -393,15 +396,37 @@ static unsigned int sane_fdtable_size(struct fdtable *fdt, struct fd_range *rang
if (last == fdt->max_fds)
return NR_OPEN_DEFAULT;
- /* Only words up to the last open descriptor can hold a kept one. */
- i = last / BITS_PER_LONG + 1;
- while (i--) {
- unsigned long dropped = dup_fd_dropped_word(i, range);
+ if (!range)
+ return ALIGN(last + 1, BITS_PER_LONG);
+
+ if (range->flags & FD_RANGE_CLOEXEC_ONLY) {
+ /* The close-on-exec bits decide what is dropped, walk the words. */
+ i = last / BITS_PER_LONG + 1;
+ while (i--) {
+ unsigned long dropped = dup_fd_dropped_word(fdt, i, range);
+
+ if (fdt->open_fds[i] & ~dropped)
+ return (i + 1) * BITS_PER_LONG;
+ }
+ return NR_OPEN_DEFAULT;
+ }
- if (fdt->open_fds[i] & ~dropped)
- return (i + 1) * BITS_PER_LONG;
+ if (range->flags & FD_RANGE_EXCEPT) {
+ /* Only the range is carried over. */
+ if (last > range->to) {
+ last = find_last_bit(fdt->open_fds, range->to + 1);
+ if (last > range->to)
+ return NR_OPEN_DEFAULT;
+ }
+ if (last < range->from)
+ return NR_OPEN_DEFAULT;
+ } else if (last >= range->from && last <= range->to) {
+ /* The last open descriptor goes, the kept ones sit below the range. */
+ last = find_last_bit(fdt->open_fds, range->from);
+ if (last == range->from)
+ return NR_OPEN_DEFAULT;
}
- return NR_OPEN_DEFAULT;
+ return ALIGN(last + 1, BITS_PER_LONG);
}
/*
@@ -487,7 +512,7 @@ struct files_struct *dup_fd(struct files_struct *oldf, struct fd_range *range)
struct file *f = rcu_dereference_raw(*old_fds++);
if (!(fd % BITS_PER_LONG))
- dropped = dup_fd_dropped_word(fd / BITS_PER_LONG, range);
+ dropped = dup_fd_dropped_word(old_fdt, fd / BITS_PER_LONG, range);
if (f && !(dropped & BIT_MASK(fd))) {
get_file(f);
} else {
diff --git a/include/linux/fdtable.h b/include/linux/fdtable.h
index d6c6c7a3400d..afdfaad381df 100644
--- a/include/linux/fdtable.h
+++ b/include/linux/fdtable.h
@@ -104,6 +104,9 @@ int unshare_files(void);
enum fd_range_flags {
/* Leave behind all descriptors outside of the specified range. */
FD_RANGE_EXCEPT = (1U << 0),
+
+ /* Only select descriptors that have close-on-exec set. */
+ FD_RANGE_CLOEXEC_ONLY = (1U << 1),
};
struct fd_range {
--
2.53.0
^ permalink raw reply related [flat|nested] 14+ messages in thread* Re: [PATCH 08/10] file: let dup_fd() drop only close-on-exec descriptors
2026-09-21 14:15 ` [PATCH 08/10] file: let dup_fd() drop only close-on-exec descriptors Christian Brauner
@ 2026-09-21 16:24 ` Jann Horn
2026-09-25 14:57 ` Christian Brauner
0 siblings, 1 reply; 14+ messages in thread
From: Jann Horn @ 2026-09-21 16:24 UTC (permalink / raw)
To: Christian Brauner
Cc: linux-fsdevel, Oleg Nesterov, Alexander Viro, Jan Kara,
Neil Brown, Jeff Layton
On Mon, Sep 21, 2026 at 4:16 PM Christian Brauner <brauner@kernel.org> wrote:
> Add FD_RANGE_CLOEXEC_ONLY. When set dup_fd() leaves everything behind
> except for fds that are close-on-exec. Without FD_RANGE_EXCEPT the
> clone loses the close-on-exec descriptors in the range. With it the
> clone keeps the range and loses the close-on-exec descriptors everywhere
> else. Descriptors without the flag are carried over either way.
[...]
> @@ -393,15 +396,37 @@ static unsigned int sane_fdtable_size(struct fdtable *fdt, struct fd_range *rang
>
> if (last == fdt->max_fds)
> return NR_OPEN_DEFAULT;
> - /* Only words up to the last open descriptor can hold a kept one. */
> - i = last / BITS_PER_LONG + 1;
> - while (i--) {
> - unsigned long dropped = dup_fd_dropped_word(i, range);
> + if (!range)
> + return ALIGN(last + 1, BITS_PER_LONG);
> +
> + if (range->flags & FD_RANGE_CLOEXEC_ONLY) {
> + /* The close-on-exec bits decide what is dropped, walk the words. */
> + i = last / BITS_PER_LONG + 1;
> + while (i--) {
> + unsigned long dropped = dup_fd_dropped_word(fdt, i, range);
> +
> + if (fdt->open_fds[i] & ~dropped)
> + return (i + 1) * BITS_PER_LONG;
> + }
> + return NR_OPEN_DEFAULT;
> + }
>
> - if (fdt->open_fds[i] & ~dropped)
> - return (i + 1) * BITS_PER_LONG;
> + if (range->flags & FD_RANGE_EXCEPT) {
> + /* Only the range is carried over. */
> + if (last > range->to) {
> + last = find_last_bit(fdt->open_fds, range->to + 1);
> + if (last > range->to)
> + return NR_OPEN_DEFAULT;
> + }
> + if (last < range->from)
> + return NR_OPEN_DEFAULT;
> + } else if (last >= range->from && last <= range->to) {
> + /* The last open descriptor goes, the kept ones sit below the range. */
> + last = find_last_bit(fdt->open_fds, range->from);
> + if (last == range->from)
> + return NR_OPEN_DEFAULT;
> }
Maybe integrate this part of the patch, except for the
FD_RANGE_CLOEXEC_ONLY branch, into "file: let dup_fd() leave the
punched hole behind", which introduced the loop in sane_fdtable_size()
that is being removed here, and "file: let dup_fd() drop everything
outside of the range", which introduces FD_RANGE_EXCEPT?
^ permalink raw reply [flat|nested] 14+ messages in thread* Re: [PATCH 08/10] file: let dup_fd() drop only close-on-exec descriptors
2026-09-21 16:24 ` Jann Horn
@ 2026-09-25 14:57 ` Christian Brauner
0 siblings, 0 replies; 14+ messages in thread
From: Christian Brauner @ 2026-09-25 14:57 UTC (permalink / raw)
To: Jann Horn
Cc: linux-fsdevel, Oleg Nesterov, Alexander Viro, Jan Kara,
Neil Brown, Jeff Layton
On Mon, Sep 21, 2026 at 06:24:25PM +0200, Jann Horn wrote:
> On Mon, Sep 21, 2026 at 4:16 PM Christian Brauner <brauner@kernel.org> wrote:
> > Add FD_RANGE_CLOEXEC_ONLY. When set dup_fd() leaves everything behind
> > except for fds that are close-on-exec. Without FD_RANGE_EXCEPT the
> > clone loses the close-on-exec descriptors in the range. With it the
> > clone keeps the range and loses the close-on-exec descriptors everywhere
> > else. Descriptors without the flag are carried over either way.
> [...]
> > @@ -393,15 +396,37 @@ static unsigned int sane_fdtable_size(struct fdtable *fdt, struct fd_range *rang
> >
> > if (last == fdt->max_fds)
> > return NR_OPEN_DEFAULT;
> > - /* Only words up to the last open descriptor can hold a kept one. */
> > - i = last / BITS_PER_LONG + 1;
> > - while (i--) {
> > - unsigned long dropped = dup_fd_dropped_word(i, range);
> > + if (!range)
> > + return ALIGN(last + 1, BITS_PER_LONG);
> > +
> > + if (range->flags & FD_RANGE_CLOEXEC_ONLY) {
> > + /* The close-on-exec bits decide what is dropped, walk the words. */
> > + i = last / BITS_PER_LONG + 1;
> > + while (i--) {
> > + unsigned long dropped = dup_fd_dropped_word(fdt, i, range);
> > +
> > + if (fdt->open_fds[i] & ~dropped)
> > + return (i + 1) * BITS_PER_LONG;
> > + }
> > + return NR_OPEN_DEFAULT;
> > + }
> >
> > - if (fdt->open_fds[i] & ~dropped)
> > - return (i + 1) * BITS_PER_LONG;
> > + if (range->flags & FD_RANGE_EXCEPT) {
> > + /* Only the range is carried over. */
> > + if (last > range->to) {
> > + last = find_last_bit(fdt->open_fds, range->to + 1);
> > + if (last > range->to)
> > + return NR_OPEN_DEFAULT;
> > + }
> > + if (last < range->from)
> > + return NR_OPEN_DEFAULT;
> > + } else if (last >= range->from && last <= range->to) {
> > + /* The last open descriptor goes, the kept ones sit below the range. */
> > + last = find_last_bit(fdt->open_fds, range->from);
> > + if (last == range->from)
> > + return NR_OPEN_DEFAULT;
> > }
>
> Maybe integrate this part of the patch, except for the
> FD_RANGE_CLOEXEC_ONLY branch, into "file: let dup_fd() leave the
> punched hole behind", which introduced the loop in sane_fdtable_size()
> that is being removed here, and "file: let dup_fd() drop everything
> outside of the range", which introduces FD_RANGE_EXCEPT?
Done in-tree!
--
^ permalink raw reply [flat|nested] 14+ messages in thread
* [PATCH 09/10] file: add CLOSE_RANGE_CLOEXEC_ONLY
2026-09-21 14:15 [PATCH 00/10] files,close_range: add CLOSE_RANGE_{CLOEXEC_ONLY,EXCEPT} Christian Brauner
` (7 preceding siblings ...)
2026-09-21 14:15 ` [PATCH 08/10] file: let dup_fd() drop only close-on-exec descriptors Christian Brauner
@ 2026-09-21 14:15 ` Christian Brauner
2026-09-21 14:15 ` [PATCH 10/10] selftests/core: test CLOSE_RANGE_CLOEXEC_ONLY Christian Brauner
2026-09-21 16:34 ` [PATCH 00/10] files,close_range: add CLOSE_RANGE_{CLOEXEC_ONLY,EXCEPT} Jann Horn
10 siblings, 0 replies; 14+ messages in thread
From: Christian Brauner @ 2026-09-21 14:15 UTC (permalink / raw)
To: Jann Horn, linux-fsdevel, Oleg Nesterov
Cc: Alexander Viro, Jan Kara, Neil Brown, Jeff Layton,
Christian Brauner (Amutable)
CLOSE_RANGE_CLOEXEC marks a range close-on-exec. Nothing happens to
these fds unless an exec happens. Add CLOSE_RANGE_CLOEXEC_ONLY which
closes the close-on-exec file descriptors in the range.
When combined with CLOSE_RANGE_EXCEPT it names the close-on-exec file
descriptors that are supposed to survive. Every close-on-exec fd outside
of the range is closed. File descriptors without that flag are left
alone. Whatever the caller deliberately passes down, stdio, LISTEN_FDS,
an inherited pipe, remains where it is.
The handful of close-on-exec descriptors the caller still needs for the
exec such as the executable, an error pipe, sit in a specific range.
That is what a task between clone(CLONE_FILES) and execve() actually
wants:
clone(CLONE_FILES | CLONE_VM | CLONE_VFORK)
child: close_range(lo, hi, CLOSE_RANGE_UNSHARE |
CLOSE_RANGE_CLOEXEC_ONLY |
CLOSE_RANGE_EXCEPT)
child: rearrange descriptors in the now private table
child: execve()
Between the clone and the close_range() the child holds no reference of
its own on any file because copy_files() only bumps the fdtable
refcount. So a close() in the parent takes effect immediately. The
unshare also never takes a reference on the fds it leaves behind either.
The range is expressed in the parent's numbering. So a caller that
cannot name the file descriptors it keeps contiguously picks a range
wide enough to cover them, unshares, and tidies up with a second
close_range() on the table it now owns alone. That one is cheap. To keep
nothing, name a range that cannot hold an open descriptor, e.g.
close_range(~0U, ~0U, ...).
CLOSE_RANGE_CLOEXEC and CLOSE_RANGE_CLOEXEC_ONLY are mutually exclusive.
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
fs/file.c | 30 +++++++++++++++++++++++++++---
include/uapi/linux/close_range.h | 5 +++++
2 files changed, 32 insertions(+), 3 deletions(-)
diff --git a/fs/file.c b/fs/file.c
index 8952efd2cf3a..37f0ba740c39 100644
--- a/fs/file.c
+++ b/fs/file.c
@@ -843,18 +843,29 @@ static inline void __range_cloexec(struct files_struct *cur_fds,
spin_unlock(&cur_fds->file_lock);
}
+/* Next open descriptor in [fd, max_fd], or the next close-on-exec one. */
+static inline unsigned int next_open_fd(struct fdtable *fdt, unsigned int fd,
+ unsigned int max_fd,
+ struct fd_range *range)
+{
+ if (range->flags & FD_RANGE_CLOEXEC_ONLY)
+ return find_next_and_bit(fdt->open_fds, fdt->close_on_exec,
+ max_fd + 1, fd);
+ return find_next_bit(fdt->open_fds, max_fd + 1, fd);
+}
+
/* Next open descriptor in [fd, max_fd] that @range selects. */
static inline unsigned int next_fd_to_close(struct fdtable *fdt,
unsigned int fd, unsigned int max_fd,
struct fd_range *range)
{
- fd = find_next_bit(fdt->open_fds, max_fd + 1, fd);
+ fd = next_open_fd(fdt, fd, max_fd, range);
/* Hop over the window the range keeps. */
if ((range->flags & FD_RANGE_EXCEPT) &&
fd >= range->from && fd <= range->to) {
if (range->to >= max_fd)
return max_fd + 1;
- fd = find_next_bit(fdt->open_fds, max_fd + 1, range->to + 1);
+ fd = next_open_fd(fdt, range->to + 1, max_fd, range);
}
return fd;
}
@@ -911,6 +922,12 @@ static inline void __range_close(struct files_struct *files,
* With CLOSE_RANGE_EXCEPT the range names what to leave alone instead:
* every open file descriptor outside of [@fd, @max_fd] is closed, or
* marked close-on-exec with CLOSE_RANGE_CLOEXEC.
+ *
+ * With CLOSE_RANGE_CLOEXEC_ONLY only file descriptors that have
+ * close-on-exec set are closed. Together with CLOSE_RANGE_EXCEPT the
+ * range names the close-on-exec file descriptors to keep. To keep none
+ * of them, name a range that cannot hold an open file descriptor, e.g.
+ * close_range(~0U, ~0U, ...).
*/
SYSCALL_DEFINE3(close_range, unsigned int, fd, unsigned int, max_fd,
unsigned int, flags)
@@ -920,7 +937,12 @@ SYSCALL_DEFINE3(close_range, unsigned int, fd, unsigned int, max_fd,
struct fd_range range = {fd, max_fd};
if (flags & ~(CLOSE_RANGE_UNSHARE | CLOSE_RANGE_CLOEXEC |
- CLOSE_RANGE_EXCEPT))
+ CLOSE_RANGE_EXCEPT | CLOSE_RANGE_CLOEXEC_ONLY))
+ return -EINVAL;
+
+ /* One marks close-on-exec, the other closes what is marked. */
+ if (hweight32(flags & (CLOSE_RANGE_CLOEXEC |
+ CLOSE_RANGE_CLOEXEC_ONLY)) > 1)
return -EINVAL;
if (fd > max_fd)
@@ -928,6 +950,8 @@ SYSCALL_DEFINE3(close_range, unsigned int, fd, unsigned int, max_fd,
if (flags & CLOSE_RANGE_EXCEPT)
range.flags |= FD_RANGE_EXCEPT;
+ if (flags & CLOSE_RANGE_CLOEXEC_ONLY)
+ range.flags |= FD_RANGE_CLOEXEC_ONLY;
if ((flags & CLOSE_RANGE_UNSHARE) && atomic_read(&cur_fds->count) > 1) {
struct fd_range *drop = ⦥
diff --git a/include/uapi/linux/close_range.h b/include/uapi/linux/close_range.h
index 39eddb3ab613..7da9ed95258a 100644
--- a/include/uapi/linux/close_range.h
+++ b/include/uapi/linux/close_range.h
@@ -9,6 +9,7 @@
#undef CLOSE_RANGE_UNSHARE
#undef CLOSE_RANGE_CLOEXEC
#undef CLOSE_RANGE_EXCEPT
+#undef CLOSE_RANGE_CLOEXEC_ONLY
enum close_range_flags {
/* Unshare the file descriptor table before closing file descriptors. */
@@ -19,12 +20,16 @@ enum close_range_flags {
/* Act on every file descriptor outside of the given range instead. */
CLOSE_RANGE_EXCEPT = (1U << 3),
+
+ /* Only close file descriptors that have the FD_CLOEXEC bit set. */
+ CLOSE_RANGE_CLOEXEC_ONLY = (1U << 4),
};
/* Keep #ifdef working and let glibc skip its own definitions. */
#define CLOSE_RANGE_UNSHARE CLOSE_RANGE_UNSHARE
#define CLOSE_RANGE_CLOEXEC CLOSE_RANGE_CLOEXEC
#define CLOSE_RANGE_EXCEPT CLOSE_RANGE_EXCEPT
+#define CLOSE_RANGE_CLOEXEC_ONLY CLOSE_RANGE_CLOEXEC_ONLY
#endif /* _UAPI_LINUX_CLOSE_RANGE_H */
--
2.53.0
^ permalink raw reply related [flat|nested] 14+ messages in thread* [PATCH 10/10] selftests/core: test CLOSE_RANGE_CLOEXEC_ONLY
2026-09-21 14:15 [PATCH 00/10] files,close_range: add CLOSE_RANGE_{CLOEXEC_ONLY,EXCEPT} Christian Brauner
` (8 preceding siblings ...)
2026-09-21 14:15 ` [PATCH 09/10] file: add CLOSE_RANGE_CLOEXEC_ONLY Christian Brauner
@ 2026-09-21 14:15 ` Christian Brauner
2026-09-21 16:34 ` [PATCH 00/10] files,close_range: add CLOSE_RANGE_{CLOEXEC_ONLY,EXCEPT} Jann Horn
10 siblings, 0 replies; 14+ messages in thread
From: Christian Brauner @ 2026-09-21 14:15 UTC (permalink / raw)
To: Jann Horn, linux-fsdevel, Oleg Nesterov
Cc: Alexander Viro, Jan Kara, Neil Brown, Jeff Layton,
Christian Brauner (Amutable)
Cover closing what is marked:
- close-on-exec descriptors in the range go, the ones outside stay, and
a range above the table closes nothing
- with CLOSE_RANGE_EXCEPT it is the other way around, and the kept ones
keep their flag, for a window in the middle, at the top and at the
bottom
- descriptors without close-on-exec are never touched, in or out of the
range
- a range that cannot hold an open descriptor keeps nothing, one that
covers everything keeps everything, in place and in a clone
- the unshare form leaves the table it was cloned from alone, with and
without CLOSE_RANGE_EXCEPT
- the unshare form keeps the descriptors without the flag that sit in
the range it drops from, and hands the dropped slots out again
- a kept range above the last descriptor without close-on-exec still
comes back, so the clone is sized off the range too
Also check that asking for CLOSE_RANGE_CLOEXEC at the same time is
refused, whatever else is asked for, and that the bounds are still
checked.
The extra cases came out of the same walk with the close-on-exec mask
on top: a marked and an unmarked descriptor on each side of every window
position, and the size of the clone when only unmarked ones are left in
the range it drops from.
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
---
tools/testing/selftests/core/close_range_test.c | 481 ++++++++++++++++++++++++
1 file changed, 481 insertions(+)
diff --git a/tools/testing/selftests/core/close_range_test.c b/tools/testing/selftests/core/close_range_test.c
index caeb2f1ea800..20ecb65e529b 100644
--- a/tools/testing/selftests/core/close_range_test.c
+++ b/tools/testing/selftests/core/close_range_test.c
@@ -1063,6 +1063,487 @@ TEST(close_range_except_unshare)
EXPECT_NE(-1, fcntl(open_fds[i], F_GETFD));
}
+TEST(close_range_cloexec_only)
+{
+ int i, ret;
+ int open_fds[101];
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++) {
+ int fd;
+
+ /* Odd slots are close-on-exec, even ones are not. */
+ fd = open("/dev/null", O_RDONLY | (i % 2 ? O_CLOEXEC : 0));
+ ASSERT_GE(fd, 0) {
+ if (errno == ENOENT)
+ SKIP(return, "Skipping test since /dev/null does not exist");
+ }
+
+ open_fds[i] = fd;
+ }
+
+ ret = sys_close_range(open_fds[10], open_fds[20],
+ CLOSE_RANGE_CLOEXEC_ONLY);
+ if (ret < 0) {
+ if (errno == ENOSYS)
+ SKIP(return, "close_range() syscall not supported");
+ if (errno == EINVAL)
+ SKIP(return, "close_range() doesn't support CLOSE_RANGE_CLOEXEC_ONLY");
+ }
+ ASSERT_EQ(0, ret);
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++) {
+ bool closed = i % 2 && i >= 10 && i <= 20;
+
+ EXPECT_EQ(!closed, fcntl(open_fds[i], F_GETFD) != -1);
+ }
+
+ /* A range above the table closes nothing. */
+ ASSERT_EQ(0, sys_close_range(UINT_MAX, UINT_MAX,
+ CLOSE_RANGE_CLOEXEC_ONLY));
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++) {
+ bool closed = i % 2 && i >= 10 && i <= 20;
+
+ EXPECT_EQ(!closed, fcntl(open_fds[i], F_GETFD) != -1);
+ }
+
+ /* Do what an exec would do to the rest. */
+ ASSERT_EQ(0, sys_close_range(0, UINT_MAX, CLOSE_RANGE_CLOEXEC_ONLY));
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++)
+ EXPECT_EQ(!(i % 2), fcntl(open_fds[i], F_GETFD) != -1);
+}
+
+TEST(close_range_cloexec_only_except)
+{
+ int i, ret;
+ int open_fds[101];
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++) {
+ int fd;
+
+ fd = open("/dev/null", O_RDONLY | (i % 2 ? O_CLOEXEC : 0));
+ ASSERT_GE(fd, 0) {
+ if (errno == ENOENT)
+ SKIP(return, "Skipping test since /dev/null does not exist");
+ }
+
+ open_fds[i] = fd;
+ }
+
+ ret = sys_close_range(open_fds[10], open_fds[20],
+ CLOSE_RANGE_CLOEXEC_ONLY | CLOSE_RANGE_EXCEPT);
+ if (ret < 0) {
+ if (errno == ENOSYS)
+ SKIP(return, "close_range() syscall not supported");
+ if (errno == EINVAL)
+ SKIP(return, "close_range() doesn't support CLOSE_RANGE_CLOEXEC_ONLY");
+ }
+ ASSERT_EQ(0, ret);
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++) {
+ bool kept = !(i % 2) || (i >= 10 && i <= 20);
+ int flags = i % 2 ? FD_CLOEXEC : 0;
+
+ /* The kept ones keep their flag, so exec still drops them. */
+ EXPECT_EQ(kept ? flags : -1, fcntl(open_fds[i], F_GETFD));
+ }
+
+ /* A range that cannot hold an open descriptor keeps nothing. */
+ ASSERT_EQ(0, sys_close_range(UINT_MAX, UINT_MAX,
+ CLOSE_RANGE_CLOEXEC_ONLY |
+ CLOSE_RANGE_EXCEPT));
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++)
+ EXPECT_EQ(!(i % 2), fcntl(open_fds[i], F_GETFD) != -1);
+}
+
+TEST(close_range_cloexec_only_except_bounds)
+{
+ int i, c, ret, status;
+ pid_t pid;
+ int open_fds[101];
+ struct __clone_args args = {
+ .exit_signal = SIGCHLD,
+ };
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++) {
+ int fd;
+
+ fd = open("/dev/null", O_RDONLY | (i % 2 ? O_CLOEXEC : 0));
+ ASSERT_GE(fd, 0) {
+ if (errno == ENOENT)
+ SKIP(return, "Skipping test since /dev/null does not exist");
+ }
+
+ open_fds[i] = fd;
+ }
+
+ /* A range covering everything keeps everything. */
+ ret = sys_close_range(0, UINT_MAX,
+ CLOSE_RANGE_CLOEXEC_ONLY | CLOSE_RANGE_EXCEPT);
+ if (ret < 0) {
+ if (errno == ENOSYS)
+ SKIP(return, "close_range() syscall not supported");
+ if (errno == EINVAL)
+ SKIP(return, "close_range() doesn't support CLOSE_RANGE_CLOEXEC_ONLY");
+ }
+ ASSERT_EQ(0, ret);
+
+ struct {
+ unsigned int fd, max_fd;
+ } cases[] = {
+ /* A window at the top keeps the marked ones in it. */
+ { open_fds[80], UINT_MAX },
+ /* One at the bottom keeps the marked ones in it. */
+ { 0, open_fds[20] },
+ /* One that cannot hold a descriptor keeps none of them. */
+ { UINT_MAX, UINT_MAX },
+ };
+
+ for (c = 0; c < ARRAY_SIZE(cases); c++) {
+ pid = sys_clone3(&args, sizeof(args));
+ ASSERT_GE(pid, 0);
+
+ if (pid == 0) {
+ ret = sys_close_range(cases[c].fd, cases[c].max_fd,
+ CLOSE_RANGE_CLOEXEC_ONLY |
+ CLOSE_RANGE_EXCEPT);
+ if (ret)
+ exit(EXIT_FAILURE);
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++) {
+ unsigned int fd = open_fds[i];
+ bool kept = !(i % 2) || (fd >= cases[c].fd &&
+ fd <= cases[c].max_fd);
+ int flags = i % 2 ? FD_CLOEXEC : 0;
+
+ if (fcntl(fd, F_GETFD) != (kept ? flags : -1))
+ exit(EXIT_FAILURE);
+ }
+
+ /* stdio is neither marked nor gone. */
+ if (fcntl(STDERR_FILENO, F_GETFD) & FD_CLOEXEC)
+ exit(EXIT_FAILURE);
+
+ exit(EXIT_SUCCESS);
+ }
+
+ EXPECT_EQ(waitpid(pid, &status, 0), pid);
+ EXPECT_EQ(true, WIFEXITED(status));
+ EXPECT_EQ(0, WEXITSTATUS(status));
+ }
+
+ /* Each fork had a table of its own. */
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++)
+ EXPECT_NE(-1, fcntl(open_fds[i], F_GETFD));
+}
+
+TEST(close_range_cloexec_only_unshare)
+{
+ int i, ret, status;
+ pid_t pid;
+ int open_fds[101];
+ struct __clone_args args = {
+ .flags = CLONE_FILES,
+ .exit_signal = SIGCHLD,
+ };
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++) {
+ int fd;
+
+ fd = open("/dev/null", O_RDONLY | (i % 2 ? O_CLOEXEC : 0));
+ ASSERT_GE(fd, 0) {
+ if (errno == ENOENT)
+ SKIP(return, "Skipping test since /dev/null does not exist");
+ }
+
+ open_fds[i] = fd;
+ }
+
+ /* A range covering everything keeps everything. */
+ ret = sys_close_range(0, UINT_MAX,
+ CLOSE_RANGE_CLOEXEC_ONLY | CLOSE_RANGE_EXCEPT);
+ if (ret < 0) {
+ if (errno == ENOSYS)
+ SKIP(return, "close_range() syscall not supported");
+ if (errno == EINVAL)
+ SKIP(return, "close_range() doesn't support CLOSE_RANGE_CLOEXEC_ONLY");
+ }
+ ASSERT_EQ(0, ret);
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++)
+ ASSERT_NE(-1, fcntl(open_fds[i], F_GETFD));
+
+ pid = sys_clone3(&args, sizeof(args));
+ ASSERT_GE(pid, 0);
+
+ if (pid == 0) {
+ ret = sys_close_range(open_fds[10], open_fds[20],
+ CLOSE_RANGE_UNSHARE |
+ CLOSE_RANGE_CLOEXEC_ONLY);
+ if (ret)
+ exit(EXIT_FAILURE);
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++) {
+ bool closed = i % 2 && i >= 10 && i <= 20;
+
+ if (closed == (fcntl(open_fds[i], F_GETFD) != -1))
+ exit(EXIT_FAILURE);
+ }
+
+ exit(EXIT_SUCCESS);
+ }
+
+ EXPECT_EQ(waitpid(pid, &status, 0), pid);
+ EXPECT_EQ(true, WIFEXITED(status));
+ EXPECT_EQ(0, WEXITSTATUS(status));
+
+ /* A range at the top keeps the descriptors without the flag in it. */
+ pid = sys_clone3(&args, sizeof(args));
+ ASSERT_GE(pid, 0);
+
+ if (pid == 0) {
+ ret = sys_close_range(open_fds[50], UINT_MAX,
+ CLOSE_RANGE_UNSHARE |
+ CLOSE_RANGE_CLOEXEC_ONLY);
+ if (ret)
+ exit(EXIT_FAILURE);
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++) {
+ bool closed = i % 2 && i >= 50;
+
+ if (closed == (fcntl(open_fds[i], F_GETFD) != -1))
+ exit(EXIT_FAILURE);
+ }
+
+ /* The first slot left behind is the next one handed out. */
+ if (dup(0) != open_fds[51])
+ exit(EXIT_FAILURE);
+
+ exit(EXIT_SUCCESS);
+ }
+
+ EXPECT_EQ(waitpid(pid, &status, 0), pid);
+ EXPECT_EQ(true, WIFEXITED(status));
+ EXPECT_EQ(0, WEXITSTATUS(status));
+
+ /* The shared table the child unshared from is untouched. */
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++)
+ EXPECT_NE(-1, fcntl(open_fds[i], F_GETFD));
+}
+
+TEST(close_range_cloexec_only_except_unshare)
+{
+ int i, ret, status;
+ pid_t pid;
+ int open_fds[101];
+ struct __clone_args args = {
+ .flags = CLONE_FILES,
+ .exit_signal = SIGCHLD,
+ };
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++) {
+ int fd;
+
+ fd = open("/dev/null", O_RDONLY | (i % 2 ? O_CLOEXEC : 0));
+ ASSERT_GE(fd, 0) {
+ if (errno == ENOENT)
+ SKIP(return, "Skipping test since /dev/null does not exist");
+ }
+
+ open_fds[i] = fd;
+ }
+
+ /* A range covering everything keeps everything. */
+ ret = sys_close_range(0, UINT_MAX,
+ CLOSE_RANGE_CLOEXEC_ONLY | CLOSE_RANGE_EXCEPT);
+ if (ret < 0) {
+ if (errno == ENOSYS)
+ SKIP(return, "close_range() syscall not supported");
+ if (errno == EINVAL)
+ SKIP(return, "close_range() doesn't support CLOSE_RANGE_CLOEXEC_ONLY");
+ }
+ ASSERT_EQ(0, ret);
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++)
+ ASSERT_NE(-1, fcntl(open_fds[i], F_GETFD));
+
+ pid = sys_clone3(&args, sizeof(args));
+ ASSERT_GE(pid, 0);
+
+ if (pid == 0) {
+ ret = sys_close_range(open_fds[10], open_fds[20],
+ CLOSE_RANGE_UNSHARE |
+ CLOSE_RANGE_CLOEXEC_ONLY |
+ CLOSE_RANGE_EXCEPT);
+ if (ret)
+ exit(EXIT_FAILURE);
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++) {
+ bool kept = !(i % 2) || (i >= 10 && i <= 20);
+ int flags = i % 2 ? FD_CLOEXEC : 0;
+
+ if (fcntl(open_fds[i], F_GETFD) != (kept ? flags : -1))
+ exit(EXIT_FAILURE);
+ }
+
+ exit(EXIT_SUCCESS);
+ }
+
+ EXPECT_EQ(waitpid(pid, &status, 0), pid);
+ EXPECT_EQ(true, WIFEXITED(status));
+ EXPECT_EQ(0, WEXITSTATUS(status));
+
+ /* A window that cannot hold a descriptor keeps none of the marked. */
+ pid = sys_clone3(&args, sizeof(args));
+ ASSERT_GE(pid, 0);
+
+ if (pid == 0) {
+ ret = sys_close_range(UINT_MAX, UINT_MAX,
+ CLOSE_RANGE_UNSHARE |
+ CLOSE_RANGE_CLOEXEC_ONLY |
+ CLOSE_RANGE_EXCEPT);
+ if (ret)
+ exit(EXIT_FAILURE);
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++)
+ if ((i % 2) == (fcntl(open_fds[i], F_GETFD) != -1))
+ exit(EXIT_FAILURE);
+
+ if (fcntl(STDERR_FILENO, F_GETFD) == -1)
+ exit(EXIT_FAILURE);
+
+ exit(EXIT_SUCCESS);
+ }
+
+ EXPECT_EQ(waitpid(pid, &status, 0), pid);
+ EXPECT_EQ(true, WIFEXITED(status));
+ EXPECT_EQ(0, WEXITSTATUS(status));
+
+ /* One that covers everything keeps everything, in a clone too. */
+ pid = sys_clone3(&args, sizeof(args));
+ ASSERT_GE(pid, 0);
+
+ if (pid == 0) {
+ ret = sys_close_range(0, UINT_MAX,
+ CLOSE_RANGE_UNSHARE |
+ CLOSE_RANGE_CLOEXEC_ONLY |
+ CLOSE_RANGE_EXCEPT);
+ if (ret)
+ exit(EXIT_FAILURE);
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++)
+ if (fcntl(open_fds[i], F_GETFD) != (i % 2 ? FD_CLOEXEC : 0))
+ exit(EXIT_FAILURE);
+
+ exit(EXIT_SUCCESS);
+ }
+
+ EXPECT_EQ(waitpid(pid, &status, 0), pid);
+ EXPECT_EQ(true, WIFEXITED(status));
+ EXPECT_EQ(0, WEXITSTATUS(status));
+
+ /* The shared table the child unshared from is untouched. */
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++)
+ EXPECT_NE(-1, fcntl(open_fds[i], F_GETFD));
+}
+
+TEST(close_range_cloexec_only_except_unshare_sizing)
+{
+ int i, ret, status;
+ pid_t pid;
+ int open_fds[200];
+ struct __clone_args args = {
+ .flags = CLONE_FILES,
+ .exit_signal = SIGCHLD,
+ };
+
+ /* All close-on-exec, so the kept range alone sizes the clone. */
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++) {
+ int fd;
+
+ fd = open("/dev/null", O_RDONLY | O_CLOEXEC);
+ ASSERT_GE(fd, 0) {
+ if (errno == ENOENT)
+ SKIP(return, "Skipping test since /dev/null does not exist");
+ }
+
+ open_fds[i] = fd;
+ }
+
+ ret = sys_close_range(0, UINT_MAX,
+ CLOSE_RANGE_CLOEXEC_ONLY | CLOSE_RANGE_EXCEPT);
+ if (ret < 0) {
+ if (errno == ENOSYS)
+ SKIP(return, "close_range() syscall not supported");
+ if (errno == EINVAL)
+ SKIP(return, "close_range() doesn't support CLOSE_RANGE_CLOEXEC_ONLY");
+ }
+ ASSERT_EQ(0, ret);
+
+ pid = sys_clone3(&args, sizeof(args));
+ ASSERT_GE(pid, 0);
+
+ if (pid == 0) {
+ ret = sys_close_range(open_fds[150], open_fds[160],
+ CLOSE_RANGE_UNSHARE |
+ CLOSE_RANGE_CLOEXEC_ONLY |
+ CLOSE_RANGE_EXCEPT);
+ if (ret)
+ exit(EXIT_FAILURE);
+
+ for (i = 0; i < ARRAY_SIZE(open_fds); i++) {
+ bool kept = i >= 150 && i <= 160;
+
+ if (kept != (fcntl(open_fds[i], F_GETFD) != -1))
+ exit(EXIT_FAILURE);
+ }
+
+ /* Nothing set close-on-exec on stdio. */
+ if (fcntl(STDERR_FILENO, F_GETFD) == -1)
+ exit(EXIT_FAILURE);
+
+ exit(EXIT_SUCCESS);
+ }
+
+ EXPECT_EQ(waitpid(pid, &status, 0), pid);
+ EXPECT_EQ(true, WIFEXITED(status));
+ EXPECT_EQ(0, WEXITSTATUS(status));
+}
+
+TEST(close_range_cloexec_only_einval)
+{
+ int ret;
+
+ /* A range covering everything keeps everything, so this only probes. */
+ ret = sys_close_range(0, UINT_MAX,
+ CLOSE_RANGE_CLOEXEC_ONLY | CLOSE_RANGE_EXCEPT);
+ if (ret < 0) {
+ if (errno == ENOSYS)
+ SKIP(return, "close_range() syscall not supported");
+ if (errno == EINVAL)
+ SKIP(return, "close_range() doesn't support CLOSE_RANGE_CLOEXEC_ONLY");
+ }
+ ASSERT_EQ(0, ret);
+
+ EXPECT_EQ(-1, sys_close_range(3, UINT_MAX, CLOSE_RANGE_CLOEXEC |
+ CLOSE_RANGE_CLOEXEC_ONLY));
+ EXPECT_EQ(EINVAL, errno);
+
+ /* The other flags do not make the pair acceptable. */
+ EXPECT_EQ(-1, sys_close_range(3, UINT_MAX, CLOSE_RANGE_UNSHARE |
+ CLOSE_RANGE_CLOEXEC |
+ CLOSE_RANGE_CLOEXEC_ONLY |
+ CLOSE_RANGE_EXCEPT));
+ EXPECT_EQ(EINVAL, errno);
+
+ /* The bounds are checked with the new flag too. */
+ EXPECT_EQ(-1, sys_close_range(4, 3, CLOSE_RANGE_CLOEXEC_ONLY |
+ CLOSE_RANGE_EXCEPT));
+ EXPECT_EQ(EINVAL, errno);
+}
+
TEST(close_range_bitmap_corruption)
{
pid_t pid;
--
2.53.0
^ permalink raw reply related [flat|nested] 14+ messages in thread* Re: [PATCH 00/10] files,close_range: add CLOSE_RANGE_{CLOEXEC_ONLY,EXCEPT}
2026-09-21 14:15 [PATCH 00/10] files,close_range: add CLOSE_RANGE_{CLOEXEC_ONLY,EXCEPT} Christian Brauner
` (9 preceding siblings ...)
2026-09-21 14:15 ` [PATCH 10/10] selftests/core: test CLOSE_RANGE_CLOEXEC_ONLY Christian Brauner
@ 2026-09-21 16:34 ` Jann Horn
10 siblings, 0 replies; 14+ messages in thread
From: Jann Horn @ 2026-09-21 16:34 UTC (permalink / raw)
To: Christian Brauner
Cc: linux-fsdevel, Oleg Nesterov, Alexander Viro, Jan Kara,
Neil Brown, Jeff Layton
On Mon, Sep 21, 2026 at 4:15 PM Christian Brauner <brauner@kernel.org> wrote:
> close_range(CLOSE_RANGE_CLOEXEC) marks a range of file descriptors
> close-on-exec. It only marks them. The file descriptors are only ever
> closed during an exec. This series adds two new flags to extend the
> abilities of close_range():
>
> (1) CLOSE_RANGE_EXCEPT
>
> While regular close_range() operates on all file descriptors in the
> range [fd, max_fd] and not on any others, CLOSE_RANGE_EXCEPT makes
> close_range() operate on all file descriptors outside of the range.
>
> When combined with CLOSE_RANGE_UNSHARE dup_fd() never takes a
> reference on what is dropped.
>
> A task between clone(CLONE_FILES | CLONE_VM | CLONE_VFORK) and
> execve() that has the descriptors for the child in one window can
> shed the rest of the shared table in one call:
>
> close_range(lo, hi, CLOSE_RANGE_UNSHARE | CLOSE_RANGE_EXCEPT)
>
> Nothing outside of the window is referenced by the child at any point.
>
> (2) CLOSE_RANGE_CLOEXEC_ONLY closes the close-on-exec file descriptors
> in the range
>
> Whatever the caller deliberately passes down, stdio, LISTEN_FDS, an
> inherited pipe, remains where it is.
>
> The handful of close-on-exec descriptors the caller still needs for the
> exec such as the executable, an error pipe, can sit in a specific
> range. A task between clone(CLONE_FILES) and execve() may do:
>
> clone(CLONE_FILES | CLONE_VM | CLONE_VFORK)
> child: close_range(lo, hi, CLOSE_RANGE_UNSHARE |
> CLOSE_RANGE_CLOEXEC_ONLY |
> CLOSE_RANGE_EXCEPT)
> child: rearrange descriptors in the now private table
> child: execve()
>
> Between the clone and the close_range() the child holds no reference of
> its own on any file because copy_files() only bumps the fdtable
> refcount. So a close() in the parent takes effect immediately. The
> unshare also never takes a reference on the fds it leaves behind either.
Nice, I think with this we'll finally have a way for a thread in a
multi-threaded process to fork+exec without keeping O_CLOEXEC FDs open
longer than concurrently running threads expect.
We probably can't make the UAPI better without introducing some
entirely new syscall that operates on file descriptor bitmasks /
lists.
> Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
For patches 1,3,4,5,6,8,9,10 (everything except the selftests):
Reviewed-by: Jann Horn <jannh@google.com>
^ permalink raw reply [flat|nested] 14+ messages in thread