* [PATCH V3] cgroup/cpuset: Prevent UAF in proc_cpuset_show()
@ 2024-06-26 9:41 Chen Ridong
2024-06-26 14:56 ` Waiman Long
` (2 more replies)
0 siblings, 3 replies; 6+ messages in thread
From: Chen Ridong @ 2024-06-26 9:41 UTC (permalink / raw)
To: tj, lizefan.x, hannes, longman, adityakali, sergeh, mkoutny
Cc: cgroups, linux-kernel
An UAF can happen when /proc/cpuset is read as reported in [1].
This can be reproduced by the following methods:
1.add an mdelay(1000) before acquiring the cgroup_lock In the
cgroup_path_ns function.
2.$cat /proc/<pid>/cpuset repeatly.
3.$mount -t cgroup -o cpuset cpuset /sys/fs/cgroup/cpuset/
$umount /sys/fs/cgroup/cpuset/ repeatly.
The race that cause this bug can be shown as below:
(umount) | (cat /proc/<pid>/cpuset)
css_release | proc_cpuset_show
css_release_work_fn | css = task_get_css(tsk, cpuset_cgrp_id);
css_free_rwork_fn | cgroup_path_ns(css->cgroup, ...);
cgroup_destroy_root | mutex_lock(&cgroup_mutex);
rebind_subsystems |
cgroup_free_root |
| // cgrp was freed, UAF
| cgroup_path_ns_locked(cgrp,..);
When the cpuset is initialized, the root node top_cpuset.css.cgrp
will point to &cgrp_dfl_root.cgrp. In cgroup v1, the mount operation will
allocate cgroup_root, and top_cpuset.css.cgrp will point to the allocated
&cgroup_root.cgrp. When the umount operation is executed,
top_cpuset.css.cgrp will be rebound to &cgrp_dfl_root.cgrp.
The problem is that when rebinding to cgrp_dfl_root, there are cases
where the cgroup_root allocated by setting up the root for cgroup v1
is cached. This could lead to a Use-After-Free (UAF) if it is
subsequently freed. The descendant cgroups of cgroup v1 can only be
freed after the css is released. However, the css of the root will never
be released, yet the cgroup_root should be freed when it is unmounted.
This means that obtaining a reference to the css of the root does
not guarantee that css.cgrp->root will not be freed.
Fix this problem by using rcu_read_lock in proc_cpuset_show().
As cgroup root_list is already RCU-safe, css->cgroup is safe.
This is similar to commit 9067d90006df ("cgroup: Eliminate the
need for cgroup_mutex in proc_cgroup_show()")
[1] https://syzkaller.appspot.com/bug?extid=9b1ff7be974a403aa4cd
Fixes: a79a908fd2b0 ("cgroup: introduce cgroup namespaces")
Signed-off-by: Chen Ridong <chenridong@huawei.com>
---
kernel/cgroup/cpuset.c | 12 ++++++++++--
1 file changed, 10 insertions(+), 2 deletions(-)
diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c
index c12b9fdb22a4..7f4536c9ccce 100644
--- a/kernel/cgroup/cpuset.c
+++ b/kernel/cgroup/cpuset.c
@@ -21,6 +21,7 @@
* License. See the file COPYING in the main directory of the Linux
* distribution for more details.
*/
+#include "cgroup-internal.h"
#include <linux/cpu.h>
#include <linux/cpumask.h>
@@ -5052,8 +5053,15 @@ int proc_cpuset_show(struct seq_file *m, struct pid_namespace *ns,
goto out;
css = task_get_css(tsk, cpuset_cgrp_id);
- retval = cgroup_path_ns(css->cgroup, buf, PATH_MAX,
- current->nsproxy->cgroup_ns);
+ rcu_read_lock();
+ spin_lock_irq(&css_set_lock);
+ /* In case the root has already been unmounted */
+ if (css->cgroup)
+ retval = cgroup_path_ns_locked(css->cgroup, buf, PATH_MAX,
+ current->nsproxy->cgroup_ns);
+
+ spin_unlock_irq(&css_set_lock);
+ rcu_read_unlock();
css_put(css);
if (retval == -E2BIG)
retval = -ENAMETOOLONG;
--
2.34.1
^ permalink raw reply related [flat|nested] 6+ messages in thread* Re: [PATCH V3] cgroup/cpuset: Prevent UAF in proc_cpuset_show()
2024-06-26 9:41 [PATCH V3] cgroup/cpuset: Prevent UAF in proc_cpuset_show() Chen Ridong
@ 2024-06-26 14:56 ` Waiman Long
2024-06-26 21:20 ` Tejun Heo
2024-06-27 9:46 ` Michal Koutný
2 siblings, 0 replies; 6+ messages in thread
From: Waiman Long @ 2024-06-26 14:56 UTC (permalink / raw)
To: Chen Ridong, tj, lizefan.x, hannes, adityakali, sergeh, mkoutny
Cc: cgroups, linux-kernel
On 6/26/24 05:41, Chen Ridong wrote:
> An UAF can happen when /proc/cpuset is read as reported in [1].
>
> This can be reproduced by the following methods:
> 1.add an mdelay(1000) before acquiring the cgroup_lock In the
> cgroup_path_ns function.
> 2.$cat /proc/<pid>/cpuset repeatly.
> 3.$mount -t cgroup -o cpuset cpuset /sys/fs/cgroup/cpuset/
> $umount /sys/fs/cgroup/cpuset/ repeatly.
>
> The race that cause this bug can be shown as below:
>
> (umount) | (cat /proc/<pid>/cpuset)
> css_release | proc_cpuset_show
> css_release_work_fn | css = task_get_css(tsk, cpuset_cgrp_id);
> css_free_rwork_fn | cgroup_path_ns(css->cgroup, ...);
> cgroup_destroy_root | mutex_lock(&cgroup_mutex);
> rebind_subsystems |
> cgroup_free_root |
> | // cgrp was freed, UAF
> | cgroup_path_ns_locked(cgrp,..);
>
> When the cpuset is initialized, the root node top_cpuset.css.cgrp
> will point to &cgrp_dfl_root.cgrp. In cgroup v1, the mount operation will
> allocate cgroup_root, and top_cpuset.css.cgrp will point to the allocated
> &cgroup_root.cgrp. When the umount operation is executed,
> top_cpuset.css.cgrp will be rebound to &cgrp_dfl_root.cgrp.
>
> The problem is that when rebinding to cgrp_dfl_root, there are cases
> where the cgroup_root allocated by setting up the root for cgroup v1
> is cached. This could lead to a Use-After-Free (UAF) if it is
> subsequently freed. The descendant cgroups of cgroup v1 can only be
> freed after the css is released. However, the css of the root will never
> be released, yet the cgroup_root should be freed when it is unmounted.
> This means that obtaining a reference to the css of the root does
> not guarantee that css.cgrp->root will not be freed.
>
> Fix this problem by using rcu_read_lock in proc_cpuset_show().
> As cgroup root_list is already RCU-safe, css->cgroup is safe.
> This is similar to commit 9067d90006df ("cgroup: Eliminate the
> need for cgroup_mutex in proc_cgroup_show()")
>
> [1] https://syzkaller.appspot.com/bug?extid=9b1ff7be974a403aa4cd
>
> Fixes: a79a908fd2b0 ("cgroup: introduce cgroup namespaces")
> Signed-off-by: Chen Ridong <chenridong@huawei.com>
> ---
> kernel/cgroup/cpuset.c | 12 ++++++++++--
> 1 file changed, 10 insertions(+), 2 deletions(-)
>
> diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c
> index c12b9fdb22a4..7f4536c9ccce 100644
> --- a/kernel/cgroup/cpuset.c
> +++ b/kernel/cgroup/cpuset.c
> @@ -21,6 +21,7 @@
> * License. See the file COPYING in the main directory of the Linux
> * distribution for more details.
> */
> +#include "cgroup-internal.h"
>
> #include <linux/cpu.h>
> #include <linux/cpumask.h>
> @@ -5052,8 +5053,15 @@ int proc_cpuset_show(struct seq_file *m, struct pid_namespace *ns,
> goto out;
>
> css = task_get_css(tsk, cpuset_cgrp_id);
> - retval = cgroup_path_ns(css->cgroup, buf, PATH_MAX,
> - current->nsproxy->cgroup_ns);
> + rcu_read_lock();
> + spin_lock_irq(&css_set_lock);
> + /* In case the root has already been unmounted */
> + if (css->cgroup)
> + retval = cgroup_path_ns_locked(css->cgroup, buf, PATH_MAX,
> + current->nsproxy->cgroup_ns);
> +
> + spin_unlock_irq(&css_set_lock);
> + rcu_read_unlock();
> css_put(css);
> if (retval == -E2BIG)
> retval = -ENAMETOOLONG;
Reviewed-by: Waiman Long <longman@redhat.com>
^ permalink raw reply [flat|nested] 6+ messages in thread* Re: [PATCH V3] cgroup/cpuset: Prevent UAF in proc_cpuset_show()
2024-06-26 9:41 [PATCH V3] cgroup/cpuset: Prevent UAF in proc_cpuset_show() Chen Ridong
2024-06-26 14:56 ` Waiman Long
@ 2024-06-26 21:20 ` Tejun Heo
2024-06-27 9:46 ` Michal Koutný
2 siblings, 0 replies; 6+ messages in thread
From: Tejun Heo @ 2024-06-26 21:20 UTC (permalink / raw)
To: Chen Ridong
Cc: lizefan.x, hannes, longman, adityakali, sergeh, mkoutny, cgroups,
linux-kernel
On Wed, Jun 26, 2024 at 09:41:01AM +0000, Chen Ridong wrote:
> An UAF can happen when /proc/cpuset is read as reported in [1].
>
> This can be reproduced by the following methods:
> 1.add an mdelay(1000) before acquiring the cgroup_lock In the
> cgroup_path_ns function.
> 2.$cat /proc/<pid>/cpuset repeatly.
> 3.$mount -t cgroup -o cpuset cpuset /sys/fs/cgroup/cpuset/
> $umount /sys/fs/cgroup/cpuset/ repeatly.
>
> The race that cause this bug can be shown as below:
>
> (umount) | (cat /proc/<pid>/cpuset)
> css_release | proc_cpuset_show
> css_release_work_fn | css = task_get_css(tsk, cpuset_cgrp_id);
> css_free_rwork_fn | cgroup_path_ns(css->cgroup, ...);
> cgroup_destroy_root | mutex_lock(&cgroup_mutex);
> rebind_subsystems |
> cgroup_free_root |
> | // cgrp was freed, UAF
> | cgroup_path_ns_locked(cgrp,..);
>
> When the cpuset is initialized, the root node top_cpuset.css.cgrp
> will point to &cgrp_dfl_root.cgrp. In cgroup v1, the mount operation will
> allocate cgroup_root, and top_cpuset.css.cgrp will point to the allocated
> &cgroup_root.cgrp. When the umount operation is executed,
> top_cpuset.css.cgrp will be rebound to &cgrp_dfl_root.cgrp.
>
> The problem is that when rebinding to cgrp_dfl_root, there are cases
> where the cgroup_root allocated by setting up the root for cgroup v1
> is cached. This could lead to a Use-After-Free (UAF) if it is
> subsequently freed. The descendant cgroups of cgroup v1 can only be
> freed after the css is released. However, the css of the root will never
> be released, yet the cgroup_root should be freed when it is unmounted.
> This means that obtaining a reference to the css of the root does
> not guarantee that css.cgrp->root will not be freed.
>
> Fix this problem by using rcu_read_lock in proc_cpuset_show().
> As cgroup root_list is already RCU-safe, css->cgroup is safe.
> This is similar to commit 9067d90006df ("cgroup: Eliminate the
> need for cgroup_mutex in proc_cgroup_show()")
>
> [1] https://syzkaller.appspot.com/bug?extid=9b1ff7be974a403aa4cd
>
> Fixes: a79a908fd2b0 ("cgroup: introduce cgroup namespaces")
> Signed-off-by: Chen Ridong <chenridong@huawei.com>
Applied to cgroup/for-6.10-fixes w/ stable cc added.
Thanks.
--
tejun
^ permalink raw reply [flat|nested] 6+ messages in thread* Re: [PATCH V3] cgroup/cpuset: Prevent UAF in proc_cpuset_show()
2024-06-26 9:41 [PATCH V3] cgroup/cpuset: Prevent UAF in proc_cpuset_show() Chen Ridong
2024-06-26 14:56 ` Waiman Long
2024-06-26 21:20 ` Tejun Heo
@ 2024-06-27 9:46 ` Michal Koutný
2024-06-27 18:29 ` Tejun Heo
2 siblings, 1 reply; 6+ messages in thread
From: Michal Koutný @ 2024-06-27 9:46 UTC (permalink / raw)
To: Chen Ridong
Cc: tj, lizefan.x, hannes, longman, adityakali, sergeh, cgroups,
linux-kernel
[-- Attachment #1: Type: text/plain, Size: 2153 bytes --]
On Wed, Jun 26, 2024 at 09:41:01AM GMT, Chen Ridong <chenridong@huawei.com> wrote:
> An UAF can happen when /proc/cpuset is read as reported in [1].
>
> This can be reproduced by the following methods:
> 1.add an mdelay(1000) before acquiring the cgroup_lock In the
> cgroup_path_ns function.
> 2.$cat /proc/<pid>/cpuset repeatly.
> 3.$mount -t cgroup -o cpuset cpuset /sys/fs/cgroup/cpuset/
> $umount /sys/fs/cgroup/cpuset/ repeatly.
>
> The race that cause this bug can be shown as below:
>
> (umount) | (cat /proc/<pid>/cpuset)
> css_release | proc_cpuset_show
> css_release_work_fn | css = task_get_css(tsk, cpuset_cgrp_id);
> css_free_rwork_fn | cgroup_path_ns(css->cgroup, ...);
> cgroup_destroy_root | mutex_lock(&cgroup_mutex);
> rebind_subsystems |
> cgroup_free_root |
> | // cgrp was freed, UAF
> | cgroup_path_ns_locked(cgrp,..);
Thanks for this breakdown.
> ...
> Fix this problem by using rcu_read_lock in proc_cpuset_show().
> As cgroup root_list is already RCU-safe, css->cgroup is safe.
> This is similar to commit 9067d90006df ("cgroup: Eliminate the
> need for cgroup_mutex in proc_cgroup_show()")
Apologies for misleading you in my previous message about root_list.
As I look better at proc_cpuset_show vs proc_cgroup_show, there's a
difference and task_get_css() doesn't rely on root_list synchronization.
I think it could go like this (with my extra comments)
rcu_read_lock();
spin_lock_irq(&css_set_lock);
css = task_css(tsk, cpuset_cgrp_id); // css is stable wrt task's migration thanks to css_set_lock
cgrp = css->cgroup; // whatever we see here, won't be free'd thanks to RCU lock and cgroup_free_root/kfree_rcu
retval = cgroup_path_ns_locked(cgrp, buf, PATH_MAX,
current->nsproxy->cgroup_ns);
...
Your patch should work thanks to the rcu_read_lock and
cgroup_free_root/kfree_rcu and the `if (css->cgroup)` guard is
unnecessary.
So the patch is a functional fix, the reasoning in commit message is
little off. Not sure if Tejun rebases his for-6.10-fixes (with a
possible v4), full fixup commit for this may not be worthy.
Michal
[-- Attachment #2: signature.asc --]
[-- Type: application/pgp-signature, Size: 228 bytes --]
^ permalink raw reply [flat|nested] 6+ messages in thread* Re: [PATCH V3] cgroup/cpuset: Prevent UAF in proc_cpuset_show()
2024-06-27 9:46 ` Michal Koutný
@ 2024-06-27 18:29 ` Tejun Heo
2024-06-27 18:37 ` Waiman Long
0 siblings, 1 reply; 6+ messages in thread
From: Tejun Heo @ 2024-06-27 18:29 UTC (permalink / raw)
To: Michal Koutný
Cc: Chen Ridong, lizefan.x, hannes, longman, adityakali, sergeh,
cgroups, linux-kernel
Hello,
On Thu, Jun 27, 2024 at 11:46:10AM +0200, Michal Koutný wrote:
> Your patch should work thanks to the rcu_read_lock and
> cgroup_free_root/kfree_rcu and the `if (css->cgroup)` guard is
> unnecessary.
>
> So the patch is a functional fix, the reasoning in commit message is
> little off. Not sure if Tejun rebases his for-6.10-fixes (with a
> possible v4), full fixup commit for this may not be worthy.
This one's on the top. If Chen can send me a patch with updated description,
I will replace the patch.
Thanks.
--
tejun
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH V3] cgroup/cpuset: Prevent UAF in proc_cpuset_show()
2024-06-27 18:29 ` Tejun Heo
@ 2024-06-27 18:37 ` Waiman Long
0 siblings, 0 replies; 6+ messages in thread
From: Waiman Long @ 2024-06-27 18:37 UTC (permalink / raw)
To: Tejun Heo, Michal Koutný
Cc: Chen Ridong, lizefan.x, hannes, adityakali, sergeh, cgroups,
linux-kernel
On 6/27/24 14:29, Tejun Heo wrote:
> Hello,
>
> On Thu, Jun 27, 2024 at 11:46:10AM +0200, Michal Koutný wrote:
>> Your patch should work thanks to the rcu_read_lock and
>> cgroup_free_root/kfree_rcu and the `if (css->cgroup)` guard is
>> unnecessary.
I also notice that the if (css->cgroup) guard is not needed. I didn't
reject it because it doesn't hurt.
Cheers,
Longman
>>
>> So the patch is a functional fix, the reasoning in commit message is
>> little off. Not sure if Tejun rebases his for-6.10-fixes (with a
>> possible v4), full fixup commit for this may not be worthy.
> This one's on the top. If Chen can send me a patch with updated description,
> I will replace the patch.
>
> Thanks.
>
^ permalink raw reply [flat|nested] 6+ messages in thread
end of thread, other threads:[~2024-06-27 18:37 UTC | newest]
Thread overview: 6+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2024-06-26 9:41 [PATCH V3] cgroup/cpuset: Prevent UAF in proc_cpuset_show() Chen Ridong
2024-06-26 14:56 ` Waiman Long
2024-06-26 21:20 ` Tejun Heo
2024-06-27 9:46 ` Michal Koutný
2024-06-27 18:29 ` Tejun Heo
2024-06-27 18:37 ` Waiman Long
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox