From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pj2-f12.google.com (mail-pj2-f12.google.com [74.125.227.140]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C5FE14BB266 for ; Thu, 10 Sep 2026 17:00:56 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.227.140 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789059676; cv=none; b=g6KLiOuziQUXX9B2FqeQvrdYEKsiMDs2QA4Oab4UJc8IzlkgL8iL7Tr3hPeZikMEKBAQuEQfAwyqrciOGKlxhTTjcvBr2Zzef8Jgs491QpTVxxnKSwjEb87efisOeZx/76z+UmJgLLUqsqabsSHjLIHcX/tu7InM6OoO5s6d2Po= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789059676; c=relaxed/simple; bh=Llh5B20qiT9l+CISLN3EMkjfjjA3OY1FAbjV2PjhrNc=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=hAcBr6S7bt6fYxpTthMypg81jPtrAqd7ao5cw+JuwD079KtXmwz9yLm3k0J8HdY/UwIDxScSjZH+oZ78JYktkijUdlhHJ7b+Y9upmKoGGx/OLbfay3Q3n5DpEbVuLnczJURrrF9JCfAVbFJQCx3Cw4oHF8zfdu4yuqDz99zj/rw= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=H2MQqVNz; arc=none smtp.client-ip=74.125.227.140 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="H2MQqVNz" Received: by mail-pj2-f12.google.com with SMTP id 98e67ed59e1d1-39b31b4281eso1696948a91.2 for ; Thu, 10 Sep 2026 10:00:55 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1789059652; x=1789664452; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=0/NIUjuuCPvMzukokndyP3uQpFaIojJ8VScY7Mca6Ns=; b=H2MQqVNzOtPdbQPbWmo1HS/090XyrsJvAn7/RMXXnI4H1yGLD82flKuziiQNtAosGV 775f+gJGPC8FbEIPVJ/R/XbgrtdfE+Nr5mgU5KPAUpEfCNnzQTNmxSNAv2jd2hwUogQE erhRg123odsPewAsSlQI1NG80KWlD9d70UoUCArgNuZahqrNSTObNl0EeT2+HZcZPAJK bXeCwbsB1/djxUsrzcFtG4ekZsGJKHf9l+2+YQCHVZQmP+/yI7YckjPi/LbQ8YgHZ80B 4NbAoIilIjqOBCzOvS30dR3D3tTb6VlKG+iIrEobpMkV1Rv38ZbILfA+Wqe/Y49xQlOP bFLg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1789059652; x=1789664452; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=0/NIUjuuCPvMzukokndyP3uQpFaIojJ8VScY7Mca6Ns=; b=skn2pBUoPYWuocFT5LnrJaQP97kVRnidbNo0OSu59LjQKQbtm/+WthixjQDlNMH8tO oYzTGj/9nAa023ckFNcNwOO0t7dRoo69fWEKEtFVjUfowxE7WqMecAkkKrv7mzz4/+VI +DHo2gKaaH/j0TOll2vJCVJ5gqt/0a2G9N/wPxYKO64QYXjRZBYmM4avuSaubbzxfnfO mSMrAJ+GHfbS1FtERpi8PnB6rZ3OiFsxpQtIFiyacpR/JqK0Ji4KNOq3mHMwJBKYvJ1s skYGPT1shUx1N/SmijwBLOFyelHqzb1MqpWXtpsRqrr1x0ElmUINVURkgAygFoukmGKw gzqQ== X-Forwarded-Encrypted: i=1; AKwUvBywdQv57yX7Usc+BBqNHuEeeBPDJZ9rzGW2uuYskGX3tCUbNe0yq8o2w5jH8jv0JUqnYGu2O4LjzUQ=@vger.kernel.org X-Gm-Message-State: AFuF++nw9UuMDLOqcDw3VwaBaGhTfxOCSlbdRO63V2O6NW7lTOgnAHdy EZBcw3ud+tppmXo9r1nCa7wJhud17Ztqow+t5llDqmTd9udHw5x+sB4d X-Gm-Gg: AYBFou0JzuP6oYqQQIIj5WMfh476+4OZEw9Mlwnr8uUVJSZZi8VOlreo4Y64gL3uu/o zPg0zYk5Xege/cSwSX6OZEjBj35ZsS0A3ULRc539l3FnJh4ADFlgZeawlWK9R2oLLAXPn/oNp7R nX/1NcA8x2LK2XfyY+kZvlx0DLLd9iS5arbhcDTrscFZxYM0YoRlPTUcDniaFPV2Uq3mTcMr9wk pgpk+RrL/6sJBpWZG28f7n3gEGAOLX3zkKdtLhMJjdAUM5kTgWQj+0GVmO9z2o/Ioe5IHt/c4r0 OJ4JOEwc+cv878Z8vt+DKdy8TlMwcfAmyuvq+xc8I0ApV2QPGU/0fYLxFRHekIe0CVmXklwYMJC 8+/SOF9r8588pXNsLyBgZAlGWlFdmI2HzIUz6L1nzZuOD3Jr+rBC7wjavZC2gIbXMiP/E7xDnSa MqBa+2djw+JrccrEPgEQt6rARyB8yhf/Hswp3qG5pIuvFFlzLduWvOMkvU7CYLgEHjipkrPHYs7 Km9cFa0tY127RuIZ1TCITH09JlEWFA2gwxNKOQCeoQ74QvRWLySMCekY6adx2Xs4GBS8mGfyn3q AZ50ZR9SkCeyFyz2qv9wvJbEMvIBNy7p3ZiZqiYW9lHiP342D1a6PR6Mopxl+EI= X-Received: by 2002:a17:90b:39c4:b0:398:9bd4:d18 with SMTP id 98e67ed59e1d1-39d70b7cdfcmr13420866a91.23.1789059649749; Thu, 10 Sep 2026 10:00:49 -0700 (PDT) Received: from spider.bream-herring.ts.net ([103.252.203.158]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-39d94744dfdsm568538a91.0.2026.09.10.10.00.43 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 10 Sep 2026 10:00:47 -0700 (PDT) From: Matthias Goergens To: paulmck@kernel.org, urezki@gmail.com, harry@kernel.org Cc: frederic@kernel.org, neeraj.upadhyay@kernel.org, joelagnelf@nvidia.com, josh@joshtriplett.org, boqun@kernel.org, rostedt@goodmis.org, mathieu.desnoyers@efficios.com, jiangshanlai@gmail.com, qiang.zhang@linux.dev, corbet@lwn.net, skhan@linuxfoundation.org, rdunlap@infradead.org, surenb@google.com, vbabka@kernel.org, rcu@vger.kernel.org, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org Subject: [PATCH v2 0/1] rcu: make userspace barrier hook drain kvfree_rcu work Date: Fri, 11 Sep 2026 01:00:39 +0800 Message-ID: <20260910170040.344864-1-matthias.goergens@gmail.com> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260910101112.1648978-1-matthias.goergens@gmail.com> References: <20260910101112.1648978-1-matthias.goergens@gmail.com> Precedence: bulk X-Mailing-List: linux-doc@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 7bit The rcutree.do_rcu_barrier hook currently waits for ordinary RCU callbacks, but objects may still be retained in kfree_rcu() batching or a partial per-CPU SLUB sheaf. This is consistent with the hook's documented rcu_barrier() operation, but incomplete for its intended use as a boundary between userspace tests. The immediate trigger was a false allocation-leak failure in the bcachefs ktest suite while testing performance changes. Its end check writes the hook before reading /proc/allocinfo, assuming a complete deferred-free drain. Small objects remained visible after repeated hook writes and 20 seconds of waiting, so otherwise clean tests failed their leak check. Changing the hook to drain kvfree_rcu() work let the same unmodified bcachefs workload pass its allocation check. All eight checkpoints in one VM, after 50 through 400 option changes, reported zero retained reconcile_scan objects. The retained population on the original kernel eventually fell as a sheaf filled; there is no evidence here of unbounded growth or OOM. Calling kvfree_rcu_barrier() from rcu_barrier_throttled() was proposed and agreed during review of the former API in 2024, specifically to restore a clean baseline between userspace benchmark runs: https://lore.kernel.org/all/20240820155935.1167988-1-urezki@gmail.com/ This patch implements that follow-up and documents the expanded hook. It also removes the old ordinary-barrier completion shortcut: an unrelated rcu_barrier() does not establish that kvfree_rcu() work was drained. Four counterbalanced fresh-VM pairs with the full private-cache fixture reported 60 to 60 active objects on the unpatched kernel and 60 to 59 on the patched kernel. A separate ordinary-callback regression test passed on both kernels. The simplified reproducer below removes that separate regression machinery. One additional fresh control/treatment pair with this exact 41-line source confirmed the same 60 to 60 versus 60 to 59 split. These counts reflect the slab layout in the tested configuration. Save the source as rcu_barrier_sheaf_repro.c and create a Makefile containing: obj-m := rcu_barrier_sheaf_repro.o Build it with: make -C /lib/modules/$(uname -r)/build M="$PWD" modules Then, as root on a disposable test kernel: insmod rcu_barrier_sheaf_repro.ko awk '$1 == "rcu_barrier_sheaf_repro" { print $2 }' /proc/slabinfo cat /sys/kernel/slab/rcu_barrier_sheaf_repro/sheaf_capacity echo 1 > /sys/module/rcutree/parameters/do_rcu_barrier awk '$1 == "rcu_barrier_sheaf_repro" { print $2 }' /proc/slabinfo rmmod rcu_barrier_sheaf_repro The first and second slabinfo readings are 60 and 60 without the patch, and 60 and 59 with it. kmem_cache_destroy() performs per-cache deferred-free cleanup when the module is removed, after the measurement. // SPDX-License-Identifier: GPL-2.0 #include #include #include #include struct repro_object { struct rcu_head rcu; unsigned long payload; }; static struct kmem_cache *repro_cache; static int __init rcu_barrier_sheaf_repro_init(void) { struct repro_object *object; repro_cache = kmem_cache_create("rcu_barrier_sheaf_repro", sizeof(*object), 0, SLAB_NO_MERGE, NULL); if (!repro_cache) return -ENOMEM; object = kmem_cache_alloc(repro_cache, GFP_KERNEL); if (!object) { kmem_cache_destroy(repro_cache); return -ENOMEM; } kfree_rcu(object, rcu); return 0; } static void __exit rcu_barrier_sheaf_repro_exit(void) { kmem_cache_destroy(repro_cache); } module_init(rcu_barrier_sheaf_repro_init); module_exit(rcu_barrier_sheaf_repro_exit); MODULE_LICENSE("GPL"); MODULE_DESCRIPTION("Reproduce incomplete rcutree.do_rcu_barrier drains"); --- Changes since v1: - Add the motivating bcachefs failure and the successful unmodified workload result to both the cover letter and commit message. - Drop the incorrect sheaf Fixes: tag and regression framing; describe this as a strengthening of the existing test interface. - Credit the agreed 2024 proposal for this extension. - Broaden the subject and changelog from sheaves to kvfree_rcu work. - Hard-wrap the prose for text-based mail readers. The code diff is unchanged from v1. The results above are the existing validation results; no new kernel tests were run for this prose revision. v1: https://lore.kernel.org/all/20260910101112.1648978-1-matthias.goergens@gmail.com/ Matthias Goergens (1): rcu: make userspace barrier hook drain kvfree_rcu work .../admin-guide/kernel-parameters.txt | 7 ++--- kernel/rcu/tree.c | 27 ++++++++++++------- 2 files changed, 21 insertions(+), 13 deletions(-) base-commit: 50d05c7c76c96b90462f24debacca971d2e86713 -- 2.55.0