From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 04965516148; Wed, 30 Sep 2026 16:47:33 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790786854; cv=none; b=He7IXREL5Yel24KVu1x1MSMG3HxyAxTsIvj+2TooiWVyhyhxtQ2Fgn+fzK1wHL0g3Mdfbc9PQM3fddIJr4k02H9It1+x9wqXxHEj2/3vAY9vFE+UXnV6t7+5sPRCvSbeojwenzMFNoFHsGCIrhrdmTQeGH8gWXAWh4occRsN9Ww= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790786854; c=relaxed/simple; bh=EHePmHe7vqbQR9PooRowVrO/nRv2lZLw74r2kAeBJb8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=JAid6vepDiozpaLt1S2utcAhyx9zeLUf4MhHnKjdz50VIdhezAaXqd0Um1m/X4V6d9NiK9MnpMhLqRiX7Wew/8jgewuurXwbhKdE338PtMG/0ZwrMqX7pn6SV1B0ozcGVKhOdqT15AG3a/pEDjNHiX+QZlAoKsoioTqU19Zyviw= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linuxfoundation.org header.i=@linuxfoundation.org header.b=CjbqsvKX; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linuxfoundation.org header.i=@linuxfoundation.org header.b="CjbqsvKX" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 60EE31F000FF; Wed, 30 Sep 2026 16:47:32 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linuxfoundation.org; s=korg; t=1790786852; bh=R8uCMkgShzY7pkZF6//otl4sGyM/hH7sgSUfqaktUyw=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=CjbqsvKXjB6Fr8dwH1wgf6Co8G4sK8JwHicifvk0HSbFlH1XRPzsAiYuuhHCSXQrz ePQvku5rIkwCu3mtbre/qZFPQ5F1ZpI1rcoc5zLIK0nSzMlaqfLeyTDxgpv0TqHHCy Zn4XyzrluEB31ExY2/YaciYMjYOCAoXY+1RtDnq8= From: Greg Kroah-Hartman To: stable@vger.kernel.org Cc: Greg Kroah-Hartman , patches@lists.linux.dev, "Paul E. McKenney" , Rik van Riel , "Jose Fernandez (Anthropic)" , Josef Bacik , Alexei Starovoitov , Sasha Levin Subject: [PATCH 7.2 026/457] bpf: Avoid soft lockup in __htab_map_lookup_and_delete_batch() Date: Wed, 30 Sep 2026 17:22:11 +0200 Message-ID: <20260930152346.595896089@linuxfoundation.org> X-Mailer: git-send-email 2.55.0 In-Reply-To: <20260930152346.024115587@linuxfoundation.org> References: <20260930152346.024115587@linuxfoundation.org> User-Agent: quilt/0.69 X-stable: review X-Patchwork-Hint: ignore Precedence: bulk X-Mailing-List: patches@lists.linux.dev List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit 7.2-stable review patch. If anyone has any objections, please let me know. ------------------ From: Jose Fernandez (Anthropic) [ Upstream commit 85136bf22404474a815fc0ed26ec0d1cbc1bc3f9 ] __htab_map_lookup_and_delete_batch() has no rescheduling point. The batch count bounds how many entries are copied out, not how many buckets are visited, so one BPF_MAP_LOOKUP_BATCH call can walk the map end to end. The empty-bucket fast path is worse: it stays inside a single rcu_read_lock() / bpf_disable_instrumentation() section for any run of consecutive empty buckets. That holds up on small maps, but it falls apart at scale. On a 144-CPU arm64 host running a CONFIG_PREEMPT_NONE kernel, periodic BPF_MAP_LOOKUP_BATCH calls against an LRU hash map with 16,777,216 buckets held a CPU inside the batch op for 77+ seconds and triggered the soft lockup watchdog. Commit 75134f16e7dd ("bpf: Add schedule points in batch ops") fixed this same problem in the generic batch ops, but not in this htab-native path, which every htab-based hash map variant uses for its lookup[_and_delete] batch ops. Complete that fix here. Leave the critical section after 64 consecutive empty buckets, call cond_resched_tasks_rcu_qs(), and resume at the saved bucket cursor. No locks are held at that point, and resuming from the cursor is already the function's behavior for non-empty buckets. Add the same call to the per-bucket loop after copy_to_user(), where every lock has been dropped. cond_resched_rcu() is not enough here: sleeping with bpf_prog_active elevated makes tracing programs on that CPU silently skip their invocations. Plain cond_resched() is not enough either. It is a no-op under PREEMPT and PREEMPT_LAZY, the only models arm64 and x86 have offered since commit 7dadeaa6e851 ("sched: Further restrict the preemption modes"). It is also never a Tasks RCU quiescent state, in any model: the reschedule counts as a preemption. The walking task stays a holdout and stalls every synchronize_rcu_tasks() caller, ftrace and BPF trampoline teardown included, until the syscall returns [1]. cond_resched_tasks_rcu_qs() is the usual tool for that [2]. It reports the quiescent state at each yield and still reschedules as cond_resched() does on PREEMPT_NONE and PREEMPT_VOLUNTARY kernels. Fixes: 057996380a42 ("bpf: Add batch ops to all htab bpf map") Cc: "Paul E. McKenney" Cc: Rik van Riel Link: https://lore.kernel.org/bpf/20260715215314.44423f47@fangorn/ [1] Link: https://lore.kernel.org/bpf/9d444098-7c03-4163-af12-bd0a79a51443@paulmck-laptop/ [2] Assisted-by: LLM Signed-off-by: Jose Fernandez (Anthropic) Signed-off-by: Josef Bacik Reviewed-by: Rik van Riel Link: https://lore.kernel.org/r/20260909-b4-htab-batch-resched-v2-1-0cb529d8f95a@toxicpanda.com Signed-off-by: Alexei Starovoitov Signed-off-by: Sasha Levin --- kernel/bpf/hashtab.c | 24 +++++++++++++++++++++--- 1 file changed, 21 insertions(+), 3 deletions(-) diff --git a/kernel/bpf/hashtab.c b/kernel/bpf/hashtab.c index 59406da06424d..419692f41ffde 100644 --- a/kernel/bpf/hashtab.c +++ b/kernel/bpf/hashtab.c @@ -1773,6 +1773,12 @@ static int htab_lru_percpu_map_lookup_and_delete_elem(struct bpf_map *map, flags); } +/* + * Max consecutive empty buckets to walk in one RCU + + * instrumentation-disabled section before rescheduling. + */ +#define HTAB_BATCH_EMPTY_RESCHED 64 + static int __htab_map_lookup_and_delete_batch(struct bpf_map *map, const union bpf_attr *attr, @@ -1794,6 +1800,7 @@ __htab_map_lookup_and_delete_batch(struct bpf_map *map, unsigned long flags = 0; bool locked = false; struct htab_elem *l; + u32 empty_cnt = 0; struct bucket *b; int ret = 0; @@ -1972,12 +1979,21 @@ __htab_map_lookup_and_delete_batch(struct bpf_map *map, } next_batch: - /* If we are not copying data, we can go to next bucket and avoid - * unlocking the rcu. + /* + * If we are not copying data, we can go to next bucket and avoid + * unlocking the rcu. Bound the walk though: after + * HTAB_BATCH_EMPTY_RESCHED consecutive empty buckets, fully exit + * the critical section (no locks are held here) and reschedule. */ if (!bucket_cnt && (batch + 1 < htab->n_buckets)) { batch++; - goto again_nocopy; + if (++empty_cnt < HTAB_BATCH_EMPTY_RESCHED) + goto again_nocopy; + empty_cnt = 0; + rcu_read_unlock(); + bpf_enable_instrumentation(); + cond_resched_tasks_rcu_qs(); + goto again; } rcu_read_unlock(); @@ -1991,11 +2007,13 @@ __htab_map_lookup_and_delete_batch(struct bpf_map *map, } total += bucket_cnt; + empty_cnt = 0; batch++; if (batch >= htab->n_buckets) { ret = -ENOENT; goto after_loop; } + cond_resched_tasks_rcu_qs(); goto again; after_loop: -- 2.53.0