From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wm1-f52.google.com (mail-wm1-f52.google.com [209.85.128.52]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1048C3D75D1 for ; Thu, 27 Aug 2026 22:32:34 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.52 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787869958; cv=none; b=g04+hz716uC6+pRvjzdDM5cYbAdmByrl+2zDLOzAIUoc1z/XKK6kwYzpBtPoN7tcEH3jp87CmQdQRbm9Md4HpkkBBk9QcGZ+CdSjJjxuUBOfw7S0y53bco5OaxvUiIcxEcKho3LeER4H5foAM6tY1MAWMcls4g/0O15GoYDqOf0= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787869958; c=relaxed/simple; bh=0KS9yDRF5ufQr71N4f6ELYPxUEqie2Ct2GuCnnm+4D4=; h=From:Date:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=HNOny9wrcXcod6tndCF8W/KaVVwslUbhtArXIlIHVWUpKno36u2n/M+VsSiEeVsCEkpE/CysFb655JaS3IYbLZT2IFMCWRDJnSTkv8bqoxauFJbymBeIrIBtbly61JRExMvMxgUY8Mf4WpT4dD46J/Mslge/7huBzAX1V97/lyk= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=cGjJ0P1V; arc=none smtp.client-ip=209.85.128.52 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="cGjJ0P1V" Received: by mail-wm1-f52.google.com with SMTP id 5b1f17b1804b1-4995b0343c1so2235605e9.3 for ; Thu, 27 Aug 2026 15:32:34 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787869953; x=1788474753; darn=vger.kernel.org; h=in-reply-to:content-transfer-encoding:content-disposition :content-type:mime-version:references:message-id:subject:cc:to:date :from:from:to:cc:subject:date:message-id:reply-to:content-type; bh=NfHSf/Xnpm7vI/oia3ddZqw/6ZjMqOWbQV8t0TtjrWE=; b=cGjJ0P1V/ejP6XpS6pTvqW0+Rak5crvE0dpMXglt5y5lV4dIOm4wOxzKfzGutVYTlg Nu5RP2drKVbc8WvgKPIf6pHLDwsOWAI/7lLkpoUDItpPns561uKtfUhlAjjGRlUvT8Gq VqO8oND824Q2EWFSpLKkP239YnG8LlYOb3kS8lJ/b0LGC6HuByMlkI4UMySd6PwVqwRI ZqyMZiEDRj1M1WAdHyoKVhF9ceFfQtIzez4EaN87cAyjsBjJ0hFhorRv03e/lkQnPfHp izCBohRlA80yZmwO4BIyux0jvN/pzHpd5U4218/htq9gWdGcbQNvGVs4D/wAesTQJ10j DYcw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787869953; x=1788474753; h=in-reply-to:content-transfer-encoding:content-disposition :content-type:mime-version:references:message-id:subject:cc:to:date :from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to:content-type; bh=NfHSf/Xnpm7vI/oia3ddZqw/6ZjMqOWbQV8t0TtjrWE=; b=CTZEZwsk9G/LYxcw+OeDt3yAdjJyjX/bS+2wPzWbKxmk8syhkN2KDAi6ffBnuZKJXO 3lfEQi+1NXg9Fi+t34NTwH7MGulD+xDOlD7JTENqsnOqEUBeRCuEfmD8q4oGH2AJ4Bjg d0kDH6W5Cgo+2MlwqBfYloSBp9JRdqniZhY9A71crVaeZUfA/UaNIftcvd3VDLniSDbT nzfDaBE7xYImdfhq297Dqoe1Roz9N7I74EZyT1rcUihgz0Goz6qWY7B5/ihxSGNps54e 5aS0SESTsjmhI2B15gFM8yHzJeei/zqJAMYRG0sKC9DCjvv9TsGOdT4XZ43ekDZqxfWQ F0ag== X-Forwarded-Encrypted: i=1; AHgh+RpfSN8zcFfwpk+QnTUJa6bywDWLDsZaGTtCe2O5jEyqKrZhujnN9z5Dp/k+9wzssetonNQ=@vger.kernel.org X-Gm-Message-State: AFuF++lVIZzIEaRQWra360GzogQqCfNPgZ+en7bcupVOQ4zMb6e8w+8g Wp5E9sG/NbkHu7wDHP7ZNx02bLyT+KHfu50g0EWWA/I3W12ztwB1cP01 X-Gm-Gg: AR+sD10wPv7gwlPIBdRbGccMvlN0/XsZj/8qNEwrJgGI/GNBSn/Lcbq6Lzwx4Qp46wb 7aCdo+CqsJKbEChY5Jqoku0FS26/sG/SOk8NNkndWw48OgCcg6NwJ8j3wHkqfFIp2TWld8yGnuk nExAGAteseIejozJrWmrsa69OU75fnOIla7aTZDTz70yiY9xB51hJKSWaPvkK8c/MsaK+zFNiVG xZ7TqiRrYozf4qdSo3XidCfUbmC4+BVQ8vqQ8kWcqRbN45NOdOmHy7/fQbUoh2aa07Q6zKJinq8 ykDLzlz3v6XWPF3qcLPYUgSdx+puv3gZzLYKdDfEsvkUQWlDjbyvwtjryZ20M/aQ0zQSg9UsT4f zRCQ3tCwgPErrGa4voPFBN9JTDgjUOV6YtmJWjJBv6YCyWyUaoSo0EUDDFdwdahNYQuHLTg70cL 7xZAn4ASv9gx175hdDQg1rutdCsI+SHtiFk2VuJ5D0cQU= X-Received: by 2002:a05:600c:1d11:b0:496:c1f3:e8f8 with SMTP id 5b1f17b1804b1-49b91c25414mr28676775e9.7.1787869952896; Thu, 27 Aug 2026 15:32:32 -0700 (PDT) Received: from krava ([77.78.87.81]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49b91ca11d0sm14821055e9.1.2026.08.27.15.32.29 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 27 Aug 2026 15:32:30 -0700 (PDT) From: Jiri Olsa X-Google-Original-From: Jiri Olsa Date: Fri, 28 Aug 2026 00:32:27 +0200 To: Alexei Starovoitov Cc: Andrii Nakryiko , Christian Simon , bpf , Alexei Starovoitov , Andrii Nakryiko , Daniel Borkmann , Martin KaFai Lau , Tejun Heo , Yonghong Song , stable , Jiri Olsa Subject: Re: [PATCH bpf v3 1/2] bpf: disable private stack for sleepable programs Message-ID: References: <20260822225444.2774461-1-simon@swine.de> <20260822225444.2774461-2-simon@swine.de> Precedence: bulk X-Mailing-List: bpf@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Disposition: inline Content-Transfer-Encoding: 8bit In-Reply-To: On Thu, Aug 27, 2026 at 09:55:08AM -0700, Alexei Starovoitov wrote: > On Thu, Aug 27, 2026 at 9:40 AM Andrii Nakryiko > wrote: > > > > On Thu, Aug 27, 2026 at 9:35 AM Alexei Starovoitov > > wrote: > > > > > > On Thu, Aug 27, 2026 at 7:56 AM Andrii Nakryiko > > > wrote: > > > > > > > > On Tue, Aug 25, 2026 at 6:20 PM Alexei Starovoitov > > > > wrote: > > > > > > > > > > On Sat Aug 22, 2026 at 3:54 PM PDT, Christian Simon wrote: > > > > > > A JITed BPF program can use one private stack per program and CPU. > > > > > > Sleepable programs can be preempted, allowing another task to run the > > > > > > same program on the same CPU. The second invocation then reuses and can > > > > > > overwrite the first invocation's private stack. > > > > > > > > > > I'm confused by this. sleepable progs go through __bpf_prog_enter_sleepable_recur() > > > > > which has per-prog recurison counter. So preemption of the prog > > > > > doesn't break private stack. > > > > > If the same prog attemps to execute on the same cpu it will be skipped. > > > > > > > > > > syscall prog types go via bpf_prog_run_array_sleepable() > > > > > that have per prog recursions counter. > > > > > > > > > > Looks like we're not doing it for bpf_prog_run_array_uprobe(). > > > > > I'm not sure what the right trade off here. > > > > > I feel universally checking for recursion is better > > > > > then selectively disabling private stack for uprobe. > > > > > > > > I'd really like to avoid adding this "recursion protection" to uprobe. > > > > With uprobes, there is no recursion, it's called from well defined > > > > context in the kernel and you can't have recursive uprobe BPF > > > > programs. > > > > > > > > All you can have is a very valid and possible sleepable uprobe > > > > interleaving, which the user cannot prevent or work around, they have > > > > no control over this and it's just a fact of life. > > > > > > > > E.g., a simple scenario, we attach one bpf program (let's call it U) > > > > to some USDT. BPF program U is sleepable and actually can sleep due to > > > > page faults (e.g., unwinding Python stack trace requires sleepable > > > > mode for reliably getting filename strings from Python runtime, which > > > > are not always paged in). > > > > > > > > In such a case, you can have thread A and thread B both hitting the > > > > same USDT (e.g., somewhere in memory allocator or whatnot). Let's say > > > > thread A hits it first on CPU X, BPF program U starts executing and > > > > unwinding Python stack, does bpf_copy_from_user() for string contents > > > > and causes page fault, is taken off CPU X. Meanwhile thread B hits > > > > USDT on the same CPU X, kernel runs program U, and it is supposed to > > > > work completely independently and concurrently (no shared state or > > > > whatever) from U's execution in thread A. > > > > > > > > Yet, if we add this per-CPU "recursion check", we'll just skip U's > > > > execution for thread B. This is data loss, and it's very bad in > > > > practice because it frequently just invalidates the entire data > > > > collection trustworthiness. > > > > > > ok. fair > > > > > > > great, thanks! > > > > > > So I think we should disable private stack for uprobes (sleepable or > > > > not) instead. I'm not sure private stack buys us anything for uprobe > > > > cases. > > > > > > why disable priv stack for non-sleepable uprobes? > > > While non-sleepable bpf prog is executing the same or different > > > uprobe cannot execute on the same cpu. > > > So bpf prog can be preempted by kernel execution, > > > but a user task cannot start preempt bpf prog, > > > so 2nd uprobe cannot start running, > > > no? > > > > I think that changes on preemptible kernels, this was called out in > > discussions on previous versions of this patch. So only for that > > reason. > > my understanding is that preemptable kernel doesn't mean that > bpf prog can be preempted by user space. > only by kernel. > > I looked up earlier thread, but don't understand what Jiri meant. > > Jiri, > please clarify what problem do you see with non-sleepable uprobes? hum.. non-sleepable uprobe prog is run by bpf_prog_run_array_uprobe and it disables only task migration, preemption is not disabled and holds rcu_read_lock (which seems ok for preemption) so I'm not sure why it wouldn't be preemptible by another task jirka