From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AD3826EB79; Mon, 28 Jul 2025 12:35:07 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1753706107; cv=none; b=dXJ51yN7Ry0NkLzK811vpPjxNeOvVQk/UX5NAVU2SBvdTZvNJcW504cL/j5YD0PhjF8Z7AodoT9inAlnqF+azTn6gZWIYFodWOBPW7IrjHvxDOfnmd9QBtNYGGE0nVLxsXzD9o5eODMNjgJHg0Mlq1p8dffaCd+r8qDdZ2qXTjY= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1753706107; c=relaxed/simple; bh=vsvA8ck8Us4xh8p23FEeipmCHTANlh9c3F96aksfHSU=; h=Date:From:To:Cc:Subject:Message-Id:In-Reply-To:References: Mime-Version:Content-Type; b=jObV4HO02H5WNJbIN6LGlj7DEdZtAaUdwkXShKGip8BxrcKltfdtKB2oTRft/3nuwBbIKPlQHqG2LkDHUConRIKEnVlt/uPlv9UFn5BqK+jAPpfIyF38WiyoUJ5Y08pN/NO2FeTmUJsN/jmCX2sJh1ZjPHHeji7X5aRK+rF1354= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=pFHswPcq; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="pFHswPcq" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 012C9C4CEE7; Mon, 28 Jul 2025 12:35:04 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1753706107; bh=vsvA8ck8Us4xh8p23FEeipmCHTANlh9c3F96aksfHSU=; h=Date:From:To:Cc:Subject:In-Reply-To:References:From; b=pFHswPcqw5EypQY7ShgYgiqYXCrKuml2PLpXcLxWf0JeZuzBVs2Ov+lYxF6X3gVwZ XoZIb67EixAAFDW8Qjbx7wDSEjwCwzpPOZTWmsjbZxqXHFZ6tNbtRn4joLyk6DcYo3 DBCDl01wF7nAOUf+cw+1iY/I4k0hIZrZjHp5sufVQCp9uZoParBPSwzhnCoEiUfKqM myHlTFWVSSjewN73EKVgK8726g4e0NtLu84tBbZDTm2Yi85FpwFKAbk9l1nX2bubyY YLAzZX1x4hEO4B00f5RT5gYLZ8gjSVPdJEwL4md6yC1kEJnlw8OxZoz3K17qAkl1Fc qFwGe7lerJusw== Date: Mon, 28 Jul 2025 21:35:02 +0900 From: Masami Hiramatsu (Google) To: Menglong Dong Cc: alexei.starovoitov@gmail.com, rostedt@goodmis.org, mathieu.desnoyers@efficios.com, hca@linux.ibm.com, revest@chromium.org, linux-kernel@vger.kernel.org, linux-trace-kernel@vger.kernel.org, bpf@vger.kernel.org Subject: Re: [PATCH RFC bpf-next v2 0/4] fprobe: use rhashtable for fprobe_ip_table Message-Id: <20250728213502.df49c01629de5fe9b6f15632@kernel.org> In-Reply-To: <20250728072637.1035818-1-dongml2@chinatelecom.cn> References: <20250728072637.1035818-1-dongml2@chinatelecom.cn> X-Mailer: Sylpheed 3.8.0beta1 (GTK+ 2.24.33; x86_64-pc-linux-gnu) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Hi Menglong, What are the updates from v1? Just adding RFC? Thanks, On Mon, 28 Jul 2025 15:22:49 +0800 Menglong Dong wrote: > For now, the budget of the hash table that is used for fprobe_ip_table is > fixed, which is 256, and can cause huge overhead when the hooked functions > is a huge quantity. > > In this series, we use rhltable for fprobe_ip_table to reduce the > overhead. > > Meanwhile, we also add the benchmark testcase "kprobe-multi-all", which > will hook all the kernel functions during the testing. Before this series, > the performance is: > usermode-count : 875.380 ± 0.366M/s > kernel-count : 435.924 ± 0.461M/s > syscall-count : 31.004 ± 0.017M/s > fentry : 134.076 ± 1.752M/s > fexit : 68.319 ± 0.055M/s > fmodret : 71.530 ± 0.032M/s > rawtp : 202.751 ± 0.138M/s > tp : 79.562 ± 0.084M/s > kprobe : 55.587 ± 0.028M/s > kprobe-multi : 56.481 ± 0.043M/s > kprobe-multi-all: 6.283 ± 0.005M/s << look this > kretprobe : 22.378 ± 0.028M/s > kretprobe-multi: 28.205 ± 0.025M/s > > With this series, the performance is: > usermode-count : 902.387 ± 0.762M/s > kernel-count : 427.356 ± 0.368M/s > syscall-count : 30.830 ± 0.016M/s > fentry : 135.554 ± 0.064M/s > fexit : 68.317 ± 0.218M/s > fmodret : 70.633 ± 0.275M/s > rawtp : 193.404 ± 0.346M/s > tp : 80.236 ± 0.068M/s > kprobe : 55.200 ± 0.359M/s > kprobe-multi : 54.304 ± 0.092M/s > kprobe-multi-all: 54.487 ± 0.035M/s << look this > kretprobe : 22.381 ± 0.075M/s > kretprobe-multi: 27.926 ± 0.034M/s > > The benchmark of "kprobe-multi-all" increase from 6.283M/s to 54.487M/s. > > The locking is not handled properly in the first patch. In the > fprobe_entry, we should use RCU when we access the rhlist_head. However, > we can't use RCU for __fprobe_handler, as it can sleep. In the origin > logic, it seems that the usage of hlist_for_each_entry_from_rcu() is not > protected by rcu_read_lock neither, isn't it? I don't know how to handle > this part ;( > > Menglong Dong (4): > fprobe: use rhltable for fprobe_ip_table > selftests/bpf: move get_ksyms and get_addrs to trace_helpers.c > selftests/bpf: skip recursive functions for kprobe_multi > selftests/bpf: add benchmark testing for kprobe-multi-all > > include/linux/fprobe.h | 2 +- > kernel/trace/fprobe.c | 141 ++++++----- > tools/testing/selftests/bpf/bench.c | 2 + > .../selftests/bpf/benchs/bench_trigger.c | 30 +++ > .../selftests/bpf/benchs/run_bench_trigger.sh | 2 +- > .../bpf/prog_tests/kprobe_multi_test.c | 220 +---------------- > tools/testing/selftests/bpf/trace_helpers.c | 230 ++++++++++++++++++ > tools/testing/selftests/bpf/trace_helpers.h | 3 + > 8 files changed, 348 insertions(+), 282 deletions(-) > > -- > 2.50.1 > -- Masami Hiramatsu (Google)