BPF List
 help / color / mirror / Atom feed
* [PATCH bpf-next] bpftool: Print average cycles per program run in profiler
@ 2026-08-27 15:29 Mykyta Yatsenko
  2026-08-27 16:22 ` bot+bpf-ci
                   ` (2 more replies)
  0 siblings, 3 replies; 7+ messages in thread
From: Mykyta Yatsenko @ 2026-08-27 15:29 UTC (permalink / raw)
  To: bpf, ast, andrii, daniel, kernel-team, eddyz87, memxor, qmo
  Cc: Mykyta Yatsenko

From: Mykyta Yatsenko <yatsenko@meta.com>

Total cycle counts are difficult to compare across workloads with
different run counts. Report cycles per run to expose the per-invocation
cost directly.

Example:
sudo ./bpftool prog profile name mprog duration 15 cycles instructions

            423256 run_cnt
         947413975 cycles              #  2238.39 cycles per run
         333965846 instructions        #     0.35 insns per cycle

Signed-off-by: Mykyta Yatsenko <yatsenko@meta.com>
---
 tools/bpf/bpftool/Documentation/bpftool-prog.rst |  6 ++++--
 tools/bpf/bpftool/prog.c                         | 18 +++++++++++-------
 2 files changed, 15 insertions(+), 9 deletions(-)

diff --git a/tools/bpf/bpftool/Documentation/bpftool-prog.rst b/tools/bpf/bpftool/Documentation/bpftool-prog.rst
index 90fa2a48cc26..2280dc4492c0 100644
--- a/tools/bpf/bpftool/Documentation/bpftool-prog.rst
+++ b/tools/bpf/bpftool/Documentation/bpftool-prog.rst
@@ -217,7 +217,9 @@ bpftool prog run *PROG* data_in *FILE* [data_out *FILE* [data_size_out *L*]] [ct
 bpftool prog profile *PROG* [duration *DURATION*] *METRICs*
     Profile *METRICs* for bpf program *PROG* for *DURATION* seconds or until
     user hits <Ctrl+C>. *DURATION* is optional. If *DURATION* is not specified,
-    the profiling will run up to **UINT_MAX** seconds.
+    the profiling will run up to **UINT_MAX** seconds. When **cycles** is
+    selected, plain output also reports the average number of cycles per
+    program run.
 
 bpftool prog help
     Print short help message.
@@ -360,7 +362,7 @@ EXAMPLES
 ::
 
          51397 run_cnt
-      40176203 cycles                                                 (83.05%)
+      40176203 cycles          # 781.68 cycles per run                (83.05%)
       42518139 instructions    #   1.06 insns per cycle               (83.39%)
            123 llc_misses      #   2.89 LLC misses per million insns  (83.15%)
 
diff --git a/tools/bpf/bpftool/prog.c b/tools/bpf/bpftool/prog.c
index a9f730d407a9..8935508f955b 100644
--- a/tools/bpf/bpftool/prog.c
+++ b/tools/bpf/bpftool/prog.c
@@ -2069,9 +2069,9 @@ struct profile_metric {
 	bool selected;
 
 	/* calculate ratios like instructions per cycle */
-	const int ratio_metric; /* 0 for N/A, 1 for index 0 (cycles) */
+	const int ratio_metric; /* 0 for run_cnt, 1 for index 0 (cycles) */
 	const char *ratio_desc;
-	const float ratio_mul;
+	const double ratio_mul;
 } metrics[] = {
 	{
 		.name = "cycles",
@@ -2080,6 +2080,9 @@ struct profile_metric {
 			.config = PERF_COUNT_HW_CPU_CYCLES,
 			.exclude_user = 1,
 		},
+		.ratio_metric = 0,
+		.ratio_desc = "cycles per run",
+		.ratio_mul = 1.0,
 	},
 	{
 		.name = "instructions",
@@ -2256,17 +2259,18 @@ static void profile_print_readings_plain(void)
 	for (m = 0; m < ARRAY_SIZE(metrics); m++) {
 		struct bpf_perf_event_value *val = &metrics[m].val;
 		int r;
+		__u64 ratio;
 
 		if (!metrics[m].selected)
 			continue;
 		printf("%18llu %-20s", val->counter, metrics[m].name);
 
-		r = metrics[m].ratio_metric - 1;
-		if (r >= 0 && metrics[r].selected &&
-		    metrics[r].val.counter > 0) {
+		r = metrics[m].ratio_metric;
+		/* r == 0 is a special case for run_cnt */
+		ratio = r ? metrics[r - 1].val.counter : profile_total_count;
+		if (metrics[m].ratio_desc && ratio) {
 			printf("# %8.2f %-30s",
-			       val->counter * metrics[m].ratio_mul /
-			       metrics[r].val.counter,
+			       val->counter * metrics[m].ratio_mul / ratio,
 			       metrics[m].ratio_desc);
 		} else {
 			printf("%-41s", "");

---
base-commit: 23ff631b3b8b1452dfe933ee21f96321a9c5e209
change-id: 20260827-bpftool_cyles_per_run-93169fcb30b7

Best regards,
--  
Mykyta Yatsenko <yatsenko@meta.com>


^ permalink raw reply related	[flat|nested] 7+ messages in thread

* Re: [PATCH bpf-next] bpftool: Print average cycles per program run in profiler
  2026-08-27 15:29 [PATCH bpf-next] bpftool: Print average cycles per program run in profiler Mykyta Yatsenko
@ 2026-08-27 16:22 ` bot+bpf-ci
  2026-08-28  0:30 ` Andrii Nakryiko
  2026-08-28 18:46 ` Suchit Karunakaran
  2 siblings, 0 replies; 7+ messages in thread
From: bot+bpf-ci @ 2026-08-27 16:22 UTC (permalink / raw)
  To: mykyta.yatsenko5, bpf, ast, andrii, daniel, kernel-team, eddyz87,
	memxor, qmo
  Cc: yatsenko, ast, andrii, daniel, martin.lau, eddyz87, yonghong.song,
	clm, ihor.solodrai

[-- Attachment #1: Type: text/plain, Size: 3006 bytes --]

> diff --git a/tools/bpf/bpftool/Documentation/bpftool-prog.rst b/tools/bpf/bpftool/Documentation/bpftool-prog.rst
> index 90fa2a48cc26a..2280dc4492c06 100644
> --- a/tools/bpf/bpftool/Documentation/bpftool-prog.rst
> +++ b/tools/bpf/bpftool/Documentation/bpftool-prog.rst

[ ... ]

> diff --git a/tools/bpf/bpftool/prog.c b/tools/bpf/bpftool/prog.c
> index a9f730d407a92..8935508f955b1 100644
> --- a/tools/bpf/bpftool/prog.c
> +++ b/tools/bpf/bpftool/prog.c
> @@ -2069,9 +2069,9 @@ struct profile_metric {
>  	bool selected;
>
>  	/* calculate ratios like instructions per cycle */
> -	const int ratio_metric; /* 0 for N/A, 1 for index 0 (cycles) */
> +	const int ratio_metric; /* 0 for run_cnt, 1 for index 0 (cycles) */
>  	const char *ratio_desc;
> -	const float ratio_mul;
> +	const double ratio_mul;
>  } metrics[] = {

This isn't a bug, but could the comment also mention that a NULL
ratio_desc means no ratio is printed? The l1d_loads entry now depends on
that rather than on ratio_metric == 0.

>  	{
>  		.name = "cycles",
> @@ -2080,6 +2080,9 @@ struct profile_metric {
>  			.config = PERF_COUNT_HW_CPU_CYCLES,
>  			.exclude_user = 1,
>  		},
> +		.ratio_metric = 0,
> +		.ratio_desc = "cycles per run",
> +		.ratio_mul = 1.0,
>  	},

[ ... ]

> @@ -2256,17 +2259,18 @@ static void profile_print_readings_plain(void)
>  	for (m = 0; m < ARRAY_SIZE(metrics); m++) {
>  		struct bpf_perf_event_value *val = &metrics[m].val;
>  		int r;
> +		__u64 ratio;
>
>  		if (!metrics[m].selected)
>  			continue;
>  		printf("%18llu %-20s", val->counter, metrics[m].name);
>
> -		r = metrics[m].ratio_metric - 1;
> -		if (r >= 0 && metrics[r].selected &&
> -		    metrics[r].val.counter > 0) {
> +		r = metrics[m].ratio_metric;
> +		/* r == 0 is a special case for run_cnt */
> +		ratio = r ? metrics[r - 1].val.counter : profile_total_count;
> +		if (metrics[m].ratio_desc && ratio) {
>  			printf("# %8.2f %-30s",
> -			       val->counter * metrics[m].ratio_mul /
> -			       metrics[r].val.counter,
> +			       val->counter * metrics[m].ratio_mul / ratio,
>  			       metrics[m].ratio_desc);

This isn't a bug, but would a name like 'divisor' or 'denom' read more
naturally here? The variable holds the denominator rather than the
computed ratio, which makes the expression 'val->counter * ... / ratio'
read as if dividing a ratio by a ratio.

The commit message explains the new cycles-per-run output but not the
encoding change that makes it work. This isn't a bug, but would it be
worth a sentence noting that ratio_metric == 0 now means run_cnt (and
that ratio_desc is what gates printing)? Currently readers have to derive
the new encoding from the diff.

The unrelated widening of ratio_mul from float to double is also
unmentioned in the changelog.


---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md

CI run summary: https://github.com/kernel-patches/bpf/actions/runs/33089436017

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH bpf-next] bpftool: Print average cycles per program run in profiler
  2026-08-27 15:29 [PATCH bpf-next] bpftool: Print average cycles per program run in profiler Mykyta Yatsenko
  2026-08-27 16:22 ` bot+bpf-ci
@ 2026-08-28  0:30 ` Andrii Nakryiko
  2026-08-28 11:00   ` Mykyta Yatsenko
  2026-08-28 18:46 ` Suchit Karunakaran
  2 siblings, 1 reply; 7+ messages in thread
From: Andrii Nakryiko @ 2026-08-28  0:30 UTC (permalink / raw)
  To: Mykyta Yatsenko
  Cc: bpf, ast, andrii, daniel, kernel-team, eddyz87, memxor, qmo,
	Mykyta Yatsenko

On Thu, Aug 27, 2026 at 8:29 AM Mykyta Yatsenko
<mykyta.yatsenko5@gmail.com> wrote:
>
> From: Mykyta Yatsenko <yatsenko@meta.com>
>
> Total cycle counts are difficult to compare across workloads with
> different run counts. Report cycles per run to expose the per-invocation
> cost directly.
>
> Example:
> sudo ./bpftool prog profile name mprog duration 15 cycles instructions
>
>             423256 run_cnt
>          947413975 cycles              #  2238.39 cycles per run
>          333965846 instructions        #     0.35 insns per cycle
>
> Signed-off-by: Mykyta Yatsenko <yatsenko@meta.com>
> ---
>  tools/bpf/bpftool/Documentation/bpftool-prog.rst |  6 ++++--
>  tools/bpf/bpftool/prog.c                         | 18 +++++++++++-------
>  2 files changed, 15 insertions(+), 9 deletions(-)
>
> diff --git a/tools/bpf/bpftool/Documentation/bpftool-prog.rst b/tools/bpf/bpftool/Documentation/bpftool-prog.rst
> index 90fa2a48cc26..2280dc4492c0 100644
> --- a/tools/bpf/bpftool/Documentation/bpftool-prog.rst
> +++ b/tools/bpf/bpftool/Documentation/bpftool-prog.rst
> @@ -217,7 +217,9 @@ bpftool prog run *PROG* data_in *FILE* [data_out *FILE* [data_size_out *L*]] [ct
>  bpftool prog profile *PROG* [duration *DURATION*] *METRICs*
>      Profile *METRICs* for bpf program *PROG* for *DURATION* seconds or until
>      user hits <Ctrl+C>. *DURATION* is optional. If *DURATION* is not specified,
> -    the profiling will run up to **UINT_MAX** seconds.
> +    the profiling will run up to **UINT_MAX** seconds. When **cycles** is
> +    selected, plain output also reports the average number of cycles per
> +    program run.
>
>  bpftool prog help
>      Print short help message.
> @@ -360,7 +362,7 @@ EXAMPLES
>  ::
>
>           51397 run_cnt
> -      40176203 cycles                                                 (83.05%)
> +      40176203 cycles          # 781.68 cycles per run                (83.05%)
>        42518139 instructions    #   1.06 insns per cycle               (83.39%)
>             123 llc_misses      #   2.89 LLC misses per million insns  (83.15%)
>
> diff --git a/tools/bpf/bpftool/prog.c b/tools/bpf/bpftool/prog.c
> index a9f730d407a9..8935508f955b 100644
> --- a/tools/bpf/bpftool/prog.c
> +++ b/tools/bpf/bpftool/prog.c
> @@ -2069,9 +2069,9 @@ struct profile_metric {
>         bool selected;
>
>         /* calculate ratios like instructions per cycle */
> -       const int ratio_metric; /* 0 for N/A, 1 for index 0 (cycles) */
> +       const int ratio_metric; /* 0 for run_cnt, 1 for index 0 (cycles) */
>         const char *ratio_desc;
> -       const float ratio_mul;
> +       const double ratio_mul;
>  } metrics[] = {
>         {
>                 .name = "cycles",
> @@ -2080,6 +2080,9 @@ struct profile_metric {
>                         .config = PERF_COUNT_HW_CPU_CYCLES,
>                         .exclude_user = 1,
>                 },
> +               .ratio_metric = 0,
> +               .ratio_desc = "cycles per run",
> +               .ratio_mul = 1.0,
>         },
>         {
>                 .name = "instructions",
> @@ -2256,17 +2259,18 @@ static void profile_print_readings_plain(void)
>         for (m = 0; m < ARRAY_SIZE(metrics); m++) {
>                 struct bpf_perf_event_value *val = &metrics[m].val;
>                 int r;
> +               __u64 ratio;
>
>                 if (!metrics[m].selected)
>                         continue;
>                 printf("%18llu %-20s", val->counter, metrics[m].name);
>
> -               r = metrics[m].ratio_metric - 1;
> -               if (r >= 0 && metrics[r].selected &&
> -                   metrics[r].val.counter > 0) {
> +               r = metrics[m].ratio_metric;
> +               /* r == 0 is a special case for run_cnt */
> +               ratio = r ? metrics[r - 1].val.counter : profile_total_count;
> +               if (metrics[m].ratio_desc && ratio) {

this ratio_desc-based thing looks suspect. We used to check .selected,
why did you change this?

And tbh, this whole ratio_metric would be much better done with enum,
where you can have -1 as "NO_METRIC", -2 as "RUN_CNT", 0 - cycles, 1 -
instructions, and so on.

then in definition of metrics array you can use explicit

[METRIC_CYCLES] = { .name = "cycles", ... },
[METRIC_INSNS] = { .name = "instructions", ..., .ratio_metric = METRIC_CYCLES }


makes everything consistent, explicit, easier to follow, wdyt?

pw-bot: cr


>                         printf("# %8.2f %-30s",
> -                              val->counter * metrics[m].ratio_mul /
> -                              metrics[r].val.counter,
> +                              val->counter * metrics[m].ratio_mul / ratio,
>                                metrics[m].ratio_desc);
>                 } else {
>                         printf("%-41s", "");
>
> ---
> base-commit: 23ff631b3b8b1452dfe933ee21f96321a9c5e209
> change-id: 20260827-bpftool_cyles_per_run-93169fcb30b7
>
> Best regards,
> --
> Mykyta Yatsenko <yatsenko@meta.com>
>

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH bpf-next] bpftool: Print average cycles per program run in profiler
  2026-08-28  0:30 ` Andrii Nakryiko
@ 2026-08-28 11:00   ` Mykyta Yatsenko
  2026-08-28 16:33     ` Andrii Nakryiko
  0 siblings, 1 reply; 7+ messages in thread
From: Mykyta Yatsenko @ 2026-08-28 11:00 UTC (permalink / raw)
  To: Andrii Nakryiko
  Cc: bpf, ast, andrii, daniel, kernel-team, eddyz87, memxor, qmo,
	Mykyta Yatsenko



On 8/28/26 1:30 AM, Andrii Nakryiko wrote:
> On Thu, Aug 27, 2026 at 8:29 AM Mykyta Yatsenko
> <mykyta.yatsenko5@gmail.com> wrote:
>>
>> From: Mykyta Yatsenko <yatsenko@meta.com>
>>
>> Total cycle counts are difficult to compare across workloads with
>> different run counts. Report cycles per run to expose the per-invocation
>> cost directly.
>>
>> Example:
>> sudo ./bpftool prog profile name mprog duration 15 cycles instructions
>>
>>             423256 run_cnt
>>          947413975 cycles              #  2238.39 cycles per run
>>          333965846 instructions        #     0.35 insns per cycle
>>
>> Signed-off-by: Mykyta Yatsenko <yatsenko@meta.com>
>> ---
>>  tools/bpf/bpftool/Documentation/bpftool-prog.rst |  6 ++++--
>>  tools/bpf/bpftool/prog.c                         | 18 +++++++++++-------
>>  2 files changed, 15 insertions(+), 9 deletions(-)
>>
>> diff --git a/tools/bpf/bpftool/Documentation/bpftool-prog.rst b/tools/bpf/bpftool/Documentation/bpftool-prog.rst
>> index 90fa2a48cc26..2280dc4492c0 100644
>> --- a/tools/bpf/bpftool/Documentation/bpftool-prog.rst
>> +++ b/tools/bpf/bpftool/Documentation/bpftool-prog.rst
>> @@ -217,7 +217,9 @@ bpftool prog run *PROG* data_in *FILE* [data_out *FILE* [data_size_out *L*]] [ct
>>  bpftool prog profile *PROG* [duration *DURATION*] *METRICs*
>>      Profile *METRICs* for bpf program *PROG* for *DURATION* seconds or until
>>      user hits <Ctrl+C>. *DURATION* is optional. If *DURATION* is not specified,
>> -    the profiling will run up to **UINT_MAX** seconds.
>> +    the profiling will run up to **UINT_MAX** seconds. When **cycles** is
>> +    selected, plain output also reports the average number of cycles per
>> +    program run.
>>
>>  bpftool prog help
>>      Print short help message.
>> @@ -360,7 +362,7 @@ EXAMPLES
>>  ::
>>
>>           51397 run_cnt
>> -      40176203 cycles                                                 (83.05%)
>> +      40176203 cycles          # 781.68 cycles per run                (83.05%)
>>        42518139 instructions    #   1.06 insns per cycle               (83.39%)
>>             123 llc_misses      #   2.89 LLC misses per million insns  (83.15%)
>>
>> diff --git a/tools/bpf/bpftool/prog.c b/tools/bpf/bpftool/prog.c
>> index a9f730d407a9..8935508f955b 100644
>> --- a/tools/bpf/bpftool/prog.c
>> +++ b/tools/bpf/bpftool/prog.c
>> @@ -2069,9 +2069,9 @@ struct profile_metric {
>>         bool selected;
>>
>>         /* calculate ratios like instructions per cycle */
>> -       const int ratio_metric; /* 0 for N/A, 1 for index 0 (cycles) */
>> +       const int ratio_metric; /* 0 for run_cnt, 1 for index 0 (cycles) */
>>         const char *ratio_desc;
>> -       const float ratio_mul;
>> +       const double ratio_mul;
>>  } metrics[] = {
>>         {
>>                 .name = "cycles",
>> @@ -2080,6 +2080,9 @@ struct profile_metric {
>>                         .config = PERF_COUNT_HW_CPU_CYCLES,
>>                         .exclude_user = 1,
>>                 },
>> +               .ratio_metric = 0,
>> +               .ratio_desc = "cycles per run",
>> +               .ratio_mul = 1.0,
>>         },
>>         {
>>                 .name = "instructions",
>> @@ -2256,17 +2259,18 @@ static void profile_print_readings_plain(void)
>>         for (m = 0; m < ARRAY_SIZE(metrics); m++) {
>>                 struct bpf_perf_event_value *val = &metrics[m].val;
>>                 int r;
>> +               __u64 ratio;
>>
>>                 if (!metrics[m].selected)
>>                         continue;
>>                 printf("%18llu %-20s", val->counter, metrics[m].name);
>>
>> -               r = metrics[m].ratio_metric - 1;
>> -               if (r >= 0 && metrics[r].selected &&
>> -                   metrics[r].val.counter > 0) {
>> +               r = metrics[m].ratio_metric;
>> +               /* r == 0 is a special case for run_cnt */
>> +               ratio = r ? metrics[r - 1].val.counter : profile_total_count;
>> +               if (metrics[m].ratio_desc && ratio) {
> 
> this ratio_desc-based thing looks suspect. We used to check .selected,
> why did you change this?

We checked .selected on the metrics[r] (denominator metric), with run_cnt,
it does not exist.
Checking ratio for zero, merges 2 checks into onet:
 - verify no division by zero
 - if ratio is not zero, that metric[r] has to have .selected == true, 
otherwise how did we bump it.
> 
> And tbh, this whole ratio_metric would be much better done with enum,
> where you can have -1 as "NO_METRIC", -2 as "RUN_CNT", 0 - cycles, 1 -
> instructions, and so on.

That'll do. But feels a bit awkward:
metrics[] = {
  ...
  {
    ...
    .ratio_metric = 1, /* But really mean METRIC_CYCLES which is index 0 */
  },
  {
    .ratio_metric = -1 /* But really mean METRIC_RUN_CNT which is -2 */
  }
}
The core difficulty here is that .ratio_metric default initializes with 0 and 
stands for NO_METRIC, then all indexes in .ratio_metric are shifted by one.
Alternatively we can explicitly set .ratio_metric for every element, but
that makes default initialized not safe (.ratio_metric == 0 means cycles, but
.ratio_desc is NULL).
> 
> then in definition of metrics array you can use explicit
> 
> [METRIC_CYCLES] = { .name = "cycles", ... },
> [METRIC_INSNS] = { .name = "instructions", ..., .ratio_metric = METRIC_CYCLES }
> 
> 
> makes everything consistent, explicit, easier to follow, wdyt?
> 
> pw-bot: cr
> 
> 
>>                         printf("# %8.2f %-30s",
>> -                              val->counter * metrics[m].ratio_mul /
>> -                              metrics[r].val.counter,
>> +                              val->counter * metrics[m].ratio_mul / ratio,
>>                                metrics[m].ratio_desc);
>>                 } else {
>>                         printf("%-41s", "");
>>
>> ---
>> base-commit: 23ff631b3b8b1452dfe933ee21f96321a9c5e209
>> change-id: 20260827-bpftool_cyles_per_run-93169fcb30b7
>>
>> Best regards,
>> --
>> Mykyta Yatsenko <yatsenko@meta.com>
>>


^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH bpf-next] bpftool: Print average cycles per program run in profiler
  2026-08-28 11:00   ` Mykyta Yatsenko
@ 2026-08-28 16:33     ` Andrii Nakryiko
  2026-08-28 17:11       ` Mykyta Yatsenko
  0 siblings, 1 reply; 7+ messages in thread
From: Andrii Nakryiko @ 2026-08-28 16:33 UTC (permalink / raw)
  To: Mykyta Yatsenko
  Cc: bpf, ast, andrii, daniel, kernel-team, eddyz87, memxor, qmo,
	Mykyta Yatsenko

On Fri, Aug 28, 2026 at 4:01 AM Mykyta Yatsenko
<mykyta.yatsenko5@gmail.com> wrote:
>
>
>
> On 8/28/26 1:30 AM, Andrii Nakryiko wrote:
> > On Thu, Aug 27, 2026 at 8:29 AM Mykyta Yatsenko
> > <mykyta.yatsenko5@gmail.com> wrote:
> >>
> >> From: Mykyta Yatsenko <yatsenko@meta.com>
> >>
> >> Total cycle counts are difficult to compare across workloads with
> >> different run counts. Report cycles per run to expose the per-invocation
> >> cost directly.
> >>
> >> Example:
> >> sudo ./bpftool prog profile name mprog duration 15 cycles instructions
> >>
> >>             423256 run_cnt
> >>          947413975 cycles              #  2238.39 cycles per run
> >>          333965846 instructions        #     0.35 insns per cycle
> >>
> >> Signed-off-by: Mykyta Yatsenko <yatsenko@meta.com>
> >> ---
> >>  tools/bpf/bpftool/Documentation/bpftool-prog.rst |  6 ++++--
> >>  tools/bpf/bpftool/prog.c                         | 18 +++++++++++-------
> >>  2 files changed, 15 insertions(+), 9 deletions(-)
> >>
> >> diff --git a/tools/bpf/bpftool/Documentation/bpftool-prog.rst b/tools/bpf/bpftool/Documentation/bpftool-prog.rst
> >> index 90fa2a48cc26..2280dc4492c0 100644
> >> --- a/tools/bpf/bpftool/Documentation/bpftool-prog.rst
> >> +++ b/tools/bpf/bpftool/Documentation/bpftool-prog.rst
> >> @@ -217,7 +217,9 @@ bpftool prog run *PROG* data_in *FILE* [data_out *FILE* [data_size_out *L*]] [ct
> >>  bpftool prog profile *PROG* [duration *DURATION*] *METRICs*
> >>      Profile *METRICs* for bpf program *PROG* for *DURATION* seconds or until
> >>      user hits <Ctrl+C>. *DURATION* is optional. If *DURATION* is not specified,
> >> -    the profiling will run up to **UINT_MAX** seconds.
> >> +    the profiling will run up to **UINT_MAX** seconds. When **cycles** is
> >> +    selected, plain output also reports the average number of cycles per
> >> +    program run.
> >>
> >>  bpftool prog help
> >>      Print short help message.
> >> @@ -360,7 +362,7 @@ EXAMPLES
> >>  ::
> >>
> >>           51397 run_cnt
> >> -      40176203 cycles                                                 (83.05%)
> >> +      40176203 cycles          # 781.68 cycles per run                (83.05%)
> >>        42518139 instructions    #   1.06 insns per cycle               (83.39%)
> >>             123 llc_misses      #   2.89 LLC misses per million insns  (83.15%)
> >>
> >> diff --git a/tools/bpf/bpftool/prog.c b/tools/bpf/bpftool/prog.c
> >> index a9f730d407a9..8935508f955b 100644
> >> --- a/tools/bpf/bpftool/prog.c
> >> +++ b/tools/bpf/bpftool/prog.c
> >> @@ -2069,9 +2069,9 @@ struct profile_metric {
> >>         bool selected;
> >>
> >>         /* calculate ratios like instructions per cycle */
> >> -       const int ratio_metric; /* 0 for N/A, 1 for index 0 (cycles) */
> >> +       const int ratio_metric; /* 0 for run_cnt, 1 for index 0 (cycles) */
> >>         const char *ratio_desc;
> >> -       const float ratio_mul;
> >> +       const double ratio_mul;
> >>  } metrics[] = {
> >>         {
> >>                 .name = "cycles",
> >> @@ -2080,6 +2080,9 @@ struct profile_metric {
> >>                         .config = PERF_COUNT_HW_CPU_CYCLES,
> >>                         .exclude_user = 1,
> >>                 },
> >> +               .ratio_metric = 0,
> >> +               .ratio_desc = "cycles per run",
> >> +               .ratio_mul = 1.0,
> >>         },
> >>         {
> >>                 .name = "instructions",
> >> @@ -2256,17 +2259,18 @@ static void profile_print_readings_plain(void)
> >>         for (m = 0; m < ARRAY_SIZE(metrics); m++) {
> >>                 struct bpf_perf_event_value *val = &metrics[m].val;
> >>                 int r;
> >> +               __u64 ratio;
> >>
> >>                 if (!metrics[m].selected)
> >>                         continue;
> >>                 printf("%18llu %-20s", val->counter, metrics[m].name);
> >>
> >> -               r = metrics[m].ratio_metric - 1;
> >> -               if (r >= 0 && metrics[r].selected &&
> >> -                   metrics[r].val.counter > 0) {
> >> +               r = metrics[m].ratio_metric;
> >> +               /* r == 0 is a special case for run_cnt */
> >> +               ratio = r ? metrics[r - 1].val.counter : profile_total_count;
> >> +               if (metrics[m].ratio_desc && ratio) {
> >
> > this ratio_desc-based thing looks suspect. We used to check .selected,
> > why did you change this?
>
> We checked .selected on the metrics[r] (denominator metric), with run_cnt,
> it does not exist.
> Checking ratio for zero, merges 2 checks into onet:
>  - verify no division by zero
>  - if ratio is not zero, that metric[r] has to have .selected == true,
> otherwise how did we bump it.

ok, makes sense, thanks for explaining!

> >
> > And tbh, this whole ratio_metric would be much better done with enum,
> > where you can have -1 as "NO_METRIC", -2 as "RUN_CNT", 0 - cycles, 1 -
> > instructions, and so on.
>
> That'll do. But feels a bit awkward:
> metrics[] = {
>   ...
>   {
>     ...
>     .ratio_metric = 1, /* But really mean METRIC_CYCLES which is index 0 */
>   },
>   {
>     .ratio_metric = -1 /* But really mean METRIC_RUN_CNT which is -2 */
>   }
> }
> The core difficulty here is that .ratio_metric default initializes with 0 and
> stands for NO_METRIC, then all indexes in .ratio_metric are shifted by one.
> Alternatively we can explicitly set .ratio_metric for every element, but
> that makes default initialized not safe (.ratio_metric == 0 means cycles, but
> .ratio_desc is NULL).

I personally think that explicit .ratio_metric = METRIC_NO_METRIC or
something like that is just fine to do and not a problem.

But if that's a problem, I'd still do enum, just make zero a "NO
METRIC", and shift everything else by one. So basically what we have
today, but explicitly named (and with small comment next to enum we
can explain that shift-by-one convention).

Your choice, my point is that plain numbers make it hard to follow
what's going on, and really here we have a limited set of explicitly
connected things, so enum is the way, IMO.


> >
> > then in definition of metrics array you can use explicit
> >
> > [METRIC_CYCLES] = { .name = "cycles", ... },
> > [METRIC_INSNS] = { .name = "instructions", ..., .ratio_metric = METRIC_CYCLES }
> >
> >
> > makes everything consistent, explicit, easier to follow, wdyt?
> >
> > pw-bot: cr
> >
> >
> >>                         printf("# %8.2f %-30s",
> >> -                              val->counter * metrics[m].ratio_mul /
> >> -                              metrics[r].val.counter,
> >> +                              val->counter * metrics[m].ratio_mul / ratio,
> >>                                metrics[m].ratio_desc);
> >>                 } else {
> >>                         printf("%-41s", "");
> >>
> >> ---
> >> base-commit: 23ff631b3b8b1452dfe933ee21f96321a9c5e209
> >> change-id: 20260827-bpftool_cyles_per_run-93169fcb30b7
> >>
> >> Best regards,
> >> --
> >> Mykyta Yatsenko <yatsenko@meta.com>
> >>
>

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH bpf-next] bpftool: Print average cycles per program run in profiler
  2026-08-28 16:33     ` Andrii Nakryiko
@ 2026-08-28 17:11       ` Mykyta Yatsenko
  0 siblings, 0 replies; 7+ messages in thread
From: Mykyta Yatsenko @ 2026-08-28 17:11 UTC (permalink / raw)
  To: Andrii Nakryiko
  Cc: bpf, ast, andrii, daniel, kernel-team, eddyz87, memxor, qmo,
	Mykyta Yatsenko



On 8/28/26 5:33 PM, Andrii Nakryiko wrote:
> On Fri, Aug 28, 2026 at 4:01 AM Mykyta Yatsenko
> <mykyta.yatsenko5@gmail.com> wrote:
>>
>>
>>
>> On 8/28/26 1:30 AM, Andrii Nakryiko wrote:
>>> On Thu, Aug 27, 2026 at 8:29 AM Mykyta Yatsenko
>>> <mykyta.yatsenko5@gmail.com> wrote:
>>>>
>>>> From: Mykyta Yatsenko <yatsenko@meta.com>
>>>>
>>>> Total cycle counts are difficult to compare across workloads with
>>>> different run counts. Report cycles per run to expose the per-invocation
>>>> cost directly.
>>>>
>>>> Example:
>>>> sudo ./bpftool prog profile name mprog duration 15 cycles instructions
>>>>
>>>>             423256 run_cnt
>>>>          947413975 cycles              #  2238.39 cycles per run
>>>>          333965846 instructions        #     0.35 insns per cycle
>>>>
>>>> Signed-off-by: Mykyta Yatsenko <yatsenko@meta.com>
>>>> ---
>>>>  tools/bpf/bpftool/Documentation/bpftool-prog.rst |  6 ++++--
>>>>  tools/bpf/bpftool/prog.c                         | 18 +++++++++++-------
>>>>  2 files changed, 15 insertions(+), 9 deletions(-)
>>>>
>>>> diff --git a/tools/bpf/bpftool/Documentation/bpftool-prog.rst b/tools/bpf/bpftool/Documentation/bpftool-prog.rst
>>>> index 90fa2a48cc26..2280dc4492c0 100644
>>>> --- a/tools/bpf/bpftool/Documentation/bpftool-prog.rst
>>>> +++ b/tools/bpf/bpftool/Documentation/bpftool-prog.rst
>>>> @@ -217,7 +217,9 @@ bpftool prog run *PROG* data_in *FILE* [data_out *FILE* [data_size_out *L*]] [ct
>>>>  bpftool prog profile *PROG* [duration *DURATION*] *METRICs*
>>>>      Profile *METRICs* for bpf program *PROG* for *DURATION* seconds or until
>>>>      user hits <Ctrl+C>. *DURATION* is optional. If *DURATION* is not specified,
>>>> -    the profiling will run up to **UINT_MAX** seconds.
>>>> +    the profiling will run up to **UINT_MAX** seconds. When **cycles** is
>>>> +    selected, plain output also reports the average number of cycles per
>>>> +    program run.
>>>>
>>>>  bpftool prog help
>>>>      Print short help message.
>>>> @@ -360,7 +362,7 @@ EXAMPLES
>>>>  ::
>>>>
>>>>           51397 run_cnt
>>>> -      40176203 cycles                                                 (83.05%)
>>>> +      40176203 cycles          # 781.68 cycles per run                (83.05%)
>>>>        42518139 instructions    #   1.06 insns per cycle               (83.39%)
>>>>             123 llc_misses      #   2.89 LLC misses per million insns  (83.15%)
>>>>
>>>> diff --git a/tools/bpf/bpftool/prog.c b/tools/bpf/bpftool/prog.c
>>>> index a9f730d407a9..8935508f955b 100644
>>>> --- a/tools/bpf/bpftool/prog.c
>>>> +++ b/tools/bpf/bpftool/prog.c
>>>> @@ -2069,9 +2069,9 @@ struct profile_metric {
>>>>         bool selected;
>>>>
>>>>         /* calculate ratios like instructions per cycle */
>>>> -       const int ratio_metric; /* 0 for N/A, 1 for index 0 (cycles) */
>>>> +       const int ratio_metric; /* 0 for run_cnt, 1 for index 0 (cycles) */
>>>>         const char *ratio_desc;
>>>> -       const float ratio_mul;
>>>> +       const double ratio_mul;
>>>>  } metrics[] = {
>>>>         {
>>>>                 .name = "cycles",
>>>> @@ -2080,6 +2080,9 @@ struct profile_metric {
>>>>                         .config = PERF_COUNT_HW_CPU_CYCLES,
>>>>                         .exclude_user = 1,
>>>>                 },
>>>> +               .ratio_metric = 0,
>>>> +               .ratio_desc = "cycles per run",
>>>> +               .ratio_mul = 1.0,
>>>>         },
>>>>         {
>>>>                 .name = "instructions",
>>>> @@ -2256,17 +2259,18 @@ static void profile_print_readings_plain(void)
>>>>         for (m = 0; m < ARRAY_SIZE(metrics); m++) {
>>>>                 struct bpf_perf_event_value *val = &metrics[m].val;
>>>>                 int r;
>>>> +               __u64 ratio;
>>>>
>>>>                 if (!metrics[m].selected)
>>>>                         continue;
>>>>                 printf("%18llu %-20s", val->counter, metrics[m].name);
>>>>
>>>> -               r = metrics[m].ratio_metric - 1;
>>>> -               if (r >= 0 && metrics[r].selected &&
>>>> -                   metrics[r].val.counter > 0) {
>>>> +               r = metrics[m].ratio_metric;
>>>> +               /* r == 0 is a special case for run_cnt */
>>>> +               ratio = r ? metrics[r - 1].val.counter : profile_total_count;
>>>> +               if (metrics[m].ratio_desc && ratio) {
>>>
>>> this ratio_desc-based thing looks suspect. We used to check .selected,
>>> why did you change this?
>>
>> We checked .selected on the metrics[r] (denominator metric), with run_cnt,
>> it does not exist.
>> Checking ratio for zero, merges 2 checks into onet:
>>  - verify no division by zero
>>  - if ratio is not zero, that metric[r] has to have .selected == true,
>> otherwise how did we bump it.
> 
> ok, makes sense, thanks for explaining!
> 
>>>
>>> And tbh, this whole ratio_metric would be much better done with enum,
>>> where you can have -1 as "NO_METRIC", -2 as "RUN_CNT", 0 - cycles, 1 -
>>> instructions, and so on.
>>
>> That'll do. But feels a bit awkward:
>> metrics[] = {
>>   ...
>>   {
>>     ...
>>     .ratio_metric = 1, /* But really mean METRIC_CYCLES which is index 0 */
>>   },
>>   {
>>     .ratio_metric = -1 /* But really mean METRIC_RUN_CNT which is -2 */
>>   }
>> }
>> The core difficulty here is that .ratio_metric default initializes with 0 and
>> stands for NO_METRIC, then all indexes in .ratio_metric are shifted by one.
>> Alternatively we can explicitly set .ratio_metric for every element, but
>> that makes default initialized not safe (.ratio_metric == 0 means cycles, but
>> .ratio_desc is NULL).
> 
> I personally think that explicit .ratio_metric = METRIC_NO_METRIC or
> something like that is just fine to do and not a problem.
> 
> But if that's a problem, I'd still do enum, just make zero a "NO
> METRIC", and shift everything else by one. So basically what we have
> today, but explicitly named (and with small comment next to enum we
> can explain that shift-by-one convention).
> 
> Your choice, my point is that plain numbers make it hard to follow
> what's going on, and really here we have a limited set of explicitly
> connected things, so enum is the way, IMO.
> 

I'll respin with enum, thanks to taking a look.
> 
>>>
>>> then in definition of metrics array you can use explicit
>>>
>>> [METRIC_CYCLES] = { .name = "cycles", ... },
>>> [METRIC_INSNS] = { .name = "instructions", ..., .ratio_metric = METRIC_CYCLES }
>>>
>>>
>>> makes everything consistent, explicit, easier to follow, wdyt?
>>>
>>> pw-bot: cr
>>>
>>>
>>>>                         printf("# %8.2f %-30s",
>>>> -                              val->counter * metrics[m].ratio_mul /
>>>> -                              metrics[r].val.counter,
>>>> +                              val->counter * metrics[m].ratio_mul / ratio,
>>>>                                metrics[m].ratio_desc);
>>>>                 } else {
>>>>                         printf("%-41s", "");
>>>>
>>>> ---
>>>> base-commit: 23ff631b3b8b1452dfe933ee21f96321a9c5e209
>>>> change-id: 20260827-bpftool_cyles_per_run-93169fcb30b7
>>>>
>>>> Best regards,
>>>> --
>>>> Mykyta Yatsenko <yatsenko@meta.com>
>>>>
>>


^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH bpf-next] bpftool: Print average cycles per program run in profiler
  2026-08-27 15:29 [PATCH bpf-next] bpftool: Print average cycles per program run in profiler Mykyta Yatsenko
  2026-08-27 16:22 ` bot+bpf-ci
  2026-08-28  0:30 ` Andrii Nakryiko
@ 2026-08-28 18:46 ` Suchit Karunakaran
  2 siblings, 0 replies; 7+ messages in thread
From: Suchit Karunakaran @ 2026-08-28 18:46 UTC (permalink / raw)
  To: Mykyta Yatsenko, bpf, ast, andrii, daniel, kernel-team, eddyz87,
	memxor, qmo
  Cc: Mykyta Yatsenko

Hi,

On 8/27/26 8:59 PM, Mykyta Yatsenko wrote:
> From: Mykyta Yatsenko <yatsenko@meta.com>
>
> Total cycle counts are difficult to compare across workloads with
> different run counts. Report cycles per run to expose the per-invocation
> cost directly.
>
> Example:
> sudo ./bpftool prog profile name mprog duration 15 cycles instructions
>
>              423256 run_cnt
>           947413975 cycles              #  2238.39 cycles per run
>           333965846 instructions        #     0.35 insns per cycle
>
> Signed-off-by: Mykyta Yatsenko <yatsenko@meta.com>
> ---
>   tools/bpf/bpftool/Documentation/bpftool-prog.rst |  6 ++++--
>   tools/bpf/bpftool/prog.c                         | 18 +++++++++++-------
>   2 files changed, 15 insertions(+), 9 deletions(-)
>
> diff --git a/tools/bpf/bpftool/Documentation/bpftool-prog.rst b/tools/bpf/bpftool/Documentation/bpftool-prog.rst
> index 90fa2a48cc26..2280dc4492c0 100644
> --- a/tools/bpf/bpftool/Documentation/bpftool-prog.rst
> +++ b/tools/bpf/bpftool/Documentation/bpftool-prog.rst
> @@ -217,7 +217,9 @@ bpftool prog run *PROG* data_in *FILE* [data_out *FILE* [data_size_out *L*]] [ct
>   bpftool prog profile *PROG* [duration *DURATION*] *METRICs*
>       Profile *METRICs* for bpf program *PROG* for *DURATION* seconds or until
>       user hits <Ctrl+C>. *DURATION* is optional. If *DURATION* is not specified,
> -    the profiling will run up to **UINT_MAX** seconds.
> +    the profiling will run up to **UINT_MAX** seconds. When **cycles** is
> +    selected, plain output also reports the average number of cycles per
> +    program run.
>   
>   bpftool prog help
>       Print short help message.
> @@ -360,7 +362,7 @@ EXAMPLES
>   ::
>   
>            51397 run_cnt
> -      40176203 cycles                                                 (83.05%)
> +      40176203 cycles          # 781.68 cycles per run                (83.05%)
>         42518139 instructions    #   1.06 insns per cycle               (83.39%)
>              123 llc_misses      #   2.89 LLC misses per million insns  (83.15%)
>   
> diff --git a/tools/bpf/bpftool/prog.c b/tools/bpf/bpftool/prog.c
> index a9f730d407a9..8935508f955b 100644
> --- a/tools/bpf/bpftool/prog.c
> +++ b/tools/bpf/bpftool/prog.c
> @@ -2069,9 +2069,9 @@ struct profile_metric {
>   	bool selected;
>   
>   	/* calculate ratios like instructions per cycle */
> -	const int ratio_metric; /* 0 for N/A, 1 for index 0 (cycles) */
> +	const int ratio_metric; /* 0 for run_cnt, 1 for index 0 (cycles) */
>   	const char *ratio_desc;
> -	const float ratio_mul;
> +	const double ratio_mul;
>   } metrics[] = {
>   	{
>   		.name = "cycles",
> @@ -2080,6 +2080,9 @@ struct profile_metric {
>   			.config = PERF_COUNT_HW_CPU_CYCLES,
>   			.exclude_user = 1,
>   		},
> +		.ratio_metric = 0,
> +		.ratio_desc = "cycles per run",
> +		.ratio_mul = 1.0,
>   	},
>   	{
>   		.name = "instructions",
> @@ -2256,17 +2259,18 @@ static void profile_print_readings_plain(void)
>   	for (m = 0; m < ARRAY_SIZE(metrics); m++) {
>   		struct bpf_perf_event_value *val = &metrics[m].val;
>   		int r;
> +		__u64 ratio;
>   
>   		if (!metrics[m].selected)
>   			continue;
>   		printf("%18llu %-20s", val->counter, metrics[m].name);
>   
> -		r = metrics[m].ratio_metric - 1;
> -		if (r >= 0 && metrics[r].selected &&
> -		    metrics[r].val.counter > 0) {
> +		r = metrics[m].ratio_metric;
> +		/* r == 0 is a special case for run_cnt */
> +		ratio = r ? metrics[r - 1].val.counter : profile_total_count;
> +		if (metrics[m].ratio_desc && ratio) {
>   			printf("# %8.2f %-30s",
> -			       val->counter * metrics[m].ratio_mul /
> -			       metrics[r].val.counter,
> +			       val->counter * metrics[m].ratio_mul / ratio,
>   			       metrics[m].ratio_desc);
>   		} else {
>   			printf("%-41s", "");
>
> ---
> base-commit: 23ff631b3b8b1452dfe933ee21f96321a9c5e209
> change-id: 20260827-bpftool_cyles_per_run-93169fcb30b7
>
> Best regards,
> --
> Mykyta Yatsenko <yatsenko@meta.com>
>
I think val->counter be normalized for counter multiplexing before 
dividing by profile_total_count? val->counter contains only the cycles 
observed while the event was actively scheduled on the PMU, whereas 
profile_total_count reflects all program invocations across the full 
duration. When val->running < val->enabled, dividing raw val->counter by 
profile_total_count undercounts the true cycles per run. Should we apply 
scaling before computing the ratio like below?
if (val->running > 0) {
     double scaled = (double)val->counter * val->enabled / val->running;
     double cycles_per_run = scaled / profile_total_count;
}

We can actually see this in the example:

51397 run_cnt
40176203 cycles          # 781.68 cycles per run (83.05%)
Here, 781.68 uses the raw unscaled count during the 83.05% runtime 
window. If we scale it to 100% estimated coverage, the true cost is 
closer to ~941.22 cycles per run.


^ permalink raw reply	[flat|nested] 7+ messages in thread

end of thread, other threads:[~2026-08-28 18:46 UTC | newest]

Thread overview: 7+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-27 15:29 [PATCH bpf-next] bpftool: Print average cycles per program run in profiler Mykyta Yatsenko
2026-08-27 16:22 ` bot+bpf-ci
2026-08-28  0:30 ` Andrii Nakryiko
2026-08-28 11:00   ` Mykyta Yatsenko
2026-08-28 16:33     ` Andrii Nakryiko
2026-08-28 17:11       ` Mykyta Yatsenko
2026-08-28 18:46 ` Suchit Karunakaran

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox