dev.dpdk.org archive mirror
 help / color / mirror / Atom feed
From: "Morten Brørup" <mb@smartsharesystems.com>
To: <dev@dpdk.org>, "Jerin Jacob" <jerinj@marvell.com>,
	"Kiran Kumar K" <kirankumark@marvell.com>,
	"Nithin Dabilpuram" <ndabilpuram@marvell.com>,
	"Zhirun Yan" <yanzhirun_163@163.com>,
	"Saeed Bishara" <saeed.bishara.os@gmail.com>
Cc: "Robin Jarry" <rjarry@redhat.com>
Subject: RE: [PATCH v9] graph: add optional profiling stats
Date: Thu, 13 Aug 2026 08:56:14 +0200	[thread overview]
Message-ID: <98CBD80474FA8B44BF855DF32C47DC35F659DD@smartserver.smartshare.dk> (raw)
In-Reply-To: <20260703154357.3068739-1-mb@smartsharesystems.com>

PING for review.

(And obviously, the release notes needs to be for 26.11, not 26.07.)

> -----Original Message-----
> From: Morten Brørup [mailto:mb@smartsharesystems.com]
> Sent: Friday, 3 July 2026 17.44
> 
> graph: add optional profiling stats
> 
> Added graph node profiling stats, build time configurable by enabling
> RTE_GRAPH_PROFILE in rte_config.h.
> 
> Signed-off-by: Morten Brørup <mb@smartsharesystems.com>
> ---
> v9:
> * Fixed comment still mentioning 32 objects.
> * Moved sample size array outside loop. (AI)
> * Added release note. (AI)
> v8:
> * Added static const array as local variable, instead of indexing
> directly
>   into const array. (AI)
>   This also eliminates the space required between "} [idx];" weirdness.
> * Added build time configurable RTE_GRAPH_PROFILE_BURST_SIZE to replace
>   the hardcoded burst size of 32. (AI)
> v7:
> * Use RTE_DIM() in histogram for loop.
> * Added static_assert for histogram index values.
> * Minor details to please checkpatch.
>   Although I disagree with requiring a space when indexing into
>   a constant array "(const type []){values} [idx];",
>   I have changed the code to comply.
> v6:
> * Consolidate the four histogram entries into one array. (Saeed
> Bishara)
> * Sample at 32 objs instead of a half burst. (Saeed Bishara)
> * Moved stats to different location in rte_node structure. (Jerin)
> * Minor details to please checkpatch.
> v5:
> * Added stats for a half burst and a full burst.
> v4:
> * Added documentation. (AI)
> * Added more comments. (AI)
> * Improved dump. (AI)
> * Debug shows both cycles/call and cycles/obj.
> v3:
> * Debug shows cycles/obj instead of cycles/call.
> * Fixed missing --in-reply-to.
> v2:
> * Fixed indentation.
> ---
>  config/rte_config.h                    |  2 ++
>  doc/guides/prog_guide/graph_lib.rst    |  3 ++
>  doc/guides/rel_notes/release_26_07.rst |  7 ++++
>  lib/graph/graph_debug.c                | 50 ++++++++++++++++++++++++++
>  lib/graph/node.c                       |  2 ++
>  lib/graph/rte_graph_worker_common.h    | 32 +++++++++++++++--
>  6 files changed, 93 insertions(+), 3 deletions(-)
> 
> diff --git a/config/rte_config.h b/config/rte_config.h
> index 0447cdf2ad..7703a6325b 100644
> --- a/config/rte_config.h
> +++ b/config/rte_config.h
> @@ -106,6 +106,8 @@
>  /* rte_graph defines */
>  #define RTE_GRAPH_BURST_SIZE 256
>  #define RTE_LIBRTE_GRAPH_STATS 1
> +/* RTE_GRAPH_PROFILE is not set */
> +#define RTE_GRAPH_PROFILE_BURST_SIZE 32
> 
>  /****** driver defines ********/
> 
> diff --git a/doc/guides/prog_guide/graph_lib.rst
> b/doc/guides/prog_guide/graph_lib.rst
> index 8dd49c19d2..4311c37cc2 100644
> --- a/doc/guides/prog_guide/graph_lib.rst
> +++ b/doc/guides/prog_guide/graph_lib.rst
> @@ -49,6 +49,9 @@ Performance tuning parameters
>    RTE_GRAPH_BURST_SIZE config option.
>    The testing shows, on x86 and arm64 servers, The sweet spot is 256
> burst
>    size. While on arm64 embedded SoCs, it is either 64 or 128.
> +- Enable the ``RTE_GRAPH_PROFILE`` config option for more profiling
> details.
> +  Set the ``RTE_GRAPH_PROFILE_BURST_SIZE`` config option to sample a
> specific
> +  burst size.
>  - Disable node statistics (using ``RTE_LIBRTE_GRAPH_STATS`` config
> option)
>    if not needed.
> 
> diff --git a/doc/guides/rel_notes/release_26_07.rst
> b/doc/guides/rel_notes/release_26_07.rst
> index 8b1bdada1a..cb7d48eeac 100644
> --- a/doc/guides/rel_notes/release_26_07.rst
> +++ b/doc/guides/rel_notes/release_26_07.rst
> @@ -83,6 +83,13 @@ New Features
>    * The size of the ``struct rte_mempool_cache`` was kept
>      for API/ABI compatibility purposes.
> 
> +* **Added optional graph profiling statistics.**
> +
> +  Added build-time configurable graph node profiling statistics via
> +  ``RTE_GRAPH_PROFILE`` in ``rte_config.h``. When enabled, tracks
> cycles
> +  spent processing bursts of 0, 1, ``RTE_GRAPH_PROFILE_BURST_SIZE``,
> +  and ``RTE_GRAPH_BURST_SIZE`` objects per node.
> +
>  * **Added RISC-V vector paths.**
> 
>    * Increased the default SIMD bitwidth to allow using the vector
> extension.
> diff --git a/lib/graph/graph_debug.c b/lib/graph/graph_debug.c
> index e3b8cccdc1..b999fa4140 100644
> --- a/lib/graph/graph_debug.c
> +++ b/lib/graph/graph_debug.c
> @@ -93,6 +93,56 @@ rte_graph_obj_dump(FILE *f, struct rte_graph *g,
> bool all)
>  				n->dispatch.total_sched_fail);
>  		}
>  		fprintf(f, "       total_calls=%" PRId64 "\n", n-
> >total_calls);
> +		if (rte_graph_has_stats_feature())
> +			fprintf(f, "       total_cycles=%" PRIu64 ", avg
> cycles/call=%.1f\n",
> +					n->total_cycles,
> +					n->total_calls == 0 ? 0.0 :
> +					(double)n->total_cycles / (double)n-
> >total_calls);
> +#ifdef RTE_GRAPH_PROFILE
> +		static const uint16_t profile_sample_sizes[] = {
> +				0, 1, RTE_GRAPH_PROFILE_BURST_SIZE,
> RTE_GRAPH_BURST_SIZE};
> +		static_assert(RTE_DIM(profile_sample_sizes) == RTE_DIM(n-
> >usage_stats),
> +				"usage_stats array size mismatch");
> +		int64_t calls_other = n->total_calls;
> +		int64_t cycles_other = n->total_cycles;
> +		int64_t objs_other = n->total_objs;
> +		for (int idx = 0; idx < RTE_DIM(n->usage_stats) + 1; idx++)
> {
> +			uint64_t calls;
> +			uint64_t cycles;
> +			double objs_per_call;
> +			if (idx < RTE_DIM(n->usage_stats)) {
> +				uint16_t idx_objs = profile_sample_sizes[idx];
> +				fprintf(f, "       objs[%u]\n", idx_objs);
> +				calls = n->usage_stats[idx].calls;
> +				cycles = n->usage_stats[idx].cycles;
> +				objs_per_call = (double)idx_objs;
> +				calls_other -= calls;
> +				cycles_other -= cycles;
> +				objs_other -= idx_objs * calls;
> +			} else {
> +				fprintf(f, "       objs[other]\n");
> +				if (calls_other > 0 && cycles_other > 0 &&
> objs_other > 0) {
> +					calls = calls_other;
> +					cycles = cycles_other;
> +					objs_per_call = (double)objs_other /
> (double)calls_other;
> +					fprintf(f, "         avg
> objs/call=%.1f\n", objs_per_call);
> +				} else {
> +					calls = 0;
> +					cycles = 0;
> +					objs_per_call = 0.0;
> +				}
> +			}
> +			fprintf(f, "         calls=%" PRIu64, calls);
> +			if (calls != 0)
> +				fprintf(f, ", cycles=%" PRIu64 ", avg
> cycles/call=%.1f",
> +						cycles,
> +						(double)cycles / (double)calls);
> +			if (calls != 0 && objs_per_call != 0.0)
> +				fprintf(f, ", avg cycles/obj=%.1f",
> +						(double)cycles / (double)calls /
> objs_per_call);
> +			fprintf(f, "\n");
> +		}
> +#endif
>  		for (i = 0; i < n->nb_edges; i++)
>  			fprintf(f, "          edge[%d] <%s>\n", i,
>  				n->nodes[i]->name);
> diff --git a/lib/graph/node.c b/lib/graph/node.c
> index 1fce3e6632..19b38881ae 100644
> --- a/lib/graph/node.c
> +++ b/lib/graph/node.c
> @@ -110,10 +110,12 @@ __rte_node_register(const struct
> rte_node_register *reg)
>  	rte_edge_t i;
>  	size_t sz;
> 
> +#ifndef RTE_GRAPH_PROFILE
>  	/* Limit Node specific metadata to one cacheline on 64B CL
> machine */
>  	RTE_BUILD_BUG_ON((offsetof(struct rte_node, nodes) -
>  			  offsetof(struct rte_node, ctx)) !=
>  			 RTE_CACHE_LINE_MIN_SIZE);
> +#endif
> 
>  	graph_spinlock_lock();
> 
> diff --git a/lib/graph/rte_graph_worker_common.h
> b/lib/graph/rte_graph_worker_common.h
> index 4ab53a533e..00e8f5859d 100644
> --- a/lib/graph/rte_graph_worker_common.h
> +++ b/lib/graph/rte_graph_worker_common.h
> @@ -121,6 +121,17 @@ struct __rte_cache_aligned rte_node {
>  	rte_graph_off_t xstat_off; /**< Offset to xstat counters. */
> 
>  	/** Fast path area cache line 2. */
> +#ifdef RTE_GRAPH_PROFILE
> +	/**
> +	 * Usage when this node processed 0, 1,
> RTE_GRAPH_PROFILE_BURST_SIZE,
> +	 * or RTE_GRAPH_BURST_SIZE objects.
> +	 */
> +	struct __rte_cache_aligned {
> +		uint64_t calls;     /**< Calls done. */
> +		uint64_t cycles;    /**< Cycles spent. */
> +	} usage_stats[4];
> +	/** Fast path area cache line 3. */
> +#endif
>  	__extension__ struct __rte_cache_aligned {
>  #define RTE_NODE_CTX_SZ 16
>  		union {
> @@ -148,8 +159,10 @@ struct __rte_cache_aligned rte_node {
>  	};
>  };
> 
> +#ifndef RTE_GRAPH_PROFILE
>  static_assert(offsetof(struct rte_node, nodes) - offsetof(struct
> rte_node, ctx)
>  	== RTE_CACHE_LINE_MIN_SIZE, "rte_node fast path area must fit in
> 64 bytes");
> +#endif
> 
>  /**
>   * @internal
> @@ -197,7 +210,7 @@ void __rte_node_stream_alloc_size(struct rte_graph
> *graph,
>  static __rte_always_inline void
>  __rte_node_process(struct rte_graph *graph, struct rte_node *node)
>  {
> -	uint64_t start;
> +	uint64_t cycles;
>  	uint16_t rc;
>  	void **objs;
> 
> @@ -206,11 +219,24 @@ __rte_node_process(struct rte_graph *graph,
> struct rte_node *node)
>  	rte_prefetch0(objs);
> 
>  	if (rte_graph_has_stats_feature()) {
> -		start = rte_rdtsc();
> +		cycles = -rte_rdtsc();
>  		rc = node->process(graph, node, objs, node->idx);
> -		node->total_cycles += rte_rdtsc() - start;
> +		cycles += rte_rdtsc();
> +		node->total_cycles += cycles;
>  		node->total_calls++;
>  		node->total_objs += rc;
> +#ifdef RTE_GRAPH_PROFILE
> +		if (rc <= 1) {
> +			node->usage_stats[rc].calls++;
> +			node->usage_stats[rc].cycles += cycles;
> +		} else if (rc == RTE_GRAPH_PROFILE_BURST_SIZE) {
> +			node->usage_stats[2].calls++;
> +			node->usage_stats[2].cycles += cycles;
> +		} else if (rc == RTE_GRAPH_BURST_SIZE) {
> +			node->usage_stats[3].calls++;
> +			node->usage_stats[3].cycles += cycles;
> +		}
> +#endif
>  	} else {
>  		node->process(graph, node, objs, node->idx);
>  	}
> --
> 2.43.0


      reply	other threads:[~2026-08-13  6:56 UTC|newest]

Thread overview: 20+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-06-19 20:20 [PATCH] graph: add optional profiling stats Morten Brørup
2026-06-19 20:56 ` [PATCH v3] " Morten Brørup
2026-06-21 17:55 ` [PATCH v4] " Morten Brørup
2026-06-21 18:41 ` [PATCH v5] " Morten Brørup
2026-06-23  5:13   ` Jerin Jacob
2026-06-23  6:45     ` Morten Brørup
2026-06-23  6:56       ` Jerin Jacob
2026-06-23  7:10         ` Morten Brørup
2026-06-23  9:08           ` Jerin Jacob
2026-08-10  6:28             ` Morten Brørup
2026-06-23  8:33         ` saeed bishara
2026-06-23 12:04           ` Morten Brørup
2026-06-23 14:10             ` saeed bishara
2026-06-24  7:59               ` Morten Brørup
2026-06-24 13:09                 ` saeed bishara
2026-07-03 13:18 ` [PATCH v6] " Morten Brørup
2026-07-03 13:53 ` [PATCH v7] " Morten Brørup
2026-07-03 14:22 ` [PATCH v8] " Morten Brørup
2026-07-03 15:43 ` [PATCH v9] " Morten Brørup
2026-08-13  6:56   ` Morten Brørup [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=98CBD80474FA8B44BF855DF32C47DC35F659DD@smartserver.smartshare.dk \
    --to=mb@smartsharesystems.com \
    --cc=dev@dpdk.org \
    --cc=jerinj@marvell.com \
    --cc=kirankumark@marvell.com \
    --cc=ndabilpuram@marvell.com \
    --cc=rjarry@redhat.com \
    --cc=saeed.bishara.os@gmail.com \
    --cc=yanzhirun_163@163.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).