Linux-NVME Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH v2 0/4] nvme-cli: NIC topology aware I/O queue scaling
@ 2026-09-21  7:54 Nilay Shroff
  2026-09-21  7:54 ` [PATCH v2 1/4] shared/net-util-linux: add support for retrieving egress netdev Nilay Shroff
                   ` (3 more replies)
  0 siblings, 4 replies; 5+ messages in thread
From: Nilay Shroff @ 2026-09-21  7:54 UTC (permalink / raw)
  To: linux-nvme
  Cc: dwagner, sagi, hare, kbusch, hch, gjoyce, chaitanyak,
	Nilay Shroff

Hi,

This series is a rework of the earlier patchset[1]. The main
difference is that --nr-io-queues is now calculated in nvme-cli
instead of in the kernel when establishing an NVMe/TCP connection.

This rework is based on the feedback received[2] from the netdev
maintainers.

The original patchset determined the number of NVMe/TCP I/O queues
based on the number of online CPUs and the number of hardware queues
available on the NIC in kernel driver. This series moves that logic
to nvme-cli.

When --nr-io-queues is not explicitly specified, nvme-cli determines
the egress netdev for the NVMe/TCP connection, retrieves its current
hardware queue count, and calculates the default as:

	min(nr_hw_queues, num_online_cpus)

If nr_hw_queues cannot be retrieved, or the NIC only supports one
hardware queue, do not calculate nr-io-queues and let the host
determine the value.

The remaining patches from the original series, which expose NVMe/TCP
queue and flow information through debugfs, are still useful for
subsequent CPU/NIC topology tuning. I will send those changes as a
separate patchset.

The two patchsets are independent and can be reviewed and merged
separately.

As usual, comments, feedback, and suggestions are most welcome!

Thanks!

[1] https://lore.kernel.org/all/20260731073918.614014-1-nilay@linux.ibm.com/
[2] https://lore.kernel.org/all/20260807160904.2eb1b3d1@kernel.org/

Changes since v1:
  - Added new patch 3/4 for covering unit test cases for 
    shr_route_get_egress_iface() and shr_netdev_get_hw_queues()
    helpers (Daniel Wagner)
  - If nr_hw_queues cannot be retrieved, or the NIC only supports
    one hardware queue, do not calculate nr-io-queues and let the
    host determine the value.

Nilay Shroff (4):
  shared/net-util-linux: add support for retrieving egress netdev
  shared/net-util-linux: add support for retrieving NIC h/w queues
  shared: add more test coverage for net-util
  fabrics: determine --nr-io-queues when not explicitly specified

 shared/net-util-linux.c      | 265 +++++++++++++++++++++++++++++++++++
 shared/net-util.h            |  32 +++++
 shared/tests/test-net-util.c | 117 ++++++++++++++++
 src/fabrics.c                |  79 ++++++++++-
 4 files changed, 490 insertions(+), 3 deletions(-)

-- 
2.53.0



^ permalink raw reply	[flat|nested] 5+ messages in thread

* [PATCH v2 1/4] shared/net-util-linux: add support for retrieving egress netdev
  2026-09-21  7:54 [PATCH v2 0/4] nvme-cli: NIC topology aware I/O queue scaling Nilay Shroff
@ 2026-09-21  7:54 ` Nilay Shroff
  2026-09-21  7:55 ` [PATCH v2 2/4] shared/net-util-linux: add support for retrieving NIC h/w queues Nilay Shroff
                   ` (2 subsequent siblings)
  3 siblings, 0 replies; 5+ messages in thread
From: Nilay Shroff @ 2026-09-21  7:54 UTC (permalink / raw)
  To: linux-nvme
  Cc: dwagner, sagi, hare, kbusch, hch, gjoyce, chaitanyak,
	Nilay Shroff

Add a new shared helper, shr_route_get_egress_iface(), which performs
a route lookup and retrieves the egress netdev for a given destination
IPv4/IPv6 address and, optionally, source IPv4/IPv6 address.

Use netlink to perform the route lookup.

This is a preparatory patch. A subsequent patch will use this helper
to determine the egress netdev when calculating the default
--nr-io-queues value for NVMe/TCP connections.

Signed-off-by: Nilay Shroff <nilay@linux.ibm.com>
---
 shared/net-util-linux.c | 228 ++++++++++++++++++++++++++++++++++++++++
 shared/net-util.h       |  16 +++
 2 files changed, 244 insertions(+)

diff --git a/shared/net-util-linux.c b/shared/net-util-linux.c
index 57c025b05..b44430f38 100644
--- a/shared/net-util-linux.c
+++ b/shared/net-util-linux.c
@@ -7,15 +7,23 @@
  */
 
 #include <arpa/inet.h>
+#include <asm/types.h>
 #include <errno.h>
+#include <linux/netlink.h>
+#include <linux/rtnetlink.h>
 #include <net/if.h>
 #include <netinet/in.h>
+#include <stdio.h>
 #include <stdlib.h>
 #include <string.h>
 #include <sys/socket.h>
+#include <unistd.h>
 
 #include "net-util.h"
 
+#define NETLINK_BUFFER_SIZE	4096
+#define NETLINK_SEQ_NUM	1
+
 /*
  * Parse @addr (IPv4, or IPv6 with an optional "%scope" suffix on a
  * link-local address) into @ss. @addr is never a hostname -- resolving one
@@ -181,3 +189,223 @@ bool shr_iface_primary_addr_matches(const struct ifaddrs *iface_list,
 
 	return match_found;
 }
+
+static int rtattr_append(struct nlmsghdr *nlh, char *buf, size_t buflen,
+		unsigned short type, void *attrval, int attrlen)
+{
+	struct rtattr *rta;
+	int nlen = NLMSG_ALIGN(nlh->nlmsg_len);
+	int rlen = RTA_LENGTH(attrlen);
+
+	if (nlen + rlen >= buflen)
+		return -ENOSPC;
+
+	rta = (struct rtattr *)(buf + nlen);
+	rta->rta_type = type;
+	rta->rta_len = rlen;
+	memcpy(RTA_DATA(rta), attrval, attrlen);
+
+	nlh->nlmsg_len = nlen + rlen;
+
+	return 0;
+}
+
+int shr_route_get_egress_iface(const char *saddr, const char *daddr,
+		char *ifname, size_t iflen)
+{
+	struct sockaddr_nl nl = {.nl_family = AF_NETLINK};
+	struct sockaddr_storage ss_src, ss_dst;
+	char buf[NETLINK_BUFFER_SIZE] = {};
+	struct sockaddr_in6 *src6, *dst6;
+	struct sockaddr_in *src, *dst;
+	struct msghdr msg = {};
+	struct nlmsghdr *nlh;
+	struct rtmsg *rtmsg;
+	struct rtattr *rta;
+	int nlen, attrlen;
+	struct iovec iov;
+	int fd, ret = 0;
+
+	if (!daddr)
+		return -EINVAL;
+
+	ret = parse_numeric_addr(daddr, &ss_dst);
+	if (ret < 0) {
+		fprintf(stderr, "Invalid dest address\n");
+		return ret;
+	}
+
+	if (saddr) {
+		ret = parse_numeric_addr(saddr, &ss_src);
+		if (ret < 0) {
+			fprintf(stderr, "Invalid src address\n");
+			return ret;
+		}
+
+		if (ss_src.ss_family != ss_dst.ss_family) {
+			fprintf(stderr, "src/dst address family doesn't match\n");
+			return -EINVAL;
+		}
+	}
+
+	if (iflen < IF_NAMESIZE) {
+		fprintf(stderr, "Invalid ifname length\n");
+		return -EINVAL;
+	}
+
+	fd = socket(AF_NETLINK, SOCK_DGRAM, NETLINK_ROUTE);
+	if (fd < 0) {
+		fprintf(stderr, "Opening netlink socket failed\n");
+		return -errno;
+	}
+
+	/* encode nlmsghdr */
+	nlh = (struct nlmsghdr *)buf;
+
+	nlh->nlmsg_type = RTM_GETROUTE;
+	nlh->nlmsg_flags = NLM_F_REQUEST;
+	nlh->nlmsg_pid = 0; /* to kernel */
+	nlh->nlmsg_seq = NETLINK_SEQ_NUM;
+	nlh->nlmsg_len = sizeof(struct nlmsghdr);
+
+	/* append rtmsg */
+	nlh->nlmsg_len = NLMSG_ALIGN(nlh->nlmsg_len);
+
+	rtmsg = (struct rtmsg *)(buf + nlh->nlmsg_len);
+	rtmsg->rtm_family = ss_dst.ss_family;
+	rtmsg->rtm_dst_len = (ss_dst.ss_family == AF_INET) ? 32 : 128;
+
+	nlh->nlmsg_len += sizeof(struct rtmsg);
+
+	/* append attribute RTA_DST */
+	if (ss_dst.ss_family == AF_INET) {
+		dst = (struct sockaddr_in *)&ss_dst;
+
+		ret = rtattr_append(nlh, buf, sizeof(buf), RTA_DST,
+				&dst->sin_addr, sizeof(struct in_addr));
+		if (ret)
+			goto out;
+	} else {
+		dst6 = (struct sockaddr_in6 *)&ss_dst;
+
+		ret = rtattr_append(nlh, buf, sizeof(buf), RTA_DST,
+				&dst6->sin6_addr, sizeof(struct in6_addr));
+		if (ret)
+			goto out;
+	}
+
+	if (saddr) {
+		/* append attribute RTA_SRC */
+		if (ss_src.ss_family == AF_INET) {
+			src = (struct sockaddr_in *)&ss_src;
+
+			ret = rtattr_append(nlh, buf, sizeof(buf), RTA_SRC,
+				&src->sin_addr, sizeof(struct in_addr));
+			if (ret)
+				goto out;
+		} else {
+			src6 = (struct sockaddr_in6 *)&ss_src;
+
+			ret = rtattr_append(nlh, buf, sizeof(buf), RTA_SRC,
+				&src6->sin6_addr, sizeof(struct in6_addr));
+			if (ret)
+				goto out;
+		}
+	}
+
+	/* construct sendmsg and send to kernel */
+	iov.iov_base = buf;
+	iov.iov_len = nlh->nlmsg_len;
+
+	msg.msg_iov = &iov;
+	msg.msg_iovlen = 1;
+	msg.msg_name = &nl;
+	msg.msg_namelen = sizeof(nl);
+
+	if (sendmsg(fd, &msg, 0) < 0) {
+		fprintf(stderr, "Sendmsg to kernel failed\n");
+		ret = -errno;
+		goto out;
+	}
+
+	/* construct recvmsg buffer and wait for the response from kernel */
+	memset(&msg, 0, sizeof(msg));
+	memset(buf, 0, sizeof(buf));
+
+	iov.iov_base = buf;
+	iov.iov_len = sizeof(buf);
+
+	msg.msg_iov = &iov;
+	msg.msg_iovlen = 1;
+	msg.msg_name = &nl;
+	msg.msg_namelen = sizeof(nl);
+
+	nlen = recvmsg(fd, &msg, 0);
+	if (nlen < 0) {
+		fprintf(stderr, "Recvmsg from kernel failed\n");
+		ret = -errno;
+		goto out;
+	}
+	close(fd);
+
+	nlh = msg.msg_iov->iov_base;
+
+	for (; NLMSG_OK(nlh, nlen); nlh = NLMSG_NEXT(nlh, nlen)) {
+
+		if (nlh->nlmsg_type == NLMSG_ERROR) {
+			struct nlmsgerr *err = NLMSG_DATA(nlh);
+
+			/*
+			 * Netlink error message is sent as an ack when user
+			 * explicitly requested for it and in which case
+			 * err->error is set to zero.
+			 */
+			if (!err->error)
+				continue;
+
+			fprintf(stderr, "Received nlmsg error: %s\n",
+					strerror(-err->error));
+			break;
+		}
+
+		if (nlh->nlmsg_type == NLMSG_DONE)
+			break;
+
+		/*
+		 * RTM_GETROUTE reply is expected to be of type RTM_NEWROUTE.
+		 */
+		if (nlh->nlmsg_type != RTM_NEWROUTE)
+			continue;
+
+		if (nlh->nlmsg_seq != NETLINK_SEQ_NUM)
+			continue;
+
+		/* parse and decode received netlink message */
+		rtmsg = NLMSG_DATA(nlh);
+		rta = RTM_RTA(rtmsg);
+		attrlen = RTM_PAYLOAD(nlh);
+
+		for (; RTA_OK(rta, attrlen); rta = RTA_NEXT(rta, attrlen)) {
+
+			switch (rta->rta_type) {
+			case RTA_OIF: {
+				int ifindex;
+
+				ifindex = *((int *)RTA_DATA(rta));
+				if (!if_indextoname(ifindex, ifname)) {
+					fprintf(stderr, "if_indextoname failed\n");
+					ret = -errno;
+				}
+				return ret;
+			}
+			default:
+				break;
+			}
+		}
+	}
+
+	return -ENOENT;
+out:
+	close(fd);
+	return ret;
+}
diff --git a/shared/net-util.h b/shared/net-util.h
index 577e582f6..0208923e5 100644
--- a/shared/net-util.h
+++ b/shared/net-util.h
@@ -48,3 +48,19 @@ const char *shr_iface_matching_addr(const struct ifaddrs *iface_list,
  */
 bool shr_iface_primary_addr_matches(const struct ifaddrs *iface_list,
 		const char *iface, const char *addr);
+
+/*
+ * shr_route_get_egress_iface - Lookup route table and find egress inerface
+ *
+ * @saddr: optional source address
+ * @daddr: destintion address
+ * @ifname: buffer to hold the matching iface name
+ * @iflen: length of ifname buffer (must be greater or equal to IF_NAMESIZE)
+ *
+ * Lookup route table using specified saddr and daddr and find the egress
+ * interface. The saddr is optional however daddr is mandatory.
+ *
+ * Return: 0 on success and negative errno on failure.
+ */
+int shr_route_get_egress_iface(const char *saddr, const char *daddr,
+		char *ifname, size_t iflen);
-- 
2.53.0



^ permalink raw reply related	[flat|nested] 5+ messages in thread

* [PATCH v2 2/4] shared/net-util-linux: add support for retrieving NIC h/w queues
  2026-09-21  7:54 [PATCH v2 0/4] nvme-cli: NIC topology aware I/O queue scaling Nilay Shroff
  2026-09-21  7:54 ` [PATCH v2 1/4] shared/net-util-linux: add support for retrieving egress netdev Nilay Shroff
@ 2026-09-21  7:55 ` Nilay Shroff
  2026-09-21  7:55 ` [PATCH v2 3/4] shared: add more test coverage for net-util Nilay Shroff
  2026-09-21  7:55 ` [PATCH v2 4/4] fabrics: determine --nr-io-queues when not explicitly specified Nilay Shroff
  3 siblings, 0 replies; 5+ messages in thread
From: Nilay Shroff @ 2026-09-21  7:55 UTC (permalink / raw)
  To: linux-nvme
  Cc: dwagner, sagi, hare, kbusch, hch, gjoyce, chaitanyak,
	Nilay Shroff

Add a new shared helper, shr_netdev_get_hw_queues(), which retrieves the
NIC h/w queue details using ethtool ioctl ETHTOOL_GCHANNELS.

This is a preparatory patch. A subsequent patch will use this helper
to determine the h/w queue details when calculating the default
--nr-io-queues value for NVMe/TCP connections.

Signed-off-by: Nilay Shroff <nilay@linux.ibm.com>
---
 shared/net-util-linux.c | 37 +++++++++++++++++++++++++++++++++++++
 shared/net-util.h       | 16 ++++++++++++++++
 2 files changed, 53 insertions(+)

diff --git a/shared/net-util-linux.c b/shared/net-util-linux.c
index b44430f38..110c01c13 100644
--- a/shared/net-util-linux.c
+++ b/shared/net-util-linux.c
@@ -9,13 +9,16 @@
 #include <arpa/inet.h>
 #include <asm/types.h>
 #include <errno.h>
+#include <linux/ethtool.h>
 #include <linux/netlink.h>
 #include <linux/rtnetlink.h>
+#include <linux/sockios.h>
 #include <net/if.h>
 #include <netinet/in.h>
 #include <stdio.h>
 #include <stdlib.h>
 #include <string.h>
+#include <sys/ioctl.h>
 #include <sys/socket.h>
 #include <unistd.h>
 
@@ -409,3 +412,37 @@ out:
 	close(fd);
 	return ret;
 }
+
+int shr_netdev_get_hw_queues(const char *ifname, uint32_t *combined_count,
+		uint32_t *tx_count, uint32_t *rx_count)
+{
+	struct ethtool_channels chan = {};
+	struct ifreq ifr = {};
+	int fd, ret = 0;
+
+	if (!ifname || !combined_count || !tx_count || !rx_count)
+		return -EINVAL;
+
+	fd = socket(AF_INET, SOCK_DGRAM, 0);
+	if (fd < 0) {
+		fprintf(stderr, "opening socket failed\n");
+		return -errno;
+	}
+
+	chan.cmd = ETHTOOL_GCHANNELS;
+	snprintf(ifr.ifr_name, sizeof(ifr.ifr_name), "%s", ifname);
+	ifr.ifr_data = (void *)&chan;
+
+	if (ioctl(fd, SIOCETHTOOL, &ifr) < 0) {
+		fprintf(stderr, "SIOCETHTOOL failed\n");
+		ret = -errno;
+		goto out;
+	}
+
+	*combined_count = chan.combined_count;
+	*tx_count = chan.tx_count;
+	*rx_count = chan.rx_count;
+out:
+	close(fd);
+	return ret;
+}
diff --git a/shared/net-util.h b/shared/net-util.h
index 0208923e5..59d201210 100644
--- a/shared/net-util.h
+++ b/shared/net-util.h
@@ -8,6 +8,7 @@
 #pragma once
 
 #include <ifaddrs.h>
+#include <inttypes.h>
 #include <stdbool.h>
 
 /*
@@ -64,3 +65,18 @@ bool shr_iface_primary_addr_matches(const struct ifaddrs *iface_list,
  */
 int shr_route_get_egress_iface(const char *saddr, const char *daddr,
 		char *ifname, size_t iflen);
+
+/*
+ * shr_netdev_get_hw_queues - Get h/w queues details for a netdevice
+ *
+ * @ifname: Name of network interface
+ * @combined_count: combined h/w queue count
+ * @tx_count: tx queue count
+ * @rx_count: rx queue count
+ *
+ * Retrieves the NIC h/w queue details.
+ *
+ * Return: 0 on success and negative errno on failure.
+ */
+int shr_netdev_get_hw_queues(const char *ifname, uint32_t *combined_count,
+		uint32_t *tx_count, uint32_t *rx_count);
-- 
2.53.0



^ permalink raw reply related	[flat|nested] 5+ messages in thread

* [PATCH v2 3/4] shared: add more test coverage for net-util
  2026-09-21  7:54 [PATCH v2 0/4] nvme-cli: NIC topology aware I/O queue scaling Nilay Shroff
  2026-09-21  7:54 ` [PATCH v2 1/4] shared/net-util-linux: add support for retrieving egress netdev Nilay Shroff
  2026-09-21  7:55 ` [PATCH v2 2/4] shared/net-util-linux: add support for retrieving NIC h/w queues Nilay Shroff
@ 2026-09-21  7:55 ` Nilay Shroff
  2026-09-21  7:55 ` [PATCH v2 4/4] fabrics: determine --nr-io-queues when not explicitly specified Nilay Shroff
  3 siblings, 0 replies; 5+ messages in thread
From: Nilay Shroff @ 2026-09-21  7:55 UTC (permalink / raw)
  To: linux-nvme
  Cc: dwagner, sagi, hare, kbusch, hch, gjoyce, chaitanyak,
	Nilay Shroff

Add test covergare for shr_route_get_egress_iface()
and shr_netdev_get_hw_queues().

Signed-off-by: Nilay Shroff <nilay@linux.ibm.com>
---
 shared/tests/test-net-util.c | 117 +++++++++++++++++++++++++++++++++++
 1 file changed, 117 insertions(+)

diff --git a/shared/tests/test-net-util.c b/shared/tests/test-net-util.c
index e1ec62e5f..e7ffc5e68 100644
--- a/shared/tests/test-net-util.c
+++ b/shared/tests/test-net-util.c
@@ -7,7 +7,9 @@
  */
 
 #include <arpa/inet.h>
+#include <errno.h>
 #include <ifaddrs.h>
+#include <net/if.h>
 #include <netinet/in.h>
 #include <stdbool.h>
 #include <stdio.h>
@@ -42,6 +44,17 @@ static bool check_str(const char *name, const char *got, const char *want)
 	return false;
 }
 
+static bool check_ret(const char *name, int got, int want)
+{
+	if (got == want) {
+		printf(" - %s [PASS]\n", name);
+		return true;
+	}
+
+	printf(" - %s: got %d, want %d [FAIL]\n", name, got, want);
+	return false;
+}
+
 static bool test_ipaddrs_eq(void)
 {
 	/*
@@ -178,6 +191,108 @@ static bool test_iface_primary_addr_matches(void)
 	return pass;
 }
 
+static bool test_route_get_egress_iface(void)
+{
+	char ifname[IF_NAMESIZE];
+	char small_buf[8];
+	bool pass = true;
+	int ret;
+
+	printf("test_route_get_egress_iface:\n");
+
+	/* Invalid arguments */
+	pass &= check_ret("NULL daddr fails with -EINVAL",
+			  shr_route_get_egress_iface(NULL, NULL, ifname,
+					  sizeof(ifname)),
+			  -EINVAL);
+	pass &= check_ret("invalid daddr string fails with -EINVAL",
+			  shr_route_get_egress_iface(NULL, "invalid_ip",
+					  ifname, sizeof(ifname)),
+			  -EINVAL);
+	pass &= check_ret("invalid saddr string fails with -EINVAL",
+			  shr_route_get_egress_iface("invalid_ip", "127.0.0.1",
+					  ifname, sizeof(ifname)),
+			  -EINVAL);
+	pass &= check_ret("mismatched saddr IPv4 and daddr IPv6 families fails with -EINVAL",
+			  shr_route_get_egress_iface("127.0.0.1", "::1",
+					  ifname, sizeof(ifname)),
+			  -EINVAL);
+	pass &= check_ret("mismatched saddr IPv6 and daddr IPv4 families fails with -EINVAL",
+			  shr_route_get_egress_iface("::1", "127.0.0.1",
+					  ifname, sizeof(ifname)),
+			  -EINVAL);
+	pass &= check_ret("buffer smaller than IF_NAMESIZE fails with -EINVAL",
+			  shr_route_get_egress_iface(NULL, "127.0.0.1",
+					  small_buf, sizeof(small_buf)),
+			  -EINVAL);
+
+	/* Valid lookups (loopback routes always present) */
+	memset(ifname, 0, sizeof(ifname));
+	ret = shr_route_get_egress_iface(NULL, "127.0.0.1", ifname,
+			sizeof(ifname));
+	pass &= check_ret("IPv4 loopback destination lookup returns 0", ret, 0);
+	pass &= check_str("IPv4 loopback egress interface is lo", ifname, "lo");
+
+	memset(ifname, 0, sizeof(ifname));
+	ret = shr_route_get_egress_iface("127.0.0.1", "127.0.0.1", ifname,
+			sizeof(ifname));
+	pass &= check_ret("IPv4 loopback src and dst lookup returns 0", ret, 0);
+	pass &= check_str("IPv4 loopback src/dst egress interface is lo",
+			ifname, "lo");
+
+	memset(ifname, 0, sizeof(ifname));
+	ret = shr_route_get_egress_iface(NULL, "::1", ifname, sizeof(ifname));
+	pass &= check_ret("IPv6 loopback destination lookup returns 0", ret, 0);
+	pass &= check_str("IPv6 loopback egress interface is lo", ifname, "lo");
+
+	memset(ifname, 0, sizeof(ifname));
+	ret = shr_route_get_egress_iface("::1", "::1", ifname, sizeof(ifname));
+	pass &= check_ret("IPv6 loopback src and dst lookup returns 0", ret, 0);
+	pass &= check_str("IPv6 loopback src/dst egress interface is lo",
+			ifname, "lo");
+
+	return pass;
+}
+
+static bool test_netdev_get_hw_queues(void)
+{
+	uint32_t combined, tx, rx;
+	bool pass = true;
+	int ret;
+
+	printf("test_netdev_get_hw_queues:\n");
+
+	/* Invalid arguments */
+	pass &= check_ret("NULL ifname fails with -EINVAL",
+			  shr_netdev_get_hw_queues(NULL, &combined, &tx, &rx),
+			  -EINVAL);
+	pass &= check_ret("NULL combined_count fails with -EINVAL",
+			  shr_netdev_get_hw_queues("lo", NULL, &tx, &rx),
+			  -EINVAL);
+	pass &= check_ret("NULL tx_count fails with -EINVAL",
+			  shr_netdev_get_hw_queues("lo", &combined, NULL, &rx),
+			  -EINVAL);
+	pass &= check_ret("NULL rx_count fails with -EINVAL",
+			  shr_netdev_get_hw_queues("lo", &combined, &tx, NULL),
+			  -EINVAL);
+
+	/* Non-existent interface */
+	pass &= check_ret("non-existent interface fails with -ENODEV",
+			  shr_netdev_get_hw_queues("nonexistent99", &combined,
+					  &tx, &rx),
+			  -ENODEV);
+
+	/*
+	 * Querying loopback interface ("lo") or an interface without ethtool
+	 * channel support typically returns -EOPNOTSUPP (or 0 if supported).
+	 */
+	ret = shr_netdev_get_hw_queues("lo", &combined, &tx, &rx);
+	pass &= check_bool("querying loopback interface returns 0 or -EOPNOTSUPP",
+			   (ret == 0 || ret == -EOPNOTSUPP), true);
+
+	return pass;
+}
+
 int main(void)
 {
 	bool pass = true;
@@ -185,6 +300,8 @@ int main(void)
 	pass &= test_ipaddrs_eq();
 	pass &= test_iface_matching_addr();
 	pass &= test_iface_primary_addr_matches();
+	pass &= test_route_get_egress_iface();
+	pass &= test_netdev_get_hw_queues();
 
 	fflush(stdout);
 	exit(pass ? EXIT_SUCCESS : EXIT_FAILURE);
-- 
2.53.0



^ permalink raw reply related	[flat|nested] 5+ messages in thread

* [PATCH v2 4/4] fabrics: determine --nr-io-queues when not explicitly specified
  2026-09-21  7:54 [PATCH v2 0/4] nvme-cli: NIC topology aware I/O queue scaling Nilay Shroff
                   ` (2 preceding siblings ...)
  2026-09-21  7:55 ` [PATCH v2 3/4] shared: add more test coverage for net-util Nilay Shroff
@ 2026-09-21  7:55 ` Nilay Shroff
  3 siblings, 0 replies; 5+ messages in thread
From: Nilay Shroff @ 2026-09-21  7:55 UTC (permalink / raw)
  To: linux-nvme
  Cc: dwagner, sagi, hare, kbusch, hch, gjoyce, chaitanyak,
	Nilay Shroff

When performing an NVMe/TCP connect operation, if the user does not
explicitly specify --nr-io-queues, determine the default value using:

	if (nr_hw_queues > 1)
		min(nr_hw_queues, num_online_cpus)

Here, nr_hw_queues represents the number of hardware queues currently
configured on the NIC used for the NVMe/TCP connection. If nr_hw_queues
cannot be retrieved, or the NIC only supports one hardware queue, do
not calculate nr-io-queues and let the host determine the value.

Use the shr_route_get_egress_iface() and shr_netdev_get_hw_queues()
helpers to determine the egress netdev and retrieve its hardware
queue count. Use get_nprocs() to determine the number of online CPUs.

Apply the same logic when connecting to the target using a config INI
file.

Signed-off-by: Nilay Shroff <nilay@linux.ibm.com>
---
 src/fabrics.c | 79 +++++++++++++++++++++++++++++++++++++++++++++++++--
 1 file changed, 76 insertions(+), 3 deletions(-)

diff --git a/src/fabrics.c b/src/fabrics.c
index f5f1327ad..f5e71fdb8 100644
--- a/src/fabrics.c
+++ b/src/fabrics.c
@@ -19,17 +19,20 @@
  * Fabrics specification standard.
  */
 
+#include <ccan/minmax/minmax.h>
 #include <errno.h>
 #include <fcntl.h>
 #include <getopt.h>
 #include <inttypes.h>
 #include <libgen.h>
+#include <net/if.h>
 #include <stdbool.h>
 #include <stddef.h>
 #include <stdint.h>
 #include <stdio.h>
 #include <stdlib.h>
 #include <sys/stat.h>
+#include <sys/sysinfo.h>
 #include <sys/types.h>
 #include <time.h>
 #include <unistd.h>
@@ -49,6 +52,7 @@
 #include <ccan/endian/endian.h>
 #include <ccan/str/str.h>
 #include <shared/io-util.h>
+#include <shared/net-util.h>
 #include <shared/sig-util.h>
 
 #include "cleanup.h"
@@ -424,6 +428,50 @@ static int build_conn_tid(const struct libnvmf_config_conn *conn,
 			hostnqn, hostid, tid);
 }
 
+/*
+ * Return the number of I/O queues on success. Return 0 if the number
+ * cannot be determined or the NIC has only one hardware queue, in
+ * which case the host should determine the number of I/O queues.
+ */
+static int fabrics_connect_nr_io_queues(const char *transport,
+					const char *traddr,
+					const char *host_traddr,
+					const char *host_iface)
+{
+	uint32_t combined_count, tx_count, rx_count;
+	char ifname[IF_NAMESIZE] = {};
+	int nr_cpus, nr_hw_queues = 0;
+
+	if (strcmp(transport, "tcp"))
+		return 0;
+
+	nr_cpus = get_nprocs();
+	if (nr_cpus <= 0)
+		return 0;
+
+	if (host_iface && host_iface[0]) {
+		strncpy(ifname, host_iface, IF_NAMESIZE - 1);
+	} else {
+		if (shr_route_get_egress_iface(host_traddr, traddr,
+				ifname, IF_NAMESIZE))
+			return 0;
+	}
+
+	if (shr_netdev_get_hw_queues(ifname, &combined_count,
+			&tx_count, &rx_count))
+		return 0;
+
+	if (combined_count)
+		nr_hw_queues = combined_count;
+	else if (tx_count && rx_count)
+		nr_hw_queues = min(tx_count, rx_count);
+
+	if (nr_hw_queues <= 1)
+		return 0;
+
+	return min(nr_cpus, nr_hw_queues);
+}
+
 /* libnvmf_config_conn_for_each() callback: settle addressing/identity,
  * check exclusion, then discover or connect.
  */
@@ -435,6 +483,8 @@ static void consume_conn(const struct libnvmf_config_conn *conn,
 	struct hook_fabrics_data hfd = { .flags = st->flags, .raw = st->raw };
 	__cleanup_nvmf_context struct libnvmf_context *fctx = NULL;
 	__cleanup_nvmf_tid struct libnvmf_tid *tid = NULL;
+	const struct libnvmf_params *params;
+	const char *key = "nr-io-queues";
 	int err;
 
 	if (st->mode == CONSUME_ROLE_BASED && !is_dc && !st->connect)
@@ -460,9 +510,29 @@ static void consume_conn(const struct libnvmf_config_conn *conn,
 		goto record_err;
 
 	err = libnvmf_context_set_connection_from_tid(fctx, tid);
-	if (!err)
-		err = libnvmf_context_apply_params(fctx,
-				libnvmf_config_conn_get_params(conn));
+	if (err)
+		goto record_err;
+
+	params = libnvmf_config_conn_get_params(conn);
+	if (!libnvmf_params_get(params, key)) {
+		int nr_io_queues;
+
+		nr_io_queues = fabrics_connect_nr_io_queues(
+				libnvmf_config_conn_get_transport(conn),
+				libnvmf_config_conn_get_traddr(conn),
+				libnvmf_config_conn_get_host_traddr(conn),
+				libnvmf_config_conn_get_host_iface(conn));
+
+		if (nr_io_queues) {
+			char val[32];
+
+			snprintf(val, sizeof(val), "%d", nr_io_queues);
+			libnvmf_params_set((struct libnvmf_params *)params,
+					key, val);
+		}
+	}
+
+	err = libnvmf_context_apply_params(fctx, params);
 	if (err)
 		goto record_err;
 
@@ -1064,6 +1134,9 @@ int fabrics_connect(const char *desc, int argc, char **argv)
 		return ret;
 	fa.traddr = resolved_traddr;
 
+	if (!fa.nr_io_queues)
+		fa.nr_io_queues = fabrics_connect_nr_io_queues(fa.transport,
+				      fa.traddr, fa.host_traddr, fa.host_iface);
 do_connect:
 	ret = nvme_create_global_ctx_hostnqn(&ctx,
 		fa.hostnqn, fa.hostid, &hnqn, &hid);
-- 
2.53.0



^ permalink raw reply related	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-09-21  7:56 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-21  7:54 [PATCH v2 0/4] nvme-cli: NIC topology aware I/O queue scaling Nilay Shroff
2026-09-21  7:54 ` [PATCH v2 1/4] shared/net-util-linux: add support for retrieving egress netdev Nilay Shroff
2026-09-21  7:55 ` [PATCH v2 2/4] shared/net-util-linux: add support for retrieving NIC h/w queues Nilay Shroff
2026-09-21  7:55 ` [PATCH v2 3/4] shared: add more test coverage for net-util Nilay Shroff
2026-09-21  7:55 ` [PATCH v2 4/4] fabrics: determine --nr-io-queues when not explicitly specified Nilay Shroff

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox