* [PATCH net-next v2 00/14] devlink, mlx5, iavf, ice: XLVF for iavf
@ 2026-10-09 12:04 Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 01/14] devlink, mlx5: add init/fini ops for shared devlink Przemek Kitszel
` (13 more replies)
0 siblings, 14 replies; 15+ messages in thread
From: Przemek Kitszel @ 2026-10-09 12:04 UTC (permalink / raw)
To: netdev, Jakub Kicinski, Jiri Pirko
Cc: Tony Nguyen, Aleksandr Loktionov, Michal Schmidt, intel-wired-lan,
edumazet, horms, pabeni, davem, Jonathan Corbet, skhan, rdunlap,
andrew+netdev, saeedm, tariqt, leon, mbloch, jacob.e.keller,
jedrzej.jagielski, anzaki, brett.creeley, jtornosm, ohartoov,
Przemek Kitszel
Also available here
https://github.com/pkitszel/linux/tree/xlvf-v2
The purpose of this series is to allow iavf to use more than 16 queue
pairs, in two modes: up to 64 and up to 256 queue pairs.
In order to support more queues for VF, we must give it a bigger RSS
table (GLOBAL LUT or PF LUT). There are 16 GLOBAL LUTs on E810, and one
PF LUT for every PF on given card. Both kinds could be (re)assigned to
a VF, while a PF must hold at least one of them at any given moment.
GLOBAL LUT allows VF to use up to 64 queues, PF LUT lets it use up to
256 queues.
RSS LUTs are exposed to the user for assignment via devlink resources
API, which I have extended to make it possible.
Devlink changes:
1. Extend shared devlink by two callbacks, that give the driver
a constructor/destructor for the priv data attached to the shared
devlink instance. The constructor gets an additional param, used in
"ice: represent RSS LUTs as devlink resources". mlx5 is touched only
to adjust to the API change.
ice uses the shared devlink instance as a "whole device" aggregate
over all PFs on given card, instead of its custom xarray.
2. Extend devlink resources API to allow user to assign resources.
Before, only the driver could assign resources, without any way for
the user to interact. Now the driver could register an occupancy
setter (.occ_set()) for a resource. ice RSS LUTs are exposed that
way. It was proposed as an RFC in Feb 2025, link in the patch.
There are also patches for, Admin Queue extension, GLOBAL RSS LUT
alloc/free, and two rather big "new opcodes" patches, by Ahmed (iavf)
and Brett (ice).
Finally there is a patch that adds devlink instance for VF and
registers devlink resources on it, combined with all the glue code to
make actual use of the whole series and accept the larger number of
queues for VF. This (the last) patch contains usage examples.
There is one resource added that just groups GLOBAL and PF LUTs under
it, so there are 3 rows of data for each device (PF/VF/whole-dev):
$ devlink resource show pci/0000:18:00.0
pci/0000:18:00.0:
name rss size 1 unit entry size_min 0 size_max 2 size_gran 1 dpipe_tables none
resources:
name lut_512 size 0 unit entry size_min 0 size_max 1 size_gran 1 dpipe_tables none
name lut_2048 size 1 unit entry size_min 0 size_max 1 size_gran 1 dpipe_tables none
Technically the aggregate "name rss" line could be eliminated, with
just the two last ones kept (then renamed to "rss_lut_2048" form, from
the current "rss/lut_2048"). I like it as it is, but this is just an
opinion, I'm open to discussion.
v2:
* drop "ice: simplify ice_vc_dis_qs_msg() a little", already applied
* devlink: propagate the .shd_init() error to the caller, so
devlink_shd_get() returns ERR_PTR() now; register the shared
instance only after successful .shd_init(); document the new ops
* devlink: document the .occ_set() mode of resources
* ice: prefix the shared devlink id with "ice:", handle errors of
devl_nested_devlink_set()
* iavf: stop multi-message virtchnl ops (queue config, queue-vector
mapping) at the first error or PF NACK, keep the op pending on
timeout
* iavf: fix queue and RSS (re)negotiation on init, reset and
ethtool -L
* ice: serialize devlink RSS LUT changes with ethtool -x/-X/-L, RXHASH
toggle and VSI (re)configuration, by new pf->rss_lut_lock; make ADQ
and devlink RSS LUT reassignment mutually exclusive
* ice: limit VF queue count and PF RSS size to the capacity of the
current RSS LUT
* ice: create VF devlink together with the VF, rework deferred VF reset
* many smaller fixes, mostly after Sashiko review; per-patch changelogs
are under "---", together with Sashiko suggestions that I decided
not to apply, and why
v1:
https://lore.kernel.org/netdev/20260508124208.11622-1-przemyslaw.kitszel@intel.com
Ahmed Zaki (1):
iavf: use new opcodes to request more than 16 queues
Brett Creeley (2):
ice: add VF queue ena/dis helper functions
ice: introduce handling of virtchnl LARGE VF opcodes
Przemek Kitszel (11):
devlink, mlx5: add init/fini ops for shared devlink
ice: use shared devlink to store ice_adapters instead of custom xarray
ice: add helpers for Global RSS LUT alloc, free, vsi_update
ice: rename ICE_MAX_RSS_QS_PER_VF to ICE_MAX_QS_PER_VF_VCV1
ice: bump to 256qs for VF
iavf: extend iavf_configure_queues() to support more queues
iavf: temporary rename of IAVF_MAX_REQ_QUEUES to
IAVF_MAX_REQ_QUEUES_VCV1
iavf: increase max number of queues to 256
devlink: give user option to allocate resources
ice: represent RSS LUTs as devlink resources
ice: support up to 256 VF queues
.../networking/devlink/devlink-resource.rst | 5 +
.../networking/devlink/devlink-shared.rst | 19 +-
drivers/net/ethernet/intel/ice/Makefile | 1 +
drivers/net/ethernet/intel/iavf/iavf.h | 21 +-
.../net/ethernet/intel/iavf/iavf_register.h | 4 +-
.../net/ethernet/intel/ice/devlink/devlink.h | 2 +-
.../net/ethernet/intel/ice/devlink/resource.h | 26 +
drivers/net/ethernet/intel/ice/ice.h | 3 +
drivers/net/ethernet/intel/ice/ice_adapter.h | 52 +-
.../net/ethernet/intel/ice/ice_adminq_cmd.h | 1 +
drivers/net/ethernet/intel/ice/ice_common.h | 1 +
drivers/net/ethernet/intel/ice/ice_lag.h | 2 +-
drivers/net/ethernet/intel/ice/ice_lib.h | 5 +
drivers/net/ethernet/intel/ice/ice_switch.h | 2 +
drivers/net/ethernet/intel/ice/ice_vf_lib.h | 36 +-
drivers/net/ethernet/intel/ice/virt/queues.h | 3 +
drivers/net/ethernet/intel/ice/virt/rss.h | 1 +
.../net/ethernet/intel/ice/virt/virtchnl.h | 4 +
include/linux/net/intel/virtchnl.h | 136 +++-
include/net/devlink.h | 37 +
.../net/ethernet/intel/iavf/iavf_ethtool.c | 3 +
drivers/net/ethernet/intel/iavf/iavf_main.c | 134 +++-
.../net/ethernet/intel/iavf/iavf_virtchnl.c | 334 ++++++++-
.../net/ethernet/intel/ice/devlink/devlink.c | 11 +-
.../net/ethernet/intel/ice/devlink/resource.c | 644 ++++++++++++++++++
drivers/net/ethernet/intel/ice/ice_adapter.c | 107 ++-
drivers/net/ethernet/intel/ice/ice_common.c | 2 +-
drivers/net/ethernet/intel/ice/ice_ethtool.c | 103 +--
drivers/net/ethernet/intel/ice/ice_lag.c | 6 +-
drivers/net/ethernet/intel/ice/ice_lib.c | 178 +++--
drivers/net/ethernet/intel/ice/ice_main.c | 60 +-
drivers/net/ethernet/intel/ice/ice_sriov.c | 15 +-
drivers/net/ethernet/intel/ice/ice_switch.c | 41 ++
drivers/net/ethernet/intel/ice/ice_vf_lib.c | 69 +-
.../net/ethernet/intel/ice/virt/allowlist.c | 8 +
drivers/net/ethernet/intel/ice/virt/queues.c | 468 +++++++++++--
drivers/net/ethernet/intel/ice/virt/rss.c | 37 +-
.../net/ethernet/intel/ice/virt/virtchnl.c | 42 +-
.../ethernet/mellanox/mlx5/core/sh_devlink.c | 6 +-
net/devlink/resource.c | 88 ++-
net/devlink/sh_dev.c | 57 +-
41 files changed, 2478 insertions(+), 296 deletions(-)
create mode 100644 drivers/net/ethernet/intel/ice/devlink/resource.h
create mode 100644 drivers/net/ethernet/intel/ice/devlink/resource.c
--
2.51.1
^ permalink raw reply [flat|nested] 15+ messages in thread
* [PATCH net-next v2 01/14] devlink, mlx5: add init/fini ops for shared devlink
2026-10-09 12:04 [PATCH net-next v2 00/14] devlink, mlx5, iavf, ice: XLVF for iavf Przemek Kitszel
@ 2026-10-09 12:04 ` Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 02/14] ice: use shared devlink to store ice_adapters instead of custom xarray Przemek Kitszel
` (12 subsequent siblings)
13 siblings, 0 replies; 15+ messages in thread
From: Przemek Kitszel @ 2026-10-09 12:04 UTC (permalink / raw)
To: netdev, Jakub Kicinski, Jiri Pirko
Cc: Tony Nguyen, Aleksandr Loktionov, Michal Schmidt, intel-wired-lan,
edumazet, horms, pabeni, davem, Jonathan Corbet, skhan, rdunlap,
andrew+netdev, saeedm, tariqt, leon, mbloch, jacob.e.keller,
jedrzej.jagielski, anzaki, brett.creeley, jtornosm, ohartoov,
Przemek Kitszel
Add .shd_init() and .shd_fini() ops, that will be called for the first
devlink_shd_get() (to initialize driver's priv data) and on the last
devlink_shd_put() (to allow for the cleanup). Both ops are optional.
.shd_init() could return an error, which will stop creation of shd
instance. The initializer also gets an additional, optional param,
that driver could use for any needs.
devlink_shd_get() now returns ERR_PTR() on failure instead of NULL, so
the error from .shd_init() reaches the caller. Adjust mlx5, the only
user, to that and to the new param.
If any of the callbacks will need to get devlink instance, it could
be accessed by shd_priv_to_devlink().
Both callbacks are called with devl_lock held, .shd_init() before
devlink is registered, .shd_fini() after it is unregistered.
Next commit will make use of the callbacks, a later one will make use
also of the additional param.
Reviewed-by: Aleksandr Loktionov <aleksandr.loktionov@intel.com>
Signed-off-by: Przemek Kitszel <przemyslaw.kitszel@intel.com>
---
(v0) first discussed at:
https://lore.kernel.org/netdev/20260325063143.261806-3-przemyslaw.kitszel@intel.com
v1: remove redundant added blank line (Jiri)
v2:
* call devl_register() only after .shd_init(), so there is no NEW/DEL
notification pair when the init callback fails; adjust the order of
devl_unregister() in devlink_shd_destroy() (Sashiko)
* return the error from .shd_init() to the caller instead of
collapsing every failure into -ENOMEM, devlink_shd_get() returns
ERR_PTR() now (Sashiko)
* commit message: .shd_init()/.shd_fini() are called outside of
devlink registration; the init param is used within this series
* kdoc: use @shd_init/@shd_fini member tags, fix the "shd_devlink'
priv" wording, document the calling context (locks, before
registration/after unregistration) and that a failed .shd_init()
is retried on the next devlink_shd_get() (Sashiko)
* devlink-shared.rst: document .shd_init()/.shd_fini(), init_param,
ERR_PTR() return and shd_priv_to_devlink() (Sashiko)
---
.../networking/devlink/devlink-shared.rst | 19 ++++++-
include/net/devlink.h | 30 ++++++++++
.../ethernet/mellanox/mlx5/core/sh_devlink.c | 6 +-
net/devlink/sh_dev.c | 57 ++++++++++++++++---
4 files changed, 98 insertions(+), 14 deletions(-)
diff --git a/Documentation/networking/devlink/devlink-shared.rst b/Documentation/networking/devlink/devlink-shared.rst
index 16bf6a7d25d9..f8b2a3d8aa82 100644
--- a/Documentation/networking/devlink/devlink-shared.rst
+++ b/Documentation/networking/devlink/devlink-shared.rst
@@ -45,6 +45,14 @@ The following functions are provided for managing shared devlink instances:
* ``devlink_shd_get()``: Get or create a shared devlink instance identified by a string ID
* ``devlink_shd_put()``: Release a reference on a shared devlink instance
* ``devlink_shd_get_priv()``: Get private data from shared devlink instance
+* ``shd_priv_to_devlink()``: Get shared devlink instance from its private data
+
+``devlink_shd_get()`` returns ``ERR_PTR()`` on failure.
+
+The driver may initialize and clean up its private data of the shared
+instance in the optional ``shd_init()`` and ``shd_fini()`` callbacks of
+``struct devlink_ops``. The ``init_param`` argument of ``devlink_shd_get()``
+is passed to ``shd_init()``.
Initialization Flow
-------------------
@@ -55,7 +63,9 @@ Initialization Flow
* The function looks up existing instance by identifier
* If none exists, creates new instance:
- - Allocates and registers devlink instance
+ - Allocates devlink instance
+ - Calls ``shd_init()`` (if set), on failure frees it and returns the error
+ - Registers devlink instance
- Adds to global shared instances list
- Increments reference count
@@ -67,7 +77,8 @@ Cleanup Flow
1. **Cleanup** when PF is removed
2. **Call** ``devlink_shd_put()`` to release reference (decrements reference count)
-3. **Shared instance is automatically destroyed** when the last PF removes (reference count reaches zero)
+3. **Shared instance is automatically destroyed** when the last PF removes (reference count reaches zero),
+ ``shd_fini()`` (if set) is called after unregistering it
Chip Identification
-------------------
@@ -85,6 +96,10 @@ Locking
A global mutex (``shd_mutex``) protects the shared instances list during registration/deregistration.
+``shd_init()`` and ``shd_fini()`` are called with ``shd_mutex`` and the devlink
+lock of the shared instance held, thus must not call ``devlink_shd_get()`` nor
+``devlink_shd_put()``, and must use ``devl_*()`` variants of devlink API.
+
Similarly to other nested devlink instance relationships, devlink lock of
the shared instance should be always taken after the devlink lock of PF.
diff --git a/include/net/devlink.h b/include/net/devlink.h
index 7abd23376319..98bdde4f72ca 100644
--- a/include/net/devlink.h
+++ b/include/net/devlink.h
@@ -1607,6 +1607,34 @@ struct devlink_ops {
* with other operations.
*/
bool supported_cross_device_rate_nodes;
+ /**
+ * @shd_init: Shared devlink instance initializer.
+ *
+ * Gets the driver's priv data of the shared instance (the same as
+ * devlink_shd_get_priv() returns) and the init_param passed to
+ * devlink_shd_get(). Should initialize the priv data. May be NULL.
+ *
+ * Called when devlink_shd_get() creates the shared instance, before
+ * it is registered. If it fails, the instance is freed without calling
+ * @shd_fini, and the next devlink_shd_get() for the same id will try to
+ * create it (and call @shd_init) again.
+ *
+ * Called with the global shared devlink mutex and the devlink lock of
+ * the shared instance held, thus must not call devlink_shd_get() nor
+ * devlink_shd_put(), and must use devl_*() variants of devlink API.
+ *
+ * Return: 0 on success, negative errno to fail the creation.
+ */
+ int (*shd_init)(void *priv, void *init_param);
+ /**
+ * @shd_fini: Shared devlink instance finalizer.
+ *
+ * Gets the driver's priv data of the shared instance. Called when the
+ * last reference is dropped, after the shared instance is unregistered,
+ * in the same locking context as @shd_init. Should clean up the priv
+ * data. May be NULL.
+ */
+ void (*shd_fini)(void *priv);
/**
* selftests_check() - queries if selftest is supported
* @devlink: devlink instance
@@ -1672,9 +1700,11 @@ void devlink_free(struct devlink *devlink);
struct devlink *devlink_shd_get(const char *id,
const struct devlink_ops *ops,
size_t priv_size,
+ void *init_param,
const struct device_driver *driver);
void devlink_shd_put(struct devlink *devlink);
void *devlink_shd_get_priv(struct devlink *devlink);
+struct devlink *shd_priv_to_devlink(void *priv);
/**
* struct devlink_port_ops - Port operations
diff --git a/drivers/net/ethernet/mellanox/mlx5/core/sh_devlink.c b/drivers/net/ethernet/mellanox/mlx5/core/sh_devlink.c
index b925364765ac..2c58f820b018 100644
--- a/drivers/net/ethernet/mellanox/mlx5/core/sh_devlink.c
+++ b/drivers/net/ethernet/mellanox/mlx5/core/sh_devlink.c
@@ -43,10 +43,10 @@ int mlx5_shd_init(struct mlx5_core_dev *dev)
*end = '\0';
/* Get or create shared devlink instance */
- devlink = devlink_shd_get(sn, &mlx5_shd_ops, 0, pdev->dev.driver);
+ devlink = devlink_shd_get(sn, &mlx5_shd_ops, 0, NULL, pdev->dev.driver);
kfree(sn);
- if (!devlink)
- return -ENOMEM;
+ if (IS_ERR(devlink))
+ return PTR_ERR(devlink);
dev->shd = devlink;
return 0;
diff --git a/net/devlink/sh_dev.c b/net/devlink/sh_dev.c
index 85acce97e788..4f1b2f24b755 100644
--- a/net/devlink/sh_dev.c
+++ b/net/devlink/sh_dev.c
@@ -34,43 +34,61 @@ static struct devlink_shd *devlink_shd_lookup(const char *id)
static struct devlink_shd *devlink_shd_create(const char *id,
const struct devlink_ops *ops,
size_t priv_size,
+ void *init_param,
const struct device_driver *driver)
{
struct devlink_shd *shd;
struct devlink *devlink;
+ int err;
devlink = __devlink_alloc(ops, sizeof(struct devlink_shd) + priv_size,
&init_net, NULL, driver);
if (!devlink)
- return NULL;
+ return ERR_PTR(-ENOMEM);
shd = devlink_priv(devlink);
shd->id = kstrdup(id, GFP_KERNEL);
- if (!shd->id)
+ if (!shd->id) {
+ err = -ENOMEM;
goto err_devlink_free;
+ }
shd->priv_size = priv_size;
- refcount_set(&shd->refcount, 1);
devl_lock(devlink);
+
+ if (ops->shd_init) {
+ err = ops->shd_init(shd->priv, init_param);
+ if (err)
+ goto err_unlock;
+ }
devl_register(devlink);
devl_unlock(devlink);
+ refcount_set(&shd->refcount, 1);
list_add_tail(&shd->list, &shd_list);
return shd;
+err_unlock:
+ devl_unlock(devlink);
+ kfree(shd->id);
+
err_devlink_free:
devlink_free(devlink);
- return NULL;
+ return ERR_PTR(err);
}
static void devlink_shd_destroy(struct devlink_shd *shd)
{
struct devlink *devlink = priv_to_devlink(shd);
list_del(&shd->list);
devl_lock(devlink);
devl_unregister(devlink);
+
+ if (devlink->ops->shd_fini)
+ devlink->ops->shd_fini(shd->priv);
+
devl_unlock(devlink);
kfree(shd->id);
devlink_free(devlink);
@@ -81,46 +99,53 @@ static void devlink_shd_destroy(struct devlink_shd *shd)
* @id: Identifier string (e.g., serial number) for the shared instance
* @ops: Devlink operations structure
* @priv_size: Size of private data structure
+ * @init_param: Passed to .shd_init() callback alongside driver's priv;
+ * this value need not match across all users of the shared
+ * instance
* @driver: Driver associated with the shared devlink instance
*
* Get an existing shared devlink instance identified by @id, or create
* a new one if it doesn't exist. Return the devlink instance with a
* reference held. The caller must call devlink_shd_put() when done.
*
+ * On creation, @ops->shd_init() is called (if set), see its description
+ * for the calling context. Its error is returned to the caller.
+ *
* All callers sharing the same @id must pass identical @ops, @priv_size
- * and @driver. A mismatch triggers a warning and returns NULL.
+ * and @driver. A mismatch triggers a warning and returns ERR_PTR(-EINVAL).
*
* Return: Pointer to the shared devlink instance on success,
- * NULL on failure
+ * ERR_PTR() on failure
*/
struct devlink *devlink_shd_get(const char *id,
const struct devlink_ops *ops,
size_t priv_size,
+ void *init_param,
const struct device_driver *driver)
{
struct devlink *devlink;
struct devlink_shd *shd;
mutex_lock(&shd_mutex);
shd = devlink_shd_lookup(id);
if (!shd) {
- shd = devlink_shd_create(id, ops, priv_size, driver);
+ shd = devlink_shd_create(id, ops, priv_size, init_param, driver);
goto unlock;
}
devlink = priv_to_devlink(shd);
if (WARN_ON_ONCE(devlink->ops != ops ||
shd->priv_size != priv_size ||
devlink->dev_driver != driver)) {
- shd = NULL;
+ shd = ERR_PTR(-EINVAL);
goto unlock;
}
refcount_inc(&shd->refcount);
unlock:
mutex_unlock(&shd_mutex);
- return shd ? priv_to_devlink(shd) : NULL;
+ return IS_ERR(shd) ? ERR_CAST(shd) : priv_to_devlink(shd);
}
EXPORT_SYMBOL_GPL(devlink_shd_get);
@@ -159,3 +184,17 @@ void *devlink_shd_get_priv(struct devlink *devlink)
return shd->priv;
}
EXPORT_SYMBOL_GPL(devlink_shd_get_priv);
+
+/**
+ * shd_priv_to_devlink - Get shared devlink instance from its priv data
+ * @priv: Driver's priv data, as returned by devlink_shd_get_priv()
+ *
+ * Return: pointer to shared devlink instance the @priv belongs to.
+ */
+struct devlink *shd_priv_to_devlink(void *priv)
+{
+ struct devlink_shd *shd = container_of(priv, struct devlink_shd, priv);
+
+ return priv_to_devlink(shd);
+}
+EXPORT_SYMBOL_GPL(shd_priv_to_devlink);
--
2.51.1
^ permalink raw reply related [flat|nested] 15+ messages in thread
* [PATCH net-next v2 02/14] ice: use shared devlink to store ice_adapters instead of custom xarray
2026-10-09 12:04 [PATCH net-next v2 00/14] devlink, mlx5, iavf, ice: XLVF for iavf Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 01/14] devlink, mlx5: add init/fini ops for shared devlink Przemek Kitszel
@ 2026-10-09 12:04 ` Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 03/14] ice: add VF queue ena/dis helper functions Przemek Kitszel
` (11 subsequent siblings)
13 siblings, 0 replies; 15+ messages in thread
From: Przemek Kitszel @ 2026-10-09 12:04 UTC (permalink / raw)
To: netdev, Jakub Kicinski, Jiri Pirko
Cc: Tony Nguyen, Aleksandr Loktionov, Michal Schmidt, intel-wired-lan,
edumazet, horms, pabeni, davem, Jonathan Corbet, skhan, rdunlap,
andrew+netdev, saeedm, tariqt, leon, mbloch, jacob.e.keller,
jedrzej.jagielski, anzaki, brett.creeley, jtornosm, ohartoov,
Przemek Kitszel, Sergey Temerkhanov, Grzegorz Nitka
Refactor our storage and deduplication logic of ice_adapters by moving
it to be handled by shared devlink instance, recently added by
Jiri Pirko [1].
We wanted the devlink instance for whole device anyway - later in the
series I will add devlink resources under it.
Make the shared devlink a parent (wrt. nesting) of the actual PF devices.
[1] commit 411ad0605875 ("Merge branch 'devlink-introduce-shared-devlink-instance-for-pfs-on-same-chip'")
[1] https://lore.kernel.org/all/20260312100407.551173-1-jiri@resnulli.us
Reviewed-by: Jedrzej Jagielski <jedrzej.jagielski@intel.com>
Reviewed-by: Aleksandr Loktionov <aleksandr.loktionov@intel.com>
Signed-off-by: Przemek Kitszel <przemyslaw.kitszel@intel.com>
---
CC: Sergey Temerkhanov <sergey.temerkhanov@intel.com>
CC: Grzegorz Nitka <grzegorz.nitka@intel.com>
v2:
* propagate the error from devlink_shd_get() with ERR_CAST()
* minor kdoc updates
* check devl_nested_devlink_set() return value, ice_devlink_register()
returns int now (Sashiko)
* prefix the shared devlink id with "ice:", to avoid collisions with
ids used by other drivers (Sashiko)
Sashiko suggestions I don't want to apply:
S: all PFs without DSN get index 0, and so share one instance
me: pre-existing, and DSN 0 is only on pre-production samples
---
.../net/ethernet/intel/ice/devlink/devlink.h | 2 +-
drivers/net/ethernet/intel/ice/ice_adapter.h | 12 +--
.../net/ethernet/intel/ice/devlink/devlink.c | 11 ++-
drivers/net/ethernet/intel/ice/ice_adapter.c | 99 ++++++-------------
drivers/net/ethernet/intel/ice/ice_main.c | 11 ++-
5 files changed, 57 insertions(+), 78 deletions(-)
diff --git a/drivers/net/ethernet/intel/ice/devlink/devlink.h b/drivers/net/ethernet/intel/ice/devlink/devlink.h
index 1af3b0763fbb..62af38d11fd9 100644
--- a/drivers/net/ethernet/intel/ice/devlink/devlink.h
+++ b/drivers/net/ethernet/intel/ice/devlink/devlink.h
@@ -7,7 +7,7 @@
struct ice_pf *ice_allocate_pf(struct device *dev);
struct ice_sf_priv *ice_allocate_sf(struct device *dev, struct ice_pf *pf);
-void ice_devlink_register(struct ice_pf *pf);
+int ice_devlink_register(struct ice_pf *pf);
void ice_devlink_unregister(struct ice_pf *pf);
int ice_devlink_register_params(struct ice_pf *pf);
void ice_devlink_unregister_params(struct ice_pf *pf);
diff --git a/drivers/net/ethernet/intel/ice/ice_adapter.h b/drivers/net/ethernet/intel/ice/ice_adapter.h
index 4f695f32da3d..c11d511d7a8e 100644
--- a/drivers/net/ethernet/intel/ice/ice_adapter.h
+++ b/drivers/net/ethernet/intel/ice/ice_adapter.h
@@ -7,7 +7,8 @@
#include <linux/types.h>
#include <linux/mutex.h>
#include <linux/spinlock_types.h>
-#include <linux/refcount_types.h>
+
+#include <net/devlink.h>
#include "ice_type.h"
@@ -30,31 +31,30 @@ struct ice_port_list {
/**
* struct ice_adapter - PCI adapter resources shared across PFs
- * @refcount: Reference count. struct ice_pf objects hold the references.
+ * @devlink: ice adapter's devlink (whole dev devlink)
* @ptp_gltsyn_time_lock: Spinlock protecting access to the GLTSYN_TIME
* register of the PTP clock.
* @txq_ctx_lock: Spinlock protecting access to the GLCOMM_QTX_CNTX_CTL register
* @cpi_phy_lock: Per-PHY mutex serializing CPI REQ/ACK transactions.
* Index 0 = PHY0, index 1 = PHY1. Used on E825C devices.
* @ctrl_pf: Control PF of the adapter
* @ports: Ports list
- * @index: 64-bit index cached for collision detection on 32bit systems
*/
struct ice_adapter {
- refcount_t refcount;
+ struct devlink *devlink;
+
/* For access to the GLTSYN_TIME register */
spinlock_t ptp_gltsyn_time_lock;
/* For access to GLCOMM_QTX_CNTX_CTL register */
spinlock_t txq_ctx_lock;
/* Serialize CPI REQ/ACK transactions per PHY (E825C only) */
struct mutex cpi_phy_lock[ICE_E825_MAX_PHYS];
struct ice_pf *ctrl_pf;
struct ice_port_list ports;
- u64 index;
};
struct ice_adapter *ice_adapter_get(struct pci_dev *pdev);
-void ice_adapter_put(struct pci_dev *pdev);
+void ice_adapter_put(struct ice_adapter *adapter);
#endif /* _ICE_ADAPTER_H */
diff --git a/drivers/net/ethernet/intel/ice/devlink/devlink.c b/drivers/net/ethernet/intel/ice/devlink/devlink.c
index 8c2b63eef82b..fa64b9744779 100644
--- a/drivers/net/ethernet/intel/ice/devlink/devlink.c
+++ b/drivers/net/ethernet/intel/ice/devlink/devlink.c
@@ -1747,11 +1747,20 @@ struct ice_sf_priv *ice_allocate_sf(struct device *dev, struct ice_pf *pf)
*
* Return: zero on success or an error code on failure.
*/
-void ice_devlink_register(struct ice_pf *pf)
+int ice_devlink_register(struct ice_pf *pf)
{
struct devlink *devlink = priv_to_devlink(pf);
+ struct ice_adapter *adapter = pf->adapter;
+ int err;
+ if (adapter) {
+ err = devl_nested_devlink_set(adapter->devlink, devlink);
+ if (err)
+ return err;
+ }
devl_register(devlink);
+
+ return 0;
}
/**
diff --git a/drivers/net/ethernet/intel/ice/ice_adapter.c b/drivers/net/ethernet/intel/ice/ice_adapter.c
index 2dc3629d6d0f..c6301341ccae 100644
--- a/drivers/net/ethernet/intel/ice/ice_adapter.c
+++ b/drivers/net/ethernet/intel/ice/ice_adapter.c
@@ -1,18 +1,13 @@
// SPDX-License-Identifier: GPL-2.0-only
// SPDX-FileCopyrightText: Copyright Red Hat
-#include <linux/cleanup.h>
-#include <linux/mutex.h>
#include <linux/pci.h>
#include <linux/slab.h>
#include <linux/spinlock.h>
-#include <linux/xarray.h>
+
#include "ice_adapter.h"
#include "ice.h"
-static DEFINE_XARRAY(ice_adapters);
-static DEFINE_MUTEX(ice_adapters_mutex);
-
#define ICE_ADAPTER_FIXED_INDEX BIT_ULL(63)
#define ICE_ADAPTER_INDEX_E825C \
@@ -40,48 +35,40 @@ static u64 ice_adapter_index(struct pci_dev *pdev)
}
}
-static unsigned long ice_adapter_xa_index(struct pci_dev *pdev)
-{
- u64 index = ice_adapter_index(pdev);
-
-#if BITS_PER_LONG == 64
- return index;
-#else
- return (u32)index ^ (u32)(index >> 32);
-#endif
-}
-
-static struct ice_adapter *ice_adapter_new(struct pci_dev *pdev)
+static int ice_adapter_init(void *priv, void *init_param)
{
- struct ice_adapter *adapter;
+ struct ice_adapter *adapter = priv;
+ struct devlink *devlink;
- adapter = kzalloc_obj(*adapter);
- if (!adapter)
- return NULL;
+ devlink = shd_priv_to_devlink(adapter);
+ adapter->devlink = devlink;
- adapter->index = ice_adapter_index(pdev);
spin_lock_init(&adapter->ptp_gltsyn_time_lock);
spin_lock_init(&adapter->txq_ctx_lock);
for (int i = 0; i < ARRAY_SIZE(adapter->cpi_phy_lock); i++)
mutex_init(&adapter->cpi_phy_lock[i]);
- refcount_set(&adapter->refcount, 1);
mutex_init(&adapter->ports.lock);
INIT_LIST_HEAD(&adapter->ports.ports);
- return adapter;
+ return 0;
}
-static void ice_adapter_free(struct ice_adapter *adapter)
+static void ice_adapter_fini(void *priv)
{
+ struct ice_adapter *adapter = priv;
+
WARN_ON(!list_empty(&adapter->ports.ports));
for (int i = 0; i < ARRAY_SIZE(adapter->cpi_phy_lock); i++)
mutex_destroy(&adapter->cpi_phy_lock[i]);
mutex_destroy(&adapter->ports.lock);
-
- kfree(adapter);
}
+static const struct devlink_ops ice_adapter_devlink_ops = {
+ .shd_init = ice_adapter_init,
+ .shd_fini = ice_adapter_fini,
+};
+
/**
* ice_adapter_get - Get a shared ice_adapter structure.
* @pdev: Pointer to the pci_dev whose driver is getting the ice_adapter.
@@ -93,59 +80,37 @@ static void ice_adapter_free(struct ice_adapter *adapter)
*
* Context: Process, may sleep.
* Return: Pointer to ice_adapter on success.
- * ERR_PTR() on error. -ENOMEM is the only possible error.
+ * ERR_PTR() on error.
*/
struct ice_adapter *ice_adapter_get(struct pci_dev *pdev)
{
struct ice_adapter *adapter;
- unsigned long index;
- int err;
-
- index = ice_adapter_xa_index(pdev);
- scoped_guard(mutex, &ice_adapters_mutex) {
- adapter = xa_load(&ice_adapters, index);
- if (adapter) {
- refcount_inc(&adapter->refcount);
- WARN_ON_ONCE(adapter->index != ice_adapter_index(pdev));
- return adapter;
- }
- err = xa_reserve(&ice_adapters, index, GFP_KERNEL);
- if (err)
- return ERR_PTR(err);
-
- adapter = ice_adapter_new(pdev);
- if (!adapter) {
- xa_release(&ice_adapters, index);
- return ERR_PTR(-ENOMEM);
- }
- xa_store(&ice_adapters, index, adapter, GFP_KERNEL);
- }
+ struct devlink *devlink;
+ char devlink_id[32];
+ u64 index;
+
+ index = ice_adapter_index(pdev);
+ snprintf(devlink_id, sizeof(devlink_id), "ice:%llx", index);
+ devlink = devlink_shd_get(devlink_id, &ice_adapter_devlink_ops,
+ sizeof(*adapter), NULL, pdev->dev.driver);
+ if (IS_ERR(devlink))
+ return ERR_CAST(devlink);
+
+ adapter = devlink_shd_get_priv(devlink);
+
return adapter;
}
/**
* ice_adapter_put - Release a reference to the shared ice_adapter structure.
- * @pdev: Pointer to the pci_dev whose driver is releasing the ice_adapter.
+ * @adapter: the ice_adapter to release reference to
*
* Releases the reference to ice_adapter previously obtained with
* ice_adapter_get.
*
* Context: Process, may sleep.
*/
-void ice_adapter_put(struct pci_dev *pdev)
+void ice_adapter_put(struct ice_adapter *adapter)
{
- struct ice_adapter *adapter;
- unsigned long index;
-
- index = ice_adapter_xa_index(pdev);
- scoped_guard(mutex, &ice_adapters_mutex) {
- adapter = xa_load(&ice_adapters, index);
- if (WARN_ON(!adapter))
- return;
- if (!refcount_dec_and_test(&adapter->refcount))
- return;
-
- WARN_ON(xa_erase(&ice_adapters, index) != adapter);
- }
- ice_adapter_free(adapter);
+ devlink_shd_put(adapter->devlink);
}
diff --git a/drivers/net/ethernet/intel/ice/ice_main.c b/drivers/net/ethernet/intel/ice/ice_main.c
index 454fba07d4ec..6373cd5ac288 100644
--- a/drivers/net/ethernet/intel/ice/ice_main.c
+++ b/drivers/net/ethernet/intel/ice/ice_main.c
@@ -4946,7 +4946,12 @@ static int ice_init_devlink(struct ice_pf *pf)
return err;
ice_devlink_init_regions(pf);
- ice_devlink_register(pf);
+ err = ice_devlink_register(pf);
+ if (err) {
+ ice_devlink_destroy_regions(pf);
+ ice_devlink_unregister_params(pf);
+ return err;
+ }
ice_health_init(pf);
return 0;
@@ -5291,7 +5296,7 @@ ice_probe(struct pci_dev *pdev, const struct pci_device_id __always_unused *ent)
unroll_dev_init:
need_dev_deinit = true;
unroll_adapter:
- ice_adapter_put(pdev);
+ ice_adapter_put(adapter);
unroll_hw_init:
ice_deinit_hw(hw);
if (need_dev_deinit)
@@ -5404,7 +5409,7 @@ static void ice_remove(struct pci_dev *pdev)
ice_setup_mc_magic_wake(pf);
ice_set_wake(pf);
- ice_adapter_put(pdev);
+ ice_adapter_put(pf->adapter);
ice_deinit_hw(&pf->hw);
ice_deinit_dev(pf);
--
2.51.1
^ permalink raw reply related [flat|nested] 15+ messages in thread
* [PATCH net-next v2 03/14] ice: add VF queue ena/dis helper functions
2026-10-09 12:04 [PATCH net-next v2 00/14] devlink, mlx5, iavf, ice: XLVF for iavf Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 01/14] devlink, mlx5: add init/fini ops for shared devlink Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 02/14] ice: use shared devlink to store ice_adapters instead of custom xarray Przemek Kitszel
@ 2026-10-09 12:04 ` Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 04/14] ice: add helpers for Global RSS LUT alloc, free, vsi_update Przemek Kitszel
` (10 subsequent siblings)
13 siblings, 0 replies; 15+ messages in thread
From: Przemek Kitszel @ 2026-10-09 12:04 UTC (permalink / raw)
To: netdev, Jakub Kicinski, Jiri Pirko
Cc: Tony Nguyen, Aleksandr Loktionov, Michal Schmidt, intel-wired-lan,
edumazet, horms, pabeni, davem, Jonathan Corbet, skhan, rdunlap,
andrew+netdev, saeedm, tariqt, leon, mbloch, jacob.e.keller,
jedrzej.jagielski, anzaki, brett.creeley, jtornosm, ohartoov,
Przemek Kitszel, Wojciech Drewek
From: Brett Creeley <brett.creeley@intel.com>
Add three new functions:
- ice_vf_vsi_dis_single_rxq()
- ice_vf_vsi_ena_single_rxq()
- ice_vf_vsi_ena_single_txq()
(ice_vf_vsi_dis_single_txq() was introduced earlier).
Those functions wrap operations needed in the processes of enabling
and disabling single Tx/Rx VF queue, which are:
- check if the queue is not already in desired state
- perform the dis/ena operations
- bookkeeping of the queue's state.
Future commit will use them from another callsite.
Signed-off-by: Brett Creeley <brett.creeley@intel.com>
Signed-off-by: Wojciech Drewek <wojciech.drewek@intel.com>
Reviewed-by: Jedrzej Jagielski <jedrzej.jagielski@intel.com>
Reviewed-by: Aleksandr Loktionov <aleksandr.loktionov@intel.com>
Signed-off-by: Przemek Kitszel <przemyslaw.kitszel@intel.com>
---
v2:
* convert one comment to kdoc (Alex)
* fix typos (Sashiko)
* rebase: move ice_vf_dis_rxq_interrupt(), added to ice_vc_dis_qs_msg()
by commit fb096882095e ("ice: fix VF interrupts cleanup"), into
ice_vf_vsi_dis_single_rxq()
---
drivers/net/ethernet/intel/ice/virt/queues.c | 117 ++++++++++++++-----
1 file changed, 89 insertions(+), 28 deletions(-)
diff --git a/drivers/net/ethernet/intel/ice/virt/queues.c b/drivers/net/ethernet/intel/ice/virt/queues.c
index ac7f98479b13..a22780575086 100644
--- a/drivers/net/ethernet/intel/ice/virt/queues.c
+++ b/drivers/net/ethernet/intel/ice/virt/queues.c
@@ -242,6 +242,57 @@ static void ice_vf_dis_rxq_interrupt(struct ice_vsi *vsi, u32 q_idx)
ice_flush(hw);
}
+/**
+ * ice_vf_vsi_ena_single_rxq - enable single Rx queue based on relative q_id
+ * @vf: VF to enable queue for
+ * @vsi: VSI for the VF
+ * @q_id: VSI relative (0-based) queue ID
+ *
+ * Enable the Rx queue passed in.
+ *
+ * Return: 0 on success or negative on error.
+ */
+static int ice_vf_vsi_ena_single_rxq(struct ice_vf *vf, struct ice_vsi *vsi,
+ u16 q_id)
+{
+ int err;
+
+ if (test_bit(q_id, vf->rxq_ena))
+ return 0;
+
+ err = ice_vsi_ctrl_one_rx_ring(vsi, true, q_id, true);
+ if (err) {
+ dev_err(ice_pf_to_dev(vsi->back), "Failed to enable Rx ring %d on VSI %d\n",
+ q_id, vsi->vsi_num);
+ return err;
+ }
+
+ ice_vf_ena_rxq_interrupt(vsi, q_id);
+ set_bit(q_id, vf->rxq_ena);
+
+ return 0;
+}
+
+/**
+ * ice_vf_vsi_ena_single_txq - enable single Tx queue based on relative q_id
+ * @vf: VF to enable queue for
+ * @vsi: VSI for the VF
+ * @q_id: VSI relative (0-based) queue ID
+ *
+ * Enable the Tx queue's interrupt. Note that the Tx queue(s) should have
+ * already been configured/enabled in VIRTCHNL_OP_CONFIG_VSI_QUEUES so this
+ * function only enables the interrupt associated with the q_id.
+ */
+static void ice_vf_vsi_ena_single_txq(struct ice_vf *vf, struct ice_vsi *vsi,
+ u16 q_id)
+{
+ if (test_bit(q_id, vf->txq_ena))
+ return;
+
+ ice_vf_ena_txq_interrupt(vsi, q_id);
+ set_bit(q_id, vf->txq_ena);
+}
+
/**
* ice_vc_ena_qs_msg
* @vf: pointer to the VF info
@@ -290,34 +341,20 @@ int ice_vc_ena_qs_msg(struct ice_vf *vf, u8 *msg)
goto error_param;
}
- /* Skip queue if enabled */
- if (test_bit(vf_q_id, vf->rxq_ena))
- continue;
-
- if (ice_vsi_ctrl_one_rx_ring(vsi, true, vf_q_id, true)) {
- dev_err(ice_pf_to_dev(vsi->back), "Failed to enable Rx ring %d on VSI %d\n",
- vf_q_id, vsi->vsi_num);
+ if (ice_vf_vsi_ena_single_rxq(vf, vsi, vf_q_id)) {
v_ret = VIRTCHNL_STATUS_ERR_PARAM;
goto error_param;
}
-
- ice_vf_ena_rxq_interrupt(vsi, vf_q_id);
- set_bit(vf_q_id, vf->rxq_ena);
}
q_map = vqs->tx_queues;
for_each_set_bit(vf_q_id, &q_map, ICE_MAX_RSS_QS_PER_VF) {
if (!ice_vc_isvalid_q_id(vsi, vf_q_id)) {
v_ret = VIRTCHNL_STATUS_ERR_PARAM;
goto error_param;
}
- /* Skip queue if enabled */
- if (test_bit(vf_q_id, vf->txq_ena))
- continue;
-
- ice_vf_ena_txq_interrupt(vsi, vf_q_id);
- set_bit(vf_q_id, vf->txq_ena);
+ ice_vf_vsi_ena_single_txq(vf, vsi, vf_q_id);
}
/* Set flag to indicate that queues are enabled */
@@ -369,6 +406,41 @@ int ice_vf_vsi_dis_single_txq(struct ice_vf *vf, struct ice_vsi *vsi, u16 q_id)
return 0;
}
+/**
+ * ice_vf_vsi_dis_single_rxq - disable a Rx queue for VF on relative queue ID
+ * @vf: VF to disable queue for
+ * @vsi: VSI for the VF
+ * @q_id: VSI relative (0-based) queue ID
+ *
+ * Attempt to disable the Rx queue passed in. If the Rx queue was successfully
+ * disabled then clear q_id bit in the enabled queues bitmap.
+ *
+ * Return: 0 when queue was already disabled or disabling succeeded, otherwise
+ * negative error code.
+ */
+static int ice_vf_vsi_dis_single_rxq(struct ice_vf *vf, struct ice_vsi *vsi,
+ u16 q_id)
+{
+ int err;
+
+ if (!test_bit(q_id, vf->rxq_ena))
+ return 0;
+
+ err = ice_vsi_ctrl_one_rx_ring(vsi, false, q_id, true);
+ if (err) {
+ dev_err(ice_pf_to_dev(vsi->back), "Failed to stop Rx ring %d on VSI %d\n",
+ q_id, vsi->vsi_num);
+ return err;
+ }
+
+ ice_vf_dis_rxq_interrupt(vsi, q_id);
+
+ /* Clear enabled queues flag */
+ clear_bit(q_id, vf->rxq_ena);
+
+ return 0;
+}
+
/**
* ice_vc_dis_qs_msg
* @vf: pointer to the VF info
@@ -433,21 +505,10 @@ int ice_vc_dis_qs_msg(struct ice_vf *vf, u8 *msg)
goto error_param;
}
- /* Skip queue if not enabled */
- if (!test_bit(vf_q_id, vf->rxq_ena))
- continue;
-
- if (ice_vsi_ctrl_one_rx_ring(vsi, false, vf_q_id,
- true)) {
- dev_err(ice_pf_to_dev(vsi->back), "Failed to stop Rx ring %d on VSI %d\n",
- vf_q_id, vsi->vsi_num);
+ if (ice_vf_vsi_dis_single_rxq(vf, vsi, vf_q_id)) {
v_ret = VIRTCHNL_STATUS_ERR_PARAM;
goto error_param;
}
-
- ice_vf_dis_rxq_interrupt(vsi, vf_q_id);
- /* Clear enabled queues flag */
- clear_bit(vf_q_id, vf->rxq_ena);
}
}
--
2.51.1
^ permalink raw reply related [flat|nested] 15+ messages in thread
* [PATCH net-next v2 04/14] ice: add helpers for Global RSS LUT alloc, free, vsi_update
2026-10-09 12:04 [PATCH net-next v2 00/14] devlink, mlx5, iavf, ice: XLVF for iavf Przemek Kitszel
` (2 preceding siblings ...)
2026-10-09 12:04 ` [PATCH net-next v2 03/14] ice: add VF queue ena/dis helper functions Przemek Kitszel
@ 2026-10-09 12:04 ` Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 05/14] ice: rename ICE_MAX_RSS_QS_PER_VF to ICE_MAX_QS_PER_VF_VCV1 Przemek Kitszel
` (9 subsequent siblings)
13 siblings, 0 replies; 15+ messages in thread
From: Przemek Kitszel @ 2026-10-09 12:04 UTC (permalink / raw)
To: netdev, Jakub Kicinski, Jiri Pirko
Cc: Tony Nguyen, Aleksandr Loktionov, Michal Schmidt, intel-wired-lan,
edumazet, horms, pabeni, davem, Jonathan Corbet, skhan, rdunlap,
andrew+netdev, saeedm, tariqt, leon, mbloch, jacob.e.keller,
jedrzej.jagielski, anzaki, brett.creeley, jtornosm, ohartoov,
Przemek Kitszel
Add AQ commands for RSS Global LUT allocation and free operations.
The new functions are used by subsequent commits.
Add programming code for GLOBAL LUT ID of UPDATE VSI AQ,
do the same for RSS LUT "type", also for PF LUT in case of VF VSI.
Co-developed-by: Brett Creeley <brett.creeley@intel.com>
Signed-off-by: Brett Creeley <brett.creeley@intel.com>
Reviewed-by: Aleksandr Loktionov <aleksandr.loktionov@intel.com>
Signed-off-by: Przemek Kitszel <przemyslaw.kitszel@intel.com>
---
v2: fix whitespace (Sashiko)
---
drivers/net/ethernet/intel/ice/ice.h | 1 +
.../net/ethernet/intel/ice/ice_adminq_cmd.h | 1 +
drivers/net/ethernet/intel/ice/ice_switch.h | 2 +
drivers/net/ethernet/intel/ice/ice_lib.c | 28 +++++++++++--
drivers/net/ethernet/intel/ice/ice_switch.c | 41 +++++++++++++++++++
5 files changed, 69 insertions(+), 4 deletions(-)
diff --git a/drivers/net/ethernet/intel/ice/ice.h b/drivers/net/ethernet/intel/ice/ice.h
index 432da05c48c0..89a2bdbecf60 100644
--- a/drivers/net/ethernet/intel/ice/ice.h
+++ b/drivers/net/ethernet/intel/ice/ice.h
@@ -367,6 +367,7 @@ struct ice_vsi {
u8 *rss_hkey_user; /* User configured hash keys */
u8 *rss_lut_user; /* User configured lookup table entries */
u8 rss_lut_type; /* used to configure Get/Set RSS LUT AQ call */
+ u8 global_lut_id; /* valid when lut_type == GLOBAL_LUT */
/* aRFS members only allocated for the PF VSI */
#define ICE_MAX_ARFS_LIST 1024
diff --git a/drivers/net/ethernet/intel/ice/ice_adminq_cmd.h b/drivers/net/ethernet/intel/ice/ice_adminq_cmd.h
index 42878abac9eb..90c5587642d0 100644
--- a/drivers/net/ethernet/intel/ice/ice_adminq_cmd.h
+++ b/drivers/net/ethernet/intel/ice/ice_adminq_cmd.h
@@ -169,6 +169,7 @@ struct ice_aqc_set_port_params {
#define ICE_AQC_RES_TYPE_VSI_LIST_PRUNE 0x04
#define ICE_AQC_RES_TYPE_RECIPE 0x05
#define ICE_AQC_RES_TYPE_SWID 0x07
+#define ICE_AQC_RES_TYPE_GLOBAL_RSS_HASH 0x20
#define ICE_AQC_RES_TYPE_FDIR_COUNTER_BLOCK 0x21
#define ICE_AQC_RES_TYPE_FDIR_GUARANTEED_ENTRIES 0x22
#define ICE_AQC_RES_TYPE_FDIR_SHARED_ENTRIES 0x23
diff --git a/drivers/net/ethernet/intel/ice/ice_switch.h b/drivers/net/ethernet/intel/ice/ice_switch.h
index 671d7a5f359f..99b593b68bd0 100644
--- a/drivers/net/ethernet/intel/ice/ice_switch.h
+++ b/drivers/net/ethernet/intel/ice/ice_switch.h
@@ -387,6 +387,8 @@ ice_rem_adv_rule_by_id(struct ice_hw *hw,
struct ice_rule_query_data *remove_entry);
int ice_init_def_sw_recp(struct ice_hw *hw);
+int ice_alloc_rss_global_lut(struct ice_hw *hw, u16 *global_lut_id);
+int ice_free_rss_global_lut(struct ice_hw *hw, u16 global_lut_id);
u16 ice_get_hw_vsi_num(struct ice_hw *hw, u16 vsi_handle);
int ice_replay_vsi_all_fltr(struct ice_hw *hw, u16 vsi_handle);
diff --git a/drivers/net/ethernet/intel/ice/ice_lib.c b/drivers/net/ethernet/intel/ice/ice_lib.c
index 7cd2dd066ef5..8319956754b7 100644
--- a/drivers/net/ethernet/intel/ice/ice_lib.c
+++ b/drivers/net/ethernet/intel/ice/ice_lib.c
@@ -1247,30 +1247,46 @@ static void ice_set_fd_vsi_ctx(struct ice_vsi_ctx *ctxt, struct ice_vsi *vsi)
ctxt->info.fd_report_opt = cpu_to_le16(val);
}
+/* Translate @lut_type used in most of the places to the Admin Queue
+ * Q_OPT value for RSS.
+ * Used with VSI ADD and VSI UPDATE AQs (opcodes 0x0210, 0x0211).
+ */
+static u8 ice_lut_type_to_aq_qopt_rss_val(enum ice_lut_type lut_type)
+{
+ switch (lut_type) {
+ case ICE_LUT_PF:
+ return ICE_AQ_VSI_Q_OPT_RSS_LUT_PF;
+ case ICE_LUT_GLOBAL:
+ return ICE_AQ_VSI_Q_OPT_RSS_LUT_GBL;
+ case ICE_LUT_VSI:
+ default:
+ return ICE_AQ_VSI_Q_OPT_RSS_LUT_VSI;
+ }
+}
+
/**
* ice_set_rss_vsi_ctx - Set RSS VSI context before adding a VSI
* @ctxt: the VSI context being set
* @vsi: the VSI being configured
*/
static void ice_set_rss_vsi_ctx(struct ice_vsi_ctx *ctxt, struct ice_vsi *vsi)
{
u8 lut_type, hash_type;
+ u8 global_lut_id = 0;
struct device *dev;
struct ice_pf *pf;
pf = vsi->back;
dev = ice_pf_to_dev(pf);
switch (vsi->type) {
case ICE_VSI_CHNL:
- case ICE_VSI_PF:
- /* PF VSI will inherit RSS instance of PF */
lut_type = ICE_AQ_VSI_Q_OPT_RSS_LUT_PF;
break;
+ case ICE_VSI_PF:
case ICE_VSI_VF:
case ICE_VSI_SF:
- /* VF VSI will gets a small RSS table which is a VSI LUT type */
- lut_type = ICE_AQ_VSI_Q_OPT_RSS_LUT_VSI;
+ lut_type = ice_lut_type_to_aq_qopt_rss_val(vsi->rss_lut_type);
break;
default:
dev_dbg(dev, "Unsupported VSI type %s\n",
@@ -1281,8 +1297,12 @@ static void ice_set_rss_vsi_ctx(struct ice_vsi_ctx *ctxt, struct ice_vsi *vsi)
hash_type = ICE_AQ_VSI_Q_OPT_RSS_HASH_TPLZ;
vsi->rss_hfunc = hash_type;
+ if (vsi->rss_lut_type == ICE_LUT_GLOBAL)
+ global_lut_id = vsi->global_lut_id;
+
ctxt->info.q_opt_rss =
FIELD_PREP(ICE_AQ_VSI_Q_OPT_RSS_LUT_M, lut_type) |
+ FIELD_PREP(ICE_AQ_VSI_Q_OPT_RSS_GBL_LUT_M, global_lut_id) |
FIELD_PREP(ICE_AQ_VSI_Q_OPT_RSS_HASH_M, hash_type);
}
diff --git a/drivers/net/ethernet/intel/ice/ice_switch.c b/drivers/net/ethernet/intel/ice/ice_switch.c
index 4daee252e2a6..e6f08fe9b5c2 100644
--- a/drivers/net/ethernet/intel/ice/ice_switch.c
+++ b/drivers/net/ethernet/intel/ice/ice_switch.c
@@ -1527,6 +1527,47 @@ ice_aq_get_sw_cfg(struct ice_hw *hw, struct ice_aqc_get_sw_cfg_resp_elem *buf,
return status;
}
+/* Allocate a new Global LUT for the caller.
+ * LUT ID is returned via @global_lut_id.
+ */
+int ice_alloc_rss_global_lut(struct ice_hw *hw, u16 *global_lut_id)
+{
+ DEFINE_RAW_FLEX(struct ice_aqc_alloc_free_res_elem, buf, elem, 1);
+ u16 buf_len = __struct_size(buf);
+ int err;
+
+ buf->num_elems = cpu_to_le16(1);
+ buf->res_type = cpu_to_le16(ICE_AQC_RES_TYPE_GLOBAL_RSS_HASH);
+
+ err = ice_aq_alloc_free_res(hw, buf, buf_len, ice_aqc_opc_alloc_res);
+ if (err)
+ ice_debug(hw, ICE_DBG_RES, "Failed to allocate RSS global LUT, err %d\n",
+ err);
+ else
+ *global_lut_id = le16_to_cpu(buf->elem[0].e.sw_resp);
+
+ return err;
+}
+
+/* Free Global LUT at @global_lut_id. */
+int ice_free_rss_global_lut(struct ice_hw *hw, u16 global_lut_id)
+{
+ DEFINE_RAW_FLEX(struct ice_aqc_alloc_free_res_elem, buf, elem, 1);
+ u16 buf_len = __struct_size(buf);
+ int err;
+
+ buf->num_elems = cpu_to_le16(1);
+ buf->res_type = cpu_to_le16(ICE_AQC_RES_TYPE_GLOBAL_RSS_HASH);
+ buf->elem[0].e.sw_resp = cpu_to_le16(global_lut_id);
+
+ err = ice_aq_alloc_free_res(hw, buf, buf_len, ice_aqc_opc_free_res);
+ if (err)
+ ice_debug(hw, ICE_DBG_RES, "Failed to free RSS global LUT %d, err %d\n",
+ global_lut_id, err);
+
+ return err;
+}
+
/**
* ice_aq_add_vsi
* @hw: pointer to the HW struct
--
2.51.1
^ permalink raw reply related [flat|nested] 15+ messages in thread
* [PATCH net-next v2 05/14] ice: rename ICE_MAX_RSS_QS_PER_VF to ICE_MAX_QS_PER_VF_VCV1
2026-10-09 12:04 [PATCH net-next v2 00/14] devlink, mlx5, iavf, ice: XLVF for iavf Przemek Kitszel
` (3 preceding siblings ...)
2026-10-09 12:04 ` [PATCH net-next v2 04/14] ice: add helpers for Global RSS LUT alloc, free, vsi_update Przemek Kitszel
@ 2026-10-09 12:04 ` Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 06/14] ice: bump to 256qs for VF Przemek Kitszel
` (8 subsequent siblings)
13 siblings, 0 replies; 15+ messages in thread
From: Przemek Kitszel @ 2026-10-09 12:04 UTC (permalink / raw)
To: netdev, Jakub Kicinski, Jiri Pirko
Cc: Tony Nguyen, Aleksandr Loktionov, Michal Schmidt, intel-wired-lan,
edumazet, horms, pabeni, davem, Jonathan Corbet, skhan, rdunlap,
andrew+netdev, saeedm, tariqt, leon, mbloch, jacob.e.keller,
jedrzej.jagielski, anzaki, brett.creeley, jtornosm, ohartoov,
Przemek Kitszel
Rename ICE_MAX_RSS_QS_PER_VF to ICE_MAX_QS_PER_VF_VCV1, in preparation for
the next patch that will extend the max to 256, using old value of 16 for
the "v1" variant of virtchnl opcodes.
Suggested-by: Jedrzej Jagielski <jedrzej.jagielski@intel.com>
Suggested-by: Jacob Keller <jacob.e.keller@intel.com>
Reviewed-by: Aleksandr Loktionov <aleksandr.loktionov@intel.com>
Signed-off-by: Przemek Kitszel <przemyslaw.kitszel@intel.com>
---
drivers/net/ethernet/intel/ice/ice_lag.h | 2 +-
drivers/net/ethernet/intel/ice/ice_vf_lib.h | 9 +++---
drivers/net/ethernet/intel/ice/ice_lib.c | 2 +-
drivers/net/ethernet/intel/ice/ice_sriov.c | 4 +--
drivers/net/ethernet/intel/ice/ice_vf_lib.c | 12 ++++----
drivers/net/ethernet/intel/ice/virt/queues.c | 30 ++++++++++----------
6 files changed, 30 insertions(+), 29 deletions(-)
diff --git a/drivers/net/ethernet/intel/ice/ice_lag.h b/drivers/net/ethernet/intel/ice/ice_lag.h
index f77ebcd61042..4bfffecbdc97 100644
--- a/drivers/net/ethernet/intel/ice/ice_lag.h
+++ b/drivers/net/ethernet/intel/ice/ice_lag.h
@@ -52,7 +52,7 @@ struct ice_lag {
u8 bond_lport_sec; /* lport values for secondary PF */
/* q_home keeps track of which interface the q is currently on */
- u8 q_home[ICE_MAX_SRIOV_VFS][ICE_MAX_RSS_QS_PER_VF];
+ u8 q_home[ICE_MAX_SRIOV_VFS][ICE_MAX_QS_PER_VF_VCV1];
/* placeholder VSI for hanging VF queues from on secondary interface */
struct ice_vsi *sec_vf[ICE_MAX_SRIOV_VFS];
diff --git a/drivers/net/ethernet/intel/ice/ice_vf_lib.h b/drivers/net/ethernet/intel/ice/ice_vf_lib.h
index fa436b3b1eac..4959ecee0c21 100644
--- a/drivers/net/ethernet/intel/ice/ice_vf_lib.h
+++ b/drivers/net/ethernet/intel/ice/ice_vf_lib.h
@@ -19,7 +19,8 @@
#define ICE_MAX_SRIOV_VFS 256
/* VF resource constraints */
-#define ICE_MAX_RSS_QS_PER_VF 16
+/* for "old" virtchnl opcodes that accept up to 16 queues */
+#define ICE_MAX_QS_PER_VF_VCV1 16
struct ice_pf;
struct ice_vf;
@@ -161,8 +162,8 @@ struct ice_vf {
u8 dev_lan_addr[ETH_ALEN];
u8 hw_lan_addr[ETH_ALEN];
struct ice_time_mac legacy_last_added_umac;
- DECLARE_BITMAP(txq_ena, ICE_MAX_RSS_QS_PER_VF);
- DECLARE_BITMAP(rxq_ena, ICE_MAX_RSS_QS_PER_VF);
+ DECLARE_BITMAP(txq_ena, ICE_MAX_QS_PER_VF_VCV1);
+ DECLARE_BITMAP(rxq_ena, ICE_MAX_QS_PER_VF_VCV1);
struct ice_vlan port_vlan_info; /* Port VLAN ID, QoS, and TPID */
struct virtchnl_vlan_caps vlan_v2_caps;
struct ice_mbx_vf_info mbx_info;
@@ -205,7 +206,7 @@ struct ice_vf {
u16 lldp_recipe_id;
u16 lldp_rule_id;
- struct ice_vf_qs_bw qs_bw[ICE_MAX_RSS_QS_PER_VF];
+ struct ice_vf_qs_bw qs_bw[ICE_MAX_QS_PER_VF_VCV1];
};
/* Flags for controlling behavior of ice_reset_vf */
diff --git a/drivers/net/ethernet/intel/ice/ice_lib.c b/drivers/net/ethernet/intel/ice/ice_lib.c
index 8319956754b7..e18128bce959 100644
--- a/drivers/net/ethernet/intel/ice/ice_lib.c
+++ b/drivers/net/ethernet/intel/ice/ice_lib.c
@@ -1026,7 +1026,7 @@ static void ice_vsi_set_rss_params(struct ice_vsi *vsi)
* For VSI_LUT, LUT size should be set to 64 bytes.
*/
vsi->rss_table_size = ICE_LUT_VSI_SIZE;
- vsi->rss_size = ICE_MAX_RSS_QS_PER_VF;
+ vsi->rss_size = ICE_MAX_QS_PER_VF_VCV1;
vsi->rss_lut_type = ICE_LUT_VSI;
break;
case ICE_VSI_LB:
diff --git a/drivers/net/ethernet/intel/ice/ice_sriov.c b/drivers/net/ethernet/intel/ice/ice_sriov.c
index c49f48b5ae38..dcd6d48ea38d 100644
--- a/drivers/net/ethernet/intel/ice/ice_sriov.c
+++ b/drivers/net/ethernet/intel/ice/ice_sriov.c
@@ -398,15 +398,15 @@ static int ice_set_per_vf_res(struct ice_pf *pf, u16 num_vfs)
}
num_txq = min_t(u16, num_msix_per_vf - ICE_NONQ_VECS_VF,
- ICE_MAX_RSS_QS_PER_VF);
+ ICE_MAX_QS_PER_VF_VCV1);
avail_qs = ice_get_avail_txq_count(pf) / num_vfs;
if (!avail_qs)
num_txq = 0;
else if (num_txq > avail_qs)
num_txq = rounddown_pow_of_two(avail_qs);
num_rxq = min_t(u16, num_msix_per_vf - ICE_NONQ_VECS_VF,
- ICE_MAX_RSS_QS_PER_VF);
+ ICE_MAX_QS_PER_VF_VCV1);
avail_qs = ice_get_avail_rxq_count(pf) / num_vfs;
if (!avail_qs)
num_rxq = 0;
diff --git a/drivers/net/ethernet/intel/ice/ice_vf_lib.c b/drivers/net/ethernet/intel/ice/ice_vf_lib.c
index a54cb2b8d3c7..be87a83d2599 100644
--- a/drivers/net/ethernet/intel/ice/ice_vf_lib.c
+++ b/drivers/net/ethernet/intel/ice/ice_vf_lib.c
@@ -528,8 +528,8 @@ static void ice_vf_rebuild_host_cfg(struct ice_vf *vf)
static void ice_set_vf_state_qs_dis(struct ice_vf *vf)
{
/* Clear Rx/Tx enabled queues flag */
- bitmap_zero(vf->txq_ena, ICE_MAX_RSS_QS_PER_VF);
- bitmap_zero(vf->rxq_ena, ICE_MAX_RSS_QS_PER_VF);
+ bitmap_zero(vf->txq_ena, ICE_MAX_QS_PER_VF_VCV1);
+ bitmap_zero(vf->rxq_ena, ICE_MAX_QS_PER_VF_VCV1);
clear_bit(ICE_VF_STATE_QS_ENA, vf->vf_states);
}
@@ -1238,13 +1238,13 @@ bool ice_is_vf_trusted(struct ice_vf *vf)
* ice_vf_has_no_qs_ena - check if the VF has any Rx or Tx queues enabled
* @vf: the VF to check
*
- * Returns true if the VF has no Rx and no Tx queues enabled and returns false
- * otherwise
+ * Return: true if the VF has no Rx and no Tx queues enabled and returns false
+ * otherwise.
*/
bool ice_vf_has_no_qs_ena(struct ice_vf *vf)
{
- return bitmap_empty(vf->rxq_ena, ICE_MAX_RSS_QS_PER_VF) &&
- bitmap_empty(vf->txq_ena, ICE_MAX_RSS_QS_PER_VF);
+ return bitmap_empty(vf->rxq_ena, ICE_MAX_QS_PER_VF_VCV1) &&
+ bitmap_empty(vf->txq_ena, ICE_MAX_QS_PER_VF_VCV1);
}
/**
diff --git a/drivers/net/ethernet/intel/ice/virt/queues.c b/drivers/net/ethernet/intel/ice/virt/queues.c
index a22780575086..2667634e3c8c 100644
--- a/drivers/net/ethernet/intel/ice/virt/queues.c
+++ b/drivers/net/ethernet/intel/ice/virt/queues.c
@@ -171,8 +171,8 @@ static int ice_vf_cfg_q_quanta_profile(struct ice_vf *vf, u16 quanta_size,
static bool ice_vc_validate_vqs_bitmaps(struct virtchnl_queue_select *vqs)
{
if ((!vqs->rx_queues && !vqs->tx_queues) ||
- vqs->rx_queues >= BIT(ICE_MAX_RSS_QS_PER_VF) ||
- vqs->tx_queues >= BIT(ICE_MAX_RSS_QS_PER_VF))
+ vqs->rx_queues >= BIT(ICE_MAX_QS_PER_VF_VCV1) ||
+ vqs->tx_queues >= BIT(ICE_MAX_QS_PER_VF_VCV1))
return false;
return true;
@@ -335,7 +335,7 @@ int ice_vc_ena_qs_msg(struct ice_vf *vf, u8 *msg)
* programmed using ice_vsi_cfg_txqs
*/
q_map = vqs->rx_queues;
- for_each_set_bit(vf_q_id, &q_map, ICE_MAX_RSS_QS_PER_VF) {
+ for_each_set_bit(vf_q_id, &q_map, ICE_MAX_QS_PER_VF_VCV1) {
if (!ice_vc_isvalid_q_id(vsi, vf_q_id)) {
v_ret = VIRTCHNL_STATUS_ERR_PARAM;
goto error_param;
@@ -348,7 +348,7 @@ int ice_vc_ena_qs_msg(struct ice_vf *vf, u8 *msg)
}
q_map = vqs->tx_queues;
- for_each_set_bit(vf_q_id, &q_map, ICE_MAX_RSS_QS_PER_VF) {
+ for_each_set_bit(vf_q_id, &q_map, ICE_MAX_QS_PER_VF_VCV1) {
if (!ice_vc_isvalid_q_id(vsi, vf_q_id)) {
v_ret = VIRTCHNL_STATUS_ERR_PARAM;
goto error_param;
@@ -484,7 +484,7 @@ int ice_vc_dis_qs_msg(struct ice_vf *vf, u8 *msg)
if (vqs->tx_queues) {
q_map = vqs->tx_queues;
- for_each_set_bit(vf_q_id, &q_map, ICE_MAX_RSS_QS_PER_VF) {
+ for_each_set_bit(vf_q_id, &q_map, ICE_MAX_QS_PER_VF_VCV1) {
if (!ice_vc_isvalid_q_id(vsi, vf_q_id)) {
v_ret = VIRTCHNL_STATUS_ERR_PARAM;
goto error_param;
@@ -499,7 +499,7 @@ int ice_vc_dis_qs_msg(struct ice_vf *vf, u8 *msg)
q_map = vqs->rx_queues;
if (q_map) {
- for_each_set_bit(vf_q_id, &q_map, ICE_MAX_RSS_QS_PER_VF) {
+ for_each_set_bit(vf_q_id, &q_map, ICE_MAX_QS_PER_VF_VCV1) {
if (!ice_vc_isvalid_q_id(vsi, vf_q_id)) {
v_ret = VIRTCHNL_STATUS_ERR_PARAM;
goto error_param;
@@ -542,7 +542,7 @@ ice_cfg_interrupt(struct ice_vf *vf, struct ice_vsi *vsi,
q_vector->num_ring_tx = 0;
qmap = map->rxq_map;
- for_each_set_bit(vsi_q_id_idx, &qmap, ICE_MAX_RSS_QS_PER_VF) {
+ for_each_set_bit(vsi_q_id_idx, &qmap, ICE_MAX_QS_PER_VF_VCV1) {
vsi_q_id = vsi_q_id_idx;
if (!ice_vc_isvalid_q_id(vsi, vsi_q_id))
@@ -557,7 +557,7 @@ ice_cfg_interrupt(struct ice_vf *vf, struct ice_vsi *vsi,
}
qmap = map->txq_map;
- for_each_set_bit(vsi_q_id_idx, &qmap, ICE_MAX_RSS_QS_PER_VF) {
+ for_each_set_bit(vsi_q_id_idx, &qmap, ICE_MAX_QS_PER_VF_VCV1) {
vsi_q_id = vsi_q_id_idx;
if (!ice_vc_isvalid_q_id(vsi, vsi_q_id))
@@ -681,7 +681,7 @@ int ice_vc_cfg_q_bw(struct ice_vf *vf, u8 *msg)
goto err;
}
- if (qbw->num_queues > ICE_MAX_RSS_QS_PER_VF ||
+ if (qbw->num_queues > ICE_MAX_QS_PER_VF_VCV1 ||
qbw->num_queues > min_t(u16, vsi->alloc_txq, vsi->alloc_rxq)) {
dev_err(ice_pf_to_dev(vf->pf), "VF-%d trying to configure more than allocated number of queues: %d\n",
vf->vf_id, min_t(u16, vsi->alloc_txq, vsi->alloc_rxq));
@@ -773,7 +773,7 @@ int ice_vc_cfg_q_quanta(struct ice_vf *vf, u8 *msg)
goto err;
}
- if (end_qid > ICE_MAX_RSS_QS_PER_VF ||
+ if (end_qid > ICE_MAX_QS_PER_VF_VCV1 ||
end_qid > min_t(u16, vsi->alloc_txq, vsi->alloc_rxq)) {
dev_err(ice_pf_to_dev(vf->pf), "VF-%d trying to configure more than allocated number of queues: %d\n",
vf->vf_id, min_t(u16, vsi->alloc_txq, vsi->alloc_rxq));
@@ -841,7 +841,7 @@ int ice_vc_cfg_qs_msg(struct ice_vf *vf, u8 *msg)
if (!vsi)
goto error_param;
- if (qci->num_queue_pairs > ICE_MAX_RSS_QS_PER_VF ||
+ if (qci->num_queue_pairs > ICE_MAX_QS_PER_VF_VCV1 ||
qci->num_queue_pairs > min_t(u16, vsi->alloc_txq, vsi->alloc_rxq)) {
dev_err(ice_pf_to_dev(pf), "VF-%d requesting more than supported number of queues: %d\n",
vf->vf_id, min_t(u16, vsi->alloc_txq, vsi->alloc_rxq));
@@ -1019,16 +1019,16 @@ int ice_vc_request_qs_msg(struct ice_vf *vf, u8 *msg)
if (!req_queues) {
dev_err(dev, "VF %d tried to request 0 queues. Ignoring.\n",
vf->vf_id);
- } else if (req_queues > ICE_MAX_RSS_QS_PER_VF) {
+ } else if (req_queues > ICE_MAX_QS_PER_VF_VCV1) {
dev_err(dev, "VF %d tried to request more than %d queues.\n",
- vf->vf_id, ICE_MAX_RSS_QS_PER_VF);
- vfres->num_queue_pairs = ICE_MAX_RSS_QS_PER_VF;
+ vf->vf_id, ICE_MAX_QS_PER_VF_VCV1);
+ vfres->num_queue_pairs = ICE_MAX_QS_PER_VF_VCV1;
} else if (req_queues > cur_queues &&
req_queues - cur_queues > tx_rx_queue_left) {
dev_warn(dev, "VF %d requested %u more queues, but only %u left.\n",
vf->vf_id, req_queues - cur_queues, tx_rx_queue_left);
vfres->num_queue_pairs = min_t(u16, max_allowed_vf_queues,
- ICE_MAX_RSS_QS_PER_VF);
+ ICE_MAX_QS_PER_VF_VCV1);
} else {
/* request is successful, then reset VF */
vf->num_req_qs = req_queues;
--
2.51.1
^ permalink raw reply related [flat|nested] 15+ messages in thread
* [PATCH net-next v2 06/14] ice: bump to 256qs for VF
2026-10-09 12:04 [PATCH net-next v2 00/14] devlink, mlx5, iavf, ice: XLVF for iavf Przemek Kitszel
` (4 preceding siblings ...)
2026-10-09 12:04 ` [PATCH net-next v2 05/14] ice: rename ICE_MAX_RSS_QS_PER_VF to ICE_MAX_QS_PER_VF_VCV1 Przemek Kitszel
@ 2026-10-09 12:04 ` Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 07/14] iavf: extend iavf_configure_queues() to support more queues Przemek Kitszel
` (7 subsequent siblings)
13 siblings, 0 replies; 15+ messages in thread
From: Przemek Kitszel @ 2026-10-09 12:04 UTC (permalink / raw)
To: netdev, Jakub Kicinski, Jiri Pirko
Cc: Tony Nguyen, Aleksandr Loktionov, Michal Schmidt, intel-wired-lan,
edumazet, horms, pabeni, davem, Jonathan Corbet, skhan, rdunlap,
andrew+netdev, saeedm, tariqt, leon, mbloch, jacob.e.keller,
jedrzej.jagielski, anzaki, brett.creeley, jtornosm, ohartoov,
Przemek Kitszel
Adjust all bitmaps and arrays in ice to accept 256 VF queues.
Extend struct ice_vf::num_req_qs width to allow 256 queues.
Keep old/legacy size for virtchnl opcodes that were designed to accept
only up to 16 queues.
Raise also the limits for the default VF queue count and for
VIRTCHNL_OP_REQUEST_QUEUES to 256. New virtchnl opcodes, needed to
make use of more than 16 queues, are handled by subsequent commits
of this series, the work is split for easier review.
Reviewed-by: Jedrzej Jagielski <jedrzej.jagielski@intel.com>
Signed-off-by: Przemek Kitszel <przemyslaw.kitszel@intel.com>
---
v2:
* change ice_lag allocations to kvzalloc(), as they have grown to
66KB (Sashiko)
* commit message: mention that the default and requested VF queue
limits are raised already here, while more than 16 queues are
enabled by subsequent commits
Sashiko suggestions I don't want to apply:
S: there is an existing error handling issue in ice_init_lag()
me: pre-existing and unrelated, will be fixed separately
---
drivers/net/ethernet/intel/ice/ice_lag.h | 2 +-
drivers/net/ethernet/intel/ice/ice_vf_lib.h | 9 +++++----
drivers/net/ethernet/intel/ice/ice_lag.c | 6 +++---
drivers/net/ethernet/intel/ice/ice_lib.c | 2 +-
drivers/net/ethernet/intel/ice/ice_sriov.c | 4 ++--
drivers/net/ethernet/intel/ice/ice_vf_lib.c | 8 ++++----
drivers/net/ethernet/intel/ice/virt/queues.c | 14 +++++++-------
7 files changed, 23 insertions(+), 22 deletions(-)
diff --git a/drivers/net/ethernet/intel/ice/ice_lag.h b/drivers/net/ethernet/intel/ice/ice_lag.h
index 4bfffecbdc97..39f6a6cc844d 100644
--- a/drivers/net/ethernet/intel/ice/ice_lag.h
+++ b/drivers/net/ethernet/intel/ice/ice_lag.h
@@ -52,7 +52,7 @@ struct ice_lag {
u8 bond_lport_sec; /* lport values for secondary PF */
/* q_home keeps track of which interface the q is currently on */
- u8 q_home[ICE_MAX_SRIOV_VFS][ICE_MAX_QS_PER_VF_VCV1];
+ u8 q_home[ICE_MAX_SRIOV_VFS][ICE_MAX_QS_PER_VF];
/* placeholder VSI for hanging VF queues from on secondary interface */
struct ice_vsi *sec_vf[ICE_MAX_SRIOV_VFS];
diff --git a/drivers/net/ethernet/intel/ice/ice_vf_lib.h b/drivers/net/ethernet/intel/ice/ice_vf_lib.h
index 4959ecee0c21..a9e95c85bc98 100644
--- a/drivers/net/ethernet/intel/ice/ice_vf_lib.h
+++ b/drivers/net/ethernet/intel/ice/ice_vf_lib.h
@@ -21,6 +21,7 @@
/* VF resource constraints */
/* for "old" virtchnl opcodes that accept up to 16 queues */
#define ICE_MAX_QS_PER_VF_VCV1 16
+#define ICE_MAX_QS_PER_VF 256
struct ice_pf;
struct ice_vf;
@@ -162,8 +163,8 @@ struct ice_vf {
u8 dev_lan_addr[ETH_ALEN];
u8 hw_lan_addr[ETH_ALEN];
struct ice_time_mac legacy_last_added_umac;
- DECLARE_BITMAP(txq_ena, ICE_MAX_QS_PER_VF_VCV1);
- DECLARE_BITMAP(rxq_ena, ICE_MAX_QS_PER_VF_VCV1);
+ DECLARE_BITMAP(txq_ena, ICE_MAX_QS_PER_VF);
+ DECLARE_BITMAP(rxq_ena, ICE_MAX_QS_PER_VF);
struct ice_vlan port_vlan_info; /* Port VLAN ID, QoS, and TPID */
struct virtchnl_vlan_caps vlan_v2_caps;
struct ice_mbx_vf_info mbx_info;
@@ -185,7 +186,7 @@ struct ice_vf {
DECLARE_BITMAP(vf_states, ICE_VF_STATES_NBITS); /* VF runtime states */
unsigned long vf_caps; /* VF's adv. capabilities */
- u8 num_req_qs; /* num of queue pairs requested by VF */
+ u16 num_req_qs; /* num of queue pairs requested by VF */
u16 num_mac;
u16 num_mac_lldp;
u16 num_vf_qs; /* num of queue configured per VF */
@@ -206,7 +207,7 @@ struct ice_vf {
u16 lldp_recipe_id;
u16 lldp_rule_id;
- struct ice_vf_qs_bw qs_bw[ICE_MAX_QS_PER_VF_VCV1];
+ struct ice_vf_qs_bw qs_bw[ICE_MAX_QS_PER_VF];
};
/* Flags for controlling behavior of ice_reset_vf */
diff --git a/drivers/net/ethernet/intel/ice/ice_lag.c b/drivers/net/ethernet/intel/ice/ice_lag.c
index 08a17ded0ad5..e006311c18b6 100644
--- a/drivers/net/ethernet/intel/ice/ice_lag.c
+++ b/drivers/net/ethernet/intel/ice/ice_lag.c
@@ -2577,7 +2577,7 @@ int ice_init_lag(struct ice_pf *pf)
if (!ice_is_feature_supported(pf, ICE_F_SRIOV_LAG))
return 0;
- pf->lag = kzalloc_obj(*lag);
+ pf->lag = kvzalloc_obj(*lag);
if (!pf->lag)
return -ENOMEM;
lag = pf->lag;
@@ -2651,7 +2651,7 @@ int ice_init_lag(struct ice_pf *pf)
ice_free_hw_res(&pf->hw, ICE_AQC_RES_TYPE_RECIPE, 1,
&lag->pf_recipe);
lag_error:
- kfree(lag);
+ kvfree(lag);
pf->lag = NULL;
return err;
}
@@ -2680,7 +2680,7 @@ void ice_deinit_lag(struct ice_pf *pf)
ice_free_hw_res(&pf->hw, ICE_AQC_RES_TYPE_RECIPE, 1,
&pf->lag->lport_recipe);
- kfree(lag);
+ kvfree(lag);
pf->lag = NULL;
}
diff --git a/drivers/net/ethernet/intel/ice/ice_lib.c b/drivers/net/ethernet/intel/ice/ice_lib.c
index e18128bce959..f81b310a98c3 100644
--- a/drivers/net/ethernet/intel/ice/ice_lib.c
+++ b/drivers/net/ethernet/intel/ice/ice_lib.c
@@ -1026,7 +1026,7 @@ static void ice_vsi_set_rss_params(struct ice_vsi *vsi)
* For VSI_LUT, LUT size should be set to 64 bytes.
*/
vsi->rss_table_size = ICE_LUT_VSI_SIZE;
- vsi->rss_size = ICE_MAX_QS_PER_VF_VCV1;
+ vsi->rss_size = ICE_MAX_QS_PER_VF;
vsi->rss_lut_type = ICE_LUT_VSI;
break;
case ICE_VSI_LB:
diff --git a/drivers/net/ethernet/intel/ice/ice_sriov.c b/drivers/net/ethernet/intel/ice/ice_sriov.c
index dcd6d48ea38d..d7c4a6540008 100644
--- a/drivers/net/ethernet/intel/ice/ice_sriov.c
+++ b/drivers/net/ethernet/intel/ice/ice_sriov.c
@@ -398,15 +398,15 @@ static int ice_set_per_vf_res(struct ice_pf *pf, u16 num_vfs)
}
num_txq = min_t(u16, num_msix_per_vf - ICE_NONQ_VECS_VF,
- ICE_MAX_QS_PER_VF_VCV1);
+ ICE_MAX_QS_PER_VF);
avail_qs = ice_get_avail_txq_count(pf) / num_vfs;
if (!avail_qs)
num_txq = 0;
else if (num_txq > avail_qs)
num_txq = rounddown_pow_of_two(avail_qs);
num_rxq = min_t(u16, num_msix_per_vf - ICE_NONQ_VECS_VF,
- ICE_MAX_QS_PER_VF_VCV1);
+ ICE_MAX_QS_PER_VF);
avail_qs = ice_get_avail_rxq_count(pf) / num_vfs;
if (!avail_qs)
num_rxq = 0;
diff --git a/drivers/net/ethernet/intel/ice/ice_vf_lib.c b/drivers/net/ethernet/intel/ice/ice_vf_lib.c
index be87a83d2599..320b393ddaaf 100644
--- a/drivers/net/ethernet/intel/ice/ice_vf_lib.c
+++ b/drivers/net/ethernet/intel/ice/ice_vf_lib.c
@@ -528,8 +528,8 @@ static void ice_vf_rebuild_host_cfg(struct ice_vf *vf)
static void ice_set_vf_state_qs_dis(struct ice_vf *vf)
{
/* Clear Rx/Tx enabled queues flag */
- bitmap_zero(vf->txq_ena, ICE_MAX_QS_PER_VF_VCV1);
- bitmap_zero(vf->rxq_ena, ICE_MAX_QS_PER_VF_VCV1);
+ bitmap_zero(vf->txq_ena, ICE_MAX_QS_PER_VF);
+ bitmap_zero(vf->rxq_ena, ICE_MAX_QS_PER_VF);
clear_bit(ICE_VF_STATE_QS_ENA, vf->vf_states);
}
@@ -1243,8 +1243,8 @@ bool ice_is_vf_trusted(struct ice_vf *vf)
*/
bool ice_vf_has_no_qs_ena(struct ice_vf *vf)
{
- return bitmap_empty(vf->rxq_ena, ICE_MAX_QS_PER_VF_VCV1) &&
- bitmap_empty(vf->txq_ena, ICE_MAX_QS_PER_VF_VCV1);
+ return bitmap_empty(vf->rxq_ena, ICE_MAX_QS_PER_VF) &&
+ bitmap_empty(vf->txq_ena, ICE_MAX_QS_PER_VF);
}
/**
diff --git a/drivers/net/ethernet/intel/ice/virt/queues.c b/drivers/net/ethernet/intel/ice/virt/queues.c
index 2667634e3c8c..61304493af34 100644
--- a/drivers/net/ethernet/intel/ice/virt/queues.c
+++ b/drivers/net/ethernet/intel/ice/virt/queues.c
@@ -681,7 +681,7 @@ int ice_vc_cfg_q_bw(struct ice_vf *vf, u8 *msg)
goto err;
}
- if (qbw->num_queues > ICE_MAX_QS_PER_VF_VCV1 ||
+ if (qbw->num_queues > ICE_MAX_QS_PER_VF ||
qbw->num_queues > min_t(u16, vsi->alloc_txq, vsi->alloc_rxq)) {
dev_err(ice_pf_to_dev(vf->pf), "VF-%d trying to configure more than allocated number of queues: %d\n",
vf->vf_id, min_t(u16, vsi->alloc_txq, vsi->alloc_rxq));
@@ -773,7 +773,7 @@ int ice_vc_cfg_q_quanta(struct ice_vf *vf, u8 *msg)
goto err;
}
- if (end_qid > ICE_MAX_QS_PER_VF_VCV1 ||
+ if (end_qid > ICE_MAX_QS_PER_VF ||
end_qid > min_t(u16, vsi->alloc_txq, vsi->alloc_rxq)) {
dev_err(ice_pf_to_dev(vf->pf), "VF-%d trying to configure more than allocated number of queues: %d\n",
vf->vf_id, min_t(u16, vsi->alloc_txq, vsi->alloc_rxq));
@@ -841,7 +841,7 @@ int ice_vc_cfg_qs_msg(struct ice_vf *vf, u8 *msg)
if (!vsi)
goto error_param;
- if (qci->num_queue_pairs > ICE_MAX_QS_PER_VF_VCV1 ||
+ if (qci->num_queue_pairs > ICE_MAX_QS_PER_VF ||
qci->num_queue_pairs > min_t(u16, vsi->alloc_txq, vsi->alloc_rxq)) {
dev_err(ice_pf_to_dev(pf), "VF-%d requesting more than supported number of queues: %d\n",
vf->vf_id, min_t(u16, vsi->alloc_txq, vsi->alloc_rxq));
@@ -1019,16 +1019,16 @@ int ice_vc_request_qs_msg(struct ice_vf *vf, u8 *msg)
if (!req_queues) {
dev_err(dev, "VF %d tried to request 0 queues. Ignoring.\n",
vf->vf_id);
- } else if (req_queues > ICE_MAX_QS_PER_VF_VCV1) {
+ } else if (req_queues > ICE_MAX_QS_PER_VF) {
dev_err(dev, "VF %d tried to request more than %d queues.\n",
- vf->vf_id, ICE_MAX_QS_PER_VF_VCV1);
- vfres->num_queue_pairs = ICE_MAX_QS_PER_VF_VCV1;
+ vf->vf_id, ICE_MAX_QS_PER_VF);
+ vfres->num_queue_pairs = ICE_MAX_QS_PER_VF;
} else if (req_queues > cur_queues &&
req_queues - cur_queues > tx_rx_queue_left) {
dev_warn(dev, "VF %d requested %u more queues, but only %u left.\n",
vf->vf_id, req_queues - cur_queues, tx_rx_queue_left);
vfres->num_queue_pairs = min_t(u16, max_allowed_vf_queues,
- ICE_MAX_QS_PER_VF_VCV1);
+ ICE_MAX_QS_PER_VF);
} else {
/* request is successful, then reset VF */
vf->num_req_qs = req_queues;
--
2.51.1
^ permalink raw reply related [flat|nested] 15+ messages in thread
* [PATCH net-next v2 07/14] iavf: extend iavf_configure_queues() to support more queues
2026-10-09 12:04 [PATCH net-next v2 00/14] devlink, mlx5, iavf, ice: XLVF for iavf Przemek Kitszel
` (5 preceding siblings ...)
2026-10-09 12:04 ` [PATCH net-next v2 06/14] ice: bump to 256qs for VF Przemek Kitszel
@ 2026-10-09 12:04 ` Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 08/14] iavf: temporary rename of IAVF_MAX_REQ_QUEUES to IAVF_MAX_REQ_QUEUES_VCV1 Przemek Kitszel
` (6 subsequent siblings)
13 siblings, 0 replies; 15+ messages in thread
From: Przemek Kitszel @ 2026-10-09 12:04 UTC (permalink / raw)
To: netdev, Jakub Kicinski, Jiri Pirko
Cc: Tony Nguyen, Aleksandr Loktionov, Michal Schmidt, intel-wired-lan,
edumazet, horms, pabeni, davem, Jonathan Corbet, skhan, rdunlap,
andrew+netdev, saeedm, tariqt, leon, mbloch, jacob.e.keller,
jedrzej.jagielski, anzaki, brett.creeley, jtornosm, ohartoov,
Przemek Kitszel
Extend iavf_configure_queues() to support more queue pairs than fit
into a single virtchnl message (62).
Virtchnl opcode used was already generic, but we have just not needed
more than one message before.
Send the messages one by one, waiting for the PF reply to each of them
with iavf_poll_virtchnl_response(), similar to iavf_set_mac_sync().
That keeps only one virtchnl op in flight, without the need to track
across watchdog runs what was already sent (as iavf_add_ether_addrs()
does for MAC filters). So the op becomes synchronous also for the
single message case: the watchdog, holding netdev_lock, waits up to 1s
for each reply.
Stop at the first PF NACK, without a retry, as before. On timeout keep
the op pending, as PF may still reply; once it does, the whole
configuration is resent.
Add helper, iavf_max_vc_entries(), that determines how many entries we
could fit into virtchnl message (ending by flex array member). Will be
also used on another op later in this series.
Signed-off-by: Przemek Kitszel <przemyslaw.kitszel@intel.com>
---
v2:
* fix typo (Sashiko)
* abort iavf_configure_queues() at error (Sashiko)
* report send and poll failures with dev_err(), they are not
recoverable here (me)
* cast via uintptr_t in iavf_match_vc_op_cb(), for clang W=1
-Wpointer-to-enum-cast
* commit message: the limit is the number of queue pairs that fit
into a single virtchnl message (62), not 31
* commit message: describe the switch to a synchronous op (Sashiko)
* stop at PF NACK (Sashiko)
* on timeout keep current_op and IAVF_FLAG_AQ_CONFIGURE_QUEUES set
until the late reply arrives, instead of resending right away;
clear both on other errors (no retry) (Sashiko)
* fix dev_err() continuation alignment (Sashiko)
---
.../net/ethernet/intel/iavf/iavf_virtchnl.c | 86 ++++++++++++++++---
1 file changed, 75 insertions(+), 11 deletions(-)
diff --git a/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c b/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c
index e6b7e8f82c7c..67d797a1003f 100644
--- a/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c
+++ b/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c
@@ -8,6 +8,19 @@
#include "iavf_ptp.h"
#include "iavf_prototype.h"
+/* how many Flex Array Member entries do fit into VC message of type *ptr */
+#define iavf_max_vc_entries(ptr, flex_member) \
+ ((IAVF_MAX_AQ_BUF_SIZE - virtchnl_struct_size(ptr, flex_member, 1)) / \
+ sizeof(ptr->flex_member[0]) + 1)
+
+static bool iavf_match_vc_op_cb(struct iavf_adapter *adapter, const void *data,
+ enum virtchnl_ops recv_op)
+{
+ enum virtchnl_ops wanted_op = (enum virtchnl_ops)(uintptr_t)data;
+
+ return recv_op == wanted_op;
+}
+
/**
* iavf_send_pf_msg
* @adapter: adapter structure
@@ -377,8 +390,10 @@ void iavf_configure_queues(struct iavf_adapter *adapter)
struct virtchnl_vsi_queue_config_info *vqci;
int pairs = adapter->num_active_queues;
struct virtchnl_queue_pair_info *vqpi;
- u32 i, max_frame;
+ struct iavf_arq_event_info event;
+ int max_pairs, err = 0;
u8 rx_flags = 0;
+ u32 max_frame;
size_t len;
max_frame = LIBIE_MAX_RX_FRM_LEN(adapter->rx_rings->pp->p.offset);
@@ -390,22 +405,30 @@ void iavf_configure_queues(struct iavf_adapter *adapter)
adapter->current_op);
return;
}
- adapter->current_op = VIRTCHNL_OP_CONFIG_VSI_QUEUES;
- len = virtchnl_struct_size(vqci, qpair, pairs);
+
+ max_pairs = iavf_max_vc_entries(vqci, qpair);
+ len = virtchnl_struct_size(vqci, qpair, min(pairs, max_pairs));
vqci = kzalloc(len, GFP_KERNEL);
if (!vqci)
return;
+ /* any message may arrive while polling, not only the awaited one */
+ event.buf_len = IAVF_MAX_AQ_BUF_SIZE;
+ event.msg_buf = kzalloc(IAVF_MAX_AQ_BUF_SIZE, GFP_KERNEL);
+ if (!event.msg_buf) {
+ kfree(vqci);
+ return;
+ }
+
if (iavf_ptp_cap_supported(adapter, VIRTCHNL_1588_PTP_CAP_RX_TSTAMP))
rx_flags |= VIRTCHNL_PTP_RX_TSTAMP;
vqci->vsi_id = adapter->vsi_res->vsi_id;
- vqci->num_queue_pairs = pairs;
vqpi = vqci->qpair;
- /* Size check is not needed here - HW max is 16 queue pairs, and we
- * can fit info for 31 of them into the AQ buffer before it overflows.
- */
- for (i = 0; i < pairs; i++) {
+
+ for (int i = 0, in_msg = 0; i < pairs; i++) {
+ const bool last = i + 1 == pairs;
+
vqpi->txq.vsi_id = vqci->vsi_id;
vqpi->txq.queue_id = i;
vqpi->txq.ring_len = adapter->tx_rings[i].count;
@@ -423,11 +446,52 @@ void iavf_configure_queues(struct iavf_adapter *adapter)
NETIF_F_RXFCS);
vqpi->rxq.flags = rx_flags;
vqpi++;
+ in_msg++;
+ if (last || in_msg == max_pairs) {
+ adapter->current_op = VIRTCHNL_OP_CONFIG_VSI_QUEUES;
+ vqci->num_queue_pairs = in_msg;
+
+ err = iavf_send_pf_msg(adapter,
+ VIRTCHNL_OP_CONFIG_VSI_QUEUES,
+ (u8 *)vqci,
+ virtchnl_struct_size(vqci, qpair, in_msg));
+ if (err) {
+ dev_err(&adapter->pdev->dev,
+ "sending VC msg to PF failed, err: %d\n",
+ err);
+ break;
+ }
+
+ err = iavf_poll_virtchnl_response(adapter, &event,
+ iavf_match_vc_op_cb,
+ (void *)VIRTCHNL_OP_CONFIG_VSI_QUEUES,
+ 1000);
+ if (err) {
+ dev_err(&adapter->pdev->dev,
+ "config queues poll failed, err: %d\n",
+ err);
+ break;
+ }
+
+ /* PF NACK, iavf_virtchnl_completion() has logged it */
+ if (event.desc.cookie_low) {
+ err = -EIO;
+ break;
+ }
+
+ vqpi = vqci->qpair;
+ in_msg = 0;
+ }
}
- adapter->aq_required &= ~IAVF_FLAG_AQ_CONFIGURE_QUEUES;
- iavf_send_pf_msg(adapter, VIRTCHNL_OP_CONFIG_VSI_QUEUES,
- (u8 *)vqci, len);
+ /* keep the op pending on timeout, as PF may still reply, the whole
+ * config is resent once it does
+ */
+ if (err != -EAGAIN) {
+ adapter->aq_required &= ~IAVF_FLAG_AQ_CONFIGURE_QUEUES;
+ adapter->current_op = VIRTCHNL_OP_UNKNOWN;
+ }
+ kfree(event.msg_buf);
kfree(vqci);
}
--
2.51.1
^ permalink raw reply related [flat|nested] 15+ messages in thread
* [PATCH net-next v2 08/14] iavf: temporary rename of IAVF_MAX_REQ_QUEUES to IAVF_MAX_REQ_QUEUES_VCV1
2026-10-09 12:04 [PATCH net-next v2 00/14] devlink, mlx5, iavf, ice: XLVF for iavf Przemek Kitszel
` (6 preceding siblings ...)
2026-10-09 12:04 ` [PATCH net-next v2 07/14] iavf: extend iavf_configure_queues() to support more queues Przemek Kitszel
@ 2026-10-09 12:04 ` Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 09/14] iavf: increase max number of queues to 256 Przemek Kitszel
` (5 subsequent siblings)
13 siblings, 0 replies; 15+ messages in thread
From: Przemek Kitszel @ 2026-10-09 12:04 UTC (permalink / raw)
To: netdev, Jakub Kicinski, Jiri Pirko
Cc: Tony Nguyen, Aleksandr Loktionov, Michal Schmidt, intel-wired-lan,
edumazet, horms, pabeni, davem, Jonathan Corbet, skhan, rdunlap,
andrew+netdev, saeedm, tariqt, leon, mbloch, jacob.e.keller,
jedrzej.jagielski, anzaki, brett.creeley, jtornosm, ohartoov,
Przemek Kitszel
Rename IAVF_MAX_REQ_QUEUES to IAVF_MAX_REQ_QUEUES_VCV1, in preparation for
the next patch that will extend the max to 256, using old value of 16 for
the "v1" variant of virtchnl opcodes.
Suggested-by: Jedrzej Jagielski <jedrzej.jagielski@intel.com>
Suggested-by: Jacob Keller <jacob.e.keller@intel.com>
Reviewed-by: Aleksandr Loktionov <aleksandr.loktionov@intel.com>
Signed-off-by: Przemek Kitszel <przemyslaw.kitszel@intel.com>
---
drivers/net/ethernet/intel/iavf/iavf.h | 3 ++-
drivers/net/ethernet/intel/iavf/iavf_main.c | 2 +-
drivers/net/ethernet/intel/iavf/iavf_virtchnl.c | 10 +++++-----
3 files changed, 8 insertions(+), 7 deletions(-)
diff --git a/drivers/net/ethernet/intel/iavf/iavf.h b/drivers/net/ethernet/intel/iavf/iavf.h
index 8c45536fd502..94ee18058115 100644
--- a/drivers/net/ethernet/intel/iavf/iavf.h
+++ b/drivers/net/ethernet/intel/iavf/iavf.h
@@ -87,7 +87,8 @@ struct iavf_vsi {
#define IAVF_TX_DESC(R, i) (&(((struct iavf_tx_desc *)((R)->desc))[i]))
#define IAVF_TX_CTXTDESC(R, i) \
(&(((struct iavf_tx_context_desc *)((R)->desc))[i]))
-#define IAVF_MAX_REQ_QUEUES 16
+/* for "old" virtchnl opcodes that accept up to 16 queues */
+#define IAVF_MAX_REQ_QUEUES_VCV1 16
#define IAVF_HKEY_ARRAY_SIZE ((IAVF_VFQF_HKEY_MAX_INDEX + 1) * 4)
#define IAVF_HLUT_ARRAY_SIZE ((IAVF_VFQF_HLUT_MAX_INDEX + 1) * 4)
diff --git a/drivers/net/ethernet/intel/iavf/iavf_main.c b/drivers/net/ethernet/intel/iavf/iavf_main.c
index 2f8a80a11336..bb239e99e450 100644
--- a/drivers/net/ethernet/intel/iavf/iavf_main.c
+++ b/drivers/net/ethernet/intel/iavf/iavf_main.c
@@ -5370,7 +5370,7 @@ static int iavf_probe(struct pci_dev *pdev, const struct pci_device_id *ent)
pci_set_master(pdev);
netdev = alloc_etherdev_mq(sizeof(struct iavf_adapter),
- IAVF_MAX_REQ_QUEUES);
+ IAVF_MAX_REQ_QUEUES_VCV1);
if (!netdev) {
err = -ENOMEM;
goto err_alloc_etherdev;
diff --git a/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c b/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c
index 67d797a1003f..630a551c7566 100644
--- a/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c
+++ b/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c
@@ -268,19 +268,19 @@ int iavf_send_vf_ptp_caps_msg(struct iavf_adapter *adapter)
**/
static void iavf_validate_num_queues(struct iavf_adapter *adapter)
{
- if (adapter->vf_res->num_queue_pairs > IAVF_MAX_REQ_QUEUES) {
+ if (adapter->vf_res->num_queue_pairs > IAVF_MAX_REQ_QUEUES_VCV1) {
struct virtchnl_vsi_resource *vsi_res;
int i;
dev_info(&adapter->pdev->dev, "Received %d queues, but can only have a max of %d\n",
adapter->vf_res->num_queue_pairs,
- IAVF_MAX_REQ_QUEUES);
+ IAVF_MAX_REQ_QUEUES_VCV1);
dev_info(&adapter->pdev->dev, "Fixing by reducing queues to %d\n",
- IAVF_MAX_REQ_QUEUES);
- adapter->vf_res->num_queue_pairs = IAVF_MAX_REQ_QUEUES;
+ IAVF_MAX_REQ_QUEUES_VCV1);
+ adapter->vf_res->num_queue_pairs = IAVF_MAX_REQ_QUEUES_VCV1;
for (i = 0; i < adapter->vf_res->num_vsis; i++) {
vsi_res = &adapter->vf_res->vsi_res[i];
- vsi_res->num_queue_pairs = IAVF_MAX_REQ_QUEUES;
+ vsi_res->num_queue_pairs = IAVF_MAX_REQ_QUEUES_VCV1;
}
}
}
--
2.51.1
^ permalink raw reply related [flat|nested] 15+ messages in thread
* [PATCH net-next v2 09/14] iavf: increase max number of queues to 256
2026-10-09 12:04 [PATCH net-next v2 00/14] devlink, mlx5, iavf, ice: XLVF for iavf Przemek Kitszel
` (7 preceding siblings ...)
2026-10-09 12:04 ` [PATCH net-next v2 08/14] iavf: temporary rename of IAVF_MAX_REQ_QUEUES to IAVF_MAX_REQ_QUEUES_VCV1 Przemek Kitszel
@ 2026-10-09 12:04 ` Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 10/14] iavf: use new opcodes to request more than 16 queues Przemek Kitszel
` (4 subsequent siblings)
13 siblings, 0 replies; 15+ messages in thread
From: Przemek Kitszel @ 2026-10-09 12:04 UTC (permalink / raw)
To: netdev, Jakub Kicinski, Jiri Pirko
Cc: Tony Nguyen, Aleksandr Loktionov, Michal Schmidt, intel-wired-lan,
edumazet, horms, pabeni, davem, Jonathan Corbet, skhan, rdunlap,
andrew+netdev, saeedm, tariqt, leon, mbloch, jacob.e.keller,
jedrzej.jagielski, anzaki, brett.creeley, jtornosm, ohartoov,
Przemek Kitszel
Increase the max number of queues that driver will handle to 256.
Use old legacy limit in the virtchnl handling of iavf_map_queues().
Reviewed-by: Jedrzej Jagielski <jedrzej.jagielski@intel.com>
Signed-off-by: Przemek Kitszel <przemyslaw.kitszel@intel.com>
---
v2:
* change iavf_q_vector::num_ringpairs to u16 (Sashiko)
* drop the stale "_i=0...15" range from IAVF_Q{R,T}X_TAIL1()
---
drivers/net/ethernet/intel/iavf/iavf.h | 6 ++++--
drivers/net/ethernet/intel/iavf/iavf_register.h | 4 ++--
drivers/net/ethernet/intel/iavf/iavf_main.c | 4 ++--
drivers/net/ethernet/intel/iavf/iavf_virtchnl.c | 15 ++++++++-------
4 files changed, 16 insertions(+), 13 deletions(-)
diff --git a/drivers/net/ethernet/intel/iavf/iavf.h b/drivers/net/ethernet/intel/iavf/iavf.h
index 94ee18058115..47236d5f3961 100644
--- a/drivers/net/ethernet/intel/iavf/iavf.h
+++ b/drivers/net/ethernet/intel/iavf/iavf.h
@@ -87,8 +87,10 @@ struct iavf_vsi {
#define IAVF_TX_DESC(R, i) (&(((struct iavf_tx_desc *)((R)->desc))[i]))
#define IAVF_TX_CTXTDESC(R, i) \
(&(((struct iavf_tx_context_desc *)((R)->desc))[i]))
+
/* for "old" virtchnl opcodes that accept up to 16 queues */
#define IAVF_MAX_REQ_QUEUES_VCV1 16
+#define IAVF_MAX_REQ_QUEUES 256
#define IAVF_HKEY_ARRAY_SIZE ((IAVF_VFQF_HKEY_MAX_INDEX + 1) * 4)
#define IAVF_HLUT_ARRAY_SIZE ((IAVF_VFQF_HLUT_MAX_INDEX + 1) * 4)
@@ -108,9 +110,9 @@ struct iavf_q_vector {
struct napi_struct napi;
struct iavf_ring_container rx;
struct iavf_ring_container tx;
- u32 ring_mask;
+ DECLARE_BITMAP(ring_mask, IAVF_MAX_REQ_QUEUES);
u8 itr_countdown; /* when 0 should adjust adaptive ITR */
- u8 num_ringpairs; /* total number of ring pairs in vector */
+ u16 num_ringpairs; /* total number of ring pairs in vector */
u16 v_idx; /* index in the vsi->q_vector array. */
u16 reg_idx; /* register index of the interrupt */
char name[IFNAMSIZ + 15];
diff --git a/drivers/net/ethernet/intel/iavf/iavf_register.h b/drivers/net/ethernet/intel/iavf/iavf_register.h
index a19e88898a0b..11043e330de8 100644
--- a/drivers/net/ethernet/intel/iavf/iavf_register.h
+++ b/drivers/net/ethernet/intel/iavf/iavf_register.h
@@ -56,8 +56,8 @@
#define IAVF_VFINT_ICR0_ENA1_RSVD_SHIFT 31
#define IAVF_VFINT_ICR01 0x00004800 /* Reset: CORER */
#define IAVF_VFINT_ITRN1(_i, _INTVF) (0x00002800 + ((_i) * 64 + (_INTVF) * 4)) /* _i=0...2, _INTVF=0...15 */ /* Reset: VFR */
-#define IAVF_QRX_TAIL1(_Q) (0x00002000 + ((_Q) * 4)) /* _i=0...15 */ /* Reset: CORER */
-#define IAVF_QTX_TAIL1(_Q) (0x00000000 + ((_Q) * 4)) /* _i=0...15 */ /* Reset: PFR */
+#define IAVF_QRX_TAIL1(_Q) (0x00002000 + ((_Q) * 4)) /* Reset: CORER */
+#define IAVF_QTX_TAIL1(_Q) (0x00000000 + ((_Q) * 4)) /* Reset: PFR */
#define IAVF_VFQF_HENA(_i) (0x0000C400 + ((_i) * 4)) /* _i=0...1 */ /* Reset: CORER */
#define IAVF_VFQF_HKEY(_i) (0x0000CC00 + ((_i) * 4)) /* _i=0...12 */ /* Reset: CORER */
#define IAVF_VFQF_HKEY_MAX_INDEX 12
diff --git a/drivers/net/ethernet/intel/iavf/iavf_main.c b/drivers/net/ethernet/intel/iavf/iavf_main.c
index bb239e99e450..8059e855dc09 100644
--- a/drivers/net/ethernet/intel/iavf/iavf_main.c
+++ b/drivers/net/ethernet/intel/iavf/iavf_main.c
@@ -439,7 +439,7 @@ iavf_map_vector_to_rxq(struct iavf_adapter *adapter, int v_idx, int r_idx)
q_vector->rx.count++;
q_vector->rx.next_update = jiffies + 1;
q_vector->rx.target_itr = ITR_TO_REG(rx_ring->itr_setting);
- q_vector->ring_mask |= BIT(r_idx);
+ set_bit(r_idx, q_vector->ring_mask);
wr32(hw, IAVF_VFINT_ITRN1(IAVF_RX_ITR, q_vector->reg_idx),
q_vector->rx.current_itr >> 1);
q_vector->rx.current_itr = q_vector->rx.target_itr;
@@ -5370,7 +5370,7 @@ static int iavf_probe(struct pci_dev *pdev, const struct pci_device_id *ent)
pci_set_master(pdev);
netdev = alloc_etherdev_mq(sizeof(struct iavf_adapter),
- IAVF_MAX_REQ_QUEUES_VCV1);
+ IAVF_MAX_REQ_QUEUES);
if (!netdev) {
err = -ENOMEM;
goto err_alloc_etherdev;
diff --git a/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c b/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c
index 630a551c7566..3b9545d2c021 100644
--- a/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c
+++ b/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c
@@ -268,19 +268,19 @@ int iavf_send_vf_ptp_caps_msg(struct iavf_adapter *adapter)
**/
static void iavf_validate_num_queues(struct iavf_adapter *adapter)
{
- if (adapter->vf_res->num_queue_pairs > IAVF_MAX_REQ_QUEUES_VCV1) {
+ if (adapter->vf_res->num_queue_pairs > IAVF_MAX_REQ_QUEUES) {
struct virtchnl_vsi_resource *vsi_res;
int i;
dev_info(&adapter->pdev->dev, "Received %d queues, but can only have a max of %d\n",
adapter->vf_res->num_queue_pairs,
- IAVF_MAX_REQ_QUEUES_VCV1);
+ IAVF_MAX_REQ_QUEUES);
dev_info(&adapter->pdev->dev, "Fixing by reducing queues to %d\n",
- IAVF_MAX_REQ_QUEUES_VCV1);
- adapter->vf_res->num_queue_pairs = IAVF_MAX_REQ_QUEUES_VCV1;
+ IAVF_MAX_REQ_QUEUES);
+ adapter->vf_res->num_queue_pairs = IAVF_MAX_REQ_QUEUES;
for (i = 0; i < adapter->vf_res->num_vsis; i++) {
vsi_res = &adapter->vf_res->vsi_res[i];
- vsi_res->num_queue_pairs = IAVF_MAX_REQ_QUEUES_VCV1;
+ vsi_res->num_queue_pairs = IAVF_MAX_REQ_QUEUES;
}
}
}
@@ -583,8 +583,9 @@ void iavf_map_queues(struct iavf_adapter *adapter)
vecmap->vsi_id = adapter->vsi_res->vsi_id;
vecmap->vector_id = v_idx + NONQ_VECS;
- vecmap->txq_map = q_vector->ring_mask;
- vecmap->rxq_map = q_vector->ring_mask;
+ vecmap->txq_map = bitmap_read(q_vector->ring_mask, 0,
+ IAVF_MAX_REQ_QUEUES_VCV1);
+ vecmap->rxq_map = vecmap->txq_map;
vecmap->rxitr_idx = IAVF_RX_ITR;
vecmap->txitr_idx = IAVF_TX_ITR;
}
--
2.51.1
^ permalink raw reply related [flat|nested] 15+ messages in thread
* [PATCH net-next v2 10/14] iavf: use new opcodes to request more than 16 queues
2026-10-09 12:04 [PATCH net-next v2 00/14] devlink, mlx5, iavf, ice: XLVF for iavf Przemek Kitszel
` (8 preceding siblings ...)
2026-10-09 12:04 ` [PATCH net-next v2 09/14] iavf: increase max number of queues to 256 Przemek Kitszel
@ 2026-10-09 12:04 ` Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 11/14] ice: introduce handling of virtchnl LARGE VF opcodes Przemek Kitszel
` (3 subsequent siblings)
13 siblings, 0 replies; 15+ messages in thread
From: Przemek Kitszel @ 2026-10-09 12:04 UTC (permalink / raw)
To: netdev, Jakub Kicinski, Jiri Pirko
Cc: Tony Nguyen, Aleksandr Loktionov, Michal Schmidt, intel-wired-lan,
edumazet, horms, pabeni, davem, Jonathan Corbet, skhan, rdunlap,
andrew+netdev, saeedm, tariqt, leon, mbloch, jacob.e.keller,
jedrzej.jagielski, anzaki, brett.creeley, jtornosm, ohartoov,
Przemek Kitszel, Ahmed Zaki
From: Ahmed Zaki <ahmed.zaki@intel.com>
Legacy virtchnl opcodes select queues by fixed width bitmaps: u32 for
VIRTCHNL_OP_ENABLE_QUEUES and VIRTCHNL_OP_DISABLE_QUEUES, u16 for
VIRTCHNL_OP_CONFIG_IRQ_MAP. Define new OP codes in order to enable
the VF to use a variable number of queues.
First the new capability VIRTCHNL_VF_LARGE_NUM_QPAIRS is defined. Only
if this cap is negotiated, the VF is allowed to use the new OP codes.
The VIRTCHNL_OP_GET_MAX_RSS_QREGION is exchanged right after the VF
resource message is received. This is done similar to the PTP and other
extended cap exchanges.
Enabling and disabling queues, and mapping queues to vectors is done
through the new VIRTCHNL_OP_ENABLE_QUEUES_V2,
VIRTCHNL_OP_DISABLE_QUEUES_V2, and VIRTCHNL_OP_MAP_QUEUE_VECTOR messages.
Use the new iavf_request_queues() in the iavf_set_channels().
It's not only for elegance, but it enables our brittle handcoded state
machine to go through requesting the queues before PF driver reallocates
queue-related arrays.
Signed-off-by: Ahmed Zaki <ahmed.zaki@intel.com>
Co-developed-by: Przemek Kitszel <przemyslaw.kitszel@intel.com>
Signed-off-by: Przemek Kitszel <przemyslaw.kitszel@intel.com>
---
v2:
* use proper IAVF_EXTENDED_CAP_SEND_RSS_QREGION flag when handling
iavf_init_recv_max_rss_qregion() error (Sashiko)
* update adapter->rss_{key,lut}_size and adapter->current_op only after
successful allocation (Sashiko)
* use the v2 opcodes to enable queues whenever
VIRTCHNL_VF_LARGE_NUM_QPAIRS is negotiated, as for disabling (it was
based on the queue count) (Sashiko)
* set adapter->current_op in iavf_send_max_rss_qregion() (Sashiko)
* iavf_set_channels(): stop clearing adapter->current_op before
iavf_request_queues(), it was a debugging leftover (Sashiko)
* iavf_request_queues(): undo current_op, num_req_queues and
IAVF_FLAG_REINIT_ITR_NEEDED when sending the request fails (Sashiko)
* iavf_map_queue_vector(): stop at the first send or poll error, and
report it with dev_err()
* replace rss key/lut krealloc() with simple kmalloc()s, to have an
easier job on the error path
* request RSS qregion again after reset, together with other extended
caps: send it from iavf_process_aq_command() and handle the reply in
iavf_virtchnl_completion()
* auto-request queues at init only once, and only with
VIRTCHNL_VF_LARGE_NUM_QPAIRS (and VIRTCHNL_VF_OFFLOAD_REQ_QUEUES),
where num_queue_pairs is the PF limit rather than the allocation;
without it keep the previous behavior, with no extra VF reset
* set channels locally (as before this series) unless
VIRTCHNL_VF_LARGE_NUM_QPAIRS, so ethtool -L down no longer lowers
max_combined for good
* reinit queues when the PF lowers num_queue_pairs below the number of
active queues on reset (e.g. VF RSS LUT switched from GLOBAL back to
VSI, so LARGE_NUM_QPAIRS is gone), instead of using legacy opcodes
with too many queues
* rebase over iavf_poll_virtchnl_response() changes
* commit message: explain the limits of the legacy opcodes
* iavf_map_queue_vector(): stop at PF NACK; on timeout keep current_op
and IAVF_FLAG_AQ_MAP_VECTORS set until the late reply arrives, clear
both on other errors (no retry), as in iavf_configure_queues()
(Sashiko)
Sashiko suggestions I don't want to apply:
S: iavf_fill_rss_lut() will crash if qregion->qregion_width >= 32
me: value is coming from PF, which must be trusted
---
drivers/net/ethernet/intel/iavf/iavf.h | 12 +
include/linux/net/intel/virtchnl.h | 136 +++++++++-
.../net/ethernet/intel/iavf/iavf_ethtool.c | 3 +
drivers/net/ethernet/intel/iavf/iavf_main.c | 132 +++++++++-
.../net/ethernet/intel/iavf/iavf_virtchnl.c | 243 +++++++++++++++++-
5 files changed, 510 insertions(+), 16 deletions(-)
diff --git a/drivers/net/ethernet/intel/iavf/iavf.h b/drivers/net/ethernet/intel/iavf/iavf.h
index 47236d5f3961..5aaa337d5ffa 100644
--- a/drivers/net/ethernet/intel/iavf/iavf.h
+++ b/drivers/net/ethernet/intel/iavf/iavf.h
@@ -345,13 +345,15 @@ struct iavf_adapter {
#define IAVF_FLAG_AQ_GET_SUPPORTED_RXDIDS BIT_ULL(42)
#define IAVF_FLAG_AQ_GET_PTP_CAPS BIT_ULL(43)
#define IAVF_FLAG_AQ_SEND_PTP_CMD BIT_ULL(44)
+#define IAVF_FLAG_AQ_GET_MAX_RSS_QREGION BIT_ULL(45)
/* AQ messages that must be sent after IAVF_FLAG_AQ_GET_CONFIG, in
* order to negotiated extended capabilities.
*/
#define IAVF_FLAG_AQ_EXTENDED_CAPS \
(IAVF_FLAG_AQ_GET_OFFLOAD_VLAN_V2_CAPS | \
IAVF_FLAG_AQ_GET_SUPPORTED_RXDIDS | \
+ IAVF_FLAG_AQ_GET_MAX_RSS_QREGION | \
IAVF_FLAG_AQ_GET_PTP_CAPS)
/* flags for processing extended capability messages during
@@ -368,12 +370,16 @@ struct iavf_adapter {
#define IAVF_EXTENDED_CAP_RECV_RXDID BIT_ULL(3)
#define IAVF_EXTENDED_CAP_SEND_PTP BIT_ULL(4)
#define IAVF_EXTENDED_CAP_RECV_PTP BIT_ULL(5)
+#define IAVF_EXTENDED_CAP_SEND_RSS_QREGION BIT_ULL(6)
+#define IAVF_EXTENDED_CAP_RECV_RSS_QREGION BIT_ULL(7)
#define IAVF_EXTENDED_CAPS \
(IAVF_EXTENDED_CAP_SEND_VLAN_V2 | \
IAVF_EXTENDED_CAP_RECV_VLAN_V2 | \
IAVF_EXTENDED_CAP_SEND_RXDID | \
IAVF_EXTENDED_CAP_RECV_RXDID | \
+ IAVF_EXTENDED_CAP_SEND_RSS_QREGION | \
+ IAVF_EXTENDED_CAP_RECV_RSS_QREGION | \
IAVF_EXTENDED_CAP_SEND_PTP | \
IAVF_EXTENDED_CAP_RECV_PTP)
@@ -413,6 +419,8 @@ struct iavf_adapter {
#define RSS_REG(_a) (!((_a)->vf_res->vf_cap_flags & \
(VIRTCHNL_VF_OFFLOAD_RSS_AQ | \
VIRTCHNL_VF_OFFLOAD_RSS_PF)))
+#define LARGE_NUM_QPAIRS_SUPPORT(_a) \
+ ((_a)->vf_res->vf_cap_flags & VIRTCHNL_VF_LARGE_NUM_QPAIRS)
#define VLAN_ALLOWED(_a) ((_a)->vf_res->vf_cap_flags & \
VIRTCHNL_VF_OFFLOAD_VLAN)
#define VLAN_V2_ALLOWED(_a) ((_a)->vf_res->vf_cap_flags & \
@@ -447,6 +455,7 @@ struct iavf_adapter {
struct virtchnl_vlan_caps vlan_v2_caps;
u64 supp_rxdids;
struct iavf_ptp ptp;
+ struct virtchnl_max_rss_qregion max_rss_qregion;
u16 msg_enable;
struct iavf_eth_stats current_stats;
struct virtchnl_qos_cap_list *qos_caps;
@@ -583,6 +592,8 @@ int iavf_send_vf_supported_rxdids_msg(struct iavf_adapter *adapter);
int iavf_get_vf_supported_rxdids(struct iavf_adapter *adapter);
int iavf_send_vf_ptp_caps_msg(struct iavf_adapter *adapter);
int iavf_get_vf_ptp_caps(struct iavf_adapter *adapter);
+int iavf_send_max_rss_qregion(struct iavf_adapter *adapter);
+int iavf_get_max_rss_qregion(struct iavf_adapter *adapter);
void iavf_set_queue_vlan_tag_loc(struct iavf_adapter *adapter);
u16 iavf_get_num_vlans_added(struct iavf_adapter *adapter);
void iavf_irq_enable(struct iavf_adapter *adapter, bool flush);
@@ -597,6 +608,7 @@ void iavf_add_vlans(struct iavf_adapter *adapter);
void iavf_del_vlans(struct iavf_adapter *adapter);
void iavf_set_promiscuous(struct iavf_adapter *adapter);
bool iavf_promiscuous_mode_changed(struct iavf_adapter *adapter);
+int iavf_request_queues(struct iavf_adapter *adapter, int num);
void iavf_request_stats(struct iavf_adapter *adapter);
int iavf_request_reset(struct iavf_adapter *adapter);
void iavf_get_rss_hashcfg(struct iavf_adapter *adapter);
diff --git a/include/linux/net/intel/virtchnl.h b/include/linux/net/intel/virtchnl.h
index 11bdab5522fd..ffb7d3c21334 100644
--- a/include/linux/net/intel/virtchnl.h
+++ b/include/linux/net/intel/virtchnl.h
@@ -1,5 +1,5 @@
/* SPDX-License-Identifier: GPL-2.0-only */
-/* Copyright (c) 2013-2022, Intel Corporation. */
+/* Copyright (c) 2013-2026, Intel Corporation. */
#ifndef _VIRTCHNL_H_
#define _VIRTCHNL_H_
@@ -22,8 +22,7 @@
* we must send all messages as "indirect", i.e. using an external buffer.
*
* All the VSI indexes are relative to the VF. Each VF can have maximum of
- * three VSIs. All the queue indexes are relative to the VSI. Each VF can
- * have a maximum of sixteen queues for all of its VSIs.
+ * three VSIs. All the queue indexes are relative to the VSI.
*
* The PF is required to return a status code in v_retval for all messages
* except RESET_VF, which does not require any response. The returned value
@@ -147,6 +146,7 @@ enum virtchnl_ops {
VIRTCHNL_OP_DEL_RSS_CFG = 46,
VIRTCHNL_OP_ADD_FDIR_FILTER = 47,
VIRTCHNL_OP_DEL_FDIR_FILTER = 48,
+ VIRTCHNL_OP_GET_MAX_RSS_QREGION = 50,
VIRTCHNL_OP_GET_OFFLOAD_VLAN_V2_CAPS = 51,
VIRTCHNL_OP_ADD_VLAN_V2 = 52,
VIRTCHNL_OP_DEL_VLAN_V2 = 53,
@@ -159,7 +159,11 @@ enum virtchnl_ops {
VIRTCHNL_OP_1588_PTP_GET_TIME = 61,
/* opcode 62 - 65 are reserved */
VIRTCHNL_OP_GET_QOS_CAPS = 66,
- /* opcode 68 through 111 are reserved */
+ /* opcodes 67 through 106 are reserved */
+ VIRTCHNL_OP_ENABLE_QUEUES_V2 = 107,
+ VIRTCHNL_OP_DISABLE_QUEUES_V2 = 108,
+ /* opcodes 109 and 110 are reserved */
+ VIRTCHNL_OP_MAP_QUEUE_VECTOR = 111,
VIRTCHNL_OP_CONFIG_QUEUE_BW = 112,
VIRTCHNL_OP_CONFIG_QUANTA = 113,
VIRTCHNL_OP_MAX,
@@ -257,6 +261,7 @@ VIRTCHNL_CHECK_STRUCT_LEN(16, virtchnl_vsi_resource);
#define VIRTCHNL_VF_OFFLOAD_REQ_QUEUES BIT(6)
/* used to negotiate communicating link speeds in Mbps */
#define VIRTCHNL_VF_CAP_ADV_LINK_SPEED BIT(7)
+#define VIRTCHNL_VF_LARGE_NUM_QPAIRS BIT(9)
#define VIRTCHNL_VF_OFFLOAD_CRC BIT(10)
#define VIRTCHNL_VF_OFFLOAD_TC_U32 BIT(11)
#define VIRTCHNL_VF_OFFLOAD_VLAN_V2 BIT(15)
@@ -500,6 +505,34 @@ struct virtchnl_queue_select {
VIRTCHNL_CHECK_STRUCT_LEN(12, virtchnl_queue_select);
+/* VIRTCHNL_OP_GET_MAX_RSS_QREGION
+ *
+ * if VIRTCHNL_VF_LARGE_NUM_QPAIRS was negotiated in
+ * VIRTCHNL_OP_GET_VF_RESOURCES then this op must be supported.
+ *
+ * VF sends this message in order to query the max RSS queue region
+ * size supported by PF, when VIRTCHNL_VF_LARGE_NUM_QPAIRS is enabled.
+ * This information should be used when configuring the RSS LUT and/or
+ * configuring queue region based filters.
+ *
+ * The maximum RSS queue region is 2^qregion_width. So, a qregion_width of 6
+ * would inform the VF that the PF supports a maximum RSS queue region of 64.
+ *
+ * A queue region represents a range of queues that can be used to configure
+ * a RSS LUT. For example, if a VF is given 64 queues, but only a max queue
+ * region size of 16 (i.e. 2^qregion_width = 16) then it will only be able
+ * to configure the RSS LUT with queue indices from 0 to 15. However, other
+ * filters can be used to direct packets to queues >15 via specifying a queue
+ * base/offset and queue region width.
+ */
+struct virtchnl_max_rss_qregion {
+ u16 vport_id;
+ u16 qregion_width;
+ u8 pad[4];
+};
+
+VIRTCHNL_CHECK_STRUCT_LEN(8, virtchnl_max_rss_qregion);
+
/* VIRTCHNL_OP_ADD_ETH_ADDR
* VF sends this message in order to add one or more unicast or multicast
* address filters for the specified VSI.
@@ -1663,6 +1696,70 @@ struct virtchnl_queue_chunk {
VIRTCHNL_CHECK_STRUCT_LEN(8, virtchnl_queue_chunk);
+/* VIRTCHNL_OP_ENABLE_QUEUES_V2
+ * VIRTCHNL_OP_DISABLE_QUEUES_V2
+ *
+ * These opcodes can be used if VIRTCHNL_VF_LARGE_NUM_QPAIRS was negotiated in
+ * VIRTCHNL_OP_GET_VF_RESOURCES
+ *
+ * VF sends virtchnl_ena_dis_queues struct to specify the queues to be
+ * enabled/disabled in chunks. Also applicable to single queue RX or
+ * TX. PF performs requested action and returns status.
+ */
+struct virtchnl_del_ena_dis_queues {
+ u16 vport_id;
+ u16 pad;
+ u16 num_chunks;
+ u16 rsvd;
+ struct virtchnl_queue_chunk chunks[];
+};
+
+VIRTCHNL_CHECK_STRUCT_LEN(8, virtchnl_del_ena_dis_queues);
+#define virtchnl_del_ena_dis_queues_LEGACY_SIZEOF 16
+
+/* Virtchannel interrupt throttling rate index */
+enum virtchnl_itr_idx {
+ VIRTCHNL_ITR_IDX_0 = 0,
+ VIRTCHNL_ITR_IDX_1 = 1,
+ VIRTCHNL_ITR_IDX_NO_ITR = 3,
+};
+
+/* Queue to vector mapping */
+struct virtchnl_queue_vector {
+ u16 queue_id;
+ u16 vector_id;
+ u8 pad[4];
+
+ /* see enum virtchnl_itr_idx */
+ s32 itr_idx;
+
+ /* see enum virtchnl_queue_type */
+ s32 queue_type;
+};
+
+VIRTCHNL_CHECK_STRUCT_LEN(16, virtchnl_queue_vector);
+
+/* VIRTCHNL_OP_MAP_QUEUE_VECTOR
+ *
+ * This opcode can be used only if VIRTCHNL_VF_LARGE_NUM_QPAIRS was negotiated
+ * in VIRTCHNL_OP_GET_VF_RESOURCES
+ *
+ * VF sends this message to map queues to vectors and ITR index registers.
+ * External data buffer contains virtchnl_queue_vector_maps structure
+ * that contains num_qv_maps of virtchnl_queue_vector structures.
+ * PF maps the requested queue vector maps after validating the queue and vector
+ * ids and returns a status code.
+ */
+struct virtchnl_queue_vector_maps {
+ u16 vport_id;
+ u16 num_qv_maps;
+ u8 pad[4];
+ struct virtchnl_queue_vector qv_maps[];
+};
+
+VIRTCHNL_CHECK_STRUCT_LEN(8, virtchnl_queue_vector_maps);
+#define virtchnl_queue_vector_maps_LEGACY_SIZEOF 24
+
struct virtchnl_quanta_cfg {
u16 quanta_size;
u16 pad;
@@ -1695,6 +1792,8 @@ VIRTCHNL_CHECK_STRUCT_LEN(12, virtchnl_quanta_cfg);
__vss(virtchnl_rdma_qvlist_info, __vss_byelem, p, m, c), \
__vss(virtchnl_qos_cap_list, __vss_byelem, p, m, c), \
__vss(virtchnl_queues_bw_cfg, __vss_byelem, p, m, c), \
+ __vss(virtchnl_del_ena_dis_queues, __vss_byelem, p, m, c), \
+ __vss(virtchnl_queue_vector_maps, __vss_byelem, p, m, c), \
__vss(virtchnl_rss_key, __vss_byone, p, m, c), \
__vss(virtchnl_rss_lut, __vss_byone, p, m, c))
@@ -1757,6 +1856,8 @@ virtchnl_vc_validate_vf_msg(struct virtchnl_version_info *ver, u32 v_opcode,
case VIRTCHNL_OP_DISABLE_QUEUES:
valid_len = sizeof(struct virtchnl_queue_select);
break;
+ case VIRTCHNL_OP_GET_MAX_RSS_QREGION:
+ break;
case VIRTCHNL_OP_ADD_ETH_ADDR:
case VIRTCHNL_OP_DEL_ETH_ADDR:
valid_len = virtchnl_ether_addr_list_LEGACY_SIZEOF;
@@ -1929,7 +2030,32 @@ virtchnl_vc_validate_vf_msg(struct virtchnl_version_info *ver, u32 v_opcode,
case VIRTCHNL_OP_1588_PTP_GET_TIME:
valid_len = sizeof(struct virtchnl_phc_time);
break;
- /* These are always errors coming from the VF. */
+ case VIRTCHNL_OP_ENABLE_QUEUES_V2:
+ case VIRTCHNL_OP_DISABLE_QUEUES_V2:
+ valid_len = sizeof(struct virtchnl_del_ena_dis_queues);
+ if (msglen >= valid_len) {
+ struct virtchnl_del_ena_dis_queues *qs =
+ (struct virtchnl_del_ena_dis_queues *)msg;
+
+ if (!qs->num_chunks) {
+ err_msg_format = true;
+ break;
+ }
+ valid_len = virtchnl_struct_size(qs, chunks,
+ qs->num_chunks);
+ }
+ break;
+ case VIRTCHNL_OP_MAP_QUEUE_VECTOR:
+ valid_len = virtchnl_queue_vector_maps_LEGACY_SIZEOF;
+ if (msglen >= valid_len) {
+ struct virtchnl_queue_vector_maps *v_qp = (void *)msg;
+
+ err_msg_format = !v_qp->num_qv_maps;
+ valid_len = virtchnl_struct_size(v_qp, qv_maps,
+ v_qp->num_qv_maps);
+ }
+ break;
+ /* These are always errors when coming from the VF. */
case VIRTCHNL_OP_EVENT:
case VIRTCHNL_OP_UNKNOWN:
default:
diff --git a/drivers/net/ethernet/intel/iavf/iavf_ethtool.c b/drivers/net/ethernet/intel/iavf/iavf_ethtool.c
index e7cf12eaa268..83ab3e12b10a 100644
--- a/drivers/net/ethernet/intel/iavf/iavf_ethtool.c
+++ b/drivers/net/ethernet/intel/iavf/iavf_ethtool.c
@@ -1739,6 +1739,9 @@ static int iavf_set_channels(struct net_device *netdev,
if (ch->rx_count || ch->tx_count || ch->other_count != NONQ_VECS)
return -EINVAL;
+ if (LARGE_NUM_QPAIRS_SUPPORT(adapter))
+ return iavf_request_queues(adapter, num_req);
+
adapter->num_req_queues = num_req;
adapter->flags |= IAVF_FLAG_REINIT_ITR_NEEDED;
adapter->flags |= IAVF_FLAG_RESET_NEEDED;
diff --git a/drivers/net/ethernet/intel/iavf/iavf_main.c b/drivers/net/ethernet/intel/iavf/iavf_main.c
index 8059e855dc09..f8193b0cae4a 100644
--- a/drivers/net/ethernet/intel/iavf/iavf_main.c
+++ b/drivers/net/ethernet/intel/iavf/iavf_main.c
@@ -1775,10 +1775,14 @@ int iavf_config_rss(struct iavf_adapter *adapter)
**/
static void iavf_fill_rss_lut(struct iavf_adapter *adapter)
{
- u16 i;
+ struct virtchnl_max_rss_qregion *qregion = &adapter->max_rss_qregion;
+ int max = adapter->num_active_queues;
+
+ if (LARGE_NUM_QPAIRS_SUPPORT(adapter) && qregion->qregion_width)
+ max = min(max, (int)BIT(qregion->qregion_width));
- for (i = 0; i < adapter->rss_lut_size; i++)
- adapter->rss_lut[i] = i % adapter->num_active_queues;
+ for (int i = 0; i < adapter->rss_lut_size; i++)
+ adapter->rss_lut[i] = i % max;
}
/**
@@ -2089,6 +2093,8 @@ static int iavf_process_aq_command(struct iavf_adapter *adapter)
{
if (adapter->aq_required & IAVF_FLAG_AQ_GET_CONFIG)
return iavf_send_vf_config_msg(adapter);
+ if (adapter->aq_required & IAVF_FLAG_AQ_GET_MAX_RSS_QREGION)
+ return iavf_send_max_rss_qregion(adapter);
if (adapter->aq_required & IAVF_FLAG_AQ_GET_OFFLOAD_VLAN_V2_CAPS)
return iavf_send_vf_offload_vlan_v2_msg(adapter);
if (adapter->aq_required & IAVF_FLAG_AQ_GET_SUPPORTED_RXDIDS)
@@ -2458,8 +2464,9 @@ static void iavf_init_version_check(struct iavf_adapter *adapter)
*/
int iavf_parse_vf_resource_msg(struct iavf_adapter *adapter)
{
- int i, num_req_queues = adapter->num_req_queues;
+ int i, qnum, num_req_queues = adapter->num_req_queues;
struct iavf_vsi *vsi = &adapter->vsi;
+ bool reconfig_rss = false;
for (i = 0; i < adapter->vf_res->num_vsis; i++) {
if (adapter->vf_res->vsi_res[i].vsi_type == VIRTCHNL_VSI_SRIOV)
@@ -2486,21 +2493,65 @@ int iavf_parse_vf_resource_msg(struct iavf_adapter *adapter)
return -EAGAIN;
}
+ /* in LARGE mode num_queue_pairs is the PF limit, not our allocation */
+ if (!num_req_queues && LARGE_NUM_QPAIRS_SUPPORT(adapter) &&
+ adapter->vf_res->vf_cap_flags & VIRTCHNL_VF_OFFLOAD_REQ_QUEUES) {
+ adapter->current_op = VIRTCHNL_OP_UNKNOWN;
+ qnum = min(adapter->vsi_res->num_queue_pairs,
+ num_online_cpus());
+
+ return iavf_request_queues(adapter, qnum);
+ }
+ /* PF could lower the limit on reset, e.g. when RSS LUT got smaller */
+ if (adapter->num_active_queues > adapter->vsi_res->num_queue_pairs) {
+ adapter->flags |= IAVF_FLAG_REINIT_MSIX_NEEDED;
+ adapter->num_req_queues = adapter->vsi_res->num_queue_pairs;
+ iavf_schedule_reset(adapter, IAVF_FLAG_RESET_NEEDED);
+
+ return -EAGAIN;
+ }
adapter->num_req_queues = 0;
adapter->vsi.id = adapter->vsi_res->vsi_id;
adapter->vsi.back = adapter;
adapter->vsi.base_vector = 1;
vsi->netdev = adapter->netdev;
vsi->qs_handle = adapter->vsi_res->qset_handle;
if (adapter->vf_res->vf_cap_flags & VIRTCHNL_VF_OFFLOAD_RSS_PF) {
- adapter->rss_key_size = adapter->vf_res->rss_key_size;
- adapter->rss_lut_size = adapter->vf_res->rss_lut_size;
+ if ((adapter->rss_key &&
+ adapter->rss_key_size != adapter->vf_res->rss_key_size) ||
+ (adapter->rss_lut &&
+ adapter->rss_lut_size != adapter->vf_res->rss_lut_size)) {
+ reconfig_rss = true;
+ } else {
+ adapter->rss_key_size = adapter->vf_res->rss_key_size;
+ adapter->rss_lut_size = adapter->vf_res->rss_lut_size;
+ }
} else {
adapter->rss_key_size = IAVF_HKEY_ARRAY_SIZE;
adapter->rss_lut_size = IAVF_HLUT_ARRAY_SIZE;
}
+ if (reconfig_rss) {
+ u8 *rss_key, *rss_lut;
+
+ rss_key = kmalloc(adapter->vf_res->rss_key_size, GFP_KERNEL);
+ rss_lut = kmalloc(adapter->vf_res->rss_lut_size, GFP_KERNEL);
+ if (!rss_lut || !rss_key) {
+ kfree(rss_key);
+ kfree(rss_lut);
+ return -ENOMEM;
+ }
+
+ kfree(adapter->rss_key);
+ kfree(adapter->rss_lut);
+ adapter->rss_key = rss_key;
+ adapter->rss_lut = rss_lut;
+ adapter->rss_key_size = adapter->vf_res->rss_key_size;
+ adapter->rss_lut_size = adapter->vf_res->rss_lut_size;
+ iavf_init_rss(adapter);
+ }
+
return 0;
}
@@ -2626,6 +2677,65 @@ static void iavf_init_recv_offload_vlan_v2_caps(struct iavf_adapter *adapter)
iavf_change_state(adapter, __IAVF_INIT_FAILED);
}
+/*
+ * iavf_init_send_max_rss_qregion - part of querying for RSS max queue region
+ * @adapter: board private structure
+ *
+ * Function processes send of the VIRTCHNL_OP_GET_MAX_RSS_QREGION to the PF.
+ * Must clear IAVF_EXTENDED_CAP_RECV_RSS_QREGION if the message is not sent, e.g.
+ * due to the PF not negotiating VIRTCHNL_VF_LARGE_NUM_QPAIRS.
+ */
+static void iavf_init_send_max_rss_qregion(struct iavf_adapter *adapter)
+{
+ int ret;
+
+ WARN_ON(!(adapter->extended_caps & IAVF_EXTENDED_CAP_SEND_RSS_QREGION));
+
+ ret = iavf_send_max_rss_qregion(adapter);
+ if (ret == -EOPNOTSUPP) {
+ /* PF does not support VIRTCHNL_VF_LARGE_NUM_QPAIRS. In this
+ * case, we did not send the capability exchange message and do
+ * not expect a response.
+ */
+ adapter->extended_caps &= ~IAVF_EXTENDED_CAP_RECV_RSS_QREGION;
+ }
+
+ /* We sent the message, so move on to the next step */
+ adapter->extended_caps &= ~IAVF_EXTENDED_CAP_SEND_RSS_QREGION;
+}
+
+/**
+ * iavf_init_recv_max_rss_qregion - part of querying for RSS max queue region
+ * @adapter: board private structure
+ *
+ * Function processes receipt of the RSS max qregion to be used for the LUT.
+ */
+static void iavf_init_recv_max_rss_qregion(struct iavf_adapter *adapter)
+{
+ int ret;
+
+ WARN_ON(!(adapter->extended_caps & IAVF_EXTENDED_CAP_RECV_RSS_QREGION));
+
+ memset(&adapter->max_rss_qregion, 0, sizeof(adapter->max_rss_qregion));
+
+ ret = iavf_get_max_rss_qregion(adapter);
+ if (ret)
+ goto err;
+
+ /* We've processed the PF response to the VIRTCHNL_OP_GET_MAX_RSS_QREGION
+ * message we sent previously.
+ */
+ adapter->extended_caps &= ~IAVF_EXTENDED_CAP_RECV_RSS_QREGION;
+ return;
+
+err:
+ /* We didn't receive a reply. Make sure we try sending again when
+ * __IAVF_INIT_FAILED attempts to recover.
+ */
+ adapter->extended_caps |= IAVF_EXTENDED_CAP_SEND_RSS_QREGION;
+ iavf_change_state(adapter, __IAVF_INIT_FAILED);
+}
+
/**
* iavf_init_send_supported_rxdids - part of querying for supported RXDID
* formats
@@ -2774,6 +2884,16 @@ static void iavf_init_process_extended_caps(struct iavf_adapter *adapter)
return;
}
+ /* Process capability exchange for RSS max qregion */
+ if (adapter->extended_caps & IAVF_EXTENDED_CAP_SEND_RSS_QREGION) {
+ iavf_init_send_max_rss_qregion(adapter);
+ return;
+ }
+ if (adapter->extended_caps & IAVF_EXTENDED_CAP_RECV_RSS_QREGION) {
+ iavf_init_recv_max_rss_qregion(adapter);
+ return;
+ }
+
/* When we reach here, no further extended capabilities exchanges are
* necessary, so we finally transition into __IAVF_INIT_CONFIG_ADAPTER
*/
diff --git a/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c b/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c
index 3b9545d2c021..89d95d3b564c 100644
--- a/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c
+++ b/drivers/net/ethernet/intel/iavf/iavf_virtchnl.c
@@ -178,6 +178,7 @@ int iavf_send_vf_config_msg(struct iavf_adapter *adapter)
VIRTCHNL_VF_OFFLOAD_VLAN_V2 |
VIRTCHNL_VF_OFFLOAD_RX_FLEX_DESC |
VIRTCHNL_VF_OFFLOAD_CRC |
+ VIRTCHNL_VF_LARGE_NUM_QPAIRS |
VIRTCHNL_VF_OFFLOAD_ENCAP_CSUM |
VIRTCHNL_VF_OFFLOAD_REQ_QUEUES |
VIRTCHNL_VF_CAP_PTP |
@@ -259,28 +260,45 @@ int iavf_send_vf_ptp_caps_msg(struct iavf_adapter *adapter)
(u8 *)&hw_caps, sizeof(hw_caps));
}
+int iavf_send_max_rss_qregion(struct iavf_adapter *adapter)
+{
+ adapter->aq_required &= ~IAVF_FLAG_AQ_GET_MAX_RSS_QREGION;
+
+ if (!LARGE_NUM_QPAIRS_SUPPORT(adapter))
+ return -EOPNOTSUPP;
+
+ adapter->current_op = VIRTCHNL_OP_GET_MAX_RSS_QREGION;
+ iavf_send_pf_msg(adapter, VIRTCHNL_OP_GET_MAX_RSS_QREGION, NULL, 0);
+ return 0;
+}
+
/**
* iavf_validate_num_queues
* @adapter: adapter structure
*
* Validate that the number of queues the PF has sent in
* VIRTCHNL_OP_GET_VF_RESOURCES is not larger than the VF can handle.
**/
static void iavf_validate_num_queues(struct iavf_adapter *adapter)
{
- if (adapter->vf_res->num_queue_pairs > IAVF_MAX_REQ_QUEUES) {
+ u32 max_req_queues = IAVF_MAX_REQ_QUEUES;
+
+ if (!LARGE_NUM_QPAIRS_SUPPORT(adapter))
+ max_req_queues = IAVF_MAX_VSI_QP;
+
+ if (adapter->vf_res->num_queue_pairs > max_req_queues) {
struct virtchnl_vsi_resource *vsi_res;
int i;
dev_info(&adapter->pdev->dev, "Received %d queues, but can only have a max of %d\n",
adapter->vf_res->num_queue_pairs,
- IAVF_MAX_REQ_QUEUES);
+ max_req_queues);
dev_info(&adapter->pdev->dev, "Fixing by reducing queues to %d\n",
- IAVF_MAX_REQ_QUEUES);
- adapter->vf_res->num_queue_pairs = IAVF_MAX_REQ_QUEUES;
+ max_req_queues);
+ adapter->vf_res->num_queue_pairs = max_req_queues;
for (i = 0; i < adapter->vf_res->num_vsis; i++) {
vsi_res = &adapter->vf_res->vsi_res[i];
- vsi_res->num_queue_pairs = IAVF_MAX_REQ_QUEUES;
+ vsi_res->num_queue_pairs = max_req_queues;
}
}
}
@@ -379,6 +397,30 @@ int iavf_get_vf_ptp_caps(struct iavf_adapter *adapter)
return err;
}
+int iavf_get_max_rss_qregion(struct iavf_adapter *adapter)
+{
+ struct iavf_arq_event_info event;
+ int err;
+ u16 len;
+
+ len = sizeof(struct virtchnl_max_rss_qregion);
+ event.buf_len = len;
+ event.msg_buf = kzalloc(len, GFP_KERNEL);
+ if (!event.msg_buf)
+ return -ENOMEM;
+
+ err = iavf_poll_virtchnl_msg(&adapter->hw, &event,
+ VIRTCHNL_OP_GET_MAX_RSS_QREGION);
+ if (!err)
+ memcpy(&adapter->max_rss_qregion, event.msg_buf,
+ min(event.msg_len, len));
+
+ adapter->current_op = VIRTCHNL_OP_UNKNOWN;
+
+ kfree(event.msg_buf);
+ return err;
+}
+
/**
* iavf_configure_queues
* @adapter: adapter structure
@@ -495,6 +537,50 @@ void iavf_configure_queues(struct iavf_adapter *adapter)
kfree(vqci);
}
+/**
+ * iavf_enable_disable_queues_v2 - send V2 messages of ENABLE/DISABLE queues ops
+ * @adapter: private adapter structure
+ * @enable: true to enable and false to disable the queues
+ */
+static void iavf_enable_disable_queues_v2(struct iavf_adapter *adapter, bool enable)
+{
+ enum virtchnl_ops op = VIRTCHNL_OP_ENABLE_QUEUES_V2;
+ struct virtchnl_del_ena_dis_queues *msg;
+ u64 flag = IAVF_FLAG_AQ_ENABLE_QUEUES;
+ struct virtchnl_queue_chunk *chunk;
+ int len;
+
+ if (!enable) {
+ op = VIRTCHNL_OP_DISABLE_QUEUES_V2;
+ flag = IAVF_FLAG_AQ_DISABLE_QUEUES;
+ }
+
+ /* We need 2 chunks (one Tx and one Rx). */
+ len = virtchnl_struct_size(msg, chunks, 2);
+ msg = kzalloc(len, GFP_KERNEL);
+ if (!msg)
+ return;
+
+ adapter->current_op = op;
+
+ msg->vport_id = adapter->vsi_res->vsi_id;
+ msg->num_chunks = 2;
+
+ chunk = &msg->chunks[0];
+ chunk->type = VIRTCHNL_QUEUE_TYPE_RX;
+ chunk->start_queue_id = 0;
+ chunk->num_queues = adapter->num_active_queues;
+
+ chunk++;
+ chunk->type = VIRTCHNL_QUEUE_TYPE_TX;
+ chunk->start_queue_id = 0;
+ chunk->num_queues = adapter->num_active_queues;
+
+ adapter->aq_required &= ~flag;
+ iavf_send_pf_msg(adapter, op, (u8 *)msg, len);
+ kfree(msg);
+}
+
/**
* iavf_enable_queues
* @adapter: adapter structure
@@ -511,6 +597,12 @@ void iavf_enable_queues(struct iavf_adapter *adapter)
adapter->current_op);
return;
}
+
+ if (LARGE_NUM_QPAIRS_SUPPORT(adapter)) {
+ iavf_enable_disable_queues_v2(adapter, true);
+ return;
+ }
+
adapter->current_op = VIRTCHNL_OP_ENABLE_QUEUES;
vqs.vsi_id = adapter->vsi_res->vsi_id;
vqs.tx_queues = BIT(adapter->num_active_queues) - 1;
@@ -536,15 +628,108 @@ void iavf_disable_queues(struct iavf_adapter *adapter)
adapter->current_op);
return;
}
+
+ if (LARGE_NUM_QPAIRS_SUPPORT(adapter)) {
+ iavf_enable_disable_queues_v2(adapter, false);
+ return;
+ }
+
adapter->current_op = VIRTCHNL_OP_DISABLE_QUEUES;
vqs.vsi_id = adapter->vsi_res->vsi_id;
vqs.tx_queues = BIT(adapter->num_active_queues) - 1;
vqs.rx_queues = vqs.tx_queues;
adapter->aq_required &= ~IAVF_FLAG_AQ_DISABLE_QUEUES;
iavf_send_pf_msg(adapter, VIRTCHNL_OP_DISABLE_QUEUES,
(u8 *)&vqs, sizeof(vqs));
}
+static void iavf_map_queue_vector(struct iavf_adapter *adapter)
+{
+ struct virtchnl_queue_vector_maps *qvmaps;
+ int qnum = adapter->num_active_queues;
+ struct virtchnl_queue_vector *qv;
+ struct iavf_arq_event_info event;
+ int len, max_pairs, err = 0;
+
+ max_pairs = iavf_max_vc_entries(qvmaps, qv_maps) / 2;
+ len = virtchnl_struct_size(qvmaps, qv_maps, 2 * min(qnum, max_pairs));
+ qvmaps = kzalloc(len, GFP_KERNEL);
+ if (!qvmaps)
+ return;
+
+ /* any message may arrive while polling, not only the awaited one */
+ event.buf_len = IAVF_MAX_AQ_BUF_SIZE;
+ event.msg_buf = kzalloc(IAVF_MAX_AQ_BUF_SIZE, GFP_KERNEL);
+ if (!event.msg_buf) {
+ kfree(qvmaps);
+ return;
+ }
+
+ qvmaps->vport_id = adapter->vsi_res->vsi_id;
+ qv = qvmaps->qv_maps;
+ for (int qid = 0, in_msg = 0; qid < qnum; qid++) {
+ const bool last = qid + 1 == qnum;
+ struct iavf_q_vector *q_vector;
+
+ q_vector = adapter->tx_rings[qid].q_vector;
+ qv->queue_id = qid;
+ qv->vector_id = NONQ_VECS + q_vector->v_idx;
+ qv->itr_idx = IAVF_TX_ITR;
+ qv->queue_type = VIRTCHNL_QUEUE_TYPE_TX;
+ qv++;
+
+ q_vector = adapter->rx_rings[qid].q_vector;
+ qv->queue_id = qid;
+ qv->vector_id = NONQ_VECS + q_vector->v_idx;
+ qv->itr_idx = IAVF_RX_ITR;
+ qv->queue_type = VIRTCHNL_QUEUE_TYPE_RX;
+ qv++;
+
+ in_msg++;
+ if (last || in_msg == max_pairs) {
+ qvmaps->num_qv_maps = 2 * in_msg;
+ adapter->current_op = VIRTCHNL_OP_MAP_QUEUE_VECTOR;
+ err = iavf_send_pf_msg(adapter,
+ VIRTCHNL_OP_MAP_QUEUE_VECTOR,
+ (u8 *)qvmaps,
+ virtchnl_struct_size(qvmaps, qv_maps,
+ 2 * in_msg));
+ if (err) {
+ dev_err(&adapter->pdev->dev,
+ "sending mapping queue vectors to PF failed, err: %d\n",
+ err);
+ break;
+ }
+ err = iavf_poll_virtchnl_response(adapter, &event,
+ iavf_match_vc_op_cb,
+ (void *)VIRTCHNL_OP_MAP_QUEUE_VECTOR,
+ 1000);
+ if (err) {
+ dev_err(&adapter->pdev->dev,
+ "polling response of mapping queue vectors failed, err: %d\n",
+ err);
+ break;
+ }
+
+ /* PF NACK, iavf_virtchnl_completion() has logged it */
+ if (event.desc.cookie_low) {
+ err = -EIO;
+ break;
+ }
+
+ in_msg = 0;
+ qv = qvmaps->qv_maps;
+ }
+ }
+ /* on timeout keep the op pending, see iavf_configure_queues() */
+ if (err != -EAGAIN) {
+ adapter->aq_required &= ~IAVF_FLAG_AQ_MAP_VECTORS;
+ adapter->current_op = VIRTCHNL_OP_UNKNOWN;
+ }
+ kfree(event.msg_buf);
+ kfree(qvmaps);
+}
+
/**
* iavf_map_queues
* @adapter: adapter structure
@@ -566,6 +751,12 @@ void iavf_map_queues(struct iavf_adapter *adapter)
adapter->current_op);
return;
}
+
+ if (LARGE_NUM_QPAIRS_SUPPORT(adapter)) {
+ iavf_map_queue_vector(adapter);
+ return;
+ }
+
adapter->current_op = VIRTCHNL_OP_CONFIG_IRQ_MAP;
q_vectors = adapter->num_msix_vectors - NONQ_VECS;
@@ -602,6 +793,40 @@ void iavf_map_queues(struct iavf_adapter *adapter)
kfree(vimi);
}
+/**
+ * iavf_request_queues - request queues via virtchnl
+ * @adapter: adapter structure
+ * @num: number of requested queues
+ *
+ * Return: 0 on success, negative on failure.
+ */
+int iavf_request_queues(struct iavf_adapter *adapter, int num)
+{
+ struct virtchnl_vf_res_request vfres = { num };
+ int err;
+
+ if (adapter->current_op != VIRTCHNL_OP_UNKNOWN) {
+ /* bail because we already have a command pending */
+ dev_err(&adapter->pdev->dev, "Cannot request queues, command %d pending\n",
+ adapter->current_op);
+ return -EBUSY;
+ }
+
+ adapter->current_op = VIRTCHNL_OP_REQUEST_QUEUES;
+ adapter->flags |= IAVF_FLAG_REINIT_ITR_NEEDED;
+ adapter->num_req_queues = num;
+
+ err = iavf_send_pf_msg(adapter, VIRTCHNL_OP_REQUEST_QUEUES,
+ (u8 *)&vfres, sizeof(vfres));
+ if (err) {
+ adapter->current_op = VIRTCHNL_OP_UNKNOWN;
+ adapter->flags &= ~IAVF_FLAG_REINIT_ITR_NEEDED;
+ adapter->num_req_queues = 0;
+ }
+
+ return err;
+}
+
/**
* iavf_set_mac_addr_type - Set the correct request type from the filter type
* @virtchnl_ether_addr: pointer to requested list element
@@ -2758,6 +2983,11 @@ void iavf_virtchnl_completion(struct iavf_adapter *adapter,
aq_required;
}
break;
+ case VIRTCHNL_OP_GET_MAX_RSS_QREGION:
+ if (!v_retval)
+ memcpy(&adapter->max_rss_qregion, msg,
+ sizeof(adapter->max_rss_qregion));
+ break;
case VIRTCHNL_OP_GET_SUPPORTED_RXDIDS:
if (msglen != sizeof(u64))
return;
@@ -2778,20 +3008,23 @@ void iavf_virtchnl_completion(struct iavf_adapter *adapter,
iavf_virtchnl_ptp_get_time(adapter, msg, msglen);
break;
case VIRTCHNL_OP_ENABLE_QUEUES:
+ case VIRTCHNL_OP_ENABLE_QUEUES_V2:
/* enable transmits */
iavf_irq_enable(adapter, true);
adapter->flags &= ~IAVF_FLAG_QUEUES_DISABLED;
break;
case VIRTCHNL_OP_DISABLE_QUEUES:
+ case VIRTCHNL_OP_DISABLE_QUEUES_V2:
iavf_free_all_tx_resources(adapter);
iavf_free_all_rx_resources(adapter);
if (adapter->state == __IAVF_DOWN_PENDING) {
iavf_change_state(adapter, __IAVF_DOWN);
wake_up(&adapter->down_waitqueue);
}
break;
case VIRTCHNL_OP_VERSION:
case VIRTCHNL_OP_CONFIG_IRQ_MAP:
+ case VIRTCHNL_OP_MAP_QUEUE_VECTOR:
/* Don't display an error if we get these out of sequence.
* If the firmware needed to get kicked, we'll get these and
* it's no problem.
--
2.51.1
^ permalink raw reply related [flat|nested] 15+ messages in thread
* [PATCH net-next v2 11/14] ice: introduce handling of virtchnl LARGE VF opcodes
2026-10-09 12:04 [PATCH net-next v2 00/14] devlink, mlx5, iavf, ice: XLVF for iavf Przemek Kitszel
` (9 preceding siblings ...)
2026-10-09 12:04 ` [PATCH net-next v2 10/14] iavf: use new opcodes to request more than 16 queues Przemek Kitszel
@ 2026-10-09 12:04 ` Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 12/14] devlink: give user option to allocate resources Przemek Kitszel
` (2 subsequent siblings)
13 siblings, 0 replies; 15+ messages in thread
From: Przemek Kitszel @ 2026-10-09 12:04 UTC (permalink / raw)
To: netdev, Jakub Kicinski, Jiri Pirko
Cc: Tony Nguyen, Aleksandr Loktionov, Michal Schmidt, intel-wired-lan,
edumazet, horms, pabeni, davem, Jonathan Corbet, skhan, rdunlap,
andrew+netdev, saeedm, tariqt, leon, mbloch, jacob.e.keller,
jedrzej.jagielski, anzaki, brett.creeley, jtornosm, ohartoov,
Przemek Kitszel
From: Brett Creeley <brett.creeley@intel.com>
With new virtchnl offload/capability VFs are able to make use of more than
16 queues. But the old opcodes were designed with a max of 16 queues,
so new ones were added (by iavf/virtchnl commit of this series).
Handle VIRTCHNL_OP_ENABLE_QUEUES_V2, VIRTCHNL_OP_DISABLE_QUEUES_V2 and
VIRTCHNL_OP_MAP_QUEUE_VECTOR here. VIRTCHNL_OP_GET_MAX_RSS_QREGION is
only allowlisted, its handler comes with bigger RSS LUT support later
in the series.
If a VF wishes to request >16 queues it should first make sure that the
PF supports the VIRTCHNL_VF_LARGE_NUM_QPAIRS capability.
Signed-off-by: Brett Creeley <brett.creeley@intel.com>
Co-developed-by: Aleksandr Loktionov <aleksandr.loktionov@intel.com> # msglen val
Signed-off-by: Aleksandr Loktionov <aleksandr.loktionov@intel.com>
Co-developed-by: Przemek Kitszel <przemyslaw.kitszel@intel.com>
Signed-off-by: Przemek Kitszel <przemyslaw.kitszel@intel.com>
---
v2:
* fix error response in one case of ice_vc_map_q_vector_msg() (Alex)
* use ice_get_vf_vsi() instead of plain array access (Sashiko)
* return VIRTCHNL_STATUS_ERR_PARAM on error (Sashiko)
* commit message: GET_MAX_RSS_QREGION is handled only later
---
drivers/net/ethernet/intel/ice/ice_vf_lib.h | 1 +
drivers/net/ethernet/intel/ice/virt/queues.h | 3 +
.../net/ethernet/intel/ice/virt/allowlist.c | 8 +
drivers/net/ethernet/intel/ice/virt/queues.c | 314 ++++++++++++++++++
4 files changed, 326 insertions(+)
diff --git a/drivers/net/ethernet/intel/ice/ice_vf_lib.h b/drivers/net/ethernet/intel/ice/ice_vf_lib.h
index a9e95c85bc98..396dfb2b2d97 100644
--- a/drivers/net/ethernet/intel/ice/ice_vf_lib.h
+++ b/drivers/net/ethernet/intel/ice/ice_vf_lib.h
@@ -125,6 +125,7 @@ struct ice_vf_ops {
void (*clear_reset_trigger)(struct ice_vf *vf);
void (*irq_close)(struct ice_vf *vf);
void (*post_vsi_rebuild)(struct ice_vf *vf);
+ struct ice_q_vector *(*get_q_vector)(struct ice_vsi *vsi, u16 vec_id);
};
/* Virtchnl/SR-IOV config info */
diff --git a/drivers/net/ethernet/intel/ice/virt/queues.h b/drivers/net/ethernet/intel/ice/virt/queues.h
index c4a792cecea1..223f609dd4f3 100644
--- a/drivers/net/ethernet/intel/ice/virt/queues.h
+++ b/drivers/net/ethernet/intel/ice/virt/queues.h
@@ -16,5 +16,8 @@ int ice_vc_cfg_q_bw(struct ice_vf *vf, u8 *msg);
int ice_vc_cfg_q_quanta(struct ice_vf *vf, u8 *msg);
int ice_vc_cfg_qs_msg(struct ice_vf *vf, u8 *msg);
int ice_vc_request_qs_msg(struct ice_vf *vf, u8 *msg);
+int ice_vc_ena_qs_v2_msg(struct ice_vf *vf, u8 *msg, u16 msglen);
+int ice_vc_dis_qs_v2_msg(struct ice_vf *vf, u8 *msg, u16 msglen);
+int ice_vc_map_q_vector_msg(struct ice_vf *vf, u8 *msg, u16 msglen);
#endif /* _ICE_VIRT_QUEUES_H_ */
diff --git a/drivers/net/ethernet/intel/ice/virt/allowlist.c b/drivers/net/ethernet/intel/ice/virt/allowlist.c
index a07efec19c45..ef769b843c6f 100644
--- a/drivers/net/ethernet/intel/ice/virt/allowlist.c
+++ b/drivers/net/ethernet/intel/ice/virt/allowlist.c
@@ -95,6 +95,13 @@ static const u32 tc_allowlist_opcodes[] = {
VIRTCHNL_OP_CONFIG_QUANTA,
};
+static const u32 large_num_qpairs_allowlist_opcodes[] = {
+ VIRTCHNL_OP_GET_MAX_RSS_QREGION,
+ VIRTCHNL_OP_ENABLE_QUEUES_V2,
+ VIRTCHNL_OP_DISABLE_QUEUES_V2,
+ VIRTCHNL_OP_MAP_QUEUE_VECTOR,
+};
+
struct allowlist_opcode_info {
const u32 *opcodes;
size_t size;
@@ -117,6 +124,7 @@ static const struct allowlist_opcode_info allowlist_opcodes[] = {
ALLOW_ITEM(VIRTCHNL_VF_OFFLOAD_VLAN_V2, vlan_v2_allowlist_opcodes),
ALLOW_ITEM(VIRTCHNL_VF_OFFLOAD_QOS, tc_allowlist_opcodes),
ALLOW_ITEM(VIRTCHNL_VF_CAP_PTP, ptp_allowlist_opcodes),
+ ALLOW_ITEM(VIRTCHNL_VF_LARGE_NUM_QPAIRS, large_num_qpairs_allowlist_opcodes),
};
/**
diff --git a/drivers/net/ethernet/intel/ice/virt/queues.c b/drivers/net/ethernet/intel/ice/virt/queues.c
index 61304493af34..3c5c5099a73d 100644
--- a/drivers/net/ethernet/intel/ice/virt/queues.c
+++ b/drivers/net/ethernet/intel/ice/virt/queues.c
@@ -1044,3 +1044,317 @@ int ice_vc_request_qs_msg(struct ice_vf *vf, u8 *msg)
v_ret, (u8 *)vfres, sizeof(*vfres));
}
+static bool ice_vc_supported_queue_type(s32 queue_type)
+{
+ return queue_type == VIRTCHNL_QUEUE_TYPE_RX ||
+ queue_type == VIRTCHNL_QUEUE_TYPE_TX;
+}
+
+/**
+ * ice_vc_validate_qs_v2_msg - validate all qs_msg parameters
+ * @vf: VF the message was received from
+ * @qs_msg: contents of the message from the VF
+ * @msglen: length of @qs_msg
+ *
+ * Used to validate both the VIRTCHNL_OP_ENABLE_QUEUES_V2 and
+ * VIRTCHNL_OP_DISABLE_QUEUES_V2 messages. This should always be called before
+ * attempting to enable and/or disable queues on behalf of a VF in response to
+ * the previously mentioned opcodes.
+ *
+ * Return: If all checks succeed, then return true. Otherwise return
+ * false, indicating to the caller that the qs_msg is invalid.
+ */
+static bool ice_vc_validate_qs_v2_msg(struct ice_vf *vf,
+ struct virtchnl_del_ena_dis_queues *qs_msg,
+ u16 msglen)
+{
+ if (msglen < virtchnl_struct_size(qs_msg, chunks, 0))
+ return false;
+
+ if (msglen < virtchnl_struct_size(qs_msg, chunks, qs_msg->num_chunks))
+ return false;
+
+ if (!qs_msg->num_chunks)
+ return false;
+
+ if (!test_bit(ICE_VF_STATE_ACTIVE, vf->vf_states))
+ return false;
+
+ if (!ice_vc_isvalid_vsi_id(vf, qs_msg->vport_id))
+ return false;
+
+ for (int i = 0; i < qs_msg->num_chunks; i++) {
+ u32 max_queue_in_chunk;
+
+ if (!ice_vc_supported_queue_type(qs_msg->chunks[i].type))
+ return false;
+
+ if (!qs_msg->chunks[i].num_queues)
+ return false;
+
+ max_queue_in_chunk = qs_msg->chunks[i].start_queue_id +
+ qs_msg->chunks[i].num_queues;
+ if (max_queue_in_chunk > vf->num_vf_qs)
+ return false;
+ }
+
+ return true;
+}
+
+#define ice_for_each_q_in_chunk(chunk, q_id) \
+ for ((q_id) = (chunk)->start_queue_id; \
+ (q_id) < (chunk)->start_queue_id + (chunk)->num_queues; \
+ (q_id)++)
+
+static int
+ice_vc_ena_rxq_chunk(struct ice_vf *vf, struct virtchnl_queue_chunk *chunk)
+{
+ struct ice_vsi *vsi;
+ u32 vf_qid;
+
+ ice_for_each_q_in_chunk(chunk, vf_qid) {
+ int err;
+
+ vsi = ice_get_vf_vsi(vf);
+ err = ice_vf_vsi_ena_single_rxq(vf, vsi, vf_qid);
+ if (err)
+ return err;
+ }
+
+ return 0;
+}
+
+static int
+ice_vc_ena_txq_chunk(struct ice_vf *vf, struct virtchnl_queue_chunk *chunk)
+{
+ struct ice_vsi *vsi;
+ u32 vf_qid;
+
+ ice_for_each_q_in_chunk(chunk, vf_qid) {
+ vsi = ice_get_vf_vsi(vf);
+ ice_vf_vsi_ena_single_txq(vf, vsi, vf_qid);
+ }
+
+ return 0;
+}
+
+/**
+ * ice_vc_ena_qs_v2_msg - message handling for VIRTCHNL_OP_ENABLE_QUEUES_V2
+ * @vf: source of the request
+ * @msg: message to handle
+ * @msglen: length of the @msg
+ *
+ * Return: 0 on success or negative on error.
+ */
+int ice_vc_ena_qs_v2_msg(struct ice_vf *vf, u8 *msg, u16 msglen)
+{
+ struct virtchnl_del_ena_dis_queues *ena_qs_msg =
+ (struct virtchnl_del_ena_dis_queues *)msg;
+ enum virtchnl_status_code v_ret = VIRTCHNL_STATUS_ERR_PARAM;
+
+ if (!ice_vc_validate_qs_v2_msg(vf, ena_qs_msg, msglen))
+ goto error_param;
+
+ for (int i = 0; i < ena_qs_msg->num_chunks; i++) {
+ struct virtchnl_queue_chunk *chunk = &ena_qs_msg->chunks[i];
+
+ if (chunk->type == VIRTCHNL_QUEUE_TYPE_RX &&
+ ice_vc_ena_rxq_chunk(vf, chunk))
+ goto error_param;
+ else if (chunk->type == VIRTCHNL_QUEUE_TYPE_TX &&
+ ice_vc_ena_txq_chunk(vf, chunk))
+ goto error_param;
+ }
+
+ v_ret = VIRTCHNL_STATUS_SUCCESS;
+ set_bit(ICE_VF_STATE_QS_ENA, vf->vf_states);
+
+error_param:
+ return ice_vc_send_msg_to_vf(vf, VIRTCHNL_OP_ENABLE_QUEUES_V2,
+ v_ret, NULL, 0);
+}
+
+static int
+ice_vc_dis_rxq_chunk(struct ice_vf *vf, struct virtchnl_queue_chunk *chunk)
+{
+ struct ice_vsi *vsi;
+ u32 vf_qid;
+
+ ice_for_each_q_in_chunk(chunk, vf_qid) {
+ int err;
+
+ vsi = ice_get_vf_vsi(vf);
+ err = ice_vf_vsi_dis_single_rxq(vf, vsi, vf_qid);
+ if (err)
+ return err;
+ }
+
+ return 0;
+}
+
+static int
+ice_vc_dis_txq_chunk(struct ice_vf *vf, struct virtchnl_queue_chunk *chunk)
+{
+ struct ice_vsi *vsi;
+ u32 vf_qid;
+
+ ice_for_each_q_in_chunk(chunk, vf_qid) {
+ int err;
+
+ vsi = ice_get_vf_vsi(vf);
+ err = ice_vf_vsi_dis_single_txq(vf, vsi, vf_qid);
+ if (err)
+ return err;
+ }
+
+ return 0;
+}
+
+/**
+ * ice_vc_dis_qs_v2_msg - message handling for VIRTCHNL_OP_DISABLE_QUEUES_V2
+ * @vf: source of the request
+ * @msg: message to handle
+ * @msglen: length of @msg
+ *
+ * Return: 0 on success or negative on error.
+ */
+int ice_vc_dis_qs_v2_msg(struct ice_vf *vf, u8 *msg, u16 msglen)
+{
+ struct virtchnl_del_ena_dis_queues *dis_qs_msg =
+ (struct virtchnl_del_ena_dis_queues *)msg;
+ enum virtchnl_status_code v_ret = VIRTCHNL_STATUS_ERR_PARAM;
+
+ if (!ice_vc_validate_qs_v2_msg(vf, dis_qs_msg, msglen))
+ goto error_param;
+
+ for (int i = 0; i < dis_qs_msg->num_chunks; i++) {
+ struct virtchnl_queue_chunk *chunk = &dis_qs_msg->chunks[i];
+
+ if (chunk->type == VIRTCHNL_QUEUE_TYPE_RX &&
+ ice_vc_dis_rxq_chunk(vf, chunk))
+ goto error_param;
+ else if (chunk->type == VIRTCHNL_QUEUE_TYPE_TX &&
+ ice_vc_dis_txq_chunk(vf, chunk))
+ goto error_param;
+ }
+
+ v_ret = VIRTCHNL_STATUS_SUCCESS;
+ if (ice_vf_has_no_qs_ena(vf))
+ clear_bit(ICE_VF_STATE_QS_ENA, vf->vf_states);
+
+error_param:
+ return ice_vc_send_msg_to_vf(vf, VIRTCHNL_OP_DISABLE_QUEUES_V2,
+ v_ret, NULL, 0);
+}
+
+/**
+ * ice_vc_validate_qv_maps - validate parameters sent in the qs_msg structure
+ * @vf: VF the message was received from
+ * @qv_maps: contents of the message from the VF
+ * @msglen: length of the @qv_maps
+ *
+ * Used to validate VIRTCHNL_OP_MAP_QUEUE_VECTOR messages. This should always
+ * be called before attempting map interrupts to queues. If all checks succeed,
+ * then return success indicating to the caller that the qv_maps are valid.
+ * Otherwise return false, indicating to the caller that the qv_maps are
+ * invalid.
+ *
+ * Return: true if parameters are valid, false otherwise.
+ */
+static bool ice_vc_validate_qv_maps(struct ice_vf *vf,
+ struct virtchnl_queue_vector_maps *qv_maps,
+ u16 msglen)
+{
+ struct ice_vsi *vsi;
+ int total_vectors;
+
+ vsi = ice_get_vf_vsi(vf);
+ if (!vsi)
+ return false;
+
+ if (msglen < virtchnl_struct_size(qv_maps, qv_maps, 0))
+ return false;
+
+ if (msglen < virtchnl_struct_size(qv_maps, qv_maps, qv_maps->num_qv_maps))
+ return false;
+
+ if (!qv_maps->num_qv_maps)
+ return false;
+
+ if (!test_bit(ICE_VF_STATE_ACTIVE, vf->vf_states))
+ return false;
+
+ if (!ice_vc_isvalid_vsi_id(vf, qv_maps->vport_id))
+ return false;
+
+ total_vectors = vsi->num_q_vectors + ICE_NONQ_VECS_VF;
+
+ for (int i = 0; i < qv_maps->num_qv_maps; i++) {
+ if (!ice_vc_supported_queue_type(qv_maps->qv_maps[i].queue_type))
+ return false;
+
+ if (qv_maps->qv_maps[i].queue_id >= vf->num_vf_qs)
+ return false;
+
+ if (qv_maps->qv_maps[i].vector_id >= total_vectors ||
+ qv_maps->qv_maps[i].vector_id < ICE_NONQ_VECS_VF)
+ return false;
+ }
+
+ return true;
+}
+
+/**
+ * ice_vc_map_q_vector_msg - message handling for VIRTCHNL_OP_MAP_QUEUE_VECTOR
+ * @vf: source of the request
+ * @msg: message to handle
+ * @msglen: length of @msg
+ *
+ * Return: 0 on success or negative on error
+ */
+int ice_vc_map_q_vector_msg(struct ice_vf *vf, u8 *msg, u16 msglen)
+{
+ enum virtchnl_status_code v_ret = VIRTCHNL_STATUS_ERR_PARAM;
+ struct virtchnl_queue_vector_maps *qv_maps;
+ struct ice_vsi *vsi;
+
+ qv_maps = (struct virtchnl_queue_vector_maps *)msg;
+
+ if (!ice_vc_validate_qv_maps(vf, qv_maps, msglen))
+ goto error_param;
+
+ for (int i = 0; i < qv_maps->num_qv_maps; i++) {
+ struct virtchnl_queue_vector *qv_map = &qv_maps->qv_maps[i];
+ struct ice_q_vector *q_vector;
+ u16 vector_id;
+ int vsi_q_id;
+
+ vsi_q_id = qv_map->queue_id;
+ vector_id = qv_map->vector_id;
+ vsi = ice_get_vf_vsi(vf);
+ if (!vsi)
+ goto error_param;
+
+ q_vector = vf->vf_ops->get_q_vector(vsi, vector_id);
+ if (!q_vector)
+ goto error_param;
+
+ if (!ice_vc_isvalid_q_id(vsi, vsi_q_id))
+ goto error_param;
+
+ if (qv_map->queue_type == VIRTCHNL_QUEUE_TYPE_RX)
+ ice_cfg_rxq_interrupt(vsi, vsi_q_id,
+ q_vector->vf_reg_idx,
+ qv_map->itr_idx);
+ else if (qv_map->queue_type == VIRTCHNL_QUEUE_TYPE_TX)
+ ice_cfg_txq_interrupt(vsi, vsi_q_id,
+ q_vector->vf_reg_idx,
+ qv_map->itr_idx);
+ }
+
+ v_ret = VIRTCHNL_STATUS_SUCCESS;
+
+error_param:
+ return ice_vc_send_msg_to_vf(vf, VIRTCHNL_OP_MAP_QUEUE_VECTOR,
+ v_ret, NULL, 0);
+}
--
2.51.1
^ permalink raw reply related [flat|nested] 15+ messages in thread
* [PATCH net-next v2 12/14] devlink: give user option to allocate resources
2026-10-09 12:04 [PATCH net-next v2 00/14] devlink, mlx5, iavf, ice: XLVF for iavf Przemek Kitszel
` (10 preceding siblings ...)
2026-10-09 12:04 ` [PATCH net-next v2 11/14] ice: introduce handling of virtchnl LARGE VF opcodes Przemek Kitszel
@ 2026-10-09 12:04 ` Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 13/14] ice: represent RSS LUTs as devlink resources Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 14/14] ice: support up to 256 VF queues Przemek Kitszel
13 siblings, 0 replies; 15+ messages in thread
From: Przemek Kitszel @ 2026-10-09 12:04 UTC (permalink / raw)
To: netdev, Jakub Kicinski, Jiri Pirko
Cc: Tony Nguyen, Aleksandr Loktionov, Michal Schmidt, intel-wired-lan,
edumazet, horms, pabeni, davem, Jonathan Corbet, skhan, rdunlap,
andrew+netdev, saeedm, tariqt, leon, mbloch, jacob.e.keller,
jedrzej.jagielski, anzaki, brett.creeley, jtornosm, ohartoov,
Przemek Kitszel
Current devlink resources are designed as a thing that user could limit,
but there is not much otherwise that could be done with them.
Perhaps that's the reason there is not much adoption despite API being
there for multiple years.
Add new mode of operation, where user could allocate/assign resources
(from a common pool) to specific devices.
That requires a new .occ_set() driver callback, called on
DEVLINK_CMD_RESOURCE_SET with the requested DEVLINK_ATTR_RESOURCE_SIZE,
instead of staging the size until the next reload. Driver registers it
together with .occ_get() by devl_resource_occ_set_get_register(). For
such resources the getter value is reported as DEVLINK_ATTR_RESOURCE_SIZE,
as opposed to DEVLINK_ATTR_RESOURCE_OCC (current occupation) in the
"legacy" mode, and there is no DEVLINK_ATTR_RESOURCE_SIZE_NEW.
Signed-off-by: Przemek Kitszel <przemyslaw.kitszel@intel.com>
---
RFC, Feb '25
https://lore.kernel.org/intel-wired-lan/20250219164410.35665-3-przemyslaw.kitszel@intel.com
I have structured code to just choose "the mode" of operation based on
the presence of .occ_set callback, instead of naming the mode, what,
even if not the most clear code, avoids the need of naming the modes.
v2:
* avoid setting DEVLINK_ATTR_RESOURCE_SIZE_VALID for leaf resources
(Sashiko)
* minor kdoc update (Sashiko)
* add more WARN_ONs (Sashiko)
* document the occ_set mode in devlink-resource.rst
* commit message: describe the .occ_set() callback and how the reported
attributes differ from the "legacy" mode
---
.../networking/devlink/devlink-resource.rst | 5 ++
include/net/devlink.h | 7 ++
net/devlink/resource.c | 88 +++++++++++++++----
3 files changed, 84 insertions(+), 16 deletions(-)
diff --git a/Documentation/networking/devlink/devlink-resource.rst b/Documentation/networking/devlink/devlink-resource.rst
index 47eec8f875b4..dad7a6d7bdcf 100644
--- a/Documentation/networking/devlink/devlink-resource.rst
+++ b/Documentation/networking/devlink/devlink-resource.rst
@@ -75,6 +75,11 @@ attribute, which represents the pending change in size. For example:
Note that changes in resource size may require a device reload to properly
take effect.
+Resources registered with ``devl_resource_occ_set_get_register()`` behave
+differently: the new size is passed to the driver and applies immediately,
+there is neither ``size_new`` nor ``occ`` reported, and ``size`` is the value
+returned by the driver's getter.
+
Port-level Resources and Full Dump
==================================
diff --git a/include/net/devlink.h b/include/net/devlink.h
index 98bdde4f72ca..d01128ef393d 100644
--- a/include/net/devlink.h
+++ b/include/net/devlink.h
@@ -420,6 +420,8 @@ devlink_resource_size_params_init(struct devlink_resource_size_params *size_para
}
typedef u64 devlink_resource_occ_get_t(void *priv);
+typedef int devlink_resource_occ_set_t(u64 size, struct netlink_ext_ack *extack,
+ void *priv);
#define DEVLINK_RESOURCE_ID_PARENT_TOP 0
@@ -1959,6 +1961,11 @@ void devl_resource_occ_get_register(struct devlink *devlink,
void *occ_get_priv);
void devl_resource_occ_get_unregister(struct devlink *devlink,
u64 resource_id);
+void devl_resource_occ_set_get_register(struct devlink *devlink,
+ u64 resource_id,
+ devlink_resource_occ_set_t *occ_set,
+ devlink_resource_occ_get_t *occ_get,
+ void *occ_priv);
int devl_params_register(struct devlink *devlink,
const struct devlink_param *params,
size_t params_count);
diff --git a/net/devlink/resource.c b/net/devlink/resource.c
index c3cfda7ea070..a13fdaddcb19 100644
--- a/net/devlink/resource.c
+++ b/net/devlink/resource.c
@@ -19,7 +19,8 @@
* @list: parent list
* @resource_list: list of child resources
* @occ_get: occupancy getter callback
- * @occ_get_priv: occupancy getter callback priv
+ * @occ_set: occupancy setter callback
+ * @occ_priv: occupancy callbacks priv
*/
struct devlink_resource {
const char *name;
@@ -32,7 +33,8 @@ struct devlink_resource {
struct list_head list;
struct list_head resource_list;
devlink_resource_occ_get_t *occ_get;
- void *occ_get_priv;
+ devlink_resource_occ_set_t *occ_set;
+ void *occ_priv;
};
static struct devlink_resource *
@@ -137,6 +139,9 @@ int devlink_nl_resource_set_doit(struct sk_buff *skb, struct genl_info *info)
if (err)
return err;
+ if (resource->occ_set)
+ return resource->occ_set(size, info->extack, resource->occ_priv);
+
resource->size_new = size;
devlink_resource_validate_children(resource);
if (resource->parent)
@@ -162,13 +167,34 @@ devlink_resource_size_params_put(struct devlink_resource *resource,
return 0;
}
-static int devlink_resource_occ_put(struct devlink_resource *resource,
- struct sk_buff *skb)
+static int
+devlink_resource_occ_size_put_legacy(struct devlink_resource *resource,
+ struct sk_buff *skb)
{
- if (!resource->occ_get)
- return 0;
- return devlink_nl_put_u64(skb, DEVLINK_ATTR_RESOURCE_OCC,
- resource->occ_get(resource->occ_get_priv));
+ if (devlink_nl_put_u64(skb, DEVLINK_ATTR_RESOURCE_SIZE, resource->size))
+ goto err;
+ if (resource->size != resource->size_new &&
+ devlink_nl_put_u64(skb, DEVLINK_ATTR_RESOURCE_SIZE_NEW,
+ resource->size_new))
+ goto err;
+ if (resource->occ_get &&
+ devlink_nl_put_u64(skb, DEVLINK_ATTR_RESOURCE_OCC,
+ resource->occ_get(resource->occ_priv)))
+ goto err;
+
+ return 0;
+err:
+ return -EMSGSIZE;
+}
+
+static int devlink_resource_occ_size_put(struct devlink_resource *resource,
+ struct sk_buff *skb)
+{
+ if (!resource->occ_get || !resource->occ_set)
+ return devlink_resource_occ_size_put_legacy(resource, skb);
+
+ return devlink_nl_put_u64(skb, DEVLINK_ATTR_RESOURCE_SIZE,
+ resource->occ_get(resource->occ_priv));
}
static int devlink_resource_put(struct devlink *devlink, struct sk_buff *skb,
@@ -183,14 +209,9 @@ static int devlink_resource_put(struct devlink *devlink, struct sk_buff *skb,
return -EMSGSIZE;
if (nla_put_string(skb, DEVLINK_ATTR_RESOURCE_NAME, resource->name) ||
- devlink_nl_put_u64(skb, DEVLINK_ATTR_RESOURCE_SIZE, resource->size) ||
devlink_nl_put_u64(skb, DEVLINK_ATTR_RESOURCE_ID, resource->id))
goto nla_put_failure;
- if (resource->size != resource->size_new &&
- devlink_nl_put_u64(skb, DEVLINK_ATTR_RESOURCE_SIZE_NEW,
- resource->size_new))
- goto nla_put_failure;
- if (devlink_resource_occ_put(resource, skb))
+ if (devlink_resource_occ_size_put(resource, skb))
goto nla_put_failure;
if (devlink_resource_size_params_put(resource, skb))
goto nla_put_failure;
@@ -640,6 +661,39 @@ int devl_resource_size_get(struct devlink *devlink,
}
EXPORT_SYMBOL_GPL(devl_resource_size_get);
+/**
+ * devl_resource_occ_set_get_register - register occupancy getter and setter
+ *
+ * @devlink: devlink
+ * @resource_id: resource id
+ * @occ_set: occupancy setter callback
+ * @occ_get: occupancy getter callback
+ * @occ_priv: occupancy callbacks priv
+ *
+ * Setter will be called when the user wants to change the resource size,
+ * getter is called to show the user what is current size of the resource.
+ */
+void devl_resource_occ_set_get_register(struct devlink *devlink,
+ u64 resource_id,
+ devlink_resource_occ_set_t *occ_set,
+ devlink_resource_occ_get_t *occ_get,
+ void *occ_priv)
+{
+ struct devlink_resource *resource;
+
+ lockdep_assert_held(&devlink->lock);
+
+ resource = devlink_resource_find(devlink, NULL, resource_id);
+ if (WARN_ON(!resource))
+ return;
+ WARN_ON(resource->occ_get || resource->occ_set);
+
+ resource->occ_set = occ_set;
+ resource->occ_get = occ_get;
+ resource->occ_priv = occ_priv;
+}
+EXPORT_SYMBOL_GPL(devl_resource_occ_set_get_register);
+
/**
* devl_resource_occ_get_register - register occupancy getter
*
@@ -661,9 +715,10 @@ void devl_resource_occ_get_register(struct devlink *devlink,
if (WARN_ON(!resource))
return;
WARN_ON(resource->occ_get);
+ WARN_ON(resource->occ_set);
resource->occ_get = occ_get;
- resource->occ_get_priv = occ_get_priv;
+ resource->occ_priv = occ_get_priv;
}
EXPORT_SYMBOL_GPL(devl_resource_occ_get_register);
@@ -684,9 +739,10 @@ void devl_resource_occ_get_unregister(struct devlink *devlink,
if (WARN_ON(!resource))
return;
WARN_ON(!resource->occ_get);
+ WARN_ON(resource->occ_set);
resource->occ_get = NULL;
- resource->occ_get_priv = NULL;
+ resource->occ_priv = NULL;
}
EXPORT_SYMBOL_GPL(devl_resource_occ_get_unregister);
--
2.51.1
^ permalink raw reply related [flat|nested] 15+ messages in thread
* [PATCH net-next v2 13/14] ice: represent RSS LUTs as devlink resources
2026-10-09 12:04 [PATCH net-next v2 00/14] devlink, mlx5, iavf, ice: XLVF for iavf Przemek Kitszel
` (11 preceding siblings ...)
2026-10-09 12:04 ` [PATCH net-next v2 12/14] devlink: give user option to allocate resources Przemek Kitszel
@ 2026-10-09 12:04 ` Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 14/14] ice: support up to 256 VF queues Przemek Kitszel
13 siblings, 0 replies; 15+ messages in thread
From: Przemek Kitszel @ 2026-10-09 12:04 UTC (permalink / raw)
To: netdev, Jakub Kicinski, Jiri Pirko
Cc: Tony Nguyen, Aleksandr Loktionov, Michal Schmidt, intel-wired-lan,
edumazet, horms, pabeni, davem, Jonathan Corbet, skhan, rdunlap,
andrew+netdev, saeedm, tariqt, leon, mbloch, jacob.e.keller,
jedrzej.jagielski, anzaki, brett.creeley, jtornosm, ohartoov,
Przemek Kitszel
E800 family offers three kinds of RSS LUTs: VSI LUT (sized 64), GLOBAL LUT
(sized 512), and PF LUT (sized 2048). Until now the GLOBAL kind was not
used at all. There are two possible usages for it, subsequent commit will
give VF option to acquire it, and this one enables PF to switch between PF
LUT and GLOBAL LUT - switching to smaller one is, again, to make it
possible for VF to then acquire the former one.
Devlink resources are used to let user show current usage and change the
allocation, see examples below.
Default state on 8-port card, asking for aggregate "whole device" usage,
note that there are as many PF LUTs as there is PFs, and, for e810, there
are 16 GLOBAL LUTs:
$ devlink resource show devlink_index/11
devlink_index/11:
name rss size 8 unit entry size_min 0 size_max 24 size_gran 1 dpipe_tables none
resources:
name lut_512 size 0 unit entry size_min 0 size_max 16 size_gran 1 dpipe_tables none
name lut_2048 size 8 unit entry size_min 0 size_max 8 size_gran 1 dpipe_tables none
Now let's add GLOBAL LUT for a single PF (on one-port NIC):
$ sudo devlink resource set pci/0000:18:00.0 path rss/lut_512 size 1
And show it's resources after that:
$ devlink resource show pci/0000:18:00.0
pci/0000:18:00.0:
name rss size 2 unit entry size_min 0 size_max 2 size_gran 1 dpipe_tables none
resources:
name lut_512 size 1 unit entry size_min 0 size_max 1 size_gran 1 dpipe_tables none
name lut_2048 size 1 unit entry size_min 0 size_max 1 size_gran 1 dpipe_tables none
Let's take the PF LUT out of that PF afterwards:
$ sudo devlink resource set pci/0000:18:00.0 path rss/lut_2048 size 0
now `ethtool -x $ifacename` will report smaller RSS table.
Signed-off-by: Przemek Kitszel <przemyslaw.kitszel@intel.com>
---
v2:
* replace -ENOANO with proper WARN_ON for "assumption error" (Alex)
* fix goto label (Sashiko)
* avoid queues being 0 in ADQ mode (Sashiko)
* undo ice_take_rss_lut_pf() on error path (Sashiko)
* program the PF VSI context with the new LUT type and global LUT ID
after writing the LUT, so the PF really switches to the new table;
publish the global LUT ID to the VSI only after a successful switch
* drop the cached user RSS LUT when the RSS table size changes, on
devlink LUT switch and on reset
* use logical_pf_id for all PF bookkeeping of devlink resources, as the
PF LUT resource is sized by the number of functions
* warn when FW fails to free a global RSS LUT
* add pf->rss_lut_lock to serialize devlink RSS LUT switch with
ethtool -x/-X/-L, RXHASH toggle and VSI RSS config
* return -EAGAIN from get/set_rxfh when the LUT was resized after
the core sized the indirection table
* reject ADQ when RSS LUTs are reassigned via devlink, and reject
devlink RSS LUT changes when ADQ is configured (-EOPNOTSUPP)
* remove trailing whitespace
* commit message: GLOBAL LUT is sized 512, not 256
---
drivers/net/ethernet/intel/ice/Makefile | 1 +
.../net/ethernet/intel/ice/devlink/resource.h | 23 +
drivers/net/ethernet/intel/ice/ice.h | 1 +
drivers/net/ethernet/intel/ice/ice_adapter.h | 40 ++
drivers/net/ethernet/intel/ice/ice_common.h | 1 +
drivers/net/ethernet/intel/ice/ice_lib.h | 3 +
.../net/ethernet/intel/ice/devlink/resource.c | 527 ++++++++++++++++++
drivers/net/ethernet/intel/ice/ice_adapter.c | 12 +-
drivers/net/ethernet/intel/ice/ice_common.c | 2 +-
drivers/net/ethernet/intel/ice/ice_ethtool.c | 91 +--
drivers/net/ethernet/intel/ice/ice_lib.c | 129 +++--
drivers/net/ethernet/intel/ice/ice_main.c | 24 +-
12 files changed, 777 insertions(+), 77 deletions(-)
create mode 100644 drivers/net/ethernet/intel/ice/devlink/resource.h
create mode 100644 drivers/net/ethernet/intel/ice/devlink/resource.c
diff --git a/drivers/net/ethernet/intel/ice/Makefile b/drivers/net/ethernet/intel/ice/Makefile
index a952bacb7ec7..e50e02efbf7a 100644
--- a/drivers/net/ethernet/intel/ice/Makefile
+++ b/drivers/net/ethernet/intel/ice/Makefile
@@ -34,6 +34,7 @@ ice-y := ice_main.o \
devlink/devlink.o \
devlink/health.o \
devlink/port.o \
+ devlink/resource.o \
ice_sf_eth.o \
ice_sf_vsi_vlan_ops.o \
ice_ddp.o \
diff --git a/drivers/net/ethernet/intel/ice/devlink/resource.h b/drivers/net/ethernet/intel/ice/devlink/resource.h
new file mode 100644
index 000000000000..8999e0413c51
--- /dev/null
+++ b/drivers/net/ethernet/intel/ice/devlink/resource.h
@@ -0,0 +1,23 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/* Copyright (c) 2026, Intel Corporation. */
+
+#ifndef _ICE_DEVL_RESOURCE_H_
+#define _ICE_DEVL_RESOURCE_H_
+
+#include <linux/types.h>
+
+struct devlink;
+struct ice_adapter;
+struct ice_hw;
+struct ice_pf;
+
+void ice_devl_pf_resources_register(struct ice_pf *pf);
+void ice_devl_whole_dev_resources_register(const struct ice_hw *hw,
+ struct ice_adapter *adapter);
+
+bool ice_rss_lut_is_reassigned(struct ice_pf *pf);
+int ice_take_rss_lut_pf(struct ice_pf *pf);
+void ice_release_rss_lut_pf(struct ice_pf *pf);
+void ice_free_rss_lut_flr(struct ice_pf *pf);
+
+#endif /* _ICE_DEVL_RESOURCE_H_ */
diff --git a/drivers/net/ethernet/intel/ice/ice.h b/drivers/net/ethernet/intel/ice/ice.h
index 89a2bdbecf60..d28ff576a342 100644
--- a/drivers/net/ethernet/intel/ice/ice.h
+++ b/drivers/net/ethernet/intel/ice/ice.h
@@ -596,6 +596,7 @@ struct ice_pf {
struct mutex tc_mutex; /* lock to protect TC changes */
struct mutex adev_mutex; /* lock to protect aux device access */
struct mutex lag_mutex; /* protect ice_lag struct in PF */
+ struct mutex rss_lut_lock; /* protect VSI RSS LUT config */
u32 msg_enable;
struct ice_ptp ptp;
struct gnss_serial *gnss_serial;
diff --git a/drivers/net/ethernet/intel/ice/ice_adapter.h b/drivers/net/ethernet/intel/ice/ice_adapter.h
index c11d511d7a8e..95592e7e10d7 100644
--- a/drivers/net/ethernet/intel/ice/ice_adapter.h
+++ b/drivers/net/ethernet/intel/ice/ice_adapter.h
@@ -15,6 +15,40 @@
struct pci_dev;
struct ice_pf;
+enum ice_devl_resource_id {
+ /* keep parent IDs prior to children, because we register in order */
+ ICE_TOP_RESOURCE = DEVLINK_RESOURCE_ID_PARENT_TOP,
+ ICE_RSS_LUT_BOTH,
+ ICE_RSS_LUT_GLOBAL,
+ ICE_RSS_LUT_PF,
+ ICE_DEVL_RESOURCES_COUNT
+};
+
+#define ICE_MAX_DEVL_RESOURCE_UNITS 16
+
+/**
+ * struct ice_devl_resource - driver data for devlink resource, config & runtime
+ *
+ * @owner: entity that owns given resource (like ptr to ice_pf or ice_vf)
+ * @pf_id: on which PF the VF is (or just PF id when PF is the owner)
+ * @name: name of the resource to register it with
+ * @get: occ getter callback
+ * @set: occ setter callback
+ * @start_size: starting size of the resource
+ * @max_size: max size of the resource, to present in the uAPI/validate against
+ * @parent_id: ID of the parent resource
+ */
+struct ice_devl_resource {
+ void *owner[ICE_MAX_DEVL_RESOURCE_UNITS];
+ u8 pf_id[ICE_MAX_DEVL_RESOURCE_UNITS];
+ const char *name;
+ devlink_resource_occ_get_t *get;
+ devlink_resource_occ_set_t *set;
+ u32 start_size;
+ u32 max_size;
+ u32 parent_id;
+};
+
/**
* struct ice_port_list - data used to store the list of adapter ports
*
@@ -39,6 +73,7 @@ struct ice_port_list {
* Index 0 = PHY0, index 1 = PHY1. Used on E825C devices.
* @ctrl_pf: Control PF of the adapter
* @ports: Ports list
+ * @resources: array of ice's data for devlink resources
*/
struct ice_adapter {
struct devlink *devlink;
@@ -52,9 +87,14 @@ struct ice_adapter {
struct ice_pf *ctrl_pf;
struct ice_port_list ports;
+
+ struct ice_devl_resource resources[ICE_DEVL_RESOURCES_COUNT];
};
struct ice_adapter *ice_adapter_get(struct pci_dev *pdev);
void ice_adapter_put(struct ice_adapter *adapter);
+DEFINE_GUARD(ice_adapter_devl, struct ice_adapter *,
+ devl_lock((_T)->devlink), devl_unlock((_T)->devlink))
+
#endif /* _ICE_ADAPTER_H */
diff --git a/drivers/net/ethernet/intel/ice/ice_common.h b/drivers/net/ethernet/intel/ice/ice_common.h
index d1d674ca644f..d019ec759eda 100644
--- a/drivers/net/ethernet/intel/ice/ice_common.h
+++ b/drivers/net/ethernet/intel/ice/ice_common.h
@@ -125,6 +125,7 @@ int ice_read_txq_ctx(struct ice_hw *hw, struct ice_tlan_ctx *tlan_ctx,
int ice_write_txq_ctx(struct ice_hw *hw, struct ice_tlan_ctx *tlan_ctx,
u32 txq_index);
+enum ice_lut_size ice_lut_type_to_size(enum ice_lut_type type);
int
ice_aq_get_rss_lut(struct ice_hw *hw, struct ice_aq_get_set_rss_lut_params *get_params);
int
diff --git a/drivers/net/ethernet/intel/ice/ice_lib.h b/drivers/net/ethernet/intel/ice/ice_lib.h
index 49454d98dcfe..15b65a757beb 100644
--- a/drivers/net/ethernet/intel/ice/ice_lib.h
+++ b/drivers/net/ethernet/intel/ice/ice_lib.h
@@ -47,6 +47,9 @@ int ice_vsi_cfg_tc(struct ice_vsi *vsi, u8 ena_tc);
int ice_vsi_cfg_rss_lut_key(struct ice_vsi *vsi);
+int ice_vsi_update_rss_lut(struct ice_vsi *vsi, enum ice_lut_type lut_type,
+ u8 global_lut_id);
+
void ice_vsi_cfg_netdev_tc(struct ice_vsi *vsi, u8 ena_tc);
struct ice_vsi *
diff --git a/drivers/net/ethernet/intel/ice/devlink/resource.c b/drivers/net/ethernet/intel/ice/devlink/resource.c
new file mode 100644
index 000000000000..09672da3ad39
--- /dev/null
+++ b/drivers/net/ethernet/intel/ice/devlink/resource.c
@@ -0,0 +1,527 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) 2026, Intel Corporation. */
+
+#include "resource.h"
+#include "ice_adapter.h"
+#include "ice.h"
+#include "ice_lib.h"
+
+#define ICE_NUM_GLOBAL_LUTS 16
+#define ICE_ANY_SLOT -1
+
+static u32 ice_devl_res_cnt(const struct ice_adapter *adapter,
+ enum ice_devl_resource_id res_id)
+{
+ const struct ice_devl_resource *res = &adapter->resources[res_id];
+ u32 sum = 0;
+
+ for (int i = 0; i < res->max_size; i++)
+ sum += res->owner[i] != NULL;
+
+ return sum;
+}
+
+static int ice_devl_res_take(struct ice_pf *pf,
+ enum ice_devl_resource_id res_id, int slot,
+ void *owner)
+{
+ struct ice_devl_resource *res = &pf->adapter->resources[res_id];
+ int beg, end, new_id = ICE_ANY_SLOT;
+ u16 lut_id;
+
+ if (res_id == ICE_RSS_LUT_GLOBAL) {
+ int err;
+
+ WARN_ON_ONCE(slot != ICE_ANY_SLOT);
+
+ err = ice_alloc_rss_global_lut(&pf->hw, &lut_id);
+ if (err)
+ return err;
+
+ slot = lut_id;
+ }
+
+ end = slot == ICE_ANY_SLOT ? res->max_size : slot + 1;
+ beg = slot == ICE_ANY_SLOT ? 0 : slot;
+ for (int id = beg; id < end; id++) {
+ if (!res->owner[id]) {
+ new_id = id;
+ break;
+ }
+ }
+ if (new_id == ICE_ANY_SLOT) {
+ if (WARN_ON(res_id == ICE_RSS_LUT_GLOBAL))
+ ice_free_rss_global_lut(&pf->hw, lut_id);
+
+ return -ENOSPC;
+ }
+
+ res->owner[new_id] = owner;
+ res->pf_id[new_id] = pf->hw.logical_pf_id;
+ return new_id;
+}
+
+static int ice_devl_res_free(struct ice_pf *pf,
+ enum ice_devl_resource_id res_id, void *owner)
+{
+ struct ice_devl_resource *res = &pf->adapter->resources[res_id];
+ int err = 0, id_to_free = ICE_ANY_SLOT;
+
+ for (int i = 0; i < res->max_size; i++) {
+ if (res->owner[i] == owner) {
+ id_to_free = i;
+ break;
+ }
+ }
+ if (id_to_free == ICE_ANY_SLOT)
+ return 0;
+
+ if (res_id == ICE_RSS_LUT_GLOBAL) {
+ err = ice_free_rss_global_lut(&pf->hw, id_to_free);
+ if (err)
+ dev_warn(ice_pf_to_dev(pf), "failed to free global RSS LUT %d, leaked until PF reset: %d\n",
+ id_to_free, err);
+ }
+
+ res->owner[id_to_free] = NULL;
+ return err;
+}
+
+void ice_free_rss_lut_flr(struct ice_pf *pf)
+{
+ struct ice_devl_resource *res, *resources = pf->adapter->resources;
+ int pf_id = pf->hw.logical_pf_id;
+
+ scoped_guard(ice_adapter_devl, pf->adapter) {
+ resources[ICE_RSS_LUT_PF].owner[pf_id] = NULL;
+
+ res = &resources[ICE_RSS_LUT_GLOBAL];
+ for (int i = 0; i < res->max_size; i++) {
+ /* On FLR/PFR resources assigned to PF are cleared by
+ * FW, reflect that in the SW table.
+ * VFs on given PF must be de-programmed too.
+ */
+ if (res->pf_id[i] == pf_id)
+ res->owner[i] = NULL;
+ }
+ }
+}
+
+static int ice_devl_res_owned_idx(struct ice_adapter *adapter,
+ enum ice_devl_resource_id res_id, void *owner)
+{
+ const struct ice_devl_resource *res = &adapter->resources[res_id];
+
+ for (int i = 0; i < res->max_size; i++) {
+ if (res->owner[i] == owner)
+ return i;
+ }
+
+ return -ENXIO;
+}
+
+static bool ice_is_devl_res_owned_by(struct ice_adapter *adapter,
+ enum ice_devl_resource_id res_id,
+ void *owner)
+{
+ return ice_devl_res_owned_idx(adapter, res_id, owner) >= 0;
+}
+
+static u64 ice_rss_lut_whole_dev_occ_get_global(void *priv)
+{
+ struct ice_adapter *adapter = priv;
+
+ return ice_devl_res_cnt(adapter, ICE_RSS_LUT_GLOBAL);
+}
+
+static u64 ice_rss_lut_whole_dev_occ_get_pf(void *priv)
+{
+ struct ice_adapter *adapter = priv;
+
+ return ice_devl_res_cnt(adapter, ICE_RSS_LUT_PF);
+}
+
+static u64 ice_rss_lut_whole_dev_occ_get_both(void *priv)
+{
+ return ice_rss_lut_whole_dev_occ_get_global(priv) +
+ ice_rss_lut_whole_dev_occ_get_pf(priv);
+}
+
+static u64 ice_rss_lut_pf_occ_get_global(void *priv)
+{
+ struct ice_adapter *adapter;
+ struct ice_pf *pf = priv;
+
+ adapter = pf->adapter;
+ scoped_guard(ice_adapter_devl, adapter)
+ return ice_is_devl_res_owned_by(adapter, ICE_RSS_LUT_GLOBAL, pf);
+}
+
+static u64 ice_rss_lut_pf_occ_get_pf(void *priv)
+{
+ struct ice_adapter *adapter;
+ struct ice_pf *pf = priv;
+
+ adapter = pf->adapter;
+ scoped_guard(ice_adapter_devl, adapter)
+ return ice_is_devl_res_owned_by(adapter, ICE_RSS_LUT_PF, pf);
+}
+
+static u64 ice_rss_lut_pf_occ_get_both(void *priv)
+{
+ struct ice_adapter *adapter;
+ struct ice_pf *pf = priv;
+
+ adapter = pf->adapter;
+ scoped_guard(ice_adapter_devl, adapter)
+ return ice_is_devl_res_owned_by(adapter, ICE_RSS_LUT_PF, pf) +
+ ice_is_devl_res_owned_by(adapter, ICE_RSS_LUT_GLOBAL, pf);
+}
+
+static int ice_devl_resource_deny_occ_set(u64 size,
+ struct netlink_ext_ack *extack,
+ void *priv)
+{
+ NL_SET_ERR_MSG_MOD(extack,
+ "can not change directly, parent/aggregate resource just adds up children data");
+ return -EPERM;
+}
+
+enum ice_rss_lut_resource_state {
+ ICE_HAS_NO_LUT = 0,
+ ICE_HAS_GLOBAL_LUT = BIT(ICE_RSS_LUT_GLOBAL),
+ ICE_HAS_PF_LUT = BIT(ICE_RSS_LUT_PF),
+ ICE_HAS_BOTH_LUTS = ICE_HAS_GLOBAL_LUT | ICE_HAS_PF_LUT,
+};
+
+/** ice_rss_lut_resource_state - compute opaque resource state for given owner
+ * @adapter: the adapter the @owner is on
+ * @owner: the entity to compute state of resources for
+ *
+ * Return: computed current state of the RSS resources the @owner has.
+ */
+static enum ice_rss_lut_resource_state
+ice_rss_lut_resource_state(struct ice_adapter *adapter, void *owner)
+{
+ enum ice_rss_lut_resource_state ret = ICE_HAS_NO_LUT;
+
+ if (ice_is_devl_res_owned_by(adapter, ICE_RSS_LUT_GLOBAL, owner))
+ ret |= ICE_HAS_GLOBAL_LUT;
+ if (ice_is_devl_res_owned_by(adapter, ICE_RSS_LUT_PF, owner))
+ ret |= ICE_HAS_PF_LUT;
+
+ return ret;
+}
+
+static int ice_maybe_change_rss_lut(struct ice_pf *pf, void *owner,
+ enum ice_rss_lut_resource_state old,
+ enum ice_rss_lut_resource_state new,
+ struct netlink_ext_ack *extack)
+{
+ struct ice_aq_get_set_rss_lut_params params = {};
+ struct ice_adapter *adapter = pf->adapter;
+ struct ice_hw *hw = &pf->hw;
+ enum ice_lut_type lut_type;
+ int err, lut_size, lut_id;
+ struct ice_vsi *vsi;
+ u8 *lut;
+
+ if (old & new & ICE_HAS_PF_LUT)
+ return 0;
+
+ if (new & ICE_HAS_PF_LUT) {
+ lut_type = ICE_LUT_PF;
+ } else if (new & ICE_HAS_GLOBAL_LUT) {
+ lut_id = ice_devl_res_owned_idx(adapter, ICE_RSS_LUT_GLOBAL, owner);
+ if (lut_id < 0)
+ return lut_id;
+
+ lut_type = ICE_LUT_GLOBAL;
+ params.global_lut_id = lut_id;
+ } else {
+ lut_type = ICE_LUT_VSI;
+ if (owner == pf) {
+ NL_SET_ERR_MSG_FMT(extack, "cannot change PF to use LUT_VSI (sized 64)");
+ return -EDOM;
+ }
+ }
+
+ if (pf == owner) {
+ vsi = ice_get_main_vsi(pf);
+ } else {
+ return -EOPNOTSUPP;
+ }
+
+ lut_size = ice_lut_type_to_size(lut_type);
+ lut = kmalloc(lut_size, GFP_KERNEL);
+ if (!lut)
+ return -ENOMEM;
+ ice_fill_rss_lut(lut, lut_size, vsi->rss_size);
+ params.lut = lut;
+ params.lut_size = lut_size;
+ params.lut_type = lut_type;
+ params.vsi_handle = vsi->idx;
+ mutex_lock(&pf->rss_lut_lock);
+ err = ice_aq_set_rss_lut(hw, ¶ms);
+ if (err) {
+ NL_SET_ERR_MSG_FMT(extack, "AQ failed: %s", libie_aq_str(hw->adminq.sq_last_status));
+ goto out;
+ }
+
+ err = ice_vsi_update_rss_lut(vsi, lut_type, params.global_lut_id);
+ if (err)
+ goto out;
+
+ devm_kfree(ice_pf_to_dev(pf), vsi->rss_lut_user);
+ vsi->rss_lut_user = NULL;
+ if (lut_type == ICE_LUT_GLOBAL)
+ vsi->global_lut_id = params.global_lut_id;
+ vsi->rss_table_size = lut_size;
+ vsi->rss_lut_type = lut_type;
+out:
+ mutex_unlock(&pf->rss_lut_lock);
+ kfree(lut);
+ return err;
+}
+
+static int ice_devl_res_change(bool take, enum ice_devl_resource_id res_id,
+ struct ice_pf *pf, void *owner, int slot,
+ struct netlink_ext_ack *extack)
+{
+ enum ice_rss_lut_resource_state old, new, change;
+ struct ice_adapter *adapter = pf->adapter;
+ int err;
+
+ change = BIT(res_id);
+ old = ice_rss_lut_resource_state(adapter, owner);
+ new = old;
+ if (take)
+ new |= change;
+ else
+ new &= ~change;
+ if (new == old)
+ return 0;
+
+ /* pairs with set_bit() under adapter devl in ice_setup_tc_mqprio_qdisc() */
+ if (test_bit(ICE_FLAG_TC_MQPRIO, pf->flags)) {
+ NL_SET_ERR_MSG_MOD(extack, "ADQ configured, can't change RSS LUTs");
+ return -EOPNOTSUPP;
+ }
+
+ if (pf == owner && !take && old != ICE_HAS_BOTH_LUTS) {
+ NL_SET_ERR_MSG_MOD(extack,
+ "at least one of 512+ sized LUTs must be assigned to PF device at all times");
+ return -EDOM;
+ }
+
+ if (take) {
+ int slot_id;
+
+ slot_id = ice_devl_res_take(pf, res_id, slot, owner);
+ if (slot_id < 0)
+ return slot_id;
+ }
+
+ err = ice_maybe_change_rss_lut(pf, owner, old, new, extack);
+ if (err) {
+ NL_SET_ERR_MSG_FMT(extack, "failed to change RSS LUT, err: %d", err);
+ if (!take)
+ /* We have not released the resource, so we don't free
+ * it, to avoid freeing LUT still in use or assigning it
+ * to other entity.
+ */
+ return -EBUSY;
+
+ goto undo_res_take;
+ }
+
+ if (!take) {
+undo_res_take:
+ int rel_err;
+
+ rel_err = ice_devl_res_free(pf, res_id, owner);
+ if (rel_err) {
+ NL_SET_ERR_MSG_FMT(extack,
+ "could not free resource, err: %d", rel_err);
+ if (!err)
+ err = rel_err;
+ }
+ }
+
+ return err;
+}
+
+static int ice_rss_lut_pf_occ_set_pf(u64 size, struct netlink_ext_ack *extack,
+ void *priv)
+{
+ struct ice_pf *pf = priv;
+ int pf_id;
+
+ pf_id = pf->hw.logical_pf_id;
+ scoped_guard(ice_adapter_devl, pf->adapter)
+ return ice_devl_res_change(size, ICE_RSS_LUT_PF, pf, pf, pf_id,
+ extack);
+}
+
+static int ice_rss_lut_pf_occ_set_global(u64 size,
+ struct netlink_ext_ack *extack,
+ void *priv)
+{
+ struct ice_pf *pf = priv;
+
+ scoped_guard(ice_adapter_devl, pf->adapter)
+ return ice_devl_res_change(size, ICE_RSS_LUT_GLOBAL, pf, pf,
+ ICE_ANY_SLOT, extack);
+}
+
+/**
+ * ice_rss_lut_is_reassigned - check if RSS LUTs of PF are in non-default state
+ * @pf: the PF to check
+ *
+ * Return: true if @pf lost its PF LUT, or @pf or any of its VFs owns a
+ * global LUT. Caller must hold the adapter devlink lock.
+ */
+bool ice_rss_lut_is_reassigned(struct ice_pf *pf)
+{
+ struct ice_devl_resource *res, *resources = pf->adapter->resources;
+ int pf_id = pf->hw.logical_pf_id;
+
+ devl_assert_locked(pf->adapter->devlink);
+ if (resources[ICE_RSS_LUT_PF].owner[pf_id] != pf)
+ return true;
+
+ res = &resources[ICE_RSS_LUT_GLOBAL];
+ for (int i = 0; i < res->max_size; i++)
+ if (res->owner[i] && res->pf_id[i] == pf_id)
+ return true;
+
+ return false;
+}
+
+/**
+ * ice_take_rss_lut_pf - allocate PF RSS LUT for PF
+ * @pf: the PF device that PF LUT is physically on, and to allocate it for
+ *
+ * Acquire PF RSS LUT for the caller.
+ *
+ * Return: 0 on success, negative on error.
+ */
+int ice_take_rss_lut_pf(struct ice_pf *pf)
+{
+ int ret, pf_id = pf->hw.logical_pf_id;
+
+ scoped_guard(ice_adapter_devl, pf->adapter) {
+ ret = ice_devl_res_take(pf, ICE_RSS_LUT_PF, pf_id, pf);
+ if (ret >= 0)
+ return 0;
+
+ return ret;
+ }
+}
+
+/**
+ * ice_release_rss_lut_pf - release PF RSS LUT of PF
+ * @pf: the PF device that PF LUT is physically on, and to release it from
+ *
+ * Release PF RSS LUT for the caller.
+ *
+ * No need to call this function in normal operation, provided for error path.
+ */
+void ice_release_rss_lut_pf(struct ice_pf *pf)
+{
+ scoped_guard(ice_adapter_devl, pf->adapter)
+ ice_devl_res_free(pf, ICE_RSS_LUT_PF, pf);
+}
+
+static void ice_devl_res_register(struct devlink *devlink,
+ struct ice_devl_resource *resources,
+ void *occ_priv)
+{
+ struct devlink_resource_size_params size_params;
+
+ devlink_resource_size_params_init(&size_params, 0, 0, 1,
+ DEVLINK_RESOURCE_UNIT_ENTRY);
+ for (int i = 0; i < ICE_DEVL_RESOURCES_COUNT; i++) {
+ struct ice_devl_resource *res = &resources[i];
+ int err, resource_id = i;
+
+ if (!res->name)
+ continue; /* skip empty entries in config table */
+
+ size_params.size_max = res->max_size;
+ err = devl_resource_register(devlink, res->name,
+ res->start_size, resource_id,
+ res->parent_id, &size_params);
+ if (WARN_ONCE(err, "not all resource handlers registered, err: %d, resname: %s\n",
+ err, res->name))
+ break;
+
+ devl_resource_occ_set_get_register(devlink, resource_id,
+ res->set, res->get, occ_priv);
+ }
+}
+
+void ice_devl_whole_dev_resources_register(const struct ice_hw *hw,
+ struct ice_adapter *adapter)
+{
+ struct devlink *devlink = adapter->devlink;
+ int pf_lut_cnt = hw->dev_caps.num_funcs;
+
+ devl_assert_locked(devlink);
+
+ adapter->resources[ICE_RSS_LUT_GLOBAL] = (struct ice_devl_resource) {
+ .name = "lut_512",
+ .parent_id = ICE_RSS_LUT_BOTH,
+ .max_size = ICE_NUM_GLOBAL_LUTS,
+ .get = ice_rss_lut_whole_dev_occ_get_global,
+ .set = ice_devl_resource_deny_occ_set,
+ };
+ adapter->resources[ICE_RSS_LUT_PF] = (struct ice_devl_resource) {
+ .name = "lut_2048",
+ .parent_id = ICE_RSS_LUT_BOTH,
+ .max_size = pf_lut_cnt,
+ .get = ice_rss_lut_whole_dev_occ_get_pf,
+ .set = ice_devl_resource_deny_occ_set,
+ };
+ adapter->resources[ICE_RSS_LUT_BOTH] = (struct ice_devl_resource) {
+ .name = "rss",
+ .parent_id = ICE_TOP_RESOURCE,
+ .max_size = pf_lut_cnt + ICE_NUM_GLOBAL_LUTS,
+ .get = ice_rss_lut_whole_dev_occ_get_both,
+ .set = ice_devl_resource_deny_occ_set,
+ };
+
+ ice_devl_res_register(devlink, adapter->resources, adapter);
+}
+
+void ice_devl_pf_resources_register(struct ice_pf *pf)
+{
+ struct ice_devl_resource pf_resources[ICE_DEVL_RESOURCES_COUNT] = {
+ [ICE_RSS_LUT_GLOBAL] = {
+ .name = "lut_512",
+ .parent_id = ICE_RSS_LUT_BOTH,
+ .max_size = 1,
+ .get = ice_rss_lut_pf_occ_get_global,
+ .set = ice_rss_lut_pf_occ_set_global,
+ },
+ [ICE_RSS_LUT_PF] = {
+ .name = "lut_2048",
+ .parent_id = ICE_RSS_LUT_BOTH,
+ .max_size = 1,
+ .get = ice_rss_lut_pf_occ_get_pf,
+ .set = ice_rss_lut_pf_occ_set_pf,
+ .start_size = 1,
+ },
+ [ICE_RSS_LUT_BOTH] = {
+ .name = "rss",
+ .parent_id = ICE_TOP_RESOURCE,
+ .max_size = 2,
+ .get = ice_rss_lut_pf_occ_get_both,
+ .set = ice_devl_resource_deny_occ_set,
+ },
+ };
+ struct devlink *devlink = priv_to_devlink(pf);
+
+ devl_assert_locked(devlink);
+ ice_devl_res_register(devlink, pf_resources, pf);
+}
diff --git a/drivers/net/ethernet/intel/ice/ice_adapter.c b/drivers/net/ethernet/intel/ice/ice_adapter.c
index c6301341ccae..8d85f07e3e69 100644
--- a/drivers/net/ethernet/intel/ice/ice_adapter.c
+++ b/drivers/net/ethernet/intel/ice/ice_adapter.c
@@ -8,6 +8,8 @@
#include "ice_adapter.h"
#include "ice.h"
+#include "devlink/resource.h"
+
#define ICE_ADAPTER_FIXED_INDEX BIT_ULL(63)
#define ICE_ADAPTER_INDEX_E825C \
@@ -37,6 +39,7 @@ static u64 ice_adapter_index(struct pci_dev *pdev)
static int ice_adapter_init(void *priv, void *init_param)
{
+ const struct ice_hw *hw = init_param;
struct ice_adapter *adapter = priv;
struct devlink *devlink;
@@ -51,12 +54,18 @@ static int ice_adapter_init(void *priv, void *init_param)
mutex_init(&adapter->ports.lock);
INIT_LIST_HEAD(&adapter->ports.ports);
+ ice_devl_whole_dev_resources_register(hw, adapter);
+
return 0;
}
static void ice_adapter_fini(void *priv)
{
struct ice_adapter *adapter = priv;
+ struct devlink *devlink;
+
+ devlink = shd_priv_to_devlink(adapter);
+ devl_resources_unregister(devlink);
WARN_ON(!list_empty(&adapter->ports.ports));
for (int i = 0; i < ARRAY_SIZE(adapter->cpi_phy_lock); i++)
@@ -84,15 +93,16 @@ static const struct devlink_ops ice_adapter_devlink_ops = {
*/
struct ice_adapter *ice_adapter_get(struct pci_dev *pdev)
{
+ struct ice_pf *pf = pci_get_drvdata(pdev);
struct ice_adapter *adapter;
struct devlink *devlink;
char devlink_id[32];
u64 index;
index = ice_adapter_index(pdev);
snprintf(devlink_id, sizeof(devlink_id), "ice:%llx", index);
devlink = devlink_shd_get(devlink_id, &ice_adapter_devlink_ops,
- sizeof(*adapter), NULL, pdev->dev.driver);
+ sizeof(*adapter), &pf->hw, pdev->dev.driver);
if (IS_ERR(devlink))
return ERR_CAST(devlink);
diff --git a/drivers/net/ethernet/intel/ice/ice_common.c b/drivers/net/ethernet/intel/ice/ice_common.c
index 04633103e3e6..3794555cd849 100644
--- a/drivers/net/ethernet/intel/ice/ice_common.c
+++ b/drivers/net/ethernet/intel/ice/ice_common.c
@@ -4452,7 +4452,7 @@ ice_aq_sff_eeprom(struct ice_hw *hw, u16 lport, u8 bus_addr,
return status;
}
-static enum ice_lut_size ice_lut_type_to_size(enum ice_lut_type type)
+enum ice_lut_size ice_lut_type_to_size(enum ice_lut_type type)
{
switch (type) {
case ICE_LUT_VSI:
diff --git a/drivers/net/ethernet/intel/ice/ice_ethtool.c b/drivers/net/ethernet/intel/ice/ice_ethtool.c
index dffa213734f1..5b68e6802c74 100644
--- a/drivers/net/ethernet/intel/ice/ice_ethtool.c
+++ b/drivers/net/ethernet/intel/ice/ice_ethtool.c
@@ -3645,21 +3645,27 @@ ice_get_rxfh(struct net_device *netdev, struct ethtool_rxfh_param *rxfh)
if (!rxfh->indir)
return 0;
- lut = kzalloc(vsi->rss_table_size, GFP_KERNEL);
- if (!lut)
- return -ENOMEM;
+ scoped_guard(mutex, &pf->rss_lut_lock) {
+ /* devlink may resize the LUT after the core sized @indir */
+ if (rxfh->indir_size != vsi->rss_table_size)
+ return -EAGAIN;
+
+ lut = kzalloc(rxfh->indir_size, GFP_KERNEL);
+ if (!lut)
+ return -ENOMEM;
- err = ice_get_rss(vsi, rxfh->key, lut, vsi->rss_table_size);
+ err = ice_get_rss(vsi, rxfh->key, lut, rxfh->indir_size);
+ }
if (err)
goto out;
if (ice_is_adq_active(pf)) {
- for (i = 0; i < vsi->rss_table_size; i++)
+ for (i = 0; i < rxfh->indir_size; i++)
rxfh->indir[i] = offset + lut[i] % qcount;
goto out;
}
- for (i = 0; i < vsi->rss_table_size; i++)
+ for (i = 0; i < rxfh->indir_size; i++)
rxfh->indir[i] = lut[i];
out:
@@ -3707,7 +3713,9 @@ ice_set_rxfh(struct net_device *netdev, struct ethtool_rxfh_param *rxfh,
if (rxfh->input_xfrm & RXH_XFRM_SYM_XOR)
hfunc = ICE_AQ_VSI_Q_OPT_RSS_HASH_SYM_TPLZ;
- err = ice_set_rss_hfunc(vsi, hfunc);
+ /* q_opt_rss also carries the RSS LUT type, that devlink may change */
+ scoped_guard(mutex, &pf->rss_lut_lock)
+ err = ice_set_rss_hfunc(vsi, hfunc);
if (err)
return err;
@@ -3727,29 +3735,35 @@ ice_set_rxfh(struct net_device *netdev, struct ethtool_rxfh_param *rxfh,
return err;
}
- if (!vsi->rss_lut_user) {
- vsi->rss_lut_user = devm_kzalloc(dev, vsi->rss_table_size,
- GFP_KERNEL);
- if (!vsi->rss_lut_user)
- return -ENOMEM;
- }
+ scoped_guard(mutex, &pf->rss_lut_lock) {
+ /* devlink may resize the LUT after the core sized @indir */
+ if (rxfh->indir && rxfh->indir_size != vsi->rss_table_size)
+ return -EAGAIN;
- /* Each 32 bits pointed by 'indir' is stored with a lut entry */
- if (rxfh->indir) {
- int i;
+ if (!vsi->rss_lut_user) {
+ vsi->rss_lut_user = devm_kzalloc(dev,
+ vsi->rss_table_size,
+ GFP_KERNEL);
+ if (!vsi->rss_lut_user)
+ return -ENOMEM;
+ }
- for (i = 0; i < vsi->rss_table_size; i++)
- vsi->rss_lut_user[i] = (u8)(rxfh->indir[i]);
- } else {
- ice_fill_rss_lut(vsi->rss_lut_user, vsi->rss_table_size,
- vsi->rss_size);
- }
+ /* Each 32 bits pointed by 'indir' is stored with a lut entry */
+ if (rxfh->indir) {
+ int i;
- err = ice_set_rss_lut(vsi, vsi->rss_lut_user, vsi->rss_table_size);
- if (err)
- return err;
+ for (i = 0; i < vsi->rss_table_size; i++)
+ vsi->rss_lut_user[i] = (u8)(rxfh->indir[i]);
+ } else {
+ ice_fill_rss_lut(vsi->rss_lut_user, vsi->rss_table_size,
+ vsi->rss_size);
+ }
- return 0;
+ err = ice_set_rss_lut(vsi, vsi->rss_lut_user,
+ vsi->rss_table_size);
+ }
+
+ return err;
}
static int
@@ -3856,19 +3870,22 @@ static int ice_vsi_set_dflt_rss_lut(struct ice_vsi *vsi, int req_rss_size)
if (!req_rss_size)
return -EINVAL;
- lut = kzalloc(vsi->rss_table_size, GFP_KERNEL);
- if (!lut)
- return -ENOMEM;
+ scoped_guard(mutex, &pf->rss_lut_lock) {
+ lut = kzalloc(vsi->rss_table_size, GFP_KERNEL);
+ if (!lut)
+ return -ENOMEM;
- /* set RSS LUT parameters */
- if (!test_bit(ICE_FLAG_RSS_ENA, pf->flags))
- vsi->rss_size = 1;
- else
- vsi->rss_size = ice_get_valid_rss_size(hw, req_rss_size);
+ /* set RSS LUT parameters */
+ if (!test_bit(ICE_FLAG_RSS_ENA, pf->flags))
+ vsi->rss_size = 1;
+ else
+ vsi->rss_size = ice_get_valid_rss_size(hw,
+ req_rss_size);
- /* create/set RSS LUT */
- ice_fill_rss_lut(lut, vsi->rss_table_size, vsi->rss_size);
- err = ice_set_rss_lut(vsi, lut, vsi->rss_table_size);
+ /* create/set RSS LUT */
+ ice_fill_rss_lut(lut, vsi->rss_table_size, vsi->rss_size);
+ err = ice_set_rss_lut(vsi, lut, vsi->rss_table_size);
+ }
if (err)
dev_err(dev, "Cannot set RSS lut, err %d aq_err %s\n", err,
libie_aq_str(hw->adminq.sq_last_status));
diff --git a/drivers/net/ethernet/intel/ice/ice_lib.c b/drivers/net/ethernet/intel/ice/ice_lib.c
index f81b310a98c3..ad43d52ea83c 100644
--- a/drivers/net/ethernet/intel/ice/ice_lib.c
+++ b/drivers/net/ethernet/intel/ice/ice_lib.c
@@ -10,6 +10,8 @@
#include "ice_type.h"
#include "ice_vsi_vlan_ops.h"
+#include "devlink/resource.h"
+
/**
* ice_vsi_type_str - maps VSI type enum to string equivalents
* @vsi_type: VSI type enum
@@ -986,10 +988,10 @@ static void ice_rss_clean(struct ice_vsi *vsi)
}
/**
- * ice_vsi_set_rss_params - Setup RSS capabilities per VSI type
+ * ice_vsi_set_dflt_rss_params - Setup default RSS capabilities per VSI type
* @vsi: the VSI being configured
*/
-static void ice_vsi_set_rss_params(struct ice_vsi *vsi)
+static void ice_vsi_set_dflt_rss_params(struct ice_vsi *vsi)
{
struct ice_hw_common_caps *cap;
struct ice_pf *pf = vsi->back;
@@ -1004,6 +1006,9 @@ static void ice_vsi_set_rss_params(struct ice_vsi *vsi)
max_rss_size = BIT(cap->rss_table_entry_width);
switch (vsi->type) {
case ICE_VSI_CHNL:
+ if (!vsi->num_rxq)
+ vsi->num_rxq = 1;
+ fallthrough;
case ICE_VSI_PF:
/* PF VSI will inherit RSS instance of PF */
vsi->rss_table_size = (u16)cap->rss_table_size;
@@ -1306,6 +1311,32 @@ static void ice_set_rss_vsi_ctx(struct ice_vsi_ctx *ctxt, struct ice_vsi *vsi)
FIELD_PREP(ICE_AQ_VSI_Q_OPT_RSS_HASH_M, hash_type);
}
+int ice_vsi_update_rss_lut(struct ice_vsi *vsi, enum ice_lut_type lut_type,
+ u8 global_lut_id)
+{
+ struct ice_vsi_ctx *ctx;
+ int err;
+
+ ctx = kzalloc_obj(*ctx);
+ if (!ctx)
+ return -ENOMEM;
+
+ ctx->info.valid_sections = cpu_to_le16(ICE_AQ_VSI_PROP_Q_OPT_VALID);
+ ctx->info.q_opt_rss = vsi->info.q_opt_rss &
+ ~(ICE_AQ_VSI_Q_OPT_RSS_LUT_M | ICE_AQ_VSI_Q_OPT_RSS_GBL_LUT_M);
+ ctx->info.q_opt_rss |= FIELD_PREP(ICE_AQ_VSI_Q_OPT_RSS_LUT_M,
+ ice_lut_type_to_aq_qopt_rss_val(lut_type)) |
+ FIELD_PREP(ICE_AQ_VSI_Q_OPT_RSS_GBL_LUT_M, global_lut_id);
+ ctx->info.q_opt_tc = vsi->info.q_opt_tc;
+ ctx->info.q_opt_flags = vsi->info.q_opt_flags;
+
+ err = ice_update_vsi(&vsi->back->hw, vsi->idx, ctx, NULL);
+ if (!err)
+ vsi->info.q_opt_rss = ctx->info.q_opt_rss;
+ kfree(ctx);
+ return err;
+}
+
static void
ice_chnl_vsi_setup_q_map(struct ice_vsi *vsi, struct ice_vsi_ctx *ctxt)
{
@@ -1581,19 +1612,22 @@ void ice_vsi_manage_rss_lut(struct ice_vsi *vsi, bool ena)
{
u8 *lut;
- lut = kzalloc(vsi->rss_table_size, GFP_KERNEL);
- if (!lut)
- return;
+ scoped_guard(mutex, &vsi->back->rss_lut_lock) {
+ lut = kzalloc(vsi->rss_table_size, GFP_KERNEL);
+ if (!lut)
+ return;
- if (ena) {
- if (vsi->rss_lut_user)
- memcpy(lut, vsi->rss_lut_user, vsi->rss_table_size);
- else
- ice_fill_rss_lut(lut, vsi->rss_table_size,
- vsi->rss_size);
- }
+ if (ena) {
+ if (vsi->rss_lut_user)
+ memcpy(lut, vsi->rss_lut_user,
+ vsi->rss_table_size);
+ else
+ ice_fill_rss_lut(lut, vsi->rss_table_size,
+ vsi->rss_size);
+ }
- ice_set_rss_lut(vsi, lut, vsi->rss_table_size);
+ ice_set_rss_lut(vsi, lut, vsi->rss_table_size);
+ }
kfree(lut);
}
@@ -1645,16 +1679,19 @@ int ice_vsi_cfg_rss_lut_key(struct ice_vsi *vsi)
}
}
- lut = kzalloc(vsi->rss_table_size, GFP_KERNEL);
- if (!lut)
- return -ENOMEM;
+ scoped_guard(mutex, &pf->rss_lut_lock) {
+ lut = kzalloc(vsi->rss_table_size, GFP_KERNEL);
+ if (!lut)
+ return -ENOMEM;
- if (vsi->rss_lut_user)
- memcpy(lut, vsi->rss_lut_user, vsi->rss_table_size);
- else
- ice_fill_rss_lut(lut, vsi->rss_table_size, vsi->rss_size);
+ if (vsi->rss_lut_user)
+ memcpy(lut, vsi->rss_lut_user, vsi->rss_table_size);
+ else
+ ice_fill_rss_lut(lut, vsi->rss_table_size,
+ vsi->rss_size);
- err = ice_set_rss_lut(vsi, lut, vsi->rss_table_size);
+ err = ice_set_rss_lut(vsi, lut, vsi->rss_table_size);
+ }
if (err) {
dev_err(dev, "set_rss_lut failed, error %d\n", err);
goto ice_vsi_cfg_rss_exit;
@@ -2427,6 +2464,18 @@ static int ice_vsi_cfg_def(struct ice_vsi *vsi)
vsi->vsw = pf->first_sw;
+ if (vsi->flags & ICE_VSI_FLAG_INIT) {
+ scoped_guard(mutex, &pf->rss_lut_lock) {
+ u16 lut_size = vsi->rss_table_size;
+
+ ice_vsi_set_dflt_rss_params(vsi);
+ if (vsi->rss_table_size != lut_size) {
+ devm_kfree(dev, vsi->rss_lut_user);
+ vsi->rss_lut_user = NULL;
+ }
+ }
+ }
+
ret = ice_vsi_alloc_def(vsi, vsi->ch);
if (ret)
return ret;
@@ -2446,15 +2495,23 @@ static int ice_vsi_cfg_def(struct ice_vsi *vsi)
}
/* set RSS capabilities */
- ice_vsi_set_rss_params(vsi);
+ if ((vsi->flags & ICE_VSI_FLAG_INIT) && vsi->type == ICE_VSI_PF) {
+ ret = ice_take_rss_lut_pf(pf);
+ if (ret) {
+ dev_err(dev, "Failed to allocate RSS LUT for PF: %d\n",
+ ret);
+ goto unroll_get_qs;
+ }
+ }
/* set TC configuration */
ice_vsi_set_tc_cfg(vsi);
- /* create the VSI */
- ret = ice_vsi_init(vsi, vsi->flags);
+ /* create the VSI, its context carries RSS LUT type and global LUT id */
+ scoped_guard(mutex, &pf->rss_lut_lock)
+ ret = ice_vsi_init(vsi, vsi->flags);
if (ret)
- goto unroll_get_qs;
+ goto unroll_pf_rss_lut;
ice_vsi_init_vlan_ops(vsi);
@@ -2566,6 +2623,9 @@ static int ice_vsi_cfg_def(struct ice_vsi *vsi)
ice_vsi_free_q_vectors(vsi);
unroll_vsi_init:
ice_vsi_delete_from_hw(vsi);
+unroll_pf_rss_lut:
+ if ((vsi->flags & ICE_VSI_FLAG_INIT) && vsi->type == ICE_VSI_PF)
+ ice_release_rss_lut_pf(pf);
unroll_get_qs:
ice_vsi_put_qs(vsi);
unroll_vsi_alloc_stat:
@@ -2634,15 +2694,16 @@ void ice_vsi_decfg(struct ice_vsi *vsi)
ice_vsi_put_qs(vsi);
ice_vsi_free_arrays(vsi);
- /* SR-IOV determines needed MSIX resources all at once instead of per
- * VSI since when VFs are spawned we know how many VFs there are and how
- * many interrupts each VF needs. SR-IOV MSIX resources are also
- * cleared in the same manner.
- */
-
- if (vsi->type == ICE_VSI_VF &&
- vsi->agg_node && vsi->agg_node->valid)
- vsi->agg_node->num_vsis--;
+ if (vsi->flags & ICE_VSI_FLAG_INIT) {
+ if (vsi->type == ICE_VSI_PF) {
+ ice_free_rss_lut_flr(pf);
+ /* reset restores PF LUT, user LUT has wrong size */
+ if (vsi->rss_lut_type != ICE_LUT_PF) {
+ devm_kfree(ice_pf_to_dev(pf), vsi->rss_lut_user);
+ vsi->rss_lut_user = NULL;
+ }
+ }
+ }
}
/**
diff --git a/drivers/net/ethernet/intel/ice/ice_main.c b/drivers/net/ethernet/intel/ice/ice_main.c
index 6373cd5ac288..fec25a362cb1 100644
--- a/drivers/net/ethernet/intel/ice/ice_main.c
+++ b/drivers/net/ethernet/intel/ice/ice_main.c
@@ -15,6 +15,7 @@
#include "ice_dcb_nl.h"
#include "devlink/devlink.h"
#include "devlink/port.h"
+#include "devlink/resource.h"
#include "ice_sf_eth.h"
#include "ice_hwmon.h"
/* Including ice_trace.h with CREATE_TRACE_POINTS defined will generate the
@@ -3873,6 +3874,7 @@ void ice_deinit_pf(struct ice_pf *pf)
{
/* note that we unroll also on ice_init_pf() failure here */
+ mutex_destroy(&pf->rss_lut_lock);
mutex_destroy(&pf->lag_mutex);
mutex_destroy(&pf->adev_mutex);
mutex_destroy(&pf->sw_mutex);
@@ -3978,6 +3980,7 @@ int ice_init_pf(struct ice_pf *pf)
mutex_init(&pf->tc_mutex);
mutex_init(&pf->adev_mutex);
mutex_init(&pf->lag_mutex);
+ mutex_init(&pf->rss_lut_lock);
INIT_HLIST_HEAD(&pf->aq_wait_list);
spin_lock_init(&pf->aq_wait_lock);
@@ -4953,6 +4956,7 @@ static int ice_init_devlink(struct ice_pf *pf)
return err;
}
ice_health_init(pf);
+ ice_devl_pf_resources_register(pf);
return 0;
}
@@ -4963,6 +4967,7 @@ static void ice_deinit_devlink(struct ice_pf *pf)
ice_devlink_unregister(pf);
ice_devlink_destroy_regions(pf);
ice_devlink_unregister_params(pf);
+ devl_resources_unregister(priv_to_devlink(pf));
}
static int ice_init(struct ice_pf *pf)
@@ -7926,6 +7931,8 @@ int ice_set_rss_lut(struct ice_vsi *vsi, u8 *lut, u16 lut_size)
params.lut_size = lut_size;
params.lut_type = vsi->rss_lut_type;
params.lut = lut;
+ if (params.lut_type == ICE_LUT_GLOBAL)
+ params.global_lut_id = vsi->global_lut_id;
status = ice_aq_set_rss_lut(hw, ¶ms);
if (status)
@@ -7979,11 +7986,14 @@ int ice_get_rss_lut(struct ice_vsi *vsi, u8 *lut, u16 lut_size)
params.lut_size = lut_size;
params.lut_type = vsi->rss_lut_type;
params.lut = lut;
+ if (params.lut_type == ICE_LUT_GLOBAL)
+ params.global_lut_id = vsi->global_lut_id;
status = ice_aq_get_rss_lut(hw, ¶ms);
- if (status)
- dev_err(ice_pf_to_dev(vsi->back), "Cannot get RSS lut, err %d aq_err %s\n",
- status, libie_aq_str(hw->adminq.sq_last_status));
+ if (status) {
+ dev_err(ice_pf_to_dev(vsi->back), "Cannot get RSS lut, err %d aq_err %s, luttype: %d\n",
+ status, libie_aq_str(hw->adminq.sq_last_status), params.lut_type);
+ }
return status;
}
@@ -9193,8 +9203,14 @@ static int ice_setup_tc_mqprio_qdisc(struct net_device *netdev, void *type_data)
ret);
return ret;
}
+ scoped_guard(ice_adapter_devl, pf->adapter) {
+ if (ice_rss_lut_is_reassigned(pf)) {
+ dev_err(dev, "RSS LUTs reassigned via devlink, can't configure ADQ\n");
+ return -EOPNOTSUPP;
+ }
+ set_bit(ICE_FLAG_TC_MQPRIO, pf->flags);
+ }
memcpy(&vsi->mqprio_qopt, mqprio_qopt, sizeof(*mqprio_qopt));
- set_bit(ICE_FLAG_TC_MQPRIO, pf->flags);
/* don't assume state of hw_tc_offload during driver load
* and set the flag for TC flower filter if hw_tc_offload
* already ON
--
2.51.1
^ permalink raw reply related [flat|nested] 15+ messages in thread
* [PATCH net-next v2 14/14] ice: support up to 256 VF queues
2026-10-09 12:04 [PATCH net-next v2 00/14] devlink, mlx5, iavf, ice: XLVF for iavf Przemek Kitszel
` (12 preceding siblings ...)
2026-10-09 12:04 ` [PATCH net-next v2 13/14] ice: represent RSS LUTs as devlink resources Przemek Kitszel
@ 2026-10-09 12:04 ` Przemek Kitszel
13 siblings, 0 replies; 15+ messages in thread
From: Przemek Kitszel @ 2026-10-09 12:04 UTC (permalink / raw)
To: netdev, Jakub Kicinski, Jiri Pirko
Cc: Tony Nguyen, Aleksandr Loktionov, Michal Schmidt, intel-wired-lan,
edumazet, horms, pabeni, davem, Jonathan Corbet, skhan, rdunlap,
andrew+netdev, saeedm, tariqt, leon, mbloch, jacob.e.keller,
jedrzej.jagielski, anzaki, brett.creeley, jtornosm, ohartoov,
Przemek Kitszel
Add support for up to 256 VF queues. To enable more than usual 16, user
needs to assign GLOBAL LUT (devlink resource named "rss/lut_512") for 64
queues or PF LUT (name "rss/lut_2048") for maximum of 256 queues. There is
a need to assign the GLOBAL LUT to PF first, then release the PF LUT
(initially assigned to PF) from PF, to finally assign it to VF, see
examples further below. Usual default of number of CPU cores does still
apply, but queues could be requested later by usual ethtool -L command.
Add devlink instance of VF device to track RSS LUT resources under
a respective PF devlink instance.
Number of MSI-X vectors assigned to VF is orthogonal to this.
Note that VF reset is deferred to service task.
How to use:
1. Up to 64 queues
a. assign one of 16 GLOBAL LUTs to VF:
sudo devlink resource set pci/0000:18:01.0 path rss/lut_512 size 1
b. if more queues than the default (num of vCPU/CPU cores) are wanted:
sudo ethtool -L $vfiface combined $more
2. Up to 256 queues
a. assign a GLOBAL LUT to PF:
sudo devlink resource set pci/0000:18:00.0 path rss/lut_512 size 1
b. free the PF LUT from PF:
sudo devlink resource set pci/0000:18:00.0 path rss/lut_2048 size 0
c. assign the PF LUT to VF:
sudo devlink resource set pci/0000:18:01.0 path rss/lut_2048 size 1
d. if more queues than the default (num of vCPU/CPU cores) are wanted:
sudo ethtool -L $vfiface combined $more
3. display current RSS LUT or RSS table:
a. see if RSS is mapped correctly (e.g. for lut_512 there are 512
entries expected):
ethtool -x $vfiface
b. see what devlink devices are present:
devlink dev show # note the "whole dev aggregate" device over your pci netdevs
c. see PF, VF, and "whole device aggregate" resources:
devlink resource show pci/0000:18:00.0 # PF
devlink resource show pci/0000:18:01.0 # VF
devlink resource show devlink_index/3 # whole dev
output will look like below, for a device with one PF LUT and one GLOBAL LUT:
pci/0000:18:00.0:
name rss size 2 unit entry size_min 0 size_max 2 size_gran 1 dpipe_tables none
resources:
name lut_512 size 1 unit entry size_min 0 size_max 1 size_gran 1 dpipe_tables none
name lut_2048 size 1 unit entry size_min 0 size_max 1 size_gran 1 dpipe_tables none
Big thanks to Mateusz for working together on the whole story for long
time! Big thanks to Alex Loktionov for pointing the reason of one nasty
bug with the message size!
Co-developed-by: Mateusz Polchlopek <mateusz.polchlopek@intel.com>
Signed-off-by: Mateusz Polchlopek <mateusz.polchlopek@intel.com>
Signed-off-by: Aleksandr Loktionov <aleksandr.loktionov@intel.com>
Reviewed-by: Jedrzej Jagielski <jedrzej.jagielski@intel.com>
Signed-off-by: Przemek Kitszel <przemyslaw.kitszel@intel.com>
---
v2:
* drop getting VF ref around vf->needs_deferred_reset, it only opened
a window for resource leak and was not needed anyway (Sashiko)
* replace needs_deferred_reset bitfield with ICE_VF_STATE_NEEDS_RESET
bit, consumed by test_and_clear_bit() before the reset, so a request
coming in during the reset is not lost
* take vfs.table_lock in service task only when a deferred VF reset
is pending, so freshly enabled VFs do not time out on
VIRTCHNL_OP_VERSION
* add !CONFIG_PCI_IOV stub for ice_schedule_vf_reset()
* create VF devlink in ice_start_vfs() under vfs.table_lock instead
of lazily from the virtchnl handler; take PF devlink lock around
devl_nested_devlink_set(); handle VF without devlink on teardown
* take cfg_lock before the adapter devlink lock in VF resource setters,
and reject changes while the VF is not ready (ice_vf_is_ready())
* switch VF VSI context to the new LUT synchronously, so HW no longer
uses the old LUT when it is released
* use logical_pf_id for VF allocation of PF LUT slots
* report the RSS qregion width from LUT capacity, and reject VF queue
requests above it
* keep clamping rss_size to num_rxq for non-VF VSIs, so after
ethtool -L and PF reset the LUT does not point at absent queues
* limit PF rss_size to the queue capacity of its current RSS LUT
(64 for GLOBAL), on devlink LUT switch and on ethtool -L; restore it
from num_rxq when switching back to the PF LUT
* drop ICE_VSI_FLAG_RELOAD, it was set but never tested (and
overwritten on VF reset anyway)
* use ice_get_vf_vsi() instead of open coding it (Sashiko)
* commit message: the whole device aggregate is a shared devlink
instance (devlink_index/N), not a faux device
* check devl_nested_devlink_set() return value (Sashiko)
---
.../net/ethernet/intel/ice/devlink/resource.h | 3 +
drivers/net/ethernet/intel/ice/ice.h | 1 +
drivers/net/ethernet/intel/ice/ice_lib.h | 2 +
drivers/net/ethernet/intel/ice/ice_vf_lib.h | 23 ++++
| 1 +
.../net/ethernet/intel/ice/virt/virtchnl.h | 4 +
.../net/ethernet/intel/ice/devlink/resource.c | 121 +++++++++++++++++-
drivers/net/ethernet/intel/ice/ice_ethtool.c | 14 +-
drivers/net/ethernet/intel/ice/ice_lib.c | 31 ++++-
drivers/net/ethernet/intel/ice/ice_main.c | 25 ++++
drivers/net/ethernet/intel/ice/ice_sriov.c | 11 ++
drivers/net/ethernet/intel/ice/ice_vf_lib.c | 57 +++++++++
drivers/net/ethernet/intel/ice/virt/queues.c | 7 +
| 37 +++++-
.../net/ethernet/intel/ice/virt/virtchnl.c | 42 +++++-
15 files changed, 357 insertions(+), 22 deletions(-)
diff --git a/drivers/net/ethernet/intel/ice/devlink/resource.h b/drivers/net/ethernet/intel/ice/devlink/resource.h
index 8999e0413c51..956e0f7368a6 100644
--- a/drivers/net/ethernet/intel/ice/devlink/resource.h
+++ b/drivers/net/ethernet/intel/ice/devlink/resource.h
@@ -10,14 +10,17 @@ struct devlink;
struct ice_adapter;
struct ice_hw;
struct ice_pf;
+struct ice_vf;
+void ice_devlink_vf_resources_register(struct ice_vf *vf);
void ice_devl_pf_resources_register(struct ice_pf *pf);
void ice_devl_whole_dev_resources_register(const struct ice_hw *hw,
struct ice_adapter *adapter);
bool ice_rss_lut_is_reassigned(struct ice_pf *pf);
int ice_take_rss_lut_pf(struct ice_pf *pf);
void ice_release_rss_lut_pf(struct ice_pf *pf);
void ice_free_rss_lut_flr(struct ice_pf *pf);
+void ice_free_rss_lut_vf(struct ice_vf *vf);
#endif /* _ICE_DEVL_RESOURCE_H_ */
diff --git a/drivers/net/ethernet/intel/ice/ice.h b/drivers/net/ethernet/intel/ice/ice.h
index d28ff576a342..fac1cb124d11 100644
--- a/drivers/net/ethernet/intel/ice/ice.h
+++ b/drivers/net/ethernet/intel/ice/ice.h
@@ -297,6 +297,7 @@ enum ice_pf_state {
ICE_SIDEBANDQ_EVENT_PENDING,
ICE_MDD_EVENT_PENDING,
ICE_VFLR_EVENT_PENDING,
+ ICE_VF_RESET_PENDING,
ICE_FLTR_OVERFLOW_PROMISC,
ICE_VF_DIS,
ICE_CFG_BUSY,
diff --git a/drivers/net/ethernet/intel/ice/ice_lib.h b/drivers/net/ethernet/intel/ice/ice_lib.h
index 15b65a757beb..320edbdece3d 100644
--- a/drivers/net/ethernet/intel/ice/ice_lib.h
+++ b/drivers/net/ethernet/intel/ice/ice_lib.h
@@ -18,6 +18,8 @@ enum ice_l2tsel {
ICE_L2TSEL_EXTRACT_FIRST_TAG_L2TAG1,
};
+u16 ice_lut_type_to_qs_num(enum ice_lut_type lut_type);
+
const char *ice_vsi_type_str(enum ice_vsi_type vsi_type);
bool ice_pf_state_is_nominal(struct ice_pf *pf);
diff --git a/drivers/net/ethernet/intel/ice/ice_vf_lib.h b/drivers/net/ethernet/intel/ice/ice_vf_lib.h
index 396dfb2b2d97..8abc8c844c96 100644
--- a/drivers/net/ethernet/intel/ice/ice_vf_lib.h
+++ b/drivers/net/ethernet/intel/ice/ice_vf_lib.h
@@ -40,6 +40,7 @@ enum ice_vf_states {
ICE_VF_STATE_DIS,
ICE_VF_STATE_MC_PROMISC,
ICE_VF_STATE_UC_PROMISC,
+ ICE_VF_STATE_NEEDS_RESET, /* reset deferred to service task */
ICE_VF_STATES_NBITS
};
@@ -145,6 +146,7 @@ struct ice_vf {
struct kref refcnt;
struct ice_pf *pf;
struct pci_dev *vfdev;
+ struct devlink *devlink;
/* Used during virtchnl message handling and NDO ops against the VF
* that will trigger a VFR
*/
@@ -243,6 +245,12 @@ static inline bool ice_vf_is_lldp_ena(struct ice_vf *vf)
return vf->num_mac_lldp && vf->trusted;
}
+static inline bool ice_vf_is_ready(struct ice_vf *vf)
+{
+ return test_bit(ICE_VF_STATE_INIT, vf->vf_states) &&
+ !test_bit(ICE_VF_STATE_DIS, vf->vf_states);
+}
+
/* VF Hash Table access functions
*
* These functions provide abstraction for interacting with the VF hash table.
@@ -319,9 +327,12 @@ int
ice_vf_clear_vsi_promisc(struct ice_vf *vf, struct ice_vsi *vsi, u8 promisc_m);
int ice_reset_vf(struct ice_vf *vf, u32 flags);
void ice_reset_all_vfs(struct ice_pf *pf);
+void ice_schedule_vf_reset(struct ice_vf *vf);
struct ice_vsi *ice_get_vf_ctrl_vsi(struct ice_pf *pf, struct ice_vsi *vsi);
void ice_vf_update_mac_lldp_num(struct ice_vf *vf, struct ice_vsi *vsi,
bool incr);
+void ice_init_vf_devlink(struct ice_vf *vf);
+void ice_deinit_vf_devlink(struct ice_vf *vf);
#else /* CONFIG_PCI_IOV */
static inline struct ice_vf *ice_get_vf_by_id(struct ice_pf *pf, u16 vf_id)
{
@@ -393,11 +404,23 @@ static inline void ice_reset_all_vfs(struct ice_pf *pf)
{
}
+static inline void ice_schedule_vf_reset(struct ice_vf *vf)
+{
+}
+
static inline struct ice_vsi *
ice_get_vf_ctrl_vsi(struct ice_pf *pf, struct ice_vsi *vsi)
{
return NULL;
}
+
+static inline void ice_init_vf_devlink(struct ice_vf *vf)
+{
+}
+
+static inline void ice_deinit_vf_devlink(struct ice_vf *vf)
+{
+}
#endif /* !CONFIG_PCI_IOV */
#endif /* _ICE_VF_LIB_H_ */
--git a/drivers/net/ethernet/intel/ice/virt/rss.h b/drivers/net/ethernet/intel/ice/virt/rss.h
index 784d4c43ce8b..388f980b4cdf 100644
--- a/drivers/net/ethernet/intel/ice/virt/rss.h
+++ b/drivers/net/ethernet/intel/ice/virt/rss.h
@@ -14,5 +14,6 @@ int ice_vc_config_rss_lut(struct ice_vf *vf, u8 *msg);
int ice_vc_config_rss_hfunc(struct ice_vf *vf, u8 *msg);
int ice_vc_get_rss_hashcfg(struct ice_vf *vf);
int ice_vc_set_rss_hashcfg(struct ice_vf *vf, u8 *msg);
+int ice_vc_get_max_rss_qregion(struct ice_vf *vf);
#endif /* _ICE_VIRT_RSS_H_ */
diff --git a/drivers/net/ethernet/intel/ice/virt/virtchnl.h b/drivers/net/ethernet/intel/ice/virt/virtchnl.h
index d11789b3ae1f..18024f13fb5d 100644
--- a/drivers/net/ethernet/intel/ice/virt/virtchnl.h
+++ b/drivers/net/ethernet/intel/ice/virt/virtchnl.h
@@ -78,6 +78,10 @@ struct ice_virtchnl_ops {
int (*get_ptp_cap)(struct ice_vf *vf,
const struct virtchnl_ptp_caps *msg);
int (*get_phc_time)(struct ice_vf *vf);
+ int (*get_max_rss_qregion)(struct ice_vf *vf);
+ int (*ena_qs_v2_msg)(struct ice_vf *vf, u8 *msg, u16 msglen);
+ int (*dis_qs_v2_msg)(struct ice_vf *vf, u8 *msg, u16 msglen);
+ int (*map_q_vector_msg)(struct ice_vf *vf, u8 *msg, u16 msglen);
};
#ifdef CONFIG_PCI_IOV
diff --git a/drivers/net/ethernet/intel/ice/devlink/resource.c b/drivers/net/ethernet/intel/ice/devlink/resource.c
index 09672da3ad39..a6c99d6c311f 100644
--- a/drivers/net/ethernet/intel/ice/devlink/resource.c
+++ b/drivers/net/ethernet/intel/ice/devlink/resource.c
@@ -107,6 +107,16 @@ void ice_free_rss_lut_flr(struct ice_pf *pf)
}
}
+void ice_free_rss_lut_vf(struct ice_vf *vf)
+{
+ struct ice_pf *pf = vf->pf;
+
+ scoped_guard(ice_adapter_devl, pf->adapter) {
+ ice_devl_res_free(pf, ICE_RSS_LUT_GLOBAL, vf);
+ ice_devl_res_free(pf, ICE_RSS_LUT_PF, vf);
+ }
+}
+
static int ice_devl_res_owned_idx(struct ice_adapter *adapter,
enum ice_devl_resource_id res_id, void *owner)
{
@@ -178,6 +188,37 @@ static u64 ice_rss_lut_pf_occ_get_both(void *priv)
ice_is_devl_res_owned_by(adapter, ICE_RSS_LUT_GLOBAL, pf);
}
+static u64 ice_rss_lut_vf_occ_get_global(void *priv)
+{
+ struct ice_adapter *adapter;
+ struct ice_vf *vf = priv;
+
+ adapter = vf->pf->adapter;
+ scoped_guard(ice_adapter_devl, adapter)
+ return ice_is_devl_res_owned_by(adapter, ICE_RSS_LUT_GLOBAL, vf);
+}
+
+static u64 ice_rss_lut_vf_occ_get_pf(void *priv)
+{
+ struct ice_adapter *adapter;
+ struct ice_vf *vf = priv;
+
+ adapter = vf->pf->adapter;
+ scoped_guard(ice_adapter_devl, adapter)
+ return ice_is_devl_res_owned_by(adapter, ICE_RSS_LUT_PF, vf);
+}
+
+static u64 ice_rss_lut_vf_occ_get_both(void *priv)
+{
+ struct ice_adapter *adapter;
+ struct ice_vf *vf = priv;
+
+ adapter = vf->pf->adapter;
+ scoped_guard(ice_adapter_devl, adapter)
+ return ice_is_devl_res_owned_by(adapter, ICE_RSS_LUT_PF, vf) +
+ ice_is_devl_res_owned_by(adapter, ICE_RSS_LUT_GLOBAL, vf);
+}
+
static int ice_devl_resource_deny_occ_set(u64 size,
struct netlink_ext_ack *extack,
void *priv)
@@ -223,7 +264,9 @@ static int ice_maybe_change_rss_lut(struct ice_pf *pf, void *owner,
struct ice_hw *hw = &pf->hw;
enum ice_lut_type lut_type;
int err, lut_size, lut_id;
+ struct ice_vf *vf = NULL;
struct ice_vsi *vsi;
+ u16 rss_size;
u8 *lut;
if (old & new & ICE_HAS_PF_LUT)
@@ -249,14 +292,16 @@ static int ice_maybe_change_rss_lut(struct ice_pf *pf, void *owner,
if (pf == owner) {
vsi = ice_get_main_vsi(pf);
} else {
- return -EOPNOTSUPP;
+ vf = owner;
+ vsi = ice_get_vf_vsi(vf);
}
lut_size = ice_lut_type_to_size(lut_type);
lut = kmalloc(lut_size, GFP_KERNEL);
if (!lut)
return -ENOMEM;
- ice_fill_rss_lut(lut, lut_size, vsi->rss_size);
+ rss_size = min(vsi->num_rxq, ice_lut_type_to_qs_num(lut_type));
+ ice_fill_rss_lut(lut, lut_size, rss_size);
params.lut = lut;
params.lut_size = lut_size;
params.lut_type = lut_type;
@@ -278,6 +323,12 @@ static int ice_maybe_change_rss_lut(struct ice_pf *pf, void *owner,
vsi->global_lut_id = params.global_lut_id;
vsi->rss_table_size = lut_size;
vsi->rss_lut_type = lut_type;
+ if (vf) {
+ vsi->rss_size = ice_lut_type_to_qs_num(lut_type);
+ ice_schedule_vf_reset(vf);
+ } else {
+ vsi->rss_size = rss_size;
+ }
out:
mutex_unlock(&pf->rss_lut_lock);
kfree(lut);
@@ -374,6 +425,41 @@ static int ice_rss_lut_pf_occ_set_global(u64 size,
ICE_ANY_SLOT, extack);
}
+static int ice_rss_lut_vf_occ_set_pf(u64 size, struct netlink_ext_ack *extack,
+ void *occ_priv)
+{
+ struct ice_vf *vf = occ_priv;
+ struct ice_pf *pf = vf->pf;
+ int pf_id;
+
+ pf_id = pf->hw.logical_pf_id;
+ scoped_guard(mutex, &vf->cfg_lock) {
+ if (!ice_vf_is_ready(vf))
+ return -EBUSY;
+
+ scoped_guard(ice_adapter_devl, pf->adapter)
+ return ice_devl_res_change(size, ICE_RSS_LUT_PF, pf, vf,
+ pf_id, extack);
+ }
+}
+
+static int ice_rss_lut_vf_occ_set_global(u64 size,
+ struct netlink_ext_ack *extack,
+ void *occ_priv)
+{
+ struct ice_vf *vf = occ_priv;
+ struct ice_pf *pf = vf->pf;
+
+ scoped_guard(mutex, &vf->cfg_lock) {
+ if (!ice_vf_is_ready(vf))
+ return -EBUSY;
+
+ scoped_guard(ice_adapter_devl, pf->adapter)
+ return ice_devl_res_change(size, ICE_RSS_LUT_GLOBAL, pf,
+ vf, ICE_ANY_SLOT, extack);
+ }
+}
+
/**
* ice_rss_lut_is_reassigned - check if RSS LUTs of PF are in non-default state
* @pf: the PF to check
@@ -525,3 +611,34 @@ void ice_devl_pf_resources_register(struct ice_pf *pf)
devl_assert_locked(devlink);
ice_devl_res_register(devlink, pf_resources, pf);
}
+
+void ice_devlink_vf_resources_register(struct ice_vf *vf)
+{
+ struct ice_devl_resource vf_resources[ICE_DEVL_RESOURCES_COUNT] = {
+ [ICE_RSS_LUT_GLOBAL] = {
+ .name = "lut_512",
+ .parent_id = ICE_RSS_LUT_BOTH,
+ .max_size = 1,
+ .get = ice_rss_lut_vf_occ_get_global,
+ .set = ice_rss_lut_vf_occ_set_global,
+ },
+ [ICE_RSS_LUT_PF] = {
+ .name = "lut_2048",
+ .parent_id = ICE_RSS_LUT_BOTH,
+ .max_size = 1,
+ .get = ice_rss_lut_vf_occ_get_pf,
+ .set = ice_rss_lut_vf_occ_set_pf,
+ },
+ [ICE_RSS_LUT_BOTH] = {
+ .name = "rss",
+ .parent_id = ICE_TOP_RESOURCE,
+ .max_size = 2,
+ .get = ice_rss_lut_vf_occ_get_both,
+ .set = ice_devl_resource_deny_occ_set,
+ },
+ };
+ struct devlink *devlink = vf->devlink;
+
+ scoped_guard(devl, devlink)
+ ice_devl_res_register(devlink, vf_resources, vf);
+}
diff --git a/drivers/net/ethernet/intel/ice/ice_ethtool.c b/drivers/net/ethernet/intel/ice/ice_ethtool.c
index 5b68e6802c74..1f910058d8c8 100644
--- a/drivers/net/ethernet/intel/ice/ice_ethtool.c
+++ b/drivers/net/ethernet/intel/ice/ice_ethtool.c
@@ -3839,14 +3839,15 @@ ice_get_channels(struct net_device *dev, struct ethtool_channels *ch)
/**
* ice_get_valid_rss_size - return valid number of RSS queues
- * @hw: pointer to the HW structure
+ * @vsi: VSI to get the RSS LUT type of
* @new_size: requested RSS queues
*/
-static int ice_get_valid_rss_size(struct ice_hw *hw, int new_size)
+static int ice_get_valid_rss_size(struct ice_vsi *vsi, int new_size)
{
- struct ice_hw_common_caps *caps = &hw->func_caps.common_cap;
+ struct ice_hw_common_caps *caps = &vsi->back->hw.func_caps.common_cap;
- return min_t(int, new_size, BIT(caps->rss_table_entry_width));
+ new_size = min_t(int, new_size, BIT(caps->rss_table_entry_width));
+ return min_t(int, new_size, ice_lut_type_to_qs_num(vsi->rss_lut_type));
}
/**
@@ -3879,7 +3880,7 @@ static int ice_vsi_set_dflt_rss_lut(struct ice_vsi *vsi, int req_rss_size)
if (!test_bit(ICE_FLAG_RSS_ENA, pf->flags))
vsi->rss_size = 1;
else
- vsi->rss_size = ice_get_valid_rss_size(hw,
+ vsi->rss_size = ice_get_valid_rss_size(vsi,
req_rss_size);
/* create/set RSS LUT */
@@ -3976,7 +3977,8 @@ static int ice_set_channels(struct net_device *dev, struct ethtool_channels *ch)
}
/* Update rss_size due to change in Rx queues */
- vsi->rss_size = ice_get_valid_rss_size(&pf->hw, new_rx);
+ scoped_guard(mutex, &pf->rss_lut_lock)
+ vsi->rss_size = ice_get_valid_rss_size(vsi, new_rx);
adev_unlock:
if (locked) {
diff --git a/drivers/net/ethernet/intel/ice/ice_lib.c b/drivers/net/ethernet/intel/ice/ice_lib.c
index ad43d52ea83c..43c8bfedbea2 100644
--- a/drivers/net/ethernet/intel/ice/ice_lib.c
+++ b/drivers/net/ethernet/intel/ice/ice_lib.c
@@ -12,6 +12,23 @@
#include "devlink/resource.h"
+#define ICE_LUT_VSI_MAX_QS 16
+#define ICE_LUT_GLOBAL_MAX_QS 64
+#define ICE_LUT_PF_MAX_QS 256
+
+u16 ice_lut_type_to_qs_num(enum ice_lut_type lut_type)
+{
+ switch (lut_type) {
+ case ICE_LUT_PF:
+ return ICE_LUT_PF_MAX_QS;
+ case ICE_LUT_GLOBAL:
+ return ICE_LUT_GLOBAL_MAX_QS;
+ case ICE_LUT_VSI:
+ default:
+ return ICE_LUT_VSI_MAX_QS;
+ }
+}
+
/**
* ice_vsi_type_str - maps VSI type enum to string equivalents
* @vsi_type: VSI type enum
@@ -1663,7 +1680,9 @@ int ice_vsi_cfg_rss_lut_key(struct ice_vsi *vsi)
(test_bit(ICE_FLAG_TC_MQPRIO, pf->flags))) {
vsi->rss_size = min_t(u16, vsi->rss_size, vsi->ch_rss_size);
} else {
- vsi->rss_size = min_t(u16, vsi->rss_size, vsi->num_rxq);
+ /* VF rss_size is the queue capacity of its LUT, keep it */
+ if (vsi->type != ICE_VSI_VF)
+ vsi->rss_size = min_t(u16, vsi->rss_size, vsi->num_rxq);
/* If orig_rss_size is valid and it is less than determined
* main VSI's rss_size, update main VSI's rss_size to be
@@ -2697,11 +2716,11 @@ void ice_vsi_decfg(struct ice_vsi *vsi)
if (vsi->flags & ICE_VSI_FLAG_INIT) {
if (vsi->type == ICE_VSI_PF) {
ice_free_rss_lut_flr(pf);
- /* reset restores PF LUT, user LUT has wrong size */
- if (vsi->rss_lut_type != ICE_LUT_PF) {
- devm_kfree(ice_pf_to_dev(pf), vsi->rss_lut_user);
- vsi->rss_lut_user = NULL;
- }
+ } else if (vsi->type == ICE_VSI_VF) {
+ struct ice_vf *vf = vsi->vf;
+
+ vf->num_req_qs = 0;
+ vf->num_vf_qs = min(vf->num_vf_qs, pf->vfs.num_qps_per);
}
}
}
diff --git a/drivers/net/ethernet/intel/ice/ice_main.c b/drivers/net/ethernet/intel/ice/ice_main.c
index fec25a362cb1..f86f4d7d7f55 100644
--- a/drivers/net/ethernet/intel/ice/ice_main.c
+++ b/drivers/net/ethernet/intel/ice/ice_main.c
@@ -2274,6 +2274,30 @@ static void ice_check_media_subtask(struct ice_pf *pf)
}
}
+static void ice_handle_deferred_vf_reset(struct ice_pf *pf)
+{
+ struct ice_vf *vf;
+ unsigned int bkt;
+ int err;
+
+ if (!test_and_clear_bit(ICE_VF_RESET_PENDING, pf->state))
+ return;
+
+ mutex_lock(&pf->vfs.table_lock);
+ ice_for_each_vf(pf, bkt, vf) {
+ if (!test_and_clear_bit(ICE_VF_STATE_NEEDS_RESET, vf->vf_states))
+ continue;
+
+ dev_info(ice_pf_to_dev(pf), "doing deferred reset of VF %d\n",
+ vf->vf_id);
+ err = ice_reset_vf(vf, ICE_VF_RESET_NOTIFY | ICE_VF_RESET_LOCK);
+ if (err)
+ dev_warn(ice_pf_to_dev(pf), "deferred reset of VF %d failed: %d\n",
+ vf->vf_id, err);
+ }
+ mutex_unlock(&pf->vfs.table_lock);
+}
+
static void ice_service_task_recovery_mode(struct work_struct *work)
{
struct ice_pf *pf = container_of(work, struct ice_pf, serv_task);
@@ -2356,6 +2380,7 @@ static void ice_service_task(struct work_struct *work)
return;
}
+ ice_handle_deferred_vf_reset(pf);
ice_process_vflr_event(pf);
ice_clean_mailboxq_subtask(pf);
ice_clean_sbq_subtask(pf);
diff --git a/drivers/net/ethernet/intel/ice/ice_sriov.c b/drivers/net/ethernet/intel/ice/ice_sriov.c
index d7c4a6540008..c86c4b96a016 100644
--- a/drivers/net/ethernet/intel/ice/ice_sriov.c
+++ b/drivers/net/ethernet/intel/ice/ice_sriov.c
@@ -494,6 +494,7 @@ static int ice_start_vfs(struct ice_pf *pf)
}
}
+ ice_init_vf_devlink(vf);
set_bit(ICE_VF_STATE_INIT, vf->vf_states);
ice_ena_vf_mappings(vf);
wr32(hw, VFGEN_RSTAT(vf->vf_id), VIRTCHNL_VFR_VFACTIVE);
@@ -657,6 +658,15 @@ static void ice_sriov_post_vsi_rebuild(struct ice_vf *vf)
wr32(&vf->pf->hw, VFGEN_RSTAT(vf->vf_id), VIRTCHNL_VFR_VFACTIVE);
}
+static struct ice_q_vector *ice_sriov_get_q_vector(struct ice_vsi *vsi,
+ u16 vector_id)
+{
+ /* Subtract non queue vector from vector_id passed by VF
+ * to get actual number of VSI queue vector array index
+ */
+ return vsi->q_vectors[vector_id - ICE_NONQ_VECS_VF];
+}
+
static const struct ice_vf_ops ice_sriov_vf_ops = {
.reset_type = ICE_VF_RESET,
.free = ice_sriov_free_vf,
@@ -667,6 +677,7 @@ static const struct ice_vf_ops ice_sriov_vf_ops = {
.clear_reset_trigger = ice_sriov_clear_reset_trigger,
.irq_close = NULL,
.post_vsi_rebuild = ice_sriov_post_vsi_rebuild,
+ .get_q_vector = ice_sriov_get_q_vector,
};
/**
diff --git a/drivers/net/ethernet/intel/ice/ice_vf_lib.c b/drivers/net/ethernet/intel/ice/ice_vf_lib.c
index 320b393ddaaf..6bc71cdc624c 100644
--- a/drivers/net/ethernet/intel/ice/ice_vf_lib.c
+++ b/drivers/net/ethernet/intel/ice/ice_vf_lib.c
@@ -6,6 +6,7 @@
#include "ice_lib.h"
#include "ice_fltr.h"
#include "virt/allowlist.h"
+#include "devlink/resource.h"
/* Public functions which may be accessed by all driver files */
@@ -1010,6 +1011,18 @@ int ice_reset_vf(struct ice_vf *vf, u32 flags)
return err;
}
+/**
+ * ice_schedule_vf_reset - reset VF, deferred to next service_task context
+ * @vf: VF to reset
+ */
+void ice_schedule_vf_reset(struct ice_vf *vf)
+{
+ set_bit(ICE_VF_STATE_NEEDS_RESET, vf->vf_states);
+ /* pairs with test_and_clear_bit() in ice_handle_deferred_vf_reset() */
+ smp_mb__after_atomic();
+ set_bit(ICE_VF_RESET_PENDING, vf->pf->state);
+}
+
/**
* ice_set_vf_state_dis - Set VF state to disabled
* @vf: pointer to the VF structure
@@ -1064,6 +1077,9 @@ void ice_deinitialize_vf_entry(struct ice_vf *vf)
{
struct ice_pf *pf = vf->pf;
+ ice_free_rss_lut_vf(vf);
+ ice_deinit_vf_devlink(vf);
+
if (!ice_is_feature_supported(pf, ICE_F_MBX_LIMIT))
list_del(&vf->mbx_info.list_entry);
}
@@ -1450,3 +1466,44 @@ void ice_vf_update_mac_lldp_num(struct ice_vf *vf, struct ice_vsi *vsi,
if (was_ena != is_ena)
ice_vsi_cfg_sw_lldp(vsi, false, is_ena);
}
+
+void ice_init_vf_devlink(struct ice_vf *vf)
+{
+ struct devlink *pf_devlink = priv_to_devlink(vf->pf);
+ static const struct devlink_ops noop = {};
+ struct devlink *devlink;
+ int err;
+
+ lockdep_assert_held(&vf->pf->vfs.table_lock);
+
+ devlink = devlink_alloc(&noop, 0, &vf->vfdev->dev);
+ if (!devlink)
+ return;
+
+ scoped_guard(devl, pf_devlink)
+ err = devl_nested_devlink_set(pf_devlink, devlink);
+ if (err) {
+ devlink_free(devlink);
+ return;
+ }
+
+ vf->devlink = devlink;
+ devlink_register(devlink);
+
+ ice_devlink_vf_resources_register(vf);
+}
+
+void ice_deinit_vf_devlink(struct ice_vf *vf)
+{
+ struct devlink *devlink = vf->devlink;
+
+ lockdep_assert_held(&vf->pf->vfs.table_lock);
+
+ if (!devlink)
+ return;
+
+ vf->devlink = NULL;
+ devlink_resources_unregister(devlink);
+ devlink_unregister(devlink);
+ devlink_free(devlink);
+}
diff --git a/drivers/net/ethernet/intel/ice/virt/queues.c b/drivers/net/ethernet/intel/ice/virt/queues.c
index 3c5c5099a73d..57ef0696775b 100644
--- a/drivers/net/ethernet/intel/ice/virt/queues.c
+++ b/drivers/net/ethernet/intel/ice/virt/queues.c
@@ -1004,18 +1004,25 @@ int ice_vc_request_qs_msg(struct ice_vf *vf, u8 *msg)
u16 max_allowed_vf_queues;
u16 tx_rx_queue_left;
struct device *dev;
+ u16 max_lut_queues;
u16 cur_queues;
dev = ice_pf_to_dev(pf);
if (!test_bit(ICE_VF_STATE_ACTIVE, vf->vf_states)) {
v_ret = VIRTCHNL_STATUS_ERR_PARAM;
goto error_param;
}
cur_queues = vf->num_vf_qs;
+ max_lut_queues = ice_lut_type_to_qs_num(ice_get_vf_vsi(vf)->rss_lut_type);
tx_rx_queue_left = min_t(u16, ice_get_avail_txq_count(pf),
ice_get_avail_rxq_count(pf));
max_allowed_vf_queues = tx_rx_queue_left + cur_queues;
+ if (req_queues > max_lut_queues) {
+ vfres->num_queue_pairs = max_lut_queues;
+ goto error_param;
+ }
+
if (!req_queues) {
dev_err(dev, "VF %d tried to request 0 queues. Ignoring.\n",
vf->vf_id);
--git a/drivers/net/ethernet/intel/ice/virt/rss.c b/drivers/net/ethernet/intel/ice/virt/rss.c
index 960012ca91b5..0fb79ec1f9f7 100644
--- a/drivers/net/ethernet/intel/ice/virt/rss.c
+++ b/drivers/net/ethernet/intel/ice/virt/rss.c
@@ -3,6 +3,7 @@
#include "rss.h"
#include "ice_vf_lib_private.h"
+#include "ice_lib.h"
#include "ice.h"
#define FIELD_SELECTOR(proto_hdr_field) \
@@ -1746,7 +1747,13 @@ int ice_vc_config_rss_lut(struct ice_vf *vf, u8 *msg)
goto error_param;
}
- if (vrl->lut_entries != ICE_LUT_VSI_SIZE) {
+ vsi = ice_get_vf_vsi(vf);
+ if (!vsi) {
+ v_ret = VIRTCHNL_STATUS_ERR_PARAM;
+ goto error_param;
+ }
+
+ if (vrl->lut_entries != vsi->rss_table_size) {
v_ret = VIRTCHNL_STATUS_ERR_PARAM;
goto error_param;
}
@@ -1762,7 +1769,7 @@ int ice_vc_config_rss_lut(struct ice_vf *vf, u8 *msg)
goto error_param;
}
- if (ice_set_rss_lut(vsi, vrl->lut, ICE_LUT_VSI_SIZE))
+ if (ice_set_rss_lut(vsi, vrl->lut, vrl->lut_entries))
v_ret = VIRTCHNL_STATUS_ERR_ADMIN_QUEUE_ERROR;
error_param:
return ice_vc_send_msg_to_vf(vf, VIRTCHNL_OP_CONFIG_RSS_LUT, v_ret,
@@ -1920,3 +1927,29 @@ int ice_vc_set_rss_hashcfg(struct ice_vf *vf, u8 *msg)
NULL, 0);
}
+int ice_vc_get_max_rss_qregion(struct ice_vf *vf)
+{
+ enum virtchnl_status_code v_ret = VIRTCHNL_STATUS_SUCCESS;
+ struct virtchnl_max_rss_qregion max_rss_qregion = {};
+ struct ice_vsi *vsi;
+ int err, len = 0;
+
+ if (!test_bit(ICE_VF_STATE_ACTIVE, vf->vf_states)) {
+ v_ret = VIRTCHNL_STATUS_ERR_PARAM;
+ goto reply;
+ }
+
+ vsi = ice_get_vf_vsi(vf);
+ if (!vsi) {
+ v_ret = VIRTCHNL_STATUS_ERR_PARAM;
+ goto reply;
+ }
+
+ len = sizeof(max_rss_qregion);
+ max_rss_qregion.vport_id = vsi->vsi_num;
+ max_rss_qregion.qregion_width = ilog2(ice_lut_type_to_qs_num(vsi->rss_lut_type));
+reply:
+ err = ice_vc_send_msg_to_vf(vf, VIRTCHNL_OP_GET_MAX_RSS_QREGION, v_ret,
+ (u8 *)&max_rss_qregion, len);
+ return err;
+}
diff --git a/drivers/net/ethernet/intel/ice/virt/virtchnl.c b/drivers/net/ethernet/intel/ice/virt/virtchnl.c
index ca8018e3dd42..d3b63032e402 100644
--- a/drivers/net/ethernet/intel/ice/virt/virtchnl.c
+++ b/drivers/net/ethernet/intel/ice/virt/virtchnl.c
@@ -246,10 +246,10 @@ static int ice_vc_get_vf_res_msg(struct ice_vf *vf, u8 *msg)
{
enum virtchnl_status_code v_ret = VIRTCHNL_STATUS_SUCCESS;
struct virtchnl_vf_resource *vfres = NULL;
+ int ret, allowed_queues, len = 0;
struct ice_hw *hw = &vf->pf->hw;
+ enum ice_lut_type lut_type;
struct ice_vsi *vsi;
- int len = 0;
- int ret;
if (ice_check_vf_init(vf)) {
v_ret = VIRTCHNL_STATUS_ERR_PARAM;
@@ -330,16 +330,24 @@ static int ice_vc_get_vf_res_msg(struct ice_vf *vf, u8 *msg)
vfres->vf_cap_flags |= VIRTCHNL_VF_CAP_PTP;
vfres->num_vsis = 1;
- /* Tx and Rx queue are equal for VF */
- vfres->num_queue_pairs = vsi->num_txq;
+
+ lut_type = vsi->rss_lut_type;
+ if (vf->driver_caps & VIRTCHNL_VF_LARGE_NUM_QPAIRS &&
+ lut_type != ICE_LUT_VSI) {
+ vfres->vf_cap_flags |= VIRTCHNL_VF_LARGE_NUM_QPAIRS;
+ allowed_queues = ice_lut_type_to_qs_num(lut_type);
+ } else {
+ allowed_queues = vsi->num_txq;
+ }
+ vfres->num_queue_pairs = allowed_queues;
vfres->max_vectors = vf->num_msix;
vfres->rss_key_size = ICE_VSIQF_HKEY_ARRAY_SIZE;
- vfres->rss_lut_size = ICE_LUT_VSI_SIZE;
+ vfres->rss_lut_size = vsi->rss_table_size;
vfres->max_mtu = ice_vc_get_max_frame_size(vf);
vfres->vsi_res[0].vsi_id = ICE_VF_VSI_ID;
vfres->vsi_res[0].vsi_type = VIRTCHNL_VSI_SRIOV;
- vfres->vsi_res[0].num_queue_pairs = vsi->num_txq;
+ vfres->vsi_res[0].num_queue_pairs = allowed_queues;
ether_addr_copy(vfres->vsi_res[0].default_mac_addr,
vf->hw_lan_addr);
@@ -2533,6 +2541,10 @@ static const struct ice_virtchnl_ops ice_virtchnl_dflt_ops = {
.cfg_q_quanta = ice_vc_cfg_q_quanta,
.get_ptp_cap = ice_vc_get_ptp_cap,
.get_phc_time = ice_vc_get_phc_time,
+ .get_max_rss_qregion = ice_vc_get_max_rss_qregion,
+ .ena_qs_v2_msg = ice_vc_ena_qs_v2_msg,
+ .dis_qs_v2_msg = ice_vc_dis_qs_v2_msg,
+ .map_q_vector_msg = ice_vc_map_q_vector_msg,
/* If you add a new op here please make sure to add it to
* ice_virtchnl_repr_ops as well.
*/
@@ -2670,6 +2682,10 @@ static const struct ice_virtchnl_ops ice_virtchnl_repr_ops = {
.cfg_q_quanta = ice_vc_cfg_q_quanta,
.get_ptp_cap = ice_vc_get_ptp_cap,
.get_phc_time = ice_vc_get_phc_time,
+ .get_max_rss_qregion = ice_vc_get_max_rss_qregion,
+ .ena_qs_v2_msg = ice_vc_ena_qs_v2_msg,
+ .dis_qs_v2_msg = ice_vc_dis_qs_v2_msg,
+ .map_q_vector_msg = ice_vc_map_q_vector_msg,
};
/**
@@ -2901,6 +2917,20 @@ void ice_vc_process_vf_msg(struct ice_pf *pf, struct ice_rq_event_info *event,
case VIRTCHNL_OP_GET_QOS_CAPS:
err = ops->get_qos_caps(vf);
break;
+ case VIRTCHNL_OP_GET_MAX_RSS_QREGION:
+ err = ops->get_max_rss_qregion(vf);
+ break;
+ case VIRTCHNL_OP_ENABLE_QUEUES_V2:
+ err = ops->ena_qs_v2_msg(vf, msg, msglen);
+ if (!err)
+ ice_vc_notify_vf_link_state(vf);
+ break;
+ case VIRTCHNL_OP_DISABLE_QUEUES_V2:
+ err = ops->dis_qs_v2_msg(vf, msg, msglen);
+ break;
+ case VIRTCHNL_OP_MAP_QUEUE_VECTOR:
+ err = ops->map_q_vector_msg(vf, msg, msglen);
+ break;
case VIRTCHNL_OP_CONFIG_QUEUE_BW:
err = ops->cfg_q_bw(vf, msg);
break;
--
2.51.1
^ permalink raw reply related [flat|nested] 15+ messages in thread
end of thread, other threads:[~2026-10-09 12:15 UTC | newest]
Thread overview: 15+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-09 12:04 [PATCH net-next v2 00/14] devlink, mlx5, iavf, ice: XLVF for iavf Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 01/14] devlink, mlx5: add init/fini ops for shared devlink Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 02/14] ice: use shared devlink to store ice_adapters instead of custom xarray Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 03/14] ice: add VF queue ena/dis helper functions Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 04/14] ice: add helpers for Global RSS LUT alloc, free, vsi_update Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 05/14] ice: rename ICE_MAX_RSS_QS_PER_VF to ICE_MAX_QS_PER_VF_VCV1 Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 06/14] ice: bump to 256qs for VF Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 07/14] iavf: extend iavf_configure_queues() to support more queues Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 08/14] iavf: temporary rename of IAVF_MAX_REQ_QUEUES to IAVF_MAX_REQ_QUEUES_VCV1 Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 09/14] iavf: increase max number of queues to 256 Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 10/14] iavf: use new opcodes to request more than 16 queues Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 11/14] ice: introduce handling of virtchnl LARGE VF opcodes Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 12/14] devlink: give user option to allocate resources Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 13/14] ice: represent RSS LUTs as devlink resources Przemek Kitszel
2026-10-09 12:04 ` [PATCH net-next v2 14/14] ice: support up to 256 VF queues Przemek Kitszel
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox