Intel-XE Archive on lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH v2 0/2] drm/xe: Expose retired VRAM pages via drm-ras
@ 2026-10-07  8:03 Aravind Iddamsetty
  2026-10-07  8:03 ` [PATCH v2 1/2] " Aravind Iddamsetty
                   ` (5 more replies)
  0 siblings, 6 replies; 9+ messages in thread
From: Aravind Iddamsetty @ 2026-10-07  8:03 UTC (permalink / raw)
  To: intel-xe, dri-devel, netdev, amd-gfx
  Cc: simona.vetter, airlied, tejas.upadhyay, himal.prasad.ghimiray,
	rodrigo.vivi, riana.tauro, raag.jadav, joshua.santosh.ranjan,
	ashwin.kumar.kulkarni, pratik.bari, Hawking.Zhang, tao.zhou1,
	YiPeng.Chai, jinzhou.su, cesun102, lijo.lazar, alexander.deucher,
	christian.koenig

The memory page offlining support tracks bad VRAM pages and currently
only exposes them through debugfs (vram_bad_pages), which is not a
stable ABI. Add a proper userspace interface using the drm-ras
generic netlink family.

Introduce a new node type DRM_RAS_NODE_TYPE_RETIRED_RESOURCES which
enumerates hardware resources that have been permanently taken out of
service. Each node reports a single resource type, identified by its
node-name, so the resource type is implied by the node rather than
carried per entry. The per-entry payload stays extensible: each entry
carries a single type-specific nested attribute, so future resource
types can be added without touching existing consumers. VRAM pages are
the first supported type, reported via the vram-page nest as
{address, size, status} where status is one of retired/pending/failed.

A single operation is added on the node:
 - GET_RETIRED_RESOURCES: dump the resources retired on the node. The
   dump opens with a leading capacity summary entry (max-pages
   inside the vram-page nest, i.e. the maximum number of pages that
   can be offlined, with no address/size/status), followed by one
   entry per retired resource. Userspace derives the retired/queued
   counts by counting entries.

The netlink interface is described by the YAML spec
Documentation/netlink/specs/drm_ras.yaml. The following files are
generated from that spec with tools/net/ynl/ynl-regen.sh and must not
be edited by hand:
 - include/uapi/drm/drm_ras.h        (uapi attribute/command enums)
 - drivers/gpu/drm/drm_ras_nl.c      (generated policy and op table)
 - drivers/gpu/drm/drm_ras_nl.h      (generated declarations)

Eg:
$ sudo ./tools/net/ynl/pyynl/cli.py \
     --spec Documentation/netlink/specs/drm_ras.yaml \
     --dump list-nodes

[{'device-name': '0000:03:00.0', 'node-id': 0, 'node-name':'correctable-errors',
 'node-type': 'error-counter'},
 {'device-name': '0000:03:00.0', 'node-id': 1, 'node-name':'uncorrectable-errors
', 'node-type': 'error-counter'},
 {'device-name': '0000:03:00.0', 'node-id': 2, 'node-name':'vram-retired-pages',
 'node-type': 'retired-resources'}]

$ sudo ./tools/net/ynl/pyynl/cli.py --spec \
Documentation/netlink/specs/drm_ras.yaml  --dump  get-retired-resources \
--json '{"node-id": 2}'

[{'node-id': 2, 'vram-page': {'max-pages': 100}},
 {'node-id': 2, 'vram-page': {'address': 12807041024, 'size': 4096,
 'status': 'retired'}},
 {'node-id': 2, 'vram-page': {'address': 12809400320, 'size': 4096,
 'status': 'retired'}}]

The first entry is the capacity summary (max-pages); subsequent entries
are the retired pages.

v2:
 - Rebase on drm-tip
 - Drop the per-entry resource-type discriminator attribute; the type is
   now implied by the node (node-name), which reports a single type.
 - Fold all VRAM-page fields, including the retirement status, into the
   vram-page nested attribute; the top-level entry is just
   {node-id, vram-page}. (Rodrigo)
 - Drop the separate GET_RETIRED_RESOURCES_INFO operation. The capacity
   (max-pages) is now emitted as a leading summary entry inside
   GET_RETIRED_RESOURCES (max-pages in the vram-page nest, no status);
   offlined/queued counts are derived by counting entries.

Cc: Tejas Upadhyay <tejas.upadhyay@intel.com>
Cc: Himal Prasad Ghimiray <himal.prasad.ghimiray@intel.com>
Cc: Rodrigo Vivi <rodrigo.vivi@intel.com>
Cc: Riana Tauro <riana.tauro@intel.com>
Cc: Raag Jadav <raag.jadav@intel.com>
Cc: Joshua Santhosh Ranjan <joshua.santosh.ranjan@intel.com>
Cc: Ashwin Kumar Kulkarni <ashwin.kumar.kulkarni@intel.com>
Cc: Pratik Bari <pratik.bari@intel.com>

Aravind Iddamsetty (2):
  drm/xe: Expose retired VRAM pages via drm-ras
  drm/xe: Remove bad VRAM pages debugfs interface

 Documentation/netlink/specs/drm_ras.yaml |  91 ++++++++++-
 drivers/gpu/drm/drm_ras.c                | 174 ++++++++++++++++++++-
 drivers/gpu/drm/drm_ras_nl.c             |  12 ++
 drivers/gpu/drm/drm_ras_nl.h             |   2 +
 drivers/gpu/drm/xe/xe_debugfs.c          |   2 -
 drivers/gpu/drm/xe/xe_drm_ras.c          |  77 +++++++++
 drivers/gpu/drm/xe/xe_drm_ras_types.h    |   3 +
 drivers/gpu/drm/xe/xe_ttm_vram_mgr.c     | 190 ++++++++++++++---------
 drivers/gpu/drm/xe/xe_ttm_vram_mgr.h     |   4 +-
 include/drm/drm_ras.h                    |  74 +++++++++
 include/uapi/drm/drm_ras.h               |  48 +++++-
 11 files changed, 592 insertions(+), 85 deletions(-)

-- 
2.25.1


^ permalink raw reply	[flat|nested] 9+ messages in thread

end of thread, other threads:[~2026-10-07 13:55 UTC | newest]

Thread overview: 9+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-07  8:03 [PATCH v2 0/2] drm/xe: Expose retired VRAM pages via drm-ras Aravind Iddamsetty
2026-10-07  8:03 ` [PATCH v2 1/2] " Aravind Iddamsetty
2026-10-07  9:28   ` Christian König
2026-10-07 13:55     ` Rodrigo Vivi
2026-10-07  8:03 ` [PATCH v2 2/2] drm/xe: Remove bad VRAM pages debugfs interface Aravind Iddamsetty
2026-10-07  8:38 ` ✗ CI.checkpatch: warning for drm/xe: Expose retired VRAM pages via drm-ras (rev2) Patchwork
2026-10-07  8:40 ` ✓ CI.KUnit: success " Patchwork
2026-10-07  9:27 ` ✓ Xe.CI.BAT: " Patchwork
2026-10-07 10:42 ` ✓ Xe.CI.FULL: " Patchwork

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox