* [PATCH 1/3 v5] drm/amd/amdgpu: Increase max rings to enable SDMA page ring @ 2025-02-27 11:47 Jesse.zhang@amd.com 2025-02-27 11:47 ` [PATCH 2/3 v5] drm/amdgpu: Optimize VM invalidation engine allocation and synchronize GPU TLB flush Jesse.zhang@amd.com 2025-02-27 11:47 ` [PATCH 3/3] drm/amdgpu/sdma_v4_4_2: update VM flush implementation for SDMA Jesse.zhang@amd.com 0 siblings, 2 replies; 6+ messages in thread From: Jesse.zhang@amd.com @ 2025-02-27 11:47 UTC (permalink / raw) To: amd-gfx Cc: Alexander.Deucher, Christian Koenig, Lijo Lazar, Jiadong Zhu, Jesse.zhang@amd.com, Jesse Zhang From: "Jesse.zhang@amd.com" <Jesse.zhang@amd.com> Increase the maximum number of rings supported by the AMDGPU driver from 132 to 148. This change is necessary to enable support for the SDMA page ring. Signed-off-by: Jesse Zhang <jesse.zhang@amd.com> --- drivers/gpu/drm/amd/amdgpu/amdgpu_ring.h | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_ring.h b/drivers/gpu/drm/amd/amdgpu/amdgpu_ring.h index 52f7a9a79e7b..4224f8fa1614 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_ring.h +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_ring.h @@ -37,7 +37,7 @@ struct amdgpu_job; struct amdgpu_vm; /* max number of rings */ -#define AMDGPU_MAX_RINGS 132 +#define AMDGPU_MAX_RINGS 148 #define AMDGPU_MAX_HWIP_RINGS 64 #define AMDGPU_MAX_GFX_RINGS 2 #define AMDGPU_MAX_SW_GFX_RINGS 2 -- 2.25.1 ^ permalink raw reply related [flat|nested] 6+ messages in thread
* [PATCH 2/3 v5] drm/amdgpu: Optimize VM invalidation engine allocation and synchronize GPU TLB flush 2025-02-27 11:47 [PATCH 1/3 v5] drm/amd/amdgpu: Increase max rings to enable SDMA page ring Jesse.zhang@amd.com @ 2025-02-27 11:47 ` Jesse.zhang@amd.com 2025-02-27 16:00 ` Christian König 2025-02-27 11:47 ` [PATCH 3/3] drm/amdgpu/sdma_v4_4_2: update VM flush implementation for SDMA Jesse.zhang@amd.com 1 sibling, 1 reply; 6+ messages in thread From: Jesse.zhang@amd.com @ 2025-02-27 11:47 UTC (permalink / raw) To: amd-gfx Cc: Alexander.Deucher, Christian Koenig, Lijo Lazar, Jiadong Zhu, Jesse.zhang@amd.com, Jesse Zhang From: "Jesse.zhang@amd.com" <Jesse.zhang@amd.com> - Modify the VM invalidation engine allocation logic to handle SDMA page rings. SDMA page rings now share the VM invalidation engine with SDMA gfx rings instead of allocating a separate engine. This change ensures efficient resource management and avoids the issue of insufficient VM invalidation engines. - Add synchronization for GPU TLB flush operations in gmc_v9_0.c. Use spin_lock and spin_unlock to ensure thread safety and prevent race conditions during TLB flush operations. This improves the stability and reliability of the driver, especially in multi-threaded environments. v2: replace the sdma ring check with a function `amdgpu_sdma_is_page_queue` to check if a ring is an SDMA page queue.(Lijo) v3: Add GC version check, only enabled on GC9.4.3/9.4.4/9.5.0 Suggested-by: Lijo Lazar <lijo.lazar@amd.com> Signed-off-by: Jesse Zhang <jesse.zhang@amd.com> --- drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c | 7 +++++++ drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.c | 23 +++++++++++++++++++++++ drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h | 1 + 3 files changed, 31 insertions(+) diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c index c6e5c50a3322..68088d731c23 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c @@ -602,8 +602,15 @@ int amdgpu_gmc_allocate_vm_inv_eng(struct amdgpu_device *adev) return -EINVAL; } + if(amdgpu_sdma_is_shared_inv_eng(adev, ring)) { + /* Do not allocate a separate VM invalidation engine for SDMA page rings. + * Shared VM invalid engine with sdma gfx ring. + */ + ring->vm_inv_eng = inv_eng - 1; + } else { ring->vm_inv_eng = inv_eng - 1; vm_inv_engs[vmhub] &= ~(1 << ring->vm_inv_eng); + } dev_info(adev->dev, "ring %s uses VM inv eng %u on hub %u\n", ring->name, ring->vm_inv_eng, ring->vm_hub); diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.c index 39669f8788a7..019f670edc29 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.c +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.c @@ -504,6 +504,29 @@ void amdgpu_sdma_sysfs_reset_mask_fini(struct amdgpu_device *adev) } } +/** +* amdgpu_sdma_is_shared_inv_eng - Check if a ring is an SDMA ring that shares a VM invalidation engine +* @adev: Pointer to the AMDGPU device structure +* @ring: Pointer to the ring structure to check +* +* This function checks if the given ring is an SDMA ring that shares a VM invalidation engine. +* It returns true if the ring is such an SDMA ring, false otherwise. +*/ +bool amdgpu_sdma_is_shared_inv_eng(struct amdgpu_device *adev, struct amdgpu_ring* ring) +{ + int i = ring->me; + + if (!adev->sdma.has_page_queue || i >= adev->sdma.num_instances) + return false; + + if (amdgpu_ip_version(adev, GC_HWIP, 0) == IP_VERSION(9, 4, 3) || + amdgpu_ip_version(adev, GC_HWIP, 0) == IP_VERSION(9, 4, 4) || + amdgpu_ip_version(adev, GC_HWIP, 0) == IP_VERSION(9, 5, 0)) + return (ring == &adev->sdma.instance[i].ring); + else + return false; +} + /** * amdgpu_sdma_register_on_reset_callbacks - Register SDMA reset callbacks * @funcs: Pointer to the callback structure containing pre_reset and post_reset functions diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h b/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h index 965169320065..dcc8fd7a6784 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h @@ -194,4 +194,5 @@ int amdgpu_sdma_ras_sw_init(struct amdgpu_device *adev); void amdgpu_debugfs_sdma_sched_mask_init(struct amdgpu_device *adev); int amdgpu_sdma_sysfs_reset_mask_init(struct amdgpu_device *adev); void amdgpu_sdma_sysfs_reset_mask_fini(struct amdgpu_device *adev); +bool amdgpu_sdma_is_shared_inv_eng(struct amdgpu_device *adev, struct amdgpu_ring* ring); #endif -- 2.25.1 ^ permalink raw reply related [flat|nested] 6+ messages in thread
* Re: [PATCH 2/3 v5] drm/amdgpu: Optimize VM invalidation engine allocation and synchronize GPU TLB flush 2025-02-27 11:47 ` [PATCH 2/3 v5] drm/amdgpu: Optimize VM invalidation engine allocation and synchronize GPU TLB flush Jesse.zhang@amd.com @ 2025-02-27 16:00 ` Christian König 0 siblings, 0 replies; 6+ messages in thread From: Christian König @ 2025-02-27 16:00 UTC (permalink / raw) To: Jesse.zhang@amd.com, amd-gfx; +Cc: Alexander.Deucher, Lijo Lazar, Jiadong Zhu Am 27.02.25 um 12:47 schrieb Jesse.zhang@amd.com: > From: "Jesse.zhang@amd.com" <Jesse.zhang@amd.com> > > - Modify the VM invalidation engine allocation logic to handle SDMA page rings. > SDMA page rings now share the VM invalidation engine with SDMA gfx rings instead of > allocating a separate engine. This change ensures efficient resource management and > avoids the issue of insufficient VM invalidation engines. > > - Add synchronization for GPU TLB flush operations in gmc_v9_0.c. > Use spin_lock and spin_unlock to ensure thread safety and prevent race conditions > during TLB flush operations. This improves the stability and reliability of the driver, > especially in multi-threaded environments. > > v2: replace the sdma ring check with a function `amdgpu_sdma_is_page_queue` > to check if a ring is an SDMA page queue.(Lijo) > > v3: Add GC version check, only enabled on GC9.4.3/9.4.4/9.5.0 This needs to be the last patch in the series and not the second. Otherwise you have a broken state in between. > > Suggested-by: Lijo Lazar <lijo.lazar@amd.com> > Signed-off-by: Jesse Zhang <jesse.zhang@amd.com> > --- > drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c | 7 +++++++ > drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.c | 23 +++++++++++++++++++++++ > drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h | 1 + > 3 files changed, 31 insertions(+) > > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c > index c6e5c50a3322..68088d731c23 100644 > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_gmc.c > @@ -602,8 +602,15 @@ int amdgpu_gmc_allocate_vm_inv_eng(struct amdgpu_device *adev) > return -EINVAL; > } > > + if(amdgpu_sdma_is_shared_inv_eng(adev, ring)) { > + /* Do not allocate a separate VM invalidation engine for SDMA page rings. > + * Shared VM invalid engine with sdma gfx ring. > + */ First of all that comment has style issues, please use checkpatch.pl. Then you need to describe why that is done and what are the pre-requisites to make it work. E.g. something like "SDMA has a special packet which allows it to use the same invalidation engine for all the rings in one instance." Christian. > + ring->vm_inv_eng = inv_eng - 1; > + } else { > ring->vm_inv_eng = inv_eng - 1; > vm_inv_engs[vmhub] &= ~(1 << ring->vm_inv_eng); > + } > > dev_info(adev->dev, "ring %s uses VM inv eng %u on hub %u\n", > ring->name, ring->vm_inv_eng, ring->vm_hub); > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.c > index 39669f8788a7..019f670edc29 100644 > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.c > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.c > @@ -504,6 +504,29 @@ void amdgpu_sdma_sysfs_reset_mask_fini(struct amdgpu_device *adev) > } > } > > +/** > +* amdgpu_sdma_is_shared_inv_eng - Check if a ring is an SDMA ring that shares a VM invalidation engine > +* @adev: Pointer to the AMDGPU device structure > +* @ring: Pointer to the ring structure to check > +* > +* This function checks if the given ring is an SDMA ring that shares a VM invalidation engine. > +* It returns true if the ring is such an SDMA ring, false otherwise. > +*/ > +bool amdgpu_sdma_is_shared_inv_eng(struct amdgpu_device *adev, struct amdgpu_ring* ring) > +{ > + int i = ring->me; > + > + if (!adev->sdma.has_page_queue || i >= adev->sdma.num_instances) > + return false; > + > + if (amdgpu_ip_version(adev, GC_HWIP, 0) == IP_VERSION(9, 4, 3) || > + amdgpu_ip_version(adev, GC_HWIP, 0) == IP_VERSION(9, 4, 4) || > + amdgpu_ip_version(adev, GC_HWIP, 0) == IP_VERSION(9, 5, 0)) > + return (ring == &adev->sdma.instance[i].ring); > + else > + return false; > +} > + > /** > * amdgpu_sdma_register_on_reset_callbacks - Register SDMA reset callbacks > * @funcs: Pointer to the callback structure containing pre_reset and post_reset functions > diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h b/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h > index 965169320065..dcc8fd7a6784 100644 > --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h > +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_sdma.h > @@ -194,4 +194,5 @@ int amdgpu_sdma_ras_sw_init(struct amdgpu_device *adev); > void amdgpu_debugfs_sdma_sched_mask_init(struct amdgpu_device *adev); > int amdgpu_sdma_sysfs_reset_mask_init(struct amdgpu_device *adev); > void amdgpu_sdma_sysfs_reset_mask_fini(struct amdgpu_device *adev); > +bool amdgpu_sdma_is_shared_inv_eng(struct amdgpu_device *adev, struct amdgpu_ring* ring); > #endif ^ permalink raw reply [flat|nested] 6+ messages in thread
* [PATCH 3/3] drm/amdgpu/sdma_v4_4_2: update VM flush implementation for SDMA 2025-02-27 11:47 [PATCH 1/3 v5] drm/amd/amdgpu: Increase max rings to enable SDMA page ring Jesse.zhang@amd.com 2025-02-27 11:47 ` [PATCH 2/3 v5] drm/amdgpu: Optimize VM invalidation engine allocation and synchronize GPU TLB flush Jesse.zhang@amd.com @ 2025-02-27 11:47 ` Jesse.zhang@amd.com 2025-02-27 15:56 ` Christian König 1 sibling, 1 reply; 6+ messages in thread From: Jesse.zhang@amd.com @ 2025-02-27 11:47 UTC (permalink / raw) To: amd-gfx Cc: Alexander.Deucher, Christian Koenig, Lijo Lazar, Jiadong Zhu, Jesse.zhang@amd.com This commit updates the VM flush implementation for the SDMA engine. - Added a new function `sdma_v4_4_2_get_invalidate_req` to construct the VM_INVALIDATE_ENG0_REQ register value for the specified VMID and flush type. This function ensures that all relevant page table cache levels (L1 PTEs, L2 PTEs, and L2 PDEs) are invalidated. - Modified the `sdma_v4_4_2_ring_emit_vm_flush` function to use the new `sdma_v4_4_2_get_invalidate_req` function. The updated function emits the necessary register writes and waits to perform a VM flush for the specified VMID. It updates the PTB address registers and issues a VM invalidation request using the specified VM invalidation engine. - Included the necessary header file `gc/gc_9_0_sh_mask.h` to provide access to the required register definitions. v2: vm flush by the vm inalidation packet (Lijo) Suggested-by: Lijo Lazar <lijo.lazar@amd.com> Signed-off-by: Jesse Zhang <jesse.zhang@amd.com> --- drivers/gpu/drm/amd/amdgpu/sdma_v4_4_2.c | 81 ++++++++++++++++++++---- 1 file changed, 67 insertions(+), 14 deletions(-) diff --git a/drivers/gpu/drm/amd/amdgpu/sdma_v4_4_2.c b/drivers/gpu/drm/amd/amdgpu/sdma_v4_4_2.c index ba43c8f46f45..a9e46a4ed7a8 100644 --- a/drivers/gpu/drm/amd/amdgpu/sdma_v4_4_2.c +++ b/drivers/gpu/drm/amd/amdgpu/sdma_v4_4_2.c @@ -31,6 +31,7 @@ #include "amdgpu_ucode.h" #include "amdgpu_trace.h" #include "amdgpu_reset.h" +#include "gc/gc_9_0_sh_mask.h" #include "sdma/sdma_4_4_2_offset.h" #include "sdma/sdma_4_4_2_sh_mask.h" @@ -1292,21 +1293,75 @@ static void sdma_v4_4_2_ring_emit_pipeline_sync(struct amdgpu_ring *ring) seq, 0xffffffff, 4); } - -/** - * sdma_v4_4_2_ring_emit_vm_flush - vm flush using sDMA +/* + * sdma_v4_4_2_get_invalidate_req - Construct the VM_INVALIDATE_ENG0_REQ register value + * @vmid: The VMID to invalidate + * @flush_type: The type of flush (0 = legacy, 1 = lightweight, 2 = heavyweight) * - * @ring: amdgpu_ring pointer - * @vmid: vmid number to use - * @pd_addr: address + * This function constructs the VM_INVALIDATE_ENG0_REQ register value for the specified VMID + * and flush type. It ensures that all relevant page table cache levels (L1 PTEs, L2 PTEs, and + * L2 PDEs) are invalidated. + */ +static uint32_t sdma_v4_4_2_get_invalidate_req(unsigned int vmid, + uint32_t flush_type) +{ + u32 req = 0; + + req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, + PER_VMID_INVALIDATE_REQ, 1 << vmid); + req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, FLUSH_TYPE, flush_type); + req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PTES, 1); + req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE0, 1); + req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE1, 1); + req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE2, 1); + req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L1_PTES, 1); + req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, + CLEAR_PROTECTION_FAULT_STATUS_ADDR, 0); + + return req; +} +/* The vm validate packet is only available for GC9.4.3/GC9.4.4/GC9.5.0 */ +#define SDMA_OP_VM_INVALIDATE 0x8 +#define SDMA_SUBOP_VM_INVALIDATE 0x4 + +/* + * sdma_v4_4_2_ring_emit_vm_flush - Emit VM flush commands for SDMA + * @ring: The SDMA ring + * @vmid: The VMID to flush + * @pd_addr: The page directory address * - * Update the page table base and flush the VM TLB - * using sDMA. + * This function emits the necessary register writes and waits to perform a VM flush for the + * specified VMID. It updates the PTB address registers and issues a VM invalidation request + * using the specified VM invalidation engine. */ static void sdma_v4_4_2_ring_emit_vm_flush(struct amdgpu_ring *ring, - unsigned vmid, uint64_t pd_addr) + unsigned int vmid, uint64_t pd_addr) { - amdgpu_gmc_emit_flush_gpu_tlb(ring, vmid, pd_addr); + struct amdgpu_device *adev = ring->adev; + uint32_t req = sdma_v4_4_2_get_invalidate_req(vmid, 0); + unsigned int eng = ring->vm_inv_eng; + struct amdgpu_vmhub *hub = &adev->vmhub[ring->vm_hub]; + + amdgpu_ring_emit_wreg(ring, hub->ctx0_ptb_addr_lo32 + + (hub->ctx_addr_distance * vmid), + lower_32_bits(pd_addr)); + + amdgpu_ring_emit_wreg(ring, hub->ctx0_ptb_addr_hi32 + + (hub->ctx_addr_distance * vmid), + upper_32_bits(pd_addr)); + /* + * Construct and emit the VM invalidation packet: + * DW0: OP, Sub OP, Engine IDs (XCC0, XCC1, MMHUB) + * DW1: Invalidation request + * DW2: Lower 32 bits of page directory address + * DW3: Upper 32 bits of page directory address and INVALIDATEACK + */ + amdgpu_ring_write(ring, SDMA_PKT_HEADER_OP(SDMA_OP_VM_INVALIDATE) | + SDMA_PKT_HEADER_SUB_OP(SDMA_SUBOP_VM_INVALIDATE) | + (0x1f << 16) | (0x1f << 21) | (eng << 26)); + amdgpu_ring_write(ring, req); + amdgpu_ring_write(ring, 0x0); + amdgpu_ring_write(ring, (0x1 << vmid)); } static void sdma_v4_4_2_ring_emit_wreg(struct amdgpu_ring *ring, @@ -2112,8 +2167,7 @@ static const struct amdgpu_ring_funcs sdma_v4_4_2_ring_funcs = { 3 + /* hdp invalidate */ 6 + /* sdma_v4_4_2_ring_emit_pipeline_sync */ /* sdma_v4_4_2_ring_emit_vm_flush */ - SOC15_FLUSH_GPU_TLB_NUM_WREG * 3 + - SOC15_FLUSH_GPU_TLB_NUM_REG_WAIT * 6 + + 4 + 2 * 3 + 10 + 10 + 10, /* sdma_v4_4_2_ring_emit_fence x3 for user fence, vm fence */ .emit_ib_size = 7 + 6, /* sdma_v4_4_2_ring_emit_ib */ .emit_ib = sdma_v4_4_2_ring_emit_ib, @@ -2145,8 +2199,7 @@ static const struct amdgpu_ring_funcs sdma_v4_4_2_page_ring_funcs = { 3 + /* hdp invalidate */ 6 + /* sdma_v4_4_2_ring_emit_pipeline_sync */ /* sdma_v4_4_2_ring_emit_vm_flush */ - SOC15_FLUSH_GPU_TLB_NUM_WREG * 3 + - SOC15_FLUSH_GPU_TLB_NUM_REG_WAIT * 6 + + 4 + 2 * 3 + 10 + 10 + 10, /* sdma_v4_4_2_ring_emit_fence x3 for user fence, vm fence */ .emit_ib_size = 7 + 6, /* sdma_v4_4_2_ring_emit_ib */ .emit_ib = sdma_v4_4_2_ring_emit_ib, -- 2.25.1 ^ permalink raw reply related [flat|nested] 6+ messages in thread
* Re: [PATCH 3/3] drm/amdgpu/sdma_v4_4_2: update VM flush implementation for SDMA 2025-02-27 11:47 ` [PATCH 3/3] drm/amdgpu/sdma_v4_4_2: update VM flush implementation for SDMA Jesse.zhang@amd.com @ 2025-02-27 15:56 ` Christian König 2025-02-28 9:36 ` Lazar, Lijo 0 siblings, 1 reply; 6+ messages in thread From: Christian König @ 2025-02-27 15:56 UTC (permalink / raw) To: Jesse.zhang@amd.com, amd-gfx; +Cc: Alexander.Deucher, Lijo Lazar, Jiadong Zhu Am 27.02.25 um 12:47 schrieb Jesse.zhang@amd.com: > This commit updates the VM flush implementation for the SDMA engine. > > - Added a new function `sdma_v4_4_2_get_invalidate_req` to construct the VM_INVALIDATE_ENG0_REQ > register value for the specified VMID and flush type. This function ensures that all relevant > page table cache levels (L1 PTEs, L2 PTEs, and L2 PDEs) are invalidated. > > - Modified the `sdma_v4_4_2_ring_emit_vm_flush` function to use the new `sdma_v4_4_2_get_invalidate_req` > function. The updated function emits the necessary register writes and waits to perform a VM flush > for the specified VMID. It updates the PTB address registers and issues a VM invalidation request > using the specified VM invalidation engine. > > - Included the necessary header file `gc/gc_9_0_sh_mask.h` to provide access to the required register > definitions. > > v2: vm flush by the vm inalidation packet (Lijo) > > Suggested-by: Lijo Lazar <lijo.lazar@amd.com> > Signed-off-by: Jesse Zhang <jesse.zhang@amd.com> > --- > drivers/gpu/drm/amd/amdgpu/sdma_v4_4_2.c | 81 ++++++++++++++++++++---- > 1 file changed, 67 insertions(+), 14 deletions(-) > > diff --git a/drivers/gpu/drm/amd/amdgpu/sdma_v4_4_2.c b/drivers/gpu/drm/amd/amdgpu/sdma_v4_4_2.c > index ba43c8f46f45..a9e46a4ed7a8 100644 > --- a/drivers/gpu/drm/amd/amdgpu/sdma_v4_4_2.c > +++ b/drivers/gpu/drm/amd/amdgpu/sdma_v4_4_2.c > @@ -31,6 +31,7 @@ > #include "amdgpu_ucode.h" > #include "amdgpu_trace.h" > #include "amdgpu_reset.h" > +#include "gc/gc_9_0_sh_mask.h" > > #include "sdma/sdma_4_4_2_offset.h" > #include "sdma/sdma_4_4_2_sh_mask.h" > @@ -1292,21 +1293,75 @@ static void sdma_v4_4_2_ring_emit_pipeline_sync(struct amdgpu_ring *ring) > seq, 0xffffffff, 4); > } > > - > -/** > - * sdma_v4_4_2_ring_emit_vm_flush - vm flush using sDMA > +/* > + * sdma_v4_4_2_get_invalidate_req - Construct the VM_INVALIDATE_ENG0_REQ register value > + * @vmid: The VMID to invalidate > + * @flush_type: The type of flush (0 = legacy, 1 = lightweight, 2 = heavyweight) > * > - * @ring: amdgpu_ring pointer > - * @vmid: vmid number to use > - * @pd_addr: address > + * This function constructs the VM_INVALIDATE_ENG0_REQ register value for the specified VMID > + * and flush type. It ensures that all relevant page table cache levels (L1 PTEs, L2 PTEs, and > + * L2 PDEs) are invalidated. > + */ > +static uint32_t sdma_v4_4_2_get_invalidate_req(unsigned int vmid, > + uint32_t flush_type) > +{ > + u32 req = 0; > + > + req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, > + PER_VMID_INVALIDATE_REQ, 1 << vmid); > + req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, FLUSH_TYPE, flush_type); > + req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PTES, 1); > + req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE0, 1); > + req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE1, 1); > + req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE2, 1); > + req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L1_PTES, 1); > + req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, > + CLEAR_PROTECTION_FAULT_STATUS_ADDR, 0); > + > + return req; > +} > +/* The vm validate packet is only available for GC9.4.3/GC9.4.4/GC9.5.0 */ > +#define SDMA_OP_VM_INVALIDATE 0x8 > +#define SDMA_SUBOP_VM_INVALIDATE 0x4 That needs to be in a header. > + > +/* > + * sdma_v4_4_2_ring_emit_vm_flush - Emit VM flush commands for SDMA > + * @ring: The SDMA ring > + * @vmid: The VMID to flush > + * @pd_addr: The page directory address > * > - * Update the page table base and flush the VM TLB > - * using sDMA. > + * This function emits the necessary register writes and waits to perform a VM flush for the > + * specified VMID. It updates the PTB address registers and issues a VM invalidation request > + * using the specified VM invalidation engine. > */ > static void sdma_v4_4_2_ring_emit_vm_flush(struct amdgpu_ring *ring, > - unsigned vmid, uint64_t pd_addr) > + unsigned int vmid, uint64_t pd_addr) > { > - amdgpu_gmc_emit_flush_gpu_tlb(ring, vmid, pd_addr); > + struct amdgpu_device *adev = ring->adev; > + uint32_t req = sdma_v4_4_2_get_invalidate_req(vmid, 0); > + unsigned int eng = ring->vm_inv_eng; > + struct amdgpu_vmhub *hub = &adev->vmhub[ring->vm_hub]; > + > + amdgpu_ring_emit_wreg(ring, hub->ctx0_ptb_addr_lo32 + > + (hub->ctx_addr_distance * vmid), > + lower_32_bits(pd_addr)); > + > + amdgpu_ring_emit_wreg(ring, hub->ctx0_ptb_addr_hi32 + > + (hub->ctx_addr_distance * vmid), > + upper_32_bits(pd_addr)); That is unecessary. > + /* > + * Construct and emit the VM invalidation packet: > + * DW0: OP, Sub OP, Engine IDs (XCC0, XCC1, MMHUB) > + * DW1: Invalidation request > + * DW2: Lower 32 bits of page directory address > + * DW3: Upper 32 bits of page directory address and INVALIDATEACK How are upper bits and invalidateack mixed together here? > + */ > + amdgpu_ring_write(ring, SDMA_PKT_HEADER_OP(SDMA_OP_VM_INVALIDATE) | > + SDMA_PKT_HEADER_SUB_OP(SDMA_SUBOP_VM_INVALIDATE) | > + (0x1f << 16) | (0x1f << 21) | (eng << 26)); What does those magic numbers mean? > + amdgpu_ring_write(ring, req); > + amdgpu_ring_write(ring, 0x0); > + amdgpu_ring_write(ring, (0x1 << vmid)); Either drop the () and the 0x or even better use the bit macro. And it looks like you completely missed the upper and lower bits of the page directory. Christian. > } > > static void sdma_v4_4_2_ring_emit_wreg(struct amdgpu_ring *ring, > @@ -2112,8 +2167,7 @@ static const struct amdgpu_ring_funcs sdma_v4_4_2_ring_funcs = { > 3 + /* hdp invalidate */ > 6 + /* sdma_v4_4_2_ring_emit_pipeline_sync */ > /* sdma_v4_4_2_ring_emit_vm_flush */ > - SOC15_FLUSH_GPU_TLB_NUM_WREG * 3 + > - SOC15_FLUSH_GPU_TLB_NUM_REG_WAIT * 6 + > + 4 + 2 * 3 + > 10 + 10 + 10, /* sdma_v4_4_2_ring_emit_fence x3 for user fence, vm fence */ > .emit_ib_size = 7 + 6, /* sdma_v4_4_2_ring_emit_ib */ > .emit_ib = sdma_v4_4_2_ring_emit_ib, > @@ -2145,8 +2199,7 @@ static const struct amdgpu_ring_funcs sdma_v4_4_2_page_ring_funcs = { > 3 + /* hdp invalidate */ > 6 + /* sdma_v4_4_2_ring_emit_pipeline_sync */ > /* sdma_v4_4_2_ring_emit_vm_flush */ > - SOC15_FLUSH_GPU_TLB_NUM_WREG * 3 + > - SOC15_FLUSH_GPU_TLB_NUM_REG_WAIT * 6 + > + 4 + 2 * 3 + > 10 + 10 + 10, /* sdma_v4_4_2_ring_emit_fence x3 for user fence, vm fence */ > .emit_ib_size = 7 + 6, /* sdma_v4_4_2_ring_emit_ib */ > + .emit_ib = sdma_v4_4_2_ring_emit_ib, ^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH 3/3] drm/amdgpu/sdma_v4_4_2: update VM flush implementation for SDMA 2025-02-27 15:56 ` Christian König @ 2025-02-28 9:36 ` Lazar, Lijo 0 siblings, 0 replies; 6+ messages in thread From: Lazar, Lijo @ 2025-02-28 9:36 UTC (permalink / raw) To: Christian König, Jesse.zhang@amd.com, amd-gfx Cc: Alexander.Deucher, Jiadong Zhu On 2/27/2025 9:26 PM, Christian König wrote: > Am 27.02.25 um 12:47 schrieb Jesse.zhang@amd.com: >> This commit updates the VM flush implementation for the SDMA engine. >> >> - Added a new function `sdma_v4_4_2_get_invalidate_req` to construct the VM_INVALIDATE_ENG0_REQ >> register value for the specified VMID and flush type. This function ensures that all relevant >> page table cache levels (L1 PTEs, L2 PTEs, and L2 PDEs) are invalidated. >> >> - Modified the `sdma_v4_4_2_ring_emit_vm_flush` function to use the new `sdma_v4_4_2_get_invalidate_req` >> function. The updated function emits the necessary register writes and waits to perform a VM flush >> for the specified VMID. It updates the PTB address registers and issues a VM invalidation request >> using the specified VM invalidation engine. >> >> - Included the necessary header file `gc/gc_9_0_sh_mask.h` to provide access to the required register >> definitions. >> >> v2: vm flush by the vm inalidation packet (Lijo) >> >> Suggested-by: Lijo Lazar <lijo.lazar@amd.com> >> Signed-off-by: Jesse Zhang <jesse.zhang@amd.com> >> --- >> drivers/gpu/drm/amd/amdgpu/sdma_v4_4_2.c | 81 ++++++++++++++++++++---- >> 1 file changed, 67 insertions(+), 14 deletions(-) >> >> diff --git a/drivers/gpu/drm/amd/amdgpu/sdma_v4_4_2.c b/drivers/gpu/drm/amd/amdgpu/sdma_v4_4_2.c >> index ba43c8f46f45..a9e46a4ed7a8 100644 >> --- a/drivers/gpu/drm/amd/amdgpu/sdma_v4_4_2.c >> +++ b/drivers/gpu/drm/amd/amdgpu/sdma_v4_4_2.c >> @@ -31,6 +31,7 @@ >> #include "amdgpu_ucode.h" >> #include "amdgpu_trace.h" >> #include "amdgpu_reset.h" >> +#include "gc/gc_9_0_sh_mask.h" >> >> #include "sdma/sdma_4_4_2_offset.h" >> #include "sdma/sdma_4_4_2_sh_mask.h" >> @@ -1292,21 +1293,75 @@ static void sdma_v4_4_2_ring_emit_pipeline_sync(struct amdgpu_ring *ring) >> seq, 0xffffffff, 4); >> } >> >> - >> -/** >> - * sdma_v4_4_2_ring_emit_vm_flush - vm flush using sDMA >> +/* >> + * sdma_v4_4_2_get_invalidate_req - Construct the VM_INVALIDATE_ENG0_REQ register value >> + * @vmid: The VMID to invalidate >> + * @flush_type: The type of flush (0 = legacy, 1 = lightweight, 2 = heavyweight) >> * >> - * @ring: amdgpu_ring pointer >> - * @vmid: vmid number to use >> - * @pd_addr: address >> + * This function constructs the VM_INVALIDATE_ENG0_REQ register value for the specified VMID >> + * and flush type. It ensures that all relevant page table cache levels (L1 PTEs, L2 PTEs, and >> + * L2 PDEs) are invalidated. >> + */ >> +static uint32_t sdma_v4_4_2_get_invalidate_req(unsigned int vmid, >> + uint32_t flush_type) >> +{ >> + u32 req = 0; >> + >> + req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, >> + PER_VMID_INVALIDATE_REQ, 1 << vmid); >> + req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, FLUSH_TYPE, flush_type); >> + req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PTES, 1); >> + req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE0, 1); >> + req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE1, 1); >> + req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L2_PDE2, 1); >> + req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, INVALIDATE_L1_PTES, 1); >> + req = REG_SET_FIELD(req, VM_INVALIDATE_ENG0_REQ, >> + CLEAR_PROTECTION_FAULT_STATUS_ADDR, 0); >> + >> + return req; >> +} >> +/* The vm validate packet is only available for GC9.4.3/GC9.4.4/GC9.5.0 */ >> +#define SDMA_OP_VM_INVALIDATE 0x8 >> +#define SDMA_SUBOP_VM_INVALIDATE 0x4 > > That needs to be in a header. > >> + >> +/* >> + * sdma_v4_4_2_ring_emit_vm_flush - Emit VM flush commands for SDMA >> + * @ring: The SDMA ring >> + * @vmid: The VMID to flush >> + * @pd_addr: The page directory address >> * >> - * Update the page table base and flush the VM TLB >> - * using sDMA. >> + * This function emits the necessary register writes and waits to perform a VM flush for the >> + * specified VMID. It updates the PTB address registers and issues a VM invalidation request >> + * using the specified VM invalidation engine. >> */ >> static void sdma_v4_4_2_ring_emit_vm_flush(struct amdgpu_ring *ring, >> - unsigned vmid, uint64_t pd_addr) >> + unsigned int vmid, uint64_t pd_addr) >> { >> - amdgpu_gmc_emit_flush_gpu_tlb(ring, vmid, pd_addr); >> + struct amdgpu_device *adev = ring->adev; >> + uint32_t req = sdma_v4_4_2_get_invalidate_req(vmid, 0); >> + unsigned int eng = ring->vm_inv_eng; >> + struct amdgpu_vmhub *hub = &adev->vmhub[ring->vm_hub]; >> + > >> + amdgpu_ring_emit_wreg(ring, hub->ctx0_ptb_addr_lo32 + >> + (hub->ctx_addr_distance * vmid), >> + lower_32_bits(pd_addr)); >> + >> + amdgpu_ring_emit_wreg(ring, hub->ctx0_ptb_addr_hi32 + >> + (hub->ctx_addr_distance * vmid), >> + upper_32_bits(pd_addr)); > > That is unecessary. > Probably you got confused by the description below. The description below is wrong. >> + /* >> + * Construct and emit the VM invalidation packet: >> + * DW0: OP, Sub OP, Engine IDs (XCC0, XCC1, MMHUB) >> + * DW1: Invalidation request >> + * DW2: Lower 32 bits of page directory address >> + * DW3: Upper 32 bits of page directory address and INVALIDATEACK > This description is not correct. DW2 and DW3 are for the logical address ranges as filled in inv_eng_addr_lo/hi. They are not meant for PD address. > How are upper bits and invalidateack mixed together here? > >> + */ >> + amdgpu_ring_write(ring, SDMA_PKT_HEADER_OP(SDMA_OP_VM_INVALIDATE) | >> + SDMA_PKT_HEADER_SUB_OP(SDMA_SUBOP_VM_INVALIDATE) | >> + (0x1f << 16) | (0x1f << 21) | (eng << 26)); > > What does those magic numbers mean? Yes, this means the packet definition needs to be there somewhere in the header too. Thanks, Lijo > >> + amdgpu_ring_write(ring, req); >> + amdgpu_ring_write(ring, 0x0); >> + amdgpu_ring_write(ring, (0x1 << vmid)); > > Either drop the () and the 0x or even better use the bit macro. > > And it looks like you completely missed the upper and lower bits of the page directory. > > Christian. > >> } >> >> static void sdma_v4_4_2_ring_emit_wreg(struct amdgpu_ring *ring, >> @@ -2112,8 +2167,7 @@ static const struct amdgpu_ring_funcs sdma_v4_4_2_ring_funcs = { >> 3 + /* hdp invalidate */ >> 6 + /* sdma_v4_4_2_ring_emit_pipeline_sync */ >> /* sdma_v4_4_2_ring_emit_vm_flush */ >> - SOC15_FLUSH_GPU_TLB_NUM_WREG * 3 + >> - SOC15_FLUSH_GPU_TLB_NUM_REG_WAIT * 6 + >> + 4 + 2 * 3 + >> 10 + 10 + 10, /* sdma_v4_4_2_ring_emit_fence x3 for user fence, vm fence */ >> .emit_ib_size = 7 + 6, /* sdma_v4_4_2_ring_emit_ib */ >> .emit_ib = sdma_v4_4_2_ring_emit_ib, >> @@ -2145,8 +2199,7 @@ static const struct amdgpu_ring_funcs sdma_v4_4_2_page_ring_funcs = { >> 3 + /* hdp invalidate */ >> 6 + /* sdma_v4_4_2_ring_emit_pipeline_sync */ >> /* sdma_v4_4_2_ring_emit_vm_flush */ >> - SOC15_FLUSH_GPU_TLB_NUM_WREG * 3 + >> - SOC15_FLUSH_GPU_TLB_NUM_REG_WAIT * 6 + >> + 4 + 2 * 3 + >> 10 + 10 + 10, /* sdma_v4_4_2_ring_emit_fence x3 for user fence, vm fence */ >> .emit_ib_size = 7 + 6, /* sdma_v4_4_2_ring_emit_ib */ >> + .emit_ib = sdma_v4_4_2_ring_emit_ib, > ^ permalink raw reply [flat|nested] 6+ messages in thread
end of thread, other threads:[~2025-02-28 9:37 UTC | newest] Thread overview: 6+ messages (download: mbox.gz follow: Atom feed -- links below jump to the message on this page -- 2025-02-27 11:47 [PATCH 1/3 v5] drm/amd/amdgpu: Increase max rings to enable SDMA page ring Jesse.zhang@amd.com 2025-02-27 11:47 ` [PATCH 2/3 v5] drm/amdgpu: Optimize VM invalidation engine allocation and synchronize GPU TLB flush Jesse.zhang@amd.com 2025-02-27 16:00 ` Christian König 2025-02-27 11:47 ` [PATCH 3/3] drm/amdgpu/sdma_v4_4_2: update VM flush implementation for SDMA Jesse.zhang@amd.com 2025-02-27 15:56 ` Christian König 2025-02-28 9:36 ` Lazar, Lijo
This is an external index of several public inboxes, see mirroring instructions on how to clone and mirror all data and code used by this external index.