* [PATCH 1/2] drm/amd/ras: report how far a CPER record query got
@ 2026-08-26 10:27 Xiang Liu
2026-08-26 10:27 ` [PATCH 2/2] drm/amd/ras: take a log batch id only once there is something to log Xiang Liu
0 siblings, 1 reply; 3+ messages in thread
From: Xiang Liu @ 2026-08-26 10:27 UTC (permalink / raw)
To: amd-gfx; +Cc: Hawking.Zhang, Tao.Zhou1, Stanley.Yang, YiPeng.Chai, Xiang Liu
The caller walks the log with cper_start_id and resumes at
cper_start_id + real_cper_num, but the reply only counts the ids that
held a record. A batch id that holds nothing therefore never moves the
caller forward, and every such id makes it read the next populated batch
one more time. Count the ids covered instead.
The rewind to the last populated batch when the query starts at the end
has the same effect once the caller has drained the log, and the latest
id is already reported by the snapshot query, so drop it.
Signed-off-by: Xiang Liu <xiang.liu@amd.com>
---
drivers/gpu/drm/amd/ras/core/cmd.c | 12 +++++++-----
1 file changed, 7 insertions(+), 5 deletions(-)
diff --git a/drivers/gpu/drm/amd/ras/core/cmd.c b/drivers/gpu/drm/amd/ras/core/cmd.c
index 66bd62a819b2..9f874800f38d 100644
--- a/drivers/gpu/drm/amd/ras/core/cmd.c
+++ b/drivers/gpu/drm/amd/ras/core/cmd.c
@@ -206,7 +206,7 @@ static int ras_cmd_get_cper_records(struct ras_core_context *ras_core,
uint32_t offset = 0, real_data_len = 0;
u64 batch_id, start_batch_id;
uint8_t *buf_ptr = (uint8_t *)(uintptr_t)req->buf_ptr;
- int ret = 0, i, count, valid_batch_count = 0;
+ int ret = 0, i, count, read_batch_count = 0;
if ((cmd->input_size != sizeof(struct ras_cmd_cper_record_req)) ||
(cmd->output_buf_size < sizeof(*rsp)))
@@ -225,8 +225,6 @@ static int ras_cmd_get_cper_records(struct ras_core_context *ras_core,
ras_log_ring_get_batch_overview(ras_core, &overview);
start_batch_id = req->cper_start_id;
- if (overview.logged_batch_count && start_batch_id == overview.last_batch_id)
- start_batch_id = overview.last_batch_id - 1;
for (i = 0; i < req->cper_num; i++) {
batch_id = start_batch_id + i;
@@ -246,9 +244,13 @@ static int ras_cmd_get_cper_records(struct ras_core_context *ras_core,
if (ret)
break;
- valid_batch_count++;
offset += real_data_len;
}
+
+ /* The caller resumes at cper_start_id + real_cper_num, so an id
+ * that held nothing still has to be counted here.
+ */
+ read_batch_count++;
}
if ((ret && (ret != -ENOMEM))) {
@@ -257,7 +259,7 @@ static int ras_cmd_get_cper_records(struct ras_core_context *ras_core,
}
rsp->real_data_size = offset;
- rsp->real_cper_num = valid_batch_count;
+ rsp->real_cper_num = read_batch_count;
rsp->remain_num = (ret == -ENOMEM) ? (req->cper_num - i) : 0;
rsp->version = 0;
--
2.34.1
^ permalink raw reply related [flat|nested] 3+ messages in thread
* [PATCH 2/2] drm/amd/ras: take a log batch id only once there is something to log
2026-08-26 10:27 [PATCH 1/2] drm/amd/ras: report how far a CPER record query got Xiang Liu
@ 2026-08-26 10:27 ` Xiang Liu
2026-08-27 3:14 ` Zhang, Hawking
0 siblings, 1 reply; 3+ messages in thread
From: Xiang Liu @ 2026-08-26 10:27 UTC (permalink / raw)
To: amd-gfx; +Cc: Hawking.Zhang, Tao.Zhou1, Stanley.Yang, YiPeng.Chai, Xiang Liu
The id is claimed before the banks are read, so a poll that ends up
logging nothing still burns one and leaves a gap that every reader of
the log then has to walk over. Claim it when the first record of the
batch is about to go in.
Signed-off-by: Xiang Liu <xiang.liu@amd.com>
---
drivers/gpu/drm/amd/ras/core/aca.c | 8 ++++----
1 file changed, 4 insertions(+), 4 deletions(-)
diff --git a/drivers/gpu/drm/amd/ras/core/aca.c b/drivers/gpu/drm/amd/ras/core/aca.c
index 881681ea9f3f..4289e968d331 100644
--- a/drivers/gpu/drm/amd/ras/core/aca.c
+++ b/drivers/gpu/drm/amd/ras/core/aca.c
@@ -394,10 +394,6 @@ static int aca_banks_update(struct ras_core_context *ras_core,
if (!count)
goto out;
- /* Only one MCE error is logged for each batch */
- if (ecc_type != RAS_ERR_TYPE__MCE)
- batch_tag = ras_log_ring_create_batch_tag(ras_core);
-
for (i = 0; i < count; i++) {
memset(&bank, 0, sizeof(bank));
ret = aca_dump_bank(ras_core, ecc_type, i, &bank);
@@ -420,6 +416,10 @@ static int aca_banks_update(struct ras_core_context *ras_core,
bank.seq_no = aca_get_bank_seqno(ras_core, ecc_type, aca_blk, &bank_ecc);
+ /* Only one MCE error is logged for each batch */
+ if (ecc_type != RAS_ERR_TYPE__MCE && !batch_tag)
+ batch_tag = ras_log_ring_create_batch_tag(ras_core);
+
aca_log_bank_data(ras_core, &bank, &bank_ecc, batch_tag);
aca_bank_log(ras_core, i, count, &bank, &bank_ecc);
--
2.34.1
^ permalink raw reply related [flat|nested] 3+ messages in thread
* RE: [PATCH 2/2] drm/amd/ras: take a log batch id only once there is something to log
2026-08-26 10:27 ` [PATCH 2/2] drm/amd/ras: take a log batch id only once there is something to log Xiang Liu
@ 2026-08-27 3:14 ` Zhang, Hawking
0 siblings, 0 replies; 3+ messages in thread
From: Zhang, Hawking @ 2026-08-27 3:14 UTC (permalink / raw)
To: Liu, Xiang(Dean), amd-gfx@lists.freedesktop.org
Cc: Zhou1, Tao, Yang, Stanley, Chai, Thomas
AMD General
Series is
Reviewed-by: Hawking Zhang <Hawking.Zhang@amd.com>
Regards,
Hawking
-----Original Message-----
From: Liu, Xiang(Dean) <Xiang.Liu@amd.com>
Sent: Wednesday, August 26, 2026 6:28 PM
To: amd-gfx@lists.freedesktop.org
Cc: Zhang, Hawking <Hawking.Zhang@amd.com>; Zhou1, Tao <Tao.Zhou1@amd.com>; Yang, Stanley <Stanley.Yang@amd.com>; Chai, Thomas <YiPeng.Chai@amd.com>; Liu, Xiang(Dean) <Xiang.Liu@amd.com>
Subject: [PATCH 2/2] drm/amd/ras: take a log batch id only once there is something to log
The id is claimed before the banks are read, so a poll that ends up logging nothing still burns one and leaves a gap that every reader of the log then has to walk over. Claim it when the first record of the batch is about to go in.
Signed-off-by: Xiang Liu <xiang.liu@amd.com>
---
drivers/gpu/drm/amd/ras/core/aca.c | 8 ++++----
1 file changed, 4 insertions(+), 4 deletions(-)
diff --git a/drivers/gpu/drm/amd/ras/core/aca.c b/drivers/gpu/drm/amd/ras/core/aca.c
index 881681ea9f3f..4289e968d331 100644
--- a/drivers/gpu/drm/amd/ras/core/aca.c
+++ b/drivers/gpu/drm/amd/ras/core/aca.c
@@ -394,10 +394,6 @@ static int aca_banks_update(struct ras_core_context *ras_core,
if (!count)
goto out;
- /* Only one MCE error is logged for each batch */
- if (ecc_type != RAS_ERR_TYPE__MCE)
- batch_tag = ras_log_ring_create_batch_tag(ras_core);
-
for (i = 0; i < count; i++) {
memset(&bank, 0, sizeof(bank));
ret = aca_dump_bank(ras_core, ecc_type, i, &bank); @@ -420,6 +416,10 @@ static int aca_banks_update(struct ras_core_context *ras_core,
bank.seq_no = aca_get_bank_seqno(ras_core, ecc_type, aca_blk, &bank_ecc);
+ /* Only one MCE error is logged for each batch */
+ if (ecc_type != RAS_ERR_TYPE__MCE && !batch_tag)
+ batch_tag = ras_log_ring_create_batch_tag(ras_core);
+
aca_log_bank_data(ras_core, &bank, &bank_ecc, batch_tag);
aca_bank_log(ras_core, i, count, &bank, &bank_ecc);
--
2.34.1
^ permalink raw reply related [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-08-27 3:15 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-26 10:27 [PATCH 1/2] drm/amd/ras: report how far a CPER record query got Xiang Liu
2026-08-26 10:27 ` [PATCH 2/2] drm/amd/ras: take a log batch id only once there is something to log Xiang Liu
2026-08-27 3:14 ` Zhang, Hawking
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.