* [PATCH] ppc/fadump: collect dump if the collected size is lesser than reserved
@ 2026-07-14 17:30 Shivang Upadhyay
2026-07-19 7:09 ` Sourabh Jain
0 siblings, 1 reply; 4+ messages in thread
From: Shivang Upadhyay @ 2026-07-14 17:30 UTC (permalink / raw)
To: linuxppc-dev, linux-kernel
Cc: maddy, mpe, npiggin, chleroy, sourabhjain, adri.vero.dev, adityag,
anushree.mathur, Shivang Upadhyay
When a machine is subjected to CPUs add/remove, using dlpar
operations, the number of collected CPU_NOTES can change.
As per PAPR, collected dump size should not be more than
allocated size. Reflecting the same in source.
Signed-off-by: Shivang Upadhyay <shivangu@linux.ibm.com>
---
arch/powerpc/platforms/pseries/rtas-fadump.c | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
diff --git a/arch/powerpc/platforms/pseries/rtas-fadump.c b/arch/powerpc/platforms/pseries/rtas-fadump.c
index 3bb4ac2ab6cc..19a5adaf326b 100644
--- a/arch/powerpc/platforms/pseries/rtas-fadump.c
+++ b/arch/powerpc/platforms/pseries/rtas-fadump.c
@@ -469,7 +469,8 @@ static int __init rtas_fadump_process(struct fw_dump *fadump_conf)
pr_err("Dump taken by platform is not valid (%d)\n", i);
rc = -EINVAL;
}
- if (fdm_active->rgn[i].bytes_dumped != fdm_active->rgn[i].source_len) {
+ if (be64_to_cpu(fdm_active->rgn[i].bytes_dumped)
+ > be64_to_cpu(fdm_active->rgn[i].source_len)) {
pr_err("Dump taken by platform is incomplete (%d)\n", i);
rc = -EINVAL;
}
--
2.54.0
^ permalink raw reply related [flat|nested] 4+ messages in thread
* Re: [PATCH] ppc/fadump: collect dump if the collected size is lesser than reserved
2026-07-14 17:30 [PATCH] ppc/fadump: collect dump if the collected size is lesser than reserved Shivang Upadhyay
@ 2026-07-19 7:09 ` Sourabh Jain
2026-07-22 9:11 ` Shivang Upadhyay
0 siblings, 1 reply; 4+ messages in thread
From: Sourabh Jain @ 2026-07-19 7:09 UTC (permalink / raw)
To: Shivang Upadhyay, linuxppc-dev, linux-kernel
Cc: maddy, mpe, npiggin, chleroy, adri.vero.dev, adityag,
anushree.mathur
Could you please reword the commit title to make it a bit clearer?
On 14/07/26 23:00, Shivang Upadhyay wrote:
> When a machine is subjected to CPUs add/remove, using dlpar
> operations, the number of collected CPU_NOTES can change.
> As per PAPR, collected dump size should not be more than
> allocated size. Reflecting the same in source.
Can you add more details about the problem you are trying to solve
with this patch and how.
Adding the error message and scenario would be really helpful in
understanding the problem.
Can you add Closes tag if it is reported upstream and if possible fixes
tag too.
> Signed-off-by: Shivang Upadhyay <shivangu@linux.ibm.com>
> ---
> arch/powerpc/platforms/pseries/rtas-fadump.c | 3 ++-
> 1 file changed, 2 insertions(+), 1 deletion(-)
>
> diff --git a/arch/powerpc/platforms/pseries/rtas-fadump.c b/arch/powerpc/platforms/pseries/rtas-fadump.c
> index 3bb4ac2ab6cc..19a5adaf326b 100644
> --- a/arch/powerpc/platforms/pseries/rtas-fadump.c
> +++ b/arch/powerpc/platforms/pseries/rtas-fadump.c
> @@ -469,7 +469,8 @@ static int __init rtas_fadump_process(struct fw_dump *fadump_conf)
> pr_err("Dump taken by platform is not valid (%d)\n", i);
> rc = -EINVAL;
> }
> - if (fdm_active->rgn[i].bytes_dumped != fdm_active->rgn[i].source_len) {
> + if (be64_to_cpu(fdm_active->rgn[i].bytes_dumped)
> + > be64_to_cpu(fdm_active->rgn[i].source_len)) {
Can you please share your observations about `bytes_dump` for both QEMU
and a
real system (LPAR) where the number of online CPUs is not equal to the
maximum
number of CPUs?
I agree that you're making these changes to comply with PAPR, but since
this has been
working for so long, it would be good to have all the data before making
this change.
- Sourabh Jain
> pr_err("Dump taken by platform is incomplete (%d)\n", i);
> rc = -EINVAL;
> }
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH] ppc/fadump: collect dump if the collected size is lesser than reserved
2026-07-19 7:09 ` Sourabh Jain
@ 2026-07-22 9:11 ` Shivang Upadhyay
2026-07-29 3:44 ` Sourabh Jain
0 siblings, 1 reply; 4+ messages in thread
From: Shivang Upadhyay @ 2026-07-22 9:11 UTC (permalink / raw)
To: Sourabh Jain, linuxppc-dev, linux-kernel
Cc: maddy, mpe, npiggin, chleroy, adri.vero.dev, adityag,
anushree.mathur
On Sun, 2026-07-19 at 12:39 +0530, Sourabh Jain wrote:
> Could you please reword the commit title to make it a bit clearer?
>
>
> On 14/07/26 23:00, Shivang Upadhyay wrote:
> > When a machine is subjected to CPUs add/remove, using dlpar
> > operations, the number of collected CPU_NOTES can change.
> > As per PAPR, collected dump size should not be more than
> > allocated size. Reflecting the same in source.
>
> Can you add more details about the problem you are trying to solve
> with this patch and how.
>
> Adding the error message and scenario would be really helpful in
> understanding the problem.
>
Hi Sourabh,
When a qemu ppc machine is booted with fadump=on and `-
smp=x,maxcpus=y`, on the panic kernel, /proc/vmcore is not generated
because dump_bytes and source_len does'nt match for CPU_STATE_DATA
region in fadump, then we just give up on parsing rest of the data.
> Can you add Closes tag if it is reported upstream and if possible
> fixes
> tag too.
>
> > Signed-off-by: Shivang Upadhyay <shivangu@linux.ibm.com>
> > ---
> > arch/powerpc/platforms/pseries/rtas-fadump.c | 3 ++-
> > 1 file changed, 2 insertions(+), 1 deletion(-)
> >
> > diff --git a/arch/powerpc/platforms/pseries/rtas-fadump.c
> > b/arch/powerpc/platforms/pseries/rtas-fadump.c
> > index 3bb4ac2ab6cc..19a5adaf326b 100644
> > --- a/arch/powerpc/platforms/pseries/rtas-fadump.c
> > +++ b/arch/powerpc/platforms/pseries/rtas-fadump.c
> > @@ -469,7 +469,8 @@ static int __init rtas_fadump_process(struct
> > fw_dump *fadump_conf)
> > pr_err("Dump taken by platform is
> > not valid (%d)\n", i);
> > rc = -EINVAL;
> > }
> > - if (fdm_active->rgn[i].bytes_dumped !=
> > fdm_active->rgn[i].source_len) {
> > + if (be64_to_cpu(fdm_active-
> > >rgn[i].bytes_dumped)
> > + > be64_to_cpu(fdm_active-
> > >rgn[i].source_len)) {
>
> Can you please share your observations about `bytes_dump` for both
> QEMU
> and a
> real system (LPAR) where the number of online CPUs is not equal to
> the
> maximum
I have the following observation.
I booted LPAR with 8 cpus. After crashing it I saw that fadump
CPU_STATE_DATA has notes for total 16 cpus, and only top 8 notes have
valid entries. when trying with 16, I see all NOTES have entried
filled.
I previously proposed this fake entries fix in qemu [1]. But that is
not compliant with PAPR, which states that cpu notes should be
collected for current_cpus only.
As per current states of things, reservation is always done for
maxcpus, (LPAR and QEMU both), NOTES collections is done for current
cpus on QEMU, and maxcpus on LPARs.
My proposal here is that we make the check for bytes_dumped and
source_len, less restrictive to let QEMU also generate /proc/vmcore,
when current_cpus != maxcpus.
Regards
~Shivang.
[1]
https://lore.kernel.org/qemu-devel/20260520092857.1079193-1-shivangu@linux.ibm.com/
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH] ppc/fadump: collect dump if the collected size is lesser than reserved
2026-07-22 9:11 ` Shivang Upadhyay
@ 2026-07-29 3:44 ` Sourabh Jain
0 siblings, 0 replies; 4+ messages in thread
From: Sourabh Jain @ 2026-07-29 3:44 UTC (permalink / raw)
To: Shivang Upadhyay, linuxppc-dev, linux-kernel
Cc: maddy, mpe, npiggin, chleroy, adri.vero.dev, adityag,
anushree.mathur
On 22/07/26 14:41, Shivang Upadhyay wrote:
> On Sun, 2026-07-19 at 12:39 +0530, Sourabh Jain wrote:
>> Could you please reword the commit title to make it a bit clearer?
>>
>>
>> On 14/07/26 23:00, Shivang Upadhyay wrote:
>>> When a machine is subjected to CPUs add/remove, using dlpar
>>> operations, the number of collected CPU_NOTES can change.
>>> As per PAPR, collected dump size should not be more than
>>> allocated size. Reflecting the same in source.
>> Can you add more details about the problem you are trying to solve
>> with this patch and how.
>>
>> Adding the error message and scenario would be really helpful in
>> understanding the problem.
>>
> Hi Sourabh,
>
> When a qemu ppc machine is booted with fadump=on and `-
> smp=x,maxcpus=y`, on the panic kernel, /proc/vmcore is not generated
> because dump_bytes and source_len does'nt match for CPU_STATE_DATA
> region in fadump, then we just give up on parsing rest of the data.
>
>
>> Can you add Closes tag if it is reported upstream and if possible
>> fixes
>> tag too.
>>
>>> Signed-off-by: Shivang Upadhyay <shivangu@linux.ibm.com>
>>> ---
>>> arch/powerpc/platforms/pseries/rtas-fadump.c | 3 ++-
>>> 1 file changed, 2 insertions(+), 1 deletion(-)
>>>
>>> diff --git a/arch/powerpc/platforms/pseries/rtas-fadump.c
>>> b/arch/powerpc/platforms/pseries/rtas-fadump.c
>>> index 3bb4ac2ab6cc..19a5adaf326b 100644
>>> --- a/arch/powerpc/platforms/pseries/rtas-fadump.c
>>> +++ b/arch/powerpc/platforms/pseries/rtas-fadump.c
>>> @@ -469,7 +469,8 @@ static int __init rtas_fadump_process(struct
>>> fw_dump *fadump_conf)
>>> pr_err("Dump taken by platform is
>>> not valid (%d)\n", i);
>>> rc = -EINVAL;
>>> }
>>> - if (fdm_active->rgn[i].bytes_dumped !=
>>> fdm_active->rgn[i].source_len) {
>>> + if (be64_to_cpu(fdm_active-
>>>> rgn[i].bytes_dumped)
>>> + > be64_to_cpu(fdm_active-
>>>> rgn[i].source_len)) {
>> Can you please share your observations about `bytes_dump` for both
>> QEMU
>> and a
>> real system (LPAR) where the number of online CPUs is not equal to
>> the
>> maximum
> I have the following observation.
>
> I booted LPAR with 8 cpus. After crashing it I saw that fadump
> CPU_STATE_DATA has notes for total 16 cpus, and only top 8 notes have
> valid entries.when trying with 16, I see all NOTES have entried
> filled.
Yes even I noticed the same on a LPAR with (Min=1 Desired=1 Max=2 with
SMT=8 CPUs) maxcpus as 16 and online CPUs as 8.
[ 0.036128] rtas fadump: --------CPU State Data------------
[ 0.036129] rtas fadump: Magic Number: 5245475341564500
[ 0.036131] rtas fadump: NumCpuOffset: 1c
[ 0.036132] rtas fadump: NumCpus : 16
[ 0.036135] fadump: Allocated buffer for cpu notes of size 65536 at
0xc000000006a30000
[ 0.036138] rtas fadump: Reading register data for cpu 0...
[ 0.036173] rtas fadump: Reading register data for cpu 1...
[ 0.036178] rtas fadump: Reading register data for cpu 2...
[ 0.036209] rtas fadump: Reading register data for cpu 3...
[ 0.036240] rtas fadump: Reading register data for cpu 4...
[ 0.036272] rtas fadump: Reading register data for cpu 5...
[ 0.036304] rtas fadump: Reading register data for cpu 6...
[ 0.036335] rtas fadump: Reading register data for cpu 7...
[ 0.036390] rtas fadump: Updating elfcore header (c000000006a20000)
with cpu notes
The NumCpus is populated to be 16 CPUs by the firmware (RTAS/PHYP) even
though only 8 CPUs were online.
Kernel avoid processing reg entries of CPUs which were offline using
below condition.
code snippet from rtas_fadump_build_cpu_notes()/rtas-faudmp.c
if (fdh && !cpumask_test_cpu(cpu, &fdh->cpu_mask)) {
RTAS_FADUMP_SKIP_TO_NEXT_CPU(reg_entry);
continue;
}
But when I removed the avoid condition kernel failed to process the reg
entries and
below logs were printed.
[ 0.037877] rtas fadump: --------CPU State Data------------
[ 0.037879] rtas fadump: Magic Number: 5245475341564500
[ 0.037881] rtas fadump: NumCpuOffset: 1c
[ 0.037883] rtas fadump: NumCpus : 16
[ 0.037889] fadump: Allocated buffer for cpu notes of size 65536 at
0xc000000007a50000
[ 0.037891] rtas fadump: Reading register data for cpu 0...
[ 0.037937] rtas fadump: Reading register data for cpu 1...
[ 0.037974] rtas fadump: Reading register data for cpu 2...
[ 0.038009] rtas fadump: Reading register data for cpu 3...
[ 0.038044] rtas fadump: Reading register data for cpu 4...
[ 0.038079] rtas fadump: Reading register data for cpu 5...
[ 0.038082] rtas fadump: Reading register data for cpu 6...
[ 0.038117] rtas fadump: Reading register data for cpu 7...
[ 0.038155] rtas fadump: CPU 8 was offline
[ 0.038157] rtas fadump: Reading register data for cpu 8...
[ 0.038195] rtas fadump: CPU 10 was offline
[ 0.038197] rtas fadump: Reading register data for cpu 10...
[ 0.038235] rtas fadump: CPU 12 was offline
[ 0.038236] rtas fadump: Reading register data for cpu 12...
[ 0.038273] rtas fadump: CPU 14 was offline
[ 0.038275] rtas fadump: Reading register data for cpu 14...
[ 0.038311] rtas fadump: Unable to read CPU state data
[ 0.038319] fadump: Invalidating firmware-assisted dump registration
[ 0.038370] rtas fadump: Firmware busy during fadump invalidate,
waiting 1ms (total 0ms)
[ 0.039403] fadump: reserved_memory_range[0]
[0x00000040000000-0x000000800305c7], 0x400305c8 bytes
[ 0.039408] fadump: freeing reserved memory (0x80030000 - 0xa00000000)
For some reason, only CPUs with even-numbered IDs were present from
offline CPUs list. This suggests that although
NumCpus was 16, only 8 CPUs had valid reg entries. I think we should
discuss this case with the firmware team before
finalizing the solution.
On the other hand when I increased the CPUs count on the same system
(Min=2 Desired=4 and max=8 with SMT 8) the
system was booted with 32 possible CPUs instead of 64 CPUs. And NumCpus
was 32 in fadump kernel.
>
> I previously proposed this fake entries fix in qemu [1]. But that is
> not compliant with PAPR, which states that cpu notes should be
> collected for current_cpus only.
Yes, that solution does not appear to be PAPR-compliant. However, based
on our experiments, it is unclear why the firmware reports NumCpus as 16
when only 8 CPUs are actually online.
> As per current states of things, reservation is always done for
> maxcpus, (LPAR and QEMU both), NOTES collections is done for current
> cpus on QEMU, and maxcpus on LPARs.
>
> My proposal here is that we make the check for bytes_dumped and
> source_len, less restrictive to let QEMU also generate /proc/vmcore,
> when current_cpus != maxcpus.
By the way, the changes proposed in this patch also apply to other regions,
such as HPTE and REAL_MODE. That is not the intended behavior, right?
- Sourabh Jain
>
>
> Regards
> ~Shivang.
>
> [1]
> https://lore.kernel.org/qemu-devel/20260520092857.1079193-1-shivangu@linux.ibm.com/
>
^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-07-29 3:45 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-07-14 17:30 [PATCH] ppc/fadump: collect dump if the collected size is lesser than reserved Shivang Upadhyay
2026-07-19 7:09 ` Sourabh Jain
2026-07-22 9:11 ` Shivang Upadhyay
2026-07-29 3:44 ` Sourabh Jain
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox