From: Sourabh Jain <sourabhjain@linux.ibm.com>
To: Shivang Upadhyay <shivangu@linux.ibm.com>,
linuxppc-dev@lists.ozlabs.org, linux-kernel@vger.kernel.org
Cc: maddy@linux.ibm.com, mpe@ellerman.id.au, npiggin@gmail.com,
chleroy@kernel.org, adri.vero.dev@gmail.com,
adityag@linux.ibm.com, anushree.mathur@linux.vnet.ibm.com
Subject: Re: [PATCH] ppc/fadump: collect dump if the collected size is lesser than reserved
Date: Wed, 29 Jul 2026 09:14:49 +0530 [thread overview]
Message-ID: <72589eca-2e97-4a2f-8665-b2a817a04722@linux.ibm.com> (raw)
In-Reply-To: <afd516619eea338039a6defb123ffaf38cf82949.camel@linux.ibm.com>
On 22/07/26 14:41, Shivang Upadhyay wrote:
> On Sun, 2026-07-19 at 12:39 +0530, Sourabh Jain wrote:
>> Could you please reword the commit title to make it a bit clearer?
>>
>>
>> On 14/07/26 23:00, Shivang Upadhyay wrote:
>>> When a machine is subjected to CPUs add/remove, using dlpar
>>> operations, the number of collected CPU_NOTES can change.
>>> As per PAPR, collected dump size should not be more than
>>> allocated size. Reflecting the same in source.
>> Can you add more details about the problem you are trying to solve
>> with this patch and how.
>>
>> Adding the error message and scenario would be really helpful in
>> understanding the problem.
>>
> Hi Sourabh,
>
> When a qemu ppc machine is booted with fadump=on and `-
> smp=x,maxcpus=y`, on the panic kernel, /proc/vmcore is not generated
> because dump_bytes and source_len does'nt match for CPU_STATE_DATA
> region in fadump, then we just give up on parsing rest of the data.
>
>
>> Can you add Closes tag if it is reported upstream and if possible
>> fixes
>> tag too.
>>
>>> Signed-off-by: Shivang Upadhyay <shivangu@linux.ibm.com>
>>> ---
>>> arch/powerpc/platforms/pseries/rtas-fadump.c | 3 ++-
>>> 1 file changed, 2 insertions(+), 1 deletion(-)
>>>
>>> diff --git a/arch/powerpc/platforms/pseries/rtas-fadump.c
>>> b/arch/powerpc/platforms/pseries/rtas-fadump.c
>>> index 3bb4ac2ab6cc..19a5adaf326b 100644
>>> --- a/arch/powerpc/platforms/pseries/rtas-fadump.c
>>> +++ b/arch/powerpc/platforms/pseries/rtas-fadump.c
>>> @@ -469,7 +469,8 @@ static int __init rtas_fadump_process(struct
>>> fw_dump *fadump_conf)
>>> pr_err("Dump taken by platform is
>>> not valid (%d)\n", i);
>>> rc = -EINVAL;
>>> }
>>> - if (fdm_active->rgn[i].bytes_dumped !=
>>> fdm_active->rgn[i].source_len) {
>>> + if (be64_to_cpu(fdm_active-
>>>> rgn[i].bytes_dumped)
>>> + > be64_to_cpu(fdm_active-
>>>> rgn[i].source_len)) {
>> Can you please share your observations about `bytes_dump` for both
>> QEMU
>> and a
>> real system (LPAR) where the number of online CPUs is not equal to
>> the
>> maximum
> I have the following observation.
>
> I booted LPAR with 8 cpus. After crashing it I saw that fadump
> CPU_STATE_DATA has notes for total 16 cpus, and only top 8 notes have
> valid entries.when trying with 16, I see all NOTES have entried
> filled.
Yes even I noticed the same on a LPAR with (Min=1 Desired=1 Max=2 with
SMT=8 CPUs) maxcpus as 16 and online CPUs as 8.
[ 0.036128] rtas fadump: --------CPU State Data------------
[ 0.036129] rtas fadump: Magic Number: 5245475341564500
[ 0.036131] rtas fadump: NumCpuOffset: 1c
[ 0.036132] rtas fadump: NumCpus : 16
[ 0.036135] fadump: Allocated buffer for cpu notes of size 65536 at
0xc000000006a30000
[ 0.036138] rtas fadump: Reading register data for cpu 0...
[ 0.036173] rtas fadump: Reading register data for cpu 1...
[ 0.036178] rtas fadump: Reading register data for cpu 2...
[ 0.036209] rtas fadump: Reading register data for cpu 3...
[ 0.036240] rtas fadump: Reading register data for cpu 4...
[ 0.036272] rtas fadump: Reading register data for cpu 5...
[ 0.036304] rtas fadump: Reading register data for cpu 6...
[ 0.036335] rtas fadump: Reading register data for cpu 7...
[ 0.036390] rtas fadump: Updating elfcore header (c000000006a20000)
with cpu notes
The NumCpus is populated to be 16 CPUs by the firmware (RTAS/PHYP) even
though only 8 CPUs were online.
Kernel avoid processing reg entries of CPUs which were offline using
below condition.
code snippet from rtas_fadump_build_cpu_notes()/rtas-faudmp.c
if (fdh && !cpumask_test_cpu(cpu, &fdh->cpu_mask)) {
RTAS_FADUMP_SKIP_TO_NEXT_CPU(reg_entry);
continue;
}
But when I removed the avoid condition kernel failed to process the reg
entries and
below logs were printed.
[ 0.037877] rtas fadump: --------CPU State Data------------
[ 0.037879] rtas fadump: Magic Number: 5245475341564500
[ 0.037881] rtas fadump: NumCpuOffset: 1c
[ 0.037883] rtas fadump: NumCpus : 16
[ 0.037889] fadump: Allocated buffer for cpu notes of size 65536 at
0xc000000007a50000
[ 0.037891] rtas fadump: Reading register data for cpu 0...
[ 0.037937] rtas fadump: Reading register data for cpu 1...
[ 0.037974] rtas fadump: Reading register data for cpu 2...
[ 0.038009] rtas fadump: Reading register data for cpu 3...
[ 0.038044] rtas fadump: Reading register data for cpu 4...
[ 0.038079] rtas fadump: Reading register data for cpu 5...
[ 0.038082] rtas fadump: Reading register data for cpu 6...
[ 0.038117] rtas fadump: Reading register data for cpu 7...
[ 0.038155] rtas fadump: CPU 8 was offline
[ 0.038157] rtas fadump: Reading register data for cpu 8...
[ 0.038195] rtas fadump: CPU 10 was offline
[ 0.038197] rtas fadump: Reading register data for cpu 10...
[ 0.038235] rtas fadump: CPU 12 was offline
[ 0.038236] rtas fadump: Reading register data for cpu 12...
[ 0.038273] rtas fadump: CPU 14 was offline
[ 0.038275] rtas fadump: Reading register data for cpu 14...
[ 0.038311] rtas fadump: Unable to read CPU state data
[ 0.038319] fadump: Invalidating firmware-assisted dump registration
[ 0.038370] rtas fadump: Firmware busy during fadump invalidate,
waiting 1ms (total 0ms)
[ 0.039403] fadump: reserved_memory_range[0]
[0x00000040000000-0x000000800305c7], 0x400305c8 bytes
[ 0.039408] fadump: freeing reserved memory (0x80030000 - 0xa00000000)
For some reason, only CPUs with even-numbered IDs were present from
offline CPUs list. This suggests that although
NumCpus was 16, only 8 CPUs had valid reg entries. I think we should
discuss this case with the firmware team before
finalizing the solution.
On the other hand when I increased the CPUs count on the same system
(Min=2 Desired=4 and max=8 with SMT 8) the
system was booted with 32 possible CPUs instead of 64 CPUs. And NumCpus
was 32 in fadump kernel.
>
> I previously proposed this fake entries fix in qemu [1]. But that is
> not compliant with PAPR, which states that cpu notes should be
> collected for current_cpus only.
Yes, that solution does not appear to be PAPR-compliant. However, based
on our experiments, it is unclear why the firmware reports NumCpus as 16
when only 8 CPUs are actually online.
> As per current states of things, reservation is always done for
> maxcpus, (LPAR and QEMU both), NOTES collections is done for current
> cpus on QEMU, and maxcpus on LPARs.
>
> My proposal here is that we make the check for bytes_dumped and
> source_len, less restrictive to let QEMU also generate /proc/vmcore,
> when current_cpus != maxcpus.
By the way, the changes proposed in this patch also apply to other regions,
such as HPTE and REAL_MODE. That is not the intended behavior, right?
- Sourabh Jain
>
>
> Regards
> ~Shivang.
>
> [1]
> https://lore.kernel.org/qemu-devel/20260520092857.1079193-1-shivangu@linux.ibm.com/
>
next prev parent reply other threads:[~2026-07-29 3:45 UTC|newest]
Thread overview: 5+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-14 17:30 [PATCH] ppc/fadump: collect dump if the collected size is lesser than reserved Shivang Upadhyay
2026-07-19 7:09 ` Sourabh Jain
2026-07-22 9:11 ` Shivang Upadhyay
2026-07-29 3:44 ` Sourabh Jain [this message]
2026-07-29 6:06 ` Shivang Upadhyay
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=72589eca-2e97-4a2f-8665-b2a817a04722@linux.ibm.com \
--to=sourabhjain@linux.ibm.com \
--cc=adityag@linux.ibm.com \
--cc=adri.vero.dev@gmail.com \
--cc=anushree.mathur@linux.vnet.ibm.com \
--cc=chleroy@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linuxppc-dev@lists.ozlabs.org \
--cc=maddy@linux.ibm.com \
--cc=mpe@ellerman.id.au \
--cc=npiggin@gmail.com \
--cc=shivangu@linux.ibm.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox