* Re: [PATCH v4] powerpc/pseries: read the lpar name from the firmware
From: Laurent Dufour @ 2022-01-06 9:17 UTC (permalink / raw)
To: Tyrel Datwyler, Nathan Lynch, Michael Ellerman; +Cc: linuxppc-dev, linux-kernel
In-Reply-To: <ac208963-d334-1f46-0db2-4a8d073b2963@linux.ibm.com>
On 06/01/2022, 02:17:21, Tyrel Datwyler wrote:
> On 1/5/22 3:19 PM, Nathan Lynch wrote:
>> Laurent Dufour <ldufour@linux.ibm.com> writes:
>>> On 07/12/2021, 18:11:09, Laurent Dufour wrote:
>>>> The LPAR name may be changed after the LPAR has been started in the HMC.
>>>> In that case lparstat command is not reporting the updated value because it
>>>> reads it from the device tree which is read at boot time.
>>>>
>>>> However this value could be read from RTAS.
>>>>
>>>> Adding this value in the /proc/powerpc/lparcfg output allows to read the
>>>> updated value.
>>>
>>> Do you consider taking that patch soon?
>>
>> This version prints an error on non-PowerVM guests the first time
>> lparcfg is read.
>
> I assume because QEMU doesn't implement the LPAR_NAME token for get_sysparm.
>
>>
>> And I still contend that having this function fall back to reporting the
>> partition name in the DT would provide a beneficial consistency in the
>> user-facing API, allowing programs to avoid hypervisor-specific branches
>> in their code.
>
> Agreed, if the get_sysparm fails just report the lpar-name from the device tree.
My aim is to not do in the kernel what can be easily done in user space but
avoiding user space program hypervisor-specific branches is a good point.
Note that if the RTAS call has been available to unprivileged user, all
that stuff would have been made in user space, so hypervisor-specific...
Anyway, I'll work on a new version fetching the DT value in the case the
RTAS call is failing.
Thanks,
Laurent.
^ permalink raw reply
* Re: [PATCH]powerpc/xmon: Dump XIVE information for online-only processors.
From: Sachin Sant @ 2022-01-06 5:29 UTC (permalink / raw)
To: Cédric Le Goater; +Cc: linuxppc-dev
In-Reply-To: <c28f88f4-c2f8-fc91-2a0f-eb9b96889274@kaod.org>
> On 05-Jan-2022, at 8:09 PM, Cédric Le Goater <clg@kaod.org> wrote:
>
> On 1/5/22 15:17, Sachin Sant wrote:
>> dxa command in XMON debugger iterates through all possible processors.
>> As a result, empty lines are printed even for processors which are not
>> online.
>> CPU 47:pp=00 CPPR=ff IPI=0x0040002f PQ=-- EQ idx=699 T=0 00000000 00000000
>> CPU 48:
>> CPU 49:
>> Restrict XIVE information(dxa) to be displayed for online processors only.
>> Signed-off-by: Sachin Sant <sachinp@linux.vnet.ibm.com>
>
> Looks good to me. We should do the same for :
>
> /sys/kernel/debug/powerpc/xive/ipis
Thanks for the pointer. Will send a separate patch for this issue.
- Sachin
>
> Reviewed-by: Cédric Le Goater <clg@kaod.org>
>
> Thanks,
>
> C.
>
^ permalink raw reply
* Re: [5.16.0-rc5][ppc][net] kernel oops when hotplug remove of vNIC interface
From: Michael Ellerman @ 2022-01-06 4:19 UTC (permalink / raw)
To: Jakub Kicinski, Abdul Haleem
Cc: dumazet, netdev, linux-kernel, Dany Madden, alexandr.lobakin,
brian King, Sukadev Bhattiprolu, linuxppc-dev
In-Reply-To: <20220105102625.2738186e@kicinski-fedora-pc1c0hjn.dhcp.thefacebook.com>
Jakub Kicinski <kuba@kernel.org> writes:
> On Wed, 5 Jan 2022 13:56:53 +0530 Abdul Haleem wrote:
>> Greeting's
>>
>> Mainline kernel 5.16.0-rc5 panics when DLPAR ADD of vNIC device on my
>> Powerpc LPAR
>>
>> Perform below dlpar commands in a loop from linux OS
>>
>> drmgr -r -c slot -s U9080.HEX.134C488-V1-C3 -w 5 -d 1
>> drmgr -a -c slot -s U9080.HEX.134C488-V1-C3 -w 5 -d 1
>>
>> after 7th iteration, the kernel panics with below messages
>>
>> console messages:
>> [102056] ibmvnic 30000003 env3: Sending CRQ: 801e000864000000
>> 0060000000000000
>> <intr> ibmvnic 30000003 env3: Handling CRQ: 809e000800000000
>> 0000000000000000
>> [102056] ibmvnic 30000003 env3: Disabling tx_scrq[0] irq
>> [102056] ibmvnic 30000003 env3: Disabling tx_scrq[1] irq
>> [102056] ibmvnic 30000003 env3: Disabling rx_scrq[0] irq
>> [102056] ibmvnic 30000003 env3: Disabling rx_scrq[1] irq
>> [102056] ibmvnic 30000003 env3: Disabling rx_scrq[2] irq
>> [102056] ibmvnic 30000003 env3: Disabling rx_scrq[3] irq
>> [102056] ibmvnic 30000003 env3: Disabling rx_scrq[4] irq
>> [102056] ibmvnic 30000003 env3: Disabling rx_scrq[5] irq
>> [102056] ibmvnic 30000003 env3: Disabling rx_scrq[6] irq
>> [102056] ibmvnic 30000003 env3: Disabling rx_scrq[7] irq
>> [102056] ibmvnic 30000003 env3: Replenished 8 pools
>> Kernel attempted to read user page (10) - exploit attempt? (uid: 0)
>> BUG: Kernel NULL pointer dereference on read at 0x00000010
>> Faulting instruction address: 0xc000000000a3c840
>> Oops: Kernel access of bad area, sig: 11 [#1]
>> LE PAGE_SIZE=64K MMU=Radix SMP NR_CPUS=2048 NUMA pSeries
>> Modules linked in: bridge stp llc ib_core rpadlpar_io rpaphp nfnetlink
>> tcp_diag udp_diag inet_diag unix_diag af_packet_diag netlink_diag
>> bonding rfkill ibmvnic sunrpc pseries_rng xts vmx_crypto gf128mul
>> sch_fq_codel binfmt_misc ip_tables ext4 mbcache jbd2 dm_service_time
>> sd_mod t10_pi sg ibmvfc scsi_transport_fc ibmveth dm_multipath dm_mirror
>> dm_region_hash dm_log dm_mod fuse
>> CPU: 9 PID: 102056 Comm: kworker/9:2 Kdump: loaded Not tainted
>> 5.16.0-rc5-autotest-g6441998e2e37 #1
>> Workqueue: events_long __ibmvnic_reset [ibmvnic]
>> NIP: c000000000a3c840 LR: c0080000029b5378 CTR: c000000000a3c820
>> REGS: c0000000548e37e0 TRAP: 0300 Not tainted
>> (5.16.0-rc5-autotest-g6441998e2e37)
>> MSR: 8000000000009033 <SF,EE,ME,IR,DR,RI,LE> CR: 28248484 XER: 00000004
>> CFAR: c0080000029bdd24 DAR: 0000000000000010 DSISR: 40000000 IRQMASK: 0
>> GPR00: c0080000029b55d0 c0000000548e3a80 c0000000028f0200 0000000000000000
>> GPR04: c000000c7d1a7e00 fffffffffffffff6 0000000000000027 c000000c7d1a7e08
>> GPR08: 0000000000000023 0000000000000000 0000000000000010 c0080000029bdd10
>> GPR12: c000000000a3c820 c000000c7fca6680 0000000000000000 c000000133016bf8
>> GPR16: 00000000000003fe 0000000000001000 0000000000000002 0000000000000008
>> GPR20: c000000133016eb0 0000000000000000 0000000000000000 0000000000000003
>> GPR24: c000000133016000 c000000133017168 0000000020000000 c000000133016a00
>> GPR28: 0000000000000006 c000000133016a00 0000000000000001 c000000133016000
>> NIP [c000000000a3c840] napi_enable+0x20/0xc0
>> LR [c0080000029b5378] __ibmvnic_open+0xf0/0x430 [ibmvnic]
>> Call Trace:
>> [c0000000548e3a80] [0000000000000006] 0x6 (unreliable)
>> [c0000000548e3ab0] [c0080000029b55d0] __ibmvnic_open+0x348/0x430 [ibmvnic]
>> [c0000000548e3b40] [c0080000029bcc28] __ibmvnic_reset+0x500/0xdf0 [ibmvnic]
>> [c0000000548e3c60] [c000000000176228] process_one_work+0x288/0x570
>> [c0000000548e3d00] [c000000000176588] worker_thread+0x78/0x660
>> [c0000000548e3da0] [c0000000001822f0] kthread+0x1c0/0x1d0
>> [c0000000548e3e10] [c00000000000cf64] ret_from_kernel_thread+0x5c/0x64
>> Instruction dump:
>> 7d2948f8 792307e0 4e800020 60000000 3c4c01eb 384239e0 f821ffd1 39430010
>> 38a0fff6 e92d1100 f9210028 39200000 <e9030010> f9010020 60420000 e9210020
>> ---[ end trace 5f8033b08fd27706 ]---
>> radix-mmu: Page sizes from device-tree:
>>
>> the fault instruction points to
>>
>> [root@ltcden11-lp1 boot]# gdb -batch
>> vmlinuz-5.16.0-rc5-autotest-g6441998e2e37 -ex 'list *(0xc000000000a3c840)'
>> 0xc000000000a3c840 is in napi_enable (net/core/dev.c:6966).
>> 6961 void napi_enable(struct napi_struct *n)
>> 6962 {
>> 6963 unsigned long val, new;
>> 6964
>> 6965 do {
>> 6966 val = READ_ONCE(n->state);
>
> If n is NULL here that's gotta be a driver problem.
Definitely looks like it, the disassembly is:
not r9,r9
clrldi r3,r9,63
blr # end of previous function
nop
addis r2,r12,491 # function entry
addi r2,r2,14816
stdu r1,-48(r1) # stack frame creation
li r5,-10
ld r9,4352(r13)
std r9,40(r1)
li r9,0
ld r8,16(r3) # load from r3 (n) + 16
The register dump shows that r3 is NULL, and it comes directly from the
caller. So we've been called with n = NULL.
cheers
^ permalink raw reply
* Re: ppc64el kernel bug?
From: Michael Ellerman @ 2022-01-06 4:03 UTC (permalink / raw)
To: Nathan Lynch, Kip Warner
Cc: Aneesh Kumar K.V, Tulio Magno Quites Machado Filho, linuxppc-dev
In-Reply-To: <87bl0pvcx9.fsf@li-e15d104c-2135-11b2-a85c-d7ef17e56be6.ibm.com>
Nathan Lynch <nathanl@linux.ibm.com> writes:
> Kip Warner <kip@thevertigo.com> writes:
>> Dec 25 06:52:52 romulus-server kernel: [28835.277591] BUG: Unable to handle kernel data access on write at 0x132b47d38499fd58
>> Dec 25 06:52:52 romulus-server kernel: [28835.277624] Faulting instruction address: 0xc0000000004d0434
>> Dec 25 06:52:52 romulus-server kernel: [28835.277636] Oops: Kernel access of bad area, sig: 11 [#150]
>> Dec 25 06:52:52 romulus-server kernel: [28835.277656] LE PAGE_SIZE=64K MMU=Radix SMP NR_CPUS=2048 NUMA PowerNV
>> Dec 25 06:52:52 romulus-server kernel: [28835.277669] Modules linked in: veth nft_masq zfs(PO) zunicode(PO) zzstd(O) zlua(O) zcommon(PO) znvpair(PO) zavl(PO) icp(PO) spl(O) vhost_vsock vmw_vsock_virtio_transport_common vhost vhost_iotlb vsock xt_CHECKSUM nft_chain_nat xt_MASQUERADE nf_nat nf_conntrack nf_defrag_ipv6 nf_defrag_ipv4 nft_counter xt_tcpudp nft_compat bridge stp llc nf_tables nfnetlink binfmt_misc dm_multipath scsi_dh_rdac scsi_dh_emc scsi_dh_alua joydev input_leds ipmi_powernv mac_hid ipmi_devintf ipmi_msghandler ofpart cmdlinepart at24 powernv_flash mtd uio_pdrv_genirq opal_prd uio ibmpowernv vmx_crypto sch_fq_codel jc42 ip_tables x_tables autofs4 xfs btrfs blake2b_generic raid10 raid456 async_raid6_recov async_memcpy async_pq async_xor async_tx xor hid_generic usbhid hid raid6_pq libcrc32c raid1 raid0 multipath linear nouveau ses enclosure scsi_transport_sas ast drm_vram_helper i2c_algo_bit drm_ttm_helper ttm drm_kms_helper syscopyarea sysfillrect sysimgblt fb_sys_fops cec rc_core drm crct10dif_vpmsum
>> Dec 25 06:52:52 romulus-server kernel: [28835.277776] crc32c_vpmsum xhci_pci tg3 aacraid xhci_pci_renesas drm_panel_orientation_quirks
>> Dec 25 06:52:52 romulus-server kernel: [28835.277918] CPU: 26 PID: 144937 Comm: postgres Tainted: P D O 5.11.0-41-generic #45-Ubuntu
>> Dec 25 06:52:52 romulus-server kernel: [28835.277943] NIP: c0000000004d0434 LR: c0000000004d032c CTR: c0000000010a90e0
>> Dec 25 06:52:52 romulus-server kernel: [28835.277975] REGS: c000000056b9f6b0 TRAP: 0380 Tainted: P D O (5.11.0-41-generic)
>> Dec 25 06:52:52 romulus-server kernel: [28835.278008] MSR: 9000000000009033 <SF,HV,EE,ME,IR,DR,RI,LE> CR: 88002281 XER: 0000008c
>> Dec 25 06:52:52 romulus-server kernel: [28835.278050] CFAR: c0000000004d041c IRQMASK: 0
>> Dec 25 06:52:52 romulus-server kernel: [28835.278050] GPR00: c0000000004d032c c000000056b9f950 c000000002409a00 0000000000000000
>> Dec 25 06:52:52 romulus-server kernel: [28835.278050] GPR04: 0000000000400cc0 0000000000000097 ffffffffffffffff c000000ffda9d0d0
>> Dec 25 06:52:52 romulus-server kernel: [28835.278050] GPR08: 0000000ffbd90000 132b47d38499fce8 0000000000000070 d4ff277338704e25
>> Dec 25 06:52:52 romulus-server kernel: [28835.278050] GPR12: 0000000000002000 c000000ffffd2c00 0000000000000000 c000000116c512d0
>> Dec 25 06:52:52 romulus-server kernel: [28835.278050] GPR16: 0000000000000154 c000000116c51570 c000000056b9fc88 0000000000000154
>> Dec 25 06:52:52 romulus-server kernel: [28835.278050] GPR20: 0000000000000000 0000000000000000 0000000000000000 0000000000000000
>> Dec 25 06:52:52 romulus-server kernel: [28835.278050] GPR24: c000000000ecccc0 0000000000000001 c0000000024588fc c000000000ec9954
>> Dec 25 06:52:52 romulus-server kernel: [28835.278050] GPR28: ffffffffffffffff c00000001d597e40 0000000000400cc0 c000000003018880
>> Dec 25 06:52:52 romulus-server kernel: [28835.278213] NIP [c0000000004d0434] kmem_cache_alloc_node+0x1d4/0x490
>> Dec 25 06:52:52 romulus-server kernel: [28835.278237] LR [c0000000004d032c] kmem_cache_alloc_node+0xcc/0x490
>> Dec 25 06:52:52 romulus-server kernel: [28835.278268] Call Trace:
>> Dec 25 06:52:52 romulus-server kernel: [28835.278283] [c000000056b9f950] [c0000000004d032c] kmem_cache_alloc_node+0xcc/0x490 (unreliable)
>> Dec 25 06:52:52 romulus-server kernel: [28835.278328] [c000000056b9f9c0] [c000000000ec9954] __alloc_skb+0x74/0x2d0
>> Dec 25 06:52:52 romulus-server kernel: [28835.278369] [c000000056b9fa20] [c000000000ecccc0] alloc_skb_with_frags+0x70/0x2e0
>> Dec 25 06:52:52 romulus-server kernel: [28835.278403] [c000000056b9faa0] [c000000000ec0f38] sock_alloc_send_pskb+0x1d8/0x200
>> Dec 25 06:52:52 romulus-server kernel: [28835.278436] [c000000056b9fb10] [c0000000010a93a8] unix_stream_sendmsg+0x2c8/0x710
>> Dec 25 06:52:52 romulus-server kernel: [28835.278471] [c000000056b9fc10] [c000000000eb64e0] sock_sendmsg+0x80/0xb0
>> Dec 25 06:52:52 romulus-server kernel: [28835.278494] [c000000056b9fc40] [c000000000ebab88] __sys_sendto+0xf8/0x1a0
>> Dec 25 06:52:52 romulus-server kernel: [28835.278526] [c000000056b9fd90] [c000000000ebaca0] sys_send+0x30/0x40
>> Dec 25 06:52:52 romulus-server kernel: [28835.278558] [c000000056b9fdb0] [c000000000036ffc] system_call_exception+0x10c/0x230
>> Dec 25 06:52:52 romulus-server kernel: [28835.278601] [c000000056b9fe10] [c00000000000d374] system_call_vectored_common+0xf4/0x26c
>> Dec 25 06:52:52 romulus-server kernel: [28835.278634] --- interrupt: 3000 at 0x7ec638a194f4
>> Dec 25 06:52:52 romulus-server kernel: [28835.278654] NIP: 00007ec638a194f4 LR: 0000000000000000 CTR: 0000000000000000
>> Dec 25 06:52:52 romulus-server kernel: [28835.278685] REGS: c000000056b9fe80 TRAP: 3000 Tainted: P D O (5.11.0-41-generic)
>> Dec 25 06:52:52 romulus-server kernel: [28835.278719] MSR: 900000000280f033 <SF,HV,VEC,VSX,EE,PR,FP,ME,IR,DR,RI,LE> CR: 48008281 XER: 00000000
>> Dec 25 06:52:52 romulus-server kernel: [28835.278766] IRQMASK: 0
>> Dec 25 06:52:52 romulus-server kernel: [28835.278766] GPR00: 000000000000014e 00007fffe99c1800 00007ec638a47f00 0000000000000009
>> Dec 25 06:52:52 romulus-server kernel: [28835.278766] GPR04: 00000043809d1148 0000000000000154 0000000000000000 0000000000001ae8
>> Dec 25 06:52:52 romulus-server kernel: [28835.278766] GPR08: 0000004362347d00 0000000000000000 0000000000000000 0000000000000000
>> Dec 25 06:52:52 romulus-server kernel: [28835.278766] GPR12: 0000000000000000 00007ec6348e0890 0000000000000000 ffffffffffffffff
>> Dec 25 06:52:52 romulus-server kernel: [28835.278766] GPR16: 0000000000000000 000000436233f7a0 0000000000000001 0000000000000000
>> Dec 25 06:52:52 romulus-server kernel: [28835.278766] GPR20: 00007fffe99c18ac 0000004362344f48 0000000000000004 00007fffe99c18b0
>> Dec 25 06:52:52 romulus-server kernel: [28835.278766] GPR24: 0000000006000001 0000000000000000 0000000000000154 00000043809d1148
>> Dec 25 06:52:52 romulus-server kernel: [28835.278766] GPR28: 0000000000000000 00007ec6348d9938 00000043809ceb00 000000000000000b
>> Dec 25 06:52:52 romulus-server kernel: [28835.278992] NIP [00007ec638a194f4] 0x7ec638a194f4
>> Dec 25 06:52:52 romulus-server kernel: [28835.279020] LR [0000000000000000] 0x0
>> Dec 25 06:52:52 romulus-server kernel: [28835.279038] --- interrupt: 3000
>> Dec 25 06:52:52 romulus-server kernel: [28835.279054] Instruction dump:
>> Dec 25 06:52:52 romulus-server kernel: [28835.279072] f9210020 41820098 2e1cffff 3b200001 2c2a0000 41820088 41920010 894a0007
>> Dec 25 06:52:52 romulus-server kernel: [28835.279110] 7c1c5000 40820078 815f0028 e97f00b8 <7ce9502a> 7c095214 886d0988 9b2d0988
>> Dec 25 06:52:52 romulus-server kernel: [28835.279141] ---[ end trace fe7ee98d0b7beb6a ]---
>
> Perhaps slab corruption, but the 'D' taint flag (TAINT_DIE) means the
> kernel oopsed at least once before this. Probably best to look at that
> one first.
You also have the 'P' taint for a proprietary module loaded, so we
(upstream) can't really help with that, you're better off reporting to
your distro.
If it's easily reproducible you could boot with slub_debug=FZP and see
if that catches the slab corruption earlier, that might help us identify
the actual problem.
cheers
^ permalink raw reply
* Re: [PATCH v4] powerpc/pseries: read the lpar name from the firmware
From: Tyrel Datwyler @ 2022-01-06 2:27 UTC (permalink / raw)
To: Nathan Lynch; +Cc: Laurent Dufour, linuxppc-dev, linux-kernel
In-Reply-To: <87ee5lve64.fsf@li-e15d104c-2135-11b2-a85c-d7ef17e56be6.ibm.com>
On 1/5/22 5:36 PM, Nathan Lynch wrote:
> Tyrel Datwyler <tyreld@linux.ibm.com> writes:
>> On 1/5/22 3:19 PM, Nathan Lynch wrote:
>>>
>>
>> Is there benefit of adding a partition_name field/value pair to lparcfg? The
>> lparstat utility can just as easily make the get_sysparm call via librtas.
>> Further, rtas_filters allows this particular RTAS call from userspace.
>
> The RTAS syscall is root-only, but we want the partition name (whether
> supplied by RTAS or the device tree) to be available to unprivileged
> programs.
>
Ah, right. I recall this discussion now from previous iterations.
-Tyrel
^ permalink raw reply
* Re: ppc64el kernel bug?
From: Nathan Lynch @ 2022-01-06 2:02 UTC (permalink / raw)
To: Kip Warner
Cc: Aneesh Kumar K.V, Tulio Magno Quites Machado Filho, linuxppc-dev
In-Reply-To: <ab1cf0bfad75e06ee2a56ebcf435a977f463b2d6.camel@thevertigo.com>
Kip Warner <kip@thevertigo.com> writes:
> Dec 25 06:52:52 romulus-server kernel: [28835.277591] BUG: Unable to handle kernel data access on write at 0x132b47d38499fd58
> Dec 25 06:52:52 romulus-server kernel: [28835.277624] Faulting instruction address: 0xc0000000004d0434
> Dec 25 06:52:52 romulus-server kernel: [28835.277636] Oops: Kernel access of bad area, sig: 11 [#150]
> Dec 25 06:52:52 romulus-server kernel: [28835.277656] LE PAGE_SIZE=64K MMU=Radix SMP NR_CPUS=2048 NUMA PowerNV
> Dec 25 06:52:52 romulus-server kernel: [28835.277669] Modules linked in: veth nft_masq zfs(PO) zunicode(PO) zzstd(O) zlua(O) zcommon(PO) znvpair(PO) zavl(PO) icp(PO) spl(O) vhost_vsock vmw_vsock_virtio_transport_common vhost vhost_iotlb vsock xt_CHECKSUM nft_chain_nat xt_MASQUERADE nf_nat nf_conntrack nf_defrag_ipv6 nf_defrag_ipv4 nft_counter xt_tcpudp nft_compat bridge stp llc nf_tables nfnetlink binfmt_misc dm_multipath scsi_dh_rdac scsi_dh_emc scsi_dh_alua joydev input_leds ipmi_powernv mac_hid ipmi_devintf ipmi_msghandler ofpart cmdlinepart at24 powernv_flash mtd uio_pdrv_genirq opal_prd uio ibmpowernv vmx_crypto sch_fq_codel jc42 ip_tables x_tables autofs4 xfs btrfs blake2b_generic raid10 raid456 async_raid6_recov async_memcpy async_pq async_xor async_tx xor hid_generic usbhid hid raid6_pq libcrc32c raid1 raid0 multipath linear nouveau ses enclosure scsi_transport_sas ast drm_vram_helper i2c_algo_bit drm_ttm_helper ttm drm_kms_helper syscopyarea sysfillrect sysimgblt fb_sys_fops cec rc_core drm crct10dif_vpmsum
> Dec 25 06:52:52 romulus-server kernel: [28835.277776] crc32c_vpmsum xhci_pci tg3 aacraid xhci_pci_renesas drm_panel_orientation_quirks
> Dec 25 06:52:52 romulus-server kernel: [28835.277918] CPU: 26 PID: 144937 Comm: postgres Tainted: P D O 5.11.0-41-generic #45-Ubuntu
> Dec 25 06:52:52 romulus-server kernel: [28835.277943] NIP: c0000000004d0434 LR: c0000000004d032c CTR: c0000000010a90e0
> Dec 25 06:52:52 romulus-server kernel: [28835.277975] REGS: c000000056b9f6b0 TRAP: 0380 Tainted: P D O (5.11.0-41-generic)
> Dec 25 06:52:52 romulus-server kernel: [28835.278008] MSR: 9000000000009033 <SF,HV,EE,ME,IR,DR,RI,LE> CR: 88002281 XER: 0000008c
> Dec 25 06:52:52 romulus-server kernel: [28835.278050] CFAR: c0000000004d041c IRQMASK: 0
> Dec 25 06:52:52 romulus-server kernel: [28835.278050] GPR00: c0000000004d032c c000000056b9f950 c000000002409a00 0000000000000000
> Dec 25 06:52:52 romulus-server kernel: [28835.278050] GPR04: 0000000000400cc0 0000000000000097 ffffffffffffffff c000000ffda9d0d0
> Dec 25 06:52:52 romulus-server kernel: [28835.278050] GPR08: 0000000ffbd90000 132b47d38499fce8 0000000000000070 d4ff277338704e25
> Dec 25 06:52:52 romulus-server kernel: [28835.278050] GPR12: 0000000000002000 c000000ffffd2c00 0000000000000000 c000000116c512d0
> Dec 25 06:52:52 romulus-server kernel: [28835.278050] GPR16: 0000000000000154 c000000116c51570 c000000056b9fc88 0000000000000154
> Dec 25 06:52:52 romulus-server kernel: [28835.278050] GPR20: 0000000000000000 0000000000000000 0000000000000000 0000000000000000
> Dec 25 06:52:52 romulus-server kernel: [28835.278050] GPR24: c000000000ecccc0 0000000000000001 c0000000024588fc c000000000ec9954
> Dec 25 06:52:52 romulus-server kernel: [28835.278050] GPR28: ffffffffffffffff c00000001d597e40 0000000000400cc0 c000000003018880
> Dec 25 06:52:52 romulus-server kernel: [28835.278213] NIP [c0000000004d0434] kmem_cache_alloc_node+0x1d4/0x490
> Dec 25 06:52:52 romulus-server kernel: [28835.278237] LR [c0000000004d032c] kmem_cache_alloc_node+0xcc/0x490
> Dec 25 06:52:52 romulus-server kernel: [28835.278268] Call Trace:
> Dec 25 06:52:52 romulus-server kernel: [28835.278283] [c000000056b9f950] [c0000000004d032c] kmem_cache_alloc_node+0xcc/0x490 (unreliable)
> Dec 25 06:52:52 romulus-server kernel: [28835.278328] [c000000056b9f9c0] [c000000000ec9954] __alloc_skb+0x74/0x2d0
> Dec 25 06:52:52 romulus-server kernel: [28835.278369] [c000000056b9fa20] [c000000000ecccc0] alloc_skb_with_frags+0x70/0x2e0
> Dec 25 06:52:52 romulus-server kernel: [28835.278403] [c000000056b9faa0] [c000000000ec0f38] sock_alloc_send_pskb+0x1d8/0x200
> Dec 25 06:52:52 romulus-server kernel: [28835.278436] [c000000056b9fb10] [c0000000010a93a8] unix_stream_sendmsg+0x2c8/0x710
> Dec 25 06:52:52 romulus-server kernel: [28835.278471] [c000000056b9fc10] [c000000000eb64e0] sock_sendmsg+0x80/0xb0
> Dec 25 06:52:52 romulus-server kernel: [28835.278494] [c000000056b9fc40] [c000000000ebab88] __sys_sendto+0xf8/0x1a0
> Dec 25 06:52:52 romulus-server kernel: [28835.278526] [c000000056b9fd90] [c000000000ebaca0] sys_send+0x30/0x40
> Dec 25 06:52:52 romulus-server kernel: [28835.278558] [c000000056b9fdb0] [c000000000036ffc] system_call_exception+0x10c/0x230
> Dec 25 06:52:52 romulus-server kernel: [28835.278601] [c000000056b9fe10] [c00000000000d374] system_call_vectored_common+0xf4/0x26c
> Dec 25 06:52:52 romulus-server kernel: [28835.278634] --- interrupt: 3000 at 0x7ec638a194f4
> Dec 25 06:52:52 romulus-server kernel: [28835.278654] NIP: 00007ec638a194f4 LR: 0000000000000000 CTR: 0000000000000000
> Dec 25 06:52:52 romulus-server kernel: [28835.278685] REGS: c000000056b9fe80 TRAP: 3000 Tainted: P D O (5.11.0-41-generic)
> Dec 25 06:52:52 romulus-server kernel: [28835.278719] MSR: 900000000280f033 <SF,HV,VEC,VSX,EE,PR,FP,ME,IR,DR,RI,LE> CR: 48008281 XER: 00000000
> Dec 25 06:52:52 romulus-server kernel: [28835.278766] IRQMASK: 0
> Dec 25 06:52:52 romulus-server kernel: [28835.278766] GPR00: 000000000000014e 00007fffe99c1800 00007ec638a47f00 0000000000000009
> Dec 25 06:52:52 romulus-server kernel: [28835.278766] GPR04: 00000043809d1148 0000000000000154 0000000000000000 0000000000001ae8
> Dec 25 06:52:52 romulus-server kernel: [28835.278766] GPR08: 0000004362347d00 0000000000000000 0000000000000000 0000000000000000
> Dec 25 06:52:52 romulus-server kernel: [28835.278766] GPR12: 0000000000000000 00007ec6348e0890 0000000000000000 ffffffffffffffff
> Dec 25 06:52:52 romulus-server kernel: [28835.278766] GPR16: 0000000000000000 000000436233f7a0 0000000000000001 0000000000000000
> Dec 25 06:52:52 romulus-server kernel: [28835.278766] GPR20: 00007fffe99c18ac 0000004362344f48 0000000000000004 00007fffe99c18b0
> Dec 25 06:52:52 romulus-server kernel: [28835.278766] GPR24: 0000000006000001 0000000000000000 0000000000000154 00000043809d1148
> Dec 25 06:52:52 romulus-server kernel: [28835.278766] GPR28: 0000000000000000 00007ec6348d9938 00000043809ceb00 000000000000000b
> Dec 25 06:52:52 romulus-server kernel: [28835.278992] NIP [00007ec638a194f4] 0x7ec638a194f4
> Dec 25 06:52:52 romulus-server kernel: [28835.279020] LR [0000000000000000] 0x0
> Dec 25 06:52:52 romulus-server kernel: [28835.279038] --- interrupt: 3000
> Dec 25 06:52:52 romulus-server kernel: [28835.279054] Instruction dump:
> Dec 25 06:52:52 romulus-server kernel: [28835.279072] f9210020 41820098 2e1cffff 3b200001 2c2a0000 41820088 41920010 894a0007
> Dec 25 06:52:52 romulus-server kernel: [28835.279110] 7c1c5000 40820078 815f0028 e97f00b8 <7ce9502a> 7c095214 886d0988 9b2d0988
> Dec 25 06:52:52 romulus-server kernel: [28835.279141] ---[ end trace fe7ee98d0b7beb6a ]---
Perhaps slab corruption, but the 'D' taint flag (TAINT_DIE) means the
kernel oopsed at least once before this. Probably best to look at that
one first.
^ permalink raw reply
* Re: [PATCH v4] powerpc/pseries: read the lpar name from the firmware
From: Nathan Lynch @ 2022-01-06 1:36 UTC (permalink / raw)
To: Tyrel Datwyler; +Cc: Laurent Dufour, linuxppc-dev, linux-kernel
In-Reply-To: <ac208963-d334-1f46-0db2-4a8d073b2963@linux.ibm.com>
Tyrel Datwyler <tyreld@linux.ibm.com> writes:
> On 1/5/22 3:19 PM, Nathan Lynch wrote:
>> Laurent Dufour <ldufour@linux.ibm.com> writes:
>>> On 07/12/2021, 18:11:09, Laurent Dufour wrote:
>>>> The LPAR name may be changed after the LPAR has been started in the HMC.
>>>> In that case lparstat command is not reporting the updated value because it
>>>> reads it from the device tree which is read at boot time.
>>>>
>>>> However this value could be read from RTAS.
>>>>
>>>> Adding this value in the /proc/powerpc/lparcfg output allows to read the
>>>> updated value.
>>>
>>> Do you consider taking that patch soon?
>>
>> This version prints an error on non-PowerVM guests the first time
>> lparcfg is read.
>
> I assume because QEMU doesn't implement the LPAR_NAME token for
> get_sysparm.
Correct.
>> And I still contend that having this function fall back to reporting the
>> partition name in the DT would provide a beneficial consistency in the
>> user-facing API, allowing programs to avoid hypervisor-specific branches
>> in their code.
>
> Agreed, if the get_sysparm fails just report the lpar-name from the device tree.
>
>> I don't understand the resistance I've encountered here.
>> The fallback I'm suggesting (a root node property lookup) is certainly
>> not more complex than the RTAS call sequence you've already implemented.
>>
>
> Is there benefit of adding a partition_name field/value pair to lparcfg? The
> lparstat utility can just as easily make the get_sysparm call via librtas.
> Further, rtas_filters allows this particular RTAS call from userspace.
The RTAS syscall is root-only, but we want the partition name (whether
supplied by RTAS or the device tree) to be available to unprivileged
programs.
^ permalink raw reply
* Re: [PATCH v4] powerpc/pseries: read the lpar name from the firmware
From: Tyrel Datwyler @ 2022-01-06 1:17 UTC (permalink / raw)
To: Nathan Lynch, Laurent Dufour, Michael Ellerman; +Cc: linuxppc-dev, linux-kernel
In-Reply-To: <8735m1ixd6.fsf@li-e15d104c-2135-11b2-a85c-d7ef17e56be6.ibm.com>
On 1/5/22 3:19 PM, Nathan Lynch wrote:
> Laurent Dufour <ldufour@linux.ibm.com> writes:
>> On 07/12/2021, 18:11:09, Laurent Dufour wrote:
>>> The LPAR name may be changed after the LPAR has been started in the HMC.
>>> In that case lparstat command is not reporting the updated value because it
>>> reads it from the device tree which is read at boot time.
>>>
>>> However this value could be read from RTAS.
>>>
>>> Adding this value in the /proc/powerpc/lparcfg output allows to read the
>>> updated value.
>>
>> Do you consider taking that patch soon?
>
> This version prints an error on non-PowerVM guests the first time
> lparcfg is read.
I assume because QEMU doesn't implement the LPAR_NAME token for get_sysparm.
>
> And I still contend that having this function fall back to reporting the
> partition name in the DT would provide a beneficial consistency in the
> user-facing API, allowing programs to avoid hypervisor-specific branches
> in their code.
Agreed, if the get_sysparm fails just report the lpar-name from the device tree.
I don't understand the resistance I've encountered here.
> The fallback I'm suggesting (a root node property lookup) is certainly
> not more complex than the RTAS call sequence you've already implemented.
>
Is there benefit of adding a partition_name field/value pair to lparcfg? The
lparstat utility can just as easily make the get_sysparm call via librtas.
Further, rtas_filters allows this particular RTAS call from userspace.
-Tyrel
^ permalink raw reply
* Re: [PATCH v4] powerpc/pseries: read the lpar name from the firmware
From: Nathan Lynch @ 2022-01-05 23:19 UTC (permalink / raw)
To: Laurent Dufour, Michael Ellerman; +Cc: linuxppc-dev, linux-kernel
In-Reply-To: <25527544-b0ac-596c-3876-560493b99f6b@linux.ibm.com>
Laurent Dufour <ldufour@linux.ibm.com> writes:
> On 07/12/2021, 18:11:09, Laurent Dufour wrote:
>> The LPAR name may be changed after the LPAR has been started in the HMC.
>> In that case lparstat command is not reporting the updated value because it
>> reads it from the device tree which is read at boot time.
>>
>> However this value could be read from RTAS.
>>
>> Adding this value in the /proc/powerpc/lparcfg output allows to read the
>> updated value.
>
> Do you consider taking that patch soon?
This version prints an error on non-PowerVM guests the first time
lparcfg is read.
And I still contend that having this function fall back to reporting the
partition name in the DT would provide a beneficial consistency in the
user-facing API, allowing programs to avoid hypervisor-specific branches
in their code. I don't understand the resistance I've encountered here.
The fallback I'm suggesting (a root node property lookup) is certainly
not more complex than the RTAS call sequence you've already implemented.
^ permalink raw reply
* [RFC PATCH v3 5/8] mm: page_isolation: check specified range for unmovable pages during isolation.
From: Zi Yan @ 2022-01-05 21:47 UTC (permalink / raw)
To: David Hildenbrand, linux-mm
Cc: Mel Gorman, Zi Yan, Robin Murphy, linux-kernel, iommu, Eric Ren,
virtualization, linuxppc-dev, Christoph Hellwig, Vlastimil Babka,
Marek Szyprowski
In-Reply-To: <20220105214756.91065-1-zi.yan@sent.com>
From: Zi Yan <ziy@nvidia.com>
Enable set_migratetype_isolate() to check specified sub-range for
unmovable pages during isolation. Page isolation is done
at max(MAX_ORDER_NR_PAEGS, pageblock_nr_pages) granularity, but not all
pages within that granularity are intended to be isolated. For example,
alloc_contig_range(), which uses page isolation, allows ranges without
alignment. This commit makes unmovable page check only look for
interesting pages, so that page isolation can succeed for any
non-overlapping ranges.
has_unmovable_pages() is moved to mm/page_isolation.c since it is only
used by page isolation.
Signed-off-by: Zi Yan <ziy@nvidia.com>
---
include/linux/page-isolation.h | 3 +-
mm/memory_hotplug.c | 12 ++-
mm/page_alloc.c | 122 +--------------------------
mm/page_isolation.c | 148 +++++++++++++++++++++++++++++++--
4 files changed, 153 insertions(+), 132 deletions(-)
diff --git a/include/linux/page-isolation.h b/include/linux/page-isolation.h
index 572458016331..a4d2687ed4e6 100644
--- a/include/linux/page-isolation.h
+++ b/include/linux/page-isolation.h
@@ -33,8 +33,6 @@ static inline bool is_migrate_isolate(int migratetype)
#define MEMORY_OFFLINE 0x1
#define REPORT_FAILURE 0x2
-struct page *has_unmovable_pages(struct zone *zone, struct page *page,
- int migratetype, int flags);
void set_pageblock_migratetype(struct page *page, int migratetype);
int move_freepages_block(struct zone *zone, struct page *page,
int migratetype, int *num_movable);
@@ -44,6 +42,7 @@ int move_freepages_block(struct zone *zone, struct page *page,
*/
int
start_isolate_page_range(unsigned long start_pfn, unsigned long end_pfn,
+ unsigned long isolate_start, unsigned long isolate_end,
unsigned migratetype, int flags);
/*
diff --git a/mm/memory_hotplug.c b/mm/memory_hotplug.c
index 0139b77c51d5..5db84c3fa882 100644
--- a/mm/memory_hotplug.c
+++ b/mm/memory_hotplug.c
@@ -1901,8 +1901,18 @@ int __ref offline_pages(unsigned long start_pfn, unsigned long nr_pages,
zone_pcp_disable(zone);
lru_cache_disable();
- /* set above range as isolated */
+ /*
+ * set above range as isolated
+ *
+ * start_pfn and end_pfn are the same as isolate_start and isolate_end,
+ * because start_pfn and end_pfn are already PAGES_PER_SECTION
+ * (>= MAX_ORDER_NR_PAGES) aligned; if start_pfn is
+ * pageblock_nr_pages aligned in memmap_on_memory case, there is no
+ * need to isolate pages before start_pfn, since they are used by
+ * memmap thus not user visible.
+ */
ret = start_isolate_page_range(start_pfn, end_pfn,
+ start_pfn, end_pfn,
MIGRATE_MOVABLE,
MEMORY_OFFLINE | REPORT_FAILURE);
if (ret) {
diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index e1c09ae54e31..faee7637740a 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -8864,125 +8864,6 @@ void *__init alloc_large_system_hash(const char *tablename,
return table;
}
-/*
- * This function checks whether pageblock includes unmovable pages or not.
- *
- * PageLRU check without isolation or lru_lock could race so that
- * MIGRATE_MOVABLE block might include unmovable pages. And __PageMovable
- * check without lock_page also may miss some movable non-lru pages at
- * race condition. So you can't expect this function should be exact.
- *
- * Returns a page without holding a reference. If the caller wants to
- * dereference that page (e.g., dumping), it has to make sure that it
- * cannot get removed (e.g., via memory unplug) concurrently.
- *
- */
-struct page *has_unmovable_pages(struct zone *zone, struct page *page,
- int migratetype, int flags)
-{
- unsigned long iter = 0;
- unsigned long pfn = page_to_pfn(page);
- unsigned long offset = pfn % pageblock_nr_pages;
-
- if (is_migrate_cma_page(page)) {
- /*
- * CMA allocations (alloc_contig_range) really need to mark
- * isolate CMA pageblocks even when they are not movable in fact
- * so consider them movable here.
- */
- if (is_migrate_cma(migratetype))
- return NULL;
-
- return page;
- }
-
- for (; iter < pageblock_nr_pages - offset; iter++) {
- page = pfn_to_page(pfn + iter);
-
- /*
- * Both, bootmem allocations and memory holes are marked
- * PG_reserved and are unmovable. We can even have unmovable
- * allocations inside ZONE_MOVABLE, for example when
- * specifying "movablecore".
- */
- if (PageReserved(page))
- return page;
-
- /*
- * If the zone is movable and we have ruled out all reserved
- * pages then it should be reasonably safe to assume the rest
- * is movable.
- */
- if (zone_idx(zone) == ZONE_MOVABLE)
- continue;
-
- /*
- * Hugepages are not in LRU lists, but they're movable.
- * THPs are on the LRU, but need to be counted as #small pages.
- * We need not scan over tail pages because we don't
- * handle each tail page individually in migration.
- */
- if (PageHuge(page) || PageTransCompound(page)) {
- struct page *head = compound_head(page);
- unsigned int skip_pages;
-
- if (PageHuge(page)) {
- if (!hugepage_migration_supported(page_hstate(head)))
- return page;
- } else if (!PageLRU(head) && !__PageMovable(head)) {
- return page;
- }
-
- skip_pages = compound_nr(head) - (page - head);
- iter += skip_pages - 1;
- continue;
- }
-
- /*
- * We can't use page_count without pin a page
- * because another CPU can free compound page.
- * This check already skips compound tails of THP
- * because their page->_refcount is zero at all time.
- */
- if (!page_ref_count(page)) {
- if (PageBuddy(page))
- iter += (1 << buddy_order(page)) - 1;
- continue;
- }
-
- /*
- * The HWPoisoned page may be not in buddy system, and
- * page_count() is not 0.
- */
- if ((flags & MEMORY_OFFLINE) && PageHWPoison(page))
- continue;
-
- /*
- * We treat all PageOffline() pages as movable when offlining
- * to give drivers a chance to decrement their reference count
- * in MEM_GOING_OFFLINE in order to indicate that these pages
- * can be offlined as there are no direct references anymore.
- * For actually unmovable PageOffline() where the driver does
- * not support this, we will fail later when trying to actually
- * move these pages that still have a reference count > 0.
- * (false negatives in this function only)
- */
- if ((flags & MEMORY_OFFLINE) && PageOffline(page))
- continue;
-
- if (__PageMovable(page) || PageLRU(page))
- continue;
-
- /*
- * If there are RECLAIMABLE pages, we need to check
- * it. But now, memory offline itself doesn't call
- * shrink_node_slabs() and it still to be fixed.
- */
- return page;
- }
- return NULL;
-}
-
#ifdef CONFIG_CONTIG_ALLOC
static unsigned long pfn_max_align_down(unsigned long pfn)
{
@@ -9226,7 +9107,8 @@ int alloc_contig_range(unsigned long start, unsigned long end,
* put back to page allocator so that buddy can use them.
*/
- ret = start_isolate_page_range(isolate_start, isolate_end, migratetype, 0);
+ ret = start_isolate_page_range(start, end, isolate_start, isolate_end,
+ migratetype, 0);
if (ret)
goto done;
diff --git a/mm/page_isolation.c b/mm/page_isolation.c
index 6a0ddda6b3c5..7a7991460eb9 100644
--- a/mm/page_isolation.c
+++ b/mm/page_isolation.c
@@ -15,12 +15,143 @@
#define CREATE_TRACE_POINTS
#include <trace/events/page_isolation.h>
-static int set_migratetype_isolate(struct page *page, int migratetype, int isol_flags)
+/*
+ * This function checks whether pageblock within [start_pfn, end_pfn) includes
+ * unmovable pages or not.
+ *
+ * PageLRU check without isolation or lru_lock could race so that
+ * MIGRATE_MOVABLE block might include unmovable pages. And __PageMovable
+ * check without lock_page also may miss some movable non-lru pages at
+ * race condition. So you can't expect this function should be exact.
+ *
+ * Returns a page without holding a reference. If the caller wants to
+ * dereference that page (e.g., dumping), it has to make sure that it
+ * cannot get removed (e.g., via memory unplug) concurrently.
+ *
+ */
+static struct page *has_unmovable_pages(struct zone *zone, struct page *page,
+ int migratetype, int flags,
+ unsigned long start_pfn, unsigned long end_pfn)
+{
+ unsigned long first_pfn = max(page_to_pfn(page), start_pfn);
+ unsigned long pfn = first_pfn;
+ unsigned long last_pfn = min(ALIGN(pfn + 1, pageblock_nr_pages), end_pfn);
+
+ page = pfn_to_page(pfn);
+
+ if (is_migrate_cma_page(page)) {
+ /*
+ * CMA allocations (alloc_contig_range) really need to mark
+ * isolate CMA pageblocks even when they are not movable in fact
+ * so consider them movable here.
+ */
+ if (is_migrate_cma(migratetype))
+ return NULL;
+
+ return page;
+ }
+
+ for (pfn = first_pfn; pfn < last_pfn; pfn++) {
+ page = pfn_to_page(pfn);
+
+ /*
+ * Both, bootmem allocations and memory holes are marked
+ * PG_reserved and are unmovable. We can even have unmovable
+ * allocations inside ZONE_MOVABLE, for example when
+ * specifying "movablecore".
+ */
+ if (PageReserved(page))
+ return page;
+
+ /*
+ * If the zone is movable and we have ruled out all reserved
+ * pages then it should be reasonably safe to assume the rest
+ * is movable.
+ */
+ if (zone_idx(zone) == ZONE_MOVABLE)
+ continue;
+
+ /*
+ * Hugepages are not in LRU lists, but they're movable.
+ * THPs are on the LRU, but need to be counted as #small pages.
+ * We need not scan over tail pages because we don't
+ * handle each tail page individually in migration.
+ */
+ if (PageHuge(page) || PageTransCompound(page)) {
+ struct page *head = compound_head(page);
+ unsigned int skip_pages;
+
+ if (PageHuge(page)) {
+ if (!hugepage_migration_supported(page_hstate(head)))
+ return page;
+ } else if (!PageLRU(head) && !__PageMovable(head)) {
+ return page;
+ }
+
+ skip_pages = compound_nr(head) - (page - head);
+ pfn += skip_pages - 1;
+ continue;
+ }
+
+ /*
+ * We can't use page_count without pin a page
+ * because another CPU can free compound page.
+ * This check already skips compound tails of THP
+ * because their page->_refcount is zero at all time.
+ */
+ if (!page_ref_count(page)) {
+ if (PageBuddy(page))
+ pfn += (1 << buddy_order(page)) - 1;
+ continue;
+ }
+
+ /*
+ * The HWPoisoned page may be not in buddy system, and
+ * page_count() is not 0.
+ */
+ if ((flags & MEMORY_OFFLINE) && PageHWPoison(page))
+ continue;
+
+ /*
+ * We treat all PageOffline() pages as movable when offlining
+ * to give drivers a chance to decrement their reference count
+ * in MEM_GOING_OFFLINE in order to indicate that these pages
+ * can be offlined as there are no direct references anymore.
+ * For actually unmovable PageOffline() where the driver does
+ * not support this, we will fail later when trying to actually
+ * move these pages that still have a reference count > 0.
+ * (false negatives in this function only)
+ */
+ if ((flags & MEMORY_OFFLINE) && PageOffline(page))
+ continue;
+
+ if (__PageMovable(page) || PageLRU(page))
+ continue;
+
+ /*
+ * If there are RECLAIMABLE pages, we need to check
+ * it. But now, memory offline itself doesn't call
+ * shrink_node_slabs() and it still to be fixed.
+ */
+ return page;
+ }
+ return NULL;
+}
+
+/*
+ * This function set pageblock migratetype to isolate if no unmovable page is
+ * present in [start_pfn, end_pfn). The pageblock must be within
+ * [start_pfn, end_pfn).
+ */
+static int set_migratetype_isolate(struct page *page, int migratetype, int isol_flags,
+ unsigned long start_pfn, unsigned long end_pfn)
{
struct zone *zone = page_zone(page);
struct page *unmovable;
unsigned long flags;
+ VM_BUG_ON(page_to_pfn(page) < start_pfn || page_to_pfn(page) >= end_pfn);
+
spin_lock_irqsave(&zone->lock, flags);
/*
@@ -37,7 +168,7 @@ static int set_migratetype_isolate(struct page *page, int migratetype, int isol_
* FIXME: Now, memory hotplug doesn't call shrink_slab() by itself.
* We just check MOVABLE pages.
*/
- unmovable = has_unmovable_pages(zone, page, migratetype, isol_flags);
+ unmovable = has_unmovable_pages(zone, page, migratetype, isol_flags, start_pfn, end_pfn);
if (!unmovable) {
unsigned long nr_pages;
int mt = get_pageblock_migratetype(page);
@@ -185,20 +316,19 @@ __first_valid_page(unsigned long pfn, unsigned long nr_pages)
* Return: 0 on success and -EBUSY if any part of range cannot be isolated.
*/
int start_isolate_page_range(unsigned long start_pfn, unsigned long end_pfn,
+ unsigned long isolate_start, unsigned long isolate_end,
unsigned migratetype, int flags)
{
unsigned long pfn;
struct page *page;
- BUG_ON(!IS_ALIGNED(start_pfn, pageblock_nr_pages));
- BUG_ON(!IS_ALIGNED(end_pfn, pageblock_nr_pages));
-
- for (pfn = start_pfn;
- pfn < end_pfn;
+ for (pfn = isolate_start;
+ pfn < isolate_end;
pfn += pageblock_nr_pages) {
page = __first_valid_page(pfn, pageblock_nr_pages);
- if (page && set_migratetype_isolate(page, migratetype, flags)) {
- undo_isolate_page_range(start_pfn, pfn, migratetype);
+ if (page && set_migratetype_isolate(page, migratetype, flags,
+ start_pfn, end_pfn)) {
+ undo_isolate_page_range(isolate_start, pfn, migratetype);
return -EBUSY;
}
}
--
2.34.1
^ permalink raw reply related
* [RFC PATCH v3 7/8] drivers: virtio_mem: use pageblock size as the minimum virtio_mem size.
From: Zi Yan @ 2022-01-05 21:47 UTC (permalink / raw)
To: David Hildenbrand, linux-mm
Cc: Mel Gorman, Zi Yan, Robin Murphy, linux-kernel, iommu, Eric Ren,
virtualization, linuxppc-dev, Christoph Hellwig, Vlastimil Babka,
Marek Szyprowski
In-Reply-To: <20220105214756.91065-1-zi.yan@sent.com>
From: Zi Yan <ziy@nvidia.com>
alloc_contig_range() now only needs to be aligned to pageblock_order,
drop virtio_mem size requirement that it needs to be the max of
pageblock_order and MAX_ORDER.
Signed-off-by: Zi Yan <ziy@nvidia.com>
---
drivers/virtio/virtio_mem.c | 3 +--
1 file changed, 1 insertion(+), 2 deletions(-)
diff --git a/drivers/virtio/virtio_mem.c b/drivers/virtio/virtio_mem.c
index a6a78685cfbe..2664dc16d0f9 100644
--- a/drivers/virtio/virtio_mem.c
+++ b/drivers/virtio/virtio_mem.c
@@ -2481,8 +2481,7 @@ static int virtio_mem_init_hotplug(struct virtio_mem *vm)
* - Is required for now for alloc_contig_range() to work reliably -
* it doesn't properly handle smaller granularity on ZONE_NORMAL.
*/
- sb_size = max_t(uint64_t, MAX_ORDER_NR_PAGES,
- pageblock_nr_pages) * PAGE_SIZE;
+ sb_size = pageblock_nr_pages * PAGE_SIZE;
sb_size = max_t(uint64_t, vm->device_block_size, sb_size);
if (sb_size < memory_block_size_bytes() && !force_bbm) {
--
2.34.1
^ permalink raw reply related
* [RFC PATCH v3 8/8] arch: powerpc: adjust fadump alignment to be pageblock aligned.
From: Zi Yan @ 2022-01-05 21:47 UTC (permalink / raw)
To: David Hildenbrand, linux-mm
Cc: Mel Gorman, Zi Yan, Robin Murphy, linux-kernel, iommu, Eric Ren,
virtualization, linuxppc-dev, Christoph Hellwig, Vlastimil Babka,
Marek Szyprowski
In-Reply-To: <20220105214756.91065-1-zi.yan@sent.com>
From: Zi Yan <ziy@nvidia.com>
CMA only requires pageblock alignment now. Change CMA alignment in
fadump too.
Signed-off-by: Zi Yan <ziy@nvidia.com>
---
arch/powerpc/include/asm/fadump-internal.h | 4 +---
1 file changed, 1 insertion(+), 3 deletions(-)
diff --git a/arch/powerpc/include/asm/fadump-internal.h b/arch/powerpc/include/asm/fadump-internal.h
index 52189928ec08..fbfca85b4200 100644
--- a/arch/powerpc/include/asm/fadump-internal.h
+++ b/arch/powerpc/include/asm/fadump-internal.h
@@ -20,9 +20,7 @@
#define memblock_num_regions(memblock_type) (memblock.memblock_type.cnt)
/* Alignment per CMA requirement. */
-#define FADUMP_CMA_ALIGNMENT (PAGE_SIZE << \
- max_t(unsigned long, MAX_ORDER - 1, \
- pageblock_order))
+#define FADUMP_CMA_ALIGNMENT (PAGE_SIZE << pageblock_order)
/* FAD commands */
#define FADUMP_REGISTER 1
--
2.34.1
^ permalink raw reply related
* [RFC PATCH v3 6/8] mm: cma: use pageblock_order as the single alignment
From: Zi Yan @ 2022-01-05 21:47 UTC (permalink / raw)
To: David Hildenbrand, linux-mm
Cc: Mel Gorman, Zi Yan, Robin Murphy, linux-kernel, iommu, Eric Ren,
virtualization, linuxppc-dev, Christoph Hellwig, Vlastimil Babka,
Marek Szyprowski
In-Reply-To: <20220105214756.91065-1-zi.yan@sent.com>
From: Zi Yan <ziy@nvidia.com>
Now alloc_contig_range() works at pageblock granularity. Change CMA
allocation, which uses alloc_contig_range(), to use pageblock_order
alignment.
Signed-off-by: Zi Yan <ziy@nvidia.com>
---
include/linux/mmzone.h | 5 +----
kernel/dma/contiguous.c | 2 +-
mm/cma.c | 6 ++----
mm/page_alloc.c | 6 +++---
4 files changed, 7 insertions(+), 12 deletions(-)
diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h
index 0aa549653e4e..d28a02a893d6 100644
--- a/include/linux/mmzone.h
+++ b/include/linux/mmzone.h
@@ -54,10 +54,7 @@ enum migratetype {
*
* The way to use it is to change migratetype of a range of
* pageblocks to MIGRATE_CMA which can be done by
- * __free_pageblock_cma() function. What is important though
- * is that a range of pageblocks must be aligned to
- * MAX_ORDER_NR_PAGES should biggest page be bigger than
- * a single pageblock.
+ * __free_pageblock_cma() function.
*/
MIGRATE_CMA,
#endif
diff --git a/kernel/dma/contiguous.c b/kernel/dma/contiguous.c
index 3d63d91cba5c..ac35b14b0786 100644
--- a/kernel/dma/contiguous.c
+++ b/kernel/dma/contiguous.c
@@ -399,7 +399,7 @@ static const struct reserved_mem_ops rmem_cma_ops = {
static int __init rmem_cma_setup(struct reserved_mem *rmem)
{
- phys_addr_t align = PAGE_SIZE << max(MAX_ORDER - 1, pageblock_order);
+ phys_addr_t align = PAGE_SIZE << pageblock_order;
phys_addr_t mask = align - 1;
unsigned long node = rmem->fdt_node;
bool default_cma = of_get_flat_dt_prop(node, "linux,cma-default", NULL);
diff --git a/mm/cma.c b/mm/cma.c
index bc9ca8f3c487..d171158bd418 100644
--- a/mm/cma.c
+++ b/mm/cma.c
@@ -180,8 +180,7 @@ int __init cma_init_reserved_mem(phys_addr_t base, phys_addr_t size,
return -EINVAL;
/* ensure minimal alignment required by mm core */
- alignment = PAGE_SIZE <<
- max_t(unsigned long, MAX_ORDER - 1, pageblock_order);
+ alignment = PAGE_SIZE << pageblock_order;
/* alignment should be aligned with order_per_bit */
if (!IS_ALIGNED(alignment >> PAGE_SHIFT, 1 << order_per_bit))
@@ -268,8 +267,7 @@ int __init cma_declare_contiguous_nid(phys_addr_t base,
* migratetype page by page allocator's buddy algorithm. In the case,
* you couldn't get a contiguous memory, which is not what we want.
*/
- alignment = max(alignment, (phys_addr_t)PAGE_SIZE <<
- max_t(unsigned long, MAX_ORDER - 1, pageblock_order));
+ alignment = max(alignment, (phys_addr_t)PAGE_SIZE << pageblock_order);
if (fixed && base & (alignment - 1)) {
ret = -EINVAL;
pr_err("Region at %pa must be aligned to %pa bytes\n",
diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index faee7637740a..63d76f436ed1 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -9013,8 +9013,8 @@ static inline void split_free_page_into_pageblocks(struct page *free_page,
* be either of the two.
* @gfp_mask: GFP mask to use during compaction
*
- * The PFN range does not have to be pageblock or MAX_ORDER_NR_PAGES
- * aligned. The PFN range must belong to a single zone.
+ * The PFN range does not have to be pageblock aligned. The PFN range must
+ * belong to a single zone.
*
* The first thing this routine does is attempt to MIGRATE_ISOLATE all
* pageblocks in the range. Once isolated, the pageblocks should not
@@ -9130,7 +9130,7 @@ int alloc_contig_range(unsigned long start, unsigned long end,
ret = 0;
/*
- * Pages from [start, end) are within a MAX_ORDER_NR_PAGES
+ * Pages from [start, end) are within a pageblock_nr_pages
* aligned blocks that are marked as MIGRATE_ISOLATE. What's
* more, all pages in [start, end) are free in page allocator.
* What we are going to do is to allocate all pages from
--
2.34.1
^ permalink raw reply related
* [RFC PATCH v3 2/8] mm: compaction: handle non-lru compound pages properly in isolate_migratepages_block().
From: Zi Yan @ 2022-01-05 21:47 UTC (permalink / raw)
To: David Hildenbrand, linux-mm
Cc: Mel Gorman, Zi Yan, Robin Murphy, linux-kernel, iommu, Eric Ren,
virtualization, linuxppc-dev, Christoph Hellwig, Vlastimil Babka,
Marek Szyprowski
In-Reply-To: <20220105214756.91065-1-zi.yan@sent.com>
From: Zi Yan <ziy@nvidia.com>
In isolate_migratepages_block(), a !PageLRU tail page can be encountered
when the page is larger than a pageblock. Use compound head page for the
checks inside and skip the entire compound page when isolation succeeds.
Signed-off-by: Zi Yan <ziy@nvidia.com>
---
mm/compaction.c | 10 +++++++---
1 file changed, 7 insertions(+), 3 deletions(-)
diff --git a/mm/compaction.c b/mm/compaction.c
index b4e94cda3019..ad9053fbbe06 100644
--- a/mm/compaction.c
+++ b/mm/compaction.c
@@ -979,19 +979,23 @@ isolate_migratepages_block(struct compact_control *cc, unsigned long low_pfn,
* Skip any other type of page
*/
if (!PageLRU(page)) {
+ struct page *head = compound_head(page);
/*
* __PageMovable can return false positive so we need
* to verify it under page_lock.
*/
- if (unlikely(__PageMovable(page)) &&
- !PageIsolated(page)) {
+ if (unlikely(__PageMovable(head)) &&
+ !PageIsolated(head)) {
if (locked) {
unlock_page_lruvec_irqrestore(locked, flags);
locked = NULL;
}
- if (!isolate_movable_page(page, isolate_mode))
+ if (!isolate_movable_page(head, isolate_mode)) {
+ low_pfn += (1 << compound_order(head)) - 1 - (page - head);
+ page = head;
goto isolate_success;
+ }
}
goto isolate_fail;
--
2.34.1
^ permalink raw reply related
* [RFC PATCH v3 1/8] mm: page_alloc: avoid merging non-fallbackable pageblocks with others.
From: Zi Yan @ 2022-01-05 21:47 UTC (permalink / raw)
To: David Hildenbrand, linux-mm
Cc: Mel Gorman, Zi Yan, Robin Murphy, linux-kernel, iommu, Eric Ren,
virtualization, linuxppc-dev, Christoph Hellwig, Vlastimil Babka,
Marek Szyprowski
In-Reply-To: <20220105214756.91065-1-zi.yan@sent.com>
From: Zi Yan <ziy@nvidia.com>
This is done in addition to MIGRATE_ISOLATE pageblock merge avoidance.
It prepares for the upcoming removal of the MAX_ORDER-1 alignment
requirement for CMA and alloc_contig_range().
MIGRARTE_HIGHATOMIC should not merge with other migratetypes like
MIGRATE_ISOLATE and MIGRARTE_CMA[1], so this commit prevents that too.
Also add MIGRARTE_HIGHATOMIC to fallbacks array for completeness.
[1] https://lore.kernel.org/linux-mm/20211130100853.GP3366@techsingularity.net/
Signed-off-by: Zi Yan <ziy@nvidia.com>
---
include/linux/mmzone.h | 6 ++++++
mm/page_alloc.c | 28 ++++++++++++++++++----------
2 files changed, 24 insertions(+), 10 deletions(-)
diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h
index aed44e9b5d89..0aa549653e4e 100644
--- a/include/linux/mmzone.h
+++ b/include/linux/mmzone.h
@@ -83,6 +83,12 @@ static inline bool is_migrate_movable(int mt)
return is_migrate_cma(mt) || mt == MIGRATE_MOVABLE;
}
+/* See fallbacks[MIGRATE_TYPES][3] in page_alloc.c */
+static inline bool migratetype_has_fallback(int mt)
+{
+ return mt < MIGRATE_PCPTYPES;
+}
+
#define for_each_migratetype_order(order, type) \
for (order = 0; order < MAX_ORDER; order++) \
for (type = 0; type < MIGRATE_TYPES; type++)
diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index 8dd6399bafb5..5193c953dbf8 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -1042,6 +1042,12 @@ buddy_merge_likely(unsigned long pfn, unsigned long buddy_pfn,
return page_is_buddy(higher_page, higher_buddy, order + 1);
}
+static inline bool has_non_fallback_pageblock(struct zone *zone)
+{
+ return has_isolate_pageblock(zone) || zone_cma_pages(zone) != 0 ||
+ zone->nr_reserved_highatomic != 0;
+}
+
/*
* Freeing function for a buddy system allocator.
*
@@ -1117,14 +1123,15 @@ static inline void __free_one_page(struct page *page,
}
if (order < MAX_ORDER - 1) {
/* If we are here, it means order is >= pageblock_order.
- * We want to prevent merge between freepages on isolate
- * pageblock and normal pageblock. Without this, pageblock
- * isolation could cause incorrect freepage or CMA accounting.
+ * We want to prevent merge between freepages on pageblock
+ * without fallbacks and normal pageblock. Without this,
+ * pageblock isolation could cause incorrect freepage or CMA
+ * accounting or HIGHATOMIC accounting.
*
* We don't want to hit this code for the more frequent
* low-order merging.
*/
- if (unlikely(has_isolate_pageblock(zone))) {
+ if (unlikely(has_non_fallback_pageblock(zone))) {
int buddy_mt;
buddy_pfn = __find_buddy_pfn(pfn, order);
@@ -1132,8 +1139,8 @@ static inline void __free_one_page(struct page *page,
buddy_mt = get_pageblock_migratetype(buddy);
if (migratetype != buddy_mt
- && (is_migrate_isolate(migratetype) ||
- is_migrate_isolate(buddy_mt)))
+ && (!migratetype_has_fallback(migratetype) ||
+ !migratetype_has_fallback(buddy_mt)))
goto done_merging;
}
max_order = order + 1;
@@ -2484,6 +2491,7 @@ static int fallbacks[MIGRATE_TYPES][3] = {
[MIGRATE_UNMOVABLE] = { MIGRATE_RECLAIMABLE, MIGRATE_MOVABLE, MIGRATE_TYPES },
[MIGRATE_MOVABLE] = { MIGRATE_RECLAIMABLE, MIGRATE_UNMOVABLE, MIGRATE_TYPES },
[MIGRATE_RECLAIMABLE] = { MIGRATE_UNMOVABLE, MIGRATE_MOVABLE, MIGRATE_TYPES },
+ [MIGRATE_HIGHATOMIC] = { MIGRATE_TYPES }, /* Never used */
#ifdef CONFIG_CMA
[MIGRATE_CMA] = { MIGRATE_TYPES }, /* Never used */
#endif
@@ -2795,8 +2803,8 @@ static void reserve_highatomic_pageblock(struct page *page, struct zone *zone,
/* Yoink! */
mt = get_pageblock_migratetype(page);
- if (!is_migrate_highatomic(mt) && !is_migrate_isolate(mt)
- && !is_migrate_cma(mt)) {
+ /* Only reserve normal pageblock */
+ if (migratetype_has_fallback(mt)) {
zone->nr_reserved_highatomic += pageblock_nr_pages;
set_pageblock_migratetype(page, MIGRATE_HIGHATOMIC);
move_freepages_block(zone, page, MIGRATE_HIGHATOMIC, NULL);
@@ -3545,8 +3553,8 @@ int __isolate_free_page(struct page *page, unsigned int order)
struct page *endpage = page + (1 << order) - 1;
for (; page < endpage; page += pageblock_nr_pages) {
int mt = get_pageblock_migratetype(page);
- if (!is_migrate_isolate(mt) && !is_migrate_cma(mt)
- && !is_migrate_highatomic(mt))
+ /* Only change normal pageblock */
+ if (migratetype_has_fallback(mt))
set_pageblock_migratetype(page,
MIGRATE_MOVABLE);
}
--
2.34.1
^ permalink raw reply related
* [RFC PATCH v3 3/8] mm: migrate: allocate the right size of non hugetlb or THP compound pages.
From: Zi Yan @ 2022-01-05 21:47 UTC (permalink / raw)
To: David Hildenbrand, linux-mm
Cc: Mel Gorman, Zi Yan, Robin Murphy, linux-kernel, iommu, Eric Ren,
virtualization, linuxppc-dev, Christoph Hellwig, Vlastimil Babka,
Marek Szyprowski
In-Reply-To: <20220105214756.91065-1-zi.yan@sent.com>
From: Zi Yan <ziy@nvidia.com>
alloc_migration_target() is used by alloc_contig_range() and non-LRU
movable compound pages can be migrated. Current code does not allocate the
right page size for such pages. Check THP precisely using
is_transparent_huge() and add allocation support for non-LRU compound
pages.
Signed-off-by: Zi Yan <ziy@nvidia.com>
---
mm/migrate.c | 11 +++++++----
1 file changed, 7 insertions(+), 4 deletions(-)
diff --git a/mm/migrate.c b/mm/migrate.c
index c7da064b4781..b1851ffb8576 100644
--- a/mm/migrate.c
+++ b/mm/migrate.c
@@ -1546,9 +1546,7 @@ struct page *alloc_migration_target(struct page *page, unsigned long private)
gfp_mask = htlb_modify_alloc_mask(h, gfp_mask);
return alloc_huge_page_nodemask(h, nid, mtc->nmask, gfp_mask);
- }
-
- if (PageTransHuge(page)) {
+ } else if (is_transparent_hugepage(page)) {
/*
* clear __GFP_RECLAIM to make the migration callback
* consistent with regular THP allocations.
@@ -1556,14 +1554,19 @@ struct page *alloc_migration_target(struct page *page, unsigned long private)
gfp_mask &= ~__GFP_RECLAIM;
gfp_mask |= GFP_TRANSHUGE;
order = HPAGE_PMD_ORDER;
+ } else if (PageCompound(page)) {
+ /* for non-LRU movable compound pages */
+ gfp_mask |= __GFP_COMP;
+ order = compound_order(page);
}
+
zidx = zone_idx(page_zone(page));
if (is_highmem_idx(zidx) || zidx == ZONE_MOVABLE)
gfp_mask |= __GFP_HIGHMEM;
new_page = __alloc_pages(gfp_mask, order, nid, mtc->nmask);
- if (new_page && PageTransHuge(new_page))
+ if (new_page && is_transparent_hugepage(page))
prep_transhuge_page(new_page);
return new_page;
--
2.34.1
^ permalink raw reply related
* [RFC PATCH v3 0/8] Use pageblock_order for cma and alloc_contig_range alignment.
From: Zi Yan @ 2022-01-05 21:47 UTC (permalink / raw)
To: David Hildenbrand, linux-mm
Cc: Mel Gorman, Zi Yan, Robin Murphy, linux-kernel, iommu, Eric Ren,
virtualization, linuxppc-dev, Christoph Hellwig, Vlastimil Babka,
Marek Szyprowski
From: Zi Yan <ziy@nvidia.com>
Hi all,
This patchset tries to remove the MAX_ORDER - 1 alignment requirement for CMA
and alloc_contig_range(). It prepares for my upcoming changes to make MAX_ORDER
adjustable at boot time[1]. It is on top of mmotm-2021-12-29-20-07.
The MAX_ORDER - 1 alignment requirement comes from that alloc_contig_range()
isolates pageblocks to remove free memory from buddy allocator but isolating
only a subset of pageblocks within a page spanning across multiple pageblocks
causes free page accounting issues. Isolated page might not be put into the
right free list, since the code assumes the migratetype of the first pageblock
as the whole free page migratetype. This is based on the discussion at [2].
To remove the requirement, this patchset:
1. still isolates pageblocks at MAX_ORDER - 1 granularity;
2. but saves the pageblock migratetypes outside the specified range of
alloc_contig_range() and restores them after all pages within the range
become free after __alloc_contig_migrate_range();
3. only checks unmovable pages within the range instead of MAX_ORDER - 1 aligned
range during isolation to avoid alloc_contig_range() failure when pageblocks
within a MAX_ORDER - 1 aligned range are allocated separately.
3. splits free pages spanning multiple pageblocks at the beginning and the end
of the range and puts the split pages to the right migratetype free lists
based on the pageblock migratetypes;
4. returns pages not in the range as it did before.
Isolation needs to be done at MAX_ORDER - 1 granularity, because otherwise
either 1) it is needed to detect to-be-isolated page size (free, PageHuge, THP,
or other PageCompound) to make sure all pageblocks belonging to a single page
are isolated together and later restore pageblock migratetypes outside the
range, or 2) assuming isolation happens at pageblock granularity, a free page
with multi-migratetype pageblocks can seen in free page path and needs
to be split and freed at pageblock granularity.
One optimization might come later:
1. make MIGRATE_ISOLATE a separate bit to avoid saving and restoring existing
migratetypes before and after isolation respectively.
Feel free to give comments and suggestions. Thanks.
[1] https://lore.kernel.org/linux-mm/20210805190253.2795604-1-zi.yan@sent.com/
[2] https://lore.kernel.org/linux-mm/d19fb078-cb9b-f60f-e310-fdeea1b947d2@redhat.com/
Zi Yan (8):
mm: page_alloc: avoid merging non-fallbackable pageblocks with others.
mm: compaction: handle non-lru compound pages properly in
isolate_migratepages_block().
mm: migrate: allocate the right size of non hugetlb or THP compound
pages.
mm: make alloc_contig_range work at pageblock granularity
mm: page_isolation: check specified range for unmovable pages during
isolation.
mm: cma: use pageblock_order as the single alignment
drivers: virtio_mem: use pageblock size as the minimum virtio_mem
size.
arch: powerpc: adjust fadump alignment to be pageblock aligned.
arch/powerpc/include/asm/fadump-internal.h | 4 +-
drivers/virtio/virtio_mem.c | 3 +-
include/linux/mmzone.h | 11 +-
include/linux/page-isolation.h | 3 +-
kernel/dma/contiguous.c | 2 +-
mm/cma.c | 6 +-
mm/compaction.c | 10 +-
mm/memory_hotplug.c | 12 +-
mm/migrate.c | 11 +-
mm/page_alloc.c | 328 +++++++++++----------
mm/page_isolation.c | 148 +++++++++-
11 files changed, 353 insertions(+), 185 deletions(-)
--
2.34.1
^ permalink raw reply
* [RFC PATCH v3 4/8] mm: make alloc_contig_range work at pageblock granularity
From: Zi Yan @ 2022-01-05 21:47 UTC (permalink / raw)
To: David Hildenbrand, linux-mm
Cc: Mel Gorman, Zi Yan, Robin Murphy, linux-kernel, iommu, Eric Ren,
virtualization, linuxppc-dev, Christoph Hellwig, Vlastimil Babka,
Marek Szyprowski
In-Reply-To: <20220105214756.91065-1-zi.yan@sent.com>
From: Zi Yan <ziy@nvidia.com>
alloc_contig_range() worked at MAX_ORDER-1 granularity to avoid merging
pageblocks with different migratetypes. It might unnecessarily convert
extra pageblocks at the beginning and at the end of the range. Change
alloc_contig_range() to work at pageblock granularity.
It is done by restoring pageblock types and split >pageblock_order free
pages after isolating at MAX_ORDER-1 granularity and migrating pages
away at pageblock granularity. The reason for this process is that
during isolation, some pages, either free or in-use, might have >pageblock
sizes and isolating part of them can cause free accounting issues.
Restoring the migratetypes of the pageblocks not in the interesting
range later is much easier.
Signed-off-by: Zi Yan <ziy@nvidia.com>
---
mm/page_alloc.c | 174 ++++++++++++++++++++++++++++++++++++++++++------
1 file changed, 154 insertions(+), 20 deletions(-)
diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index 5193c953dbf8..e1c09ae54e31 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -8986,8 +8986,8 @@ struct page *has_unmovable_pages(struct zone *zone, struct page *page,
#ifdef CONFIG_CONTIG_ALLOC
static unsigned long pfn_max_align_down(unsigned long pfn)
{
- return pfn & ~(max_t(unsigned long, MAX_ORDER_NR_PAGES,
- pageblock_nr_pages) - 1);
+ return ALIGN_DOWN(pfn, max_t(unsigned long, MAX_ORDER_NR_PAGES,
+ pageblock_nr_pages));
}
static unsigned long pfn_max_align_up(unsigned long pfn)
@@ -9076,6 +9076,52 @@ static int __alloc_contig_migrate_range(struct compact_control *cc,
return 0;
}
+static inline int save_migratetypes(unsigned char *migratetypes,
+ unsigned long start_pfn, unsigned long end_pfn)
+{
+ unsigned long pfn = start_pfn;
+ int num = 0;
+
+ while (pfn < end_pfn) {
+ migratetypes[num] = get_pageblock_migratetype(pfn_to_page(pfn));
+ num++;
+ pfn += pageblock_nr_pages;
+ }
+ return num;
+}
+
+static inline int restore_migratetypes(unsigned char *migratetypes,
+ unsigned long start_pfn, unsigned long end_pfn)
+{
+ unsigned long pfn = start_pfn;
+ int num = 0;
+
+ while (pfn < end_pfn) {
+ set_pageblock_migratetype(pfn_to_page(pfn), migratetypes[num]);
+ num++;
+ pfn += pageblock_nr_pages;
+ }
+ return num;
+}
+
+static inline void split_free_page_into_pageblocks(struct page *free_page,
+ int order, struct zone *zone)
+{
+ unsigned long pfn;
+
+ spin_lock(&zone->lock);
+ del_page_from_free_list(free_page, zone, order);
+ for (pfn = page_to_pfn(free_page);
+ pfn < page_to_pfn(free_page) + (1UL << order);
+ pfn += pageblock_nr_pages) {
+ int mt = get_pfnblock_migratetype(pfn_to_page(pfn), pfn);
+
+ __free_one_page(pfn_to_page(pfn), pfn, zone, pageblock_order,
+ mt, FPI_NONE);
+ }
+ spin_unlock(&zone->lock);
+}
+
/**
* alloc_contig_range() -- tries to allocate given range of pages
* @start: start PFN to allocate
@@ -9101,8 +9147,15 @@ int alloc_contig_range(unsigned long start, unsigned long end,
unsigned migratetype, gfp_t gfp_mask)
{
unsigned long outer_start, outer_end;
+ unsigned long isolate_start = pfn_max_align_down(start);
+ unsigned long isolate_end = pfn_max_align_up(end);
+ unsigned long alloc_start = ALIGN_DOWN(start, pageblock_nr_pages);
+ unsigned long alloc_end = ALIGN(end, pageblock_nr_pages);
+ unsigned long num_pageblock_to_save;
unsigned int order;
int ret = 0;
+ unsigned char *saved_mt;
+ int num;
struct compact_control cc = {
.nr_migratepages = 0,
@@ -9116,11 +9169,30 @@ int alloc_contig_range(unsigned long start, unsigned long end,
};
INIT_LIST_HEAD(&cc.migratepages);
+ /*
+ * TODO: make MIGRATE_ISOLATE a standalone bit to avoid overwriting
+ * the exiting migratetype. Then, we will not need the save and restore
+ * process here.
+ */
+
+ /* Save the migratepages of the pageblocks before start and after end */
+ num_pageblock_to_save = (alloc_start - isolate_start) / pageblock_nr_pages
+ + (isolate_end - alloc_end) / pageblock_nr_pages;
+ saved_mt =
+ kmalloc_array(num_pageblock_to_save,
+ sizeof(unsigned char), GFP_KERNEL);
+ if (!saved_mt)
+ return -ENOMEM;
+
+ num = save_migratetypes(saved_mt, isolate_start, alloc_start);
+
+ num = save_migratetypes(&saved_mt[num], alloc_end, isolate_end);
+
/*
* What we do here is we mark all pageblocks in range as
* MIGRATE_ISOLATE. Because pageblock and max order pages may
* have different sizes, and due to the way page allocator
- * work, we align the range to biggest of the two pages so
+ * work, we align the isolation range to biggest of the two so
* that page allocator won't try to merge buddies from
* different pageblocks and change MIGRATE_ISOLATE to some
* other migration type.
@@ -9130,6 +9202,20 @@ int alloc_contig_range(unsigned long start, unsigned long end,
* we are interested in). This will put all the pages in
* range back to page allocator as MIGRATE_ISOLATE.
*
+ * Afterwards, we restore the migratetypes of the pageblocks not
+ * in range, split free pages spanning outside the range,
+ * and put split free pages (at pageblock_order) to the right
+ * migratetype list.
+ *
+ * NOTE: the above approach is used because it can cause free
+ * page accounting issues during isolation, if a page, either
+ * free or in-use, contains multiple pageblocks and we only
+ * isolate a subset of them. For example, if only the second
+ * pageblock is isolated from a page with 2 pageblocks, after
+ * the page is free, it will be put in the first pageblock
+ * migratetype list instead of having 2 pageblocks in two
+ * separate migratetype lists.
+ *
* When this is done, we take the pages in range from page
* allocator removing them from the buddy system. This way
* page allocator will never consider using them.
@@ -9140,10 +9226,9 @@ int alloc_contig_range(unsigned long start, unsigned long end,
* put back to page allocator so that buddy can use them.
*/
- ret = start_isolate_page_range(pfn_max_align_down(start),
- pfn_max_align_up(end), migratetype, 0);
+ ret = start_isolate_page_range(isolate_start, isolate_end, migratetype, 0);
if (ret)
- return ret;
+ goto done;
drain_all_pages(cc.zone);
@@ -9179,6 +9264,19 @@ int alloc_contig_range(unsigned long start, unsigned long end,
* isolated thus they won't get removed from buddy.
*/
+ /*
+ * Restore migratetypes of pageblocks outside [start, end)
+ * TODO: remove it when MIGRATE_ISOLATE becomes a standalone bit
+ */
+
+ num = restore_migratetypes(saved_mt, isolate_start, alloc_start);
+
+ num = restore_migratetypes(&saved_mt[num], alloc_end, isolate_end);
+
+ /*
+ * Split free page spanning [isolate_start, alloc_start) and put the
+ * pageblocks in the right migratetype lists.
+ */
order = 0;
outer_start = start;
while (!PageBuddy(pfn_to_page(outer_start))) {
@@ -9193,37 +9291,73 @@ int alloc_contig_range(unsigned long start, unsigned long end,
order = buddy_order(pfn_to_page(outer_start));
/*
- * outer_start page could be small order buddy page and
- * it doesn't include start page. Adjust outer_start
- * in this case to report failed page properly
- * on tracepoint in test_pages_isolated()
+ * split the free page has start page and put the pageblocks
+ * in the right migratetype list
*/
- if (outer_start + (1UL << order) <= start)
- outer_start = start;
+ if (outer_start + (1UL << order) > start) {
+ struct page *free_page = pfn_to_page(outer_start);
+
+ split_free_page_into_pageblocks(free_page, order, cc.zone);
+ }
+ }
+
+ /*
+ * Split free page spanning [alloc_end, isolate_end) and put the
+ * pageblocks in the right migratetype list
+ */
+ for (outer_end = alloc_end; outer_end < isolate_end;) {
+ unsigned long begin_pfn = outer_end;
+
+ order = 0;
+ while (!PageBuddy(pfn_to_page(outer_end))) {
+ if (++order >= MAX_ORDER) {
+ outer_end = begin_pfn;
+ break;
+ }
+ outer_end &= ~0UL << order;
+ }
+
+ if (outer_end != begin_pfn) {
+ order = buddy_order(pfn_to_page(outer_end));
+
+ /*
+ * split the free page has start page and put the pageblocks
+ * in the right migratetype list
+ */
+ VM_BUG_ON(outer_end + (1UL << order) <= begin_pfn);
+ {
+ struct page *free_page = pfn_to_page(outer_end);
+
+ split_free_page_into_pageblocks(free_page, order, cc.zone);
+ }
+ outer_end += 1UL << order;
+ } else
+ outer_end = begin_pfn + 1;
}
/* Make sure the range is really isolated. */
- if (test_pages_isolated(outer_start, end, 0)) {
+ if (test_pages_isolated(alloc_start, alloc_end, 0)) {
ret = -EBUSY;
goto done;
}
/* Grab isolated pages from freelists. */
- outer_end = isolate_freepages_range(&cc, outer_start, end);
+ outer_end = isolate_freepages_range(&cc, alloc_start, alloc_end);
if (!outer_end) {
ret = -EBUSY;
goto done;
}
/* Free head and tail (if any) */
- if (start != outer_start)
- free_contig_range(outer_start, start - outer_start);
- if (end != outer_end)
- free_contig_range(end, outer_end - end);
+ if (start != alloc_start)
+ free_contig_range(alloc_start, start - alloc_start);
+ if (end != alloc_end)
+ free_contig_range(end, alloc_end - end);
done:
- undo_isolate_page_range(pfn_max_align_down(start),
- pfn_max_align_up(end), migratetype);
+ kfree(saved_mt);
+ undo_isolate_page_range(alloc_start,
+ alloc_end, migratetype);
return ret;
}
EXPORT_SYMBOL(alloc_contig_range);
--
2.34.1
^ permalink raw reply related
* Re: [PATCH] ethernet: ibmveth: use default_groups in kobj_type
From: Tyrel Datwyler @ 2022-01-05 19:23 UTC (permalink / raw)
To: Greg Kroah-Hartman, linux-kernel
Cc: Cristobal Forno, netdev, Paul Mackerras, Jakub Kicinski,
linuxppc-dev, David S. Miller
In-Reply-To: <20220105184101.2859410-1-gregkh@linuxfoundation.org>
On 1/5/22 10:41 AM, Greg Kroah-Hartman wrote:
> There are currently 2 ways to create a set of sysfs files for a
> kobj_type, through the default_attrs field, and the default_groups
> field. Move the ibmveth sysfs code to use default_groups
> field which has been the preferred way since aa30f47cf666 ("kobject: Add
> support for default attribute groups to kobj_type") so that we can soon
> get rid of the obsolete default_attrs field.
>
> Cc: Michael Ellerman <mpe@ellerman.id.au>
> Cc: Benjamin Herrenschmidt <benh@kernel.crashing.org>
> Cc: Paul Mackerras <paulus@samba.org>
> Cc: Cristobal Forno <cforno12@linux.ibm.com>
> Cc: "David S. Miller" <davem@davemloft.net>
> Cc: Jakub Kicinski <kuba@kernel.org>
> Cc: linuxppc-dev@lists.ozlabs.org
> Cc: netdev@vger.kernel.org
> Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
> ---
Reviewed-by: Tyrel Datwyler <tyreld@linux.ibm.com>
^ permalink raw reply
* [PATCH] ethernet: ibmveth: use default_groups in kobj_type
From: Greg Kroah-Hartman @ 2022-01-05 18:41 UTC (permalink / raw)
To: linux-kernel
Cc: Cristobal Forno, Greg Kroah-Hartman, netdev, Paul Mackerras,
Jakub Kicinski, linuxppc-dev, David S. Miller
There are currently 2 ways to create a set of sysfs files for a
kobj_type, through the default_attrs field, and the default_groups
field. Move the ibmveth sysfs code to use default_groups
field which has been the preferred way since aa30f47cf666 ("kobject: Add
support for default attribute groups to kobj_type") so that we can soon
get rid of the obsolete default_attrs field.
Cc: Michael Ellerman <mpe@ellerman.id.au>
Cc: Benjamin Herrenschmidt <benh@kernel.crashing.org>
Cc: Paul Mackerras <paulus@samba.org>
Cc: Cristobal Forno <cforno12@linux.ibm.com>
Cc: "David S. Miller" <davem@davemloft.net>
Cc: Jakub Kicinski <kuba@kernel.org>
Cc: linuxppc-dev@lists.ozlabs.org
Cc: netdev@vger.kernel.org
Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
---
drivers/net/ethernet/ibm/ibmveth.c | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c
index 45ba40cf4d07..22fb0d109a68 100644
--- a/drivers/net/ethernet/ibm/ibmveth.c
+++ b/drivers/net/ethernet/ibm/ibmveth.c
@@ -1890,6 +1890,7 @@ static struct attribute *veth_pool_attrs[] = {
&veth_size_attr,
NULL,
};
+ATTRIBUTE_GROUPS(veth_pool);
static const struct sysfs_ops veth_pool_ops = {
.show = veth_pool_show,
@@ -1899,7 +1900,7 @@ static const struct sysfs_ops veth_pool_ops = {
static struct kobj_type ktype_veth_pool = {
.release = NULL,
.sysfs_ops = &veth_pool_ops,
- .default_attrs = veth_pool_attrs,
+ .default_groups = veth_pool_groups,
};
static int ibmveth_resume(struct device *dev)
--
2.34.1
^ permalink raw reply related
* Re: [5.16.0-rc5][ppc][net] kernel oops when hotplug remove of vNIC interface
From: Jakub Kicinski @ 2022-01-05 18:26 UTC (permalink / raw)
To: Abdul Haleem
Cc: dumazet, netdev, linux-kernel, Dany Madden, alexandr.lobakin,
brian King, Sukadev Bhattiprolu, linuxppc-dev
In-Reply-To: <63380c22-a163-2664-62be-2cf401065e73@linux.vnet.ibm.com>
On Wed, 5 Jan 2022 13:56:53 +0530 Abdul Haleem wrote:
> Greeting's
>
> Mainline kernel 5.16.0-rc5 panics when DLPAR ADD of vNIC device on my
> Powerpc LPAR
>
> Perform below dlpar commands in a loop from linux OS
>
> drmgr -r -c slot -s U9080.HEX.134C488-V1-C3 -w 5 -d 1
> drmgr -a -c slot -s U9080.HEX.134C488-V1-C3 -w 5 -d 1
>
> after 7th iteration, the kernel panics with below messages
>
> console messages:
> [102056] ibmvnic 30000003 env3: Sending CRQ: 801e000864000000
> 0060000000000000
> <intr> ibmvnic 30000003 env3: Handling CRQ: 809e000800000000
> 0000000000000000
> [102056] ibmvnic 30000003 env3: Disabling tx_scrq[0] irq
> [102056] ibmvnic 30000003 env3: Disabling tx_scrq[1] irq
> [102056] ibmvnic 30000003 env3: Disabling rx_scrq[0] irq
> [102056] ibmvnic 30000003 env3: Disabling rx_scrq[1] irq
> [102056] ibmvnic 30000003 env3: Disabling rx_scrq[2] irq
> [102056] ibmvnic 30000003 env3: Disabling rx_scrq[3] irq
> [102056] ibmvnic 30000003 env3: Disabling rx_scrq[4] irq
> [102056] ibmvnic 30000003 env3: Disabling rx_scrq[5] irq
> [102056] ibmvnic 30000003 env3: Disabling rx_scrq[6] irq
> [102056] ibmvnic 30000003 env3: Disabling rx_scrq[7] irq
> [102056] ibmvnic 30000003 env3: Replenished 8 pools
> Kernel attempted to read user page (10) - exploit attempt? (uid: 0)
> BUG: Kernel NULL pointer dereference on read at 0x00000010
> Faulting instruction address: 0xc000000000a3c840
> Oops: Kernel access of bad area, sig: 11 [#1]
> LE PAGE_SIZE=64K MMU=Radix SMP NR_CPUS=2048 NUMA pSeries
> Modules linked in: bridge stp llc ib_core rpadlpar_io rpaphp nfnetlink
> tcp_diag udp_diag inet_diag unix_diag af_packet_diag netlink_diag
> bonding rfkill ibmvnic sunrpc pseries_rng xts vmx_crypto gf128mul
> sch_fq_codel binfmt_misc ip_tables ext4 mbcache jbd2 dm_service_time
> sd_mod t10_pi sg ibmvfc scsi_transport_fc ibmveth dm_multipath dm_mirror
> dm_region_hash dm_log dm_mod fuse
> CPU: 9 PID: 102056 Comm: kworker/9:2 Kdump: loaded Not tainted
> 5.16.0-rc5-autotest-g6441998e2e37 #1
> Workqueue: events_long __ibmvnic_reset [ibmvnic]
> NIP: c000000000a3c840 LR: c0080000029b5378 CTR: c000000000a3c820
> REGS: c0000000548e37e0 TRAP: 0300 Not tainted
> (5.16.0-rc5-autotest-g6441998e2e37)
> MSR: 8000000000009033 <SF,EE,ME,IR,DR,RI,LE> CR: 28248484 XER: 00000004
> CFAR: c0080000029bdd24 DAR: 0000000000000010 DSISR: 40000000 IRQMASK: 0
> GPR00: c0080000029b55d0 c0000000548e3a80 c0000000028f0200 0000000000000000
> GPR04: c000000c7d1a7e00 fffffffffffffff6 0000000000000027 c000000c7d1a7e08
> GPR08: 0000000000000023 0000000000000000 0000000000000010 c0080000029bdd10
> GPR12: c000000000a3c820 c000000c7fca6680 0000000000000000 c000000133016bf8
> GPR16: 00000000000003fe 0000000000001000 0000000000000002 0000000000000008
> GPR20: c000000133016eb0 0000000000000000 0000000000000000 0000000000000003
> GPR24: c000000133016000 c000000133017168 0000000020000000 c000000133016a00
> GPR28: 0000000000000006 c000000133016a00 0000000000000001 c000000133016000
> NIP [c000000000a3c840] napi_enable+0x20/0xc0
> LR [c0080000029b5378] __ibmvnic_open+0xf0/0x430 [ibmvnic]
> Call Trace:
> [c0000000548e3a80] [0000000000000006] 0x6 (unreliable)
> [c0000000548e3ab0] [c0080000029b55d0] __ibmvnic_open+0x348/0x430 [ibmvnic]
> [c0000000548e3b40] [c0080000029bcc28] __ibmvnic_reset+0x500/0xdf0 [ibmvnic]
> [c0000000548e3c60] [c000000000176228] process_one_work+0x288/0x570
> [c0000000548e3d00] [c000000000176588] worker_thread+0x78/0x660
> [c0000000548e3da0] [c0000000001822f0] kthread+0x1c0/0x1d0
> [c0000000548e3e10] [c00000000000cf64] ret_from_kernel_thread+0x5c/0x64
> Instruction dump:
> 7d2948f8 792307e0 4e800020 60000000 3c4c01eb 384239e0 f821ffd1 39430010
> 38a0fff6 e92d1100 f9210028 39200000 <e9030010> f9010020 60420000 e9210020
> ---[ end trace 5f8033b08fd27706 ]---
> radix-mmu: Page sizes from device-tree:
>
> the fault instruction points to
>
> [root@ltcden11-lp1 boot]# gdb -batch
> vmlinuz-5.16.0-rc5-autotest-g6441998e2e37 -ex 'list *(0xc000000000a3c840)'
> 0xc000000000a3c840 is in napi_enable (net/core/dev.c:6966).
> 6961 void napi_enable(struct napi_struct *n)
> 6962 {
> 6963 unsigned long val, new;
> 6964
> 6965 do {
> 6966 val = READ_ONCE(n->state);
If n is NULL here that's gotta be a driver problem.
Adding Dany & Suka.
> 6967 BUG_ON(!test_bit(NAPI_STATE_SCHED, &val));
> 6968
> 6969 new = val & ~(NAPIF_STATE_SCHED | NAPIF_STATE_NPSVC);
> 6970 if (n->dev->threaded && n->thread)
>
^ permalink raw reply
* Re: [PATCH v4] powerpc/pseries: read the lpar name from the firmware
From: Laurent Dufour @ 2022-01-05 16:13 UTC (permalink / raw)
To: Michael Ellerman; +Cc: Nathan Lynch, linuxppc-dev, linux-kernel
In-Reply-To: <20211207171109.22793-1-ldufour@linux.ibm.com>
Happy New Year, Michael!
Do you consider taking that patch soon?
Thanks,
Laurent.
On 07/12/2021, 18:11:09, Laurent Dufour wrote:
> The LPAR name may be changed after the LPAR has been started in the HMC.
> In that case lparstat command is not reporting the updated value because it
> reads it from the device tree which is read at boot time.
>
> However this value could be read from RTAS.
>
> Adding this value in the /proc/powerpc/lparcfg output allows to read the
> updated value.
>
> Cc: Nathan Lynch <nathanl@linux.ibm.com>
> Signed-off-by: Laurent Dufour <ldufour@linux.ibm.com>
> ---
> v4:
> address Nathan's new comments limiting size of the buffer.
> v3:
> address Michael's comments.
> v2:
> address Nathan's comments.
> change title to partition_name aligning with existing partition_id
> ---
> arch/powerpc/platforms/pseries/lparcfg.c | 54 ++++++++++++++++++++++++
> 1 file changed, 54 insertions(+)
>
> diff --git a/arch/powerpc/platforms/pseries/lparcfg.c b/arch/powerpc/platforms/pseries/lparcfg.c
> index f71eac74ea92..058d9a5fe545 100644
> --- a/arch/powerpc/platforms/pseries/lparcfg.c
> +++ b/arch/powerpc/platforms/pseries/lparcfg.c
> @@ -311,6 +311,59 @@ static void parse_mpp_x_data(struct seq_file *m)
> seq_printf(m, "coalesce_pool_spurr=%ld\n", mpp_x_data.pool_spurr_cycles);
> }
>
> +/*
> + * PAPR defines, in section "7.3.16 System Parameters Option", the token 55 to
> + * read the LPAR name, and the largest output data to 4000 + 2 bytes length.
> + */
> +#define SPLPAR_LPAR_NAME_TOKEN 55
> +#define GET_SYS_PARM_BUF_SIZE 4002
> +#if GET_SYS_PARM_BUF_SIZE > RTAS_DATA_BUF_SIZE
> +#error "GET_SYS_PARM_BUF_SIZE is larger than RTAS_DATA_BUF_SIZE"
> +#endif
> +static void read_lpar_name(struct seq_file *m)
> +{
> + int rc, len, token;
> + union {
> + char raw_buffer[GET_SYS_PARM_BUF_SIZE];
> + struct {
> + __be16 len;
> + char name[GET_SYS_PARM_BUF_SIZE-2];
> + };
> + } *local_buffer;
> +
> + token = rtas_token("ibm,get-system-parameter");
> + if (token == RTAS_UNKNOWN_SERVICE)
> + return;
> +
> + local_buffer = kmalloc(sizeof(*local_buffer), GFP_KERNEL);
> + if (!local_buffer)
> + return;
> +
> + do {
> + spin_lock(&rtas_data_buf_lock);
> + memset(rtas_data_buf, 0, sizeof(*local_buffer));
> + rc = rtas_call(token, 3, 1, NULL, SPLPAR_LPAR_NAME_TOKEN,
> + __pa(rtas_data_buf), sizeof(*local_buffer));
> + if (!rc)
> + memcpy(local_buffer->raw_buffer, rtas_data_buf,
> + sizeof(local_buffer->raw_buffer));
> + spin_unlock(&rtas_data_buf_lock);
> + } while (rtas_busy_delay(rc));
> +
> + if (!rc) {
> + /* Force end of string */
> + len = min((int) be16_to_cpu(local_buffer->len),
> + (int) sizeof(local_buffer->name)-1);
> + local_buffer->name[len] = '\0';
> +
> + seq_printf(m, "partition_name=%s\n", local_buffer->name);
> + } else
> + pr_err_once("Error calling get-system-parameter (0x%x)\n", rc);
> +
> + kfree(local_buffer);
> +}
> +
> +
> #define SPLPAR_CHARACTERISTICS_TOKEN 20
> #define SPLPAR_MAXLENGTH 1026*(sizeof(char))
>
> @@ -496,6 +549,7 @@ static int pseries_lparcfg_data(struct seq_file *m, void *v)
>
> if (firmware_has_feature(FW_FEATURE_SPLPAR)) {
> /* this call handles the ibm,get-system-parameter contents */
> + read_lpar_name(m);
> parse_system_parameter_string(m);
> parse_ppp_data(m);
> parse_mpp_data(m);
^ permalink raw reply
* Re: [PATCH]powerpc/xmon: Dump XIVE information for online-only processors.
From: Cédric Le Goater @ 2022-01-05 14:39 UTC (permalink / raw)
To: Sachin Sant, linuxppc-dev
In-Reply-To: <164139226833.12930.272224382183014664.sendpatchset@MacBook-Pro.local>
On 1/5/22 15:17, Sachin Sant wrote:
> dxa command in XMON debugger iterates through all possible processors.
> As a result, empty lines are printed even for processors which are not
> online.
>
> CPU 47:pp=00 CPPR=ff IPI=0x0040002f PQ=-- EQ idx=699 T=0 00000000 00000000
> CPU 48:
> CPU 49:
>
> Restrict XIVE information(dxa) to be displayed for online processors only.
>
> Signed-off-by: Sachin Sant <sachinp@linux.vnet.ibm.com>
Looks good to me. We should do the same for :
/sys/kernel/debug/powerpc/xive/ipis
Reviewed-by: Cédric Le Goater <clg@kaod.org>
Thanks,
C.
> ---
> diff -Naurp a/arch/powerpc/xmon/xmon.c b/arch/powerpc/xmon/xmon.c
> --- a/arch/powerpc/xmon/xmon.c 2022-01-05 08:52:59.480118166 -0500
> +++ b/arch/powerpc/xmon/xmon.c 2022-01-05 08:56:18.469589555 -0500
> @@ -2817,12 +2817,12 @@ static void dump_all_xives(void)
> {
> int cpu;
>
> - if (num_possible_cpus() == 0) {
> + if (num_online_cpus() == 0) {
> printf("No possible cpus, use 'dx #' to dump individual cpus\n");
> return;
> }
>
> - for_each_possible_cpu(cpu)
> + for_each_online_cpu(cpu)
> dump_one_xive(cpu);
> }
>
>
^ permalink raw reply
* [PATCH]powerpc/xmon: Dump XIVE information for online-only processors.
From: Sachin Sant @ 2022-01-05 14:17 UTC (permalink / raw)
To: linuxppc-dev; +Cc: Sachin Sant, clg
dxa command in XMON debugger iterates through all possible processors.
As a result, empty lines are printed even for processors which are not
online.
CPU 47:pp=00 CPPR=ff IPI=0x0040002f PQ=-- EQ idx=699 T=0 00000000 00000000
CPU 48:
CPU 49:
Restrict XIVE information(dxa) to be displayed for online processors only.
Signed-off-by: Sachin Sant <sachinp@linux.vnet.ibm.com>
---
diff -Naurp a/arch/powerpc/xmon/xmon.c b/arch/powerpc/xmon/xmon.c
--- a/arch/powerpc/xmon/xmon.c 2022-01-05 08:52:59.480118166 -0500
+++ b/arch/powerpc/xmon/xmon.c 2022-01-05 08:56:18.469589555 -0500
@@ -2817,12 +2817,12 @@ static void dump_all_xives(void)
{
int cpu;
- if (num_possible_cpus() == 0) {
+ if (num_online_cpus() == 0) {
printf("No possible cpus, use 'dx #' to dump individual cpus\n");
return;
}
- for_each_possible_cpu(cpu)
+ for_each_online_cpu(cpu)
dump_one_xive(cpu);
}
^ permalink raw reply
* [PATCH] ASoC: fsl_asrc: refine the check of available clock divider
From: Shengjiu Wang @ 2022-01-05 11:08 UTC (permalink / raw)
To: timur, nicoleotsuka, Xiubo.Lee, festevam, broonie, alsa-devel,
lgirdwood, perex, tiwai
Cc: linuxppc-dev, linux-kernel
According to RM, the clock divider range is from 1 to 8, clock
prescaling ratio may be any power of 2 from 1 to 128.
So the supported divider is not all the value between
1 and 1024, just limited value in that range.
Create table for the supported divder and add function to
check the clock divider is available by comparing with
the table.
Fixes: d0250cf4f2ab ("ASoC: fsl_asrc: Add an option to select internal ratio mode")
Signed-off-by: Shengjiu Wang <shengjiu.wang@nxp.com>
---
sound/soc/fsl/fsl_asrc.c | 69 +++++++++++++++++++++++++++++++++-------
1 file changed, 58 insertions(+), 11 deletions(-)
diff --git a/sound/soc/fsl/fsl_asrc.c b/sound/soc/fsl/fsl_asrc.c
index 24b41881a68f..d7d1536a4f37 100644
--- a/sound/soc/fsl/fsl_asrc.c
+++ b/sound/soc/fsl/fsl_asrc.c
@@ -19,6 +19,7 @@
#include "fsl_asrc.h"
#define IDEAL_RATIO_DECIMAL_DEPTH 26
+#define DIVIDER_NUM 64
#define pair_err(fmt, ...) \
dev_err(&asrc->pdev->dev, "Pair %c: " fmt, 'A' + index, ##__VA_ARGS__)
@@ -101,6 +102,55 @@ static unsigned char clk_map_imx8qxp[2][ASRC_CLK_MAP_LEN] = {
},
};
+/*
+ * According to RM, the divider range is 1 ~ 8,
+ * prescaler is power of 2 from 1 ~ 128.
+ */
+static int asrc_clk_divider[DIVIDER_NUM] = {
+ 1, 2, 4, 8, 16, 32, 64, 128, /* divider = 1 */
+ 2, 4, 8, 16, 32, 64, 128, 256, /* divider = 2 */
+ 3, 6, 12, 24, 48, 96, 192, 384, /* divider = 3 */
+ 4, 8, 16, 32, 64, 128, 256, 512, /* divider = 4 */
+ 5, 10, 20, 40, 80, 160, 320, 640, /* divider = 5 */
+ 6, 12, 24, 48, 96, 192, 384, 768, /* divider = 6 */
+ 7, 14, 28, 56, 112, 224, 448, 896, /* divider = 7 */
+ 8, 16, 32, 64, 128, 256, 512, 1024, /* divider = 8 */
+};
+
+/*
+ * Check if the divider is available for internal ratio mode
+ */
+static bool fsl_asrc_divider_avail(int clk_rate, int rate, int *div)
+{
+ u32 rem, i;
+ u64 n;
+
+ if (div)
+ *div = 0;
+
+ if (clk_rate == 0 || rate == 0)
+ return false;
+
+ n = clk_rate;
+ rem = do_div(n, rate);
+
+ if (div)
+ *div = n;
+
+ if (rem != 0)
+ return false;
+
+ for (i = 0; i < DIVIDER_NUM; i++) {
+ if (n == asrc_clk_divider[i])
+ break;
+ }
+
+ if (i == DIVIDER_NUM)
+ return false;
+
+ return true;
+}
+
/**
* fsl_asrc_sel_proc - Select the pre-processing and post-processing options
* @inrate: input sample rate
@@ -330,12 +380,12 @@ static int fsl_asrc_config_pair(struct fsl_asrc_pair *pair, bool use_ideal_rate)
enum asrc_word_width input_word_width;
enum asrc_word_width output_word_width;
u32 inrate, outrate, indiv, outdiv;
- u32 clk_index[2], div[2], rem[2];
+ u32 clk_index[2], div[2];
u64 clk_rate;
int in, out, channels;
int pre_proc, post_proc;
struct clk *clk;
- bool ideal;
+ bool ideal, div_avail;
if (!config) {
pair_err("invalid pair config\n");
@@ -415,8 +465,7 @@ static int fsl_asrc_config_pair(struct fsl_asrc_pair *pair, bool use_ideal_rate)
clk = asrc_priv->asrck_clk[clk_index[ideal ? OUT : IN]];
clk_rate = clk_get_rate(clk);
- rem[IN] = do_div(clk_rate, inrate);
- div[IN] = (u32)clk_rate;
+ div_avail = fsl_asrc_divider_avail(clk_rate, inrate, &div[IN]);
/*
* The divider range is [1, 1024], defined by the hardware. For non-
@@ -425,7 +474,7 @@ static int fsl_asrc_config_pair(struct fsl_asrc_pair *pair, bool use_ideal_rate)
* only result in different converting speeds. So remainder does not
* matter, as long as we keep the divider within its valid range.
*/
- if (div[IN] == 0 || (!ideal && (div[IN] > 1024 || rem[IN] != 0))) {
+ if (div[IN] == 0 || (!ideal && !div_avail)) {
pair_err("failed to support input sample rate %dHz by asrck_%x\n",
inrate, clk_index[ideal ? OUT : IN]);
return -EINVAL;
@@ -436,13 +485,12 @@ static int fsl_asrc_config_pair(struct fsl_asrc_pair *pair, bool use_ideal_rate)
clk = asrc_priv->asrck_clk[clk_index[OUT]];
clk_rate = clk_get_rate(clk);
if (ideal && use_ideal_rate)
- rem[OUT] = do_div(clk_rate, IDEAL_RATIO_RATE);
+ div_avail = fsl_asrc_divider_avail(clk_rate, IDEAL_RATIO_RATE, &div[OUT]);
else
- rem[OUT] = do_div(clk_rate, outrate);
- div[OUT] = clk_rate;
+ div_avail = fsl_asrc_divider_avail(clk_rate, outrate, &div[OUT]);
/* Output divider has the same limitation as the input one */
- if (div[OUT] == 0 || (!ideal && (div[OUT] > 1024 || rem[OUT] != 0))) {
+ if (div[OUT] == 0 || (!ideal && !div_avail)) {
pair_err("failed to support output sample rate %dHz by asrck_%x\n",
outrate, clk_index[OUT]);
return -EINVAL;
@@ -621,8 +669,7 @@ static void fsl_asrc_select_clk(struct fsl_asrc_priv *asrc_priv,
clk_index = asrc_priv->clk_map[j][i];
clk_rate = clk_get_rate(asrc_priv->asrck_clk[clk_index]);
/* Only match a perfect clock source with no remainder */
- if (clk_rate != 0 && (clk_rate / rate[j]) <= 1024 &&
- (clk_rate % rate[j]) == 0)
+ if (fsl_asrc_divider_avail(clk_rate, rate[j], NULL))
break;
}
--
2.17.1
^ permalink raw reply related
page: next (older) | prev (newer) | latest
- recent:[subjects (threaded)|topics (new)|topics (active)]
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox