* [PATCH 0/2] macintosh/rack-meter: Adjustments for rackmeter_probe()
From: SF Markus Elfring @ 2018-01-16 20:38 UTC (permalink / raw)
To: linuxppc-dev, Benjamin Herrenschmidt; +Cc: LKML, kernel-janitors
From: Markus Elfring <elfring@users.sourceforge.net>
Date: Tue, 16 Jan 2018 21:36:56 +0100
Two update suggestions were taken into account
from static source code analysis.
Markus Elfring (2):
Delete an error message for a failed memory allocation
Improve a size determination
drivers/macintosh/rack-meter.c | 3 +--
1 file changed, 1 insertion(+), 2 deletions(-)
--
2.15.1
^ permalink raw reply
* Re: [PATCH] [RESEND] spufs: use timespec64 for timestamps
From: Arnd Bergmann @ 2018-01-16 19:58 UTC (permalink / raw)
To: Jeremy Kerr
Cc: Benjamin Herrenschmidt, Paul Mackerras, Michael Ellerman, Al Viro,
Andrew Morton, linuxppc-dev, Linux Kernel Mailing List
In-Reply-To: <1dc0e4ce-2190-eb83-166f-b8ba7cdacede@ozlabs.org>
On Tue, Jan 16, 2018 at 8:07 PM, Jeremy Kerr <jk@ozlabs.org> wrote:
> Hi Arnd,
>
>> The switch log prints the tv_sec portion of timespec as a 32-bit
>> number, while overflows in 2106. It also uses the timespec type,
>> which is safe on 64-bit architectures, but deprecated because
>> it causes overflows in 2038 elsewhere.
>>
>> This changes it to timespec64 and printing a 64-bit number for
>> consistency.
>
> If we still have spufs in the tree in 2038 I'd be worried :)
Agreed. My hope is to get rid of 'timespec' in 2018 though, this is just
one of many parts of the puzzle.
Arnd
^ permalink raw reply
* Re: [PATCH] [RESEND] spufs: use timespec64 for timestamps
From: Jeremy Kerr @ 2018-01-16 19:07 UTC (permalink / raw)
To: Arnd Bergmann, Benjamin Herrenschmidt, Paul Mackerras,
Michael Ellerman
Cc: Al Viro, Andrew Morton, linuxppc-dev, linux-kernel
In-Reply-To: <20180116170053.2557047-1-arnd@arndb.de>
Hi Arnd,
> The switch log prints the tv_sec portion of timespec as a 32-bit
> number, while overflows in 2106. It also uses the timespec type,
> which is safe on 64-bit architectures, but deprecated because
> it causes overflows in 2038 elsewhere.
>
> This changes it to timespec64 and printing a 64-bit number for
> consistency.
If we still have spufs in the tree in 2038 I'd be worried :) But good
to keep things consistent.
Acked-by: Jeremy Kerr <jk@ozlabs.org>
Michael: want to take this directly through your tree?
Cheers,
Jeremy
^ permalink raw reply
* Re: DPAA Ethernet traffice troubles with Linux kernel
From: mad skateman @ 2018-01-16 18:44 UTC (permalink / raw)
To: Joakim Tjernlund
Cc: andrew@lunn.ch, linuxppc-dev@lists.ozlabs.org,
madalin.bucur@nxp.com
In-Reply-To: <1516125454.18795.87.camel@infinera.com>
[-- Attachment #1: Type: text/plain, Size: 1956 bytes --]
Hi,
I have been looking deeper into my wireshark packet captures and found
something that could be helpfull.
I can see that the Ethernet NIC at least does something. ARP BROADCAST
traffic is seen.
But i also found some packets...which have LG BITs set to 1 .. when i think
0 should be the correct value..something with Octets.
*"the first 3 octets/first-half of a MAC-48/EUI-48 Address, correspond to
the OUI (e.g.: MAC = 06:00:00:xx:xx:xx, OUI = 06:00:00), and the 2nd least
significant bit of its first octet is used to differentiate "Universally"
and "Locally" administered addresses). In other words, if we convert 06
(Hex) to 00000110 (Binary), we can see that the U/L bit is set to one,
which means that it a locally administered address.*
*Given this if we disable that bit, we get the matching "Universally
Administered Address" 00000100 (Binary), 04 (Hex) -> "04:00:00", hence my
question:"*
This has something to do with the MAC adresses being locally administered
.. and since whe can use Uboot and choose any Mac addr we want, this could
make sense..
These types of Logs also apear in my Wireshark Capture files... ( these are
not my org. logs)
Ethernet II, Src: Microchi_8f:c6:a8 (d8:80:39:8f:c6:a8), Dst: Broadcast
(ff:ff:ff:ff:ff:ff)
Destination: Broadcast (ff:ff:ff:ff:ff:ff)
Address: Broadcast (ff:ff:ff:ff:ff:ff)
.... ..1. .... .... .... .... = LG bit: Locally administered address (this
is NOT the factory default)
.... ...1 .... .... .... .... = IG bit: Group address (multicast/broadcast)
Source: Microchi_8f:c6:a8 (d8:80:39:8f:c6:a8)
Address: Microchi_8f:c6:a8 (d8:80:39:8f:c6:a8)
.... ..0. .... .... .... .... = LG bit: Globally unique address (factory
default)
.... ...0 .... .... .... .... = IG bit: Individual address (unicast)
Type: IP (0x0800)
Internet Protocol Version 4, Src: 0.0.0.0 (0.0.0.0), Dst: 255.255.255.255
(255.255.255.255)
In the link below some similair Logs and problems regarding DHCP for
example.
Dave
[-- Attachment #2: Type: text/html, Size: 4356 bytes --]
^ permalink raw reply
* Re: DPAA Ethernet traffice troubles with Linux kernel
From: mad skateman @ 2018-01-16 18:39 UTC (permalink / raw)
To: Joakim Tjernlund
Cc: andrew@lunn.ch, linuxppc-dev@lists.ozlabs.org,
netdev@vger.kernel.org, madalin.bucur@nxp.com
In-Reply-To: <1516125454.18795.87.camel@infinera.com>
[-- Attachment #1: Type: text/plain, Size: 4415 bytes --]
Hi,
I have been looking deeper into my wireshark packet captures and found
something that could be helpfull.
I can see using wireshark that the Ethernet NIC at least does something.
ARP BROADCAST traffic is seen.
But i also found some packets...which have LG BITs set to 1 .. when i think
0 should be the correct value..
This has something to do with the MAC adresses being locally administered
.. and since whe can use Uboot and choose any Mac addr we want, this could
make sense..
These types of Logs also apear in my Wireshark Capture files... ( these are
not my org. logs)
Ethernet II, Src: Microchi_8f:c6:a8 (d8:80:39:8f:c6:a8), Dst: Broadcast
(ff:ff:ff:ff:ff:ff)
Destination: Broadcast (ff:ff:ff:ff:ff:ff)
Address: Broadcast (ff:ff:ff:ff:ff:ff)
.... ..1. .... .... .... .... = LG bit: Locally administered address (this
is NOT the factory default)
.... ...1 .... .... .... .... = IG bit: Group address (multicast/broadcast)
Source: Microchi_8f:c6:a8 (d8:80:39:8f:c6:a8)
Address: Microchi_8f:c6:a8 (d8:80:39:8f:c6:a8)
.... ..0. .... .... .... .... = LG bit: Globally unique address (factory
default)
.... ...0 .... .... .... .... = IG bit: Individual address (unicast)
Type: IP (0x0800)
Internet Protocol Version 4, Src: 0.0.0.0 (0.0.0.0), Dst: 255.255.255.255
(255.255.255.255)
In the link below some similair Logs and problems regarding DHCP for
example.
Dave
On Tue, Jan 16, 2018 at 6:57 PM, Joakim Tjernlund <
Joakim.Tjernlund@infinera.com> wrote:
> On Thu, 1970-01-01 at 00:00 +0000, Andrew Lunn wrote:
> > CAUTION: This email originated from outside of the organization. Do not
> click links or open attachments unless you recognize the sender and know
> the content is safe.
> >
> >
> > > Hi, just saw this and thought of a small patch I just wrote for mdio
> bus, o idea
> > > if it is relevant but here goes:
> > >
> > > From fe0b98d54a79779482700676331b4d10a0f3cada Mon Sep 17 00:00:00 2001
> > > From: Joakim Tjernlund <joakim.tjernlund@infinera.com>
> > > Date: Sun, 14 Jan 2018 21:27:20 +0100
> > > Subject: [PATCH] of_mdiobus_register: Continue after error
> > >
> > > of_mdiobus_register unregister itself if one phy fails to register
> > > which is bad for system having all its PHYs on the same MDIO bus.
> > > Just log the error and continue with the remaining PHYs instead.
> > >
> > > Signed-off-by: Joakim Tjernlund <joakim.tjernlund@infinera.com>
> >
> > Hi Joakim
> >
> > You appear to be using an old kernel. Take a look at:
>
> Not really, I am using 4.14.x and I don't think that is old. Seems like
> this
> patch hasn't been sent to 4.14.x.
>
> I wonder if I might be missing something else, we just moved to 4.14 and
> notic that all
> our fixed PHYs are non functioning:
> fsl_mac ffe4e2000.ethernet: FMan MEMAC
> fsl_mac ffe4e2000.ethernet: FMan MAC address: 00:06:9c:0b:06:20
> fsl_mac dpaa-ethernet.0: __devm_request_mem_region(mac) failed
> fsl_mac: probe of dpaa-ethernet.0 failed with error -16
> fsl_mac ffe4e4000.ethernet: FMan MEMAC
> fsl_mac ffe4e4000.ethernet: FMan MAC address: 00:06:9c:0b:06:21
> fsl_mac dpaa-ethernet.1: __devm_request_mem_region(mac) failed
> fsl_mac: probe of dpaa-ethernet.1 failed with error -16
> fsl_mac ffe4e6000.ethernet: FMan MEMAC
> fsl_mac ffe4e6000.ethernet: FMan MAC address: 00:06:9c:0b:06:22
> fsl_mac dpaa-ethernet.2: __devm_request_mem_region(mac) failed
> fsl_mac: probe of dpaa-ethernet.2 failed with error -16
> fsl_mac ffe4e8000.ethernet: FMan MEMAC
> fsl_mac ffe4e8000.ethernet: FMan MAC address: 00:06:9c:0b:06:23
> fsl_mac dpaa-ethernet.3: __devm_request_mem_region(mac) failed
> fsl_mac: probe of dpaa-ethernet.3 failed with error -16
>
> Feels like FMAN still think there are real PHYs there ?
> >
> > commit 95f566de0269a0c59fd6a737a147731302136429
> > Author: Madalin Bucur <madalin.bucur@nxp.com>
> > Date: Tue Jan 9 14:43:34 2018 +0200
> >
> > of_mdio: avoid MDIO bus removal when a PHY is missing
> >
> > If one of the child devices is missing the of_mdiobus_register_phy()
> > call will return -ENODEV. When a missing device is encountered the
> > registration of the remaining PHYs is stopped and the MDIO bus will
> > fail to register. Propagate all errors except ENODEV to avoid it.
> >
> > Signed-off-by: Madalin Bucur <madalin.bucur@nxp.com>
> > Reviewed-by: Andrew Lunn <andrew@lunn.ch>
> > Signed-off-by: David S. Miller <davem@davemloft.net>
> >
> >
> > Andrew
>
[-- Attachment #2: Type: text/html, Size: 6388 bytes --]
^ permalink raw reply
* Re: DPAA Ethernet traffice troubles with Linux kernel
From: mad skateman @ 2018-01-16 18:38 UTC (permalink / raw)
To: Joakim Tjernlund
Cc: andrew@lunn.ch, linuxppc-dev@lists.ozlabs.org,
netdev@vger.kernel.org, madalin.bucur@nxp.com
In-Reply-To: <1516125454.18795.87.camel@infinera.com>
[-- Attachment #1: Type: text/plain, Size: 4477 bytes --]
Hi,
I have been looking deeper into my wireshark packet captures and found
something that could be helpfull.
I can see using wireshark that the Ethernet NIC at least does something.
ARP BROADCAST traffic is seen.
But i also found some packets...which have LG BITs set to 1 .. when i think
0 should be the correct value..
This has something to do with the MAC adresses being locally administered
.. and since whe can use Uboot and choose any Mac addr we want, this could
make sense..
These types of Logs also apear in my Wireshark Capture files... ( these are
not my org. logs)
Ethernet II, Src: Microchi_8f:c6:a8 (d8:80:39:8f:c6:a8), Dst: Broadcast
(ff:ff:ff:ff:ff:ff)
Destination: Broadcast (ff:ff:ff:ff:ff:ff)
Address: Broadcast (ff:ff:ff:ff:ff:ff)
.... ..1. .... .... .... .... = LG bit: Locally administered address (this
is NOT the factory default)
.... ...1 .... .... .... .... = IG bit: Group address (multicast/broadcast)
Source: Microchi_8f:c6:a8 (d8:80:39:8f:c6:a8)
Address: Microchi_8f:c6:a8 (d8:80:39:8f:c6:a8)
.... ..0. .... .... .... .... = LG bit: Globally unique address (factory
default)
.... ...0 .... .... .... .... = IG bit: Individual address (unicast)
Type: IP (0x0800)
Internet Protocol Version 4, Src: 0.0.0.0 (0.0.0.0), Dst: 255.255.255.255
(255.255.255.255)
In the link below some similair Logs and problems regarding DHCP for
example.
Dave
https://www.microchip.com/forums/m/tm.aspx?m=956881&fp=1&p=2
On Tue, Jan 16, 2018 at 6:57 PM, Joakim Tjernlund <
Joakim.Tjernlund@infinera.com> wrote:
> On Thu, 1970-01-01 at 00:00 +0000, Andrew Lunn wrote:
> > CAUTION: This email originated from outside of the organization. Do not
> click links or open attachments unless you recognize the sender and know
> the content is safe.
> >
> >
> > > Hi, just saw this and thought of a small patch I just wrote for mdio
> bus, o idea
> > > if it is relevant but here goes:
> > >
> > > From fe0b98d54a79779482700676331b4d10a0f3cada Mon Sep 17 00:00:00 2001
> > > From: Joakim Tjernlund <joakim.tjernlund@infinera.com>
> > > Date: Sun, 14 Jan 2018 21:27:20 +0100
> > > Subject: [PATCH] of_mdiobus_register: Continue after error
> > >
> > > of_mdiobus_register unregister itself if one phy fails to register
> > > which is bad for system having all its PHYs on the same MDIO bus.
> > > Just log the error and continue with the remaining PHYs instead.
> > >
> > > Signed-off-by: Joakim Tjernlund <joakim.tjernlund@infinera.com>
> >
> > Hi Joakim
> >
> > You appear to be using an old kernel. Take a look at:
>
> Not really, I am using 4.14.x and I don't think that is old. Seems like
> this
> patch hasn't been sent to 4.14.x.
>
> I wonder if I might be missing something else, we just moved to 4.14 and
> notic that all
> our fixed PHYs are non functioning:
> fsl_mac ffe4e2000.ethernet: FMan MEMAC
> fsl_mac ffe4e2000.ethernet: FMan MAC address: 00:06:9c:0b:06:20
> fsl_mac dpaa-ethernet.0: __devm_request_mem_region(mac) failed
> fsl_mac: probe of dpaa-ethernet.0 failed with error -16
> fsl_mac ffe4e4000.ethernet: FMan MEMAC
> fsl_mac ffe4e4000.ethernet: FMan MAC address: 00:06:9c:0b:06:21
> fsl_mac dpaa-ethernet.1: __devm_request_mem_region(mac) failed
> fsl_mac: probe of dpaa-ethernet.1 failed with error -16
> fsl_mac ffe4e6000.ethernet: FMan MEMAC
> fsl_mac ffe4e6000.ethernet: FMan MAC address: 00:06:9c:0b:06:22
> fsl_mac dpaa-ethernet.2: __devm_request_mem_region(mac) failed
> fsl_mac: probe of dpaa-ethernet.2 failed with error -16
> fsl_mac ffe4e8000.ethernet: FMan MEMAC
> fsl_mac ffe4e8000.ethernet: FMan MAC address: 00:06:9c:0b:06:23
> fsl_mac dpaa-ethernet.3: __devm_request_mem_region(mac) failed
> fsl_mac: probe of dpaa-ethernet.3 failed with error -16
>
> Feels like FMAN still think there are real PHYs there ?
> >
> > commit 95f566de0269a0c59fd6a737a147731302136429
> > Author: Madalin Bucur <madalin.bucur@nxp.com>
> > Date: Tue Jan 9 14:43:34 2018 +0200
> >
> > of_mdio: avoid MDIO bus removal when a PHY is missing
> >
> > If one of the child devices is missing the of_mdiobus_register_phy()
> > call will return -ENODEV. When a missing device is encountered the
> > registration of the remaining PHYs is stopped and the MDIO bus will
> > fail to register. Propagate all errors except ENODEV to avoid it.
> >
> > Signed-off-by: Madalin Bucur <madalin.bucur@nxp.com>
> > Reviewed-by: Andrew Lunn <andrew@lunn.ch>
> > Signed-off-by: David S. Miller <davem@davemloft.net>
> >
> >
> > Andrew
>
[-- Attachment #2: Type: text/html, Size: 6495 bytes --]
^ permalink raw reply
* Re: DPAA Ethernet traffice troubles with Linux kernel
From: mad skateman @ 2018-01-16 18:16 UTC (permalink / raw)
To: Joakim Tjernlund
Cc: andrew@lunn.ch, linuxppc-dev@lists.ozlabs.org,
netdev@vger.kernel.org, madalin.bucur@nxp.com
In-Reply-To: <1516125454.18795.87.camel@infinera.com>
[-- Attachment #1: Type: text/plain, Size: 3997 bytes --]
Hi,
I have been looking deeper into my wireshark packet captures and found
something that could be helpfull.
I can see that the Ethernet NICS at least do something. ARP BROADCAST
traffic is seen.
But i also found some packets...which have LG BITs set to 1 .. when i think
0 should be the correct value..
This has something to do with the MAC adresses being non Authorative.. and
since whe can use Uboot and choose any Mac addr we want, this could make
sense.. More info about this
https://osqa-ask.wireshark.org/questions/59761/oui-lookup-tool-does-not-recognize-local-addresses
Logs like these appear: not the original.
Destination: Broadcast (ff:ff:ff:ff:ff:ff)
Address: Broadcast (ff:ff:ff:ff:ff:ff)
.... ..1. .... .... .... .... = LG bit: Locally administered address (this
is NOT the factory default)
.... ...1 .... .... .... .... = IG bit: Group address (multicast/broadcast)
Will try to get the picture more clear... but think about this..
On Tue, Jan 16, 2018 at 6:57 PM, Joakim Tjernlund <
Joakim.Tjernlund@infinera.com> wrote:
> On Thu, 1970-01-01 at 00:00 +0000, Andrew Lunn wrote:
> > CAUTION: This email originated from outside of the organization. Do not
> click links or open attachments unless you recognize the sender and know
> the content is safe.
> >
> >
> > > Hi, just saw this and thought of a small patch I just wrote for mdio
> bus, o idea
> > > if it is relevant but here goes:
> > >
> > > From fe0b98d54a79779482700676331b4d10a0f3cada Mon Sep 17 00:00:00 2001
> > > From: Joakim Tjernlund <joakim.tjernlund@infinera.com>
> > > Date: Sun, 14 Jan 2018 21:27:20 +0100
> > > Subject: [PATCH] of_mdiobus_register: Continue after error
> > >
> > > of_mdiobus_register unregister itself if one phy fails to register
> > > which is bad for system having all its PHYs on the same MDIO bus.
> > > Just log the error and continue with the remaining PHYs instead.
> > >
> > > Signed-off-by: Joakim Tjernlund <joakim.tjernlund@infinera.com>
> >
> > Hi Joakim
> >
> > You appear to be using an old kernel. Take a look at:
>
> Not really, I am using 4.14.x and I don't think that is old. Seems like
> this
> patch hasn't been sent to 4.14.x.
>
> I wonder if I might be missing something else, we just moved to 4.14 and
> notic that all
> our fixed PHYs are non functioning:
> fsl_mac ffe4e2000.ethernet: FMan MEMAC
> fsl_mac ffe4e2000.ethernet: FMan MAC address: 00:06:9c:0b:06:20
> fsl_mac dpaa-ethernet.0: __devm_request_mem_region(mac) failed
> fsl_mac: probe of dpaa-ethernet.0 failed with error -16
> fsl_mac ffe4e4000.ethernet: FMan MEMAC
> fsl_mac ffe4e4000.ethernet: FMan MAC address: 00:06:9c:0b:06:21
> fsl_mac dpaa-ethernet.1: __devm_request_mem_region(mac) failed
> fsl_mac: probe of dpaa-ethernet.1 failed with error -16
> fsl_mac ffe4e6000.ethernet: FMan MEMAC
> fsl_mac ffe4e6000.ethernet: FMan MAC address: 00:06:9c:0b:06:22
> fsl_mac dpaa-ethernet.2: __devm_request_mem_region(mac) failed
> fsl_mac: probe of dpaa-ethernet.2 failed with error -16
> fsl_mac ffe4e8000.ethernet: FMan MEMAC
> fsl_mac ffe4e8000.ethernet: FMan MAC address: 00:06:9c:0b:06:23
> fsl_mac dpaa-ethernet.3: __devm_request_mem_region(mac) failed
> fsl_mac: probe of dpaa-ethernet.3 failed with error -16
>
> Feels like FMAN still think there are real PHYs there ?
> >
> > commit 95f566de0269a0c59fd6a737a147731302136429
> > Author: Madalin Bucur <madalin.bucur@nxp.com>
> > Date: Tue Jan 9 14:43:34 2018 +0200
> >
> > of_mdio: avoid MDIO bus removal when a PHY is missing
> >
> > If one of the child devices is missing the of_mdiobus_register_phy()
> > call will return -ENODEV. When a missing device is encountered the
> > registration of the remaining PHYs is stopped and the MDIO bus will
> > fail to register. Propagate all errors except ENODEV to avoid it.
> >
> > Signed-off-by: Madalin Bucur <madalin.bucur@nxp.com>
> > Reviewed-by: Andrew Lunn <andrew@lunn.ch>
> > Signed-off-by: David S. Miller <davem@davemloft.net>
> >
> >
> > Andrew
>
[-- Attachment #2: Type: text/html, Size: 6177 bytes --]
^ permalink raw reply
* Re: DPAA Ethernet traffice troubles with Linux kernel
From: Joakim Tjernlund @ 2018-01-16 17:57 UTC (permalink / raw)
To: andrew@lunn.ch
Cc: linuxppc-dev@lists.ozlabs.org, netdev@vger.kernel.org,
madalin.bucur@nxp.com, madskateman@gmail.com
In-Reply-To: <20180116143836.GC22903@lunn.ch>
T24gVGh1LCAxOTcwLTAxLTAxIGF0IDAwOjAwICswMDAwLCBBbmRyZXcgTHVubiB3cm90ZToNCj4g
Q0FVVElPTjogVGhpcyBlbWFpbCBvcmlnaW5hdGVkIGZyb20gb3V0c2lkZSBvZiB0aGUgb3JnYW5p
emF0aW9uLiBEbyBub3QgY2xpY2sgbGlua3Mgb3Igb3BlbiBhdHRhY2htZW50cyB1bmxlc3MgeW91
IHJlY29nbml6ZSB0aGUgc2VuZGVyIGFuZCBrbm93IHRoZSBjb250ZW50IGlzIHNhZmUuDQo+IA0K
PiANCj4gPiBIaSwganVzdCBzYXcgdGhpcyBhbmQgdGhvdWdodCBvZiBhIHNtYWxsIHBhdGNoIEkg
anVzdCB3cm90ZSBmb3IgbWRpbyBidXMsIG8gaWRlYQ0KPiA+IGlmIGl0IGlzIHJlbGV2YW50IGJ1
dCBoZXJlIGdvZXM6DQo+ID4gDQo+ID4gRnJvbSBmZTBiOThkNTRhNzk3Nzk0ODI3MDA2NzYzMzFi
NGQxMGEwZjNjYWRhIE1vbiBTZXAgMTcgMDA6MDA6MDAgMjAwMQ0KPiA+IEZyb206IEpvYWtpbSBU
amVybmx1bmQgPGpvYWtpbS50amVybmx1bmRAaW5maW5lcmEuY29tPg0KPiA+IERhdGU6IFN1biwg
MTQgSmFuIDIwMTggMjE6Mjc6MjAgKzAxMDANCj4gPiBTdWJqZWN0OiBbUEFUQ0hdIG9mX21kaW9i
dXNfcmVnaXN0ZXI6IENvbnRpbnVlIGFmdGVyIGVycm9yDQo+ID4gDQo+ID4gb2ZfbWRpb2J1c19y
ZWdpc3RlciB1bnJlZ2lzdGVyIGl0c2VsZiBpZiBvbmUgcGh5IGZhaWxzIHRvIHJlZ2lzdGVyDQo+
ID4gd2hpY2ggaXMgYmFkIGZvciBzeXN0ZW0gaGF2aW5nIGFsbCBpdHMgUEhZcyBvbiB0aGUgc2Ft
ZSBNRElPIGJ1cy4NCj4gPiBKdXN0IGxvZyB0aGUgZXJyb3IgYW5kIGNvbnRpbnVlIHdpdGggdGhl
IHJlbWFpbmluZyBQSFlzIGluc3RlYWQuDQo+ID4gDQo+ID4gU2lnbmVkLW9mZi1ieTogSm9ha2lt
IFRqZXJubHVuZCA8am9ha2ltLnRqZXJubHVuZEBpbmZpbmVyYS5jb20+DQo+IA0KPiBIaSBKb2Fr
aW0NCj4gDQo+IFlvdSBhcHBlYXIgdG8gYmUgdXNpbmcgYW4gb2xkIGtlcm5lbC4gVGFrZSBhIGxv
b2sgYXQ6DQoNCk5vdCByZWFsbHksIEkgYW0gdXNpbmcgNC4xNC54IGFuZCBJIGRvbid0IHRoaW5r
IHRoYXQgaXMgb2xkLiBTZWVtcyBsaWtlIHRoaXMNCnBhdGNoIGhhc24ndCBiZWVuIHNlbnQgdG8g
NC4xNC54Lg0KDQpJIHdvbmRlciBpZiBJIG1pZ2h0IGJlIG1pc3Npbmcgc29tZXRoaW5nIGVsc2Us
IHdlIGp1c3QgbW92ZWQgdG8gNC4xNCBhbmQgbm90aWMgdGhhdCBhbGwNCm91ciBmaXhlZCBQSFlz
IGFyZSBub24gZnVuY3Rpb25pbmc6DQpmc2xfbWFjIGZmZTRlMjAwMC5ldGhlcm5ldDogRk1hbiBN
RU1BQw0KZnNsX21hYyBmZmU0ZTIwMDAuZXRoZXJuZXQ6IEZNYW4gTUFDIGFkZHJlc3M6IDAwOjA2
OjljOjBiOjA2OjIwDQpmc2xfbWFjIGRwYWEtZXRoZXJuZXQuMDogX19kZXZtX3JlcXVlc3RfbWVt
X3JlZ2lvbihtYWMpIGZhaWxlZA0KZnNsX21hYzogcHJvYmUgb2YgZHBhYS1ldGhlcm5ldC4wIGZh
aWxlZCB3aXRoIGVycm9yIC0xNg0KZnNsX21hYyBmZmU0ZTQwMDAuZXRoZXJuZXQ6IEZNYW4gTUVN
QUMNCmZzbF9tYWMgZmZlNGU0MDAwLmV0aGVybmV0OiBGTWFuIE1BQyBhZGRyZXNzOiAwMDowNjo5
YzowYjowNjoyMQ0KZnNsX21hYyBkcGFhLWV0aGVybmV0LjE6IF9fZGV2bV9yZXF1ZXN0X21lbV9y
ZWdpb24obWFjKSBmYWlsZWQNCmZzbF9tYWM6IHByb2JlIG9mIGRwYWEtZXRoZXJuZXQuMSBmYWls
ZWQgd2l0aCBlcnJvciAtMTYNCmZzbF9tYWMgZmZlNGU2MDAwLmV0aGVybmV0OiBGTWFuIE1FTUFD
DQpmc2xfbWFjIGZmZTRlNjAwMC5ldGhlcm5ldDogRk1hbiBNQUMgYWRkcmVzczogMDA6MDY6OWM6
MGI6MDY6MjINCmZzbF9tYWMgZHBhYS1ldGhlcm5ldC4yOiBfX2Rldm1fcmVxdWVzdF9tZW1fcmVn
aW9uKG1hYykgZmFpbGVkDQpmc2xfbWFjOiBwcm9iZSBvZiBkcGFhLWV0aGVybmV0LjIgZmFpbGVk
IHdpdGggZXJyb3IgLTE2DQpmc2xfbWFjIGZmZTRlODAwMC5ldGhlcm5ldDogRk1hbiBNRU1BQw0K
ZnNsX21hYyBmZmU0ZTgwMDAuZXRoZXJuZXQ6IEZNYW4gTUFDIGFkZHJlc3M6IDAwOjA2OjljOjBi
OjA2OjIzDQpmc2xfbWFjIGRwYWEtZXRoZXJuZXQuMzogX19kZXZtX3JlcXVlc3RfbWVtX3JlZ2lv
bihtYWMpIGZhaWxlZA0KZnNsX21hYzogcHJvYmUgb2YgZHBhYS1ldGhlcm5ldC4zIGZhaWxlZCB3
aXRoIGVycm9yIC0xNg0KDQpGZWVscyBsaWtlIEZNQU4gc3RpbGwgdGhpbmsgdGhlcmUgYXJlIHJl
YWwgUEhZcyB0aGVyZSA/DQo+IA0KPiBjb21taXQgOTVmNTY2ZGUwMjY5YTBjNTlmZDZhNzM3YTE0
NzczMTMwMjEzNjQyOQ0KPiBBdXRob3I6IE1hZGFsaW4gQnVjdXIgPG1hZGFsaW4uYnVjdXJAbnhw
LmNvbT4NCj4gRGF0ZTogICBUdWUgSmFuIDkgMTQ6NDM6MzQgMjAxOCArMDIwMA0KPiANCj4gICAg
IG9mX21kaW86IGF2b2lkIE1ESU8gYnVzIHJlbW92YWwgd2hlbiBhIFBIWSBpcyBtaXNzaW5nDQo+
IA0KPiAgICAgSWYgb25lIG9mIHRoZSBjaGlsZCBkZXZpY2VzIGlzIG1pc3NpbmcgdGhlIG9mX21k
aW9idXNfcmVnaXN0ZXJfcGh5KCkNCj4gICAgIGNhbGwgd2lsbCByZXR1cm4gLUVOT0RFVi4gV2hl
biBhIG1pc3NpbmcgZGV2aWNlIGlzIGVuY291bnRlcmVkIHRoZQ0KPiAgICAgcmVnaXN0cmF0aW9u
IG9mIHRoZSByZW1haW5pbmcgUEhZcyBpcyBzdG9wcGVkIGFuZCB0aGUgTURJTyBidXMgd2lsbA0K
PiAgICAgZmFpbCB0byByZWdpc3Rlci4gUHJvcGFnYXRlIGFsbCBlcnJvcnMgZXhjZXB0IEVOT0RF
ViB0byBhdm9pZCBpdC4NCj4gDQo+ICAgICBTaWduZWQtb2ZmLWJ5OiBNYWRhbGluIEJ1Y3VyIDxt
YWRhbGluLmJ1Y3VyQG54cC5jb20+DQo+ICAgICBSZXZpZXdlZC1ieTogQW5kcmV3IEx1bm4gPGFu
ZHJld0BsdW5uLmNoPg0KPiAgICAgU2lnbmVkLW9mZi1ieTogRGF2aWQgUy4gTWlsbGVyIDxkYXZl
bUBkYXZlbWxvZnQubmV0Pg0KPiANCj4gDQo+ICAgICBBbmRyZXcNCg==
^ permalink raw reply
* RE: DPAA Ethernet problems with mainstream Linux kernels
From: Madalin-cristian Bucur @ 2018-01-16 17:33 UTC (permalink / raw)
To: Darren Stevens
Cc: Jamie Krueger, linuxppc-dev@lists.ozlabs.org,
netdev@vger.kernel.org
In-Reply-To: <4b50ee845f3.281d6f13@auth.smtp.1and1.co.uk>
> -----Original Message-----
> From: Darren Stevens [mailto:darren@stevens-zone.net]
> Sent: Tuesday, January 16, 2018 12:40 AM
> To: Madalin-cristian Bucur <madalin.bucur@nxp.com>
> Cc: Jamie Krueger <jamie@bitbybitsoftwaregroup.com>; linuxppc-
> dev@lists.ozlabs.org; netdev@vger.kernel.org
> Subject: Re: DPAA Ethernet problems with mainstream Linux kernels
>=20
> Hello Madalin-cristian
>=20
> On 15/01/2018, Madalin-cristian Bucur wrote:
> >> > The device tree that you mention, cyrus_p5020.eth.dts is not found i=
n
> >> > the Linux kernel sources. The cyrus_p5020.dts file from the fsl ppc
> >> > device tree folder does not include the PHY information for the DPAA
> >> > interfaces. The problems that you experience may be caused by some
> >> > issues with the PHY configuration (i.e. internal delay).
> >> The cyrus_p5020.eth.dts is a modified version of the cyrus_p5020.dts,
> >> which of course was based off the original p5020ds.dts file. As you
> >> noted, the current cyrus_p5020.dts file is incomplete, and does not
> >> map the Ethernet connections properly.
>=20
> This is because the current linux kernel version of cyrus_p5020.dts
> includes 'p5020si-pre.dtsi' and 'p5020si-post.dtsi' include files, which
> orginally gave us working ethernet (when we used the SDK kernel) However
> at some point you moved the ethernet aliases from the board dts file to
> the p5020si-pre.dtsi file breaking the linkages for our board.
>=20
> cyrus-pre.dtsi is simply p5020si-pre.dtsi with only the correct aliases
> in.
>=20
> >> ** I have attached both the cyrus_p5020.eth.dts and cyrus-pre.dtsi
> >> =A0=A0=A0=A0 files with this email for comparison. Please let me know=
if you
> see
> >> =A0=A0=A0=A0 any corrections that should be made to either file.
> >
> > At a first glance they look fine to me.
>=20
> That's good to know.
>=20
> >> I have started testing along that line, using Wireshark to view the
> >> traffic on the X5000/20 itself, and from another machine connected
> >> on the same subnet. So far (as indicated by some details of in my
> >> initial email), I can see outgoing broadcast requests (for DHCP)
> >> being sent out from the X5000/20, and these requests are correctly
> >> constructed and visible outside the X5000/20.
> >>
> >> However, no responses to the DHCP broadcasts appear to reach
> >> to X5000/20's DPAA Ethernet. I will need to setup some further
> >> tests to determine if the DHCP server saw the requests and responded
> >> to them. (I assume the DHCP server is getting them, and responding,
> >> as I can always get a successful DHCP response to the X5000/20
> >> when using an add-on Ethernet PICe card on the same subnet).
>=20
> This matches what I see, the interface I have connected to the network
> shows an increasing number of transmitted packets, but no received ones.
>=20
> Jamie also noticed the following error in dmesg (right after the ethernet
> port becomes active)
>=20
> [ 4.112165] fsl_dpa dpaa-ethernet.0 eth0: Probed interface eth0
> [ 4.116616] fsl_dpa dpaa-ethernet.1 eth1: Probed interface eth1
> [ ... ]
> [ 106.501055] IPv6: ADDRCONF(NETDEV_UP): eth1: link is not ready
> [ 106.570944] IPv6: ADDRCONF(NETDEV_UP): eth1: link is not ready
> [ 106.605044] IPv6: ADDRCONF(NETDEV_UP): eth0: link is not ready
> [ 106.674918] IPv6: ADDRCONF(NETDEV_UP): eth0: link is not ready
> [ 108.702771] IPv6: ADDRCONF(NETDEV_CHANGE): eth0: link becomes ready
> [ 109.032798] fsl-pamu: pamu_av_isr: access violation interrupt
> [ 109.032806] fsl-pamu: pamu_av_isr: POES1=3D00000000
> [ 109.032808] fsl-pamu: pamu_av_isr: POES2=3D00000000
> [ 109.032809] fsl-pamu: pamu_av_isr: AVS1=3D002d0081
> [ 109.032811] fsl-pamu: pamu_av_isr: AVS2=3D00000081
> [ 109.032813] fsl-pamu: pamu_av_isr: AVA=3D00000001f1328000
> [ 109.032815] fsl-pamu: pamu_av_isr: UDAD=3D002d0001
> [ 109.032817] fsl-pamu: pamu_av_isr: POEA=3D0000000000000000
>=20
> I haven't seen this anywhere else, and wondered if it is relevant.
>=20
> Regards
> Darren
The PAMU related errors may be relevant to the issue, if you have incorrect
settings you may have no traffic passing through. The PAMU configuration
should be made by the bootloader. Can you try to disable CONFIG_FSL_PAMU?
Madalin
^ permalink raw reply
* [PATCH] [RESEND] powerpc: mpic_timer: avoid struct timeval
From: Arnd Bergmann @ 2018-01-16 17:01 UTC (permalink / raw)
To: Benjamin Herrenschmidt, Paul Mackerras, Michael Ellerman
Cc: Arnd Bergmann, Tyrel Datwyler, linuxppc-dev, linux-kernel
In an effort to remove all instances of 'struct timeval'
from the kernel, I'm changing the powerpc mpic_timer interface
to use plain seconds instead. There is only one user of this
interface, and that doesn't use the microseconds portion, so
the code gets noticeably simpler in the process.
Signed-off-by: Arnd Bergmann <arnd@arndb.de>
---
Submitted in November 2017, no reply, resending. Please apply.
---
arch/powerpc/include/asm/mpic_timer.h | 8 ++---
arch/powerpc/sysdev/fsl_mpic_timer_wakeup.c | 16 ++++-----
arch/powerpc/sysdev/mpic_timer.c | 55 ++++++-----------------------
3 files changed, 21 insertions(+), 58 deletions(-)
diff --git a/arch/powerpc/include/asm/mpic_timer.h b/arch/powerpc/include/asm/mpic_timer.h
index 0e23cd4ac8aa..13e6702ec458 100644
--- a/arch/powerpc/include/asm/mpic_timer.h
+++ b/arch/powerpc/include/asm/mpic_timer.h
@@ -29,17 +29,17 @@ struct mpic_timer {
#ifdef CONFIG_MPIC_TIMER
struct mpic_timer *mpic_request_timer(irq_handler_t fn, void *dev,
- const struct timeval *time);
+ time64_t time);
void mpic_start_timer(struct mpic_timer *handle);
void mpic_stop_timer(struct mpic_timer *handle);
-void mpic_get_remain_time(struct mpic_timer *handle, struct timeval *time);
+void mpic_get_remain_time(struct mpic_timer *handle, time64_t *time);
void mpic_free_timer(struct mpic_timer *handle);
#else
struct mpic_timer *mpic_request_timer(irq_handler_t fn, void *dev,
- const struct timeval *time) { return NULL; }
+ time64_t time) { return NULL; }
void mpic_start_timer(struct mpic_timer *handle) { }
void mpic_stop_timer(struct mpic_timer *handle) { }
-void mpic_get_remain_time(struct mpic_timer *handle, struct timeval *time) { }
+void mpic_get_remain_time(struct mpic_timer *handle, time64_t *time) { }
void mpic_free_timer(struct mpic_timer *handle) { }
#endif
diff --git a/arch/powerpc/sysdev/fsl_mpic_timer_wakeup.c b/arch/powerpc/sysdev/fsl_mpic_timer_wakeup.c
index 1707bf04dec6..94278e8af192 100644
--- a/arch/powerpc/sysdev/fsl_mpic_timer_wakeup.c
+++ b/arch/powerpc/sysdev/fsl_mpic_timer_wakeup.c
@@ -56,17 +56,16 @@ static ssize_t fsl_timer_wakeup_show(struct device *dev,
struct device_attribute *attr,
char *buf)
{
- struct timeval interval;
- int val = 0;
+ time64_t interval = 0;
mutex_lock(&sysfs_lock);
if (fsl_wakeup->timer) {
mpic_get_remain_time(fsl_wakeup->timer, &interval);
- val = interval.tv_sec + 1;
+ interval++;
}
mutex_unlock(&sysfs_lock);
- return sprintf(buf, "%d\n", val);
+ return sprintf(buf, "%lld\n", interval);
}
static ssize_t fsl_timer_wakeup_store(struct device *dev,
@@ -74,11 +73,10 @@ static ssize_t fsl_timer_wakeup_store(struct device *dev,
const char *buf,
size_t count)
{
- struct timeval interval;
+ time64_t interval;
int ret;
- interval.tv_usec = 0;
- if (kstrtol(buf, 0, &interval.tv_sec))
+ if (kstrtoll(buf, 0, &interval))
return -EINVAL;
mutex_lock(&sysfs_lock);
@@ -89,13 +87,13 @@ static ssize_t fsl_timer_wakeup_store(struct device *dev,
fsl_wakeup->timer = NULL;
}
- if (!interval.tv_sec) {
+ if (!interval) {
mutex_unlock(&sysfs_lock);
return count;
}
fsl_wakeup->timer = mpic_request_timer(fsl_mpic_timer_irq,
- fsl_wakeup, &interval);
+ fsl_wakeup, interval);
if (!fsl_wakeup->timer) {
mutex_unlock(&sysfs_lock);
return -EINVAL;
diff --git a/arch/powerpc/sysdev/mpic_timer.c b/arch/powerpc/sysdev/mpic_timer.c
index a418579591be..87e7c42777a8 100644
--- a/arch/powerpc/sysdev/mpic_timer.c
+++ b/arch/powerpc/sysdev/mpic_timer.c
@@ -47,9 +47,6 @@
#define MAX_TICKS_CASCADE (~0U)
#define TIMER_OFFSET(num) (1 << (TIMERS_PER_GROUP - 1 - num))
-/* tv_usec should be less than ONE_SECOND, otherwise use tv_sec */
-#define ONE_SECOND 1000000
-
struct timer_regs {
u32 gtccr;
u32 res0[3];
@@ -90,51 +87,23 @@ static struct cascade_priv cascade_timer[] = {
static LIST_HEAD(timer_group_list);
static void convert_ticks_to_time(struct timer_group_priv *priv,
- const u64 ticks, struct timeval *time)
+ const u64 ticks, time64_t *time)
{
- u64 tmp_sec;
-
- time->tv_sec = (__kernel_time_t)div_u64(ticks, priv->timerfreq);
- tmp_sec = (u64)time->tv_sec * (u64)priv->timerfreq;
-
- time->tv_usec = 0;
-
- if (tmp_sec <= ticks)
- time->tv_usec = (__kernel_suseconds_t)
- div_u64((ticks - tmp_sec) * 1000000, priv->timerfreq);
-
- return;
+ *time = (u64)div_u64(ticks, priv->timerfreq);
}
/* the time set by the user is converted to "ticks" */
static int convert_time_to_ticks(struct timer_group_priv *priv,
- const struct timeval *time, u64 *ticks)
+ time64_t time, u64 *ticks)
{
u64 max_value; /* prevent u64 overflow */
- u64 tmp = 0;
-
- u64 tmp_sec;
- u64 tmp_ms;
- u64 tmp_us;
max_value = div_u64(ULLONG_MAX, priv->timerfreq);
- if (time->tv_sec > max_value ||
- (time->tv_sec == max_value && time->tv_usec > 0))
+ if (time > max_value)
return -EINVAL;
- tmp_sec = (u64)time->tv_sec * (u64)priv->timerfreq;
- tmp += tmp_sec;
-
- tmp_ms = time->tv_usec / 1000;
- tmp_ms = div_u64((u64)tmp_ms * (u64)priv->timerfreq, 1000);
- tmp += tmp_ms;
-
- tmp_us = time->tv_usec % 1000;
- tmp_us = div_u64((u64)tmp_us * (u64)priv->timerfreq, 1000000);
- tmp += tmp_us;
-
- *ticks = tmp;
+ *ticks = (u64)time * (u64)priv->timerfreq;
return 0;
}
@@ -223,7 +192,7 @@ static struct mpic_timer *get_cascade_timer(struct timer_group_priv *priv,
return allocated_timer;
}
-static struct mpic_timer *get_timer(const struct timeval *time)
+static struct mpic_timer *get_timer(time64_t time)
{
struct timer_group_priv *priv;
struct mpic_timer *timer;
@@ -277,7 +246,7 @@ static struct mpic_timer *get_timer(const struct timeval *time)
* @handle: the timer to be started.
*
* It will do ->fn(->dev) callback from the hardware interrupt at
- * the ->timeval point in the future.
+ * the 'time64_t' point in the future.
*/
void mpic_start_timer(struct mpic_timer *handle)
{
@@ -319,7 +288,7 @@ EXPORT_SYMBOL(mpic_stop_timer);
*
* Query timer remaining time.
*/
-void mpic_get_remain_time(struct mpic_timer *handle, struct timeval *time)
+void mpic_get_remain_time(struct mpic_timer *handle, time64_t *time)
{
struct timer_group_priv *priv = container_of(handle,
struct timer_group_priv, timer[handle->num]);
@@ -391,7 +360,7 @@ EXPORT_SYMBOL(mpic_free_timer);
* else "handle" on success.
*/
struct mpic_timer *mpic_request_timer(irq_handler_t fn, void *dev,
- const struct timeval *time)
+ time64_t time)
{
struct mpic_timer *allocated_timer;
int ret;
@@ -399,11 +368,7 @@ struct mpic_timer *mpic_request_timer(irq_handler_t fn, void *dev,
if (list_empty(&timer_group_list))
return NULL;
- if (!(time->tv_sec + time->tv_usec) ||
- time->tv_sec < 0 || time->tv_usec < 0)
- return NULL;
-
- if (time->tv_usec > ONE_SECOND)
+ if (time < 0)
return NULL;
allocated_timer = get_timer(time);
--
2.9.0
^ permalink raw reply related
* RE: DPAA Ethernet traffice troubles with Linux kernel
From: Madalin-cristian Bucur @ 2018-01-16 17:07 UTC (permalink / raw)
To: Andrew Lunn, mad skateman
Cc: Christian Zigotzky, Joakim Tjernlund,
linuxppc-dev@lists.ozlabs.org, netdev@vger.kernel.org
In-Reply-To: <20180116150411.GG22903@lunn.ch>
> -----Original Message-----
> From: netdev-owner@vger.kernel.org [mailto:netdev-owner@vger.kernel.org]
> On Behalf Of Andrew Lunn
> Sent: Tuesday, January 16, 2018 5:04 PM
> To: mad skateman <madskateman@gmail.com>
> Cc: Christian Zigotzky <chzigotzky@xenosoft.de>; Joakim Tjernlund
> <Joakim.Tjernlund@infinera.com>; linuxppc-dev@lists.ozlabs.org; Madalin-
> cristian Bucur <madalin.bucur@nxp.com>; netdev@vger.kernel.org
> Subject: Re: DPAA Ethernet traffice troubles with Linux kernel
>=20
> > When i use mii-tool too Kick the tranceiver... it comes alive.. i can
> > ping the eth0 itself
> >
> > root@X5000LNX:/home/skateman# mii-tool -R eth0
> > resetting the transceiver...
> > root@X5000LNX:/home/skateman# ping 192.168.22.44
> > PING 192.168.22.44 (192.168.22.44) 56(84) bytes of data.
> > 64 bytes from 192.168.22.44: icmp_seq=3D1 ttl=3D64 time=3D0.045 ms
> > 64 bytes from 192.168.22.44: icmp_seq=3D2 ttl=3D64 time=3D0.046 ms
> > 64 bytes from 192.168.22.44: icmp_seq=3D3 ttl=3D64 time=3D0.047 ms
> > 64 bytes from 192.168.22.44: icmp_seq=3D4 ttl=3D64 time=3D0.048 ms
>=20
> What PHY driver are you using?
>=20
> This smells a bit like an RGMII-ID problem.
>=20
> Andrew
Hi Andrew,
>From another thread[1] on the same topic:
> I am not sure what PHY hardware/configuration you are using on the
> DS and RDB platforms, but I can confirm that AmigaONE X5000/20
> (Cyrus Motherboard with p5020 SoC), has dTSEC 4 and dTSEC 5
> wired to two Micrel KSZ9021RN Gigabit Ethernet PHYs, using the
> RGMII protocol.
[1] https://www.spinics.net/lists/netdev/msg478062.html
^ permalink raw reply
* [PATCH] [RESEND] spufs: use timespec64 for timestamps
From: Arnd Bergmann @ 2018-01-16 17:00 UTC (permalink / raw)
To: Jeremy Kerr, Arnd Bergmann, Benjamin Herrenschmidt,
Paul Mackerras, Michael Ellerman
Cc: Al Viro, Andrew Morton, linuxppc-dev, linux-kernel
The switch log prints the tv_sec portion of timespec as a 32-bit
number, while overflows in 2106. It also uses the timespec type,
which is safe on 64-bit architectures, but deprecated because
it causes overflows in 2038 elsewhere.
This changes it to timespec64 and printing a 64-bit number for
consistency.
Signed-off-by: Arnd Bergmann <arnd@arndb.de>
---
Submitted in November 2017, no reply, resending. Please apply.
---
arch/powerpc/platforms/cell/spufs/file.c | 6 +++---
arch/powerpc/platforms/cell/spufs/spufs.h | 2 +-
2 files changed, 4 insertions(+), 4 deletions(-)
diff --git a/arch/powerpc/platforms/cell/spufs/file.c b/arch/powerpc/platforms/cell/spufs/file.c
index fc7772c3d068..c1be486da899 100644
--- a/arch/powerpc/platforms/cell/spufs/file.c
+++ b/arch/powerpc/platforms/cell/spufs/file.c
@@ -2375,8 +2375,8 @@ static int switch_log_sprint(struct spu_context *ctx, char *tbuf, int n)
p = ctx->switch_log->log + ctx->switch_log->tail % SWITCH_LOG_BUFSIZE;
- return snprintf(tbuf, n, "%u.%09u %d %u %u %llu\n",
- (unsigned int) p->tstamp.tv_sec,
+ return snprintf(tbuf, n, "%llu.%09u %d %u %u %llu\n",
+ (unsigned long long) p->tstamp.tv_sec,
(unsigned int) p->tstamp.tv_nsec,
p->spu_id,
(unsigned int) p->type,
@@ -2499,7 +2499,7 @@ void spu_switch_log_notify(struct spu *spu, struct spu_context *ctx,
struct switch_log_entry *p;
p = ctx->switch_log->log + ctx->switch_log->head;
- ktime_get_ts(&p->tstamp);
+ ktime_get_ts64(&p->tstamp);
p->timebase = get_tb();
p->spu_id = spu ? spu->number : -1;
p->type = type;
diff --git a/arch/powerpc/platforms/cell/spufs/spufs.h b/arch/powerpc/platforms/cell/spufs/spufs.h
index 2d0479ad3af4..b5fc1b3fe538 100644
--- a/arch/powerpc/platforms/cell/spufs/spufs.h
+++ b/arch/powerpc/platforms/cell/spufs/spufs.h
@@ -69,7 +69,7 @@ struct switch_log {
unsigned long head;
unsigned long tail;
struct switch_log_entry {
- struct timespec tstamp;
+ struct timespec64 tstamp;
s32 spu_id;
u32 type;
u32 val;
--
2.9.0
^ permalink raw reply related
* Re: [PATCH 1/3] powerpc/32: Fix hugepage allocation on 8xx at hint address
From: Christophe LEROY @ 2018-01-16 16:57 UTC (permalink / raw)
To: Aneesh Kumar K.V, Benjamin Herrenschmidt, Paul Mackerras,
Michael Ellerman, Scott Wood, Nicholas Piggin
Cc: linux-kernel, linuxppc-dev
In-Reply-To: <945affcd-b25c-bc6e-68e5-8bbbcd31c0fd@linux.vnet.ibm.com>
Le 16/01/2018 à 17:41, Aneesh Kumar K.V a écrit :
>
>
> On 01/16/2018 10:01 PM, Christophe LEROY wrote:
>>
>>>> diff --git a/arch/powerpc/include/asm/page_64.h
>>>> b/arch/powerpc/include/asm/page_64.h
>>>> index 56234c6fcd61..a7baef5bbe5f 100644
>>>> --- a/arch/powerpc/include/asm/page_64.h
>>>> +++ b/arch/powerpc/include/asm/page_64.h
>>>> @@ -91,30 +91,13 @@ extern u64 ppc64_pft_size;
>>>> #define SLICE_LOW_SHIFT 28
>>>> #define SLICE_HIGH_SHIFT 40
>>>> -#define SLICE_LOW_TOP (0x100000000ul)
>>>> -#define SLICE_NUM_LOW (SLICE_LOW_TOP >> SLICE_LOW_SHIFT)
>>>> +#define SLICE_LOW_TOP (0xfffffffful)
>>>> +#define SLICE_NUM_LOW ((SLICE_LOW_TOP >> SLICE_LOW_SHIFT) + 1)
>>>> #define SLICE_NUM_HIGH (H_PGTABLE_RANGE >> SLICE_HIGH_SHIFT)
>>>
>>>
>>> Why are you changing this? is this a bug fix?
>>
>> That's because 0x100000000ul is out of range of unsigned long on PPC32.
>
> Ok that detail was important. I missed that.
>
>>
>>>
>>>> #define GET_LOW_SLICE_INDEX(addr) ((addr) >> SLICE_LOW_SHIFT)
>>>> #define GET_HIGH_SLICE_INDEX(addr) ((addr) >> SLICE_HIGH_SHIFT)
>>>> -#ifndef __ASSEMBLY__
>>>> -struct mm_struct;
>>>> -
>>>> -extern unsigned long slice_get_unmapped_area(unsigned long addr,
>>>> - unsigned long len,
>>>> - unsigned long flags,
>>>> - unsigned int psize,
>>>> - int topdown);
>>>> -
>>>> -extern unsigned int get_slice_psize(struct mm_struct *mm,
>>>> - unsigned long addr);
>>>> -
>>>> -extern void slice_set_user_psize(struct mm_struct *mm, unsigned int
>>>> psize);
>>>> -extern void slice_set_range_psize(struct mm_struct *mm, unsigned
>>>> long start,
>>>> - unsigned long len, unsigned int psize);
>>>> -
>>>> -#endif /* __ASSEMBLY__ */
>>>> #else
>>>> #define slice_init()
>>>> #ifdef CONFIG_PPC_BOOK3S_64
>>>> diff --git a/arch/powerpc/kernel/setup-common.c
>>>> b/arch/powerpc/kernel/setup-common.c
>>>> index 9d213542a48b..a285e1067713 100644
>>>> --- a/arch/powerpc/kernel/setup-common.c
>>>> +++ b/arch/powerpc/kernel/setup-common.c
>>>> @@ -928,7 +928,7 @@ void __init setup_arch(char **cmdline_p)
>>>> if (!radix_enabled())
>>>> init_mm.context.slb_addr_limit = DEFAULT_MAP_WINDOW_USER64;
>>>> #else
>>>> -#error "context.addr_limit not initialized."
>>>> + init_mm.context.slb_addr_limit = DEFAULT_MAP_WINDOW;
>>>> #endif
>>>
>>>
>>> May be put this within #ifdef 8XX and retain the error?
>>
>> Is this error really worth it ?
>> I wanted to avoid spreading too many #ifdef PPC_8xx, but ok I can do
>> that.
>>
>>>
>>>> #endif
>>>> diff --git a/arch/powerpc/mm/8xx_mmu.c b/arch/powerpc/mm/8xx_mmu.c
>>>> index f29212e40f40..0be77709446c 100644
>>>> --- a/arch/powerpc/mm/8xx_mmu.c
>>>> +++ b/arch/powerpc/mm/8xx_mmu.c
>>>> @@ -192,7 +192,7 @@ void set_context(unsigned long id, pgd_t *pgd)
>>>> mtspr(SPRN_M_TW, __pa(pgd) - offset);
>>>> /* Update context */
>>>> - mtspr(SPRN_M_CASID, id);
>>>> + mtspr(SPRN_M_CASID, id - 1);
>>>> /* sync */
>>>> mb();
>>>> }
>>>> diff --git a/arch/powerpc/mm/hash_utils_64.c
>>>> b/arch/powerpc/mm/hash_utils_64.c
>>>> index 655a5a9a183d..3266b3326088 100644
>>>> --- a/arch/powerpc/mm/hash_utils_64.c
>>>> +++ b/arch/powerpc/mm/hash_utils_64.c
>>>> @@ -1101,7 +1101,7 @@ static unsigned int get_paca_psize(unsigned
>>>> long addr)
>>>> unsigned char *hpsizes;
>>>> unsigned long index, mask_index;
>>>> - if (addr < SLICE_LOW_TOP) {
>>>> + if (addr <= SLICE_LOW_TOP) {
>>>
>>> If this is part of bug fix, please do it as part of seperate patch
>>> with details
>>
>> As explained above, in order to allow comparison to work on PPC32,
>> SLICE_LOW_TOP has to be 0xffffffff instead of 0x100000000
>>
>> How should I split in separate patches ? Something like ?
>> 1/ Slice support for PPC32 > 2/ Activate slice for 8xx
>
> Yes something like that. Will you be able to avoid that
> if (SLICE_NUM_HIGH) from the code? That makes the code ugly. Right now
> i don't have definite suggestion on what we could do though.
>
Could use #ifdefs instead, but in my mind it would be even more ugly.
I would have liked just doing nothing, but the issue is that at the
moment bitmap_xxx() functions are not prepared to handle bitmaps of size
zero. Should we try to change that ? Any chance to succeed ?
Christophe
^ permalink raw reply
* Re: [PATCH 1/3] powerpc/32: Fix hugepage allocation on 8xx at hint address
From: Christophe LEROY @ 2018-01-16 16:53 UTC (permalink / raw)
To: Aneesh Kumar K.V, Benjamin Herrenschmidt, Paul Mackerras,
Michael Ellerman, Scott Wood, Nicholas Piggin
Cc: linux-kernel, linuxppc-dev
In-Reply-To: <cc2106e9-6ba1-e937-3e19-99572fa014d2@linux.vnet.ibm.com>
Le 16/01/2018 à 17:43, Aneesh Kumar K.V a écrit :
>
>
> On 01/16/2018 10:01 PM, Christophe LEROY wrote:
>>
>>
>> Le 16/01/2018 à 16:49, Aneesh Kumar K.V a écrit :
>>> Christophe Leroy <christophe.leroy@c-s.fr> writes:
>>>
>>>> When an app has some regular pages allocated (e.g. see below) and tries
>>>> to mmap() a huge page at a hint address covered by the same PMD entry,
>>>> the kernel accepts the hint allthough the 8xx cannot handle different
>>>> page sizes in the same PMD entry.
>>>
>>>
>>> So that is a bug in get_unmapped_area function that you are using and
>>> you want to fix that by using the slice code. Can you describe here what
>>> the allocation restrictions are w.r.t 8xx? Do they have segments and
>>> base page size like hash64?
>>
>> I don't think it is a bug in get_unmapped_area() that is used by
>> default. It is that some HW do support mixing any page size in the
>> same page table (eg BOOK3E ?), but the 8xx doesn't.
>> In the 8xx, the page size is defined in the PGD entry, then all pages
>> defined in a given page table pointed by a PGD entry have the same size.
>>
>> So it is similar to segments if you consider each PGD entry as a kind
>> of segment
>>
>
> so IIUC, hugepd format encodes the page size details and that require us
> to ensure that all the address range mapped at that hupge_pd entry is of
> same page size? Hence we want to avoid mmap handing over an address in
> that range when we already have a hugetlb mapping in that range?
Exactly
And also avoid hugetlb_get_unmapped_area() accepting an hint address in
that range when we already have a regular mapping in that range.
Christophe
^ permalink raw reply
* Re: [PATCH v2] powerpc/mm: Fix growth direction for hugepages mmaps with slice
From: Christophe LEROY @ 2018-01-16 16:48 UTC (permalink / raw)
To: Aneesh Kumar K.V, Benjamin Herrenschmidt, Paul Mackerras,
Michael Ellerman, Scott Wood
Cc: linuxppc-dev, linux-kernel
In-Reply-To: <87wp0haizf.fsf@linux.vnet.ibm.com>
Le 16/01/2018 à 17:03, Aneesh Kumar K.V a écrit :
> Christophe Leroy <christophe.leroy@c-s.fr> writes:
>
>> An application running with libhugetlbfs fails to allocate
>> additional pages to HEAP due to the hugemap being done
>> inconditionally as topdown mapping:
>>
>> mmap(0x10080000, 1572864, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS|0x40000, -1, 0) = 0x73e80000
>> [...]
>> mmap(0x74000000, 1048576, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS|0x40000, -1, 0x180000) = 0x73d80000
>> munmap(0x73d80000, 1048576) = 0
>> [...]
>> mmap(0x74000000, 1572864, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS|0x40000, -1, 0x180000) = 0x73d00000
>> munmap(0x73d00000, 1572864) = 0
>> [...]
>> mmap(0x74000000, 1572864, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS|0x40000, -1, 0x180000) = 0x73d00000
>> munmap(0x73d00000, 1572864) = 0
>> [...]
>>
>
> Can you explain the failure details above. I am not sure I understand
> what to read from the above output.
libhugetlbfs first requests an area of size 1.5Mbytes, at address 0x10080000
mmap() returns an area at address 0x73e80000
Then libhugetlbfs requests an additional area on top of that, ie at
address 0x74000000, to expand the heap.
But mmap() returns an area at address 0x73d80000, ie under the previous
area.
This is not the behaviour when using the generic (ie without mm_slices)
hugepages code, and this is not what libhugetlbfs expects for expending
the heap.
>
>> As one can see from the above strace log, mmap() allocates further
>> pages below the initial one.
>>
>> This patch fixes it by taking into account MAP_GROWSDOWN flag.
>
> Rest of the kernel don't depend on that flag to select a topdown search
> or not. So what is special with hugetlb? IF we select legacy mmap that
> is when we select a bottomup search. Hugetlb on ppc64 always did a
> topdown search.
The generic hugepage code does a bottomup search. First page is
allocated at address 0x30000000 and following pages are allocated at
requested addresses when requested, then libhugetlbfs has no issue
expanding the heap when required.
>
>>
>> Fixes: d0f13e3c20b6f ("[POWERPC] Introduce address space "slices" ")
>> Signed-off-by: Christophe Leroy <christophe.leroy@c-s.fr>
>> ---
>> v2: Added missing include
>>
>> arch/powerpc/mm/hugetlbpage.c | 4 +++-
>> 1 file changed, 3 insertions(+), 1 deletion(-)
>>
>> diff --git a/arch/powerpc/mm/hugetlbpage.c b/arch/powerpc/mm/hugetlbpage.c
>> index 79e1378ee303..0eadf9f199de 100644
>> --- a/arch/powerpc/mm/hugetlbpage.c
>> +++ b/arch/powerpc/mm/hugetlbpage.c
>> @@ -19,6 +19,7 @@
>> #include <linux/moduleparam.h>
>> #include <linux/swap.h>
>> #include <linux/swapops.h>
>> +#include <linux/mman.h>
>> #include <asm/pgtable.h>
>> #include <asm/pgalloc.h>
>> #include <asm/tlb.h>
>> @@ -558,7 +559,8 @@ unsigned long hugetlb_get_unmapped_area(struct file *file, unsigned long addr,
>> return radix__hugetlb_get_unmapped_area(file, addr, len,
>> pgoff, flags);
>> #endif
>> - return slice_get_unmapped_area(addr, len, flags, mmu_psize, 1);
>> + return slice_get_unmapped_area(addr, len, flags, mmu_psize,
>> + flags & MAP_GROWSDOWN);
>> }
>> #endif
>>
>> --
>> 2.13.3
^ permalink raw reply
* Re: [PATCH 1/3] powerpc/32: Fix hugepage allocation on 8xx at hint address
From: Aneesh Kumar K.V @ 2018-01-16 16:43 UTC (permalink / raw)
To: Christophe LEROY, Benjamin Herrenschmidt, Paul Mackerras,
Michael Ellerman, Scott Wood, Nicholas Piggin
Cc: linux-kernel, linuxppc-dev
In-Reply-To: <b5e34a5b-94b8-9405-929d-d35dde3fcb3b@c-s.fr>
On 01/16/2018 10:01 PM, Christophe LEROY wrote:
>
>
> Le 16/01/2018 à 16:49, Aneesh Kumar K.V a écrit :
>> Christophe Leroy <christophe.leroy@c-s.fr> writes:
>>
>>> When an app has some regular pages allocated (e.g. see below) and tries
>>> to mmap() a huge page at a hint address covered by the same PMD entry,
>>> the kernel accepts the hint allthough the 8xx cannot handle different
>>> page sizes in the same PMD entry.
>>
>>
>> So that is a bug in get_unmapped_area function that you are using and
>> you want to fix that by using the slice code. Can you describe here what
>> the allocation restrictions are w.r.t 8xx? Do they have segments and
>> base page size like hash64?
>
> I don't think it is a bug in get_unmapped_area() that is used by
> default. It is that some HW do support mixing any page size in the same
> page table (eg BOOK3E ?), but the 8xx doesn't.
> In the 8xx, the page size is defined in the PGD entry, then all pages
> defined in a given page table pointed by a PGD entry have the same size.
>
> So it is similar to segments if you consider each PGD entry as a kind of
> segment
>
so IIUC, hugepd format encodes the page size details and that require us
to ensure that all the address range mapped at that hupge_pd entry is of
same page size? Hence we want to avoid mmap handing over an address in
that range when we already have a hugetlb mapping in that range?
-aneesh
^ permalink raw reply
* Re: [PATCH 1/3] powerpc/32: Fix hugepage allocation on 8xx at hint address
From: Aneesh Kumar K.V @ 2018-01-16 16:41 UTC (permalink / raw)
To: Christophe LEROY, Benjamin Herrenschmidt, Paul Mackerras,
Michael Ellerman, Scott Wood, Nicholas Piggin
Cc: linux-kernel, linuxppc-dev
In-Reply-To: <b5e34a5b-94b8-9405-929d-d35dde3fcb3b@c-s.fr>
On 01/16/2018 10:01 PM, Christophe LEROY wrote:
>
>>> diff --git a/arch/powerpc/include/asm/page_64.h
>>> b/arch/powerpc/include/asm/page_64.h
>>> index 56234c6fcd61..a7baef5bbe5f 100644
>>> --- a/arch/powerpc/include/asm/page_64.h
>>> +++ b/arch/powerpc/include/asm/page_64.h
>>> @@ -91,30 +91,13 @@ extern u64 ppc64_pft_size;
>>> #define SLICE_LOW_SHIFT 28
>>> #define SLICE_HIGH_SHIFT 40
>>> -#define SLICE_LOW_TOP (0x100000000ul)
>>> -#define SLICE_NUM_LOW (SLICE_LOW_TOP >> SLICE_LOW_SHIFT)
>>> +#define SLICE_LOW_TOP (0xfffffffful)
>>> +#define SLICE_NUM_LOW ((SLICE_LOW_TOP >> SLICE_LOW_SHIFT) + 1)
>>> #define SLICE_NUM_HIGH (H_PGTABLE_RANGE >> SLICE_HIGH_SHIFT)
>>
>>
>> Why are you changing this? is this a bug fix?
>
> That's because 0x100000000ul is out of range of unsigned long on PPC32.
Ok that detail was important. I missed that.
>
>>
>>> #define GET_LOW_SLICE_INDEX(addr) ((addr) >> SLICE_LOW_SHIFT)
>>> #define GET_HIGH_SLICE_INDEX(addr) ((addr) >> SLICE_HIGH_SHIFT)
>>> -#ifndef __ASSEMBLY__
>>> -struct mm_struct;
>>> -
>>> -extern unsigned long slice_get_unmapped_area(unsigned long addr,
>>> - unsigned long len,
>>> - unsigned long flags,
>>> - unsigned int psize,
>>> - int topdown);
>>> -
>>> -extern unsigned int get_slice_psize(struct mm_struct *mm,
>>> - unsigned long addr);
>>> -
>>> -extern void slice_set_user_psize(struct mm_struct *mm, unsigned int
>>> psize);
>>> -extern void slice_set_range_psize(struct mm_struct *mm, unsigned
>>> long start,
>>> - unsigned long len, unsigned int psize);
>>> -
>>> -#endif /* __ASSEMBLY__ */
>>> #else
>>> #define slice_init()
>>> #ifdef CONFIG_PPC_BOOK3S_64
>>> diff --git a/arch/powerpc/kernel/setup-common.c
>>> b/arch/powerpc/kernel/setup-common.c
>>> index 9d213542a48b..a285e1067713 100644
>>> --- a/arch/powerpc/kernel/setup-common.c
>>> +++ b/arch/powerpc/kernel/setup-common.c
>>> @@ -928,7 +928,7 @@ void __init setup_arch(char **cmdline_p)
>>> if (!radix_enabled())
>>> init_mm.context.slb_addr_limit = DEFAULT_MAP_WINDOW_USER64;
>>> #else
>>> -#error "context.addr_limit not initialized."
>>> + init_mm.context.slb_addr_limit = DEFAULT_MAP_WINDOW;
>>> #endif
>>
>>
>> May be put this within #ifdef 8XX and retain the error?
>
> Is this error really worth it ?
> I wanted to avoid spreading too many #ifdef PPC_8xx, but ok I can do that.
>
>>
>>> #endif
>>> diff --git a/arch/powerpc/mm/8xx_mmu.c b/arch/powerpc/mm/8xx_mmu.c
>>> index f29212e40f40..0be77709446c 100644
>>> --- a/arch/powerpc/mm/8xx_mmu.c
>>> +++ b/arch/powerpc/mm/8xx_mmu.c
>>> @@ -192,7 +192,7 @@ void set_context(unsigned long id, pgd_t *pgd)
>>> mtspr(SPRN_M_TW, __pa(pgd) - offset);
>>> /* Update context */
>>> - mtspr(SPRN_M_CASID, id);
>>> + mtspr(SPRN_M_CASID, id - 1);
>>> /* sync */
>>> mb();
>>> }
>>> diff --git a/arch/powerpc/mm/hash_utils_64.c
>>> b/arch/powerpc/mm/hash_utils_64.c
>>> index 655a5a9a183d..3266b3326088 100644
>>> --- a/arch/powerpc/mm/hash_utils_64.c
>>> +++ b/arch/powerpc/mm/hash_utils_64.c
>>> @@ -1101,7 +1101,7 @@ static unsigned int get_paca_psize(unsigned
>>> long addr)
>>> unsigned char *hpsizes;
>>> unsigned long index, mask_index;
>>> - if (addr < SLICE_LOW_TOP) {
>>> + if (addr <= SLICE_LOW_TOP) {
>>
>> If this is part of bug fix, please do it as part of seperate patch
>> with details
>
> As explained above, in order to allow comparison to work on PPC32,
> SLICE_LOW_TOP has to be 0xffffffff instead of 0x100000000
>
> How should I split in separate patches ? Something like ?
> 1/ Slice support for PPC32 > 2/ Activate slice for 8xx
Yes something like that. Will you be able to avoid that
if (SLICE_NUM_HIGH) from the code? That makes the code ugly. Right now
i don't have definite suggestion on what we could do though.
-aneesh
^ permalink raw reply
* Re: [PATCH 2/3] powerpc/mm: Allow more than 16 low slices
From: Christophe LEROY @ 2018-01-16 16:37 UTC (permalink / raw)
To: Aneesh Kumar K.V, Benjamin Herrenschmidt, Paul Mackerras,
Michael Ellerman, Scott Wood, Nicholas Piggin
Cc: linux-kernel, linuxppc-dev
In-Reply-To: <873735by2e.fsf@linux.vnet.ibm.com>
Le 16/01/2018 à 16:52, Aneesh Kumar K.V a écrit :
> Christophe Leroy <christophe.leroy@c-s.fr> writes:
>
>> While the implementation of the "slices" address space allows
>> a significant amount of high slices, it limits the number of
>> low slices to 16 due to the use of a single u64 low_slices element
>> in struct slice_mask.
>
> It is not really slice_mask. it is mm_context_t.low_slice_psize which is
> of type 64 and we need 4 bits per each slice to store the segment base
> page size details. What is that you want to achieve here. For book3s,
> we have 256MB segments upto 1TB and beyound that we use 1TB segments.
> But for 256MB segments in the range from 4G - 1TB they all use the same
> base page size because that is tracked by one slice in high slice.
>
> Can you state the 8xx requirement here?
On the 8xx, we need each slice to be one or a multiple of PGD entries.
In 4k page size mode, each PGD entry covers 4M. So it may be interesting
to have slices of size 4M.
In 16k page size mode, each PGD entry covers 64M, So it may be
interesting to have slices of size 64M.
>
>
>>
>> In order to override this limitation, this patch switches the
>> handling of low_slices to BITMAPs as done already for high_slices.
>>
>> Signed-off-by: Christophe Leroy <christophe.leroy@c-s.fr>
>> ---
>> arch/powerpc/include/asm/book3s/64/mmu.h | 2 +-
>> arch/powerpc/include/asm/mmu-8xx.h | 2 +-
>> arch/powerpc/include/asm/paca.h | 2 +-
>> arch/powerpc/kernel/paca.c | 3 +-
>> arch/powerpc/mm/hash_utils_64.c | 13 ++--
>> arch/powerpc/mm/slb_low.S | 8 ++-
>> arch/powerpc/mm/slice.c | 102 +++++++++++++++++--------------
>> 7 files changed, 73 insertions(+), 59 deletions(-)
>>
>> diff --git a/arch/powerpc/include/asm/book3s/64/mmu.h b/arch/powerpc/include/asm/book3s/64/mmu.h
>> index c9448e19847a..27e7e9732ea1 100644
>> --- a/arch/powerpc/include/asm/book3s/64/mmu.h
>> +++ b/arch/powerpc/include/asm/book3s/64/mmu.h
>> @@ -91,7 +91,7 @@ typedef struct {
>> struct npu_context *npu_context;
>>
>> #ifdef CONFIG_PPC_MM_SLICES
>> - u64 low_slices_psize; /* SLB page size encodings */
>> + unsigned char low_slices_psize[8]; /* SLB page size encodings */
>> unsigned char high_slices_psize[SLICE_ARRAY_SIZE];
>> unsigned long slb_addr_limit;
>> #else
>> diff --git a/arch/powerpc/include/asm/mmu-8xx.h b/arch/powerpc/include/asm/mmu-8xx.h
>> index 5f89b6010453..d669d0062da4 100644
>> --- a/arch/powerpc/include/asm/mmu-8xx.h
>> +++ b/arch/powerpc/include/asm/mmu-8xx.h
>> @@ -171,7 +171,7 @@ typedef struct {
>> unsigned long vdso_base;
>> #ifdef CONFIG_PPC_MM_SLICES
>> u16 user_psize; /* page size index */
>> - u64 low_slices_psize; /* page size encodings */
>> + unsigned char low_slices_psize[8]; /* 16 slices */
>> unsigned char high_slices_psize[0];
>> unsigned long slb_addr_limit;
>> #endif
>> diff --git a/arch/powerpc/include/asm/paca.h b/arch/powerpc/include/asm/paca.h
>> index 3892db93b837..612017054825 100644
>> --- a/arch/powerpc/include/asm/paca.h
>> +++ b/arch/powerpc/include/asm/paca.h
>> @@ -141,7 +141,7 @@ struct paca_struct {
>> #ifdef CONFIG_PPC_BOOK3S
>> mm_context_id_t mm_ctx_id;
>> #ifdef CONFIG_PPC_MM_SLICES
>> - u64 mm_ctx_low_slices_psize;
>> + unsigned char mm_ctx_low_slices_psize[8];
>> unsigned char mm_ctx_high_slices_psize[SLICE_ARRAY_SIZE];
>> unsigned long mm_ctx_slb_addr_limit;
>> #else
>> diff --git a/arch/powerpc/kernel/paca.c b/arch/powerpc/kernel/paca.c
>> index d6597038931d..8e1566bf82b8 100644
>> --- a/arch/powerpc/kernel/paca.c
>> +++ b/arch/powerpc/kernel/paca.c
>> @@ -264,7 +264,8 @@ void copy_mm_to_paca(struct mm_struct *mm)
>> #ifdef CONFIG_PPC_MM_SLICES
>> VM_BUG_ON(!mm->context.slb_addr_limit);
>> get_paca()->mm_ctx_slb_addr_limit = mm->context.slb_addr_limit;
>> - get_paca()->mm_ctx_low_slices_psize = context->low_slices_psize;
>> + memcpy(&get_paca()->mm_ctx_low_slices_psize,
>> + &context->low_slices_psize, sizeof(context->low_slices_psize));
>> memcpy(&get_paca()->mm_ctx_high_slices_psize,
>> &context->high_slices_psize, TASK_SLICE_ARRAY_SZ(mm));
>> #else /* CONFIG_PPC_MM_SLICES */
>> diff --git a/arch/powerpc/mm/hash_utils_64.c b/arch/powerpc/mm/hash_utils_64.c
>> index 3266b3326088..2f0c6b527a83 100644
>> --- a/arch/powerpc/mm/hash_utils_64.c
>> +++ b/arch/powerpc/mm/hash_utils_64.c
>> @@ -1097,19 +1097,18 @@ unsigned int hash_page_do_lazy_icache(unsigned int pp, pte_t pte, int trap)
>> #ifdef CONFIG_PPC_MM_SLICES
>> static unsigned int get_paca_psize(unsigned long addr)
>> {
>> - u64 lpsizes;
>> - unsigned char *hpsizes;
>> + unsigned char *psizes;
>> unsigned long index, mask_index;
>>
>> if (addr <= SLICE_LOW_TOP) {
>> - lpsizes = get_paca()->mm_ctx_low_slices_psize;
>> + psizes = get_paca()->mm_ctx_low_slices_psize;
>> index = GET_LOW_SLICE_INDEX(addr);
>> - return (lpsizes >> (index * 4)) & 0xF;
>> + } else {
>> + psizes = get_paca()->mm_ctx_high_slices_psize;
>> + index = GET_HIGH_SLICE_INDEX(addr);
>> }
>> - hpsizes = get_paca()->mm_ctx_high_slices_psize;
>> - index = GET_HIGH_SLICE_INDEX(addr);
>> mask_index = index & 0x1;
>> - return (hpsizes[index >> 1] >> (mask_index * 4)) & 0xF;
>> + return (psizes[index >> 1] >> (mask_index * 4)) & 0xF;
>> }
>>
>> #else
>> diff --git a/arch/powerpc/mm/slb_low.S b/arch/powerpc/mm/slb_low.S
>> index 2cf5ef3fc50d..2c7c717fd2ea 100644
>> --- a/arch/powerpc/mm/slb_low.S
>> +++ b/arch/powerpc/mm/slb_low.S
>> @@ -200,10 +200,12 @@ END_MMU_FTR_SECTION_IFCLR(MMU_FTR_1T_SEGMENT)
>> 5:
>> /*
>> * Handle lpsizes
>> - * r9 is get_paca()->context.low_slices_psize, r11 is index
>> + * r9 is get_paca()->context.low_slices_psize[index], r11 is mask_index
>> */
>> - ld r9,PACALOWSLICESPSIZE(r13)
>> - mr r11,r10
>> + srdi r11,r10,1 /* index */
>> + addi r9,r11,PACALOWSLICESPSIZE
>> + lbzx r9,r13,r9 /* r9 is lpsizes[r11] */
>> + rldicl r11,r10,0,63 /* r11 = r10 & 0x1 */
>> 6:
>> sldi r11,r11,2 /* index * 4 */
>> /* Extract the psize and multiply to get an array offset */
>> diff --git a/arch/powerpc/mm/slice.c b/arch/powerpc/mm/slice.c
>> index 1a66fafc3e45..e01ea72f21c6 100644
>> --- a/arch/powerpc/mm/slice.c
>> +++ b/arch/powerpc/mm/slice.c
>> @@ -43,7 +43,7 @@ static DEFINE_SPINLOCK(slice_convert_lock);
>> * in 1TB size.
>> */
>> struct slice_mask {
>> - u64 low_slices;
>> + DECLARE_BITMAP(low_slices, SLICE_NUM_LOW);
>> DECLARE_BITMAP(high_slices, SLICE_NUM_HIGH);
>> };
>>
>> @@ -54,7 +54,8 @@ static void slice_print_mask(const char *label, struct slice_mask mask)
>> {
>> if (!_slice_debug)
>> return;
>> - pr_devel("%s low_slice: %*pbl\n", label, (int)SLICE_NUM_LOW, &mask.low_slices);
>> + pr_devel("%s low_slice: %*pbl\n", label, (int)SLICE_NUM_LOW,
>> + mask.low_slices);
>> pr_devel("%s high_slice: %*pbl\n", label, (int)SLICE_NUM_HIGH, mask.high_slices);
>> }
>>
>> @@ -72,15 +73,18 @@ static void slice_range_to_mask(unsigned long start, unsigned long len,
>> {
>> unsigned long end = start + len - 1;
>>
>> - ret->low_slices = 0;
>> + bitmap_zero(ret->low_slices, SLICE_NUM_LOW);
>> if (SLICE_NUM_HIGH)
>> bitmap_zero(ret->high_slices, SLICE_NUM_HIGH);
>>
>> if (start <= SLICE_LOW_TOP) {
>> unsigned long mend = min(end, SLICE_LOW_TOP);
>> + unsigned long start_index = GET_LOW_SLICE_INDEX(start);
>> + unsigned long align_end = ALIGN(mend, (1UL << SLICE_LOW_SHIFT));
>> + unsigned long count = GET_LOW_SLICE_INDEX(align_end) -
>> + start_index;
>>
>> - ret->low_slices = (1u << (GET_LOW_SLICE_INDEX(mend) + 1))
>> - - (1u << GET_LOW_SLICE_INDEX(start));
>> + bitmap_set(ret->low_slices, start_index, count);
>> }
>>
>> if ((start + len) > SLICE_LOW_TOP) {
>> @@ -128,13 +132,13 @@ static void slice_mask_for_free(struct mm_struct *mm, struct slice_mask *ret,
>> {
>> unsigned long i;
>>
>> - ret->low_slices = 0;
>> + bitmap_zero(ret->low_slices, SLICE_NUM_LOW);
>> if (SLICE_NUM_HIGH)
>> bitmap_zero(ret->high_slices, SLICE_NUM_HIGH);
>>
>> for (i = 0; i < SLICE_NUM_LOW; i++)
>> if (!slice_low_has_vma(mm, i))
>> - ret->low_slices |= 1u << i;
>> + __set_bit(i, ret->low_slices);
>>
>> if (high_limit <= SLICE_LOW_TOP)
>> return;
>> @@ -147,19 +151,21 @@ static void slice_mask_for_free(struct mm_struct *mm, struct slice_mask *ret,
>> static void slice_mask_for_size(struct mm_struct *mm, int psize, struct slice_mask *ret,
>> unsigned long high_limit)
>> {
>> - unsigned char *hpsizes;
>> + unsigned char *hpsizes, *lpsizes;
>> int index, mask_index;
>> unsigned long i;
>> - u64 lpsizes;
>>
>> - ret->low_slices = 0;
>> + bitmap_zero(ret->low_slices, SLICE_NUM_LOW);
>> if (SLICE_NUM_HIGH)
>> bitmap_zero(ret->high_slices, SLICE_NUM_HIGH);
>>
>> lpsizes = mm->context.low_slices_psize;
>> - for (i = 0; i < SLICE_NUM_LOW; i++)
>> - if (((lpsizes >> (i * 4)) & 0xf) == psize)
>> - ret->low_slices |= 1u << i;
>> + for (i = 0; i < SLICE_NUM_LOW; i++) {
>> + mask_index = i & 0x1;
>> + index = i >> 1;
>> + if (((lpsizes[index] >> (mask_index * 4)) & 0xf) == psize)
>> + __set_bit(i, ret->low_slices);
>> + }
>>
>> if (high_limit <= SLICE_LOW_TOP)
>> return;
>> @@ -176,6 +182,7 @@ static void slice_mask_for_size(struct mm_struct *mm, int psize, struct slice_ma
>> static int slice_check_fit(struct mm_struct *mm,
>> struct slice_mask mask, struct slice_mask available)
>> {
>> + DECLARE_BITMAP(result_low, SLICE_NUM_LOW);
>> DECLARE_BITMAP(result, SLICE_NUM_HIGH);
>> /*
>> * Make sure we just do bit compare only to the max
>> @@ -183,11 +190,13 @@ static int slice_check_fit(struct mm_struct *mm,
>> */
>> unsigned long slice_count = GET_HIGH_SLICE_INDEX(mm->context.slb_addr_limit);
>>
>> + bitmap_and(result_low, mask.low_slices,
>> + available.low_slices, SLICE_NUM_LOW);
>> if (SLICE_NUM_HIGH)
>> bitmap_and(result, mask.high_slices,
>> available.high_slices, slice_count);
>>
>> - return (mask.low_slices & available.low_slices) == mask.low_slices &&
>> + return bitmap_equal(result_low, mask.low_slices, SLICE_NUM_LOW) &&
>> (!slice_count ||
>> bitmap_equal(result, mask.high_slices, slice_count));
>> }
>> @@ -213,8 +222,7 @@ static void slice_convert(struct mm_struct *mm, struct slice_mask mask, int psiz
>> {
>> int index, mask_index;
>> /* Write the new slice psize bits */
>> - unsigned char *hpsizes;
>> - u64 lpsizes;
>> + unsigned char *hpsizes, *lpsizes;
>> unsigned long i, flags;
>>
>> slice_dbg("slice_convert(mm=%p, psize=%d)\n", mm, psize);
>> @@ -226,13 +234,14 @@ static void slice_convert(struct mm_struct *mm, struct slice_mask mask, int psiz
>> spin_lock_irqsave(&slice_convert_lock, flags);
>>
>> lpsizes = mm->context.low_slices_psize;
>> - for (i = 0; i < SLICE_NUM_LOW; i++)
>> - if (mask.low_slices & (1u << i))
>> - lpsizes = (lpsizes & ~(0xful << (i * 4))) |
>> - (((unsigned long)psize) << (i * 4));
>> -
>> - /* Assign the value back */
>> - mm->context.low_slices_psize = lpsizes;
>> + for (i = 0; i < SLICE_NUM_LOW; i++) {
>> + mask_index = i & 0x1;
>> + index = i >> 1;
>> + if (test_bit(i, mask.low_slices))
>> + lpsizes[index] = (lpsizes[index] &
>> + ~(0xf << (mask_index * 4))) |
>> + (((unsigned long)psize) << (mask_index * 4));
>> + }
>>
>> hpsizes = mm->context.high_slices_psize;
>> for (i = 0; i < GET_HIGH_SLICE_INDEX(mm->context.slb_addr_limit); i++) {
>> @@ -269,7 +278,7 @@ static bool slice_scan_available(unsigned long addr,
>> if (addr <= SLICE_LOW_TOP) {
>> slice = GET_LOW_SLICE_INDEX(addr);
>> *boundary_addr = (slice + end) << SLICE_LOW_SHIFT;
>> - return !!(available.low_slices & (1u << slice));
>> + return !!test_bit(slice, available.low_slices);
>> } else {
>> slice = GET_HIGH_SLICE_INDEX(addr);
>> *boundary_addr = (slice + end) ?
>> @@ -397,7 +406,8 @@ static inline void slice_or_mask(struct slice_mask *dst, struct slice_mask *src)
>> {
>> DECLARE_BITMAP(result, SLICE_NUM_HIGH);
>>
>> - dst->low_slices |= src->low_slices;
>> + bitmap_or(dst->low_slices, dst->low_slices, src->low_slices,
>> + SLICE_NUM_LOW);
>> if (SLICE_NUM_HIGH) {
>> bitmap_or(result, dst->high_slices, src->high_slices,
>> SLICE_NUM_HIGH);
>> @@ -409,7 +419,8 @@ static inline void slice_andnot_mask(struct slice_mask *dst, struct slice_mask *
>> {
>> DECLARE_BITMAP(result, SLICE_NUM_HIGH);
>>
>> - dst->low_slices &= ~src->low_slices;
>> + bitmap_andnot(dst->low_slices, dst->low_slices, src->low_slices,
>> + SLICE_NUM_LOW);
>>
>> if (SLICE_NUM_HIGH) {
>> bitmap_andnot(result, dst->high_slices, src->high_slices,
>> @@ -464,16 +475,16 @@ unsigned long slice_get_unmapped_area(unsigned long addr, unsigned long len,
>> /*
>> * init different masks
>> */
>> - mask.low_slices = 0;
>> + bitmap_zero(mask.low_slices, SLICE_NUM_LOW);
>> if (SLICE_NUM_HIGH)
>> bitmap_zero(mask.high_slices, SLICE_NUM_HIGH);
>>
>> /* silence stupid warning */;
>> - potential_mask.low_slices = 0;
>> + bitmap_zero(potential_mask.low_slices, SLICE_NUM_LOW);
>> if (SLICE_NUM_HIGH)
>> bitmap_zero(potential_mask.high_slices, SLICE_NUM_HIGH);
>>
>> - compat_mask.low_slices = 0;
>> + bitmap_zero(compat_mask.low_slices, SLICE_NUM_LOW);
>> if (SLICE_NUM_HIGH)
>> bitmap_zero(compat_mask.high_slices, SLICE_NUM_HIGH);
>>
>> @@ -613,7 +624,7 @@ unsigned long slice_get_unmapped_area(unsigned long addr, unsigned long len,
>> convert:
>> slice_andnot_mask(&mask, &good_mask);
>> slice_andnot_mask(&mask, &compat_mask);
>> - if (mask.low_slices ||
>> + if (!bitmap_empty(mask.low_slices, SLICE_NUM_LOW) ||
>> (SLICE_NUM_HIGH &&
>> !bitmap_empty(mask.high_slices, SLICE_NUM_HIGH))) {
>> slice_convert(mm, mask, psize);
>> @@ -647,7 +658,7 @@ unsigned long arch_get_unmapped_area_topdown(struct file *filp,
>>
>> unsigned int get_slice_psize(struct mm_struct *mm, unsigned long addr)
>> {
>> - unsigned char *hpsizes;
>> + unsigned char *psizes;
>> int index, mask_index;
>>
>> /*
>> @@ -661,15 +672,14 @@ unsigned int get_slice_psize(struct mm_struct *mm, unsigned long addr)
>> #endif
>> }
>> if (addr <= SLICE_LOW_TOP) {
>> - u64 lpsizes;
>> - lpsizes = mm->context.low_slices_psize;
>> + psizes = mm->context.low_slices_psize;
>> index = GET_LOW_SLICE_INDEX(addr);
>> - return (lpsizes >> (index * 4)) & 0xf;
>> + } else {
>> + psizes = mm->context.high_slices_psize;
>> + index = GET_HIGH_SLICE_INDEX(addr);
>> }
>> - hpsizes = mm->context.high_slices_psize;
>> - index = GET_HIGH_SLICE_INDEX(addr);
>> mask_index = index & 0x1;
>> - return (hpsizes[index >> 1] >> (mask_index * 4)) & 0xf;
>> + return (psizes[index >> 1] >> (mask_index * 4)) & 0xf;
>> }
>> EXPORT_SYMBOL_GPL(get_slice_psize);
>>
>> @@ -690,8 +700,8 @@ EXPORT_SYMBOL_GPL(get_slice_psize);
>> void slice_set_user_psize(struct mm_struct *mm, unsigned int psize)
>> {
>> int index, mask_index;
>> - unsigned char *hpsizes;
>> - unsigned long flags, lpsizes;
>> + unsigned char *hpsizes, *lpsizes;
>> + unsigned long flags;
>> unsigned int old_psize;
>> int i;
>>
>> @@ -709,12 +719,14 @@ void slice_set_user_psize(struct mm_struct *mm, unsigned int psize)
>> wmb();
>>
>> lpsizes = mm->context.low_slices_psize;
>> - for (i = 0; i < SLICE_NUM_LOW; i++)
>> - if (((lpsizes >> (i * 4)) & 0xf) == old_psize)
>> - lpsizes = (lpsizes & ~(0xful << (i * 4))) |
>> - (((unsigned long)psize) << (i * 4));
>> - /* Assign the value back */
>> - mm->context.low_slices_psize = lpsizes;
>> + for (i = 0; i < SLICE_NUM_LOW; i++) {
>> + mask_index = i & 0x1;
>> + index = i >> 1;
>> + if (((lpsizes[index] >> (mask_index * 4)) & 0xf) == old_psize)
>> + lpsizes[index] = (lpsizes[index] &
>> + ~(0xf << (mask_index * 4))) |
>> + (((unsigned long)psize) << (mask_index * 4));
>> + }
>>
>> hpsizes = mm->context.high_slices_psize;
>> for (i = 0; i < SLICE_NUM_HIGH; i++) {
>> --
>> 2.13.3
^ permalink raw reply
* Re: [PATCH 1/3] powerpc/32: Fix hugepage allocation on 8xx at hint address
From: Christophe LEROY @ 2018-01-16 16:31 UTC (permalink / raw)
To: Aneesh Kumar K.V, Benjamin Herrenschmidt, Paul Mackerras,
Michael Ellerman, Scott Wood, Nicholas Piggin
Cc: linux-kernel, linuxppc-dev
In-Reply-To: <876081by7g.fsf@linux.vnet.ibm.com>
Le 16/01/2018 à 16:49, Aneesh Kumar K.V a écrit :
> Christophe Leroy <christophe.leroy@c-s.fr> writes:
>
>> When an app has some regular pages allocated (e.g. see below) and tries
>> to mmap() a huge page at a hint address covered by the same PMD entry,
>> the kernel accepts the hint allthough the 8xx cannot handle different
>> page sizes in the same PMD entry.
>
>
> So that is a bug in get_unmapped_area function that you are using and
> you want to fix that by using the slice code. Can you describe here what
> the allocation restrictions are w.r.t 8xx? Do they have segments and
> base page size like hash64?
I don't think it is a bug in get_unmapped_area() that is used by
default. It is that some HW do support mixing any page size in the same
page table (eg BOOK3E ?), but the 8xx doesn't.
In the 8xx, the page size is defined in the PGD entry, then all pages
defined in a given page table pointed by a PGD entry have the same size.
So it is similar to segments if you consider each PGD entry as a kind of
segment
>
>>
>> 10000000-10001000 r-xp 00000000 00:0f 2597 /root/malloc
>> 10010000-10011000 rwxp 00000000 00:0f 2597 /root/malloc
>>
>> mmap(0x10080000, 524288, PROT_READ|PROT_WRITE,
>> MAP_PRIVATE|MAP_ANONYMOUS|0x40000, -1, 0) = 0x10080000
>>
>> This results in the following warning, and the app remains forever in
>> do_page_fault()/hugetlb_fault()
>>
>> [162980.035629] WARNING: CPU: 0 PID: 2777 at arch/powerpc/mm/hugetlbpage.c:354 hugetlb_free_pgd_range+0xc8/0x1e4
>> [162980.035699] CPU: 0 PID: 2777 Comm: malloc Tainted: G W 4.14.6 #85
>> [162980.035744] task: c67e2c00 task.stack: c668e000
>> [162980.035783] NIP: c000fe18 LR: c00e1eec CTR: c00f90c0
>> [162980.035830] REGS: c668fc20 TRAP: 0700 Tainted: G W (4.14.6)
>> [162980.035854] MSR: 00029032 <EE,ME,IR,DR,RI> CR: 24044224 XER: 20000000
>> [162980.036003]
>> [162980.036003] GPR00: c00e1eec c668fcd0 c67e2c00 00000010 c6869410 10080000 00000000 77fb4000
>> [162980.036003] GPR08: ffff0001 0683c001 00000000 ffffff80 44028228 10018a34 00004008 418004fc
>> [162980.036003] GPR16: c668e000 00040100 c668e000 c06c0000 c668fe78 c668e000 c6835ba0 c668fd48
>> [162980.036003] GPR24: 00000000 73ffffff 74000000 00000001 77fb4000 100fffff 10100000 10100000
>> [162980.036743] NIP [c000fe18] hugetlb_free_pgd_range+0xc8/0x1e4
>> [162980.036839] LR [c00e1eec] free_pgtables+0x12c/0x150
>> [162980.036861] Call Trace:
>> [162980.036939] [c668fcd0] [c00f0774] unlink_anon_vmas+0x1c4/0x214 (unreliable)
>> [162980.037040] [c668fd10] [c00e1eec] free_pgtables+0x12c/0x150
>> [162980.037118] [c668fd40] [c00eabac] exit_mmap+0xe8/0x1b4
>> [162980.037210] [c668fda0] [c0019710] mmput.part.9+0x20/0xd8
>> [162980.037301] [c668fdb0] [c001ecb0] do_exit+0x1f0/0x93c
>> [162980.037386] [c668fe00] [c001f478] do_group_exit+0x40/0xcc
>> [162980.037479] [c668fe10] [c002a76c] get_signal+0x47c/0x614
>> [162980.037570] [c668fe70] [c0007840] do_signal+0x54/0x244
>> [162980.037654] [c668ff30] [c0007ae8] do_notify_resume+0x34/0x88
>> [162980.037744] [c668ff40] [c000dae8] do_user_signal+0x74/0xc4
>> [162980.037781] Instruction dump:
>> [162980.037821] 7fdff378 81370000 54a3463a 80890020 7d24182e 7c841a14 712a0004 4082ff94
>> [162980.038014] 2f890000 419e0010 712a0ff0 408200e0 <0fe00000> 54a9000a 7f984840 419d0094
>> [162980.038216] ---[ end trace c0ceeca8e7a5800a ]---
>> [162980.038754] BUG: non-zero nr_ptes on freeing mm: 1
>> [162985.363322] BUG: non-zero nr_ptes on freeing mm: -1
>>
>> In order to fix this, the address space "slices" implemented
>> for BOOK3S/64 is reused.
>>
>> This patch:
>> 1/ Modifies the "slices" implementation to support 32 bits CPUs,
>> based on using only the low slices.
>> 2/ Moves "slices" functions prototypes from page64.h to page.h
>> 3/ Modifies the context.id on the 8xx to be in the range [1:16]
>> instead of [0:15] in order to identify context.id == 0 as
>> not initialised contexts
>> 4/ Activates CONFIG_PPC_MM_SLICES when CONFIG_HUGETLB_PAGE is
>> selected for the 8xx
>>
>> Alltough we could in theory have as many slices as PMD entries, the current
>> slices implementation limits the number of low slices to 16.
>
> Can you explain this more?
As you commented in your other mail, mm_context_t.low_slice_psize which
is of type 64 and need 4 bits per each slice to store the segment base
page size details, so the maximum number of low slices is 64/4=16
>
>
>>
>> Fixes: 4b91428699477 ("powerpc/8xx: Implement support of hugepages")
>> Signed-off-by: Christophe Leroy <christophe.leroy@c-s.fr>
>> ---
>> arch/powerpc/include/asm/mmu-8xx.h | 6 ++++
>> arch/powerpc/include/asm/page.h | 14 ++++++++
>> arch/powerpc/include/asm/page_32.h | 19 +++++++++++
>> arch/powerpc/include/asm/page_64.h | 21 ++----------
>> arch/powerpc/kernel/setup-common.c | 2 +-
>> arch/powerpc/mm/8xx_mmu.c | 2 +-
>> arch/powerpc/mm/hash_utils_64.c | 2 +-
>> arch/powerpc/mm/hugetlbpage.c | 2 ++
>> arch/powerpc/mm/mmu_context_nohash.c | 11 +++++--
>> arch/powerpc/mm/slice.c | 58 +++++++++++++++++++++++-----------
>> arch/powerpc/platforms/Kconfig.cputype | 1 +
>> 11 files changed, 95 insertions(+), 43 deletions(-)
>>
>> diff --git a/arch/powerpc/include/asm/mmu-8xx.h b/arch/powerpc/include/asm/mmu-8xx.h
>> index 5bb3dbede41a..5f89b6010453 100644
>> --- a/arch/powerpc/include/asm/mmu-8xx.h
>> +++ b/arch/powerpc/include/asm/mmu-8xx.h
>> @@ -169,6 +169,12 @@ typedef struct {
>> unsigned int id;
>> unsigned int active;
>> unsigned long vdso_base;
>> +#ifdef CONFIG_PPC_MM_SLICES
>> + u16 user_psize; /* page size index */
>> + u64 low_slices_psize; /* page size encodings */
>> + unsigned char high_slices_psize[0];
>> + unsigned long slb_addr_limit;
>> +#endif
>> } mm_context_t;
>>
>> #define PHYS_IMMR_BASE (mfspr(SPRN_IMMR) & 0xfff80000)
>> diff --git a/arch/powerpc/include/asm/page.h b/arch/powerpc/include/asm/page.h
>> index 8da5d4c1cab2..d0384f9db9eb 100644
>> --- a/arch/powerpc/include/asm/page.h
>> +++ b/arch/powerpc/include/asm/page.h
>> @@ -342,6 +342,20 @@ typedef struct page *pgtable_t;
>> #endif
>> #endif
>>
>> +#ifdef CONFIG_PPC_MM_SLICES
>> +struct mm_struct;
>> +
>> +unsigned long slice_get_unmapped_area(unsigned long addr, unsigned long len,
>> + unsigned long flags, unsigned int psize,
>> + int topdown);
>> +
>> +unsigned int get_slice_psize(struct mm_struct *mm, unsigned long addr);
>> +
>> +void slice_set_user_psize(struct mm_struct *mm, unsigned int psize);
>> +void slice_set_range_psize(struct mm_struct *mm, unsigned long start,
>> + unsigned long len, unsigned int psize);
>> +#endif
>> +
>> #include <asm-generic/memory_model.h>
>> #endif /* __ASSEMBLY__ */
>>
>> diff --git a/arch/powerpc/include/asm/page_32.h b/arch/powerpc/include/asm/page_32.h
>> index 5c378e9b78c8..f7d1bd1183c8 100644
>> --- a/arch/powerpc/include/asm/page_32.h
>> +++ b/arch/powerpc/include/asm/page_32.h
>> @@ -60,4 +60,23 @@ extern void copy_page(void *to, void *from);
>>
>> #endif /* __ASSEMBLY__ */
>>
>> +#ifdef CONFIG_PPC_MM_SLICES
>> +
>> +#define SLICE_LOW_SHIFT 28
>> +#define SLICE_HIGH_SHIFT 0
>> +
>> +#define SLICE_LOW_TOP (0xfffffffful)
>> +#define SLICE_NUM_LOW ((SLICE_LOW_TOP >> SLICE_LOW_SHIFT) + 1)
>> +#define SLICE_NUM_HIGH 0ul
>> +
>> +#define GET_LOW_SLICE_INDEX(addr) ((addr) >> SLICE_LOW_SHIFT)
>> +#define GET_HIGH_SLICE_INDEX(addr) (addr & 0)
>> +
>> +#ifdef CONFIG_HUGETLB_PAGE
>> +#define HAVE_ARCH_HUGETLB_UNMAPPED_AREA
>> +#endif
>> +#define HAVE_ARCH_UNMAPPED_AREA
>> +#define HAVE_ARCH_UNMAPPED_AREA_TOPDOWN
>> +
>> +#endif
>> #endif /* _ASM_POWERPC_PAGE_32_H */
>> diff --git a/arch/powerpc/include/asm/page_64.h b/arch/powerpc/include/asm/page_64.h
>> index 56234c6fcd61..a7baef5bbe5f 100644
>> --- a/arch/powerpc/include/asm/page_64.h
>> +++ b/arch/powerpc/include/asm/page_64.h
>> @@ -91,30 +91,13 @@ extern u64 ppc64_pft_size;
>> #define SLICE_LOW_SHIFT 28
>> #define SLICE_HIGH_SHIFT 40
>>
>> -#define SLICE_LOW_TOP (0x100000000ul)
>> -#define SLICE_NUM_LOW (SLICE_LOW_TOP >> SLICE_LOW_SHIFT)
>> +#define SLICE_LOW_TOP (0xfffffffful)
>> +#define SLICE_NUM_LOW ((SLICE_LOW_TOP >> SLICE_LOW_SHIFT) + 1)
>> #define SLICE_NUM_HIGH (H_PGTABLE_RANGE >> SLICE_HIGH_SHIFT)
>
>
> Why are you changing this? is this a bug fix?
That's because 0x100000000ul is out of range of unsigned long on PPC32.
>
>>
>> #define GET_LOW_SLICE_INDEX(addr) ((addr) >> SLICE_LOW_SHIFT)
>> #define GET_HIGH_SLICE_INDEX(addr) ((addr) >> SLICE_HIGH_SHIFT)
>>
>> -#ifndef __ASSEMBLY__
>> -struct mm_struct;
>> -
>> -extern unsigned long slice_get_unmapped_area(unsigned long addr,
>> - unsigned long len,
>> - unsigned long flags,
>> - unsigned int psize,
>> - int topdown);
>> -
>> -extern unsigned int get_slice_psize(struct mm_struct *mm,
>> - unsigned long addr);
>> -
>> -extern void slice_set_user_psize(struct mm_struct *mm, unsigned int psize);
>> -extern void slice_set_range_psize(struct mm_struct *mm, unsigned long start,
>> - unsigned long len, unsigned int psize);
>> -
>> -#endif /* __ASSEMBLY__ */
>> #else
>> #define slice_init()
>> #ifdef CONFIG_PPC_BOOK3S_64
>> diff --git a/arch/powerpc/kernel/setup-common.c b/arch/powerpc/kernel/setup-common.c
>> index 9d213542a48b..a285e1067713 100644
>> --- a/arch/powerpc/kernel/setup-common.c
>> +++ b/arch/powerpc/kernel/setup-common.c
>> @@ -928,7 +928,7 @@ void __init setup_arch(char **cmdline_p)
>> if (!radix_enabled())
>> init_mm.context.slb_addr_limit = DEFAULT_MAP_WINDOW_USER64;
>> #else
>> -#error "context.addr_limit not initialized."
>> + init_mm.context.slb_addr_limit = DEFAULT_MAP_WINDOW;
>> #endif
>
>
> May be put this within #ifdef 8XX and retain the error?
Is this error really worth it ?
I wanted to avoid spreading too many #ifdef PPC_8xx, but ok I can do that.
>
>> #endif
>>
>> diff --git a/arch/powerpc/mm/8xx_mmu.c b/arch/powerpc/mm/8xx_mmu.c
>> index f29212e40f40..0be77709446c 100644
>> --- a/arch/powerpc/mm/8xx_mmu.c
>> +++ b/arch/powerpc/mm/8xx_mmu.c
>> @@ -192,7 +192,7 @@ void set_context(unsigned long id, pgd_t *pgd)
>> mtspr(SPRN_M_TW, __pa(pgd) - offset);
>>
>> /* Update context */
>> - mtspr(SPRN_M_CASID, id);
>> + mtspr(SPRN_M_CASID, id - 1);
>> /* sync */
>> mb();
>> }
>> diff --git a/arch/powerpc/mm/hash_utils_64.c b/arch/powerpc/mm/hash_utils_64.c
>> index 655a5a9a183d..3266b3326088 100644
>> --- a/arch/powerpc/mm/hash_utils_64.c
>> +++ b/arch/powerpc/mm/hash_utils_64.c
>> @@ -1101,7 +1101,7 @@ static unsigned int get_paca_psize(unsigned long addr)
>> unsigned char *hpsizes;
>> unsigned long index, mask_index;
>>
>> - if (addr < SLICE_LOW_TOP) {
>> + if (addr <= SLICE_LOW_TOP) {
>
> If this is part of bug fix, please do it as part of seperate patch with details
As explained above, in order to allow comparison to work on PPC32,
SLICE_LOW_TOP has to be 0xffffffff instead of 0x100000000
How should I split in separate patches ? Something like ?
1/ Slice support for PPC32
2/ Activate slice for 8xx
>
>
>> lpsizes = get_paca()->mm_ctx_low_slices_psize;
>> index = GET_LOW_SLICE_INDEX(addr);
>> return (lpsizes >> (index * 4)) & 0xF;
>> diff --git a/arch/powerpc/mm/hugetlbpage.c b/arch/powerpc/mm/hugetlbpage.c
>> index a9b9083c5e49..79e1378ee303 100644
>> --- a/arch/powerpc/mm/hugetlbpage.c
>> +++ b/arch/powerpc/mm/hugetlbpage.c
>> @@ -553,9 +553,11 @@ unsigned long hugetlb_get_unmapped_area(struct file *file, unsigned long addr,
>> struct hstate *hstate = hstate_file(file);
>> int mmu_psize = shift_to_mmu_psize(huge_page_shift(hstate));
>>
>> +#ifdef CONFIG_PPC_RADIX_MMU
>> if (radix_enabled())
>> return radix__hugetlb_get_unmapped_area(file, addr, len,
>> pgoff, flags);
>> +#endif
>> return slice_get_unmapped_area(addr, len, flags, mmu_psize, 1);
>> }
>> #endif
>> diff --git a/arch/powerpc/mm/mmu_context_nohash.c b/arch/powerpc/mm/mmu_context_nohash.c
>> index 4554d6527682..c1e1bf186871 100644
>> --- a/arch/powerpc/mm/mmu_context_nohash.c
>> +++ b/arch/powerpc/mm/mmu_context_nohash.c
>> @@ -331,6 +331,13 @@ int init_new_context(struct task_struct *t, struct mm_struct *mm)
>> {
>> pr_hard("initing context for mm @%p\n", mm);
>>
>> +#ifdef CONFIG_PPC_MM_SLICES
>> + if (!mm->context.slb_addr_limit)
>> + mm->context.slb_addr_limit = DEFAULT_MAP_WINDOW;
>> + if (!mm->context.id)
>> + slice_set_user_psize(mm, mmu_virtual_psize);
>> +#endif
>> +
>> mm->context.id = MMU_NO_CONTEXT;
>> mm->context.active = 0;
>> return 0;
>> @@ -428,8 +435,8 @@ void __init mmu_context_init(void)
>> * -- BenH
>> */
>> if (mmu_has_feature(MMU_FTR_TYPE_8xx)) {
>> - first_context = 0;
>> - last_context = 15;
>> + first_context = 1;
>> + last_context = 16;
>> no_selective_tlbil = true;
>> } else if (mmu_has_feature(MMU_FTR_TYPE_47x)) {
>> first_context = 1;
>> diff --git a/arch/powerpc/mm/slice.c b/arch/powerpc/mm/slice.c
>> index 23ec2c5e3b78..1a66fafc3e45 100644
>> --- a/arch/powerpc/mm/slice.c
>> +++ b/arch/powerpc/mm/slice.c
>> @@ -73,10 +73,11 @@ static void slice_range_to_mask(unsigned long start, unsigned long len,
>> unsigned long end = start + len - 1;
>>
>> ret->low_slices = 0;
>> - bitmap_zero(ret->high_slices, SLICE_NUM_HIGH);
>> + if (SLICE_NUM_HIGH)
>> + bitmap_zero(ret->high_slices, SLICE_NUM_HIGH);
>
> So you don't want to use high slices but just low slice? If so can you
> add that as a different patch which implements just that.
high slices are over 0xffffffff, so pointless on PPC32.
>
>
>>
>> - if (start < SLICE_LOW_TOP) {
>> - unsigned long mend = min(end, (SLICE_LOW_TOP - 1));
>> + if (start <= SLICE_LOW_TOP) {
>> + unsigned long mend = min(end, SLICE_LOW_TOP);
>>
>> ret->low_slices = (1u << (GET_LOW_SLICE_INDEX(mend) + 1))
>> - (1u << GET_LOW_SLICE_INDEX(start));
>> @@ -117,7 +118,7 @@ static int slice_high_has_vma(struct mm_struct *mm, unsigned long slice)
>> * of the high or low area bitmaps, the first high area starts
>> * at 4GB, not 0 */
>> if (start == 0)
>> - start = SLICE_LOW_TOP;
>> + start = SLICE_LOW_TOP + 1;
>>
>> return !slice_area_is_free(mm, start, end - start);
>> }
>> @@ -128,7 +129,8 @@ static void slice_mask_for_free(struct mm_struct *mm, struct slice_mask *ret,
>> unsigned long i;
>>
>> ret->low_slices = 0;
>> - bitmap_zero(ret->high_slices, SLICE_NUM_HIGH);
>> + if (SLICE_NUM_HIGH)
>> + bitmap_zero(ret->high_slices, SLICE_NUM_HIGH);
>>
>> for (i = 0; i < SLICE_NUM_LOW; i++)
>> if (!slice_low_has_vma(mm, i))
>> @@ -151,7 +153,8 @@ static void slice_mask_for_size(struct mm_struct *mm, int psize, struct slice_ma
>> u64 lpsizes;
>>
>> ret->low_slices = 0;
>> - bitmap_zero(ret->high_slices, SLICE_NUM_HIGH);
>> + if (SLICE_NUM_HIGH)
>> + bitmap_zero(ret->high_slices, SLICE_NUM_HIGH);
>>
>> lpsizes = mm->context.low_slices_psize;
>> for (i = 0; i < SLICE_NUM_LOW; i++)
>> @@ -180,15 +183,18 @@ static int slice_check_fit(struct mm_struct *mm,
>> */
>> unsigned long slice_count = GET_HIGH_SLICE_INDEX(mm->context.slb_addr_limit);
>>
>> - bitmap_and(result, mask.high_slices,
>> - available.high_slices, slice_count);
>> + if (SLICE_NUM_HIGH)
>> + bitmap_and(result, mask.high_slices,
>> + available.high_slices, slice_count);
>>
>> return (mask.low_slices & available.low_slices) == mask.low_slices &&
>> - bitmap_equal(result, mask.high_slices, slice_count);
>> + (!slice_count ||
>> + bitmap_equal(result, mask.high_slices, slice_count));
>> }
>>
>> static void slice_flush_segments(void *parm)
>> {
>> +#ifdef CONFIG_PPC_BOOK3S_64
>> struct mm_struct *mm = parm;
>> unsigned long flags;
>>
>> @@ -200,6 +206,7 @@ static void slice_flush_segments(void *parm)
>> local_irq_save(flags);
>> slb_flush_and_rebolt();
>> local_irq_restore(flags);
>> +#endif
>> }
>>
>> static void slice_convert(struct mm_struct *mm, struct slice_mask mask, int psize)
>> @@ -259,7 +266,7 @@ static bool slice_scan_available(unsigned long addr,
>> unsigned long *boundary_addr)
>> {
>> unsigned long slice;
>> - if (addr < SLICE_LOW_TOP) {
>> + if (addr <= SLICE_LOW_TOP) {
>> slice = GET_LOW_SLICE_INDEX(addr);
>> *boundary_addr = (slice + end) << SLICE_LOW_SHIFT;
>> return !!(available.low_slices & (1u << slice));
>> @@ -391,8 +398,11 @@ static inline void slice_or_mask(struct slice_mask *dst, struct slice_mask *src)
>> DECLARE_BITMAP(result, SLICE_NUM_HIGH);
>>
>> dst->low_slices |= src->low_slices;
>> - bitmap_or(result, dst->high_slices, src->high_slices, SLICE_NUM_HIGH);
>> - bitmap_copy(dst->high_slices, result, SLICE_NUM_HIGH);
>> + if (SLICE_NUM_HIGH) {
>> + bitmap_or(result, dst->high_slices, src->high_slices,
>> + SLICE_NUM_HIGH);
>> + bitmap_copy(dst->high_slices, result, SLICE_NUM_HIGH);
>> + }
>> }
>>
>> static inline void slice_andnot_mask(struct slice_mask *dst, struct slice_mask *src)
>> @@ -401,12 +411,17 @@ static inline void slice_andnot_mask(struct slice_mask *dst, struct slice_mask *
>>
>> dst->low_slices &= ~src->low_slices;
>>
>> - bitmap_andnot(result, dst->high_slices, src->high_slices, SLICE_NUM_HIGH);
>> - bitmap_copy(dst->high_slices, result, SLICE_NUM_HIGH);
>> + if (SLICE_NUM_HIGH) {
>> + bitmap_andnot(result, dst->high_slices, src->high_slices,
>> + SLICE_NUM_HIGH);
>> + bitmap_copy(dst->high_slices, result, SLICE_NUM_HIGH);
>> + }
>> }
>>
>> #ifdef CONFIG_PPC_64K_PAGES
>> #define MMU_PAGE_BASE MMU_PAGE_64K
>> +#elif defined(CONFIG_PPC_16K_PAGES)
>> +#define MMU_PAGE_BASE MMU_PAGE_16K
>> #else
>> #define MMU_PAGE_BASE MMU_PAGE_4K
>> #endif
>
> I am not sure we want them based on page size. The rule is we flush
> segments on book3s63 if the page size is different from MMU_PAGE_BASE
Do you mean this definition is just useless for the 8xx as it doesn't
have real segments ?
>
>
>> @@ -450,14 +465,17 @@ unsigned long slice_get_unmapped_area(unsigned long addr, unsigned long len,
>> * init different masks
>> */
>> mask.low_slices = 0;
>> - bitmap_zero(mask.high_slices, SLICE_NUM_HIGH);
>> + if (SLICE_NUM_HIGH)
>> + bitmap_zero(mask.high_slices, SLICE_NUM_HIGH);
>>
>> /* silence stupid warning */;
>> potential_mask.low_slices = 0;
>> - bitmap_zero(potential_mask.high_slices, SLICE_NUM_HIGH);
>> + if (SLICE_NUM_HIGH)
>> + bitmap_zero(potential_mask.high_slices, SLICE_NUM_HIGH);
>>
>> compat_mask.low_slices = 0;
>> - bitmap_zero(compat_mask.high_slices, SLICE_NUM_HIGH);
>> + if (SLICE_NUM_HIGH)
>> + bitmap_zero(compat_mask.high_slices, SLICE_NUM_HIGH);
>>
>> /* Sanity checks */
>> BUG_ON(mm->task_size == 0);
>> @@ -595,7 +613,9 @@ unsigned long slice_get_unmapped_area(unsigned long addr, unsigned long len,
>> convert:
>> slice_andnot_mask(&mask, &good_mask);
>> slice_andnot_mask(&mask, &compat_mask);
>> - if (mask.low_slices || !bitmap_empty(mask.high_slices, SLICE_NUM_HIGH)) {
>> + if (mask.low_slices ||
>> + (SLICE_NUM_HIGH &&
>> + !bitmap_empty(mask.high_slices, SLICE_NUM_HIGH))) {
>> slice_convert(mm, mask, psize);
>> if (psize > MMU_PAGE_BASE)
>> on_each_cpu(slice_flush_segments, mm, 1);
>> @@ -640,7 +660,7 @@ unsigned int get_slice_psize(struct mm_struct *mm, unsigned long addr)
>> return MMU_PAGE_4K;
>> #endif
>> }
>> - if (addr < SLICE_LOW_TOP) {
>> + if (addr <= SLICE_LOW_TOP) {
>> u64 lpsizes;
>> lpsizes = mm->context.low_slices_psize;
>> index = GET_LOW_SLICE_INDEX(addr);
>> diff --git a/arch/powerpc/platforms/Kconfig.cputype b/arch/powerpc/platforms/Kconfig.cputype
>> index ae07470fde3c..73a7ea333e9e 100644
>> --- a/arch/powerpc/platforms/Kconfig.cputype
>> +++ b/arch/powerpc/platforms/Kconfig.cputype
>> @@ -334,6 +334,7 @@ config PPC_BOOK3E_MMU
>> config PPC_MM_SLICES
>> bool
>> default y if PPC_BOOK3S_64
>> + default y if PPC_8xx && HUGETLB_PAGE
>> default n
>>
>> config PPC_HAVE_PMU_SUPPORT
>> --
>> 2.13.3
^ permalink raw reply
* Re: [PATCH 3/3] powerpc/8xx: Increase the number of mm slices
From: Christophe LEROY @ 2018-01-16 16:16 UTC (permalink / raw)
To: Aneesh Kumar K.V, Benjamin Herrenschmidt, Paul Mackerras,
Michael Ellerman, Scott Wood, Nicholas Piggin
Cc: linux-kernel, linuxppc-dev
In-Reply-To: <87zi5dajgj.fsf@linux.vnet.ibm.com>
Le 16/01/2018 à 16:53, Aneesh Kumar K.V a écrit :
> Christophe Leroy <christophe.leroy@c-s.fr> writes:
>
>> On the 8xx, we can have as many slices as PMD entries.
>> This means we could have 1024 slices in 4k size pages mode
>> and 64 slices in 16k size pages.
>>
>> However, due to a stack overflow in slice_get_unmapped_area(),
>> we limit to 512 slices.
>>
>> Signed-off-by: Christophe Leroy <christophe.leroy@c-s.fr>
>> ---
>> arch/powerpc/include/asm/mmu-8xx.h | 6 +++++-
>> arch/powerpc/include/asm/page_32.h | 3 ++-
>> 2 files changed, 7 insertions(+), 2 deletions(-)
>>
>> diff --git a/arch/powerpc/include/asm/mmu-8xx.h b/arch/powerpc/include/asm/mmu-8xx.h
>> index d669d0062da4..40aa7b0cd0dc 100644
>> --- a/arch/powerpc/include/asm/mmu-8xx.h
>> +++ b/arch/powerpc/include/asm/mmu-8xx.h
>> @@ -171,7 +171,11 @@ typedef struct {
>> unsigned long vdso_base;
>> #ifdef CONFIG_PPC_MM_SLICES
>> u16 user_psize; /* page size index */
>> - unsigned char low_slices_psize[8]; /* 16 slices */
>> +#if defined(CONFIG_PPC_16K_PAGES)
>> + unsigned char low_slices_psize[32]; /* 64 slices */
>> +#else
>> + unsigned char low_slices_psize[256]; /* 512 slices */
>> +#endif
>
> These #ifdef should be 8xx and then 16K.
We are in file asm/mmu-8xx.h so it obviously only applies to 8xx
Christophe
>
>
>> unsigned char high_slices_psize[0];
>> unsigned long slb_addr_limit;
>> #endif
>> diff --git a/arch/powerpc/include/asm/page_32.h b/arch/powerpc/include/asm/page_32.h
>> index f7d1bd1183c8..43695ce7ee07 100644
>> --- a/arch/powerpc/include/asm/page_32.h
>> +++ b/arch/powerpc/include/asm/page_32.h
>> @@ -62,7 +62,8 @@ extern void copy_page(void *to, void *from);
>>
>> #ifdef CONFIG_PPC_MM_SLICES
>>
>> -#define SLICE_LOW_SHIFT 28
>> +/* SLICE_LOW_SHIFT >= 23 to avoid stack overflow in slice_get_unmapped_area() */
>> +#define SLICE_LOW_SHIFT (PMD_SHIFT > 23 ? PMD_SHIFT : 23)
>> #define SLICE_HIGH_SHIFT 0
>>
>> #define SLICE_LOW_TOP (0xfffffffful)
>> --
>> 2.13.3
^ permalink raw reply
* Re: [PATCH 3/3] powerpc/8xx: Increase the number of mm slices
From: Aneesh Kumar K.V @ 2018-01-16 15:53 UTC (permalink / raw)
To: Christophe Leroy, Benjamin Herrenschmidt, Paul Mackerras,
Michael Ellerman, Scott Wood, Nicholas Piggin
Cc: linux-kernel, linuxppc-dev
In-Reply-To: <aed2799c8b369bd2ffba9967ed052e21d33e76f0.1515169256.git.christophe.leroy@c-s.fr>
Christophe Leroy <christophe.leroy@c-s.fr> writes:
> On the 8xx, we can have as many slices as PMD entries.
> This means we could have 1024 slices in 4k size pages mode
> and 64 slices in 16k size pages.
>
> However, due to a stack overflow in slice_get_unmapped_area(),
> we limit to 512 slices.
>
> Signed-off-by: Christophe Leroy <christophe.leroy@c-s.fr>
> ---
> arch/powerpc/include/asm/mmu-8xx.h | 6 +++++-
> arch/powerpc/include/asm/page_32.h | 3 ++-
> 2 files changed, 7 insertions(+), 2 deletions(-)
>
> diff --git a/arch/powerpc/include/asm/mmu-8xx.h b/arch/powerpc/include/asm/mmu-8xx.h
> index d669d0062da4..40aa7b0cd0dc 100644
> --- a/arch/powerpc/include/asm/mmu-8xx.h
> +++ b/arch/powerpc/include/asm/mmu-8xx.h
> @@ -171,7 +171,11 @@ typedef struct {
> unsigned long vdso_base;
> #ifdef CONFIG_PPC_MM_SLICES
> u16 user_psize; /* page size index */
> - unsigned char low_slices_psize[8]; /* 16 slices */
> +#if defined(CONFIG_PPC_16K_PAGES)
> + unsigned char low_slices_psize[32]; /* 64 slices */
> +#else
> + unsigned char low_slices_psize[256]; /* 512 slices */
> +#endif
These #ifdef should be 8xx and then 16K.
> unsigned char high_slices_psize[0];
> unsigned long slb_addr_limit;
> #endif
> diff --git a/arch/powerpc/include/asm/page_32.h b/arch/powerpc/include/asm/page_32.h
> index f7d1bd1183c8..43695ce7ee07 100644
> --- a/arch/powerpc/include/asm/page_32.h
> +++ b/arch/powerpc/include/asm/page_32.h
> @@ -62,7 +62,8 @@ extern void copy_page(void *to, void *from);
>
> #ifdef CONFIG_PPC_MM_SLICES
>
> -#define SLICE_LOW_SHIFT 28
> +/* SLICE_LOW_SHIFT >= 23 to avoid stack overflow in slice_get_unmapped_area() */
> +#define SLICE_LOW_SHIFT (PMD_SHIFT > 23 ? PMD_SHIFT : 23)
> #define SLICE_HIGH_SHIFT 0
>
> #define SLICE_LOW_TOP (0xfffffffful)
> --
> 2.13.3
^ permalink raw reply
* Re: [PATCH 2/3] powerpc/mm: Allow more than 16 low slices
From: Aneesh Kumar K.V @ 2018-01-16 15:52 UTC (permalink / raw)
To: Christophe Leroy, Benjamin Herrenschmidt, Paul Mackerras,
Michael Ellerman, Scott Wood, Nicholas Piggin
Cc: linux-kernel, linuxppc-dev
In-Reply-To: <eed37907f78d84c973434b90a137a351bc233651.1515169256.git.christophe.leroy@c-s.fr>
Christophe Leroy <christophe.leroy@c-s.fr> writes:
> While the implementation of the "slices" address space allows
> a significant amount of high slices, it limits the number of
> low slices to 16 due to the use of a single u64 low_slices element
> in struct slice_mask.
It is not really slice_mask. it is mm_context_t.low_slice_psize which is
of type 64 and we need 4 bits per each slice to store the segment base
page size details. What is that you want to achieve here. For book3s,
we have 256MB segments upto 1TB and beyound that we use 1TB segments.
But for 256MB segments in the range from 4G - 1TB they all use the same
base page size because that is tracked by one slice in high slice.
Can you state the 8xx requirement here?
>
> In order to override this limitation, this patch switches the
> handling of low_slices to BITMAPs as done already for high_slices.
>
> Signed-off-by: Christophe Leroy <christophe.leroy@c-s.fr>
> ---
> arch/powerpc/include/asm/book3s/64/mmu.h | 2 +-
> arch/powerpc/include/asm/mmu-8xx.h | 2 +-
> arch/powerpc/include/asm/paca.h | 2 +-
> arch/powerpc/kernel/paca.c | 3 +-
> arch/powerpc/mm/hash_utils_64.c | 13 ++--
> arch/powerpc/mm/slb_low.S | 8 ++-
> arch/powerpc/mm/slice.c | 102 +++++++++++++++++--------------
> 7 files changed, 73 insertions(+), 59 deletions(-)
>
> diff --git a/arch/powerpc/include/asm/book3s/64/mmu.h b/arch/powerpc/include/asm/book3s/64/mmu.h
> index c9448e19847a..27e7e9732ea1 100644
> --- a/arch/powerpc/include/asm/book3s/64/mmu.h
> +++ b/arch/powerpc/include/asm/book3s/64/mmu.h
> @@ -91,7 +91,7 @@ typedef struct {
> struct npu_context *npu_context;
>
> #ifdef CONFIG_PPC_MM_SLICES
> - u64 low_slices_psize; /* SLB page size encodings */
> + unsigned char low_slices_psize[8]; /* SLB page size encodings */
> unsigned char high_slices_psize[SLICE_ARRAY_SIZE];
> unsigned long slb_addr_limit;
> #else
> diff --git a/arch/powerpc/include/asm/mmu-8xx.h b/arch/powerpc/include/asm/mmu-8xx.h
> index 5f89b6010453..d669d0062da4 100644
> --- a/arch/powerpc/include/asm/mmu-8xx.h
> +++ b/arch/powerpc/include/asm/mmu-8xx.h
> @@ -171,7 +171,7 @@ typedef struct {
> unsigned long vdso_base;
> #ifdef CONFIG_PPC_MM_SLICES
> u16 user_psize; /* page size index */
> - u64 low_slices_psize; /* page size encodings */
> + unsigned char low_slices_psize[8]; /* 16 slices */
> unsigned char high_slices_psize[0];
> unsigned long slb_addr_limit;
> #endif
> diff --git a/arch/powerpc/include/asm/paca.h b/arch/powerpc/include/asm/paca.h
> index 3892db93b837..612017054825 100644
> --- a/arch/powerpc/include/asm/paca.h
> +++ b/arch/powerpc/include/asm/paca.h
> @@ -141,7 +141,7 @@ struct paca_struct {
> #ifdef CONFIG_PPC_BOOK3S
> mm_context_id_t mm_ctx_id;
> #ifdef CONFIG_PPC_MM_SLICES
> - u64 mm_ctx_low_slices_psize;
> + unsigned char mm_ctx_low_slices_psize[8];
> unsigned char mm_ctx_high_slices_psize[SLICE_ARRAY_SIZE];
> unsigned long mm_ctx_slb_addr_limit;
> #else
> diff --git a/arch/powerpc/kernel/paca.c b/arch/powerpc/kernel/paca.c
> index d6597038931d..8e1566bf82b8 100644
> --- a/arch/powerpc/kernel/paca.c
> +++ b/arch/powerpc/kernel/paca.c
> @@ -264,7 +264,8 @@ void copy_mm_to_paca(struct mm_struct *mm)
> #ifdef CONFIG_PPC_MM_SLICES
> VM_BUG_ON(!mm->context.slb_addr_limit);
> get_paca()->mm_ctx_slb_addr_limit = mm->context.slb_addr_limit;
> - get_paca()->mm_ctx_low_slices_psize = context->low_slices_psize;
> + memcpy(&get_paca()->mm_ctx_low_slices_psize,
> + &context->low_slices_psize, sizeof(context->low_slices_psize));
> memcpy(&get_paca()->mm_ctx_high_slices_psize,
> &context->high_slices_psize, TASK_SLICE_ARRAY_SZ(mm));
> #else /* CONFIG_PPC_MM_SLICES */
> diff --git a/arch/powerpc/mm/hash_utils_64.c b/arch/powerpc/mm/hash_utils_64.c
> index 3266b3326088..2f0c6b527a83 100644
> --- a/arch/powerpc/mm/hash_utils_64.c
> +++ b/arch/powerpc/mm/hash_utils_64.c
> @@ -1097,19 +1097,18 @@ unsigned int hash_page_do_lazy_icache(unsigned int pp, pte_t pte, int trap)
> #ifdef CONFIG_PPC_MM_SLICES
> static unsigned int get_paca_psize(unsigned long addr)
> {
> - u64 lpsizes;
> - unsigned char *hpsizes;
> + unsigned char *psizes;
> unsigned long index, mask_index;
>
> if (addr <= SLICE_LOW_TOP) {
> - lpsizes = get_paca()->mm_ctx_low_slices_psize;
> + psizes = get_paca()->mm_ctx_low_slices_psize;
> index = GET_LOW_SLICE_INDEX(addr);
> - return (lpsizes >> (index * 4)) & 0xF;
> + } else {
> + psizes = get_paca()->mm_ctx_high_slices_psize;
> + index = GET_HIGH_SLICE_INDEX(addr);
> }
> - hpsizes = get_paca()->mm_ctx_high_slices_psize;
> - index = GET_HIGH_SLICE_INDEX(addr);
> mask_index = index & 0x1;
> - return (hpsizes[index >> 1] >> (mask_index * 4)) & 0xF;
> + return (psizes[index >> 1] >> (mask_index * 4)) & 0xF;
> }
>
> #else
> diff --git a/arch/powerpc/mm/slb_low.S b/arch/powerpc/mm/slb_low.S
> index 2cf5ef3fc50d..2c7c717fd2ea 100644
> --- a/arch/powerpc/mm/slb_low.S
> +++ b/arch/powerpc/mm/slb_low.S
> @@ -200,10 +200,12 @@ END_MMU_FTR_SECTION_IFCLR(MMU_FTR_1T_SEGMENT)
> 5:
> /*
> * Handle lpsizes
> - * r9 is get_paca()->context.low_slices_psize, r11 is index
> + * r9 is get_paca()->context.low_slices_psize[index], r11 is mask_index
> */
> - ld r9,PACALOWSLICESPSIZE(r13)
> - mr r11,r10
> + srdi r11,r10,1 /* index */
> + addi r9,r11,PACALOWSLICESPSIZE
> + lbzx r9,r13,r9 /* r9 is lpsizes[r11] */
> + rldicl r11,r10,0,63 /* r11 = r10 & 0x1 */
> 6:
> sldi r11,r11,2 /* index * 4 */
> /* Extract the psize and multiply to get an array offset */
> diff --git a/arch/powerpc/mm/slice.c b/arch/powerpc/mm/slice.c
> index 1a66fafc3e45..e01ea72f21c6 100644
> --- a/arch/powerpc/mm/slice.c
> +++ b/arch/powerpc/mm/slice.c
> @@ -43,7 +43,7 @@ static DEFINE_SPINLOCK(slice_convert_lock);
> * in 1TB size.
> */
> struct slice_mask {
> - u64 low_slices;
> + DECLARE_BITMAP(low_slices, SLICE_NUM_LOW);
> DECLARE_BITMAP(high_slices, SLICE_NUM_HIGH);
> };
>
> @@ -54,7 +54,8 @@ static void slice_print_mask(const char *label, struct slice_mask mask)
> {
> if (!_slice_debug)
> return;
> - pr_devel("%s low_slice: %*pbl\n", label, (int)SLICE_NUM_LOW, &mask.low_slices);
> + pr_devel("%s low_slice: %*pbl\n", label, (int)SLICE_NUM_LOW,
> + mask.low_slices);
> pr_devel("%s high_slice: %*pbl\n", label, (int)SLICE_NUM_HIGH, mask.high_slices);
> }
>
> @@ -72,15 +73,18 @@ static void slice_range_to_mask(unsigned long start, unsigned long len,
> {
> unsigned long end = start + len - 1;
>
> - ret->low_slices = 0;
> + bitmap_zero(ret->low_slices, SLICE_NUM_LOW);
> if (SLICE_NUM_HIGH)
> bitmap_zero(ret->high_slices, SLICE_NUM_HIGH);
>
> if (start <= SLICE_LOW_TOP) {
> unsigned long mend = min(end, SLICE_LOW_TOP);
> + unsigned long start_index = GET_LOW_SLICE_INDEX(start);
> + unsigned long align_end = ALIGN(mend, (1UL << SLICE_LOW_SHIFT));
> + unsigned long count = GET_LOW_SLICE_INDEX(align_end) -
> + start_index;
>
> - ret->low_slices = (1u << (GET_LOW_SLICE_INDEX(mend) + 1))
> - - (1u << GET_LOW_SLICE_INDEX(start));
> + bitmap_set(ret->low_slices, start_index, count);
> }
>
> if ((start + len) > SLICE_LOW_TOP) {
> @@ -128,13 +132,13 @@ static void slice_mask_for_free(struct mm_struct *mm, struct slice_mask *ret,
> {
> unsigned long i;
>
> - ret->low_slices = 0;
> + bitmap_zero(ret->low_slices, SLICE_NUM_LOW);
> if (SLICE_NUM_HIGH)
> bitmap_zero(ret->high_slices, SLICE_NUM_HIGH);
>
> for (i = 0; i < SLICE_NUM_LOW; i++)
> if (!slice_low_has_vma(mm, i))
> - ret->low_slices |= 1u << i;
> + __set_bit(i, ret->low_slices);
>
> if (high_limit <= SLICE_LOW_TOP)
> return;
> @@ -147,19 +151,21 @@ static void slice_mask_for_free(struct mm_struct *mm, struct slice_mask *ret,
> static void slice_mask_for_size(struct mm_struct *mm, int psize, struct slice_mask *ret,
> unsigned long high_limit)
> {
> - unsigned char *hpsizes;
> + unsigned char *hpsizes, *lpsizes;
> int index, mask_index;
> unsigned long i;
> - u64 lpsizes;
>
> - ret->low_slices = 0;
> + bitmap_zero(ret->low_slices, SLICE_NUM_LOW);
> if (SLICE_NUM_HIGH)
> bitmap_zero(ret->high_slices, SLICE_NUM_HIGH);
>
> lpsizes = mm->context.low_slices_psize;
> - for (i = 0; i < SLICE_NUM_LOW; i++)
> - if (((lpsizes >> (i * 4)) & 0xf) == psize)
> - ret->low_slices |= 1u << i;
> + for (i = 0; i < SLICE_NUM_LOW; i++) {
> + mask_index = i & 0x1;
> + index = i >> 1;
> + if (((lpsizes[index] >> (mask_index * 4)) & 0xf) == psize)
> + __set_bit(i, ret->low_slices);
> + }
>
> if (high_limit <= SLICE_LOW_TOP)
> return;
> @@ -176,6 +182,7 @@ static void slice_mask_for_size(struct mm_struct *mm, int psize, struct slice_ma
> static int slice_check_fit(struct mm_struct *mm,
> struct slice_mask mask, struct slice_mask available)
> {
> + DECLARE_BITMAP(result_low, SLICE_NUM_LOW);
> DECLARE_BITMAP(result, SLICE_NUM_HIGH);
> /*
> * Make sure we just do bit compare only to the max
> @@ -183,11 +190,13 @@ static int slice_check_fit(struct mm_struct *mm,
> */
> unsigned long slice_count = GET_HIGH_SLICE_INDEX(mm->context.slb_addr_limit);
>
> + bitmap_and(result_low, mask.low_slices,
> + available.low_slices, SLICE_NUM_LOW);
> if (SLICE_NUM_HIGH)
> bitmap_and(result, mask.high_slices,
> available.high_slices, slice_count);
>
> - return (mask.low_slices & available.low_slices) == mask.low_slices &&
> + return bitmap_equal(result_low, mask.low_slices, SLICE_NUM_LOW) &&
> (!slice_count ||
> bitmap_equal(result, mask.high_slices, slice_count));
> }
> @@ -213,8 +222,7 @@ static void slice_convert(struct mm_struct *mm, struct slice_mask mask, int psiz
> {
> int index, mask_index;
> /* Write the new slice psize bits */
> - unsigned char *hpsizes;
> - u64 lpsizes;
> + unsigned char *hpsizes, *lpsizes;
> unsigned long i, flags;
>
> slice_dbg("slice_convert(mm=%p, psize=%d)\n", mm, psize);
> @@ -226,13 +234,14 @@ static void slice_convert(struct mm_struct *mm, struct slice_mask mask, int psiz
> spin_lock_irqsave(&slice_convert_lock, flags);
>
> lpsizes = mm->context.low_slices_psize;
> - for (i = 0; i < SLICE_NUM_LOW; i++)
> - if (mask.low_slices & (1u << i))
> - lpsizes = (lpsizes & ~(0xful << (i * 4))) |
> - (((unsigned long)psize) << (i * 4));
> -
> - /* Assign the value back */
> - mm->context.low_slices_psize = lpsizes;
> + for (i = 0; i < SLICE_NUM_LOW; i++) {
> + mask_index = i & 0x1;
> + index = i >> 1;
> + if (test_bit(i, mask.low_slices))
> + lpsizes[index] = (lpsizes[index] &
> + ~(0xf << (mask_index * 4))) |
> + (((unsigned long)psize) << (mask_index * 4));
> + }
>
> hpsizes = mm->context.high_slices_psize;
> for (i = 0; i < GET_HIGH_SLICE_INDEX(mm->context.slb_addr_limit); i++) {
> @@ -269,7 +278,7 @@ static bool slice_scan_available(unsigned long addr,
> if (addr <= SLICE_LOW_TOP) {
> slice = GET_LOW_SLICE_INDEX(addr);
> *boundary_addr = (slice + end) << SLICE_LOW_SHIFT;
> - return !!(available.low_slices & (1u << slice));
> + return !!test_bit(slice, available.low_slices);
> } else {
> slice = GET_HIGH_SLICE_INDEX(addr);
> *boundary_addr = (slice + end) ?
> @@ -397,7 +406,8 @@ static inline void slice_or_mask(struct slice_mask *dst, struct slice_mask *src)
> {
> DECLARE_BITMAP(result, SLICE_NUM_HIGH);
>
> - dst->low_slices |= src->low_slices;
> + bitmap_or(dst->low_slices, dst->low_slices, src->low_slices,
> + SLICE_NUM_LOW);
> if (SLICE_NUM_HIGH) {
> bitmap_or(result, dst->high_slices, src->high_slices,
> SLICE_NUM_HIGH);
> @@ -409,7 +419,8 @@ static inline void slice_andnot_mask(struct slice_mask *dst, struct slice_mask *
> {
> DECLARE_BITMAP(result, SLICE_NUM_HIGH);
>
> - dst->low_slices &= ~src->low_slices;
> + bitmap_andnot(dst->low_slices, dst->low_slices, src->low_slices,
> + SLICE_NUM_LOW);
>
> if (SLICE_NUM_HIGH) {
> bitmap_andnot(result, dst->high_slices, src->high_slices,
> @@ -464,16 +475,16 @@ unsigned long slice_get_unmapped_area(unsigned long addr, unsigned long len,
> /*
> * init different masks
> */
> - mask.low_slices = 0;
> + bitmap_zero(mask.low_slices, SLICE_NUM_LOW);
> if (SLICE_NUM_HIGH)
> bitmap_zero(mask.high_slices, SLICE_NUM_HIGH);
>
> /* silence stupid warning */;
> - potential_mask.low_slices = 0;
> + bitmap_zero(potential_mask.low_slices, SLICE_NUM_LOW);
> if (SLICE_NUM_HIGH)
> bitmap_zero(potential_mask.high_slices, SLICE_NUM_HIGH);
>
> - compat_mask.low_slices = 0;
> + bitmap_zero(compat_mask.low_slices, SLICE_NUM_LOW);
> if (SLICE_NUM_HIGH)
> bitmap_zero(compat_mask.high_slices, SLICE_NUM_HIGH);
>
> @@ -613,7 +624,7 @@ unsigned long slice_get_unmapped_area(unsigned long addr, unsigned long len,
> convert:
> slice_andnot_mask(&mask, &good_mask);
> slice_andnot_mask(&mask, &compat_mask);
> - if (mask.low_slices ||
> + if (!bitmap_empty(mask.low_slices, SLICE_NUM_LOW) ||
> (SLICE_NUM_HIGH &&
> !bitmap_empty(mask.high_slices, SLICE_NUM_HIGH))) {
> slice_convert(mm, mask, psize);
> @@ -647,7 +658,7 @@ unsigned long arch_get_unmapped_area_topdown(struct file *filp,
>
> unsigned int get_slice_psize(struct mm_struct *mm, unsigned long addr)
> {
> - unsigned char *hpsizes;
> + unsigned char *psizes;
> int index, mask_index;
>
> /*
> @@ -661,15 +672,14 @@ unsigned int get_slice_psize(struct mm_struct *mm, unsigned long addr)
> #endif
> }
> if (addr <= SLICE_LOW_TOP) {
> - u64 lpsizes;
> - lpsizes = mm->context.low_slices_psize;
> + psizes = mm->context.low_slices_psize;
> index = GET_LOW_SLICE_INDEX(addr);
> - return (lpsizes >> (index * 4)) & 0xf;
> + } else {
> + psizes = mm->context.high_slices_psize;
> + index = GET_HIGH_SLICE_INDEX(addr);
> }
> - hpsizes = mm->context.high_slices_psize;
> - index = GET_HIGH_SLICE_INDEX(addr);
> mask_index = index & 0x1;
> - return (hpsizes[index >> 1] >> (mask_index * 4)) & 0xf;
> + return (psizes[index >> 1] >> (mask_index * 4)) & 0xf;
> }
> EXPORT_SYMBOL_GPL(get_slice_psize);
>
> @@ -690,8 +700,8 @@ EXPORT_SYMBOL_GPL(get_slice_psize);
> void slice_set_user_psize(struct mm_struct *mm, unsigned int psize)
> {
> int index, mask_index;
> - unsigned char *hpsizes;
> - unsigned long flags, lpsizes;
> + unsigned char *hpsizes, *lpsizes;
> + unsigned long flags;
> unsigned int old_psize;
> int i;
>
> @@ -709,12 +719,14 @@ void slice_set_user_psize(struct mm_struct *mm, unsigned int psize)
> wmb();
>
> lpsizes = mm->context.low_slices_psize;
> - for (i = 0; i < SLICE_NUM_LOW; i++)
> - if (((lpsizes >> (i * 4)) & 0xf) == old_psize)
> - lpsizes = (lpsizes & ~(0xful << (i * 4))) |
> - (((unsigned long)psize) << (i * 4));
> - /* Assign the value back */
> - mm->context.low_slices_psize = lpsizes;
> + for (i = 0; i < SLICE_NUM_LOW; i++) {
> + mask_index = i & 0x1;
> + index = i >> 1;
> + if (((lpsizes[index] >> (mask_index * 4)) & 0xf) == old_psize)
> + lpsizes[index] = (lpsizes[index] &
> + ~(0xf << (mask_index * 4))) |
> + (((unsigned long)psize) << (mask_index * 4));
> + }
>
> hpsizes = mm->context.high_slices_psize;
> for (i = 0; i < SLICE_NUM_HIGH; i++) {
> --
> 2.13.3
^ permalink raw reply
* Re: [PATCH v2] powerpc/mm: Fix growth direction for hugepages mmaps with slice
From: Aneesh Kumar K.V @ 2018-01-16 16:03 UTC (permalink / raw)
To: Christophe Leroy, Benjamin Herrenschmidt, Paul Mackerras,
Michael Ellerman, Scott Wood
Cc: linuxppc-dev, linux-kernel
In-Reply-To: <20180109101810.2471D6C6CF@localhost.localdomain>
Christophe Leroy <christophe.leroy@c-s.fr> writes:
> An application running with libhugetlbfs fails to allocate
> additional pages to HEAP due to the hugemap being done
> inconditionally as topdown mapping:
>
> mmap(0x10080000, 1572864, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS|0x40000, -1, 0) = 0x73e80000
> [...]
> mmap(0x74000000, 1048576, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS|0x40000, -1, 0x180000) = 0x73d80000
> munmap(0x73d80000, 1048576) = 0
> [...]
> mmap(0x74000000, 1572864, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS|0x40000, -1, 0x180000) = 0x73d00000
> munmap(0x73d00000, 1572864) = 0
> [...]
> mmap(0x74000000, 1572864, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS|0x40000, -1, 0x180000) = 0x73d00000
> munmap(0x73d00000, 1572864) = 0
> [...]
>
Can you explain the failure details above. I am not sure I understand
what to read from the above output.
> As one can see from the above strace log, mmap() allocates further
> pages below the initial one.
>
> This patch fixes it by taking into account MAP_GROWSDOWN flag.
Rest of the kernel don't depend on that flag to select a topdown search
or not. So what is special with hugetlb? IF we select legacy mmap that
is when we select a bottomup search. Hugetlb on ppc64 always did a
topdown search.
>
> Fixes: d0f13e3c20b6f ("[POWERPC] Introduce address space "slices" ")
> Signed-off-by: Christophe Leroy <christophe.leroy@c-s.fr>
> ---
> v2: Added missing include
>
> arch/powerpc/mm/hugetlbpage.c | 4 +++-
> 1 file changed, 3 insertions(+), 1 deletion(-)
>
> diff --git a/arch/powerpc/mm/hugetlbpage.c b/arch/powerpc/mm/hugetlbpage.c
> index 79e1378ee303..0eadf9f199de 100644
> --- a/arch/powerpc/mm/hugetlbpage.c
> +++ b/arch/powerpc/mm/hugetlbpage.c
> @@ -19,6 +19,7 @@
> #include <linux/moduleparam.h>
> #include <linux/swap.h>
> #include <linux/swapops.h>
> +#include <linux/mman.h>
> #include <asm/pgtable.h>
> #include <asm/pgalloc.h>
> #include <asm/tlb.h>
> @@ -558,7 +559,8 @@ unsigned long hugetlb_get_unmapped_area(struct file *file, unsigned long addr,
> return radix__hugetlb_get_unmapped_area(file, addr, len,
> pgoff, flags);
> #endif
> - return slice_get_unmapped_area(addr, len, flags, mmu_psize, 1);
> + return slice_get_unmapped_area(addr, len, flags, mmu_psize,
> + flags & MAP_GROWSDOWN);
> }
> #endif
>
> --
> 2.13.3
^ permalink raw reply
* Re: [PATCH 1/3] powerpc/32: Fix hugepage allocation on 8xx at hint address
From: Aneesh Kumar K.V @ 2018-01-16 15:49 UTC (permalink / raw)
To: Christophe Leroy, Benjamin Herrenschmidt, Paul Mackerras,
Michael Ellerman, Scott Wood, Nicholas Piggin
Cc: linux-kernel, linuxppc-dev
In-Reply-To: <9a5dadc10f88e2fc0ac9fb5d18c5424df33f3f4c.1515169256.git.christophe.leroy@c-s.fr>
Christophe Leroy <christophe.leroy@c-s.fr> writes:
> When an app has some regular pages allocated (e.g. see below) and tries
> to mmap() a huge page at a hint address covered by the same PMD entry,
> the kernel accepts the hint allthough the 8xx cannot handle different
> page sizes in the same PMD entry.
So that is a bug in get_unmapped_area function that you are using and
you want to fix that by using the slice code. Can you describe here what
the allocation restrictions are w.r.t 8xx? Do they have segments and
base page size like hash64?
>
> 10000000-10001000 r-xp 00000000 00:0f 2597 /root/malloc
> 10010000-10011000 rwxp 00000000 00:0f 2597 /root/malloc
>
> mmap(0x10080000, 524288, PROT_READ|PROT_WRITE,
> MAP_PRIVATE|MAP_ANONYMOUS|0x40000, -1, 0) = 0x10080000
>
> This results in the following warning, and the app remains forever in
> do_page_fault()/hugetlb_fault()
>
> [162980.035629] WARNING: CPU: 0 PID: 2777 at arch/powerpc/mm/hugetlbpage.c:354 hugetlb_free_pgd_range+0xc8/0x1e4
> [162980.035699] CPU: 0 PID: 2777 Comm: malloc Tainted: G W 4.14.6 #85
> [162980.035744] task: c67e2c00 task.stack: c668e000
> [162980.035783] NIP: c000fe18 LR: c00e1eec CTR: c00f90c0
> [162980.035830] REGS: c668fc20 TRAP: 0700 Tainted: G W (4.14.6)
> [162980.035854] MSR: 00029032 <EE,ME,IR,DR,RI> CR: 24044224 XER: 20000000
> [162980.036003]
> [162980.036003] GPR00: c00e1eec c668fcd0 c67e2c00 00000010 c6869410 10080000 00000000 77fb4000
> [162980.036003] GPR08: ffff0001 0683c001 00000000 ffffff80 44028228 10018a34 00004008 418004fc
> [162980.036003] GPR16: c668e000 00040100 c668e000 c06c0000 c668fe78 c668e000 c6835ba0 c668fd48
> [162980.036003] GPR24: 00000000 73ffffff 74000000 00000001 77fb4000 100fffff 10100000 10100000
> [162980.036743] NIP [c000fe18] hugetlb_free_pgd_range+0xc8/0x1e4
> [162980.036839] LR [c00e1eec] free_pgtables+0x12c/0x150
> [162980.036861] Call Trace:
> [162980.036939] [c668fcd0] [c00f0774] unlink_anon_vmas+0x1c4/0x214 (unreliable)
> [162980.037040] [c668fd10] [c00e1eec] free_pgtables+0x12c/0x150
> [162980.037118] [c668fd40] [c00eabac] exit_mmap+0xe8/0x1b4
> [162980.037210] [c668fda0] [c0019710] mmput.part.9+0x20/0xd8
> [162980.037301] [c668fdb0] [c001ecb0] do_exit+0x1f0/0x93c
> [162980.037386] [c668fe00] [c001f478] do_group_exit+0x40/0xcc
> [162980.037479] [c668fe10] [c002a76c] get_signal+0x47c/0x614
> [162980.037570] [c668fe70] [c0007840] do_signal+0x54/0x244
> [162980.037654] [c668ff30] [c0007ae8] do_notify_resume+0x34/0x88
> [162980.037744] [c668ff40] [c000dae8] do_user_signal+0x74/0xc4
> [162980.037781] Instruction dump:
> [162980.037821] 7fdff378 81370000 54a3463a 80890020 7d24182e 7c841a14 712a0004 4082ff94
> [162980.038014] 2f890000 419e0010 712a0ff0 408200e0 <0fe00000> 54a9000a 7f984840 419d0094
> [162980.038216] ---[ end trace c0ceeca8e7a5800a ]---
> [162980.038754] BUG: non-zero nr_ptes on freeing mm: 1
> [162985.363322] BUG: non-zero nr_ptes on freeing mm: -1
>
> In order to fix this, the address space "slices" implemented
> for BOOK3S/64 is reused.
>
> This patch:
> 1/ Modifies the "slices" implementation to support 32 bits CPUs,
> based on using only the low slices.
> 2/ Moves "slices" functions prototypes from page64.h to page.h
> 3/ Modifies the context.id on the 8xx to be in the range [1:16]
> instead of [0:15] in order to identify context.id == 0 as
> not initialised contexts
> 4/ Activates CONFIG_PPC_MM_SLICES when CONFIG_HUGETLB_PAGE is
> selected for the 8xx
>
> Alltough we could in theory have as many slices as PMD entries, the current
> slices implementation limits the number of low slices to 16.
Can you explain this more?
>
> Fixes: 4b91428699477 ("powerpc/8xx: Implement support of hugepages")
> Signed-off-by: Christophe Leroy <christophe.leroy@c-s.fr>
> ---
> arch/powerpc/include/asm/mmu-8xx.h | 6 ++++
> arch/powerpc/include/asm/page.h | 14 ++++++++
> arch/powerpc/include/asm/page_32.h | 19 +++++++++++
> arch/powerpc/include/asm/page_64.h | 21 ++----------
> arch/powerpc/kernel/setup-common.c | 2 +-
> arch/powerpc/mm/8xx_mmu.c | 2 +-
> arch/powerpc/mm/hash_utils_64.c | 2 +-
> arch/powerpc/mm/hugetlbpage.c | 2 ++
> arch/powerpc/mm/mmu_context_nohash.c | 11 +++++--
> arch/powerpc/mm/slice.c | 58 +++++++++++++++++++++++-----------
> arch/powerpc/platforms/Kconfig.cputype | 1 +
> 11 files changed, 95 insertions(+), 43 deletions(-)
>
> diff --git a/arch/powerpc/include/asm/mmu-8xx.h b/arch/powerpc/include/asm/mmu-8xx.h
> index 5bb3dbede41a..5f89b6010453 100644
> --- a/arch/powerpc/include/asm/mmu-8xx.h
> +++ b/arch/powerpc/include/asm/mmu-8xx.h
> @@ -169,6 +169,12 @@ typedef struct {
> unsigned int id;
> unsigned int active;
> unsigned long vdso_base;
> +#ifdef CONFIG_PPC_MM_SLICES
> + u16 user_psize; /* page size index */
> + u64 low_slices_psize; /* page size encodings */
> + unsigned char high_slices_psize[0];
> + unsigned long slb_addr_limit;
> +#endif
> } mm_context_t;
>
> #define PHYS_IMMR_BASE (mfspr(SPRN_IMMR) & 0xfff80000)
> diff --git a/arch/powerpc/include/asm/page.h b/arch/powerpc/include/asm/page.h
> index 8da5d4c1cab2..d0384f9db9eb 100644
> --- a/arch/powerpc/include/asm/page.h
> +++ b/arch/powerpc/include/asm/page.h
> @@ -342,6 +342,20 @@ typedef struct page *pgtable_t;
> #endif
> #endif
>
> +#ifdef CONFIG_PPC_MM_SLICES
> +struct mm_struct;
> +
> +unsigned long slice_get_unmapped_area(unsigned long addr, unsigned long len,
> + unsigned long flags, unsigned int psize,
> + int topdown);
> +
> +unsigned int get_slice_psize(struct mm_struct *mm, unsigned long addr);
> +
> +void slice_set_user_psize(struct mm_struct *mm, unsigned int psize);
> +void slice_set_range_psize(struct mm_struct *mm, unsigned long start,
> + unsigned long len, unsigned int psize);
> +#endif
> +
> #include <asm-generic/memory_model.h>
> #endif /* __ASSEMBLY__ */
>
> diff --git a/arch/powerpc/include/asm/page_32.h b/arch/powerpc/include/asm/page_32.h
> index 5c378e9b78c8..f7d1bd1183c8 100644
> --- a/arch/powerpc/include/asm/page_32.h
> +++ b/arch/powerpc/include/asm/page_32.h
> @@ -60,4 +60,23 @@ extern void copy_page(void *to, void *from);
>
> #endif /* __ASSEMBLY__ */
>
> +#ifdef CONFIG_PPC_MM_SLICES
> +
> +#define SLICE_LOW_SHIFT 28
> +#define SLICE_HIGH_SHIFT 0
> +
> +#define SLICE_LOW_TOP (0xfffffffful)
> +#define SLICE_NUM_LOW ((SLICE_LOW_TOP >> SLICE_LOW_SHIFT) + 1)
> +#define SLICE_NUM_HIGH 0ul
> +
> +#define GET_LOW_SLICE_INDEX(addr) ((addr) >> SLICE_LOW_SHIFT)
> +#define GET_HIGH_SLICE_INDEX(addr) (addr & 0)
> +
> +#ifdef CONFIG_HUGETLB_PAGE
> +#define HAVE_ARCH_HUGETLB_UNMAPPED_AREA
> +#endif
> +#define HAVE_ARCH_UNMAPPED_AREA
> +#define HAVE_ARCH_UNMAPPED_AREA_TOPDOWN
> +
> +#endif
> #endif /* _ASM_POWERPC_PAGE_32_H */
> diff --git a/arch/powerpc/include/asm/page_64.h b/arch/powerpc/include/asm/page_64.h
> index 56234c6fcd61..a7baef5bbe5f 100644
> --- a/arch/powerpc/include/asm/page_64.h
> +++ b/arch/powerpc/include/asm/page_64.h
> @@ -91,30 +91,13 @@ extern u64 ppc64_pft_size;
> #define SLICE_LOW_SHIFT 28
> #define SLICE_HIGH_SHIFT 40
>
> -#define SLICE_LOW_TOP (0x100000000ul)
> -#define SLICE_NUM_LOW (SLICE_LOW_TOP >> SLICE_LOW_SHIFT)
> +#define SLICE_LOW_TOP (0xfffffffful)
> +#define SLICE_NUM_LOW ((SLICE_LOW_TOP >> SLICE_LOW_SHIFT) + 1)
> #define SLICE_NUM_HIGH (H_PGTABLE_RANGE >> SLICE_HIGH_SHIFT)
Why are you changing this? is this a bug fix?
>
> #define GET_LOW_SLICE_INDEX(addr) ((addr) >> SLICE_LOW_SHIFT)
> #define GET_HIGH_SLICE_INDEX(addr) ((addr) >> SLICE_HIGH_SHIFT)
>
> -#ifndef __ASSEMBLY__
> -struct mm_struct;
> -
> -extern unsigned long slice_get_unmapped_area(unsigned long addr,
> - unsigned long len,
> - unsigned long flags,
> - unsigned int psize,
> - int topdown);
> -
> -extern unsigned int get_slice_psize(struct mm_struct *mm,
> - unsigned long addr);
> -
> -extern void slice_set_user_psize(struct mm_struct *mm, unsigned int psize);
> -extern void slice_set_range_psize(struct mm_struct *mm, unsigned long start,
> - unsigned long len, unsigned int psize);
> -
> -#endif /* __ASSEMBLY__ */
> #else
> #define slice_init()
> #ifdef CONFIG_PPC_BOOK3S_64
> diff --git a/arch/powerpc/kernel/setup-common.c b/arch/powerpc/kernel/setup-common.c
> index 9d213542a48b..a285e1067713 100644
> --- a/arch/powerpc/kernel/setup-common.c
> +++ b/arch/powerpc/kernel/setup-common.c
> @@ -928,7 +928,7 @@ void __init setup_arch(char **cmdline_p)
> if (!radix_enabled())
> init_mm.context.slb_addr_limit = DEFAULT_MAP_WINDOW_USER64;
> #else
> -#error "context.addr_limit not initialized."
> + init_mm.context.slb_addr_limit = DEFAULT_MAP_WINDOW;
> #endif
May be put this within #ifdef 8XX and retain the error?
> #endif
>
> diff --git a/arch/powerpc/mm/8xx_mmu.c b/arch/powerpc/mm/8xx_mmu.c
> index f29212e40f40..0be77709446c 100644
> --- a/arch/powerpc/mm/8xx_mmu.c
> +++ b/arch/powerpc/mm/8xx_mmu.c
> @@ -192,7 +192,7 @@ void set_context(unsigned long id, pgd_t *pgd)
> mtspr(SPRN_M_TW, __pa(pgd) - offset);
>
> /* Update context */
> - mtspr(SPRN_M_CASID, id);
> + mtspr(SPRN_M_CASID, id - 1);
> /* sync */
> mb();
> }
> diff --git a/arch/powerpc/mm/hash_utils_64.c b/arch/powerpc/mm/hash_utils_64.c
> index 655a5a9a183d..3266b3326088 100644
> --- a/arch/powerpc/mm/hash_utils_64.c
> +++ b/arch/powerpc/mm/hash_utils_64.c
> @@ -1101,7 +1101,7 @@ static unsigned int get_paca_psize(unsigned long addr)
> unsigned char *hpsizes;
> unsigned long index, mask_index;
>
> - if (addr < SLICE_LOW_TOP) {
> + if (addr <= SLICE_LOW_TOP) {
If this is part of bug fix, please do it as part of seperate patch with details
> lpsizes = get_paca()->mm_ctx_low_slices_psize;
> index = GET_LOW_SLICE_INDEX(addr);
> return (lpsizes >> (index * 4)) & 0xF;
> diff --git a/arch/powerpc/mm/hugetlbpage.c b/arch/powerpc/mm/hugetlbpage.c
> index a9b9083c5e49..79e1378ee303 100644
> --- a/arch/powerpc/mm/hugetlbpage.c
> +++ b/arch/powerpc/mm/hugetlbpage.c
> @@ -553,9 +553,11 @@ unsigned long hugetlb_get_unmapped_area(struct file *file, unsigned long addr,
> struct hstate *hstate = hstate_file(file);
> int mmu_psize = shift_to_mmu_psize(huge_page_shift(hstate));
>
> +#ifdef CONFIG_PPC_RADIX_MMU
> if (radix_enabled())
> return radix__hugetlb_get_unmapped_area(file, addr, len,
> pgoff, flags);
> +#endif
> return slice_get_unmapped_area(addr, len, flags, mmu_psize, 1);
> }
> #endif
> diff --git a/arch/powerpc/mm/mmu_context_nohash.c b/arch/powerpc/mm/mmu_context_nohash.c
> index 4554d6527682..c1e1bf186871 100644
> --- a/arch/powerpc/mm/mmu_context_nohash.c
> +++ b/arch/powerpc/mm/mmu_context_nohash.c
> @@ -331,6 +331,13 @@ int init_new_context(struct task_struct *t, struct mm_struct *mm)
> {
> pr_hard("initing context for mm @%p\n", mm);
>
> +#ifdef CONFIG_PPC_MM_SLICES
> + if (!mm->context.slb_addr_limit)
> + mm->context.slb_addr_limit = DEFAULT_MAP_WINDOW;
> + if (!mm->context.id)
> + slice_set_user_psize(mm, mmu_virtual_psize);
> +#endif
> +
> mm->context.id = MMU_NO_CONTEXT;
> mm->context.active = 0;
> return 0;
> @@ -428,8 +435,8 @@ void __init mmu_context_init(void)
> * -- BenH
> */
> if (mmu_has_feature(MMU_FTR_TYPE_8xx)) {
> - first_context = 0;
> - last_context = 15;
> + first_context = 1;
> + last_context = 16;
> no_selective_tlbil = true;
> } else if (mmu_has_feature(MMU_FTR_TYPE_47x)) {
> first_context = 1;
> diff --git a/arch/powerpc/mm/slice.c b/arch/powerpc/mm/slice.c
> index 23ec2c5e3b78..1a66fafc3e45 100644
> --- a/arch/powerpc/mm/slice.c
> +++ b/arch/powerpc/mm/slice.c
> @@ -73,10 +73,11 @@ static void slice_range_to_mask(unsigned long start, unsigned long len,
> unsigned long end = start + len - 1;
>
> ret->low_slices = 0;
> - bitmap_zero(ret->high_slices, SLICE_NUM_HIGH);
> + if (SLICE_NUM_HIGH)
> + bitmap_zero(ret->high_slices, SLICE_NUM_HIGH);
So you don't want to use high slices but just low slice? If so can you
add that as a different patch which implements just that.
>
> - if (start < SLICE_LOW_TOP) {
> - unsigned long mend = min(end, (SLICE_LOW_TOP - 1));
> + if (start <= SLICE_LOW_TOP) {
> + unsigned long mend = min(end, SLICE_LOW_TOP);
>
> ret->low_slices = (1u << (GET_LOW_SLICE_INDEX(mend) + 1))
> - (1u << GET_LOW_SLICE_INDEX(start));
> @@ -117,7 +118,7 @@ static int slice_high_has_vma(struct mm_struct *mm, unsigned long slice)
> * of the high or low area bitmaps, the first high area starts
> * at 4GB, not 0 */
> if (start == 0)
> - start = SLICE_LOW_TOP;
> + start = SLICE_LOW_TOP + 1;
>
> return !slice_area_is_free(mm, start, end - start);
> }
> @@ -128,7 +129,8 @@ static void slice_mask_for_free(struct mm_struct *mm, struct slice_mask *ret,
> unsigned long i;
>
> ret->low_slices = 0;
> - bitmap_zero(ret->high_slices, SLICE_NUM_HIGH);
> + if (SLICE_NUM_HIGH)
> + bitmap_zero(ret->high_slices, SLICE_NUM_HIGH);
>
> for (i = 0; i < SLICE_NUM_LOW; i++)
> if (!slice_low_has_vma(mm, i))
> @@ -151,7 +153,8 @@ static void slice_mask_for_size(struct mm_struct *mm, int psize, struct slice_ma
> u64 lpsizes;
>
> ret->low_slices = 0;
> - bitmap_zero(ret->high_slices, SLICE_NUM_HIGH);
> + if (SLICE_NUM_HIGH)
> + bitmap_zero(ret->high_slices, SLICE_NUM_HIGH);
>
> lpsizes = mm->context.low_slices_psize;
> for (i = 0; i < SLICE_NUM_LOW; i++)
> @@ -180,15 +183,18 @@ static int slice_check_fit(struct mm_struct *mm,
> */
> unsigned long slice_count = GET_HIGH_SLICE_INDEX(mm->context.slb_addr_limit);
>
> - bitmap_and(result, mask.high_slices,
> - available.high_slices, slice_count);
> + if (SLICE_NUM_HIGH)
> + bitmap_and(result, mask.high_slices,
> + available.high_slices, slice_count);
>
> return (mask.low_slices & available.low_slices) == mask.low_slices &&
> - bitmap_equal(result, mask.high_slices, slice_count);
> + (!slice_count ||
> + bitmap_equal(result, mask.high_slices, slice_count));
> }
>
> static void slice_flush_segments(void *parm)
> {
> +#ifdef CONFIG_PPC_BOOK3S_64
> struct mm_struct *mm = parm;
> unsigned long flags;
>
> @@ -200,6 +206,7 @@ static void slice_flush_segments(void *parm)
> local_irq_save(flags);
> slb_flush_and_rebolt();
> local_irq_restore(flags);
> +#endif
> }
>
> static void slice_convert(struct mm_struct *mm, struct slice_mask mask, int psize)
> @@ -259,7 +266,7 @@ static bool slice_scan_available(unsigned long addr,
> unsigned long *boundary_addr)
> {
> unsigned long slice;
> - if (addr < SLICE_LOW_TOP) {
> + if (addr <= SLICE_LOW_TOP) {
> slice = GET_LOW_SLICE_INDEX(addr);
> *boundary_addr = (slice + end) << SLICE_LOW_SHIFT;
> return !!(available.low_slices & (1u << slice));
> @@ -391,8 +398,11 @@ static inline void slice_or_mask(struct slice_mask *dst, struct slice_mask *src)
> DECLARE_BITMAP(result, SLICE_NUM_HIGH);
>
> dst->low_slices |= src->low_slices;
> - bitmap_or(result, dst->high_slices, src->high_slices, SLICE_NUM_HIGH);
> - bitmap_copy(dst->high_slices, result, SLICE_NUM_HIGH);
> + if (SLICE_NUM_HIGH) {
> + bitmap_or(result, dst->high_slices, src->high_slices,
> + SLICE_NUM_HIGH);
> + bitmap_copy(dst->high_slices, result, SLICE_NUM_HIGH);
> + }
> }
>
> static inline void slice_andnot_mask(struct slice_mask *dst, struct slice_mask *src)
> @@ -401,12 +411,17 @@ static inline void slice_andnot_mask(struct slice_mask *dst, struct slice_mask *
>
> dst->low_slices &= ~src->low_slices;
>
> - bitmap_andnot(result, dst->high_slices, src->high_slices, SLICE_NUM_HIGH);
> - bitmap_copy(dst->high_slices, result, SLICE_NUM_HIGH);
> + if (SLICE_NUM_HIGH) {
> + bitmap_andnot(result, dst->high_slices, src->high_slices,
> + SLICE_NUM_HIGH);
> + bitmap_copy(dst->high_slices, result, SLICE_NUM_HIGH);
> + }
> }
>
> #ifdef CONFIG_PPC_64K_PAGES
> #define MMU_PAGE_BASE MMU_PAGE_64K
> +#elif defined(CONFIG_PPC_16K_PAGES)
> +#define MMU_PAGE_BASE MMU_PAGE_16K
> #else
> #define MMU_PAGE_BASE MMU_PAGE_4K
> #endif
I am not sure we want them based on page size. The rule is we flush
segments on book3s63 if the page size is different from MMU_PAGE_BASE
> @@ -450,14 +465,17 @@ unsigned long slice_get_unmapped_area(unsigned long addr, unsigned long len,
> * init different masks
> */
> mask.low_slices = 0;
> - bitmap_zero(mask.high_slices, SLICE_NUM_HIGH);
> + if (SLICE_NUM_HIGH)
> + bitmap_zero(mask.high_slices, SLICE_NUM_HIGH);
>
> /* silence stupid warning */;
> potential_mask.low_slices = 0;
> - bitmap_zero(potential_mask.high_slices, SLICE_NUM_HIGH);
> + if (SLICE_NUM_HIGH)
> + bitmap_zero(potential_mask.high_slices, SLICE_NUM_HIGH);
>
> compat_mask.low_slices = 0;
> - bitmap_zero(compat_mask.high_slices, SLICE_NUM_HIGH);
> + if (SLICE_NUM_HIGH)
> + bitmap_zero(compat_mask.high_slices, SLICE_NUM_HIGH);
>
> /* Sanity checks */
> BUG_ON(mm->task_size == 0);
> @@ -595,7 +613,9 @@ unsigned long slice_get_unmapped_area(unsigned long addr, unsigned long len,
> convert:
> slice_andnot_mask(&mask, &good_mask);
> slice_andnot_mask(&mask, &compat_mask);
> - if (mask.low_slices || !bitmap_empty(mask.high_slices, SLICE_NUM_HIGH)) {
> + if (mask.low_slices ||
> + (SLICE_NUM_HIGH &&
> + !bitmap_empty(mask.high_slices, SLICE_NUM_HIGH))) {
> slice_convert(mm, mask, psize);
> if (psize > MMU_PAGE_BASE)
> on_each_cpu(slice_flush_segments, mm, 1);
> @@ -640,7 +660,7 @@ unsigned int get_slice_psize(struct mm_struct *mm, unsigned long addr)
> return MMU_PAGE_4K;
> #endif
> }
> - if (addr < SLICE_LOW_TOP) {
> + if (addr <= SLICE_LOW_TOP) {
> u64 lpsizes;
> lpsizes = mm->context.low_slices_psize;
> index = GET_LOW_SLICE_INDEX(addr);
> diff --git a/arch/powerpc/platforms/Kconfig.cputype b/arch/powerpc/platforms/Kconfig.cputype
> index ae07470fde3c..73a7ea333e9e 100644
> --- a/arch/powerpc/platforms/Kconfig.cputype
> +++ b/arch/powerpc/platforms/Kconfig.cputype
> @@ -334,6 +334,7 @@ config PPC_BOOK3E_MMU
> config PPC_MM_SLICES
> bool
> default y if PPC_BOOK3S_64
> + default y if PPC_8xx && HUGETLB_PAGE
> default n
>
> config PPC_HAVE_PMU_SUPPORT
> --
> 2.13.3
^ permalink raw reply
* Re: [PATCH v3 02/10] include: Move compat_timespec/ timeval to compat_time.h
From: Steven Rostedt @ 2018-01-16 15:34 UTC (permalink / raw)
To: Deepa Dinamani
Cc: tglx, john.stultz, linux-kernel, arnd, y2038, acme, benh,
borntraeger, catalin.marinas, cmetcalf, cohuck, davem, deller,
devel, gerald.schaefer, gregkh, heiko.carstens, hoeppner, hpa,
jejb, jwi, linux-mips, linux-parisc, linuxppc-dev, linux-s390,
mark.rutland, mingo, mpe, oberpar, oprofile-list, paulus, peterz,
ralf, rric, schwidefsky, sebott, sparclinux, sth, ubraun,
will.deacon, x86
In-Reply-To: <20180116021818.24791-3-deepa.kernel@gmail.com>
On Mon, 15 Jan 2018 18:18:10 -0800
Deepa Dinamani <deepa.kernel@gmail.com> wrote:
> diff --git a/arch/x86/include/asm/ftrace.h b/arch/x86/include/asm/ftrace.h
> index 09ad88572746..db25aa15b705 100644
> --- a/arch/x86/include/asm/ftrace.h
> +++ b/arch/x86/include/asm/ftrace.h
Acked-by: Steven Rostedt (VMware) <rostedt@goodmis.org>
-- Steve
> @@ -49,7 +49,7 @@ int ftrace_int3_handler(struct pt_regs *regs);
> #if !defined(__ASSEMBLY__) && !defined(COMPILE_OFFSETS)
>
> #if defined(CONFIG_FTRACE_SYSCALLS) && defined(CONFIG_IA32_EMULATION)
> -#include <asm/compat.h>
> +#include <linux/compat.h>
>
> /*
> * Because ia32 syscalls do not map to x86_64 syscall numbers
^ permalink raw reply
page: next (older) | prev (newer) | latest
- recent:[subjects (threaded)|topics (new)|topics (active)]
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox