* Re: [PATCH] mm: fix hugetlb page unmap count balance issue
[not found] ` <20230512232947.GA3927@monkey>
@ 2023-05-15 17:04 ` Mike Kravetz
2023-05-16 22:34 ` Mike Kravetz
2023-06-19 12:27 ` Gerd Hoffmann
0 siblings, 2 replies; 10+ messages in thread
From: Mike Kravetz @ 2023-05-15 17:04 UTC (permalink / raw)
To: James Houghton
Cc: mhocko, jmarchan, Dongwon Kim, Junxiao Chang, muchun.song,
linux-kernel, dri-devel, Vivek Kasireddy, linux-mm, Gerd Hoffmann,
akpm, kirill.shutemov
On 05/12/23 16:29, Mike Kravetz wrote:
> On 05/12/23 14:26, James Houghton wrote:
> > On Fri, May 12, 2023 at 12:20 AM Junxiao Chang <junxiao.chang@intel.com> wrote:
> >
> > This alone doesn't fix mapcounting for PTE-mapped HugeTLB pages. You
> > need something like [1]. I can resend it if that's what we should be
> > doing, but this mapcounting scheme doesn't work when the page structs
> > have been freed.
> >
> > It seems like it was a mistake to include support for hugetlb memfds in udmabuf.
>
> IIUC, it was added with commit 16c243e99d33 udmabuf: Add support for mapping
> hugepages (v4). Looks like it was never sent to linux-mm? That is unfortunate
> as hugetlb vmemmap freeing went in at about the same time. And, as you have
> noted udmabuf will not work if hugetlb vmemmap freeing is enabled.
>
> Sigh!
>
> Trying to think of a way forward.
> --
> Mike Kravetz
>
> >
> > [1]: https://lore.kernel.org/linux-mm/20230306230004.1387007-2-jthoughton@google.com/
> >
> > - James
Adding people and list on Cc: involved with commit 16c243e99d33.
There are several issues with trying to map tail pages of hugetllb pages
not taken into account with udmabuf. James spent quite a bit of time trying
to understand and address all the issues with the HGM code. While using
the scheme proposed by James, may be an approach to the mapcount issue there
are also other issues that need attention. For example, I do not see how
the fault code checks the state of the hugetlb page (such as poison) as none
of that state is carried in tail pages.
The more I think about it, the more I think udmabuf should treat hugetlb
pages as hugetlb pages. They should be mapped at the appropriate level
in the page table. Of course, this would impose new restrictions on the
API (mmap and ioctl) that may break existing users. I have no idea how
extensively udmabuf is being used with hugetlb mappings.
--
Mike Kravetz
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH] mm: fix hugetlb page unmap count balance issue
2023-05-15 17:04 ` [PATCH] mm: fix hugetlb page unmap count balance issue Mike Kravetz
@ 2023-05-16 22:34 ` Mike Kravetz
2023-06-07 19:03 ` Andrew Morton
2023-06-07 19:27 ` David Hildenbrand
2023-06-19 12:27 ` Gerd Hoffmann
1 sibling, 2 replies; 10+ messages in thread
From: Mike Kravetz @ 2023-05-16 22:34 UTC (permalink / raw)
To: Gerd Hoffmann
Cc: mhocko, jmarchan, Dongwon Kim, Junxiao Chang, muchun.song,
linux-kernel, dri-devel, Vivek Kasireddy, linux-mm,
James Houghton, akpm, kirill.shutemov
On 05/15/23 10:04, Mike Kravetz wrote:
> On 05/12/23 16:29, Mike Kravetz wrote:
> > On 05/12/23 14:26, James Houghton wrote:
> > > On Fri, May 12, 2023 at 12:20 AM Junxiao Chang <junxiao.chang@intel.com> wrote:
> > >
> > > This alone doesn't fix mapcounting for PTE-mapped HugeTLB pages. You
> > > need something like [1]. I can resend it if that's what we should be
> > > doing, but this mapcounting scheme doesn't work when the page structs
> > > have been freed.
> > >
> > > It seems like it was a mistake to include support for hugetlb memfds in udmabuf.
> >
> > IIUC, it was added with commit 16c243e99d33 udmabuf: Add support for mapping
> > hugepages (v4). Looks like it was never sent to linux-mm? That is unfortunate
> > as hugetlb vmemmap freeing went in at about the same time. And, as you have
> > noted udmabuf will not work if hugetlb vmemmap freeing is enabled.
> >
> > Sigh!
> >
> > Trying to think of a way forward.
> > --
> > Mike Kravetz
> >
> > >
> > > [1]: https://lore.kernel.org/linux-mm/20230306230004.1387007-2-jthoughton@google.com/
> > >
> > > - James
>
> Adding people and list on Cc: involved with commit 16c243e99d33.
>
> There are several issues with trying to map tail pages of hugetllb pages
> not taken into account with udmabuf. James spent quite a bit of time trying
> to understand and address all the issues with the HGM code. While using
> the scheme proposed by James, may be an approach to the mapcount issue there
> are also other issues that need attention. For example, I do not see how
> the fault code checks the state of the hugetlb page (such as poison) as none
> of that state is carried in tail pages.
>
> The more I think about it, the more I think udmabuf should treat hugetlb
> pages as hugetlb pages. They should be mapped at the appropriate level
> in the page table. Of course, this would impose new restrictions on the
> API (mmap and ioctl) that may break existing users. I have no idea how
> extensively udmabuf is being used with hugetlb mappings.
Verified that using udmabug on a hugetlb mapping with vmemmap optimization will
BUG as:
[14106.812312] BUG: unable to handle page fault for address: ffffea000a7c4030
[14106.813704] #PF: supervisor write access in kernel mode
[14106.814791] #PF: error_code(0x0003) - permissions violation
[14106.815921] PGD 27fff9067 P4D 27fff9067 PUD 27fff8067 PMD 17ec34067 PTE 8000000285dab021
[14106.818489] Oops: 0003 [#1] PREEMPT SMP PTI
[14106.819345] CPU: 2 PID: 2313 Comm: udmabuf Not tainted 6.4.0-rc1-next-20230508+ #44
[14106.820906] Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.2-1.fc37 04/01/2014
[14106.822679] RIP: 0010:page_add_file_rmap+0x2e/0x270
I started looking more closely at the driver and I do not fully understand the
usage model. I took clues from the selftest and driver. It seems the first
step is to create a buffer via the UDMABUF_CREATE ioctl. This will copy 4K
pages from the page cache to an array associated with a file. I did note that
hugetlb and shm behavior is different here as the driver can not add missing
hugetlb pages to the cache as it does with shm. However, what seems more
concerning is that there is nothing to prevent the pages from being replaced
in the cache before being added to a udmabuf mapping. This means udmabuf
mapping and original memfd could be operating on a different set of pages.
Is this acceptable, or somehow prevented?
In my role, I am more interested in udmabuf handling of hugetlb pages.
Trying to use individual 4K pages of hugetlb pages is something that
should be avoided here. Would it be acceptable to change code so that
only whole hugetlb pages are used by udmabuf? If not, then perhaps the
existing hugetlb support can be removed?
--
Mike Kravetz
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH] mm: fix hugetlb page unmap count balance issue
2023-05-16 22:34 ` Mike Kravetz
@ 2023-06-07 19:03 ` Andrew Morton
2023-06-07 20:53 ` Mike Kravetz
2023-06-07 19:27 ` David Hildenbrand
1 sibling, 1 reply; 10+ messages in thread
From: Andrew Morton @ 2023-06-07 19:03 UTC (permalink / raw)
To: Mike Kravetz
Cc: mhocko, jmarchan, Dongwon Kim, Junxiao Chang, muchun.song,
linux-kernel, dri-devel, Vivek Kasireddy, linux-mm,
James Houghton, Gerd Hoffmann, kirill.shutemov
On Tue, 16 May 2023 15:34:40 -0700 Mike Kravetz <mike.kravetz@oracle.com> wrote:
> On 05/15/23 10:04, Mike Kravetz wrote:
> > On 05/12/23 16:29, Mike Kravetz wrote:
> > > On 05/12/23 14:26, James Houghton wrote:
> > > > On Fri, May 12, 2023 at 12:20 AM Junxiao Chang <junxiao.chang@intel.com> wrote:
> > > >
> > > > This alone doesn't fix mapcounting for PTE-mapped HugeTLB pages. You
> > > > need something like [1]. I can resend it if that's what we should be
> > > > doing, but this mapcounting scheme doesn't work when the page structs
> > > > have been freed.
> > > >
> > > > It seems like it was a mistake to include support for hugetlb memfds in udmabuf.
> > >
> > > IIUC, it was added with commit 16c243e99d33 udmabuf: Add support for mapping
> > > hugepages (v4). Looks like it was never sent to linux-mm? That is unfortunate
> > > as hugetlb vmemmap freeing went in at about the same time. And, as you have
> > > noted udmabuf will not work if hugetlb vmemmap freeing is enabled.
> > >
> > > Sigh!
> > >
> > > Trying to think of a way forward.
> > > --
> > > Mike Kravetz
> > >
> > > >
> > > > [1]: https://lore.kernel.org/linux-mm/20230306230004.1387007-2-jthoughton@google.com/
> > > >
> > > > - James
> >
> > Adding people and list on Cc: involved with commit 16c243e99d33.
> >
> > There are several issues with trying to map tail pages of hugetllb pages
> > not taken into account with udmabuf. James spent quite a bit of time trying
> > to understand and address all the issues with the HGM code. While using
> > the scheme proposed by James, may be an approach to the mapcount issue there
> > are also other issues that need attention. For example, I do not see how
> > the fault code checks the state of the hugetlb page (such as poison) as none
> > of that state is carried in tail pages.
> >
> > The more I think about it, the more I think udmabuf should treat hugetlb
> > pages as hugetlb pages. They should be mapped at the appropriate level
> > in the page table. Of course, this would impose new restrictions on the
> > API (mmap and ioctl) that may break existing users. I have no idea how
> > extensively udmabuf is being used with hugetlb mappings.
>
> Verified that using udmabug on a hugetlb mapping with vmemmap optimization will
> BUG as:
BUGs aren't good. Can we please find a way to push this along?
Have we heard anything from any udmabuf people?
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH] mm: fix hugetlb page unmap count balance issue
2023-05-16 22:34 ` Mike Kravetz
2023-06-07 19:03 ` Andrew Morton
@ 2023-06-07 19:27 ` David Hildenbrand
1 sibling, 0 replies; 10+ messages in thread
From: David Hildenbrand @ 2023-06-07 19:27 UTC (permalink / raw)
To: Mike Kravetz, Gerd Hoffmann
Cc: mhocko, jmarchan, Dongwon Kim, Junxiao Chang, muchun.song,
linux-kernel, dri-devel, Vivek Kasireddy, linux-mm,
James Houghton, akpm, kirill.shutemov
On 17.05.23 00:34, Mike Kravetz wrote:
> On 05/15/23 10:04, Mike Kravetz wrote:
>> On 05/12/23 16:29, Mike Kravetz wrote:
>>> On 05/12/23 14:26, James Houghton wrote:
>>>> On Fri, May 12, 2023 at 12:20 AM Junxiao Chang <junxiao.chang@intel.com> wrote:
>>>>
>>>> This alone doesn't fix mapcounting for PTE-mapped HugeTLB pages. You
>>>> need something like [1]. I can resend it if that's what we should be
>>>> doing, but this mapcounting scheme doesn't work when the page structs
>>>> have been freed.
>>>>
>>>> It seems like it was a mistake to include support for hugetlb memfds in udmabuf.
>>>
>>> IIUC, it was added with commit 16c243e99d33 udmabuf: Add support for mapping
>>> hugepages (v4). Looks like it was never sent to linux-mm? That is unfortunate
>>> as hugetlb vmemmap freeing went in at about the same time. And, as you have
>>> noted udmabuf will not work if hugetlb vmemmap freeing is enabled.
>>>
>>> Sigh!
>>>
>>> Trying to think of a way forward.
>>> --
>>> Mike Kravetz
>>>
>>>>
>>>> [1]: https://lore.kernel.org/linux-mm/20230306230004.1387007-2-jthoughton@google.com/
>>>>
>>>> - James
>>
>> Adding people and list on Cc: involved with commit 16c243e99d33.
>>
>> There are several issues with trying to map tail pages of hugetllb pages
>> not taken into account with udmabuf. James spent quite a bit of time trying
>> to understand and address all the issues with the HGM code. While using
>> the scheme proposed by James, may be an approach to the mapcount issue there
>> are also other issues that need attention. For example, I do not see how
>> the fault code checks the state of the hugetlb page (such as poison) as none
>> of that state is carried in tail pages.
>>
>> The more I think about it, the more I think udmabuf should treat hugetlb
>> pages as hugetlb pages. They should be mapped at the appropriate level
>> in the page table. Of course, this would impose new restrictions on the
>> API (mmap and ioctl) that may break existing users. I have no idea how
>> extensively udmabuf is being used with hugetlb mappings.
>
> Verified that using udmabug on a hugetlb mapping with vmemmap optimization will
> BUG as:
>
> [14106.812312] BUG: unable to handle page fault for address: ffffea000a7c4030
> [14106.813704] #PF: supervisor write access in kernel mode
> [14106.814791] #PF: error_code(0x0003) - permissions violation
> [14106.815921] PGD 27fff9067 P4D 27fff9067 PUD 27fff8067 PMD 17ec34067 PTE 8000000285dab021
> [14106.818489] Oops: 0003 [#1] PREEMPT SMP PTI
> [14106.819345] CPU: 2 PID: 2313 Comm: udmabuf Not tainted 6.4.0-rc1-next-20230508+ #44
> [14106.820906] Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.2-1.fc37 04/01/2014
> [14106.822679] RIP: 0010:page_add_file_rmap+0x2e/0x270
>
> I started looking more closely at the driver and I do not fully understand the
> usage model. I took clues from the selftest and driver. It seems the first
> step is to create a buffer via the UDMABUF_CREATE ioctl. This will copy 4K
> pages from the page cache to an array associated with a file. I did note that
> hugetlb and shm behavior is different here as the driver can not add missing
> hugetlb pages to the cache as it does with shm. However, what seems more
> concerning is that there is nothing to prevent the pages from being replaced
> in the cache before being added to a udmabuf mapping. This means udmabuf
> mapping and original memfd could be operating on a different set of pages.
> Is this acceptable, or somehow prevented?
>
> In my role, I am more interested in udmabuf handling of hugetlb pages.
> Trying to use individual 4K pages of hugetlb pages is something that
> should be avoided here. Would it be acceptable to change code so that
> only whole hugetlb pages are used by udmabuf? If not, then perhaps the
> existing hugetlb support can be removed?
I'm wondering if that VMA shouldn't be some kind of special mapping
(VM_PFNMAP), such that the struct page is entirely ignored?
I'm quite confused and concerned when I read that code (what the hell is
it doing with shmem/hugetlb pages? why does the mapcount even get adjusted?)
This all has a bad smell to it, I hope I'm missing something important ...
--
Cheers,
David / dhildenb
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH] mm: fix hugetlb page unmap count balance issue
2023-06-07 19:03 ` Andrew Morton
@ 2023-06-07 20:53 ` Mike Kravetz
2023-06-07 21:00 ` Andrew Morton
0 siblings, 1 reply; 10+ messages in thread
From: Mike Kravetz @ 2023-06-07 20:53 UTC (permalink / raw)
To: Andrew Morton
Cc: mhocko, jmarchan, Dongwon Kim, Junxiao Chang, muchun.song,
linux-kernel, dri-devel, Vivek Kasireddy, linux-mm,
James Houghton, Gerd Hoffmann, kirill.shutemov
On 06/07/23 12:03, Andrew Morton wrote:
> On Tue, 16 May 2023 15:34:40 -0700 Mike Kravetz <mike.kravetz@oracle.com> wrote:
>
> > On 05/15/23 10:04, Mike Kravetz wrote:
> > > On 05/12/23 16:29, Mike Kravetz wrote:
> > > > On 05/12/23 14:26, James Houghton wrote:
> > > > > On Fri, May 12, 2023 at 12:20 AM Junxiao Chang <junxiao.chang@intel.com> wrote:
> > > > >
> > > > > This alone doesn't fix mapcounting for PTE-mapped HugeTLB pages. You
> > > > > need something like [1]. I can resend it if that's what we should be
> > > > > doing, but this mapcounting scheme doesn't work when the page structs
> > > > > have been freed.
> > > > >
> > > > > It seems like it was a mistake to include support for hugetlb memfds in udmabuf.
> > > >
> > > > IIUC, it was added with commit 16c243e99d33 udmabuf: Add support for mapping
> > > > hugepages (v4). Looks like it was never sent to linux-mm? That is unfortunate
> > > > as hugetlb vmemmap freeing went in at about the same time. And, as you have
> > > > noted udmabuf will not work if hugetlb vmemmap freeing is enabled.
> > > >
> > > > Sigh!
> > > >
> > > > Trying to think of a way forward.
> > > > --
> > > > Mike Kravetz
> > > >
> > > > >
> > > > > [1]: https://lore.kernel.org/linux-mm/20230306230004.1387007-2-jthoughton@google.com/
> > > > >
> > > > > - James
> > >
> > > Adding people and list on Cc: involved with commit 16c243e99d33.
> > >
> > > There are several issues with trying to map tail pages of hugetllb pages
> > > not taken into account with udmabuf. James spent quite a bit of time trying
> > > to understand and address all the issues with the HGM code. While using
> > > the scheme proposed by James, may be an approach to the mapcount issue there
> > > are also other issues that need attention. For example, I do not see how
> > > the fault code checks the state of the hugetlb page (such as poison) as none
> > > of that state is carried in tail pages.
> > >
> > > The more I think about it, the more I think udmabuf should treat hugetlb
> > > pages as hugetlb pages. They should be mapped at the appropriate level
> > > in the page table. Of course, this would impose new restrictions on the
> > > API (mmap and ioctl) that may break existing users. I have no idea how
> > > extensively udmabuf is being used with hugetlb mappings.
> >
> > Verified that using udmabug on a hugetlb mapping with vmemmap optimization will
> > BUG as:
>
> BUGs aren't good. Can we please find a way to push this along?
>
> Have we heard anything from any udmabuf people?
>
I have not heard anything. When this issue popped up, it took me by surprise.
udmabuf maintainer (Gerd Hoffmann), the people who added hugetlb support and
the list where udmabuf was developed (dri-devel@lists.freedesktop.org) have
been on cc.
My 'gut reaction' would be to remove hugetlb support from udmabuf. From a
quick look, if we really want this support then there will need to be some
API changes. For example UDMABUF_CREATE should be hugetlb page aligned
and a multiple of hugetlb page size if using a hugetlb mapping.
It would be good to know about users of the driver.
--
Mike Kravetz
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH] mm: fix hugetlb page unmap count balance issue
2023-06-07 20:53 ` Mike Kravetz
@ 2023-06-07 21:00 ` Andrew Morton
2023-06-07 21:16 ` Mike Kravetz
2023-06-08 7:59 ` Greg Kroah-Hartman
0 siblings, 2 replies; 10+ messages in thread
From: Andrew Morton @ 2023-06-07 21:00 UTC (permalink / raw)
To: Mike Kravetz
Cc: mhocko, jmarchan, Dongwon Kim, Junxiao Chang, muchun.song,
linux-kernel, dri-devel, Vivek Kasireddy, linux-mm,
James Houghton, Gerd Hoffmann, Greg Kroah-Hartman,
kirill.shutemov
On Wed, 7 Jun 2023 13:53:10 -0700 Mike Kravetz <mike.kravetz@oracle.com> wrote:
> >
> > BUGs aren't good. Can we please find a way to push this along?
> >
> > Have we heard anything from any udmabuf people?
> >
>
> I have not heard anything. When this issue popped up, it took me by surprise.
>
> udmabuf maintainer (Gerd Hoffmann), the people who added hugetlb support and
> the list where udmabuf was developed (dri-devel@lists.freedesktop.org) have
> been on cc.
Maybe Greg can suggest a way forward.
> My 'gut reaction' would be to remove hugetlb support from udmabuf. From a
> quick look, if we really want this support then there will need to be some
> API changes. For example UDMABUF_CREATE should be hugetlb page aligned
> and a multiple of hugetlb page size if using a hugetlb mapping.
>
> It would be good to know about users of the driver.
So disabling "hugetlb=on" (and adding an explanatory printk) would
suffice for now?
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH] mm: fix hugetlb page unmap count balance issue
2023-06-07 21:00 ` Andrew Morton
@ 2023-06-07 21:16 ` Mike Kravetz
2023-06-08 7:59 ` Greg Kroah-Hartman
1 sibling, 0 replies; 10+ messages in thread
From: Mike Kravetz @ 2023-06-07 21:16 UTC (permalink / raw)
To: Andrew Morton
Cc: mhocko, jmarchan, Dongwon Kim, Junxiao Chang, muchun.song,
linux-kernel, dri-devel, Vivek Kasireddy, linux-mm,
James Houghton, Gerd Hoffmann, Greg Kroah-Hartman,
kirill.shutemov
On 06/07/23 14:00, Andrew Morton wrote:
> On Wed, 7 Jun 2023 13:53:10 -0700 Mike Kravetz <mike.kravetz@oracle.com> wrote:
>
> > >
> > > BUGs aren't good. Can we please find a way to push this along?
> > >
> > > Have we heard anything from any udmabuf people?
> > >
> >
> > I have not heard anything. When this issue popped up, it took me by surprise.
> >
> > udmabuf maintainer (Gerd Hoffmann), the people who added hugetlb support and
> > the list where udmabuf was developed (dri-devel@lists.freedesktop.org) have
> > been on cc.
>
> Maybe Greg can suggest a way forward.
>
> > My 'gut reaction' would be to remove hugetlb support from udmabuf. From a
> > quick look, if we really want this support then there will need to be some
> > API changes. For example UDMABUF_CREATE should be hugetlb page aligned
> > and a multiple of hugetlb page size if using a hugetlb mapping.
> >
> > It would be good to know about users of the driver.
>
> So disabling "hugetlb=on" (and adding an explanatory printk) would
> suffice for now?
I can put together a patch to do that.
--
Mike Kravetz
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH] mm: fix hugetlb page unmap count balance issue
2023-06-07 21:00 ` Andrew Morton
2023-06-07 21:16 ` Mike Kravetz
@ 2023-06-08 7:59 ` Greg Kroah-Hartman
1 sibling, 0 replies; 10+ messages in thread
From: Greg Kroah-Hartman @ 2023-06-08 7:59 UTC (permalink / raw)
To: Andrew Morton
Cc: mhocko, jmarchan, Dongwon Kim, Junxiao Chang, muchun.song,
linux-kernel, dri-devel, Vivek Kasireddy, linux-mm,
James Houghton, Gerd Hoffmann, kirill.shutemov, Mike Kravetz
On Wed, Jun 07, 2023 at 02:00:01PM -0700, Andrew Morton wrote:
> On Wed, 7 Jun 2023 13:53:10 -0700 Mike Kravetz <mike.kravetz@oracle.com> wrote:
>
> > >
> > > BUGs aren't good. Can we please find a way to push this along?
> > >
> > > Have we heard anything from any udmabuf people?
> > >
> >
> > I have not heard anything. When this issue popped up, it took me by surprise.
> >
> > udmabuf maintainer (Gerd Hoffmann), the people who added hugetlb support and
> > the list where udmabuf was developed (dri-devel@lists.freedesktop.org) have
> > been on cc.
>
> Maybe Greg can suggest a way forward.
I'm guessing that no one is using this code then, so why don't we just
remove it entirely?
thanks,
greg k-h
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH] mm: fix hugetlb page unmap count balance issue
2023-05-15 17:04 ` [PATCH] mm: fix hugetlb page unmap count balance issue Mike Kravetz
2023-05-16 22:34 ` Mike Kravetz
@ 2023-06-19 12:27 ` Gerd Hoffmann
2023-06-20 6:23 ` Kasireddy, Vivek
1 sibling, 1 reply; 10+ messages in thread
From: Gerd Hoffmann @ 2023-06-19 12:27 UTC (permalink / raw)
To: Mike Kravetz
Cc: James Houghton, jmarchan, Dongwon Kim, Junxiao Chang, muchun.song,
linux-kernel, dri-devel, Vivek Kasireddy, linux-mm, mhocko, akpm,
kirill.shutemov
On Mon, May 15, 2023 at 10:04:42AM -0700, Mike Kravetz wrote:
> On 05/12/23 16:29, Mike Kravetz wrote:
> > On 05/12/23 14:26, James Houghton wrote:
> > > On Fri, May 12, 2023 at 12:20 AM Junxiao Chang <junxiao.chang@intel.com> wrote:
> > >
> > > This alone doesn't fix mapcounting for PTE-mapped HugeTLB pages. You
> > > need something like [1]. I can resend it if that's what we should be
> > > doing, but this mapcounting scheme doesn't work when the page structs
> > > have been freed.
> > >
> > > It seems like it was a mistake to include support for hugetlb memfds in udmabuf.
> >
> > IIUC, it was added with commit 16c243e99d33 udmabuf: Add support for mapping
> > hugepages (v4). Looks like it was never sent to linux-mm? That is unfortunate
> > as hugetlb vmemmap freeing went in at about the same time. And, as you have
> > noted udmabuf will not work if hugetlb vmemmap freeing is enabled.
> >
> > Sigh!
> >
> > Trying to think of a way forward.
> > --
> > Mike Kravetz
> >
> > >
> > > [1]: https://lore.kernel.org/linux-mm/20230306230004.1387007-2-jthoughton@google.com/
> > >
> > > - James
>
> Adding people and list on Cc: involved with commit 16c243e99d33.
>
> There are several issues with trying to map tail pages of hugetllb pages
> not taken into account with udmabuf. James spent quite a bit of time trying
> to understand and address all the issues with the HGM code. While using
> the scheme proposed by James, may be an approach to the mapcount issue there
> are also other issues that need attention. For example, I do not see how
> the fault code checks the state of the hugetlb page (such as poison) as none
> of that state is carried in tail pages.
>
> The more I think about it, the more I think udmabuf should treat hugetlb
> pages as hugetlb pages. They should be mapped at the appropriate level
> in the page table. Of course, this would impose new restrictions on the
> API (mmap and ioctl) that may break existing users. I have no idea how
> extensively udmabuf is being used with hugetlb mappings.
User of this is qemu. It can use the udmabuf driver to create host
dma-bufs for guest resources (virtio-gpu buffers), to avoid copying
data when showing the guest display in a host window.
hugetlb support is needed in case qemu guest memory is backed by
hugetlbfs. That does not imply the virtio-gpu buffers are hugepage
aligned though, udmabuf would still need to operate on smaller chunks
of memory. So with additional restrictions this will not work any
more for qemu. I'd suggest to just revert hugetlb support instead
and go back to the drawing board.
Also not sure why hugetlbfs is used for guest memory in the first place.
It used to be a thing years ago, but with the arrival of transparent
hugepages there is as far I know little reason to still use hugetlbfs.
Vivek? Dongwon?
take care,
Gerd
^ permalink raw reply [flat|nested] 10+ messages in thread
* RE: [PATCH] mm: fix hugetlb page unmap count balance issue
2023-06-19 12:27 ` Gerd Hoffmann
@ 2023-06-20 6:23 ` Kasireddy, Vivek
0 siblings, 0 replies; 10+ messages in thread
From: Kasireddy, Vivek @ 2023-06-20 6:23 UTC (permalink / raw)
To: Gerd Hoffmann
Cc: James Houghton, jmarchan@redhat.com, Kim, Dongwon, Chang, Junxiao,
muchun.song@linux.dev, linux-kernel@vger.kernel.org,
dri-devel@lists.freedesktop.org, linux-mm@kvack.org,
Hocko, Michal, akpm@linux-foundation.org,
kirill.shutemov@linux.intel.com, Mike Kravetz
Hi Gerd,
>
> On Mon, May 15, 2023 at 10:04:42AM -0700, Mike Kravetz wrote:
> > On 05/12/23 16:29, Mike Kravetz wrote:
> > > On 05/12/23 14:26, James Houghton wrote:
> > > > On Fri, May 12, 2023 at 12:20 AM Junxiao Chang
> <junxiao.chang@intel.com> wrote:
> > > >
> > > > This alone doesn't fix mapcounting for PTE-mapped HugeTLB pages.
> You
> > > > need something like [1]. I can resend it if that's what we should be
> > > > doing, but this mapcounting scheme doesn't work when the page
> structs
> > > > have been freed.
> > > >
> > > > It seems like it was a mistake to include support for hugetlb memfds in
> udmabuf.
> > >
> > > IIUC, it was added with commit 16c243e99d33 udmabuf: Add support for
> mapping
> > > hugepages (v4). Looks like it was never sent to linux-mm? That is
> unfortunate
> > > as hugetlb vmemmap freeing went in at about the same time. And, as
> you have
> > > noted udmabuf will not work if hugetlb vmemmap freeing is enabled.
> > >
> > > Sigh!
> > >
> > > Trying to think of a way forward.
> > > --
> > > Mike Kravetz
> > >
> > > >
> > > > [1]: https://lore.kernel.org/linux-mm/20230306230004.1387007-2-
> jthoughton@google.com/
> > > >
> > > > - James
> >
> > Adding people and list on Cc: involved with commit 16c243e99d33.
> >
> > There are several issues with trying to map tail pages of hugetllb pages
> > not taken into account with udmabuf. James spent quite a bit of time
> trying
> > to understand and address all the issues with the HGM code. While using
> > the scheme proposed by James, may be an approach to the mapcount
> issue there
> > are also other issues that need attention. For example, I do not see how
> > the fault code checks the state of the hugetlb page (such as poison) as none
> > of that state is carried in tail pages.
> >
> > The more I think about it, the more I think udmabuf should treat hugetlb
> > pages as hugetlb pages. They should be mapped at the appropriate level
> > in the page table. Of course, this would impose new restrictions on the
> > API (mmap and ioctl) that may break existing users. I have no idea how
> > extensively udmabuf is being used with hugetlb mappings.
>
> User of this is qemu. It can use the udmabuf driver to create host
> dma-bufs for guest resources (virtio-gpu buffers), to avoid copying
> data when showing the guest display in a host window.
>
> hugetlb support is needed in case qemu guest memory is backed by
> hugetlbfs. That does not imply the virtio-gpu buffers are hugepage
> aligned though, udmabuf would still need to operate on smaller chunks
> of memory. So with additional restrictions this will not work any
> more for qemu. I'd suggest to just revert hugetlb support instead
> and go back to the drawing board.
>
> Also not sure why hugetlbfs is used for guest memory in the first place.
> It used to be a thing years ago, but with the arrival of transparent
> hugepages there is as far I know little reason to still use hugetlbfs.
The main reason why we are interested in using hugetlbfs for guest memory
is because we observed non-trivial performance improvement while running
certain 3D heavy workloads in the guest. And, we noticed this by only
switching the Guest memory backend to include hugepages (i.e, hugetlb=on)
and with no other changes.
To address the current situation, I am readying a patch for udmabuf driver that
would add back support for mapping hugepages but without making use of
the subpages directly.
Thanks,
Vivek
>
> Vivek? Dongwon?
>
> take care,
> Gerd
^ permalink raw reply [flat|nested] 10+ messages in thread
end of thread, other threads:[~2023-06-20 6:23 UTC | newest]
Thread overview: 10+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
[not found] <20230512072036.1027784-1-junxiao.chang@intel.com>
[not found] ` <CADrL8HV25JyeaT=peaR7NWhUiaBz8LzpyFosYZ3_0ACt+twU6w@mail.gmail.com>
[not found] ` <20230512232947.GA3927@monkey>
2023-05-15 17:04 ` [PATCH] mm: fix hugetlb page unmap count balance issue Mike Kravetz
2023-05-16 22:34 ` Mike Kravetz
2023-06-07 19:03 ` Andrew Morton
2023-06-07 20:53 ` Mike Kravetz
2023-06-07 21:00 ` Andrew Morton
2023-06-07 21:16 ` Mike Kravetz
2023-06-08 7:59 ` Greg Kroah-Hartman
2023-06-07 19:27 ` David Hildenbrand
2023-06-19 12:27 ` Gerd Hoffmann
2023-06-20 6:23 ` Kasireddy, Vivek
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox