All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH] userfaultfd: reset err to be 0 when move_pages_ptes succeeded
@ 2026-08-15 10:42 Foxie Flakey
  2026-08-16  9:15 ` Mike Rapoport
  0 siblings, 1 reply; 7+ messages in thread
From: Foxie Flakey @ 2026-08-15 10:42 UTC (permalink / raw)
  To: akpm, rppt, peterx; +Cc: linux-mm, linux-kernel


An fix for edge case can occur if move_pages_ptes return -EAGAIN, later
when checked and it is EAGAIN, outer loop would retry again on same page
and succeeded but the err isn't reset so the outer loop would think need
to retry again so it goes back again and move pages again. On third attempt
move_pages_ptes will fail because it already moved and returns an error
that is not EAGAIN when outer loop checks again it sees non EAGAIN so it
dont retry and break out of loop. When loop is terminated it did not update
the "moved" variable from successful 2nd iteration.

That behaviour manifested into this at userspace

Source:      [ .. unmapped  .. ][ .. mapped    ..]
Destination: [ .. mapped    .. ][ .. unmapped  ..]
                          ^     ^
                          \     Kernel moved this far in actuality
                           What is reported to userspace on struct
                           uffdio_move's move field

When the previous behaviour is
Source:      [ .. unmapped  .. ][ .. mapped    ..]
Destination: [ .. mapped    .. ][ .. unmapped  ..]
                                ^
                                Reported to user space via uffdio_move's
                                move field

Fixes: 50944692052b ("userfaultfd: opportunistic TLB-flush batching for present pages in MOVE")
Signed-off-by: Foxie Flakey <foxieflakey@gmail.com>
---
 mm/userfaultfd.c | 6 ++++--
 1 file changed, 4 insertions(+), 2 deletions(-)

diff --git a/mm/userfaultfd.c b/mm/userfaultfd.c
index c3adedaaf7d5..595e7e232f90 100644
--- a/mm/userfaultfd.c
+++ b/mm/userfaultfd.c
@@ -2069,10 +2069,12 @@ static ssize_t move_pages(struct userfaultfd_ctx *ctx, unsigned long dst_start,
 			ret = move_pages_ptes(mm, dst_pmd, src_pmd,
 					      dst_vma, src_vma, dst_addr,
 					      src_addr, src_end - src_addr, mode);
-			if (ret < 0)
+			if (ret < 0) {
 				err = ret;
-			else
+			} else {
+				err = 0;
 				step_size = ret;
+			}
 		}

 		cond_resched();

base-commit: 62cc90241548d5570ee68e01aaba6506964e9811
-- 
2.55.0



^ permalink raw reply related	[flat|nested] 7+ messages in thread

* Re: [PATCH] userfaultfd: reset err to be 0 when move_pages_ptes succeeded
  2026-08-15 10:42 [PATCH] userfaultfd: reset err to be 0 when move_pages_ptes succeeded Foxie Flakey
@ 2026-08-16  9:15 ` Mike Rapoport
  2026-08-16  9:41   ` Foxie Flakey
  2026-08-16 15:43   ` Suren Baghdasaryan
  0 siblings, 2 replies; 7+ messages in thread
From: Mike Rapoport @ 2026-08-16  9:15 UTC (permalink / raw)
  To: Foxie Flakey, Suren Baghdasaryan; +Cc: akpm, peterx, linux-mm, linux-kernel

(adding Suren)

On Sat, Aug 15, 2026 at 05:42:12PM +0700, Foxie Flakey wrote:
> 
> An fix for edge case can occur if move_pages_ptes return -EAGAIN, later
> when checked and it is EAGAIN, outer loop would retry again on same page
> and succeeded but the err isn't reset so the outer loop would think need
> to retry again so it goes back again and move pages again. On third attempt
> move_pages_ptes will fail because it already moved and returns an error
> that is not EAGAIN when outer loop checks again it sees non EAGAIN so it
> dont retry and break out of loop. When loop is terminated it did not update
> the "moved" variable from successful 2nd iteration.
> 
> That behaviour manifested into this at userspace
> 
> Source:      [ .. unmapped  .. ][ .. mapped    ..]
> Destination: [ .. mapped    .. ][ .. unmapped  ..]
>                           ^     ^
>                           \     Kernel moved this far in actuality
>                            What is reported to userspace on struct
>                            uffdio_move's move field
> 
> When the previous behaviour is
> Source:      [ .. unmapped  .. ][ .. mapped    ..]
> Destination: [ .. mapped    .. ][ .. unmapped  ..]
>                                 ^
>                                 Reported to user space via uffdio_move's
>                                 move field
> 
> Fixes: 50944692052b ("userfaultfd: opportunistic TLB-flush batching for present pages in MOVE")
> Signed-off-by: Foxie Flakey <foxieflakey@gmail.com>

Is Foxie Flakey your real name?
Signed-off-by should be using a known identity (sorry, no anonymous contributions.)

> ---
>  mm/userfaultfd.c | 6 ++++--
>  1 file changed, 4 insertions(+), 2 deletions(-)
> 
> diff --git a/mm/userfaultfd.c b/mm/userfaultfd.c
> index c3adedaaf7d5..595e7e232f90 100644
> --- a/mm/userfaultfd.c
> +++ b/mm/userfaultfd.c
> @@ -2069,10 +2069,12 @@ static ssize_t move_pages(struct userfaultfd_ctx *ctx, unsigned long dst_start,
>  			ret = move_pages_ptes(mm, dst_pmd, src_pmd,
>  					      dst_vma, src_vma, dst_addr,
>  					      src_addr, src_end - src_addr, mode);
> -			if (ret < 0)
> +			if (ret < 0) {
>  				err = ret;
> -			else
> +			} else {
> +				err = 0;
>  				step_size = ret;
> +			}
>  		}
> 
>  		cond_resched();
> 
> base-commit: 62cc90241548d5570ee68e01aaba6506964e9811
> -- 
> 2.55.0
> 

-- 
Sincerely yours,
Mike.


^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH] userfaultfd: reset err to be 0 when move_pages_ptes succeeded
  2026-08-16  9:15 ` Mike Rapoport
@ 2026-08-16  9:41   ` Foxie Flakey
  2026-08-16 11:45     ` Mike Rapoport
  2026-08-16 15:43   ` Suren Baghdasaryan
  1 sibling, 1 reply; 7+ messages in thread
From: Foxie Flakey @ 2026-08-16  9:41 UTC (permalink / raw)
  To: Mike Rapoport, Suren Baghdasaryan; +Cc: akpm, peterx, linux-mm, linux-kernel

On Sun, 16 Aug 2026, Mike Rapoport wrote:

> (adding Suren)
>
> On Sat, Aug 15, 2026 at 05:42:12PM +0700, Foxie Flakey wrote:
> >
> > Fixes: 50944692052b ("userfaultfd: opportunistic TLB-flush batching for present pages in MOVE")
> > Signed-off-by: Foxie Flakey <foxieflakey@gmail.com>
>
> Is Foxie Flakey your real name?
> Signed-off-by should be using a known identity (sorry, no anonymous contributions.)

Thanks for response. Sorry, no. A question, by "known identity" there is it
means my real/legal name? because I have read a document about Signed-off-by
and because on internet I'm known as Foxie Flakey in multiple places, I
thought its fine.



^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH] userfaultfd: reset err to be 0 when move_pages_ptes succeeded
  2026-08-16  9:41   ` Foxie Flakey
@ 2026-08-16 11:45     ` Mike Rapoport
  2026-08-16 13:58       ` Foxie Flakey
  0 siblings, 1 reply; 7+ messages in thread
From: Mike Rapoport @ 2026-08-16 11:45 UTC (permalink / raw)
  To: Foxie Flakey; +Cc: Suren Baghdasaryan, akpm, peterx, linux-mm, linux-kernel

On Sun, Aug 16, 2026 at 04:41:07PM +0700, Foxie Flakey wrote:
> On Sun, 16 Aug 2026, Mike Rapoport wrote:
> 
> > (adding Suren)
> >
> > On Sat, Aug 15, 2026 at 05:42:12PM +0700, Foxie Flakey wrote:
> > >
> > > Fixes: 50944692052b ("userfaultfd: opportunistic TLB-flush batching for present pages in MOVE")
> > > Signed-off-by: Foxie Flakey <foxieflakey@gmail.com>
> >
> > Is Foxie Flakey your real name?
> > Signed-off-by should be using a known identity (sorry, no anonymous contributions.)
> 
> Thanks for response. Sorry, no. A question, by "known identity" there is it
> means my real/legal name? because I have read a document about Signed-off-by
> and because on internet I'm known as Foxie Flakey in multiple places, I
> thought its fine.

It does not have to be the legal name, but we do prefer real/preferred
names over nicknames:

CNCF's clarification on DCO explains it well:
https://github.com/cncf/foundation/blob/659fd32c86dc/dco-guidelines.md
 

-- 
Sincerely yours,
Mike.

^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH] userfaultfd: reset err to be 0 when move_pages_ptes succeeded
  2026-08-16 11:45     ` Mike Rapoport
@ 2026-08-16 13:58       ` Foxie Flakey
  0 siblings, 0 replies; 7+ messages in thread
From: Foxie Flakey @ 2026-08-16 13:58 UTC (permalink / raw)
  To: Mike Rapoport; +Cc: Suren Baghdasaryan, akpm, peterx, linux-mm, linux-kernel

On Sun, 16 Aug 2026, Mike Rapoport wrote:
> On Sun, Aug 16, 2026 at 04:41:07PM +0700, Foxie Flakey wrote:
> > On Sun, 16 Aug 2026, Mike Rapoport wrote:
> >
> > > (adding Suren)
> > >
> > > On Sat, Aug 15, 2026 at 05:42:12PM +0700, Foxie Flakey wrote:
> > > >
> > > > Fixes: 50944692052b ("userfaultfd: opportunistic TLB-flush batching for present pages in MOVE")
> > > > Signed-off-by: Foxie Flakey <foxieflakey@gmail.com>
> > >
> > > Is Foxie Flakey your real name?
> > > Signed-off-by should be using a known identity (sorry, no anonymous contributions.)
> >
> > Thanks for response. Sorry, no. A question, by "known identity" there is it
> > means my real/legal name? because I have read a document about Signed-off-by
> > and because on internet I'm known as Foxie Flakey in multiple places, I
> > thought its fine.
>
> It does not have to be the legal name, but we do prefer real/preferred
> names over nicknames:
>
> CNCF's clarification on DCO explains it well:
> https://github.com/cncf/foundation/blob/659fd32c86dc/dco-guidelines.md

I see thank you. So the Signed-off-by would be fine as in if I prefer not
to put my legal/real/birth name. Would it fine to keep as "Foxie Flakey"?
Because I have just read that guideline, Foxie Flakey isn't anonymous id
nor false name.

A small confirmation question due not sure with results from internet about
where should I add Suren. I looked that you placed Suren in Cc. After the
initial To (your first reply to my mail), on followup replies I should
place Suren or anyone else to Cc filed after initial To, am I correct?


^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH] userfaultfd: reset err to be 0 when move_pages_ptes succeeded
  2026-08-16  9:15 ` Mike Rapoport
  2026-08-16  9:41   ` Foxie Flakey
@ 2026-08-16 15:43   ` Suren Baghdasaryan
  2026-08-16 15:58     ` Foxie Flakey
  1 sibling, 1 reply; 7+ messages in thread
From: Suren Baghdasaryan @ 2026-08-16 15:43 UTC (permalink / raw)
  To: Mike Rapoport; +Cc: Foxie Flakey, akpm, peterx, linux-mm, linux-kernel

On Sun, Aug 16, 2026 at 2:15 AM Mike Rapoport <rppt@kernel.org> wrote:
>
> (adding Suren)

Thanks Mike!

>
> On Sat, Aug 15, 2026 at 05:42:12PM +0700, Foxie Flakey wrote:
> >
> > An fix for edge case can occur if move_pages_ptes return -EAGAIN, later
> > when checked and it is EAGAIN, outer loop would retry again on same page
> > and succeeded but the err isn't reset so the outer loop would think need
> > to retry again so it goes back again and move pages again. On third attempt
> > move_pages_ptes will fail because it already moved and returns an error
> > that is not EAGAIN when outer loop checks again it sees non EAGAIN so it
> > dont retry and break out of loop. When loop is terminated it did not update
> > the "moved" variable from successful 2nd iteration.
> >
> > That behaviour manifested into this at userspace
> >
> > Source:      [ .. unmapped  .. ][ .. mapped    ..]
> > Destination: [ .. mapped    .. ][ .. unmapped  ..]
> >                           ^     ^
> >                           \     Kernel moved this far in actuality
> >                            What is reported to userspace on struct
> >                            uffdio_move's move field
> >
> > When the previous behaviour is
> > Source:      [ .. unmapped  .. ][ .. mapped    ..]
> > Destination: [ .. mapped    .. ][ .. unmapped  ..]
> >                                 ^
> >                                 Reported to user space via uffdio_move's
> >                                 move field

This description left me scratching my head. If I understand the
problem correctly, the issue is that the err is not cleared after we
decided that we need to retry. If so, how about a simpler explanation:

During move_pages() operation, when move_pages_ptes() returns EAGAIN,
the error code is not cleared even after we processed it. This leads
to a successful retry but then the same pages are retried again due to
the stale error code. This time move fails because pages are already
moved, loop is terminated and move_pages() reports a failure.
Clear the error code once we processes EAGAIN.

> >
> > Fixes: 50944692052b ("userfaultfd: opportunistic TLB-flush batching for present pages in MOVE")
> > Signed-off-by: Foxie Flakey <foxieflakey@gmail.com>
>
> Is Foxie Flakey your real name?
> Signed-off-by should be using a known identity (sorry, no anonymous contributions.)
>
> > ---
> >  mm/userfaultfd.c | 6 ++++--
> >  1 file changed, 4 insertions(+), 2 deletions(-)
> >
> > diff --git a/mm/userfaultfd.c b/mm/userfaultfd.c
> > index c3adedaaf7d5..595e7e232f90 100644
> > --- a/mm/userfaultfd.c
> > +++ b/mm/userfaultfd.c
> > @@ -2069,10 +2069,12 @@ static ssize_t move_pages(struct userfaultfd_ctx *ctx, unsigned long dst_start,
> >                       ret = move_pages_ptes(mm, dst_pmd, src_pmd,
> >                                             dst_vma, src_vma, dst_addr,
> >                                             src_addr, src_end - src_addr, mode);
> > -                     if (ret < 0)
> > +                     if (ret < 0) {
> >                               err = ret;
> > -                     else
> > +                     } else {
> > +                             err = 0;
> >                               step_size = ret;
> > +                     }

This fix is wrong. It resets the err before we process it and
determine that a retry is needed.
A proper fix is to reset it later here:

if (err) {
-        if (err == -EAGAIN)
+       if (err == -EAGAIN) {
+              err = 0;
               continue;
+       }
         break;
}

> >               }
> >
> >               cond_resched();
> >
> > base-commit: 62cc90241548d5570ee68e01aaba6506964e9811
> > --
> > 2.55.0
> >
>
> --
> Sincerely yours,
> Mike.


^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH] userfaultfd: reset err to be 0 when move_pages_ptes succeeded
  2026-08-16 15:43   ` Suren Baghdasaryan
@ 2026-08-16 15:58     ` Foxie Flakey
  0 siblings, 0 replies; 7+ messages in thread
From: Foxie Flakey @ 2026-08-16 15:58 UTC (permalink / raw)
  To: Suren Baghdasaryan; +Cc: Mike Rapoport, akpm, peterx, linux-mm, linux-kernel

[-- Attachment #1: Type: text/plain, Size: 4424 bytes --]

On Sun, 16 Aug 2026, Suren Baghdasaryan wrote:

> On Sun, Aug 16, 2026 at 2:15 AM Mike Rapoport <rppt@kernel.org> wrote:
> >
> > (adding Suren)
>
> Thanks Mike!
>
> >
> > On Sat, Aug 15, 2026 at 05:42:12PM +0700, Foxie Flakey wrote:
> > >
> > > An fix for edge case can occur if move_pages_ptes return -EAGAIN, later
> > > when checked and it is EAGAIN, outer loop would retry again on same page
> > > and succeeded but the err isn't reset so the outer loop would think need
> > > to retry again so it goes back again and move pages again. On third attempt
> > > move_pages_ptes will fail because it already moved and returns an error
> > > that is not EAGAIN when outer loop checks again it sees non EAGAIN so it
> > > dont retry and break out of loop. When loop is terminated it did not update
> > > the "moved" variable from successful 2nd iteration.
> > >
> > > That behaviour manifested into this at userspace
> > >
> > > Source:      [ .. unmapped  .. ][ .. mapped    ..]
> > > Destination: [ .. mapped    .. ][ .. unmapped  ..]
> > >                           ^     ^
> > >                           \     Kernel moved this far in actuality
> > >                            What is reported to userspace on struct
> > >                            uffdio_move's move field
> > >
> > > When the previous behaviour is
> > > Source:      [ .. unmapped  .. ][ .. mapped    ..]
> > > Destination: [ .. mapped    .. ][ .. unmapped  ..]
> > >                                 ^
> > >                                 Reported to user space via uffdio_move's
> > >                                 move field
>
> This description left me scratching my head. If I understand the
> problem correctly, the issue is that the err is not cleared after we
> decided that we need to retry. If so, how about a simpler explanation:

Sorry for the bad explanation, but that is correct. To repeat again to
make sure I understood correct, move_pages() wrongfully retries to move
again due stale err.

> During move_pages() operation, when move_pages_ptes() returns EAGAIN,
> the error code is not cleared even after we processed it. This leads
> to a successful retry but then the same pages are retried again due to
> the stale error code. This time move fails because pages are already
> moved, loop is terminated and move_pages() reports a failure.
> Clear the error code once we processes EAGAIN.

Thank you, I'll update in v2. I'm waiting for answer from Mike whether
Foxie Flakey is fine in Signed-off-by so I don't create too many revisions
when I can combine feedbacks into one.

> > >
> > > Fixes: 50944692052b ("userfaultfd: opportunistic TLB-flush batching for present pages in MOVE")
> > > Signed-off-by: Foxie Flakey <foxieflakey@gmail.com>
> >
> > Is Foxie Flakey your real name?
> > Signed-off-by should be using a known identity (sorry, no anonymous contributions.)
> >
> > > ---
> > >  mm/userfaultfd.c | 6 ++++--
> > >  1 file changed, 4 insertions(+), 2 deletions(-)
> > >
> > > diff --git a/mm/userfaultfd.c b/mm/userfaultfd.c
> > > index c3adedaaf7d5..595e7e232f90 100644
> > > --- a/mm/userfaultfd.c
> > > +++ b/mm/userfaultfd.c
> > > @@ -2069,10 +2069,12 @@ static ssize_t move_pages(struct userfaultfd_ctx *ctx, unsigned long dst_start,
> > >                       ret = move_pages_ptes(mm, dst_pmd, src_pmd,
> > >                                             dst_vma, src_vma, dst_addr,
> > >                                             src_addr, src_end - src_addr, mode);
> > > -                     if (ret < 0)
> > > +                     if (ret < 0) {
> > >                               err = ret;
> > > -                     else
> > > +                     } else {
> > > +                             err = 0;
> > >                               step_size = ret;
> > > +                     }
>
> This fix is wrong. It resets the err before we process it and
> determine that a retry is needed.
> A proper fix is to reset it later here:
>
> if (err) {
> -        if (err == -EAGAIN)
> +       if (err == -EAGAIN) {
> +              err = 0;
>                continue;
> +       }
>          break;
> }

I see, that one make more sense after thinking about it that retry should
clear err before retrying.

> > >               }
> > >
> > >               cond_resched();
> > >
> > > base-commit: 62cc90241548d5570ee68e01aaba6506964e9811
> > > --
> > > 2.55.0
> > >
> >
> > --
> > Sincerely yours,
> > Mike.
>

^ permalink raw reply	[flat|nested] 7+ messages in thread

end of thread, other threads:[~2026-08-16 15:58 UTC | newest]

Thread overview: 7+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-15 10:42 [PATCH] userfaultfd: reset err to be 0 when move_pages_ptes succeeded Foxie Flakey
2026-08-16  9:15 ` Mike Rapoport
2026-08-16  9:41   ` Foxie Flakey
2026-08-16 11:45     ` Mike Rapoport
2026-08-16 13:58       ` Foxie Flakey
2026-08-16 15:43   ` Suren Baghdasaryan
2026-08-16 15:58     ` Foxie Flakey

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.