From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by smtp.lore.kernel.org (Postfix) with ESMTP id 9C402C54F30 for ; Tue, 27 May 2025 08:37:34 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id 396226B008C; Tue, 27 May 2025 04:37:34 -0400 (EDT) Received: by kanga.kvack.org (Postfix, from userid 40) id 347286B0092; Tue, 27 May 2025 04:37:34 -0400 (EDT) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id 25C526B0093; Tue, 27 May 2025 04:37:34 -0400 (EDT) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0010.hostedemail.com [216.40.44.10]) by kanga.kvack.org (Postfix) with ESMTP id 084D46B008C for ; Tue, 27 May 2025 04:37:34 -0400 (EDT) Received: from smtpin24.hostedemail.com (a10.router.float.18 [10.200.18.1]) by unirelay02.hostedemail.com (Postfix) with ESMTP id 9F044120375 for ; Tue, 27 May 2025 08:37:33 +0000 (UTC) X-FDA: 83488033986.24.B5578B5 Received: from mail-pj1-f49.google.com (mail-pj1-f49.google.com [209.85.216.49]) by imf25.hostedemail.com (Postfix) with ESMTP id AD8BDA0009 for ; Tue, 27 May 2025 08:37:31 +0000 (UTC) Authentication-Results: imf25.hostedemail.com; dkim=pass header.d=gmail.com header.s=20230601 header.b=UoVgFCOj; spf=pass (imf25.hostedemail.com: domain of 21cnbao@gmail.com designates 209.85.216.49 as permitted sender) smtp.mailfrom=21cnbao@gmail.com; dmarc=pass (policy=none) header.from=gmail.com ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1748335051; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references:dkim-signature; bh=vG068GBhBEM6vEx1ff4IYRmbQ2abHAH9u0tmgVxbbpY=; b=RkEdhY8Vwzjk1UC3mp0RJFI0G+2CMEnGeEYc5WXYy3NCNONguKR3ibakfKMo1Q5ArzEFTI gUaHOYXKdHJdPHYntoZpc/q7RtmyjFSJ+ObMJL4KdX2Bz9b8LRWgEnMeNFGTPxTlBVnJrj 2Aed2Q6lDh/bxabAog8bMRXBtaKtmnM= ARC-Authentication-Results: i=1; imf25.hostedemail.com; dkim=pass header.d=gmail.com header.s=20230601 header.b=UoVgFCOj; spf=pass (imf25.hostedemail.com: domain of 21cnbao@gmail.com designates 209.85.216.49 as permitted sender) smtp.mailfrom=21cnbao@gmail.com; dmarc=pass (policy=none) header.from=gmail.com ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1748335051; a=rsa-sha256; cv=none; b=Ar0LHLkf6bQ8SLdPQxDd7M3s6qmaTZBF+BkKa3HIxXWIg49pKFEwgd+Ti7eq9LRXSRfyBg RkiwLJ9rq/fgPG+gLZx0HwYQekmyKxfyvR1cjutzRSUqPXtzW7W3/YKs7LEmLTcNYWZt66 1F8AuJl+3RF42sSvD8OzIMC2SwGVHIU= Received: by mail-pj1-f49.google.com with SMTP id 98e67ed59e1d1-311374f95b8so1756583a91.0 for ; Tue, 27 May 2025 01:37:31 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1748335050; x=1748939850; darn=kvack.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to; bh=vG068GBhBEM6vEx1ff4IYRmbQ2abHAH9u0tmgVxbbpY=; b=UoVgFCOjHaKaYiZg2AjCEaVuzTZQht58sQ5JgaDDWlhN/9h39vBfNzRpp5Ct27WEly UJHbyvWvHLGTxOi0v2fj5KnxaCTbHwU0perXSHIEpDmsXnyMc3ipgBrlyEP7M7DbHm7L RbqNDToctd55dzE8z6yQ5HB0LmoUNpd5zwP1y9t5wHJt7A5x7wdCbt9yGYznK1cp2z1N O5U/SiXcJ2wlgGbOdaQJPREOhwe9TqAPprij8AXzIZbZWzBm9imrJwkxryf+bEtUNv+f P1EKL0dNDaCVZoduLJGXHRtEBEEpzbllnTKl0rFApEKGYgxTng9KheQ+/8lRFc6VoC8b Pq2g== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1748335050; x=1748939850; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=vG068GBhBEM6vEx1ff4IYRmbQ2abHAH9u0tmgVxbbpY=; b=cNbR/oCQmkCEn4soq9XPLAXhgKK8Qkvut4/y5iJ6zuqZTQrE1qsrIqCdOPp1ur+Czb a1KREXaQtDcPnWlSg5YRFQ+iOC/frd9bfeymAdY4ofmcJeKiuPleheeShaV+Dqz69Wa2 PQR75oiZlMXopSYmCVfRTBWz/Bj6/YU2v0hpcXWcx20u4DgD35nB/6bnVjmvuMRBt5hE vyBR3adNom9i2kaDSRFHIdAB9bOvChjWvkTudsFrIoRFqDf7eUMzwa7XwqgmaMlHdNEH DBjxepP9pGyP2bjuo1/ilP34tABgTs5ShlpO5sX8Qqbuiz4sKJzrDiUTeHEphbV2+Be1 rjBA== X-Forwarded-Encrypted: i=1; AJvYcCWDH8IgJrfYjyKn7tfD5RjdcCB1KWKrFs9Spi62hRyS2rH4LHtBjB9F/qLqe5ccZclu6VwsP7KYqA==@kvack.org X-Gm-Message-State: AOJu0Ywxl/9xgYCpAci8MExttVflXp7ulLxSnm200H+UnAORLTe4NVli mtH8w2v4ZM5ZD6VH73BxRGu1nfkR33WU7vU1QQ/GtDYMefWHSq81h7k/ X-Gm-Gg: ASbGncsAuUuxbsJ0X2ur+LXien1D5/JenLYpnmzyjSZx5h9sg/n5xRCmnxnMObKo1lE HAgj8ONJdHedl84zqo/5s7OjN2iX5+Mac/uolWm1X1jiuu5RDmJgcoEGtb2gdoI6UNCFK2XxzLA YRRqndJFRZfYLKkHlgvVBdDglORYTw1FzjXKPLjPAmxiGPMqb0kSmRKwCGGo+vuwVT9ymTnqRwl LYGfAPmUXR1BLeCIji9kZM0FBtm/rfKR2/No1DJ3RfXfvrEnHsfz9kssbpikdXPxKEEWFimBCca AgiYhAkCZSLNH3rM02ojxErbPqjjWpL8fmI0FnyKazfEIB+/XAyrn/3Yjr1lxoUq1ipd X-Google-Smtp-Source: AGHT+IGgUOniQeowny4ikpFfBKzAvvOgmt8e+sabsiZxvOnXLLrj+o/hCy59jr5nmXu1emXHWTzEGQ== X-Received: by 2002:a17:90b:35cc:b0:306:b6f7:58ba with SMTP id 98e67ed59e1d1-3110f31c313mr16749250a91.6.1748335050418; Tue, 27 May 2025 01:37:30 -0700 (PDT) Received: from Barrys-MBP.hub ([118.92.145.159]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-30f365d0d44sm13680014a91.21.2025.05.27.01.37.26 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Tue, 27 May 2025 01:37:29 -0700 (PDT) From: Barry Song <21cnbao@gmail.com> To: david@redhat.com Cc: 21cnbao@gmail.com, aarcange@redhat.com, akpm@linux-foundation.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, lokeshgidra@google.com, peterx@redhat.com, ryncsn@gmail.com, surenb@google.com Subject: Re: [BUG]userfaultfd_move fails to move a folio when swap-in occurs concurrently with swap-out Date: Tue, 27 May 2025 20:37:22 +1200 Message-Id: <20250527083722.27309-1-21cnbao@gmail.com> X-Mailer: git-send-email 2.39.3 (Apple Git-146) In-Reply-To: References: MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit X-Rspamd-Server: rspam03 X-Rspamd-Queue-Id: AD8BDA0009 X-Stat-Signature: dydqpu4nxryk3gwz5ad7abtx745uk1ju X-Rspam-User: X-HE-Tag: 1748335051-852858 X-HE-Meta: U2FsdGVkX19grzT61jreg2zUzbhkUGusX/D5SGD+1nAUn8KADfT0e9zJd0fQsHpku+Hty+U8n8C+ZgrYEQfIrOmB5SaSt+It7CY2LQkSvbKToETo21BGunpXewOvtljAUGgNLYTlY9k+El1glvhpfchx/5eVtanzsN5ynbavJUpl0IrI1pYV3UtspH9c1xRLXX3hgqIf+JnKVsKwKFHYjU11mwddctyAuDAJXSecyzCd5AosKR6mCSnV4iDqXPCFRLYfzqHYmkBqg+a76CZhLzuiKR97d3dFf3ydEK7rq3Aza+CU0T8EDvyON9EywrCbCFXq2bo3/uiB+LNDqTEboRhXk43nYv3+uwBEu9HKDi/dbedb06Jo7/9cRBa+HLiRekmmaVyVammCHVTXhlSy/w9ocEfyM1L7uxfkszyraACB2A1H2+ct4BtqxRq/u8uZf32K46prv9Vhfr3bXrFkYtXXxjFhmqNSmj6FULmGhbp6L/9SwYWT/77eik7hj/kkA8sTPtPiCBGMMAOqaAX1MI3shqu4m9Xn1Dg1Y9zMYpy/ofTPJ9MjverlByhUPaKBEBG/wNjCE5MrZiUJxpZF8VASwDMV1qIsT2CjKeqhZqjFhetrIb/Qs827A2eeD986G5e1DbpVlCAlbIKaXrsRRiX2tM+HDUTllpam78CBSfiqrOJrwMdiwIrV8IeDPkLrwq32mYEePz9WQvC2xG6zUfvGRii+JoRD75eWHIJI8+mP2ymA+vte/n3oxqYeJ48Z09AOvhy6+oqMVLL8N7dyCrDTH6ZECRM8mT2UyeHHoIJBngbw1q8NvuYWRa7VqjP5Y1i0VmZPxA1cE+cTzJK7D+zymrLTbgaUWz3U8iFpMInmdJdyNbcEEWm64NZIQODCqrCtRg0n2ym6xWCn6X4ot8+mFewHpfTkwg2+RuTWQKW9sRyuWcv3zI5IZ8giIGBNLFUGpTC/15mX+vWrhAD yRX9/OQO 7A4rA6ln5RHh0aYWyGjiO1teFZ3CnD39F869Ow62rWlvYNpVO6IGG5+394It/r9pGgWqAjQj0Dj7zKGti6Fk3k+2YbJZBG8q81Rh3MEZ8m9G8hTmLmshNC/Tn9pYkh+RN4vWeccrImD0yqRi8Vd6UdgFGsqrjbg5pxADo1pLYSpANY8iGuwTTG/+qnfeWXJw6yR5WHLNzjGnwtOR8Zo2+N1BA0qFnDx1h0JNAghG+mnq7DXKVfZOPrjRs7ulhmr1HI902GMOIyJUz55nDixQ8ZcDD20Mtq55pg6NfX+90q4b1oCiJMtPUm8Vz8mH1EJeHqmZ+1VDUcyw8dsfiDFWWKi+dG2ureUPddMw4aBETvXjQrvUAOaOenoeKFIP4fMnfF7ZwxtU7/Hl6BhmF6nNEgYiUCjHBKmNzbakegjfCf30T+YvKrVQoQOfVILPou+3vdYMugQpvrAsPCwvv1K2GD5eS/ba5h4DsztJFpykjQ0tK0HQ= X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: List-Subscribe: List-Unsubscribe: On Tue, May 27, 2025 at 4:17 PM Barry Song <21cnbao@gmail.com> wrote: > > On Tue, May 27, 2025 at 12:39 AM David Hildenbrand wrote: > > > > On 23.05.25 01:23, Barry Song wrote: > > > Hi All, > > > > Hi! > > > > > > > > I'm encountering another bug that can be easily reproduced using the small > > > program below[1], which performs swap-out and swap-in in parallel. > > > > > > The issue occurs when a folio is being swapped out while it is accessed > > > concurrently. In this case, do_swap_page() handles the access. However, > > > because the folio is under writeback, do_swap_page() completely removes > > > its exclusive attribute. > > > > > > do_swap_page: > > >                 } else if (exclusive && folio_test_writeback(folio) && > > >                            data_race(si->flags & SWP_STABLE_WRITES)) { > > >                          ... > > >                          exclusive = false; > > > > > > As a result, userfaultfd_move() will return -EBUSY, even though the > > > folio is not shared and is in fact exclusively owned. > > > > > >                          folio = vm_normal_folio(src_vma, src_addr, > > > orig_src_pte); > > >                          if (!folio || !PageAnonExclusive(&folio->page)) { > > >                                  spin_unlock(src_ptl); > > > +                               pr_err("%s %d folio:%lx exclusive:%d > > > swapcache:%d\n", > > > +                                       __func__, __LINE__, folio, > > > PageAnonExclusive(&folio->page), > > > +                                       folio_test_swapcache(folio)); > > >                                  err = -EBUSY; > > >                                  goto out; > > >                          } > > > > > > I understand that shared folios should not be moved. However, in this > > > case, the folio is not shared, yet its exclusive flag is not set. > > > > > > Therefore, I believe PageAnonExclusive is not a reliable indicator of > > > whether a folio is truly exclusive to a process. > > > > It is. The flag *not* being set is not a reliable indicator whether it > > is really shared. ;) > > > > The reason why we have this PAE workaround (dropping the flag) in place > > is because the page must not be written to (SWP_STABLE_WRITES). CoW > > reuse is not possible. > > > > uffd moving that page -- and in that same process setting it writable, > > see move_present_pte()->pte_mkwrite() -- would be very bad. > > An alternative approach is to make the folio writable only when we are > reasonably certain it is exclusive; otherwise, it remains read-only. If the > destination is later written to and the folio has become exclusive, it can > be reused directly. If not, a copy-on-write will occur on the destination > address, transparently to userspace. This avoids Lokesh’s userspace-based > strategy, which requires forcing a write to the source address. Conceptually, I mean something like this: diff --git a/mm/userfaultfd.c b/mm/userfaultfd.c index bc473ad21202..70eaabf4f1a3 100644 --- a/mm/userfaultfd.c +++ b/mm/userfaultfd.c @@ -1047,7 +1047,8 @@ static int move_present_pte(struct mm_struct *mm, } if (folio_test_large(src_folio) || folio_maybe_dma_pinned(src_folio) || - !PageAnonExclusive(&src_folio->page)) { + (!PageAnonExclusive(&src_folio->page) && + folio_mapcount(src_folio) != 1)) { err = -EBUSY; goto out; } @@ -1070,7 +1071,8 @@ static int move_present_pte(struct mm_struct *mm, #endif if (pte_dirty(orig_src_pte)) orig_dst_pte = pte_mkdirty(orig_dst_pte); - orig_dst_pte = pte_mkwrite(orig_dst_pte, dst_vma); + if (PageAnonExclusive(&src_folio->page)) + orig_dst_pte = pte_mkwrite(orig_dst_pte, dst_vma); set_pte_at(mm, dst_addr, dst_pte, orig_dst_pte); out: @@ -1268,7 +1270,8 @@ static int move_pages_pte(struct mm_struct *mm, pmd_t *dst_pmd, pmd_t *src_pmd, } folio = vm_normal_folio(src_vma, src_addr, orig_src_pte); - if (!folio || !PageAnonExclusive(&folio->page)) { + if (!folio || (!PageAnonExclusive(&folio->page) && + folio_mapcount(folio) != 1)) { spin_unlock(src_ptl); err = -EBUSY; goto out; I'm not trying to push this approach—unless Lokesh clearly sees that it could reduce userspace noise. I'm mainly just curious how we might make the fixup transparent to userspace. :-) > > > > > > > > > The kernel log output is shown below: > > > [   23.009516] move_pages_pte 1285 folio:fffffdffc01bba40 exclusive:0 > > > swapcache:1 > > > > > > I'm still struggling to find a real fix; it seems quite challenging. > > > > PAE tells you that you can immediately write to that page without going > > through CoW. However, here, CoW is required. > > > > > Please let me know if you have any ideas. In any case It seems > > > userspace should fall back to userfaultfd_copy. > > > > We could try detecting whether the page is now exclusive, to reset PAE. > > That will only be possible after writeback completed, so it adds > > complexity without being able to move the page in all cases (during > > writeback). > > > > Letting userspace deal with that in these rate scenarios is > > significantly easier. > > Right, this appears to introduce the least change—essentially none—to the > kernel, while shifting more noise to userspace :-) > > > > > -- > > Cheers, > > > > David / dhildenb > > > Thanks Barry