From: Gao Xiang <hsiangkao@linux.alibaba.com>
To: Joanne Koong <joannelkoong@gmail.com>
Cc: miklos@szeredi.hu, linux-fsdevel@vger.kernel.org, osandov@fb.com,
kernel-team@meta.com
Subject: Re: [PATCH] fuse: fix readahead reclaim deadlock
Date: Tue, 30 Sep 2025 10:35:50 +0800 [thread overview]
Message-ID: <e5b4985d-18c3-4609-b1f7-2425f161375d@linux.alibaba.com> (raw)
In-Reply-To: <a517168d-840f-483b-b9a1-4b9c417df217@linux.alibaba.com>
On 2025/9/30 10:21, Gao Xiang wrote:
>
>
> On 2025/9/30 01:25, Joanne Koong wrote:
>> On Fri, Sep 26, 2025 at 12:19 AM Gao Xiang <hsiangkao@linux.alibaba.com> wrote:
>>>
>>> On 2025/9/26 14:51, Gao Xiang wrote:
>>>>
>>>> On 2025/9/26 06:44, Joanne Koong wrote:
>>>>> A deadlock can occur if the server triggers reclaim while servicing a
>>>>> readahead request, and reclaim attempts to evict the inode of the file
>>>>> being read ahead:
>>>>>
>>>>>>>> stack_trace(1504735)
>>>>> folio_wait_bit_common (mm/filemap.c:1308:4)
>>>>> folio_lock (./include/linux/pagemap.h:1052:3)
>>>>> truncate_inode_pages_range (mm/truncate.c:336:10)
>>>>> fuse_evict_inode (fs/fuse/inode.c:161:2)
>>>>> evict (fs/inode.c:704:3)
>>>>> dentry_unlink_inode (fs/dcache.c:412:3)
>>>>> __dentry_kill (fs/dcache.c:615:3)
>>>>> shrink_kill (fs/dcache.c:1060:12)
>>>>> shrink_dentry_list (fs/dcache.c:1087:3)
>>>>> prune_dcache_sb (fs/dcache.c:1168:2)
>>>>> super_cache_scan (fs/super.c:221:10)
>>>>> do_shrink_slab (mm/shrinker.c:435:9)
>>>>> shrink_slab (mm/shrinker.c:626:10)
>>>>> shrink_node (mm/vmscan.c:5951:2)
>>>>> shrink_zones (mm/vmscan.c:6195:3)
>>>>> do_try_to_free_pages (mm/vmscan.c:6257:3)
>>>>> do_swap_page (mm/memory.c:4136:11)
>>>>> handle_pte_fault (mm/memory.c:5562:10)
>>>>> handle_mm_fault (mm/memory.c:5870:9)
>>>>> do_user_addr_fault (arch/x86/mm/fault.c:1338:10)
>>>>> handle_page_fault (arch/x86/mm/fault.c:1481:3)
>>>>> exc_page_fault (arch/x86/mm/fault.c:1539:2)
>>>>> asm_exc_page_fault+0x22/0x27
>>>>>
>>>>> During readahead, the folio is locked. When fuse_evict_inode() is
>>>>> called, it attempts to remove all folios associated with the inode from
>>>>> the page cache (truncate_inode_pages_range()), which requires acquiring
>>>>> the folio lock. If the server triggers reclaim while servicing a
>>>>> readahead request, reclaim will block indefinitely waiting for the folio
>>>>> lock, while readahead cannot relinquish the lock because it is itself
>>>>> blocked in reclaim, resulting in a deadlock.
>>>>>
>>>>> The inode is only evicted if it has no remaining references after its
>>>>> dentry is unlinked. Since readahead is asynchronous, it is not
>>>>> guaranteed that the inode will have any references at this point.
>>>>>
>>>>> This fixes the deadlock by holding a reference on the inode while
>>>>> readahead is in progress, which prevents the inode from being evicted
>>>>> until readahead completes. Additionally, this also prevents a malicious
>>>>> or buggy server from indefinitely blocking kswapd if it never fulfills a
>>>>> readahead request.
>>>>>
>>>>> Signed-off-by: Joanne Koong <joannelkoong@gmail.com>
>>>>> Reported-by: Omar Sandoval <osandov@fb.com>
>>>>> ---
>>>>> fs/fuse/file.c | 7 +++++++
>>>>> 1 file changed, 7 insertions(+)
>>>>>
>>>>> diff --git a/fs/fuse/file.c b/fs/fuse/file.c
>>>>> index f1ef77a0be05..8e759061b843 100644
>>>>> --- a/fs/fuse/file.c
>>>>> +++ b/fs/fuse/file.c
>>>>> @@ -893,6 +893,7 @@ static void fuse_readpages_end(struct fuse_mount *fm, struct fuse_args *args,
>>>>> if (ia->ff)
>>>>> fuse_file_put(ia->ff, false);
>>>>> + iput(inode);
>>>>
>>>> It's somewhat odd to use `igrab` and `iput` in the read(ahead)
>>>> context.
>>>>
>>>> I wonder for FUSE, if it's possible to just wait ongoing
>>>> locked folios when i_count == 0 (e.g. in .drop_inode) before
>>>> adding into lru so that the final inode eviction won't wait
>>>> its readahead requests itself so that deadlock like this can
>>>> be avoided.
>>>
>>> Oh, it was in the dentry LRU list instead, I don't think it can
>>> work.
>>>
>>> Or normally the kernel filesystem uses GFP_NOFS to avoid such
>>> deadlock (see `if (!(sc->gfp_mask & __GFP_FS))` in
>>> super_cache_scan()), I wonder if the daemon should simply use
>>> prctl(PR_SET_IO_FLUSHER) so that the user daemon won't be called
>>> into the fs reclaim context again.
>>
>> Hi Gao,
>>
>> We cannot rely on the daemon to set this unfortunately. This can tie
>> up reclaim and kswapd for the entire system so I think this
>> enforcement needs to be guaranteed and at the kernel level. For
>> example, there is the possibility of malicious servers, which we
>> cannot rely on to set FR_SET_IO_FLUSHER.
>
> Hi Joanne,
>
> Yes, currently I don't have a saner way in my mind but iput()
> in such nested context sounds a new entry (e.g. I thought
> kernel page fault path should have nothing tangled with
To clarify:
^ kernel file read page fault path tangled with this particular
inode in progress (I doesn't mean random inode reclaimation).
> evict() directly but I may be wrong.)
>
> In principle, typical the kernel filesystem holds a valid `file`
> during the entire buffered read (file)/mmap (vma->vm_file)
> submission path (and of course they won't upcall to userspace
> and then do random behavior in the userspace for I/O processing).
>
> So for the kernel filesystems I think the GFP_NOFS allocation
> isn't needed since `file.f_path` always takes a valid dentry
> ref during the submission so that such dentry/inode reclaim
> above is impossible IMO.)
>
> Thanks,
> Gao Xiang
>
>>
>> Thanks,
>> Joanne
>>
next prev parent reply other threads:[~2025-09-30 2:36 UTC|newest]
Thread overview: 12+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-09-25 22:44 [PATCH] fuse: fix readahead reclaim deadlock Joanne Koong
2025-09-26 6:51 ` Gao Xiang
2025-09-26 7:19 ` Gao Xiang
2025-09-29 17:25 ` Joanne Koong
2025-09-30 2:21 ` Gao Xiang
2025-09-30 2:35 ` Gao Xiang [this message]
2025-09-30 10:08 ` Miklos Szeredi
2025-09-30 18:47 ` Joanne Koong
2025-09-30 18:55 ` Miklos Szeredi
2025-10-01 0:18 ` Joanne Koong
2025-10-07 0:37 ` Joanne Koong
2025-09-26 9:01 ` Miklos Szeredi
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=e5b4985d-18c3-4609-b1f7-2425f161375d@linux.alibaba.com \
--to=hsiangkao@linux.alibaba.com \
--cc=joannelkoong@gmail.com \
--cc=kernel-team@meta.com \
--cc=linux-fsdevel@vger.kernel.org \
--cc=miklos@szeredi.hu \
--cc=osandov@fb.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.