From: Jeff Layton <jlayton@kernel.org>
To: Li Lingfeng <lilingfeng3@huawei.com>,
chuck.lever@oracle.com, neilb@suse.de, okorniev@redhat.com,
Dai.Ngo@oracle.com, tom@talpey.com, linux-nfs@vger.kernel.org,
linux-kernel@vger.kernel.org
Cc: yukuai1@huaweicloud.com, houtao1@huawei.com, yi.zhang@huawei.com,
yangerkun@huawei.com, lilingfeng@huaweicloud.com
Subject: Re: [PATCH] nfsd: put dl_stid if fail to queue dl_recall
Date: Thu, 13 Feb 2025 06:49:02 -0500 [thread overview]
Message-ID: <7df4de2bff617dc5c2bb482df6dbc5ef21ba0d01.camel@kernel.org> (raw)
In-Reply-To: <20250213072536.69986-1-lilingfeng3@huawei.com>
On Thu, 2025-02-13 at 15:25 +0800, Li Lingfeng wrote:
> Before calling nfsd4_run_cb to queue dl_recall to the callback_wq, we
> increment the reference count of dl_stid.
> We expect that after the corresponding work_struct is processed, the
> reference count of dl_stid will be decremented through the callback
> function nfsd4_cb_recall_release.
> However, if the call to nfsd4_run_cb fails, the incremented reference
> count of dl_stid will not be decremented correspondingly, leading to the
> following nfs4_stid leak:
> unreferenced object 0xffff88812067b578 (size 344):
> comm "nfsd", pid 2761, jiffies 4295044002 (age 5541.241s)
> hex dump (first 32 bytes):
> 01 00 00 00 6b 6b 6b 6b b8 02 c0 e2 81 88 ff ff ....kkkk........
> 00 6b 6b 6b 6b 6b 6b 6b 00 00 00 00 ad 4e ad de .kkkkkkk.....N..
> backtrace:
> kmem_cache_alloc+0x4b9/0x700
> nfsd4_process_open1+0x34/0x300
> nfsd4_open+0x2d1/0x9d0
> nfsd4_proc_compound+0x7a2/0xe30
> nfsd_dispatch+0x241/0x3e0
> svc_process_common+0x5d3/0xcc0
> svc_process+0x2a3/0x320
> nfsd+0x180/0x2e0
> kthread+0x199/0x1d0
> ret_from_fork+0x30/0x50
> ret_from_fork_asm+0x1b/0x30
> unreferenced object 0xffff8881499f4d28 (size 368):
> comm "nfsd", pid 2761, jiffies 4295044005 (age 5541.239s)
> hex dump (first 32 bytes):
> 01 00 00 00 00 00 00 00 30 4d 9f 49 81 88 ff ff ........0M.I....
> 30 4d 9f 49 81 88 ff ff 20 00 00 00 01 00 00 00 0M.I.... .......
> backtrace:
> kmem_cache_alloc+0x4b9/0x700
> nfs4_alloc_stid+0x29/0x210
> alloc_init_deleg+0x92/0x2e0
> nfs4_set_delegation+0x284/0xc00
> nfs4_open_delegation+0x216/0x3f0
> nfsd4_process_open2+0x2b3/0xee0
> nfsd4_open+0x770/0x9d0
> nfsd4_proc_compound+0x7a2/0xe30
> nfsd_dispatch+0x241/0x3e0
> svc_process_common+0x5d3/0xcc0
> svc_process+0x2a3/0x320
> nfsd+0x180/0x2e0
> kthread+0x199/0x1d0
> ret_from_fork+0x30/0x50
> ret_from_fork_asm+0x1b/0x30
> Fix it by checking the result of nfsd4_run_cb and call nfs4_put_stid if
> fail to queue dl_recall.
>
> Fixes: 1da177e4c3f4 ("Linux-2.6.12-rc2")
> Signed-off-by: Li Lingfeng <lilingfeng3@huawei.com>
> ---
> fs/nfsd/nfs4state.c | 6 +++++-
> 1 file changed, 5 insertions(+), 1 deletion(-)
>
> diff --git a/fs/nfsd/nfs4state.c b/fs/nfsd/nfs4state.c
> index 153eeea2c7c9..0ccb87be47b7 100644
> --- a/fs/nfsd/nfs4state.c
> +++ b/fs/nfsd/nfs4state.c
> @@ -5414,6 +5414,7 @@ static const struct nfsd4_callback_ops nfsd4_cb_recall_ops = {
>
> static void nfsd_break_one_deleg(struct nfs4_delegation *dp)
> {
> + bool queued;
> /*
> * We're assuming the state code never drops its reference
> * without first removing the lease. Since we're in this lease
> @@ -5422,7 +5423,10 @@ static void nfsd_break_one_deleg(struct nfs4_delegation *dp)
> * we know it's safe to take a reference.
> */
> refcount_inc(&dp->dl_stid.sc_count);
> - WARN_ON_ONCE(!nfsd4_run_cb(&dp->dl_recall));
> + queued = nfsd4_run_cb(&dp->dl_recall);
> + WARN_ON_ONCE(!queued);
> + if (!queued)
> + nfs4_put_stid(&dp->dl_stid);
> }
>
> /* Called from break_lease() with flc_lock held. */
Have you actually seen the WARN_ON_ONCE() pop under normal usage, or
was the problem you reproduced done via fault injection?
Unfortunately, this won't work. nfsd_break_one_deleg() is called from
the ->break_lease callback with the flc->flc_lock held, and
nfs4_put_stid can sleep in the sc_free callbacks.
--
Jeff Layton <jlayton@kernel.org>
next prev parent reply other threads:[~2025-02-13 11:49 UTC|newest]
Thread overview: 5+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-02-13 7:25 [PATCH] nfsd: put dl_stid if fail to queue dl_recall Li Lingfeng
2025-02-13 11:49 ` Jeff Layton [this message]
2025-02-13 12:28 ` Li Lingfeng
2025-02-13 12:34 ` Jeff Layton
2025-02-13 14:21 ` Li Lingfeng
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=7df4de2bff617dc5c2bb482df6dbc5ef21ba0d01.camel@kernel.org \
--to=jlayton@kernel.org \
--cc=Dai.Ngo@oracle.com \
--cc=chuck.lever@oracle.com \
--cc=houtao1@huawei.com \
--cc=lilingfeng3@huawei.com \
--cc=lilingfeng@huaweicloud.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-nfs@vger.kernel.org \
--cc=neilb@suse.de \
--cc=okorniev@redhat.com \
--cc=tom@talpey.com \
--cc=yangerkun@huawei.com \
--cc=yi.zhang@huawei.com \
--cc=yukuai1@huaweicloud.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox