From: Zhu Yanjun <yanjun.zhu@linux.dev>
To: Nicolas Morey <nmorey@suse.com>,
Zhu Yanjun <zyjzyj2000@gmail.com>, Jason Gunthorpe <jgg@ziepe.ca>,
Leon Romanovsky <leon@kernel.org>,
Tristan Madani <tristmd@gmail.com>,
"open list:SOFT-ROCE DRIVER (rxe)" <linux-rdma@vger.kernel.org>,
open list <linux-kernel@vger.kernel.org>,
"yanjun.zhu@linux.dev" <yanjun.zhu@linux.dev>
Subject: Re: [PATCH] RDMA/rxe: Use validated num_sge in local buffer
Date: Wed, 9 Sep 2026 07:15:22 -0700 [thread overview]
Message-ID: <594ad636-369d-4f0c-8c42-c1b3fba51f5b@linux.dev> (raw)
In-Reply-To: <a8d64435-2fd8-4e2b-9b48-aea496630e07@suse.com>
在 2026/9/8 0:35, Nicolas Morey 写道:
> On 2026-09-08 05:07, Zhu Yanjun wrote:
>> 在 2026/9/7 14:44, Nicolas Morey 写道:
>>> On 2026-09-07 23:20, Zhu Yanjun wrote:
>>>> 在 2026/9/7 9:15, Nicolas Morey 写道:
>>>>> For both SRQ and non-SRQ receive paths, the WQE is copied into a local
>>>>> buffer to provide a kernel-owned, validated copy. While calculating the
>>>>> memcpy size from the validated num_sge prevents overflow during the
>>>>> copy, memcpy() itself still copies num_sge from shared memory.
>>>>>
>>>>> A concurrent userspace modification before or during memcpy() leaves
>>>>> an unvalidated num_sge in the local buffer, leading to potential
>>>>> out-of-bounds reads in rxe_resp_check_length() and copy_data().
>>>>>
>>>>
>>>> Hi Nicolas,
>>>>
>>>> Thanks for the patch. The logic makes total sense to prevent the TOCTOU race condition after memcpy.
>>>>
>>>> Just out of curiosity, do you happen to have a reproducer or a POC script that demonstrates this race in practice?
>>>>
>>>> It would be great to know if this can be reliably reproduced or integrated into testing setups (like rdma-core tests, or tools/testing/selftests/rdma) to catch similar double-read issues in the future.
>>>>
>>>
>>> No reproducer or PoC sadly. I haven't tried to make one though.
>>> This got caught by one of our AI tools when checking the backport of CVE-2026-74377.
>>
>> Hi, Nicolas
>>
>> Thanks for sharing this. Could you clarify whether your AI tool detected this purely through static analysis (and provided a suggested fix), or if it actually triggered a runtime bug with a backtrace/calltrace?
>>
>> If you have a calltrace, please share it—that would help a lot in understanding the issue. Also, if this was flagged and fixed by a specific AI tool, adding an Assisted-by: tag (or referencing the tool in the commit description) would be appropriate.
>>
>> Thanks,
>>
>> Zhu Yanjun
>>
>
> Hi Zhu,
>
> The tool is a LLM static analyser that simply detected that d6ab440240a0 ("RDMA/rxe: Copy WQE to local buffer in non-SRQ receive path") was a valid patch but did not fix the whole issue. It did not suggest any breaker PoC nor any fix.
> All the rest is human (me) made.
> Does that warrant an Assisted-by tag ?
> It seems sashiko picked the same issue for 22b8fbded65b ("RDMA/rxe: Fix TOCTOU heap overflow in get_srq_wqe")
> https://sashiko.dev/#/patchset/20260518215040.1598586-1-tristan%40talencesecurity.com
>
>> If a concurrent userspace thread modifies wqe->dma.num_sge after the
>> bounds check but before the memcpy completes, doesn't this copy the
>> unvalidated value into the kernel heap?
>
> By tweaking the reproducer from Tristan, I have been able to reproduce it once out of sheer luck I think:
Thanks a lot. Please add the following call trace into the commit log
and send out a new commit.
Except the above, I am fine with this commit.
Reviewed-by: Zhu Yanjun <yanjun.zhu@linux.dev>
Zhu Yanjun
> ==================================================================
> BUG: KASAN: slab-out-of-bounds in rxe_receiver+0x8109/0x9ec0 [rdma_rxe]
> Read of size 4 at addr ffff88812c4867f8 by task kworker/u9:6/361
>
> CPU: 0 UID: 0 PID: 361 Comm: kworker/u9:6 Tainted: G E 7.3.0-rc2-00006-g28924df2a08f #3 PREEMPT(full) 318a88bba3045f81bad31a7694727561b4bd4965
> Tainted: [E]=UNSIGNED_MODULE
> Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS rel-1.16.3-2-gc13ff2cd-prebuilt.qemu.org 04/01/2014
> Workqueue: rxe_wq do_work [rdma_rxe]
> Call Trace:
> <TASK>
> dump_stack_lvl+0x4b/0x70
> print_report+0x153/0x4b5
> ? __pfx__raw_spin_lock_irqsave+0x10/0x10
> ? stack_trace_save+0x93/0xd0
> kasan_report+0xbc/0xf0
> ? rxe_receiver+0x8109/0x9ec0 [rdma_rxe b8d153a5fa9911e345612beff278ac3241bac327]
> ? rxe_receiver+0x8109/0x9ec0 [rdma_rxe b8d153a5fa9911e345612beff278ac3241bac327]
> rxe_receiver+0x8109/0x9ec0 [rdma_rxe b8d153a5fa9911e345612beff278ac3241bac327]
> ? pick_task_fair+0x12e/0x1b50
> ? __pfx_rxe_receiver+0x10/0x10 [rdma_rxe b8d153a5fa9911e345612beff278ac3241bac327]
> ? __pfx__raw_spin_lock_irqsave+0x10/0x10
> ? hrtimer_start_range_ns+0xe7/0x320
> ? ktime_get+0xe3/0x170
> ? _raw_spin_unlock+0xe/0x30
> ? _raw_spin_lock_irqsave+0x8a/0xf0
> ? __pfx__raw_spin_lock_irqsave+0x10/0x10
> ? __pfx_rxe_receiver+0x10/0x10 [rdma_rxe b8d153a5fa9911e345612beff278ac3241bac327]
> do_work+0x149/0x610 [rdma_rxe b8d153a5fa9911e345612beff278ac3241bac327]
> process_one_work+0x726/0x10a0
> ? __pfx___schedule+0x10/0x10
> ? __pfx_process_one_work+0x10/0x10
> ? _raw_spin_lock_irq+0x85/0xe0
> ? __pfx__raw_spin_lock_irq+0x10/0x10
> worker_thread+0x500/0xd70
> ? __kthread_parkme+0x8d/0x170
> ? __pfx_worker_thread+0x10/0x10
> ? __pfx_worker_thread+0x10/0x10
> kthread+0x329/0x410
> ? recalc_sigpending+0x15c/0x200
> ? __pfx_kthread+0x10/0x10
> ret_from_fork+0x4cf/0x760
> ? __pfx_ret_from_fork+0x10/0x10
> ? __switch_to+0x575/0x11c0
> ? __pfx_kthread+0x10/0x10
> ret_from_fork_asm+0x1a/0x30
> </TASK>
>
> Allocated by task 1698:
> kasan_save_stack+0x20/0x40
> kasan_save_track+0x14/0x30
> __kasan_kmalloc+0x9a/0xb0
> __kmalloc_noprof+0x209/0x560
> create_qp.part.0+0x760/0x9d0 [ib_core]
> ib_create_qp_user+0xa2/0x500 [ib_core]
> ib_uverbs_handler_UVERBS_METHOD_QP_CREATE+0x99b/0x12af [ib_uverbs]
> ib_uverbs_cmd_verbs+0x214b/0x32d0 [ib_uverbs]
> ib_uverbs_ioctl+0x131/0x200 [ib_uverbs]
> __x64_sys_ioctl+0x13c/0x1c0
> do_syscall_64+0xba/0x530
> entry_SYSCALL_64_after_hwframe+0x76/0x7e
>
> Last potentially related work creation:
> kasan_save_stack+0x20/0x40
> kasan_record_aux_stack+0xb0/0xc0
> __queue_work+0x8c5/0x11f0
> queue_work_on+0x60/0x70
> rxe_sched_task+0x1d2/0x250 [rdma_rxe]
> rxe_rcv+0x847/0x1830 [rdma_rxe]
> rxe_xmit_packet+0x408/0x950 [rdma_rxe]
> rxe_requester+0x1878/0x5460 [rdma_rxe]
> rxe_sender+0x17/0x40 [rdma_rxe]
> do_work+0x149/0x610 [rdma_rxe]
> process_one_work+0x726/0x10a0
> worker_thread+0x500/0xd70
> kthread+0x329/0x410
> ret_from_fork+0x4cf/0x760
> ret_from_fork_asm+0x1a/0x30
>
> The buggy address belongs to the object at ffff88812c486000
> which belongs to the cache kmalloc-part-13-2k of size 2048
> The buggy address is located 0 bytes to the right of
> allocated 2040-byte region [ffff88812c486000, ffff88812c4867f8)
>
> The buggy address belongs to the physical page:
> page: refcount:0 mapcount:0 mapping:0000000000000000 index:0xffff88812c480000 pfn:0x12c480
> head: order:3 mapcount:0 entire_mapcount:0 nr_pages_mapped:0 pincount:0
> flags: 0x17ffffc0000240(workingset|head|node=0|zone=2|lastcpupid=0x1fffff)
> page_type: f5(slab)
> raw: 0017ffffc0000240 ffff888100059140 ffffea0004499010 ffffea00049a7a10
> raw: ffff88812c480000 0000000200080007 00000000f5000000 0000000000000000
> head: 0017ffffc0000240 ffff888100059140 ffffea0004499010 ffffea00049a7a10
> head: ffff88812c480000 0000000200080007 00000000f5000000 0000000000000000
> head: 0017ffffc0000003 fffffffffffffe01 00000000ffffffff 00000000ffffffff
> head: ffffffffffffffff 0000000000000000 00000000ffffffff 0000000000000008
> page dumped because: kasan: bad access detected
>
> Memory state around the buggy address:
> ffff88812c486680: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
> ffff88812c486700: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
>> ffff88812c486780: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 fc
> ^
> ffff88812c486800: fc fc fc fc fc fc fc fc fc fc fc fc fc fc fc fc
> ffff88812c486880: fc fc fc fc fc fc fc fc fc fc fc fc fc fc fc fc
> ==================================================================
prev parent reply other threads:[~2026-09-09 14:15 UTC|newest]
Thread overview: 6+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-07 16:15 [PATCH] RDMA/rxe: Use validated num_sge in local buffer Nicolas Morey
2026-09-07 21:20 ` Zhu Yanjun
2026-09-07 21:44 ` Nicolas Morey
2026-09-08 3:07 ` Zhu Yanjun
2026-09-08 7:35 ` Nicolas Morey
2026-09-09 14:15 ` Zhu Yanjun [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=594ad636-369d-4f0c-8c42-c1b3fba51f5b@linux.dev \
--to=yanjun.zhu@linux.dev \
--cc=jgg@ziepe.ca \
--cc=leon@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-rdma@vger.kernel.org \
--cc=nmorey@suse.com \
--cc=tristmd@gmail.com \
--cc=zyjzyj2000@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.