From: Cheng Xu <chengyou@linux.alibaba.com>
To: Leon Romanovsky <leon@kernel.org>
Cc: jgg@ziepe.ca, linux-rdma@vger.kernel.org, KaiShen@linux.alibaba.com
Subject: Re: [PATCH for-next v2 1/4] RDMA/erdma: Support non-contiguous kernel QP buffers
Date: Wed, 23 Sep 2026 18:17:38 +0800 [thread overview]
Message-ID: <f1b0112c-962b-ede6-fcc8-5d6dc4ecc6a7@linux.alibaba.com> (raw)
In-Reply-To: <20260918124759.GX13683@unreal>
On 9/18/26 8:47 PM, Leon Romanovsky wrote:
> On Fri, Sep 11, 2026 at 10:41:52AM +0800, Cheng Xu wrote:
>>
>>
>> On 9/10/26 11:33 PM, Leon Romanovsky wrote:
>>> On Wed, Sep 09, 2026 at 02:21:37PM +0800, Cheng Xu wrote:
>>>>
>>>>
>>>> On 9/3/26 8:38 PM, Cheng Xu wrote:
>>>>>
>>>>>
>>>>> On 9/3/26 5:10 PM, Leon Romanovsky wrote:
>>>>>> On Thu, Aug 27, 2026 at 04:25:20PM +0800, Cheng Xu wrote:
>>>>>>> A single coherent allocation for kernel QP queues can fail for large
>>>>>>> queues when memory is fragmented.
>>>>>>>
>>>>>>> Allocate page-sized coherent buffers and describe them with the existing
>>>>>>> MTT. Keep the userspace QP path unchanged.
>>>>>>>
>>>>>>> Signed-off-by: Cheng Xu <chengyou@linux.alibaba.com>
>>>>>>> ---
>>>>>>> drivers/infiniband/hw/erdma/erdma_cq.c | 4 +-
>>>>>>> drivers/infiniband/hw/erdma/erdma_qp.c | 38 +++--
>>>>>>> drivers/infiniband/hw/erdma/erdma_verbs.c | 195 +++++++++++++---------
>>>>>>> drivers/infiniband/hw/erdma/erdma_verbs.h | 44 ++++-
>>>>>>> 4 files changed, 179 insertions(+), 102 deletions(-)
>>>>>>
>>>>>> <...>
>>>>>>
>>>>
>>>> <...>
>>>>
>>>>>>
>>>>>>> +struct erdma_buf_list {
>>>>>>> + void *buf;
>>>>>>> + dma_addr_t dma_addr;
>>>>>>> +};
>>>>>>
>>>>>> This struct is very similar to scatter-gather list, why don't you use it
>>>>>> directly?
>>>>>
>>>>> Good idea. I will use struct scatterlist in the next revision.
>>>>
>>>> Hi Leon,
>>>>
>>>> I switched to scatterlist in v3, but Sashiko pointed out an issue with
>>>> using sg_set_buf() and sg_virt() on dma_alloc_coherent() memory [1].
>>>>
>>>> To handle this correctly, the driver would still need to retain the
>>>> original CPU addresses returned by dma_alloc_coherent(). Using scatterlist
>>>> does not simplify this implementation: we still need separate storage for
>>>> the CPU addresses, and the only scatterlist field we actually need is the
>>>> DMA address. A small structure holding both addresses would therefore be
>>>> simpler.
>>>
>>> Sashiko thinks that you are creating SG list to feed it to dma_map_sg()
>>> later which is not. You are using SG as simple database and you will get
>>> iterators for free.
>>>
>>> cpu_address = sg_page()
>>> dma_address = sg_dma_address()
>>>
>>
>> Hi Leon,
>>
>> Maybe I misunderstood something. My understanding of Sashiko's concern is
>> that the CPU address returned by dma_alloc_coherent() is not guaranteed to
>> be valid for virt_to_page(), as used by sg_set_buf(). Consequently, the
>> struct page later returned by sg_page() may be invalid.
>>
>> Although this works on the x86 platforms we tested, it may not be portable
>> to architectures where coherent memory is remapped outside the linear
>> mapping. From this perspective, would a small structure that explicitly
>> stores the CPU and DMA addresses, similar to those used by HNS and mlx5, be
>> more appropriate than scatterlist here?
>
> Yes, let's use them directly, without SG. Sorry for the confusion in my
> previous comment.
>
Thanks for confirming. v4 keeps the small private structure as discussed.
Thanks,
Cheng Xu
> Thanks
>
>>
>> I collected some information related to this below [1][2][3].
>>
>> [1] include/linux/scatterlist.h:
>>
>> #ifdef CONFIG_DEBUG_SG
>> BUG_ON(!virt_addr_valid(buf));
>> #endif
>> sg_set_page(sg, virt_to_page(buf), buflen, offset_in_page(buf));
>>
>> [2] kernel/dma/mapping.c:
>>
>> /*
>> * The whole dma_get_sgtable() idea is fundamentally unsafe ...
>> * 1. Not all memory allocated via the coherent DMA APIs is backed by
>> * a struct page
>> */
>>
>> [3] drivers/infiniband/hw/mthca/mthca_memfree.c:
>>
>> /* We use sg_set_buf for coherent allocs, which assumes low memory */
>
> I wouldn't hold my breath about the validity of that comment. It was written
> a long time ago.
>
>>
>> Thanks,
>> Cheng Xu
>>
>>
>>> Thanks
>>>
>>
next prev parent reply other threads:[~2026-09-23 10:17 UTC|newest]
Thread overview: 14+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-27 8:25 [PATCH for-next v2 0/4] RDMA/erdma: Support non-contiguous kernel queue buffers Cheng Xu
2026-08-27 8:25 ` [PATCH for-next v2 1/4] RDMA/erdma: Support non-contiguous kernel QP buffers Cheng Xu
2026-09-03 9:10 ` Leon Romanovsky
2026-09-03 12:38 ` Cheng Xu
2026-09-09 6:21 ` Cheng Xu
2026-09-10 15:33 ` Leon Romanovsky
2026-09-11 2:41 ` Cheng Xu
2026-09-18 12:47 ` Leon Romanovsky
2026-09-23 10:17 ` Cheng Xu [this message]
2026-08-27 8:25 ` [PATCH for-next v2 2/4] RDMA/erdma: Support non-contiguous kernel CQ buffers Cheng Xu
2026-08-27 8:25 ` [PATCH for-next v2 3/4] RDMA/erdma: Unify userspace and kernel queue buffer management Cheng Xu
2026-09-03 9:12 ` Leon Romanovsky
2026-09-03 12:40 ` Cheng Xu
2026-08-27 8:25 ` [PATCH for-next v2 4/4] RDMA/erdma: Move kernel QP helpers after memory helpers Cheng Xu
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=f1b0112c-962b-ede6-fcc8-5d6dc4ecc6a7@linux.alibaba.com \
--to=chengyou@linux.alibaba.com \
--cc=KaiShen@linux.alibaba.com \
--cc=jgg@ziepe.ca \
--cc=leon@kernel.org \
--cc=linux-rdma@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox