Linux CXL
 help / color / mirror / Atom feed
From: Dave Jiang <dave.jiang@intel.com>
To: "Lucero Palau, Alejandro" <alejandro.lucero-palau@amd.com>,
	alucerop@amd.com, linux-cxl@vger.kernel.org,
	netdev@vger.kernel.org
Cc: davem@davemloft.net, kuba@kernel.org, pabeni@redhat.com,
	edumazet@google.com, ecree.xilinx@gmail.com, icheng@nvidia.com,
	rafael@kernel.org
Subject: Re: [PATCH v2 2/4] cxl/region: Add region reference in memdev attach
Date: Thu, 8 Oct 2026 14:05:08 -0700	[thread overview]
Message-ID: <3f40d953-2e91-4492-b100-c851fdb143b5@intel.com> (raw)
In-Reply-To: <ba7731fe-f916-4e51-977b-95a0197bf52d@amd.com>



On 10/8/26 11:07 AM, Lucero Palau, Alejandro wrote:
> 
> On 08/10/2026 17:18, Dave Jiang wrote:
>>
>> On 10/8/26 6:50 AM, Lucero Palau, Alejandro wrote:
>>> On 02/10/2026 16:52, Dave Jiang wrote:
>>>> On 10/1/26 9:41 PM, Lucero Palau, Alejandro wrote:
>>>>> On 01/10/2026 22:38, Dave Jiang wrote:
>>>>>> On 10/1/26 6:20 AM, alucerop@amd.com wrote:
>>>>>>> From: Alejandro Lucero <alucerop@amd.com>
>>>>>>>
>>>>>>> Use a new field in cxl_attach_region struct for easily link it with the
>>>>>>> region the memdev is attached to.
>>>>>>>
>>>>>>> This facilitates device links creation where such a region is the supplier
>>>>>>> with non-PF0 physical functions wanting to use the CXL region being the
>>>>>>> consumers.
>>>>>>>
>>>>>>> Signed-off-by: Alejandro Lucero <alucerop@amd.com>
>>>>>>> ---
>>>>>>>     drivers/cxl/core/region.c | 1 +
>>>>>>>     drivers/cxl/cxlmem.h      | 2 ++
>>>>>>>     2 files changed, 3 insertions(+)
>>>>>>>
>>>>>>> diff --git a/drivers/cxl/core/region.c b/drivers/cxl/core/region.c
>>>>>>> index 27e63e6dab7c..78ca7ebc3e55 100644
>>>>>>> --- a/drivers/cxl/core/region.c
>>>>>>> +++ b/drivers/cxl/core/region.c
>>>>>>> @@ -4132,6 +4132,7 @@ int cxl_memdev_attach_region(struct cxl_memdev *cxlmd)
>>>>>>>         if (rc)
>>>>>>>             return rc;
>>>>>>>     +    attach->cxlr = cxlr;
>>>>>>>         attach->hpa_range = (struct range) {
>>>>>>>             .start = cxlr->params.res->start,
>>>>>>>             .end = cxlr->params.res->end,
>>>>>>> diff --git a/drivers/cxl/cxlmem.h b/drivers/cxl/cxlmem.h
>>>>>>> index c401e3a1af06..c598561b8e5f 100644
>>>>>>> --- a/drivers/cxl/cxlmem.h
>>>>>>> +++ b/drivers/cxl/cxlmem.h
>>>>>>> @@ -104,6 +104,7 @@ struct cxl_memdev_attach {
>>>>>>>     /**
>>>>>>>      * struct cxl_attach_region - coordinate mapping a region at memdev registration
>>>>>>>      * @attach: common core attachment descriptor
>>>>>>> + * @cxlr: cxl region the memdev is attached to.
>>>>>>>      * @hpa_range: physical address range of the region
>>>>>>>      *
>>>>>>>      * For the common simple case of a CXL device with private (non-general purpose
>>>>>>> @@ -112,6 +113,7 @@ struct cxl_memdev_attach {
>>>>>>>      */
>>>>>>>     struct cxl_attach_region {
>>>>>>>         struct cxl_memdev_attach attach;
>>>>>>> +    struct cxl_region *cxlr;
>>>>>>>         struct range hpa_range;
>>>>>>>     };
>>>>>>>     
>>>>>> attach->cxlr is never cleared when the region goes away. Unbinding the endpoint port runs endpoint_unregister_region(), which unregisters the region and drops its reference. PF0 stays bound until the detach work runs.
>>>>>>
>>>>>> A non-PF0 probe in that window still finds the memdev and calls device_link_add() on the freed region.
>>>>> I do not think so. This version, see next patch, relies on locking the supplier, PF0, before getting the memdev and potentially using the attach region. Only on PF0 release can such a memdev and region disappear, so I think this is enough.
>>>>>
>>> Hi Dave,
>>>
>>>
>>>> The memdev, yes. It's devm on PF0. The region, no. Unbinding the endpoint port, or cxl_mem from the memdev, runs endpoint_unregister_region() under the endpoint's lock, not PF0's. PF0's release is only queued (schedule_detach() -> detach_memdev()). Until that work runs, PF0 is bound, the memdev is found, hpa_range is valid, and attach->cxlr points at a freed region. Root teardown (kill_regions()) and delete_region also unregister the region without touching PF0.
>>>
>>> You are partially right. The problem is not that async detach_memdev() but the fact the region can be removed without the PF0 being aware ...
>>>
>>>
>>> I was relying on the memdev being unregister first then the related device released later on, with the important part being the memdev being unregister making it unavailable to the other PFs because it can not be found in the cxl bus. I'm pretty sure about that scenario but I was not counting on sysfs unbinds for the endpoint device ...
>>>
>>>
>>> I need to think further about this because I think it is wrong to remove the region in this case without PF0 or the memdev owner and original region consumer still unaware of it. One thing I tried to discuss with Dan was why we have the unbinding options for cxl_mem and cxl_port drivers. What is the point? I know this is standard linux device model, but the unbinding could do nothing if we decide so. Couldn't we? Otherwise, If there is a good reason for having this functionality where the unwinding can happen through different starting points, we should document it.
>>>
>>>
>>> What I'm going to try is to link the region removal to Type2 driver removal as well, not necessarily with device links but through devm_action_or_reset links as we are already doing for other cases. And I think to document the different objects/devices involved and its lifetime depending on the current unwinding supported would be good to have, so I will work on that as well.
>>>
>>>
>> So there are multiple paths a region can go away without PF0 knowing and sysfs is only one of the triggers.
>> 1. endpoint port teardown -> endpoint_unregister_region()
>> 2. root decoder teardown -> kill_regions()
>> 3. userspace delete region
>> 4. some other places calling unregister_region().
>>
>> Using suppress_bind_attrs only hides the sysfs files. There's also device hot-remove and module unload in addition to root decoder action and user action. So doing that is definitely the wrong way to go.
>>
>> The path of PF0 going first is fine. The memdev goes with it. The bug path is region going first. Something needs to deal with the stale attach->cxlr. You need a region-side teardown clearing attach->cxlr under the cxl_rwsem.region before unregistering. And readers need to check under the same lock. The endpoint_detach_attach_region() proposal does close that hole.
> 
> 
> I tried to discuss this with Dan unsuccessfully, so I hope I can make my point clear: having so many ways of cxl objects/devices being destroyed is, IMO, wrong. Moreover, some user actions on things created by a Type2 driver should not be so easy accessed, and definitely, having a region unregistered and released with its main consumer and owner completely unaware, should not be happening.
> 
> 
> Dan and I addressed some concerns with "these options" but it is worse after realising now port and region can also suffer from unbinding actions. We contemplated memdev unbinding and that is supported, and acpi module removal as well (all the unwinding is hopefully right for sfc driver removal), but the fact is, current Type2 support is unsound. It is likely good enough with current usage expectations but something to improve/fix.

Removing the acpi module will also cause the issue I pointed out. So that isn't safe either. The only path that's good right now is sfc driver removal or the device going away.

> 
> 
> All this user space potential actions were implemented mainly for testing (I guess you know this). I did ask Dan about it, and I was expecting use cases where HDM decoders and regions are dynamically created, which makes a lot of sense to me, but the fact is all is relying on firmware/BIOS configuration. Richard is working on adding this functionality for Type2 and pmems, and Jonathan considers it theoretically useful as well, but the way is going to be handled requires, IMO, further thinking and maybe a change before someone starts using it (does anyone know about users now?).
> 
> 
> As a summary, if we allow user space actions (at least for Type2) they need to be consistent and somehow protected.
> 
> 
> Finally, you did not answer my question: what is the point user space removing and endpoint port handled by a Type2 driver? What about the cxl region? Maybe I am missing a necessity I can not see here, so please, help me to understand this if that is the case.

Shouldn't does not mean does not exist. Sure I can agree with you that under normal operations, certain things a sane user should avoid doing for type2. But it is possible currently and those issues can be triggered. However you feel about the current CXL architecture, here we are with where it is. You can either consider the smaller changes I suggested to keep the attach->cxlr sane (or with some other means) and make what you need working now with raised the issue addressed, and come back with hashing out the larger grievances later, or keep beating this horse.... "It's silly for users to do that and therefore the issue can be ignored" is not a good enough reason for me look the other way and merge the code.

> 
> 
> Thanks,
> 
> Alejandro.
> 
> 
>>
>> DJ
>>
>>> Thank you,
>>>
>>> Alejandro
>>>
>>>
>>>> I attached an LLM generated test kernel module you can use to reproduce the KASAN complaint. Commit log provides instructions.
>>>> BUG: KASAN: slab-use-after-free in device_link_add+0x521/0xa80
>>>>    cxl_get_range_and_link+0xa9/0x120 [cxl_core]
>>>>
>>>> DJ
>>>>
>>>>> I added comments in the exported function, cxl_get_range_and_link() which does the locking before calling the internal function __cxl_get_range_and_link() which looks for the memdev and the attach region. In fact, it should not be possible to obtain the PF0 memdev reference and the attach region not there yet, but the code is still checking that possibility as a sanity check.
>>>>>
>>>>>
>>>>> Thank you,
>>>>>
>>>>> Alejandro.
>>>>>
>>>>>
>>>>>> How about something like this?
>>>>>>
>>>>>> diff --git a/drivers/cxl/core/region.c b/drivers/cxl/core/region.c
>>>>>> index 27e63e6dab7c..38ca73f12b84 100644
>>>>>> --- a/drivers/cxl/core/region.c
>>>>>> +++ b/drivers/cxl/core/region.c
>>>>>> @@ -4076,6 +4076,23 @@ static int first_mapped_decoder(struct device *dev, const void *data)
>>>>>>         return 0;
>>>>>>     }
>>>>>>     +/*
>>>>>> + * Invalidate @attach before the region goes away so that
>>>>>> + * cxl_get_range_and_link() can not pick up a stale region.
>>>>>> + */
>>>>>> +static void endpoint_detach_attach_region(void *_attach)
>>>>>> +{
>>>>>> +    struct cxl_attach_region *attach = _attach;
>>>>>> +    struct cxl_region *cxlr;
>>>>>> +
>>>>>> +    scoped_guard(rwsem_write, &cxl_rwsem.region) {
>>>>>> +        cxlr = attach->cxlr;
>>>>>> +        WRITE_ONCE(attach->cxlr, NULL);
>>>>>> +        attach->hpa_range = DEFINE_RANGE(0, -1);
>>>>>> +    }
>>>>>> +    endpoint_unregister_region(cxlr);
>>>>>> +}
>>>>>> +
>>>>>>     /*
>>>>>>      * Runs in cxl_mem_probe context after successful endpoint probe, assumes the
>>>>>>      * simple case of single mapped decoder per memdev.
>>>>>> @@ -4127,15 +4144,23 @@ int cxl_memdev_attach_region(struct cxl_memdev *cxlmd)
>>>>>>           /* Only teardown regions that pass validation, ignore the rest */
>>>>>>         get_device(&cxlr->dev);
>>>>>> -    rc = devm_add_action_or_reset(&endpoint->dev,
>>>>>> -                      endpoint_unregister_region, cxlr);
>>>>>> -    if (rc)
>>>>>> +    /*
>>>>>> +     * Not devm_add_action_or_reset(): the reset path would take
>>>>>> +     * cxl_rwsem.region for write while it is held for read here. The
>>>>>> +     * endpoint lock keeps the action from running before @attach is set.
>>>>>> +     */
>>>>>> +    rc = devm_add_action(&endpoint->dev, endpoint_detach_attach_region,
>>>>>> +                 attach);
>>>>>> +    if (rc) {
>>>>>> +        put_device(&cxlr->dev);
>>>>>>             return rc;
>>>>>> +    }
>>>>>>           attach->hpa_range = (struct range) {
>>>>>>             .start = cxlr->params.res->start,
>>>>>>             .end = cxlr->params.res->end,
>>>>>>         };
>>>>>> +    WRITE_ONCE(attach->cxlr, cxlr);
>>>>>>         return 0;
>>>>>>     }
>>>>>>     EXPORT_SYMBOL_FOR_MODULES(cxl_memdev_attach_region, "cxl_mem");
>>>>>> diff --git a/drivers/cxl/cxlmem.h b/drivers/cxl/cxlmem.h
>>>>>> index c401e3a1af06..7cd3a69cd5f5 100644
>>>>>> --- a/drivers/cxl/cxlmem.h
>>>>>> +++ b/drivers/cxl/cxlmem.h
>>>>>> @@ -104,6 +104,8 @@ struct cxl_memdev_attach {
>>>>>>     /**
>>>>>>      * struct cxl_attach_region - coordinate mapping a region at memdev registration
>>>>>>      * @attach: common core attachment descriptor
>>>>>> + * @cxlr: cxl region the memdev is attached to, cleared under cxl_rwsem.region
>>>>>> + *    before the region is unregistered
>>>>>>      * @hpa_range: physical address range of the region
>>>>>>      *
>>>>>>      * For the common simple case of a CXL device with private (non-general purpose
>>>>>> @@ -112,6 +114,7 @@ struct cxl_memdev_attach {
>>>>>>      */
>>>>>>     struct cxl_attach_region {
>>>>>>         struct cxl_memdev_attach attach;
>>>>>> +    struct cxl_region *cxlr;
>>>>>>         struct range hpa_range;
>>>>>>     };
>>>>>>     


  reply	other threads:[~2026-10-08 21:05 UTC|newest]

Thread overview: 27+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-01 13:20 [PATCH v2 0/4] Type2 multipf support alucerop
2026-10-01 13:20 ` [PATCH v2 1/4] driver core: Check for supplier requiring PM at link creation alucerop
2026-10-01 20:31   ` Dave Jiang
2026-10-02  4:32     ` Lucero Palau, Alejandro
2026-10-02 15:31       ` Dave Jiang
2026-10-02 12:02   ` sashiko-bot
2026-10-01 13:20 ` [PATCH v2 2/4] cxl/region: Add region reference in memdev attach alucerop
2026-10-01 21:38   ` Dave Jiang
2026-10-02  4:41     ` Lucero Palau, Alejandro
2026-10-02 15:52       ` Dave Jiang
2026-10-08 13:50         ` Lucero Palau, Alejandro
2026-10-08 16:18           ` Dave Jiang
2026-10-08 18:07             ` Lucero Palau, Alejandro
2026-10-08 21:05               ` Dave Jiang [this message]
2026-10-09  6:58                 ` Lucero Palau, Alejandro
2026-10-09 16:57                   ` Dave Jiang
2026-10-02 12:02   ` sashiko-bot
2026-10-01 13:20 ` [PATCH v2 3/4] cxl/memdev: Add support for multi PF devices alucerop
2026-10-01 22:11   ` Dave Jiang
2026-10-01 22:41     ` Dave Jiang
2026-10-02  4:50     ` Lucero Palau, Alejandro
2026-10-02 15:55       ` Dave Jiang
2026-10-02 12:02   ` sashiko-bot
2026-10-01 13:20 ` [PATCH v2 4/4] sfc: add multipf support alucerop
2026-10-01 22:32   ` Dave Jiang
2026-10-02  5:33     ` Lucero Palau, Alejandro
2026-10-02 12:02   ` sashiko-bot

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=3f40d953-2e91-4492-b100-c851fdb143b5@intel.com \
    --to=dave.jiang@intel.com \
    --cc=alejandro.lucero-palau@amd.com \
    --cc=alucerop@amd.com \
    --cc=davem@davemloft.net \
    --cc=ecree.xilinx@gmail.com \
    --cc=edumazet@google.com \
    --cc=icheng@nvidia.com \
    --cc=kuba@kernel.org \
    --cc=linux-cxl@vger.kernel.org \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=rafael@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox