From: Andrey Grodzovsky <Andrey.Grodzovsky@amd.com>
To: Daniel Vetter <daniel@ffwll.ch>
Cc: robh@kernel.org, gregkh@linuxfoundation.org,
ckoenig.leichtzumerken@gmail.com,
dri-devel@lists.freedesktop.org, eric@anholt.net,
ppaalanen@gmail.com, amd-gfx@lists.freedesktop.org,
daniel.vetter@ffwll.ch, Alexander.Deucher@amd.com,
yuq825@gmail.com, Harry.Wentland@amd.com, l.stach@pengutronix.de
Subject: Re: [PATCH v3 10/12] drm/amdgpu: Avoid sysfs dirs removal post device unplug
Date: Tue, 24 Nov 2020 17:27:42 -0500 [thread overview]
Message-ID: <36fdb2f8-2238-6321-201e-a25a3a828fc5@amd.com> (raw)
In-Reply-To: <20201124144938.GR401619@phenom.ffwll.local>
On 11/24/20 9:49 AM, Daniel Vetter wrote:
> On Sat, Nov 21, 2020 at 12:21:20AM -0500, Andrey Grodzovsky wrote:
>> Avoids NULL ptr due to kobj->sd being unset on device removal.
>>
>> Signed-off-by: Andrey Grodzovsky <andrey.grodzovsky@amd.com>
>> ---
>> drivers/gpu/drm/amd/amdgpu/amdgpu_ras.c | 4 +++-
>> drivers/gpu/drm/amd/amdgpu/amdgpu_ucode.c | 4 +++-
>> 2 files changed, 6 insertions(+), 2 deletions(-)
>>
>> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_ras.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_ras.c
>> index caf828a..812e592 100644
>> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_ras.c
>> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_ras.c
>> @@ -27,6 +27,7 @@
>> #include <linux/uaccess.h>
>> #include <linux/reboot.h>
>> #include <linux/syscalls.h>
>> +#include <drm/drm_drv.h>
>>
>> #include "amdgpu.h"
>> #include "amdgpu_ras.h"
>> @@ -1043,7 +1044,8 @@ static int amdgpu_ras_sysfs_remove_feature_node(struct amdgpu_device *adev)
>> .attrs = attrs,
>> };
>>
>> - sysfs_remove_group(&adev->dev->kobj, &group);
>> + if (!drm_dev_is_unplugged(&adev->ddev))
>> + sysfs_remove_group(&adev->dev->kobj, &group);
> This looks wrong. sysfs, like any other interface, should be
> unconditionally thrown out when we do the drm_dev_unregister. Whether
> hotunplugged or not should matter at all. Either this isn't needed at all,
> or something is wrong with the ordering here. But definitely fishy.
> -Daniel
So technically this is needed because kobejct's sysfs directory entry kobj->sd
is set to NULL
on device removal (from sysfs_remove_dir) but because we don't finalize the device
until last reference to drm file is dropped (which can happen later) we end up
calling sysfs_remove_file/dir after
this pointer is NULL. sysfs_remove_file checks for NULL and aborts while
sysfs_remove_dir
is not and that why I guard against calls to sysfs_remove_dir.
But indeed the whole approach in the driver is incorrect, as Greg pointed out -
we should use
default groups attributes instead of explicit calls to sysfs interface and this
would save those troubles.
But again. the issue here of scope of work, converting all of amdgpu to default
groups attributes is somewhat
lengthy process with extra testing as the entire driver is papered with sysfs
references and seems to me more of a standalone
cleanup, just like switching to devm_ and drmm_ work. To me at least it seems
that it makes more sense
to finalize and push the hot unplug patches so that this new functionality can
be part of the driver sooner
and then incrementally improve it by working on those other topics. Just as
devm_/drmm_ I also added sysfs cleanup
to my TODO list in the RFC patch.
Andrey
>
>>
>> return 0;
>> }
>> diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_ucode.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_ucode.c
>> index 2b7c90b..54331fc 100644
>> --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_ucode.c
>> +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_ucode.c
>> @@ -24,6 +24,7 @@
>> #include <linux/firmware.h>
>> #include <linux/slab.h>
>> #include <linux/module.h>
>> +#include <drm/drm_drv.h>
>>
>> #include "amdgpu.h"
>> #include "amdgpu_ucode.h"
>> @@ -464,7 +465,8 @@ int amdgpu_ucode_sysfs_init(struct amdgpu_device *adev)
>>
>> void amdgpu_ucode_sysfs_fini(struct amdgpu_device *adev)
>> {
>> - sysfs_remove_group(&adev->dev->kobj, &fw_attr_group);
>> + if (!drm_dev_is_unplugged(&adev->ddev))
>> + sysfs_remove_group(&adev->dev->kobj, &fw_attr_group);
>> }
>>
>> static int amdgpu_ucode_init_single_fw(struct amdgpu_device *adev,
>> --
>> 2.7.4
>>
_______________________________________________
amd-gfx mailing list
amd-gfx@lists.freedesktop.org
https://lists.freedesktop.org/mailman/listinfo/amd-gfx
next prev parent reply other threads:[~2020-11-24 22:27 UTC|newest]
Thread overview: 105+ messages / expand[flat|nested] mbox.gz Atom feed top
2020-11-21 5:21 [PATCH v3 00/12] RFC Support hot device unplug in amdgpu Andrey Grodzovsky
2020-11-21 5:21 ` [PATCH v3 01/12] drm: Add dummy page per device or GEM object Andrey Grodzovsky
2020-11-21 14:15 ` Christian König
2020-11-23 4:54 ` Andrey Grodzovsky
2020-11-23 8:01 ` Christian König
2021-01-05 21:04 ` Andrey Grodzovsky
2021-01-07 16:21 ` Daniel Vetter
2021-01-07 16:26 ` Andrey Grodzovsky
2021-01-07 16:28 ` Andrey Grodzovsky
2021-01-07 16:30 ` Daniel Vetter
2021-01-07 16:37 ` Andrey Grodzovsky
2021-01-08 14:26 ` Andrey Grodzovsky
2021-01-08 14:33 ` Christian König
2021-01-08 14:46 ` Andrey Grodzovsky
2021-01-08 14:52 ` Christian König
2021-01-08 16:49 ` Grodzovsky, Andrey
2021-01-11 16:13 ` Daniel Vetter
2021-01-11 16:15 ` Daniel Vetter
2021-01-11 17:41 ` Andrey Grodzovsky
2021-01-11 20:45 ` Andrey Grodzovsky
2021-01-12 9:10 ` Daniel Vetter
2021-01-12 12:32 ` Christian König
2021-01-12 15:59 ` Andrey Grodzovsky
2021-01-13 9:14 ` Christian König
2021-01-13 14:40 ` Andrey Grodzovsky
2021-01-12 15:54 ` Andrey Grodzovsky
2021-01-12 8:12 ` Christian König
2021-01-12 9:13 ` Daniel Vetter
2020-11-21 5:21 ` [PATCH v3 02/12] drm: Unamp the entire device address space on device unplug Andrey Grodzovsky
2020-11-21 14:16 ` Christian König
2020-11-24 14:44 ` Daniel Vetter
2020-11-21 5:21 ` [PATCH v3 03/12] drm/ttm: Remap all page faults to per process dummy page Andrey Grodzovsky
2020-11-21 5:21 ` [PATCH v3 04/12] drm/ttm: Set dma addr to null after freee Andrey Grodzovsky
2020-11-21 14:13 ` Christian König
2020-11-23 5:15 ` Andrey Grodzovsky
2020-11-23 8:04 ` Christian König
2020-11-21 5:21 ` [PATCH v3 05/12] drm/ttm: Expose ttm_tt_unpopulate for driver use Andrey Grodzovsky
2020-11-25 10:42 ` Christian König
2020-11-23 20:05 ` Andrey Grodzovsky
2020-11-23 20:20 ` Christian König
2020-11-23 20:38 ` Andrey Grodzovsky
2020-11-23 20:41 ` Christian König
2020-11-23 21:08 ` Andrey Grodzovsky
2020-11-24 7:41 ` Christian König
2020-11-24 16:22 ` Andrey Grodzovsky
2020-11-24 16:44 ` Christian König
2020-11-25 10:40 ` Daniel Vetter
2020-11-25 12:57 ` Christian König
2020-11-25 16:36 ` Daniel Vetter
2020-11-25 19:34 ` Andrey Grodzovsky
2020-11-27 13:10 ` Grodzovsky, Andrey
2020-11-27 14:59 ` Daniel Vetter
2020-11-27 16:04 ` Andrey Grodzovsky
2020-11-30 14:15 ` Daniel Vetter
2020-11-25 16:56 ` Michel Dänzer
2020-11-25 17:02 ` Daniel Vetter
2020-12-15 20:18 ` Andrey Grodzovsky
2020-12-16 8:04 ` Christian König
2020-12-16 14:21 ` Daniel Vetter
2020-12-16 16:13 ` Andrey Grodzovsky
2020-12-16 16:18 ` Christian König
2020-12-16 17:12 ` Daniel Vetter
2020-12-16 17:20 ` Daniel Vetter
2020-12-16 18:26 ` Andrey Grodzovsky
2020-12-16 23:15 ` Daniel Vetter
2020-12-17 0:20 ` Andrey Grodzovsky
2020-12-17 12:01 ` Daniel Vetter
2020-12-17 19:19 ` Andrey Grodzovsky
2020-12-17 20:10 ` Christian König
2020-12-17 20:38 ` Andrey Grodzovsky
2020-12-17 20:48 ` Daniel Vetter
2020-12-17 21:06 ` Andrey Grodzovsky
2020-12-18 14:30 ` Daniel Vetter
2020-12-17 20:42 ` Daniel Vetter
2020-12-17 21:13 ` Andrey Grodzovsky
2021-01-04 16:33 ` Andrey Grodzovsky
2020-11-21 5:21 ` [PATCH v3 06/12] drm/sched: Cancel and flush all oustatdning jobs before finish Andrey Grodzovsky
2020-11-22 11:56 ` Christian König
2020-11-21 5:21 ` [PATCH v3 07/12] drm/sched: Prevent any job recoveries after device is unplugged Andrey Grodzovsky
2020-11-22 11:57 ` Christian König
2020-11-23 5:37 ` Andrey Grodzovsky
2020-11-23 8:06 ` Christian König
2020-11-24 1:12 ` Luben Tuikov
2020-11-24 7:50 ` Christian König
2020-11-24 17:11 ` Luben Tuikov
2020-11-24 17:17 ` Andrey Grodzovsky
2020-11-24 17:41 ` Luben Tuikov
2020-11-24 17:40 ` Christian König
2020-11-24 17:44 ` Luben Tuikov
2020-11-21 5:21 ` [PATCH v3 08/12] drm/amdgpu: Split amdgpu_device_fini into early and late Andrey Grodzovsky
2020-11-24 14:53 ` Daniel Vetter
2020-11-24 15:51 ` Andrey Grodzovsky
2020-11-25 10:41 ` Daniel Vetter
2020-11-25 17:41 ` Andrey Grodzovsky
2020-11-21 5:21 ` [PATCH v3 09/12] drm/amdgpu: Add early fini callback Andrey Grodzovsky
2020-11-21 5:21 ` [PATCH v3 10/12] drm/amdgpu: Avoid sysfs dirs removal post device unplug Andrey Grodzovsky
2020-11-24 14:49 ` Daniel Vetter
2020-11-24 22:27 ` Andrey Grodzovsky [this message]
2020-11-25 9:04 ` Daniel Vetter
2020-11-25 17:39 ` Andrey Grodzovsky
2020-11-27 13:12 ` Grodzovsky, Andrey
2020-11-27 15:04 ` Daniel Vetter
2020-11-27 15:34 ` Andrey Grodzovsky
2020-11-21 5:21 ` [PATCH v3 11/12] drm/amdgpu: Register IOMMU topology notifier per device Andrey Grodzovsky
2020-11-21 5:21 ` [PATCH v3 12/12] drm/amdgpu: Fix a bunch of sdma code crash post device unplug Andrey Grodzovsky
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=36fdb2f8-2238-6321-201e-a25a3a828fc5@amd.com \
--to=andrey.grodzovsky@amd.com \
--cc=Alexander.Deucher@amd.com \
--cc=Harry.Wentland@amd.com \
--cc=amd-gfx@lists.freedesktop.org \
--cc=ckoenig.leichtzumerken@gmail.com \
--cc=daniel.vetter@ffwll.ch \
--cc=daniel@ffwll.ch \
--cc=dri-devel@lists.freedesktop.org \
--cc=eric@anholt.net \
--cc=gregkh@linuxfoundation.org \
--cc=l.stach@pengutronix.de \
--cc=ppaalanen@gmail.com \
--cc=robh@kernel.org \
--cc=yuq825@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox