From: Felix Kuehling <felix.kuehling-5C7GfCeVMHo@public.gmane.org>
To: Jan Vesely <jan.vesely-kgbqMDwikbSVc3sceRu5cw@public.gmane.org>,
Oded Gabbay <oded.gabbay-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org>
Cc: amd-gfx list
<amd-gfx-PD4FTy7X32lNgt0PjOBp9y5qC8QIuHrW@public.gmane.org>,
Maling list - DRI developers
<dri-devel-PD4FTy7X32lNgt0PjOBp9y5qC8QIuHrW@public.gmane.org>
Subject: Re: [PATCH 1/1] drm/amdkfd: Do not ignore requested queue size during allocation
Date: Fri, 1 Dec 2017 12:10:20 -0500 [thread overview]
Message-ID: <e9cedbd3-9a18-769b-08ed-23a1bf344c15@amd.com> (raw)
In-Reply-To: <1512085889.3631.50.camel-kgbqMDwikbSVc3sceRu5cw@public.gmane.org>
DIQ is the debug interface queue. Are you running a GPU debugger?
Otherwise I would not expect to even see a DIQ.
Are you not seeing any compute queues in mqds? If there are no compute
queues in mqds, that means your queue has been destroyed. That would
explain why the read pointer is not advancing.
Regards,
Felix
On 2017-11-30 06:51 PM, Jan Vesely wrote:
> On Wed, 2017-11-29 at 16:58 -0500, Felix Kuehling wrote:
>> You can see the state of the queues in debugfs:
>> /sys/kernel/debug/kfd/... You can look at MQDs and HQDs.
> thanks. how do I decode the information?
> The rptr always stops at pos 60 which looks like this in mqds:
>
> DIQ on device 45a2
> 00000000: c0310800 00004000 00000000 00000000 00000000 00000000 00000000 00000000
> 00000020: 00000000 00000000 00000000 00000001 00000000 00000000 00000000 00000000
> 00000040: 00000000 00000000 00000000 00000000 00000000 00000000 00000000 ffffffff
> 00000060: ffffffff 00000000 ffffffff ffffffff 00000000 00000000 00000000 00000000
>
> If I understood correctly that's the queue dump, so those fffffs look
> wrong
>
>> If your application isn't stopping queues deliberately, queues get
>> disabled by evictions, usually temporarily. You'll see kernel messages
>> when that happens.
>>
>> A VM fault will result in queues of the offending process getting
>> disabled permanently. Again, you'll see messages about that in the
>> kernel log.
>>
>> The RPTR can also stop advancing if you have an infinite loop in a
>> shader program, or just a shader that takes a very long time to execute.
>> Or maybe if you have some dependencies (barriers) in your AQL packets
>> that never get satisfied.
>>
>> The function you changed only affects the HIQ, the queue that KFD uses
>> to control the HWS. It does not affect user mode queues. If your problem
>> is with a user mode queue, your change should have no effect at all.
> It's not a userspace queue that stops. I'm using kernel dbgdev to issue
> wave_resume commands. (waves are halted after executing
> s_sendmsg_halt).
> I bumped KFD_KERNEL_QUEUE_SIZE to 16KB to make sure all 320 resume
> commads fit (otherwise I get spurious ENOMEM when the queue is full but
> still advancing).
>
> thanks,
> Jan
>
>> Regards,
>> Felix
>>
>>
>> On 2017-11-29 04:43 PM, Jan Vesely wrote:
>>> On Mon, 2017-11-20 at 14:22 -0500, Felix Kuehling wrote:
>>>> I think this patch is not correct. The EOP-mem is not associated with
>>>> the queue size. The EOP buffer is a separate buffer used by the firmware
>>>> to handle command completion. As I understand it, this allows more
>>>> concurrency, while still making it look like all commands in the queue
>>>> are completing in order.
>>> thanks for the explanation. I was looking for a source of a CP hang
>>> (rptr stops advancing), but bumping the eop size actually mode things
>>> worse. Is there a way to find out if a queue got disabled and for what
>>> reason? (I'm running ROCK-1.6.x based kernel)
>>>
>>> thanks,
>>> Jan
>>>
>>>> Regards,
>>>> Felix
>>>>
>>>>
>>>> On 2017-11-19 03:19 AM, Oded Gabbay wrote:
>>>>> On Thu, Nov 16, 2017 at 11:36 PM, Jan Vesely <jan.vesely@rutgers.edu> wrote:
>>>>>> Signed-off-by: Jan Vesely <jan.vesely@rutgers.edu>
>>>>>> ---
>>>>>> drivers/gpu/drm/amd/amdkfd/kfd_kernel_queue_vi.c | 5 +++--
>>>>>> 1 file changed, 3 insertions(+), 2 deletions(-)
>>>>>>
>>>>>> diff --git a/drivers/gpu/drm/amd/amdkfd/kfd_kernel_queue_vi.c b/drivers/gpu/drm/amd/amdkfd/kfd_kernel_queue_vi.c
>>>>>> index f1d48281e322..b3bee39661ab 100644
>>>>>> --- a/drivers/gpu/drm/amd/amdkfd/kfd_kernel_queue_vi.c
>>>>>> +++ b/drivers/gpu/drm/amd/amdkfd/kfd_kernel_queue_vi.c
>>>>>> @@ -37,15 +37,16 @@ static bool initialize_vi(struct kernel_queue *kq, struct kfd_dev *dev,
>>>>>> enum kfd_queue_type type, unsigned int queue_size)
>>>>>> {
>>>>>> int retval;
>>>>>> + unsigned int size = ALIGN(queue_size, PAGE_SIZE);
>>>>>>
>>>>>> - retval = kfd_gtt_sa_allocate(dev, PAGE_SIZE, &kq->eop_mem);
>>>>>> + retval = kfd_gtt_sa_allocate(dev, size, &kq->eop_mem);
>>>>>> if (retval != 0)
>>>>>> return false;
>>>>>>
>>>>>> kq->eop_gpu_addr = kq->eop_mem->gpu_addr;
>>>>>> kq->eop_kernel_addr = kq->eop_mem->cpu_ptr;
>>>>>>
>>>>>> - memset(kq->eop_kernel_addr, 0, PAGE_SIZE);
>>>>>> + memset(kq->eop_kernel_addr, 0, size);
>>>>>>
>>>>>> return true;
>>>>>> }
>>>>>> --
>>>>>> 2.13.6
>>>>>>
>>>>>> _______________________________________________
>>>>>> amd-gfx mailing list
>>>>>> amd-gfx@lists.freedesktop.org
>>>>>> https://lists.freedesktop.org/mailman/listinfo/amd-gfx
>>>>> Thanks!
>>>>> Applied to -next tree
>>>>> Oded
>>>>> _______________________________________________
>>>>> amd-gfx mailing list
>>>>> amd-gfx@lists.freedesktop.org
>>>>> https://lists.freedesktop.org/mailman/listinfo/amd-gfx
>>
_______________________________________________
amd-gfx mailing list
amd-gfx@lists.freedesktop.org
https://lists.freedesktop.org/mailman/listinfo/amd-gfx
next prev parent reply other threads:[~2017-12-01 17:10 UTC|newest]
Thread overview: 11+ messages / expand[flat|nested] mbox.gz Atom feed top
2017-11-16 21:36 [PATCH 1/1] drm/amdkfd: Do not ignore requested queue size during allocation Jan Vesely
[not found] ` <20171116213631.3987-1-jan.vesely-kgbqMDwikbSVc3sceRu5cw@public.gmane.org>
2017-11-19 8:19 ` Oded Gabbay
[not found] ` <CAFCwf10MTrBcGn1kejNvn9AcHDsCiW85HkC6bSmYfY=1shGJxQ-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
2017-11-20 19:22 ` Felix Kuehling
[not found] ` <21e77adc-4fbe-a3e9-0a02-5d84eb201561-5C7GfCeVMHo@public.gmane.org>
2017-11-21 11:44 ` Oded Gabbay
[not found] ` <CAFCwf12ZeO=6F-EQ_MFgSEugxMSAFNe+OkGT_gZBMp34VzFAcA-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>
2017-11-21 16:30 ` Felix Kuehling
2017-11-29 21:43 ` Jan Vesely
[not found] ` <1511991803.2978.67.camel-kgbqMDwikbSVc3sceRu5cw@public.gmane.org>
2017-11-29 21:58 ` Felix Kuehling
2017-11-30 23:51 ` Jan Vesely
[not found] ` <1512085889.3631.50.camel-kgbqMDwikbSVc3sceRu5cw@public.gmane.org>
2017-12-01 17:10 ` Felix Kuehling [this message]
2017-12-01 17:15 ` Felix Kuehling
2017-12-01 19:37 ` Felix Kuehling
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=e9cedbd3-9a18-769b-08ed-23a1bf344c15@amd.com \
--to=felix.kuehling-5c7gfcevmho@public.gmane.org \
--cc=amd-gfx-PD4FTy7X32lNgt0PjOBp9y5qC8QIuHrW@public.gmane.org \
--cc=dri-devel-PD4FTy7X32lNgt0PjOBp9y5qC8QIuHrW@public.gmane.org \
--cc=jan.vesely-kgbqMDwikbSVc3sceRu5cw@public.gmane.org \
--cc=oded.gabbay-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox