Linux USB
 help / color / mirror / Atom feed
From: Krishna Kurapati <krishna.kurapati@oss.qualcomm.com>
To: Selvarasu Ganesan <selvarasu.g@samsung.com>,
	Thinh Nguyen <Thinh.Nguyen@synopsys.com>
Cc: linux-usb@vger.kernel.org, linux-kernel@vger.kernel.org,
	Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Subject: Re: [PATCH v3] usb: dwc3: core: Fix RAM interface getting stuck during enumeration
Date: Tue, 29 Sep 2026 14:05:52 +0530	[thread overview]
Message-ID: <647d8c0e-9323-4b1e-9720-9e51ea8e3227@oss.qualcomm.com> (raw)
In-Reply-To: <e8b9b694-0ba5-4d0f-97c9-ceeb89ae4729@samsung.com>



On 9/29/2026 12:15 PM, Selvarasu Ganesan wrote:
> 
> On 9/29/2026 10:49 AM, Krishna Kurapati wrote:
>> During plug-in/plug-out test cases, it is sometimes seen that no events
>> are generated by the controller and all CSR register reads give "0" and
>> CSR_Timeout bit gets set indicating that CSR reads/writes are timing out
>> or timed out.
> Hi Krishna,
> 
> As per our discussion with synopsys we got a below feedback for CSR
> timeout in different issue.
> 
> "Note: CSR timeout is not expected to occur during functional mode. Once
> timeout is reported by controller, s/w should treat it as a fatal error
> and should identify the root cause of issuing the parallel access during
> soft reset and fix it"
> 
> As per our understanding, there is no recovery from CSR timeout.
> 
> Are you able to recover from CSR timeout in your case even after trigger
> controller recovery?.
> 

Disabling and re-enabling gadget mode is the recommended fix. I did not 
test this particular patch as mentioned below, but this error recovery 
mechanism does the same thing as a soft_disconnect followed by soft_connect.

Regards,
Krishna,

> 
>>
>> The issue comes up on different instnaces of enumeration on different
>> platforms. On SM8550, the debug log is as follows:
>>
>> Prepared a TRB on ep0out and did start transfer to get set
>> address request from host:
>>
>> <...>-7191    [000] D..1.    66.421006: dwc3_gadget_ep_cmd: ep0out:
>> cmd 'Start Transfer' [406] params 00000000 efffa000 00000000 -->
>> status: Successful
>>
>> <...>-7191    [000] D..1.    66.421196: dwc3_event: event (0000c040):
>> ep0out: Transfer Complete (sIL) [Setup Phase]
>>
>> <...>-7191    [000] D..1.    66.421197: dwc3_ctrl_req: Set
>> Address(Addr = 01)
>>
>> An XFER NRDY is received on ep0in for zero length status phase and
>> a Start Transfer was done on ep0in with 0-length packet in 2 Stage
>> status phase:
>>
>> <...>-7191    [000] D..1.    66.421249: dwc3_event: event (000020c2):
>> ep0in: Transfer Not Ready [00000000] (Not Active) [Status Phase]
>>
>> <...>-7191    [000] D..1.    66.421266: dwc3_prepare_trb: ep0in: trb
>> ffffffc00fcfd000 (E0:D0) buf 00000000efffa000 size 0 ctrl 00000c33
>> sofn 00000000 (HLcs:SC:status2)
>>
>> <...>-7191    [000] D..1.    66.421387: dwc3_gadget_ep_cmd: ep0in: cmd
>> 'Start Transfer' [406] params 00000000 efffa000 00000000 -->status:
>> Successful
>>
>> A bus reset was then received directly after 500 msec. Software never
>> got the cmd complete for the start transfer done in status phase. Here
>> the RAM interface is stuck. So host issues a bus reset as link is
>> idle for 500 msec:
>>
>> <...>-7191    [000] D..1.    66.935603: dwc3_event: event (00000101):
>> Reset [U0]
>>
>> Then software sees that it is in status phase and we issue an ENDXFER
>> on ep0in and it gets timedout waiting for the CMDACT to go '0':
>>
>> <...>-7191    [000] D..1.    66.958249: dwc3_gadget_ep_cmd: ep0in: cmd
>> 'End Transfer' [10508] params 00000000 00000000 00000000 --> status:
>> Timed Out
>>
>> Upon debug with Synopsys, the root cause is as follows:
>>
>> During any transfer, if the data is not successfully transmitted,
>> then a Done (with failure) handshake is returned, so that the BMU
>> can re-attempt the same data again by rewinding its data pointers.
>>
>> But, if the USB IN is a 0-length payload (which is what is happening
>> in this case - 2 stage status phase of set_address), then there is no
>> need to rewind the pointers and the Done (with failure) handshake is
>> not returned for failure case. This keeps the Request-Done interface
>> busy till the next Done handshake. The MAC sends the 0-length payload
>> again when the host requests. If the transmission is successful this
>> time, the Done (with success) handshake is provided back. Otherwise,
>> it repeats the same steps again.
>>
>> If the cable is disconnected or if the Host aborts the transfer on 3
>> consecutive failed attempts, the Request-Done handshake is not
>> complete. This keeps the interface busy.
>>
>> The subsequent RAM access cannot proceed until the above pending
>> transfer is complete. This results in failure of any access to RAM
>> address locations. Many of the EndPoint commands need to access the
>> RAM and they would fail to complete successfully.
>>
>> Furthermore when cable removal happens, this would not generate a
>> disconnect event and the "connected" flag remains true always blockin
>> suspend.
>>
>> Synopsys confirmed that the issue is present on all USB3 devices and
>> as a workaround, suggested to re-initialize device mode.
>>
>> Signed-off-by: Krishna Kurapati <krishna.kurapati@oss.qualcomm.com>
>> ---
>> This series has only been compile tested. The issue was reproduced
>> easily with a certain kind of cable and CDP port of AMD based Lenovo
>> laptop. I don't have access to the cable currently and hence only
>> compile testing the fix for now. But the issue has popped up on OEM
>> testing as well.
>>
>> Also, didn't add locking while calling error recovery work in
>> gadget_ep_cmd since the caller is supposed to handle it.
>>
>> Changes to v3:
>> - Using error receovery mechanism from [1].
>>
>> Link to v2:
>> https://lore.kernel.org/all/20260806-ram-interface-code-v2-1-fe4a0de31d42@oss.qualcomm.com/
>>
>> Changes in v2:
>> - Implemented gadget recovery mechanism during gadget_ep_cmd instead of
>> handling this issue only during disconnect.
>>
>> Link to RFC:
>> https://lore.kernel.org/all/20231011100214.25720-1-quic_kriskura@quicinc.com/
>>
>> [1]: https://lore.kernel.org/all/20260915110637.17658-1-jiazi.liu1984@gmail.com/
>> ---
>>    drivers/usb/dwc3/gadget.c | 19 +++++++++++++++++++
>>    1 file changed, 19 insertions(+)
>>
>> diff --git a/drivers/usb/dwc3/gadget.c b/drivers/usb/dwc3/gadget.c
>> index ee837235630a..a7c4514cf0e2 100644
>> --- a/drivers/usb/dwc3/gadget.c
>> +++ b/drivers/usb/dwc3/gadget.c
>> @@ -283,6 +283,8 @@ int dwc3_send_gadget_generic_command(struct dwc3 *dwc, unsigned int cmd,
>>    	return ret;
>>    }
>>    
>> +static void dwc3_schedule_err_recovery(struct dwc3 *dwc);
> 
> Where is the function implementation for this function?. Could you
> please point me on what you are doing in this function?
> 
> 
> Thanks,
> Selva
> 
> 
>> +
>>    /**
>>     * dwc3_send_gadget_ep_cmd - issue an endpoint command
>>     * @dep: the endpoint to which the command is going to be issued
>> @@ -432,6 +434,23 @@ int dwc3_send_gadget_ep_cmd(struct dwc3_ep *dep, unsigned int cmd,
>>    		cmd_status = -ETIMEDOUT;
>>    	}
>>    
>> +	/*
>> +	 * STAR 5001544 - In some situations, like the cable is
>> +	 * disconnected or if the Host aborts the transfer on 3
>> +	 * consecutive failed attempts, the Request-Done handshake is not
>> +	 * complete. This keeps the RAM interface busy.
>> +	 *
>> +	 * The subsequent RAM access cannot proceed until the pending
>> +	 * transfer is complete. This results in failure of any access
>> +	 * to RAM address locations. Many of the EndPoint commands need to
>> +	 * access the RAM and they would fail to complete successfully.
>> +	 *
>> +	 * If the depcmd doesn't match the actual command, trigger controller
>> +	 * recovery.
>> +	 */
>> +	if (DWC3_DEPCMD_CMD(reg) != DWC3_DEPCMD_CMD(cmd))
>> +		dwc3_schedule_err_recovery(dwc);
>> +
>>    skip_status:
>>    	trace_dwc3_gadget_ep_cmd(dep, cmd, params, cmd_status);
>>    
>>
>> ---
>> base-commit: ab29ca7714b82485ebb31d66835ccd5364221a76
>> change-id: 20260929-ram-interface-stuck-v3-aaf0c4b13b98
>>
>> Best regards,
>> --
>> Krishna Kurapati <krishna.kurapati@oss.qualcomm.com>


  reply	other threads:[~2026-09-29  8:36 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
     [not found] <CGME20260929051941epcas5p27a5c69151937b9105660941776cda819@epcas5p2.samsung.com>
2026-09-29  5:19 ` [PATCH v3] usb: dwc3: core: Fix RAM interface getting stuck during enumeration Krishna Kurapati
2026-09-29  6:45   ` Selvarasu Ganesan
2026-09-29  8:35     ` Krishna Kurapati [this message]
2026-09-29 11:02       ` Selvarasu Ganesan
2026-09-29 17:50         ` Krishna Kurapati

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=647d8c0e-9323-4b1e-9720-9e51ea8e3227@oss.qualcomm.com \
    --to=krishna.kurapati@oss.qualcomm.com \
    --cc=Thinh.Nguyen@synopsys.com \
    --cc=gregkh@linuxfoundation.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-usb@vger.kernel.org \
    --cc=selvarasu.g@samsung.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox