All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH] SUNRPC: add transport details to timeout diagnostics
@ 2026-08-19 17:30 Manjunath Patil
  2026-08-20 11:13 ` Petr Vorel
  2026-10-06 21:23 ` Anna Schumaker
  0 siblings, 2 replies; 9+ messages in thread
From: Manjunath Patil @ 2026-08-19 17:30 UTC (permalink / raw)
  To: Trond Myklebust, Anna Schumaker
  Cc: linux-nfs, Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia,
	Dai Ngo, Tom Talpey, Manjunath Patil

The existing RPC timeout messages identify only the RPC program and
server name. That makes it difficult to distinguish a server-side
response stall from a transport that is disconnected, reconnecting,
or using a different connection.

Include the RPC client id, transport id, current transport connection
cookie, and RPC task owner in the ratelimited timeout messages. These
fields provide a compact way to correlate a timeout message with
rpcctl and other transport state.

For example, an NFS timeout reports:

  nfs: server 192.168.50.1 clid=4 xprt=1 cc=1 owner=6245
  not responding, timed out

Assisted-by: Codex:gpt-5
Signed-off-by: Manjunath Patil <manjunath.b.patil@oracle.com>
---
 net/sunrpc/clnt.c | 20 +++++++++++++-------
 1 file changed, 13 insertions(+), 7 deletions(-)

diff --git a/net/sunrpc/clnt.c b/net/sunrpc/clnt.c
index efa26899bc7d..6b5e2787dc81 100644
--- a/net/sunrpc/clnt.c
+++ b/net/sunrpc/clnt.c
@@ -2534,17 +2534,21 @@ static void
 rpc_check_timeout(struct rpc_task *task)
 {
 	struct rpc_clnt	*clnt = task->tk_client;
+	struct rpc_rqst	*req;
+	struct rpc_xprt	*xprt;
 
 	if (RPC_SIGNALLED(task))
 		return;
 
-	if (xprt_adjust_timeout(task->tk_rqstp) == 0)
+	req = task->tk_rqstp;
+	if (xprt_adjust_timeout(req) == 0)
 		return;
+	xprt = req->rq_xprt;
 
 	trace_rpc_timeout_status(task);
 	task->tk_timeouts++;
 
-	if (RPC_IS_SOFTCONN(task) && !rpc_check_connected(task->tk_rqstp)) {
+	if (RPC_IS_SOFTCONN(task) && !rpc_check_connected(req)) {
 		rpc_call_rpcerror(task, -ETIMEDOUT);
 		return;
 	}
@@ -2556,14 +2560,15 @@ rpc_check_timeout(struct rpc_task *task)
 		 * connection gets terminally broken.
 		 */
 		if ((task->tk_flags & RPC_TASK_NO_RETRANS_TIMEOUT) &&
-		    rpc_check_connected(task->tk_rqstp))
+		    rpc_check_connected(req))
 			return;
 
 		if (clnt->cl_chatty) {
 			pr_notice_ratelimited(
-				"%s: server %s not responding, timed out\n",
+				"%s: server %s clid=%u xprt=%u cc=%u owner=%d not responding, timed out\n",
 				clnt->cl_program->name,
-				task->tk_xprt->servername);
+				xprt->servername, clnt->cl_clid, xprt->id,
+				READ_ONCE(xprt->connect_cookie), task->tk_owner);
 		}
 		if (task->tk_flags & RPC_TASK_TIMEOUT)
 			rpc_call_rpcerror(task, -ETIMEDOUT);
@@ -2576,9 +2581,10 @@ rpc_check_timeout(struct rpc_task *task)
 		task->tk_flags |= RPC_CALL_MAJORSEEN;
 		if (clnt->cl_chatty) {
 			pr_notice_ratelimited(
-				"%s: server %s not responding, still trying\n",
+				"%s: server %s clid=%u xprt=%u cc=%u owner=%d not responding, still trying\n",
 				clnt->cl_program->name,
-				task->tk_xprt->servername);
+				xprt->servername, clnt->cl_clid, xprt->id,
+				READ_ONCE(xprt->connect_cookie), task->tk_owner);
 		}
 	}
 	rpc_force_rebind(clnt);
-- 
2.47.3


^ permalink raw reply related	[flat|nested] 9+ messages in thread

* [PATCH] SUNRPC: add transport details to timeout diagnostics
  2026-08-19 17:30 [PATCH] SUNRPC: add transport details to timeout diagnostics Manjunath Patil
@ 2026-08-20 11:13 ` Petr Vorel
  2026-08-20 16:15   ` manjunath.b.patil
  2026-10-06 21:23 ` Anna Schumaker
  1 sibling, 1 reply; 9+ messages in thread
From: Petr Vorel @ 2026-08-20 11:13 UTC (permalink / raw)
  To: manjunath.b.patil
  Cc: Dai.Ngo, anna, cel, jlayton, linux-nfs, neil, okorniev, tom,
	trondmy, Petr Vorel

> The existing RPC timeout messages identify only the RPC program and
> server name. That makes it difficult to distinguish a server-side
> response stall from a transport that is disconnected, reconnecting,
> or using a different connection.
> 
> Include the RPC client id, transport id, current transport connection
> cookie, and RPC task owner in the ratelimited timeout messages. These
> fields provide a compact way to correlate a timeout message with
> rpcctl and other transport state.
> 
> For example, an NFS timeout reports:
> 
>   nfs: server 192.168.50.1 clid=4 xprt=1 cc=1 owner=6245
>   not responding, timed out

Nice improvement, thanks!

> 
> Assisted-by: Codex:gpt-5
> Signed-off-by: Manjunath Patil <manjunath.b.patil@oracle.com>
> ---
>  net/sunrpc/clnt.c | 20 +++++++++++++-------
>  1 file changed, 13 insertions(+), 7 deletions(-)
> 
> diff --git a/net/sunrpc/clnt.c b/net/sunrpc/clnt.c
> index efa26899bc7d..6b5e2787dc81 100644
> --- a/net/sunrpc/clnt.c
> +++ b/net/sunrpc/clnt.c
> @@ -2534,17 +2534,21 @@ static void
>  rpc_check_timeout(struct rpc_task *task)
>  {
>  	struct rpc_clnt	*clnt = task->tk_client;
> +	struct rpc_rqst	*req;
> +	struct rpc_xprt	*xprt;

nit: Would it work to assign req and task earlier (just readability?)
Or is it not safe before checking RPC_SIGNALLED(task)?

	struct rpc_clnt	*clnt = task->tk_client;
	struct rpc_rqst	*req = task->tk_rqstp;
	struct rpc_xprt	*xprt = req->rq_xprt;

The rest LGTM.
Reviewed-by: Petr Vorel <pvorel@suse.cz>

Kind regards,
Petr


>  
>  	if (RPC_SIGNALLED(task))
>  		return;
>  
> -	if (xprt_adjust_timeout(task->tk_rqstp) == 0)
> +	req = task->tk_rqstp;
> +	if (xprt_adjust_timeout(req) == 0)
>  		return;
> +	xprt = req->rq_xprt;
>  
>  	trace_rpc_timeout_status(task);
>  	task->tk_timeouts++;
>  
> -	if (RPC_IS_SOFTCONN(task) && !rpc_check_connected(task->tk_rqstp)) {
> +	if (RPC_IS_SOFTCONN(task) && !rpc_check_connected(req)) {
>  		rpc_call_rpcerror(task, -ETIMEDOUT);
>  		return;
>  	}
> @@ -2556,14 +2560,15 @@ rpc_check_timeout(struct rpc_task *task)
>  		 * connection gets terminally broken.
>  		 */
>  		if ((task->tk_flags & RPC_TASK_NO_RETRANS_TIMEOUT) &&
> -		    rpc_check_connected(task->tk_rqstp))
> +		    rpc_check_connected(req))
>  			return;
>  
>  		if (clnt->cl_chatty) {
>  			pr_notice_ratelimited(
> -				"%s: server %s not responding, timed out\n",
> +				"%s: server %s clid=%u xprt=%u cc=%u owner=%d not responding, timed out\n",
>  				clnt->cl_program->name,
> -				task->tk_xprt->servername);
> +				xprt->servername, clnt->cl_clid, xprt->id,
> +				READ_ONCE(xprt->connect_cookie), task->tk_owner);
>  		}
>  		if (task->tk_flags & RPC_TASK_TIMEOUT)
>  			rpc_call_rpcerror(task, -ETIMEDOUT);
> @@ -2576,9 +2581,10 @@ rpc_check_timeout(struct rpc_task *task)
>  		task->tk_flags |= RPC_CALL_MAJORSEEN;
>  		if (clnt->cl_chatty) {
>  			pr_notice_ratelimited(
> -				"%s: server %s not responding, still trying\n",
> +				"%s: server %s clid=%u xprt=%u cc=%u owner=%d not responding, still trying\n",
>  				clnt->cl_program->name,
> -				task->tk_xprt->servername);
> +				xprt->servername, clnt->cl_clid, xprt->id,
> +				READ_ONCE(xprt->connect_cookie), task->tk_owner);
>  		}
>  	}
>  	rpc_force_rebind(clnt);

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH] SUNRPC: add transport details to timeout diagnostics
  2026-08-20 11:13 ` Petr Vorel
@ 2026-08-20 16:15   ` manjunath.b.patil
  2026-08-20 18:59     ` Petr Vorel
  0 siblings, 1 reply; 9+ messages in thread
From: manjunath.b.patil @ 2026-08-20 16:15 UTC (permalink / raw)
  To: Petr Vorel
  Cc: Dai.Ngo, anna, cel, jlayton, linux-nfs, neil, okorniev, tom,
	trondmy



On 8/20/26 4:13 AM, Petr Vorel wrote:
>  > The existing RPC timeout messages identify only the RPC program and > 
> server name. That makes it difficult to distinguish a server-side > 
> response stall from a transport that is disconnected, reconnecting, > or 
> using a different
> 
> 
>> The existing RPC timeout messages identify only the RPC program and
>> server name. That makes it difficult to distinguish a server-side
>> response stall from a transport that is disconnected, reconnecting,
>> or using a different connection.
>> 
>> Include the RPC client id, transport id, current transport connection
>> cookie, and RPC task owner in the ratelimited timeout messages. These
>> fields provide a compact way to correlate a timeout message with
>> rpcctl and other transport state.
>> 
>> For example, an NFS timeout reports:
>> 
>>   nfs: server 192.168.50.1 clid=4 xprt=1 cc=1 owner=6245
>>   not responding, timed out
> 
> Nice improvement, thanks!
> 
>> 
>> Assisted-by: Codex:gpt-5
>> Signed-off-by: Manjunath Patil <manjunath.b.patil@oracle.com>
>> ---
>>  net/sunrpc/clnt.c | 20 +++++++++++++-------
>>  1 file changed, 13 insertions(+), 7 deletions(-)
>> 
>> diff --git a/net/sunrpc/clnt.c b/net/sunrpc/clnt.c
>> index efa26899bc7d..6b5e2787dc81 100644
>> --- a/net/sunrpc/clnt.c
>> +++ b/net/sunrpc/clnt.c
>> @@ -2534,17 +2534,21 @@ static void
>>  rpc_check_timeout(struct rpc_task *task)
>>  {
>>  	struct rpc_clnt	*clnt = task->tk_client;
>> +	struct rpc_rqst	*req;
>> +	struct rpc_xprt	*xprt;
> 
> nit: Would it work to assign req and task earlier (just readability?)
> Or is it not safe before checking RPC_SIGNALLED(task)?
> 
> 	struct rpc_clnt	*clnt = task->tk_client;
> 	struct rpc_rqst	*req = task->tk_rqstp;
> 	struct rpc_xprt	*xprt = req->rq_xprt;
> 
> The rest LGTM.
> Reviewed-by: Petr Vorel <pvorel@suse.cz>
> 
> Kind regards,
> Petr

Hi Petr,

Thanks for the review and the Reviewed-by tag.

I kept the assignments after RPC_SIGNALLED(task) intentionally, 
preserving the existing behavior where a signalled task returns before 
we dereference task->tk_rqstp. xprt is also needed only after 
xprt_adjust_timeout(req) reports an expired timeout.

I would prefer to retain the current ordering for those reasons. Also I 
am open if maintainers chose to alter it.

Thanks, 
  
                 Manjunath
> 
> 
>>  
>>  	if (RPC_SIGNALLED(task))
>>  		return;
>>  
>> -	if (xprt_adjust_timeout(task->tk_rqstp) == 0)
>> +	req = task->tk_rqstp;
>> +	if (xprt_adjust_timeout(req) == 0)
>>  		return;
>> +	xprt = req->rq_xprt;
>>  
>>  	trace_rpc_timeout_status(task);
>>  	task->tk_timeouts++;
>>  
>> -	if (RPC_IS_SOFTCONN(task) && !rpc_check_connected(task->tk_rqstp)) {
>> +	if (RPC_IS_SOFTCONN(task) && !rpc_check_connected(req)) {
>>  		rpc_call_rpcerror(task, -ETIMEDOUT);
>>  		return;
>>  	}
>> @@ -2556,14 +2560,15 @@ rpc_check_timeout(struct rpc_task *task)
>>  		 * connection gets terminally broken.
>>  		 */
>>  		if ((task->tk_flags & RPC_TASK_NO_RETRANS_TIMEOUT) &&
>> -		    rpc_check_connected(task->tk_rqstp))
>> +		    rpc_check_connected(req))
>>  			return;
>>  
>>  		if (clnt->cl_chatty) {
>>  			pr_notice_ratelimited(
>> -				"%s: server %s not responding, timed out\n",
>> +				"%s: server %s clid=%u xprt=%u cc=%u owner=%d not responding, timed out\n",
>>  				clnt->cl_program->name,
>> -				task->tk_xprt->servername);
>> +				xprt->servername, clnt->cl_clid, xprt->id,
>> +				READ_ONCE(xprt->connect_cookie), task->tk_owner);
>>  		}
>>  		if (task->tk_flags & RPC_TASK_TIMEOUT)
>>  			rpc_call_rpcerror(task, -ETIMEDOUT);
>> @@ -2576,9 +2581,10 @@ rpc_check_timeout(struct rpc_task *task)
>>  		task->tk_flags |= RPC_CALL_MAJORSEEN;
>>  		if (clnt->cl_chatty) {
>>  			pr_notice_ratelimited(
>> -				"%s: server %s not responding, still trying\n",
>> +				"%s: server %s clid=%u xprt=%u cc=%u owner=%d not responding, still trying\n",
>>  				clnt->cl_program->name,
>> -				task->tk_xprt->servername);
>> +				xprt->servername, clnt->cl_clid, xprt->id,
>> +				READ_ONCE(xprt->connect_cookie), task->tk_owner);
>>  		}
>>  	}
>>  	rpc_force_rebind(clnt);
> 


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH] SUNRPC: add transport details to timeout diagnostics
  2026-08-20 16:15   ` manjunath.b.patil
@ 2026-08-20 18:59     ` Petr Vorel
  2026-08-31 19:33       ` manjunath.b.patil
  0 siblings, 1 reply; 9+ messages in thread
From: Petr Vorel @ 2026-08-20 18:59 UTC (permalink / raw)
  To: manjunath.b.patil
  Cc: Dai.Ngo, anna, cel, jlayton, linux-nfs, neil, okorniev, tom,
	trondmy



> On 8/20/26 4:13 AM, Petr Vorel wrote:
> >  > The existing RPC timeout messages identify only the RPC program and >
> > server name. That makes it difficult to distinguish a server-side >
> > response stall from a transport that is disconnected, reconnecting, > or
> > using a different


> > > The existing RPC timeout messages identify only the RPC program and
> > > server name. That makes it difficult to distinguish a server-side
> > > response stall from a transport that is disconnected, reconnecting,
> > > or using a different connection.

> > > Include the RPC client id, transport id, current transport connection
> > > cookie, and RPC task owner in the ratelimited timeout messages. These
> > > fields provide a compact way to correlate a timeout message with
> > > rpcctl and other transport state.

> > > For example, an NFS timeout reports:

> > >   nfs: server 192.168.50.1 clid=4 xprt=1 cc=1 owner=6245
> > >   not responding, timed out

> > Nice improvement, thanks!


> > > Assisted-by: Codex:gpt-5
> > > Signed-off-by: Manjunath Patil <manjunath.b.patil@oracle.com>
> > > ---
> > >  net/sunrpc/clnt.c | 20 +++++++++++++-------
> > >  1 file changed, 13 insertions(+), 7 deletions(-)

> > > diff --git a/net/sunrpc/clnt.c b/net/sunrpc/clnt.c
> > > index efa26899bc7d..6b5e2787dc81 100644
> > > --- a/net/sunrpc/clnt.c
> > > +++ b/net/sunrpc/clnt.c
> > > @@ -2534,17 +2534,21 @@ static void
> > >  rpc_check_timeout(struct rpc_task *task)
> > >  {
> > >  	struct rpc_clnt	*clnt = task->tk_client;
> > > +	struct rpc_rqst	*req;
> > > +	struct rpc_xprt	*xprt;

> > nit: Would it work to assign req and task earlier (just readability?)
> > Or is it not safe before checking RPC_SIGNALLED(task)?

> > 	struct rpc_clnt	*clnt = task->tk_client;
> > 	struct rpc_rqst	*req = task->tk_rqstp;
> > 	struct rpc_xprt	*xprt = req->rq_xprt;

> > The rest LGTM.
> > Reviewed-by: Petr Vorel <pvorel@suse.cz>

> > Kind regards,
> > Petr

> Hi Petr,

Hi Manjunath,

> Thanks for the review and the Reviewed-by tag.

yw.

> I kept the assignments after RPC_SIGNALLED(task) intentionally, preserving
> the existing behavior where a signalled task returns before we dereference
> task->tk_rqstp. xprt is also needed only after xprt_adjust_timeout(req)
> reports an expired timeout.

Understand, fair enough.

> I would prefer to retain the current ordering for those reasons. Also I am
> open if maintainers chose to alter it.

+1

Kind regards,
Petr

> Thanks,                   Manjunath


> > >  	if (RPC_SIGNALLED(task))
> > >  		return;
> > > -	if (xprt_adjust_timeout(task->tk_rqstp) == 0)
> > > +	req = task->tk_rqstp;
> > > +	if (xprt_adjust_timeout(req) == 0)
> > >  		return;
> > > +	xprt = req->rq_xprt;
> > >  	trace_rpc_timeout_status(task);
> > >  	task->tk_timeouts++;
> > > -	if (RPC_IS_SOFTCONN(task) && !rpc_check_connected(task->tk_rqstp)) {
> > > +	if (RPC_IS_SOFTCONN(task) && !rpc_check_connected(req)) {
> > >  		rpc_call_rpcerror(task, -ETIMEDOUT);
> > >  		return;
> > >  	}
> > > @@ -2556,14 +2560,15 @@ rpc_check_timeout(struct rpc_task *task)
> > >  		 * connection gets terminally broken.
> > >  		 */
> > >  		if ((task->tk_flags & RPC_TASK_NO_RETRANS_TIMEOUT) &&
> > > -		    rpc_check_connected(task->tk_rqstp))
> > > +		    rpc_check_connected(req))
> > >  			return;
> > >  		if (clnt->cl_chatty) {
> > >  			pr_notice_ratelimited(
> > > -				"%s: server %s not responding, timed out\n",
> > > +				"%s: server %s clid=%u xprt=%u cc=%u owner=%d not responding, timed out\n",
> > >  				clnt->cl_program->name,
> > > -				task->tk_xprt->servername);
> > > +				xprt->servername, clnt->cl_clid, xprt->id,
> > > +				READ_ONCE(xprt->connect_cookie), task->tk_owner);
> > >  		}
> > >  		if (task->tk_flags & RPC_TASK_TIMEOUT)
> > >  			rpc_call_rpcerror(task, -ETIMEDOUT);
> > > @@ -2576,9 +2581,10 @@ rpc_check_timeout(struct rpc_task *task)
> > >  		task->tk_flags |= RPC_CALL_MAJORSEEN;
> > >  		if (clnt->cl_chatty) {
> > >  			pr_notice_ratelimited(
> > > -				"%s: server %s not responding, still trying\n",
> > > +				"%s: server %s clid=%u xprt=%u cc=%u owner=%d not responding, still trying\n",
> > >  				clnt->cl_program->name,
> > > -				task->tk_xprt->servername);
> > > +				xprt->servername, clnt->cl_clid, xprt->id,
> > > +				READ_ONCE(xprt->connect_cookie), task->tk_owner);
> > >  		}
> > >  	}
> > >  	rpc_force_rebind(clnt);



^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH] SUNRPC: add transport details to timeout diagnostics
  2026-08-20 18:59     ` Petr Vorel
@ 2026-08-31 19:33       ` manjunath.b.patil
  2026-08-31 21:13         ` Chuck Lever
  2026-10-06 19:55         ` manjunath.b.patil
  0 siblings, 2 replies; 9+ messages in thread
From: manjunath.b.patil @ 2026-08-31 19:33 UTC (permalink / raw)
  To: Petr Vorel
  Cc: Dai.Ngo, anna, cel, jlayton, linux-nfs, neil, okorniev, tom,
	trondmy

On 8/20/26 11:59 AM, Petr Vorel wrote:
>  > On 8/20/26 4: 13 AM, Petr Vorel wrote: > > > The existing RPC timeout 
> messages identify only the RPC program and > > > server name. That makes 
> it difficult to distinguish a server-side > > > response stall from
> 
> 
>> On 8/20/26 4:13 AM, Petr Vorel wrote:
>> >  > The existing RPC timeout messages identify only the RPC program and >
>> > server name. That makes it difficult to distinguish a server-side >
>> > response stall from a transport that is disconnected, reconnecting, > or
>> > using a different
> 
> 
>> > > The existing RPC timeout messages identify only the RPC program and
>> > > server name. That makes it difficult to distinguish a server-side
>> > > response stall from a transport that is disconnected, reconnecting,
>> > > or using a different connection.
> 
>> > > Include the RPC client id, transport id, current transport connection
>> > > cookie, and RPC task owner in the ratelimited timeout messages. These
>> > > fields provide a compact way to correlate a timeout message with
>> > > rpcctl and other transport state.
> 
>> > > For example, an NFS timeout reports:
> 
>> > >   nfs: server 192.168.50.1 clid=4 xprt=1 cc=1 owner=6245
>> > >   not responding, timed out
> 
>> > Nice improvement, thanks!
> 
> 
>> > > Assisted-by: Codex:gpt-5
>> > > Signed-off-by: Manjunath Patil <manjunath.b.patil@oracle.com>
>> > > ---
>> > >  net/sunrpc/clnt.c | 20 +++++++++++++-------
>> > >  1 file changed, 13 insertions(+), 7 deletions(-)
> 
>> > > diff --git a/net/sunrpc/clnt.c b/net/sunrpc/clnt.c
>> > > index efa26899bc7d..6b5e2787dc81 100644
>> > > --- a/net/sunrpc/clnt.c
>> > > +++ b/net/sunrpc/clnt.c
>> > > @@ -2534,17 +2534,21 @@ static void
>> > >  rpc_check_timeout(struct rpc_task *task)
>> > >  {
>> > >  	struct rpc_clnt	*clnt = task->tk_client;
>> > > +	struct rpc_rqst	*req;
>> > > +	struct rpc_xprt	*xprt;
> 
>> > nit: Would it work to assign req and task earlier (just readability?)
>> > Or is it not safe before checking RPC_SIGNALLED(task)?
> 
>> > 	struct rpc_clnt	*clnt = task->tk_client;
>> > 	struct rpc_rqst	*req = task->tk_rqstp;
>> > 	struct rpc_xprt	*xprt = req->rq_xprt;
> 
>> > The rest LGTM.
>> > Reviewed-by: Petr Vorel <pvorel@suse.cz>
> 

Hi Trond, Anna, and Chuck,

Could you please let me know the next steps for this patch submission? 
Should I wait longer for review, or are there any formatting or 
submission issues that I need to fix?

I also noticed that the Patchwork event shows the patch as "Not Applicable":

https://patchwork.kernel.org/api/1.2/events/?patch=14757794

Could you clarify what that means for this submission?

Thanks,
Manjunath

>> > Kind regards,
>> > Petr
> 
>> Hi Petr,
> 
> Hi Manjunath,
> 
>> Thanks for the review and the Reviewed-by tag.
> 
> yw.
> 
>> I kept the assignments after RPC_SIGNALLED(task) intentionally, preserving
>> the existing behavior where a signalled task returns before we dereference
>> task->tk_rqstp. xprt is also needed only after xprt_adjust_timeout(req)
>> reports an expired timeout.
> 
> Understand, fair enough.
> 
>> I would prefer to retain the current ordering for those reasons. Also I am
>> open if maintainers chose to alter it.
> 
> +1
> 
> Kind regards,
> Petr
> 
>> Thanks,                   Manjunath
> 
> 
>> > >  	if (RPC_SIGNALLED(task))
>> > >  		return;
>> > > -	if (xprt_adjust_timeout(task->tk_rqstp) == 0)
>> > > +	req = task->tk_rqstp;
>> > > +	if (xprt_adjust_timeout(req) == 0)
>> > >  		return;
>> > > +	xprt = req->rq_xprt;
>> > >  	trace_rpc_timeout_status(task);
>> > >  	task->tk_timeouts++;
>> > > -	if (RPC_IS_SOFTCONN(task) && !rpc_check_connected(task->tk_rqstp)) {
>> > > +	if (RPC_IS_SOFTCONN(task) && !rpc_check_connected(req)) {
>> > >  		rpc_call_rpcerror(task, -ETIMEDOUT);
>> > >  		return;
>> > >  	}
>> > > @@ -2556,14 +2560,15 @@ rpc_check_timeout(struct rpc_task *task)
>> > >  		 * connection gets terminally broken.
>> > >  		 */
>> > >  		if ((task->tk_flags & RPC_TASK_NO_RETRANS_TIMEOUT) &&
>> > > -		    rpc_check_connected(task->tk_rqstp))
>> > > +		    rpc_check_connected(req))
>> > >  			return;
>> > >  		if (clnt->cl_chatty) {
>> > >  			pr_notice_ratelimited(
>> > > -				"%s: server %s not responding, timed out\n",
>> > > +				"%s: server %s clid=%u xprt=%u cc=%u owner=%d not responding, timed out\n",
>> > >  				clnt->cl_program->name,
>> > > -				task->tk_xprt->servername);
>> > > +				xprt->servername, clnt->cl_clid, xprt->id,
>> > > +				READ_ONCE(xprt->connect_cookie), task->tk_owner);
>> > >  		}
>> > >  		if (task->tk_flags & RPC_TASK_TIMEOUT)
>> > >  			rpc_call_rpcerror(task, -ETIMEDOUT);
>> > > @@ -2576,9 +2581,10 @@ rpc_check_timeout(struct rpc_task *task)
>> > >  		task->tk_flags |= RPC_CALL_MAJORSEEN;
>> > >  		if (clnt->cl_chatty) {
>> > >  			pr_notice_ratelimited(
>> > > -				"%s: server %s not responding, still trying\n",
>> > > +				"%s: server %s clid=%u xprt=%u cc=%u owner=%d not responding, still trying\n",
>> > >  				clnt->cl_program->name,
>> > > -				task->tk_xprt->servername);
>> > > +				xprt->servername, clnt->cl_clid, xprt->id,
>> > > +				READ_ONCE(xprt->connect_cookie), task->tk_owner);
>> > >  		}
>> > >  	}
>> > >  	rpc_force_rebind(clnt);
> 
> 


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH] SUNRPC: add transport details to timeout diagnostics
  2026-08-31 19:33       ` manjunath.b.patil
@ 2026-08-31 21:13         ` Chuck Lever
  2026-10-06 19:55         ` manjunath.b.patil
  1 sibling, 0 replies; 9+ messages in thread
From: Chuck Lever @ 2026-08-31 21:13 UTC (permalink / raw)
  To: manjunath.b.patil, Petr Vorel
  Cc: Dai Ngo, Anna Schumaker, Jeff Layton, linux-nfs, NeilBrown,
	Olga Kornievskaia, Tom Talpey, Trond Myklebust



On Mon, Aug 31, 2026, at 3:33 PM, manjunath.b.patil@oracle.com wrote:

> I also noticed that the Patchwork event shows the patch as "Not Applicable":
>
> https://patchwork.kernel.org/api/1.2/events/?patch=14757794
>
> Could you clarify what that means for this submission?

The kernel patchwork service applies only to NFSD. Your patch is for
the Linux NFS client, so it’s “Not Applicable” to NFSD. The NFS
client maintainers need to respond to your review request.

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH] SUNRPC: add transport details to timeout diagnostics
  2026-08-31 19:33       ` manjunath.b.patil
  2026-08-31 21:13         ` Chuck Lever
@ 2026-10-06 19:55         ` manjunath.b.patil
  1 sibling, 0 replies; 9+ messages in thread
From: manjunath.b.patil @ 2026-10-06 19:55 UTC (permalink / raw)
  To: Petr Vorel
  Cc: Dai.Ngo, anna, cel, jlayton, linux-nfs, neil, okorniev, tom,
	trondmy


On 8/31/26 12:33 PM, manjunath.b.patil@oracle.com wrote:
> On 8/20/26 11:59 AM, Petr Vorel wrote:
>>  > On 8/20/26 4: 13 AM, Petr Vorel wrote: > > > The existing RPC 
>> timeout messages identify only the RPC program and > > > server name. 
>> That makes it difficult to distinguish a server-side > > > response 
>> stall from
>> 
>>
>>> On 8/20/26 4:13 AM, Petr Vorel wrote:
>>> >  > The existing RPC timeout messages identify only the RPC program 
>>> and >
>>> > server name. That makes it difficult to distinguish a server-side >
>>> > response stall from a transport that is disconnected, reconnecting, 
>>> > or
>>> > using a different
>>
>>
>>> > > The existing RPC timeout messages identify only the RPC program and
>>> > > server name. That makes it difficult to distinguish a server-side
>>> > > response stall from a transport that is disconnected, reconnecting,
>>> > > or using a different connection.
>>
>>> > > Include the RPC client id, transport id, current transport 
>>> connection
>>> > > cookie, and RPC task owner in the ratelimited timeout messages. 
>>> These
>>> > > fields provide a compact way to correlate a timeout message with
>>> > > rpcctl and other transport state.
>>
>>> > > For example, an NFS timeout reports:
>>
>>> > >   nfs: server 192.168.50.1 clid=4 xprt=1 cc=1 owner=6245
>>> > >   not responding, timed out
>>
>>> > Nice improvement, thanks!
>>
>>
>>> > > Assisted-by: Codex:gpt-5
>>> > > Signed-off-by: Manjunath Patil <manjunath.b.patil@oracle.com>
>>> > > ---
>>> > >  net/sunrpc/clnt.c | 20 +++++++++++++-------
>>> > >  1 file changed, 13 insertions(+), 7 deletions(-)
>>
>>> > > diff --git a/net/sunrpc/clnt.c b/net/sunrpc/clnt.c
>>> > > index efa26899bc7d..6b5e2787dc81 100644
>>> > > --- a/net/sunrpc/clnt.c
>>> > > +++ b/net/sunrpc/clnt.c
>>> > > @@ -2534,17 +2534,21 @@ static void
>>> > >  rpc_check_timeout(struct rpc_task *task)
>>> > >  {
>>> > >      struct rpc_clnt    *clnt = task->tk_client;
>>> > > +    struct rpc_rqst    *req;
>>> > > +    struct rpc_xprt    *xprt;
>>
>>> > nit: Would it work to assign req and task earlier (just readability?)
>>> > Or is it not safe before checking RPC_SIGNALLED(task)?
>>
>>> >     struct rpc_clnt    *clnt = task->tk_client;
>>> >     struct rpc_rqst    *req = task->tk_rqstp;
>>> >     struct rpc_xprt    *xprt = req->rq_xprt;
>>
>>> > The rest LGTM.
>>> > Reviewed-by: Petr Vorel <pvorel@suse.cz>
>>
> 
> Hi Trond, Anna, and Chuck,
> 
> Could you please let me know the next steps for this patch submission? 
> Should I wait longer for review, or are there any formatting or 
> submission issues that I need to fix?
> 
> I also noticed that the Patchwork event shows the patch as "Not 
> Applicable":
> 
> https://patchwork.kernel.org/api/1.2/events/?patch=14757794
> 
> Could you clarify what that means for this submission?
> 
> Thanks,
> Manjunath

Hi Trond and Anna,

I am following up on this SUNRPC client patch. Petr Vorel reviewed it, 
and his question about the placement of the request and transport 
assignments was resolved in the thread.

Chuck clarified that the Patchwork "Not Applicable" state reflects NFSD 
tracking and does not indicate a decision on this NFS client patch.

Could you let me know whether this patch is suitable for the NFS client 
queue, or if you would like any changes or additional testing? I am 
happy to send a revised version if needed.

Thanks,
Manjunath
> 
>>> > Kind regards,
>>> > Petr
>>
>>> Hi Petr,
>>
>> Hi Manjunath,
>>
>>> Thanks for the review and the Reviewed-by tag.
>>
>> yw.
>>
>>> I kept the assignments after RPC_SIGNALLED(task) intentionally, 
>>> preserving
>>> the existing behavior where a signalled task returns before we 
>>> dereference
>>> task->tk_rqstp. xprt is also needed only after xprt_adjust_timeout(req)
>>> reports an expired timeout.
>>
>> Understand, fair enough.
>>
>>> I would prefer to retain the current ordering for those reasons. Also 
>>> I am
>>> open if maintainers chose to alter it.
>>
>> +1
>>
>> Kind regards,
>> Petr
>>
>>> Thanks,                   Manjunath
>>
>>
>>> > >      if (RPC_SIGNALLED(task))
>>> > >          return;
>>> > > -    if (xprt_adjust_timeout(task->tk_rqstp) == 0)
>>> > > +    req = task->tk_rqstp;
>>> > > +    if (xprt_adjust_timeout(req) == 0)
>>> > >          return;
>>> > > +    xprt = req->rq_xprt;
>>> > >      trace_rpc_timeout_status(task);
>>> > >      task->tk_timeouts++;
>>> > > -    if (RPC_IS_SOFTCONN(task) && !rpc_check_connected(task- 
>>> >tk_rqstp)) {
>>> > > +    if (RPC_IS_SOFTCONN(task) && !rpc_check_connected(req)) {
>>> > >          rpc_call_rpcerror(task, -ETIMEDOUT);
>>> > >          return;
>>> > >      }
>>> > > @@ -2556,14 +2560,15 @@ rpc_check_timeout(struct rpc_task *task)
>>> > >           * connection gets terminally broken.
>>> > >           */
>>> > >          if ((task->tk_flags & RPC_TASK_NO_RETRANS_TIMEOUT) &&
>>> > > -            rpc_check_connected(task->tk_rqstp))
>>> > > +            rpc_check_connected(req))
>>> > >              return;
>>> > >          if (clnt->cl_chatty) {
>>> > >              pr_notice_ratelimited(
>>> > > -                "%s: server %s not responding, timed out\n",
>>> > > +                "%s: server %s clid=%u xprt=%u cc=%u owner=%d 
>>> not responding, timed out\n",
>>> > >                  clnt->cl_program->name,
>>> > > -                task->tk_xprt->servername);
>>> > > +                xprt->servername, clnt->cl_clid, xprt->id,
>>> > > +                READ_ONCE(xprt->connect_cookie), task->tk_owner);
>>> > >          }
>>> > >          if (task->tk_flags & RPC_TASK_TIMEOUT)
>>> > >              rpc_call_rpcerror(task, -ETIMEDOUT);
>>> > > @@ -2576,9 +2581,10 @@ rpc_check_timeout(struct rpc_task *task)
>>> > >          task->tk_flags |= RPC_CALL_MAJORSEEN;
>>> > >          if (clnt->cl_chatty) {
>>> > >              pr_notice_ratelimited(
>>> > > -                "%s: server %s not responding, still trying\n",
>>> > > +                "%s: server %s clid=%u xprt=%u cc=%u owner=%d 
>>> not responding, still trying\n",
>>> > >                  clnt->cl_program->name,
>>> > > -                task->tk_xprt->servername);
>>> > > +                xprt->servername, clnt->cl_clid, xprt->id,
>>> > > +                READ_ONCE(xprt->connect_cookie), task->tk_owner);
>>> > >          }
>>> > >      }
>>> > >      rpc_force_rebind(clnt);
>>
>>
> 


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH] SUNRPC: add transport details to timeout diagnostics
  2026-08-19 17:30 [PATCH] SUNRPC: add transport details to timeout diagnostics Manjunath Patil
  2026-08-20 11:13 ` Petr Vorel
@ 2026-10-06 21:23 ` Anna Schumaker
  2026-10-07  1:03   ` manjunath.b.patil
  1 sibling, 1 reply; 9+ messages in thread
From: Anna Schumaker @ 2026-10-06 21:23 UTC (permalink / raw)
  To: Manjunath Patil, Trond Myklebust
  Cc: linux-nfs, Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia,
	Dai Ngo, Tom Talpey

Hi Manjunath,

On Wed, Aug 19, 2026, at 1:30 PM, Manjunath Patil wrote:
> The existing RPC timeout messages identify only the RPC program and
> server name. That makes it difficult to distinguish a server-side
> response stall from a transport that is disconnected, reconnecting,
> or using a different connection.
>
> Include the RPC client id, transport id, current transport connection
> cookie, and RPC task owner in the ratelimited timeout messages. These
> fields provide a compact way to correlate a timeout message with
> rpcctl and other transport state.
>
> For example, an NFS timeout reports:
>
>   nfs: server 192.168.50.1 clid=4 xprt=1 cc=1 owner=6245
>   not responding, timed out
>
> Assisted-by: Codex:gpt-5
> Signed-off-by: Manjunath Patil <manjunath.b.patil@oracle.com>
> ---
>  net/sunrpc/clnt.c | 20 +++++++++++++-------
>  1 file changed, 13 insertions(+), 7 deletions(-)
>
> diff --git a/net/sunrpc/clnt.c b/net/sunrpc/clnt.c
> index efa26899bc7d..6b5e2787dc81 100644
> --- a/net/sunrpc/clnt.c
> +++ b/net/sunrpc/clnt.c
> @@ -2534,17 +2534,21 @@ static void
>  rpc_check_timeout(struct rpc_task *task)
>  {
>  	struct rpc_clnt	*clnt = task->tk_client;
> +	struct rpc_rqst	*req;
> +	struct rpc_xprt	*xprt;
> 
>  	if (RPC_SIGNALLED(task))
>  		return;
> 
> -	if (xprt_adjust_timeout(task->tk_rqstp) == 0)
> +	req = task->tk_rqstp;
> +	if (xprt_adjust_timeout(req) == 0)
>  		return;
> +	xprt = req->rq_xprt;
> 
>  	trace_rpc_timeout_status(task);
>  	task->tk_timeouts++;
> 
> -	if (RPC_IS_SOFTCONN(task) && !rpc_check_connected(task->tk_rqstp)) {
> +	if (RPC_IS_SOFTCONN(task) && !rpc_check_connected(req)) {
>  		rpc_call_rpcerror(task, -ETIMEDOUT);
>  		return;
>  	}
> @@ -2556,14 +2560,15 @@ rpc_check_timeout(struct rpc_task *task)
>  		 * connection gets terminally broken.
>  		 */
>  		if ((task->tk_flags & RPC_TASK_NO_RETRANS_TIMEOUT) &&
> -		    rpc_check_connected(task->tk_rqstp))
> +		    rpc_check_connected(req))
>  			return;
> 
>  		if (clnt->cl_chatty) {
>  			pr_notice_ratelimited(
> -				"%s: server %s not responding, timed out\n",
> +				"%s: server %s clid=%u xprt=%u cc=%u owner=%d not responding, 
> timed out\n",
>  				clnt->cl_program->name,
> -				task->tk_xprt->servername);
> +				xprt->servername, clnt->cl_clid, xprt->id,
> +				READ_ONCE(xprt->connect_cookie), task->tk_owner);
>  		}
>  		if (task->tk_flags & RPC_TASK_TIMEOUT)
>  			rpc_call_rpcerror(task, -ETIMEDOUT);
> @@ -2576,9 +2581,10 @@ rpc_check_timeout(struct rpc_task *task)
>  		task->tk_flags |= RPC_CALL_MAJORSEEN;
>  		if (clnt->cl_chatty) {
>  			pr_notice_ratelimited(
> -				"%s: server %s not responding, still trying\n",
> +				"%s: server %s clid=%u xprt=%u cc=%u owner=%d not responding, 
> still trying\n",

More observability around this makes sense. Any reason you're adding to the notice
here and not making it a new tracepoint for the extra information? I'm worried
people might have tooling around watching for unresponsive servers and this
might break whatever regexes they're using to look for this.

Anna

>  				clnt->cl_program->name,
> -				task->tk_xprt->servername);
> +				xprt->servername, clnt->cl_clid, xprt->id,
> +				READ_ONCE(xprt->connect_cookie), task->tk_owner);
>  		}
>  	}
>  	rpc_force_rebind(clnt);
> -- 
> 2.47.3

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [PATCH] SUNRPC: add transport details to timeout diagnostics
  2026-10-06 21:23 ` Anna Schumaker
@ 2026-10-07  1:03   ` manjunath.b.patil
  0 siblings, 0 replies; 9+ messages in thread
From: manjunath.b.patil @ 2026-10-07  1:03 UTC (permalink / raw)
  To: Anna Schumaker, Trond Myklebust
  Cc: linux-nfs, Chuck Lever, Jeff Layton, NeilBrown, Olga Kornievskaia,
	Dai Ngo, Tom Talpey

On 10/6/26 2:23 PM, Anna Schumaker wrote:
> Hi Manjunath,
> 
> On Wed, Aug 19, 2026, at 1:30 PM, Manjunath Patil wrote:
>> The existing RPC timeout messages identify only the RPC program and
>> server name. That makes it difficult to distinguish a server-side
>> response stall from a transport that is disconnected, reconnecting,
>> or using a different connection.
>>
>> Include the RPC client id, transport id, current transport connection
>> cookie, and RPC task owner in the ratelimited timeout messages. These
>> fields provide a compact way to correlate a timeout message with
>> rpcctl and other transport state.
>>
>> For example, an NFS timeout reports:
>>
>>    nfs: server 192.168.50.1 clid=4 xprt=1 cc=1 owner=6245
>>    not responding, timed out
>>
>> Assisted-by: Codex:gpt-5
>> Signed-off-by: Manjunath Patil <manjunath.b.patil@oracle.com>
>> ---
>>   net/sunrpc/clnt.c | 20 +++++++++++++-------
>>   1 file changed, 13 insertions(+), 7 deletions(-)
>>
>> diff --git a/net/sunrpc/clnt.c b/net/sunrpc/clnt.c
>> index efa26899bc7d..6b5e2787dc81 100644
>> --- a/net/sunrpc/clnt.c
>> +++ b/net/sunrpc/clnt.c
>> @@ -2534,17 +2534,21 @@ static void
>>   rpc_check_timeout(struct rpc_task *task)
>>   {
>>   	struct rpc_clnt	*clnt = task->tk_client;
>> +	struct rpc_rqst	*req;
>> +	struct rpc_xprt	*xprt;
>>
>>   	if (RPC_SIGNALLED(task))
>>   		return;
>>
>> -	if (xprt_adjust_timeout(task->tk_rqstp) == 0)
>> +	req = task->tk_rqstp;
>> +	if (xprt_adjust_timeout(req) == 0)
>>   		return;
>> +	xprt = req->rq_xprt;
>>
>>   	trace_rpc_timeout_status(task);
>>   	task->tk_timeouts++;
>>
>> -	if (RPC_IS_SOFTCONN(task) && !rpc_check_connected(task->tk_rqstp)) {
>> +	if (RPC_IS_SOFTCONN(task) && !rpc_check_connected(req)) {
>>   		rpc_call_rpcerror(task, -ETIMEDOUT);
>>   		return;
>>   	}
>> @@ -2556,14 +2560,15 @@ rpc_check_timeout(struct rpc_task *task)
>>   		 * connection gets terminally broken.
>>   		 */
>>   		if ((task->tk_flags & RPC_TASK_NO_RETRANS_TIMEOUT) &&
>> -		    rpc_check_connected(task->tk_rqstp))
>> +		    rpc_check_connected(req))
>>   			return;
>>
>>   		if (clnt->cl_chatty) {
>>   			pr_notice_ratelimited(
>> -				"%s: server %s not responding, timed out\n",
>> +				"%s: server %s clid=%u xprt=%u cc=%u owner=%d not responding,
>> timed out\n",
>>   				clnt->cl_program->name,
>> -				task->tk_xprt->servername);
>> +				xprt->servername, clnt->cl_clid, xprt->id,
>> +				READ_ONCE(xprt->connect_cookie), task->tk_owner);
>>   		}
>>   		if (task->tk_flags & RPC_TASK_TIMEOUT)
>>   			rpc_call_rpcerror(task, -ETIMEDOUT);
>> @@ -2576,9 +2581,10 @@ rpc_check_timeout(struct rpc_task *task)
>>   		task->tk_flags |= RPC_CALL_MAJORSEEN;
>>   		if (clnt->cl_chatty) {
>>   			pr_notice_ratelimited(
>> -				"%s: server %s not responding, still trying\n",
>> +				"%s: server %s clid=%u xprt=%u cc=%u owner=%d not responding,
>> still trying\n",
> 
> More observability around this makes sense. Any reason you're adding to the notice
> here and not making it a new tracepoint for the extra information? I'm worried
> people might have tooling around watching for unresponsive servers and this
> might break whatever regexes they're using to look for this.
> 
> Anna
Hi Anna,

We added these identifiers to the ratelimited notice so they would be 
available in ordinary production logs after an intermittent NFS outage. 
A tracepoint would help only if tracing had been enabled and retained 
beforehand. In our investigation, the existing notice identified the 
server but not which RPC client, transport, or connection generation was 
involved.

I understand the concern about log parsers. The patch retains the 
notice's opening and closing phrases, though some parsers may still 
depend on its exact format. I considered a separate log line, but that 
would add noise to production logs.

Thanks,
Manjunath
> 
>>   				clnt->cl_program->name,
>> -				task->tk_xprt->servername);
>> +				xprt->servername, clnt->cl_clid, xprt->id,
>> +				READ_ONCE(xprt->connect_cookie), task->tk_owner);
>>   		}
>>   	}
>>   	rpc_force_rebind(clnt);
>> -- 
>> 2.47.3


^ permalink raw reply	[flat|nested] 9+ messages in thread

end of thread, other threads:[~2026-10-07  1:04 UTC | newest]

Thread overview: 9+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-19 17:30 [PATCH] SUNRPC: add transport details to timeout diagnostics Manjunath Patil
2026-08-20 11:13 ` Petr Vorel
2026-08-20 16:15   ` manjunath.b.patil
2026-08-20 18:59     ` Petr Vorel
2026-08-31 19:33       ` manjunath.b.patil
2026-08-31 21:13         ` Chuck Lever
2026-10-06 19:55         ` manjunath.b.patil
2026-10-06 21:23 ` Anna Schumaker
2026-10-07  1:03   ` manjunath.b.patil

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.