From mboxrd@z Thu Jan 1 00:00:00 1970 From: Eric Dumazet Subject: Re: [PATCH] reduce the spinlock conflict during massive connect Date: Sun, 05 Nov 2017 21:27:15 -0800 Message-ID: <1509946035.2849.79.camel@edumazet-glaptop3.roam.corp.google.com> References: <9b38c346-0035-4c12-21e7-dbde9f961c8a@gmail.com> Mime-Version: 1.0 Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: 7bit Cc: netdev@vger.kernel.org, "\"David S. Miller\" ;Alexey Kuznetsov " ";Hideaki YOSHIFUJI" To: Liu Yu Return-path: Received: from mail-io0-f195.google.com ([209.85.223.195]:44472 "EHLO mail-io0-f195.google.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1750760AbdKFF1U (ORCPT ); Mon, 6 Nov 2017 00:27:20 -0500 Received: by mail-io0-f195.google.com with SMTP id m16so14472180iod.1 for ; Sun, 05 Nov 2017 21:27:19 -0800 (PST) In-Reply-To: <9b38c346-0035-4c12-21e7-dbde9f961c8a@gmail.com> Sender: netdev-owner@vger.kernel.org List-ID: On Mon, 2017-11-06 at 10:28 +0800, Liu Yu wrote: > From: Liu Yu > > When a mount of processes connect to the same port at the same address > simultaneously, they are likely getting the same bhash and therefore > conflict with each other. > > The more the cpu number, the worse in this case. > > Use spin_trylock instead for this scene, which seems doesn't matter > for common case. > > Signed-off-by: Liu Yu > --- > net/ipv4/inet_hashtables.c | 6 +++++- > 1 files changed, 5 insertions(+), 1 deletions(-) > > diff --git a/net/ipv4/inet_hashtables.c b/net/ipv4/inet_hashtables.c > index e7d15fb..cc11ec7 100644 > --- a/net/ipv4/inet_hashtables.c > +++ b/net/ipv4/inet_hashtables.c > @@ -581,13 +581,17 @@ int __inet_hash_connect(struct inet_timewait_death_row *death_row, > other_parity_scan: > port = low + offset; > for (i = 0; i < remaining; i += 2, port += 2) { > + int ret; > + > if (unlikely(port >= high)) > port -= remaining; > if (inet_is_local_reserved_port(net, port)) > continue; > head = &hinfo->bhash[inet_bhashfn(net, port, > hinfo->bhash_size)]; > - spin_lock_bh(&head->lock); > + ret = spin_trylock(&head->lock); > + if (unlikely(!ret)) > + continue; > > /* Does not bother with rcv_saddr checks, because > * the established check is already unique enough. This is broken. I am pretty sure you have not really tested this patch properly. Chances are very high that a connect() will miss slots and wont succeed, when table is almost full. Performance is nice, but we actually need to allocate a 4-tuple in a more deterministic fashion.