From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org X-Spam-Level: X-Spam-Status: No, score=-10.6 required=3.0 tests=DKIMWL_WL_MED,DKIM_SIGNED, DKIM_VALID,DKIM_VALID_AU,HEADER_FROM_DIFFERENT_DOMAINS,MAILING_LIST_MULTI, SPF_PASS,USER_AGENT_GIT,USER_IN_DEF_DKIM_WL autolearn=ham autolearn_force=no version=3.4.0 Received: from mail.kernel.org (mail.kernel.org [198.145.29.99]) by smtp.lore.kernel.org (Postfix) with ESMTP id 1740BC43381 for ; Fri, 22 Mar 2019 15:56:47 +0000 (UTC) Received: from vger.kernel.org (vger.kernel.org [209.132.180.67]) by mail.kernel.org (Postfix) with ESMTP id D60BE21900 for ; Fri, 22 Mar 2019 15:56:46 +0000 (UTC) Authentication-Results: mail.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="BxtYmRzu" Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1727645AbfCVP4p (ORCPT ); Fri, 22 Mar 2019 11:56:45 -0400 Received: from mail-yw1-f74.google.com ([209.85.161.74]:39498 "EHLO mail-yw1-f74.google.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1727169AbfCVP4p (ORCPT ); Fri, 22 Mar 2019 11:56:45 -0400 Received: by mail-yw1-f74.google.com with SMTP id p1so3601464ywm.6 for ; Fri, 22 Mar 2019 08:56:45 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20161025; h=date:message-id:mime-version:subject:from:to:cc; bh=3lvv2k/4oTeGKTncksBh0sC3pido8/ISsDhaxgilIkM=; b=BxtYmRzuC2GU9EyV3kE+r9+yGbuhrjcHnMzqCbNk+yt40C+4zHD2R+ZG7/IGPvHiyl +MBePUMw1KOScrvynpsWnUspvzvFxEz2Oh4KDciPy+kWjUSFIQMPIpOlIPfYJxKuZLuj 5PbMzgbblLcbc8eKqn23nihNb9gWcir/os6vBsE1DKZRv2S4pjGMeKcoTe9aQmg3YoBt vgwhGwIoZLvUpi8NZvNsGrL+RCFr1bJQ43VRcODyVMmsjq9xBXo3Q7jyt7DGC4v65DyF JEA/aHBuH2qTyC3eKXGzwM1GJymP4acVfr08XZKqYvDyAiGecvPVNet6TdZVI7Hp59Uj 1XwQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20161025; h=x-gm-message-state:date:message-id:mime-version:subject:from:to:cc; bh=3lvv2k/4oTeGKTncksBh0sC3pido8/ISsDhaxgilIkM=; b=kJmr9DO3z1rk+Q1JAzYBeF5ocEguUwnYkRQWnQ1Wk0xxSv4U9RHigwAFL1gdAOq+c7 FUx0250Br+JiipOSKRplQfx3tHY0LsrE0IBZYV1hc9UWXKBgnImFOilUHPysnFUU6CHc 4sWchLFT85/+y3W8XZ7jt28O3kB+ofxb69Al16llfHhxw6lMX5lBNOs9BP2O5EavFDMj S7HsWHsoQ2HZCRa/c5xDrQDICQST2DlLaub7aALoBPHMpNw5BN8q8AHh2f6o55b4Kz5A aTYV12vLic54C5jxK7uRf0WrCYaz8aAhONkDOMM7Bewj+FNYNUm0FJXV6TcWNtFhFXqu CVDA== X-Gm-Message-State: APjAAAVw2eoQMGS+DGgtysFMJTK//KWakmqL7MihjT9yWD7LVtUf6/fG HElNI4ClAjK93MCL5F9h3pjECViREDopig== X-Google-Smtp-Source: APXvYqxK4Ww+ZXjAWT9SV6wr2qNnT1DdfSdrapWNSWTB/UVmWMu4I55gURbmuXEZkKlWzFpurGdDGHanDJ0Xeg== X-Received: by 2002:a81:5f06:: with SMTP id t6mr8754699ywb.433.1553270204563; Fri, 22 Mar 2019 08:56:44 -0700 (PDT) Date: Fri, 22 Mar 2019 08:56:37 -0700 Message-Id: <20190322155640.248144-1-edumazet@google.com> Mime-Version: 1.0 X-Mailer: git-send-email 2.21.0.392.gf8f6787159e-goog Subject: [PATCH v3 net-next 0/3] tcp: add rx/tx cache to reduce lock contention From: Eric Dumazet To: "David S . Miller" Cc: netdev , Eric Dumazet , Eric Dumazet Content-Type: text/plain; charset="UTF-8" Sender: netdev-owner@vger.kernel.org Precedence: bulk List-ID: X-Mailing-List: netdev@vger.kernel.org On hosts with many cpus we can observe a very serious contention on spinlocks used in mm slab layer. The following can happen quite often : 1) TX path sendmsg() allocates one (fclone) skb on CPU A, sends a clone. ACK is received on CPU B, and consumes the skb that was in the retransmit queue. 2) RX path network driver allocates skb on CPU C recvmsg() happens on CPU D, freeing the skb after it has been delivered to user space. In both cases, we are hitting the asymetric alloc/free pattern for which slab has to drain alien caches. At 8 Mpps per second, this represents 16 Mpps alloc/free per second and has a huge penalty. In an interesting experiment, I tried to use a single kmem_cache for all the skbs (in skb_init() : skbuff_fclone_cache = skbuff_head_cache = kmem_cache_create("skbuff_fclone_cache", sizeof(struct sk_buff_fclones),); qnd most of the contention disappeared, since cpus could better use their local slab per-cpu cache. But we can do actually better, in the following patches. TX : at ACK time, no longer free the skb but put it back in a tcp socket cache, so that next sendmsg() can reuse it immediately. RX : at recvmsg() time, do not free the skb but put it in a tcp socket cache so that it can be freed by the cpu feeding the incoming packets in BH. This increased the performance of small RPC benchmark by about 10 % on a host with 112 hyperthreads. v2 : - Solved a race condition : sk_stream_alloc_skb() to make sure the prior clone has been freed. - Really test rps_needed in sk_eat_skb() as claimed. - Fixed rps_needed use in drivers/net/tun.c v3: Added a #ifdef CONFIG_RPS, to avoid compile error (kbuild robot) Eric Dumazet (3): net: convert rps_needed and rfs_needed to new static branch api tcp: add one skb cache for tx tcp: add one skb cache for rx drivers/net/tun.c | 2 +- include/linux/netdevice.h | 4 +-- include/net/sock.h | 17 +++++++++++- net/core/dev.c | 10 +++---- net/core/net-sysfs.c | 4 +-- net/core/sysctl_net_core.c | 8 +++--- net/ipv4/af_inet.c | 4 +++ net/ipv4/tcp.c | 54 +++++++++++++++++++------------------- net/ipv4/tcp_ipv4.c | 11 ++++++-- net/ipv6/tcp_ipv6.c | 12 ++++++--- 10 files changed, 79 insertions(+), 47 deletions(-) -- 2.21.0.392.gf8f6787159e-goog