From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pf1-f202.google.com (mail-pf1-f202.google.com [209.85.210.202]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3CEDD1E8837 for ; Thu, 24 Apr 2025 20:02:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.202 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1745524946; cv=none; b=KU10LKFBsUSSjJ3y/JDt67rIgfKE8+eYELTPM7O9NA1UTfqq5yU7QR7W/qO/1B8B5EiicbXtEfWJM9dxVqQPV4pfUXS6Ayw0exkDqsS+larFUjk0NwbQ1kKeoIEXjZpsiCZteooaLBhFLmbeQVAEXrg/1j2+nc2BVgZO+qY+XTs= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1745524946; c=relaxed/simple; bh=/fgtOE5ynXviqm9MKFV+uE789F+qbv9CXSQIZmSPVkM=; h=Date:Mime-Version:Message-ID:Subject:From:To:Cc:Content-Type; b=CYiTPRVhnBnNpsVgZszN3ssvVHmdK9Jqxo5FCGIlQbGlGP9gSAZ1IGor3y7tviPolbVYs6NREZRv4AgbatdkateeZUHTYvh5v3EYZJxYZfMLsKnH13PKQaqIdkGeHxQe5pfJu23zVp9Da7qmIQkszVX3qgBqaizsJXdnL8sK65Q= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=UUbCp2aq; arc=none smtp.client-ip=209.85.210.202 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--skhawaja.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="UUbCp2aq" Received: by mail-pf1-f202.google.com with SMTP id d2e1a72fcca58-7369c5ed395so1493659b3a.0 for ; Thu, 24 Apr 2025 13:02:23 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20230601; t=1745524943; x=1746129743; darn=vger.kernel.org; h=cc:to:from:subject:message-id:mime-version:date:from:to:cc:subject :date:message-id:reply-to; bh=u/j6qDUCSodRxM7o32g4JnZg2Q8JLOAEvPQp+KmuUzE=; b=UUbCp2aqFBWfE9PEnfip9YwOLhnBeaNRKSup53c6+O8nYyib7CCAPL3cuxuDsLYJ3E eEGusBwQYBeMQTK38o0AJAktn1ZsRl/u52C/jillMfUNmQCbReui5EOeaVY66QnG9Vte udciziCc6DFaBssfpT78/RsN48P1xt1eC2HPxJhrolO5T7pNxBQhSck/gfDQAGeJ65w4 1eoEkVice12Rh0sEruxaFidqnsTyz3vkCiG2yTB+dUyBmbhNJ2GWUKqCvylqDLVlj0+U hzhXrYt20R/Krs/WCyAP0+NSmWMcybO9oCphSp+TbyEFT43RHm8M5h7zvvlUAs8yDtap JC2w== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1745524943; x=1746129743; h=cc:to:from:subject:message-id:mime-version:date:x-gm-message-state :from:to:cc:subject:date:message-id:reply-to; bh=u/j6qDUCSodRxM7o32g4JnZg2Q8JLOAEvPQp+KmuUzE=; b=XrJHQ+X9oBF3H0w2qIOewoVcDSfdjaqGUaCfMAXZAnOwUxAbaufcc+1T+UN3h4M1x5 4cHFD5aL3Kbph+z/ewQSAry1xN5iydmZAq/Cvv00Ipe7XkayQkOa46JB6H9sIrFB1irF gk/r1X5OJp5SHdVbM0Bbk7IWGHBvdYNxJE7EKBoEv7I6KTG5TY/DSVy9LYhCujPvqFrG haCfGIYPtzowVmcV82NVQiV15jRFObrc5uJl5kEnmvB66oz3bDx/UW4XBWwxm/eCtFN8 76IFMD2z/jNRwYR+ihYxF95LjrCIfdtWAM7PnBuX//bmqaaxx5bPuEArLAu1UrD0VRK2 IhgA== X-Gm-Message-State: AOJu0YwiI0SMFSWm3Gx854bNWnglyLjuz/nFsDV0OxySz70bgahaRCYW AATtWXqAzomgphWESMLZLZHNs5wgeoxQYcXDhGTJnKopITIvTQrXd1Ic9giF2wgv6i2QKMNa/I0 N5b0po7RO/A== X-Google-Smtp-Source: AGHT+IEy408+YmvpYelHsgH4GClYrT1U52HgU2/3xTZSCvsar6MzPMZ0pjrJ/sXSZL3VsRXker78fHKabRSHyA== X-Received: from pfna17.prod.google.com ([2002:aa7:80d1:0:b0:73d:65cb:b18b]) (user=skhawaja job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a00:e16:b0:73e:598:7e5b with SMTP id d2e1a72fcca58-73e32fbb020mr1190930b3a.1.1745524943471; Thu, 24 Apr 2025 13:02:23 -0700 (PDT) Date: Thu, 24 Apr 2025 20:02:18 +0000 Precedence: bulk X-Mailing-List: netdev@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 X-Mailer: git-send-email 2.49.0.850.g28803427d3-goog Message-ID: <20250424200222.2602990-1-skhawaja@google.com> Subject: [PATCH net-next v5 0/4] Add support to do threaded napi busy poll From: Samiullah Khawaja To: Jakub Kicinski , "David S . Miller " , Eric Dumazet , Paolo Abeni , almasrymina@google.com, willemb@google.com, jdamato@fastly.com, mkarsten@uwaterloo.ca Cc: netdev@vger.kernel.org, skhawaja@google.com Content-Type: text/plain; charset="UTF-8" Extend the already existing support of threaded napi poll to do continuous busy polling. This is used for doing continuous polling of napi to fetch descriptors from backing RX/TX queues for low latency applications. Allow enabling of threaded busypoll using netlink so this can be enabled on a set of dedicated napis for low latency applications. Once enabled user can fetch the PID of the kthread doing NAPI polling and set affinity, priority and scheduler for it depending on the low-latency requirements. Currently threaded napi is only enabled at device level using sysfs. Add support to enable/disable threaded mode for a napi individually. This can be done using the netlink interface. Extend `napi-set` op in netlink spec that allows setting the `threaded` attribute of a napi. Extend the threaded attribute in napi struct to add an option to enable continuous busy polling. Extend the netlink and sysfs interface to allow enabling/disabling threaded busypolling at device or individual napi level. We use this for our AF_XDP based hard low-latency usecase with usecs level latency requirement. For our usecase we want low jitter and stable latency at P99. Following is an analysis and comparison of available (and compatible) busy poll interfaces for a low latency usecase with stable P99. Please note that the throughput and cpu efficiency is a non-goal. For analysis we use an AF_XDP based benchmarking tool `xdp_rr`. The description of the tool and how it tries to simulate the real workload is following, - It sends UDP packets between 2 machines. - The client machine sends packets at a fixed frequency. To maintain the frequency of the packet being sent, we use open-loop sampling. That is the packets are sent in a separate thread. - The server replies to the packet inline by reading the pkt from the recv ring and replies using the tx ring. - To simulate the application processing time, we use a configurable delay in usecs on the client side after a reply is received from the server. The xdp_rr tool is posted separately as an RFC for tools/testing/selftest. We use this tool with following napi polling configurations, - Interrupts only - SO_BUSYPOLL (inline in the same thread where the client receives the packet). - SO_BUSYPOLL (separate thread and separate core) - Threaded NAPI busypoll System is configured using following script in all 4 cases, ``` echo 0 | sudo tee /sys/class/net/eth0/threaded echo 0 | sudo tee /proc/sys/kernel/timer_migration echo off | sudo tee /sys/devices/system/cpu/smt/control sudo ethtool -L eth0 rx 1 tx 1 sudo ethtool -G eth0 rx 1024 echo 0 | sudo tee /proc/sys/net/core/rps_sock_flow_entries echo 0 | sudo tee /sys/class/net/eth0/queues/rx-0/rps_cpus # pin IRQs on CPU 2 IRQS="$(gawk '/eth0-(TxRx-)?1/ {match($1, /([0-9]+)/, arr); \ print arr[0]}' < /proc/interrupts)" for irq in "${IRQS}"; \ do echo 2 | sudo tee /proc/irq/$irq/smp_affinity_list; done echo -1 | sudo tee /proc/sys/kernel/sched_rt_runtime_us for i in /sys/devices/virtual/workqueue/*/cpumask; \ do echo $i; echo 1,2,3,4,5,6 > $i; done if [[ -z "$1" ]]; then echo 400 | sudo tee /proc/sys/net/core/busy_read echo 100 | sudo tee /sys/class/net/eth0/napi_defer_hard_irqs echo 15000 | sudo tee /sys/class/net/eth0/gro_flush_timeout fi sudo ethtool -C eth0 adaptive-rx off adaptive-tx off rx-usecs 0 tx-usecs 0 if [[ "$1" == "enable_threaded" ]]; then echo 0 | sudo tee /proc/sys/net/core/busy_poll echo 0 | sudo tee /proc/sys/net/core/busy_read echo 100 | sudo tee /sys/class/net/eth0/napi_defer_hard_irqs echo 15000 | sudo tee /sys/class/net/eth0/gro_flush_timeout echo 2 | sudo tee /sys/class/net/eth0/threaded NAPI_T=$(ps -ef | grep napi | grep -v grep | awk '{ print $2 }') sudo chrt -f -p 50 $NAPI_T # pin threaded poll thread to CPU 2 sudo taskset -pc 2 $NAPI_T fi if [[ "$1" == "enable_interrupt" ]]; then echo 0 | sudo tee /proc/sys/net/core/busy_read echo 0 | sudo tee /sys/class/net/eth0/napi_defer_hard_irqs echo 15000 | sudo tee /sys/class/net/eth0/gro_flush_timeout fi ``` To enable various configurations, script can be run as following, - Interrupt Only ```