From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wm1-f41.google.com (mail-wm1-f41.google.com [209.85.128.41]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 744E437F318 for ; Tue, 25 Aug 2026 16:33:26 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.41 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787675609; cv=none; b=m49GnMyC2/MK9+GsWmTEdVqIl/4y1cBcjp9Fuki485FQmZ71MWCQsz5in5T333/1X+4lU3TRT8BCmbotSneHLmUb0C3wqTvTY6e8uIQSLF08nRtV/FcXYJKRBYBpBCy4Ut3BOWR3vdobyG4iL58YOM47wz7tTKkt3gj8HG7wGvw= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787675609; c=relaxed/simple; bh=B0G9v6BhBY4Fuj6Zio2PzdtQEPK77b2tYVyr2I4VYEA=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=O9JeBYjrCfHZjq3En+xjZDirVwUNPoHnAQldySYEgFVod9RDKZDfkKDiaA3bIedzitKWK0HSgTJc9xphUPnxS3+8jFl/X/ERRsGaAAg9YR7jjIJW8Q5EHb5eaMa813/0PBUdKIteexeImfOyuzlk71IKimzPWDKO8Yhdahfx0Xk= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=KDd/qEWG; arc=none smtp.client-ip=209.85.128.41 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="KDd/qEWG" Received: by mail-wm1-f41.google.com with SMTP id 5b1f17b1804b1-499add6348fso51455e9.1 for ; Tue, 25 Aug 2026 09:33:25 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1787675604; x=1788280404; darn=vger.kernel.org; h=in-reply-to:content-transfer-encoding:content-disposition :content-type:mime-version:references:message-id:subject:cc:to:from :date:from:to:cc:subject:date:message-id:reply-to:content-type; bh=SkuHlshzOD+AxFcfYa+qONyS5a/GbEXwYVTwh7Zfzc8=; b=KDd/qEWGhDvEPOieRLxzIHmbNPoUWC80DsbNQlSfeQu6Wkz+/ObtT0FCEZfS1lgS8r /PpeFW7PdWH09J7u5QWeNykrtvTsW9LM87ldGaBCS0D2eQ/EFiAstWLYcstO8UDzivA1 yKBK477wqZcKJm8REhL4M5LHbNDU8dixqklHRFh0RxeaBn6NTE5dBH3B8hYcQpuI5GXR DuqUk91ET3nuqpmVB7cy3oiXQZjTxTV4PesQru2i+4sAtwdc+QrO6JMMhfmAlXPoZPri Lp0UrS113aniefYlxl+vsbIeQLx5Xb4GVVhcmyEWPJ4xZuY7Zp8dSFNpGJbjxia3XuMu t2HA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787675604; x=1788280404; h=in-reply-to:content-transfer-encoding:content-disposition :content-type:mime-version:references:message-id:subject:cc:to:from :date:x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to:content-type; bh=SkuHlshzOD+AxFcfYa+qONyS5a/GbEXwYVTwh7Zfzc8=; b=YWhojMjd8REruiz4oND/ekrtoy3VIN9f72Thxu/eNuBs84ySr8GYjuS2sjnxNiR7Sk BovgU6iGPWRFThKGRlXp9tUsefHBBv3l4XAbqWLDCWDk5vEqdixuGFa25W0gMLETuj9X rRcCSLnFoXHq9DKj52qSjJV/ZCY4ApU2OIEloDOAwq2UzGD0KEfbBXIBdYPIfVv33yCa 4FskELFYJ/OFmIahwUXlJh7fWaPVCgBi7scvwHJbQ4pzVzpoyakIZNjNfrMegH52Gnc1 BE/JEOUsxa40gOcJfd/YHC1Vq83e//lktLBj0jCfHTy4+sm1l8HtyXsY65ucMMEmtAw/ +LbQ== X-Forwarded-Encrypted: i=1; AHgh+Rq6w2t/OX9LD83SUT6WfC1KByHRsMUejM3MWFThNfnB4fu+h9Rdu862FQzfaXN8Z7wgHCWTF68J/RVyrn0=@vger.kernel.org X-Gm-Message-State: AFuF++kPNDDp72pH870VUh9MY4YuFh0xx3jCSsoJPQKCL6vaJMaUO3Xg q3y9nn125PUOiPRhvbyZptkajH8ZksRE91/eWSORczF83Qm0ci7wCuuL5HKV3bYidw== X-Gm-Gg: AR+sD12D/taiWENgrBSt12YJeG9U57cVtTzmth++wbCHGOB8QuDHWzrBEw6uc1UGVor HdMSbMOdUBdXcV2EknIZUNgSpt2GmsoFYXC4TOKhdSsVou+ZG2YOIxRLCgvzuzct7YSr3rifkUz ZXeqLKWbJcke2nDZEG3Eb+uzyrWzfYOt9OzG5SrV+2HR3dY0SwvQQDUP7esdgd+eqE1aOnW4ow8 2GBn1e79hK2afxTiSHPdbURDsTYHaET/cSsrta/HxYbgGnG0in8rE7iSG9zPKBIufMWAeDJtjLQ 9bwbKjv+7Sqd1RvqSyih58/s6hCuRYRrNNUaJ4uh1/cE2YIe+q0pCsdpews10MzLyzvIwd2pNJr vlvXkTTR3XulZbnYaREclYJqzqLcf37X82maymZtMU+sPJFHF9OnlLFnQRdZ4MkoybGdpVDhsxg H6u7tTEPQfvHQiZOfAaoBPm+LRPT6mOmXG6tdGBCmIeFtPQS2C6jxXOnB7uAaDs5pKyxGyCgXog fV5TfYOZedgG6u+TBsgBi+sxby7Dg== X-Received: by 2002:a05:600c:a31c:b0:499:7da3:dd0c with SMTP id 5b1f17b1804b1-499d867c424mr1003585e9.0.1787675603950; Tue, 25 Aug 2026 09:33:23 -0700 (PDT) Received: from google.com (250.192.189.35.bc.googleusercontent.com. [35.189.192.250]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-499dc960f24sm217945e9.3.2026.08.25.09.33.23 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 25 Aug 2026 09:33:23 -0700 (PDT) Date: Tue, 25 Aug 2026 16:33:19 +0000 From: Mostafa Saleh To: Dragos Tatulea Cc: Luigi Rizzo , Jakub Kicinski , rizzo.unipi@gmail.com, m.szyprowski@samsung.com, robin.murphy@arm.com, willemb@google.com, kuniyu@google.com, davem@davemloft.net, edumazet@google.com, pabeni@redhat.com, gregkh@linuxfoundation.org, rafael@kernel.org, akpm@linux-foundation.org, david@kernel.org, netdev@vger.kernel.org, linux-mm@kvack.org, iommu@lists.linux.dev, driver-core@lists.linux.dev, linux-kernel@vger.kernel.org Subject: Re: [PATCH] swiotlb: avoid double copy with swiotlb on tx socket Message-ID: References: <20260615234220.3946885-1-lrizzo@google.com> <20260615172535.080cf94f@kernel.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Disposition: inline Content-Transfer-Encoding: 8bit In-Reply-To: On Mon, Aug 24, 2026 at 10:59:21AM +0200, Dragos Tatulea wrote: > > > On 16.06.26 13:06, Mostafa Saleh wrote: > > On Tue, Jun 16, 2026 at 02:33:52AM +0200, Luigi Rizzo wrote: > >> On Tue, Jun 16, 2026 at 2:25 AM Jakub Kicinski wrote: > >>> > >>> On Mon, 15 Jun 2026 23:42:20 +0000 Luigi Rizzo wrote: > >>>> The use of swiotlb causes an extra data copy on I/O. For tx sockets, > >>>> especially with greedy senders, this has a high chance of happening in > >>>> the softirq handler for tx network interrupts, creating a significant > >>>> performance bottleneck. > >>> > >>> What's the use case? I associate swiotlb with debug / testing mostly, > >>> so it'd be useful for people like me to explain why you care. > >> > >> Ah sorry, I forgot to mention. > >> swiotlb is used in guest kernels for confidential computing VMs. > >> Ordinary memory pages are encrypted and the host or devices > >> have no way to decrypt them, so the kernel must use > >> unencrypted bounce buffers to exchange data with I/O devices. > > > > I started looking into the same problem recently, to reduce the > > bouncing in protected KVM (pKVM) confidential guests. > > My first attempt was to update dma_direct_map_phys() to skip > > bouncing and do inline memory decryption (for pKVM that is a hypercall > > which updates the stage-2 page tables), however, that was really slow > > compared to the memcpy in bouncing even for massive pages. > > My conclusion was similar that we need to solve this at construction > > by making this memory allocated from a pre-decrypted pool (which > > does not have to be part of the SWIOTLB) > > My initial idea was to teach some of the kernel subsystems (SKB, > > BLK, SLAB) about "CoCo allocators" that allocate decrypted memory, > > as this is not a net specific problem. > > > An example of this is Jiri's system_cc_shared heap which is a dma-buf > heap with decrypted memory for userspace. > > > I am still looking into this, I was planning to bring this up in the > > upcoming LPC. > > I will give this patch a try. However, I believe that we need a more > > generalised concept for CoCo pre-decrypted allocators in the kernel. > > > There is a talk at LPC in the networking track about this [2]. This is > exactly the type of discussion that I was hoping to have there. I see, thanks for point that. I plan to be in LPC, so I will aim to attend this talk. > > Besides the issues mentioned in this thread we've also found that a lot > of overhead can come only from swiotlb allocations when running many queues. > > I will add information about this series in my talk. Hopefully I will also > have time to add some numbers for comparison. I had a quick look, I am not sure how that will shape at the end, but I think we need a general solution beyond NICs. For example in pKVM the bouncing is used with virtio devices, so doing this per-driver won't really work and is not possible in some scenarios where memory is allocated from the core kernel and then passed to the driver. So, I was thinking the kernel relying on CC_ATTR_MEM_ENCRYPT/force_dma_unencrypted() could detect that and allocate pre-shared memory for those cases. Thanks, Mostafa > > Sorry for the late reply but I spotted this thread only now by > accident. > > [1] https://lore.kernel.org/all/20260325192352.437608-3-jiri@resnulli.us/ > [2] https://lpc.events/event/20/contributions/2464 > > Thanks, > Dragos