From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from bombadil.infradead.org (bombadil.infradead.org [198.137.202.133]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.lore.kernel.org (Postfix) with ESMTPS id 32FE0EE6428 for ; Wed, 31 Dec 2025 14:43:49 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=lists.infradead.org; s=bombadil.20210309; h=Sender:Cc:List-Subscribe: List-Help:List-Post:List-Archive:List-Unsubscribe:List-Id:References: Content-Type:Content-Transfer-Encoding:In-Reply-To:From:To:Subject: MIME-Version:Date:Message-ID:Reply-To:Content-ID:Content-Description: Resent-Date:Resent-From:Resent-Sender:Resent-To:Resent-Cc:Resent-Message-ID: List-Owner; bh=M8aPWywfPFOyTSGu0q6hhPXyHxVoa1F4hqfQwCmIl58=; b=anm62GnjCy+43U bu1D9SH3cOppXsm/0f8fjmEjoxWRFCrs+pQCcLra8CafWMtKUUguT7si6ETzbcPTG+myVtcF2bF06 vXDidAO83CSzvrfnROX2kI7B9XRKCvi3EL7Vr3yIYdIy21G2I+oaqgUFs12hjB9h35sz3BCpUZ/Z7 NtbHRtS0iuDuabnUQdCnV7uAEdfYaoGZKxVhyPMr9b4RfPZV97Uewa0rFbtrmCFDjY/iEVcOiUuUj IEqixt6Cd3ROe8BkwuFV8xzozIPYEQ2RTD/QCm4/Ey2onJ2rgNLSvjZUW0e1XjQYz8h1AYw4HJQfZ KzTEIqNN3iv4x3ikj8Ew==; Received: from localhost ([::1] helo=bombadil.infradead.org) by bombadil.infradead.org with esmtp (Exim 4.98.2 #2 (Red Hat Linux)) id 1vaxQK-0000000633h-2qB8; Wed, 31 Dec 2025 14:43:40 +0000 Received: from mailout1.w1.samsung.com ([210.118.77.11]) by bombadil.infradead.org with esmtps (Exim 4.98.2 #2 (Red Hat Linux)) id 1vaxQG-00000006336-2RFA for linux-arm-kernel@lists.infradead.org; Wed, 31 Dec 2025 14:43:38 +0000 Received: from eucas1p2.samsung.com (unknown [182.198.249.207]) by mailout1.w1.samsung.com (KnoxPortal) with ESMTP id 20251231144330euoutp01c58badeadda370d7c7c3cd9549da0f60~GU-Pi5Z5i3076130761euoutp013 for ; Wed, 31 Dec 2025 14:43:30 +0000 (GMT) DKIM-Filter: OpenDKIM Filter v2.11.0 mailout1.w1.samsung.com 20251231144330euoutp01c58badeadda370d7c7c3cd9549da0f60~GU-Pi5Z5i3076130761euoutp013 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=samsung.com; s=mail20170921; t=1767192210; bh=M8aPWywfPFOyTSGu0q6hhPXyHxVoa1F4hqfQwCmIl58=; h=Date:Subject:To:Cc:From:In-Reply-To:References:From; b=Alom80QxoG+L7ICMesd3Euki0jWEh5m9WrDQOO82KpT3M8GlUii6jb4NC/BUwrtla x3F2P1lkvxBq8MsCMGMLR2GRCZGVZ1OzbjGOF4McE9zgsSSTUJqP6FtdiWlME5ilT5 BOXZctdYokjY0BnnWLe9y+syt3f3BRIA9KcPecjg= Received: from eusmtip2.samsung.com (unknown [203.254.199.222]) by eucas1p2.samsung.com (KnoxPortal) with ESMTPA id 20251231144330eucas1p2e45e514c58572acda6dc94e809f9d2ee~GU-PNyoET1215412154eucas1p27; Wed, 31 Dec 2025 14:43:30 +0000 (GMT) Received: from [106.210.134.192] (unknown [106.210.134.192]) by eusmtip2.samsung.com (KnoxPortal) with ESMTPA id 20251231144321eusmtip23facba6ecea740b65bcbedd64e01dd62~GU-GtmYe61784117841eusmtip2B; Wed, 31 Dec 2025 14:43:21 +0000 (GMT) Message-ID: Date: Wed, 31 Dec 2025 15:43:19 +0100 MIME-Version: 1.0 User-Agent: Betterbird (Windows) Subject: Re: [PATCH v2 4/8] dma-mapping: Separate DMA sync issuing and completion waiting To: Barry Song <21cnbao@gmail.com>, Leon Romanovsky Content-Language: en-US From: Marek Szyprowski In-Reply-To: Content-Transfer-Encoding: 8bit X-CMS-MailID: 20251231144330eucas1p2e45e514c58572acda6dc94e809f9d2ee X-Msg-Generator: CA Content-Type: text/plain; charset="utf-8" X-RootMTR: 20251228213843eucas1p298280bb01abe59739cd6e0482f570455 X-EPHeader: CA X-CMS-RootMailID: 20251228213843eucas1p298280bb01abe59739cd6e0482f570455 References: <20251226225254.46197-1-21cnbao@gmail.com> <20251226225254.46197-5-21cnbao@gmail.com> <20251227200706.GN11869@unreal> <20251228144909.GR11869@unreal> X-CRM114-Version: 20100106-BlameMichelson ( TRE 0.8.0 (BSD) ) MR-646709E3 X-CRM114-CacheID: sfid-20251231_064337_192325_83F4BBCF X-CRM114-Status: GOOD ( 35.43 ) X-BeenThere: linux-arm-kernel@lists.infradead.org X-Mailman-Version: 2.1.34 Precedence: list List-Id: List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Cc: Juergen Gross , Tangquan Zheng , Stefano Stabellini , Ryan Roberts , will@kernel.org, Anshuman Khandual , catalin.marinas@arm.com, Joerg Roedel , linux-kernel@vger.kernel.org, Suren Baghdasaryan , iommu@lists.linux.dev, Marc Zyngier , Oleksandr Tyshchenko , xen-devel@lists.xenproject.org, robin.murphy@arm.com, Ard Biesheuvel , linux-arm-kernel@lists.infradead.org Sender: "linux-arm-kernel" Errors-To: linux-arm-kernel-bounces+linux-arm-kernel=archiver.kernel.org@lists.infradead.org On 28.12.2025 22:38, Barry Song wrote: > On Mon, Dec 29, 2025 at 3:49 AM Leon Romanovsky wrote: >> On Sun, Dec 28, 2025 at 10:45:13AM +1300, Barry Song wrote: >>> On Sun, Dec 28, 2025 at 9:07 AM Leon Romanovsky wrote: >>>> On Sat, Dec 27, 2025 at 11:52:44AM +1300, Barry Song wrote: >>>>> From: Barry Song >>>>> >>>>> Currently, arch_sync_dma_for_cpu and arch_sync_dma_for_device >>>>> always wait for the completion of each DMA buffer. That is, >>>>> issuing the DMA sync and waiting for completion is done in a >>>>> single API call. >>>>> >>>>> For scatter-gather lists with multiple entries, this means >>>>> issuing and waiting is repeated for each entry, which can hurt >>>>> performance. Architectures like ARM64 may be able to issue all >>>>> DMA sync operations for all entries first and then wait for >>>>> completion together. >>>>> >>>>> To address this, arch_sync_dma_for_* now issues DMA operations in >>>>> batch, followed by a flush. On ARM64, the flush is implemented >>>>> using a dsb instruction within arch_sync_dma_flush(). >>>>> >>>>> For now, add arch_sync_dma_flush() after each >>>>> arch_sync_dma_for_*() call. arch_sync_dma_flush() is defined as a >>>>> no-op on all architectures except arm64, so this patch does not >>>>> change existing behavior. Subsequent patches will introduce true >>>>> batching for SG DMA buffers. >>>>> >>>>> Cc: Leon Romanovsky >>>>> Cc: Catalin Marinas >>>>> Cc: Will Deacon >>>>> Cc: Marek Szyprowski >>>>> Cc: Robin Murphy >>>>> Cc: Ada Couprie Diaz >>>>> Cc: Ard Biesheuvel >>>>> Cc: Marc Zyngier >>>>> Cc: Anshuman Khandual >>>>> Cc: Ryan Roberts >>>>> Cc: Suren Baghdasaryan >>>>> Cc: Joerg Roedel >>>>> Cc: Juergen Gross >>>>> Cc: Stefano Stabellini >>>>> Cc: Oleksandr Tyshchenko >>>>> Cc: Tangquan Zheng >>>>> Signed-off-by: Barry Song >>>>> --- >>>>> arch/arm64/include/asm/cache.h | 6 ++++++ >>>>> arch/arm64/mm/dma-mapping.c | 4 ++-- >>>>> drivers/iommu/dma-iommu.c | 37 +++++++++++++++++++++++++--------- >>>>> drivers/xen/swiotlb-xen.c | 24 ++++++++++++++-------- >>>>> include/linux/dma-map-ops.h | 6 ++++++ >>>>> kernel/dma/direct.c | 8 ++++++-- >>>>> kernel/dma/direct.h | 9 +++++++-- >>>>> kernel/dma/swiotlb.c | 4 +++- >>>>> 8 files changed, 73 insertions(+), 25 deletions(-) >>>> <...> >>>> >>>>> +#ifndef arch_sync_dma_flush >>>>> +static inline void arch_sync_dma_flush(void) >>>>> +{ >>>>> +} >>>>> +#endif >>>> Over the weekend I realized a useful advantage of the ARCH_HAVE_* config >>>> options: they make it straightforward to inspect the entire DMA path simply >>>> by looking at the .config. >>> I am not quite sure how much this benefits users, as the same >>> information could also be obtained by grepping for >>> #define arch_sync_dma_flush in the source code. >> It differs slightly. Users no longer need to grep around or guess whether this >> platform used the arch_sync_dma_flush path. A simple grep for ARCH_HAVE_ in >> /proc/config.gz provides the answer. > In any case, it is only two or three lines of code, so I am fine with > either approach. Perhaps Marek, Robin, and others have a point here? If possible I would suggest to follow the already used style in the given code even if it means a bit larger patch. >>>> Thanks, >>>> Reviewed-by: Leon Romanovsky >>> Thanks very much, Leon, for reviewing this over the weekend. One thing >>> you might have missed is that I place arch_sync_dma_flush() after all >>> arch_sync_dma_for_*() calls, for both single and sg cases. I also >>> used a Python script to scan the code and verify that every >>> arch_sync_dma_for_*() is followed by arch_sync_dma_flush(), to ensure >>> that no call is left out. >>> >>> In the subsequent patches, for sg cases, the per-entry flush is >>> replaced by a single flush of the entire sg. Each sg case has >>> different characteristics: some are straightforward, while others >>> can be tricky and involve additional contexts. >> I didn't overlook it, and I understand your rationale. However, this is >> not how kernel patches should be structured. You should not introduce >> code in patch X and then move it elsewhere in patch X + Y. > I am not quite convinced by this concern. This patch only > separates DMA sync issuing from completion waiting, and it > reflects that the development is done step by step. > >> Place the code in the correct location from the start. Your patches are >> small enough to review as is. > My point is that this patch places the code in the correct locations > from the start. It splits arch_sync_dma_for_*() into > arch_sync_dma_for_*() plus arch_sync_dma_flush() everywhere, without > introducing any functional changes from the outset. > The subsequent patches clearly show which parts are truly batched. > > In the meantime, I do not have a strong preference here. If you think > it is better to move some of the straightforward batching code here, > I can follow that approach. Perhaps I could move patch 5, patch 8, > and the iommu_dma_iova_unlink_range_slow change from patch 7 here, > while keeping > > [PATCH 6] dma-mapping: Support batch mode for > dma_direct_{map,unmap}_sg > > and the IOVA link part from patch 7 as separate patches, since that > part is not straightforward. The IOVA link changes affect both > __dma_iova_link() and dma_iova_sync(), which are two separate > functions and require a deeper understanding of the contexts to > determine correctness. That part also lacks testing. > > Would that be okay with you? Yes, this will be okay. The changes are easy to understand, so we don't need to go there with such very small steps. Best regards -- Marek Szyprowski, PhD Samsung R&D Institute Poland