From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-yx2-f13.google.com (mail-yx2-f13.google.com [74.125.224.141]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 97F3349B5AF for ; Mon, 21 Sep 2026 13:09:01 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.224.141 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789996144; cv=none; b=H40re5BRL3ieClvWdt7Ydru1QKuE6j8tQzXuYA4sfldu3s74ioXvdHTE8xSz1MKj56vKlJIlosBdH5FZ5oceN1qhX5kQe1qq2sC9poubSxdsci7+bDX96s6larDjP/4zWO6h7giOEezb1MuCD1CmlV4j+vpYmosuxy04GO2zToM= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789996144; c=relaxed/simple; bh=7iRzAcHHARKyQ+ATaMDpW4DruCrP6jplXNxnuj7OqwM=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=WW3twCjgUv3cg+6eMoBGPKMQb5u39H6NKGV8eaJmxF2B7AncR+nC3GbDRV2SeAso3HjfIaL2xOWoIMd1qcZTeSThkEcprKMgAql2SO2iw0e8GZUd8PTZJg/sOuS/X/gIm+SrZBOcxiqTStM7S4l0b02A9kZTiRqAF8t79rGQXdw= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=ziepe.ca; spf=pass smtp.mailfrom=ziepe.ca; dkim=pass (2048-bit key) header.d=ziepe.ca header.i=@ziepe.ca header.b=OVbP23qH; arc=none smtp.client-ip=74.125.224.141 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=ziepe.ca Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=ziepe.ca Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ziepe.ca header.i=@ziepe.ca header.b="OVbP23qH" Received: by mail-yx2-f13.google.com with SMTP id 956f58d0204a3-66f943b286fso2487875d50.2 for ; Mon, 21 Sep 2026 06:09:01 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ziepe.ca; s=google; t=1789996141; x=1790600941; darn=vger.kernel.org; h=in-reply-to:content-transfer-encoding:content-disposition :content-type:mime-version:references:message-id:subject:cc:to:from :date:from:to:cc:subject:date:message-id:reply-to:content-type; bh=/I4CnabOFJ3V+cBV+khghf4ltJkhcHX6gZAKyxfIiuU=; b=OVbP23qHxpA1f0E25xNKG0Utb4qY1pci7lNk1dz9kSCfiFs84khz7O89QkedKxGdF/ 0BgGqFlTDcjyoq1EhB62Hs1JbrYZYD8S5vHBJOhauPNXuaX1nb3pv3wtsBQTIyaxKlHH W5RPLjfQVDytPLMyAVNbHYDlUMwBER/D7KO1XAs4nvgPFc7AW8QN2S5fU7XQBgbrQklV mea3UtzB/mAP5ERVg1OpKwyibQEtCEHIyAnlL9kblIObxvClWKjATxcupqTIfL8HDbgJ DArpRHVstouE13CeyOaGnHw7CHAQLXfCxKjIZToSseuODn/RSEXWdw31uTLaznBIBhdH ujIQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789996141; x=1790600941; h=in-reply-to:content-transfer-encoding:content-disposition :content-type:mime-version:references:message-id:subject:cc:to:from :date:x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to:content-type; bh=/I4CnabOFJ3V+cBV+khghf4ltJkhcHX6gZAKyxfIiuU=; b=LvvPR41ci1FpzwTcm1jp2T9bUA6oPXxJIu/bPm7nG3OZyxUQwMUr/499j7xi+ofXj3 pbHCjAYB1VbxIotAB8Qc6DQozILyOkakZI6PFQByU9uHTtLP6fdhVz7rCg8E4qE4fedA 5hZMN4DEOU3vpxrHFliTKP6q7RIW6VdPcEWrY1MwKmVa7aVB1G2NKu1yeFvkjipaYq5F z63a3Fgkss5Us0kt2SAJB0RCN+XwIWMUr0TU2mBPj52mwb8QeO82MK5MtV5jSM0rvvwY MhcqZIQOtNTdps53bmkSzpHrN3NqMqkL36C6/V9woyRkz92RQGoQ3W9ngypx3DuwsKK0 sCsw== X-Forwarded-Encrypted: i=1; AKwUvBy4CWBFbDikeZoKL51aUFWXeKSc9MCvpXge1f7HMsxsVWE6XQjgJaKwzt8ZnoJ/ds8MtQo=@vger.kernel.org X-Gm-Message-State: AFuF++lExsi55Uck5y1lxbIPE5LXM1UUVar16npQCv4dYELeZuQDUYPY CpffbVRjoW6aBOETIFM3XIYXSvsx+GLwp2q5ne8sLFaylaIIzpS/xLUspxyV+13Z18I= X-Gm-Gg: AYBFou33l9W/fExrBu7pV/ifq1Qw39j4Kysj8dEGdXRyg+BAydhpR97DXsfUnBz5qEP t/3vvJRi7rmktBQa4mgZLAqZd28AKju1GZ1y/yAELkvtDL5WfM2Gw8YTBVu7l1pMoNKZSCYRdYq HXtKvBLWLnuSiVcZkiPpbcmboljmdUcrJa33vnSwW2G1ei4exqYTmuob/OhRi0Dn4/SatfJPdL0 eaLeQtl5CWCXG2RPSpnUWSfy7a1UkAUUJScKL199e4Z000p6sJMYyjuFj395i4ePUDW7ccymH41 yAiyxQpeWDOP3c3ejUubRqDyt+2/5fWE0DrjwNngf36lFWTqZOhbLL8Rj0JlZ40AQnRPW5uEt1O +VFos6qooH4QRIusg3h+9m9Jj2h1K5XBuMsQK31NZAygbZ2SuxZc2kzTMNUMrXrkcPy1+dx+Es3 cXY+zSwoaJ0uODMWvRUbgHvWtNiDymB2QbJXJa+r9ejNHOx4CWZPyQ+5vcVFwjiT2sSnOFj1YQz fbmuVD+gY+4vj3paWzyirNwRalbR75ZixDm7tu1n6lvgfCTH6OERxz4 X-Received: by 2002:a05:690e:11ca:b0:672:a52c:db25 with SMTP id 956f58d0204a3-672a52ce30bmr1843524d50.109.1789996140490; Mon, 21 Sep 2026 06:09:00 -0700 (PDT) Received: from ziepe.ca (hlfxns010zw-159-2-239-150.pppoe-dynamic.high-speed.ns.bellaliant.net. [159.2.239.150]) by smtp.gmail.com with ESMTPSA id 6a1803df08f44-91260aa1db7sm67720106d6.44.2026.09.21.06.08.59 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 21 Sep 2026 06:08:59 -0700 (PDT) Received: from jgg by wakko with local (Exim 4.97) (envelope-from ) id 1x8dlS-00000006Jlj-41LR; Mon, 21 Sep 2026 10:08:58 -0300 Date: Mon, 21 Sep 2026 10:08:58 -0300 From: Jason Gunthorpe To: Christian =?utf-8?B?S8O2bmln?= Cc: Thomas =?utf-8?Q?Hellstr=C3=B6m?= , Christoph Hellwig , Leon Romanovsky , Bjorn Helgaas , Logan Gunthorpe , Chaitanya Kulkarni , Greg Kroah-Hartman , Jens Axboe , Alex Williamson , Ankit Agrawal , Jonathan Corbet , Shuah Khan , "Joerg Roedel (AMD)" , Will Deacon , Robin Murphy , Randy Dunlap , Sumit Semwal , linux-pci@vger.kernel.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, iommu@lists.linux.dev, Tushar Dave , linux-media@vger.kernel.org, dri-devel@lists.freedesktop.org, linaro-mm-sig@lists.linaro.org, linux-rdma@vger.kernel.org, kvm@vger.kernel.org Subject: Re: [PATCH v6 18/18] RDMA/mlx5: Ask P2PDMA whether ATS takes a direct peer-to-peer route Message-ID: <20260921130858.GO11599@ziepe.ca> References: <20260914-fix-p2p-acs-v4-0-v6-0-5ef07ec9ef06@nvidia.com> <20260914-fix-p2p-acs-v4-0-v6-18-5ef07ec9ef06@nvidia.com> <321890690ce83d1943b2f678bd9bee9b8c895b66.camel@linux.intel.com> <20260918121500.GV13683@unreal> <20260918170524.GH11599@ziepe.ca> <9656f2f9-2e39-4007-b065-3b763b53c6de@amd.com> Precedence: bulk X-Mailing-List: kvm@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Disposition: inline Content-Transfer-Encoding: 8bit In-Reply-To: <9656f2f9-2e39-4007-b065-3b763b53c6de@amd.com> On Mon, Sep 21, 2026 at 08:43:49AM +0200, Christian König wrote: > On 9/18/26 19:05, Jason Gunthorpe wrote: > > On Fri, Sep 18, 2026 at 03:42:28PM +0200, Thomas Hellström wrote: > >> > >> 1) Xe attachment check if pci_p2pdma_distance() returns OK for the > >> path. Then Xe always sets up dma-addresses using dma_map_resource(). > > > > Open coding pci_p2pdma_distance() in drivers is a hack. Using > > dma_map_resource() like this was never "allowed". > > > > We've fixed things so these hacks are not needed, the drivers need to > > move over to things like dma_buf_phys_vec_to_sgt() and the hmm helpers > > to use the DMA API correctly. > > That is a completely broken approach as well since it limits the > exported resources to addresses the CPU can reach. Yes, of course it does. The DMA API only works on phy_addr_t. If you have a struct p2pdma_provider * then you have a phys_addr_t for it. If the exporter knows it is working with a MMIO mapping on a PCI device it gets to acquire a p2pdma_provider and use the helper. None of this is suposed to solve your "resources the CPU cannot reach" problem, that has nothing to do with DMA API or P2P. It is not broken just because it doesn't solve every problem. > Christoph Hellwig is right that drivers should never use that stuff > directly, not even through that dma_buf_phys_vec_to_sgt() function. Hellwig's point was that the subsystem needs a mapping helper that goes from the subsystem address representation to the HW representation and hides these details from the drivers. Look at what he built in nvme around biovec. The dma_buf_phys_vec_to_sgt() is the dmabuf version of the same idea, the subystem provides the mapping helper. We can try to do better, but better is not making the exporters touch the mapping algorithm. Ultimately I want to see something in lib/ handle this with a non-scatterlist datastructure, but there is a huge gulf between where dmabuf is now and it being able to work with a non-scatterlist datastructure. Look, it is easy to complain you don't like how it looks, but this stuff is hard there are lots of competing concerns, if you have a better idea now is a good time to present it. Maybe if you look closely you will appreciate how much work has gone into even getting things this far. > >> It seems to me that a pci-device settable flag "ATS always enabled" > >> should be enough to fix both issues? > > > > It should be be per-mapping to support the NIC workflow that isn't a > > global operation. > > I just realized what you guys are doing and I'm not sure if the > Linux PCI subsystem should support such hacks at all. I don't know how to respond to that. It is spec complaint, it is shipping in enormous volumes, of course Linux needs to support the HW that exists. > Basically from the point of view of the TA the NIC has ATS enabled > all the time, but has a per request option to use translated > addresses directly without previously translating and caching them > using ATS, correct? Yep. Very few devices in the world can use ATS for every single operation. Many have this split operating model. Some even only use ATS for DMA flows that are faultable. Mostly the OS can't tell what the device is doing and doesn't care. The P2P routing is the main issue and we have hacked around it in our systems till now. Leon is trying to fix it. Thomas needs it fixed too for Xe. So what's the issue here? DMABUF needs to learn how to do interconnect specific behaviors. PCI is an interconnect, it has lots and lots of crazy rules. An importer/exporter that chooses to use PCI for their DMA should have a way to exchange PCI specific information. So should UALink and all the other zoo of options we have now. It cannot be completely generic and meet everyones needs. Can we focus on that instead of arguing if the PCI craziness should exist or not? > If yes than that is extremely questionable behavior, I'm not sure if > that is covered by the PCIe spec. Spec doesn't say anything about when a device has to translated vs untranslated. The ATS flags only say translated is allowed to be used. > ATS is meant to be an optimization which moves the TLB from the root > complex (TA) into the devices at the cost of TLB invalidation > complexity. But what you do here is abusing that functionality as > far as I can see. ATS is for alot more than that, and there is no abuse here. Jason