From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org X-Spam-Level: X-Spam-Status: No, score=-8.7 required=3.0 tests=DKIMWL_WL_HIGH,DKIM_SIGNED, DKIM_VALID,DKIM_VALID_AU,HEADER_FROM_DIFFERENT_DOMAINS,INCLUDES_PATCH, MAILING_LIST_MULTI,SIGNED_OFF_BY,SPF_PASS,UNPARSEABLE_RELAY,USER_AGENT_MUTT autolearn=ham autolearn_force=no version=3.4.0 Received: from mail.kernel.org (mail.kernel.org [198.145.29.99]) by smtp.lore.kernel.org (Postfix) with ESMTP id 82A29C43381 for ; Sat, 23 Feb 2019 00:47:08 +0000 (UTC) Received: from vger.kernel.org (vger.kernel.org [209.132.180.67]) by mail.kernel.org (Postfix) with ESMTP id 45F92205C9 for ; Sat, 23 Feb 2019 00:47:08 +0000 (UTC) Authentication-Results: mail.kernel.org; dkim=pass (2048-bit key) header.d=oracle.com header.i=@oracle.com header.b="OR9EanjQ" Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1725900AbfBWArH (ORCPT ); Fri, 22 Feb 2019 19:47:07 -0500 Received: from userp2120.oracle.com ([156.151.31.85]:38754 "EHLO userp2120.oracle.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1725774AbfBWArH (ORCPT ); Fri, 22 Feb 2019 19:47:07 -0500 Received: from pps.filterd (userp2120.oracle.com [127.0.0.1]) by userp2120.oracle.com (8.16.0.27/8.16.0.27) with SMTP id x1N0YStb131374; Sat, 23 Feb 2019 00:46:56 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=oracle.com; h=date : from : to : cc : subject : message-id : references : mime-version : content-type : in-reply-to; s=corp-2018-07-02; bh=p3Br5qFrZ0dyJaWuLjX3unOG+h/3/rh6R0rQ3QOQ/yY=; b=OR9EanjQY8P+7rcK4zNMwBV7N5SsAtMxo5QV0V0lB6D4KczuL6SSP1HHeBMJn8Ud8vYM JyWTBtlltcsjDqGFlMMJ3pSsxofCZNgmVgc4TDjMpugoWAnu3ACgR3XtoK1zmF5JUztD XS2sKEe0dhUri1azQK66hNlJN2njDwcpZuIXUJHn8WSMLzCXZ+5UVvy+UgcHKcnzmGvm YkhZEfYjzkYLagHHSoJhV4mRyaYsPT/aoyCHi8TH8MGoLU51l7PDnCLZbvSmixHqh4pM fnxClfnaVx1DqpKv60nkk0tM+KfhtR7UDd3HFNK3j3SIl0ondOALs+ajvnjhEOpvINYY 6g== Received: from userv0022.oracle.com (userv0022.oracle.com [156.151.31.74]) by userp2120.oracle.com with ESMTP id 2qpb5s22ha-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Sat, 23 Feb 2019 00:46:56 +0000 Received: from userv0122.oracle.com (userv0122.oracle.com [156.151.31.75]) by userv0022.oracle.com (8.14.4/8.14.4) with ESMTP id x1N0kuLe009655 (version=TLSv1/SSLv3 cipher=DHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Sat, 23 Feb 2019 00:46:56 GMT Received: from abhmp0013.oracle.com (abhmp0013.oracle.com [141.146.116.19]) by userv0122.oracle.com (8.14.4/8.14.4) with ESMTP id x1N0ktXB017094; Sat, 23 Feb 2019 00:46:55 GMT Received: from localhost (/10.159.224.245) by default (Oracle Beehive Gateway v4.0) with ESMTP ; Fri, 22 Feb 2019 16:46:55 -0800 Date: Fri, 22 Feb 2019 16:46:53 -0800 From: "Darrick J. Wong" To: Dave Chinner Cc: Dan Williams , linux-nvdimm , Ross Zwisler , Vishal L Verma , xfs , linux-fsdevel Subject: Re: [RFC PATCH] pmem: advertise page alignment for pmem devices supporting fsdax Message-ID: <20190223004653.GD21626@magnolia> References: <20190222182008.GT6503@magnolia> <20190222184525.GA21626@magnolia> <20190222233038.GD23020@dastard> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20190222233038.GD23020@dastard> User-Agent: Mutt/1.9.4 (2018-02-28) X-Proofpoint-Virus-Version: vendor=nai engine=5900 definitions=9175 signatures=668685 X-Proofpoint-Spam-Details: rule=notspam policy=default score=0 priorityscore=1501 malwarescore=0 suspectscore=0 phishscore=0 bulkscore=0 spamscore=0 clxscore=1015 lowpriorityscore=0 mlxscore=0 impostorscore=0 mlxlogscore=999 adultscore=0 classifier=spam adjust=0 reason=mlx scancount=1 engine=8.0.1-1810050000 definitions=main-1902230002 Sender: linux-fsdevel-owner@vger.kernel.org Precedence: bulk List-ID: X-Mailing-List: linux-fsdevel@vger.kernel.org On Sat, Feb 23, 2019 at 10:30:38AM +1100, Dave Chinner wrote: > On Fri, Feb 22, 2019 at 10:45:25AM -0800, Darrick J. Wong wrote: > > On Fri, Feb 22, 2019 at 10:28:15AM -0800, Dan Williams wrote: > > > On Fri, Feb 22, 2019 at 10:21 AM Darrick J. Wong > > > wrote: > > > > > > > > Hi all! > > > > > > > > Uh, we have an internal customer who's been trying out MAP_SYNC > > > > on pmem, and they've observed that one has to do a fair amount of > > > > legwork (in the form of mkfs.xfs parameters) to get the kernel to set up > > > > 2M PMD mappings. They (of course) want to mmap hundreds of GB of pmem, > > > > so the PMD mappings are much more efficient. > > Are you really saying that "mkfs.xfs -d su=2MB,sw=1 " is > considered "too much legwork" to set up the filesystem for DAX and > PMD alignment? Yes. I mean ... userspace /can/ figure out the page sizes on arm64 & ppc64le (or extract it from sysfs), but why not just advertise it as a io hint on the pmem "block" device? Hmm, now having watched various xfstests blow up because they don't expect blocks to be larger than 64k, maybe I'll rethink this as a default behavior. :) > > > > I started poking around w.r.t. what mkfs.xfs was doing and realized that > > > > if the fsdax pmem device advertised iomin/ioopt of 2MB, then mkfs will > > > > set up all the parameters automatically. Below is my ham-handed attempt > > > > to teach the kernel to do this. > > Still need extent size hints so that writes that are smaller than > the PMD size are allocated correctly aligned and sized to map to > PMDs... I think we're generally planning to use the RT device where we can make 2M alignment mandatory, so for the data device the effectiveness of the extent hint doesn't really matter. > > > > Comments, flames, "WTF is this guy smoking?" are all welcome. :) > > > > > > > > --D > > > > > > > > --- > > > > Configure pmem devices to advertise the default page alignment when said > > > > block device supports fsdax. Certain filesystems use these iomin/ioopt > > > > hints to try to create aligned file extents, which makes it much easier > > > > for mmaps to take advantage of huge page table entries. > > > > > > > > Signed-off-by: Darrick J. Wong > > > > --- > > > > drivers/nvdimm/pmem.c | 5 ++++- > > > > 1 file changed, 4 insertions(+), 1 deletion(-) > > > > > > > > diff --git a/drivers/nvdimm/pmem.c b/drivers/nvdimm/pmem.c > > > > index bc2f700feef8..3eeb9dd117d5 100644 > > > > --- a/drivers/nvdimm/pmem.c > > > > +++ b/drivers/nvdimm/pmem.c > > > > @@ -441,8 +441,11 @@ static int pmem_attach_disk(struct device *dev, > > > > blk_queue_logical_block_size(q, pmem_sector_size(ndns)); > > > > blk_queue_max_hw_sectors(q, UINT_MAX); > > > > blk_queue_flag_set(QUEUE_FLAG_NONROT, q); > > > > - if (pmem->pfn_flags & PFN_MAP) > > > > + if (pmem->pfn_flags & PFN_MAP) { > > > > blk_queue_flag_set(QUEUE_FLAG_DAX, q); > > > > + blk_queue_io_min(q, PFN_DEFAULT_ALIGNMENT); > > > > + blk_queue_io_opt(q, PFN_DEFAULT_ALIGNMENT); > > > > > > The device alignment might sometimes be bigger than this default. > > > Would there be any detrimental effects for filesystems if io_min and > > > io_opt were set to 1GB? > > > > Hmmm, that's going to be a struggle on ext4 and the xfs data device > > because we'd be preferentially skipping the 1023.8MB immediately after > > each allocation group's metadata. It already does this now with a 2MB > > io hint, but losing 1.8MB here and there isn't so bad. > > > > We'd have to study it further, though; filesystems historically have > > interpreted the iomin/ioopt hints as RAID striping geometry, and I don't > > think very many people set up 1GB raid stripe units. > > Setting sunit=1GB is really going to cause havoc with things like > inode chunk allocation alignment, and the first write() will either > have to be >=1GB or use 1GB extent size hints to trigger alignment. > And, AFAICT, it will prevent us from doing 2MB alignment on other > files, even with 2MB extent size hints set. > > IOWs, I don't think 1GB alignment is a good idea as a default. > > (I doubt very many people have done 2M raid stripes either, but it seems > > to work easily where we've tried it...) > > That's been pretty common with stacked hardware raid for as long as > I've worked on XFS. e.g. a software RAID0 stripe of hardware RAID5/6 > luns was pretty common with large storage arrays in HPC environments > (i.e. huge streaming read/write bandwidth). In these cases, XFS was > set up with the RAID5/6 lun width as the stripe unit (commonly 2MB > with 8+1 and 256k raid chunk size), and the RAID 0 > width as the stripe width (commonly 8-16 wide spread across 8-16 FC > portsi w/ multipath) and it wasn't uncommon to see widths in the > 16-32MB range. > > This aligned the filesystem to the underlying RAID5/6 luns, and > allows stripe width IO to be aligned an hit every RAID5/6 lun > evenly. Ensuring applications could do this easily with large direct > IO reads and writes is where the swalloc and largeio mount > options come into their own.... > > > I'm thinking and xfs-realtime configuration might be able to support > > > 1GB mappings in the future. > > > > The xfs realtime device ought to be able to support 1g alignment pretty > > easily though. :) > > Yup, but I think that's the maximum "block" size it can support and It is; our users with 16G page size are out of luck. > DAX will have some serious long tail latency and CPU usage issues at > allocation time because each new 1GB "block" that is dynamically > allocated will have to be completely zeroed during the allocation > inside the page fault handler..... Agreed, 1G pages are most probably too unwieldly to be worth advertising. --D > Cheers, > > Dave. > -- > Dave Chinner > david@fromorbit.com