From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from kanga.kvack.org (kanga.kvack.org [205.233.56.17]) by smtp.lore.kernel.org (Postfix) with ESMTP id 18F12C3DA5D for ; Fri, 19 Jul 2024 17:02:17 +0000 (UTC) Received: by kanga.kvack.org (Postfix) id A110E6B0089; Fri, 19 Jul 2024 13:02:16 -0400 (EDT) Received: by kanga.kvack.org (Postfix, from userid 40) id 9BF8B6B0092; Fri, 19 Jul 2024 13:02:16 -0400 (EDT) X-Delivered-To: int-list-linux-mm@kvack.org Received: by kanga.kvack.org (Postfix, from userid 63042) id 8606E6B0093; Fri, 19 Jul 2024 13:02:16 -0400 (EDT) X-Delivered-To: linux-mm@kvack.org Received: from relay.hostedemail.com (smtprelay0015.hostedemail.com [216.40.44.15]) by kanga.kvack.org (Postfix) with ESMTP id 649936B0089 for ; Fri, 19 Jul 2024 13:02:16 -0400 (EDT) Received: from smtpin05.hostedemail.com (a10.router.float.18 [10.200.18.1]) by unirelay01.hostedemail.com (Postfix) with ESMTP id C70251C1486 for ; Fri, 19 Jul 2024 17:02:15 +0000 (UTC) X-FDA: 82357120230.05.11ECDB1 Received: from mail-qk1-f173.google.com (mail-qk1-f173.google.com [209.85.222.173]) by imf30.hostedemail.com (Postfix) with ESMTP id D05578000D for ; Fri, 19 Jul 2024 17:02:12 +0000 (UTC) Authentication-Results: imf30.hostedemail.com; dkim=pass header.d=cmpxchg-org.20230601.gappssmtp.com header.s=20230601 header.b=irt2gfHC; spf=pass (imf30.hostedemail.com: domain of hannes@cmpxchg.org designates 209.85.222.173 as permitted sender) smtp.mailfrom=hannes@cmpxchg.org; dmarc=pass (policy=none) header.from=cmpxchg.org ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=hostedemail.com; s=arc-20220608; t=1721408490; h=from:from:sender:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type:content-transfer-encoding: in-reply-to:in-reply-to:references:references:dkim-signature; bh=Cnk9NxbJkVDBO6rEbrTCNl0EkRG/DVVM+mWiYP6bZrA=; b=zI5e/3GBRVJ6fbMrOLFBC4gd++rn7DJdCUPRPRQnkK3zSlAsrS34WyqPMHP1JO5cM3QdBt 0lhvMMwhzpYBMYlmmUkIdcbPIy5fd/+T3jRbR2xJ/LNrO8+kFwjDtb+34ZH/uJyDuTNrkJ AerkuIpz69jZLU3xZmnr3u5MxEVLcKU= ARC-Seal: i=1; s=arc-20220608; d=hostedemail.com; t=1721408490; a=rsa-sha256; cv=none; b=gAThAOWLCuupWRsKbQOT0qESp/+xZLj0Vn7+kOtCUSkcsDewS0v7Ya2vMp0Ved59vQCNSN KwoUYkAMSsW7h/VrNQfD/3hwyxVYQINNrzv4K3tjaRF9KMpzI41DtA0ApX2VmyUH+bbrK2 D9RzkQQAQFyMLaKjmOKVP8fhOhljlGg= ARC-Authentication-Results: i=1; imf30.hostedemail.com; dkim=pass header.d=cmpxchg-org.20230601.gappssmtp.com header.s=20230601 header.b=irt2gfHC; spf=pass (imf30.hostedemail.com: domain of hannes@cmpxchg.org designates 209.85.222.173 as permitted sender) smtp.mailfrom=hannes@cmpxchg.org; dmarc=pass (policy=none) header.from=cmpxchg.org Received: by mail-qk1-f173.google.com with SMTP id af79cd13be357-79f19f19070so84564085a.0 for ; Fri, 19 Jul 2024 10:02:12 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=cmpxchg-org.20230601.gappssmtp.com; s=20230601; t=1721408531; x=1722013331; darn=kvack.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=Cnk9NxbJkVDBO6rEbrTCNl0EkRG/DVVM+mWiYP6bZrA=; b=irt2gfHC+bPJIppSa7r2kuYkLzbz9Bp0R9kFzLJKh+KuyVKF8/0vjeEeXRjUeYNCgH 8O4ZzDERPMojoSfclNxVhw9D9I819zlkSPdnJG7cp0qGam4aAUJgBTL0TmwDhuce+Tl4 8d1F7VUCBb18Z90NO0pSi5qeDnNHlbngt4D21g9c89vkvbP/kIwrAZ1UzgXoTj2mNSxk pdyyuBd5NFWWyQj0iMiCXr+LuIZd51M95L5i7VO38fbdB9kpT184VznaMHBAPl2oop1b KhUEpVSdOce+EdBbjzoB5CUXyjCM4CPcRPe0baNBDLjOai35F2SqgcXA7V1G3/TaBRBX ckfg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1721408531; x=1722013331; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=Cnk9NxbJkVDBO6rEbrTCNl0EkRG/DVVM+mWiYP6bZrA=; b=jVzXc5Sdo9oAxXWl8S224IpMmmpvVcMBbjSRuw6LAcPEC1QwcA3vimAatud2Ji/oru 822e6gICQ05sZROuZ5Wuud/sGML+iGpjsyshXhSjbZRRmugAfbfUc70OzNuW2TGigZ/q JbcGKwkqefjl4gbYzaMcOpxOZGSX5qrFJLpByok+FpRbaeeESsXix21pGhJKY0Z2aUnF vEm51teLkTTzc2QXwTcwvmmshHnMgvRoFuBCyUNmjTycfxpbQhmtyHiK/7vrFa/bC8sX OZZAE/S04jXNe7Yug5hT17OUGtNbEqxCxqs1+LXhKXwZUBBeEiftcCHF7Rm+zb3Its+r Iukg== X-Forwarded-Encrypted: i=1; AJvYcCUUxgCG8jQLMcb7aLgcTxDJ26oFgthn2Ry37cOxCL5Gcywh+MChnLDviT+RhmGZvPKIZedznyFYpXoxqgelzalpOkk= X-Gm-Message-State: AOJu0YwLSkiaIO4bxS5SHiUDXJBIhM9xdj5/8as6+8t7iCBxDcSQ49MH oqnbAIkT2T7s0wBDl2eEqFBq1x7r4lBeUNogZbDBnz+btWICPbdlZa6hFREXWr8= X-Google-Smtp-Source: AGHT+IGJpOS3XSHRy7i/fHyIF9ZOh73D7/b4jdB8R2qZ3jDtr4nw9P7TgvwN7pA3m41bG/vJBRUWFg== X-Received: by 2002:a05:6214:c41:b0:6b5:16b:6998 with SMTP id 6a1803df08f44-6b78e258c37mr118182136d6.42.1721408531253; Fri, 19 Jul 2024 10:02:11 -0700 (PDT) Received: from localhost ([2603:7000:c01:2716:da5e:d3ff:fee7:26e7]) by smtp.gmail.com with ESMTPSA id 6a1803df08f44-6b7ac7e4754sm9947136d6.51.2024.07.19.10.02.10 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 19 Jul 2024 10:02:10 -0700 (PDT) Date: Fri, 19 Jul 2024 13:02:06 -0400 From: Johannes Weiner To: Qu Wenruo Cc: linux-btrfs@vger.kernel.org, mhocko@kernel.org, roman.gushchin@linux.dev, shakeel.butt@linux.dev, muchun.song@linux.dev, cgroups@vger.kernel.org, linux-mm@kvack.org, Michal Hocko , Vlastimil Babka Subject: Re: [PATCH v7 2/3] btrfs: always uses root memcgroup for filemap_add_folio() Message-ID: <20240719170206.GA3242034@cmpxchg.org> References: <6a9ba2c8e70c7b5c4316404612f281a031f847da.1721384771.git.wqu@suse.com> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <6a9ba2c8e70c7b5c4316404612f281a031f847da.1721384771.git.wqu@suse.com> X-Stat-Signature: sgsz9gep1zpmgi85qeruxaj7mfffrbph X-Rspamd-Queue-Id: D05578000D X-Rspam-User: X-Rspamd-Server: rspam08 X-HE-Tag: 1721408532-877046 X-HE-Meta: U2FsdGVkX1/6awp8HLWm0t8+5CqGK5VCAoItuXp8TiqiapdLkNAY37cJEmao8KYXyynjanS/5MG7Xdda+mnFwvG/kWa5t5R3bL7IjmKwyhRb/4A51sof+SkGr4N5om8ghNOOottF1KQeYvR8uhxl4YN6+8GZsMroeFY5luvPZDOUC4PX0ygIZCQ4KdSNDyVVxwN686afa3yJWZu88unlwKhuw4ezKqkerGw69sOV1D82FemGIqNmuVlFl282q8YbjmSAaLmWae7Bieqi+KMNTOraceXehQhCbafphtfAMutthTWX5DchLfJla86OfdaNOmeEN1qDh330jDvPnH+FsEZz5EZxCmgu7QdcuWACiJqgG7xHDaKCSQTdqs+Atyh++Ckfx8jRa0jq7FkSIWbu5GTG2Xczw1brXujQvJYtTge/sSwlaZUccaRxY/AU8VA/4mM1V+88+oTC528+FtJKKX+3LYB+JyO9kNEz2rQP/0Zvip57kwYke6zy3eeH/OHE0DE6ucilFYpGjYxeea61pfsoPh7f1JJmN2tmkJFrj5HO4KI5Zz/n3Dm4fvz3MBb4QOlReiFJz8omGOnKWf1Ehsr/j0xbZH2srDAsBrN3l6zwyBJb2CYFD1TAMDf4w4PrnYsoBV7TS9NbftnZoepzDY6K/QC74vhsmIi8KUQjuJcmSekSnOWM8cr408JBjps2RED5HEtVuY2jJlGnp05o7aw7R0ATVXY/rmsXVCILm/vruPrLqL7uOff2bv/2ZdimI5jC0sPODZTWID7PsjL4Hdc42i9vOLI+wgjBziBN9AIHZ1jXw7QPz1bK1OaDQ+fIjM6msVCjgAl5T+S3ir4U2SEPZBiCha/b6jYWb4WPEjEB7YSvd5yNhSxTaoWNGI/qF/JjqfNI5V5iksb2KN7be5zFfJ2yYHk2avrpx1MSJ7KoHjjaL9Cg3qKN4KpG+qRl8qorm7WsOjKDVrSfb+5 rQAvjoDR TzKtTZOnxGcP97WDGvPMju43mdv+JguPoK3MVVpqai59TJQgNU+M9L84KeylAiJTHp3AbZ7EgyAeL4/g/tUPLWjF3Impx4Ep9FzBs1t1VbDEFSaJBabekGvpyqlIBluiD3E2pGw/dR2aN7WX0LmULS+Wb4PbDirtyHSSntIaD+XpRuMt2X/Ue35itQ3Slqd9itVgTTG1CjY/vPQGyMJtIUl12IbUNtS3h7YV90fXVreGeV579AuV8G4LjQcYWW9U1GHbbDjTkrrA8p31W66qmKUgLLXDrYdFREq3jUmBz30M44fLFZRP42E8axta4Bo9V6kMj/MGJwvFE1kWICb3tfCrareYzDSKp8CP61oLb/AmRW7dHEZNP2C4F9Jyp4egJonR1PvlhO1FrvrFzLtG9EI8j2sIILh21RVemXPN3/3zpeg8i91PuxCnIbWt88OHvpPzSwz5M1dRLY3TPG06krWLIrnYsDeu5jluqKnrm5tePZMn4UoFP0xQQ4u8PCayxlGIiszIJPSCOFGjf0KtFZKUuG4jhiLkg0P+yf6D+/34YJAwVAjBsLt2kc5J0E8BVl4QfnipBAGC5SeW5+bEP4rlMYO0OM/RxjnaN X-Bogosity: Ham, tests=bogofilter, spamicity=0.000000, version=1.2.4 Sender: owner-linux-mm@kvack.org Precedence: bulk X-Loop: owner-majordomo@kvack.org List-ID: List-Subscribe: List-Unsubscribe: On Fri, Jul 19, 2024 at 07:58:40PM +0930, Qu Wenruo wrote: > [BACKGROUND] > The function filemap_add_folio() charges the memory cgroup, > as we assume all page caches are accessible by user space progresses > thus needs the cgroup accounting. > > However btrfs is a special case, it has a very large metadata thanks to > its support of data csum (by default it's 4 bytes per 4K data, and can > be as large as 32 bytes per 4K data). > This means btrfs has to go page cache for its metadata pages, to take > advantage of both cache and reclaim ability of filemap. > > This has a tiny problem, that all btrfs metadata pages have to go through > the memcgroup charge, even all those metadata pages are not > accessible by the user space, and doing the charging can introduce some > latency if there is a memory limits set. > > Btrfs currently uses __GFP_NOFAIL flag as a workaround for this cgroup > charge situation so that metadata pages won't really be limited by > memcgroup. > > [ENHANCEMENT] > Instead of relying on __GFP_NOFAIL to avoid charge failure, use root > memory cgroup to attach metadata pages. > > With root memory cgroup, we directly skip the charging part, and only > rely on __GFP_NOFAIL for the real memory allocation part. > > Suggested-by: Michal Hocko > Suggested-by: Vlastimil Babka (SUSE) > Signed-off-by: Qu Wenruo > --- > fs/btrfs/extent_io.c | 10 ++++++++++ > 1 file changed, 10 insertions(+) > > diff --git a/fs/btrfs/extent_io.c b/fs/btrfs/extent_io.c > index aa7f8148cd0d..cfeed7673009 100644 > --- a/fs/btrfs/extent_io.c > +++ b/fs/btrfs/extent_io.c > @@ -2971,6 +2971,7 @@ static int attach_eb_folio_to_filemap(struct extent_buffer *eb, int i, > > struct btrfs_fs_info *fs_info = eb->fs_info; > struct address_space *mapping = fs_info->btree_inode->i_mapping; > + struct mem_cgroup *old_memcg; > const unsigned long index = eb->start >> PAGE_SHIFT; > struct folio *existing_folio = NULL; > int ret; > @@ -2981,8 +2982,17 @@ static int attach_eb_folio_to_filemap(struct extent_buffer *eb, int i, > ASSERT(eb->folios[i]); > > retry: > + /* > + * Btree inode is a btrfs internal inode, and not exposed to any > + * user. > + * Furthermore we do not want any cgroup limits on this inode. > + * So we always use root_mem_cgroup as our active memcg when attaching > + * the folios. > + */ > + old_memcg = set_active_memcg(root_mem_cgroup); > ret = filemap_add_folio(mapping, eb->folios[i], index + i, > GFP_NOFS | __GFP_NOFAIL); > + set_active_memcg(old_memcg); It looks correct. But it's going through all dance to set up current->active_memcg, then have the charge path look that up, css_get(), call try_charge() only to bail immediately, css_put(), then update current->active_memcg again. All those branches are necessary when we want to charge to a "real" other cgroup. But in this case, we always know we're not charging, so it seems uncalled for. Wouldn't it be a lot simpler (and cheaper) to have a filemap_add_folio_nocharge()?