From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wr1-f50.google.com (mail-wr1-f50.google.com [209.85.221.50]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AB5C636B900 for ; Sun, 2 Aug 2026 10:37:34 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.221.50 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785667056; cv=none; b=rv5IorRS7rxHF6bEwacOegjWCWOl/lmLv0zfxOBsifg21mL+49uAryWMdUGQu2zxzHLUoLGj7X0Ikzv41Pl0aF/d5zbml7V4eNI6b1x1m/GmW9ZAPanPz8urBlcIH3+2isvfdDNKhlvLc1zFG/YNR6X36/Bbk9od7mJ7sPutR8w= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785667056; c=relaxed/simple; bh=InKrJCk1a+fYtFFg5XL4yD/BmpaibaLkfR031TRLIf8=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=ZHzrt4XIkJ8M8SVgwcHYT7QconTLbnqkDSi9kXccHvD/f7G8o0NIV7e3F+sWh4Bj/NhWB70Tr40uI1XSyk67WfYyQ8E0XD4/uRWpg+1A4o/iSF71OzkEgO+u5V1E+30Ch41onlU5ymmxxYvLQjYM/AztsWZ2gaLX6c2tJCV3cQU= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=suse.com; spf=pass smtp.mailfrom=suse.com; dkim=pass (2048-bit key) header.d=suse.com header.i=@suse.com header.b=Z35zZedm; arc=none smtp.client-ip=209.85.221.50 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=suse.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=suse.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=suse.com header.i=@suse.com header.b="Z35zZedm" Received: by mail-wr1-f50.google.com with SMTP id ffacd0b85a97d-47c2ae992beso246872f8f.2 for ; Sun, 02 Aug 2026 03:37:34 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=suse.com; s=google; t=1785667053; x=1786271853; darn=vger.kernel.org; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:from:to:cc:subject :date:message-id:reply-to:content-type; bh=OcFGvJhCANWpgLUVs3lL6SLMUlm/Sqd+TRsG69oeBjM=; b=Z35zZedmnWUEc0EbbM8ziMlxqDRNWQnpUzGHpYoBmFAaT/sJAv1UUGeX4W7CD2a5Y+ HsHcVcAVZVeTZ/Pmv6zMiYWBbgqvRxdWkvA//5N7xddHPVbMHWAlgdUkV9/zO3iQvlyM l+hm2iVvD8GUjbJpalmaUaXAEh3brmKluPDg0DhGd/nTGkq/OPBslMnXfoif/3eUr3HF FvUKYKId0SxDQuzGzQZLM9qVu0K0u+dHc3vV5eqD9EfiQ4ZoaS5XUxrdABJuugKKZOx2 O6Tlt3DejjHoFeoTGEWS28iKPdc6oqXgjr0TGXtjD1mZud4VJqvH5pmXt+9ZEoe6ZQk1 +F4Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785667053; x=1786271853; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=OcFGvJhCANWpgLUVs3lL6SLMUlm/Sqd+TRsG69oeBjM=; b=enIgsStfd1Ld7g1l54QxkeBBB4Tla061XxHY4wrn4C9qNSj3G+dailm6se8jtnOiVA Q9o7DQ/owe/ZIRltgboJ6PpSP35tGr6PVjn3Hu89ERq/kSvw7zwEoGdhhO4kxtQQWRd4 nMTfubGeGCw5orJUFGd5QU1CrnoSqx8QbP8KrMGkrx+TztTCOOQ3UiifHxeC8ciSf2VD jVCRQFuRBLkJbK8HxQIDDX18AMBdraNVwabf127pbC/wpkodc2w7D52IzLEG9YiFRus9 4QplZEVxnToW8ZwXpivTmpf0GPlAfAjtnzXLZGMZJLk/lhw7GfxQBHJ+3FGVS83Xz3Rz t2Lg== X-Forwarded-Encrypted: i=1; AHgh+RpGm+BYLEZcK99y+9thSI84StqOw0g91oemORSFRscaJe/98e8LkhQ0uGtHPcVNtCDU8ewXyF6jvlijc1A=@vger.kernel.org X-Gm-Message-State: AOJu0YxQDqibANp/c3O64DBcetV8duis7wEG61cajdvevXLWJvP+0Pwc Qg0u/9P3llQjsoFRtqIJHG4wkVXGP1vTzn7CyIhZwQf3J1dt/OdLKRJPXlSpxpd9PzI= X-Gm-Gg: AR+sD115rK7yPnoiYgJfCLrWiAT62LKgVVqoIqr2eww1aBa3q6SYxYGDM5FE+Nc2gwy d6VMjtZf1meDDhR8Se898I1hAgSdI8hOKvpVqsg6GA+0IobKHHiN+cQF50qmchGPm84bqzpsxBG JnLbXXTuiFx4twjBzYTNpWNIh7pja1DHDhtjn/d0gsiLI87LuHfTKpT+jjTL0eFHJ4ygSZY+axI yvP7dXrV9EwxdOWqNI9ji/C3hpJhOYkhSqBuOHQfHxt+zxW71yvNqW45naNhCur6TXz/zxeOx4c Nmu0gsaNN688FZz3QDPljPMdrVni02B5/WQe0CzZyPAxynWeARYk0CCgFNCSb/WOlkDRB6Tguqr h/WZJHIpAZZx3Wjh2ClZQseTFkWzNbdpJT7wCvg746+wBdgKXsrJkyzboaTQV3ayITmz1yF38ND GdWJ1n67wXhFtfpv7xmGi5QMxC6PZiVZvioAOO+PCSzU45geLPIEfg3RXdDeo= X-Received: by 2002:a05:6000:1a8c:b0:46f:7d90:8124 with SMTP id ffacd0b85a97d-47fd7315ec8mr9835131f8f.2.1785667052846; Sun, 02 Aug 2026 03:37:32 -0700 (PDT) Received: from localhost ([202.127.77.110]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-38fb2e10549sm2696080a91.2.2026.08.02.03.37.30 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 02 Aug 2026 03:37:31 -0700 (PDT) Date: Sun, 2 Aug 2026 18:37:27 +0800 From: Heming Zhao To: Joseph Qi Cc: mark@fasheh.com, jlbec@evilplan.org, hch@lst.de, ocfs2-devel@lists.linux.dev, linux-kernel@vger.kernel.org Subject: Re: [RFC PATCH v2 3/4] ocfs2: switch dio write path from buffer_head to iomap Message-ID: References: <20260727061802.18485-1-heming.zhao@suse.com> <20260727061802.18485-4-heming.zhao@suse.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: On Tue, Jul 28, 2026 at 02:07:42PM +0800, Joseph Qi wrote: > > > On 7/27/26 2:17 PM, Heming Zhao wrote: > > This patch converts OCFS2's DIO write path from the legacy > > buffer_head infrastructure to the modern iomap framework. > > > > Key modifications and designs are as follows: > > > > 1. Dynamic Context Allocation: > > Refactor 'struct ocfs2_write_ctxt' to use a flexible array 'w_desc[]' > > instead of a fixed-size array. Dynamically allocate the context based on > > the write length ('w_clen') in 'ocfs2_alloc_write_ctxt()'. This prevents > > static limits overflow and optimizes kernel heap memory utilization. > > > > 2. Robust Mapping and Limits: > > Introduce 'ocfs2_dio_wr_map_blocks()' to allocate and map direct write > > blocks. Implement a 1 MiB cap ('OCFS2_DIO_WR_MAX_MAX_BYTES') per mapping > > call to restrict allocation granularity, preventing JBD2 transaction > > credit exhaustion during huge asynchronous sequential writes. > > > > 3. Reliable Completion Work and Fallback: > > - Implement 'ocfs2_iomap_dio_end_io_write()' to handle metadata completion. > > It converts UNWRITTEN extents, updates inode size (EOF), and deletes the > > inode from the orphan directory if it was appended. > > - Implement 'ocfs2_dio_write_end_io()' to finalize the dio lifecycle and > > release cluster locks safely. > > - Intercept '-ENOTBLK' errors from 'iomap_dio_rw()' caused by page cache > > invalidation failures (due to mmap/buffered collisions). Gracefully clear > > the error, strip the IOCB_DIRECT flag, and fall back to buffered write > > > > 4. Moved ocfs2_add_inode_to_orphan(): > > - moved ocfs2_add_inode_to_orphan() from ocfs2_dio_wr_map_blocks() to > > ocfs2_file_write_iter(). > > > > 5. Uncertain logic in code > > For the following code block in ocfs2_dio_wr_map_blocks(): > > ``` > > if (extend) { > > if (ocfs2_sparse_alloc(osb)) > > ret = ocfs2_zero_tail(inode, di_bh, pos); > > else > > ret = ocfs2_expand_nonsparse_inode(inode, di_bh, pos, > > map_len, NULL); > > if (ret < 0) { > > mlog_errno(ret); > > goto unlock; > > } > > } > > ``` > > I am not completely certain whether it is called only once per > > ocfs2_file_write_iter(), but I believe calling it multiple times will > > not introduce any side effects. Furthermore, testing across various > > scenarios showed no instances of multiple calls. > > > > Assisted-by: Gemini:gemini-3.5-flash > > Assisted-by: Claude:claude-sonnet-4-5 > > Co-developed-by: Joseph Qi > > Signed-off-by: Joseph Qi > > Signed-off-by: Heming Zhao > > --- > > fs/ocfs2/aops.c | 394 ++++++++++++++++++++++++++++++++++++-- > > fs/ocfs2/buffer_head_io.c | 7 +- > > fs/ocfs2/file.c | 91 +++++++-- > > fs/ocfs2/ocfs2.h | 2 + > > 4 files changed, 459 insertions(+), 35 deletions(-) > > > > diff --git a/fs/ocfs2/aops.c b/fs/ocfs2/aops.c > > index 12f5f2e3530a..9a079436c9c0 100644 > > --- a/fs/ocfs2/aops.c > > +++ b/fs/ocfs2/aops.c > > ... > > > + > > + ocfs2_free_unwritten_list(inode, &wc->w_unwritten_list); > > + ret = ocfs2_write_end_nolock(inode->i_mapping, pos, map_len, map_len, wc); > > + BUG_ON(ret != map_len); > > Under memory pressure, folio allocation may fail. In this case, > ocfs2_write_end_nolock() can return a short count. > So we must handle this case gracefully. I simply copied the code logic from ocfs2_dio_wr_get_block(). IIUC, folio allocation mainly happens in ocfs2_write_begin_nolock => ocfs2_grab_pages_for_write. while ocfs2_write_end_nolock is responsible for marking pages dirty. If the code encounters memory pressure issue, the existing error handling should be enough (I mean the error handling after ocfs2_write_begin_nolock()). > > > + ret = 0; > > + > > +unlock: > > + up_write(&oi->ip_alloc_sem); > > + ocfs2_inode_unlock(inode, 1); > > + brelse(di_bh); > > + > > +out: > > + return ret; > > +} > > + > > static int ocfs2_dio_end_io_write(struct inode *inode, > > struct ocfs2_dio_write_ctxt *dwc, > > loff_t offset, > > ... > > > +static int ocfs2_iomap_dio_end_io_write(struct inode *inode, > > + loff_t offset, > > + ssize_t bytes) > > +{ > > + struct ocfs2_cached_dealloc_ctxt dealloc; > > + struct ocfs2_extent_tree et; > > + struct ocfs2_super *osb = OCFS2_SB(inode->i_sb); > > + struct ocfs2_inode_info *oi = OCFS2_I(inode); > > + struct buffer_head *di_bh = NULL; > > + struct ocfs2_dinode *di; > > + struct ocfs2_alloc_context *data_ac = NULL; > > + struct ocfs2_alloc_context *meta_ac = NULL; > > + handle_t *handle = NULL; > > + loff_t end = offset + bytes; > > + int ret = 0, credits = 0; > > + struct ocfs2_map_block map; > > + unsigned int blkbits = inode->i_blkbits; > > + unsigned int max_blocks; > > + unsigned int ue_cpos = 0, ue_phys = 0, ue_len = 0; > > + unsigned int curr_lblk, end_lblk; > > + > > + map.lblk = offset >> blkbits; > > + max_blocks = (bytes + offset) >> osb->s_clustersize_bits; > > Seems unused. > Yes. > ... > > > + curr_lblk = offset >> blkbits; > > + /* > > + * Round the end up so the final partial block (sub-block direct I/O) > > + * is included; otherwise the last, partially-written cluster is left > > + * unwritten and reads back as zero. > > + */ > > + end_lblk = (offset + bytes + (1 << blkbits) - 1) >> blkbits; > > + while (ret >= 0 && curr_lblk < end_lblk) { > > + memset(&map, 0, sizeof(map)); > > + map.lblk += curr_lblk; > > Since map is memset just now, so here we can use "map.lblk = curr_lblk" > directly. Agree. > > > + map.len = end_lblk - curr_lblk; > > + > > + ret = ocfs2_assure_trans_credits(handle, credits); > > + if (ret < 0) { > > + mlog_errno(ret); > > + break; > > + } > > + > > ... > > > + > > +static int ocfs2_dio_write_end_io(struct kiocb *iocb, ssize_t size, > > + int error, unsigned int flags, int level) > > +{ > > + struct inode *inode = file_inode(iocb->ki_filp); > > + loff_t offset = iocb->ki_pos; > > + int ret = 0; > > + > > + if (error) > > + mlog_ratelimited(ML_ERROR, "Direct IO failed, bytes = %lld errno:%d", > > + (long long)size, error); > > + > > + if (size && ((flags & IOMAP_DIO_UNWRITTEN) || > > + (offset + size > i_size_read(inode)))) { > > Is it safe to do i_size_read() in case async dio completion? > I set IOMAP_DIO_FORCE_WAIT in ocfs2_file_write_iter() when the inode extends the i_size. Without IOMAP_DIO_FORCE_WAIT, xfstests will fail in many cases. Thanks, Heming