From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 455FA1F5821 for ; Fri, 21 Aug 2026 00:55:30 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787273733; cv=none; b=JK1wWtjgL4+ef5FT9Tme8Oke46XGlG3TrKHfebKru91nxasqX5tFPC8fVyWEAUVBTC4ipykpxMDV3kIYQdWumJab8QKl6oSxhLL7Zf8sOsTICjPlR8G9Sd0tHc89GdLf9qBfihJh8iClxzS+THBNxq47S6SxCsyPGbqa+ptCikg= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787273733; c=relaxed/simple; bh=YMGgr3qAKY/ouHrYFvIDolsF0m9d2XnRepYFXYD9+vY=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=UBIn8x8gjxBFtfLiPY5BWH8Y96e3DOmzld4VSBm5NS4qY4Wz5hmzCV9p6zqJwPhthoV3PTx9D/xgt0Q9+OF9h2PqmsjUwjmrGe5+J90wyW4aLXoIHKTtJXhedGIrLrJx2MftEhOuE/wx/UXjIYT6332BYsPNfHchGfMuWTRm4C8= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Ooc9lDL1; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Ooc9lDL1" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 9C9211F000E9; Fri, 21 Aug 2026 00:55:30 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel.org; s=k20260515; t=1787273730; bh=ya0Rjel8eZCmLznav0Kn62T/kM8xoxEPcb3L5crht1Q=; h=Date:From:To:Cc:Subject:References:In-Reply-To; b=Ooc9lDL1hCxUwLf44tqFmDBweBW7a363AkZM0PFUBODspvS3jjZs6HWs29VVMAl+n h5ysyB3Jg7aR5kOVm1K9EkQyShCZCcC983Guyxb/u/6aKj0hITbJqaH9bTwdPGUONV qJOIqtt9hEd0Qqvivnc5IQsgpKbeDHvvM8zwkPgcxMzc7ZNR1Lkj7JXkLjxUwTZZBw I1iqdj+3b3GY9ix/SUszQnekPm1OZ667SSf1jaB/kLt6WIdin04mD2Nckaq4niB8pz 2f6/7TBnMoev+o2kahz+o4ca2eNbXk0SvK5PRUOmqXXkf3dy5QozTIWoCpIxToGHT6 0ogfAH9DpQiPw== Date: Fri, 21 Aug 2026 00:55:29 +0000 From: Jaegeuk Kim To: Daeho Jeong Cc: linux-kernel@vger.kernel.org, linux-f2fs-devel@lists.sourceforge.net, kernel-team@android.com, Daeho Jeong , Chao Yu Subject: Re: [PATCH v8] f2fs: support dynamic reserve/release for device aliasing Message-ID: References: <20260820211641.1273552-1-daeho43@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20260820211641.1273552-1-daeho43@gmail.com> Hi Daeho, Could you please post another patch to fix the original one? Thanks, On 08/20, Daeho Jeong wrote: > From: Daeho Jeong > > This patch adds a dynamic management feature to the existing device > aliasing functionality. It allows users to dynamically reserve or > release specific devices from the filesystem's free pool at runtime > through new ioctls. > > To support this, three new ioctls are introduced: > - F2FS_IOC_RESERVE_DEV_ALIAS: This reclaims the space occupied by a > device aliasing file. It first performs a capacity check, resets GC > victim information for the target range, marks the segments as in-use > to prevent new allocations, and then triggers GC to migrate existing > valid data out of the range. Finally, it reserves these blocks in the > SIT to effectively exclude the device from the usable capacity. > > - F2FS_IOC_RELEASE_DEV_ALIAS: This releases the reserved space of a > previously reserved device aliasing file. It truncates the blocks > associated with the file, which makes them available for general > filesystem allocation again. > > - F2FS_IOC_GET_DEV_ALIAS_STATUS: This retrieves the current aliasing > status of a device aliasing file, returning whether the file is > released (inactive alias) or reserved (active alias, with blocks > fully allocated on the device). > > Signed-off-by: Daeho Jeong > Reviewed-by: Chao Yu > --- > v8: fixed device aliasing checks in f2fs_rename and f2fs_unlink. > v7: fixed extent alignment check to accomodate zoned storage alignment. > v6: fixed duplicated block reservations and racing w/ mmap. > added the extent alignment check for device aliasing areas. > v5: prevented new segment allocation on devices undergoing reservation > or with active aliases. > used in-memory struct to reserve blocks for device aliasing. > cleared prefree (PRE) dirty segment bitmap. > v4: renamed interfaces. > fixed race conditions between checkpoint=disable mount and ioctls. > refactored segment reservation part. > modified lock usage. > v3: add CAP_SYS_ADMIN and checkpoint=disabled check. > remove a f2fs specific flag exposed with getflags. > v2: prevent operations during checkpoint=disabled. > --- > Documentation/filesystems/f2fs.rst | 35 ++++ > fs/f2fs/data.c | 4 +- > fs/f2fs/extent_cache.c | 9 + > fs/f2fs/f2fs.h | 19 +- > fs/f2fs/file.c | 273 ++++++++++++++++++++++++++++- > fs/f2fs/gc.c | 30 ++-- > fs/f2fs/namei.c | 18 ++ > fs/f2fs/segment.c | 176 +++++++++++++------ > fs/f2fs/segment.h | 22 +++ > fs/f2fs/super.c | 36 ++++ > include/uapi/linux/f2fs.h | 7 + > 11 files changed, 561 insertions(+), 68 deletions(-) > > diff --git a/Documentation/filesystems/f2fs.rst b/Documentation/filesystems/f2fs.rst > index 8c4a14ae444f..1a5fd4afe609 100644 > --- a/Documentation/filesystems/f2fs.rst > +++ b/Documentation/filesystems/f2fs.rst > @@ -1045,6 +1045,41 @@ So, the key idea is, user can do any file operations on /dev/vdc, and > reclaim the space after the use, while the space is counted as /data. > That doesn't require modifying partition size and filesystem format. > > +Dynamic Device Aliasing Management > +---------------------------------- > + > +In addition to static device aliasing by deleting the aliasing file, F2FS > +supports dynamic management of device aliasing. This mechanism allows the system > +to dynamically transition partition ownership between F2FS userdata and external > +entities (e.g., zRAM, raw partition) based on system requirements without > +deleting the master aliasing file or requiring unmount/remount. > + > +The master aliasing file is created during the initial format of the file system > +and remains as a persistent control entity (ioctl gateway) in the root directory. > + > +- Partition Reservation (In-service to Aliased) > + When a specific partition needs to be dedicated to external services (e.g., zRAM), > + a user can reserve the device alias range via ioctl. The kernel resets GC victim > + information for the target range, marks segments as in-use to prevent new > + allocations, and triggers forced GC to migrate existing valid data out of the > + range. Finally, it reserves these blocks in the SIT to effectively exclude the > + device from the usable capacity. > + > +- Partition Release (Aliased to In-service) > + When external usage concludes, the space is reclaimed not by deleting the file, > + but through the release ioctl. The kernel truncates blocks associated with > + the file, releasing them back to general filesystem allocation. > + > +.. code-block:: > + > + # f2fs_io dev_alias release /mnt/f2fs/vdc.file > + # df -h > + /dev/vdb 64G 753M 64G 2% /mnt/f2fs > + > + # f2fs_io dev_alias reserve /mnt/f2fs/vdc.file > + # df -h > + /dev/vdb 64G 33G 32G 52% /mnt/f2fs > + > Per-file Read-Only Large Folio Support > -------------------------------------- > > diff --git a/fs/f2fs/data.c b/fs/f2fs/data.c > index 30e8084da313..c219ea76a3a7 100644 > --- a/fs/f2fs/data.c > +++ b/fs/f2fs/data.c > @@ -1270,7 +1270,7 @@ int f2fs_reserve_new_blocks(struct dnode_of_data *dn, blkcnt_t count) > > if (unlikely(is_inode_flag_set(dn->inode, FI_NO_ALLOC))) > return -EPERM; > - err = inc_valid_block_count(sbi, dn->inode, &count, true); > + err = inc_valid_block_count(sbi, dn->inode, &count, true, false); > if (unlikely(err)) > return err; > > @@ -1543,7 +1543,7 @@ static int __allocate_data_block(struct dnode_of_data *dn, int seg_type) > > dn->data_blkaddr = f2fs_data_blkaddr(dn); > if (dn->data_blkaddr == NULL_ADDR) { > - err = inc_valid_block_count(sbi, dn->inode, &count, true); > + err = inc_valid_block_count(sbi, dn->inode, &count, true, false); > if (unlikely(err)) > return err; > } > diff --git a/fs/f2fs/extent_cache.c b/fs/f2fs/extent_cache.c > index 61f6b9714366..37cf9fa8d537 100644 > --- a/fs/f2fs/extent_cache.c > +++ b/fs/f2fs/extent_cache.c > @@ -17,6 +17,7 @@ > > #include "f2fs.h" > #include "node.h" > +#include "segment.h" > #include > > bool sanity_check_extent_cache(struct inode *inode, struct folio *ifolio) > @@ -62,6 +63,14 @@ bool sanity_check_extent_cache(struct inode *inode, struct folio *ifolio) > __func__, inode->i_ino, ei.blk, ei.fofs, ei.len); > return false; > } > + > + if ((GET_SEGOFF_FROM_SEG0(sbi, ei.blk) % BLKS_PER_SEC(sbi)) || > + (ei.len % BLKS_PER_SEC(sbi))) { > + f2fs_warn(sbi, "%s: device alias inode (ino=%llx)'s extent info [%u, %u, %u] is not aligned to section size %u", > + __func__, inode->i_ino, ei.blk, ei.fofs, ei.len, > + BLKS_PER_SEC(sbi)); > + return false; > + } > return true; > } > > diff --git a/fs/f2fs/f2fs.h b/fs/f2fs/f2fs.h > index 8e2fb0bda467..a380b8be8819 100644 > --- a/fs/f2fs/f2fs.h > +++ b/fs/f2fs/f2fs.h > @@ -1404,6 +1404,8 @@ struct f2fs_dev_info { > unsigned int total_segments; > block_t start_blk; > block_t end_blk; > + bool has_alias; > + bool is_reserving; > #ifdef CONFIG_BLK_DEV_ZONED > unsigned int nr_blkz; /* Total number of zones */ > unsigned long *blkz_seq; /* Bitmap indicating sequential zones */ > @@ -1884,6 +1886,7 @@ struct f2fs_sb_info { > block_t last_valid_block_count; /* for recovery */ > block_t reserved_blocks; /* configurable reserved blocks */ > block_t current_reserved_blocks; /* current reserved blocks */ > + block_t alias_reserved_blocks; /* reserved blocks for device alias */ > > /* Additional tracking for no checkpoint mode */ > block_t unusable_block_count; /* # of blocks saved by last cp */ > @@ -2586,7 +2589,8 @@ static inline unsigned int get_available_block_count(struct f2fs_sb_info *sbi, > block_t avail_user_block_count; > > avail_user_block_count = sbi->user_block_count - > - sbi->current_reserved_blocks; > + sbi->current_reserved_blocks - > + sbi->alias_reserved_blocks; > > if (test_opt(sbi, RESERVE_ROOT) && !__allow_reserved_root(sbi, inode, cap)) > avail_user_block_count -= F2FS_OPTION(sbi).root_reserved_blocks; > @@ -2603,7 +2607,8 @@ static inline unsigned int get_available_block_count(struct f2fs_sb_info *sbi, > > static inline void f2fs_i_blocks_write(struct inode *, block_t, bool, bool); > static inline int inc_valid_block_count(struct f2fs_sb_info *sbi, > - struct inode *inode, blkcnt_t *count, bool partial) > + struct inode *inode, blkcnt_t *count, > + bool partial, bool alias_reserved) > { > long long diff = 0, release = 0; > block_t avail_user_block_count; > @@ -2626,10 +2631,16 @@ static inline int inc_valid_block_count(struct f2fs_sb_info *sbi, > > spin_lock(&sbi->stat_lock); > > + if (alias_reserved) > + sbi->alias_reserved_blocks -= *count; > + > avail_user_block_count = get_available_block_count(sbi, inode, true); > diff = (long long)sbi->total_valid_block_count + *count - > avail_user_block_count; > if (unlikely(diff > 0)) { > + if (alias_reserved) > + sbi->alias_reserved_blocks += *count; > + > if (!partial) { > spin_unlock(&sbi->stat_lock); > release = *count; > @@ -4037,6 +4048,8 @@ int f2fs_flush_device_cache(struct f2fs_sb_info *sbi); > void f2fs_destroy_flush_cmd_control(struct f2fs_sb_info *sbi, bool free); > void f2fs_invalidate_blocks(struct f2fs_sb_info *sbi, block_t addr, > unsigned int len); > +void f2fs_reserve_device_alias(struct f2fs_sb_info *sbi, block_t addr, > + unsigned int len); > bool f2fs_is_checkpointed_data(struct f2fs_sb_info *sbi, block_t blkaddr); > int f2fs_start_discard_thread(struct f2fs_sb_info *sbi); > void f2fs_drop_discard_cmd(struct f2fs_sb_info *sbi); > @@ -4258,6 +4271,8 @@ void f2fs_build_gc_manager(struct f2fs_sb_info *sbi); > int f2fs_gc_range(struct f2fs_sb_info *sbi, > unsigned int start_seg, unsigned int end_seg, > bool dry_run, unsigned int dry_run_sections); > +void f2fs_reset_gc_victim_resource(struct f2fs_sb_info *sbi, > + unsigned int start, unsigned int end); > int f2fs_resize_fs(struct file *filp, __u64 block_count); > int __init f2fs_create_garbage_collection_cache(void); > void f2fs_destroy_garbage_collection_cache(void); > diff --git a/fs/f2fs/file.c b/fs/f2fs/file.c > index b99d9cdf9ba7..56529a82e027 100644 > --- a/fs/f2fs/file.c > +++ b/fs/f2fs/file.c > @@ -813,13 +813,19 @@ int f2fs_do_truncate_blocks(struct inode *inode, u64 from, bool lock) > > if (IS_DEVICE_ALIASING(inode)) { > struct extent_tree *et = F2FS_I(inode)->extent_tree[EX_READ]; > - struct extent_info ei = et->largest; > + struct extent_info ei; > + > + read_lock(&et->lock); > + ei = et->largest; > + read_unlock(&et->lock); > > f2fs_invalidate_blocks(sbi, ei.blk, ei.len); > > dec_valid_block_count(sbi, inode, ei.len); > f2fs_update_time(sbi, REQ_TIME); > > + f2fs_drop_extent_tree(inode); > + > f2fs_folio_put(ifolio, true); > goto out; > } > @@ -1100,8 +1106,9 @@ int f2fs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, > if ((attr->ia_valid & ATTR_SIZE)) { > if (mapping_large_folio_support(inode->i_mapping)) > return -EOPNOTSUPP; > - if (!f2fs_is_compress_backend_ready(inode) || > - IS_DEVICE_ALIASING(inode)) > + if (IS_DEVICE_ALIASING(inode)) > + return -EPERM; > + if (!f2fs_is_compress_backend_ready(inode)) > return -EOPNOTSUPP; > if (is_inode_flag_set(inode, FI_COMPRESS_RELEASED) && > !IS_ALIGNED(attr->ia_size, > @@ -2139,6 +2146,9 @@ static int f2fs_setflags_common(struct inode *inode, u32 iflags, u32 mask) > if (IS_NOQUOTA(inode)) > return -EPERM; > > + if (IS_DEVICE_ALIASING(inode)) > + return -EPERM; > + > if ((iflags ^ masked_flags) & F2FS_CASEFOLD_FL) { > if (!f2fs_sb_has_casefold(F2FS_I_SB(inode))) > return -EOPNOTSUPP; > @@ -2687,6 +2697,17 @@ static int f2fs_ioc_get_encryption_policy(struct file *filp, unsigned long arg) > return fscrypt_ioctl_get_policy(filp, (void __user *)arg); > } > > +static int f2fs_ioc_get_dev_alias_status(struct file *filp, unsigned long arg) > +{ > + struct inode *inode = file_inode(filp); > + > + if (!IS_DEVICE_ALIASING(inode)) > + return -EINVAL; > + > + return put_user(F2FS_HAS_BLOCKS(inode) ? F2FS_DEV_ALIAS_STATUS_RESERVED : > + F2FS_DEV_ALIAS_STATUS_RELEASED, (u32 __user *)arg); > +} > + > static int f2fs_ioc_get_encryption_pwsalt(struct file *filp, unsigned long arg) > { > struct inode *inode = file_inode(filp); > @@ -3637,6 +3658,241 @@ static int f2fs_ioc_get_dev_alias_file(struct file *filp, unsigned long arg) > (u32 __user *)arg); > } > > +static bool f2fs_get_dev_alias_extent(struct f2fs_sb_info *sbi, > + struct dentry *dentry, > + struct extent_info *ei) > +{ > + int i; > + > + for (i = 1; i < sbi->s_ndevs; i++) { > + char *name = strrchr(FDEV(i).path, '/'); > + > + name = name ? name + 1 : FDEV(i).path; > + if (strcmp(name, dentry->d_name.name)) > + continue; > + > + ei->blk = FDEV(i).start_blk; > + ei->len = FDEV(i).total_segments << sbi->log_blocks_per_seg; > + ei->fofs = 0; > + return true; > + } > + return false; > +} > + > +static int f2fs_ioc_reserve_dev_alias(struct file *filp) > +{ > + struct inode *inode = file_inode(filp); > + struct f2fs_sb_info *sbi = F2FS_I_SB(inode); > + struct extent_tree *et = F2FS_I(inode)->extent_tree[EX_READ]; > + struct extent_info ei; > + struct cp_control cpc = { CP_SYNC, 0, 0, 0 }; > + struct f2fs_lock_context lc, glc; > + blkcnt_t count; > + unsigned int start, end; > + int type, err; > + > + if (!capable(CAP_SYS_ADMIN)) > + return -EPERM; > + > + if (unlikely(is_sbi_flag_set(sbi, SBI_CP_DISABLED))) > + return -EINVAL; > + > + err = mnt_want_write_file(filp); > + if (err) > + return err; > + > + inode_lock(inode); > + > + if (!IS_DEVICE_ALIASING(inode)) { > + err = -EINVAL; > + goto out_inode_unlock; > + } > + > + if (F2FS_HAS_BLOCKS(inode)) { > + err = 0; > + goto out_inode_unlock; > + } > + > + if (!f2fs_get_dev_alias_extent(sbi, filp->f_path.dentry, &ei)) { > + f2fs_warn(sbi, "device alias file (%s, ino=%llu) has no matching device", > + filp->f_path.dentry->d_name.name, > + (unsigned long long)inode->i_ino); > + set_sbi_flag(sbi, SBI_NEED_FSCK); > + f2fs_handle_error(sbi, ERROR_CORRUPTED_INODE); > + err = -EFSCORRUPTED; > + goto out_inode_unlock; > + } > + > + spin_lock(&sbi->stat_lock); > + if (sbi->total_valid_block_count + ei.len > > + get_available_block_count(sbi, inode, true)) { > + spin_unlock(&sbi->stat_lock); > + err = -ENOSPC; > + goto out_inode_unlock; > + } > + sbi->alias_reserved_blocks += ei.len; > + spin_unlock(&sbi->stat_lock); > + > + spin_lock(&FREE_I(sbi)->segmap_lock); > + FDEV(f2fs_target_device_index(sbi, ei.blk)).is_reserving = true; > + spin_unlock(&FREE_I(sbi)->segmap_lock); > + > + start = GET_SEGNO(sbi, ei.blk); > + end = GET_SEGNO(sbi, ei.blk + ei.len - 1); > + > + /* Acquire gc_lock for victim reset, curseg resize, and range GC */ > + f2fs_down_write_trace(&sbi->gc_lock, &glc); > + > + /* Reset the victim information to prevent GC from targeting the range */ > + f2fs_reset_gc_victim_resource(sbi, start, end); > + > + /* Move out cursegs from the target range */ > + for (type = CURSEG_HOT_DATA; type < NR_CURSEG_PERSIST_TYPE; type++) { > + err = f2fs_allocate_segment_for_resize(sbi, type, start, end); > + if (err) > + goto out_gc_unlock; > + } > + > + f2fs_lock_op(sbi, &lc); > + > + if (unlikely(is_sbi_flag_set(sbi, SBI_CP_DISABLED))) { > + err = -EINVAL; > + f2fs_unlock_op(sbi, &lc); > + goto out_gc_unlock; > + } > + > + /* do GC to move out valid blocks in the range all at once! */ > + err = f2fs_gc_range(sbi, start, end, false, 0); > + if (err) { > + f2fs_unlock_op(sbi, &lc); > + goto out_gc_unlock; > + } > + > + count = ei.len; > + err = inc_valid_block_count(sbi, inode, &count, false, true); > + if (err) { > + f2fs_unlock_op(sbi, &lc); > + goto out_gc_unlock; > + } > + > + write_lock(&et->lock); > + et->largest = ei; > + write_unlock(&et->lock); > + clear_inode_flag(inode, FI_NO_EXTENT); > + > + f2fs_reserve_device_alias(sbi, ei.blk, ei.len); > + > + i_size_write(inode, (loff_t)ei.len << sbi->log_blocksize); > + f2fs_update_inode_page(inode); > + > + spin_lock(&FREE_I(sbi)->segmap_lock); > + FDEV(f2fs_target_device_index(sbi, ei.blk)).is_reserving = false; > + spin_unlock(&FREE_I(sbi)->segmap_lock); > + > + f2fs_unlock_op(sbi, &lc); > + f2fs_up_write_trace(&sbi->gc_lock, &glc); > + > + inode_unlock(inode); > + mnt_drop_write_file(filp); > + > + return f2fs_write_checkpoint(sbi, &cpc); > + > +out_gc_unlock: > + spin_lock(&sbi->stat_lock); > + sbi->alias_reserved_blocks -= ei.len; > + spin_unlock(&sbi->stat_lock); > + > + spin_lock(&FREE_I(sbi)->segmap_lock); > + FDEV(f2fs_target_device_index(sbi, ei.blk)).is_reserving = false; > + spin_unlock(&FREE_I(sbi)->segmap_lock); > + f2fs_up_write_trace(&sbi->gc_lock, &glc); > + > +out_inode_unlock: > + inode_unlock(inode); > + mnt_drop_write_file(filp); > + return err; > +} > + > +static int f2fs_ioc_release_dev_alias(struct file *filp) > +{ > + struct inode *inode = file_inode(filp); > + struct f2fs_sb_info *sbi = F2FS_I_SB(inode); > + struct extent_tree *et = F2FS_I(inode)->extent_tree[EX_READ]; > + struct extent_info ei = {0, }; > + struct cp_control cpc = { CP_SYNC, 0, 0, 0 }; > + struct f2fs_lock_context lc, glc; > + int err; > + > + if (!capable(CAP_SYS_ADMIN)) > + return -EPERM; > + > + if (unlikely(is_sbi_flag_set(sbi, SBI_CP_DISABLED))) > + return -EINVAL; > + > + err = mnt_want_write_file(filp); > + if (err) > + return err; > + > + inode_lock(inode); > + > + if (!IS_DEVICE_ALIASING(inode)) { > + err = -EINVAL; > + goto out_inode_unlock; > + } > + > + if (!F2FS_HAS_BLOCKS(inode)) { > + err = 0; > + goto out_inode_unlock; > + } > + > + err = filemap_write_and_wait(inode->i_mapping); > + if (err) > + goto out_inode_unlock; > + > + read_lock(&et->lock); > + ei = et->largest; > + read_unlock(&et->lock); > + > + f2fs_down_write_trace(&sbi->gc_lock, &glc); > + f2fs_lock_op(sbi, &lc); > + > + if (unlikely(is_sbi_flag_set(sbi, SBI_CP_DISABLED))) { > + err = -EINVAL; > + f2fs_unlock_op(sbi, &lc); > + f2fs_up_write_trace(&sbi->gc_lock, &glc); > + goto out_inode_unlock; > + } > + > + filemap_invalidate_lock(inode->i_mapping); > + truncate_setsize(inode, 0); > + > + err = f2fs_truncate_blocks(inode, 0, false); > + if (err) > + i_size_write(inode, (loff_t)ei.len << sbi->log_blocksize); > + filemap_invalidate_unlock(inode->i_mapping); > + > + if (err) { > + f2fs_unlock_op(sbi, &lc); > + f2fs_up_write_trace(&sbi->gc_lock, &glc); > + goto out_inode_unlock; > + } > + > + f2fs_update_inode_page(inode); > + > + f2fs_unlock_op(sbi, &lc); > + f2fs_up_write_trace(&sbi->gc_lock, &glc); > + > + inode_unlock(inode); > + mnt_drop_write_file(filp); > + > + return f2fs_write_checkpoint(sbi, &cpc); > + > +out_inode_unlock: > + inode_unlock(inode); > + mnt_drop_write_file(filp); > + return err; > +} > + > static int f2fs_ioc_io_prio(struct file *filp, unsigned long arg) > { > struct inode *inode = file_inode(filp); > @@ -4062,7 +4318,7 @@ static int reserve_compress_blocks(struct dnode_of_data *dn, pgoff_t count, > } > > ret = inc_valid_block_count(sbi, dn->inode, > - &to_reserved, false); > + &to_reserved, false, false); > if (unlikely(ret)) > return ret; > > @@ -4763,8 +5019,14 @@ static long __f2fs_ioctl(struct file *filp, unsigned int cmd, unsigned long arg) > return f2fs_ioc_compress_file(filp); > case F2FS_IOC_GET_DEV_ALIAS_FILE: > return f2fs_ioc_get_dev_alias_file(filp, arg); > + case F2FS_IOC_GET_DEV_ALIAS_STATUS: > + return f2fs_ioc_get_dev_alias_status(filp, arg); > case F2FS_IOC_IO_PRIO: > return f2fs_ioc_io_prio(filp, arg); > + case F2FS_IOC_RESERVE_DEV_ALIAS: > + return f2fs_ioc_reserve_dev_alias(filp); > + case F2FS_IOC_RELEASE_DEV_ALIAS: > + return f2fs_ioc_release_dev_alias(filp); > default: > return -ENOTTY; > } > @@ -5551,7 +5813,10 @@ long f2fs_compat_ioctl(struct file *file, unsigned int cmd, unsigned long arg) > case F2FS_IOC_DECOMPRESS_FILE: > case F2FS_IOC_COMPRESS_FILE: > case F2FS_IOC_GET_DEV_ALIAS_FILE: > + case F2FS_IOC_GET_DEV_ALIAS_STATUS: > case F2FS_IOC_IO_PRIO: > + case F2FS_IOC_RESERVE_DEV_ALIAS: > + case F2FS_IOC_RELEASE_DEV_ALIAS: > break; > default: > return -ENOIOCTLCMD; > diff --git a/fs/f2fs/gc.c b/fs/f2fs/gc.c > index d04633f872ef..51fc8c69eb25 100644 > --- a/fs/f2fs/gc.c > +++ b/fs/f2fs/gc.c > @@ -2191,29 +2191,37 @@ int f2fs_gc_range(struct f2fs_sb_info *sbi, > return 0; > } > > +void f2fs_reset_gc_victim_resource(struct f2fs_sb_info *sbi, > + unsigned int start, unsigned int end) > +{ > + int i; > + > + mutex_lock(&DIRTY_I(sbi)->seglist_lock); > + for (i = 0; i < MAX_GC_POLICY; i++) > + if (SIT_I(sbi)->last_victim[i] >= start && > + SIT_I(sbi)->last_victim[i] <= end) > + SIT_I(sbi)->last_victim[i] = 0; > + > + for (i = BG_GC; i <= FG_GC; i++) > + if (sbi->next_victim_seg[i] >= start && > + sbi->next_victim_seg[i] <= end) > + sbi->next_victim_seg[i] = NULL_SEGNO; > + mutex_unlock(&DIRTY_I(sbi)->seglist_lock); > +} > + > static int free_segment_range(struct f2fs_sb_info *sbi, > unsigned int secs, bool dry_run) > { > unsigned int next_inuse, start, end; > struct cp_control cpc = { CP_RESIZE, 0, 0, 0 }; > - int gc_mode, gc_type; > int err = 0; > int type; > > - /* Force block allocation for GC */ > MAIN_SECS(sbi) -= secs; > start = MAIN_SECS(sbi) * SEGS_PER_SEC(sbi); > end = MAIN_SEGS(sbi) - 1; > > - mutex_lock(&DIRTY_I(sbi)->seglist_lock); > - for (gc_mode = 0; gc_mode < MAX_GC_POLICY; gc_mode++) > - if (SIT_I(sbi)->last_victim[gc_mode] >= start) > - SIT_I(sbi)->last_victim[gc_mode] = 0; > - > - for (gc_type = BG_GC; gc_type <= FG_GC; gc_type++) > - if (sbi->next_victim_seg[gc_type] >= start) > - sbi->next_victim_seg[gc_type] = NULL_SEGNO; > - mutex_unlock(&DIRTY_I(sbi)->seglist_lock); > + f2fs_reset_gc_victim_resource(sbi, start, end); > > /* Move out cursegs from the target range */ > for (type = CURSEG_HOT_DATA; type < NR_CURSEG_PERSIST_TYPE; type++) { > diff --git a/fs/f2fs/namei.c b/fs/f2fs/namei.c > index 7ffdf23cea5e..b9b15c5d28de 100644 > --- a/fs/f2fs/namei.c > +++ b/fs/f2fs/namei.c > @@ -425,6 +425,9 @@ static int f2fs_link(struct dentry *old_dentry, struct inode *dir, > if (!f2fs_is_checkpoint_ready(sbi)) > return -ENOSPC; > > + if (IS_DEVICE_ALIASING(inode)) > + return -EPERM; > + > err = fscrypt_prepare_link(old_dentry, dir, dentry); > if (err) > return err; > @@ -568,6 +571,11 @@ static int f2fs_unlink(struct inode *dir, struct dentry *dentry) > > trace_f2fs_unlink_enter(dir, dentry); > > + if (IS_DEVICE_ALIASING(inode)) { > + err = -EPERM; > + goto out; > + } > + > if (unlikely(f2fs_cp_error(sbi))) { > err = -EIO; > goto out; > @@ -946,6 +954,9 @@ static int f2fs_rename(struct mnt_idmap *idmap, struct inode *old_dir, > bool old_is_dir = S_ISDIR(old_inode->i_mode); > int err; > > + if (IS_DEVICE_ALIASING(old_inode)) > + return -EPERM; > + > if (unlikely(f2fs_cp_error(sbi))) > return -EIO; > if (!f2fs_is_checkpoint_ready(sbi)) > @@ -1016,6 +1027,10 @@ static int f2fs_rename(struct mnt_idmap *idmap, struct inode *old_dir, > } > > if (new_inode) { > + if (IS_DEVICE_ALIASING(new_inode)) { > + err = -EPERM; > + goto out_dir; > + } > > err = -ENOTEMPTY; > if (old_is_dir && !f2fs_empty_dir(new_inode)) > @@ -1143,6 +1158,9 @@ static int f2fs_cross_rename(struct inode *old_dir, struct dentry *old_dentry, > int old_nlink = 0, new_nlink = 0; > int err; > > + if (IS_DEVICE_ALIASING(old_inode) || IS_DEVICE_ALIASING(new_inode)) > + return -EPERM; > + > if (unlikely(f2fs_cp_error(sbi))) > return -EIO; > if (!f2fs_is_checkpoint_ready(sbi)) > diff --git a/fs/f2fs/segment.c b/fs/f2fs/segment.c > index 00cc2af45ff9..b81245ddd490 100644 > --- a/fs/f2fs/segment.c > +++ b/fs/f2fs/segment.c > @@ -261,7 +261,7 @@ static int __replace_atomic_write_block(struct inode *inode, pgoff_t index, > } else { > blkcnt_t count = 1; > > - err = inc_valid_block_count(sbi, inode, &count, true); > + err = inc_valid_block_count(sbi, inode, &count, true, false); > if (err) { > f2fs_put_dnode(&dn); > return err; > @@ -2539,35 +2539,42 @@ static int update_sit_entry_for_alloc(struct f2fs_sb_info *sbi, struct seg_entry > unsigned int segno, block_t blkaddr, unsigned int offset, int del) > { > bool exist; > + int del_count = del; > + int i; > > - exist = f2fs_test_and_set_bit(offset, se->cur_valid_map); > - if (unlikely(exist)) { > - f2fs_err(sbi, "Bitmap was wrongly set, blk:%u", blkaddr); > - f2fs_bug_on(sbi, 1); > - se->valid_blocks--; > - del = 0; > - } > + f2fs_bug_on(sbi, GET_SEGNO(sbi, blkaddr) != GET_SEGNO(sbi, blkaddr + del_count - 1)); > > - if (f2fs_block_unit_discard(sbi) && > - !f2fs_test_and_set_bit(offset, se->discard_map)) > - sbi->discard_blks--; > + for (i = 0; i < del_count; i++) { > + exist = f2fs_test_and_set_bit(offset + i, se->cur_valid_map); > + if (unlikely(exist)) { > + f2fs_err(sbi, "Bitmap was wrongly set, blk:%u", blkaddr + i); > + f2fs_bug_on(sbi, 1); > + se->valid_blocks--; > + del -= 1; > + continue; > + } > > - /* > - * SSR should never reuse block which is checkpointed > - * or newly invalidated. > - */ > - if (!is_sbi_flag_set(sbi, SBI_CP_DISABLED)) { > - if (!f2fs_test_and_set_bit(offset, se->ckpt_valid_map)) { > - se->ckpt_valid_blocks++; > - if (__is_large_section(sbi)) > - get_sec_entry(sbi, segno)->ckpt_valid_blocks++; > + if (f2fs_block_unit_discard(sbi) && > + !f2fs_test_and_set_bit(offset + i, se->discard_map)) > + sbi->discard_blks--; > + > + /* > + * SSR should never reuse block which is checkpointed > + * or newly invalidated. > + */ > + if (!is_sbi_flag_set(sbi, SBI_CP_DISABLED)) { > + if (!f2fs_test_and_set_bit(offset + i, se->ckpt_valid_map)) { > + se->ckpt_valid_blocks++; > + if (__is_large_section(sbi)) > + get_sec_entry(sbi, segno)->ckpt_valid_blocks++; > + } > } > - } > > - if (!f2fs_test_bit(offset, se->ckpt_valid_map)) { > - se->ckpt_valid_blocks += del; > - if (__is_large_section(sbi)) > - get_sec_entry(sbi, segno)->ckpt_valid_blocks += del; > + if (!f2fs_test_bit(offset + i, se->ckpt_valid_map)) { > + se->ckpt_valid_blocks += 1; > + if (__is_large_section(sbi)) > + get_sec_entry(sbi, segno)->ckpt_valid_blocks += 1; > + } > } > > if (__is_large_section(sbi)) > @@ -2622,9 +2629,14 @@ void f2fs_invalidate_blocks(struct f2fs_sb_info *sbi, block_t addr, > unsigned int segno = GET_SEGNO(sbi, addr); > struct sit_info *sit_i = SIT_I(sbi); > block_t addr_start = addr, addr_end = addr + len - 1; > - unsigned int seg_num = GET_SEGNO(sbi, addr_end) - segno + 1; > + unsigned int seg_num; > unsigned int i = 1, max_blocks = sbi->blocks_per_seg, cnt; > > + if (len == 0) > + return; > + > + seg_num = GET_SEGNO(sbi, addr_end) - segno + 1; > + > f2fs_bug_on(sbi, addr == NULL_ADDR); > if (addr == NEW_ADDR || addr == COMPRESS_ADDR) > return; > @@ -2657,6 +2669,52 @@ void f2fs_invalidate_blocks(struct f2fs_sb_info *sbi, block_t addr, > up_write(&sit_i->sentry_lock); > } > > +void f2fs_reserve_device_alias(struct f2fs_sb_info *sbi, block_t addr, > + unsigned int len) > +{ > + unsigned int segno = GET_SEGNO(sbi, addr); > + struct sit_info *sit_i = SIT_I(sbi); > + block_t addr_start = addr, addr_end = addr + len - 1; > + unsigned int seg_num; > + unsigned int i = 1, max_blocks = sbi->blocks_per_seg, cnt; > + > + if (len == 0) > + return; > + > + seg_num = GET_SEGNO(sbi, addr_end) - segno + 1; > + > + down_write(&sit_i->sentry_lock); > + > + if (seg_num == 1) > + cnt = len; > + else > + cnt = max_blocks - GET_BLKOFF_FROM_SEG0(sbi, addr); > + > + do { > + update_segment_mtime(sbi, addr_start, 0); > + update_sit_entry(sbi, addr_start, cnt); > + __set_test_and_inuse(sbi, segno); > + > + /* Remove the segment from PRE (prefree) to prevent checkpoint from freeing it! */ > + mutex_lock(&DIRTY_I(sbi)->seglist_lock); > + if (test_and_clear_bit(segno, DIRTY_I(sbi)->dirty_segmap[PRE])) > + DIRTY_I(sbi)->nr_dirty[PRE]--; > + mutex_unlock(&DIRTY_I(sbi)->seglist_lock); > + > + /* add it into dirty seglist */ > + locate_dirty_segment(sbi, segno); > + > + /* update @addr_start and @cnt and @segno */ > + addr_start = START_BLOCK(sbi, ++segno); > + if (++i == seg_num) > + cnt = GET_BLKOFF_FROM_SEG0(sbi, addr_end) + 1; > + else > + cnt = max_blocks; > + } while (i <= seg_num); > + > + up_write(&sit_i->sentry_lock); > +} > + > bool f2fs_is_checkpointed_data(struct f2fs_sb_info *sbi, block_t blkaddr) > { > struct sit_info *sit_i = SIT_I(sbi); > @@ -2795,8 +2853,13 @@ static int is_next_segment_free(struct f2fs_sb_info *sbi, > unsigned int segno = curseg->segno + 1; > struct free_segmap_info *free_i = FREE_I(sbi); > > - if (segno < MAIN_SEGS(sbi) && segno % SEGS_PER_SEC(sbi)) > + if (segno < MAIN_SEGS(sbi) && segno % SEGS_PER_SEC(sbi)) { > + int devi = f2fs_target_device_index(sbi, START_BLOCK(sbi, segno)); > + > + if (f2fs_dev_is_reserving(sbi, devi)) > + return 0; > return !test_bit(segno, free_i->free_segmap); > + } > return 0; > } > > @@ -2815,7 +2878,8 @@ static int get_new_segment(struct f2fs_sb_info *sbi, > unsigned int alloc_policy = sbi->allocate_section_policy; > unsigned int alloc_hint = sbi->allocate_section_hint; > bool init = true; > - int i; > + bool looped = false; > + int i, devi; > int ret = 0; > > spin_lock(&free_i->segmap_lock); > @@ -2828,8 +2892,13 @@ static int get_new_segment(struct f2fs_sb_info *sbi, > if (!new_sec && ((*newseg + 1) % SEGS_PER_SEC(sbi))) { > segno = find_next_zero_bit(free_i->free_segmap, > GET_SEG_FROM_SEC(sbi, hint + 1), *newseg + 1); > - if (segno < GET_SEG_FROM_SEC(sbi, hint + 1)) > + if (segno < GET_SEG_FROM_SEC(sbi, hint + 1)) { > + devi = f2fs_target_device_index(sbi, START_BLOCK(sbi, segno)); > + > + if (f2fs_dev_is_alloc_blocked(sbi, devi, pinning)) > + goto find_other_zone; > goto got_it; > + } > } > > #ifdef CONFIG_BLK_DEV_ZONED > @@ -2865,33 +2934,42 @@ static int get_new_segment(struct f2fs_sb_info *sbi, > find_other_zone: > secno = find_next_zero_bit(free_i->free_secmap, MAIN_SECS(sbi), hint); > > -#ifdef CONFIG_BLK_DEV_ZONED > - if (secno >= MAIN_SECS(sbi) && f2fs_sb_has_blkzoned(sbi)) { > - /* Write only to sequential zones */ > - if (sbi->blkzone_alloc_policy == BLKZONE_ALLOC_ONLY_SEQ) { > - hint = GET_SEC_FROM_SEG(sbi, sbi->first_seq_zone_segno); > - secno = find_next_zero_bit(free_i->free_secmap, MAIN_SECS(sbi), hint); > - } else > - secno = find_first_zero_bit(free_i->free_secmap, > - MAIN_SECS(sbi)); > - if (secno >= MAIN_SECS(sbi)) { > - ret = -ENOSPC; > - f2fs_bug_on(sbi, 1); > - goto out_unlock; > - } > - } > -#endif > - > if (secno >= MAIN_SECS(sbi)) { > - secno = find_first_zero_bit(free_i->free_secmap, > - MAIN_SECS(sbi)); > - if (secno >= MAIN_SECS(sbi)) { > + if (looped) { > ret = -ENOSPC; > f2fs_bug_on(sbi, !pinning); > goto out_unlock; > } > + hint = 0; > +#ifdef CONFIG_BLK_DEV_ZONED > + /* Write only to sequential zones */ > + if (f2fs_sb_has_blkzoned(sbi) && > + sbi->blkzone_alloc_policy == BLKZONE_ALLOC_ONLY_SEQ) > + hint = GET_SEC_FROM_SEG(sbi, sbi->first_seq_zone_segno); > +#endif > + looped = true; > + goto find_other_zone; > } > + > segno = GET_SEG_FROM_SEC(sbi, secno); > + > + devi = f2fs_target_device_index(sbi, START_BLOCK(sbi, segno)); > + > + if (f2fs_dev_is_alloc_blocked(sbi, devi, pinning)) { > + while (devi < sbi->s_ndevs && > + f2fs_dev_is_alloc_blocked(sbi, devi, pinning)) { > + unsigned int end_segno = GET_SEGNO(sbi, FDEV(devi).end_blk); > + > + hint = GET_SEC_FROM_SEG(sbi, end_segno) + 1; > + devi++; > + } > + goto find_other_zone; > + } > + > + if (sec_usage_check(sbi, secno)) { > + hint = secno + 1; > + goto find_other_zone; > + } > zoneno = GET_ZONE_FROM_SEC(sbi, secno); > > /* give up on finding another zone */ > diff --git a/fs/f2fs/segment.h b/fs/f2fs/segment.h > index 33a2257da1e6..db1079169a23 100644 > --- a/fs/f2fs/segment.h > +++ b/fs/f2fs/segment.h > @@ -940,10 +940,32 @@ static inline block_t sum_blk_addr(struct f2fs_sb_info *sbi, int base, int type) > - (base + 1) + type; > } > > +static inline bool f2fs_dev_is_reserving(struct f2fs_sb_info *sbi, int devi) > +{ > + if (!f2fs_sb_has_device_alias(sbi)) > + return false; > + return FDEV(devi).is_reserving; > +} > + > +static inline bool f2fs_dev_is_alloc_blocked(struct f2fs_sb_info *sbi, > + int devi, bool pinning) > +{ > + if (!f2fs_sb_has_device_alias(sbi)) > + return false; > + return (pinning && FDEV(devi).has_alias) || FDEV(devi).is_reserving; > +} > + > static inline bool sec_usage_check(struct f2fs_sb_info *sbi, unsigned int secno) > { > if (is_cursec(sbi, secno) || (sbi->cur_victim_sec == secno)) > return true; > + if (f2fs_sb_has_device_alias(sbi)) { > + block_t start_blk = START_BLOCK(sbi, GET_SEG_FROM_SEC(sbi, secno)); > + int devi = f2fs_target_device_index(sbi, start_blk); > + > + if (f2fs_dev_is_reserving(sbi, devi)) > + return true; > + } > return false; > } > > diff --git a/fs/f2fs/super.c b/fs/f2fs/super.c > index 5902b2da7ea4..ed9e0ba9cf0b 100644 > --- a/fs/f2fs/super.c > +++ b/fs/f2fs/super.c > @@ -5005,6 +5005,39 @@ static void f2fs_tuning_parameters(struct f2fs_sb_info *sbi) > sbi->readdir_ra = true; > } > > +static void f2fs_restore_device_alias(struct f2fs_sb_info *sbi) > +{ > + struct inode *root = d_inode(sbi->sb->s_root); > + struct f2fs_dir_entry *de; > + struct folio *folio; > + int i; > + > + if (!f2fs_sb_has_device_alias(sbi)) > + return; > + > + for (i = 1; i < sbi->s_ndevs; i++) { > + char *name = strrchr(FDEV(i).path, '/'); > + struct inode *inode; > + struct qstr qstr; > + > + name = name ? name + 1 : FDEV(i).path; > + qstr.name = name; > + qstr.len = strlen(name); > + > + de = f2fs_find_entry(root, &qstr, &folio); > + if (!de) > + continue; > + > + inode = f2fs_iget(sbi->sb, le32_to_cpu(de->ino)); > + if (!IS_ERR(inode)) { > + if (IS_DEVICE_ALIASING(inode)) > + FDEV(i).has_alias = true; > + iput(inode); > + } > + f2fs_folio_put(folio, 0); > + } > +} > + > static int f2fs_fill_super(struct super_block *sb, struct fs_context *fc) > { > struct f2fs_fs_context *ctx = fc->fs_private; > @@ -5209,6 +5242,7 @@ static int f2fs_fill_super(struct super_block *sb, struct fs_context *fc) > sbi->last_valid_block_count = sbi->total_valid_block_count; > sbi->reserved_blocks = 0; > sbi->current_reserved_blocks = 0; > + sbi->alias_reserved_blocks = 0; > limit_reserve_root(sbi); > adjust_unusable_cap_perc(sbi); > > @@ -5436,6 +5470,8 @@ static int f2fs_fill_super(struct super_block *sb, struct fs_context *fc) > f2fs_update_time(sbi, REQ_TIME); > clear_sbi_flag(sbi, SBI_CP_DISABLED_QUICK); > > + f2fs_restore_device_alias(sbi); > + > sbi->umount_lock_holder = NULL; > return 0; > > diff --git a/include/uapi/linux/f2fs.h b/include/uapi/linux/f2fs.h > index 795e26258355..4409ada2fecb 100644 > --- a/include/uapi/linux/f2fs.h > +++ b/include/uapi/linux/f2fs.h > @@ -45,6 +45,9 @@ > #define F2FS_IOC_START_ATOMIC_REPLACE _IO(F2FS_IOCTL_MAGIC, 25) > #define F2FS_IOC_GET_DEV_ALIAS_FILE _IOR(F2FS_IOCTL_MAGIC, 26, __u32) > #define F2FS_IOC_IO_PRIO _IOW(F2FS_IOCTL_MAGIC, 27, __u32) > +#define F2FS_IOC_RESERVE_DEV_ALIAS _IO(F2FS_IOCTL_MAGIC, 28) > +#define F2FS_IOC_RELEASE_DEV_ALIAS _IO(F2FS_IOCTL_MAGIC, 29) > +#define F2FS_IOC_GET_DEV_ALIAS_STATUS _IOR(F2FS_IOCTL_MAGIC, 30, __u32) > > /* > * should be same as XFS_IOC_GOINGDOWN. > @@ -70,6 +73,10 @@ enum { > F2FS_IOPRIO_MAX, > }; > > +/* for F2FS_IOC_GET_DEV_ALIAS_STATUS */ > +#define F2FS_DEV_ALIAS_STATUS_RELEASED 0 > +#define F2FS_DEV_ALIAS_STATUS_RESERVED 1 > + > struct f2fs_gc_range { > __u32 sync; > __u64 start; > -- > 2.55.0.766.g2966f0265a-goog >