Linux Btrfs filesystem development
 help / color / mirror / Atom feed
From: Ioannis Angelakopoulos <iangelak@fb.com>
To: Wang Yugui <wangyugui@e16-tech.com>
Cc: "linux-btrfs@vger.kernel.org" <linux-btrfs@vger.kernel.org>,
	Kernel Team <Kernel-team@fb.com>
Subject: Re: [PATCH v3 1/6] btrfs: Add a lockdep model for the num_writers wait event
Date: Thu, 21 Jul 2022 16:37:09 +0000	[thread overview]
Message-ID: <3d874e37-43ee-876a-6328-52bc61cd07d8@fb.com> (raw)
In-Reply-To: <20220721084245.2C65.409509F4@e16-tech.com>

On 7/20/22 5:42 PM, Wang Yugui wrote:
> Hi,
> 
> 
>> Annotate the num_writers wait event in fs/btrfs/transaction.c with lockdep
>> in order to catch deadlocks involving this wait event.
>>
>> Use a read/write lockdep map for the annotation. A thread starting/joining
>> the transaction acquires the map as a reader when it increments
>> cur_trans->num_writers and it acquires the map as a writer before it
>> blocks on the wait event.
>>
>> Signed-off-by: Ioannis Angelakopoulos <iangelak@fb.com>
>> ---
>>   fs/btrfs/ctree.h       | 47 ++++++++++++++++++++++++++++++++++++++++++
>>   fs/btrfs/disk-io.c     |  2 ++
>>   fs/btrfs/transaction.c | 37 ++++++++++++++++++++++++++++-----
>>   3 files changed, 81 insertions(+), 5 deletions(-)
>>
>> diff --git a/fs/btrfs/ctree.h b/fs/btrfs/ctree.h
>> index 202496172059..d4d69c0e001e 100644
>> --- a/fs/btrfs/ctree.h
>> +++ b/fs/btrfs/ctree.h
>> @@ -1095,6 +1095,8 @@ struct btrfs_fs_info {
>>   	/* Updates are not protected by any lock */
>>   	struct btrfs_commit_stats commit_stats;
>>   
>> +	struct lockdep_map btrfs_trans_num_writers_map;
>> +
>>   #ifdef CONFIG_BTRFS_FS_REF_VERIFY
>>   	spinlock_t ref_verify_lock;
>>   	struct rb_root block_tree;
>> @@ -1175,6 +1177,51 @@ enum {
>>   	BTRFS_ROOT_UNFINISHED_DROP,
>>   };
>>   
>> +/*
>> + * Lockdep annotation for wait events.
>> + *
>> + * @b: The struct where the lockdep map is defined
>> + * @lock: The lockdep map corresponding to a wait event
>> + *
>> + * This macro is used to annotate a wait event. In this case a thread acquires
>> + * the lockdep map as writer (exclusive lock) because it has to block until all
>> + * the threads that hold the lock as readers signal the condition for the wait
>> + * event and release their locks.
>> + */
>> +#define btrfs_might_wait_for_event(b, lock)					\
>> +	do {									\
>> +		rwsem_acquire(&b->lock##_map, 0, 0, _THIS_IP_);			\
>> +		rwsem_release(&b->lock##_map, _THIS_IP_);			\
>> +	} while (0)
>> +
>> +/*
>> + * Protection for the resource/condition of a wait event.
>> + *
>> + * @b: The struct where the lockdep map is defined
>> + * @lock: The lockdep map corresponding to a wait event
>> + *
>> + * Many threads can modify the condition for the wait event at the same time
>> + * and signal the threads that block on the wait event. The threads that
>> + * modify the condition and do the signaling acquire the lock as readers
>> + * (shared lock).
>> + */
>> +#define btrfs_lockdep_acquire(b, lock)						\
>> +	rwsem_acquire_read(&b->lock##_map, 0, 0, _THIS_IP_)
>> +
>> +/*
>> + * Used after signaling the condition for a wait event to release the
>> + * lockdep map held by a reader thread.
>> + */
>> +#define btrfs_lockdep_release(b, lock)						\
>> +	rwsem_release(&b->lock##_map, _THIS_IP_)
>> +
>> +/* Initialization of the lockdep map */
>> +#define btrfs_lockdep_init_map(b, lock)                                        \
>> +	do {									\
>> +		static struct lock_class_key lock##_key;			\
>> +		lockdep_init_map(&b->lock##_map, #lock, &lock##_key, 0);	\
>> +	} while (0)
>> +
>>   static inline void btrfs_wake_unfinished_drop(struct btrfs_fs_info *fs_info)
>>   {
>>   	clear_and_wake_up_bit(BTRFS_FS_UNFINISHED_DROPS, &fs_info->flags);
>> diff --git a/fs/btrfs/disk-io.c b/fs/btrfs/disk-io.c
>> index 3fac429cf8a4..38831c730d61 100644
>> --- a/fs/btrfs/disk-io.c
>> +++ b/fs/btrfs/disk-io.c
>> @@ -3074,6 +3074,8 @@ void btrfs_init_fs_info(struct btrfs_fs_info *fs_info)
>>   	mutex_init(&fs_info->zoned_data_reloc_io_lock);
>>   	seqlock_init(&fs_info->profiles_lock);
>>   
>> +	btrfs_lockdep_init_map(fs_info, btrfs_trans_num_writers);
>> +
>>   	INIT_LIST_HEAD(&fs_info->dirty_cowonly_roots);
>>   	INIT_LIST_HEAD(&fs_info->space_info);
>>   	INIT_LIST_HEAD(&fs_info->tree_mod_seq_list);
>> diff --git a/fs/btrfs/transaction.c b/fs/btrfs/transaction.c
>> index 0bec10740ad3..d8287ec890bc 100644
>> --- a/fs/btrfs/transaction.c
>> +++ b/fs/btrfs/transaction.c
>> @@ -313,6 +313,7 @@ static noinline int join_transaction(struct btrfs_fs_info *fs_info,
>>   		atomic_inc(&cur_trans->num_writers);
>>   		extwriter_counter_inc(cur_trans, type);
>>   		spin_unlock(&fs_info->trans_lock);
>> +		btrfs_lockdep_acquire(fs_info, btrfs_trans_num_writers);
>>   		return 0;
>>   	}
>>   	spin_unlock(&fs_info->trans_lock);
>> @@ -334,16 +335,20 @@ static noinline int join_transaction(struct btrfs_fs_info *fs_info,
>>   	if (!cur_trans)
>>   		return -ENOMEM;
>>   
>> +	btrfs_lockdep_acquire(fs_info, btrfs_trans_num_writers);
>> +
>>   	spin_lock(&fs_info->trans_lock);
>>   	if (fs_info->running_transaction) {
>>   		/*
>>   		 * someone started a transaction after we unlocked.  Make sure
>>   		 * to redo the checks above
>>   		 */
>> +		btrfs_lockdep_release(fs_info, btrfs_trans_num_writers);
>>   		kfree(cur_trans);
>>   		goto loop;
>>   	} else if (BTRFS_FS_ERROR(fs_info)) {
>>   		spin_unlock(&fs_info->trans_lock);
>> +		btrfs_lockdep_release(fs_info, btrfs_trans_num_writers);
>>   		kfree(cur_trans);
>>   		return -EROFS;
>>   	}
>> @@ -1022,6 +1027,9 @@ static int __btrfs_end_transaction(struct btrfs_trans_handle *trans,
>>   	extwriter_counter_dec(cur_trans, trans->type);
>>   
>>   	cond_wake_up(&cur_trans->writer_wait);
>> +
>> +	btrfs_lockdep_release(info, btrfs_trans_num_writers);
>> +
>>   	btrfs_put_transaction(cur_trans);
>>   
>>   	if (current->journal_info == trans)
>> @@ -1994,6 +2002,12 @@ static void cleanup_transaction(struct btrfs_trans_handle *trans, int err)
>>   	if (cur_trans == fs_info->running_transaction) {
>>   		cur_trans->state = TRANS_STATE_COMMIT_DOING;
>>   		spin_unlock(&fs_info->trans_lock);
>> +
>> +		/*
>> +		 * The thread has already released the lockdep map as reader already in
>> +		 * btrfs_commit_transaction().
>> +		 */
>> +		btrfs_might_wait_for_event(fs_info, btrfs_trans_num_writers);
>>   		wait_event(cur_trans->writer_wait,
>>   			   atomic_read(&cur_trans->num_writers) == 1);
>>   
>> @@ -2222,7 +2236,7 @@ int btrfs_commit_transaction(struct btrfs_trans_handle *trans)
>>   
>>   			btrfs_put_transaction(prev_trans);
>>   			if (ret)
>> -				goto cleanup_transaction;
>> +				goto lockdep_release;
>>   		} else {
>>   			spin_unlock(&fs_info->trans_lock);
>>   		}
>> @@ -2236,7 +2250,7 @@ int btrfs_commit_transaction(struct btrfs_trans_handle *trans)
>>   		 */
>>   		if (BTRFS_FS_ERROR(fs_info)) {
>>   			ret = -EROFS;
>> -			goto cleanup_transaction;
>> +			goto lockdep_release;
>>   		}
>>   	}
>>   
>> @@ -2250,19 +2264,21 @@ int btrfs_commit_transaction(struct btrfs_trans_handle *trans)
>>   
>>   	ret = btrfs_start_delalloc_flush(fs_info);
>>   	if (ret)
>> -		goto cleanup_transaction;
>> +		goto lockdep_release;
>>   
>>   	ret = btrfs_run_delayed_items(trans);
>>   	if (ret)
>> -		goto cleanup_transaction;
>> +		goto lockdep_release;
>>   
>>   	wait_event(cur_trans->writer_wait,
>>   		   extwriter_counter_read(cur_trans) == 0);
>>   
>>   	/* some pending stuffs might be added after the previous flush. */
>>   	ret = btrfs_run_delayed_items(trans);
>> -	if (ret)
>> +	if (ret) {
>> +		btrfs_lockdep_release(fs_info, btrfs_trans_num_writers);
>>   		goto cleanup_transaction;
>> +	}
>>   
>>   	btrfs_wait_delalloc_flush(fs_info);
>>   
>> @@ -2284,6 +2300,14 @@ int btrfs_commit_transaction(struct btrfs_trans_handle *trans)
>>   	add_pending_snapshot(trans);
>>   	cur_trans->state = TRANS_STATE_COMMIT_DOING;
>>   	spin_unlock(&fs_info->trans_lock);
>> +
>> +	/*
>> +	 * The thread has started/joined the transaction thus it holds the lockdep
>> +	 * map as a reader. It has to release it before acquiring the lockdep map
>> +	 * as a writer.
>> +	 */
>> +	btrfs_lockdep_release(fs_info, btrfs_trans_num_writers);
>> +	btrfs_might_wait_for_event(fs_info, btrfs_trans_num_writers);
>>   	wait_event(cur_trans->writer_wait,
>>   		   atomic_read(&cur_trans->num_writers) == 1);
>>   
>> @@ -2515,6 +2539,9 @@ int btrfs_commit_transaction(struct btrfs_trans_handle *trans)
>>   	cleanup_transaction(trans, ret);
>>   
>>   	return ret;
>> +lockdep_release:
>> +	btrfs_lockdep_release(fs_info, btrfs_trans_num_writers);
>> +	goto cleanup_transaction;
>>   }
> 
> 
> Could we rename 'lockdep_release' to 'cleanup_transaction_with_lockdep_release',
> and put it just before 'cleanup_transaction:'?
> 
> Best Regards
> Wang Yugui (wangyugui@e16-tech.com)
> 2022/07/21
>
Unfortunately no, since there are error paths within 
btrfs_commit_transaction() that jump to labels prior to 
cleanup_transaction after the btrfs_trans_num_writers lock is already 
released. Moving the lockdep_release label above cleanup_transaction 
would result in lockdep complaining about double lock release.

Thanks,
Ioannis
> 
> 


  reply	other threads:[~2022-07-21 16:37 UTC|newest]

Thread overview: 10+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2022-07-20 23:38 [PATCH v3 0/6] btrfs: Annotate wait events with lockdep Ioannis Angelakopoulos
2022-07-20 23:38 ` [PATCH v3 1/6] btrfs: Add a lockdep model for the num_writers wait event Ioannis Angelakopoulos
2022-07-21  0:42   ` Wang Yugui
2022-07-21 16:37     ` Ioannis Angelakopoulos [this message]
2022-07-20 23:38 ` [PATCH v3 2/6] btrfs: Add a lockdep model for the num_extwriters " Ioannis Angelakopoulos
2022-07-20 23:38 ` [PATCH v3 3/6] btrfs: Add lockdep models for the transaction states wait events Ioannis Angelakopoulos
2022-07-20 23:38 ` [PATCH v3 4/6] btrfs: Add a lockdep model for the pending_ordered wait event Ioannis Angelakopoulos
2022-07-20 23:38 ` [PATCH v3 5/6] btrfs: Change the lockdep class of struct inode's invalidate_lock Ioannis Angelakopoulos
2022-07-20 23:38 ` [PATCH v3 6/6] btrfs: Add a lockdep model for the ordered extents wait event Ioannis Angelakopoulos
2022-07-22 13:36 ` [PATCH v3 0/6] btrfs: Annotate wait events with lockdep Josef Bacik

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=3d874e37-43ee-876a-6328-52bc61cd07d8@fb.com \
    --to=iangelak@fb.com \
    --cc=Kernel-team@fb.com \
    --cc=linux-btrfs@vger.kernel.org \
    --cc=wangyugui@e16-tech.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox