Linux Documentation
 help / color / mirror / Atom feed
* [PATCH v2 0/3] blk-cgroup: store blkcg in bio before blkcg_mutex conversion
@ 2026-09-13  6:54 Yu Kuai
  2026-09-13  6:54 ` [PATCH v2 1/3] blk-cgroup: use a request_queue rhashtable for blkg lookup Yu Kuai
                   ` (2 more replies)
  0 siblings, 3 replies; 6+ messages in thread
From: Yu Kuai @ 2026-09-13  6:54 UTC (permalink / raw)
  To: Jens Axboe, Josef Bacik, Tejun Heo, Johannes Weiner,
	Michal Koutný, Jonathan Corbet, Shuah Khan, Randy Dunlap,
	Coly Li, Kent Overstreet, Alasdair Kergon, Mike Snitzer,
	Mikulas Patocka, Benjamin Marzinski, Song Liu, Li Nan, Xiao Ni,
	Andreas Gruenbacher, Matthew Wilcox, Jan Kara, Andrew Morton,
	Chris Li, Kairui Song, Kemeng Shi, Nhat Pham, Baoquan He,
	Barry Song, Youngjun Park, Nathan Chancellor, Nick Desaulniers,
	Bill Wendling, Justin Stitt
  Cc: Yu Kuai, Christoph Hellwig, Nilay Shroff, Tao Cui, linux-block,
	cgroups, linux-doc, linux-kernel, linux-bcache, dm-devel,
	linux-raid, gfs2, linux-fsdevel, linux-mm, llvm

From: Yu Kuai <yukuai@fygo.io>

This is v2 of the preparatory series for the blkcg_mutex conversion
proposed in the related series [1]. That conversion moves queue-local
blkg topology synchronization from q->queue_lock to q->blkcg_mutex,
which is awkward while bios directly store queue-local blkg references.

Storing a blkg also ties a bio to one request_queue. A remapped bio must
drop that reference and recreate the association for the new queue even
when no blkcg policy needs it. Dying blkgs add another complication:
they disappear from the per-blkcg lookup tree before all pinned bio
references drain.

Patch 1 makes request_queue the authoritative lookup owner. A
queue-owned rhashtable keyed by the blkcg CSS ID keeps dying blkgs
discoverable until their references drain, while q->blkg_list remains
available for ordered walks.

Patch 2 stores and pins the queue-independent blkcg CSS in each bio.
Queue-local blkgs are created lazily by policy users, then pinned until
the bio changes devices or releases its cgroup state. Completion and
accounting paths which do not create a blkg remain lookup-only.

Patch 3 moves async bio punt state from blkg to blkcg, so punting alone
does not instantiate a queue-local blkg.

Changes since v1:

  - Rebase onto the latest block-7.3 branch.
  - In patch 2, require every blkg rhashtable lookup to run inside an
    explicit RCU read-side critical section. Annotate blkg_lookup_any(),
    blkg_lookup(), and recursive wrappers with __must_hold_shared(RCU)
    so Clang thread-safety analysis verifies the contract.
  - In patch 2, exclude BIO_BLKG_REF from the flags saved by
    dm_bio_record(), since restoring an ownership bit cannot recreate
    the associated blkg reference.
  - In patch 3, use scoped spinlock guards for async_bios and annotate
    the list with __guarded_by().
  - Add Nilay Shroff's Reviewed-by to patch 1.

Changes since RFC v3:

  - Drop the RFC prefix.
  - Add Christoph Hellwig's Reviewed-by to patch 1.
  - Add Tao Cui's Reviewed-by to patch 2.
  - In patch 2, relax blkg_lookup_any()/blkg_lookup() to also allow
    q->queue_lock as an alternative to the RCU read lock.

Changes since RFC v2:

  - Drop the old patch 1 and send it separately as a bugfix for 7.3 and
    -stable, as suggested by Christoph Hellwig.
  - Rebase onto the latest for-7.3/block branch.
  - In the new patch 1, note that the remaining q->blkg_list walkers are
    cgroupfs/sysfs slow paths and can move to rhashtable iteration after
    the queue_lock-to-blkcg_mutex conversion.
  - In the new patch 1, explain the list_empty case in blkg_release(), and
    add an RCU lockdep assertion plus documentation that blkg_lookup_any()
    does not acquire a reference.

Changes since RFC v1:

  - Add patch 1 to wait for every old blkg to leave q->blkg_list before a
    shared request_queue is rebound, instead of treating root_blkg == NULL
    as completion of asynchronous blkg teardown.
  - Add patch 2 to replace the per-blkcg radix tree and lookup hint with a
    request_queue rhashtable keyed by blkcg->css.id, as suggested by
    Christoph Hellwig. Keep dying pinned blkgs in the hash until
    blkg_release() so bio-owned references remain discoverable.
  - Fold the v1 helper-only patch into patch 3, as suggested by Jan Kara and
    Christoph Hellwig, and make bio_blkcg() naturally return NULL for an
    unassociated bio.
  - Rework patch 3 so blkg_lookup_create() acquires the bio-owned reference,
    falls back to a live parent when creation or tryget fails, and updates
    bi_blkcg when the returned blkg belongs to an ancestor.
  - Make bio_blkg_lookup() lookup-only: it returns NULL unless BIO_BLKG_REF
    is already set. Use bio_blkg() in the BFQ and IOCOST merge paths which
    may need to create a blkg, while keeping blk_cgroup_bio_start() and
    completion paths lookup-only.
  - Move CSS online-reference handling into
    bio_associate_blkcg_from_css(), including fallback to the root blkcg, so
    bio_associate_blkcg() does not take a redundant reference.
  - Keep the v1 async bio punt conversion as patch 4 and document that
    async_bio_lock protects async_bios.

Previous versions:
  v1: https://lore.kernel.org/r/20260823133045.970199-1-yukuai@kernel.org
  RFC v3: https://lore.kernel.org/r/20260818070641.756747-1-yukuai@kernel.org
  RFC v2: https://lore.kernel.org/r/20260811064744.1139446-1-yukuai@kernel.org
  RFC v1: https://lore.kernel.org/r/20260804065313.2092022-1-yukuai@kernel.org

Related series:
  [1] RFC v2 blk-cgroup: protect blkgs with blkcg_mutex
      https://lore.kernel.org/r/20260724123037.3004560-1-yukuai@kernel.org

Yu Kuai (3):
  blk-cgroup: use a request_queue rhashtable for blkg lookup
  blk-cgroup: store blkcg in bio instead of blkg
  blk-cgroup: move async bio punt state to blkcg

 Documentation/admin-guide/cgroup-v2.rst |   2 +-
 block/bfq-cgroup.c                      |  18 +-
 block/bfq-iosched.c                     |  19 +-
 block/bio.c                             |  22 +-
 block/blk-cgroup-fc-appid.c             |  10 +-
 block/blk-cgroup-rwstat.h               |   2 +-
 block/blk-cgroup.c                      | 358 ++++++++++++++----------
 block/blk-cgroup.h                      |  93 ++++--
 block/blk-core.c                        |   9 +-
 block/blk-crypto-fallback.c             |   2 +-
 block/blk-iocost.c                      |  12 +-
 block/blk-iolatency.c                   |  11 +-
 block/blk-ioprio.c                      |   2 +-
 block/blk-throttle.c                    |   3 +-
 block/blk-throttle.h                    |   2 +-
 drivers/md/bcache/request.c             |   2 +-
 drivers/md/dm-bio-record.h              |   3 +-
 drivers/md/dm.c                         |   2 +-
 drivers/md/md.c                         |   2 +-
 fs/gfs2/lops.c                          |   3 +-
 include/linux/bio.h                     |  26 +-
 include/linux/blk_types.h               |   9 +-
 include/linux/blkdev.h                  |   2 +
 include/linux/writeback.h               |   2 +-
 mm/page_io.c                            |  12 +-
 25 files changed, 391 insertions(+), 237 deletions(-)


base-commit: fc34aff88b3bd22146745c75cff0abb3c067cf42
-- 
2.51.0

^ permalink raw reply	[flat|nested] 6+ messages in thread

end of thread, other threads:[~2026-09-13 13:31 UTC | newest]

Thread overview: 6+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-13  6:54 [PATCH v2 0/3] blk-cgroup: store blkcg in bio before blkcg_mutex conversion Yu Kuai
2026-09-13  6:54 ` [PATCH v2 1/3] blk-cgroup: use a request_queue rhashtable for blkg lookup Yu Kuai
2026-09-13  6:54 ` [PATCH v2 2/3] blk-cgroup: store blkcg in bio instead of blkg Yu Kuai
2026-09-13 13:06   ` Nilay Shroff
2026-09-13  6:54 ` [PATCH v2 3/3] blk-cgroup: move async bio punt state to blkcg Yu Kuai
2026-09-13 13:29   ` Nilay Shroff

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox