All of lore.kernel.org
 help / color / mirror / Atom feed
* [PATCH v7 00/14] ceph: add binary logging (BLOG) for CephFS
@ 2026-09-24 15:30 Alex Markuze
  2026-09-24 15:30 ` [PATCH v7 01/14] ceph: add BLOG private headers Alex Markuze
                   ` (14 more replies)
  0 siblings, 15 replies; 16+ messages in thread
From: Alex Markuze @ 2026-09-24 15:30 UTC (permalink / raw)
  To: ceph-devel; +Cc: idryomov, xiubo.li

This series adds a per-mount binary flight recorder for CephFS debug
messages. It stores typed arguments in per-task page-fragment buffers
and reconstructs text through debugfs. Runtime enablement is per mount
and defaults to off; CONFIG_DEBUG_FS remains the build gate.

Thanks to Xiubo for the detailed v6 testing and reproducer. This revision
addresses the two xattr logging hazards and the empty-magazine growth
reported in that review. Further code reviews and a 32-bit build also
found issues in cache lookup, counted-name logging, debugfs controls,
dump completion, record timestamps and stack usage, fixed here. The
complete 14-patch series is based on testing tip 6a8449a3f814.

Changes since v6:

  - In __ceph_destroy_xattrs(), log the node pointers without reading
    xattr->name. Snapshot encoding can free the blob backing the old
    index before its destruction; a bounded string read is still unsafe.
  - In __copy_xattr_names(), log the NUL-terminated destination and xattr
    pointer. The borrowed source name is length-delimited and cannot be
    passed to an unbounded %s.
  - Return exhausted allocation magazines directly to log_batch's empty
    list, where blog_batch_put() can refill them. Do this for both task
    context allocation and atomic buffer rotation. Detach the per-CPU
    magazine before publishing it under the destination empty-list lock.
    Both batches use the same magazine slab cache.
  - Load the per-CPU cached context once and check that same pointer.
    A remote migration or retirement can clear the slot between the old
    check and reload, causing a NULL dereference despite local preemption
    being disabled.
  - On a long encrypted snapshot-name lookup failure, log the existing
    terminated copy, restoring the leading underscore in the format.
    The original input is length-delimited.
  - Bound binary capture of the base64-encoded ciphertext name with
    BLOG_STR(p, elen). base64_encode() does not append a terminator;
    %.*s alone only bounds the text path.
  - Make the BLOG read files root-readable (0400), matching the other
    Ceph debugfs data files. Recorded names and xattr values must not be
    exposed when the debugfs root is traversable by other users.
  - Check the mount's enabled flag during context lookup. Disabling one
    mount must stop subsequent capture even if another enabled mount
    keeps the global static key active.
  - Bound an entries dump by a maximum context ID, preserving that bound
    across seq_file overflow retries. Stop formatting on overflow so
    sustained rotation cannot keep a reader chasing new contexts.
  - Store record base times as u64 from get_jiffies_64(), including the
    delta calculation. Protect published base updates with the pagefrag
    lock used by readers. This preserves the epoch on 32-bit kernels,
    whose low jiffies word first wraps about five minutes after boot.
    Use one clock sample for the overflow check and stored delta, so a
    tick between them cannot wrap an exactly U32_MAX delta to zero.
  - Prepare each argument directly in its existing TLS scratch slot.
    Returning a temporary struct for every argument pushed the GCC i386
    __ceph_setattr() frame over its 1280-byte build limit. Direct filling
    brings reported stack usage to 300 bytes without changing evaluation
    order, string bounds or the recursive arity macros.
  - Rebase onto testing with the subsequent CephFS fixes already applied.
    Fold the fixes into patches 01, 04, 06, 07, 10 and 14; retain 14 patches.

Local validation:

  - GCC 11.4 and Clang 18.1 builds of fs/ceph/ceph.o and
    net/ceph/libceph.o with DEBUG_FS=y and DYNAMIC_DEBUG=y; GCC also
    builds both objects with DEBUG_FS=n. A GCC i386 build of both objects
    also passes with the default 1280-byte frame limit.
  - ASan/UBSan host tests using the actual xattr functions, BLOG argument
    serializer and kernel string-reading loop reproduce the v6 UAF and
    over-read. The v7 cases pass with BLOG and dynamic debug off/on.
  - Host tests using the actual magazine get/put/cleanup and rebalance
    code check both allocation call sites, sustained churn, reuse with
    new magazine allocations disabled, and 64,000 cycles on eight
    pthread workers. v6 exposes the growth; v7 reuses the magazines and
    releases all remaining objects at cleanup.
  - Additional ASan/UBSan host checks inject a remote cache clear and use
    exact-sized, unterminated encrypted names. They reproduce all three
    additional faults in the earlier v7 draft and pass with these fixes.
    The decoder passes 200,000 randomized malformed-payload cases with
    exact-sized input and output allocations under ASan/UBSan.
  - Before/after reader tests cover continuous context-ID churn and a
    stable bound across overflow retries. A cache test covers disabling
    a mount with an active context. Native freestanding i386 checks of
    the actual timestamp expressions cover crossing the first wrap,
    creating a context after it, detecting an expired u32 delta, and a
    clock tick at the exact U32_MAX boundary.
  - GCC/Clang serializer, request-context and initialization-failure
    host tests pass. The review-fix diff has no strict checkpatch errors,
    warnings or checks.

These are composite-object builds and host regression checks. Kernel
primitives are substituted in the host harnesses; they do not establish
kernel scheduling, interrupt, KASAN or lockdep behavior. A v7 booted
kernel and Ceph-cluster rerun remains for Xiubo; KASAN is not available
in the local setup. Xiubo's reported cluster results were against v6
with the destructor logging fix.

Known follow-up: existing %ptSp arguments still decode as pointer values
on the binary path. Timestamp-by-value serialization and further style
cleanup remain deferred as in v6.

AI assistance was used for the v7 fixes, code review, regression harnesses
and this cover letter. The changed implementation commits carry
Assisted-by attribution.

Alex Markuze (14):
  ceph: add BLOG private headers
  ceph: add BLOG deserialization support
  ceph: add BLOG page-fragment allocator
  ceph: add BLOG magazine batch allocator
  ceph: add BLOG logger core
  ceph: add BLOG per-module context management
  ceph: add Ceph BLOG scaffolding
  ceph: add boutc wrappers for BLOG
  ceph: switch MDS request plumbing to struct ceph_journal_info
  ceph: add BLOG debugfs interface
  ceph: convert VFS inode and directory paths to BLOG logging
  ceph: convert VFS data I/O paths to BLOG logging
  ceph: convert capability and snapshot paths to BLOG logging
  ceph: convert remaining helper paths to BLOG logging

 fs/ceph/Makefile                |    3 +
 fs/ceph/addr.c                  |  202 ++++---
 fs/ceph/blog.h                  |  216 +++++++
 fs/ceph/blog_batch.c            |  267 ++++++++
 fs/ceph/blog_batch.h            |   44 ++
 fs/ceph/blog_client.c           |  644 ++++++++++++++++++++
 fs/ceph/blog_core.c             |  297 +++++++++
 fs/ceph/blog_debugfs.c          |  702 +++++++++++++++++++++
 fs/ceph/blog_des.c              |  335 +++++++++++
 fs/ceph/blog_des.h              |   16 +
 fs/ceph/blog_module.c           | 1003 +++++++++++++++++++++++++++++++
 fs/ceph/blog_module.h           |   42 ++
 fs/ceph/blog_pagefrag.c         |   62 ++
 fs/ceph/blog_pagefrag.h         |   27 +
 fs/ceph/blog_ser.h              |  404 +++++++++++++
 fs/ceph/caps.c                  |  116 ++--
 fs/ceph/crypto.c                |   19 +-
 fs/ceph/debugfs.c               |   11 +-
 fs/ceph/dir.c                   |  329 +++++++---
 fs/ceph/export.c                |   83 ++-
 fs/ceph/file.c                  |  287 ++++++---
 fs/ceph/inode.c                 |  256 ++++----
 fs/ceph/locks.c                 |   66 +-
 fs/ceph/mds_client.c            |  369 +++++++-----
 fs/ceph/snap.c                  |   17 +-
 fs/ceph/super.c                 |   61 +-
 fs/ceph/super.h                 |    7 +
 fs/ceph/xattr.c                 |  101 ++--
 include/linux/ceph/ceph_blog.h  |  292 +++++++++
 include/linux/ceph/ceph_debug.h |   81 ++-
 include/linux/ceph/libceph.h    |    2 +
 31 files changed, 5705 insertions(+), 656 deletions(-)
 create mode 100644 fs/ceph/blog.h
 create mode 100644 fs/ceph/blog_batch.c
 create mode 100644 fs/ceph/blog_batch.h
 create mode 100644 fs/ceph/blog_client.c
 create mode 100644 fs/ceph/blog_core.c
 create mode 100644 fs/ceph/blog_debugfs.c
 create mode 100644 fs/ceph/blog_des.c
 create mode 100644 fs/ceph/blog_des.h
 create mode 100644 fs/ceph/blog_module.c
 create mode 100644 fs/ceph/blog_module.h
 create mode 100644 fs/ceph/blog_pagefrag.c
 create mode 100644 fs/ceph/blog_pagefrag.h
 create mode 100644 fs/ceph/blog_ser.h
 create mode 100644 include/linux/ceph/ceph_blog.h


base-commit: 6a8449a3f81444b6a25adc41cfd2c0ac33f8f96a
-- 
2.34.1


^ permalink raw reply	[flat|nested] 16+ messages in thread

* [PATCH v7 01/14] ceph: add BLOG private headers
  2026-09-24 15:30 [PATCH v7 00/14] ceph: add binary logging (BLOG) for CephFS Alex Markuze
@ 2026-09-24 15:30 ` Alex Markuze
  2026-09-24 15:30 ` [PATCH v7 02/14] ceph: add BLOG deserialization support Alex Markuze
                   ` (13 subsequent siblings)
  14 siblings, 0 replies; 16+ messages in thread
From: Alex Markuze @ 2026-09-24 15:30 UTC (permalink / raw)
  To: ceph-devel; +Cc: idryomov, xiubo.li

Add the private header files for the binary logging (BLOG) subsystem:

  blog.h          - core types: blog_logger, blog_log_entry,
                    blog_tls_ctx, blog_source_info, blog_log_iter,
                    blog_source_id_cache; core API prototypes
  blog_ser.h      - single-evaluation argument packing and bounded,
                    type-aware serialization helpers
  blog_des.h      - deserialization prototypes
  blog_batch.h    - per-CPU magazine batch allocator types
  blog_pagefrag.h - page-fragment allocator types
  blog_module.h   - per-module context and task-entry declarations

Private headers live under fs/ceph/; only the CephFS integration
surface remains in include/linux/ceph/.

CEPH_BLOG_LOG_CLIENT stages packed arguments in the per-task
blog_tls_ctx scratch and emits from a noinline helper so VFS frames
do not grow a struct blog_arg[] at each call site. Same-width values
go through unsigned long before u64 so 32-bit sparse does not see
pointer-to-u64.

Fill TLS scratch with recursive macros that capture each destination
and argument once, in argument order. Keep the fill definitions short
without reintroducing a stack argument array. Provide a private
static-key context selector for logging-only preparation helpers.

Fill each existing scratch slot directly. Returning a temporary
struct blog_arg per argument made the GCC i386 __ceph_setattr()
frame exceed the default 1280-byte build limit. Direct filling
keeps destination and argument evaluation single and ordered,
while reducing reported stack usage to 300 bytes.

Keep the record base timestamp in u64 so it can preserve the full
get_jiffies_64() epoch on both 32-bit and 64-bit kernels.

Reported-by: kernel test robot <lkp@intel.com>
Closes: https://lore.kernel.org/oe-kbuild-all/202609080414.m3ZvUYmq-lkp@intel.com/
Signed-off-by: Alex Markuze <amarkuze@redhat.com>
Assisted-by: LLM
---
 fs/ceph/blog.h          | 216 +++++++++++++++++++++
 fs/ceph/blog_batch.h    |  44 +++++
 fs/ceph/blog_des.h      |  16 ++
 fs/ceph/blog_module.h   |  42 +++++
 fs/ceph/blog_pagefrag.h |  27 +++
 fs/ceph/blog_ser.h      | 404 ++++++++++++++++++++++++++++++++++++++++
 6 files changed, 749 insertions(+)
 create mode 100644 fs/ceph/blog.h
 create mode 100644 fs/ceph/blog_batch.h
 create mode 100644 fs/ceph/blog_des.h
 create mode 100644 fs/ceph/blog_module.h
 create mode 100644 fs/ceph/blog_pagefrag.h
 create mode 100644 fs/ceph/blog_ser.h

diff --git a/fs/ceph/blog.h b/fs/ceph/blog.h
new file mode 100644
index 000000000000..6bfc942eeb4e
--- /dev/null
+++ b/fs/ceph/blog.h
@@ -0,0 +1,216 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * CephFS binary logging (BLOG) private types.
+ */
+#ifndef _FS_CEPH_BLOG_H
+#define _FS_CEPH_BLOG_H
+
+#include <linux/types.h>
+#include <linux/limits.h>
+#include <linux/build_bug.h>
+#include <linux/sched.h>
+#include <linux/list.h>
+#include <linux/mutex.h>
+#include <linux/spinlock.h>
+#include <linux/seqlock.h>
+#include <linux/atomic.h>
+#include <linux/workqueue.h>
+#include <linux/rhashtable-types.h>
+#include <linux/ceph/ceph_blog.h>
+#include "blog_batch.h"
+#include "blog_pagefrag.h"
+#include "blog_ser.h"
+#include "blog_des.h"
+
+struct blog_module_context;
+
+/* Absolute caps for module parameters (load-time only). */
+#define BLOG_MAX_SOURCE_IDS_CAP	65535
+#define BLOG_MAX_CLIENT_IDS_CAP	256
+#define BLOG_DEFAULT_MAX_SOURCES 4096
+#define BLOG_DEFAULT_MAX_CLIENTS 256
+#define BLOG_MAX_PAYLOAD U16_MAX
+
+struct blog_source_info {
+	const char *file;
+	const char *func;
+	unsigned int line;
+	const char *fmt;
+	int warn_count;
+};
+
+struct blog_log_entry {
+	u32 ts_delta;
+	u16 source_id;
+	u16 len;
+	u8 client_id;
+	u8 flags;
+	char buffer[];
+};
+
+#define BLOG_CTX_NEEDS_RESET	0
+#define BLOG_CTX_TASK_COUNTED	1
+
+struct blog_tls_ctx {
+	struct list_head list;
+	void (*release)(void *data);
+	atomic_t refcount;
+	struct task_struct *task;
+	pid_t pid;
+	char comm[TASK_COMM_LEN];
+	u64 id;
+	u64 base_jiffies;
+	struct blog_logger *logger;
+	int pending_offset;
+	size_t pending_size;
+	unsigned long flags;
+	/* The CPU cache is usable only while inside ceph_blog_enter*(). */
+	unsigned int enter_depth;
+	/* CPU whose per-CPU cache slot bind published; unbind clears it after migration. */
+	int cache_cpu;
+	/* Matches logger->clear_seq after a reset; readers treat mismatch as empty. */
+	atomic64_t clear_seq;
+	/*
+	 * Staging for boutc() argument packing. Lives here so VFS frames
+	 * do not grow a struct blog_arg[] per call site. Fill+emit in one
+	 * boutc() is not re-entrant: argument expressions must not log.
+	 */
+	struct blog_arg arg_scratch[BLOG_MAX_ARGS];
+};
+
+struct blog_tls_pagefrag {
+	struct blog_tls_ctx ctx;
+	struct blog_pagefrag pf;
+	unsigned char buf[];
+};
+
+#define BLOG_TLS_PAGEFRAG_ALLOC_SIZE BLOG_PAGEFRAG_SIZE
+#define BLOG_TLS_PAGEFRAG_BUFFER_SIZE \
+	(BLOG_PAGEFRAG_SIZE - sizeof(struct blog_tls_pagefrag))
+
+struct blog_logger {
+	struct list_head contexts;
+	spinlock_t lock;
+	struct mutex snapshot_mutex;
+	struct rhashtable task_map;
+	struct blog_batch alloc_batch;
+	struct blog_batch log_batch;
+	struct kmem_cache *magazine_cache;
+	struct blog_source_info *source_map;
+	u32 *source_hash;	/* open-addressed, stores source IDs (0 = empty) */
+	u32 source_hash_mask;
+	u32 max_source_ids;
+	u32 next_source_id;	/* serialized by source_lock */
+	spinlock_t source_lock;
+	unsigned long total_contexts_allocated;
+	u64 next_ctx_id;
+	spinlock_t ctx_id_lock;
+	struct blog_module_context *owner_ctx;
+	u64 generation;
+	/* Bumped by debugfs "clear"; writers/readers sync via ctx->clear_seq. */
+	atomic64_t clear_seq;
+	struct delayed_work task_gc_work;
+	bool task_gc_stopping;
+	/*
+	 * Atomic rotate (boutc under spinlock) cannot take snapshot_mutex.
+	 * Failed log-batch puts land on reclaim_list; reclaim_work drains
+	 * them and rebalances under sleepable context.
+	 */
+	struct list_head reclaim_list;
+	struct work_struct reclaim_work;
+};
+
+/**
+ * struct blog_source_id_cache - per-callsite source ID fast cache
+ * @logger: logger that owns the cached ID
+ * @generation: logger generation at cache-fill time
+ * @id: cached source ID (0 = not yet cached)
+ * @seq: validates lockless cache snapshots
+ * @lock: serializes cache updates shared by all mounts
+ */
+struct blog_source_id_cache {
+	struct blog_logger *logger;
+	u64 generation;
+	u32 id;
+	seqcount_spinlock_t seq;
+	spinlock_t lock;
+};
+
+#define BLOG_SOURCE_ID_CACHE_INIT(name) { \
+	.seq = SEQCNT_SPINLOCK_ZERO(name.seq, &(name).lock), \
+	.lock = __SPIN_LOCK_UNLOCKED(name.lock), \
+}
+
+struct blog_log_iter {
+	struct blog_pagefrag *pf;
+	u64 current_offset;
+	u64 end_offset;
+	u64 prev_offset;
+	u64 steps;
+};
+
+typedef int (*blog_client_des_fn)(char *buf, size_t size, u8 client_id);
+
+int blog_param_max_sources(void);
+int blog_param_max_clients(void);
+
+u32 blog_get_source_id(struct blog_logger *logger, const char *file,
+		       const char *func, unsigned int line, const char *fmt);
+u32 blog_get_source_id_cached(struct blog_logger *logger,
+			      struct blog_source_id_cache *cache,
+			      const char *file, const char *func,
+			      unsigned int line, const char *fmt);
+struct blog_source_info *blog_get_source_info(struct blog_logger *logger,
+					      u32 id);
+void blog_log_iter_init(struct blog_log_iter *iter, struct blog_pagefrag *pf,
+			u64 head_snapshot);
+struct blog_log_entry *blog_log_iter_next(struct blog_log_iter *iter);
+int blog_des_entry(struct blog_logger *logger, struct blog_log_entry *entry,
+		   char *output, size_t out_size,
+		   blog_client_des_fn client_cb);
+void *blog_log_with_ctx(struct blog_logger *logger,
+			struct blog_tls_ctx *tls_ctx,
+			u32 source_id, u8 client_id, size_t needed_size);
+int blog_log_commit_with_ctx(struct blog_logger *logger,
+			     struct blog_tls_ctx *tls_ctx,
+			     size_t actual_size);
+void blog_log_client_emit(struct blog_tls_ctx *ctx, struct ceph_client *client,
+			  struct blog_source_id_cache *cache,
+			  const char *file, const char *func,
+			  unsigned int line, const char *fmt, size_t nargs);
+static inline struct blog_tls_pagefrag *blog_ctx_container(struct blog_tls_ctx *ctx)
+{
+	return container_of(ctx, struct blog_tls_pagefrag, ctx);
+}
+
+static inline struct blog_pagefrag *blog_ctx_pf(struct blog_tls_ctx *ctx)
+{
+	return &blog_ctx_container(ctx)->pf;
+}
+
+/* Select the binary path before preparing logging-only snapshots or strings. */
+static inline struct blog_tls_ctx *ceph_blog_get_ctx(struct ceph_fs_client *fsc)
+{
+#ifdef CONFIG_DEBUG_FS
+	if (static_branch_unlikely(&ceph_blog_key))
+		return ceph_blog_get_cached_ctx(fsc);
+#endif
+	return NULL;
+}
+
+#define CEPH_BLOG_LOG_CLIENT(ctx, client, fmt, ...) \
+	do { \
+		static struct blog_source_id_cache __source_cache = \
+			BLOG_SOURCE_ID_CACHE_INIT(__source_cache); \
+		struct blog_tls_ctx *__blog_ctx = (ctx); \
+		size_t __nargs = blog_narg(__VA_ARGS__); \
+		if (unlikely(!__blog_ctx) || unlikely(!__blog_ctx->logger)) \
+			break; \
+		BUILD_BUG_ON(blog_narg(__VA_ARGS__) > BLOG_MAX_ARGS); \
+		BLOG_FILL_ARGS(__blog_ctx->arg_scratch, ##__VA_ARGS__); \
+		blog_log_client_emit(__blog_ctx, client, &__source_cache, \
+				     kbasename(__FILE__), __func__, __LINE__, \
+				     fmt, __nargs); \
+	} while (0)
+
+#endif /* _FS_CEPH_BLOG_H */
diff --git a/fs/ceph/blog_batch.h b/fs/ceph/blog_batch.h
new file mode 100644
index 000000000000..a38c8d50ffc8
--- /dev/null
+++ b/fs/ceph/blog_batch.h
@@ -0,0 +1,44 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * BLOG magazine batcher.
+ */
+#ifndef _FS_CEPH_BLOG_BATCH_H
+#define _FS_CEPH_BLOG_BATCH_H
+
+#include <linux/types.h>
+#include <linux/percpu.h>
+#include <linux/spinlock.h>
+#include <linux/list.h>
+
+#define BLOG_MAGAZINE_SIZE 16
+
+struct blog_magazine {
+	struct list_head list;
+	unsigned int count;
+	void *elements[BLOG_MAGAZINE_SIZE];
+};
+
+struct blog_cpu_magazine {
+	struct blog_magazine *mag;
+};
+
+struct blog_batch {
+	struct list_head full_magazines;
+	struct list_head empty_magazines;
+	raw_spinlock_t full_lock;
+	raw_spinlock_t empty_lock;
+	unsigned int nr_full;
+	unsigned int nr_empty;
+	struct blog_cpu_magazine __percpu *cpu_magazines;
+	struct kmem_cache *magazine_cache;
+	bool external_cache;
+	unsigned int retain_limit;
+};
+
+int blog_batch_init(struct blog_batch *batch, struct kmem_cache *mag_cache,
+		    unsigned int nr_prealloc, unsigned int retain_limit);
+void blog_batch_cleanup(struct blog_batch *batch);
+void *blog_batch_get(struct blog_batch *batch, struct blog_batch *recycle);
+bool blog_batch_put(struct blog_batch *batch, void *element);
+
+#endif /* _FS_CEPH_BLOG_BATCH_H */
diff --git a/fs/ceph/blog_des.h b/fs/ceph/blog_des.h
new file mode 100644
index 000000000000..a67bec43ab2d
--- /dev/null
+++ b/fs/ceph/blog_des.h
@@ -0,0 +1,16 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * BLOG record deserialization.
+ */
+#ifndef _FS_CEPH_BLOG_DES_H
+#define _FS_CEPH_BLOG_DES_H
+
+#include <linux/types.h>
+
+struct blog_log_entry;
+struct blog_logger;
+
+int blog_des_reconstruct(const char *fmt, const void *buffer,
+			 size_t size, char *out, size_t out_size);
+
+#endif /* _FS_CEPH_BLOG_DES_H */
diff --git a/fs/ceph/blog_module.h b/fs/ceph/blog_module.h
new file mode 100644
index 000000000000..adb24a8ee1ac
--- /dev/null
+++ b/fs/ceph/blog_module.h
@@ -0,0 +1,42 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Per-superblock BLOG context and task-keyed rhashtable.
+ */
+#ifndef _FS_CEPH_BLOG_MODULE_H
+#define _FS_CEPH_BLOG_MODULE_H
+
+#include "blog.h"
+#include <linux/rhashtable.h>
+#include <linux/refcount.h>
+
+struct blog_task_entry {
+	struct rhash_head	node;
+	struct task_struct	*task;
+	pid_t			pid;
+	char			comm[TASK_COMM_LEN];
+	struct blog_tls_ctx	*ctx;
+	unsigned long		flags;
+	struct rcu_head		rcu;
+};
+
+struct blog_module_context {
+	char name[32];
+	struct blog_logger *logger;
+	void *module_private;
+	refcount_t refcount;
+	atomic_t allocated_contexts;
+	struct work_struct free_work;
+	bool initialized;
+};
+
+struct blog_module_context *blog_module_init(const char *module_name);
+void blog_module_put(struct blog_module_context *ctx);
+void blog_module_flush_frees(void);
+int blog_module_wq_init(void);
+void blog_module_wq_exit(void);
+
+struct blog_tls_ctx *blog_lookup_tls_ctx(struct blog_module_context *ctx);
+struct blog_tls_ctx *blog_get_tls_ctx_ctx(struct blog_module_context *ctx,
+					  gfp_t gfp);
+
+#endif /* _FS_CEPH_BLOG_MODULE_H */
diff --git a/fs/ceph/blog_pagefrag.h b/fs/ceph/blog_pagefrag.h
new file mode 100644
index 000000000000..d6a77268cf12
--- /dev/null
+++ b/fs/ceph/blog_pagefrag.h
@@ -0,0 +1,27 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * BLOG page-fragment allocator.
+ */
+#ifndef _FS_CEPH_BLOG_PAGEFRAG_H
+#define _FS_CEPH_BLOG_PAGEFRAG_H
+
+#include <linux/types.h>
+#include <linux/sizes.h>
+#include <linux/spinlock.h>
+
+#define BLOG_PAGEFRAG_SIZE  SZ_4K
+#define BLOG_PAGEFRAG_MASK (BLOG_PAGEFRAG_SIZE - 1)
+
+struct blog_pagefrag {
+	void *buffer;
+	size_t capacity;
+	spinlock_t lock;
+	unsigned int head;
+};
+
+int blog_pagefrag_reserve(struct blog_pagefrag *pf, unsigned int n);
+void blog_pagefrag_publish(struct blog_pagefrag *pf, unsigned int publish_head);
+void blog_pagefrag_reset(struct blog_pagefrag *pf);
+void *blog_pagefrag_get_ptr(struct blog_pagefrag *pf, u64 val);
+
+#endif /* _FS_CEPH_BLOG_PAGEFRAG_H */
diff --git a/fs/ceph/blog_ser.h b/fs/ceph/blog_ser.h
new file mode 100644
index 000000000000..7c1d6b774439
--- /dev/null
+++ b/fs/ceph/blog_ser.h
@@ -0,0 +1,404 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * BLOG argument packing.
+ */
+#ifndef _FS_CEPH_BLOG_SER_H
+#define _FS_CEPH_BLOG_SER_H
+
+#include <linux/limits.h>
+#include <linux/math.h>
+#include <linux/minmax.h>
+#include <linux/string.h>
+#include <linux/types.h>
+#include <linux/unaligned.h>
+
+#define IS_STR_PTR(t) \
+	(__builtin_types_compatible_p(typeof(t), const char *) || \
+	 __builtin_types_compatible_p(typeof(t), char *) || \
+	 __builtin_types_compatible_p(typeof(t), const unsigned char *) || \
+	 __builtin_types_compatible_p(typeof(t), unsigned char *))
+
+#define IS_STR_ARRAY(t) \
+	(__builtin_types_compatible_p(typeof(t), const char []) || \
+	 __builtin_types_compatible_p(typeof(t), char []) || \
+	 __builtin_types_compatible_p(typeof(t), const unsigned char []) || \
+	 __builtin_types_compatible_p(typeof(t), unsigned char []))
+
+#define IS_STR(t) (IS_STR_PTR(t) || IS_STR_ARRAY(t))
+
+struct blog_bounded_string {
+	const char *str;
+	size_t len;
+};
+
+#define BLOG_STR(__str, __len) \
+	(&(const struct blog_bounded_string){ \
+		.str = (const char *)(__str), \
+		.len = (__len), \
+	})
+
+#define IS_BOUNDED_STR(t) \
+	(__builtin_types_compatible_p(typeof(t), \
+				      const struct blog_bounded_string *) || \
+	 __builtin_types_compatible_p(typeof(t), struct blog_bounded_string *))
+
+#define __suppress_cast_warning(type, value) \
+({ \
+	_Pragma("GCC diagnostic push") \
+	_Pragma("GCC diagnostic ignored \"-Wint-to-pointer-cast\"") \
+	_Pragma("GCC diagnostic ignored \"-Wpointer-to-int-cast\"") \
+	type __scw_result; \
+	__scw_result = ((type)(value)); \
+	_Pragma("GCC diagnostic pop") \
+	__scw_result; \
+})
+
+#define ___blog_concat(__a, __b) __a ## __b
+#define ___blog_apply(__fn, __n) ___blog_concat(__fn, __n)
+
+#define ___blog_nth(_, __1, __2, __3, __4, __5, __6, __7, __8, __9, \
+	__10, __11, __12, __13, __14, __15, __16, __17, __18, __19, __20, \
+	__21, __22, __23, __24, __25, __26, __27, __28, __29, __30, __31, \
+	__32, __N, ...) __N
+#define ___blog_narg(...) ___blog_nth(_, ##__VA_ARGS__, \
+	32, 31, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, \
+	16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1, 0)
+#define blog_narg(...) ___blog_narg(__VA_ARGS__)
+
+#define STR_MAX_SIZE 255
+#define BLOG_MAX_ARGS 32
+
+/**
+ * struct blog_arg - an evaluated argument and its serialization reservation
+ * @value: cached scalar value
+ * @str: cached string pointer
+ * @string_len: measured number of source bytes, excluding the terminator
+ * @reserved: bytes reserved for this argument, including string padding
+ * @is_string: @str and @string_len are valid instead of @value
+ *
+ * Logging macros build these records before reserving pagefrag space.  String
+ * expressions are therefore evaluated once, and serialization cannot copy
+ * beyond the length used to calculate that string's reservation.
+ */
+struct blog_arg {
+	union {
+		u64 value;
+		const char *str;
+	};
+	u16 string_len;
+	u16 reserved;
+	bool is_string;
+};
+
+static inline void blog_arg_set_string(struct blog_arg *arg, const char *str,
+				       size_t limit)
+{
+	size_t len = 0;
+
+	arg->is_string = true;
+	arg->str = str;
+	if (!str) {
+		arg->string_len = 0;
+		arg->reserved = sizeof("(NULL) ");
+		return;
+	}
+
+	/*
+	 * @limit is the caller's scan bound.  Unbounded %s passes
+	 * STR_MAX_SIZE; BLOG_STR() passes the caller length (paths,
+	 * NAME_MAX dentries) and must not be silently shrunk to 254.
+	 * Cap only so reserved (= round_up(len + 1, 4)) fits in u16.
+	 */
+	limit = min_t(size_t, limit, (size_t)U16_MAX - 4);
+	while (len < limit && str[len])
+		len++;
+	arg->string_len = len;
+	arg->reserved = round_up(len + 1, 4);
+}
+
+#define const_char_ptr(str) __suppress_cast_warning(const char *, (str))
+/* Integer-width round-trip so both if-branches type-check for any arg. */
+#define bounded_string_ptr(str) \
+	((const struct blog_bounded_string *)(unsigned long)(str))
+
+#define BLOG_ARG(__dst, __arg) \
+do { \
+	struct blog_arg *__blog_arg = (__dst); \
+	__auto_type __blog_value = (__arg); \
+	*__blog_arg = (struct blog_arg){}; \
+	if (IS_BOUNDED_STR(__blog_value)) { \
+		const struct blog_bounded_string *__blog_str = \
+			bounded_string_ptr(__blog_value); \
+		blog_arg_set_string(__blog_arg, __blog_str->str, \
+				    __blog_str->len); \
+	} else if (IS_STR(__blog_value)) { \
+		blog_arg_set_string(__blog_arg, const_char_ptr(__blog_value), \
+				    STR_MAX_SIZE); \
+	} else { \
+		/* Same-width integer first so 32-bit sparse does not see \
+		 * pointer-to-u64. Wider scalars still go through u64. \
+		 */ \
+		if (sizeof(__blog_value) == sizeof(void *)) \
+			__blog_arg->value = (u64)__suppress_cast_warning( \
+				unsigned long, __blog_value); \
+		else \
+			__blog_arg->value = __suppress_cast_warning( \
+				u64, __blog_value); \
+		__blog_arg->reserved = sizeof(__blog_value) < 4 ? \
+			4 : sizeof(__blog_value); \
+	} \
+} while (0)
+
+/*
+ * Fill the per-task scratch in argument order, evaluating each expression once.
+ * Each recursive step captures its destination before writing the next slot.
+ * Fill each slot directly: by-value struct temporaries can grow VFS stack frames.
+ */
+#define ___blog_fill0(__dst, ...) ((void)(__dst))
+#define ___blog_fill1(__dst, __arg) BLOG_ARG(__dst, __arg)
+#define ___blog_fill2(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst2 = (__dst); \
+		___blog_fill1(__blog_dst2, __arg); \
+		___blog_fill1(__blog_dst2 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill3(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst3 = (__dst); \
+		___blog_fill1(__blog_dst3, __arg); \
+		___blog_fill2(__blog_dst3 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill4(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst4 = (__dst); \
+		___blog_fill1(__blog_dst4, __arg); \
+		___blog_fill3(__blog_dst4 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill5(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst5 = (__dst); \
+		___blog_fill1(__blog_dst5, __arg); \
+		___blog_fill4(__blog_dst5 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill6(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst6 = (__dst); \
+		___blog_fill1(__blog_dst6, __arg); \
+		___blog_fill5(__blog_dst6 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill7(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst7 = (__dst); \
+		___blog_fill1(__blog_dst7, __arg); \
+		___blog_fill6(__blog_dst7 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill8(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst8 = (__dst); \
+		___blog_fill1(__blog_dst8, __arg); \
+		___blog_fill7(__blog_dst8 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill9(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst9 = (__dst); \
+		___blog_fill1(__blog_dst9, __arg); \
+		___blog_fill8(__blog_dst9 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill10(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst10 = (__dst); \
+		___blog_fill1(__blog_dst10, __arg); \
+		___blog_fill9(__blog_dst10 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill11(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst11 = (__dst); \
+		___blog_fill1(__blog_dst11, __arg); \
+		___blog_fill10(__blog_dst11 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill12(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst12 = (__dst); \
+		___blog_fill1(__blog_dst12, __arg); \
+		___blog_fill11(__blog_dst12 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill13(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst13 = (__dst); \
+		___blog_fill1(__blog_dst13, __arg); \
+		___blog_fill12(__blog_dst13 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill14(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst14 = (__dst); \
+		___blog_fill1(__blog_dst14, __arg); \
+		___blog_fill13(__blog_dst14 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill15(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst15 = (__dst); \
+		___blog_fill1(__blog_dst15, __arg); \
+		___blog_fill14(__blog_dst15 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill16(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst16 = (__dst); \
+		___blog_fill1(__blog_dst16, __arg); \
+		___blog_fill15(__blog_dst16 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill17(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst17 = (__dst); \
+		___blog_fill1(__blog_dst17, __arg); \
+		___blog_fill16(__blog_dst17 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill18(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst18 = (__dst); \
+		___blog_fill1(__blog_dst18, __arg); \
+		___blog_fill17(__blog_dst18 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill19(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst19 = (__dst); \
+		___blog_fill1(__blog_dst19, __arg); \
+		___blog_fill18(__blog_dst19 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill20(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst20 = (__dst); \
+		___blog_fill1(__blog_dst20, __arg); \
+		___blog_fill19(__blog_dst20 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill21(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst21 = (__dst); \
+		___blog_fill1(__blog_dst21, __arg); \
+		___blog_fill20(__blog_dst21 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill22(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst22 = (__dst); \
+		___blog_fill1(__blog_dst22, __arg); \
+		___blog_fill21(__blog_dst22 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill23(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst23 = (__dst); \
+		___blog_fill1(__blog_dst23, __arg); \
+		___blog_fill22(__blog_dst23 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill24(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst24 = (__dst); \
+		___blog_fill1(__blog_dst24, __arg); \
+		___blog_fill23(__blog_dst24 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill25(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst25 = (__dst); \
+		___blog_fill1(__blog_dst25, __arg); \
+		___blog_fill24(__blog_dst25 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill26(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst26 = (__dst); \
+		___blog_fill1(__blog_dst26, __arg); \
+		___blog_fill25(__blog_dst26 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill27(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst27 = (__dst); \
+		___blog_fill1(__blog_dst27, __arg); \
+		___blog_fill26(__blog_dst27 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill28(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst28 = (__dst); \
+		___blog_fill1(__blog_dst28, __arg); \
+		___blog_fill27(__blog_dst28 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill29(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst29 = (__dst); \
+		___blog_fill1(__blog_dst29, __arg); \
+		___blog_fill28(__blog_dst29 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill30(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst30 = (__dst); \
+		___blog_fill1(__blog_dst30, __arg); \
+		___blog_fill29(__blog_dst30 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill31(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst31 = (__dst); \
+		___blog_fill1(__blog_dst31, __arg); \
+		___blog_fill30(__blog_dst31 + 1, __VA_ARGS__); \
+	} while (0)
+#define ___blog_fill32(__dst, __arg, ...) \
+	do { \
+		struct blog_arg *__blog_dst32 = (__dst); \
+		___blog_fill1(__blog_dst32, __arg); \
+		___blog_fill31(__blog_dst32 + 1, __VA_ARGS__); \
+	} while (0)
+#define BLOG_FILL_ARGS(__dst, ...) \
+	___blog_apply(___blog_fill, blog_narg(__VA_ARGS__))(__dst, __VA_ARGS__)
+
+static inline size_t blog_args_size(const struct blog_arg *args,
+				    size_t nr_args)
+{
+	size_t size = 0;
+	size_t i;
+
+	for (i = 0; i < nr_args; i++)
+		size += args[i].reserved;
+	return size;
+}
+
+static inline size_t blog_serialize_string(char *dst,
+					   const struct blog_arg *arg)
+{
+	static const char null_str[] = "(NULL) ";
+	size_t limit;
+	size_t count;
+
+	if (!arg->str) {
+		memcpy(dst, null_str, min(sizeof(null_str), arg->reserved));
+		return arg->reserved;
+	}
+
+	limit = min(arg->string_len, arg->reserved - 1);
+	for (count = 0; count < limit; count++) {
+		dst[count] = arg->str[count];
+		if (!dst[count])
+			return round_up(count + 1, 4);
+	}
+	dst[count] = '\0';
+	return round_up(count + 1, 4);
+}
+
+static inline void *blog_serialize_args(void *buffer,
+					const struct blog_arg *args,
+					size_t nr_args)
+{
+	char *dst = buffer;
+	size_t i;
+
+	for (i = 0; i < nr_args; i++) {
+		const struct blog_arg *arg = &args[i];
+
+		if (arg->is_string) {
+			dst += blog_serialize_string(dst, arg);
+		} else if (arg->reserved == 8) {
+			put_unaligned(arg->value, (u64 *)dst);
+			dst += 8;
+		} else {
+			put_unaligned((u32)arg->value, (u32 *)dst);
+			dst += 4;
+		}
+	}
+	return dst;
+}
+
+#endif /* _FS_CEPH_BLOG_SER_H */
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 16+ messages in thread

* [PATCH v7 02/14] ceph: add BLOG deserialization support
  2026-09-24 15:30 [PATCH v7 00/14] ceph: add binary logging (BLOG) for CephFS Alex Markuze
  2026-09-24 15:30 ` [PATCH v7 01/14] ceph: add BLOG private headers Alex Markuze
@ 2026-09-24 15:30 ` Alex Markuze
  2026-09-24 15:30 ` [PATCH v7 03/14] ceph: add BLOG page-fragment allocator Alex Markuze
                   ` (12 subsequent siblings)
  14 siblings, 0 replies; 16+ messages in thread
From: Alex Markuze @ 2026-09-24 15:30 UTC (permalink / raw)
  To: ceph-devel; +Cc: idryomov, xiubo.li

Add blog_des.c implementing the binary record deserialization engine
used by the debugfs read paths.  Supports %d, %i, %u, %o, %x, %X,
%s, %p, %c, and length modifiers (l, ll, h, hh, z).

Signed-off-by: Alex Markuze <amarkuze@redhat.com>
---
 fs/ceph/blog_des.c | 335 +++++++++++++++++++++++++++++++++++++++++++++
 1 file changed, 335 insertions(+)
 create mode 100644 fs/ceph/blog_des.c

diff --git a/fs/ceph/blog_des.c b/fs/ceph/blog_des.c
new file mode 100644
index 000000000000..b83212dee856
--- /dev/null
+++ b/fs/ceph/blog_des.c
@@ -0,0 +1,335 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * BLOG record deserialization.
+ */
+
+#include "blog_des.h"
+#include "blog.h"
+#include <linux/string.h>
+#include <linux/ctype.h>
+#include <linux/types.h>
+#include <linux/kernel.h>
+#include <linux/printk.h>
+#include <linux/align.h>
+#include <linux/unaligned.h>
+
+static inline void des_advance(int ret, char **out, size_t *remaining)
+{
+	if (ret > 0) {
+		if ((size_t)ret > *remaining)
+			ret = *remaining;
+		*out += ret;
+		*remaining -= ret;
+	}
+}
+
+/**
+ * blog_des_reconstruct - Reconstructs a formatted string from serialized values
+ * @fmt: Format string containing % specifiers
+ * @buffer: Buffer containing serialized values
+ * @size: Size of the buffer in bytes
+ * @out: Buffer to store the reconstructed string
+ * @out_size: Size of the output buffer
+ */
+int blog_des_reconstruct(const char *fmt, const void *buffer,
+			 size_t size, char *out, size_t out_size)
+{
+	const char *buf_start = (const char *)buffer;
+	const char *buf_ptr = buf_start;
+	const char *buf_end = buf_start + size;
+	const char *fmt_ptr = fmt;
+	char *out_ptr = out;
+	/* Payload budget; snprintf() is called with remaining + 1 so the
+	 * last character that still fits is not dropped for the NUL. */
+	size_t remaining = out_size - 1;
+	size_t arg_count = 0;
+	int ret;
+
+	if (!fmt || !buffer || !out || !out_size) {
+		pr_err_ratelimited("blog_des_reconstruct: invalid params fmt=%p buffer=%p out=%p out_size=%zu\n",
+			fmt, buffer, out, out_size);
+		return -EINVAL;
+	}
+
+	*out_ptr = '\0';
+
+	while (*fmt_ptr && remaining > 0) {
+		int is_long;
+		int is_long_long;
+		int alt_form = 0;
+		int dyn_width = -1;
+		int dyn_precision = -1;
+
+		if (*fmt_ptr != '%') {
+			*out_ptr++ = *fmt_ptr++;
+			remaining--;
+			continue;
+		}
+
+		fmt_ptr++;
+
+		if (*fmt_ptr == '%') {
+			*out_ptr++ = '%';
+			fmt_ptr++;
+			remaining--;
+			continue;
+		}
+
+		/* Skip flags (-+#0 space); remember '#' for hex/octal. */
+		while (*fmt_ptr && (*fmt_ptr == '-' || *fmt_ptr == '+' || *fmt_ptr == '#' ||
+				   *fmt_ptr == '0' || *fmt_ptr == ' ')) {
+			if (*fmt_ptr == '#')
+				alt_form = 1;
+			fmt_ptr++;
+		}
+
+		/* Consume field width: digits are format-only, * is serialized */
+		while (*fmt_ptr && (*fmt_ptr >= '0' && *fmt_ptr <= '9'))
+			fmt_ptr++;
+		if (*fmt_ptr == '*') {
+			if (buf_ptr + sizeof(int) > buf_end)
+				return -EBADMSG;
+			dyn_width = get_unaligned((int *)buf_ptr);
+			buf_ptr += sizeof(int);
+			fmt_ptr++;
+		}
+
+		/* Consume precision: digits are format-only, * is serialized */
+		if (*fmt_ptr == '.') {
+			fmt_ptr++;
+			while (*fmt_ptr && (*fmt_ptr >= '0' && *fmt_ptr <= '9'))
+				fmt_ptr++;
+			if (*fmt_ptr == '*') {
+				if (buf_ptr + sizeof(int) > buf_end)
+					return -EBADMSG;
+				dyn_precision = get_unaligned((int *)buf_ptr);
+				buf_ptr += sizeof(int);
+				fmt_ptr++;
+			}
+		}
+
+		/*
+		 * dyn_width: consumed from the buffer to keep offsets
+		 * aligned; output width-padding is not implemented.
+		 */
+		(void)dyn_width;
+
+		/* Parse length modifiers (l, ll, h, hh, z) */
+		is_long = 0;
+		is_long_long = 0;
+
+		if (*fmt_ptr == 'l') {
+			fmt_ptr++;
+			is_long = 1;
+			if (*fmt_ptr == 'l') {
+				fmt_ptr++;
+				is_long_long = 1;
+				is_long = 0;
+			}
+		} else if (*fmt_ptr == 'h') {
+			fmt_ptr++;
+			if (*fmt_ptr == 'h')
+				fmt_ptr++;
+		} else if (*fmt_ptr == 'z') {
+			fmt_ptr++;
+			if (sizeof(size_t) == sizeof(long long))
+				is_long_long = 1;
+			else
+				is_long = 1;
+		}
+
+		switch (*fmt_ptr) {
+		case 's': {
+			const char *str;
+			size_t str_len;
+			size_t out_len;
+			size_t max_scan_len;
+
+			if (buf_ptr >= buf_end) {
+				pr_err_ratelimited("blog_des_reconstruct: string arg %zu overruns buffer (no space)\n",
+					       arg_count);
+				return -EBADMSG;
+			}
+
+			str = buf_ptr;
+			max_scan_len = buf_end - buf_ptr;
+
+			str_len = strnlen(str, max_scan_len);
+			if (str_len == max_scan_len && str[str_len - 1] != '\0') {
+				pr_err_ratelimited("blog_des_reconstruct: unterminated string at arg %zu (fmt=%s)\n",
+					       arg_count, fmt);
+				return -EBADMSG;
+			}
+
+			buf_ptr += round_up(str_len + 1, 4);
+			if (buf_ptr > buf_end) {
+				pr_err_ratelimited("blog_des_reconstruct: string arg %zu overruns buffer after copy (fmt=%s)\n",
+					       arg_count, fmt);
+				return -EBADMSG;
+			}
+
+			out_len = str_len;
+			if (dyn_precision >= 0 && (size_t)dyn_precision < out_len)
+				out_len = (size_t)dyn_precision;
+			if (out_len > remaining)
+				out_len = remaining;
+			memcpy(out_ptr, str, out_len);
+			out_ptr += out_len;
+			remaining -= out_len;
+			break;
+		}
+		case 'd':
+		case 'i': {
+			if (is_long_long) {
+				long long val;
+
+				if (buf_ptr + sizeof(long long) > buf_end)
+					return -EBADMSG;
+				val = get_unaligned((long long *)buf_ptr);
+				buf_ptr += sizeof(long long);
+				ret = snprintf(out_ptr, remaining + 1, "%lld", val);
+			} else if (is_long) {
+				long val;
+
+				if (buf_ptr + sizeof(long) > buf_end)
+					return -EBADMSG;
+				val = get_unaligned((long *)buf_ptr);
+				buf_ptr += sizeof(long);
+				ret = snprintf(out_ptr, remaining + 1, "%ld", val);
+			} else {
+				int val;
+
+				if (buf_ptr + sizeof(int) > buf_end)
+					return -EBADMSG;
+				val = get_unaligned((int *)buf_ptr);
+				buf_ptr += sizeof(int);
+				ret = snprintf(out_ptr, remaining + 1, "%d", val);
+			}
+			des_advance(ret, &out_ptr, &remaining);
+			break;
+		}
+		case 'u': {
+			if (is_long_long) {
+				unsigned long long val;
+
+				if (buf_ptr + sizeof(unsigned long long) > buf_end)
+					return -EBADMSG;
+				val = get_unaligned((unsigned long long *)buf_ptr);
+				buf_ptr += sizeof(unsigned long long);
+				ret = snprintf(out_ptr, remaining + 1, "%llu", val);
+			} else if (is_long) {
+				unsigned long val;
+
+				if (buf_ptr + sizeof(unsigned long) > buf_end)
+					return -EBADMSG;
+				val = get_unaligned((unsigned long *)buf_ptr);
+				buf_ptr += sizeof(unsigned long);
+				ret = snprintf(out_ptr, remaining + 1, "%lu", val);
+			} else {
+				unsigned int val;
+
+				if (buf_ptr + sizeof(unsigned int) > buf_end)
+					return -EBADMSG;
+				val = get_unaligned((unsigned int *)buf_ptr);
+				buf_ptr += sizeof(unsigned int);
+				ret = snprintf(out_ptr, remaining + 1, "%u", val);
+			}
+			des_advance(ret, &out_ptr, &remaining);
+			break;
+		}
+		case 'o':
+		case 'x':
+		case 'X': {
+			const char *num_fmt;
+
+			if (*fmt_ptr == 'o')
+				num_fmt = is_long_long ? (alt_form ? "%#llo" : "%llo") :
+					  is_long ? (alt_form ? "%#lo" : "%lo") :
+					  (alt_form ? "%#o" : "%o");
+			else if (*fmt_ptr == 'x')
+				num_fmt = is_long_long ? (alt_form ? "%#llx" : "%llx") :
+					  is_long ? (alt_form ? "%#lx" : "%lx") :
+					  (alt_form ? "%#x" : "%x");
+			else
+				num_fmt = is_long_long ? (alt_form ? "%#llX" : "%llX") :
+					  is_long ? (alt_form ? "%#lX" : "%lX") :
+					  (alt_form ? "%#X" : "%X");
+
+			if (is_long_long) {
+				unsigned long long val;
+
+				if (buf_ptr + sizeof(unsigned long long) > buf_end)
+					return -EBADMSG;
+				val = get_unaligned((unsigned long long *)buf_ptr);
+				buf_ptr += sizeof(unsigned long long);
+				ret = snprintf(out_ptr, remaining + 1, num_fmt, val);
+			} else if (is_long) {
+				unsigned long val;
+
+				if (buf_ptr + sizeof(unsigned long) > buf_end)
+					return -EBADMSG;
+				val = get_unaligned((unsigned long *)buf_ptr);
+				buf_ptr += sizeof(unsigned long);
+				ret = snprintf(out_ptr, remaining + 1, num_fmt, val);
+			} else {
+				unsigned int val;
+
+				if (buf_ptr + sizeof(unsigned int) > buf_end)
+					return -EBADMSG;
+				val = get_unaligned((unsigned int *)buf_ptr);
+				buf_ptr += sizeof(unsigned int);
+				ret = snprintf(out_ptr, remaining + 1, num_fmt, val);
+			}
+			des_advance(ret, &out_ptr, &remaining);
+			break;
+		}
+		case 'p': {
+			void *ptr;
+
+			if (buf_ptr + sizeof(void *) > buf_end)
+				return -EBADMSG;
+
+			ptr = (void *)(unsigned long)get_unaligned((unsigned long *)buf_ptr);
+			buf_ptr += sizeof(void *);
+
+			/*
+			 * Skip kernel %p sub-specifiers (U, I, d, D, etc.).
+				 * boutc does not support %p extensions; call sites
+			 * must pre-format them with snprintf and pass %s.
+			 */
+			while (fmt_ptr[1] && isalnum(fmt_ptr[1]))
+				fmt_ptr++;
+
+			ret = snprintf(out_ptr, remaining + 1, "%p", ptr);
+			des_advance(ret, &out_ptr, &remaining);
+			break;
+		}
+		case 'c': {
+			char val;
+
+			if (buf_ptr + sizeof(int) > buf_end)
+				return -EBADMSG;
+
+			val = (char)get_unaligned((int *)buf_ptr);
+			buf_ptr += sizeof(int);
+
+			if (remaining > 0) {
+				*out_ptr++ = val;
+				remaining--;
+			}
+			break;
+		}
+		default:
+			pr_err_ratelimited("%s: unsupported format specifier '%%%c' at argument %zu\n",
+				       __func__, *fmt_ptr, arg_count);
+			return -EINVAL;
+		}
+
+		fmt_ptr++;
+		arg_count++;
+	}
+
+	*out_ptr = '\0';
+
+	return out_ptr - out;
+}
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 16+ messages in thread

* [PATCH v7 03/14] ceph: add BLOG page-fragment allocator
  2026-09-24 15:30 [PATCH v7 00/14] ceph: add binary logging (BLOG) for CephFS Alex Markuze
  2026-09-24 15:30 ` [PATCH v7 01/14] ceph: add BLOG private headers Alex Markuze
  2026-09-24 15:30 ` [PATCH v7 02/14] ceph: add BLOG deserialization support Alex Markuze
@ 2026-09-24 15:30 ` Alex Markuze
  2026-09-24 15:30 ` [PATCH v7 04/14] ceph: add BLOG magazine batch allocator Alex Markuze
                   ` (11 subsequent siblings)
  14 siblings, 0 replies; 16+ messages in thread
From: Alex Markuze @ 2026-09-24 15:30 UTC (permalink / raw)
  To: ceph-devel; +Cc: idryomov, xiubo.li

Add blog_pagefrag.c: lockless page-fragment buffer allocator for
per-task BLOG contexts.

Signed-off-by: Alex Markuze <amarkuze@redhat.com>
---
 fs/ceph/blog_pagefrag.c | 62 +++++++++++++++++++++++++++++++++++++++++
 1 file changed, 62 insertions(+)
 create mode 100644 fs/ceph/blog_pagefrag.c

diff --git a/fs/ceph/blog_pagefrag.c b/fs/ceph/blog_pagefrag.c
new file mode 100644
index 000000000000..20f57082fdab
--- /dev/null
+++ b/fs/ceph/blog_pagefrag.c
@@ -0,0 +1,62 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * BLOG page-fragment allocator.
+ */
+
+#include "blog_pagefrag.h"
+
+/**
+ * blog_pagefrag_reserve - reserve space without publishing it
+ * @pf: pagefrag allocator
+ * @n: number of bytes
+ *
+ * Returns the current head offset without advancing. Caller writes, then
+ * blog_pagefrag_publish(). Single-writer (per-task).
+ */
+int blog_pagefrag_reserve(struct blog_pagefrag *pf, unsigned int n)
+{
+	if (n > pf->capacity - pf->head)
+		return -ENOMEM;
+	return pf->head;
+}
+
+/**
+ * blog_pagefrag_publish - make reserved bytes visible to readers
+ * @pf: pagefrag allocator
+ * @publish_head: new head (offset + bytes_written)
+ *
+ * Store-release so a reader that observes the new head (load-acquire or
+ * under pf->lock) sees the preceding entry writes. Single-writer.
+ */
+void blog_pagefrag_publish(struct blog_pagefrag *pf, unsigned int publish_head)
+{
+	/* Release-store pairs with the acquire in readers (pf->lock). */
+	smp_store_release(&pf->head, publish_head);
+}
+
+void *blog_pagefrag_get_ptr(struct blog_pagefrag *pf, u64 val)
+{
+	char *base = (char *)pf->buffer;
+	void *rc = base + val;
+
+	if (unlikely(rc < pf->buffer ||
+		     rc >= (void *)(base + pf->capacity))) {
+		WARN_ON_ONCE(1);
+		return NULL;
+	}
+	return rc;
+}
+
+/**
+ * blog_pagefrag_reset - discard stored entries
+ * @pf: pagefrag allocator to reset
+ *
+ * Callers must ensure no writer is mid-reservation. The pagefrag is
+ * single-writer and writers do not take pf->lock.
+ */
+void blog_pagefrag_reset(struct blog_pagefrag *pf)
+{
+	spin_lock(&pf->lock);
+	pf->head = 0;
+	spin_unlock(&pf->lock);
+}
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 16+ messages in thread

* [PATCH v7 04/14] ceph: add BLOG magazine batch allocator
  2026-09-24 15:30 [PATCH v7 00/14] ceph: add binary logging (BLOG) for CephFS Alex Markuze
                   ` (2 preceding siblings ...)
  2026-09-24 15:30 ` [PATCH v7 03/14] ceph: add BLOG page-fragment allocator Alex Markuze
@ 2026-09-24 15:30 ` Alex Markuze
  2026-09-24 15:30 ` [PATCH v7 05/14] ceph: add BLOG logger core Alex Markuze
                   ` (10 subsequent siblings)
  14 siblings, 0 replies; 16+ messages in thread
From: Alex Markuze @ 2026-09-24 15:30 UTC (permalink / raw)
  To: ceph-devel; +Cc: idryomov, xiubo.li

Add blog_batch.c: per-CPU magazine batching for TLS context recycling.
Freed composites go to a local magazine; subsequent acquisitions
reclaim from the magazine without allocating.

Return exhausted magazines to an explicit refill batch. When full
magazines move from the log batch to the allocation batch, returning
empties to the allocation batch strands them: only the log batch
puts elements back. Detach the per-CPU magazine before publishing it
under the refill batch's empty-list lock. Both batches must share
the magazine slab cache.

Reported-by: Xiubo Li <xiubo.li@clyso.com>
Link: https://lore.kernel.org/ceph-devel/CAOJNxRJTiUkSAW6diKfZA51mGMCeGdKsbcb5mdkYxqf+-+e8fQ@mail.gmail.com/
Signed-off-by: Alex Markuze <amarkuze@redhat.com>
Assisted-by: LLM
---
 fs/ceph/blog_batch.c | 267 +++++++++++++++++++++++++++++++++++++++++++
 1 file changed, 267 insertions(+)
 create mode 100644 fs/ceph/blog_batch.c

diff --git a/fs/ceph/blog_batch.c b/fs/ceph/blog_batch.c
new file mode 100644
index 000000000000..c4d2a69e6621
--- /dev/null
+++ b/fs/ceph/blog_batch.c
@@ -0,0 +1,267 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Per-CPU magazine batching for BLOG TLS context recycling.
+ */
+
+#include <linux/slab.h>
+#include <linux/module.h>
+#include <linux/percpu.h>
+#include <linux/preempt.h>
+#include <linux/spinlock.h>
+#include <linux/list.h>
+#include <linux/vmalloc.h>
+#include "blog_batch.h"
+#include "blog.h"
+
+static struct blog_magazine *alloc_magazine(struct blog_batch *batch, gfp_t gfp)
+{
+	struct blog_magazine *mag;
+
+	mag = kmem_cache_zalloc(batch->magazine_cache, gfp);
+	if (!mag)
+		return NULL;
+
+	INIT_LIST_HEAD(&mag->list);
+	mag->count = 0;
+	return mag;
+}
+
+static void free_magazine(struct blog_batch *batch, struct blog_magazine *mag)
+{
+	int i;
+	struct blog_tls_pagefrag *composite;
+
+	for (i = 0; i < mag->count; i++) {
+		composite = mag->elements[i];
+		if (composite)
+			kvfree_atomic(composite);
+	}
+
+	kmem_cache_free(batch->magazine_cache, mag);
+}
+
+/**
+ * blog_batch_init - Initialize the batching system
+ * @batch: Batch structure to initialize
+ * @mag_cache: Slab cache for magazine structs, or NULL to create one
+ * @nr_prealloc: Number of composites to preallocate (0 = none)
+ * @retain_limit: Max composites to retain on put; excess are freed (0 = unlimited)
+ *
+ * Pass nr_prealloc = 0 for batches that start empty (e.g. log_batch).
+ */
+int blog_batch_init(struct blog_batch *batch, struct kmem_cache *mag_cache,
+		    unsigned int nr_prealloc, unsigned int retain_limit)
+{
+	unsigned int nr_mags, i, j;
+	int cpu;
+	struct blog_cpu_magazine *cpu_mag;
+	struct blog_magazine *mag;
+	struct blog_tls_pagefrag *composite;
+
+	batch->nr_full = 0;
+	batch->nr_empty = 0;
+	batch->retain_limit = retain_limit;
+
+	if (mag_cache) {
+		batch->magazine_cache = mag_cache;
+		batch->external_cache = true;
+	} else {
+		batch->magazine_cache = kmem_cache_create("blog_magazine",
+						       sizeof(struct blog_magazine),
+						       0, SLAB_HWCACHE_ALIGN, NULL);
+		if (!batch->magazine_cache)
+			return -ENOMEM;
+		batch->external_cache = false;
+	}
+
+	INIT_LIST_HEAD(&batch->full_magazines);
+	INIT_LIST_HEAD(&batch->empty_magazines);
+	raw_spin_lock_init(&batch->full_lock);
+	raw_spin_lock_init(&batch->empty_lock);
+
+	batch->cpu_magazines = alloc_percpu(struct blog_cpu_magazine);
+	if (!batch->cpu_magazines)
+		goto cleanup_cache;
+
+	for_each_possible_cpu(cpu) {
+		cpu_mag = per_cpu_ptr(batch->cpu_magazines, cpu);
+		cpu_mag->mag = NULL;
+	}
+
+	nr_mags = DIV_ROUND_UP(nr_prealloc, BLOG_MAGAZINE_SIZE);
+	for (i = 0; i < nr_mags; i++) {
+		mag = alloc_magazine(batch, GFP_KERNEL);
+		if (!mag)
+			goto cleanup;
+
+		for (j = 0; j < BLOG_MAGAZINE_SIZE; j++) {
+			composite = kvzalloc(BLOG_TLS_PAGEFRAG_ALLOC_SIZE,
+					      GFP_KERNEL);
+			if (!composite) {
+				free_magazine(batch, mag);
+				goto cleanup;
+			}
+			mag->elements[j] = composite;
+			mag->count++;
+		}
+
+		raw_spin_lock(&batch->full_lock);
+		list_add(&mag->list, &batch->full_magazines);
+		batch->nr_full++;
+		raw_spin_unlock(&batch->full_lock);
+	}
+
+	return 0;
+
+cleanup:
+	blog_batch_cleanup(batch);
+	return -ENOMEM;
+
+cleanup_cache:
+	if (!batch->external_cache && batch->magazine_cache)
+		kmem_cache_destroy(batch->magazine_cache);
+	return -ENOMEM;
+}
+
+void blog_batch_cleanup(struct blog_batch *batch)
+{
+	int cpu;
+	struct blog_magazine *mag, *tmp;
+	struct blog_cpu_magazine *cpu_mag;
+
+	if (batch->cpu_magazines) {
+		for_each_possible_cpu(cpu) {
+			cpu_mag = per_cpu_ptr(batch->cpu_magazines, cpu);
+			if (cpu_mag->mag)
+				free_magazine(batch, cpu_mag->mag);
+		}
+		free_percpu(batch->cpu_magazines);
+	}
+
+	raw_spin_lock(&batch->full_lock);
+	list_for_each_entry_safe(mag, tmp, &batch->full_magazines, list) {
+		list_del(&mag->list);
+		batch->nr_full--;
+		free_magazine(batch, mag);
+	}
+	raw_spin_unlock(&batch->full_lock);
+
+	raw_spin_lock(&batch->empty_lock);
+	list_for_each_entry_safe(mag, tmp, &batch->empty_magazines, list) {
+		list_del(&mag->list);
+		batch->nr_empty--;
+		free_magazine(batch, mag);
+	}
+	raw_spin_unlock(&batch->empty_lock);
+
+	if (!batch->external_cache && batch->magazine_cache)
+		kmem_cache_destroy(batch->magazine_cache);
+
+	batch->magazine_cache = NULL;
+	batch->external_cache = false;
+}
+
+/**
+ * blog_batch_get - Take an element and return exhausted magazines for reuse
+ * @batch: Batch supplying elements
+ * @recycle: Batch that will refill exhausted magazines
+ *
+ * The batches must share a magazine cache. Pass the same batch to reuse
+ * magazines locally, or its producer when elements move between batches.
+ */
+void *blog_batch_get(struct blog_batch *batch, struct blog_batch *recycle)
+{
+	struct blog_cpu_magazine *cpu_mag;
+	struct blog_magazine *old_mag, *new_mag;
+	void *element = NULL;
+
+	preempt_disable();
+	cpu_mag = this_cpu_ptr(batch->cpu_magazines);
+
+	if (cpu_mag->mag && cpu_mag->mag->count > 0) {
+		element = cpu_mag->mag->elements[--cpu_mag->mag->count];
+		goto out;
+	}
+
+	old_mag = cpu_mag->mag;
+
+	if (old_mag) {
+		cpu_mag->mag = NULL;
+		raw_spin_lock(&recycle->empty_lock);
+		list_add(&old_mag->list, &recycle->empty_magazines);
+		recycle->nr_empty++;
+		raw_spin_unlock(&recycle->empty_lock);
+	}
+
+	if (READ_ONCE(batch->nr_full) > 0) {
+		raw_spin_lock(&batch->full_lock);
+		if (!list_empty(&batch->full_magazines)) {
+			new_mag = list_first_entry(&batch->full_magazines,
+						   struct blog_magazine, list);
+			list_del(&new_mag->list);
+			batch->nr_full--;
+			raw_spin_unlock(&batch->full_lock);
+
+			cpu_mag->mag = new_mag;
+			if (new_mag->count > 0)
+				element = new_mag->elements[--new_mag->count];
+		} else {
+			raw_spin_unlock(&batch->full_lock);
+		}
+	}
+out:
+	preempt_enable();
+	return element;
+}
+
+bool blog_batch_put(struct blog_batch *batch, void *element)
+{
+	struct blog_cpu_magazine *cpu_mag;
+	struct blog_magazine *mag;
+	bool stored = true;
+
+	/* Trim: if over retention limit, decline to store the element */
+	if (batch->retain_limit &&
+	    READ_ONCE(batch->nr_full) * BLOG_MAGAZINE_SIZE >= batch->retain_limit)
+		return false;
+
+	preempt_disable();
+	cpu_mag = this_cpu_ptr(batch->cpu_magazines);
+
+	if (likely(cpu_mag->mag && cpu_mag->mag->count < BLOG_MAGAZINE_SIZE)) {
+		cpu_mag->mag->elements[cpu_mag->mag->count++] = element;
+		goto out;
+	}
+
+	if (likely(cpu_mag->mag && cpu_mag->mag->count >= BLOG_MAGAZINE_SIZE)) {
+		raw_spin_lock(&batch->full_lock);
+		list_add_tail(&cpu_mag->mag->list, &batch->full_magazines);
+		batch->nr_full++;
+		raw_spin_unlock(&batch->full_lock);
+		cpu_mag->mag = NULL;
+	}
+
+	if (likely(!cpu_mag->mag)) {
+		raw_spin_lock(&batch->empty_lock);
+		if (!list_empty(&batch->empty_magazines)) {
+			mag = list_first_entry(&batch->empty_magazines,
+					       struct blog_magazine, list);
+			list_del(&mag->list);
+			batch->nr_empty--;
+			raw_spin_unlock(&batch->empty_lock);
+			cpu_mag->mag = mag;
+		} else {
+			raw_spin_unlock(&batch->empty_lock);
+			cpu_mag->mag = alloc_magazine(batch, GFP_ATOMIC);
+		}
+
+		if (unlikely(!cpu_mag->mag)) {
+			stored = false;
+			goto out;
+		}
+	}
+	cpu_mag->mag->elements[cpu_mag->mag->count++] = element;
+out:
+	preempt_enable();
+	return stored;
+}
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 16+ messages in thread

* [PATCH v7 05/14] ceph: add BLOG logger core
  2026-09-24 15:30 [PATCH v7 00/14] ceph: add binary logging (BLOG) for CephFS Alex Markuze
                   ` (3 preceding siblings ...)
  2026-09-24 15:30 ` [PATCH v7 04/14] ceph: add BLOG magazine batch allocator Alex Markuze
@ 2026-09-24 15:30 ` Alex Markuze
  2026-09-24 15:30 ` [PATCH v7 06/14] ceph: add BLOG per-module context management Alex Markuze
                   ` (9 subsequent siblings)
  14 siblings, 0 replies; 16+ messages in thread
From: Alex Markuze @ 2026-09-24 15:30 UTC (permalink / raw)
  To: ceph-devel; +Cc: idryomov, xiubo.li

Add blog_core.c: central logger, source-ID registry with per-callsite
caching (smp_store_release/smp_load_acquire plus generation counter),
circular entry buffer, and iteration API for debugfs consumers.

blog_log_client_emit() is the noinline helper that writes a record
from TLS argument scratch, so boutc() call sites do not allocate
struct blog_arg[] on the VFS stack.

Signed-off-by: Alex Markuze <amarkuze@redhat.com>
---
 fs/ceph/blog_core.c | 297 ++++++++++++++++++++++++++++++++++++++++++++
 1 file changed, 297 insertions(+)
 create mode 100644 fs/ceph/blog_core.c

diff --git a/fs/ceph/blog_core.c b/fs/ceph/blog_core.c
new file mode 100644
index 000000000000..3b5bca332625
--- /dev/null
+++ b/fs/ceph/blog_core.c
@@ -0,0 +1,297 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * BLOG logger: source-ID registry and log iteration.
+ */
+
+#include <linux/module.h>
+#include <linux/kernel.h>
+#include <linux/init.h>
+#include <linux/slab.h>
+#include <linux/string.h>
+#include <linux/printk.h>
+#include <linux/time.h>
+#include <linux/percpu.h>
+#include <linux/spinlock.h>
+#include <linux/list.h>
+#include <linux/sched.h>
+#include <linux/hash.h>
+
+#include <linux/rhashtable.h>
+#include "blog.h"
+#include "blog_batch.h"
+#include "blog_pagefrag.h"
+#include "blog_ser.h"
+#include "blog_des.h"
+#include "blog_module.h"
+
+static bool blog_source_matches(const struct blog_source_info *info,
+				const char *file, const char *func,
+				unsigned int line, const char *fmt)
+{
+	return info->file && info->func && info->fmt &&
+	       info->line == line && info->fmt == fmt &&
+	       !strcmp(info->file, file) && !strcmp(info->func, func);
+}
+
+static u32 blog_source_hash(const char *fmt, unsigned int line, u32 mask)
+{
+	return (hash_ptr((void *)fmt, 32) ^ line) & mask;
+}
+
+/**
+ * blog_get_source_id - Get or create a source ID for the given location
+ * @logger: Logger instance to use
+ * @file: Source file name
+ * @func: Function name
+ * @line: Line number
+ * @fmt: Format string
+ */
+u32 blog_get_source_id(struct blog_logger *logger, const char *file,
+		       const char *func, unsigned int line, const char *fmt)
+{
+	struct blog_source_info *info;
+	u32 id, slot, first;
+
+	if (!logger || !logger->source_hash)
+		return 0;
+
+	spin_lock(&logger->source_lock);
+	slot = blog_source_hash(fmt, line, logger->source_hash_mask);
+	first = slot;
+	do {
+		id = logger->source_hash[slot];
+		if (!id)
+			break;
+		info = &logger->source_map[id];
+		if (blog_source_matches(info, file, func, line, fmt))
+			goto out_unlock;
+		slot = (slot + 1) & logger->source_hash_mask;
+	} while (slot != first);
+
+	id = logger->next_source_id;
+	if (id >= logger->max_source_ids) {
+		spin_unlock(&logger->source_lock);
+		pr_warn_once("blog: source ID overflow\n");
+		return 0;
+	}
+
+	/* No empty slot in the hash table. */
+	if (logger->source_hash[slot]) {
+		spin_unlock(&logger->source_lock);
+		pr_warn_once("blog: source hash full\n");
+		return 0;
+	}
+
+	logger->next_source_id = id + 1;
+	info = &logger->source_map[id];
+	info->file = file;
+	info->func = func;
+	info->line = line;
+	info->fmt = fmt;
+	info->warn_count = 0;
+	logger->source_hash[slot] = id;
+
+out_unlock:
+	spin_unlock(&logger->source_lock);
+	return id;
+}
+
+u32 blog_get_source_id_cached(struct blog_logger *logger,
+			      struct blog_source_id_cache *cache,
+			      const char *file, const char *func,
+			      unsigned int line, const char *fmt)
+{
+	struct blog_logger *cached_logger;
+	u64 generation;
+	unsigned int seq;
+	u32 sid;
+
+	if (!logger)
+		return 0;
+	if (!cache)
+		return blog_get_source_id(logger, file, func, line, fmt);
+
+	/* Callsites are shared by mounts, so read one lockless snapshot. */
+	do {
+		seq = read_seqcount_begin(&cache->seq);
+		sid = cache->id;
+		cached_logger = cache->logger;
+		generation = cache->generation;
+	} while (read_seqcount_retry(&cache->seq, seq));
+	if (sid && cached_logger == logger &&
+	    generation == logger->generation)
+		return sid;
+
+	spin_lock(&cache->lock);
+	sid = cache->id;
+	if (sid && cache->logger == logger &&
+	    cache->generation == logger->generation)
+		goto out;
+
+	sid = blog_get_source_id(logger, file, func, line, fmt);
+	if (sid) {
+		write_seqcount_begin(&cache->seq);
+		cache->logger = logger;
+		cache->generation = logger->generation;
+		cache->id = sid;
+		write_seqcount_end(&cache->seq);
+	}
+
+out:
+	spin_unlock(&cache->lock);
+	return sid;
+}
+
+struct blog_source_info *blog_get_source_info(struct blog_logger *logger, u32 id)
+{
+	if (!logger || unlikely(id == 0 || id >= logger->max_source_ids))
+		return NULL;
+	return &logger->source_map[id];
+}
+
+void blog_log_iter_init(struct blog_log_iter *iter, struct blog_pagefrag *pf,
+			u64 head_snapshot)
+{
+	if (!iter || !pf)
+		return;
+
+	iter->pf = pf;
+	iter->current_offset = 0;
+	iter->end_offset = head_snapshot;
+	iter->prev_offset = 0;
+	iter->steps = 0;
+}
+
+struct blog_log_entry *blog_log_iter_next(struct blog_log_iter *iter)
+{
+	struct blog_log_entry *entry;
+
+	if (!iter || iter->current_offset >= iter->end_offset)
+		return NULL;
+
+	/* Ensure the entry header itself fits within the snapshot. */
+	if (iter->current_offset + sizeof(struct blog_log_entry) >
+	    iter->end_offset)
+		return NULL;
+
+	entry = blog_pagefrag_get_ptr(iter->pf, iter->current_offset);
+	if (!entry)
+		return NULL;
+
+	/* Reject truncated / corrupt payloads before deserializing. */
+	if (iter->current_offset + sizeof(*entry) + entry->len >
+	    iter->end_offset)
+		return NULL;
+
+	iter->prev_offset = iter->current_offset;
+	iter->current_offset +=
+		round_up(sizeof(struct blog_log_entry) + entry->len, 8);
+	iter->steps++;
+
+	/*
+	 * Clamp to the snapshot boundary: a corrupted entry->len could
+	 * push current_offset past end_offset into garbage memory.
+	 */
+	if (iter->current_offset > iter->end_offset)
+		iter->current_offset = iter->end_offset;
+
+	return entry;
+}
+
+int blog_des_entry(struct blog_logger *logger, struct blog_log_entry *entry,
+		   char *output, size_t out_size, blog_client_des_fn client_cb)
+{
+	int len = 0;
+	struct blog_source_info *source;
+
+	if (!entry || !output)
+		return -EINVAL;
+
+	if (client_cb) {
+		len = client_cb(output, out_size, entry->client_id);
+		if (len < 0)
+			return len;
+		if (len >= out_size)
+			return len;
+	}
+
+	source = blog_get_source_info(logger, entry->source_id);
+	if (!source) {
+		len += scnprintf(output + len, out_size - len,
+				 "[unknown source %u]", entry->source_id);
+		return len;
+	}
+
+	/* Snapshot under source_lock (same pattern as blog_sources_show). */
+	{
+		const char *file, *func, *fmt;
+		unsigned int line;
+		int ret;
+
+		spin_lock(&logger->source_lock);
+		file = source->file;
+		func = source->func;
+		line = source->line;
+		fmt = source->fmt;
+		spin_unlock(&logger->source_lock);
+		if (!file) {
+			len += scnprintf(output + len, out_size - len,
+					 "[unknown source %u]",
+					 entry->source_id);
+			return len;
+		}
+
+		len += scnprintf(output + len, out_size - len, "[%s:%s:%u] ",
+				 file, func, line);
+		if (len >= out_size)
+			return len;
+
+		ret = blog_des_reconstruct(fmt, entry->buffer, entry->len,
+					   output + len, out_size - len);
+		if (ret < 0)
+			return ret;
+		len += ret;
+	}
+
+	return len;
+}
+
+noinline void blog_log_client_emit(struct blog_tls_ctx *ctx,
+				   struct ceph_client *client,
+				   struct blog_source_id_cache *cache,
+				   const char *file, const char *func,
+				   unsigned int line, const char *fmt,
+				   size_t nargs)
+{
+	struct blog_logger *logger;
+	struct blog_arg *args;
+	size_t size;
+	void *buffer;
+	void *tmp;
+	u32 sid;
+	u32 client_id;
+
+	BUILD_BUG_ON(sizeof(struct blog_tls_pagefrag) >= BLOG_PAGEFRAG_SIZE);
+
+	if (unlikely(!ctx || !cache || nargs > BLOG_MAX_ARGS))
+		return;
+
+	logger = ctx->logger;
+	if (unlikely(!logger))
+		return;
+
+	args = ctx->arg_scratch;
+	sid = blog_get_source_id_cached(logger, cache, file, func, line, fmt);
+	if (unlikely(!sid))
+		return;
+
+	client_id = ceph_blog_get_client_id(client);
+	size = blog_args_size(args, nargs);
+	buffer = blog_log_with_ctx(logger, ctx, sid, client_id, size);
+	if (!buffer)
+		return;
+
+	tmp = buffer;
+	buffer = blog_serialize_args(buffer, args, nargs);
+	blog_log_commit_with_ctx(logger, ctx, buffer - tmp);
+}
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 16+ messages in thread

* [PATCH v7 06/14] ceph: add BLOG per-module context management
  2026-09-24 15:30 [PATCH v7 00/14] ceph: add binary logging (BLOG) for CephFS Alex Markuze
                   ` (4 preceding siblings ...)
  2026-09-24 15:30 ` [PATCH v7 05/14] ceph: add BLOG logger core Alex Markuze
@ 2026-09-24 15:30 ` Alex Markuze
  2026-09-24 15:30 ` [PATCH v7 07/14] ceph: add Ceph BLOG scaffolding Alex Markuze
                   ` (8 subsequent siblings)
  14 siblings, 0 replies; 16+ messages in thread
From: Alex Markuze @ 2026-09-24 15:30 UTC (permalink / raw)
  To: ceph-devel; +Cc: idryomov, xiubo.li

Add blog_module.c: per-superblock rhashtable mapping tasks to
blog_task_entry structures, with context acquire/release and
retirement that clears the RETIRED bit on removal failure.

Recycle exhausted allocation magazines back to log_batch from both
task-context allocation and atomic buffer rotation. This closes the
full/empty magazine cycle, so context churn reuses empty containers
instead of retaining them until unmount and allocating replacements.

Capture record base times with get_jiffies_64() and calculate the
delta at full width before storing it in u32. On 32-bit kernels,
the low jiffies word first wraps about five minutes after boot;
an unsigned long base would lose that epoch. Protect published
base resets with the page-fragment lock used by readers.

Use one clock sample for both the delta overflow check and the
encoded timestamp. After rotation, use the new base. A tick
between separate samples must not wrap an exactly U32_MAX delta.

Reported-by: Xiubo Li <xiubo.li@clyso.com>
Link: https://lore.kernel.org/ceph-devel/CAOJNxRJTiUkSAW6diKfZA51mGMCeGdKsbcb5mdkYxqf+-+e8fQ@mail.gmail.com/
Signed-off-by: Alex Markuze <amarkuze@redhat.com>
Assisted-by: LLM
---
 fs/ceph/blog_module.c | 1003 +++++++++++++++++++++++++++++++++++++++++
 1 file changed, 1003 insertions(+)
 create mode 100644 fs/ceph/blog_module.c

diff --git a/fs/ceph/blog_module.c b/fs/ceph/blog_module.c
new file mode 100644
index 000000000000..68f1aa770249
--- /dev/null
+++ b/fs/ceph/blog_module.c
@@ -0,0 +1,1003 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Per-superblock BLOG context. Tasks map to TLS contexts via rhashtable.
+ * Task references keep keys stable until exit reaping removes them.
+ */
+
+#include <linux/module.h>
+#include <linux/slab.h>
+#include <linux/spinlock.h>
+#include <linux/list.h>
+#include <linux/atomic.h>
+#include <linux/log2.h>
+#include <linux/refcount.h>
+#include <linux/gfp.h>
+#include <linux/sched.h>
+#include <linux/sched/signal.h>
+#include <linux/rhashtable.h>
+#include <linux/vmalloc.h>
+#include <linux/workqueue.h>
+#include "blog.h"
+#include "blog_module.h"
+
+static atomic64_t blog_logger_gen = ATOMIC64_INIT(1);
+static struct workqueue_struct *blog_free_wq;
+
+static void blog_module_free_workfn(struct work_struct *work);
+
+#define BLOG_LOG_BATCH_MAX_FULL 16
+#define BLOG_LOG_BATCH_MAX_FULL_CAP 1024
+#define BLOG_MAX_TASK_CONTEXTS 64
+#define BLOG_TASK_GC_INTERVAL (5 * HZ)
+#define BLOG_TASK_ENTRY_RETIRED 0
+
+static int blog_max_full = BLOG_LOG_BATCH_MAX_FULL;
+module_param_named(blog_max_full, blog_max_full, int, 0644);
+MODULE_PARM_DESC(blog_max_full,
+		 "Visible full-magazine count for BLOG entries (default 16)");
+
+static unsigned int blog_param_max_full(void)
+{
+	int n = READ_ONCE(blog_max_full);
+
+	if (n < 1)
+		n = 1;
+	if (n > BLOG_LOG_BATCH_MAX_FULL_CAP)
+		n = BLOG_LOG_BATCH_MAX_FULL_CAP;
+	return n;
+}
+
+static const struct rhashtable_params blog_task_ht_params = {
+	.key_offset = offsetof(struct blog_task_entry, task),
+	.key_len = sizeof(struct task_struct *),
+	.head_offset = offsetof(struct blog_task_entry, node),
+	.automatic_shrinking = true,
+};
+
+static void blog_module_free_magazine(struct blog_logger *logger,
+				      struct blog_magazine *mag)
+{
+	int i;
+
+	for (i = 0; i < mag->count; i++)
+		kvfree_atomic(mag->elements[i]);
+	kmem_cache_free(logger->magazine_cache, mag);
+}
+
+static void blog_module_hide_magazine(struct blog_logger *logger,
+				      struct blog_magazine *mag)
+{
+	struct blog_tls_pagefrag *composite;
+	int i;
+
+	spin_lock(&logger->lock);
+	for (i = 0; i < mag->count; i++) {
+		composite = mag->elements[i];
+		if (!list_empty(&composite->ctx.list))
+			list_del_init(&composite->ctx.list);
+		/*
+		 * Drop the publication ID on pool return so the next
+		 * rotate/reuse takes a fresh monotonic ID.  Reusing the
+		 * old ID lets blog_entries_show()'s cursor skip the
+		 * newly published snapshot (silent loss) or reprint it.
+		 */
+		composite->ctx.id = 0;
+	}
+	spin_unlock(&logger->lock);
+}
+
+static void blog_module_rebalance_log_batch(struct blog_logger *logger)
+{
+	struct blog_magazine *mag;
+	bool retain;
+
+	if (!logger || logger->log_batch.nr_full <= blog_param_max_full())
+		return;
+
+	raw_spin_lock(&logger->log_batch.full_lock);
+	if (list_empty(&logger->log_batch.full_magazines)) {
+		raw_spin_unlock(&logger->log_batch.full_lock);
+		return;
+	}
+	mag = list_first_entry(&logger->log_batch.full_magazines,
+			       struct blog_magazine, list);
+	list_del(&mag->list);
+	logger->log_batch.nr_full--;
+	raw_spin_unlock(&logger->log_batch.full_lock);
+
+	/* Keep retained records visible until this magazine is reclaimed. */
+	mutex_lock(&logger->snapshot_mutex);
+	blog_module_hide_magazine(logger, mag);
+
+	raw_spin_lock(&logger->alloc_batch.full_lock);
+	retain = !logger->alloc_batch.retain_limit ||
+		 logger->alloc_batch.nr_full * BLOG_MAGAZINE_SIZE <
+		 logger->alloc_batch.retain_limit;
+	if (retain) {
+		list_add(&mag->list, &logger->alloc_batch.full_magazines);
+		logger->alloc_batch.nr_full++;
+	}
+	raw_spin_unlock(&logger->alloc_batch.full_lock);
+	if (!retain)
+		blog_module_free_magazine(logger, mag);
+	mutex_unlock(&logger->snapshot_mutex);
+}
+
+static void blog_module_schedule_log_reclaim(struct blog_logger *logger)
+{
+	if (logger && blog_free_wq && !READ_ONCE(logger->task_gc_stopping))
+		queue_work(blog_free_wq, &logger->reclaim_work);
+}
+
+static void blog_module_reclaim_workfn(struct work_struct *work)
+{
+	struct blog_logger *logger =
+		container_of(work, struct blog_logger, reclaim_work);
+	LIST_HEAD(to_free);
+	struct blog_tls_ctx *ctx, *tmp;
+	bool again;
+
+	mutex_lock(&logger->snapshot_mutex);
+	spin_lock(&logger->lock);
+	list_splice_init(&logger->reclaim_list, &to_free);
+	spin_unlock(&logger->lock);
+
+	list_for_each_entry_safe(ctx, tmp, &to_free, list) {
+		list_del_init(&ctx->list);
+		kvfree_atomic(blog_ctx_container(ctx));
+	}
+	mutex_unlock(&logger->snapshot_mutex);
+
+	while (logger->log_batch.nr_full > blog_param_max_full())
+		blog_module_rebalance_log_batch(logger);
+
+	spin_lock(&logger->lock);
+	again = !list_empty(&logger->reclaim_list);
+	spin_unlock(&logger->lock);
+	if (again)
+		blog_module_schedule_log_reclaim(logger);
+}
+
+static void blog_module_queue_to_log_batch(struct blog_logger *logger,
+					   struct blog_tls_ctx *ctx)
+{
+	struct blog_tls_pagefrag *composite;
+
+	if (!logger || !ctx)
+		return;
+
+	/*
+	 * Only task-map contexts are charged against allocated_contexts.
+	 * Rotate-on-full snap buffers share this recycle path but were
+	 * never counted. Do not let them underflow the 64-task cap.
+	 */
+	if (test_and_clear_bit(BLOG_CTX_TASK_COUNTED, &ctx->flags) &&
+	    logger->owner_ctx)
+		atomic_dec(&logger->owner_ctx->allocated_contexts);
+	composite = blog_ctx_container(ctx);
+	atomic_set(&ctx->refcount, 0);
+	ctx->pending_offset = 0;
+	ctx->pending_size = 0;
+	if (!blog_batch_put(&logger->log_batch, composite)) {
+		mutex_lock(&logger->snapshot_mutex);
+		spin_lock(&logger->lock);
+		if (!list_empty(&ctx->list))
+			list_del_init(&ctx->list);
+		spin_unlock(&logger->lock);
+		kvfree_atomic(composite);
+		mutex_unlock(&logger->snapshot_mutex);
+	}
+	blog_module_rebalance_log_batch(logger);
+}
+
+static void blog_module_clear_task(struct blog_tls_ctx *ctx)
+{
+	if (ctx) {
+		ceph_blog_cpu_clear(ctx);
+		WRITE_ONCE(ctx->task, NULL);
+	}
+}
+
+static void blog_module_tls_release(void *ptr)
+{
+	struct blog_tls_ctx *ctx = ptr;
+	struct blog_logger *logger;
+
+	if (!ctx)
+		return;
+	logger = ctx->logger;
+	if (!logger) {
+		pr_err("BUG: TLS context id=%llu has no logger\n", ctx->id);
+		return;
+	}
+	blog_module_clear_task(ctx);
+	blog_module_queue_to_log_batch(logger, ctx);
+}
+
+/*
+ * Retire a stale rhashtable entry: remove it from the hash table, retain its
+ * TLS context on the reader-visible list while the log batch owns it, and
+ * schedule the entry for RCU-deferred freeing.
+ */
+static bool blog_retire_entry_locked(struct blog_logger *logger,
+				     struct blog_task_entry *entry,
+				     struct blog_tls_ctx **tls_ctx)
+{
+	if (test_and_set_bit(BLOG_TASK_ENTRY_RETIRED, &entry->flags))
+		return false;
+	if (rhashtable_remove_fast(&logger->task_map, &entry->node,
+				   blog_task_ht_params)) {
+		clear_bit(BLOG_TASK_ENTRY_RETIRED, &entry->flags);
+		return false;
+	}
+
+	*tls_ctx = entry->ctx;
+	return true;
+}
+
+static bool blog_retire_stale_task(struct blog_logger *logger,
+				   struct task_struct *task)
+{
+	struct blog_task_entry *entry;
+	struct blog_tls_ctx *tls_ctx = NULL;
+	bool retired = false;
+
+	spin_lock(&logger->lock);
+	rcu_read_lock();
+	entry = rhashtable_lookup_fast(&logger->task_map, &task,
+				       blog_task_ht_params);
+	if (entry && entry->pid != task->pid)
+		retired = blog_retire_entry_locked(logger, entry, &tls_ctx);
+	rcu_read_unlock();
+	spin_unlock(&logger->lock);
+
+	if (tls_ctx) {
+		blog_module_clear_task(tls_ctx);
+		blog_module_queue_to_log_batch(logger, tls_ctx);
+	}
+	if (retired) {
+		put_task_struct(entry->task);
+		kfree_rcu(entry, rcu);
+	}
+	return retired;
+}
+
+static bool blog_retire_dead_task(struct blog_logger *logger,
+				  struct task_struct *task, pid_t pid)
+{
+	struct blog_task_entry *entry;
+	struct blog_tls_ctx *tls_ctx = NULL;
+	bool retired = false;
+
+	spin_lock(&logger->lock);
+	rcu_read_lock();
+	entry = rhashtable_lookup_fast(&logger->task_map, &task,
+				       blog_task_ht_params);
+	if (entry && entry->pid == pid && !pid_alive(entry->task))
+		retired = blog_retire_entry_locked(logger, entry, &tls_ctx);
+	rcu_read_unlock();
+	spin_unlock(&logger->lock);
+
+	if (tls_ctx) {
+		blog_module_clear_task(tls_ctx);
+		blog_module_queue_to_log_batch(logger, tls_ctx);
+	}
+	if (retired) {
+		put_task_struct(entry->task);
+		kfree_rcu(entry, rcu);
+	}
+	return retired;
+}
+
+struct blog_gc_item {
+	struct list_head list;
+	struct task_struct *task;
+	pid_t pid;
+};
+
+static void blog_gc_dead_tasks(struct blog_logger *logger)
+{
+	struct rhashtable_iter iter;
+	struct blog_task_entry *entry;
+	struct blog_gc_item *item, *tmp;
+	LIST_HEAD(dead);
+
+	rhashtable_walk_enter(&logger->task_map, &iter);
+	rhashtable_walk_start(&iter);
+	for (;;) {
+		entry = rhashtable_walk_next(&iter);
+		if (IS_ERR(entry)) {
+			if (PTR_ERR(entry) == -EAGAIN)
+				continue;
+			break;
+		}
+		if (!entry)
+			break;
+		if (!pid_alive(entry->task)) {
+			item = kmalloc(sizeof(*item), GFP_ATOMIC);
+			if (!item)
+				continue;
+			item->task = entry->task;
+			item->pid = entry->pid;
+			list_add_tail(&item->list, &dead);
+		}
+	}
+	rhashtable_walk_stop(&iter);
+	rhashtable_walk_exit(&iter);
+
+	list_for_each_entry_safe(item, tmp, &dead, list) {
+		blog_retire_dead_task(logger, item->task, item->pid);
+		list_del(&item->list);
+		kfree(item);
+	}
+}
+
+static void blog_task_gc_workfn(struct work_struct *work)
+{
+	struct blog_logger *logger =
+		container_of(to_delayed_work(work), struct blog_logger,
+			     task_gc_work);
+
+	blog_gc_dead_tasks(logger);
+	if (!READ_ONCE(logger->task_gc_stopping))
+		schedule_delayed_work(&logger->task_gc_work,
+				      BLOG_TASK_GC_INTERVAL);
+}
+
+/*
+ * Allocate a fresh TLS context (composite) from the magazine batch or
+ * the page allocator, initialize it, and link it into the logger.
+ */
+static struct blog_tls_ctx *blog_alloc_tls_ctx(struct blog_logger *logger,
+						 gfp_t gfp)
+{
+	struct blog_tls_pagefrag *composite;
+	struct blog_tls_ctx *tls_ctx;
+	struct blog_pagefrag *pf;
+	struct task_struct *task = current;
+
+	composite = blog_batch_get(&logger->alloc_batch, &logger->log_batch);
+	if (!composite)
+		composite = kvzalloc(BLOG_TLS_PAGEFRAG_ALLOC_SIZE, gfp);
+	if (!composite)
+		return NULL;
+
+	tls_ctx = &composite->ctx;
+
+	if (tls_ctx->id == 0) {
+		INIT_LIST_HEAD(&tls_ctx->list);
+		spin_lock(&logger->ctx_id_lock);
+		tls_ctx->id = logger->next_ctx_id++;
+		spin_unlock(&logger->ctx_id_lock);
+	}
+
+	atomic_set(&tls_ctx->refcount, 1);
+	tls_ctx->task = task;
+	tls_ctx->pid = task->pid;
+	get_task_comm(tls_ctx->comm, task);
+	tls_ctx->base_jiffies = get_jiffies_64();
+	tls_ctx->release = blog_module_tls_release;
+	tls_ctx->logger = logger;
+	tls_ctx->flags = 0;
+	tls_ctx->pending_offset = 0;
+	tls_ctx->pending_size = 0;
+	WRITE_ONCE(tls_ctx->enter_depth, 0);
+	WRITE_ONCE(tls_ctx->cache_cpu, -1);
+	atomic64_set(&tls_ctx->clear_seq, atomic64_read(&logger->clear_seq));
+
+	pf = &composite->pf;
+	pf->buffer = composite->buf;
+	pf->capacity = BLOG_TLS_PAGEFRAG_BUFFER_SIZE;
+	spin_lock_init(&pf->lock);
+	pf->head = 0;
+
+	spin_lock(&logger->lock);
+	if (list_empty(&tls_ctx->list)) {
+		list_add(&tls_ctx->list, &logger->contexts);
+		logger->total_contexts_allocated++;
+	}
+	spin_unlock(&logger->lock);
+
+	return tls_ctx;
+}
+
+struct blog_module_context *blog_module_init(const char *module_name)
+{
+	struct blog_module_context *ctx;
+	struct blog_logger *logger;
+	char cache_name[48];
+	int ret;
+
+	if (!module_name || !*module_name)
+		return NULL;
+	if (strlen(module_name) >= sizeof(ctx->name))
+		return NULL;
+
+	ctx = kzalloc(sizeof(*ctx), GFP_KERNEL);
+	if (!ctx)
+		return NULL;
+
+	logger = kzalloc(sizeof(*logger), GFP_KERNEL);
+	if (!logger)
+		goto err_ctx;
+
+	logger->generation = atomic64_inc_return(&blog_logger_gen);
+	snprintf(cache_name, sizeof(cache_name), "blog_magazine_%llu",
+		 (unsigned long long)logger->generation);
+	logger->magazine_cache = kmem_cache_create(cache_name,
+						   sizeof(struct blog_magazine),
+						   0, SLAB_HWCACHE_ALIGN, NULL);
+	if (!logger->magazine_cache)
+		goto err_logger;
+
+	logger->max_source_ids = blog_param_max_sources();
+	logger->source_map = kvcalloc(logger->max_source_ids,
+				      sizeof(struct blog_source_info),
+				      GFP_KERNEL);
+	if (!logger->source_map)
+		goto err_cache;
+
+	{
+		u32 hash_size = roundup_pow_of_two(logger->max_source_ids) * 2;
+
+		if (hash_size < 4)
+			hash_size = 4;
+		logger->source_hash = kvcalloc(hash_size, sizeof(u32), GFP_KERNEL);
+		if (!logger->source_hash)
+			goto err_source_map;
+		logger->source_hash_mask = hash_size - 1;
+	}
+
+	strscpy(ctx->name, module_name, sizeof(ctx->name));
+	ctx->logger = logger;
+	refcount_set(&ctx->refcount, 1);
+	atomic_set(&ctx->allocated_contexts, 0);
+	INIT_WORK(&ctx->free_work, blog_module_free_workfn);
+
+	INIT_LIST_HEAD(&logger->contexts);
+	INIT_LIST_HEAD(&logger->reclaim_list);
+	spin_lock_init(&logger->lock);
+	mutex_init(&logger->snapshot_mutex);
+	spin_lock_init(&logger->source_lock);
+	spin_lock_init(&logger->ctx_id_lock);
+	logger->next_source_id = 1;
+	logger->next_ctx_id = 1;
+	logger->total_contexts_allocated = 0;
+	logger->owner_ctx = ctx;
+	atomic64_set(&logger->clear_seq, 0);
+	INIT_DELAYED_WORK(&logger->task_gc_work, blog_task_gc_workfn);
+	INIT_WORK(&logger->reclaim_work, blog_module_reclaim_workfn);
+	logger->task_gc_stopping = false;
+
+	ret = rhashtable_init(&logger->task_map, &blog_task_ht_params);
+	if (ret)
+		goto err_source_hash;
+
+	ret = blog_batch_init(&logger->alloc_batch, logger->magazine_cache,
+			      0,
+			      num_possible_cpus() + 32);
+	if (ret)
+		goto err_ht;
+
+	ret = blog_batch_init(&logger->log_batch, logger->magazine_cache, 0, 0);
+	if (ret)
+		goto err_batch_alloc;
+
+	schedule_delayed_work(&logger->task_gc_work, BLOG_TASK_GC_INTERVAL);
+
+	ctx->initialized = true;
+	pr_debug("BLOG: module '%s' initialized\n", module_name);
+	return ctx;
+
+err_batch_alloc:
+	blog_batch_cleanup(&logger->alloc_batch);
+err_ht:
+	rhashtable_destroy(&logger->task_map);
+err_source_hash:
+	kvfree(logger->source_hash);
+err_source_map:
+	kvfree(logger->source_map);
+err_cache:
+	kmem_cache_destroy(logger->magazine_cache);
+err_logger:
+	kfree(logger);
+err_ctx:
+	kfree(ctx);
+	return NULL;
+}
+
+/*
+ * Walk callback for rhashtable_free_and_destroy -- release each
+ * task entry and its associated TLS context.
+ */
+static void blog_task_entry_free_cb(void *ptr, void *arg)
+{
+	struct blog_task_entry *entry = ptr;
+	struct blog_logger *logger = arg;
+
+	if (entry->ctx) {
+		spin_lock(&logger->lock);
+		if (!list_empty(&entry->ctx->list))
+			list_del_init(&entry->ctx->list);
+		spin_unlock(&logger->lock);
+
+		blog_module_clear_task(entry->ctx);
+		blog_module_queue_to_log_batch(logger, entry->ctx);
+	}
+	put_task_struct(entry->task);
+	kfree(entry);
+}
+
+static void blog_module_free(struct blog_module_context *ctx)
+{
+	struct blog_logger *logger;
+	struct blog_tls_ctx *tls_ctx, *tmp;
+	LIST_HEAD(pending);
+	LIST_HEAD(reclaim);
+
+	if (!ctx || !ctx->initialized)
+		return;
+	logger = ctx->logger;
+	if (!logger)
+		return;
+
+	WRITE_ONCE(logger->task_gc_stopping, true);
+	cancel_delayed_work_sync(&logger->task_gc_work);
+	cancel_work_sync(&logger->reclaim_work);
+
+	rhashtable_free_and_destroy(&logger->task_map,
+				    blog_task_entry_free_cb, logger);
+
+	/* Detach retained log-batch contexts from the reader-visible list. */
+	spin_lock(&logger->lock);
+	list_for_each_entry_safe(tls_ctx, tmp, &logger->contexts, list)
+		list_move(&tls_ctx->list, &pending);
+	list_for_each_entry_safe(tls_ctx, tmp, &logger->reclaim_list, list)
+		list_move(&tls_ctx->list, &reclaim);
+	spin_unlock(&logger->lock);
+
+	/* Failed atomic puts were never owned by log_batch. */
+	list_for_each_entry_safe(tls_ctx, tmp, &reclaim, list) {
+		list_del_init(&tls_ctx->list);
+		kvfree_atomic(blog_ctx_container(tls_ctx));
+	}
+
+	list_for_each_entry_safe(tls_ctx, tmp, &pending, list) {
+		list_del_init(&tls_ctx->list);
+		if (!READ_ONCE(tls_ctx->task))
+			continue;
+		blog_module_clear_task(tls_ctx);
+		if (tls_ctx->release)
+			tls_ctx->release(tls_ctx);
+		else
+			blog_module_queue_to_log_batch(logger, tls_ctx);
+	}
+
+	/*
+	 * rhashtable callbacks and the drain above can put into log_batch
+	 * and, before task_gc_stopping, would re-queue reclaim_work.  Cancel
+	 * again so a late queue cannot run after we free logger.
+	 */
+	cancel_work_sync(&logger->reclaim_work);
+
+	blog_batch_cleanup(&logger->alloc_batch);
+	blog_batch_cleanup(&logger->log_batch);
+
+	if (logger->magazine_cache)
+		kmem_cache_destroy(logger->magazine_cache);
+	kvfree(logger->source_hash);
+	kvfree(logger->source_map);
+
+	pr_debug("BLOG: module '%s' cleaned up\n", ctx->name);
+
+	kfree(logger);
+	ctx->logger = NULL;
+	ctx->initialized = false;
+	kfree(ctx);
+}
+
+static void blog_module_free_workfn(struct work_struct *work)
+{
+	struct blog_module_context *ctx =
+		container_of(work, struct blog_module_context, free_work);
+
+	blog_module_free(ctx);
+}
+
+static void blog_discard_unmapped_ctx(struct blog_logger *logger,
+				      struct blog_tls_ctx *tls_ctx)
+{
+	spin_lock(&logger->lock);
+	if (!list_empty(&tls_ctx->list))
+		list_del_init(&tls_ctx->list);
+	spin_unlock(&logger->lock);
+
+	blog_module_clear_task(tls_ctx);
+	blog_module_queue_to_log_batch(logger, tls_ctx);
+}
+
+struct blog_tls_ctx *blog_lookup_tls_ctx(struct blog_module_context *ctx)
+{
+	struct blog_logger *logger;
+	struct blog_task_entry *entry;
+	struct blog_tls_ctx *tls_ctx = NULL;
+	struct task_struct *task = current;
+
+	if (!ctx || !ctx->logger)
+		return NULL;
+	logger = ctx->logger;
+
+	rcu_read_lock();
+	entry = rhashtable_lookup_fast(&logger->task_map, &task,
+				       blog_task_ht_params);
+	if (entry && entry->pid == task->pid)
+		tls_ctx = entry->ctx;
+	rcu_read_unlock();
+	return tls_ctx;
+}
+
+struct blog_tls_ctx *blog_get_tls_ctx_ctx(struct blog_module_context *ctx,
+					  gfp_t gfp)
+{
+	struct blog_logger *logger;
+	struct blog_task_entry *entry;
+	struct blog_tls_ctx *tls_ctx;
+	struct task_struct *task = current;
+	bool stale;
+	int err;
+
+	if (!ctx || !ctx->logger)
+		return NULL;
+	logger = ctx->logger;
+
+	if (!gfpflags_allow_blocking(gfp))
+		return blog_lookup_tls_ctx(ctx);
+
+retry:
+	tls_ctx = blog_lookup_tls_ctx(ctx);
+	if (tls_ctx)
+		return tls_ctx;
+
+	rcu_read_lock();
+	entry = rhashtable_lookup_fast(&logger->task_map, &task,
+				       blog_task_ht_params);
+	stale = entry && entry->pid != task->pid;
+	rcu_read_unlock();
+
+	if (stale) {
+		if (!blog_retire_stale_task(logger, task))
+			return NULL;
+		goto retry;
+	}
+
+	entry = kzalloc(sizeof(*entry), gfp);
+	if (!entry)
+		return NULL;
+	if (atomic_inc_return(&ctx->allocated_contexts) >
+	    BLOG_MAX_TASK_CONTEXTS) {
+		atomic_dec(&ctx->allocated_contexts);
+		blog_gc_dead_tasks(logger);
+		if (atomic_inc_return(&ctx->allocated_contexts) >
+		    BLOG_MAX_TASK_CONTEXTS) {
+			atomic_dec(&ctx->allocated_contexts);
+			kfree(entry);
+			return NULL;
+		}
+	}
+
+	tls_ctx = blog_alloc_tls_ctx(logger, gfp);
+	if (!tls_ctx) {
+		atomic_dec(&ctx->allocated_contexts);
+		kfree(entry);
+		return NULL;
+	}
+	__set_bit(BLOG_CTX_TASK_COUNTED, &tls_ctx->flags);
+
+	entry->task = task;
+	entry->pid = task->pid;
+	get_task_comm(entry->comm, task);
+	entry->ctx = tls_ctx;
+	entry->flags = 0;
+	get_task_struct(task);
+
+	err = rhashtable_lookup_insert_fast(&logger->task_map, &entry->node,
+					    blog_task_ht_params);
+	if (err) {
+		put_task_struct(task);
+		kfree(entry);
+		blog_discard_unmapped_ctx(logger, tls_ctx);
+		if (err != -EEXIST)
+			return NULL;
+		goto retry;
+	}
+
+	return tls_ctx;
+}
+
+void blog_module_put(struct blog_module_context *ctx)
+{
+	if (ctx && refcount_dec_and_test(&ctx->refcount)) {
+		if (blog_free_wq)
+			queue_work(blog_free_wq, &ctx->free_work);
+		else
+			blog_module_free(ctx);
+	}
+}
+
+void blog_module_flush_frees(void)
+{
+	if (blog_free_wq)
+		flush_workqueue(blog_free_wq);
+}
+
+int blog_module_wq_init(void)
+{
+	if (blog_free_wq)
+		return 0;
+	blog_free_wq = alloc_workqueue("ceph_blog_free",
+				       WQ_MEM_RECLAIM | WQ_UNBOUND, 0);
+	return blog_free_wq ? 0 : -ENOMEM;
+}
+
+void blog_module_wq_exit(void)
+{
+	if (!blog_free_wq)
+		return;
+	destroy_workqueue(blog_free_wq);
+	blog_free_wq = NULL;
+}
+
+/*
+ * Retire the live buffer's published records into a reader-visible
+ * snapshot, then reset the live pagefrag so logging can continue.
+ * Returns true if the snapshot is on the reader list.  Returns false if
+ * allocation or log-batch put failed: live is left unpublished so the
+ * caller can drop only the new record.
+ *
+ * blog/entries walks contexts by ascending ID.  The snapshot inherits
+ * the live ID (older records) and live takes next_ctx_id (newer), so a
+ * task's retired chunk sorts before its new tail.  Id swap and
+ * contexts-list insertion happen under logger->lock *before* log_batch
+ * put and before live head is cleared, so a concurrent dump cannot
+ * advance past N while the snapshot is still invisible.  A dump may
+ * briefly see the old bytes under the new live id (duplicate); that is
+ * preferred to silently skipping the snapshot.
+ *
+ * Must not sleep: boutc can run under GFP_ATOMIC enters.
+ * Do not take snapshot_mutex here.
+ */
+static bool blog_retire_full_live_buffer(struct blog_logger *logger,
+					 struct blog_tls_ctx *live)
+{
+	struct blog_tls_pagefrag *snap;
+	struct blog_tls_ctx *snap_ctx;
+	struct blog_pagefrag *live_pf = blog_ctx_pf(live);
+	struct blog_pagefrag *snap_pf;
+	unsigned int head;
+	u64 new_live_id;
+
+	snap = blog_batch_get(&logger->alloc_batch, &logger->log_batch);
+	if (!snap)
+		snap = kvzalloc(BLOG_TLS_PAGEFRAG_ALLOC_SIZE, GFP_ATOMIC);
+	if (!snap)
+		return false;
+
+	snap_ctx = &snap->ctx;
+	if (snap_ctx->id == 0)
+		INIT_LIST_HEAD(&snap_ctx->list);
+
+	atomic_set(&snap_ctx->refcount, 0);
+	snap_ctx->task = NULL;
+	snap_ctx->pid = live->pid;
+	memcpy(snap_ctx->comm, live->comm, sizeof(snap_ctx->comm));
+	WRITE_ONCE(snap_ctx->base_jiffies, READ_ONCE(live->base_jiffies));
+	snap_ctx->release = blog_module_tls_release;
+	snap_ctx->logger = logger;
+	snap_ctx->flags = 0;
+	snap_ctx->pending_offset = 0;
+	snap_ctx->pending_size = 0;
+	WRITE_ONCE(snap_ctx->enter_depth, 0);
+	atomic64_set(&snap_ctx->clear_seq, atomic64_read(&live->clear_seq));
+
+	snap_pf = &snap->pf;
+	snap_pf->buffer = snap->buf;
+	snap_pf->capacity = BLOG_TLS_PAGEFRAG_BUFFER_SIZE;
+	spin_lock_init(&snap_pf->lock);
+
+	spin_lock(&live_pf->lock);
+	head = live_pf->head;
+	if (head)
+		memcpy(snap->buf, live_pf->buffer, head);
+	snap_pf->head = head;
+	spin_unlock(&live_pf->lock);
+
+	spin_lock(&logger->ctx_id_lock);
+	new_live_id = logger->next_ctx_id++;
+	spin_unlock(&logger->ctx_id_lock);
+
+	/*
+	 * Publish before put and before clearing live.  Otherwise a
+	 * dump can copy live (id=N), verify id==N, advance the cursor,
+	 * then miss the snapshot that later appears with id=N.
+	 */
+	spin_lock(&logger->lock);
+	snap_ctx->id = live->id;
+	live->id = new_live_id;
+	if (list_empty(&snap_ctx->list)) {
+		list_add(&snap_ctx->list, &logger->contexts);
+		logger->total_contexts_allocated++;
+	}
+	spin_unlock(&logger->lock);
+
+	if (!blog_batch_put(&logger->log_batch, snap)) {
+		spin_lock(&logger->lock);
+		/*
+		 * Leave live->id at new_live_id.  Restoring the old id
+		 * can hide this buffer behind a dump cursor that already
+		 * walked past new_live_id.  The unused snap is reclaimed.
+		 */
+		snap_ctx->id = 0;
+		if (!list_empty(&snap_ctx->list)) {
+			list_del_init(&snap_ctx->list);
+			if (logger->total_contexts_allocated)
+				logger->total_contexts_allocated--;
+		}
+		list_add(&snap_ctx->list, &logger->reclaim_list);
+		spin_unlock(&logger->lock);
+		blog_module_schedule_log_reclaim(logger);
+		return false;
+	}
+
+	spin_lock(&logger->lock);
+	spin_lock(&live_pf->lock);
+	WRITE_ONCE(live->base_jiffies, get_jiffies_64());
+	smp_store_release(&live_pf->head, 0);
+	spin_unlock(&live_pf->lock);
+	spin_unlock(&logger->lock);
+
+	blog_module_schedule_log_reclaim(logger);
+	live->pending_offset = 0;
+	live->pending_size = 0;
+	return true;
+}
+
+/**
+ * blog_log_with_ctx - Reserve buffer for a binary log message (explicit ctx)
+ * @logger: Logger instance
+ * @tls_ctx: TLS context to log into
+ * @source_id: Source ID for this location
+ * @client_id: Client ID for this message
+ * @needed_size: Size needed for the message
+ *
+ * Only one reservation may be outstanding per context at a time.
+ * The caller must call blog_log_commit_with_ctx() before issuing
+ * another reservation on the same context.
+ *
+ * Returns a buffer to write the message into, or NULL on failure
+ */
+void *blog_log_with_ctx(struct blog_logger *logger,
+			struct blog_tls_ctx *tls_ctx,
+			u32 source_id, u8 client_id, size_t needed_size)
+{
+	struct blog_pagefrag *pf;
+	struct blog_log_entry *entry;
+	int alloc;
+	size_t total_size;
+	u64 now;
+
+	if (!logger || !tls_ctx)
+		return NULL;
+
+	if (needed_size > BLOG_MAX_PAYLOAD)
+		return NULL;
+
+	total_size = round_up(sizeof(*entry) + needed_size, 8);
+	pf = blog_ctx_pf(tls_ctx);
+
+	if (test_and_clear_bit(BLOG_CTX_NEEDS_RESET, &tls_ctx->flags) ||
+	    atomic64_read(&tls_ctx->clear_seq) !=
+	    atomic64_read(&logger->clear_seq)) {
+		blog_pagefrag_reset(pf);
+		tls_ctx->pending_offset = 0;
+		tls_ctx->pending_size = 0;
+		atomic64_set(&tls_ctx->clear_seq,
+			     atomic64_read(&logger->clear_seq));
+	}
+
+	/*
+	 * Records store get_jiffies_64() - base_jiffies in a u32.  Long-lived
+	 * lightly-logging tasks can exceed U32_MAX without a natural
+	 * rotate; force one (or reset base) before the delta truncates.
+	 */
+	now = get_jiffies_64();
+	if (unlikely((now - tls_ctx->base_jiffies) > U32_MAX)) {
+		if (pf->head) {
+			if (!blog_retire_full_live_buffer(logger, tls_ctx)) {
+				blog_pagefrag_reset(pf);
+				tls_ctx->pending_offset = 0;
+				tls_ctx->pending_size = 0;
+				spin_lock(&pf->lock);
+				WRITE_ONCE(tls_ctx->base_jiffies, get_jiffies_64());
+				spin_unlock(&pf->lock);
+			}
+		} else {
+			spin_lock(&pf->lock);
+			WRITE_ONCE(tls_ctx->base_jiffies, get_jiffies_64());
+			spin_unlock(&pf->lock);
+		}
+		now = tls_ctx->base_jiffies;
+	}
+
+	alloc = blog_pagefrag_reserve(pf, total_size);
+	if (alloc == -ENOMEM) {
+		/* Message larger than an empty buffer cannot fit after rotate. */
+		if (!pf->head)
+			return NULL;
+		/*
+		 * Retire failed (GFP_ATOMIC OOM or log-batch put): keep
+		 * the full live window and drop only this record.  Do
+		 * not wipe in place.
+		 */
+		if (!blog_retire_full_live_buffer(logger, tls_ctx)) {
+			pr_warn_ratelimited(
+				"blog: rotate-on-full alloc failed, dropping\n");
+			return NULL;
+		}
+		now = tls_ctx->base_jiffies;
+		alloc = blog_pagefrag_reserve(pf, total_size);
+	}
+	if (alloc < 0)
+		return NULL;
+
+	entry = blog_pagefrag_get_ptr(pf, alloc);
+	if (!entry)
+		return NULL;
+
+	if (WARN_ON_ONCE(tls_ctx->pending_size != 0))
+		return NULL;
+	tls_ctx->pending_offset = alloc;
+	tls_ctx->pending_size = total_size;
+
+	entry->ts_delta = now - tls_ctx->base_jiffies;
+	entry->source_id = source_id;
+	entry->len = 0;
+	entry->client_id = client_id;
+	entry->flags = 0;
+
+	return entry->buffer;
+}
+
+int blog_log_commit_with_ctx(struct blog_logger *logger,
+			     struct blog_tls_ctx *tls_ctx,
+			     size_t actual_size)
+{
+	struct blog_pagefrag *pf;
+	struct blog_log_entry *entry;
+	size_t total_size;
+
+	if (!logger || !tls_ctx)
+		return -EINVAL;
+
+	total_size = round_up(sizeof(struct blog_log_entry) + actual_size, 8);
+	if (total_size > tls_ctx->pending_size) {
+		tls_ctx->pending_offset = 0;
+		tls_ctx->pending_size = 0;
+		return -ENOSPC;
+	}
+
+	pf = blog_ctx_pf(tls_ctx);
+
+	entry = blog_pagefrag_get_ptr(pf, tls_ctx->pending_offset);
+	if (!entry) {
+		tls_ctx->pending_offset = 0;
+		tls_ctx->pending_size = 0;
+		return -EFAULT;
+	}
+	entry->len = (u16)actual_size;
+
+	blog_pagefrag_publish(pf, tls_ctx->pending_offset + total_size);
+	tls_ctx->pending_offset = 0;
+	tls_ctx->pending_size = 0;
+
+	return 0;
+}
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 16+ messages in thread

* [PATCH v7 07/14] ceph: add Ceph BLOG scaffolding
  2026-09-24 15:30 [PATCH v7 00/14] ceph: add binary logging (BLOG) for CephFS Alex Markuze
                   ` (5 preceding siblings ...)
  2026-09-24 15:30 ` [PATCH v7 06/14] ceph: add BLOG per-module context management Alex Markuze
@ 2026-09-24 15:30 ` Alex Markuze
  2026-09-24 15:30 ` [PATCH v7 08/14] ceph: add boutc wrappers for BLOG Alex Markuze
                   ` (7 subsequent siblings)
  14 siblings, 0 replies; 16+ messages in thread
From: Alex Markuze @ 2026-09-24 15:30 UTC (permalink / raw)
  To: ceph-devel; +Cc: idryomov, xiubo.li

Wire BLOG into CephFS: ceph_blog.h (journal_info, enter/exit, logging
macros), blog_client.c (per-fsc context lifecycle, client-ID mapping,
and blog_max_sources/blog_max_clients module parameters), super.h
(blog_enabled, blog_ctx, debugfs_blog), libceph.h (blog_client_id),
and Makefile (gate all BLOG objects on CONFIG_DEBUG_FS).

BLOG initialization failures must not prevent CephFS registration.
Warn and leave binary logging unavailable if its global resources
cannot be allocated; attempts to enable it then return -ENODEV.
MDS request tracking remains independent of BLOG availability.

Read the per-CPU cached context once and check that same pointer
before dereferencing it. Migration or retirement on another CPU
can clear the slot even with local preemption disabled; separate
loads for the check and assignment can otherwise yield NULL.

Check the per-mount enabled flag before resolving either a cached
or task-mapped context. Another enabled mount can keep the global
static key active after this mount has disabled capture. A call
that has already passed the check may finish normally.

Signed-off-by: Alex Markuze <amarkuze@redhat.com>
Assisted-by: LLM
---
 fs/ceph/Makefile               |   3 +
 fs/ceph/blog_client.c          | 644 +++++++++++++++++++++++++++++++++
 fs/ceph/super.c                |  61 +++-
 fs/ceph/super.h                |   7 +
 include/linux/ceph/ceph_blog.h | 292 +++++++++++++++
 include/linux/ceph/libceph.h   |   2 +
 6 files changed, 996 insertions(+), 13 deletions(-)
 create mode 100644 fs/ceph/blog_client.c
 create mode 100644 include/linux/ceph/ceph_blog.h

diff --git a/fs/ceph/Makefile b/fs/ceph/Makefile
index ebb29d11ac22..2330229d9783 100644
--- a/fs/ceph/Makefile
+++ b/fs/ceph/Makefile
@@ -10,6 +10,9 @@ ceph-y := super.o inode.o dir.o file.o locks.o addr.o ioctl.o \
 	mds_client.o mdsmap.o strings.o ceph_frag.o \
 	debugfs.o util.o metric.o subvolume_metrics.o
 
+ceph-$(CONFIG_DEBUG_FS) += blog_core.o blog_module.o blog_batch.o \
+	blog_pagefrag.o blog_des.o blog_client.o blog_debugfs.o
+
 ceph-$(CONFIG_CEPH_FSCACHE) += cache.o
 ceph-$(CONFIG_CEPH_FS_POSIX_ACL) += acl.o
 ceph-$(CONFIG_FS_ENCRYPTION) += crypto.o
diff --git a/fs/ceph/blog_client.c b/fs/ceph/blog_client.c
new file mode 100644
index 000000000000..40bc827b9cd3
--- /dev/null
+++ b/fs/ceph/blog_client.c
@@ -0,0 +1,644 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Ceph client ID management for BLOG integration
+ *
+ * Maintains mapping between Ceph's fsid/global_id and BLOG client IDs
+ */
+
+#include <linux/ceph/ceph_debug.h>
+#include <linux/module.h>
+#include <linux/slab.h>
+#include <linux/spinlock.h>
+#include <linux/string.h>
+#include <linux/gfp.h>
+#include <linux/preempt.h>
+#include <linux/sched.h>
+#include <linux/jump_label.h>
+#include <linux/ceph/libceph.h>
+#include <linux/ceph/ceph_blog.h>
+#include "blog.h"
+#include "blog_module.h"
+
+#include "super.h"
+
+DEFINE_STATIC_KEY_FALSE(ceph_blog_key);
+
+static int blog_max_sources = BLOG_DEFAULT_MAX_SOURCES;
+static int blog_max_clients = BLOG_DEFAULT_MAX_CLIENTS;
+
+module_param_named(blog_max_sources, blog_max_sources, int, 0444);
+MODULE_PARM_DESC(blog_max_sources,
+		 "Maximum BLOG source IDs per logger (load-time, default 4096)");
+module_param_named(blog_max_clients, blog_max_clients, int, 0444);
+MODULE_PARM_DESC(blog_max_clients,
+		 "Maximum BLOG client IDs (load-time, default 256)");
+
+int blog_param_max_sources(void)
+{
+	int n = READ_ONCE(blog_max_sources);
+
+	if (n < 2)
+		n = 2;
+	if (n > BLOG_MAX_SOURCE_IDS_CAP)
+		n = BLOG_MAX_SOURCE_IDS_CAP;
+	return n;
+}
+
+int blog_param_max_clients(void)
+{
+	int n = READ_ONCE(blog_max_clients);
+
+	if (n < 2)
+		n = 2;
+	if (n > BLOG_MAX_CLIENT_IDS_CAP)
+		n = BLOG_MAX_CLIENT_IDS_CAP;
+	return n;
+}
+
+/* Global client mapping state */
+static struct {
+	struct ceph_blog_client_info *client_map;
+	/* Parallel to client_map: which ceph_client owns each slot. */
+	struct ceph_client **owners;
+	u32 max_clients;
+	u32 next_client_id;
+	spinlock_t lock;  /* protects client_map */
+	bool initialized;
+} ceph_blog_state = {
+	.next_client_id = 1,  /* Start from 1, 0 is reserved */
+	.lock = __SPIN_LOCK_UNLOCKED(ceph_blog_state.lock),
+	.initialized = false,
+};
+
+static bool ceph_blog_ids_match(const struct ceph_blog_client_info *entry,
+				     const char *fsid, u64 global_id)
+{
+	if (!entry)
+		return false;
+	if (entry->global_id != global_id)
+		return false;
+	return !memcmp(entry->fsid, fsid, sizeof(entry->fsid));
+}
+
+static bool ceph_blog_client_slot_free(const struct ceph_blog_client_info *entry)
+{
+	return !data_race(entry->global_id) &&
+	       !data_race(memchr_inv(entry->fsid, 0, sizeof(entry->fsid)));
+}
+
+int ceph_blog_init(void)
+{
+	u32 max_clients;
+	int ret;
+
+	if (ceph_blog_state.initialized)
+		return 0;
+
+	ret = blog_module_wq_init();
+	if (ret)
+		return ret;
+
+	max_clients = blog_param_max_clients();
+	ceph_blog_state.client_map = kcalloc(max_clients,
+					     sizeof(*ceph_blog_state.client_map),
+					     GFP_KERNEL);
+	if (!ceph_blog_state.client_map) {
+		blog_module_wq_exit();
+		return -ENOMEM;
+	}
+	ceph_blog_state.owners = kcalloc(max_clients,
+					 sizeof(*ceph_blog_state.owners),
+					 GFP_KERNEL);
+	if (!ceph_blog_state.owners) {
+		kfree(ceph_blog_state.client_map);
+		ceph_blog_state.client_map = NULL;
+		blog_module_wq_exit();
+		return -ENOMEM;
+	}
+
+	ceph_blog_state.max_clients = max_clients;
+	ceph_blog_state.next_client_id = 1;
+	ceph_blog_state.initialized = true;
+
+	pr_debug("ceph: BLOG client mapping initialized (max_clients=%u)\n",
+		 max_clients);
+	return 0;
+}
+
+void ceph_blog_cleanup(void)
+{
+	void *client_map = NULL;
+	void *owners = NULL;
+
+	blog_module_flush_frees();
+
+	if (ceph_blog_state.initialized) {
+		spin_lock(&ceph_blog_state.lock);
+		client_map = ceph_blog_state.client_map;
+		ceph_blog_state.client_map = NULL;
+		owners = ceph_blog_state.owners;
+		ceph_blog_state.owners = NULL;
+		ceph_blog_state.max_clients = 0;
+		ceph_blog_state.next_client_id = 1;
+		ceph_blog_state.initialized = false;
+		spin_unlock(&ceph_blog_state.lock);
+		kfree(client_map);
+		kfree(owners);
+		pr_debug("ceph: BLOG client mapping cleaned up\n");
+	}
+
+	blog_module_wq_exit();
+}
+
+int ceph_blog_fsc_init(struct ceph_fs_client *fsc)
+{
+	if (!fsc)
+		return -EINVAL;
+
+	mutex_init(&fsc->blog_mutex);
+	RCU_INIT_POINTER(fsc->blog_ctx, NULL);
+	WRITE_ONCE(fsc->blog_enabled, false);
+	return 0;
+}
+
+int ceph_blog_set_enabled(struct ceph_fs_client *fsc, bool enabled)
+{
+	struct blog_module_context *ctx;
+	bool was_enabled;
+	int ret = 0;
+
+	if (!fsc)
+		return -EINVAL;
+	if (enabled && !READ_ONCE(ceph_blog_state.initialized))
+		return -ENODEV;
+
+	mutex_lock(&fsc->blog_mutex);
+	was_enabled = READ_ONCE(fsc->blog_enabled);
+	ctx = rcu_dereference_protected(fsc->blog_ctx,
+					lockdep_is_held(&fsc->blog_mutex));
+	if (enabled && !ctx) {
+		ctx = blog_module_init("ceph");
+		if (!ctx) {
+			pr_err("ceph: failed to initialize BLOG context for fs client\n");
+			ret = -ENOMEM;
+			goto out;
+		}
+		rcu_assign_pointer(fsc->blog_ctx, ctx);
+	}
+	WRITE_ONCE(fsc->blog_enabled, enabled);
+	if (enabled && !was_enabled)
+		static_branch_inc(&ceph_blog_key);
+	else if (!enabled && was_enabled)
+		static_branch_dec(&ceph_blog_key);
+out:
+	mutex_unlock(&fsc->blog_mutex);
+	return ret;
+}
+
+void ceph_blog_fsc_cleanup(struct ceph_fs_client *fsc)
+{
+	struct blog_module_context *ctx;
+	bool was_enabled;
+
+	if (!fsc)
+		return;
+
+	mutex_lock(&fsc->blog_mutex);
+	was_enabled = READ_ONCE(fsc->blog_enabled);
+	WRITE_ONCE(fsc->blog_enabled, false);
+	if (was_enabled)
+		static_branch_dec(&ceph_blog_key);
+	ctx = rcu_replace_pointer(fsc->blog_ctx, NULL,
+				  lockdep_is_held(&fsc->blog_mutex));
+	mutex_unlock(&fsc->blog_mutex);
+
+	if (ctx) {
+		synchronize_rcu();
+		blog_module_put(ctx);
+		blog_module_flush_frees();
+	}
+}
+
+bool ceph_blog_is_enabled(struct ceph_fs_client *fsc)
+{
+	struct blog_module_context *ctx;
+	bool enabled = false;
+
+	if (!fsc || !READ_ONCE(fsc->blog_enabled))
+		return false;
+
+	rcu_read_lock();
+	ctx = rcu_dereference(fsc->blog_ctx);
+	if (ctx && READ_ONCE(ctx->logger))
+		enabled = true;
+	rcu_read_unlock();
+
+	return enabled;
+}
+
+struct blog_tls_ctx *ceph_blog_acquire_ctx(struct ceph_fs_client *fsc,
+					   gfp_t gfp,
+					   struct blog_module_context **held_mod)
+{
+	struct blog_module_context *ctx;
+	struct blog_tls_ctx *tls_ctx = NULL;
+
+	if (held_mod)
+		*held_mod = NULL;
+
+	if (!fsc || !READ_ONCE(fsc->blog_enabled))
+		return NULL;
+
+	rcu_read_lock();
+	ctx = rcu_dereference(fsc->blog_ctx);
+	if (!ctx || !READ_ONCE(ctx->logger)) {
+		rcu_read_unlock();
+		return NULL;
+	}
+
+	if (!gfpflags_allow_blocking(gfp)) {
+		/*
+		 * Non-blocking: reuse an existing per-task ctx only.  Do not
+		 * pin the module until lookup hits. A miss must not call
+		 * blog_module_put(), which can sleep in blog_module_free().
+		 */
+		tls_ctx = blog_lookup_tls_ctx(ctx);
+		if (!tls_ctx || !refcount_inc_not_zero(&ctx->refcount)) {
+			rcu_read_unlock();
+			return NULL;
+		}
+		rcu_read_unlock();
+		goto hold;
+	}
+
+	if (!refcount_inc_not_zero(&ctx->refcount)) {
+		rcu_read_unlock();
+		return NULL;
+	}
+	rcu_read_unlock();
+
+	tls_ctx = blog_get_tls_ctx_ctx(ctx, gfp);
+	if (!tls_ctx) {
+		blog_module_put(ctx);
+		return NULL;
+	}
+
+hold:
+	/* Keep the module ref for the enter-exit window; put via blog_mod. */
+	if (held_mod)
+		*held_mod = ctx;
+	else
+		blog_module_put(ctx);
+
+	return tls_ctx;
+}
+
+void ceph_blog_module_put(struct blog_module_context *ctx)
+{
+	blog_module_put(ctx);
+}
+
+struct ceph_blog_cpu_cache {
+	struct task_struct *task;
+	struct blog_tls_ctx *ctx;
+};
+
+static DEFINE_PER_CPU(struct ceph_blog_cpu_cache, ceph_blog_cpu_cache);
+
+static void blog_cpu_cache_clear_slot(struct blog_tls_ctx *ctx, int cpu)
+{
+	struct ceph_blog_cpu_cache *c;
+
+	if (cpu < 0)
+		return;
+	c = per_cpu_ptr(&ceph_blog_cpu_cache, cpu);
+	if (READ_ONCE(c->ctx) == ctx) {
+		WRITE_ONCE(c->task, NULL);
+		WRITE_ONCE(c->ctx, NULL);
+	}
+}
+
+/*
+ * Drop published per-CPU slots for @ctx before GC/retire can free it.
+ * With the single-slot invariant, clearing cache_cpu (and this CPU) is enough.
+ */
+void ceph_blog_cpu_clear(struct blog_tls_ctx *ctx)
+{
+	int cpu;
+
+	if (!ctx)
+		return;
+
+	preempt_disable();
+	cpu = READ_ONCE(ctx->cache_cpu);
+	blog_cpu_cache_clear_slot(ctx, cpu);
+	if (cpu != smp_processor_id())
+		blog_cpu_cache_clear_slot(ctx, smp_processor_id());
+	WRITE_ONCE(ctx->cache_cpu, -1);
+	preempt_enable();
+}
+
+void ceph_blog_cpu_bind(struct blog_tls_ctx *ctx)
+{
+	struct ceph_blog_cpu_cache *c;
+	int cpu, prev_cpu;
+
+	if (!ctx)
+		return;
+
+	/*
+	 * Publish into the per-CPU cache under preempt_disable only.
+	 * Do not migrate_disable() across the enter-exit window: on !RT
+	 * that is preempt_disable and would leave preempt elevated for
+	 * the whole VFS call.  Keep each ctx on at most one CPU slot:
+	 * clear the prior publish before installing the new one.
+	 */
+	preempt_disable();
+	WRITE_ONCE(ctx->enter_depth, READ_ONCE(ctx->enter_depth) + 1);
+	cpu = smp_processor_id();
+	prev_cpu = READ_ONCE(ctx->cache_cpu);
+	if (prev_cpu >= 0 && prev_cpu != cpu)
+		blog_cpu_cache_clear_slot(ctx, prev_cpu);
+	c = this_cpu_ptr(&ceph_blog_cpu_cache);
+	WRITE_ONCE(c->task, current);
+	WRITE_ONCE(c->ctx, ctx);
+	WRITE_ONCE(ctx->cache_cpu, cpu);
+	preempt_enable();
+}
+
+void ceph_blog_cpu_unbind(struct blog_tls_ctx *ctx)
+{
+	int cpu;
+
+	if (!ctx || !READ_ONCE(ctx->enter_depth))
+		return;
+
+	preempt_disable();
+	WRITE_ONCE(ctx->enter_depth, READ_ONCE(ctx->enter_depth) - 1);
+	if (!READ_ONCE(ctx->enter_depth)) {
+		cpu = READ_ONCE(ctx->cache_cpu);
+		blog_cpu_cache_clear_slot(ctx, cpu);
+		if (cpu != smp_processor_id())
+			blog_cpu_cache_clear_slot(ctx, smp_processor_id());
+		WRITE_ONCE(ctx->cache_cpu, -1);
+	}
+	preempt_enable();
+}
+
+struct blog_tls_ctx *ceph_blog_get_cached_ctx(struct ceph_fs_client *fsc)
+{
+	struct ceph_blog_cpu_cache *c;
+	struct blog_tls_ctx *ctx = NULL;
+	struct blog_module_context *mod;
+	struct ceph_journal_info *ji;
+	struct blog_logger *want_logger;
+	int cpu, prev_cpu;
+
+	if (!fsc || !READ_ONCE(fsc->blog_enabled))
+		return NULL;
+
+	rcu_read_lock();
+	mod = rcu_dereference(fsc->blog_ctx);
+	want_logger = (mod && mod->logger) ? mod->logger : NULL;
+	if (!want_logger) {
+		rcu_read_unlock();
+		return NULL;
+	}
+
+	preempt_disable();
+	c = this_cpu_ptr(&ceph_blog_cpu_cache);
+	if (likely(READ_ONCE(c->task) == current)) {
+		/* Remote migration/retirement can clear this slot. Read it once. */
+		ctx = READ_ONCE(c->ctx);
+		if (ctx && READ_ONCE(ctx->task) == current &&
+		    READ_ONCE(ctx->enter_depth) &&
+		    ctx->logger == want_logger)
+			; /* hit. Mount-scoped */
+		else {
+			WRITE_ONCE(c->task, NULL);
+			WRITE_ONCE(c->ctx, NULL);
+			ctx = NULL;
+		}
+	}
+	preempt_enable();
+	if (ctx) {
+		rcu_read_unlock();
+		return ctx;
+	}
+
+	/*
+	 * Cache miss after preemption.  Prefer the mount-scoped
+	 * journal_info ctx when it matches @fsc so nested multi-mount
+	 * work does not attach to another mount's logger.  Plain
+	 * ceph_blog_enter() never installs journal_info. Recover only
+	 * from @fsc's own task map (not a cross-mount module-list walk).
+	 */
+	ji = ceph_ji_from_current();
+	if (ji && ceph_ji_matches_fsc(ji, fsc)) {
+		if (ji->blog_ctx &&
+		    READ_ONCE(ji->blog_ctx->enter_depth) &&
+		    READ_ONCE(ji->blog_ctx->task) == current &&
+		    ji->blog_ctx->logger == want_logger)
+			ctx = ji->blog_ctx;
+	} else {
+		ctx = blog_lookup_tls_ctx(mod);
+		if (ctx && !(READ_ONCE(ctx->enter_depth) &&
+			     READ_ONCE(ctx->task) == current))
+			ctx = NULL;
+	}
+	rcu_read_unlock();
+
+	if (ctx) {
+		preempt_disable();
+		cpu = smp_processor_id();
+		prev_cpu = READ_ONCE(ctx->cache_cpu);
+		if (prev_cpu >= 0 && prev_cpu != cpu)
+			blog_cpu_cache_clear_slot(ctx, prev_cpu);
+		c = this_cpu_ptr(&ceph_blog_cpu_cache);
+		WRITE_ONCE(c->task, current);
+		WRITE_ONCE(c->ctx, ctx);
+		WRITE_ONCE(ctx->cache_cpu, cpu);
+		preempt_enable();
+		return ctx;
+	}
+
+	return NULL;
+}
+
+/**
+ * ceph_blog_check_client_id - Check if a client ID matches the given fsid:global_id pair
+ * @id: Client ID to check
+ * @fsid: Client FSID to compare
+ * @global_id: Client global ID to compare
+ *
+ * Returns the actual ID of the pair. If the given ID doesn't match, scans for
+ * existing matches or allocates a new ID if no match is found.
+ */
+u32 ceph_blog_check_client_id(u32 id, const char *fsid, u64 global_id)
+{
+	u32 found_id = 0;
+	u32 max_clients;
+	struct ceph_blog_client_info *entry;
+
+	if (unlikely(!ceph_blog_state.initialized)) {
+		WARN_ON_ONCE(1);
+		return 0;
+	}
+
+	spin_lock(&ceph_blog_state.lock);
+	max_clients = ceph_blog_state.max_clients;
+
+	if (id != 0 && id < max_clients) {
+		entry = &ceph_blog_state.client_map[id];
+		if (ceph_blog_ids_match(entry, fsid, global_id)) {
+			found_id = id;
+			goto out;
+		}
+	}
+
+	for (id = 1; id < max_clients; id++) {
+		entry = &ceph_blog_state.client_map[id];
+		if (ceph_blog_ids_match(entry, fsid, global_id)) {
+			found_id = id;
+			goto out;
+		}
+	}
+
+	if (ceph_blog_state.next_client_id < max_clients) {
+		found_id = ceph_blog_state.next_client_id++;
+	} else {
+		found_id = 0;
+		for (id = 1; id < max_clients; id++) {
+			entry = &ceph_blog_state.client_map[id];
+			if (ceph_blog_client_slot_free(entry)) {
+				found_id = id;
+				break;
+			}
+		}
+		if (!found_id) {
+			pr_warn_once("ceph: BLOG client ID space exhausted\n");
+			goto out;
+		}
+	}
+
+	entry = &ceph_blog_state.client_map[found_id];
+	memset(entry, 0, sizeof(*entry));
+	memcpy(entry->fsid, fsid, sizeof(entry->fsid));
+	entry->global_id = global_id;
+
+out:
+	spin_unlock(&ceph_blog_state.lock);
+	return found_id;
+}
+
+/**
+ * ceph_blog_get_client_info - Get client info for a given ID
+ * @id: Client ID
+ *
+ * Reads client_map[] without holding ceph_blog_state.lock.
+ * Writers store fields under the lock. Callers accept the benign
+ * race: a concurrent slot release and reuse may cause old log
+ * entries to show a new client's identity; the impact is cosmetic.
+ */
+const struct ceph_blog_client_info *ceph_blog_get_client_info(u32 id)
+{
+	const struct ceph_blog_client_info *entry;
+
+	if (!READ_ONCE(ceph_blog_state.initialized) ||
+	    id == 0 || id >= READ_ONCE(ceph_blog_state.max_clients))
+		return NULL;
+	entry = &ceph_blog_state.client_map[id];
+	/* Freed/zeroed slots must not deserialize as a valid client. */
+	if (ceph_blog_client_slot_free(entry))
+		return NULL;
+	return entry;
+}
+
+int ceph_blog_client_des_callback(char *buf, size_t size, u8 client_id)
+{
+	const struct ceph_blog_client_info *info;
+	char fsid[16];
+	u64 global_id;
+
+	if (!buf || !size)
+		return -EINVAL;
+	if (client_id == 0)
+		return 0;
+
+	info = ceph_blog_get_client_info(client_id);
+	if (!info)
+		return snprintf(buf, size, "[unknown_client_%u]", client_id);
+
+	global_id = data_race(info->global_id);
+	data_race(memcpy(fsid, info->fsid, sizeof(fsid)));
+	return snprintf(buf, size, "[%pU %llu] ", fsid, global_id);
+}
+
+u32 ceph_blog_get_client_id(struct ceph_client *client)
+{
+	u32 cached;
+	u32 id;
+
+	if (!client)
+		return 0;
+	if (!client->monc.auth)
+		return 0;
+
+	cached = READ_ONCE(client->blog_client_id);
+
+	id = ceph_blog_check_client_id(cached,
+				       client->fsid.fsid,
+				       client->monc.auth->global_id);
+	if (!id)
+		return 0;
+
+	/*
+	 * Record ownership of the (possibly new) slot.  On auth rekey do
+	 * not clear the prior map entry: buffered records still carry the
+	 * old client_id and resolve it lazily on readback.  All of this
+	 * client's slots are freed in ceph_blog_release_client_id() before
+	 * fsc cleanup tears down the buffers.
+	 */
+	spin_lock(&ceph_blog_state.lock);
+	if (ceph_blog_state.initialized &&
+	    id < ceph_blog_state.max_clients)
+		ceph_blog_state.owners[id] = client;
+	spin_unlock(&ceph_blog_state.lock);
+
+	if (cached != id)
+		WRITE_ONCE(client->blog_client_id, id);
+
+	return id;
+}
+
+/**
+ * ceph_blog_release_client_id - Free a client's BLOG ID mapping slots
+ * @client: Ceph client being torn down
+ *
+ * Clears the cached ID on @client and zeroes every global map entry this
+ * client owns (including slots left behind by auth rekey) so the 8-bit
+ * namespace can be reused after remounts / new auth sessions.
+ */
+void ceph_blog_release_client_id(struct ceph_client *client)
+{
+	u32 id;
+
+	if (!client)
+		return;
+
+	WRITE_ONCE(client->blog_client_id, 0);
+
+	spin_lock(&ceph_blog_state.lock);
+	if (!ceph_blog_state.initialized) {
+		spin_unlock(&ceph_blog_state.lock);
+		return;
+	}
+	for (id = 1; id < ceph_blog_state.max_clients; id++) {
+		if (ceph_blog_state.owners[id] != client)
+			continue;
+		ceph_blog_state.owners[id] = NULL;
+		memset(&ceph_blog_state.client_map[id], 0,
+		       sizeof(*ceph_blog_state.client_map));
+	}
+	spin_unlock(&ceph_blog_state.lock);
+}
diff --git a/fs/ceph/super.c b/fs/ceph/super.c
index d3c97a1d0ad2..40be4b16e486 100644
--- a/fs/ceph/super.c
+++ b/fs/ceph/super.c
@@ -49,11 +49,15 @@ static LIST_HEAD(ceph_fsc_list);
 static void ceph_put_super(struct super_block *s)
 {
 	struct ceph_fs_client *fsc = ceph_sb_to_fs_client(s);
+	struct ceph_journal_info __ji;
 
-	doutc(fsc->client, "begin\n");
+	ceph_blog_enter(fsc, &__ji);
+
+	boutc(fsc->client, "begin\n");
 	ceph_fscrypt_free_dummy_policy(fsc);
 	ceph_mdsc_close_sessions(fsc->mdsc);
-	doutc(fsc->client, "done\n");
+	boutc(fsc->client, "done\n");
+	ceph_blog_exit(&__ji);
 }
 
 static int ceph_statfs(struct dentry *dentry, struct kstatfs *buf)
@@ -64,8 +68,11 @@ static int ceph_statfs(struct dentry *dentry, struct kstatfs *buf)
 	struct ceph_statfs st;
 	int i, err;
 	u64 data_pool;
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(fsc, &__ji);
 
-	doutc(fsc->client, "begin\n");
+	boutc(fsc->client, "begin\n");
 	if (fsc->mdsc->mdsmap->m_num_data_pg_pools == 1) {
 		data_pool = fsc->mdsc->mdsmap->m_data_pg_pools[0];
 	} else {
@@ -73,8 +80,10 @@ static int ceph_statfs(struct dentry *dentry, struct kstatfs *buf)
 	}
 
 	err = ceph_monc_do_statfs(monc, data_pool, &st);
-	if (err < 0)
+	if (err < 0) {
+		ceph_blog_exit(&__ji);
 		return err;
+	}
 
 	/* fill in kstatfs */
 	buf->f_type = CEPH_SUPER_MAGIC;  /* ?? */
@@ -121,7 +130,8 @@ static int ceph_statfs(struct dentry *dentry, struct kstatfs *buf)
 	/* fold the fs_cluster_id into the upper bits */
 	buf->f_fsid.val[1] = monc->fs_cluster_id;
 
-	doutc(fsc->client, "done\n");
+	boutc(fsc->client, "done\n");
+	ceph_blog_exit(&__ji);
 	return 0;
 }
 
@@ -129,19 +139,24 @@ static int ceph_sync_fs(struct super_block *sb, int wait)
 {
 	struct ceph_fs_client *fsc = ceph_sb_to_fs_client(sb);
 	struct ceph_client *cl = fsc->client;
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(fsc, &__ji);
 
 	if (!wait) {
-		doutc(cl, "(non-blocking)\n");
+		boutc(cl, "(non-blocking)\n");
 		ceph_flush_dirty_caps(fsc->mdsc);
 		ceph_flush_cap_releases(fsc->mdsc);
-		doutc(cl, "(non-blocking) done\n");
+		boutc(cl, "(non-blocking) done\n");
+		ceph_blog_exit(&__ji);
 		return 0;
 	}
 
-	doutc(cl, "(blocking)\n");
+	boutc(cl, "(blocking)\n");
 	ceph_osdc_sync(&fsc->client->osdc);
 	ceph_mdsc_sync(fsc->mdsc);
-	doutc(cl, "(blocking) done\n");
+	boutc(cl, "(blocking) done\n");
+	ceph_blog_exit(&__ji);
 	return 0;
 }
 
@@ -887,12 +902,18 @@ static struct ceph_fs_client *create_fs_client(struct ceph_mount_options *fsopt,
 	hash_init(fsc->async_unlink_conflict);
 	spin_lock_init(&fsc->async_unlink_conflict_lock);
 
+	err = ceph_blog_fsc_init(fsc);
+	if (err)
+		goto fail_cap_wq;
+
 	spin_lock(&ceph_fsc_lock);
 	list_add_tail(&fsc->metric_wakeup, &ceph_fsc_list);
 	spin_unlock(&ceph_fsc_lock);
 
 	return fsc;
 
+fail_cap_wq:
+	destroy_workqueue(fsc->cap_wq);
 fail_inode_wq:
 	destroy_workqueue(fsc->inode_wq);
 fail_client:
@@ -922,6 +943,8 @@ static void destroy_fs_client(struct ceph_fs_client *fsc)
 	ceph_mdsc_destroy(fsc);
 	destroy_workqueue(fsc->inode_wq);
 	destroy_workqueue(fsc->cap_wq);
+	ceph_blog_release_client_id(fsc->client);
+	ceph_blog_fsc_cleanup(fsc);
 
 	destroy_mount_options(fsc->mount_options);
 
@@ -1056,11 +1079,15 @@ static void __ceph_umount_begin(struct ceph_fs_client *fsc)
 void ceph_umount_begin(struct super_block *sb)
 {
 	struct ceph_fs_client *fsc = ceph_sb_to_fs_client(sb);
+	struct ceph_journal_info __ji;
 
-	doutc(fsc->client, "starting forced umount\n");
+	ceph_blog_enter(fsc, &__ji);
+
+	boutc(fsc->client, "starting forced umount\n");
 
 	fsc->mount_state = CEPH_MOUNT_SHUTDOWN;
 	__ceph_umount_begin(fsc);
+	ceph_blog_exit(&__ji);
 }
 
 static const struct super_operations ceph_super_ops = {
@@ -1683,9 +1710,15 @@ int ceph_force_reconnect(struct super_block *sb)
 
 static int __init init_ceph(void)
 {
-	int ret = init_caches();
+	int ret;
+
+	ret = ceph_blog_init();
 	if (ret)
-		goto out;
+		pr_warn("BLOG initialization failed (%d); binary logging unavailable\n", ret);
+
+	ret = init_caches();
+	if (ret)
+		goto out_blog;
 
 	ceph_flock_init();
 	ret = register_filesystem(&ceph_fs_type);
@@ -1698,7 +1731,8 @@ static int __init init_ceph(void)
 
 out_caches:
 	destroy_caches();
-out:
+out_blog:
+	ceph_blog_cleanup();
 	return ret;
 }
 
@@ -1707,6 +1741,7 @@ static void __exit exit_ceph(void)
 	dout("exit_ceph\n");
 	unregister_filesystem(&ceph_fs_type);
 	destroy_caches();
+	ceph_blog_cleanup();
 }
 
 static int param_set_metrics(const char *val, const struct kernel_param *kp)
diff --git a/fs/ceph/super.h b/fs/ceph/super.h
index e00df8000427..458fd63dd79f 100644
--- a/fs/ceph/super.h
+++ b/fs/ceph/super.h
@@ -3,8 +3,11 @@
 #define _FS_CEPH_SUPER_H
 
 #include <linux/ceph/ceph_debug.h>
+#include <linux/ceph/ceph_blog.h>
 #include <linux/ceph/osd_client.h>
 
+#include "blog.h"
+
 #include <linux/unaligned.h>
 #include <linux/backing-dev.h>
 #include <linux/completion.h>
@@ -174,6 +177,9 @@ struct ceph_fs_client {
 	spinlock_t async_unlink_conflict_lock;
 
 #ifdef CONFIG_DEBUG_FS
+	bool blog_enabled;
+	struct mutex blog_mutex;
+	struct blog_module_context __rcu *blog_ctx;
 	struct dentry *debugfs_dentry_lru, *debugfs_caps;
 	struct dentry *debugfs_congestion_kb;
 	struct dentry *debugfs_bdi;
@@ -182,6 +188,7 @@ struct ceph_fs_client {
 	struct dentry *debugfs_mds_sessions;
 	struct dentry *debugfs_metrics_dir;
 	struct dentry *debugfs_reset_dir;
+	struct dentry *debugfs_blog;
 	struct dentry *debugfs_subvolume_metrics;
 #endif
 
diff --git a/include/linux/ceph/ceph_blog.h b/include/linux/ceph/ceph_blog.h
new file mode 100644
index 000000000000..82c64bc98972
--- /dev/null
+++ b/include/linux/ceph/ceph_blog.h
@@ -0,0 +1,292 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Ceph integration with BLOG (Binary LOGging)
+ *
+ * Provides a shared per-call context (struct ceph_journal_info) used for
+ * the narrow MDS fill-trace window (mds_req in current->journal_info) and
+ * optional BLOG TLS binding via Ceph-private per-task storage / CPU cache.
+ */
+#ifndef CEPH_BLOG_H
+#define CEPH_BLOG_H
+
+#include <linux/sched/mm.h>
+#include <linux/gfp.h>
+#include <linux/jump_label.h>
+
+/* ---------- shared journal_info carrier ---------- */
+
+#define CEPH_JI_MAGIC	0xCE9B7081UL
+#define CEPH_JI_TAG	3UL
+#define CEPH_JI_TAG_MASK	3UL
+
+struct ceph_fs_client;
+struct ceph_mds_request;
+struct ceph_client;
+struct blog_module_context;
+struct blog_tls_ctx;
+
+/**
+ * struct ceph_journal_info - per-call context stashed in journal_info
+ * @magic: CEPH_JI_MAGIC, for safe type-checking when reading journal_info
+ * @saved_ji: previous value of current->journal_info (restored on exit)
+ * @fsc: filesystem client for this enter
+ * @blog_ctx: BLOG TLS context for binary logging, or NULL
+ * @blog_mod: module context ref held for @blog_ctx when this enter acquired
+ *            it (NULL when ctx was inherited from a nested parent)
+ * @mds_req: MDS request during ceph_fill_trace / readdir_prepopulate,
+ *           or NULL.  Read by xattr.c to avoid deadlocking RPCs and
+ *           to discover which capabilities were already fetched.
+ *
+ * Allocated on the caller's stack at every Ceph VFS entry point.
+ * ceph_blog_enter_req() installs it in current->journal_info for MDS
+ * fill-trace; plain ceph_blog_enter() binds BLOG without touching
+ * journal_info.  ceph_blog_exit() restores journal_info when installed.
+ */
+struct ceph_journal_info {
+	unsigned long		magic;
+	void			*saved_ji;
+	struct ceph_fs_client	*fsc;
+	struct blog_tls_ctx	*blog_ctx;
+	struct blog_module_context *blog_mod;
+	struct ceph_mds_request	*mds_req;
+};
+
+/**
+ * ceph_ji_from_current - safely retrieve ceph_journal_info from journal_info
+ *
+ * Ceph tags its journal_info pointer so foreign filesystem state can be
+ * rejected without dereferencing it.  Returns the decoded pointer if the
+ * tag and magic match, NULL otherwise.
+ */
+static inline struct ceph_journal_info *ceph_ji_from_current(void)
+{
+	void *journal_info = current->journal_info;
+	struct ceph_journal_info *ji;
+
+	if (((unsigned long)journal_info & CEPH_JI_TAG_MASK) != CEPH_JI_TAG)
+		return NULL;
+	ji = (void *)((unsigned long)journal_info & ~CEPH_JI_TAG_MASK);
+	if (!ji)
+		return NULL;
+	if (READ_ONCE(ji->magic) == CEPH_JI_MAGIC)
+		return ji;
+	return NULL;
+}
+
+static inline void *ceph_ji_encode(struct ceph_journal_info *ji)
+{
+	WARN_ON_ONCE((unsigned long)ji & CEPH_JI_TAG_MASK);
+	return (void *)((unsigned long)ji | CEPH_JI_TAG);
+}
+
+static inline bool ceph_ji_matches_fsc(const struct ceph_journal_info *ji,
+				       struct ceph_fs_client *fsc)
+{
+	return ji && ji->fsc == fsc;
+}
+
+/**
+ * ceph_current_mds_request - get this mount's in-flight MDS request
+ *
+ * Same-fsc helper for request introspection (e.g. getattr mask).
+ * Returns NULL outside fill-trace or when the tagged journal_info
+ * belongs to a different ceph_fs_client.
+ */
+static inline struct ceph_mds_request *
+ceph_current_mds_request(struct ceph_fs_client *fsc)
+{
+	struct ceph_journal_info *ji = ceph_ji_from_current();
+
+	return ceph_ji_matches_fsc(ji, fsc) ? ji->mds_req : NULL;
+}
+
+/**
+ * ceph_current_fill_trace_request - any in-flight Ceph fill-trace on this task
+ *
+ * Sync xattr recursion must stay task-scoped: a security hook on a
+ * second Ceph mount during handle_reply() -> ceph_fill_trace() still
+ * has to return -EBUSY.  Do not use the same-fsc helper for that guard.
+ */
+static inline struct ceph_mds_request *
+ceph_current_fill_trace_request(void)
+{
+	struct ceph_journal_info *ji = ceph_ji_from_current();
+
+	return ji ? ji->mds_req : NULL;
+}
+
+/* ---------- client ID mapping ---------- */
+
+struct ceph_blog_client_info {
+	char fsid[16];
+	u64 global_id;
+};
+
+#ifdef CONFIG_DEBUG_FS
+extern struct static_key_false ceph_blog_key;
+
+int  ceph_blog_init(void);
+void ceph_blog_cleanup(void);
+int  ceph_blog_fsc_init(struct ceph_fs_client *fsc);
+void ceph_blog_fsc_cleanup(struct ceph_fs_client *fsc);
+int  ceph_blog_set_enabled(struct ceph_fs_client *fsc, bool enabled);
+u32  ceph_blog_check_client_id(u32 id, const char *fsid, u64 global_id);
+u32  ceph_blog_get_client_id(struct ceph_client *client);
+void ceph_blog_release_client_id(struct ceph_client *client);
+const struct ceph_blog_client_info *ceph_blog_get_client_info(u32 id);
+int  ceph_blog_client_des_callback(char *buf, size_t size, u8 client_id);
+bool ceph_blog_is_enabled(struct ceph_fs_client *fsc);
+struct blog_tls_ctx *ceph_blog_acquire_ctx(struct ceph_fs_client *fsc,
+					   gfp_t gfp,
+					   struct blog_module_context **held_mod);
+void ceph_blog_module_put(struct blog_module_context *ctx);
+void ceph_blog_cpu_bind(struct blog_tls_ctx *ctx);
+void ceph_blog_cpu_unbind(struct blog_tls_ctx *ctx);
+void ceph_blog_cpu_clear(struct blog_tls_ctx *ctx);
+struct blog_tls_ctx *ceph_blog_get_cached_ctx(struct ceph_fs_client *fsc);
+#else
+/* CONFIG_DEBUG_FS=n: BLOG objects are not linked; stubs below. */
+
+static inline int ceph_blog_init(void) { return 0; }
+static inline void ceph_blog_cleanup(void) {}
+static inline int ceph_blog_fsc_init(struct ceph_fs_client *fsc) { return 0; }
+static inline void ceph_blog_fsc_cleanup(struct ceph_fs_client *fsc) {}
+static inline int ceph_blog_set_enabled(struct ceph_fs_client *fsc, bool enabled)
+{
+	return 0;
+}
+static inline u32 ceph_blog_check_client_id(u32 id, const char *fsid,
+					    u64 global_id)
+{
+	return 0;
+}
+static inline u32 ceph_blog_get_client_id(struct ceph_client *client)
+{
+	return 0;
+}
+static inline void ceph_blog_release_client_id(struct ceph_client *client) {}
+static inline const struct ceph_blog_client_info *
+ceph_blog_get_client_info(u32 id)
+{
+	return NULL;
+}
+static inline int ceph_blog_client_des_callback(char *buf, size_t size,
+						u8 client_id)
+{
+	return 0;
+}
+static inline bool ceph_blog_is_enabled(struct ceph_fs_client *fsc)
+{
+	return false;
+}
+static inline struct blog_tls_ctx *
+ceph_blog_acquire_ctx(struct ceph_fs_client *fsc, gfp_t gfp,
+		      struct blog_module_context **held_mod)
+{
+	if (held_mod)
+		*held_mod = NULL;
+	return NULL;
+}
+static inline void ceph_blog_module_put(struct blog_module_context *ctx) {}
+static inline void ceph_blog_cpu_bind(struct blog_tls_ctx *ctx) {}
+static inline void ceph_blog_cpu_unbind(struct blog_tls_ctx *ctx) {}
+static inline void ceph_blog_cpu_clear(struct blog_tls_ctx *ctx) {}
+static inline struct blog_tls_ctx *
+ceph_blog_get_cached_ctx(struct ceph_fs_client *fsc)
+{
+	return NULL;
+}
+#endif
+
+/* ---------- entry / exit helpers ---------- */
+
+/**
+ * ceph_blog_enter_req_gfp - bind optional BLOG ctx; install journal_info for MDS
+ * @gfp: GFP_NOFS for sleepable VFS paths; GFP_ATOMIC (or any non-blocking
+ *       combination) for callbacks that must not sleep.  Non-blocking
+ *       acquires only reuse an existing per-task context; first-touch
+ *       allocation is skipped and logging is a no-op for that enter.
+ *
+ * BLOG state lives in the per-task map and CPU cache.  Plain VFS enters
+ * never publish into current->journal_info.  Only enter_req (MDS
+ * fill-trace, which already runs under memalloc_nofs_save) installs the
+ * tagged carrier so xattr paths can see mds_req.
+ */
+static inline void ceph_blog_enter_req_gfp(struct ceph_fs_client *fsc,
+					   struct ceph_journal_info *ji,
+					   struct ceph_mds_request *req,
+					   gfp_t gfp)
+{
+	struct ceph_journal_info *parent = ceph_ji_from_current();
+
+	ji->magic    = CEPH_JI_MAGIC;
+	ji->saved_ji = current->journal_info;
+	ji->fsc      = fsc;
+	ji->mds_req  = req;
+	ji->blog_mod = NULL;
+	ji->blog_ctx = ceph_ji_matches_fsc(parent, fsc) ? parent->blog_ctx : NULL;
+
+	if (!ji->blog_ctx && ceph_blog_is_enabled(fsc))
+		ji->blog_ctx = ceph_blog_acquire_ctx(fsc, gfp, &ji->blog_mod);
+
+	if (ji->blog_ctx)
+		ceph_blog_cpu_bind(ji->blog_ctx);
+
+	/* MDS fill-trace only: keep journal_info off reclaimable VFS paths. */
+	if (req)
+		current->journal_info = ceph_ji_encode(ji);
+}
+
+static inline void ceph_blog_enter_req(struct ceph_fs_client *fsc,
+				       struct ceph_journal_info *ji,
+				       struct ceph_mds_request *req)
+{
+	ceph_blog_enter_req_gfp(fsc, ji, req, GFP_NOFS);
+}
+
+static inline void ceph_blog_enter_gfp(struct ceph_fs_client *fsc,
+				       struct ceph_journal_info *ji,
+				       gfp_t gfp)
+{
+	struct ceph_journal_info *parent = ceph_ji_from_current();
+	struct ceph_mds_request *req =
+		ceph_ji_matches_fsc(parent, fsc) ? parent->mds_req : NULL;
+
+	ceph_blog_enter_req_gfp(fsc, ji, req, gfp);
+}
+
+static inline void ceph_blog_enter(struct ceph_fs_client *fsc,
+				   struct ceph_journal_info *ji)
+{
+	ceph_blog_enter_gfp(fsc, ji, GFP_NOFS);
+}
+
+/**
+ * ceph_blog_exit - call at every Ceph VFS exit point
+ * @ji: the same struct passed to ceph_blog_enter()
+ */
+static inline void ceph_blog_exit(struct ceph_journal_info *ji)
+{
+	if (ji->blog_ctx)
+		ceph_blog_cpu_unbind(ji->blog_ctx);
+
+	if (ji->blog_mod) {
+		ceph_blog_module_put(ji->blog_mod);
+		ji->blog_mod = NULL;
+	}
+
+	if (current->journal_info == ceph_ji_encode(ji))
+		current->journal_info = ji->saved_ji;
+}
+
+/* ---------- debugfs ---------- */
+
+#ifdef CONFIG_DEBUG_FS
+int  ceph_blog_debugfs_init(struct ceph_fs_client *fsc);
+void ceph_blog_debugfs_cleanup(struct ceph_fs_client *fsc);
+#else
+static inline int  ceph_blog_debugfs_init(struct ceph_fs_client *fsc) { return 0; }
+static inline void ceph_blog_debugfs_cleanup(struct ceph_fs_client *fsc) {}
+#endif
+
+#endif /* CEPH_BLOG_H */
diff --git a/include/linux/ceph/libceph.h b/include/linux/ceph/libceph.h
index d800b30425bf..4d493a08faf5 100644
--- a/include/linux/ceph/libceph.h
+++ b/include/linux/ceph/libceph.h
@@ -135,6 +135,8 @@ struct ceph_client {
 	struct ceph_osd_client osdc;
 
 #ifdef CONFIG_DEBUG_FS
+	u32 blog_client_id;
+
 	struct dentry *debugfs_dir;
 	struct dentry *debugfs_monmap;
 	struct dentry *debugfs_osdmap;
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 16+ messages in thread

* [PATCH v7 08/14] ceph: add boutc wrappers for BLOG
  2026-09-24 15:30 [PATCH v7 00/14] ceph: add binary logging (BLOG) for CephFS Alex Markuze
                   ` (6 preceding siblings ...)
  2026-09-24 15:30 ` [PATCH v7 07/14] ceph: add Ceph BLOG scaffolding Alex Markuze
@ 2026-09-24 15:30 ` Alex Markuze
  2026-09-24 15:30 ` [PATCH v7 09/14] ceph: switch MDS request plumbing to struct ceph_journal_info Alex Markuze
                   ` (6 subsequent siblings)
  14 siblings, 0 replies; 16+ messages in thread
From: Alex Markuze @ 2026-09-24 15:30 UTC (permalink / raw)
  To: ceph-devel; +Cc: idryomov, xiubo.li

Add boutc()/boutc_bounded()/boutc_formats() macros to ceph_debug.h as
drop-in replacements for doutc().  When BLOG is enabled they serialize
to binary; when disabled they fall back to doutc().

Resolve the optional binary context before choosing the backend.
Expand the doutc fallback only once per call site to avoid duplicate
dynamic-debug metadata and excessive VFS stack frames.

Signed-off-by: Alex Markuze <amarkuze@redhat.com>
Assisted-by: LLM
---
 include/linux/ceph/ceph_debug.h | 81 ++++++++++++++++++++++++++++-----
 1 file changed, 69 insertions(+), 12 deletions(-)

diff --git a/include/linux/ceph/ceph_debug.h b/include/linux/ceph/ceph_debug.h
index 5f904591fa5f..e055091d5c3c 100644
--- a/include/linux/ceph/ceph_debug.h
+++ b/include/linux/ceph/ceph_debug.h
@@ -5,28 +5,22 @@
 #define pr_fmt(fmt) KBUILD_MODNAME ": " fmt
 
 #include <linux/string.h>
+#include <linux/jump_label.h>
 
 #ifdef CONFIG_CEPH_LIB_PRETTYDEBUG
 
-/*
- * wrap pr_debug to include a filename:lineno prefix on each line.
- * this incurs some overhead (kernel size and execution time) due to
- * the extra function call at each call site.
- */
-
 # if defined(DEBUG) || defined(CONFIG_DYNAMIC_DEBUG)
 #  define dout(fmt, ...)						\
 	pr_debug("%.*s %12.12s:%-4d : " fmt,				\
 		 8 - (int)sizeof(KBUILD_MODNAME), "    ",		\
 		 kbasename(__FILE__), __LINE__, ##__VA_ARGS__)
 #  define doutc(client, fmt, ...)					\
-	pr_debug("%.*s %12.12s:%-4d : [%pU %llu] " fmt,			\
+	pr_debug("%.*s %12.12s:%-4d : [%pU %llu] " fmt,		\
 		 8 - (int)sizeof(KBUILD_MODNAME), "    ",		\
 		 kbasename(__FILE__), __LINE__,				\
 		 &client->fsid, client->monc.auth->global_id,		\
 		 ##__VA_ARGS__)
 # else
-/* faux printk call just to see any compiler warnings. */
 #  define dout(fmt, ...)					\
 		no_printk(KERN_DEBUG fmt, ##__VA_ARGS__)
 #  define doutc(client, fmt, ...)				\
@@ -38,16 +32,79 @@
 
 #else
 
-/*
- * or, just wrap pr_debug
- */
 # define dout(fmt, ...)	pr_debug(" " fmt, ##__VA_ARGS__)
 # define doutc(client, fmt, ...)					\
-	pr_debug(" [%pU %llu] %s: " fmt, &client->fsid,			\
+	pr_debug(" [%pU %llu] %s: " fmt, &client->fsid,		\
 		 client->monc.auth->global_id, __func__, ##__VA_ARGS__)
 
 #endif
 
+/*
+ * boutc* -- binary-logging variants of doutc*.
+ *
+ * When Ceph BLOG tracing has been explicitly enabled and a BLOG context
+ * has been bound for the current task (via ceph_blog_enter()),
+ * these route through the BLOG serialization path. Otherwise they fall
+ * back to the traditional text-based doutc macros so that existing
+ * debug semantics remain unchanged.
+ *
+ * Must only be used from fs/ceph/ where client->private is a
+ * ceph_fs_client *.  Do not call from net/ceph or RBD; ->private is not
+ * an fsc there (ceph_blog_key stays false, but the cast is still wrong).
+ */
+#define __ceph_blog_args(...) __VA_ARGS__
+#if IS_ENABLED(CONFIG_CEPH_FS) && IS_ENABLED(CONFIG_DEBUG_FS)
+# include <linux/ceph/ceph_blog.h>
+# define boutc(client, fmt, ...) \
+	do { \
+		struct blog_tls_ctx *__ctx = NULL; \
+		if (static_branch_unlikely(&ceph_blog_key)) \
+			__ctx = ceph_blog_get_cached_ctx( \
+				(struct ceph_fs_client *)(client)->private); \
+		if (__ctx) \
+			CEPH_BLOG_LOG_CLIENT(__ctx, client, fmt, ##__VA_ARGS__); \
+		else \
+			doutc(client, fmt, ##__VA_ARGS__); \
+	} while (0)
+# define boutc_bounded(client, fmt, blog_args, text_args) \
+	do { \
+		struct blog_tls_ctx *__ctx = NULL; \
+		if (static_branch_unlikely(&ceph_blog_key)) \
+			__ctx = ceph_blog_get_cached_ctx( \
+				(struct ceph_fs_client *)(client)->private); \
+		if (__ctx) \
+			CEPH_BLOG_LOG_CLIENT(__ctx, client, fmt, \
+					     __ceph_blog_args blog_args); \
+		else \
+			doutc(client, fmt, __ceph_blog_args text_args); \
+	} while (0)
+# define boutc_formats(client, blog_fmt, text_fmt, blog_args, text_args) \
+	do { \
+		struct blog_tls_ctx *__ctx = NULL; \
+		if (static_branch_unlikely(&ceph_blog_key)) \
+			__ctx = ceph_blog_get_cached_ctx( \
+				(struct ceph_fs_client *)(client)->private); \
+		if (__ctx) \
+			CEPH_BLOG_LOG_CLIENT(__ctx, client, blog_fmt, \
+					     __ceph_blog_args blog_args); \
+		else \
+			doutc(client, text_fmt, __ceph_blog_args text_args); \
+	} while (0)
+#else
+# define boutc(client, fmt, ...) doutc(client, fmt, ##__VA_ARGS__)
+# define boutc_bounded(client, fmt, blog_args, text_args) \
+	do { \
+		(void)sizeof(#blog_args); \
+		doutc(client, fmt, __ceph_blog_args text_args); \
+	} while (0)
+# define boutc_formats(client, blog_fmt, text_fmt, blog_args, text_args) \
+	do { \
+		(void)sizeof(blog_fmt); \
+		(void)sizeof(#blog_args); \
+		doutc(client, text_fmt, __ceph_blog_args text_args); \
+	} while (0)
+#endif
+
 #define pr_notice_client(client, fmt, ...)				\
 	pr_notice("[%pU %llu]: " fmt, &client->fsid,			\
 		  client->monc.auth->global_id, ##__VA_ARGS__)
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 16+ messages in thread

* [PATCH v7 09/14] ceph: switch MDS request plumbing to struct ceph_journal_info
  2026-09-24 15:30 [PATCH v7 00/14] ceph: add binary logging (BLOG) for CephFS Alex Markuze
                   ` (7 preceding siblings ...)
  2026-09-24 15:30 ` [PATCH v7 08/14] ceph: add boutc wrappers for BLOG Alex Markuze
@ 2026-09-24 15:30 ` Alex Markuze
  2026-09-24 15:30 ` [PATCH v7 10/14] ceph: add BLOG debugfs interface Alex Markuze
                   ` (5 subsequent siblings)
  14 siblings, 0 replies; 16+ messages in thread
From: Alex Markuze @ 2026-09-24 15:30 UTC (permalink / raw)
  To: ceph-devel; +Cc: idryomov, xiubo.li

Replace ad-hoc journal_info setup in handle_reply() with
ceph_blog_enter_req(mdsc->fsc, &ji, req) / ceph_blog_exit(&ji) so nested
VFS callbacks inherit the BLOG context only when the owning fsc
matches.

Snapshot unlink-conflict names only when a binary context is present.
Keep that preparation in an out-of-line helper and retain the lazy
%pd text fallback.

Signed-off-by: Alex Markuze <amarkuze@redhat.com>
Assisted-by: LLM
---
 fs/ceph/mds_client.c | 369 +++++++++++++++++++++++++------------------
 1 file changed, 217 insertions(+), 152 deletions(-)

diff --git a/fs/ceph/mds_client.c b/fs/ceph/mds_client.c
index 0185c3c5edab..8ae3fa099b45 100644
--- a/fs/ceph/mds_client.c
+++ b/fs/ceph/mds_client.c
@@ -527,7 +527,9 @@ static int parse_reply_info_readdir(void **p, void *end,
 		ceph_decode_need(p, end, _name_len, bad);
 		_name = *p;
 		*p += _name_len;
-		doutc(cl, "parsed dir dname '%.*s'\n", _name_len, _name);
+		boutc_bounded(cl, "parsed dir dname '%.*s'\n",
+			      (_name_len, BLOG_STR(_name, _name_len)),
+			      (_name_len, (const char *)_name));
 
 		if (info->hash_order)
 			rde->raw_hash = ceph_str_hash(ci->i_dir_layout.dl_dir_hash,
@@ -677,7 +679,7 @@ static int ceph_parse_deleg_inos(void **p, void *end,
 	u32 sets;
 
 	ceph_decode_32_safe(p, end, sets, bad);
-	doutc(cl, "got %u sets of delegated inodes\n", sets);
+	boutc(cl, "got %u sets of delegated inodes\n", sets);
 	while (sets--) {
 		u64 start, len;
 
@@ -711,7 +713,7 @@ static int ceph_parse_deleg_inos(void **p, void *end,
 			int err = ceph_insert_deleg_ino(s, start++);
 
 			if (!err) {
-				doutc(cl, "added delegated inode 0x%llx\n", start - 1);
+				boutc(cl, "added delegated inode 0x%llx\n", start - 1);
 			} else if (err == -EBUSY) {
 				pr_warn_client(cl,
 					"MDS delegated inode 0x%llx more than once.\n",
@@ -935,6 +937,23 @@ static void destroy_reply_info(struct ceph_mds_reply_info_parsed *info)
 	free_pages((unsigned long)info->dir_entries, get_order(info->dir_buf_size));
 }
 
+static noinline void ceph_blog_unlink_conflict(struct blog_tls_ctx *blog_ctx,
+					       struct ceph_client *cl,
+					       struct dentry *dentry,
+					       struct dentry *found)
+{
+	struct name_snapshot dname_snap, found_snap;
+
+	take_dentry_name_snapshot(&dname_snap, dentry);
+	take_dentry_name_snapshot(&found_snap, found);
+	CEPH_BLOG_LOG_CLIENT(blog_ctx, cl,
+			     "dentry %p:%s conflict with old %p:%s\n",
+			     dentry, BLOG_STR(dname_snap.name.name, dname_snap.name.len),
+			     found, BLOG_STR(found_snap.name.name, found_snap.name.len));
+	release_dentry_name_snapshot(&dname_snap);
+	release_dentry_name_snapshot(&found_snap);
+}
+
 /*
  * In async unlink case the kclient won't wait for the first reply
  * from MDS and just drop all the links and unhash the dentry and then
@@ -960,6 +979,7 @@ int ceph_wait_on_conflict_unlink(struct dentry *dentry)
 	struct ceph_fs_client *fsc = ceph_sb_to_fs_client(dentry->d_sb);
 	struct ceph_client *cl = fsc->client;
 	struct dentry *pdentry = dentry->d_parent;
+	struct blog_tls_ctx *blog_ctx;
 	struct dentry *udentry, *found = NULL;
 	struct ceph_dentry_info *di;
 	struct qstr dname;
@@ -1000,8 +1020,12 @@ int ceph_wait_on_conflict_unlink(struct dentry *dentry)
 	if (likely(!found))
 		return 0;
 
-	doutc(cl, "dentry %p:%pd conflict with old %p:%pd\n", dentry, dentry,
-	      found, found);
+	blog_ctx = ceph_blog_get_ctx(fsc);
+	if (blog_ctx)
+		ceph_blog_unlink_conflict(blog_ctx, cl, dentry, found);
+	else
+		doutc(cl, "dentry %p:%pd conflict with old %p:%pd\n",
+		      dentry, dentry, found, found);
 
 	err = wait_on_bit(&di->flags, CEPH_DENTRY_ASYNC_UNLINK_BIT,
 			  TASK_KILLABLE);
@@ -1103,7 +1127,7 @@ static struct ceph_mds_session *register_session(struct ceph_mds_client *mdsc,
 		struct ceph_mds_session **sa;
 		size_t ptr_size = sizeof(struct ceph_mds_session *);
 
-		doutc(cl, "realloc to %d\n", newmax);
+		boutc(cl, "realloc to %d\n", newmax);
 		sa = kcalloc(newmax, ptr_size, GFP_NOFS);
 		if (!sa)
 			goto fail_realloc;
@@ -1116,7 +1140,7 @@ static struct ceph_mds_session *register_session(struct ceph_mds_client *mdsc,
 		mdsc->max_sessions = newmax;
 	}
 
-	doutc(cl, "mds%d\n", mds);
+	boutc(cl, "mds%d\n", mds);
 	s->s_mdsc = mdsc;
 	s->s_mds = mds;
 	s->s_state = CEPH_MDS_SESSION_NEW;
@@ -1160,7 +1184,7 @@ static struct ceph_mds_session *register_session(struct ceph_mds_client *mdsc,
 static void __unregister_session(struct ceph_mds_client *mdsc,
 			       struct ceph_mds_session *s)
 {
-	doutc(mdsc->fsc->client, "mds%d %p\n", s->s_mds, s);
+	boutc(mdsc->fsc->client, "mds%d %p\n", s->s_mds, s);
 	BUG_ON(mdsc->sessions[s->s_mds] != s);
 	mdsc->sessions[s->s_mds] = NULL;
 	ceph_con_close(&s->s_con);
@@ -1303,7 +1327,7 @@ static void __register_request(struct ceph_mds_client *mdsc,
 			return;
 		}
 	}
-	doutc(cl, "%p tid %lld\n", req, req->r_tid);
+	boutc(cl, "%p tid %lld\n", req, req->r_tid);
 	ceph_mdsc_get_request(req);
 	insert_request(&mdsc->request_tree, req);
 
@@ -1328,7 +1352,7 @@ static void __register_request(struct ceph_mds_client *mdsc,
 static void __unregister_request(struct ceph_mds_client *mdsc,
 				 struct ceph_mds_request *req)
 {
-	doutc(mdsc->fsc->client, "%p tid %lld\n", req, req->r_tid);
+	boutc(mdsc->fsc->client, "%p tid %lld\n", req, req->r_tid);
 
 	/* Never leave an unregistered request on an unsafe list! */
 	list_del_init(&req->r_unsafe_item);
@@ -1426,7 +1450,7 @@ static int __choose_mds(struct ceph_mds_client *mdsc,
 	if (req->r_resend_mds >= 0 &&
 	    (__have_session(mdsc, req->r_resend_mds) ||
 	     ceph_mdsmap_get_state(mdsc->mdsmap, req->r_resend_mds) > 0)) {
-		doutc(cl, "using resend_mds mds%d\n", req->r_resend_mds);
+		boutc(cl, "using resend_mds mds%d\n", req->r_resend_mds);
 		return req->r_resend_mds;
 	}
 
@@ -1443,7 +1467,7 @@ static int __choose_mds(struct ceph_mds_client *mdsc,
 			rcu_read_lock();
 			inode = get_nonsnap_parent(req->r_dentry);
 			rcu_read_unlock();
-			doutc(cl, "using snapdir's parent %p %llx.%llx\n",
+			boutc(cl, "using snapdir's parent %p %llx.%llx\n",
 			      inode, ceph_vinop(inode));
 		}
 	} else if (req->r_dentry) {
@@ -1464,7 +1488,7 @@ static int __choose_mds(struct ceph_mds_client *mdsc,
 			/* direct snapped/virtual snapdir requests
 			 * based on parent dir inode */
 			inode = get_nonsnap_parent(parent);
-			doutc(cl, "using nonsnap parent %p %llx.%llx\n",
+			boutc(cl, "using nonsnap parent %p %llx.%llx\n",
 			      inode, ceph_vinop(inode));
 		} else {
 			/* dentry target */
@@ -1484,7 +1508,7 @@ static int __choose_mds(struct ceph_mds_client *mdsc,
 	if (!inode)
 		goto random;
 
-	doutc(cl, "%p %llx.%llx is_hash=%d (0x%x) mode %d\n", inode,
+	boutc(cl, "%p %llx.%llx is_hash=%d (0x%x) mode %d\n", inode,
 	      ceph_vinop(inode), (int)is_hash, hash, mode);
 	ci = ceph_inode(inode);
 
@@ -1501,7 +1525,7 @@ static int __choose_mds(struct ceph_mds_client *mdsc,
 				get_random_bytes(&r, 1);
 				r %= frag.ndist;
 				mds = frag.dist[r];
-				doutc(cl, "%p %llx.%llx frag %u mds%d (%d/%d)\n",
+				boutc(cl, "%p %llx.%llx frag %u mds%d (%d/%d)\n",
 				      inode, ceph_vinop(inode), frag.frag,
 				      mds, (int)r, frag.ndist);
 				if (ceph_mdsmap_get_state(mdsc->mdsmap, mds) >=
@@ -1516,7 +1540,7 @@ static int __choose_mds(struct ceph_mds_client *mdsc,
 			if (frag.mds >= 0) {
 				/* choose auth mds */
 				mds = frag.mds;
-				doutc(cl, "%p %llx.%llx frag %u mds%d (auth)\n",
+				boutc(cl, "%p %llx.%llx frag %u mds%d (auth)\n",
 				      inode, ceph_vinop(inode), frag.frag, mds);
 				if (ceph_mdsmap_get_state(mdsc->mdsmap, mds) >=
 				    CEPH_MDS_STATE_ACTIVE) {
@@ -1541,7 +1565,7 @@ static int __choose_mds(struct ceph_mds_client *mdsc,
 		goto random;
 	}
 	mds = cap->session->s_mds;
-	doutc(cl, "%p %llx.%llx mds%d (%scap %p)\n", inode,
+	boutc(cl, "%p %llx.%llx mds%d (%scap %p)\n", inode,
 	      ceph_vinop(inode), mds,
 	      cap == ci->i_auth_cap ? "auth " : "", cap);
 	spin_unlock(&ci->i_ceph_lock);
@@ -1554,7 +1578,7 @@ static int __choose_mds(struct ceph_mds_client *mdsc,
 		*random = true;
 
 	mds = ceph_mdsmap_get_random_mds(mdsc->mdsmap);
-	doutc(cl, "chose random mds%d\n", mds);
+	boutc(cl, "chose random mds%d\n", mds);
 	return mds;
 }
 
@@ -1794,7 +1818,7 @@ static int __open_session(struct ceph_mds_client *mdsc,
 
 	/* wait for mds to go active? */
 	mstate = ceph_mdsmap_get_state(mdsc->mdsmap, mds);
-	doutc(mdsc->fsc->client, "open_session to mds%d (%s)\n", mds,
+	boutc(mdsc->fsc->client, "open_session to mds%d (%s)\n", mds,
 	      ceph_mds_state_name(mstate));
 	session->s_state = CEPH_MDS_SESSION_OPENING;
 	session->s_renew_requested = jiffies;
@@ -1843,7 +1867,7 @@ ceph_mdsc_open_export_target_session(struct ceph_mds_client *mdsc, int target)
 	struct ceph_mds_session *session;
 	struct ceph_client *cl = mdsc->fsc->client;
 
-	doutc(cl, "to mds%d\n", target);
+	boutc(cl, "to mds%d\n", target);
 
 	mutex_lock(&mdsc->mutex);
 	session = __open_export_target_session(mdsc, target);
@@ -1864,7 +1888,7 @@ static void __open_export_target_sessions(struct ceph_mds_client *mdsc,
 		return;
 
 	mi = &mdsc->mdsmap->m_info[mds];
-	doutc(cl, "for mds%d (%d targets)\n", session->s_mds,
+	boutc(cl, "for mds%d (%d targets)\n", session->s_mds,
 	      mi->num_export_targets);
 
 	for (i = 0; i < mi->num_export_targets; i++) {
@@ -1887,7 +1911,7 @@ static int detach_cap_releases(struct ceph_mds_session *session,
 
 	list_splice_init(&session->s_cap_releases, target);
 	session->s_num_cap_releases = 0;
-	doutc(cl, "mds%d\n", session->s_mds);
+	boutc(cl, "mds%d\n", session->s_mds);
 
 	return num_cap_releases;
 }
@@ -1911,7 +1935,7 @@ static void cleanup_session_requests(struct ceph_mds_client *mdsc,
 	struct ceph_mds_request *req;
 	struct rb_node *p;
 
-	doutc(cl, "mds%d\n", session->s_mds);
+	boutc(cl, "mds%d\n", session->s_mds);
 	mutex_lock(&mdsc->mutex);
 	while (!list_empty(&session->s_unsafe)) {
 		req = list_first_entry(&session->s_unsafe,
@@ -1953,7 +1977,7 @@ int ceph_iterate_session_caps(struct ceph_mds_session *session,
 	struct ceph_cap *old_cap = NULL;
 	int ret;
 
-	doutc(cl, "%p mds%d\n", session, session->s_mds);
+	boutc(cl, "%p mds%d\n", session, session->s_mds);
 	spin_lock(&session->s_cap_lock);
 	p = session->s_caps.next;
 	while (p != &session->s_caps) {
@@ -1984,7 +2008,7 @@ int ceph_iterate_session_caps(struct ceph_mds_session *session,
 		spin_lock(&session->s_cap_lock);
 		p = p->next;
 		if (ceph_cap_is_removed(cap)) {
-			doutc(cl, "finishing cap %p removal\n", cap);
+			boutc(cl, "finishing cap %p removal\n", cap);
 			BUG_ON(cap->session != session);
 			cap->session = NULL;
 			list_del_init(&cap->session_caps);
@@ -2021,7 +2045,7 @@ static int remove_session_caps_cb(struct inode *inode, int mds, void *arg)
 	spin_lock(&ci->i_ceph_lock);
 	cap = __get_cap_for_mds(ci, mds);
 	if (cap) {
-		doutc(cl, " removing cap %p, ci is %p, inode is %p\n",
+		boutc(cl, " removing cap %p, ci is %p, inode is %p\n",
 		      cap, ci, &ci->netfs.inode);
 
 		iputs = ceph_purge_inode_cap(inode, cap, &invalidate);
@@ -2046,7 +2070,7 @@ static void remove_session_caps(struct ceph_mds_session *session)
 	struct super_block *sb = fsc->sb;
 	LIST_HEAD(dispose);
 
-	doutc(fsc->client, "on %p\n", session);
+	boutc(fsc->client, "on %p\n", session);
 	ceph_iterate_session_caps(session, remove_session_caps_cb, fsc);
 
 	wake_up_all(&fsc->mdsc->cap_flushing_wq);
@@ -2129,7 +2153,7 @@ static void wake_up_session_caps(struct ceph_mds_session *session, int ev)
 {
 	struct ceph_client *cl = session->s_mdsc->fsc->client;
 
-	doutc(cl, "session %p mds%d\n", session, session->s_mds);
+	boutc(cl, "session %p mds%d\n", session, session->s_mds);
 	ceph_iterate_session_caps(session, wake_up_session_cb,
 				  (void *)(unsigned long)ev);
 }
@@ -2156,12 +2180,12 @@ static int send_renew_caps(struct ceph_mds_client *mdsc,
 	 * with its clients. */
 	state = ceph_mdsmap_get_state(mdsc->mdsmap, session->s_mds);
 	if (state < CEPH_MDS_STATE_RECONNECT) {
-		doutc(cl, "ignoring mds%d (%s)\n", session->s_mds,
+		boutc(cl, "ignoring mds%d (%s)\n", session->s_mds,
 		      ceph_mds_state_name(state));
 		return 0;
 	}
 
-	doutc(cl, "to mds%d (%s)\n", session->s_mds,
+	boutc(cl, "to mds%d (%s)\n", session->s_mds,
 	      ceph_mds_state_name(state));
 	msg = create_session_full_msg(mdsc, CEPH_SESSION_REQUEST_RENEWCAPS,
 				      ++session->s_renew_seq);
@@ -2177,7 +2201,7 @@ static int send_flushmsg_ack(struct ceph_mds_client *mdsc,
 	struct ceph_client *cl = mdsc->fsc->client;
 	struct ceph_msg *msg;
 
-	doutc(cl, "to mds%d (%s)s seq %lld\n", session->s_mds,
+	boutc(cl, "to mds%d (%s)s seq %lld\n", session->s_mds,
 	      ceph_session_state_name(session->s_state), seq);
 	msg = ceph_create_session_msg(CEPH_SESSION_FLUSHMSG_ACK, seq);
 	if (!msg)
@@ -2215,7 +2239,7 @@ static void renewed_caps(struct ceph_mds_client *mdsc,
 				       session->s_mds);
 		}
 	}
-	doutc(cl, "mds%d ttl now %lu, was %s, now %s\n", session->s_mds,
+	boutc(cl, "mds%d ttl now %lu, was %s, now %s\n", session->s_mds,
 	      session->s_cap_ttl, was_stale ? "stale" : "fresh",
 	      time_before(jiffies, session->s_cap_ttl) ? "stale" : "fresh");
 	spin_unlock(&session->s_cap_lock);
@@ -2232,7 +2256,7 @@ static int request_close_session(struct ceph_mds_session *session)
 	struct ceph_client *cl = session->s_mdsc->fsc->client;
 	struct ceph_msg *msg;
 
-	doutc(cl, "mds%d state %s seq %lld\n", session->s_mds,
+	boutc(cl, "mds%d state %s seq %lld\n", session->s_mds,
 	      ceph_session_state_name(session->s_state), session->s_seq);
 	msg = ceph_create_session_msg(CEPH_SESSION_REQUEST_CLOSE,
 				      session->s_seq);
@@ -2311,7 +2335,7 @@ static int trim_caps_cb(struct inode *inode, int mds, void *arg)
 	wanted = __ceph_caps_file_wanted(ci);
 	oissued = __ceph_caps_issued_other(ci, cap);
 
-	doutc(cl, "%p %llx.%llx cap %p mine %s oissued %s used %s wanted %s\n",
+	boutc(cl, "%p %llx.%llx cap %p mine %s oissued %s used %s wanted %s\n",
 	      inode, ceph_vinop(inode), cap, ceph_cap_string(mine),
 	      ceph_cap_string(oissued), ceph_cap_string(used),
 	      ceph_cap_string(wanted));
@@ -2380,7 +2404,7 @@ static int trim_caps_cb(struct inode *inode, int mds, void *arg)
 			count = icount_read_once(inode);
 			if (count == 1)
 				(*remaining)--;
-			doutc(cl, "%p %llx.%llx cap %p pruned, count now %d\n",
+			boutc(cl, "%p %llx.%llx cap %p pruned, count now %d\n",
 			      inode, ceph_vinop(inode), cap, count);
 		}
 		return 0;
@@ -2401,13 +2425,13 @@ int ceph_trim_caps(struct ceph_mds_client *mdsc,
 	struct ceph_client *cl = mdsc->fsc->client;
 	int trim_caps = session->s_nr_caps - max_caps;
 
-	doutc(cl, "mds%d start: %d / %d, trim %d\n", session->s_mds,
+	boutc(cl, "mds%d start: %d / %d, trim %d\n", session->s_mds,
 	      session->s_nr_caps, max_caps, trim_caps);
 	if (trim_caps > 0) {
 		int remaining = trim_caps;
 
 		ceph_iterate_session_caps(session, trim_caps_cb, &remaining);
-		doutc(cl, "mds%d done: %d / %d, trimmed %d\n",
+		boutc(cl, "mds%d done: %d / %d, trimmed %d\n",
 		      session->s_mds, session->s_nr_caps, max_caps,
 		      trim_caps - remaining);
 	}
@@ -2428,7 +2452,7 @@ static int check_caps_flush(struct ceph_mds_client *mdsc,
 			list_first_entry(&mdsc->cap_flush_list,
 					 struct ceph_cap_flush, g_list);
 		if (cf->tid <= want_flush_tid) {
-			doutc(cl, "still flushing tid %llu <= %llu\n",
+			boutc(cl, "still flushing tid %llu <= %llu\n",
 			      cf->tid, want_flush_tid);
 			ret = 0;
 		}
@@ -2528,7 +2552,7 @@ static void wait_caps_flush(struct ceph_mds_client *mdsc,
 	int i = 0;
 	long ret;
 
-	doutc(cl, "want %llu\n", want_flush_tid);
+	boutc(cl, "want %llu\n", want_flush_tid);
 
 	do {
 		/* 60 * HZ fits in a long on all supported architectures. */
@@ -2545,7 +2569,7 @@ static void wait_caps_flush(struct ceph_mds_client *mdsc,
 		}
 	} while (ret == 0);
 
-	doutc(cl, "ok, flushed thru %llu\n", want_flush_tid);
+	boutc(cl, "ok, flushed thru %llu\n", want_flush_tid);
 }
 
 /*
@@ -2611,7 +2635,7 @@ static void ceph_send_cap_releases(struct ceph_mds_client *mdsc,
 			msg->front.iov_len += sizeof(*cap_barrier);
 
 			msg->hdr.front_len = cpu_to_le32(msg->front.iov_len);
-			doutc(cl, "mds%d %p\n", session->s_mds, msg);
+			boutc(cl, "mds%d %p\n", session->s_mds, msg);
 			ceph_con_send(&session->s_con, msg);
 			msg = NULL;
 		}
@@ -2631,7 +2655,7 @@ static void ceph_send_cap_releases(struct ceph_mds_client *mdsc,
 		msg->front.iov_len += sizeof(*cap_barrier);
 
 		msg->hdr.front_len = cpu_to_le32(msg->front.iov_len);
-		doutc(cl, "mds%d %p\n", session->s_mds, msg);
+		boutc(cl, "mds%d %p\n", session->s_mds, msg);
 		ceph_con_send(&session->s_con, msg);
 	}
 	return;
@@ -2667,10 +2691,10 @@ void ceph_flush_session_cap_releases(struct ceph_mds_client *mdsc,
 	ceph_get_mds_session(session);
 	if (queue_work(mdsc->fsc->cap_wq,
 		       &session->s_cap_release_work)) {
-		doutc(cl, "cap release work queued\n");
+		boutc(cl, "cap release work queued\n");
 	} else {
 		ceph_put_mds_session(session);
-		doutc(cl, "failed to queue cap release work\n");
+		boutc(cl, "failed to queue cap release work\n");
 	}
 }
 
@@ -2691,9 +2715,14 @@ static void ceph_cap_reclaim_work(struct work_struct *work)
 {
 	struct ceph_mds_client *mdsc =
 		container_of(work, struct ceph_mds_client, cap_reclaim_work);
-	int ret = ceph_trim_dentries(mdsc);
+	struct ceph_journal_info __ji;
+	int ret;
+
+	ceph_blog_enter(mdsc->fsc, &__ji);
+	ret = ceph_trim_dentries(mdsc);
 	if (ret == -EAGAIN)
 		ceph_queue_cap_reclaim_work(mdsc);
+	ceph_blog_exit(&__ji);
 }
 
 void ceph_queue_cap_reclaim_work(struct ceph_mds_client *mdsc)
@@ -2702,11 +2731,10 @@ void ceph_queue_cap_reclaim_work(struct ceph_mds_client *mdsc)
 	if (mdsc->stopping)
 		return;
 
-        if (queue_work(mdsc->fsc->cap_wq, &mdsc->cap_reclaim_work)) {
-                doutc(cl, "caps reclaim work queued\n");
-        } else {
-                doutc(cl, "failed to queue caps release work\n");
-        }
+	if (queue_work(mdsc->fsc->cap_wq, &mdsc->cap_reclaim_work))
+		doutc(cl, "caps reclaim work queued\n");
+	else
+		doutc(cl, "failed to queue caps release work\n");
 }
 
 void ceph_reclaim_caps_nr(struct ceph_mds_client *mdsc, int nr)
@@ -2727,11 +2755,10 @@ void ceph_queue_cap_unlink_work(struct ceph_mds_client *mdsc)
 	if (mdsc->stopping)
 		return;
 
-        if (queue_work(mdsc->fsc->cap_wq, &mdsc->cap_unlink_work)) {
-                doutc(cl, "caps unlink work queued\n");
-        } else {
-                doutc(cl, "failed to queue caps unlink work\n");
-        }
+	if (queue_work(mdsc->fsc->cap_wq, &mdsc->cap_unlink_work))
+		doutc(cl, "caps unlink work queued\n");
+	else
+		doutc(cl, "failed to queue caps unlink work\n");
 }
 
 static void ceph_cap_unlink_work(struct work_struct *work)
@@ -2739,8 +2766,11 @@ static void ceph_cap_unlink_work(struct work_struct *work)
 	struct ceph_mds_client *mdsc =
 		container_of(work, struct ceph_mds_client, cap_unlink_work);
 	struct ceph_client *cl = mdsc->fsc->client;
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(mdsc->fsc, &__ji);
 
-	doutc(cl, "begin\n");
+	boutc(cl, "begin\n");
 	spin_lock(&mdsc->cap_delay_lock);
 	while (!list_empty(&mdsc->cap_unlink_delay_list)) {
 		struct ceph_inode_info *ci;
@@ -2754,7 +2784,7 @@ static void ceph_cap_unlink_work(struct work_struct *work)
 		inode = igrab(&ci->netfs.inode);
 		if (inode) {
 			spin_unlock(&mdsc->cap_delay_lock);
-			doutc(cl, "on %p %llx.%llx\n", inode,
+			boutc(cl, "on %p %llx.%llx\n", inode,
 			      ceph_vinop(inode));
 			ceph_check_caps(ci, CHECK_CAPS_FLUSH);
 			iput(inode);
@@ -2762,7 +2792,8 @@ static void ceph_cap_unlink_work(struct work_struct *work)
 		}
 	}
 	spin_unlock(&mdsc->cap_delay_lock);
-	doutc(cl, "done\n");
+	boutc(cl, "done\n");
+	ceph_blog_exit(&__ji);
 }
 
 /*
@@ -2977,7 +3008,7 @@ char *ceph_mdsc_build_path(struct ceph_mds_client *mdsc, struct dentry *dentry,
 		spin_lock(&cur->d_lock);
 		inode = d_inode(cur);
 		if (inode && ceph_snap(inode) == CEPH_SNAPDIR) {
-			doutc(cl, "path+%d: %p SNAPDIR\n", pos, cur);
+			boutc(cl, "path+%d: %p SNAPDIR\n", pos, cur);
 			spin_unlock(&cur->d_lock);
 			parent = dget_parent(cur);
 		} else if (for_wire && inode && dentry != cur &&
@@ -3076,8 +3107,11 @@ char *ceph_mdsc_build_path(struct ceph_mds_client *mdsc, struct dentry *dentry,
 	else
 		path_info->vino.snap = CEPH_NOSNAP;
 
-	doutc(cl, "on %p %d built %llx '%.*s'\n", dentry, d_count(dentry),
-	      base, PATH_MAX - 1 - pos, path + pos);
+	boutc_bounded(cl, "on %p %d built %llx '%.*s'\n",
+		      (dentry, d_count(dentry), base, PATH_MAX - 1 - pos,
+		       BLOG_STR(path + pos, PATH_MAX - 1 - pos)),
+		      (dentry, d_count(dentry), base, PATH_MAX - 1 - pos,
+		       (const char *)(path + pos)));
 	return path + pos;
 }
 
@@ -3154,12 +3188,15 @@ static int set_request_path_attr(struct ceph_mds_client *mdsc, struct inode *rin
 
 	if (rinode) {
 		r = build_inode_path(rinode, path_info);
-		doutc(cl, " inode %p %llx.%llx\n", rinode, ceph_ino(rinode),
+		boutc(cl, " inode %p %llx.%llx\n", rinode, ceph_ino(rinode),
 		      ceph_snap(rinode));
 	} else if (rdentry) {
 		r = build_dentry_path(mdsc, rdentry, rdiri, path_info, parent_locked);
-		doutc(cl, " dentry %p %llx/%.*s\n", rdentry, path_info->vino.ino,
-		      path_info->pathlen, path_info->path);
+		boutc_bounded(cl, " dentry %p %llx/%.*s\n",
+			      (rdentry, path_info->vino.ino, path_info->pathlen,
+			       BLOG_STR(path_info->path, path_info->pathlen)),
+			      (rdentry, path_info->vino.ino, path_info->pathlen,
+			       (const char *)path_info->path));
 	} else if (rpath || rino) {
 		path_info->vino.ino = rino;
 		path_info->vino.snap = CEPH_NOSNAP;
@@ -3167,7 +3204,10 @@ static int set_request_path_attr(struct ceph_mds_client *mdsc, struct inode *rin
 		path_info->pathlen = rpath ? strlen(rpath) : 0;
 		path_info->freepath = false;
 
-		doutc(cl, " path %.*s\n", path_info->pathlen, rpath);
+		boutc_bounded(cl, " path %.*s\n",
+			      (path_info->pathlen,
+			       BLOG_STR(rpath, path_info->pathlen)),
+			      (path_info->pathlen, (const char *)rpath));
 	}
 
 	return r;
@@ -3586,7 +3626,7 @@ static int __prepare_send_request(struct ceph_mds_session *session,
 		else
 			req->r_sent_on_mseq = -1;
 	}
-	doutc(cl, "%p tid %lld %s (attempt %d)\n", req, req->r_tid,
+	boutc(cl, "%p tid %lld %s (attempt %d)\n", req, req->r_tid,
 	      ceph_mds_op_name(req->r_op), req->r_attempts);
 
 	if (test_bit(CEPH_MDS_R_GOT_UNSAFE, &req->r_req_flags)) {
@@ -3655,7 +3695,7 @@ static int __prepare_send_request(struct ceph_mds_session *session,
 		nhead->ext_num_retry = cpu_to_le32(req->r_attempts - 1);
 	}
 
-	doutc(cl, " r_parent = %p\n", req->r_parent);
+	boutc(cl, " r_parent = %p\n", req->r_parent);
 	return 0;
 }
 
@@ -3761,7 +3801,7 @@ static void __do_request(struct ceph_mds_client *mdsc,
 	}
 	req->r_session = ceph_get_mds_session(session);
 
-	doutc(cl, "mds%d session %p state %s\n", mds, session,
+	boutc(cl, "mds%d session %p state %s\n", mds, session,
 	      ceph_session_state_name(session->s_state));
 
 	/*
@@ -3858,7 +3898,7 @@ static void __do_request(struct ceph_mds_client *mdsc,
 		cap = ci->i_auth_cap;
 		if (test_bit(CEPH_I_ASYNC_CREATE_BIT, &ci->i_ceph_flags) &&
 		    mds != cap->mds) {
-			doutc(cl, "session changed for auth cap %d -> %d\n",
+			boutc(cl, "session changed for auth cap %d -> %d\n",
 			      cap->session->s_mds, session->s_mds);
 
 			/* Remove the auth cap from old session */
@@ -3886,7 +3926,7 @@ static void __do_request(struct ceph_mds_client *mdsc,
 	ceph_put_mds_session(session);
 finish:
 	if (err) {
-		doutc(cl, "early error %d\n", err);
+		boutc(cl, "early error %d\n", err);
 		req->r_err = err;
 		complete_request(mdsc, req);
 		__unregister_request(mdsc, req);
@@ -3927,7 +3967,7 @@ static void kick_requests(struct ceph_mds_client *mdsc, int mds)
 	struct ceph_mds_request *req;
 	struct rb_node *p = rb_first(&mdsc->request_tree);
 
-	doutc(cl, "kick_requests mds%d\n", mds);
+	boutc(cl, "kick_requests mds%d\n", mds);
 	while (p) {
 		req = rb_entry(p, struct ceph_mds_request, r_node);
 		p = rb_next(p);
@@ -3986,7 +4026,7 @@ int ceph_mdsc_submit_request(struct ceph_mds_client *mdsc, struct inode *dir,
 	if (req->r_inode) {
 		err = ceph_wait_on_async_create(req->r_inode);
 		if (err) {
-			doutc(cl, "wait for async create returned: %d\n", err);
+			boutc(cl, "wait for async create returned: %d\n", err);
 			return err;
 		}
 	}
@@ -4017,7 +4057,7 @@ int ceph_mdsc_wait_request(struct ceph_mds_client *mdsc,
 	int err;
 
 	/* wait */
-	doutc(cl, "do_request waiting\n");
+	boutc(cl, "do_request waiting\n");
 	if (wait_func) {
 		err = wait_func(mdsc, req);
 	} else {
@@ -4031,14 +4071,14 @@ int ceph_mdsc_wait_request(struct ceph_mds_client *mdsc,
 		else
 			err = timeleft;  /* killed */
 	}
-	doutc(cl, "do_request waited, got %d\n", err);
+	boutc(cl, "do_request waited, got %d\n", err);
 	mutex_lock(&mdsc->mutex);
 
 	/* only abort if we didn't race with a real reply */
 	if (test_bit(CEPH_MDS_R_GOT_RESULT, &req->r_req_flags)) {
 		err = le32_to_cpu(req->r_reply_info.head->result);
 	} else if (err < 0) {
-		doutc(cl, "aborted request %lld with %d\n", req->r_tid, err);
+		boutc(cl, "aborted request %lld with %d\n", req->r_tid, err);
 
 		/*
 		 * ensure we aren't running concurrently with
@@ -4072,13 +4112,13 @@ int ceph_mdsc_do_request(struct ceph_mds_client *mdsc,
 	struct ceph_client *cl = mdsc->fsc->client;
 	int err;
 
-	doutc(cl, "do_request on %p\n", req);
+	boutc(cl, "do_request on %p\n", req);
 
 	/* issue */
 	err = ceph_mdsc_submit_request(mdsc, dir, req);
 	if (!err)
 		err = ceph_mdsc_wait_request(mdsc, req, NULL);
-	doutc(cl, "do_request %p done, result %d\n", req, err);
+	boutc(cl, "do_request %p done, result %d\n", req, err);
 	return err;
 }
 
@@ -4092,7 +4132,7 @@ void ceph_invalidate_dir_request(struct ceph_mds_request *req)
 	struct inode *old_dir = req->r_old_dentry_dir;
 	struct ceph_client *cl = req->r_mdsc->fsc->client;
 
-	doutc(cl, "invalidate_dir_request %p %p (complete, lease(s))\n",
+	boutc(cl, "invalidate_dir_request %p %p (complete, lease(s))\n",
 	      dir, old_dir);
 
 	ceph_dir_clear_complete(dir);
@@ -4124,10 +4164,14 @@ static void handle_reply(struct ceph_mds_session *session, struct ceph_msg *msg)
 	int err, result;
 	int mds = session->s_mds;
 	bool close_sessions = false;
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(mdsc->fsc, &__ji);
 
 	if (msg->front.iov_len < sizeof(*head)) {
 		pr_err_client(cl, "got corrupt (short) reply\n");
 		ceph_msg_dump(msg);
+		ceph_blog_exit(&__ji);
 		return;
 	}
 
@@ -4136,11 +4180,12 @@ static void handle_reply(struct ceph_mds_session *session, struct ceph_msg *msg)
 	mutex_lock(&mdsc->mutex);
 	req = lookup_get_request(mdsc, tid);
 	if (!req) {
-		doutc(cl, "on unknown tid %llu\n", tid);
+		boutc(cl, "on unknown tid %llu\n", tid);
 		mutex_unlock(&mdsc->mutex);
+		ceph_blog_exit(&__ji);
 		return;
 	}
-	doutc(cl, "handle_reply %p\n", req);
+	boutc(cl, "handle_reply %p\n", req);
 
 	/* correct session? */
 	if (req->r_session != session) {
@@ -4184,7 +4229,7 @@ static void handle_reply(struct ceph_mds_session *session, struct ceph_msg *msg)
 			 * response.  And even if it did, there is nothing
 			 * useful we could do with a revised return value.
 			 */
-			doutc(cl, "got safe reply %llu, mds%d\n", tid, mds);
+			boutc(cl, "got safe reply %llu, mds%d\n", tid, mds);
 
 			mutex_unlock(&mdsc->mutex);
 			goto out;
@@ -4201,7 +4246,7 @@ static void handle_reply(struct ceph_mds_session *session, struct ceph_msg *msg)
 	 */
 	mutex_unlock(&mdsc->mutex);
 
-	doutc(cl, "tid %lld result %d\n", tid, result);
+	boutc(cl, "tid %lld result %d\n", tid, result);
 	if (test_bit(CEPHFS_FEATURE_REPLY_ENCODING, &session->s_features))
 		err = parse_reply_info(session, msg, req, (u64)-1);
 	else
@@ -4270,21 +4315,20 @@ static void handle_reply(struct ceph_mds_session *session, struct ceph_msg *msg)
 	/* insert trace into our cache */
 	mutex_lock(&req->r_fill_mutex);
 
-	/* disable fs reclaim while we are using current->journal_info
-	 * for our own purposes, or else shrinkers of other
-	 * filesystems might dereference this pointer as a different
-	 * type
-	 */
+	/* Disable filesystem reclaim while journal_info carries Ceph state. */
 	nofs_flags = memalloc_nofs_save();
-
-	current->journal_info = req;
-	err = ceph_fill_trace(mdsc->fsc->sb, req);
-	if (err == 0) {
-		if (result == 0 && (req->r_op == CEPH_MDS_OP_READDIR ||
-				    req->r_op == CEPH_MDS_OP_LSSNAP))
-			err = ceph_readdir_prepopulate(req, req->r_session);
+	{
+		struct ceph_journal_info ji;
+
+		ceph_blog_enter_req(mdsc->fsc, &ji, req);
+		err = ceph_fill_trace(mdsc->fsc->sb, req);
+		if (err == 0) {
+			if (result == 0 && (req->r_op == CEPH_MDS_OP_READDIR ||
+					    req->r_op == CEPH_MDS_OP_LSSNAP))
+				err = ceph_readdir_prepopulate(req, req->r_session);
+		}
+		ceph_blog_exit(&ji);
 	}
-	current->journal_info = NULL;
 	memalloc_nofs_restore(nofs_flags);
 	mutex_unlock(&req->r_fill_mutex);
 
@@ -4315,7 +4359,7 @@ static void handle_reply(struct ceph_mds_session *session, struct ceph_msg *msg)
 			set_bit(CEPH_MDS_R_GOT_RESULT, &req->r_req_flags);
 		}
 	} else {
-		doutc(cl, "reply arrived after request %lld was aborted\n", tid);
+		boutc(cl, "reply arrived after request %lld was aborted\n", tid);
 	}
 	mutex_unlock(&mdsc->mutex);
 
@@ -4332,6 +4376,7 @@ static void handle_reply(struct ceph_mds_session *session, struct ceph_msg *msg)
 	/* Defer closing the sessions after s_mutex lock being released */
 	if (close_sessions)
 		ceph_mdsc_close_sessions(mdsc);
+	ceph_blog_exit(&__ji);
 	return;
 }
 
@@ -4353,6 +4398,9 @@ static void handle_forward(struct ceph_mds_client *mdsc,
 	void *p = msg->front.iov_base;
 	void *end = p + msg->front.iov_len;
 	bool aborted = false;
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(mdsc->fsc, &__ji);
 
 	ceph_decode_need(&p, end, 2*sizeof(u32), bad);
 	next_mds = ceph_decode_32(&p);
@@ -4362,12 +4410,13 @@ static void handle_forward(struct ceph_mds_client *mdsc,
 	req = lookup_get_request(mdsc, tid);
 	if (!req) {
 		mutex_unlock(&mdsc->mutex);
-		doutc(cl, "forward tid %llu to mds%d - req dne\n", tid, next_mds);
+		boutc(cl, "forward tid %llu to mds%d - req dne\n", tid, next_mds);
+		ceph_blog_exit(&__ji);
 		return;  /* dup reply? */
 	}
 
 	if (test_bit(CEPH_MDS_R_ABORTED, &req->r_req_flags)) {
-		doutc(cl, "forward tid %llu aborted, unregistering\n", tid);
+		boutc(cl, "forward tid %llu aborted, unregistering\n", tid);
 		__unregister_request(mdsc, req);
 	} else if (fwd_seq <= req->r_num_fwd || (uint32_t)fwd_seq >= U32_MAX) {
 		/*
@@ -4387,7 +4436,7 @@ static void handle_forward(struct ceph_mds_client *mdsc,
 					   tid);
 	} else {
 		/* resend. forward race not possible; mds would drop */
-		doutc(cl, "forward tid %llu to mds%d (we resend)\n", tid, next_mds);
+		boutc(cl, "forward tid %llu to mds%d (we resend)\n", tid, next_mds);
 		BUG_ON(req->r_err);
 		BUG_ON(test_bit(CEPH_MDS_R_GOT_RESULT, &req->r_req_flags));
 		req->r_attempts = 0;
@@ -4402,11 +4451,13 @@ static void handle_forward(struct ceph_mds_client *mdsc,
 	if (aborted)
 		complete_request(mdsc, req);
 	ceph_mdsc_put_request(req);
+	ceph_blog_exit(&__ji);
 	return;
 
 bad:
 	pr_err_client(cl, "decode error err=%d\n", err);
 	ceph_msg_dump(msg);
+	ceph_blog_exit(&__ji);
 }
 
 static int __decode_session_metadata(void **p, void *end,
@@ -4456,7 +4507,9 @@ static void handle_session(struct ceph_mds_session *session,
 	int wake = 0;
 	bool blocklisted = false;
 	u32 i;
+	struct ceph_journal_info __ji;
 
+	ceph_blog_enter(mdsc->fsc, &__ji);
 
 	/* decode */
 	ceph_decode_need(&p, end, sizeof(*h), bad);
@@ -4503,7 +4556,7 @@ static void handle_session(struct ceph_mds_session *session,
 
 	if (msg_version >= 6) {
 		ceph_decode_32_safe(&p, end, cap_auths_num, bad);
-		doutc(cl, "cap_auths_num %d\n", cap_auths_num);
+		boutc(cl, "cap_auths_num %d\n", cap_auths_num);
 
 		if (cap_auths_num && op != CEPH_SESSION_OPEN) {
 			WARN_ON_ONCE(op != CEPH_SESSION_OPEN);
@@ -4514,6 +4567,7 @@ static void handle_session(struct ceph_mds_session *session,
 					 cap_auths_num);
 		if (!cap_auths) {
 			pr_err_client(cl, "No memory for cap_auths\n");
+			ceph_blog_exit(&__ji);
 			return;
 		}
 
@@ -4577,7 +4631,7 @@ static void handle_session(struct ceph_mds_session *session,
 			ceph_decode_8_safe(&p, end, cap_auths[i].match.root_squash, bad);
 			ceph_decode_8_safe(&p, end, cap_auths[i].readable, bad);
 			ceph_decode_8_safe(&p, end, cap_auths[i].writeable, bad);
-			doutc(cl, "uid %lld, num_gids %u, path %s, fs_name %s, root_squash %d, readable %d, writeable %d\n",
+			boutc(cl, "uid %lld, num_gids %u, path %s, fs_name %s, root_squash %d, readable %d, writeable %d\n",
 			      cap_auths[i].match.uid, cap_auths[i].match.num_gids,
 			      cap_auths[i].match.path, cap_auths[i].match.fs_name,
 			      cap_auths[i].match.root_squash,
@@ -4614,7 +4668,7 @@ static void handle_session(struct ceph_mds_session *session,
 
 	mutex_lock(&session->s_mutex);
 
-	doutc(cl, "mds%d %s %p state %s seq %llu\n", mds,
+	boutc(cl, "mds%d %s %p state %s seq %llu\n", mds,
 	      ceph_session_op_name(op), session,
 	      ceph_session_state_name(session->s_state), seq);
 
@@ -4697,7 +4751,7 @@ static void handle_session(struct ceph_mds_session *session,
 		break;
 
 	case CEPH_SESSION_FORCE_RO:
-		doutc(cl, "force_session_readonly %p\n", session);
+		boutc(cl, "force_session_readonly %p\n", session);
 		spin_lock(&session->s_cap_lock);
 		session->s_readonly = true;
 		spin_unlock(&session->s_cap_lock);
@@ -4736,6 +4790,7 @@ static void handle_session(struct ceph_mds_session *session,
 	}
 	if (op == CEPH_SESSION_CLOSE)
 		ceph_put_mds_session(session);
+	ceph_blog_exit(&__ji);
 	return;
 
 bad:
@@ -4749,6 +4804,7 @@ static void handle_session(struct ceph_mds_session *session,
 		kfree(cap_auths[i].match.fs_name);
 	}
 	kfree(cap_auths);
+	ceph_blog_exit(&__ji);
 	return;
 }
 
@@ -4759,7 +4815,7 @@ void ceph_mdsc_release_dir_caps(struct ceph_mds_request *req)
 
 	dcaps = xchg(&req->r_dir_caps, 0);
 	if (dcaps) {
-		doutc(cl, "releasing r_dir_caps=%s\n", ceph_cap_string(dcaps));
+		boutc(cl, "releasing r_dir_caps=%s\n", ceph_cap_string(dcaps));
 		ceph_put_cap_refs(ceph_inode(req->r_parent), dcaps);
 	}
 }
@@ -4771,7 +4827,7 @@ void ceph_mdsc_release_dir_caps_async(struct ceph_mds_request *req)
 
 	dcaps = xchg(&req->r_dir_caps, 0);
 	if (dcaps) {
-		doutc(cl, "releasing r_dir_caps=%s\n", ceph_cap_string(dcaps));
+		boutc(cl, "releasing r_dir_caps=%s\n", ceph_cap_string(dcaps));
 		ceph_put_cap_refs_async(ceph_inode(req->r_parent), dcaps);
 	}
 }
@@ -4785,7 +4841,7 @@ static void replay_unsafe_requests(struct ceph_mds_client *mdsc,
 	struct ceph_mds_request *req, *nreq;
 	struct rb_node *p;
 
-	doutc(mdsc->fsc->client, "mds%d\n", session->s_mds);
+	boutc(mdsc->fsc->client, "mds%d\n", session->s_mds);
 
 	mutex_lock(&mdsc->mutex);
 	list_for_each_entry_safe(req, nreq, &session->s_unsafe, r_unsafe_item)
@@ -4963,7 +5019,7 @@ static int reconnect_caps_cb(struct inode *inode, int mds, void *arg)
 		err = 0;
 		goto out_err;
 	}
-	doutc(cl, " adding %p ino %llx.%llx cap %p %lld %s\n", inode,
+	boutc(cl, " adding %p ino %llx.%llx cap %p %lld %s\n", inode,
 	      ceph_vinop(inode), cap, cap->cap_id,
 	      ceph_cap_string(cap->issued));
 
@@ -5167,7 +5223,7 @@ static int encode_snap_realms(struct ceph_mds_client *mdsc,
 			ceph_pagelist_encode_32(pagelist, sizeof(sr_rec));
 		}
 
-		doutc(cl, " adding snap realm %llx seq %lld parent %llx\n",
+		boutc(cl, " adding snap realm %llx seq %lld parent %llx\n",
 		      realm->ino, realm->seq, realm->parent_ino);
 		sr_rec.ino = cpu_to_le64(realm->ino);
 		sr_rec.seq = cpu_to_le64(realm->seq);
@@ -5250,7 +5306,7 @@ static int send_mds_reconnect(struct ceph_mds_client *mdsc,
 	session->s_state = CEPH_MDS_SESSION_RECONNECTING;
 	session->s_seq = 0;
 
-	doutc(cl, "session %p state %s\n", session,
+	boutc(cl, "session %p state %s\n", session,
 	      ceph_session_state_name(session->s_state));
 
 	atomic_inc(&session->s_cap_gen);
@@ -5902,7 +5958,7 @@ static void check_new_map(struct ceph_mds_client *mdsc,
 	unsigned long targets[DIV_ROUND_UP(CEPH_MAX_MDS, sizeof(unsigned long))] = {0};
 	struct ceph_client *cl = mdsc->fsc->client;
 
-	doutc(cl, "new %u old %u\n", newmap->m_epoch, oldmap->m_epoch);
+	boutc(cl, "new %u old %u\n", newmap->m_epoch, oldmap->m_epoch);
 
 	if (newmap->m_info) {
 		for (i = 0; i < newmap->possible_max_rank; i++) {
@@ -5918,7 +5974,7 @@ static void check_new_map(struct ceph_mds_client *mdsc,
 		oldstate = ceph_mdsmap_get_state(oldmap, i);
 		newstate = ceph_mdsmap_get_state(newmap, i);
 
-		doutc(cl, "mds%d state %s%s -> %s%s (session %s)\n",
+		boutc(cl, "mds%d state %s%s -> %s%s (session %s)\n",
 		      i, ceph_mds_state_name(oldstate),
 		      ceph_mdsmap_is_laggy(oldmap, i) ? " (laggy)" : "",
 		      ceph_mds_state_name(newstate),
@@ -6039,7 +6095,7 @@ static void check_new_map(struct ceph_mds_client *mdsc,
 				continue;
 			}
 		}
-		doutc(cl, "send reconnect to export target mds.%d\n", i);
+		boutc(cl, "send reconnect to export target mds.%d\n", i);
 		mutex_unlock(&mdsc->mutex);
 		err = send_mds_reconnect(mdsc, s);
 		if (err)
@@ -6059,7 +6115,7 @@ static void check_new_map(struct ceph_mds_client *mdsc,
 		if (s->s_state == CEPH_MDS_SESSION_OPEN ||
 		    s->s_state == CEPH_MDS_SESSION_HUNG ||
 		    s->s_state == CEPH_MDS_SESSION_CLOSING) {
-			doutc(cl, " connecting to export targets of laggy mds%d\n", i);
+			boutc(cl, " connecting to export targets of laggy mds%d\n", i);
 			__open_export_target_sessions(mdsc, s);
 		}
 	}
@@ -6098,7 +6154,7 @@ static void handle_lease(struct ceph_mds_client *mdsc,
 	struct qstr dname;
 	int release = 0;
 
-	doutc(cl, "from mds%d\n", mds);
+	boutc(cl, "from mds%d\n", mds);
 
 	if (!ceph_inc_mds_stopping_blocker(mdsc, session))
 		return;
@@ -6116,19 +6172,22 @@ static void handle_lease(struct ceph_mds_client *mdsc,
 
 	/* lookup inode */
 	inode = ceph_find_inode(sb, vino);
-	doutc(cl, "%s, ino %llx %p %.*s\n", ceph_lease_op_name(h->action),
-	      vino.ino, inode, dname.len, dname.name);
+	boutc_bounded(cl, "%s, ino %llx %p %.*s\n",
+		      (ceph_lease_op_name(h->action), vino.ino, inode, dname.len,
+		       BLOG_STR(dname.name, dname.len)),
+		      (ceph_lease_op_name(h->action), vino.ino, inode, dname.len,
+		       (const char *)dname.name));
 
 	mutex_lock(&session->s_mutex);
 	if (!inode) {
-		doutc(cl, "no inode %llx\n", vino.ino);
+		boutc(cl, "no inode %llx\n", vino.ino);
 		goto release;
 	}
 
 	/* dentry */
 	parent = d_find_alias(inode);
 	if (!parent) {
-		doutc(cl, "no parent dentry on inode %p\n", inode);
+		boutc(cl, "no parent dentry on inode %p\n", inode);
 		WARN_ON(1);
 		goto release;  /* hrm... */
 	}
@@ -6202,7 +6261,7 @@ void ceph_mdsc_lease_send_msg(struct ceph_mds_session *session,
 	struct inode *dir;
 	int len = sizeof(*lease) + sizeof(u32) + NAME_MAX;
 
-	doutc(cl, "identry %p %s to mds%d\n", dentry, ceph_lease_op_name(action),
+	boutc(cl, "identry %p %s to mds%d\n", dentry, ceph_lease_op_name(action),
 	      session->s_mds);
 
 	msg = ceph_msg_new(CEPH_MSG_CLIENT_LEASE, len, GFP_NOFS, false);
@@ -6289,7 +6348,7 @@ void inc_session_sequence(struct ceph_mds_session *s)
 	if (s->s_state == CEPH_MDS_SESSION_CLOSING) {
 		int ret;
 
-		doutc(cl, "resending session close request for mds%d\n", s->s_mds);
+		boutc(cl, "resending session close request for mds%d\n", s->s_mds);
 		ret = request_close_session(s);
 		if (ret < 0)
 			pr_err_client(cl, "unable to close session to mds%d: %d\n",
@@ -6317,15 +6376,20 @@ static void delayed_work(struct work_struct *work)
 {
 	struct ceph_mds_client *mdsc =
 		container_of(work, struct ceph_mds_client, delayed_work.work);
+	struct ceph_journal_info __ji;
 	unsigned long delay;
 	int renew_interval;
 	int renew_caps;
 	int i;
 
-	doutc(mdsc->fsc->client, "mdsc delayed_work\n");
+	ceph_blog_enter(mdsc->fsc, &__ji);
 
-	if (mdsc->stopping >= CEPH_MDSC_STOPPING_FLUSHED)
+	boutc(mdsc->fsc->client, "mdsc delayed_work\n");
+
+	if (mdsc->stopping >= CEPH_MDSC_STOPPING_FLUSHED) {
+		ceph_blog_exit(&__ji);
 		return;
+	}
 
 	mutex_lock(&mdsc->mutex);
 	renew_interval = mdsc->mdsmap->m_session_timeout >> 2;
@@ -6371,6 +6435,7 @@ static void delayed_work(struct work_struct *work)
 	maybe_recover_session(mdsc);
 
 	schedule_delayed(mdsc, delay);
+	ceph_blog_exit(&__ji);
 }
 
 int ceph_mdsc_init(struct ceph_fs_client *fsc)
@@ -6478,20 +6543,20 @@ static void wait_requests(struct ceph_mds_client *mdsc)
 	if (__get_oldest_req(mdsc)) {
 		mutex_unlock(&mdsc->mutex);
 
-		doutc(cl, "waiting for requests\n");
+		boutc(cl, "waiting for requests\n");
 		wait_for_completion_timeout(&mdsc->safe_umount_waiters,
 				    ceph_timeout_jiffies(opts->mount_timeout));
 
 		/* tear down remaining requests */
 		mutex_lock(&mdsc->mutex);
 		while ((req = __get_oldest_req(mdsc))) {
-			doutc(cl, "timed out on tid %llu\n", req->r_tid);
+			boutc(cl, "timed out on tid %llu\n", req->r_tid);
 			list_del_init(&req->r_wait);
 			__unregister_request(mdsc, req);
 		}
 	}
 	mutex_unlock(&mdsc->mutex);
-	doutc(cl, "done\n");
+	boutc(cl, "done\n");
 }
 
 void send_flush_mdlog(struct ceph_mds_session *s)
@@ -6506,7 +6571,7 @@ void send_flush_mdlog(struct ceph_mds_session *s)
 		return;
 
 	mutex_lock(&s->s_mutex);
-	doutc(cl, "request mdlog flush to mds%d (%s)s seq %lld\n",
+	boutc(cl, "request mdlog flush to mds%d (%s)s seq %lld\n",
 	      s->s_mds, ceph_session_state_name(s->s_state), s->s_seq);
 	msg = ceph_create_session_msg(CEPH_SESSION_REQUEST_FLUSH_MDLOG,
 				      s->s_seq);
@@ -6581,7 +6646,7 @@ static int ceph_mds_auth_match(struct ceph_mds_client *mdsc,
 			bool free_tpath = false;
 			int m, n;
 
-			doutc(cl, "server path %s, tpath %s, match.path %s\n",
+			boutc(cl, "server path %s, tpath %s, match.path %s\n",
 			      spath, tpath, auth->match.path);
 			if (spath && (m = strlen(spath)) != 1) {
 				/* mount path + '/' + tpath + an extra space */
@@ -6605,7 +6670,7 @@ static int ceph_mds_auth_match(struct ceph_mds_client *mdsc,
 				_tpath[tlen - 1] = '\0';
 				tlen -= 1;
 			}
-			doutc(cl, "_tpath %s\n", _tpath);
+			boutc(cl, "_tpath %s\n", _tpath);
 
 			/*
 			 * In case first == _tpath && tlen == len:
@@ -6635,7 +6700,7 @@ static int ceph_mds_auth_match(struct ceph_mds_client *mdsc,
 		}
 	}
 
-	doutc(cl, "matched\n");
+	boutc(cl, "matched\n");
 	return 1;
 }
 
@@ -6649,7 +6714,7 @@ int ceph_mds_check_access(struct ceph_mds_client *mdsc, char *tpath, int mask)
 	bool root_squash_perms = true;
 	int i, err;
 
-	doutc(cl, "tpath '%s', mask %d, caller_uid %d, caller_gid %d\n",
+	boutc(cl, "tpath '%s', mask %d, caller_uid %d, caller_gid %d\n",
 	      tpath, mask, caller_uid, caller_gid);
 
 	mutex_lock(&mdsc->mutex);
@@ -6677,24 +6742,24 @@ int ceph_mds_check_access(struct ceph_mds_client *mdsc, char *tpath, int mask)
 
 	put_cred(cred);
 
-	doutc(cl, "root_squash_perms %d, rw_perms_s %p\n", root_squash_perms,
+	boutc(cl, "root_squash_perms %d, rw_perms_s %p\n", root_squash_perms,
 	      rw_perms_s);
 	if (root_squash_perms && rw_perms_s == NULL) {
 		mutex_unlock(&mdsc->mutex);
-		doutc(cl, "access allowed\n");
+		boutc(cl, "access allowed\n");
 		return 0;
 	}
 
 	if (!root_squash_perms) {
-		doutc(cl, "root_squash is enabled and user(%d %d) isn't allowed to write",
+		boutc(cl, "root_squash is enabled and user(%d %d) isn't allowed to write",
 		      caller_uid, caller_gid);
 	}
 	if (rw_perms_s) {
-		doutc(cl, "mds auth caps readable/writeable %d/%d while request r/w %d/%d",
+		boutc(cl, "mds auth caps readable/writeable %d/%d while request r/w %d/%d",
 		      rw_perms_s->readable, rw_perms_s->writeable,
 		      !!(mask & MAY_READ), !!(mask & MAY_WRITE));
 	}
-	doutc(cl, "access denied\n");
+	boutc(cl, "access denied\n");
 	mutex_unlock(&mdsc->mutex);
 	return -EACCES;
 }
@@ -6705,7 +6770,7 @@ int ceph_mds_check_access(struct ceph_mds_client *mdsc, char *tpath, int mask)
  */
 void ceph_mdsc_pre_umount(struct ceph_mds_client *mdsc)
 {
-	doutc(mdsc->fsc->client, "begin\n");
+	boutc(mdsc->fsc->client, "begin\n");
 	mdsc->stopping = CEPH_MDSC_STOPPING_BEGIN;
 
 	ceph_mdsc_iterate_sessions(mdsc, send_flush_mdlog, true);
@@ -6720,7 +6785,7 @@ void ceph_mdsc_pre_umount(struct ceph_mds_client *mdsc)
 	ceph_msgr_flush();
 
 	ceph_cleanup_quotarealms_inodes(mdsc);
-	doutc(mdsc->fsc->client, "done\n");
+	boutc(mdsc->fsc->client, "done\n");
 }
 
 /*
@@ -6735,7 +6800,7 @@ static void flush_mdlog_and_wait_mdsc_unsafe_requests(struct ceph_mds_client *md
 	struct rb_node *n;
 
 	mutex_lock(&mdsc->mutex);
-	doutc(cl, "want %lld\n", want_tid);
+	boutc(cl, "want %lld\n", want_tid);
 restart:
 	req = __get_oldest_req(mdsc);
 	while (req && req->r_tid <= want_tid) {
@@ -6769,7 +6834,7 @@ static void flush_mdlog_and_wait_mdsc_unsafe_requests(struct ceph_mds_client *md
 			} else {
 				ceph_put_mds_session(s);
 			}
-			doutc(cl, "wait on %llu (want %llu)\n",
+			boutc(cl, "wait on %llu (want %llu)\n",
 			      req->r_tid, want_tid);
 			wait_for_completion(&req->r_safe_completion);
 
@@ -6788,7 +6853,7 @@ static void flush_mdlog_and_wait_mdsc_unsafe_requests(struct ceph_mds_client *md
 	}
 	mutex_unlock(&mdsc->mutex);
 	ceph_put_mds_session(last_session);
-	doutc(cl, "done\n");
+	boutc(cl, "done\n");
 }
 
 void ceph_mdsc_sync(struct ceph_mds_client *mdsc)
@@ -6799,7 +6864,7 @@ void ceph_mdsc_sync(struct ceph_mds_client *mdsc)
 	if (READ_ONCE(mdsc->fsc->mount_state) >= CEPH_MOUNT_SHUTDOWN)
 		return;
 
-	doutc(cl, "sync\n");
+	boutc(cl, "sync\n");
 	mutex_lock(&mdsc->mutex);
 	want_tid = mdsc->last_tid;
 	mutex_unlock(&mdsc->mutex);
@@ -6816,7 +6881,7 @@ void ceph_mdsc_sync(struct ceph_mds_client *mdsc)
 	}
 	spin_unlock(&mdsc->cap_dirty_lock);
 
-	doutc(cl, "sync want tid %lld flush_seq %lld\n", want_tid, want_flush);
+	boutc(cl, "sync want tid %lld flush_seq %lld\n", want_tid, want_flush);
 
 	flush_mdlog_and_wait_mdsc_unsafe_requests(mdsc, want_tid);
 	wait_caps_flush(mdsc, want_flush);
@@ -6843,7 +6908,7 @@ void ceph_mdsc_close_sessions(struct ceph_mds_client *mdsc)
 	int i;
 	int skipped = 0;
 
-	doutc(cl, "begin\n");
+	boutc(cl, "begin\n");
 
 	/* close sessions */
 	mutex_lock(&mdsc->mutex);
@@ -6861,7 +6926,7 @@ void ceph_mdsc_close_sessions(struct ceph_mds_client *mdsc)
 	}
 	mutex_unlock(&mdsc->mutex);
 
-	doutc(cl, "waiting for sessions to close\n");
+	boutc(cl, "waiting for sessions to close\n");
 	wait_event_timeout(mdsc->session_close_wq,
 			   done_closing_sessions(mdsc, skipped),
 			   ceph_timeout_jiffies(opts->mount_timeout));
@@ -6890,7 +6955,7 @@ void ceph_mdsc_close_sessions(struct ceph_mds_client *mdsc)
 	cancel_work_sync(&mdsc->cap_unlink_work);
 	cancel_delayed_work_sync(&mdsc->delayed_work); /* cancel timer */
 
-	doutc(cl, "done\n");
+	boutc(cl, "done\n");
 }
 
 void ceph_mdsc_force_umount(struct ceph_mds_client *mdsc)
@@ -6898,7 +6963,7 @@ void ceph_mdsc_force_umount(struct ceph_mds_client *mdsc)
 	struct ceph_mds_session *session;
 	int mds;
 
-	doutc(mdsc->fsc->client, "force umount\n");
+	boutc(mdsc->fsc->client, "force umount\n");
 
 	mutex_lock(&mdsc->mutex);
 	for (mds = 0; mds < mdsc->max_sessions; mds++) {
@@ -6929,7 +6994,7 @@ void ceph_mdsc_force_umount(struct ceph_mds_client *mdsc)
 
 static void ceph_mdsc_stop(struct ceph_mds_client *mdsc)
 {
-	doutc(mdsc->fsc->client, "stop\n");
+	boutc(mdsc->fsc->client, "stop\n");
 	/*
 	 * Make sure the delayed work stopped before releasing
 	 * the resources.
@@ -6962,7 +7027,7 @@ static void ceph_mdsc_stop(struct ceph_mds_client *mdsc)
 void ceph_mdsc_destroy(struct ceph_fs_client *fsc)
 {
 	struct ceph_mds_client *mdsc = fsc->mdsc;
-	doutc(fsc->client, "%p\n", mdsc);
+	boutc(fsc->client, "%p\n", mdsc);
 
 	if (!mdsc)
 		return;
@@ -6995,7 +7060,7 @@ void ceph_mdsc_destroy(struct ceph_fs_client *fsc)
 
 	fsc->mdsc = NULL;
 	kfree(mdsc);
-	doutc(fsc->client, "%p done\n", mdsc);
+	boutc(fsc->client, "%p done\n", mdsc);
 }
 
 void ceph_mdsc_handle_fsmap(struct ceph_mds_client *mdsc, struct ceph_msg *msg)
@@ -7013,7 +7078,7 @@ void ceph_mdsc_handle_fsmap(struct ceph_mds_client *mdsc, struct ceph_msg *msg)
 	ceph_decode_need(&p, end, sizeof(u32), bad);
 	epoch = ceph_decode_32(&p);
 
-	doutc(cl, "epoch %u\n", epoch);
+	boutc(cl, "epoch %u\n", epoch);
 
 	/* struct_v, struct_cv, map_len, epoch, legacy_client_fscid */
 	ceph_decode_skip_n(&p, end, 2 + sizeof(u32) * 3, bad);
@@ -7089,12 +7154,12 @@ void ceph_mdsc_handle_mdsmap(struct ceph_mds_client *mdsc, struct ceph_msg *msg)
 		return;
 	epoch = ceph_decode_32(&p);
 	maplen = ceph_decode_32(&p);
-	doutc(cl, "epoch %u len %d\n", epoch, (int)maplen);
+	boutc(cl, "epoch %u len %d\n", epoch, (int)maplen);
 
 	/* do we need it? */
 	mutex_lock(&mdsc->mutex);
 	if (mdsc->mdsmap && epoch <= mdsc->mdsmap->m_epoch) {
-		doutc(cl, "epoch %u <= our %u\n", epoch, mdsc->mdsmap->m_epoch);
+		boutc(cl, "epoch %u <= our %u\n", epoch, mdsc->mdsmap->m_epoch);
 		mutex_unlock(&mdsc->mutex);
 		return;
 	}
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 16+ messages in thread

* [PATCH v7 10/14] ceph: add BLOG debugfs interface
  2026-09-24 15:30 [PATCH v7 00/14] ceph: add binary logging (BLOG) for CephFS Alex Markuze
                   ` (8 preceding siblings ...)
  2026-09-24 15:30 ` [PATCH v7 09/14] ceph: switch MDS request plumbing to struct ceph_journal_info Alex Markuze
@ 2026-09-24 15:30 ` Alex Markuze
  2026-09-24 15:30 ` [PATCH v7 11/14] ceph: convert VFS inode and directory paths to BLOG logging Alex Markuze
                   ` (4 subsequent siblings)
  14 siblings, 0 replies; 16+ messages in thread
From: Alex Markuze @ 2026-09-24 15:30 UTC (permalink / raw)
  To: ceph-devel; +Cc: idryomov, xiubo.li

Add blog_debugfs.c: per-superblock debugfs files under
/sys/kernel/debug/ceph/<fsid>/blog/ (entries, stats, sources,
clients, enabled, clear).  Wire lifecycle hooks in super.c and
debugfs.c.

Document blog_entries_show() parameters.

Make all BLOG data files root-readable, matching the other Ceph
debugfs data files. Entries can contain filenames and xattr
values, even when the debugfs root is traversable by other users.

Bound each entries dump by a context-ID ceiling, retaining it in
file-private state across seq_file overflow retries. Stop both
formatting loops on overflow and reset the ceiling after a
successful dump. Continuous rotation must not extend a read
indefinitely. Copy the u64 record base under the page-fragment
lock along with the published bytes.

Signed-off-by: Alex Markuze <amarkuze@redhat.com>
Assisted-by: LLM
---
 fs/ceph/blog_debugfs.c | 702 +++++++++++++++++++++++++++++++++++++++++
 fs/ceph/debugfs.c      |  11 +-
 2 files changed, 710 insertions(+), 3 deletions(-)
 create mode 100644 fs/ceph/blog_debugfs.c

diff --git a/fs/ceph/blog_debugfs.c b/fs/ceph/blog_debugfs.c
new file mode 100644
index 000000000000..517d2bae3823
--- /dev/null
+++ b/fs/ceph/blog_debugfs.c
@@ -0,0 +1,702 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Ceph BLOG debugfs.
+ */
+
+#include <linux/ceph/ceph_debug.h>
+#include <linux/module.h>
+#include <linux/debugfs.h>
+#include <linux/seq_file.h>
+#include <linux/slab.h>
+#include <linux/string.h>
+#include <linux/jiffies.h>
+#include <linux/timekeeping.h>
+#include <linux/ceph/ceph_blog.h>
+#include "blog.h"
+#include "blog_des.h"
+#include "blog_module.h"
+
+#include "super.h"
+
+static int jiffies_to_formatted_time(u64 jiffies_value, char *buffer,
+	size_t buffer_len);
+
+struct blog_dbg_file {
+	struct ceph_fs_client *fsc;
+	struct blog_module_context *ctx;
+	u64 entries_end_id;
+	bool entries_end_valid;
+};
+
+static int blog_dbg_pin(struct ceph_fs_client *fsc,
+			struct blog_dbg_file *priv)
+{
+	struct blog_module_context *ctx;
+
+	if (!fsc)
+		return -ENODEV;
+
+	/*
+	 * Module + blog_ctx refs only.  No open-count wait on fsc:
+	 * debugfs_create_file() already get/puts around read/write, so
+	 * remove waits for in-flight ops; release must not touch fsc.
+	 */
+	if (!try_module_get(THIS_MODULE))
+		return -ENODEV;
+
+	rcu_read_lock();
+	ctx = rcu_dereference(fsc->blog_ctx);
+	if (ctx && !refcount_inc_not_zero(&ctx->refcount))
+		ctx = NULL;
+	rcu_read_unlock();
+
+	priv->fsc = fsc;
+	priv->ctx = ctx;
+	return 0;
+}
+
+static void blog_dbg_unpin(struct blog_dbg_file *priv)
+{
+	if (priv->ctx)
+		blog_module_put(priv->ctx);
+	module_put(THIS_MODULE);
+}
+
+static struct blog_logger *blog_dbg_logger(struct blog_dbg_file *priv)
+{
+	if (!priv || !priv->ctx)
+		return NULL;
+	return READ_ONCE(priv->ctx->logger);
+}
+
+/**
+ * blog_entries_show - dump decoded records
+ * @s: seq_file to write
+ * @p: unused iterator cookie
+ *
+ * ID-cursor walk: each pass finds the next context by ascending ID, drops
+ * logger->lock, and snapshots only its published bytes under pf->lock.
+ * snapshot_mutex keeps retained contexts from being reclaimed between the
+ * list lookup and the snapshot.  Keep one upper ID bound across seq_file
+ * overflow retries so concurrent rotation cannot extend the dump forever.
+ */
+static int blog_entries_show(struct seq_file *s, void *p)
+{
+	struct blog_dbg_file *priv = s->private;
+	struct blog_logger *logger;
+	struct blog_tls_ctx *ctx;
+	void *buf_copy;
+	char *output_buf;
+	int entry_count = 0;
+	u64 cursor_id = 0;
+
+	logger = blog_dbg_logger(priv);
+	if (!logger) {
+		seq_puts(s, "Ceph BLOG context not initialized\n");
+		return 0;
+	}
+
+	buf_copy = kmalloc(BLOG_TLS_PAGEFRAG_BUFFER_SIZE, GFP_KERNEL);
+	if (!buf_copy)
+		return -ENOMEM;
+	/* Match live buffer capacity so long reconstructed lines are not cut. */
+	output_buf = kmalloc(BLOG_TLS_PAGEFRAG_BUFFER_SIZE, GFP_KERNEL);
+	if (!output_buf) {
+		kfree(buf_copy);
+		return -ENOMEM;
+	}
+
+	if (!priv->entries_end_valid) {
+		spin_lock(&logger->ctx_id_lock);
+		priv->entries_end_id = logger->next_ctx_id - 1;
+		spin_unlock(&logger->ctx_id_lock);
+		priv->entries_end_valid = true;
+	}
+
+	while (!seq_has_overflowed(s)) {
+		struct blog_tls_ctx *best = NULL;
+		struct blog_pagefrag *pf;
+		u64 base_jiffies;
+		struct blog_pagefrag tmp_pf;
+		struct blog_log_iter iter;
+		struct blog_log_entry *entry;
+		u64 best_id = U64_MAX;
+		u64 head;
+		int ret;
+
+		mutex_lock(&logger->snapshot_mutex);
+		spin_lock(&logger->lock);
+		list_for_each_entry(ctx, &logger->contexts, list) {
+			if (ctx->id > cursor_id &&
+			    ctx->id <= priv->entries_end_id && ctx->id < best_id) {
+				best = ctx;
+				best_id = ctx->id;
+			}
+		}
+
+		if (!best) {
+			spin_unlock(&logger->lock);
+			mutex_unlock(&logger->snapshot_mutex);
+			break;
+		}
+
+		pf = blog_ctx_pf(best);
+		/* Idle contexts lagging a clear look empty immediately. */
+		if (atomic64_read(&best->clear_seq) !=
+		    atomic64_read(&logger->clear_seq)) {
+			spin_unlock(&logger->lock);
+			cursor_id = best_id;
+			mutex_unlock(&logger->snapshot_mutex);
+			continue;
+		}
+		spin_unlock(&logger->lock);
+
+		/*
+		 * Snapshot published prefix under pf->lock.  Buffer is SZ_4K,
+		 * so holding the lock across the copy is cheaper and safer
+		 * than an unbounded head/head2 retry loop.
+		 */
+		spin_lock(&pf->lock);
+		head = smp_load_acquire(&pf->head);
+		base_jiffies = READ_ONCE(best->base_jiffies);
+		if (head)
+			memcpy(buf_copy, pf->buffer, head);
+		spin_unlock(&pf->lock);
+
+		/*
+		 * Rotate publishes the snapshot (id swap + contexts insert)
+		 * before clearing live and without snapshot_mutex.  After
+		 * we dropped logger->lock it can move best_id onto a
+		 * snapshot and give this ctx a new id.  Any id mismatch
+		 * means this copy is not the retired window; leave the
+		 * cursor so the next walk finds the snapshot.  Do not
+		 * advance just because head is non-zero. That may be
+		 * the *new* live tail.
+		 */
+		spin_lock(&logger->lock);
+		if (best->id != best_id) {
+			spin_unlock(&logger->lock);
+			mutex_unlock(&logger->snapshot_mutex);
+			continue;
+		}
+		cursor_id = best_id;
+		spin_unlock(&logger->lock);
+		mutex_unlock(&logger->snapshot_mutex);
+
+		if (!head)
+			continue;
+
+		/* Deserialize and output outside any lock */
+		memset(&tmp_pf, 0, sizeof(tmp_pf));
+		tmp_pf.buffer = buf_copy;
+		tmp_pf.capacity = head;
+		tmp_pf.head = head;
+
+		blog_log_iter_init(&iter, &tmp_pf, head);
+
+		while (!seq_has_overflowed(s) &&
+		       (entry = blog_log_iter_next(&iter)) != NULL) {
+			char time_buf[64];
+			u64 entry_jiffies;
+
+			entry_count++;
+			memset(output_buf, 0, BLOG_TLS_PAGEFRAG_BUFFER_SIZE);
+			ret = blog_des_entry(logger, entry,
+					     output_buf,
+					     BLOG_TLS_PAGEFRAG_BUFFER_SIZE,
+					     ceph_blog_client_des_callback);
+			if (ret < 0) {
+				seq_printf(s,
+					   "[Error deserializing entry %d: %d]\n",
+					   entry_count, ret);
+				continue;
+			}
+			entry_jiffies = base_jiffies + entry->ts_delta;
+			if (jiffies_to_formatted_time(entry_jiffies, time_buf,
+						      sizeof(time_buf)) < 0)
+				strscpy(time_buf, "(invalid)", sizeof(time_buf));
+			if (ret > 0 && output_buf[ret - 1] == '\n')
+				output_buf[ret - 1] = '\0';
+			seq_printf(s, "%s %s\n", time_buf, output_buf);
+		}
+	}
+
+	if (!seq_has_overflowed(s))
+		priv->entries_end_valid = false;
+	kfree(output_buf);
+	kfree(buf_copy);
+	return 0;
+}
+
+static int blog_entries_open(struct inode *inode, struct file *file)
+{
+	struct blog_dbg_file *priv;
+	int ret;
+
+	priv = kzalloc(sizeof(*priv), GFP_KERNEL);
+	if (!priv)
+		return -ENOMEM;
+
+	ret = blog_dbg_pin(inode->i_private, priv);
+	if (ret) {
+		kfree(priv);
+		return ret;
+	}
+
+	ret = single_open(file, blog_entries_show, priv);
+	if (ret) {
+		blog_dbg_unpin(priv);
+		kfree(priv);
+	}
+	return ret;
+}
+
+static int blog_dbg_release(struct inode *inode, struct file *file)
+{
+	struct seq_file *seq = file->private_data;
+	struct blog_dbg_file *priv = seq ? seq->private : NULL;
+	int ret = single_release(inode, file);
+
+	if (priv) {
+		blog_dbg_unpin(priv);
+		kfree(priv);
+	}
+	return ret;
+}
+
+static const struct file_operations blog_entries_fops = {
+	.owner = THIS_MODULE,
+	.open = blog_entries_open,
+	.read = seq_read,
+	.llseek = seq_lseek,
+	.release = blog_dbg_release,
+};
+
+static int blog_stats_show(struct seq_file *s, void *p)
+{
+	struct blog_dbg_file *priv = s->private;
+	struct blog_logger *logger = blog_dbg_logger(priv);
+
+	seq_puts(s, "Ceph BLOG Statistics\n");
+	seq_puts(s, "====================\n\n");
+
+	if (!logger) {
+		seq_puts(s, "Ceph BLOG context not initialized\n");
+		return 0;
+	}
+
+	seq_puts(s, "Ceph Module Logger State:\n");
+	seq_printf(s, "  Total contexts allocated: %lu\n",
+		   logger->total_contexts_allocated);
+	seq_printf(s, "  Next context ID: %llu\n",
+		   READ_ONCE(logger->next_ctx_id));
+	seq_printf(s, "  Next source ID: %u\n",
+		   READ_ONCE(logger->next_source_id));
+
+	seq_puts(s, "\nAllocation Batch:\n");
+	seq_printf(s, "  Full magazines: %u\n",
+		   READ_ONCE(logger->alloc_batch.nr_full));
+	seq_printf(s, "  Empty magazines: %u\n",
+		   READ_ONCE(logger->alloc_batch.nr_empty));
+
+	seq_puts(s, "\nLog Batch:\n");
+	seq_printf(s, "  Full magazines: %u\n",
+		   READ_ONCE(logger->log_batch.nr_full));
+	seq_printf(s, "  Empty magazines: %u\n",
+		   READ_ONCE(logger->log_batch.nr_empty));
+
+	return 0;
+}
+
+static int blog_stats_open(struct inode *inode, struct file *file)
+{
+	struct blog_dbg_file *priv;
+	int ret;
+
+	priv = kzalloc(sizeof(*priv), GFP_KERNEL);
+	if (!priv)
+		return -ENOMEM;
+
+	ret = blog_dbg_pin(inode->i_private, priv);
+	if (ret) {
+		kfree(priv);
+		return ret;
+	}
+
+	ret = single_open(file, blog_stats_show, priv);
+	if (ret) {
+		blog_dbg_unpin(priv);
+		kfree(priv);
+	}
+	return ret;
+}
+
+static const struct file_operations blog_stats_fops = {
+	.owner = THIS_MODULE,
+	.open = blog_stats_open,
+	.read = seq_read,
+	.llseek = seq_lseek,
+	.release = blog_dbg_release,
+};
+
+static int blog_sources_show(struct seq_file *s, void *p)
+{
+	struct blog_dbg_file *priv = s->private;
+	struct blog_logger *logger = blog_dbg_logger(priv);
+	struct blog_source_info *source;
+	const char *file, *func, *fmt;
+	unsigned int line;
+	int warn_count;
+	u32 id;
+	int count = 0;
+
+	seq_puts(s, "Ceph BLOG Source Locations\n");
+	seq_puts(s, "===========================\n\n");
+
+	if (!logger) {
+		seq_puts(s, "Ceph BLOG context not initialized\n");
+		return 0;
+	}
+
+	for (id = 1; id < logger->max_source_ids; id++) {
+		source = blog_get_source_info(logger, id);
+		if (!source)
+			continue;
+
+		spin_lock(&logger->source_lock);
+		file = source->file;
+		func = source->func;
+		line = source->line;
+		fmt = source->fmt;
+		warn_count = source->warn_count;
+		spin_unlock(&logger->source_lock);
+		if (!file)
+			continue;
+
+		count++;
+		seq_printf(s, "ID %u: %s:%s:%u\n", id, file, func, line);
+		seq_printf(s, "  Format: %s\n", fmt ? fmt : "(null)");
+		seq_printf(s, "  Warnings: %d\n", warn_count);
+
+		seq_puts(s, "\n");
+	}
+
+	seq_printf(s, "Total registered sources: %d\n", count);
+
+	return 0;
+}
+
+static int blog_sources_open(struct inode *inode, struct file *file)
+{
+	struct blog_dbg_file *priv;
+	int ret;
+
+	priv = kzalloc(sizeof(*priv), GFP_KERNEL);
+	if (!priv)
+		return -ENOMEM;
+
+	ret = blog_dbg_pin(inode->i_private, priv);
+	if (ret) {
+		kfree(priv);
+		return ret;
+	}
+
+	ret = single_open(file, blog_sources_show, priv);
+	if (ret) {
+		blog_dbg_unpin(priv);
+		kfree(priv);
+	}
+	return ret;
+}
+
+static const struct file_operations blog_sources_fops = {
+	.owner = THIS_MODULE,
+	.open = blog_sources_open,
+	.read = seq_read,
+	.llseek = seq_lseek,
+	.release = blog_dbg_release,
+};
+
+static int blog_clients_show(struct seq_file *s, void *p)
+{
+	struct blog_dbg_file *priv = s->private;
+	struct ceph_fs_client *fsc = priv->fsc;
+	struct ceph_client *client;
+	u32 client_id;
+
+	seq_puts(s, "Ceph BLOG Mount Client\n");
+	seq_puts(s, "======================\n\n");
+
+	if (!fsc || !fsc->client) {
+		seq_puts(s, "client unavailable\n");
+		return 0;
+	}
+
+	client = fsc->client;
+	client_id = READ_ONCE(client->blog_client_id);
+
+	seq_printf(s, "FSID: %pU\n", &client->fsid);
+	if (client->monc.auth)
+		seq_printf(s, "Global ID: %llu\n", client->monc.auth->global_id);
+	else
+		seq_puts(s, "Global ID: (unavailable)\n");
+	if (client_id)
+		seq_printf(s, "Cached BLOG client ID: %u\n", client_id);
+	else
+		seq_puts(s, "Cached BLOG client ID: (unassigned)\n");
+	if (ceph_blog_is_enabled(fsc))
+		seq_puts(s, "BLOG enabled: yes\n");
+	else
+		seq_puts(s, "BLOG enabled: no\n");
+
+	return 0;
+}
+
+static int blog_clients_open(struct inode *inode, struct file *file)
+{
+	struct blog_dbg_file *priv;
+	int ret;
+
+	priv = kzalloc(sizeof(*priv), GFP_KERNEL);
+	if (!priv)
+		return -ENOMEM;
+
+	ret = blog_dbg_pin(inode->i_private, priv);
+	if (ret) {
+		kfree(priv);
+		return ret;
+	}
+
+	ret = single_open(file, blog_clients_show, priv);
+	if (ret) {
+		blog_dbg_unpin(priv);
+		kfree(priv);
+	}
+	return ret;
+}
+
+static const struct file_operations blog_clients_fops = {
+	.owner = THIS_MODULE,
+	.open = blog_clients_open,
+	.read = seq_read,
+	.llseek = seq_lseek,
+	.release = blog_dbg_release,
+};
+
+static ssize_t blog_clear_write(struct file *file, const char __user *buf,
+				size_t count, loff_t *ppos)
+{
+	struct blog_dbg_file *priv = file->private_data;
+	struct blog_logger *logger = blog_dbg_logger(priv);
+	char cmd[16];
+
+	if (count >= sizeof(cmd))
+		return -EINVAL;
+
+	if (copy_from_user(cmd, buf, count))
+		return -EFAULT;
+
+	cmd[count] = '\0';
+
+	/* Only accept exact "clear" (optional trailing newline) */
+	if (strncmp(cmd, "clear", 5) != 0 ||
+	    (cmd[5] != '\0' && cmd[5] != '\n'))
+		return -EINVAL;
+
+	/*
+	 * Bump clear_seq so readers treat lagging contexts as empty until
+	 * their next write resets the pagefrag.  Also set NEEDS_RESET so a
+	 * concurrent writer notices promptly.  Hold snapshot_mutex so a
+	 * concurrent entries snapshot cannot copy pre-clear data after
+	 * this write returns.
+	 */
+	if (logger) {
+		struct blog_tls_ctx *tls_ctx;
+
+		mutex_lock(&logger->snapshot_mutex);
+		spin_lock(&logger->lock);
+		atomic64_inc(&logger->clear_seq);
+		list_for_each_entry(tls_ctx, &logger->contexts, list)
+			set_bit(BLOG_CTX_NEEDS_RESET, &tls_ctx->flags);
+		spin_unlock(&logger->lock);
+		mutex_unlock(&logger->snapshot_mutex);
+		pr_debug("ceph: BLOG entries cleared via debugfs\n");
+	}
+
+	return count;
+}
+
+static int blog_clear_open(struct inode *inode, struct file *file)
+{
+	struct blog_dbg_file *priv;
+	int ret;
+
+	priv = kzalloc(sizeof(*priv), GFP_KERNEL);
+	if (!priv)
+		return -ENOMEM;
+
+	ret = blog_dbg_pin(inode->i_private, priv);
+	if (ret) {
+		kfree(priv);
+		return ret;
+	}
+
+	file->private_data = priv;
+	return 0;
+}
+
+static int blog_clear_release(struct inode *inode, struct file *file)
+{
+	struct blog_dbg_file *priv = file->private_data;
+
+	if (priv) {
+		blog_dbg_unpin(priv);
+		kfree(priv);
+	}
+	return 0;
+}
+
+static const struct file_operations blog_clear_fops = {
+	.owner = THIS_MODULE,
+	.open = blog_clear_open,
+	.write = blog_clear_write,
+	.release = blog_clear_release,
+	.llseek = noop_llseek,
+};
+
+static ssize_t blog_enabled_read(struct file *file, char __user *buf,
+				 size_t count, loff_t *ppos)
+{
+	struct blog_dbg_file *priv = file->private_data;
+	char tmp[32];
+	int len;
+
+	len = scnprintf(tmp, sizeof(tmp), "%llu\n",
+			(u64)READ_ONCE(priv->fsc->blog_enabled));
+	return simple_read_from_buffer(buf, count, ppos, tmp, len);
+}
+
+static ssize_t blog_enabled_write(struct file *file, const char __user *buf,
+				  size_t count, loff_t *ppos)
+{
+	struct blog_dbg_file *priv = file->private_data;
+	u64 val;
+	int ret;
+
+	ret = kstrtoull_from_user(buf, count, 0, &val);
+	if (ret)
+		return ret;
+	if (val > 1)
+		return -EINVAL;
+
+	ret = ceph_blog_set_enabled(priv->fsc, val);
+	return ret ? ret : count;
+}
+
+static int blog_enabled_open(struct inode *inode, struct file *file)
+{
+	struct blog_dbg_file *priv;
+	int ret;
+
+	priv = kzalloc(sizeof(*priv), GFP_KERNEL);
+	if (!priv)
+		return -ENOMEM;
+
+	ret = blog_dbg_pin(inode->i_private, priv);
+	if (ret) {
+		kfree(priv);
+		return ret;
+	}
+
+	file->private_data = priv;
+	return 0;
+}
+
+static int blog_enabled_release(struct inode *inode, struct file *file)
+{
+	struct blog_dbg_file *priv = file->private_data;
+
+	if (priv) {
+		blog_dbg_unpin(priv);
+		kfree(priv);
+	}
+	return 0;
+}
+
+static const struct file_operations blog_enabled_fops = {
+	.owner = THIS_MODULE,
+	.open = blog_enabled_open,
+	.release = blog_enabled_release,
+	.read = blog_enabled_read,
+	.write = blog_enabled_write,
+	.llseek = generic_file_llseek,
+};
+
+int ceph_blog_debugfs_init(struct ceph_fs_client *fsc)
+{
+	struct dentry *dir;
+
+	if (!fsc || !fsc->client || !fsc->client->debugfs_dir)
+		return -EINVAL;
+	if (fsc->debugfs_blog)
+		return 0;
+
+	dir = debugfs_create_dir("blog", fsc->client->debugfs_dir);
+	if (IS_ERR(dir))
+		return PTR_ERR(dir);
+	fsc->debugfs_blog = dir;
+
+	debugfs_create_file("enabled", 0600, fsc->debugfs_blog, fsc,
+			    &blog_enabled_fops);
+	debugfs_create_file("entries", 0400, fsc->debugfs_blog, fsc,
+			    &blog_entries_fops);
+
+	debugfs_create_file("stats", 0400, fsc->debugfs_blog, fsc,
+			    &blog_stats_fops);
+
+	debugfs_create_file("sources", 0400, fsc->debugfs_blog, fsc,
+			    &blog_sources_fops);
+
+	debugfs_create_file("clients", 0400, fsc->debugfs_blog, fsc,
+			    &blog_clients_fops);
+
+	debugfs_create_file("clear", 0200, fsc->debugfs_blog, fsc,
+			    &blog_clear_fops);
+
+	pr_debug("ceph: BLOG debugfs initialized\n");
+	return 0;
+}
+
+void ceph_blog_debugfs_cleanup(struct ceph_fs_client *fsc)
+{
+	if (!fsc || !fsc->debugfs_blog)
+		return;
+
+	debugfs_remove_recursive(fsc->debugfs_blog);
+	fsc->debugfs_blog = NULL;
+	pr_debug("ceph: BLOG debugfs cleaned up\n");
+}
+
+static int jiffies_to_formatted_time(u64 jiffies_value, char *buffer,
+	size_t buffer_len)
+{
+	u64 now_ns = ktime_get_real_ns();
+	u64 now_jiffies = get_jiffies_64();
+	u64 delta_jiffies = (now_jiffies > jiffies_value) ?
+		now_jiffies - jiffies_value : 0;
+	u64 delta_ns = jiffies64_to_nsecs(delta_jiffies);
+	u64 event_ns = (delta_ns > now_ns) ? 0 : now_ns - delta_ns;
+	struct timespec64 event_ts = ns_to_timespec64(event_ns);
+	struct tm tm_time;
+
+	if (!buffer || !buffer_len)
+		return -EINVAL;
+
+	time64_to_tm(event_ts.tv_sec, 0, &tm_time);
+
+	return scnprintf(buffer, buffer_len,
+			"%04ld-%02d-%02d %02d:%02d:%02d.%03lu",
+			tm_time.tm_year + 1900, tm_time.tm_mon + 1, tm_time.tm_mday,
+			tm_time.tm_hour, tm_time.tm_min, tm_time.tm_sec,
+			(unsigned long)(event_ts.tv_nsec / NSEC_PER_MSEC));
+}
diff --git a/fs/ceph/debugfs.c b/fs/ceph/debugfs.c
index 18eb5da03411..6a9a5551d9d7 100644
--- a/fs/ceph/debugfs.c
+++ b/fs/ceph/debugfs.c
@@ -11,12 +11,12 @@
 #include <linux/ktime.h>
 #include <linux/uaccess.h>
 #include <linux/atomic.h>
-
 #include <linux/ceph/libceph.h>
 #include <linux/ceph/mon_client.h>
 #include <linux/ceph/auth.h>
 #include <linux/ceph/debugfs.h>
 #include <linux/ceph/decode.h>
+#include <linux/ceph/ceph_blog.h>
 
 #include "super.h"
 
@@ -650,6 +650,9 @@ void ceph_fs_debugfs_cleanup(struct ceph_fs_client *fsc)
 	debugfs_remove_recursive(fsc->debugfs_reset_dir);
 	debugfs_remove(fsc->debugfs_subvolume_metrics);
 	debugfs_remove_recursive(fsc->debugfs_metrics_dir);
+
+	ceph_blog_debugfs_cleanup(fsc);
+
 	doutc(fsc->client, "done\n");
 }
 
@@ -728,10 +731,12 @@ void ceph_fs_debugfs_init(struct ceph_fs_client *fsc)
 		debugfs_create_file("subvolumes", 0400,
 				    fsc->debugfs_metrics_dir, fsc,
 				    &subvolume_metrics_fops);
+
+	ceph_blog_debugfs_init(fsc);
+
 	doutc(fsc->client, "done\n");
 }
 
-
 #else  /* CONFIG_DEBUG_FS */
 
 void ceph_fs_debugfs_init(struct ceph_fs_client *fsc)
@@ -742,4 +747,4 @@ void ceph_fs_debugfs_cleanup(struct ceph_fs_client *fsc)
 {
 }
 
-#endif  /* CONFIG_DEBUG_FS */
+#endif	/* CONFIG_DEBUG_FS */
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 16+ messages in thread

* [PATCH v7 11/14] ceph: convert VFS inode and directory paths to BLOG logging
  2026-09-24 15:30 [PATCH v7 00/14] ceph: add binary logging (BLOG) for CephFS Alex Markuze
                   ` (9 preceding siblings ...)
  2026-09-24 15:30 ` [PATCH v7 10/14] ceph: add BLOG debugfs interface Alex Markuze
@ 2026-09-24 15:30 ` Alex Markuze
  2026-09-24 15:30 ` [PATCH v7 12/14] ceph: convert VFS data I/O " Alex Markuze
                   ` (3 subsequent siblings)
  14 siblings, 0 replies; 16+ messages in thread
From: Alex Markuze @ 2026-09-24 15:30 UTC (permalink / raw)
  To: ceph-devel; +Cc: idryomov, xiubo.li

Replace text debug logging with the boutc() wrappers in dir.c (readdir,
lookup, mkdir, unlink, rename), inode.c (getattr, setattr, fill_inode),
and export.c (NFS export ops).

Take a referenced dentry-name snapshot in a binary-only helper
instead of copying every readdir name into a VFS stack buffer.
The ordinary text path continues to use lazy %pd formatting.

Signed-off-by: Alex Markuze <amarkuze@redhat.com>
Assisted-by: LLM
---
 fs/ceph/dir.c    | 329 +++++++++++++++++++++++++++++++++++------------
 fs/ceph/export.c |  83 +++++++++---
 fs/ceph/inode.c  | 256 +++++++++++++++++++++---------------
 3 files changed, 466 insertions(+), 202 deletions(-)

diff --git a/fs/ceph/dir.c b/fs/ceph/dir.c
index 8acf308dfbf5..f4b5a1f84a8a 100644
--- a/fs/ceph/dir.c
+++ b/fs/ceph/dir.c
@@ -122,7 +122,7 @@ static int note_last_dentry(struct ceph_fs_client *fsc,
 	memcpy(dfi->last_name, name, len);
 	dfi->last_name[len] = 0;
 	dfi->next_offset = next_offset;
-	doutc(fsc->client, "'%s'\n", dfi->last_name);
+	boutc(fsc->client, "'%s'\n", dfi->last_name);
 	return 0;
 }
 
@@ -146,7 +146,7 @@ __dcache_find_get_entry(struct dentry *parent, u64 idx,
 		cache_ctl->folio = filemap_lock_folio(&dir->i_data, ptr_pgoff);
 		if (IS_ERR(cache_ctl->folio)) {
 			cache_ctl->folio = NULL;
-			doutc(cl, " folio %lu not found\n", ptr_pgoff);
+			boutc(cl, " folio %lu not found\n", ptr_pgoff);
 			return ERR_PTR(-EAGAIN);
 		}
 		/* reading/filling the cache are serialized by
@@ -172,6 +172,20 @@ __dcache_find_get_entry(struct dentry *parent, u64 idx,
 	return dentry ? : ERR_PTR(-EAGAIN);
 }
 
+static noinline void ceph_blog_readdir(struct blog_tls_ctx *blog_ctx,
+				       struct ceph_client *cl, u64 offset,
+				       struct dentry *dentry)
+{
+	struct name_snapshot name;
+
+	take_dentry_name_snapshot(&name, dentry);
+	CEPH_BLOG_LOG_CLIENT(blog_ctx, cl, " %llx dentry %p %s %p\n",
+			     offset, dentry,
+			     BLOG_STR(name.name.name, name.name.len),
+			     d_inode(dentry));
+	release_dentry_name_snapshot(&name);
+}
+
 /*
  * When possible, we try to satisfy a readdir by peeking at the
  * dcache.  We make this work by carefully ordering dentries on
@@ -197,7 +211,7 @@ static int __dcache_readdir(struct file *file,  struct dir_context *ctx,
 	u64 idx = 0;
 	int err = 0;
 
-	doutc(cl, "%p %llx.%llx v%u at %llx\n", dir, ceph_vinop(dir),
+	boutc(cl, "%p %llx.%llx v%u at %llx\n", dir, ceph_vinop(dir),
 	      (unsigned)shared_gen, ctx->pos);
 
 	/* search start position */
@@ -228,7 +242,7 @@ static int __dcache_readdir(struct file *file,  struct dir_context *ctx,
 			dput(dentry);
 		}
 
-		doutc(cl, "%p %llx.%llx cache idx %llu\n", dir,
+		boutc(cl, "%p %llx.%llx cache idx %llu\n", dir,
 		      ceph_vinop(dir), idx);
 	}
 
@@ -265,8 +279,13 @@ static int __dcache_readdir(struct file *file,  struct dir_context *ctx,
 		spin_unlock(&dentry->d_lock);
 
 		if (emit_dentry) {
-			doutc(cl, " %llx dentry %p %pd %p\n", di->offset,
-			      dentry, dentry, d_inode(dentry));
+			struct blog_tls_ctx *blog_ctx = ceph_blog_get_ctx(fsc);
+
+			if (blog_ctx)
+				ceph_blog_readdir(blog_ctx, cl, di->offset, dentry);
+			else
+				doutc(cl, " %llx dentry %p %pd %p\n",
+				      di->offset, dentry, dentry, d_inode(dentry));
 			ctx->pos = di->offset;
 			if (!dir_emit(ctx, dentry->d_name.name,
 				      dentry->d_name.len, ceph_present_inode(d_inode(dentry)),
@@ -326,19 +345,26 @@ static int ceph_readdir(struct file *file, struct dir_context *ctx)
 	int err;
 	unsigned frag = -1;
 	struct ceph_mds_reply_info_parsed *rinfo;
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(fsc, &__ji);
 
-	doutc(cl, "%p %llx.%llx file %p pos %llx\n", inode,
+	boutc(cl, "%p %llx.%llx file %p pos %llx\n", inode,
 	      ceph_vinop(inode), file, ctx->pos);
-	if (dfi->file_info.flags & CEPH_F_ATEND)
+	if (dfi->file_info.flags & CEPH_F_ATEND) {
+		ceph_blog_exit(&__ji);
 		return 0;
+	}
 
 	/* always start with . and .. */
 	if (ctx->pos == 0) {
-		doutc(cl, "%p %llx.%llx off 0 -> '.'\n", inode,
+		boutc(cl, "%p %llx.%llx off 0 -> '.'\n", inode,
 		      ceph_vinop(inode));
 		if (!dir_emit(ctx, ".", 1, ceph_present_inode(inode),
-			    inode->i_mode >> 12))
+			    inode->i_mode >> 12)) {
+			ceph_blog_exit(&__ji);
 			return 0;
+		}
 		ctx->pos = 1;
 	}
 	if (ctx->pos == 1) {
@@ -349,16 +375,20 @@ static int ceph_readdir(struct file *file, struct dir_context *ctx)
 		ino = ceph_present_inode(dentry->d_parent->d_inode);
 		spin_unlock(&dentry->d_lock);
 
-		doutc(cl, "%p %llx.%llx off 1 -> '..'\n", inode,
+		boutc(cl, "%p %llx.%llx off 1 -> '..'\n", inode,
 		      ceph_vinop(inode));
-		if (!dir_emit(ctx, "..", 2, ino, inode->i_mode >> 12))
+		if (!dir_emit(ctx, "..", 2, ino, inode->i_mode >> 12)) {
+			ceph_blog_exit(&__ji);
 			return 0;
+		}
 		ctx->pos = 2;
 	}
 
 	err = ceph_fscrypt_prepare_readdir(inode);
-	if (err < 0)
+	if (err < 0) {
+		ceph_blog_exit(&__ji);
 		return err;
+	}
 
 	spin_lock(&ci->i_ceph_lock);
 	/* request Fx cap. if have Fx, we don't need to release Fs cap
@@ -374,8 +404,10 @@ static int ceph_readdir(struct file *file, struct dir_context *ctx)
 
 		spin_unlock(&ci->i_ceph_lock);
 		err = __dcache_readdir(file, ctx, shared_gen);
-		if (err != -EAGAIN)
+		if (err != -EAGAIN) {
+			ceph_blog_exit(&__ji);
 			return err;
+		}
 	} else {
 		spin_unlock(&ci->i_ceph_lock);
 	}
@@ -404,15 +436,18 @@ static int ceph_readdir(struct file *file, struct dir_context *ctx)
 			frag = fpos_frag(ctx->pos);
 		}
 
-		doutc(cl, "fetching %p %llx.%llx frag %x offset '%s'\n",
+		boutc(cl, "fetching %p %llx.%llx frag %x offset '%s'\n",
 		      inode, ceph_vinop(inode), frag, dfi->last_name);
 		req = ceph_mdsc_create_request(mdsc, op, USE_AUTH_MDS);
-		if (IS_ERR(req))
+		if (IS_ERR(req)) {
+			ceph_blog_exit(&__ji);
 			return PTR_ERR(req);
+		}
 
 		err = ceph_alloc_readdir_reply_buffer(req, inode);
 		if (err) {
 			ceph_mdsc_put_request(req);
+			ceph_blog_exit(&__ji);
 			return err;
 		}
 		/* hints to request -> mds selection code */
@@ -428,6 +463,7 @@ static int ceph_readdir(struct file *file, struct dir_context *ctx)
 			req->r_path2 = kzalloc(NAME_MAX + 1, GFP_KERNEL);
 			if (!req->r_path2) {
 				ceph_mdsc_put_request(req);
+				ceph_blog_exit(&__ji);
 				return -ENOMEM;
 			}
 			memcpy(req->r_path2, dfi->last_name, len);
@@ -435,6 +471,7 @@ static int ceph_readdir(struct file *file, struct dir_context *ctx)
 			err = ceph_encode_encrypted_dname(inode, req->r_path2, len);
 			if (err < 0) {
 				ceph_mdsc_put_request(req);
+				ceph_blog_exit(&__ji);
 				return err;
 			}
 		} else if (is_hash_order(ctx->pos)) {
@@ -456,9 +493,10 @@ static int ceph_readdir(struct file *file, struct dir_context *ctx)
 		err = ceph_mdsc_do_request(mdsc, NULL, req);
 		if (err < 0) {
 			ceph_mdsc_put_request(req);
+			ceph_blog_exit(&__ji);
 			return err;
 		}
-		doutc(cl, "%p %llx.%llx got and parsed readdir result=%d"
+		boutc(cl, "%p %llx.%llx got and parsed readdir result=%d"
 		      "on frag %x, end=%d, complete=%d, hash_order=%d\n",
 		      inode, ceph_vinop(inode), err, frag,
 		      (int)req->r_reply_info.dir_end,
@@ -493,7 +531,7 @@ static int ceph_readdir(struct file *file, struct dir_context *ctx)
 				dfi->dir_ordered_count = req->r_dir_ordered_cnt;
 			}
 		} else {
-			doutc(cl, "%p %llx.%llx !did_prepopulate\n", inode,
+			boutc(cl, "%p %llx.%llx !did_prepopulate\n", inode,
 			      ceph_vinop(inode));
 			/* disable readdir cache */
 			dfi->readdir_cache_idx = -1;
@@ -512,6 +550,7 @@ static int ceph_readdir(struct file *file, struct dir_context *ctx)
 			if (err) {
 				ceph_mdsc_put_request(dfi->last_readdir);
 				dfi->last_readdir = NULL;
+				ceph_blog_exit(&__ji);
 				return err;
 			}
 		} else if (req->r_reply_info.dir_end) {
@@ -521,7 +560,7 @@ static int ceph_readdir(struct file *file, struct dir_context *ctx)
 	}
 
 	rinfo = &dfi->last_readdir->r_reply_info;
-	doutc(cl, "%p %llx.%llx frag %x num %d pos %llx chunk first %llx\n",
+	boutc(cl, "%p %llx.%llx frag %x num %d pos %llx chunk first %llx\n",
 	      inode, ceph_vinop(inode), dfi->frag, rinfo->dir_nr, ctx->pos,
 	      rinfo->dir_nr ? rinfo->dir_entries[0].offset : 0LL);
 
@@ -548,19 +587,26 @@ static int ceph_readdir(struct file *file, struct dir_context *ctx)
 				inode, ceph_vinop(inode), rde->offset, ctx->pos);
 			ceph_mdsc_put_request(dfi->last_readdir);
 			dfi->last_readdir = NULL;
+			ceph_blog_exit(&__ji);
 			return -EIO;
 		}
 
 		if (WARN_ON_ONCE(!rde->inode.in)) {
 			ceph_mdsc_put_request(dfi->last_readdir);
 			dfi->last_readdir = NULL;
+			ceph_blog_exit(&__ji);
 			return -EIO;
 		}
 
 		ctx->pos = rde->offset;
-		doutc(cl, "%p %llx.%llx (%d/%d) -> %llx '%.*s' %p\n", inode,
-		      ceph_vinop(inode), i, rinfo->dir_nr, ctx->pos,
-		      rde->name_len, rde->name, &rde->inode.in);
+		boutc_bounded(cl,
+			      "%p %llx.%llx (%d/%d) -> %llx '%.*s' %p\n",
+			      (inode, ceph_vinop(inode), i, rinfo->dir_nr,
+			       ctx->pos, rde->name_len,
+			       BLOG_STR(rde->name, rde->name_len), &rde->inode.in),
+			      (inode, ceph_vinop(inode), i, rinfo->dir_nr,
+			       ctx->pos, rde->name_len, (const char *)rde->name,
+			       &rde->inode.in));
 
 		if (!dir_emit(ctx, rde->name, rde->name_len,
 			      ceph_present_ino(inode->i_sb, le64_to_cpu(rde->inode.in->ino)),
@@ -571,7 +617,8 @@ static int ceph_readdir(struct file *file, struct dir_context *ctx)
 			 * doesn't have enough memory, etc. So for next readdir
 			 * it will continue.
 			 */
-			doutc(cl, "filldir stopping us...\n");
+			boutc(cl, "filldir stopping us...\n");
+			ceph_blog_exit(&__ji);
 			return 0;
 		}
 
@@ -602,7 +649,7 @@ static int ceph_readdir(struct file *file, struct dir_context *ctx)
 			kfree(dfi->last_name);
 			dfi->last_name = NULL;
 		}
-		doutc(cl, "%p %llx.%llx next frag is %x\n", inode,
+		boutc(cl, "%p %llx.%llx next frag is %x\n", inode,
 		      ceph_vinop(inode), frag);
 		goto more;
 	}
@@ -618,7 +665,7 @@ static int ceph_readdir(struct file *file, struct dir_context *ctx)
 		spin_lock(&ci->i_ceph_lock);
 		if (dfi->dir_ordered_count ==
 				atomic64_read(&ci->i_ordered_count)) {
-			doutc(cl, " marking %p %llx.%llx complete and ordered\n",
+			boutc(cl, " marking %p %llx.%llx complete and ordered\n",
 			      inode, ceph_vinop(inode));
 			/* use i_size to track number of entries in
 			 * readdir cache */
@@ -626,15 +673,16 @@ static int ceph_readdir(struct file *file, struct dir_context *ctx)
 			i_size_write(inode, dfi->readdir_cache_idx *
 				     sizeof(struct dentry*));
 		} else {
-			doutc(cl, " marking %llx.%llx complete\n",
+			boutc(cl, " marking %llx.%llx complete\n",
 			      ceph_vinop(inode));
 		}
 		__ceph_dir_set_complete(ci, dfi->dir_release_count,
 					dfi->dir_ordered_count);
 		spin_unlock(&ci->i_ceph_lock);
 	}
-	doutc(cl, "%p %llx.%llx file %p done.\n", inode, ceph_vinop(inode),
+	boutc(cl, "%p %llx.%llx file %p done.\n", inode, ceph_vinop(inode),
 	      file);
+	ceph_blog_exit(&__ji);
 	return 0;
 }
 
@@ -680,8 +728,12 @@ static loff_t ceph_dir_llseek(struct file *file, loff_t offset, int whence)
 {
 	struct ceph_dir_file_info *dfi = file->private_data;
 	struct inode *inode = file->f_mapping->host;
+	struct ceph_fs_client *fsc = ceph_inode_to_fs_client(inode);
 	struct ceph_client *cl = ceph_inode_to_client(inode);
 	loff_t retval;
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(fsc, &__ji);
 
 	inode_lock(inode);
 	retval = -EINVAL;
@@ -700,7 +752,7 @@ static loff_t ceph_dir_llseek(struct file *file, loff_t offset, int whence)
 
 	if (offset >= 0) {
 		if (need_reset_readdir(dfi, offset)) {
-			doutc(cl, "%p %llx.%llx dropping %p content\n",
+			boutc(cl, "%p %llx.%llx dropping %p content\n",
 			      inode, ceph_vinop(inode), file);
 			reset_readdir(dfi);
 		} else if (is_hash_order(offset) && offset > file->f_pos) {
@@ -718,6 +770,7 @@ static loff_t ceph_dir_llseek(struct file *file, loff_t offset, int whence)
 	}
 out:
 	inode_unlock(inode);
+	ceph_blog_exit(&__ji);
 	return retval;
 }
 
@@ -738,9 +791,12 @@ struct dentry *ceph_handle_snapdir(struct ceph_mds_request *req,
 		struct inode *inode = ceph_get_snapdir(parent);
 
 		res = d_splice_alias(inode, dentry);
-		doutc(cl, "ENOENT on snapdir %p '%pd', linking to "
-		      "snapdir %p %llx.%llx. Spliced dentry %p\n",
-		      dentry, dentry, inode, ceph_vinop(inode), res);
+		boutc_formats(cl,
+			      "ENOENT on snapdir %p '%s', linking to snapdir %p %llx.%llx. Spliced dentry %p\n",
+			      "ENOENT on snapdir %p '%pd', linking to snapdir %p %llx.%llx. Spliced dentry %p\n",
+			      (dentry, dentry->d_name.name, inode,
+			       ceph_vinop(inode), res),
+			      (dentry, dentry, inode, ceph_vinop(inode), res));
 		if (res)
 			dentry = res;
 	}
@@ -767,7 +823,7 @@ struct dentry *ceph_finish_lookup(struct ceph_mds_request *req,
 		/* no trace? */
 		err = 0;
 		if (!req->r_reply_info.head->is_dentry) {
-			doutc(cl,
+			boutc(cl,
 			      "ENOENT and no trace, dentry %p inode %llx.%llx\n",
 			      dentry, ceph_vinop(d_inode(dentry)));
 			if (d_really_is_positive(dentry)) {
@@ -813,19 +869,28 @@ static struct dentry *ceph_lookup(struct inode *dir, struct dentry *dentry,
 	int op;
 	int mask;
 	int err;
+	struct ceph_journal_info __ji;
 
-	doutc(cl, "%p %llx.%llx/'%pd' dentry %p\n", dir, ceph_vinop(dir),
-	      dentry, dentry);
+	ceph_blog_enter(fsc, &__ji);
 
-	if (dentry->d_name.len > NAME_MAX)
+	boutc_formats(cl, "%p %llx.%llx/'%s' dentry %p\n",
+		      "%p %llx.%llx/'%pd' dentry %p\n",
+		      (dir, ceph_vinop(dir), dentry->d_name.name, dentry),
+		      (dir, ceph_vinop(dir), dentry, dentry));
+
+	if (dentry->d_name.len > NAME_MAX) {
+		ceph_blog_exit(&__ji);
 		return ERR_PTR(-ENAMETOOLONG);
+	}
 
 	if (IS_ENCRYPTED(dir)) {
 		bool had_key = fscrypt_has_encryption_key(dir);
 
 		err = fscrypt_prepare_lookup_partial(dir, dentry);
-		if (err < 0)
+		if (err < 0) {
+			ceph_blog_exit(&__ji);
 			return ERR_PTR(err);
+		}
 
 		/* mark directory as incomplete if it has been unlocked */
 		if (!had_key && fscrypt_has_encryption_key(dir))
@@ -838,7 +903,7 @@ static struct dentry *ceph_lookup(struct inode *dir, struct dentry *dentry,
 		struct ceph_dentry_info *di = ceph_dentry(dentry);
 
 		spin_lock(&ci->i_ceph_lock);
-		doutc(cl, " dir %llx.%llx flags are 0x%lx\n",
+		boutc(cl, " dir %llx.%llx flags are 0x%lx\n",
 		      ceph_vinop(dir), ci->i_ceph_flags);
 		if (strncmp(dentry->d_name.name,
 			    fsc->mount_options->snapdir_name,
@@ -850,11 +915,12 @@ static struct dentry *ceph_lookup(struct inode *dir, struct dentry *dentry,
 		    __ceph_caps_issued_mask_metric(ci, CEPH_CAP_FILE_SHARED, 1)) {
 			__ceph_touch_fmode(ci, mdsc, CEPH_FILE_MODE_RD);
 			spin_unlock(&ci->i_ceph_lock);
-			doutc(cl, " dir %llx.%llx complete, -ENOENT\n",
+			boutc(cl, " dir %llx.%llx complete, -ENOENT\n",
 			      ceph_vinop(dir));
 			if (d_unhashed(dentry))
 				d_add(dentry, NULL);
 			di->lease_shared_gen = atomic_read(&ci->i_shared_gen);
+			ceph_blog_exit(&__ji);
 			return NULL;
 		}
 		spin_unlock(&ci->i_ceph_lock);
@@ -863,8 +929,10 @@ static struct dentry *ceph_lookup(struct inode *dir, struct dentry *dentry,
 	op = ceph_snap(dir) == CEPH_SNAPDIR ?
 		CEPH_MDS_OP_LOOKUPSNAP : CEPH_MDS_OP_LOOKUP;
 	req = ceph_mdsc_create_request(mdsc, op, USE_ANY_MDS);
-	if (IS_ERR(req))
+	if (IS_ERR(req)) {
+		ceph_blog_exit(&__ji);
 		return ERR_CAST(req);
+	}
 	req->r_dentry = dget(dentry);
 	req->r_num_caps = 2;
 
@@ -890,7 +958,8 @@ static struct dentry *ceph_lookup(struct inode *dir, struct dentry *dentry,
 	}
 	dentry = ceph_finish_lookup(req, dentry, err);
 	ceph_mdsc_put_request(req);  /* will dput(dentry) */
-	doutc(cl, "result=%p\n", dentry);
+	boutc(cl, "result=%p\n", dentry);
+	ceph_blog_exit(&__ji);
 	return dentry;
 }
 
@@ -925,25 +994,37 @@ static int ceph_mknod(struct mnt_idmap *idmap, struct inode *dir,
 		      struct dentry *dentry, umode_t mode, dev_t rdev)
 {
 	struct ceph_mds_client *mdsc = ceph_sb_to_mdsc(dir->i_sb);
+	struct ceph_fs_client *fsc = mdsc->fsc;
 	struct ceph_client *cl = mdsc->fsc->client;
 	struct ceph_mds_request *req;
 	struct ceph_acl_sec_ctx as_ctx = {};
 	int err;
+	struct ceph_journal_info __ji;
 
-	if (ceph_in_snap(dir))
+	ceph_blog_enter(fsc, &__ji);
+
+	if (ceph_in_snap(dir)) {
+		ceph_blog_exit(&__ji);
 		return -EROFS;
+	}
 
 	err = ceph_wait_on_conflict_unlink(dentry);
-	if (err)
+	if (err) {
+		ceph_blog_exit(&__ji);
 		return err;
+	}
 
 	if (ceph_quota_is_max_files_exceeded(dir)) {
 		err = -EDQUOT;
 		goto out;
 	}
 
-	doutc(cl, "%p %llx.%llx/'%pd' dentry %p mode 0%ho rdev %d\n",
-	      dir, ceph_vinop(dir), dentry, dentry, mode, rdev);
+	boutc_formats(cl,
+		      "%p %llx.%llx/'%s' dentry %p mode 0%ho rdev %d\n",
+		      "%p %llx.%llx/'%pd' dentry %p mode 0%ho rdev %d\n",
+		      (dir, ceph_vinop(dir), dentry->d_name.name, dentry,
+		       mode, rdev),
+		      (dir, ceph_vinop(dir), dentry, dentry, mode, rdev));
 	req = ceph_mdsc_create_request(mdsc, CEPH_MDS_OP_MKNOD, USE_AUTH_MDS);
 	if (IS_ERR(req)) {
 		err = PTR_ERR(req);
@@ -985,6 +1066,7 @@ static int ceph_mknod(struct mnt_idmap *idmap, struct inode *dir,
 	else
 		d_drop(dentry);
 	ceph_release_acl_sec_ctx(&as_ctx);
+	ceph_blog_exit(&__ji);
 	return err;
 }
 
@@ -1036,26 +1118,36 @@ static int ceph_symlink(struct mnt_idmap *idmap, struct inode *dir,
 			struct dentry *dentry, const char *dest)
 {
 	struct ceph_mds_client *mdsc = ceph_sb_to_mdsc(dir->i_sb);
+	struct ceph_fs_client *fsc = mdsc->fsc;
 	struct ceph_client *cl = mdsc->fsc->client;
 	struct ceph_mds_request *req;
 	struct ceph_acl_sec_ctx as_ctx = {};
 	umode_t mode = S_IFLNK | 0777;
 	int err;
+	struct ceph_journal_info __ji;
 
-	if (ceph_in_snap(dir))
+	ceph_blog_enter(fsc, &__ji);
+
+	if (ceph_in_snap(dir)) {
+		ceph_blog_exit(&__ji);
 		return -EROFS;
+	}
 
 	err = ceph_wait_on_conflict_unlink(dentry);
-	if (err)
+	if (err) {
+		ceph_blog_exit(&__ji);
 		return err;
+	}
 
 	if (ceph_quota_is_max_files_exceeded(dir)) {
 		err = -EDQUOT;
 		goto out;
 	}
 
-	doutc(cl, "%p %llx.%llx/'%pd' to '%s'\n", dir, ceph_vinop(dir), dentry,
-	      dest);
+	boutc_formats(cl, "%p %llx.%llx/'%s' to '%s'\n",
+		      "%p %llx.%llx/'%pd' to '%s'\n",
+		      (dir, ceph_vinop(dir), dentry->d_name.name, dest),
+		      (dir, ceph_vinop(dir), dentry, dest));
 	req = ceph_mdsc_create_request(mdsc, CEPH_MDS_OP_SYMLINK, USE_AUTH_MDS);
 	if (IS_ERR(req)) {
 		err = PTR_ERR(req);
@@ -1103,6 +1195,7 @@ static int ceph_symlink(struct mnt_idmap *idmap, struct inode *dir,
 	if (err)
 		d_drop(dentry);
 	ceph_release_acl_sec_ctx(&as_ctx);
+	ceph_blog_exit(&__ji);
 	return err;
 }
 
@@ -1110,25 +1203,36 @@ static struct dentry *ceph_mkdir(struct mnt_idmap *idmap, struct inode *dir,
 				 struct dentry *dentry, umode_t mode)
 {
 	struct ceph_mds_client *mdsc = ceph_sb_to_mdsc(dir->i_sb);
+	struct ceph_fs_client *fsc = mdsc->fsc;
 	struct ceph_client *cl = mdsc->fsc->client;
 	struct ceph_mds_request *req;
 	struct ceph_acl_sec_ctx as_ctx = {};
 	struct dentry *ret;
 	int err;
 	int op;
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(fsc, &__ji);
 
 	err = ceph_wait_on_conflict_unlink(dentry);
-	if (err)
+	if (err) {
+		ceph_blog_exit(&__ji);
 		return ERR_PTR(err);
+	}
 
 	if (ceph_snap(dir) == CEPH_SNAPDIR) {
 		/* mkdir .snap/foo is a MKSNAP */
 		op = CEPH_MDS_OP_MKSNAP;
-		doutc(cl, "mksnap %llx.%llx/'%pd' dentry %p\n",
-		      ceph_vinop(dir), dentry, dentry);
+		boutc_formats(cl, "mksnap %llx.%llx/'%s' dentry %p\n",
+			      "mksnap %llx.%llx/'%pd' dentry %p\n",
+			      (ceph_vinop(dir), dentry->d_name.name, dentry),
+			      (ceph_vinop(dir), dentry, dentry));
 	} else if (!ceph_in_snap(dir)) {
-		doutc(cl, "mkdir %llx.%llx/'%pd' dentry %p mode 0%ho\n",
-		      ceph_vinop(dir), dentry, dentry, mode);
+		boutc_formats(cl, "mkdir %llx.%llx/'%s' dentry %p mode 0%ho\n",
+			      "mkdir %llx.%llx/'%pd' dentry %p mode 0%ho\n",
+			      (ceph_vinop(dir), dentry->d_name.name, dentry,
+			       mode),
+			      (ceph_vinop(dir), dentry, dentry, mode));
 		op = CEPH_MDS_OP_MKDIR;
 	} else {
 		ret = ERR_PTR(-EROFS);
@@ -1194,6 +1298,7 @@ static struct dentry *ceph_mkdir(struct mnt_idmap *idmap, struct inode *dir,
 		d_drop(dentry);
 	}
 	ceph_release_acl_sec_ctx(&as_ctx);
+	ceph_blog_exit(&__ji);
 	return ret;
 }
 
@@ -1201,29 +1306,45 @@ static int ceph_link(struct dentry *old_dentry, struct inode *dir,
 		     struct dentry *dentry)
 {
 	struct ceph_mds_client *mdsc = ceph_sb_to_mdsc(dir->i_sb);
+	struct ceph_fs_client *fsc = mdsc->fsc;
 	struct ceph_client *cl = mdsc->fsc->client;
 	struct ceph_mds_request *req;
 	int err;
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(fsc, &__ji);
 
-	if (dentry->d_flags & DCACHE_DISCONNECTED)
+	if (dentry->d_flags & DCACHE_DISCONNECTED) {
+		ceph_blog_exit(&__ji);
 		return -EINVAL;
+	}
 
 	err = ceph_wait_on_conflict_unlink(dentry);
-	if (err)
+	if (err) {
+		ceph_blog_exit(&__ji);
 		return err;
+	}
 
-	if (ceph_in_snap(dir))
+	if (ceph_in_snap(dir)) {
+		ceph_blog_exit(&__ji);
 		return -EROFS;
+	}
 
 	err = fscrypt_prepare_link(old_dentry, dir, dentry);
-	if (err)
+	if (err) {
+		ceph_blog_exit(&__ji);
 		return err;
+	}
 
-	doutc(cl, "%p %llx.%llx/'%pd' to '%pd'\n", dir, ceph_vinop(dir),
-	      old_dentry, dentry);
+	boutc_formats(cl, "%p %llx.%llx/'%s' to '%s'\n",
+		      "%p %llx.%llx/'%pd' to '%pd'\n",
+		      (dir, ceph_vinop(dir), old_dentry->d_name.name,
+		       dentry->d_name.name),
+		      (dir, ceph_vinop(dir), old_dentry, dentry));
 	req = ceph_mdsc_create_request(mdsc, CEPH_MDS_OP_LINK, USE_AUTH_MDS);
 	if (IS_ERR(req)) {
 		d_drop(dentry);
+		ceph_blog_exit(&__ji);
 		return PTR_ERR(req);
 	}
 	req->r_dentry = dget(dentry);
@@ -1250,6 +1371,7 @@ static int ceph_link(struct dentry *old_dentry, struct inode *dir,
 		d_instantiate(dentry, d_inode(old_dentry));
 	}
 	ceph_mdsc_put_request(req);
+	ceph_blog_exit(&__ji);
 	return err;
 }
 
@@ -1358,15 +1480,24 @@ static int ceph_unlink(struct inode *dir, struct dentry *dentry)
 	int err = -EROFS;
 	int op;
 	char *path;
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(fsc, &__ji);
 
 	if (ceph_snap(dir) == CEPH_SNAPDIR) {
 		/* rmdir .snap/foo is RMSNAP */
-		doutc(cl, "rmsnap %llx.%llx/'%pd' dn\n", ceph_vinop(dir),
-		      dentry);
+		boutc_formats(cl, "rmsnap %llx.%llx/'%s' dn\n",
+			      "rmsnap %llx.%llx/'%pd' dn\n",
+			      (ceph_vinop(dir), dentry->d_name.name),
+			      (ceph_vinop(dir), dentry));
 		op = CEPH_MDS_OP_RMSNAP;
 	} else if (!ceph_in_snap(dir)) {
-		doutc(cl, "unlink/rmdir %llx.%llx/'%pd' inode %llx.%llx\n",
-		      ceph_vinop(dir), dentry, ceph_vinop(inode));
+		boutc_formats(cl,
+			      "unlink/rmdir %llx.%llx/'%s' inode %llx.%llx\n",
+			      "unlink/rmdir %llx.%llx/'%pd' inode %llx.%llx\n",
+			      (ceph_vinop(dir), dentry->d_name.name,
+			       ceph_vinop(inode)),
+			      (ceph_vinop(dir), dentry, ceph_vinop(inode)));
 		op = d_is_dir(dentry) ?
 			CEPH_MDS_OP_RMDIR : CEPH_MDS_OP_UNLINK;
 	} else
@@ -1389,6 +1520,7 @@ static int ceph_unlink(struct inode *dir, struct dentry *dentry)
 
 		/* For none EACCES cases will let the MDS do the mds auth check */
 		if (err == -EACCES) {
+			ceph_blog_exit(&__ji);
 			return err;
 		} else if (err < 0) {
 			try_async = false;
@@ -1414,9 +1546,12 @@ static int ceph_unlink(struct inode *dir, struct dentry *dentry)
 	    (req->r_dir_caps = get_caps_for_async_unlink(dir, dentry))) {
 		struct ceph_dentry_info *di = ceph_dentry(dentry);
 
-		doutc(cl, "async unlink on %llx.%llx/'%pd' caps=%s",
-		      ceph_vinop(dir), dentry,
-		      ceph_cap_string(req->r_dir_caps));
+		boutc_formats(cl, "async unlink on %llx.%llx/'%s' caps=%s",
+			      "async unlink on %llx.%llx/'%pd' caps=%s",
+			      (ceph_vinop(dir), dentry->d_name.name,
+			       ceph_cap_string(req->r_dir_caps)),
+			      (ceph_vinop(dir), dentry,
+			       ceph_cap_string(req->r_dir_caps)));
 		set_bit(CEPH_MDS_R_ASYNC, &req->r_req_flags);
 		req->r_callback = ceph_async_unlink_cb;
 		req->r_old_inode = d_inode(dentry);
@@ -1475,6 +1610,7 @@ static int ceph_unlink(struct inode *dir, struct dentry *dentry)
 
 	ceph_mdsc_put_request(req);
 out:
+	ceph_blog_exit(&__ji);
 	return err;
 }
 
@@ -1483,42 +1619,63 @@ static int ceph_rename(struct mnt_idmap *idmap, struct inode *old_dir,
 		       struct dentry *new_dentry, unsigned int flags)
 {
 	struct ceph_mds_client *mdsc = ceph_sb_to_mdsc(old_dir->i_sb);
+	struct ceph_fs_client *fsc = mdsc->fsc;
 	struct ceph_client *cl = mdsc->fsc->client;
 	struct ceph_mds_request *req;
 	int op = CEPH_MDS_OP_RENAME;
 	int err;
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(fsc, &__ji);
 
-	if (flags)
+	if (flags) {
+		ceph_blog_exit(&__ji);
 		return -EINVAL;
+	}
 
-	if (ceph_snap(old_dir) != ceph_snap(new_dir))
+	if (ceph_snap(old_dir) != ceph_snap(new_dir)) {
+		ceph_blog_exit(&__ji);
 		return -EXDEV;
+	}
 	if (ceph_in_snap(old_dir)) {
 		if (old_dir == new_dir && ceph_snap(old_dir) == CEPH_SNAPDIR)
 			op = CEPH_MDS_OP_RENAMESNAP;
-		else
+		else {
+			ceph_blog_exit(&__ji);
 			return -EROFS;
+		}
 	}
 	/* don't allow cross-quota renames */
 	if ((old_dir != new_dir) &&
-	    (!ceph_quota_is_same_realm(old_dir, new_dir)))
+	    (!ceph_quota_is_same_realm(old_dir, new_dir))) {
+		ceph_blog_exit(&__ji);
 		return -EXDEV;
+	}
 
 	err = ceph_wait_on_conflict_unlink(new_dentry);
-	if (err)
+	if (err) {
+		ceph_blog_exit(&__ji);
 		return err;
+	}
 
 	err = fscrypt_prepare_rename(old_dir, old_dentry, new_dir, new_dentry,
 				     flags);
-	if (err)
+	if (err) {
+		ceph_blog_exit(&__ji);
 		return err;
+	}
 
-	doutc(cl, "%llx.%llx/'%pd' to %llx.%llx/'%pd'\n",
-	      ceph_vinop(old_dir), old_dentry, ceph_vinop(new_dir),
-	      new_dentry);
+	boutc_formats(cl, "%llx.%llx/'%s' to %llx.%llx/'%s'\n",
+		      "%llx.%llx/'%pd' to %llx.%llx/'%pd'\n",
+		      (ceph_vinop(old_dir), old_dentry->d_name.name,
+		       ceph_vinop(new_dir), new_dentry->d_name.name),
+		      (ceph_vinop(old_dir), old_dentry,
+		       ceph_vinop(new_dir), new_dentry));
 	req = ceph_mdsc_create_request(mdsc, op, USE_AUTH_MDS);
-	if (IS_ERR(req))
+	if (IS_ERR(req)) {
+		ceph_blog_exit(&__ji);
 		return PTR_ERR(req);
+	}
 	ihold(old_dir);
 	req->r_dentry = dget(new_dentry);
 	req->r_num_caps = 2;
@@ -1547,6 +1704,7 @@ static int ceph_rename(struct mnt_idmap *idmap, struct inode *old_dir,
 		d_move(old_dentry, new_dentry);
 	}
 	ceph_mdsc_put_request(req);
+	ceph_blog_exit(&__ji);
 	return err;
 }
 
@@ -1563,7 +1721,8 @@ void __ceph_dentry_lease_touch(struct ceph_dentry_info *di)
 	struct ceph_mds_client *mdsc = ceph_sb_to_fs_client(dn->d_sb)->mdsc;
 	struct ceph_client *cl = mdsc->fsc->client;
 
-	doutc(cl, "%p %p '%pd'\n", di, dn, dn);
+	boutc_formats(cl, "%p %p '%s'\n", "%p %p '%pd'\n",
+		      (di, dn, dn->d_name.name), (di, dn, dn));
 
 	di->flags |= CEPH_DENTRY_LEASE_LIST;
 	if (di->flags & CEPH_DENTRY_SHRINK_LIST) {
@@ -1597,7 +1756,10 @@ void __ceph_dentry_dir_lease_touch(struct ceph_dentry_info *di)
 	struct ceph_mds_client *mdsc = ceph_sb_to_fs_client(dn->d_sb)->mdsc;
 	struct ceph_client *cl = mdsc->fsc->client;
 
-	doutc(cl, "%p %p '%pd' (offset 0x%llx)\n", di, dn, dn, di->offset);
+	boutc_formats(cl, "%p %p '%s' (offset 0x%llx)\n",
+		      "%p %p '%pd' (offset 0x%llx)\n",
+		      (di, dn, dn->d_name.name, di->offset),
+		      (di, dn, dn, di->offset));
 
 	if (!list_empty(&di->lease_list)) {
 		if (di->flags & CEPH_DENTRY_LEASE_LIST) {
@@ -1904,7 +2066,7 @@ static int dentry_lease_is_valid(struct dentry *dentry, unsigned int flags)
 					 CEPH_MDS_LEASE_RENEW, seq);
 		ceph_put_mds_session(session);
 	}
-	doutc(cl, "dentry %p = %d\n", dentry, valid);
+	boutc(cl, "dentry %p = %d\n", dentry, valid);
 	return valid;
 }
 
@@ -1969,8 +2131,9 @@ static int dir_lease_is_valid(struct inode *dir, struct dentry *dentry,
 			valid = 0;
 		spin_unlock(&dentry->d_lock);
 	}
-	doutc(cl, "dir %p %llx.%llx v%u dentry %p '%pd' = %d\n", dir,
-	      ceph_vinop(dir), (unsigned)atomic_read(&ci->i_shared_gen),
+	doutc(cl, "dir %p %llx.%llx v%u dentry %p '%pd' = %d\n",
+	      dir, ceph_vinop(dir),
+	      (unsigned)atomic_read(&ci->i_shared_gen),
 	      dentry, dentry, valid);
 	return valid;
 }
@@ -1992,6 +2155,7 @@ static int ceph_d_revalidate(struct inode *dir, const struct qstr *name,
 
 	inode = d_inode_rcu(dentry);
 
+	/* No BLOG enter under LOOKUP_RCU; keep %pd text logging. */
 	doutc(cl, "%p '%pd' inode %p offset 0x%llx nokey %d\n",
 	      dentry, dentry, inode, ceph_dentry(dentry)->offset,
 	      !!(dentry->d_flags & DCACHE_NOKEY_NAME));
@@ -2065,7 +2229,8 @@ static int ceph_d_revalidate(struct inode *dir, const struct qstr *name,
 		percpu_counter_inc(&mdsc->metric.d_lease_hit);
 	}
 
-	doutc(cl, "%p '%pd' %s\n", dentry, dentry, valid ? "valid" : "invalid");
+	doutc(cl, "%p '%pd' %s\n", dentry, dentry,
+	      valid ? "valid" : "invalid");
 	if (!valid)
 		ceph_dir_clear_complete(dir);
 	return valid;
diff --git a/fs/ceph/export.c b/fs/ceph/export.c
index 1466c46f0691..a698f5bb11a0 100644
--- a/fs/ceph/export.c
+++ b/fs/ceph/export.c
@@ -87,32 +87,42 @@ static int ceph_encode_snapfh(struct inode *inode, u32 *rawfh, int *max_len,
 	*max_len = snap_handle_length;
 	ret = FILEID_BTRFS_WITH_PARENT;
 out:
-	doutc(cl, "%p %llx.%llx ret=%d\n", inode, ceph_vinop(inode), ret);
+	boutc(cl, "%p %llx.%llx ret=%d\n", inode, ceph_vinop(inode), ret);
 	return ret;
 }
 
 static int ceph_encode_fh(struct inode *inode, u32 *rawfh, int *max_len,
 			  struct inode *parent_inode)
 {
+	struct ceph_fs_client *fsc = ceph_inode_to_fs_client(inode);
 	struct ceph_client *cl = ceph_inode_to_client(inode);
 	static const int handle_length = CEPH_FH_BASIC_SIZE;
 	static const int connected_handle_length = CEPH_FH_WITH_PARENT_SIZE;
 	int type;
+	struct ceph_journal_info __ji;
 
-	if (ceph_in_snap(inode))
-		return ceph_encode_snapfh(inode, rawfh, max_len, parent_inode);
+	ceph_blog_enter(fsc, &__ji);
+
+	if (ceph_in_snap(inode)) {
+		int ret = ceph_encode_snapfh(inode, rawfh, max_len,
+					     parent_inode);
+		ceph_blog_exit(&__ji);
+		return ret;
+	}
 
 	if (parent_inode && (*max_len < connected_handle_length)) {
 		*max_len = connected_handle_length;
+		ceph_blog_exit(&__ji);
 		return FILEID_INVALID;
 	} else if (*max_len < handle_length) {
 		*max_len = handle_length;
+		ceph_blog_exit(&__ji);
 		return FILEID_INVALID;
 	}
 
 	if (parent_inode) {
 		struct ceph_nfs_confh *cfh = (void *)rawfh;
-		doutc(cl, "%p %llx.%llx with parent %p %llx.%llx\n", inode,
+		boutc(cl, "%p %llx.%llx with parent %p %llx.%llx\n", inode,
 		      ceph_vinop(inode), parent_inode, ceph_vinop(parent_inode));
 		cfh->ino = ceph_ino(inode);
 		cfh->parent_ino = ceph_ino(parent_inode);
@@ -120,11 +130,12 @@ static int ceph_encode_fh(struct inode *inode, u32 *rawfh, int *max_len,
 		type = FILEID_INO32_GEN_PARENT;
 	} else {
 		struct ceph_nfs_fh *fh = (void *)rawfh;
-		doutc(cl, "%p %llx.%llx\n", inode, ceph_vinop(inode));
+		boutc(cl, "%p %llx.%llx\n", inode, ceph_vinop(inode));
 		fh->ino = ceph_ino(inode);
 		*max_len = handle_length;
 		type = FILEID_INO32_GEN;
 	}
+	ceph_blog_exit(&__ji);
 	return type;
 }
 
@@ -286,9 +297,9 @@ static struct dentry *__snapfh_to_dentry(struct super_block *sb,
 	ceph_mdsc_put_request(req);
 
 	if (want_parent) {
-		doutc(cl, "%llx.%llx\n err=%d\n", vino.ino, vino.snap, err);
+		boutc(cl, "%llx.%llx\n err=%d\n", vino.ino, vino.snap, err);
 	} else {
-		doutc(cl, "%llx.%llx parent %llx hash %x err=%d", vino.ino,
+		boutc(cl, "%llx.%llx parent %llx hash %x err=%d", vino.ino,
 		      vino.snap, sfh->parent_ino, sfh->hash, err);
 	}
 	/* see comments in ceph_get_parent() */
@@ -304,20 +315,31 @@ static struct dentry *ceph_fh_to_dentry(struct super_block *sb,
 {
 	struct ceph_fs_client *fsc = ceph_sb_to_fs_client(sb);
 	struct ceph_nfs_fh *fh = (void *)fid->raw;
+	struct dentry *dentry;
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(fsc, &__ji);
 
 	if (fh_type == FILEID_BTRFS_WITH_PARENT) {
 		struct ceph_nfs_snapfh *sfh = (void *)fid->raw;
+		ceph_blog_exit(&__ji);
 		return __snapfh_to_dentry(sb, sfh, false);
 	}
 
 	if (fh_type != FILEID_INO32_GEN  &&
-	    fh_type != FILEID_INO32_GEN_PARENT)
+	    fh_type != FILEID_INO32_GEN_PARENT) {
+		ceph_blog_exit(&__ji);
 		return NULL;
-	if (fh_len < sizeof(*fh) / BYTES_PER_U32)
+	}
+	if (fh_len < sizeof(*fh) / BYTES_PER_U32) {
+		ceph_blog_exit(&__ji);
 		return NULL;
+	}
 
-	doutc(fsc->client, "%llx\n", fh->ino);
-	return __fh_to_dentry(sb, fh->ino);
+	boutc(fsc->client, "%llx\n", fh->ino);
+	dentry = __fh_to_dentry(sb, fh->ino);
+	ceph_blog_exit(&__ji);
+	return dentry;
 }
 
 static struct dentry *__get_parent(struct super_block *sb,
@@ -369,8 +391,12 @@ static struct dentry *__get_parent(struct super_block *sb,
 static struct dentry *ceph_get_parent(struct dentry *child)
 {
 	struct inode *inode = d_inode(child);
+	struct ceph_fs_client *fsc = ceph_inode_to_fs_client(inode);
 	struct ceph_client *cl = ceph_inode_to_client(inode);
 	struct dentry *dn;
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(fsc, &__ji);
 
 	if (ceph_in_snap(inode)) {
 		struct inode* dir;
@@ -409,8 +435,9 @@ static struct dentry *ceph_get_parent(struct dentry *child)
 		dn = __get_parent(child->d_sb, child, 0);
 	}
 out:
-	doutc(cl, "child %p %p %llx.%llx err=%ld\n", child, inode,
+	boutc(cl, "child %p %p %llx.%llx err=%ld\n", child, inode,
 	      ceph_vinop(inode), (long)PTR_ERR_OR_ZERO(dn));
+	ceph_blog_exit(&__ji);
 	return dn;
 }
 
@@ -424,21 +451,30 @@ static struct dentry *ceph_fh_to_parent(struct super_block *sb,
 	struct ceph_fs_client *fsc = ceph_sb_to_fs_client(sb);
 	struct ceph_nfs_confh *cfh = (void *)fid->raw;
 	struct dentry *dentry;
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(fsc, &__ji);
 
 	if (fh_type == FILEID_BTRFS_WITH_PARENT) {
 		struct ceph_nfs_snapfh *sfh = (void *)fid->raw;
+		ceph_blog_exit(&__ji);
 		return __snapfh_to_dentry(sb, sfh, true);
 	}
 
-	if (fh_type != FILEID_INO32_GEN_PARENT)
+	if (fh_type != FILEID_INO32_GEN_PARENT) {
+		ceph_blog_exit(&__ji);
 		return NULL;
-	if (fh_len < sizeof(*cfh) / BYTES_PER_U32)
+	}
+	if (fh_len < sizeof(*cfh) / BYTES_PER_U32) {
+		ceph_blog_exit(&__ji);
 		return NULL;
+	}
 
-	doutc(fsc->client, "%llx\n", cfh->parent_ino);
+	boutc(fsc->client, "%llx\n", cfh->parent_ino);
 	dentry = __get_parent(sb, NULL, cfh->ino);
 	if (unlikely(dentry == ERR_PTR(-ENOENT)))
 		dentry = __fh_to_dentry(sb, cfh->parent_ino);
+	ceph_blog_exit(&__ji);
 	return dentry;
 }
 
@@ -549,7 +585,7 @@ static int __get_snap_name(struct dentry *parent, char *name,
 	if (req)
 		ceph_mdsc_put_request(req);
 	kfree(last_name);
-	doutc(fsc->client, "child dentry %p %p %llx.%llx err=%d\n", child,
+	boutc(fsc->client, "child dentry %p %p %llx.%llx err=%d\n", child,
 	      inode, ceph_vinop(inode), err);
 	return err;
 }
@@ -561,17 +597,25 @@ static int ceph_get_name(struct dentry *parent, char *name,
 	struct ceph_mds_request *req;
 	struct inode *dir = d_inode(parent);
 	struct inode *inode = d_inode(child);
+	struct ceph_fs_client *fsc = ceph_inode_to_fs_client(inode);
 	struct ceph_mds_reply_info_parsed *rinfo;
 	int err;
+	struct ceph_journal_info __ji;
 
-	if (ceph_in_snap(inode))
+	ceph_blog_enter(fsc, &__ji);
+
+	if (ceph_in_snap(inode)) {
+		ceph_blog_exit(&__ji);
 		return __get_snap_name(parent, name, child);
+	}
 
 	mdsc = ceph_inode_to_fs_client(inode)->mdsc;
 	req = ceph_mdsc_create_request(mdsc, CEPH_MDS_OP_LOOKUPNAME,
 				       USE_ANY_MDS);
-	if (IS_ERR(req))
+	if (IS_ERR(req)) {
+		ceph_blog_exit(&__ji);
 		return PTR_ERR(req);
+	}
 
 	inode_lock(dir);
 	req->r_inode = inode;
@@ -610,10 +654,11 @@ static int ceph_get_name(struct dentry *parent, char *name,
 		ceph_fname_free_buffer(dir, &oname);
 	}
 out:
-	doutc(mdsc->fsc->client, "child dentry %p %p %llx.%llx err %d %s%s\n",
+	boutc(mdsc->fsc->client, "child dentry %p %p %llx.%llx err %d %s%s\n",
 	      child, inode, ceph_vinop(inode), err, err ? "" : "name ",
 	      err ? "" : name);
 	ceph_mdsc_put_request(req);
+	ceph_blog_exit(&__ji);
 	return err;
 }
 
diff --git a/fs/ceph/inode.c b/fs/ceph/inode.c
index 13c42c301e02..94f5215f8eed 100644
--- a/fs/ceph/inode.c
+++ b/fs/ceph/inode.c
@@ -264,7 +264,7 @@ struct inode *ceph_get_snapdir(struct inode *parent)
 			inode->i_flags |= S_ENCRYPTED;
 			ci->fscrypt_auth_len = pci->fscrypt_auth_len;
 		} else {
-			doutc(cl, "Failed to alloc snapdir fscrypt_auth\n");
+			boutc(cl, "Failed to alloc snapdir fscrypt_auth\n");
 			ret = -ENOMEM;
 			goto err;
 		}
@@ -342,7 +342,7 @@ static struct ceph_inode_frag *__get_or_create_frag(struct ceph_inode_info *ci,
 	rb_link_node(&frag->node, parent, p);
 	rb_insert_color(&frag->node, &ci->i_fragtree);
 
-	doutc(cl, "added %p %llx.%llx frag %x\n", inode, ceph_vinop(inode), f);
+	boutc(cl, "added %p %llx.%llx frag %x\n", inode, ceph_vinop(inode), f);
 	return frag;
 }
 
@@ -399,7 +399,7 @@ static u32 __ceph_choose_frag(struct ceph_inode_info *ci, u32 v,
 
 		/* choose child */
 		nway = 1 << frag->split_by;
-		doutc(cl, "frag(%x) %x splits by %d (%d ways)\n", v, t,
+		boutc(cl, "frag(%x) %x splits by %d (%d ways)\n", v, t,
 		      frag->split_by, nway);
 		for (i = 0; i < nway; i++) {
 			n = ceph_frag_make_child(t, frag->split_by, i);
@@ -410,7 +410,7 @@ static u32 __ceph_choose_frag(struct ceph_inode_info *ci, u32 v,
 		}
 		BUG_ON(i == nway);
 	}
-	doutc(cl, "frag(%x) = %x\n", v, t);
+	boutc(cl, "frag(%x) = %x\n", v, t);
 
 	return t;
 }
@@ -459,13 +459,13 @@ static int ceph_fill_dirfrag(struct inode *inode,
 			goto out;
 		if (frag->split_by == 0) {
 			/* tree leaf, remove */
-			doutc(cl, "removed %p %llx.%llx frag %x (no ref)\n",
+			boutc(cl, "removed %p %llx.%llx frag %x (no ref)\n",
 			      inode, ceph_vinop(inode), id);
 			rb_erase(&frag->node, &ci->i_fragtree);
 			kfree(frag);
 		} else {
 			/* tree branch, keep and clear */
-			doutc(cl, "cleared %p %llx.%llx frag %x referral\n",
+			boutc(cl, "cleared %p %llx.%llx frag %x referral\n",
 			      inode, ceph_vinop(inode), id);
 			frag->mds = -1;
 			frag->ndist = 0;
@@ -490,7 +490,7 @@ static int ceph_fill_dirfrag(struct inode *inode,
 	frag->ndist = min_t(u32, ndist, CEPH_MAX_DIRFRAG_REP);
 	for (i = 0; i < frag->ndist; i++)
 		frag->dist[i] = le32_to_cpu(dirinfo->dist[i]);
-	doutc(cl, "%p %llx.%llx frag %x ndist=%d\n", inode,
+	boutc(cl, "%p %llx.%llx frag %x ndist=%d\n", inode,
 	      ceph_vinop(inode), frag->frag, frag->ndist);
 
 out:
@@ -555,7 +555,7 @@ static int ceph_fill_fragtree(struct inode *inode,
 		     frag_tree_split_cmp, NULL);
 	}
 
-	doutc(cl, "%p %llx.%llx\n", inode, ceph_vinop(inode));
+	boutc(cl, "%p %llx.%llx\n", inode, ceph_vinop(inode));
 	rb_node = rb_first(&ci->i_fragtree);
 	for (i = 0; i < nsplits; i++) {
 		id = le32_to_cpu(fragtree->splits[i].frag);
@@ -595,7 +595,7 @@ static int ceph_fill_fragtree(struct inode *inode,
 		if (frag->split_by == 0)
 			ci->i_fragtree_nsplits++;
 		frag->split_by = split_by;
-		doutc(cl, " frag %x split by %d\n", frag->frag, frag->split_by);
+		boutc(cl, " frag %x split by %d\n", frag->frag, frag->split_by);
 		prev_frag = frag;
 	}
 	while (rb_node) {
@@ -744,13 +744,17 @@ void ceph_free_inode(struct inode *inode)
 
 void ceph_evict_inode(struct inode *inode)
 {
+	struct ceph_fs_client *fsc = ceph_inode_to_fs_client(inode);
 	struct ceph_inode_info *ci = ceph_inode(inode);
 	struct ceph_mds_client *mdsc = ceph_sb_to_mdsc(inode->i_sb);
 	struct ceph_client *cl = ceph_inode_to_client(inode);
 	struct ceph_inode_frag *frag;
 	struct rb_node *n;
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(fsc, &__ji);
 
-	doutc(cl, "%p ino %llx.%llx\n", inode, ceph_vinop(inode));
+	boutc(cl, "%p ino %llx.%llx\n", inode, ceph_vinop(inode));
 
 	percpu_counter_dec(&mdsc->metric.total_inodes);
 
@@ -776,7 +780,7 @@ void ceph_evict_inode(struct inode *inode)
 	 */
 	if (ci->i_snap_realm) {
 		if (!ceph_in_snap(inode)) {
-			doutc(cl, " dropping residual ref to snap realm %p\n",
+			boutc(cl, " dropping residual ref to snap realm %p\n",
 			      ci->i_snap_realm);
 			ceph_change_snap_realm(inode, NULL);
 		} else {
@@ -800,6 +804,7 @@ void ceph_evict_inode(struct inode *inode)
 
 	ceph_put_string(rcu_dereference_raw(ci->i_layout.pool_ns));
 	ceph_put_string(rcu_dereference_raw(ci->i_cached_layout.pool_ns));
+	ceph_blog_exit(&__ji);
 }
 
 /*
@@ -884,7 +889,7 @@ int ceph_fill_file_size(struct inode *inode, int issued,
 
 	if (ceph_seq_cmp(truncate_seq, ci->i_truncate_seq) > 0 ||
 	    (truncate_seq == ci->i_truncate_seq && size > isize)) {
-		doutc(cl, "size %lld -> %llu\n", isize, size);
+		boutc(cl, "size %lld -> %llu\n", isize, size);
 		if (size > 0 && S_ISDIR(inode->i_mode)) {
 			pr_err_client(cl, "non-zero size for directory\n");
 			size = 0;
@@ -899,7 +904,7 @@ int ceph_fill_file_size(struct inode *inode, int issued,
 			ceph_fscache_update(inode);
 		ci->i_reported_size = size;
 		if (truncate_seq != ci->i_truncate_seq) {
-			doutc(cl, "truncate_seq %u -> %u\n",
+			boutc(cl, "truncate_seq %u -> %u\n",
 			      ci->i_truncate_seq, truncate_seq);
 			ci->i_truncate_seq = truncate_seq;
 
@@ -929,7 +934,7 @@ int ceph_fill_file_size(struct inode *inode, int issued,
 	 * anyway.
 	 */
 	if (ceph_seq_cmp(truncate_seq, ci->i_truncate_seq) >= 0) {
-		doutc(cl, "truncate_size %lld -> %llu, encrypted %d\n",
+		boutc(cl, "truncate_size %lld -> %llu, encrypted %d\n",
 		      ci->i_truncate_size, truncate_size,
 		      !!IS_ENCRYPTED(inode));
 
@@ -937,7 +942,7 @@ int ceph_fill_file_size(struct inode *inode, int issued,
 
 #ifdef CONFIG_FS_ENCRYPTION
 		if (IS_ENCRYPTED(inode)) {
-			doutc(cl, "truncate_pagecache_size %lld -> %llu\n",
+			boutc(cl, "truncate_pagecache_size %lld -> %llu\n",
 			      ci->i_truncate_pagecache_size, size);
 			ci->i_truncate_pagecache_size = size;
 		} else {
@@ -1000,13 +1005,13 @@ void ceph_fill_file_time(struct inode *inode, int issued,
 		      CEPH_CAP_XATTR_EXCL)) {
 		if (ci->i_version == 0 ||
 		    timespec64_compare(ctime, &ictime) > 0) {
-			doutc(cl, "ctime %ptSp -> %ptSp inc w/ cap\n", &ictime, ctime);
+			boutc(cl, "ctime %ptSp -> %ptSp inc w/ cap\n", &ictime, ctime);
 			inode_set_ctime_to_ts(inode, *ctime);
 		}
 		if (ci->i_version == 0 ||
 		    ceph_seq_cmp(time_warp_seq, ci->i_time_warp_seq) > 0) {
 			/* the MDS did a utimes() */
-			doutc(cl, "mtime %ptSp -> %ptSp tw %d -> %d\n", &imtime, mtime,
+			boutc(cl, "mtime %ptSp -> %ptSp tw %d -> %d\n", &imtime, mtime,
 			      ci->i_time_warp_seq, time_warp_seq);
 
 			inode_set_mtime_to_ts(inode, *mtime);
@@ -1015,11 +1020,11 @@ void ceph_fill_file_time(struct inode *inode, int issued,
 		} else if (time_warp_seq == ci->i_time_warp_seq) {
 			/* nobody did utimes(); take the max */
 			if (timespec64_compare(mtime, &imtime) > 0) {
-				doutc(cl, "mtime %ptSp -> %ptSp inc\n", &imtime, mtime);
+				boutc(cl, "mtime %ptSp -> %ptSp inc\n", &imtime, mtime);
 				inode_set_mtime_to_ts(inode, *mtime);
 			}
 			if (timespec64_compare(atime, &iatime) > 0) {
-				doutc(cl, "atime %ptSp -> %ptSp inc\n", &iatime, atime);
+				boutc(cl, "atime %ptSp -> %ptSp inc\n", &iatime, atime);
 				inode_set_atime_to_ts(inode, *atime);
 			}
 		} else if (issued & CEPH_CAP_FILE_EXCL) {
@@ -1039,7 +1044,7 @@ void ceph_fill_file_time(struct inode *inode, int issued,
 		}
 	}
 	if (warn) /* time_warp_seq shouldn't go backwards */
-		doutc(cl, "%p mds time_warp_seq %u < %u\n", inode,
+		boutc(cl, "%p mds time_warp_seq %u < %u\n", inode,
 		      time_warp_seq, ci->i_time_warp_seq);
 }
 
@@ -1107,7 +1112,7 @@ int ceph_fill_inode(struct inode *inode, struct folio *locked_folio,
 
 	lockdep_assert_held(&mdsc->snap_rwsem);
 
-	doutc(cl, "%p ino %llx.%llx v %llu had %llu\n", inode, ceph_vinop(inode),
+	boutc(cl, "%p ino %llx.%llx v %llu had %llu\n", inode, ceph_vinop(inode),
 	      le64_to_cpu(info->version), ci->i_version);
 
 	/* Once I_NEW is cleared, we can't change type or dev numbers */
@@ -1204,7 +1209,7 @@ int ceph_fill_inode(struct inode *inode, struct folio *locked_folio,
 		inode->i_mode = mode;
 		inode->i_uid = make_kuid(&init_user_ns, le32_to_cpu(info->uid));
 		inode->i_gid = make_kgid(&init_user_ns, le32_to_cpu(info->gid));
-		doutc(cl, "%p %llx.%llx mode 0%o uid.gid %d.%d\n", inode,
+		boutc(cl, "%p %llx.%llx mode 0%o uid.gid %d.%d\n", inode,
 		      ceph_vinop(inode), inode->i_mode,
 		      from_kuid(&init_user_ns, inode->i_uid),
 		      from_kgid(&init_user_ns, inode->i_gid));
@@ -1276,7 +1281,7 @@ int ceph_fill_inode(struct inode *inode, struct folio *locked_folio,
 		/* only update max_size on auth cap */
 		if ((info->cap.flags & CEPH_CAP_FLAG_AUTH) &&
 		    ci->i_max_size != le64_to_cpu(info->max_size)) {
-			doutc(cl, "max_size %lld -> %llu\n",
+			boutc(cl, "max_size %lld -> %llu\n",
 			    ci->i_max_size, le64_to_cpu(info->max_size));
 			ci->i_max_size = le64_to_cpu(info->max_size);
 		}
@@ -1417,7 +1422,7 @@ int ceph_fill_inode(struct inode *inode, struct folio *locked_folio,
 			    (info_caps & CEPH_CAP_FILE_SHARED) &&
 			    (issued & CEPH_CAP_FILE_EXCL) == 0 &&
 			    !__ceph_dir_is_complete(ci)) {
-				doutc(cl, " marking %p complete (empty)\n",
+				boutc(cl, " marking %p complete (empty)\n",
 				      inode);
 				i_size_write(inode, 0);
 				__ceph_dir_set_complete(ci,
@@ -1427,7 +1432,7 @@ int ceph_fill_inode(struct inode *inode, struct folio *locked_folio,
 
 			wake = true;
 		} else {
-			doutc(cl, " %p got snap_caps %s\n", inode,
+			boutc(cl, " %p got snap_caps %s\n", inode,
 			      ceph_cap_string(info_caps));
 			ci->i_snap_caps |= info_caps;
 		}
@@ -1498,7 +1503,7 @@ static void __update_dentry_lease(struct inode *dir, struct dentry *dentry,
 	long unsigned ttl = from_time + (duration * HZ) / 1000;
 	long unsigned half_ttl = from_time + (duration * HZ / 2) / 1000;
 
-	doutc(cl, "%p duration %lu ms ttl %lu\n", dentry, duration, ttl);
+	boutc(cl, "%p duration %lu ms ttl %lu\n", dentry, duration, ttl);
 
 	/* only track leases on regular dentries */
 	if (ceph_in_snap(dir))
@@ -1637,14 +1642,14 @@ static int splice_dentry(struct dentry **pdn, struct inode *in)
 	}
 
 	if (realdn) {
-		doutc(cl, "dn %p (%d) spliced with %p (%d) inode %p ino %llx.%llx\n",
+		boutc(cl, "dn %p (%d) spliced with %p (%d) inode %p ino %llx.%llx\n",
 		      dn, d_count(dn), realdn, d_count(realdn),
 		      d_inode(realdn), ceph_vinop(d_inode(realdn)));
 		dput(dn);
 		*pdn = realdn;
 	} else {
 		BUG_ON(!ceph_dentry(dn));
-		doutc(cl, "dn %p attached to %p ino %llx.%llx\n", dn,
+		boutc(cl, "dn %p attached to %p ino %llx.%llx\n", dn,
 		      d_inode(dn), ceph_vinop(d_inode(dn)));
 	}
 	return 0;
@@ -1672,11 +1677,11 @@ int ceph_fill_trace(struct super_block *sb, struct ceph_mds_request *req)
 	struct inode *parent_dir = NULL;
 	int err = 0;
 
-	doutc(cl, "%p is_dentry %d is_target %d\n", req,
+	boutc(cl, "%p is_dentry %d is_target %d\n", req,
 	      rinfo->head->is_dentry, rinfo->head->is_target);
 
 	if (!rinfo->head->is_target && !rinfo->head->is_dentry) {
-		doutc(cl, "reply is empty!\n");
+		boutc(cl, "reply is empty!\n");
 		if (rinfo->head->result == 0 && req->r_parent)
 			ceph_invalidate_dir_request(req);
 		return 0;
@@ -1742,13 +1747,20 @@ int ceph_fill_trace(struct super_block *sb, struct ceph_mds_request *req)
 			tvino.snap = le64_to_cpu(rinfo->targeti.in->snapid);
 retry_lookup:
 			dn = d_lookup(parent, &dname);
-			doutc(cl, "d_lookup on parent=%p name=%.*s got %p\n",
-			      parent, dname.len, dname.name, dn);
+			boutc_bounded(cl,
+				      "d_lookup on parent=%p name=%.*s got %p\n",
+				      (parent, dname.len,
+				       BLOG_STR(dname.name, dname.len), dn),
+				      (parent, dname.len,
+				       (const char *)dname.name, dn));
 
 			if (!dn) {
 				dn = d_alloc(parent, &dname);
-				doutc(cl, "d_alloc %p '%.*s' = %p\n", parent,
-				      dname.len, dname.name, dn);
+				boutc_bounded(cl, "d_alloc %p '%.*s' = %p\n",
+					      (parent, dname.len,
+					       BLOG_STR(dname.name, dname.len), dn),
+					      (parent, dname.len,
+					       (const char *)dname.name, dn));
 				if (!dn) {
 					dput(parent);
 					ceph_fname_free_buffer(parent_dir, &oname);
@@ -1764,7 +1776,7 @@ int ceph_fill_trace(struct super_block *sb, struct ceph_mds_request *req)
 			} else if (d_really_is_positive(dn) &&
 				   (ceph_ino(d_inode(dn)) != tvino.ino ||
 				    ceph_snap(d_inode(dn)) != tvino.snap)) {
-				doutc(cl, " dn %p points to wrong inode %p\n",
+				boutc(cl, " dn %p points to wrong inode %p\n",
 				      dn, d_inode(dn));
 				ceph_dir_clear_ordered(parent_dir);
 				d_delete(dn);
@@ -1842,30 +1854,30 @@ int ceph_fill_trace(struct super_block *sb, struct ceph_mds_request *req)
 		have_lease = have_dir_cap ||
 			le32_to_cpu(rinfo->dlease->duration_ms);
 		if (!have_lease)
-			doutc(cl, "no dentry lease or dir cap\n");
+			boutc(cl, "no dentry lease or dir cap\n");
 
 		/* rename? */
 		if (req->r_old_dentry && req->r_op == CEPH_MDS_OP_RENAME) {
 			struct inode *olddir = req->r_old_dentry_dir;
 			BUG_ON(!olddir);
 
-			doutc(cl, " src %p '%pd' dst %p '%pd'\n",
-			      req->r_old_dentry, req->r_old_dentry, dn, dn);
-			doutc(cl, "doing d_move %p -> %p\n", req->r_old_dentry, dn);
+			boutc(cl, " src %p '%s' dst %p '%s'\n",
+			      req->r_old_dentry, req->r_old_dentry->d_name.name, dn, dn->d_name.name);
+			boutc(cl, "doing d_move %p -> %p\n", req->r_old_dentry, dn);
 
 			/* d_move screws up sibling dentries' offsets */
 			ceph_dir_clear_ordered(dir);
 			ceph_dir_clear_ordered(olddir);
 
 			d_move(req->r_old_dentry, dn);
-			doutc(cl, " src %p '%pd' dst %p '%pd'\n",
-			      req->r_old_dentry, req->r_old_dentry, dn, dn);
+			boutc(cl, " src %p '%s' dst %p '%s'\n",
+			      req->r_old_dentry, req->r_old_dentry->d_name.name, dn, dn->d_name.name);
 
 			/* ensure target dentry is invalidated, despite
 			   rehashing bug in vfs_rename_dir */
 			ceph_invalidate_dentry_lease(dn);
 
-			doutc(cl, "dn %p gets new offset %lld\n",
+			boutc(cl, "dn %p gets new offset %lld\n",
 			      req->r_old_dentry,
 			      ceph_dentry(req->r_old_dentry)->offset);
 
@@ -1879,9 +1891,9 @@ int ceph_fill_trace(struct super_block *sb, struct ceph_mds_request *req)
 
 		/* null dentry? */
 		if (!rinfo->head->is_target) {
-			doutc(cl, "null dentry\n");
+			boutc(cl, "null dentry\n");
 			if (d_really_is_positive(dn)) {
-				doutc(cl, "d_delete %p\n", dn);
+				boutc(cl, "d_delete %p\n", dn);
 				ceph_dir_clear_ordered(dir);
 				d_delete(dn);
 			} else if (have_lease) {
@@ -1911,7 +1923,7 @@ int ceph_fill_trace(struct super_block *sb, struct ceph_mds_request *req)
 				goto done;
 			dn = req->r_dentry;  /* may have spliced */
 		} else if (d_really_is_positive(dn) && d_inode(dn) != in) {
-			doutc(cl, " %p links to %p %llx.%llx, not %llx.%llx\n",
+			boutc(cl, " %p links to %p %llx.%llx, not %llx.%llx\n",
 			      dn, d_inode(dn), ceph_vinop(d_inode(dn)),
 			      ceph_vinop(in));
 			d_invalidate(dn);
@@ -1923,7 +1935,7 @@ int ceph_fill_trace(struct super_block *sb, struct ceph_mds_request *req)
 					    rinfo->dlease, session,
 					    req->r_request_started);
 		}
-		doutc(cl, " final dn %p\n", dn);
+		boutc(cl, " final dn %p\n", dn);
 	} else if ((req->r_op == CEPH_MDS_OP_LOOKUPSNAP ||
 		    req->r_op == CEPH_MDS_OP_MKSNAP) &&
 	           test_bit(CEPH_MDS_R_PARENT_LOCKED, &req->r_req_flags) &&
@@ -1934,7 +1946,7 @@ int ceph_fill_trace(struct super_block *sb, struct ceph_mds_request *req)
 		BUG_ON(!dir);
 		BUG_ON(ceph_snap(dir) != CEPH_SNAPDIR);
 		BUG_ON(!req->r_dentry);
-		doutc(cl, " linking snapped dir %p to dn %p\n", in,
+		boutc(cl, " linking snapped dir %p to dn %p\n", in,
 		      req->r_dentry);
 		ceph_dir_clear_ordered(dir);
 
@@ -1966,7 +1978,7 @@ int ceph_fill_trace(struct super_block *sb, struct ceph_mds_request *req)
 	/* Drop extra ref from ceph_get_reply_dir() if it returned a new inode */
 	if (unlikely(!IS_ERR_OR_NULL(parent_dir) && parent_dir != req->r_parent))
 		iput(parent_dir);
-	doutc(cl, "done err=%d\n", err);
+	boutc(cl, "done err=%d\n", err);
 	return err;
 }
 
@@ -1992,7 +2004,7 @@ static int readdir_prepopulate_inodes_only(struct ceph_mds_request *req,
 		in = ceph_get_inode(req->r_dentry->d_sb, vino, NULL);
 		if (IS_ERR(in)) {
 			err = PTR_ERR(in);
-			doutc(cl, "badness got %d\n", err);
+			boutc(cl, "badness got %d\n", err);
 			continue;
 		}
 		rc = ceph_fill_inode(in, NULL, &rde->inode, NULL, session,
@@ -2059,11 +2071,11 @@ static int fill_readdir_cache(struct inode *dir, struct dentry *dn,
 
 	if (req->r_dir_release_cnt == atomic64_read(&ci->i_release_count) &&
 	    req->r_dir_ordered_cnt == atomic64_read(&ci->i_ordered_count)) {
-		doutc(cl, "dn %p idx %d\n", dn, ctl->index);
+		boutc(cl, "dn %p idx %d\n", dn, ctl->index);
 		ctl->dentries[idx] = dn;
 		ctl->index++;
 	} else {
-		doutc(cl, "disable readdir cache\n");
+		boutc(cl, "disable readdir cache\n");
 		ctl->index = -1;
 	}
 	return 0;
@@ -2104,7 +2116,7 @@ int ceph_readdir_prepopulate(struct ceph_mds_request *req,
 
 	if (rinfo->dir_dir &&
 	    le32_to_cpu(rinfo->dir_dir->frag) != frag) {
-		doutc(cl, "got new frag %x -> %x\n", frag,
+		boutc(cl, "got new frag %x -> %x\n", frag,
 			    le32_to_cpu(rinfo->dir_dir->frag));
 		frag = le32_to_cpu(rinfo->dir_dir->frag);
 		if (!rinfo->hash_order)
@@ -2112,10 +2124,10 @@ int ceph_readdir_prepopulate(struct ceph_mds_request *req,
 	}
 
 	if (le32_to_cpu(rinfo->head->op) == CEPH_MDS_OP_LSSNAP) {
-		doutc(cl, "%d items under SNAPDIR dn %p\n",
+		boutc(cl, "%d items under SNAPDIR dn %p\n",
 		      rinfo->dir_nr, parent);
 	} else {
-		doutc(cl, "%d items under dn %p\n", rinfo->dir_nr, parent);
+		boutc(cl, "%d items under dn %p\n", rinfo->dir_nr, parent);
 		if (rinfo->dir_dir)
 			ceph_fill_dirfrag(d_inode(parent), rinfo->dir_dir);
 
@@ -2159,15 +2171,20 @@ int ceph_readdir_prepopulate(struct ceph_mds_request *req,
 
 retry_lookup:
 		dn = d_lookup(parent, &dname);
-		doutc(cl, "d_lookup on parent=%p name=%.*s got %p\n",
-		      parent, dname.len, dname.name, dn);
+		boutc_bounded(cl, "d_lookup on parent=%p name=%.*s got %p\n",
+			      (parent, dname.len,
+			       BLOG_STR(dname.name, dname.len), dn),
+			      (parent, dname.len, (const char *)dname.name, dn));
 
 		if (!dn) {
 			dn = d_alloc(parent, &dname);
-			doutc(cl, "d_alloc %p '%.*s' = %p\n", parent,
-			      dname.len, dname.name, dn);
+			boutc_bounded(cl, "d_alloc %p '%.*s' = %p\n",
+				      (parent, dname.len,
+				       BLOG_STR(dname.name, dname.len), dn),
+				      (parent, dname.len,
+				       (const char *)dname.name, dn));
 			if (!dn) {
-				doutc(cl, "d_alloc badness\n");
+				boutc(cl, "d_alloc badness\n");
 				err = -ENOMEM;
 				goto out;
 			}
@@ -2180,7 +2197,7 @@ int ceph_readdir_prepopulate(struct ceph_mds_request *req,
 			   (ceph_ino(d_inode(dn)) != tvino.ino ||
 			    ceph_snap(d_inode(dn)) != tvino.snap)) {
 			struct ceph_dentry_info *di = ceph_dentry(dn);
-			doutc(cl, " dn %p points to wrong inode %p\n",
+			boutc(cl, " dn %p points to wrong inode %p\n",
 			      dn, d_inode(dn));
 
 			spin_lock(&dn->d_lock);
@@ -2203,7 +2220,7 @@ int ceph_readdir_prepopulate(struct ceph_mds_request *req,
 		} else {
 			in = ceph_get_inode(parent->d_sb, tvino, NULL);
 			if (IS_ERR(in)) {
-				doutc(cl, "new_inode badness\n");
+				boutc(cl, "new_inode badness\n");
 				d_drop(dn);
 				dput(dn);
 				err = PTR_ERR(in);
@@ -2232,7 +2249,7 @@ int ceph_readdir_prepopulate(struct ceph_mds_request *req,
 
 		if (d_really_is_negative(dn)) {
 			if (ceph_security_xattr_deadlock(in)) {
-				doutc(cl, " skip splicing dn %p to inode %p"
+				boutc(cl, " skip splicing dn %p to inode %p"
 				      " (security xattr deadlock)\n", dn, in);
 				iput(in);
 				skipped++;
@@ -2265,7 +2282,7 @@ int ceph_readdir_prepopulate(struct ceph_mds_request *req,
 		req->r_readdir_cache_idx = cache_ctl.index;
 	}
 	ceph_readdir_cache_release(&cache_ctl);
-	doutc(cl, "done\n");
+	boutc(cl, "done\n");
 	return err;
 }
 
@@ -2276,7 +2293,7 @@ bool ceph_inode_set_size(struct inode *inode, loff_t size)
 	bool ret;
 
 	spin_lock(&ci->i_ceph_lock);
-	doutc(cl, "set_size %p %llu -> %llu\n", inode, i_size_read(inode), size);
+	boutc(cl, "set_size %p %llu -> %llu\n", inode, i_size_read(inode), size);
 	i_size_write(inode, size);
 	ceph_fscache_update(inode);
 	inode->i_blocks = calc_inode_blocks(size);
@@ -2328,7 +2345,7 @@ static void ceph_do_invalidate_pages(struct inode *inode)
 	}
 
 	spin_lock(&ci->i_ceph_lock);
-	doutc(cl, "%p %llx.%llx gen %d revoking %d\n", inode,
+	boutc(cl, "%p %llx.%llx gen %d revoking %d\n", inode,
 	      ceph_vinop(inode), ci->i_rdcache_gen, ci->i_rdcache_revoking);
 	if (ci->i_rdcache_revoking != ci->i_rdcache_gen) {
 		if (__ceph_caps_revoking_other(ci, NULL, CEPH_CAP_FILE_CACHE))
@@ -2348,12 +2365,12 @@ static void ceph_do_invalidate_pages(struct inode *inode)
 	spin_lock(&ci->i_ceph_lock);
 	if (orig_gen == ci->i_rdcache_gen &&
 	    orig_gen == ci->i_rdcache_revoking) {
-		doutc(cl, "%p %llx.%llx gen %d successful\n", inode,
+		boutc(cl, "%p %llx.%llx gen %d successful\n", inode,
 		      ceph_vinop(inode), ci->i_rdcache_gen);
 		ci->i_rdcache_revoking--;
 		check = 1;
 	} else {
-		doutc(cl, "%p %llx.%llx gen %d raced, now %d revoking %d\n",
+		boutc(cl, "%p %llx.%llx gen %d raced, now %d revoking %d\n",
 		      inode, ceph_vinop(inode), orig_gen, ci->i_rdcache_gen,
 		      ci->i_rdcache_revoking);
 		if (__ceph_caps_revoking_other(ci, NULL, CEPH_CAP_FILE_CACHE))
@@ -2381,7 +2398,7 @@ void __ceph_do_pending_vmtruncate(struct inode *inode)
 retry:
 	spin_lock(&ci->i_ceph_lock);
 	if (ci->i_truncate_pending == 0) {
-		doutc(cl, "%p %llx.%llx none pending\n", inode,
+		boutc(cl, "%p %llx.%llx none pending\n", inode,
 		      ceph_vinop(inode));
 		spin_unlock(&ci->i_ceph_lock);
 		mutex_unlock(&ci->i_truncate_mutex);
@@ -2394,7 +2411,7 @@ void __ceph_do_pending_vmtruncate(struct inode *inode)
 	 */
 	if (ci->i_wrbuffer_ref_head < ci->i_wrbuffer_ref) {
 		spin_unlock(&ci->i_ceph_lock);
-		doutc(cl, "%p %llx.%llx flushing snaps first\n", inode,
+		boutc(cl, "%p %llx.%llx flushing snaps first\n", inode,
 		      ceph_vinop(inode));
 		filemap_write_and_wait_range(&inode->i_data, 0,
 					     inode->i_sb->s_maxbytes);
@@ -2406,7 +2423,7 @@ void __ceph_do_pending_vmtruncate(struct inode *inode)
 
 	to = ceph_get_truncate_pagecache_size(ci);
 	wrbuffer_refs = ci->i_wrbuffer_ref;
-	doutc(cl, "%p %llx.%llx (%d) to %lld\n", inode, ceph_vinop(inode),
+	boutc(cl, "%p %llx.%llx (%d) to %lld\n", inode, ceph_vinop(inode),
 	      ci->i_truncate_pending, to);
 	spin_unlock(&ci->i_ceph_lock);
 
@@ -2435,10 +2452,14 @@ static void ceph_inode_work(struct work_struct *work)
 	struct ceph_inode_info *ci = container_of(work, struct ceph_inode_info,
 						 i_work);
 	struct inode *inode = &ci->netfs.inode;
+	struct ceph_fs_client *fsc = ceph_inode_to_fs_client(inode);
 	struct ceph_client *cl = ceph_inode_to_client(inode);
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(fsc, &__ji);
 
 	if (test_and_clear_bit(CEPH_I_WORK_WRITEBACK, &ci->i_work_mask)) {
-		doutc(cl, "writeback %p %llx.%llx\n", inode, ceph_vinop(inode));
+		boutc(cl, "writeback %p %llx.%llx\n", inode, ceph_vinop(inode));
 		filemap_fdatawrite(&inode->i_data);
 	}
 	if (test_and_clear_bit(CEPH_I_WORK_INVALIDATE_PAGES, &ci->i_work_mask))
@@ -2453,6 +2474,7 @@ static void ceph_inode_work(struct work_struct *work)
 	if (test_and_clear_bit(CEPH_I_WORK_FLUSH_SNAPS, &ci->i_work_mask))
 		ceph_flush_snaps(ci, NULL);
 
+	ceph_blog_exit(&__ji);
 	iput(inode);
 }
 
@@ -2533,7 +2555,7 @@ static int fill_fscrypt_truncate(struct inode *inode,
 
 	issued = __ceph_caps_issued(ci, NULL);
 
-	doutc(cl, "size %lld -> %lld got cap refs on %s, issued %s\n",
+	boutc(cl, "size %lld -> %lld got cap refs on %s, issued %s\n",
 	      i_size, attr->ia_size, ceph_cap_string(got),
 	      ceph_cap_string(issued));
 
@@ -2590,7 +2612,7 @@ static int fill_fscrypt_truncate(struct inode *inode,
 	 * If the Rados object doesn't exist, it will be set to 0.
 	 */
 	if (!objver) {
-		doutc(cl, "hit hole, ppos %lld < size %lld\n", pos, i_size);
+		boutc(cl, "hit hole, ppos %lld < size %lld\n", pos, i_size);
 
 		header.data_len = cpu_to_le32(8 + 8 + 4);
 		header.file_offset = 0;
@@ -2599,7 +2621,7 @@ static int fill_fscrypt_truncate(struct inode *inode,
 		header.data_len = cpu_to_le32(8 + 8 + 4 + CEPH_FSCRYPT_BLOCK_SIZE);
 		header.file_offset = cpu_to_le64(orig_pos);
 
-		doutc(cl, "encrypt block boff/bsize %d/%lu\n", boff,
+		boutc(cl, "encrypt block boff/bsize %d/%lu\n", boff,
 		      CEPH_FSCRYPT_BLOCK_SIZE);
 
 		/* truncate and zero out the extra contents for the last block */
@@ -2627,7 +2649,7 @@ static int fill_fscrypt_truncate(struct inode *inode,
 	}
 	req->r_pagelist = pagelist;
 out:
-	doutc(cl, "%p %llx.%llx size dropping cap refs on %s\n", inode,
+	boutc(cl, "%p %llx.%llx size dropping cap refs on %s\n", inode,
 	      ceph_vinop(inode), ceph_cap_string(got));
 	ceph_put_cap_refs(ci, got);
 	if (iov.iov_base)
@@ -2712,7 +2734,7 @@ int __ceph_setattr(struct mnt_idmap *idmap, struct inode *inode,
 		}
 	}
 
-	doutc(cl, "%p %llx.%llx issued %s\n", inode, ceph_vinop(inode),
+	boutc(cl, "%p %llx.%llx issued %s\n", inode, ceph_vinop(inode),
 	      ceph_cap_string(issued));
 #if IS_ENABLED(CONFIG_FS_ENCRYPTION)
 	if (cia && cia->fscrypt_auth) {
@@ -2724,7 +2746,7 @@ int __ceph_setattr(struct mnt_idmap *idmap, struct inode *inode,
 			goto out;
 		}
 
-		doutc(cl, "%p %llx.%llx fscrypt_auth len %u to %u)\n", inode,
+		boutc(cl, "%p %llx.%llx fscrypt_auth len %u to %u)\n", inode,
 		      ceph_vinop(inode), ci->fscrypt_auth_len, len);
 
 		/* It should never be re-set once set */
@@ -2755,7 +2777,7 @@ int __ceph_setattr(struct mnt_idmap *idmap, struct inode *inode,
 	if (ia_valid & ATTR_UID) {
 		kuid_t fsuid = from_vfsuid(idmap, i_user_ns(inode), attr->ia_vfsuid);
 
-		doutc(cl, "%p %llx.%llx uid %d -> %d\n", inode,
+		boutc(cl, "%p %llx.%llx uid %d -> %d\n", inode,
 		      ceph_vinop(inode),
 		      from_kuid(&init_user_ns, inode->i_uid),
 		      from_kuid(&init_user_ns, attr->ia_uid));
@@ -2773,7 +2795,7 @@ int __ceph_setattr(struct mnt_idmap *idmap, struct inode *inode,
 	if (ia_valid & ATTR_GID) {
 		kgid_t fsgid = from_vfsgid(idmap, i_user_ns(inode), attr->ia_vfsgid);
 
-		doutc(cl, "%p %llx.%llx gid %d -> %d\n", inode,
+		boutc(cl, "%p %llx.%llx gid %d -> %d\n", inode,
 		      ceph_vinop(inode),
 		      from_kgid(&init_user_ns, inode->i_gid),
 		      from_kgid(&init_user_ns, attr->ia_gid));
@@ -2789,7 +2811,7 @@ int __ceph_setattr(struct mnt_idmap *idmap, struct inode *inode,
 		}
 	}
 	if (ia_valid & ATTR_MODE) {
-		doutc(cl, "%p %llx.%llx mode 0%o -> 0%o\n", inode,
+		boutc(cl, "%p %llx.%llx mode 0%o -> 0%o\n", inode,
 		      ceph_vinop(inode), inode->i_mode, attr->ia_mode);
 		if (!do_sync && (issued & CEPH_CAP_AUTH_EXCL)) {
 			inode->i_mode = attr->ia_mode;
@@ -2806,7 +2828,7 @@ int __ceph_setattr(struct mnt_idmap *idmap, struct inode *inode,
 	if (ia_valid & ATTR_ATIME) {
 		struct timespec64 atime = inode_get_atime(inode);
 
-		doutc(cl, "%p %llx.%llx atime %ptSp -> %ptSp\n",
+		boutc(cl, "%p %llx.%llx atime %ptSp -> %ptSp\n",
 		      inode, ceph_vinop(inode), &atime, &attr->ia_atime);
 		if (!do_sync && (issued & CEPH_CAP_FILE_EXCL)) {
 			ci->i_time_warp_seq++;
@@ -2827,7 +2849,7 @@ int __ceph_setattr(struct mnt_idmap *idmap, struct inode *inode,
 		}
 	}
 	if (ia_valid & ATTR_SIZE) {
-		doutc(cl, "%p %llx.%llx size %lld -> %lld\n", inode,
+		boutc(cl, "%p %llx.%llx size %lld -> %lld\n", inode,
 		      ceph_vinop(inode), isize, attr->ia_size);
 		/*
 		 * Only when the new size is smaller and not aligned to
@@ -2881,7 +2903,7 @@ int __ceph_setattr(struct mnt_idmap *idmap, struct inode *inode,
 	if (ia_valid & ATTR_MTIME) {
 		struct timespec64 mtime = inode_get_mtime(inode);
 
-		doutc(cl, "%p %llx.%llx mtime %ptSp -> %ptSp\n",
+		boutc(cl, "%p %llx.%llx mtime %ptSp -> %ptSp\n",
 		      inode, ceph_vinop(inode), &mtime, &attr->ia_mtime);
 		if (!do_sync && (issued & CEPH_CAP_FILE_EXCL)) {
 			ci->i_time_warp_seq++;
@@ -2906,7 +2928,7 @@ int __ceph_setattr(struct mnt_idmap *idmap, struct inode *inode,
 		struct timespec64 ictime = inode_get_ctime(inode);
 		bool only = (ia_valid & (ATTR_SIZE|ATTR_MTIME|ATTR_ATIME|
 					 ATTR_MODE|ATTR_UID|ATTR_GID)) == 0;
-		doutc(cl, "%p %llx.%llx ctime %ptSp -> %ptSp (%s)\n",
+		boutc(cl, "%p %llx.%llx ctime %ptSp -> %ptSp (%s)\n",
 		      inode, ceph_vinop(inode), &ictime, &attr->ia_ctime,
 		      only ? "ctime only" : "ignored");
 		if (only) {
@@ -2926,7 +2948,7 @@ int __ceph_setattr(struct mnt_idmap *idmap, struct inode *inode,
 		}
 	}
 	if (ia_valid & ATTR_FILE)
-		doutc(cl, "%p %llx.%llx ATTR_FILE ... hrm!\n", inode,
+		boutc(cl, "%p %llx.%llx ATTR_FILE ... hrm!\n", inode,
 		      ceph_vinop(inode));
 
 	if (dirtied) {
@@ -2968,7 +2990,7 @@ int __ceph_setattr(struct mnt_idmap *idmap, struct inode *inode,
 		 */
 		err = ceph_mdsc_do_request(mdsc, NULL, req);
 		if (err == -EAGAIN && truncate_retry--) {
-			doutc(cl, "%p %llx.%llx result=%d (%s locally, %d remote), retry it!\n",
+			boutc(cl, "%p %llx.%llx result=%d (%s locally, %d remote), retry it!\n",
 			      inode, ceph_vinop(inode), err,
 			      ceph_cap_string(dirtied), mask);
 			ceph_mdsc_put_request(req);
@@ -2977,7 +2999,7 @@ int __ceph_setattr(struct mnt_idmap *idmap, struct inode *inode,
 		}
 	}
 out:
-	doutc(cl, "%p %llx.%llx result=%d (%s locally, %d remote)\n", inode,
+	boutc(cl, "%p %llx.%llx result=%d (%s locally, %d remote)\n", inode,
 	      ceph_vinop(inode), err, ceph_cap_string(dirtied), mask);
 
 	ceph_mdsc_put_request(req);
@@ -2998,34 +3020,50 @@ int ceph_setattr(struct mnt_idmap *idmap, struct dentry *dentry,
 	struct inode *inode = d_inode(dentry);
 	struct ceph_fs_client *fsc = ceph_inode_to_fs_client(inode);
 	int err;
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(fsc, &__ji);
 
-	if (ceph_in_snap(inode))
+	if (ceph_in_snap(inode)) {
+		ceph_blog_exit(&__ji);
 		return -EROFS;
+	}
 
-	if (ceph_inode_is_shutdown(inode))
+	if (ceph_inode_is_shutdown(inode)) {
+		ceph_blog_exit(&__ji);
 		return -ESTALE;
+	}
 
 	err = fscrypt_prepare_setattr(dentry, attr);
-	if (err)
+	if (err) {
+		ceph_blog_exit(&__ji);
 		return err;
+	}
 
 	err = setattr_prepare(idmap, dentry, attr);
-	if (err != 0)
+	if (err != 0) {
+		ceph_blog_exit(&__ji);
 		return err;
+	}
 
 	if ((attr->ia_valid & ATTR_SIZE) &&
-	    attr->ia_size > max(i_size_read(inode), fsc->max_file_size))
+	    attr->ia_size > max(i_size_read(inode), fsc->max_file_size)) {
+		ceph_blog_exit(&__ji);
 		return -EFBIG;
+	}
 
 	if ((attr->ia_valid & ATTR_SIZE) &&
-	    ceph_quota_is_max_bytes_exceeded(inode, attr->ia_size))
+	    ceph_quota_is_max_bytes_exceeded(inode, attr->ia_size)) {
+		ceph_blog_exit(&__ji);
 		return -EDQUOT;
+	}
 
 	err = __ceph_setattr(idmap, inode, attr, NULL);
 
 	if (err >= 0 && (attr->ia_valid & ATTR_MODE))
 		err = posix_acl_chmod(idmap, dentry, attr->ia_mode);
 
+	ceph_blog_exit(&__ji);
 	return err;
 }
 
@@ -3074,12 +3112,12 @@ int __ceph_do_getattr(struct inode *inode, struct folio *locked_folio,
 	int err;
 
 	if (ceph_snap(inode) == CEPH_SNAPDIR) {
-		doutc(cl, "inode %p %llx.%llx SNAPDIR\n", inode,
+		boutc(cl, "inode %p %llx.%llx SNAPDIR\n", inode,
 		      ceph_vinop(inode));
 		return 0;
 	}
 
-	doutc(cl, "inode %p %llx.%llx mask %s mode 0%o\n", inode,
+	boutc(cl, "inode %p %llx.%llx mask %s mode 0%o\n", inode,
 	      ceph_vinop(inode), ceph_cap_string(mask), inode->i_mode);
 	if (!force && ceph_caps_issued_mask_metric(ceph_inode(inode), mask, 1))
 			return 0;
@@ -3107,7 +3145,7 @@ int __ceph_do_getattr(struct inode *inode, struct folio *locked_folio,
 		}
 	}
 	ceph_mdsc_put_request(req);
-	doutc(cl, "result=%d\n", err);
+	boutc(cl, "result=%d\n", err);
 	return err;
 }
 
@@ -3145,7 +3183,7 @@ int ceph_do_getvxattr(struct inode *inode, const char *name, void *value,
 	xattr_value = req->r_reply_info.xattr_info.xattr_value;
 	xattr_value_len = req->r_reply_info.xattr_info.xattr_value_len;
 
-	doutc(cl, "xattr_value_len:%zu, size:%zu\n", xattr_value_len, size);
+	boutc(cl, "xattr_value_len:%zu, size:%zu\n", xattr_value_len, size);
 
 	err = (int)xattr_value_len;
 	if (size == 0)
@@ -3160,7 +3198,7 @@ int ceph_do_getvxattr(struct inode *inode, const char *name, void *value,
 put:
 	ceph_mdsc_put_request(req);
 out:
-	doutc(cl, "result=%d\n", err);
+	boutc(cl, "result=%d\n", err);
 	return err;
 }
 
@@ -3172,15 +3210,20 @@ int ceph_do_getvxattr(struct inode *inode, const char *name, void *value,
 int ceph_permission(struct mnt_idmap *idmap, struct inode *inode,
 		    int mask)
 {
+	struct ceph_fs_client *fsc = ceph_inode_to_fs_client(inode);
 	int err;
+	struct ceph_journal_info __ji;
 
 	if (mask & MAY_NOT_BLOCK)
 		return -ECHILD;
 
+	ceph_blog_enter(fsc, &__ji);
+
 	err = ceph_do_getattr(inode, CEPH_CAP_AUTH_SHARED, false);
 
 	if (!err)
 		err = generic_permission(idmap, inode, mask);
+	ceph_blog_exit(&__ji);
 	return err;
 }
 
@@ -3220,21 +3263,29 @@ int ceph_getattr(struct mnt_idmap *idmap, const struct path *path,
 		 struct kstat *stat, u32 request_mask, unsigned int flags)
 {
 	struct inode *inode = d_inode(path->dentry);
+	struct ceph_fs_client *fsc = ceph_inode_to_fs_client(inode);
 	struct super_block *sb = inode->i_sb;
 	struct ceph_inode_info *ci = ceph_inode(inode);
 	u32 valid_mask = STATX_BASIC_STATS;
 	int err = 0;
+	struct ceph_journal_info __ji;
 
-	if (ceph_inode_is_shutdown(inode))
+	ceph_blog_enter(fsc, &__ji);
+
+	if (ceph_inode_is_shutdown(inode)) {
+		ceph_blog_exit(&__ji);
 		return -ESTALE;
+	}
 
 	/* Skip the getattr altogether if we're asked not to sync */
 	if ((flags & AT_STATX_SYNC_TYPE) != AT_STATX_DONT_SYNC) {
 		err = ceph_do_getattr(inode,
 				statx_to_caps(request_mask, inode->i_mode),
 				flags & AT_STATX_FORCE_SYNC);
-		if (err)
+		if (err) {
+			ceph_blog_exit(&__ji);
 			return err;
+		}
 	}
 
 	generic_fillattr(idmap, request_mask, inode, stat);
@@ -3268,8 +3319,10 @@ int ceph_getattr(struct mnt_idmap *idmap, const struct path *path,
 			struct inode *parent;
 
 			parent = ceph_lookup_inode(sb, ceph_ino(inode));
-			if (IS_ERR(parent))
+			if (IS_ERR(parent)) {
+				ceph_blog_exit(&__ji);
 				return PTR_ERR(parent);
+			}
 
 			pci = ceph_inode(parent);
 			spin_lock(&pci->i_ceph_lock);
@@ -3302,6 +3355,7 @@ int ceph_getattr(struct mnt_idmap *idmap, const struct path *path,
 				  STATX_ATTR_ENCRYPTED);
 
 	stat->result_mask = request_mask & valid_mask;
+	ceph_blog_exit(&__ji);
 	return err;
 }
 
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 16+ messages in thread

* [PATCH v7 12/14] ceph: convert VFS data I/O paths to BLOG logging
  2026-09-24 15:30 [PATCH v7 00/14] ceph: add binary logging (BLOG) for CephFS Alex Markuze
                   ` (10 preceding siblings ...)
  2026-09-24 15:30 ` [PATCH v7 11/14] ceph: convert VFS inode and directory paths to BLOG logging Alex Markuze
@ 2026-09-24 15:30 ` Alex Markuze
  2026-09-24 15:30 ` [PATCH v7 13/14] ceph: convert capability and snapshot " Alex Markuze
                   ` (2 subsequent siblings)
  14 siblings, 0 replies; 16+ messages in thread
From: Alex Markuze @ 2026-09-24 15:30 UTC (permalink / raw)
  To: ceph-devel; +Cc: idryomov, xiubo.li

Replace text debug logging with the boutc() wrappers in addr.c
(readpage, writepage, direct-IO) and file.c (open, read_iter,
write_iter, fsync, fallocate).

Format cross-cluster UUIDs only inside the selected binary
helper. The text fallback keeps the original %pU arguments.

Signed-off-by: Alex Markuze <amarkuze@redhat.com>
Assisted-by: LLM
---
 fs/ceph/addr.c | 202 +++++++++++++++++++++-------------
 fs/ceph/file.c | 287 ++++++++++++++++++++++++++++++++++---------------
 2 files changed, 327 insertions(+), 162 deletions(-)

diff --git a/fs/ceph/addr.c b/fs/ceph/addr.c
index a185849a0490..b4ea5ed660f9 100644
--- a/fs/ceph/addr.c
+++ b/fs/ceph/addr.c
@@ -81,15 +81,21 @@ struct ceph_snap_context *ceph_folio_snap_context(const struct folio *folio)
 static bool ceph_dirty_folio(struct address_space *mapping, struct folio *folio)
 {
 	struct inode *inode = mapping->host;
+	struct ceph_fs_client *fsc = ceph_inode_to_fs_client(inode);
 	struct ceph_client *cl = ceph_inode_to_client(inode);
 	struct ceph_mds_client *mdsc = ceph_sb_to_mdsc(inode->i_sb);
 	struct ceph_inode_info *ci;
 	struct ceph_snap_context *snapc;
+	struct ceph_journal_info __ji;
+
+	/* May run under page-table locks; only reuse an existing BLOG ctx. */
+	ceph_blog_enter_gfp(fsc, &__ji, GFP_ATOMIC);
 
 	if (folio_test_dirty(folio)) {
-		doutc(cl, "%llx.%llx %p idx %lu -- already dirty\n",
+		boutc(cl, "%llx.%llx %p idx %lu -- already dirty\n",
 		      ceph_vinop(inode), folio, folio->index);
 		VM_BUG_ON_FOLIO(!folio_test_private(folio), folio);
+		ceph_blog_exit(&__ji);
 		return false;
 	}
 
@@ -114,7 +120,7 @@ static bool ceph_dirty_folio(struct address_space *mapping, struct folio *folio)
 	if (ci->i_wrbuffer_ref == 0)
 		ihold(inode);
 	++ci->i_wrbuffer_ref;
-	doutc(cl, "%llx.%llx %p idx %lu head %d/%d -> %d/%d "
+	boutc(cl, "%llx.%llx %p idx %lu head %d/%d -> %d/%d "
 	      "snapc %p seq %lld (%d snaps)\n",
 	      ceph_vinop(inode), folio, folio->index,
 	      ci->i_wrbuffer_ref-1, ci->i_wrbuffer_ref_head-1,
@@ -129,6 +135,7 @@ static bool ceph_dirty_folio(struct address_space *mapping, struct folio *folio)
 	VM_WARN_ON_FOLIO(folio->private, folio);
 	folio_attach_private(folio, snapc);
 
+	ceph_blog_exit(&__ji);
 	return ceph_fscache_dirty_folio(mapping, folio);
 }
 
@@ -141,20 +148,25 @@ static void ceph_invalidate_folio(struct folio *folio, size_t offset,
 				size_t length)
 {
 	struct inode *inode = folio->mapping->host;
+	struct ceph_fs_client *fsc = ceph_inode_to_fs_client(inode);
 	struct ceph_client *cl = ceph_inode_to_client(inode);
 	struct ceph_inode_info *ci = ceph_inode(inode);
 	struct ceph_snap_context *snapc;
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter_gfp(fsc, &__ji, GFP_ATOMIC);
 
 
 	if (offset != 0 || length != folio_size(folio)) {
-		doutc(cl, "%llx.%llx idx %lu partial dirty page %zu~%zu\n",
+		boutc(cl, "%llx.%llx idx %lu partial dirty page %zu~%zu\n",
 		      ceph_vinop(inode), folio->index, offset, length);
+		ceph_blog_exit(&__ji);
 		return;
 	}
 
 	WARN_ON(!folio_test_locked(folio));
 	if (folio_test_private(folio)) {
-		doutc(cl, "%llx.%llx idx %lu full dirty page\n",
+		boutc(cl, "%llx.%llx idx %lu full dirty page\n",
 		      ceph_vinop(inode), folio->index);
 
 		snapc = folio_detach_private(folio);
@@ -163,6 +175,7 @@ static void ceph_invalidate_folio(struct folio *folio, size_t offset,
 	}
 
 	netfs_invalidate_folio(folio, offset, length);
+	ceph_blog_exit(&__ji);
 }
 
 static void ceph_netfs_expand_readahead(struct netfs_io_request *rreq)
@@ -218,11 +231,14 @@ static void finish_netfs_read(struct ceph_osd_request *req)
 	struct ceph_osd_req_op *op = &req->r_ops[0];
 	int err = req->r_result;
 	bool sparse = (op->op == CEPH_OSD_OP_SPARSE_READ);
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter_gfp(fsc, &__ji, GFP_ATOMIC);
 
 	ceph_update_read_metrics(&fsc->mdsc->metric, req->r_start_latency,
 				 req->r_end_latency, osd_data->length, err);
 
-	doutc(cl, "result %d subreq->len=%zu i_size=%lld\n", req->r_result,
+	boutc(cl, "result %d subreq->len=%zu i_size=%lld\n", req->r_result,
 	      subreq->len, i_size_read(req->r_inode));
 
 	/* no object means success but no data */
@@ -274,6 +290,7 @@ static void finish_netfs_read(struct ceph_osd_request *req)
 	netfs_read_subreq_terminated(subreq);
 	iput(req->r_inode);
 	ceph_dec_osd_stopping_blocker(fsc->mdsc);
+	ceph_blog_exit(&__ji);
 }
 
 #ifdef CONFIG_CEPH_FS_INLINE_DATA
@@ -414,7 +431,7 @@ static void ceph_netfs_issue_read(struct netfs_io_subrequest *subreq)
 			goto out;
 	}
 
-	doutc(cl, "%llx.%llx pos=%llu orig_len=%zu len=%llu\n",
+	boutc(cl, "%llx.%llx pos=%llu orig_len=%zu len=%llu\n",
 	      ceph_vinop(inode), subreq->start, subreq->len, len);
 
 	/*
@@ -439,7 +456,7 @@ static void ceph_netfs_issue_read(struct netfs_io_subrequest *subreq)
 
 		err = iov_iter_get_pages_alloc2(&subreq->io_iter, &pages, len, &page_off);
 		if (err < 0) {
-			doutc(cl, "%llx.%llx failed to allocate pages, %d\n",
+			boutc(cl, "%llx.%llx failed to allocate pages, %d\n",
 			      ceph_vinop(inode), err);
 			goto out;
 		}
@@ -482,7 +499,7 @@ static void ceph_netfs_issue_read(struct netfs_io_subrequest *subreq)
 		subreq->error = err;
 		netfs_read_subreq_terminated(subreq);
 	}
-	doutc(cl, "%llx.%llx result %d\n", ceph_vinop(inode), err);
+	boutc(cl, "%llx.%llx result %d\n", ceph_vinop(inode), err);
 }
 
 static int ceph_init_request(struct netfs_io_request *rreq, struct file *file)
@@ -532,13 +549,13 @@ static int ceph_init_request(struct netfs_io_request *rreq, struct file *file)
 		ret = ceph_try_get_caps(inode, CEPH_CAP_FILE_RD, want, true,
 					&got);
 		if (ret < 0) {
-			doutc(cl, "%llx.%llx, error getting cap\n",
+			boutc(cl, "%llx.%llx, error getting cap\n",
 			      ceph_vinop(inode));
 			goto out;
 		}
 
 		if (!(got & want)) {
-			doutc(cl, "%llx.%llx, no cache cap\n", ceph_vinop(inode));
+			boutc(cl, "%llx.%llx, no cache cap\n", ceph_vinop(inode));
 			ret = -EACCES;
 			goto out;
 		}
@@ -557,7 +574,7 @@ static int ceph_init_request(struct netfs_io_request *rreq, struct file *file)
 		ret = __ceph_get_caps(inode, fi, CEPH_CAP_FILE_RD, want, -1,
 				      &got);
 		if (ret < 0) {
-			doutc(cl, "%llx.%llx, error getting cap\n",
+			boutc(cl, "%llx.%llx, error getting cap\n",
 			      ceph_vinop(inode));
 			goto out;
 		}
@@ -693,7 +710,7 @@ get_oldest_context(struct inode *inode, struct ceph_writeback_ctl *ctl,
 
 	spin_lock(&ci->i_ceph_lock);
 	list_for_each_entry(capsnap, &ci->i_cap_snaps, ci_item) {
-		doutc(cl, " capsnap %p snapc %p has %d dirty pages\n",
+		boutc(cl, " capsnap %p snapc %p has %d dirty pages\n",
 		      capsnap, capsnap->context, capsnap->dirty_pages);
 		if (!capsnap->dirty_pages)
 			continue;
@@ -726,7 +743,7 @@ get_oldest_context(struct inode *inode, struct ceph_writeback_ctl *ctl,
 	}
 	if (!snapc && ci->i_wrbuffer_ref_head) {
 		snapc = ceph_get_snap_context(ci->i_head_snapc);
-		doutc(cl, " head snapc %p has %d dirty pages\n", snapc,
+		boutc(cl, " head snapc %p has %d dirty pages\n", snapc,
 		      ci->i_wrbuffer_ref_head);
 		if (ctl) {
 			ctl->i_size = i_size_read(inode);
@@ -797,7 +814,7 @@ static int write_folio_nounlock(struct folio *folio,
 	bool caching = ceph_is_cache_enabled(inode);
 	struct page *bounce_page = NULL;
 
-	doutc(cl, "%llx.%llx folio %p idx %lu\n", ceph_vinop(inode), folio,
+	boutc(cl, "%llx.%llx folio %p idx %lu\n", ceph_vinop(inode), folio,
 	      folio->index);
 
 	if (ceph_inode_is_shutdown(inode))
@@ -806,13 +823,13 @@ static int write_folio_nounlock(struct folio *folio,
 	/* verify this is a writeable snap context */
 	snapc = ceph_folio_snap_context(folio);
 	if (!snapc) {
-		doutc(cl, "%llx.%llx folio %p not dirty?\n", ceph_vinop(inode),
+		boutc(cl, "%llx.%llx folio %p not dirty?\n", ceph_vinop(inode),
 		      folio);
 		return 0;
 	}
 	oldest = get_oldest_context(inode, &ceph_wbc, snapc);
 	if (snapc->seq > oldest->seq) {
-		doutc(cl, "%llx.%llx folio %p snapc %p not writeable - noop\n",
+		boutc(cl, "%llx.%llx folio %p snapc %p not writeable - noop\n",
 		      ceph_vinop(inode), folio, snapc);
 		/* we should only noop if called by kswapd */
 		WARN_ON(!(current->flags & PF_MEMALLOC));
@@ -824,7 +841,7 @@ static int write_folio_nounlock(struct folio *folio,
 
 	/* is this a partial page at end of file? */
 	if (page_off >= ceph_wbc.i_size) {
-		doutc(cl, "%llx.%llx folio at %lu beyond eof %llu\n",
+		boutc(cl, "%llx.%llx folio at %lu beyond eof %llu\n",
 		      ceph_vinop(inode), folio->index, ceph_wbc.i_size);
 		folio_invalidate(folio, 0, folio_size(folio));
 		return 0;
@@ -834,7 +851,7 @@ static int write_folio_nounlock(struct folio *folio,
 		len = ceph_wbc.i_size - page_off;
 
 	wlen = IS_ENCRYPTED(inode) ? round_up(len, CEPH_FSCRYPT_BLOCK_SIZE) : len;
-	doutc(cl, "%llx.%llx folio %p index %lu on %llu~%llu snapc %p seq %lld\n",
+	boutc(cl, "%llx.%llx folio %p index %lu on %llu~%llu snapc %p seq %lld\n",
 	      ceph_vinop(inode), folio, folio->index, page_off, wlen, snapc,
 	      snapc->seq);
 
@@ -885,7 +902,7 @@ static int write_folio_nounlock(struct folio *folio,
 	WARN_ON_ONCE(len > folio_size(folio));
 	page = bounce_page ? bounce_page : &folio->page;
 	osd_req_op_extent_osd_data_pages(req, 0, &page, wlen, 0, false, false);
-	doutc(cl, "%llx.%llx %llu~%llu (%llu bytes, %sencrypted)\n",
+	boutc(cl, "%llx.%llx %llu~%llu (%llu bytes, %sencrypted)\n",
 	      ceph_vinop(inode), page_off, len, wlen,
 	      IS_ENCRYPTED(inode) ? "" : "not ");
 
@@ -910,7 +927,7 @@ static int write_folio_nounlock(struct folio *folio,
 			wbc = &tmp_wbc;
 		if (err == -ERESTARTSYS) {
 			/* killed by SIGKILL */
-			doutc(cl, "%llx.%llx interrupted page %p\n",
+			boutc(cl, "%llx.%llx interrupted page %p\n",
 			      ceph_vinop(inode), folio);
 			folio_redirty_for_writepage(wbc, folio);
 			folio_end_writeback(folio);
@@ -921,12 +938,12 @@ static int write_folio_nounlock(struct folio *folio,
 		}
 		if (err == -EBLOCKLISTED)
 			fsc->blocklisted = true;
-		doutc(cl, "%llx.%llx setting mapping error %d %p\n",
+		boutc(cl, "%llx.%llx setting mapping error %d %p\n",
 		      ceph_vinop(inode), err, folio);
 		mapping_set_error(&inode->i_data, err);
 		wbc->pages_skipped++;
 	} else {
-		doutc(cl, "%llx.%llx cleaned page %p\n",
+		boutc(cl, "%llx.%llx cleaned page %p\n",
 		      ceph_vinop(inode), folio);
 		err = 0;  /* vfs expects us to return 0 */
 	}
@@ -964,8 +981,12 @@ static void writepages_finish(struct ceph_osd_request *req)
 	struct ceph_mds_client *mdsc = ceph_sb_to_mdsc(inode->i_sb);
 	unsigned int len = 0;
 	bool remove_page;
+	struct ceph_journal_info __ji;
 
-	doutc(cl, "%llx.%llx rc %d\n", ceph_vinop(inode), rc);
+	/* OSD callback: reuse existing ctx only. */
+	ceph_blog_enter_gfp(fsc, &__ji, GFP_ATOMIC);
+
+	boutc(cl, "%llx.%llx rc %d\n", ceph_vinop(inode), rc);
 	if (rc < 0) {
 		mapping_set_error(mapping, rc);
 		ceph_set_error_write(ci);
@@ -1023,7 +1044,7 @@ static void writepages_finish(struct ceph_osd_request *req)
 				WARN_ON(atomic64_read(&mdsc->dirty_folios) < 0);
 			}
 
-			doutc(cl, "unlocking %p\n", folio);
+			boutc(cl, "unlocking %p\n", folio);
 
 			if (remove_page)
 				generic_error_remove_folio(inode->i_mapping,
@@ -1031,7 +1052,7 @@ static void writepages_finish(struct ceph_osd_request *req)
 
 			folio_unlock(folio);
 		}
-		doutc(cl, "%llx.%llx wrote %llu bytes cleaned %d pages\n",
+		boutc(cl, "%llx.%llx wrote %llu bytes cleaned %d pages\n",
 		      ceph_vinop(inode), osd_data->length,
 		      rc >= 0 ? num_pages : 0);
 
@@ -1055,6 +1076,7 @@ static void writepages_finish(struct ceph_osd_request *req)
 		kfree(osd_data->pages);
 	ceph_osdc_put_request(req);
 	ceph_dec_osd_stopping_blocker(fsc->mdsc);
+	ceph_blog_exit(&__ji);
 }
 
 static inline
@@ -1157,11 +1179,11 @@ int ceph_define_writeback_range(struct address_space *mapping,
 	if (!ceph_wbc->snapc) {
 		/* hmm, why does writepages get called when there
 		   is no dirty data? */
-		doutc(cl, " no snap context with dirty data?\n");
+		boutc(cl, " no snap context with dirty data?\n");
 		return -ENODATA;
 	}
 
-	doutc(cl, " oldest snapc is %p seq %lld (%d snaps)\n",
+	boutc(cl, " oldest snapc is %p seq %lld (%d snaps)\n",
 	      ceph_wbc->snapc, ceph_wbc->snapc->seq,
 	      ceph_wbc->snapc->num_snaps);
 
@@ -1174,13 +1196,13 @@ int ceph_define_writeback_range(struct address_space *mapping,
 			ceph_wbc->end = -1;
 			if (ceph_wbc->index > 0)
 				ceph_wbc->should_loop = true;
-			doutc(cl, " cyclic, start at %lu\n", ceph_wbc->index);
+			boutc(cl, " cyclic, start at %lu\n", ceph_wbc->index);
 		} else {
 			ceph_wbc->index = wbc->range_start >> PAGE_SHIFT;
 			ceph_wbc->end = wbc->range_end >> PAGE_SHIFT;
 			if (wbc->range_start == 0 && wbc->range_end == LLONG_MAX)
 				ceph_wbc->range_whole = true;
-			doutc(cl, " not cyclic, %lu to %lu\n",
+			boutc(cl, " not cyclic, %lu to %lu\n",
 				ceph_wbc->index, ceph_wbc->end);
 		}
 	} else if (!ceph_wbc->head_snapc) {
@@ -1190,7 +1212,7 @@ int ceph_define_writeback_range(struct address_space *mapping,
 		 * associated with 'snapc' get written */
 		if (ceph_wbc->index > 0)
 			ceph_wbc->should_loop = true;
-		doutc(cl, " non-head snapc, range whole\n");
+		boutc(cl, " non-head snapc, range whole\n");
 	}
 
 	ceph_put_snap_context(ceph_wbc->last_snapc);
@@ -1226,14 +1248,14 @@ int ceph_check_folio_before_write(struct address_space *mapping,
 
 	/* only dirty folios, or our accounting breaks */
 	if (unlikely(!folio_test_dirty(folio) || folio->mapping != mapping)) {
-		doutc(cl, "!dirty or !mapping %p\n", folio);
+		boutc(cl, "!dirty or !mapping %p\n", folio);
 		return -ENODATA;
 	}
 
 	/* only if matching snap context */
 	pgsnapc = ceph_folio_snap_context(folio);
 	if (pgsnapc != ceph_wbc->snapc) {
-		doutc(cl, "folio snapc %p %lld != oldest %p %lld\n",
+		boutc(cl, "folio snapc %p %lld != oldest %p %lld\n",
 		      pgsnapc, pgsnapc->seq,
 		      ceph_wbc->snapc, ceph_wbc->snapc->seq);
 
@@ -1245,7 +1267,7 @@ int ceph_check_folio_before_write(struct address_space *mapping,
 	}
 
 	if (folio_pos(folio) >= ceph_wbc->i_size) {
-		doutc(cl, "folio at %lu beyond eof %llu\n",
+		boutc(cl, "folio at %lu beyond eof %llu\n",
 		      folio->index, ceph_wbc->i_size);
 
 		if ((ceph_wbc->size_stable ||
@@ -1258,7 +1280,7 @@ int ceph_check_folio_before_write(struct address_space *mapping,
 
 	if (ceph_wbc->strip_unit_end &&
 	    (folio->index > ceph_wbc->strip_unit_end)) {
-		doutc(cl, "end of strip unit %p\n", folio);
+		boutc(cl, "end of strip unit %p\n", folio);
 		return -E2BIG;
 	}
 
@@ -1382,7 +1404,7 @@ void ceph_process_folio_batch(struct address_space *mapping,
 		if (!folio)
 			continue;
 
-		doutc(cl, "? %p idx %lu, folio_test_writeback %#x, "
+		boutc(cl, "? %p idx %lu, folio_test_writeback %#x, "
 			"folio_test_dirty %#x, folio_test_locked %#x\n",
 			folio, folio->index, folio_test_writeback(folio),
 			folio_test_dirty(folio),
@@ -1390,7 +1412,7 @@ void ceph_process_folio_batch(struct address_space *mapping,
 
 		if (folio_test_writeback(folio) ||
 		    folio_test_private_2(folio) /* [DEPRECATED] */) {
-			doutc(cl, "waiting on writeback %p\n", folio);
+			boutc(cl, "waiting on writeback %p\n", folio);
 			folio_wait_writeback(folio);
 			folio_wait_private_2(folio); /* [DEPRECATED] */
 			continue;
@@ -1414,7 +1436,7 @@ void ceph_process_folio_batch(struct address_space *mapping,
 		}
 
 		if (!folio_clear_dirty_for_io(folio)) {
-			doutc(cl, "%p !folio_clear_dirty_for_io\n", folio);
+			boutc(cl, "%p !folio_clear_dirty_for_io\n", folio);
 			folio_unlock(folio);
 			folio_put(folio);
 			ceph_wbc->fbatch.folios[i] = NULL;
@@ -1442,7 +1464,7 @@ void ceph_process_folio_batch(struct address_space *mapping,
 		}
 
 		/* note position of first page in fbatch */
-		doutc(cl, "%llx.%llx will write folio %p idx %lu\n",
+		boutc(cl, "%llx.%llx will write folio %p idx %lu\n",
 		      ceph_vinop(inode), folio, folio->index);
 
 		rc = move_dirty_folio_in_page_array(mapping, wbc, ceph_wbc,
@@ -1602,7 +1624,7 @@ int ceph_submit_write(struct address_space *mapping,
 			osd_req_op_extent_dup_last(req, ceph_wbc->op_idx,
 						   cur_offset - offset);
 
-			doutc(cl, "got pages at %llu~%llu\n", offset, len);
+			boutc(cl, "got pages at %llu~%llu\n", offset, len);
 
 			osd_req_op_extent_osd_data_pages(req, ceph_wbc->op_idx,
 							 ceph_wbc->data_pages,
@@ -1643,7 +1665,7 @@ int ceph_submit_write(struct address_space *mapping,
 	if (IS_ENCRYPTED(inode))
 		len = round_up(len, CEPH_FSCRYPT_BLOCK_SIZE);
 
-	doutc(cl, "got pages at %llu~%llu\n", offset, len);
+	boutc(cl, "got pages at %llu~%llu\n", offset, len);
 
 	if (IS_ENCRYPTED(inode) &&
 	    ((offset | len) & ~CEPH_FSCRYPT_BLOCK_MASK)) {
@@ -1732,16 +1754,21 @@ static int ceph_writepages_start(struct address_space *mapping,
 	struct ceph_client *cl = fsc->client;
 	struct ceph_writeback_ctl ceph_wbc;
 	int rc = 0;
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(fsc, &__ji);
 
-	if (wbc->sync_mode == WB_SYNC_NONE && fsc->write_congested)
+	if (wbc->sync_mode == WB_SYNC_NONE && fsc->write_congested) {
+		ceph_blog_exit(&__ji);
 		return 0;
+	}
 
-	doutc(cl, "%llx.%llx (mode=%s)\n", ceph_vinop(inode),
+	boutc(cl, "%llx.%llx (mode=%s)\n", ceph_vinop(inode),
 	      wbc->sync_mode == WB_SYNC_NONE ? "NONE" :
 	      (wbc->sync_mode == WB_SYNC_ALL ? "ALL" : "HOLD"));
 
 	if (is_forced_umount(mapping)) {
-		/* we're in a forced umount, don't write! */
+		ceph_blog_exit(&__ji);
 		return -EIO;
 	}
 
@@ -1778,7 +1805,7 @@ static int ceph_writepages_start(struct address_space *mapping,
 							    ceph_wbc.end,
 							    ceph_wbc.tag,
 							    &ceph_wbc.fbatch);
-		doutc(cl, "pagevec_lookup_range_tag for tag %#x got %d\n",
+		boutc(cl, "pagevec_lookup_range_tag for tag %#x got %d\n",
 			ceph_wbc.tag, ceph_wbc.nr_folios);
 
 		if (!ceph_wbc.nr_folios && !ceph_wbc.locked_pages)
@@ -1795,7 +1822,7 @@ static int ceph_writepages_start(struct address_space *mapping,
 		if (ceph_wbc.processed_in_fbatch) {
 			if (folio_batch_count(&ceph_wbc.fbatch) == 0 &&
 			    ceph_wbc.locked_pages < ceph_wbc.max_pages) {
-				doutc(cl, "reached end fbatch, trying for more\n");
+				boutc(cl, "reached end fbatch, trying for more\n");
 				goto get_more_pages;
 			}
 		}
@@ -1826,7 +1853,7 @@ static int ceph_writepages_start(struct address_space *mapping,
 			ceph_wbc.done = true;
 
 release_folios:
-		doutc(cl, "folio_batch release on %d folios (%p)\n",
+		boutc(cl, "folio_batch release on %d folios (%p)\n",
 		      (int)ceph_wbc.fbatch.nr,
 		      ceph_wbc.fbatch.nr ? ceph_wbc.fbatch.folios[0] : NULL);
 		folio_batch_release(&ceph_wbc.fbatch);
@@ -1834,7 +1861,7 @@ static int ceph_writepages_start(struct address_space *mapping,
 
 	if (ceph_wbc.should_loop && !ceph_wbc.done) {
 		/* more to do; loop back to beginning of file */
-		doutc(cl, "looping back to beginning of file\n");
+		boutc(cl, "looping back to beginning of file\n");
 		/* OK even when start_index == 0 */
 		ceph_wbc.end = ceph_wbc.start_index - 1;
 
@@ -1855,9 +1882,10 @@ static int ceph_writepages_start(struct address_space *mapping,
 
 out:
 	ceph_put_snap_context(ceph_wbc.last_snapc);
-	doutc(cl, "%llx.%llx dend - startone, rc = %d\n", ceph_vinop(inode),
+	boutc(cl, "%llx.%llx dend - startone, rc = %d\n", ceph_vinop(inode),
 	      rc);
 
+	ceph_blog_exit(&__ji);
 	return rc;
 }
 
@@ -1893,7 +1921,7 @@ ceph_find_incompatible(struct folio *folio)
 	struct ceph_inode_info *ci = ceph_inode(inode);
 
 	if (ceph_inode_is_shutdown(inode)) {
-		doutc(cl, " %llx.%llx folio %p is shutdown\n",
+		boutc(cl, " %llx.%llx folio %p is shutdown\n",
 		      ceph_vinop(inode), folio);
 		return ERR_PTR(-ESTALE);
 	}
@@ -1915,14 +1943,14 @@ ceph_find_incompatible(struct folio *folio)
 		if (snapc->seq > oldest->seq) {
 			/* not writeable -- return it for the caller to deal with */
 			ceph_put_snap_context(oldest);
-			doutc(cl, " %llx.%llx folio %p snapc %p not current or oldest\n",
+			boutc(cl, " %llx.%llx folio %p snapc %p not current or oldest\n",
 			      ceph_vinop(inode), folio, snapc);
 			return ceph_get_snap_context(snapc);
 		}
 		ceph_put_snap_context(oldest);
 
 		/* yay, writeable, do it now (without dropping folio lock) */
-		doutc(cl, " %llx.%llx folio %p snapc %p not current, but oldest\n",
+		boutc(cl, " %llx.%llx folio %p snapc %p not current, but oldest\n",
 		      ceph_vinop(inode), folio, snapc);
 		if (folio_clear_dirty_for_io(folio)) {
 			int r = write_folio_nounlock(folio, NULL);
@@ -1970,15 +1998,22 @@ static int ceph_write_begin(const struct kiocb *iocb,
 {
 	struct file *file = iocb->ki_filp;
 	struct inode *inode = file_inode(file);
+	struct ceph_fs_client *fsc = ceph_inode_to_fs_client(inode);
 	struct ceph_inode_info *ci = ceph_inode(inode);
 	int r;
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(fsc, &__ji);
 
 	r = netfs_write_begin(&ci->netfs, file, inode->i_mapping, pos, len, foliop, NULL);
-	if (r < 0)
+	if (r < 0) {
+		ceph_blog_exit(&__ji);
 		return r;
+	}
 
 	folio_wait_private_2(*foliop); /* [DEPRECATED] */
 	WARN_ON_ONCE(!folio_test_locked(*foliop));
+	ceph_blog_exit(&__ji);
 	return 0;
 }
 
@@ -1993,10 +2028,14 @@ static int ceph_write_end(const struct kiocb *iocb,
 {
 	struct file *file = iocb->ki_filp;
 	struct inode *inode = file_inode(file);
+	struct ceph_fs_client *fsc = ceph_inode_to_fs_client(inode);
 	struct ceph_client *cl = ceph_inode_to_client(inode);
 	bool check_cap = false;
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter_gfp(fsc, &__ji, GFP_ATOMIC);
 
-	doutc(cl, "%llx.%llx file %p folio %p %d~%d (%d)\n", ceph_vinop(inode),
+	boutc(cl, "%llx.%llx file %p folio %p %d~%d (%d)\n", ceph_vinop(inode),
 	      file, folio, (int)pos, (int)copied, (int)len);
 
 	if (!folio_test_uptodate(folio)) {
@@ -2021,6 +2060,7 @@ static int ceph_write_end(const struct kiocb *iocb,
 	if (check_cap)
 		ceph_check_caps(ceph_inode(inode), CHECK_CAPS_AUTHONLY);
 
+	ceph_blog_exit(&__ji);
 	return copied;
 }
 
@@ -2093,19 +2133,22 @@ static vm_fault_t ceph_filemap_fault(struct vm_fault *vmf)
 	struct vm_area_struct *vma = vmf->vma;
 	struct inode *inode = file_inode(vma->vm_file);
 	struct ceph_inode_info *ci = ceph_inode(inode);
-	struct ceph_client *cl = ceph_inode_to_client(inode);
+	struct ceph_fs_client *fsc = ceph_inode_to_fs_client(inode);
+	struct ceph_client *cl = fsc->client;
 	struct ceph_file_info *fi = vma->vm_file->private_data;
 	loff_t off = (loff_t)vmf->pgoff << PAGE_SHIFT;
 	int want, got, err;
 	sigset_t oldset;
 	vm_fault_t ret = VM_FAULT_SIGBUS;
+	struct ceph_journal_info __ji;
 
 	if (ceph_inode_is_shutdown(inode))
 		return ret;
 
 	ceph_block_sigs(&oldset);
+	ceph_blog_enter(fsc, &__ji);
 
-	doutc(cl, "%llx.%llx %llu trying to get caps\n",
+	boutc(cl, "%llx.%llx %llu trying to get caps\n",
 	      ceph_vinop(inode), off);
 	if (fi->fmode & CEPH_FILE_MODE_LAZY)
 		want = CEPH_CAP_FILE_CACHE | CEPH_CAP_FILE_LAZYIO;
@@ -2117,7 +2160,7 @@ static vm_fault_t ceph_filemap_fault(struct vm_fault *vmf)
 	if (err < 0)
 		goto out_restore;
 
-	doutc(cl, "%llx.%llx %llu got cap refs on %s\n", ceph_vinop(inode),
+	boutc(cl, "%llx.%llx %llu got cap refs on %s\n", ceph_vinop(inode),
 	      off, ceph_cap_string(got));
 
 	/*
@@ -2163,7 +2206,7 @@ static vm_fault_t ceph_filemap_fault(struct vm_fault *vmf)
 		ceph_add_rw_context(fi, &rw_ctx);
 		ret = filemap_fault(vmf);
 		ceph_del_rw_context(fi, &rw_ctx);
-		doutc(cl, "%llx.%llx %llu drop cap refs %s ret %x\n",
+		boutc(cl, "%llx.%llx %llu drop cap refs %s ret %x\n",
 		      ceph_vinop(inode), off, ceph_cap_string(got), ret);
 	} else
 		err = -EAGAIN;
@@ -2206,10 +2249,11 @@ static vm_fault_t ceph_filemap_fault(struct vm_fault *vmf)
 		ret = VM_FAULT_MAJOR | VM_FAULT_LOCKED;
 out_inline:
 		filemap_invalidate_unlock_shared(mapping);
-		doutc(cl, "%llx.%llx %llu read inline data ret %x\n",
+		boutc(cl, "%llx.%llx %llu read inline data ret %x\n",
 		      ceph_vinop(inode), off, ret);
 	}
 out_restore:
+	ceph_blog_exit(&__ji);
 	ceph_restore_sigs(&oldset);
 	if (err < 0)
 		ret = vmf_error(err);
@@ -2245,7 +2289,8 @@ static vm_fault_t ceph_page_mkwrite(struct vm_fault *vmf)
 {
 	struct vm_area_struct *vma = vmf->vma;
 	struct inode *inode = file_inode(vma->vm_file);
-	struct ceph_client *cl = ceph_inode_to_client(inode);
+	struct ceph_fs_client *fsc = ceph_inode_to_fs_client(inode);
+	struct ceph_client *cl = fsc->client;
 	struct ceph_inode_info *ci = ceph_inode(inode);
 	struct ceph_file_info *fi = vma->vm_file->private_data;
 	struct ceph_cap_flush *prealloc_cf;
@@ -2256,6 +2301,7 @@ static vm_fault_t ceph_page_mkwrite(struct vm_fault *vmf)
 	int want, got, err;
 	sigset_t oldset;
 	vm_fault_t ret = VM_FAULT_SIGBUS;
+	struct ceph_journal_info __ji;
 
 	if (ceph_inode_is_shutdown(inode))
 		return ret;
@@ -2266,13 +2312,14 @@ static vm_fault_t ceph_page_mkwrite(struct vm_fault *vmf)
 
 	sb_start_pagefault(inode->i_sb);
 	ceph_block_sigs(&oldset);
+	ceph_blog_enter(fsc, &__ji);
 
 	if (off + folio_size(folio) <= size)
 		len = folio_size(folio);
 	else
 		len = offset_in_folio(folio, size);
 
-	doutc(cl, "%llx.%llx %llu~%zd getting caps i_size %llu\n",
+	boutc(cl, "%llx.%llx %llu~%zd getting caps i_size %llu\n",
 	      ceph_vinop(inode), off, len, size);
 	if (fi->fmode & CEPH_FILE_MODE_LAZY)
 		want = CEPH_CAP_FILE_BUFFER | CEPH_CAP_FILE_LAZYIO;
@@ -2285,7 +2332,7 @@ static vm_fault_t ceph_page_mkwrite(struct vm_fault *vmf)
 	if (err < 0)
 		goto out_free;
 
-	doutc(cl, "%llx.%llx %llu~%zd got cap refs on %s\n", ceph_vinop(inode),
+	boutc(cl, "%llx.%llx %llu~%zd got cap refs on %s\n", ceph_vinop(inode),
 	      off, len, ceph_cap_string(got));
 
 	/*
@@ -2370,10 +2417,11 @@ static vm_fault_t ceph_page_mkwrite(struct vm_fault *vmf)
 			__mark_inode_dirty(inode, dirty);
 	}
 
-	doutc(cl, "%llx.%llx %llu~%zd dropping cap refs on %s ret %x\n",
+	boutc(cl, "%llx.%llx %llu~%zd dropping cap refs on %s ret %x\n",
 	      ceph_vinop(inode), off, len, ceph_cap_string(got), ret);
 	ceph_put_cap_refs_async(ci, got);
 out_free:
+	ceph_blog_exit(&__ji);
 	ceph_restore_sigs(&oldset);
 	sb_end_pagefault(inode->i_sb);
 	ceph_free_cap_flush(prealloc_cf);
@@ -2407,7 +2455,7 @@ void ceph_fill_inline_data(struct inode *inode, struct folio *locked_folio,
 		}
 	}
 
-	doutc(cl, "%p %llx.%llx len %zu locked_folio %p\n", inode,
+	boutc(cl, "%p %llx.%llx len %zu locked_folio %p\n", inode,
 	      ceph_vinop(inode), len, locked_folio);
 
 	if (len > folio_size(folio)) {
@@ -2450,7 +2498,7 @@ int ceph_uninline_data(struct file *file)
 	inline_version = ceph_inline_version(ci);
 	spin_unlock(&ci->i_ceph_lock);
 
-	doutc(cl, "%llx.%llx inline_version %llu\n", ceph_vinop(inode),
+	boutc(cl, "%llx.%llx inline_version %llu\n", ceph_vinop(inode),
 	      inline_version);
 
 	if (ceph_inode_is_shutdown(inode)) {
@@ -2582,7 +2630,7 @@ int ceph_uninline_data(struct file *file)
 out:
 	ceph_put_snap_context(snapc);
 	ceph_free_cap_flush(prealloc_cf);
-	doutc(cl, "%llx.%llx inline_version %llu = %d\n",
+	boutc(cl, "%llx.%llx inline_version %llu = %d\n",
 	      ceph_vinop(inode), inline_version, err);
 	return err;
 }
@@ -2648,10 +2696,13 @@ static int __ceph_pool_perm_get(struct ceph_inode_info *ci,
 		goto out;
 
 	if (pool_ns)
-		doutc(cl, "pool %lld ns %.*s no perm cached\n", pool,
-		      (int)pool_ns->len, pool_ns->str);
+		boutc_bounded(cl, "pool %lld ns %.*s no perm cached\n",
+			      (pool, (int)pool_ns->len,
+			       BLOG_STR(pool_ns->str, pool_ns->len)),
+			      (pool, (int)pool_ns->len,
+			       (const char *)pool_ns->str));
 	else
-		doutc(cl, "pool %lld no perm cached\n", pool);
+		boutc(cl, "pool %lld no perm cached\n", pool);
 
 	down_write(&mdsc->pool_perm_rwsem);
 	p = &mdsc->pool_perm_tree.rb_node;
@@ -2717,7 +2768,7 @@ static int __ceph_pool_perm_get(struct ceph_inode_info *ci,
 		goto out_unlock;
 
 	/* one page should be large enough for STAT data */
-	pages = ceph_alloc_page_vector(1, GFP_KERNEL);
+	pages = ceph_alloc_page_vector(1, GFP_NOFS);
 	if (IS_ERR(pages)) {
 		err = PTR_ERR(pages);
 		goto out_unlock;
@@ -2776,10 +2827,13 @@ static int __ceph_pool_perm_get(struct ceph_inode_info *ci,
 	if (!err)
 		err = have;
 	if (pool_ns)
-		doutc(cl, "pool %lld ns %.*s result = %d\n", pool,
-		      (int)pool_ns->len, pool_ns->str, err);
+		boutc_bounded(cl, "pool %lld ns %.*s result = %d\n",
+			      (pool, (int)pool_ns->len,
+			       BLOG_STR(pool_ns->str, pool_ns->len), err),
+			      (pool, (int)pool_ns->len,
+			       (const char *)pool_ns->str, err));
 	else
-		doutc(cl, "pool %lld result = %d\n", pool, err);
+		boutc(cl, "pool %lld result = %d\n", pool, err);
 	return err;
 }
 
@@ -2816,11 +2870,11 @@ int ceph_pool_perm_check(struct inode *inode, int need)
 check:
 	if (flags & CEPH_I_POOL_PERM) {
 		if ((need & CEPH_CAP_FILE_RD) && !(flags & CEPH_I_POOL_RD)) {
-			doutc(cl, "pool %lld no read perm\n", pool);
+			boutc(cl, "pool %lld no read perm\n", pool);
 			return -EPERM;
 		}
 		if ((need & CEPH_CAP_FILE_WR) && !(flags & CEPH_I_POOL_WR)) {
-			doutc(cl, "pool %lld no write perm\n", pool);
+			boutc(cl, "pool %lld no write perm\n", pool);
 			return -EPERM;
 		}
 		return 0;
diff --git a/fs/ceph/file.c b/fs/ceph/file.c
index fdf5b6918c09..856710c63765 100644
--- a/fs/ceph/file.c
+++ b/fs/ceph/file.c
@@ -69,7 +69,7 @@ static __le32 ceph_flags_sys2wire(struct ceph_mds_client *mdsc, u32 flags)
 #undef ceph_sys2wire
 
 	if (flags)
-		doutc(cl, "unused open flags: %x\n", flags);
+		boutc(cl, "unused open flags: %x\n", flags);
 
 	return cpu_to_le32(wire_flags);
 }
@@ -208,7 +208,7 @@ static int ceph_init_file_info(struct inode *inode, struct file *file,
 	struct ceph_file_info *fi;
 	int ret;
 
-	doutc(cl, "%p %llx.%llx %p 0%o (%s)\n", inode, ceph_vinop(inode),
+	boutc(cl, "%p %llx.%llx %p 0%o (%s)\n", inode, ceph_vinop(inode),
 	      file, inode->i_mode, isdir ? "dir" : "regular");
 	BUG_ON(inode->i_fop->release != ceph_release);
 
@@ -276,12 +276,12 @@ static int ceph_init_file(struct inode *inode, struct file *file, int fmode)
 		break;
 
 	case S_IFLNK:
-		doutc(cl, "%p %llx.%llx %p 0%o (symlink)\n", inode,
+		boutc(cl, "%p %llx.%llx %p 0%o (symlink)\n", inode,
 		      ceph_vinop(inode), file, inode->i_mode);
 		break;
 
 	default:
-		doutc(cl, "%p %llx.%llx %p 0%o (special)\n", inode,
+		boutc(cl, "%p %llx.%llx %p 0%o (special)\n", inode,
 		      ceph_vinop(inode), file, inode->i_mode);
 		/*
 		 * we need to drop the open ref now, since we don't
@@ -314,7 +314,7 @@ int ceph_renew_caps(struct inode *inode, int fmode)
 	    (!(wanted & CEPH_CAP_ANY_WR) || ci->i_auth_cap) &&
 	    (issued & wanted) == wanted) {
 		spin_unlock(&ci->i_ceph_lock);
-		doutc(cl, "%p %llx.%llx want %s issued %s updating mds_wanted\n",
+		boutc(cl, "%p %llx.%llx want %s issued %s updating mds_wanted\n",
 		      inode, ceph_vinop(inode), ceph_cap_string(wanted),
 		      ceph_cap_string(issued));
 		ceph_check_caps(ci, 0);
@@ -347,7 +347,7 @@ int ceph_renew_caps(struct inode *inode, int fmode)
 	err = ceph_mdsc_do_request(mdsc, NULL, req);
 	ceph_mdsc_put_request(req);
 out:
-	doutc(cl, "%p %llx.%llx open result=%d\n", inode, ceph_vinop(inode),
+	boutc(cl, "%p %llx.%llx open result=%d\n", inode, ceph_vinop(inode),
 	      err);
 	return err < 0 ? err : 0;
 }
@@ -372,9 +372,13 @@ int ceph_open(struct inode *inode, struct file *file)
 	char *path;
 	bool do_sync = false;
 	int mask = MAY_READ;
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(fsc, &__ji);
 
 	if (fi) {
-		doutc(cl, "file %p is already opened\n", file);
+		boutc(cl, "file %p is already opened\n", file);
+		ceph_blog_exit(&__ji);
 		return 0;
 	}
 
@@ -384,11 +388,13 @@ int ceph_open(struct inode *inode, struct file *file)
 		flags = O_DIRECTORY;  /* mds likes to know */
 	} else if (S_ISREG(inode->i_mode)) {
 		err = fscrypt_file_open(inode, file);
-		if (err)
+		if (err) {
+			ceph_blog_exit(&__ji);
 			return err;
+		}
 	}
 
-	doutc(cl, "%p %llx.%llx file %p flags %d (%d)\n", inode,
+	boutc(cl, "%p %llx.%llx file %p flags %d (%d)\n", inode,
 	      ceph_vinop(inode), file, flags, file->f_flags);
 	fmode = ceph_flags_to_mode(flags);
 
@@ -425,6 +431,7 @@ int ceph_open(struct inode *inode, struct file *file)
 
 		/* For none EACCES cases will let the MDS do the mds auth check */
 		if (err == -EACCES) {
+			ceph_blog_exit(&__ji);
 			return err;
 		} else if (err < 0) {
 			do_sync = true;
@@ -433,11 +440,14 @@ int ceph_open(struct inode *inode, struct file *file)
 	}
 
 	/* snapped files are read-only */
-	if (ceph_in_snap(inode) && (file->f_mode & FMODE_WRITE))
+	if (ceph_in_snap(inode) && (file->f_mode & FMODE_WRITE)) {
+		ceph_blog_exit(&__ji);
 		return -EROFS;
+	}
 
 	/* trivially open snapdir */
 	if (ceph_snap(inode) == CEPH_SNAPDIR) {
+		ceph_blog_exit(&__ji);
 		return ceph_init_file(inode, file, fmode);
 	}
 
@@ -452,7 +462,7 @@ int ceph_open(struct inode *inode, struct file *file)
 		int mds_wanted = __ceph_caps_mds_wanted(ci, true);
 		int issued = __ceph_caps_issued(ci, NULL);
 
-		doutc(cl, "open %p fmode %d want %s issued %s using existing\n",
+		boutc(cl, "open %p fmode %d want %s issued %s using existing\n",
 		      inode, fmode, ceph_cap_string(wanted),
 		      ceph_cap_string(issued));
 		__ceph_touch_fmode(ci, mdsc, fmode);
@@ -464,17 +474,19 @@ int ceph_open(struct inode *inode, struct file *file)
 		    ceph_snap(inode) != CEPH_SNAPDIR)
 			ceph_check_caps(ci, 0);
 
+		ceph_blog_exit(&__ji);
 		return ceph_init_file(inode, file, fmode);
 	} else if (!do_sync && ceph_in_snap(inode) &&
 		   (ci->i_snap_caps & wanted) == wanted) {
 		__ceph_touch_fmode(ci, mdsc, fmode);
 		spin_unlock(&ci->i_ceph_lock);
+		ceph_blog_exit(&__ji);
 		return ceph_init_file(inode, file, fmode);
 	}
 
 	spin_unlock(&ci->i_ceph_lock);
 
-	doutc(cl, "open fmode %d wants %s\n", fmode, ceph_cap_string(wanted));
+	boutc(cl, "open fmode %d wants %s\n", fmode, ceph_cap_string(wanted));
 	req = prepare_open_request(inode->i_sb, flags, 0);
 	if (IS_ERR(req)) {
 		err = PTR_ERR(req);
@@ -491,8 +503,9 @@ int ceph_open(struct inode *inode, struct file *file)
 	if (!err)
 		err = ceph_init_file(inode, file, fmode);
 	ceph_mdsc_put_request(req);
-	doutc(cl, "open result=%d on %llx.%llx\n", err, ceph_vinop(inode));
+	boutc(cl, "open result=%d on %llx.%llx\n", err, ceph_vinop(inode));
 out:
+	ceph_blog_exit(&__ji);
 	return err;
 }
 
@@ -746,7 +759,7 @@ static int ceph_finish_async_create(struct inode *dir, struct inode *inode,
 			      req->r_fmode, NULL);
 	up_read(&mdsc->snap_rwsem);
 	if (ret) {
-		doutc(cl, "failed to fill inode: %d\n", ret);
+		boutc(cl, "failed to fill inode: %d\n", ret);
 		ceph_dir_clear_complete(dir);
 		if (!d_unhashed(dentry))
 			d_drop(dentry);
@@ -754,7 +767,7 @@ static int ceph_finish_async_create(struct inode *dir, struct inode *inode,
 	} else {
 		struct dentry *dn;
 
-		doutc(cl, "d_adding new inode 0x%llx to 0x%llx/%s\n",
+		boutc(cl, "d_adding new inode 0x%llx to 0x%llx/%s\n",
 		      vino.ino, ceph_ino(dir), dentry->d_name.name);
 		ceph_dir_clear_ordered(dir);
 		ceph_init_inode_acls(inode, as_ctx);
@@ -805,17 +818,24 @@ int ceph_atomic_open(struct inode *dir, struct dentry *dentry,
 	int mask;
 	int err;
 	char *path;
+	struct ceph_journal_info __ji;
 
-	doutc(cl, "%p %llx.%llx dentry %p '%pd' %s flags %d mode 0%o\n",
-	      dir, ceph_vinop(dir), dentry, dentry,
+	ceph_blog_enter(fsc, &__ji);
+
+	boutc(cl, "%p %llx.%llx dentry %p '%s' %s flags %d mode 0%o\n",
+	      dir, ceph_vinop(dir), dentry, dentry->d_name.name,
 	      d_unhashed(dentry) ? "unhashed" : "hashed", flags, mode);
 
-	if (dentry->d_name.len > NAME_MAX)
+	if (dentry->d_name.len > NAME_MAX) {
+		ceph_blog_exit(&__ji);
 		return -ENAMETOOLONG;
+	}
 
 	err = ceph_wait_on_conflict_unlink(dentry);
-	if (err)
+	if (err) {
+		ceph_blog_exit(&__ji);
 		return err;
+	}
 	/*
 	 * Do not truncate the file, since atomic_open is called before the
 	 * permission check. The caller will do the truncation afterward.
@@ -847,6 +867,7 @@ int ceph_atomic_open(struct inode *dir, struct dentry *dentry,
 
 		/* For none EACCES cases will let the MDS do the mds auth check */
 		if (err == -EACCES) {
+			ceph_blog_exit(&__ji);
 			return err;
 		} else if (err < 0) {
 			try_async = false;
@@ -856,8 +877,10 @@ int ceph_atomic_open(struct inode *dir, struct dentry *dentry,
 
 retry:
 	if (flags & O_CREAT) {
-		if (ceph_quota_is_max_files_exceeded(dir))
+		if (ceph_quota_is_max_files_exceeded(dir)) {
+			ceph_blog_exit(&__ji);
 			return -EDQUOT;
+		}
 
 		new_inode = ceph_new_inode(dir, dentry, &mode, &as_ctx);
 		if (IS_ERR(new_inode)) {
@@ -870,6 +893,7 @@ int ceph_atomic_open(struct inode *dir, struct dentry *dentry,
 			try_async = false;
 	} else if (!d_in_lookup(dentry)) {
 		/* If it's not being looked up, it's negative */
+		ceph_blog_exit(&__ji);
 		return -ENOENT;
 	}
 
@@ -984,7 +1008,7 @@ int ceph_atomic_open(struct inode *dir, struct dentry *dentry,
 		goto out_req;
 	if (dn || d_really_is_negative(dentry) || d_is_symlink(dentry)) {
 		/* make vfs retry on splice, ENOENT, or symlink */
-		doutc(cl, "finish_no_open on dn %p\n", dn);
+		boutc(cl, "finish_no_open on dn %p\n", dn);
 		err = finish_no_open(file, dn);
 	} else {
 		if (IS_ENCRYPTED(dir) &&
@@ -995,7 +1019,7 @@ int ceph_atomic_open(struct inode *dir, struct dentry *dentry,
 			goto out_req;
 		}
 
-		doutc(cl, "finish_open on dn %p\n", dn);
+		boutc(cl, "finish_open on dn %p\n", dn);
 		if (req->r_op == CEPH_MDS_OP_CREATE && req->r_reply_info.has_create_ino) {
 			struct inode *newino = d_inode(dentry);
 
@@ -1014,18 +1038,23 @@ int ceph_atomic_open(struct inode *dir, struct dentry *dentry,
 	iput(new_inode);
 out_ctx:
 	ceph_release_acl_sec_ctx(&as_ctx);
-	doutc(cl, "result=%d\n", err);
+	boutc(cl, "result=%d\n", err);
+	ceph_blog_exit(&__ji);
 	return err;
 }
 
 int ceph_release(struct inode *inode, struct file *file)
 {
+	struct ceph_fs_client *fsc = ceph_inode_to_fs_client(inode);
 	struct ceph_client *cl = ceph_inode_to_client(inode);
 	struct ceph_inode_info *ci = ceph_inode(inode);
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(fsc, &__ji);
 
 	if (S_ISDIR(inode->i_mode)) {
 		struct ceph_dir_file_info *dfi = file->private_data;
-		doutc(cl, "%p %llx.%llx dir file %p\n", inode,
+		boutc(cl, "%p %llx.%llx dir file %p\n", inode,
 		      ceph_vinop(inode), file);
 		WARN_ON(!list_empty(&dfi->file_info.rw_contexts));
 
@@ -1038,7 +1067,7 @@ int ceph_release(struct inode *inode, struct file *file)
 		kmem_cache_free(ceph_dir_file_cachep, dfi);
 	} else {
 		struct ceph_file_info *fi = file->private_data;
-		doutc(cl, "%p %llx.%llx regular file %p\n", inode,
+		boutc(cl, "%p %llx.%llx regular file %p\n", inode,
 		      ceph_vinop(inode), file);
 		WARN_ON(!list_empty(&fi->rw_contexts));
 
@@ -1050,6 +1079,7 @@ int ceph_release(struct inode *inode, struct file *file)
 
 	/* wake up anyone waiting for caps on this inode */
 	wake_up_all(&ci->i_cap_wq);
+	ceph_blog_exit(&__ji);
 	return 0;
 }
 
@@ -1084,7 +1114,7 @@ ssize_t __ceph_sync_read(struct inode *inode, loff_t *ki_pos,
 	bool sparse = IS_ENCRYPTED(inode) || ceph_test_mount_opt(fsc, SPARSEREAD);
 	u64 objver = 0;
 
-	doutc(cl, "on inode %p %llx.%llx %llx~%llx\n", inode,
+	boutc(cl, "on inode %p %llx.%llx %llx~%llx\n", inode,
 	      ceph_vinop(inode), *ki_pos, len);
 
 	if (ceph_inode_is_shutdown(inode))
@@ -1120,7 +1150,7 @@ ssize_t __ceph_sync_read(struct inode *inode, loff_t *ki_pos,
 		/* determine new offset/length if encrypted */
 		ceph_fscrypt_adjust_off_and_len(inode, &read_off, &read_len);
 
-		doutc(cl, "orig %llu~%llu reading %llu~%llu", off, len,
+		boutc(cl, "orig %llu~%llu reading %llu~%llu", off, len,
 		      read_off, read_len);
 
 		req = ceph_osdc_new_request(osdc, &ci->i_layout,
@@ -1184,7 +1214,7 @@ ssize_t __ceph_sync_read(struct inode *inode, loff_t *ki_pos,
 			objver = req->r_version;
 
 		i_size = i_size_read(inode);
-		doutc(cl, "%llu~%llu got %zd i_size %llu%s\n", off, len,
+		boutc(cl, "%llu~%llu got %zd i_size %llu%s\n", off, len,
 		      ret, i_size, (more ? " MORE" : ""));
 
 		/* Fix it to go to end of extent map */
@@ -1231,7 +1261,7 @@ ssize_t __ceph_sync_read(struct inode *inode, loff_t *ki_pos,
 			int zlen = min(len - ret, i_size - off - ret);
 			int zoff = page_off + ret;
 
-			doutc(cl, "zero gap %llu~%llu\n", off + ret,
+			boutc(cl, "zero gap %llu~%llu\n", off + ret,
 			      off + ret + zlen);
 			ceph_zero_page_vector_range(zoff, zlen, pages);
 			ret += zlen;
@@ -1277,7 +1307,7 @@ ssize_t __ceph_sync_read(struct inode *inode, loff_t *ki_pos,
 		if (last_objver)
 			*last_objver = objver;
 	}
-	doutc(cl, "result %zd retry_op %d\n", ret, *retry_op);
+	boutc(cl, "result %zd retry_op %d\n", ret, *retry_op);
 	return ret;
 }
 
@@ -1288,7 +1318,7 @@ static ssize_t ceph_sync_read(struct kiocb *iocb, struct iov_iter *to,
 	struct inode *inode = file_inode(file);
 	struct ceph_client *cl = ceph_inode_to_client(inode);
 
-	doutc(cl, "on file %p %llx~%zx %s\n", file, iocb->ki_pos,
+	boutc(cl, "on file %p %llx~%zx %s\n", file, iocb->ki_pos,
 	      iov_iter_count(to),
 	      (file->f_flags & O_DIRECT) ? "O_DIRECT" : "");
 
@@ -1333,7 +1363,7 @@ static void ceph_aio_complete(struct inode *inode,
 	if (!ret)
 		ret = aio_req->total_len;
 
-	doutc(cl, "%p %llx.%llx rc %d\n", inode, ceph_vinop(inode), ret);
+	boutc(cl, "%p %llx.%llx rc %d\n", inode, ceph_vinop(inode), ret);
 
 	if (ret >= 0 && aio_req->write) {
 		int dirty;
@@ -1377,7 +1407,7 @@ static void ceph_aio_complete_req(struct ceph_osd_request *req)
 	BUG_ON(osd_data->type != CEPH_OSD_DATA_TYPE_BVECS);
 	BUG_ON(!osd_data->num_bvecs);
 
-	doutc(cl, "req %p inode %p %llx.%llx, rc %d bytes %u\n", req,
+	boutc(cl, "req %p inode %p %llx.%llx, rc %d bytes %u\n", req,
 	      inode, ceph_vinop(inode), rc, len);
 
 	if (rc == -EOLDSNAPC) {
@@ -1550,7 +1580,7 @@ ceph_direct_read_write(struct kiocb *iocb, struct iov_iter *iter,
 	if (write && ceph_in_snap(file_inode(file)))
 		return -EROFS;
 
-	doutc(cl, "sync_direct_%s on file %p %lld~%u snapc %p seq %lld\n",
+	boutc(cl, "sync_direct_%s on file %p %lld~%u snapc %p seq %lld\n",
 	      (write ? "write" : "read"), file, pos, (unsigned)count,
 	      snapc, snapc ? snapc->seq : 0);
 
@@ -1563,7 +1593,7 @@ ceph_direct_read_write(struct kiocb *iocb, struct iov_iter *iter,
 					pos >> PAGE_SHIFT,
 					(pos + count - 1) >> PAGE_SHIFT);
 		if (ret2 < 0)
-			doutc(cl, "invalidate_inode_pages2_range returned %d\n",
+			boutc(cl, "invalidate_inode_pages2_range returned %d\n",
 			      ret2);
 
 		flags = /* CEPH_OSD_FLAG_ORDERSNAP | */ CEPH_OSD_FLAG_WRITE;
@@ -1787,7 +1817,7 @@ ceph_sync_write(struct kiocb *iocb, struct iov_iter *from, loff_t pos,
 	if (ceph_in_snap(file_inode(file)))
 		return -EROFS;
 
-	doutc(cl, "on file %p %lld~%u snapc %p seq %lld\n", file, pos,
+	boutc(cl, "on file %p %lld~%u snapc %p seq %lld\n", file, pos,
 	      (unsigned)count, snapc, snapc->seq);
 
 	ret = filemap_write_and_wait_range(inode->i_mapping,
@@ -1832,7 +1862,7 @@ ceph_sync_write(struct kiocb *iocb, struct iov_iter *from, loff_t pos,
 		last = (pos + len) != (write_pos + write_len);
 		rmw = first || last;
 
-		doutc(cl, "ino %llx %lld~%llu adjusted %lld~%llu -- %srmw\n",
+		boutc(cl, "ino %llx %lld~%llu adjusted %lld~%llu -- %srmw\n",
 		      ceph_ino(inode), pos, len, write_pos, write_len,
 		      rmw ? "" : "no ");
 
@@ -2048,7 +2078,7 @@ ceph_sync_write(struct kiocb *iocb, struct iov_iter *from, loff_t pos,
 			left -= ret;
 		}
 		if (ret < 0) {
-			doutc(cl, "write failed with %d\n", ret);
+			boutc(cl, "write failed with %d\n", ret);
 			ceph_release_page_vector(pages, num_pages);
 			break;
 		}
@@ -2057,7 +2087,7 @@ ceph_sync_write(struct kiocb *iocb, struct iov_iter *from, loff_t pos,
 			ret = ceph_fscrypt_encrypt_pages(inode, pages,
 							 write_pos, write_len);
 			if (ret < 0) {
-				doutc(cl, "encryption failed with %d\n", ret);
+				boutc(cl, "encryption failed with %d\n", ret);
 				ceph_release_page_vector(pages, num_pages);
 				break;
 			}
@@ -2076,7 +2106,7 @@ ceph_sync_write(struct kiocb *iocb, struct iov_iter *from, loff_t pos,
 			break;
 		}
 
-		doutc(cl, "write op %lld~%llu\n", write_pos, write_len);
+		boutc(cl, "write op %lld~%llu\n", write_pos, write_len);
 		osd_req_op_extent_osd_data_pages(req, rmw ? 1 : 0, pages, write_len,
 						 offset_in_page(write_pos), false,
 						 true);
@@ -2112,7 +2142,7 @@ ceph_sync_write(struct kiocb *iocb, struct iov_iter *from, loff_t pos,
 						 write_len);
 		ceph_osdc_put_request(req);
 		if (ret != 0) {
-			doutc(cl, "osd write returned %d\n", ret);
+			boutc(cl, "osd write returned %d\n", ret);
 			/* Version changed! Must re-do the rmw cycle */
 			if ((assert_ver && (ret == -ERANGE || ret == -EOVERFLOW)) ||
 			    (!assert_ver && ret == -EEXIST)) {
@@ -2142,13 +2172,13 @@ ceph_sync_write(struct kiocb *iocb, struct iov_iter *from, loff_t pos,
 				pos >> PAGE_SHIFT,
 				(pos + len - 1) >> PAGE_SHIFT);
 		if (ret < 0) {
-			doutc(cl, "invalidate_inode_pages2_range returned %d\n",
+			boutc(cl, "invalidate_inode_pages2_range returned %d\n",
 			      ret);
 			ret = 0;
 		}
 		pos += len;
 		written += len;
-		doutc(cl, "written %d\n", written);
+		boutc(cl, "written %d\n", written);
 		if (pos > i_size_read(inode)) {
 			check_caps = ceph_inode_set_size(inode, pos);
 			if (check_caps)
@@ -2162,7 +2192,7 @@ ceph_sync_write(struct kiocb *iocb, struct iov_iter *from, loff_t pos,
 		ret = written;
 		iocb->ki_pos = pos;
 	}
-	doutc(cl, "returning %d\n", ret);
+	boutc(cl, "returning %d\n", ret);
 	return ret;
 }
 
@@ -2181,22 +2211,30 @@ static ssize_t ceph_read_iter(struct kiocb *iocb, struct iov_iter *to)
 	struct inode *inode = file_inode(filp);
 	struct ceph_inode_info *ci = ceph_inode(inode);
 	bool direct_lock = iocb->ki_flags & IOCB_DIRECT;
+	struct ceph_fs_client *fsc = ceph_inode_to_fs_client(inode);
 	struct ceph_client *cl = ceph_inode_to_client(inode);
 	ssize_t ret;
 	int want = 0, got = 0;
 	int retry_op = 0, read = 0;
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(fsc, &__ji);
 
 again:
-	doutc(cl, "%llu~%u trying to get caps on %p %llx.%llx\n",
+	boutc(cl, "%llu~%u trying to get caps on %p %llx.%llx\n",
 	      iocb->ki_pos, (unsigned)len, inode, ceph_vinop(inode));
 
-	if (ceph_inode_is_shutdown(inode))
+	if (ceph_inode_is_shutdown(inode)) {
+		ceph_blog_exit(&__ji);
 		return -ESTALE;
+	}
 
 	ret = direct_lock ? ceph_start_io_direct(inode) :
 			    ceph_start_io_read(inode);
-	if (ret)
+	if (ret) {
+		ceph_blog_exit(&__ji);
 		return ret;
+	}
 
 	if (!(fi->flags & CEPH_F_SYNC) && !direct_lock)
 		want |= CEPH_CAP_FILE_CACHE;
@@ -2209,6 +2247,7 @@ static ssize_t ceph_read_iter(struct kiocb *iocb, struct iov_iter *to)
 			ceph_end_io_direct(inode);
 		else
 			ceph_end_io_read(inode);
+		ceph_blog_exit(&__ji);
 		return ret;
 	}
 
@@ -2216,7 +2255,7 @@ static ssize_t ceph_read_iter(struct kiocb *iocb, struct iov_iter *to)
 	    (iocb->ki_flags & IOCB_DIRECT) ||
 	    (fi->flags & CEPH_F_SYNC)) {
 
-		doutc(cl, "sync %p %llx.%llx %llu~%u got cap refs on %s\n",
+		boutc(cl, "sync %p %llx.%llx %llu~%u got cap refs on %s\n",
 		      inode, ceph_vinop(inode), iocb->ki_pos, (unsigned)len,
 		      ceph_cap_string(got));
 
@@ -2236,7 +2275,7 @@ static ssize_t ceph_read_iter(struct kiocb *iocb, struct iov_iter *to)
 		}
 	} else {
 		CEPH_DEFINE_RW_CONTEXT(rw_ctx, got);
-		doutc(cl, "async %p %llx.%llx %llu~%u got cap refs on %s\n",
+		boutc(cl, "async %p %llx.%llx %llu~%u got cap refs on %s\n",
 		      inode, ceph_vinop(inode), iocb->ki_pos, (unsigned)len,
 		      ceph_cap_string(got));
 		ceph_add_rw_context(fi, &rw_ctx);
@@ -2244,7 +2283,7 @@ static ssize_t ceph_read_iter(struct kiocb *iocb, struct iov_iter *to)
 		ceph_del_rw_context(fi, &rw_ctx);
 	}
 
-	doutc(cl, "%p %llx.%llx dropping cap refs on %s = %d\n",
+	boutc(cl, "%p %llx.%llx dropping cap refs on %s = %d\n",
 	      inode, ceph_vinop(inode), ceph_cap_string(got), (int)ret);
 	ceph_put_cap_refs(ci, got);
 
@@ -2260,8 +2299,10 @@ static ssize_t ceph_read_iter(struct kiocb *iocb, struct iov_iter *to)
 		int mask = CEPH_STAT_CAP_SIZE;
 		if (retry_op == READ_INLINE) {
 			folio = folio_alloc(GFP_KERNEL, 0);
-			if (!folio)
+			if (!folio) {
+				ceph_blog_exit(&__ji);
 				return -ENOMEM;
+			}
 
 			mask = CEPH_STAT_CAP_INLINE_DATA;
 		}
@@ -2274,6 +2315,7 @@ static ssize_t ceph_read_iter(struct kiocb *iocb, struct iov_iter *to)
 				BUG_ON(retry_op != READ_INLINE);
 				goto again;
 			}
+			ceph_blog_exit(&__ji);
 			return statret;
 		}
 
@@ -2301,13 +2343,14 @@ static ssize_t ceph_read_iter(struct kiocb *iocb, struct iov_iter *to)
 				read += ret;
 			}
 			folio_put(folio);
+			ceph_blog_exit(&__ji);
 			return read;
 		}
 
 		/* hit EOF or hole? */
 		if (retry_op == CHECK_EOF && iocb->ki_pos < i_size &&
 		    ret < len) {
-			doutc(cl, "may hit hole, ppos %lld < size %lld, reading more\n",
+			boutc(cl, "may hit hole, ppos %lld < size %lld, reading more\n",
 			      iocb->ki_pos, i_size);
 
 			read += ret;
@@ -2320,6 +2363,7 @@ static ssize_t ceph_read_iter(struct kiocb *iocb, struct iov_iter *to)
 	if (ret >= 0)
 		ret += read;
 
+	ceph_blog_exit(&__ji);
 	return ret;
 }
 
@@ -2334,24 +2378,35 @@ static ssize_t ceph_splice_read(struct file *in, loff_t *ppos,
 {
 	struct ceph_file_info *fi = in->private_data;
 	struct inode *inode = file_inode(in);
+	struct ceph_fs_client *fsc = ceph_inode_to_fs_client(inode);
+	struct ceph_client *cl = fsc->client;
 	struct ceph_inode_info *ci = ceph_inode(inode);
 	ssize_t ret;
 	int want = 0, got = 0;
 	CEPH_DEFINE_RW_CONTEXT(rw_ctx, 0);
+	struct ceph_journal_info __ji;
 
-	dout("splice_read %p %llx.%llx %llu~%zu trying to get caps on %p\n",
+	ceph_blog_enter(fsc, &__ji);
+
+	boutc(cl, "splice_read %p %llx.%llx %llu~%zu trying to get caps on %p\n",
 	     inode, ceph_vinop(inode), *ppos, len, inode);
 
-	if (ceph_inode_is_shutdown(inode))
+	if (ceph_inode_is_shutdown(inode)) {
+		ceph_blog_exit(&__ji);
 		return -ESTALE;
+	}
 
 	if (ceph_has_inline_data(ci) ||
-	    (fi->flags & CEPH_F_SYNC))
+	    (fi->flags & CEPH_F_SYNC)) {
+		ceph_blog_exit(&__ji);
 		return copy_splice_read(in, ppos, pipe, len, flags);
+	}
 
 	ret = ceph_start_io_read(inode);
-	if (ret)
+	if (ret) {
+		ceph_blog_exit(&__ji);
 		return ret;
+	}
 
 	want = CEPH_CAP_FILE_CACHE;
 	if (fi->fmode & CEPH_FILE_MODE_LAZY)
@@ -2362,16 +2417,17 @@ static ssize_t ceph_splice_read(struct file *in, loff_t *ppos,
 		goto out_end;
 
 	if ((got & (CEPH_CAP_FILE_CACHE | CEPH_CAP_FILE_LAZYIO)) == 0) {
-		dout("splice_read/sync %p %llx.%llx %llu~%zu got cap refs on %s\n",
+		boutc(cl, "splice_read/sync %p %llx.%llx %llu~%zu got cap refs on %s\n",
 		     inode, ceph_vinop(inode), *ppos, len,
 		     ceph_cap_string(got));
 
 		ceph_put_cap_refs(ci, got);
 		ceph_end_io_read(inode);
+		ceph_blog_exit(&__ji);
 		return copy_splice_read(in, ppos, pipe, len, flags);
 	}
 
-	dout("splice_read %p %llx.%llx %llu~%zu got cap refs on %s\n",
+	boutc(cl, "splice_read %p %llx.%llx %llu~%zu got cap refs on %s\n",
 	     inode, ceph_vinop(inode), *ppos, len, ceph_cap_string(got));
 
 	rw_ctx.caps = got;
@@ -2379,12 +2435,13 @@ static ssize_t ceph_splice_read(struct file *in, loff_t *ppos,
 	ret = filemap_splice_read(in, ppos, pipe, len, flags);
 	ceph_del_rw_context(fi, &rw_ctx);
 
-	dout("splice_read %p %llx.%llx dropping cap refs on %s = %zd\n",
+	boutc(cl, "splice_read %p %llx.%llx dropping cap refs on %s = %zd\n",
 	     inode, ceph_vinop(inode), ceph_cap_string(got), ret);
 
 	ceph_put_cap_refs(ci, got);
 out_end:
 	ceph_end_io_read(inode);
+	ceph_blog_exit(&__ji);
 	return ret;
 }
 
@@ -2416,16 +2473,25 @@ static ssize_t ceph_write_iter(struct kiocb *iocb, struct iov_iter *from)
 	u64 pool_flags;
 	loff_t pos;
 	loff_t limit = max(i_size_read(inode), fsc->max_file_size);
+	struct ceph_journal_info __ji;
 
-	if (ceph_inode_is_shutdown(inode))
+	ceph_blog_enter(fsc, &__ji);
+
+	if (ceph_inode_is_shutdown(inode)) {
+		ceph_blog_exit(&__ji);
 		return -ESTALE;
+	}
 
-	if (ceph_in_snap(inode))
+	if (ceph_in_snap(inode)) {
+		ceph_blog_exit(&__ji);
 		return -EROFS;
+	}
 
 	prealloc_cf = ceph_alloc_cap_flush();
-	if (!prealloc_cf)
+	if (!prealloc_cf) {
+		ceph_blog_exit(&__ji);
 		return -ENOMEM;
+	}
 
 	if ((iocb->ki_flags & (IOCB_DIRECT | IOCB_APPEND)) == IOCB_DIRECT)
 		direct_lock = true;
@@ -2474,7 +2540,7 @@ static ssize_t ceph_write_iter(struct kiocb *iocb, struct iov_iter *from)
 	if (err)
 		goto out;
 
-	doutc(cl, "%p %llx.%llx %llu~%zd getting caps. i_size %llu\n",
+	boutc(cl, "%p %llx.%llx %llu~%zd getting caps. i_size %llu\n",
 	      inode, ceph_vinop(inode), pos, count,
 	      i_size_read(inode));
 	if (!(fi->flags & CEPH_F_SYNC) && !direct_lock)
@@ -2498,7 +2564,7 @@ static ssize_t ceph_write_iter(struct kiocb *iocb, struct iov_iter *from)
 		loff_t cur_eof = i_size_read(inode);
 
 		if (cur_eof != pos) {
-			doutc(cl,
+			boutc(cl,
 			      "%p %llx.%llx O_APPEND: pos adjusted %lld -> %lld\n",
 			      inode, ceph_vinop(inode), pos, cur_eof);
 			iocb->ki_pos = cur_eof;
@@ -2540,7 +2606,7 @@ static ssize_t ceph_write_iter(struct kiocb *iocb, struct iov_iter *from)
 
 	inode_inc_iversion_raw(inode);
 
-	doutc(cl, "%p %llx.%llx %llu~%zd got cap refs on %s\n",
+	boutc(cl, "%p %llx.%llx %llu~%zd got cap refs on %s\n",
 	      inode, ceph_vinop(inode), pos, count, ceph_cap_string(got));
 
 	if ((got & (CEPH_CAP_FILE_BUFFER|CEPH_CAP_FILE_LAZYIO)) == 0 ||
@@ -2601,13 +2667,13 @@ static ssize_t ceph_write_iter(struct kiocb *iocb, struct iov_iter *from)
 			ceph_check_caps(ci, CHECK_CAPS_FLUSH);
 	}
 
-	doutc(cl, "%p %llx.%llx %llu~%u  dropping cap refs on %s\n",
+	boutc(cl, "%p %llx.%llx %llu~%u  dropping cap refs on %s\n",
 	      inode, ceph_vinop(inode), pos, (unsigned)count,
 	      ceph_cap_string(got));
 	ceph_put_cap_refs(ci, got);
 
 	if (written == -EOLDSNAPC) {
-		doutc(cl, "%p %llx.%llx %llu~%u" "got EOLDSNAPC, retrying\n",
+		boutc(cl, "%p %llx.%llx %llu~%u" "got EOLDSNAPC, retrying\n",
 		      inode, ceph_vinop(inode), pos, (unsigned)count);
 		goto retry_snap;
 	}
@@ -2630,6 +2696,7 @@ static ssize_t ceph_write_iter(struct kiocb *iocb, struct iov_iter *from)
 		ceph_end_io_write(inode);
 out_unlocked:
 	ceph_free_cap_flush(prealloc_cf);
+	ceph_blog_exit(&__ji);
 	return written ? written : err;
 }
 
@@ -2638,14 +2705,22 @@ static ssize_t ceph_write_iter(struct kiocb *iocb, struct iov_iter *from)
  */
 static loff_t ceph_llseek(struct file *file, loff_t offset, int whence)
 {
+	struct inode *inode = file_inode(file);
+	struct ceph_fs_client *fsc = ceph_inode_to_fs_client(inode);
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(fsc, &__ji);
+
 	if (whence == SEEK_END || whence == SEEK_DATA || whence == SEEK_HOLE) {
-		struct inode *inode = file_inode(file);
 		int ret;
 
 		ret = ceph_do_getattr(inode, CEPH_STAT_CAP_SIZE, false);
-		if (ret < 0)
+		if (ret < 0) {
+			ceph_blog_exit(&__ji);
 			return ret;
+		}
 	}
+	ceph_blog_exit(&__ji);
 	return generic_file_llseek(file, offset, whence);
 }
 
@@ -2802,22 +2877,34 @@ static long ceph_fallocate(struct file *file, int mode,
 	int ret = 0;
 	loff_t endoff = 0;
 	loff_t size;
+	struct ceph_fs_client *fsc = ceph_inode_to_fs_client(inode);
+	struct ceph_journal_info __ji;
 
-	doutc(cl, "%p %llx.%llx mode %x, offset %llu length %llu\n",
+	ceph_blog_enter(fsc, &__ji);
+
+	boutc(cl, "%p %llx.%llx mode %x, offset %llu length %llu\n",
 	      inode, ceph_vinop(inode), mode, offset, length);
 
-	if (mode != (FALLOC_FL_KEEP_SIZE | FALLOC_FL_PUNCH_HOLE))
+	if (mode != (FALLOC_FL_KEEP_SIZE | FALLOC_FL_PUNCH_HOLE)) {
+		ceph_blog_exit(&__ji);
 		return -EOPNOTSUPP;
+	}
 
-	if (!S_ISREG(inode->i_mode))
+	if (!S_ISREG(inode->i_mode)) {
+		ceph_blog_exit(&__ji);
 		return -EOPNOTSUPP;
+	}
 
-	if (IS_ENCRYPTED(inode))
+	if (IS_ENCRYPTED(inode)) {
+		ceph_blog_exit(&__ji);
 		return -EOPNOTSUPP;
+	}
 
 	prealloc_cf = ceph_alloc_cap_flush();
-	if (!prealloc_cf)
+	if (!prealloc_cf) {
+		ceph_blog_exit(&__ji);
 		return -ENOMEM;
+	}
 
 	inode_lock(inode);
 
@@ -2867,6 +2954,7 @@ static long ceph_fallocate(struct file *file, int mode,
 unlock:
 	inode_unlock(inode);
 	ceph_free_cap_flush(prealloc_cf);
+	ceph_blog_exit(&__ji);
 	return ret;
 }
 
@@ -2944,7 +3032,7 @@ static int is_file_size_ok(struct inode *src_inode, struct inode *dst_inode,
 	 * inode.
 	 */
 	if (src_off + len > size) {
-		doutc(cl, "Copy beyond EOF (%llu + %zu > %llu)\n", src_off,
+		boutc(cl, "Copy beyond EOF (%llu + %zu > %llu)\n", src_off,
 		      len, size);
 		return -EOPNOTSUPP;
 	}
@@ -3066,7 +3154,7 @@ static ssize_t ceph_do_objects_copy(struct ceph_inode_info *src_ci, u64 *src_off
 				pr_notice_client(cl,
 					"OSDs don't support copy-from2; disabling copy offload\n");
 			}
-			doutc(cl, "returned %d\n", ret);
+			boutc(cl, "returned %d\n", ret);
 			if (bytes <= 0)
 				bytes = ret;
 			goto out;
@@ -3083,6 +3171,19 @@ static ssize_t ceph_do_objects_copy(struct ceph_inode_info *src_ci, u64 *src_off
 	return bytes;
 }
 
+static noinline void ceph_blog_copy_clusters(struct blog_tls_ctx *blog_ctx,
+					     struct ceph_client *src,
+					     struct ceph_client *dst)
+{
+	char src_uuid[40], dst_uuid[40];
+
+	snprintf(src_uuid, sizeof(src_uuid), "%pU", &src->fsid);
+	snprintf(dst_uuid, sizeof(dst_uuid), "%pU", &dst->fsid);
+	CEPH_BLOG_LOG_CLIENT(blog_ctx, src,
+			     "Copying files across clusters: src: %s dst: %s\n",
+			     src_uuid, dst_uuid);
+}
+
 static ssize_t __ceph_copy_file_range(struct file *src_file, loff_t src_off,
 				      struct file *dst_file, loff_t dst_off,
 				      size_t len, unsigned int flags)
@@ -3105,8 +3206,13 @@ static ssize_t __ceph_copy_file_range(struct file *src_file, loff_t src_off,
 
 		if (ceph_fsid_compare(&src_fsc->client->fsid,
 				      &dst_fsc->client->fsid)) {
-			dout("Copying files across clusters: src: %pU dst: %pU\n",
-			     &src_fsc->client->fsid, &dst_fsc->client->fsid);
+			struct blog_tls_ctx *blog_ctx = ceph_blog_get_ctx(src_fsc);
+
+			if (blog_ctx)
+				ceph_blog_copy_clusters(blog_ctx, cl, dst_fsc->client);
+			else
+				doutc(cl, "Copying files across clusters: src: %pU dst: %pU\n",
+				      &cl->fsid, &dst_fsc->client->fsid);
 			return -EXDEV;
 		}
 	}
@@ -3136,7 +3242,7 @@ static ssize_t __ceph_copy_file_range(struct file *src_file, loff_t src_off,
 	    (src_ci->i_layout.stripe_count != 1) ||
 	    (dst_ci->i_layout.stripe_count != 1) ||
 	    (src_ci->i_layout.object_size != dst_ci->i_layout.object_size)) {
-		doutc(cl, "Invalid src/dst files layout\n");
+		boutc(cl, "Invalid src/dst files layout\n");
 		return -EOPNOTSUPP;
 	}
 
@@ -3154,12 +3260,12 @@ static ssize_t __ceph_copy_file_range(struct file *src_file, loff_t src_off,
 	/* Start by sync'ing the source and destination files */
 	ret = file_write_and_wait_range(src_file, src_off, (src_off + len));
 	if (ret < 0) {
-		doutc(cl, "failed to write src file (%zd)\n", ret);
+		boutc(cl, "failed to write src file (%zd)\n", ret);
 		goto out;
 	}
 	ret = file_write_and_wait_range(dst_file, dst_off, (dst_off + len));
 	if (ret < 0) {
-		doutc(cl, "failed to write dst file (%zd)\n", ret);
+		boutc(cl, "failed to write dst file (%zd)\n", ret);
 		goto out;
 	}
 
@@ -3171,7 +3277,7 @@ static ssize_t __ceph_copy_file_range(struct file *src_file, loff_t src_off,
 	err = get_rd_wr_caps(src_file, &src_got,
 			     dst_file, (dst_off + len), &dst_got);
 	if (err < 0) {
-		doutc(cl, "get_rd_wr_caps returned %d\n", err);
+		boutc(cl, "get_rd_wr_caps returned %d\n", err);
 		ret = -EOPNOTSUPP;
 		goto out;
 	}
@@ -3186,7 +3292,7 @@ static ssize_t __ceph_copy_file_range(struct file *src_file, loff_t src_off,
 					    dst_off >> PAGE_SHIFT,
 					    (dst_off + len) >> PAGE_SHIFT);
 	if (ret < 0) {
-		doutc(cl, "Failed to invalidate inode pages (%zd)\n",
+		boutc(cl, "Failed to invalidate inode pages (%zd)\n",
 			    ret);
 		ret = 0; /* XXX */
 	}
@@ -3208,7 +3314,7 @@ static ssize_t __ceph_copy_file_range(struct file *src_file, loff_t src_off,
 	 * starting at the src_off
 	 */
 	if (src_objoff) {
-		doutc(cl, "Initial partial copy of %u bytes\n", src_objlen);
+		boutc(cl, "Initial partial copy of %u bytes\n", src_objlen);
 
 		/*
 		 * we need to temporarily drop all caps as we'll be calling
@@ -3219,7 +3325,7 @@ static ssize_t __ceph_copy_file_range(struct file *src_file, loff_t src_off,
 					src_objlen);
 		/* Abort on short copies or on error */
 		if (ret < (long)src_objlen) {
-			doutc(cl, "Failed partial copy (%zd)\n", ret);
+			boutc(cl, "Failed partial copy (%zd)\n", ret);
 			goto out;
 		}
 		len -= ret;
@@ -3241,7 +3347,7 @@ static ssize_t __ceph_copy_file_range(struct file *src_file, loff_t src_off,
 			ret = bytes;
 		goto out_caps;
 	}
-	doutc(cl, "Copied %zu bytes out of %zu\n", bytes, len);
+	boutc(cl, "Copied %zu bytes out of %zu\n", bytes, len);
 	len -= bytes;
 	ret += bytes;
 
@@ -3269,13 +3375,13 @@ static ssize_t __ceph_copy_file_range(struct file *src_file, loff_t src_off,
 	 * there were errors in remote object copies (len >= object_size).
 	 */
 	if (len && (len < src_ci->i_layout.object_size)) {
-		doutc(cl, "Final partial copy of %zu bytes\n", len);
+		boutc(cl, "Final partial copy of %zu bytes\n", len);
 		bytes = splice_file_range(src_file, &src_off, dst_file,
 					  &dst_off, len);
 		if (bytes > 0)
 			ret += bytes;
 		else
-			doutc(cl, "Failed partial copy (%zd)\n", bytes);
+			boutc(cl, "Failed partial copy (%zd)\n", bytes);
 	}
 
 out:
@@ -3288,7 +3394,11 @@ static ssize_t ceph_copy_file_range(struct file *src_file, loff_t src_off,
 				    struct file *dst_file, loff_t dst_off,
 				    size_t len, unsigned int flags)
 {
+	struct ceph_fs_client *fsc = ceph_inode_to_fs_client(file_inode(src_file));
 	ssize_t ret;
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(fsc, &__ji);
 
 	ret = __ceph_copy_file_range(src_file, src_off, dst_file, dst_off,
 				     len, flags);
@@ -3296,6 +3406,7 @@ static ssize_t ceph_copy_file_range(struct file *src_file, loff_t src_off,
 	if (ret == -EOPNOTSUPP || ret == -EXDEV)
 		ret = splice_copy_file_range(src_file, src_off, dst_file,
 					     dst_off, len);
+	ceph_blog_exit(&__ji);
 	return ret;
 }
 
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 16+ messages in thread

* [PATCH v7 13/14] ceph: convert capability and snapshot paths to BLOG logging
  2026-09-24 15:30 [PATCH v7 00/14] ceph: add binary logging (BLOG) for CephFS Alex Markuze
                   ` (11 preceding siblings ...)
  2026-09-24 15:30 ` [PATCH v7 12/14] ceph: convert VFS data I/O " Alex Markuze
@ 2026-09-24 15:30 ` Alex Markuze
  2026-09-24 15:30 ` [PATCH v7 14/14] ceph: convert remaining helper " Alex Markuze
  2026-10-01 13:25 ` [PATCH v7 00/14] ceph: add binary logging (BLOG) for CephFS Xiubo Li
  14 siblings, 0 replies; 16+ messages in thread
From: Alex Markuze @ 2026-09-24 15:30 UTC (permalink / raw)
  To: ceph-devel; +Cc: idryomov, xiubo.li

Replace text debug logging with the boutc() wrappers only on MDS/VFS
paths that already bind a BLOG context via ceph_blog_enter()
(ceph_handle_caps and its handle_cap_* helpers,
ceph_handle_snap/flush_snaps, ceph_fsync, ceph_write_inode).

Signed-off-by: Alex Markuze <amarkuze@redhat.com>
---
 fs/ceph/caps.c | 116 ++++++++++++++++++++++++++++++-------------------
 fs/ceph/snap.c |  17 +++++---
 2 files changed, 83 insertions(+), 50 deletions(-)

diff --git a/fs/ceph/caps.c b/fs/ceph/caps.c
index 16e13ba9c9ac..9d1dcbd38d37 100644
--- a/fs/ceph/caps.c
+++ b/fs/ceph/caps.c
@@ -334,6 +334,7 @@ struct ceph_cap *ceph_get_cap(struct ceph_mds_client *mdsc,
 {
 	struct ceph_client *cl = mdsc->fsc->client;
 	struct ceph_cap *cap = NULL;
+	int count, total, use, resv, avail;
 
 	/* temporary, until we do something about cap import/export */
 	if (!ctx) {
@@ -364,9 +365,11 @@ struct ceph_cap *ceph_get_cap(struct ceph_mds_client *mdsc,
 	}
 
 	spin_lock(&mdsc->caps_list_lock);
-	doutc(cl, "ctx=%p (%d) %d = %d used + %d resv + %d avail\n", ctx,
-	      ctx->count, mdsc->caps_total_count, mdsc->caps_use_count,
-	      mdsc->caps_reserve_count, mdsc->caps_avail_count);
+	count = ctx->count;
+	total = mdsc->caps_total_count;
+	use = mdsc->caps_use_count;
+	resv = mdsc->caps_reserve_count;
+	avail = mdsc->caps_avail_count;
 	BUG_ON(!ctx->count);
 	BUG_ON(ctx->count > mdsc->caps_reserve_count);
 	BUG_ON(list_empty(&mdsc->caps_list));
@@ -382,17 +385,21 @@ struct ceph_cap *ceph_get_cap(struct ceph_mds_client *mdsc,
 	BUG_ON(mdsc->caps_total_count != mdsc->caps_use_count +
 	       mdsc->caps_reserve_count + mdsc->caps_avail_count);
 	spin_unlock(&mdsc->caps_list_lock);
+	doutc(cl, "ctx=%p (%d) %d = %d used + %d resv + %d avail\n", ctx,
+	      count, total, use, resv, avail);
 	return cap;
 }
 
 void ceph_put_cap(struct ceph_mds_client *mdsc, struct ceph_cap *cap)
 {
 	struct ceph_client *cl = mdsc->fsc->client;
+	int total, use, resv, avail;
 
 	spin_lock(&mdsc->caps_list_lock);
-	doutc(cl, "%p %d = %d used + %d resv + %d avail\n", cap,
-	      mdsc->caps_total_count, mdsc->caps_use_count,
-	      mdsc->caps_reserve_count, mdsc->caps_avail_count);
+	total = mdsc->caps_total_count;
+	use = mdsc->caps_use_count;
+	resv = mdsc->caps_reserve_count;
+	avail = mdsc->caps_avail_count;
 	mdsc->caps_use_count--;
 	/*
 	 * Keep some preallocated caps around (ceph_min_count), to
@@ -410,6 +417,8 @@ void ceph_put_cap(struct ceph_mds_client *mdsc, struct ceph_cap *cap)
 	BUG_ON(mdsc->caps_total_count != mdsc->caps_use_count +
 	       mdsc->caps_reserve_count + mdsc->caps_avail_count);
 	spin_unlock(&mdsc->caps_list_lock);
+	doutc(cl, "%p %d = %d used + %d resv + %d avail\n", cap,
+	      total, use, resv, avail);
 }
 
 void ceph_reservation_status(struct ceph_fs_client *fsc,
@@ -547,7 +556,7 @@ static void __cap_delay_requeue_front(struct ceph_mds_client *mdsc,
 {
 	struct inode *inode = &ci->netfs.inode;
 
-	doutc(mdsc->fsc->client, "%p %llx.%llx\n", inode, ceph_vinop(inode));
+	boutc(mdsc->fsc->client, "%p %llx.%llx\n", inode, ceph_vinop(inode));
 	spin_lock(&mdsc->cap_delay_lock);
 	set_bit(CEPH_I_FLUSH_BIT, &ci->i_ceph_flags);
 	if (!list_empty(&ci->i_cap_delay_list))
@@ -872,6 +881,7 @@ static void __touch_cap(struct ceph_inode_info *ci, struct ceph_cap *cap)
 	struct ceph_mds_session *s = cap->session;
 	struct ceph_client *cl = s->s_mdsc->fsc->client;
 	static u8 skip_counter;
+	bool iterating;
 
 	if (data_race(++skip_counter))
 		/* skip this call most of the time to reduce lock
@@ -881,15 +891,17 @@ static void __touch_cap(struct ceph_inode_info *ci, struct ceph_cap *cap)
 		return;
 
 	spin_lock(&s->s_cap_lock);
-	if (!s->s_cap_iterator) {
+	iterating = !!s->s_cap_iterator;
+	if (!iterating)
+		list_move_tail(&cap->session_caps, &s->s_caps);
+	spin_unlock(&s->s_cap_lock);
+
+	if (!iterating)
 		doutc(cl, "%p %llx.%llx cap %p mds%d\n", inode,
 		      ceph_vinop(inode), cap, s->s_mds);
-		list_move_tail(&cap->session_caps, &s->s_caps);
-	} else {
+	else
 		doutc(cl, "%p %llx.%llx cap %p mds%d NOP, iterating over caps\n",
 		      inode, ceph_vinop(inode), cap, s->s_mds);
-	}
-	spin_unlock(&s->s_cap_lock);
 }
 
 /*
@@ -1281,7 +1293,7 @@ void ceph_remove_cap(struct ceph_mds_client *mdsc, struct ceph_cap *cap,
 	struct ceph_fs_client *fsc;
 
 	if (ceph_cap_is_removed(cap)) {
-		doutc(mdsc->fsc->client, "inode is NULL\n");
+		doutc(mdsc->fsc->client, "cap already removed\n");
 		return;
 	}
 
@@ -2575,7 +2587,7 @@ static int flush_mdlog_and_wait_inode_unsafe_requests(struct inode *inode)
 		kfree(sessions);
 	}
 
-	doutc(cl, "%p %llx.%llx wait on tid %llu %llu\n", inode,
+	boutc(cl, "%p %llx.%llx wait on tid %llu %llu\n", inode,
 	      ceph_vinop(inode), req1 ? req1->r_tid : 0ULL,
 	      req2 ? req2->r_tid : 0ULL);
 	if (req1) {
@@ -2602,13 +2614,17 @@ static int flush_mdlog_and_wait_inode_unsafe_requests(struct inode *inode)
 int ceph_fsync(struct file *file, loff_t start, loff_t end, int datasync)
 {
 	struct inode *inode = file->f_mapping->host;
+	struct ceph_fs_client *fsc = ceph_inode_to_fs_client(inode);
 	struct ceph_inode_info *ci = ceph_inode(inode);
 	struct ceph_client *cl = ceph_inode_to_client(inode);
 	u64 flush_tid;
 	int ret, err;
 	int dirty;
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(fsc, &__ji);
 
-	doutc(cl, "%p %llx.%llx%s\n", inode, ceph_vinop(inode),
+	boutc(cl, "%p %llx.%llx%s\n", inode, ceph_vinop(inode),
 	      datasync ? " datasync" : "");
 
 	ret = file_write_and_wait_range(file, start, end);
@@ -2620,7 +2636,7 @@ int ceph_fsync(struct file *file, loff_t start, loff_t end, int datasync)
 		goto out;
 
 	dirty = try_flush_caps(inode, &flush_tid);
-	doutc(cl, "dirty caps are %s\n", ceph_cap_string(dirty));
+	boutc(cl, "dirty caps are %s\n", ceph_cap_string(dirty));
 
 	err = flush_mdlog_and_wait_inode_unsafe_requests(inode);
 
@@ -2641,8 +2657,9 @@ int ceph_fsync(struct file *file, loff_t start, loff_t end, int datasync)
 	if (err < 0)
 		ret = err;
 out:
-	doutc(cl, "%p %llx.%llx%s result=%d\n", inode, ceph_vinop(inode),
+	boutc(cl, "%p %llx.%llx%s result=%d\n", inode, ceph_vinop(inode),
 	      datasync ? " datasync" : "", ret);
+	ceph_blog_exit(&__ji);
 	return ret;
 }
 
@@ -2654,19 +2671,25 @@ int ceph_fsync(struct file *file, loff_t start, loff_t end, int datasync)
  */
 int ceph_write_inode(struct inode *inode, struct writeback_control *wbc)
 {
+	struct ceph_fs_client *fsc = ceph_inode_to_fs_client(inode);
 	struct ceph_inode_info *ci = ceph_inode(inode);
 	struct ceph_client *cl = ceph_inode_to_client(inode);
 	u64 flush_tid;
 	int err = 0;
 	int dirty;
 	int wait = (wbc->sync_mode == WB_SYNC_ALL && !wbc->for_sync);
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(fsc, &__ji);
 
-	doutc(cl, "%p %llx.%llx wait=%d\n", inode, ceph_vinop(inode), wait);
+	boutc(cl, "%p %llx.%llx wait=%d\n", inode, ceph_vinop(inode), wait);
 	ceph_fscache_unpin_writeback(inode, wbc);
 	if (wait) {
 		err = ceph_wait_on_async_create(inode);
-		if (err)
+		if (err) {
+			ceph_blog_exit(&__ji);
 			return err;
+		}
 		dirty = try_flush_caps(inode, &flush_tid);
 		if (dirty)
 			err = wait_event_interruptible(ci->i_cap_wq,
@@ -2680,6 +2703,7 @@ int ceph_write_inode(struct inode *inode, struct writeback_control *wbc)
 			__cap_delay_requeue_front(mdsc, ci);
 		spin_unlock(&ci->i_ceph_lock);
 	}
+	ceph_blog_exit(&__ji);
 	return err;
 }
 
@@ -3627,7 +3651,7 @@ static void invalidate_aliases(struct inode *inode)
 	struct ceph_client *cl = ceph_inode_to_client(inode);
 	struct dentry *dn, *prev = NULL;
 
-	doutc(cl, "%p %llx.%llx\n", inode, ceph_vinop(inode));
+	boutc(cl, "%p %llx.%llx\n", inode, ceph_vinop(inode));
 	d_prune_aliases(inode);
 	/*
 	 * For non-directory inode, d_find_alias() only returns
@@ -3738,10 +3762,10 @@ static void handle_cap_grant(struct inode *inode,
 	if (IS_ENCRYPTED(inode) && size)
 		size = extra_info->fscrypt_file_size;
 
-	doutc(cl, "%p %llx.%llx cap %p mds%d seq %d %s\n", inode,
+	boutc(cl, "%p %llx.%llx cap %p mds%d seq %d %s\n", inode,
 	      ceph_vinop(inode), cap, session->s_mds, seq,
 	      ceph_cap_string(newcaps));
-	doutc(cl, " size %llu max_size %llu, i_size %llu\n", size,
+	boutc(cl, " size %llu max_size %llu, i_size %llu\n", size,
 	      max_size, i_size_read(inode));
 
 
@@ -3859,7 +3883,7 @@ static void handle_cap_grant(struct inode *inode,
 		inode->i_uid = make_kuid(&init_user_ns, le32_to_cpu(grant->uid));
 		inode->i_gid = make_kgid(&init_user_ns, le32_to_cpu(grant->gid));
 		ci->i_btime = extra_info->btime;
-		doutc(cl, "%p %llx.%llx mode 0%o uid.gid %d.%d\n", inode,
+		boutc(cl, "%p %llx.%llx mode 0%o uid.gid %d.%d\n", inode,
 		      ceph_vinop(inode), inode->i_mode,
 		      from_kuid(&init_user_ns, inode->i_uid),
 		      from_kgid(&init_user_ns, inode->i_gid));
@@ -3887,7 +3911,7 @@ static void handle_cap_grant(struct inode *inode,
 		u64 version = le64_to_cpu(grant->xattr_version);
 
 		if (version > ci->i_xattrs.version) {
-			doutc(cl, " got new xattrs v%llu on %p %llx.%llx len %d\n",
+			boutc(cl, " got new xattrs v%llu on %p %llx.%llx len %d\n",
 			      version, inode, ceph_vinop(inode), len);
 			if (ci->i_xattrs.blob)
 				ceph_buffer_put(ci->i_xattrs.blob);
@@ -3939,7 +3963,7 @@ static void handle_cap_grant(struct inode *inode,
 
 	if (ci->i_auth_cap == cap && (newcaps & CEPH_CAP_ANY_FILE_WR)) {
 		if (max_size != ci->i_max_size) {
-			doutc(cl, "max_size %lld -> %llu\n", ci->i_max_size,
+			boutc(cl, "max_size %lld -> %llu\n", ci->i_max_size,
 			      max_size);
 			ci->i_max_size = max_size;
 			if (max_size >= ci->i_wanted_max_size) {
@@ -3956,7 +3980,7 @@ static void handle_cap_grant(struct inode *inode,
 	used = ceph_adjust_caps_used_for_lazyio(ci, used, cap->issued,
 						cap->implemented);
 	dirty = __ceph_caps_dirty(ci);
-	doutc(cl, " my wanted = %s, used = %s, dirty %s\n",
+	boutc(cl, " my wanted = %s, used = %s, dirty %s\n",
 	      ceph_cap_string(wanted), ceph_cap_string(used),
 	      ceph_cap_string(dirty));
 
@@ -3979,7 +4003,7 @@ static void handle_cap_grant(struct inode *inode,
 	if (cap->issued & ~newcaps) {
 		int revoking = cap->issued & ~newcaps;
 
-		doutc(cl, "revocation: %s -> %s (revoking %s)\n",
+		boutc(cl, "revocation: %s -> %s (revoking %s)\n",
 		      ceph_cap_string(cap->issued), ceph_cap_string(newcaps),
 		      ceph_cap_string(revoking));
 		/*
@@ -4014,11 +4038,11 @@ static void handle_cap_grant(struct inode *inode,
 		cap->issued = newcaps;
 		cap->implemented |= newcaps;
 	} else if (cap->issued == newcaps) {
-		doutc(cl, "caps unchanged: %s -> %s\n",
+		boutc(cl, "caps unchanged: %s -> %s\n",
 		      ceph_cap_string(cap->issued),
 		      ceph_cap_string(newcaps));
 	} else {
-		doutc(cl, "grant: %s -> %s\n", ceph_cap_string(cap->issued),
+		boutc(cl, "grant: %s -> %s\n", ceph_cap_string(cap->issued),
 		      ceph_cap_string(newcaps));
 		/* non-auth MDS is revoking the newly grant caps ? */
 		if (cap == ci->i_auth_cap &&
@@ -4170,7 +4194,7 @@ static void handle_cap_flush_ack(struct inode *inode, u64 flush_tid,
 		}
 	}
 
-	doutc(cl, "%p %llx.%llx mds%d seq %d on %s cleaned %s, flushing %s -> %s\n",
+	boutc(cl, "%p %llx.%llx mds%d seq %d on %s cleaned %s, flushing %s -> %s\n",
 	      inode, ceph_vinop(inode), session->s_mds, seq,
 	      ceph_cap_string(dirty), ceph_cap_string(cleaned),
 	      ceph_cap_string(ci->i_flushing_caps),
@@ -4194,16 +4218,16 @@ static void handle_cap_flush_ack(struct inode *inode, u64 flush_tid,
 					    &list_first_entry(&session->s_cap_flushing,
 							      struct ceph_inode_info,
 							      i_flushing_item)->netfs.inode;
-				doutc(cl, " mds%d still flushing cap on %p %llx.%llx\n",
+				boutc(cl, " mds%d still flushing cap on %p %llx.%llx\n",
 				      session->s_mds, inode, ceph_vinop(inode));
 			}
 		}
 		mdsc->num_cap_flushing--;
-		doutc(cl, " %p %llx.%llx now !flushing\n", inode,
+		boutc(cl, " %p %llx.%llx now !flushing\n", inode,
 		      ceph_vinop(inode));
 
 		if (ci->i_dirty_caps == 0) {
-			doutc(cl, " %p %llx.%llx now clean\n", inode,
+			boutc(cl, " %p %llx.%llx now clean\n", inode,
 			      ceph_vinop(inode));
 			BUG_ON(!list_empty(&ci->i_dirty_item));
 			drop = true;
@@ -4295,14 +4319,14 @@ static void handle_cap_flushsnap_ack(struct inode *inode, u64 flush_tid,
 	bool wake_ci = false;
 	bool wake_mdsc = false;
 
-	doutc(cl, "%p %llx.%llx ci %p mds%d follows %lld\n", inode,
+	boutc(cl, "%p %llx.%llx ci %p mds%d follows %lld\n", inode,
 	      ceph_vinop(inode), ci, session->s_mds, follows);
 
 	spin_lock(&ci->i_ceph_lock);
 	list_for_each_entry(iter, &ci->i_cap_snaps, ci_item) {
 		if (iter->follows == follows) {
 			if (iter->cap_flush.tid != flush_tid) {
-				doutc(cl, " cap_snap %p follows %lld "
+				boutc(cl, " cap_snap %p follows %lld "
 				      "tid %lld != %lld\n", iter,
 				      follows, flush_tid,
 				      iter->cap_flush.tid);
@@ -4311,7 +4335,7 @@ static void handle_cap_flushsnap_ack(struct inode *inode, u64 flush_tid,
 			capsnap = iter;
 			break;
 		} else {
-			doutc(cl, " skipping cap_snap %p follows %lld\n",
+			boutc(cl, " skipping cap_snap %p follows %lld\n",
 			      iter, iter->follows);
 		}
 	}
@@ -4364,7 +4388,7 @@ static bool handle_cap_trunc(struct inode *inode,
 	if (IS_ENCRYPTED(inode) && size)
 		size = extra_info->fscrypt_file_size;
 
-	doutc(cl, "%p %llx.%llx mds%d seq %d to %lld truncate seq %d\n",
+	boutc(cl, "%p %llx.%llx mds%d seq %d to %lld truncate seq %d\n",
 	      inode, ceph_vinop(inode), mds, seq, truncate_size, truncate_seq);
 	queue_trunc = ceph_fill_file_size(inode, issued,
 					  truncate_seq, truncate_size, size);
@@ -4403,7 +4427,7 @@ static void handle_cap_export(struct inode *inode, struct ceph_mds_caps *ex,
 		target = -1;
 	}
 
-	doutc(cl, " cap %llx.%llx export to peer %d piseq %u pmseq %u\n",
+	boutc(cl, " cap %llx.%llx export to peer %d piseq %u pmseq %u\n",
 	      ceph_vinop(inode), target, t_issue_seq, t_mseq);
 retry:
 	down_read(&mdsc->snap_rwsem);
@@ -4438,7 +4462,7 @@ static void handle_cap_export(struct inode *inode, struct ceph_mds_caps *ex,
 		/* already have caps from the target */
 		if (tcap->cap_id == t_cap_id &&
 		    ceph_seq_cmp(tcap->seq, t_issue_seq) < 0) {
-			doutc(cl, " updating import cap %p mds%d\n", tcap,
+			boutc(cl, " updating import cap %p mds%d\n", tcap,
 			      target);
 			tcap->cap_id = t_cap_id;
 			tcap->seq = t_issue_seq - 1;
@@ -4545,7 +4569,7 @@ static void handle_cap_import(struct ceph_mds_client *mdsc,
 		peer = -1;
 	}
 
-	doutc(cl, " cap %llx.%llx import from peer %d piseq %u pmseq %u\n",
+	boutc(cl, " cap %llx.%llx import from peer %d piseq %u pmseq %u\n",
 	      ceph_vinop(inode), peer, piseq, pmseq);
 retry:
 	cap = __get_cap_for_mds(ci, mds);
@@ -4572,7 +4596,7 @@ static void handle_cap_import(struct ceph_mds_client *mdsc,
 
 	ocap = peer >= 0 ? __get_cap_for_mds(ci, peer) : NULL;
 	if (ocap && ocap->cap_id == p_cap_id) {
-		doutc(cl, " remove export cap %p mds%d flags %d\n",
+		boutc(cl, " remove export cap %p mds%d flags %d\n",
 		      ocap, peer, ph->flags);
 		if ((ph->flags & CEPH_CAP_FLAG_AUTH) &&
 		    (ocap->seq != piseq ||
@@ -4665,10 +4689,13 @@ void ceph_handle_caps(struct ceph_mds_session *session,
 	bool queue_trunc;
 	bool close_sessions = false;
 	bool do_cap_release = false;
+	struct ceph_journal_info __ji;
 
 	if (!ceph_inc_mds_stopping_blocker(mdsc, session))
 		return;
 
+	ceph_blog_enter(mdsc->fsc, &__ji);
+
 	/* decode */
 	end = msg->front.iov_base + msg->front.iov_len;
 	if (msg->front.iov_len < sizeof(*h))
@@ -4765,7 +4792,7 @@ void ceph_handle_caps(struct ceph_mds_session *session,
 
 	/* lookup ino */
 	inode = ceph_find_inode(mdsc->fsc->sb, vino);
-	doutc(cl, " caps mds%d op %s ino %llx.%llx inode %p seq %u iseq %u mseq %u\n",
+	boutc(cl, " caps mds%d op %s ino %llx.%llx inode %p seq %u iseq %u mseq %u\n",
 	      session->s_mds, ceph_cap_op_name(op), vino.ino, vino.snap, inode,
 	      seq, issue_seq, mseq);
 
@@ -4775,7 +4802,7 @@ void ceph_handle_caps(struct ceph_mds_session *session,
 	mutex_lock(&session->s_mutex);
 
 	if (!inode) {
-		doutc(cl, " i don't have ino %llx\n", vino.ino);
+		boutc(cl, " i don't have ino %llx\n", vino.ino);
 
 		switch (op) {
 		case CEPH_CAP_OP_IMPORT:
@@ -4830,7 +4857,7 @@ void ceph_handle_caps(struct ceph_mds_session *session,
 	spin_lock(&ci->i_ceph_lock);
 	cap = __get_cap_for_mds(ceph_inode(inode), session->s_mds);
 	if (!cap) {
-		doutc(cl, " no cap on %p ino %llx.%llx from mds%d\n",
+		boutc(cl, " no cap on %p ino %llx.%llx from mds%d\n",
 		      inode, ceph_ino(inode), ceph_snap(inode),
 		      session->s_mds);
 		spin_unlock(&ci->i_ceph_lock);
@@ -4888,6 +4915,7 @@ void ceph_handle_caps(struct ceph_mds_session *session,
 		ceph_mdsc_close_sessions(mdsc);
 
 	kfree(extra_info.fscrypt_auth);
+	ceph_blog_exit(&__ji);
 	return;
 
 flush_cap_releases:
diff --git a/fs/ceph/snap.c b/fs/ceph/snap.c
index c1a0c645567d..f59b83a97d52 100644
--- a/fs/ceph/snap.c
+++ b/fs/ceph/snap.c
@@ -949,7 +949,7 @@ static void flush_snaps(struct ceph_mds_client *mdsc)
 	struct inode *inode;
 	struct ceph_mds_session *session = NULL;
 
-	doutc(cl, "begin\n");
+	boutc(cl, "begin\n");
 	spin_lock(&mdsc->snap_flush_lock);
 	while (!list_empty(&mdsc->snap_flush_list)) {
 		ci = list_first_entry(&mdsc->snap_flush_list,
@@ -964,7 +964,7 @@ static void flush_snaps(struct ceph_mds_client *mdsc)
 	spin_unlock(&mdsc->snap_flush_lock);
 
 	ceph_put_mds_session(session);
-	doutc(cl, "done\n");
+	boutc(cl, "done\n");
 }
 
 /**
@@ -1035,10 +1035,13 @@ void ceph_handle_snap(struct ceph_mds_client *mdsc,
 	int i;
 	int locked_rwsem = 0;
 	bool close_sessions = false;
+	struct ceph_journal_info __ji;
 
 	if (!ceph_inc_mds_stopping_blocker(mdsc, session))
 		return;
 
+	ceph_blog_enter(mdsc->fsc, &__ji);
+
 	/* decode */
 	if (msg->front.iov_len < sizeof(*h))
 		goto bad;
@@ -1051,7 +1054,7 @@ void ceph_handle_snap(struct ceph_mds_client *mdsc,
 	trace_len = le32_to_cpu(h->trace_len);
 	p += sizeof(*h);
 
-	doutc(cl, "from mds%d op %s split %llx tracelen %d\n", mds,
+	boutc(cl, "from mds%d op %s split %llx tracelen %d\n", mds,
 	      ceph_snap_op_name(op), split, trace_len);
 
 	down_write(&mdsc->snap_rwsem);
@@ -1083,7 +1086,7 @@ void ceph_handle_snap(struct ceph_mds_client *mdsc,
 				goto out;
 		}
 
-		doutc(cl, "splitting snap_realm %llx %p\n", realm->ino, realm);
+		boutc(cl, "splitting snap_realm %llx %p\n", realm->ino, realm);
 		for (i = 0; i < num_split_inos; i++) {
 			struct ceph_vino vino = {
 				.ino = le64_to_cpu(split_inos[i]),
@@ -1108,12 +1111,12 @@ void ceph_handle_snap(struct ceph_mds_client *mdsc,
 			 */
 			if (ci->i_snap_realm->created >
 			    le64_to_cpu(ri->created)) {
-				doutc(cl, " leaving %p %llx.%llx in newer realm %llx %p\n",
+				boutc(cl, " leaving %p %llx.%llx in newer realm %llx %p\n",
 				      inode, ceph_vinop(inode), ci->i_snap_realm->ino,
 				      ci->i_snap_realm);
 				goto skip_inode;
 			}
-			doutc(cl, " will move %p %llx.%llx to split realm %llx %p\n",
+			boutc(cl, " will move %p %llx.%llx to split realm %llx %p\n",
 			      inode, ceph_vinop(inode), realm->ino, realm);
 
 			ceph_get_snap_realm(mdsc, realm);
@@ -1172,6 +1175,7 @@ void ceph_handle_snap(struct ceph_mds_client *mdsc,
 
 	flush_snaps(mdsc);
 	ceph_dec_mds_stopping_blocker(mdsc);
+	ceph_blog_exit(&__ji);
 	return;
 
 bad:
@@ -1185,6 +1189,7 @@ void ceph_handle_snap(struct ceph_mds_client *mdsc,
 
 	if (close_sessions)
 		ceph_mdsc_close_sessions(mdsc);
+	ceph_blog_exit(&__ji);
 	return;
 }
 
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 16+ messages in thread

* [PATCH v7 14/14] ceph: convert remaining helper paths to BLOG logging
  2026-09-24 15:30 [PATCH v7 00/14] ceph: add binary logging (BLOG) for CephFS Alex Markuze
                   ` (12 preceding siblings ...)
  2026-09-24 15:30 ` [PATCH v7 13/14] ceph: convert capability and snapshot " Alex Markuze
@ 2026-09-24 15:30 ` Alex Markuze
  2026-10-01 13:25 ` [PATCH v7 00/14] ceph: add binary logging (BLOG) for CephFS Xiubo Li
  14 siblings, 0 replies; 16+ messages in thread
From: Alex Markuze @ 2026-09-24 15:30 UTC (permalink / raw)
  To: ceph-devel; +Cc: idryomov, xiubo.li

Replace text debug logging with the boutc() wrappers in locks.c,
xattr.c, and crypto.c.

Do not log borrowed xattr names while destroying an invalidated
index. Snapshot encoding can replace and free the backing blob
before the index is destroyed, making even a bounded name read
unsafe. Log only the live index-node pointers at that site.

When copying xattr names, log the terminated destination and the
xattr pointer. Names borrowed from the blob are length-delimited
and must not be passed to an unbounded %s. Both text logging bugs
predate BLOG; eager binary argument capture also exposes them.

Keep encrypted-name logging within the supplied name bounds too.
On a long snapshot-name lookup failure, log the existing terminated
copy with its leading underscore restored, rather than the counted
input. Base64 encoding does not append a terminator: use BLOG_STR
with the encoded length, since format precision only bounds the
text path and does not constrain binary argument capture.

Reported-by: Xiubo Li <xiubo.li@clyso.com>
Link: https://lore.kernel.org/ceph-devel/CAOJNxRJTiUkSAW6diKfZA51mGMCeGdKsbcb5mdkYxqf+-+e8fQ@mail.gmail.com/
Signed-off-by: Alex Markuze <amarkuze@redhat.com>
Assisted-by: LLM
---
 fs/ceph/crypto.c |  19 ++++-----
 fs/ceph/locks.c  |  66 ++++++++++++++++++++++---------
 fs/ceph/xattr.c  | 101 +++++++++++++++++++++++++++++++----------------
 3 files changed, 124 insertions(+), 62 deletions(-)

diff --git a/fs/ceph/crypto.c b/fs/ceph/crypto.c
index a7883336e6a1..de349397fc47 100644
--- a/fs/ceph/crypto.c
+++ b/fs/ceph/crypto.c
@@ -181,7 +181,7 @@ static struct inode *parse_longname(const struct inode *parent,
 		return ERR_PTR(-ENOMEM);
 	name_end = strrchr(str, '_');
 	if (!name_end) {
-		doutc(cl, "failed to parse long snapshot name: %s\n", str);
+		boutc(cl, "failed to parse long snapshot name: %s\n", str);
 		return ERR_PTR(-EIO);
 	}
 	*name_len = (name_end - str);
@@ -194,7 +194,7 @@ static struct inode *parse_longname(const struct inode *parent,
 	inode_number = name_end + 1;
 	ret = kstrtou64(inode_number, 10, &vino.ino);
 	if (ret) {
-		doutc(cl, "failed to parse inode number: %s\n", str);
+		boutc(cl, "failed to parse inode number: %s\n", str);
 		return ERR_PTR(ret);
 	}
 
@@ -204,7 +204,7 @@ static struct inode *parse_longname(const struct inode *parent,
 		/* This can happen if we're not mounting cephfs on the root */
 		dir = ceph_get_inode(parent->i_sb, vino, NULL);
 		if (IS_ERR(dir))
-			doutc(cl, "can't find inode %s (%s)\n", inode_number, name);
+			boutc(cl, "can't find inode %s (_%s)\n", inode_number, str);
 	}
 	return dir;
 }
@@ -275,7 +275,8 @@ int ceph_encode_encrypted_dname(struct inode *parent, char *buf, int elen)
 
 	/* base64 encode the encrypted name */
 	elen = base64_encode(cryptbuf, len, p, false, BASE64_IMAP);
-	doutc(cl, "base64-encoded ciphertext name = %.*s\n", elen, p);
+	boutc_bounded(cl, "base64-encoded ciphertext name = %.*s\n",
+		      (elen, BLOG_STR(p, elen)), (elen, (const char *)p));
 
 	/* To understand the 240 limit, see CEPH_NOHASH_NAME_MAX comments */
 	WARN_ON(elen > 240);
@@ -491,7 +492,7 @@ int ceph_fscrypt_decrypt_block_inplace(const struct inode *inode,
 {
 	struct ceph_client *cl = ceph_inode_to_client(inode);
 
-	doutc(cl, "%p %llx.%llx len %u offs %u blk %llu\n", inode,
+	boutc(cl, "%p %llx.%llx len %u offs %u blk %llu\n", inode,
 	      ceph_vinop(inode), len, offs, lblk_num);
 	return fscrypt_decrypt_block_inplace(inode, page, len, offs, lblk_num);
 }
@@ -502,7 +503,7 @@ int ceph_fscrypt_encrypt_block_inplace(const struct inode *inode,
 {
 	struct ceph_client *cl = ceph_inode_to_client(inode);
 
-	doutc(cl, "%p %llx.%llx len %u offs %u blk %llu\n", inode,
+	boutc(cl, "%p %llx.%llx len %u offs %u blk %llu\n", inode,
 	      ceph_vinop(inode), len, offs, lblk_num);
 	return fscrypt_encrypt_block_inplace(inode, page, len, offs, lblk_num);
 }
@@ -579,7 +580,7 @@ int ceph_fscrypt_decrypt_extents(struct inode *inode, struct page **page,
 
 	/* Nothing to do for empty array */
 	if (ext_cnt == 0) {
-		doutc(cl, "%p %llx.%llx empty array, ret 0\n", inode,
+		boutc(cl, "%p %llx.%llx empty array, ret 0\n", inode,
 		      ceph_vinop(inode));
 		return 0;
 	}
@@ -603,7 +604,7 @@ int ceph_fscrypt_decrypt_extents(struct inode *inode, struct page **page,
 		}
 		fret = ceph_fscrypt_decrypt_pages(inode, &page[pgidx],
 						 off + pgsoff, ext->len);
-		doutc(cl, "%p %llx.%llx [%d] 0x%llx~0x%llx fret %d\n", inode,
+		boutc(cl, "%p %llx.%llx [%d] 0x%llx~0x%llx fret %d\n", inode,
 		      ceph_vinop(inode), i, ext->off, ext->len, fret);
 		if (fret < 0) {
 			if (ret == 0)
@@ -612,7 +613,7 @@ int ceph_fscrypt_decrypt_extents(struct inode *inode, struct page **page,
 		}
 		ret = pgsoff + fret;
 	}
-	doutc(cl, "ret %d\n", ret);
+	boutc(cl, "ret %d\n", ret);
 	return ret;
 }
 
diff --git a/fs/ceph/locks.c b/fs/ceph/locks.c
index 677221bd64e0..a0c946e5890e 100644
--- a/fs/ceph/locks.c
+++ b/fs/ceph/locks.c
@@ -110,7 +110,7 @@ static int ceph_lock_message(u8 lock_type, u16 operation, struct inode *inode,
 
 	owner = secure_addr(fl->c.flc_owner);
 
-	doutc(cl, "rule: %d, op: %d, owner: %llx, pid: %llu, "
+	boutc(cl, "rule: %d, op: %d, owner: %llx, pid: %llu, "
 		    "start: %llu, length: %llu, wait: %d, type: %d\n",
 		    (int)lock_type, (int)operation, owner,
 		    (u64) fl->c.flc_pid,
@@ -147,7 +147,7 @@ static int ceph_lock_message(u8 lock_type, u16 operation, struct inode *inode,
 
 	}
 	ceph_mdsc_put_request(req);
-	doutc(cl, "rule: %d, op: %d, pid: %llu, start: %llu, "
+	boutc(cl, "rule: %d, op: %d, pid: %llu, start: %llu, "
 	      "length: %llu, wait: %d, type: %d, err code %d\n",
 	      (int)lock_type, (int)operation, (u64) fl->c.flc_pid,
 	      fl->fl_start, length, wait, fl->c.flc_type, err);
@@ -175,7 +175,7 @@ static int ceph_lock_wait_for_completion(struct ceph_mds_client *mdsc,
 	if (!err)
 		return 0;
 
-	doutc(cl, "request %llu was interrupted\n", req->r_tid);
+	boutc(cl, "request %llu was interrupted\n", req->r_tid);
 
 	mutex_lock(&mdsc->mutex);
 	if (test_bit(CEPH_MDS_R_GOT_RESULT, &req->r_req_flags)) {
@@ -248,6 +248,7 @@ static int try_unlock_file(struct file *file, struct file_lock *fl)
 int ceph_lock(struct file *file, int cmd, struct file_lock *fl)
 {
 	struct inode *inode = file_inode(file);
+	struct ceph_fs_client *fsc = ceph_inode_to_fs_client(inode);
 	struct ceph_inode_info *ci = ceph_inode(inode);
 	struct ceph_mds_client *mdsc = ceph_sb_to_mdsc(inode->i_sb);
 	struct ceph_client *cl = ceph_inode_to_client(inode);
@@ -255,14 +256,21 @@ int ceph_lock(struct file *file, int cmd, struct file_lock *fl)
 	u16 op = CEPH_MDS_OP_SETFILELOCK;
 	u8 wait = 0;
 	u8 lock_cmd;
+	struct ceph_journal_info __ji;
 
-	if (!(fl->c.flc_flags & FL_POSIX))
+	ceph_blog_enter(fsc, &__ji);
+
+	if (!(fl->c.flc_flags & FL_POSIX)) {
+		ceph_blog_exit(&__ji);
 		return -ENOLCK;
+	}
 
-	if (ceph_inode_is_shutdown(inode))
+	if (ceph_inode_is_shutdown(inode)) {
+		ceph_blog_exit(&__ji);
 		return -ESTALE;
+	}
 
-	doutc(cl, "fl_owner: %p\n", fl->c.flc_owner);
+	boutc(cl, "fl_owner: %p\n", fl->c.flc_owner);
 
 	/* set wait bit as appropriate, then make command as Ceph expects it*/
 	if (IS_GETLK(cmd))
@@ -273,14 +281,17 @@ int ceph_lock(struct file *file, int cmd, struct file_lock *fl)
 	if (test_bit(CEPH_I_ERROR_FILELOCK_BIT, &ci->i_ceph_flags)) {
 		if (op == CEPH_MDS_OP_SETFILELOCK && lock_is_unlock(fl))
 			posix_lock_file(file, fl, NULL);
+		ceph_blog_exit(&__ji);
 		return -EIO;
 	}
 
 	/* Wait for reset to complete before acquiring new locks */
 	if (op == CEPH_MDS_OP_SETFILELOCK && !lock_is_unlock(fl)) {
 		err = ceph_mdsc_wait_for_reset(mdsc);
-		if (err)
+		if (err) {
+			ceph_blog_exit(&__ji);
 			return err;
+		}
 	}
 
 	if (lock_is_read(fl))
@@ -292,14 +303,16 @@ int ceph_lock(struct file *file, int cmd, struct file_lock *fl)
 
 	if (op == CEPH_MDS_OP_SETFILELOCK && lock_is_unlock(fl)) {
 		err = try_unlock_file(file, fl);
-		if (err <= 0)
+		if (err <= 0) {
+			ceph_blog_exit(&__ji);
 			return err;
+		}
 	}
 
 	err = ceph_lock_message(CEPH_LOCK_FCNTL, op, inode, lock_cmd, wait, fl);
 	if (!err) {
 		if (op == CEPH_MDS_OP_SETFILELOCK && F_UNLCK != fl->c.flc_type) {
-			doutc(cl, "locking locally\n");
+			boutc(cl, "locking locally\n");
 			err = posix_lock_file(file, fl, NULL);
 			if (err) {
 				/* undo! This should only happen if
@@ -307,43 +320,55 @@ int ceph_lock(struct file *file, int cmd, struct file_lock *fl)
 				 * deadlock. */
 				ceph_lock_message(CEPH_LOCK_FCNTL, op, inode,
 						  CEPH_LOCK_UNLOCK, 0, fl);
-				doutc(cl, "got %d on posix_lock_file, undid lock\n",
+				boutc(cl, "got %d on posix_lock_file, undid lock\n",
 				      err);
 			}
 		}
 	}
+	ceph_blog_exit(&__ji);
 	return err;
 }
 
 int ceph_flock(struct file *file, int cmd, struct file_lock *fl)
 {
 	struct inode *inode = file_inode(file);
+	struct ceph_fs_client *fsc = ceph_inode_to_fs_client(inode);
 	struct ceph_inode_info *ci = ceph_inode(inode);
 	struct ceph_mds_client *mdsc = ceph_sb_to_mdsc(inode->i_sb);
 	struct ceph_client *cl = ceph_inode_to_client(inode);
 	int err = 0;
 	u8 wait = 0;
 	u8 lock_cmd;
+	struct ceph_journal_info __ji;
 
-	if (!(fl->c.flc_flags & FL_FLOCK))
+	ceph_blog_enter(fsc, &__ji);
+
+	if (!(fl->c.flc_flags & FL_FLOCK)) {
+		ceph_blog_exit(&__ji);
 		return -ENOLCK;
+	}
 
-	if (ceph_inode_is_shutdown(inode))
+	if (ceph_inode_is_shutdown(inode)) {
+		ceph_blog_exit(&__ji);
 		return -ESTALE;
+	}
 
-	doutc(cl, "fl_file: %p\n", fl->c.flc_file);
+	boutc(cl, "fl_file: %p\n", fl->c.flc_file);
 
 	if (test_bit(CEPH_I_ERROR_FILELOCK_BIT, &ci->i_ceph_flags)) {
 		if (lock_is_unlock(fl))
 			locks_lock_file_wait(file, fl);
+		ceph_blog_exit(&__ji);
 		return -EIO;
 	}
 
 	/* Wait for reset to complete before acquiring new locks */
 	if (!lock_is_unlock(fl)) {
 		err = ceph_mdsc_wait_for_reset(mdsc);
-		if (err)
+		if (err) {
+			ceph_blog_exit(&__ji);
 			return err;
+		}
 	}
 
 	if (IS_SETLKW(cmd))
@@ -358,8 +383,10 @@ int ceph_flock(struct file *file, int cmd, struct file_lock *fl)
 
 	if (lock_is_unlock(fl)) {
 		err = try_unlock_file(file, fl);
-		if (err <= 0)
+		if (err <= 0) {
+			ceph_blog_exit(&__ji);
 			return err;
+		}
 	}
 
 	err = ceph_lock_message(CEPH_LOCK_FLOCK, CEPH_MDS_OP_SETFILELOCK,
@@ -370,10 +397,11 @@ int ceph_flock(struct file *file, int cmd, struct file_lock *fl)
 			ceph_lock_message(CEPH_LOCK_FLOCK,
 					  CEPH_MDS_OP_SETFILELOCK,
 					  inode, CEPH_LOCK_UNLOCK, 0, fl);
-			doutc(cl, "got %d on locks_lock_file_wait, undid lock\n",
+			boutc(cl, "got %d on locks_lock_file_wait, undid lock\n",
 			      err);
 		}
 	}
+	ceph_blog_exit(&__ji);
 	return err;
 }
 
@@ -399,7 +427,7 @@ void ceph_count_locks(struct inode *inode, int *fcntl_count, int *flock_count)
 			++(*flock_count);
 		spin_unlock(&ctx->flc_lock);
 	}
-	doutc(cl, "counted %d flock locks and %d fcntl locks\n",
+	boutc(cl, "counted %d flock locks and %d fcntl locks\n",
 	      *flock_count, *fcntl_count);
 }
 
@@ -430,7 +458,7 @@ static int lock_to_ceph_filelock(struct inode *inode,
 		cephlock->type = CEPH_LOCK_UNLOCK;
 		break;
 	default:
-		doutc(cl, "Have unknown lock type %d\n",
+		boutc(cl, "Have unknown lock type %d\n",
 		      lock->c.flc_type);
 		err = -EINVAL;
 	}
@@ -455,7 +483,7 @@ int ceph_encode_locks_to_buffer(struct inode *inode,
 	int seen_flock = 0;
 	int l = 0;
 
-	doutc(cl, "encoding %d flock and %d fcntl locks\n", num_flock_locks,
+	boutc(cl, "encoding %d flock and %d fcntl locks\n", num_flock_locks,
 	      num_fcntl_locks);
 
 	if (!ctx)
diff --git a/fs/ceph/xattr.c b/fs/ceph/xattr.c
index 1da6deea2c24..39dbbcc287c0 100644
--- a/fs/ceph/xattr.c
+++ b/fs/ceph/xattr.c
@@ -70,7 +70,7 @@ static ssize_t ceph_vxattrcb_layout(struct ceph_inode_info *ci, char *val,
 
 	pool_ns = ceph_try_get_string(ci->i_layout.pool_ns);
 
-	doutc(cl, "%p\n", &ci->netfs.inode);
+	boutc(cl, "%p\n", &ci->netfs.inode);
 	down_read(&osdc->lock);
 	pool_name = ceph_pg_pool_name_by_id(osdc->osdmap, pool);
 	if (pool_name) {
@@ -627,7 +627,7 @@ static int __set_xattr(struct ceph_inode_info *ci,
 		xattr->should_free_name = update_xattr;
 
 		ci->i_xattrs.count++;
-		doutc(cl, "count=%d\n", ci->i_xattrs.count);
+		boutc(cl, "count=%d\n", ci->i_xattrs.count);
 	} else {
 		kfree(*newxattr);
 		*newxattr = NULL;
@@ -655,13 +655,19 @@ static int __set_xattr(struct ceph_inode_info *ci,
 	if (new) {
 		rb_link_node(&xattr->node, parent, p);
 		rb_insert_color(&xattr->node, &ci->i_xattrs.index);
-		doutc(cl, "p=%p\n", p);
+		boutc(cl, "p=%p\n", p);
 	}
 
-	doutc(cl, "added %p %llx.%llx xattr %p %.*s=%.*s%s\n", inode,
-	      ceph_vinop(inode), xattr, name_len, name, min(val_len,
-	      MAX_XATTR_VAL_PRINT_LEN), val,
-	      val_len > MAX_XATTR_VAL_PRINT_LEN ? "..." : "");
+	boutc_bounded(cl, "added %p %llx.%llx xattr %p %.*s=%.*s%s\n",
+		      (inode, ceph_vinop(inode), xattr, name_len,
+		       BLOG_STR(name, name_len),
+		       min(val_len, MAX_XATTR_VAL_PRINT_LEN),
+		       BLOG_STR(val, min(val_len, MAX_XATTR_VAL_PRINT_LEN)),
+		       val_len > MAX_XATTR_VAL_PRINT_LEN ? "..." : ""),
+		      (inode, ceph_vinop(inode), xattr, name_len,
+		       (const char *)name,
+		       min(val_len, MAX_XATTR_VAL_PRINT_LEN), (const char *)val,
+		       val_len > MAX_XATTR_VAL_PRINT_LEN ? "..." : ""));
 
 	return 0;
 }
@@ -690,13 +696,16 @@ static struct ceph_inode_xattr *__get_xattr(struct ceph_inode_info *ci,
 		else {
 			int len = min(xattr->val_len, MAX_XATTR_VAL_PRINT_LEN);
 
-			doutc(cl, "%s found %.*s%s\n", name, len, xattr->val,
-			      xattr->val_len > len ? "..." : "");
+			boutc_bounded(cl, "%s found %.*s%s\n",
+				      (name, len, BLOG_STR(xattr->val, len),
+				       xattr->val_len > len ? "..." : ""),
+				      (name, len, (const char *)xattr->val,
+				       xattr->val_len > len ? "..." : ""));
 			return xattr;
 		}
 	}
 
-	doutc(cl, "%s not found\n", name);
+	boutc(cl, "%s not found\n", name);
 
 	return NULL;
 }
@@ -742,14 +751,14 @@ static char *__copy_xattr_names(struct ceph_inode_info *ci,
 	struct ceph_inode_xattr *xattr = NULL;
 
 	p = rb_first(&ci->i_xattrs.index);
-	doutc(cl, "count=%d\n", ci->i_xattrs.count);
+	boutc(cl, "count=%d\n", ci->i_xattrs.count);
 
 	while (p) {
 		xattr = rb_entry(p, struct ceph_inode_xattr, node);
 		memcpy(dest, xattr->name, xattr->name_len);
 		dest[xattr->name_len] = '\0';
 
-		doutc(cl, "dest=%s %p (%s) (%d/%d)\n", dest, xattr, xattr->name,
+		boutc(cl, "dest=%s xattr=%p (%d/%d)\n", dest, xattr,
 		      xattr->name_len, ci->i_xattrs.names_size);
 
 		dest += xattr->name_len + 1;
@@ -767,13 +776,14 @@ void __ceph_destroy_xattrs(struct ceph_inode_info *ci)
 
 	p = rb_first(&ci->i_xattrs.index);
 
-	doutc(cl, "p=%p\n", p);
+	boutc(cl, "p=%p\n", p);
 
 	while (p) {
 		xattr = rb_entry(p, struct ceph_inode_xattr, node);
 		tmp = p;
 		p = rb_next(tmp);
-		doutc(cl, "next p=%p (%.*s)\n", p, xattr->name_len, xattr->name);
+		/* Names borrowed from an older blob may already be stale. */
+		boutc(cl, "next p=%p xattr=%p\n", p, xattr);
 		rb_erase(tmp, &ci->i_xattrs.index);
 
 		__free_xattr(xattr);
@@ -802,7 +812,7 @@ static int __build_xattrs(struct inode *inode)
 	int err = 0;
 	int i;
 
-	doutc(cl, "len=%d\n",
+	boutc(cl, "len=%d\n",
 	      ci->i_xattrs.blob ? (int)ci->i_xattrs.blob->vec.iov_len : 0);
 
 	if (ci->i_xattrs.index_version >= ci->i_xattrs.version)
@@ -888,7 +898,7 @@ static int __get_required_blob_size(struct ceph_inode_info *ci, int name_size,
 	int size = 4 + ci->i_xattrs.count*(4 + 4) +
 			     ci->i_xattrs.names_size +
 			     ci->i_xattrs.vals_size;
-	doutc(cl, "c=%d names.size=%d vals.size=%d\n", ci->i_xattrs.count,
+	boutc(cl, "c=%d names.size=%d vals.size=%d\n", ci->i_xattrs.count,
 	      ci->i_xattrs.names_size, ci->i_xattrs.vals_size);
 
 	if (name_size)
@@ -912,7 +922,7 @@ struct ceph_buffer *__ceph_build_xattrs_blob(struct ceph_inode_info *ci)
 	struct ceph_buffer *old_blob = NULL;
 	void *dest;
 
-	doutc(cl, "%p %llx.%llx\n", inode, ceph_vinop(inode));
+	boutc(cl, "%p %llx.%llx\n", inode, ceph_vinop(inode));
 	if (ci->i_xattrs.dirty) {
 		int need = __get_required_blob_size(ci, 0, 0);
 
@@ -951,7 +961,8 @@ struct ceph_buffer *__ceph_build_xattrs_blob(struct ceph_inode_info *ci)
 }
 
 static inline int __get_request_mask(struct inode *in) {
-	struct ceph_mds_request *req = current->journal_info;
+	struct ceph_mds_request *req =
+		ceph_current_mds_request(ceph_sb_to_fs_client(in->i_sb));
 	int mask = 0;
 	if (req && req->r_target_inode == in) {
 		if (req->r_op == CEPH_MDS_OP_LOOKUP ||
@@ -1009,7 +1020,7 @@ ssize_t __ceph_getxattr(struct inode *inode, const char *name, void *value,
 	req_mask = __get_request_mask(inode);
 
 	spin_lock(&ci->i_ceph_lock);
-	doutc(cl, "%p %llx.%llx name '%s' ver=%lld index_ver=%lld\n", inode,
+	boutc(cl, "%p %llx.%llx name '%s' ver=%lld index_ver=%lld\n", inode,
 	      ceph_vinop(inode), name, ci->i_xattrs.version,
 	      ci->i_xattrs.index_version);
 
@@ -1019,7 +1030,7 @@ ssize_t __ceph_getxattr(struct inode *inode, const char *name, void *value,
 		spin_unlock(&ci->i_ceph_lock);
 
 		/* security module gets xattr while filling trace */
-		if (current->journal_info) {
+		if (ceph_current_fill_trace_request()) {
 			pr_warn_ratelimited_client(cl,
 				"sync %p %llx.%llx during filling trace\n",
 				inode, ceph_vinop(inode));
@@ -1052,7 +1063,7 @@ ssize_t __ceph_getxattr(struct inode *inode, const char *name, void *value,
 
 	memcpy(value, xattr->val, xattr->val_len);
 
-	if (current->journal_info &&
+	if (ceph_current_fill_trace_request() &&
 	    !strncmp(name, XATTR_SECURITY_PREFIX, XATTR_SECURITY_PREFIX_LEN) &&
 	    security_ismaclabel(name + XATTR_SECURITY_PREFIX_LEN))
 		set_bit(CEPH_I_SEC_INITED_BIT, &ci->i_ceph_flags);
@@ -1064,14 +1075,18 @@ ssize_t __ceph_getxattr(struct inode *inode, const char *name, void *value,
 ssize_t ceph_listxattr(struct dentry *dentry, char *names, size_t size)
 {
 	struct inode *inode = d_inode(dentry);
+	struct ceph_fs_client *fsc = ceph_sb_to_fs_client(inode->i_sb);
 	struct ceph_client *cl = ceph_inode_to_client(inode);
 	struct ceph_inode_info *ci = ceph_inode(inode);
 	bool len_only = (size == 0);
 	u32 namelen;
 	int err;
+	struct ceph_journal_info __ji;
+
+	ceph_blog_enter(fsc, &__ji);
 
 	spin_lock(&ci->i_ceph_lock);
-	doutc(cl, "%p %llx.%llx ver=%lld index_ver=%lld\n", inode,
+	boutc(cl, "%p %llx.%llx ver=%lld index_ver=%lld\n", inode,
 	      ceph_vinop(inode), ci->i_xattrs.version,
 	      ci->i_xattrs.index_version);
 
@@ -1079,8 +1094,10 @@ ssize_t ceph_listxattr(struct dentry *dentry, char *names, size_t size)
 	    !__ceph_caps_issued_mask_metric(ci, CEPH_CAP_XATTR_SHARED, 1)) {
 		spin_unlock(&ci->i_ceph_lock);
 		err = ceph_do_getattr(inode, CEPH_STAT_CAP_XATTR, true);
-		if (err)
+		if (err) {
+			ceph_blog_exit(&__ji);
 			return err;
+		}
 		spin_lock(&ci->i_ceph_lock);
 	}
 
@@ -1101,6 +1118,7 @@ ssize_t ceph_listxattr(struct dentry *dentry, char *names, size_t size)
 	err = namelen;
 out:
 	spin_unlock(&ci->i_ceph_lock);
+	ceph_blog_exit(&__ji);
 	return err;
 }
 
@@ -1133,7 +1151,7 @@ static int ceph_sync_setxattr(struct inode *inode, const char *name,
 			flags |= CEPH_XATTR_REMOVE;
 	}
 
-	doutc(cl, "name %s value size %zu\n", name, size);
+	boutc(cl, "name %s value size %zu\n", name, size);
 
 	/* do request */
 	req = ceph_mdsc_create_request(mdsc, op, USE_AUTH_MDS);
@@ -1162,10 +1180,10 @@ static int ceph_sync_setxattr(struct inode *inode, const char *name,
 	req->r_num_caps = 1;
 	req->r_inode_drop = CEPH_CAP_XATTR_SHARED;
 
-	doutc(cl, "xattr.ver (before): %lld\n", ci->i_xattrs.version);
+	boutc(cl, "xattr.ver (before): %lld\n", ci->i_xattrs.version);
 	err = ceph_mdsc_do_request(mdsc, NULL, req);
 	ceph_mdsc_put_request(req);
-	doutc(cl, "xattr.ver (after): %lld\n", ci->i_xattrs.version);
+	boutc(cl, "xattr.ver (after): %lld\n", ci->i_xattrs.version);
 
 out:
 	if (pagelist)
@@ -1176,10 +1194,11 @@ static int ceph_sync_setxattr(struct inode *inode, const char *name,
 int __ceph_setxattr(struct inode *inode, const char *name,
 			const void *value, size_t size, int flags)
 {
+	struct ceph_fs_client *fsc = ceph_sb_to_fs_client(inode->i_sb);
 	struct ceph_client *cl = ceph_inode_to_client(inode);
 	struct ceph_vxattr *vxattr;
 	struct ceph_inode_info *ci = ceph_inode(inode);
-	struct ceph_mds_client *mdsc = ceph_sb_to_fs_client(inode->i_sb)->mdsc;
+	struct ceph_mds_client *mdsc = fsc->mdsc;
 	struct ceph_cap_flush *prealloc_cf = NULL;
 	struct ceph_buffer *old_blob = NULL;
 	int issued;
@@ -1235,7 +1254,7 @@ int __ceph_setxattr(struct inode *inode, const char *name,
 	required_blob_size = __get_required_blob_size(ci, name_len, val_len);
 	if ((ci->i_xattrs.version == 0) || !(issued & CEPH_CAP_XATTR_EXCL) ||
 	    (required_blob_size > mdsc->mdsmap->m_max_xattr_size)) {
-		doutc(cl, "sync version: %llu size: %d max: %llu\n",
+		boutc(cl, "sync version: %llu size: %d max: %llu\n",
 		      ci->i_xattrs.version, required_blob_size,
 		      mdsc->mdsmap->m_max_xattr_size);
 		goto do_sync;
@@ -1251,7 +1270,7 @@ int __ceph_setxattr(struct inode *inode, const char *name,
 		}
 	}
 
-	doutc(cl, "%p %llx.%llx name '%s' issued %s\n", inode,
+	boutc(cl, "%p %llx.%llx name '%s' issued %s\n", inode,
 	      ceph_vinop(inode), name, ceph_cap_string(issued));
 	__build_xattrs(inode);
 
@@ -1266,7 +1285,7 @@ int __ceph_setxattr(struct inode *inode, const char *name,
 	 */
 	required_blob_size = __get_required_blob_size(ci, name_len, val_len);
 	if (required_blob_size > mdsc->mdsmap->m_max_xattr_size) {
-		doutc(cl, "sync (size too large): %d > %llu\n",
+		boutc(cl, "sync (size too large): %d > %llu\n",
 		      required_blob_size, mdsc->mdsmap->m_max_xattr_size);
 		goto do_sync;
 	}
@@ -1277,7 +1296,7 @@ int __ceph_setxattr(struct inode *inode, const char *name,
 
 		spin_unlock(&ci->i_ceph_lock);
 		ceph_buffer_put(old_blob); /* Shouldn't be required */
-		doutc(cl, " pre-allocating new blob size=%d\n",
+		boutc(cl, " pre-allocating new blob size=%d\n",
 		      required_blob_size);
 		blob = ceph_buffer_new(required_blob_size, GFP_NOFS);
 		if (!blob)
@@ -1317,7 +1336,7 @@ int __ceph_setxattr(struct inode *inode, const char *name,
 		up_read(&mdsc->snap_rwsem);
 
 	/* security module set xattr while filling trace */
-	if (current->journal_info) {
+	if (ceph_current_fill_trace_request()) {
 		pr_warn_ratelimited_client(cl,
 				"sync %p %llx.%llx during filling trace\n",
 				inode, ceph_vinop(inode));
@@ -1346,9 +1365,16 @@ static int ceph_get_xattr_handler(const struct xattr_handler *handler,
 				  struct dentry *dentry, struct inode *inode,
 				  const char *name, void *value, size_t size)
 {
+	struct ceph_fs_client *fsc = ceph_sb_to_fs_client(inode->i_sb);
+	struct ceph_journal_info __ji;
+	int ret;
+
 	if (!ceph_is_valid_xattr(name))
 		return -EOPNOTSUPP;
-	return __ceph_getxattr(inode, name, value, size);
+	ceph_blog_enter(fsc, &__ji);
+	ret = __ceph_getxattr(inode, name, value, size);
+	ceph_blog_exit(&__ji);
+	return ret;
 }
 
 static int ceph_set_xattr_handler(const struct xattr_handler *handler,
@@ -1357,9 +1383,16 @@ static int ceph_set_xattr_handler(const struct xattr_handler *handler,
 				  const char *name, const void *value,
 				  size_t size, int flags)
 {
+	struct ceph_fs_client *fsc = ceph_sb_to_fs_client(inode->i_sb);
+	struct ceph_journal_info __ji;
+	int ret;
+
 	if (!ceph_is_valid_xattr(name))
 		return -EOPNOTSUPP;
-	return __ceph_setxattr(inode, name, value, size, flags);
+	ceph_blog_enter(fsc, &__ji);
+	ret = __ceph_setxattr(inode, name, value, size, flags);
+	ceph_blog_exit(&__ji);
+	return ret;
 }
 
 static const struct xattr_handler ceph_other_xattr_handler = {
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 16+ messages in thread

* Re: [PATCH v7 00/14] ceph: add binary logging (BLOG) for CephFS
  2026-09-24 15:30 [PATCH v7 00/14] ceph: add binary logging (BLOG) for CephFS Alex Markuze
                   ` (13 preceding siblings ...)
  2026-09-24 15:30 ` [PATCH v7 14/14] ceph: convert remaining helper " Alex Markuze
@ 2026-10-01 13:25 ` Xiubo Li
  14 siblings, 0 replies; 16+ messages in thread
From: Xiubo Li @ 2026-10-01 13:25 UTC (permalink / raw)
  To: Alex Markuze; +Cc: ceph-devel, idryomov

Thanks Alex.

This whole series LGTM and tested it. Just feel free to add:

Tested-by: Xiubo Li <xiubo.li@clyso.com>
Reviewed-by: Xiubo Li <xiubo.li@clyso.com>

On Thu, 24 Sept 2026 at 17:30, Alex Markuze <amarkuze@redhat.com> wrote:
>
> This series adds a per-mount binary flight recorder for CephFS debug
> messages. It stores typed arguments in per-task page-fragment buffers
> and reconstructs text through debugfs. Runtime enablement is per mount
> and defaults to off; CONFIG_DEBUG_FS remains the build gate.
>
> Thanks to Xiubo for the detailed v6 testing and reproducer. This revision
> addresses the two xattr logging hazards and the empty-magazine growth
> reported in that review. Further code reviews and a 32-bit build also
> found issues in cache lookup, counted-name logging, debugfs controls,
> dump completion, record timestamps and stack usage, fixed here. The
> complete 14-patch series is based on testing tip 6a8449a3f814.
>
> Changes since v6:
>
>   - In __ceph_destroy_xattrs(), log the node pointers without reading
>     xattr->name. Snapshot encoding can free the blob backing the old
>     index before its destruction; a bounded string read is still unsafe.
>   - In __copy_xattr_names(), log the NUL-terminated destination and xattr
>     pointer. The borrowed source name is length-delimited and cannot be
>     passed to an unbounded %s.
>   - Return exhausted allocation magazines directly to log_batch's empty
>     list, where blog_batch_put() can refill them. Do this for both task
>     context allocation and atomic buffer rotation. Detach the per-CPU
>     magazine before publishing it under the destination empty-list lock.
>     Both batches use the same magazine slab cache.
>   - Load the per-CPU cached context once and check that same pointer.
>     A remote migration or retirement can clear the slot between the old
>     check and reload, causing a NULL dereference despite local preemption
>     being disabled.
>   - On a long encrypted snapshot-name lookup failure, log the existing
>     terminated copy, restoring the leading underscore in the format.
>     The original input is length-delimited.
>   - Bound binary capture of the base64-encoded ciphertext name with
>     BLOG_STR(p, elen). base64_encode() does not append a terminator;
>     %.*s alone only bounds the text path.
>   - Make the BLOG read files root-readable (0400), matching the other
>     Ceph debugfs data files. Recorded names and xattr values must not be
>     exposed when the debugfs root is traversable by other users.
>   - Check the mount's enabled flag during context lookup. Disabling one
>     mount must stop subsequent capture even if another enabled mount
>     keeps the global static key active.
>   - Bound an entries dump by a maximum context ID, preserving that bound
>     across seq_file overflow retries. Stop formatting on overflow so
>     sustained rotation cannot keep a reader chasing new contexts.
>   - Store record base times as u64 from get_jiffies_64(), including the
>     delta calculation. Protect published base updates with the pagefrag
>     lock used by readers. This preserves the epoch on 32-bit kernels,
>     whose low jiffies word first wraps about five minutes after boot.
>     Use one clock sample for the overflow check and stored delta, so a
>     tick between them cannot wrap an exactly U32_MAX delta to zero.
>   - Prepare each argument directly in its existing TLS scratch slot.
>     Returning a temporary struct for every argument pushed the GCC i386
>     __ceph_setattr() frame over its 1280-byte build limit. Direct filling
>     brings reported stack usage to 300 bytes without changing evaluation
>     order, string bounds or the recursive arity macros.
>   - Rebase onto testing with the subsequent CephFS fixes already applied.
>     Fold the fixes into patches 01, 04, 06, 07, 10 and 14; retain 14 patches.
>
> Local validation:
>
>   - GCC 11.4 and Clang 18.1 builds of fs/ceph/ceph.o and
>     net/ceph/libceph.o with DEBUG_FS=y and DYNAMIC_DEBUG=y; GCC also
>     builds both objects with DEBUG_FS=n. A GCC i386 build of both objects
>     also passes with the default 1280-byte frame limit.
>   - ASan/UBSan host tests using the actual xattr functions, BLOG argument
>     serializer and kernel string-reading loop reproduce the v6 UAF and
>     over-read. The v7 cases pass with BLOG and dynamic debug off/on.
>   - Host tests using the actual magazine get/put/cleanup and rebalance
>     code check both allocation call sites, sustained churn, reuse with
>     new magazine allocations disabled, and 64,000 cycles on eight
>     pthread workers. v6 exposes the growth; v7 reuses the magazines and
>     releases all remaining objects at cleanup.
>   - Additional ASan/UBSan host checks inject a remote cache clear and use
>     exact-sized, unterminated encrypted names. They reproduce all three
>     additional faults in the earlier v7 draft and pass with these fixes.
>     The decoder passes 200,000 randomized malformed-payload cases with
>     exact-sized input and output allocations under ASan/UBSan.
>   - Before/after reader tests cover continuous context-ID churn and a
>     stable bound across overflow retries. A cache test covers disabling
>     a mount with an active context. Native freestanding i386 checks of
>     the actual timestamp expressions cover crossing the first wrap,
>     creating a context after it, detecting an expired u32 delta, and a
>     clock tick at the exact U32_MAX boundary.
>   - GCC/Clang serializer, request-context and initialization-failure
>     host tests pass. The review-fix diff has no strict checkpatch errors,
>     warnings or checks.
>
> These are composite-object builds and host regression checks. Kernel
> primitives are substituted in the host harnesses; they do not establish
> kernel scheduling, interrupt, KASAN or lockdep behavior. A v7 booted
> kernel and Ceph-cluster rerun remains for Xiubo; KASAN is not available
> in the local setup. Xiubo's reported cluster results were against v6
> with the destructor logging fix.
>
> Known follow-up: existing %ptSp arguments still decode as pointer values
> on the binary path. Timestamp-by-value serialization and further style
> cleanup remain deferred as in v6.
>
> AI assistance was used for the v7 fixes, code review, regression harnesses
> and this cover letter. The changed implementation commits carry
> Assisted-by attribution.
>
> Alex Markuze (14):
>   ceph: add BLOG private headers
>   ceph: add BLOG deserialization support
>   ceph: add BLOG page-fragment allocator
>   ceph: add BLOG magazine batch allocator
>   ceph: add BLOG logger core
>   ceph: add BLOG per-module context management
>   ceph: add Ceph BLOG scaffolding
>   ceph: add boutc wrappers for BLOG
>   ceph: switch MDS request plumbing to struct ceph_journal_info
>   ceph: add BLOG debugfs interface
>   ceph: convert VFS inode and directory paths to BLOG logging
>   ceph: convert VFS data I/O paths to BLOG logging
>   ceph: convert capability and snapshot paths to BLOG logging
>   ceph: convert remaining helper paths to BLOG logging
>
>  fs/ceph/Makefile                |    3 +
>  fs/ceph/addr.c                  |  202 ++++---
>  fs/ceph/blog.h                  |  216 +++++++
>  fs/ceph/blog_batch.c            |  267 ++++++++
>  fs/ceph/blog_batch.h            |   44 ++
>  fs/ceph/blog_client.c           |  644 ++++++++++++++++++++
>  fs/ceph/blog_core.c             |  297 +++++++++
>  fs/ceph/blog_debugfs.c          |  702 +++++++++++++++++++++
>  fs/ceph/blog_des.c              |  335 +++++++++++
>  fs/ceph/blog_des.h              |   16 +
>  fs/ceph/blog_module.c           | 1003 +++++++++++++++++++++++++++++++
>  fs/ceph/blog_module.h           |   42 ++
>  fs/ceph/blog_pagefrag.c         |   62 ++
>  fs/ceph/blog_pagefrag.h         |   27 +
>  fs/ceph/blog_ser.h              |  404 +++++++++++++
>  fs/ceph/caps.c                  |  116 ++--
>  fs/ceph/crypto.c                |   19 +-
>  fs/ceph/debugfs.c               |   11 +-
>  fs/ceph/dir.c                   |  329 +++++++---
>  fs/ceph/export.c                |   83 ++-
>  fs/ceph/file.c                  |  287 ++++++---
>  fs/ceph/inode.c                 |  256 ++++----
>  fs/ceph/locks.c                 |   66 +-
>  fs/ceph/mds_client.c            |  369 +++++++-----
>  fs/ceph/snap.c                  |   17 +-
>  fs/ceph/super.c                 |   61 +-
>  fs/ceph/super.h                 |    7 +
>  fs/ceph/xattr.c                 |  101 ++--
>  include/linux/ceph/ceph_blog.h  |  292 +++++++++
>  include/linux/ceph/ceph_debug.h |   81 ++-
>  include/linux/ceph/libceph.h    |    2 +
>  31 files changed, 5705 insertions(+), 656 deletions(-)
>  create mode 100644 fs/ceph/blog.h
>  create mode 100644 fs/ceph/blog_batch.c
>  create mode 100644 fs/ceph/blog_batch.h
>  create mode 100644 fs/ceph/blog_client.c
>  create mode 100644 fs/ceph/blog_core.c
>  create mode 100644 fs/ceph/blog_debugfs.c
>  create mode 100644 fs/ceph/blog_des.c
>  create mode 100644 fs/ceph/blog_des.h
>  create mode 100644 fs/ceph/blog_module.c
>  create mode 100644 fs/ceph/blog_module.h
>  create mode 100644 fs/ceph/blog_pagefrag.c
>  create mode 100644 fs/ceph/blog_pagefrag.h
>  create mode 100644 fs/ceph/blog_ser.h
>  create mode 100644 include/linux/ceph/ceph_blog.h
>
>
> base-commit: 6a8449a3f81444b6a25adc41cfd2c0ac33f8f96a
> --
> 2.34.1
>

^ permalink raw reply	[flat|nested] 16+ messages in thread

end of thread, other threads:[~2026-10-01 13:25 UTC | newest]

Thread overview: 16+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-09-24 15:30 [PATCH v7 00/14] ceph: add binary logging (BLOG) for CephFS Alex Markuze
2026-09-24 15:30 ` [PATCH v7 01/14] ceph: add BLOG private headers Alex Markuze
2026-09-24 15:30 ` [PATCH v7 02/14] ceph: add BLOG deserialization support Alex Markuze
2026-09-24 15:30 ` [PATCH v7 03/14] ceph: add BLOG page-fragment allocator Alex Markuze
2026-09-24 15:30 ` [PATCH v7 04/14] ceph: add BLOG magazine batch allocator Alex Markuze
2026-09-24 15:30 ` [PATCH v7 05/14] ceph: add BLOG logger core Alex Markuze
2026-09-24 15:30 ` [PATCH v7 06/14] ceph: add BLOG per-module context management Alex Markuze
2026-09-24 15:30 ` [PATCH v7 07/14] ceph: add Ceph BLOG scaffolding Alex Markuze
2026-09-24 15:30 ` [PATCH v7 08/14] ceph: add boutc wrappers for BLOG Alex Markuze
2026-09-24 15:30 ` [PATCH v7 09/14] ceph: switch MDS request plumbing to struct ceph_journal_info Alex Markuze
2026-09-24 15:30 ` [PATCH v7 10/14] ceph: add BLOG debugfs interface Alex Markuze
2026-09-24 15:30 ` [PATCH v7 11/14] ceph: convert VFS inode and directory paths to BLOG logging Alex Markuze
2026-09-24 15:30 ` [PATCH v7 12/14] ceph: convert VFS data I/O " Alex Markuze
2026-09-24 15:30 ` [PATCH v7 13/14] ceph: convert capability and snapshot " Alex Markuze
2026-09-24 15:30 ` [PATCH v7 14/14] ceph: convert remaining helper " Alex Markuze
2026-10-01 13:25 ` [PATCH v7 00/14] ceph: add binary logging (BLOG) for CephFS Xiubo Li

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.