* [RFC] module: init-failure path can free a module with live try_module_get() users
@ 2026-08-24 6:13 Mahanta Jambigi
2026-08-28 12:03 ` Petr Pavlu
0 siblings, 1 reply; 2+ messages in thread
From: Mahanta Jambigi @ 2026-08-24 6:13 UTC (permalink / raw)
To: Luis Chamberlain, Petr Pavlu, Daniel Gomez, Sami Tolvanen,
Aaron Tomlin
Cc: linux-modules, linux-kernel, netdev, linux-s390, D. Wythe,
Dust Li, Sidraya Jayagond, Tony Lu, Tony Lu, Wen Gu,
Alexandra Winter, Halil Pasic, Hidayath Khan
Hi Luis, Petr, Daniel, Sami, Aaron,
I'm writing to ask about what looks like a generic module-init failure
lifetime problem in the module loader. I ran into it while working on
the SMC networking module (net/smc/), but after several patch
iterations, it seems the root issue may belong in kernel/module/main.c
rather than in SMC itself. I'd appreciate your guidance on whether this
reading is correct, and if so, what fix direction would be preferred.
THE ISSUE IN do_init_module()
=============================
include/linux/module.h has a long-standing FIXME in module_is_live():
/* FIXME: It'd be nice to isolate modules during init, too, so they
aren't used before they (may) fail. But presently too much code
(IDE & SCSI) require entry into the module during init. */
static inline bool module_is_live(struct module *mod)
{
return mod->state != MODULE_STATE_GOING;
}
Because MODULE_STATE_COMING is not MODULE_STATE_GOING, try_module_get()
can succeed once a module's __init is executing. If __init makes the
module externally reachable partway through and then later fails, the
failure path in do_init_module() appears to do:
fail:
mod->state = MODULE_STATE_GOING;
synchronize_rcu();
module_put(mod);
...
free_module(mod);
synchronize_rcu() waits for RCU readers, but not for threads that
already obtained a module reference via try_module_get() and are still
executing module text.
By contrast, the normal unload path in try_stop_module() refuses to
proceed while the refcount is non-zero.
So the asymmetry seems to be that the normal unload path waits for
references to drain, while the init-failure path does not.
A concrete race would look like:
1. Module __init registers an externally reachable interface.
2. User space enters through that interface and try_module_get()
succeeds while the module is still COMING.
3. A later __init step fails.
4. do_init_module() frees the module.
5. The in-flight caller is still executing module text.
SMC AS A CONCRETE EXAMPLE
=========================
In SMC, simply moving registration later does not appear to eliminate
the window, because there are two separate registration points that can
make the module reachable via socket():
1. sock_register(&smc_sock_family_ops)
After this, socket(AF_SMC, ...) can succeed and reach
try_module_get() via __sock_create().
2. smc_inet_init() -> inet_register_protosw()
After this, socket(AF_INET, SOCK_STREAM, IPPROTO_SMC) can succeed
and again reach try_module_get().
Either registration point can succeed before a later init step fails.
This may not be specific to SMC; other protocol modules that become
reachable during init, such as Bluetooth, may have similar exposure and
appear worth auditing as well.
ON THE FIXME'S IDE/SCSI CONCERN
===============================
The FIXME mentions IDE and SCSI as reasons not to isolate modules
during init.
1. IDE was removed in Linux 5.14, so that half of the concern no
longer applies.
2. SCSI still appears to self-reference during init
(scsi_device_get() -> try_module_get(hostt->module) during
scsi_scan_host()), so a blanket wait-for-refcount-to-drain
approach in the failure path may deadlock there.
Also, strong_try_module_get() already rejects MODULE_STATE_COMING with
-EBUSY, so the infrastructure for refusing callers during init already
exists in some form.
QUESTIONS
=========
First, is my reading of this init-failure refcount/lifetime asymmetry
correct?
If so, would one of the following directions be acceptable?
1. An opt-in mechanism (for example, a module flag) for modules that
are safe to isolate during init and whose init-failure path should
wait for external references to drain.
2. Treating MODULE_STATE_COMING as non-live for normal
try_module_get() users, with some explicit escape hatch for the
remaining subsystems that genuinely need self-entry during init.
Any guidance on the preferred direction would be much appreciated.
Best regards,
Mahanta Jambigi
^ permalink raw reply [flat|nested] 2+ messages in thread* Re: [RFC] module: init-failure path can free a module with live try_module_get() users
2026-08-24 6:13 [RFC] module: init-failure path can free a module with live try_module_get() users Mahanta Jambigi
@ 2026-08-28 12:03 ` Petr Pavlu
0 siblings, 0 replies; 2+ messages in thread
From: Petr Pavlu @ 2026-08-28 12:03 UTC (permalink / raw)
To: Mahanta Jambigi
Cc: Luis Chamberlain, Daniel Gomez, Sami Tolvanen, Aaron Tomlin,
linux-modules, linux-kernel, netdev, linux-s390, D. Wythe,
Dust Li, Sidraya Jayagond, Tony Lu, Wen Gu, Alexandra Winter,
Halil Pasic, Hidayath Khan
On 8/24/26 8:13 AM, Mahanta Jambigi wrote:
> Hi Luis, Petr, Daniel, Sami, Aaron,
>
> I'm writing to ask about what looks like a generic module-init failure
> lifetime problem in the module loader. I ran into it while working on
> the SMC networking module (net/smc/), but after several patch
> iterations, it seems the root issue may belong in kernel/module/main.c
> rather than in SMC itself. I'd appreciate your guidance on whether this
> reading is correct, and if so, what fix direction would be preferred.
>
> THE ISSUE IN do_init_module()
> =============================
>
> include/linux/module.h has a long-standing FIXME in module_is_live():
>
> /* FIXME: It'd be nice to isolate modules during init, too, so they
> aren't used before they (may) fail. But presently too much code
> (IDE & SCSI) require entry into the module during init. */
> static inline bool module_is_live(struct module *mod)
> {
> return mod->state != MODULE_STATE_GOING;
> }
>
> Because MODULE_STATE_COMING is not MODULE_STATE_GOING, try_module_get()
> can succeed once a module's __init is executing. If __init makes the
> module externally reachable partway through and then later fails, the
> failure path in do_init_module() appears to do:
>
> fail:
> mod->state = MODULE_STATE_GOING;
> synchronize_rcu();
> module_put(mod);
> ...
> free_module(mod);
>
> synchronize_rcu() waits for RCU readers, but not for threads that
> already obtained a module reference via try_module_get() and are still
> executing module text.
>
> By contrast, the normal unload path in try_stop_module() refuses to
> proceed while the refcount is non-zero.
>
> So the asymmetry seems to be that the normal unload path waits for
> references to drain, while the init-failure path does not.
>
> A concrete race would look like:
>
> 1. Module __init registers an externally reachable interface.
> 2. User space enters through that interface and try_module_get()
> succeeds while the module is still COMING.
> 3. A later __init step fails.
> 4. do_init_module() frees the module.
> 5. The in-flight caller is still executing module text.
>
> SMC AS A CONCRETE EXAMPLE
> =========================
>
> In SMC, simply moving registration later does not appear to eliminate
> the window, because there are two separate registration points that can
> make the module reachable via socket():
>
> 1. sock_register(&smc_sock_family_ops)
> After this, socket(AF_SMC, ...) can succeed and reach
> try_module_get() via __sock_create().
>
> 2. smc_inet_init() -> inet_register_protosw()
> After this, socket(AF_INET, SOCK_STREAM, IPPROTO_SMC) can succeed
> and again reach try_module_get().
>
> Either registration point can succeed before a later init step fails.
>
> This may not be specific to SMC; other protocol modules that become
> reachable during init, such as Bluetooth, may have similar exposure and
> appear worth auditing as well.
>
> ON THE FIXME'S IDE/SCSI CONCERN
> ===============================
>
> The FIXME mentions IDE and SCSI as reasons not to isolate modules
> during init.
>
> 1. IDE was removed in Linux 5.14, so that half of the concern no
> longer applies.
>
> 2. SCSI still appears to self-reference during init
> (scsi_device_get() -> try_module_get(hostt->module) during
> scsi_scan_host()), so a blanket wait-for-refcount-to-drain
> approach in the failure path may deadlock there.
>
> Also, strong_try_module_get() already rejects MODULE_STATE_COMING with
> -EBUSY, so the infrastructure for refusing callers during init already
> exists in some form.
>
> QUESTIONS
> =========
>
> First, is my reading of this init-failure refcount/lifetime asymmetry
> correct?
Your analysis looks correct to me.
>
> If so, would one of the following directions be acceptable?
>
> 1. An opt-in mechanism (for example, a module flag) for modules that
> are safe to isolate during init and whose init-failure path should
> wait for external references to drain.
In general, it is preferred if the module loader handles all modules in
the same way.
I would say that the module loader should wait for external references
to drain after an init failure for all modules and that it should be the
responsibility of individual modules to ensure that this wait eventually
completes. Excluding some modules would mean that the module loader
could still free them while they are in use by the kernel.
Before such a wait, the module loader should cancel all idempotent
module loads. This is especially important during boot when several
udevd workers may be trying to insert the same module. In that case,
a failed module init function should block only a single udevd task, so
that the system can still boot properly.
>
> 2. Treating MODULE_STATE_COMING as non-live for normal
> try_module_get() users, with some explicit escape hatch for the
> remaining subsystems that genuinely need self-entry during init.
This might make sense, but the kernel currently has over 500 users of
try_module_get() so changing its semantics in this way would need to be
done very carefully.
I'm also not sure there is anything inherently wrong with calling
try_module_get() on a module while it is still executing its init
function. It must mean the module has reached a state in which it
properly initialized functionality to be registered with a specific
subsystem.
For instance, the __sock_create() function mentioned in your email
contains (simplified):
[...]
if (rcu_access_pointer(net_families[family]) == NULL)
request_module("net-pf-%d", family);
rcu_read_lock();
pf = rcu_dereference(net_families[family]);
err = -EAFNOSUPPORT;
if (!pf)
goto out_release;
if (!try_module_get(pf->owner))
goto out_release;
[...]
Consider two tasks that both call socket(XYZ). The first one sees
net_families[family]==NULL and calls request_module(XYZ). The module
then starts loading and its init function calls sock_register(XYZ). At
this point, the second task can see that net_families[family]!=NULL, so
it skips request_module(XYZ) and calls try_module_get(pf->owner). If
that call newly fails because the module is still in
MODULE_STATE_COMING, it breaks the autoload functionality of
__sock_create().
The scheme would need to be changed to call request_module(XYZ) if
either net_families[family]==NULL or !try_module_get(pf->owner).
>
> Any guidance on the preferred direction would be much appreciated.
Waiting for external references to drain after an init failure makes
sense to me. I'm less sure about changing the semantics of
try_module_get(), or whether strong_try_module_get() should instead be
exposed.
--
Thanks,
Petr
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-08-28 12:03 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-24 6:13 [RFC] module: init-failure path can free a module with live try_module_get() users Mahanta Jambigi
2026-08-28 12:03 ` Petr Pavlu
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox