From: kasuta@riseup.net
To: connman@lists.linux.dev
Subject: Re: [BUG UPDATE] Further isolation: Massive Wi-Fi vs Ethernet divergence on WireGuard disconnect
Date: Sat, 15 Aug 2026 14:41:11 +0000 [thread overview]
Message-ID: <19768b042fe429253e266be2b82e222a@riseup.net> (raw)
In-Reply-To: <f86a2479ff98787d74812f053764b5a4@riseup.net>
[BUG UPDATE] Upstream Root Cause: Asynchronous race condition inside
plugins/wifi.c `interface_removed` loop
Hello ConnMan Developers,
We have isolated a race condition in `plugins/wifi.c` causing a File
Descriptor (FD) leak, triggered when virtual interfaces (e.g.,
WireGuard) drop, causing `interface_removed` to skip essential cleanup.
The `if (!wifi || !wifi->device)` check causes an early return,
bypassing `g_supplicant_interface_cancel`, leading to orphaned Netlink
sockets.
Fix: Ensure `g_supplicant_interface_cancel` and socket closures run
unconditionally to prevent leaks.
Best regards,
Doemela
On 2026-08-15 16:02, kasuta@riseup.net wrote:
> [BUG UPDATE] Technical Breakdown: Netlink/RTNL descriptor aggregation
> inside `src/rtnl.c` loops on `wg` tear-down
>
> Hello ConnMan Developers,
>
> To narrow this down to the absolute shortest path for a patch, we
> performed a structural review of how ConnMan interacts with the Wi-Fi
> stack under GLib environment rules during an out-of-band interface drop
> (like a virtual WireGuard tunnel teardown).
>
> Since this resource aggregation leak is strictly limited to active Wi-Fi
> interfaces, the primary tracking error is highly likely localized within
> the cleanup tracking loops of:
> -> plugins/wifi.c
> -> src/rtnl.c (specifically under Wi-Fi/wpa_supplicant event hooks)
>
> The Likely Root Cause: Missing GLib GIOChannel Dereferencing
>
> When a routing switch occurs over wireless architecture, ConnMan
> instantiates netlink event notification routines or local telemetry
> tracking channels. In a GLib-based architecture, these descriptors are
> commonly wrapped using:
> `GIOChannel *channel = g_io_channel_unix_new(fd);`
>
> If an unmanaged interface (like `wg0`) drops abruptly, the state machine
> triggers an early `return` or a conditional error-handling bypass inside
> the event listener loop.
>
> If this early exit path skips the mandatory cleanup sequences:
> 1. `g_io_channel_unref(channel);`
> 2. `close(fd);`
>
> The low-level socket descriptor remains completely orphaned and open in
> the active process table indefinitely. This matches our telemetry
> profile perfectly, where sequential groups of sockets are permanently
> abandoned every time the interface context switches.
>
> Checking `plugins/wifi.c` for any asynchronous socket loop callbacks or
> Netlink listeners that fail to execute an explicit
> `g_io_channel_unref()` during a non-standard interface drop should
> isolate the exact line causing this issue.
>
> Best regards,
> Doemela
>
> On 2026-08-15 15:57, kasuta@riseup.net wrote:
>> Hello ConnMan Developers,
>>
>> To expedite downstream tracking and save triage cycles, we mapped our
>> telemetry logs directly against the target subsystems. Since the
>> resource footprint continues to increase monotonically over Wi-Fi
>> channels despite shutting down user-space proxy routing and connectivity
>> verification triggers (`OnlineCheckMode = none` and `--nodnsproxy`), the
>> leakage vector is verified within the core kernel notification loops.
>>
>> The leak behavior strictly manifests during the parsing sequence for
>> `RTM_DELLINK` / `RTM_NEWLINK` message sweeps when out-of-band virtual
>> device routing domains (such as standard `wg0` kernel interfaces) are
>> recycled beneath active infrastructure.
>>
>> #### The Code Coordinates to Audit
>>
>> We highly recommend reviewing the object reference counts and lifecycle
>> tracking pathways inside `src/rtnl.c`:
>>
>> 1. `rtnl_link_cb` Allocation Boundaries:
>> Inside the primary Netlink route event engine handler callbacks,
>> check whether incoming interface event transitions are cloning reference
>> counts via structural configurations without matching dereferences when
>> device groups change state or vanish.
>>
>> 2. `g_io_channel_unref()` Missing Triggers:
>> If ConnMan triggers localized wireless scanning adjustments or link
>> quality polling sweeps over internal communication pipes following link
>> loss events, verify that the low-level file handles created to poll
>> netlink arrays are properly executing an explicit `close()` or
>> `g_io_channel_unref()` upon message resolution loop conclusions.
>>
>> 3. Sub-Interface Tracking Buffers:
>> Ensure that virtual routing nodes created dynamically under active
>> Wi-Fi backhauls clear their child reference records upon interface link
>> destruction events, instead of abandoning orphaned sockets within the
>> internal lifecycle architecture.
>>
>> Our current management suite hooks into the environment by enforcing an
>> early-stage service-level sandbox (`LimitNOFILE=512`) to act as a
>> system-safe circuit breaker while this deep-seated lifecycle bug remains
>> open upstream.
>>
>> Hopefully, these technical data logs allow you to pin down the exact
>> routing socket release mismatch in the core daemon source arrays.
>>
>> Best regards,
>> Doemela
>>
>> On 2026-08-12 17:23, kasuta@riseup.net wrote:
>>> Hello ConnMan Developers,
>>>
>>> Correction to my previous telemetry update: The file descriptor count
>>> has broken past the initial threshold and is still continuously crawling
>>> upward on Wi-Fi, even with `OnlineCheckMode=none` and `--nodnsproxy`
>>> fully operational.
>>>
>>> Here is the extended tracking timeline showing the active, creeping leak
>>> over time:
>>>
>>> * VPN Connected baseline over Wi-Fi: 18 open files
>>> * VPN Disconnected (Immediate jump): 34 open files
>>> * A few minutes later (Idle sitting): 63 open files and climbing!
>>>
>>> # ls -l /proc/$(pidof connmand)/fd | wc -l
>>> Output: 63
>>>
>>> This completely eliminates the Fritz!Box WPAD or the HTTP online check
>>> loops as the core drivers of the resource leak. The daemon is actively
>>> compounding and abandoning sockets/handles on a low-frequency cycle
>>> purely over the wireless tracking layers.
>>>
>>> Something inside ConnMan's internal Netlink event listeners or
>>> link-state polling routines is generating or duplicating sockets on
>>> every internal cycle following an out-of-band interface teardown, and it
>>> never executes a clean garbage-collection close().
>>>
>>> Best regards,
>>> Doemela
>>>
>>> On 2026-08-12 17:20, kasuta@riseup.net wrote:
>>>> Hello ConnMan Developers,
>>>>
>>>> An important update regarding the live telemetry behavior after
>>>> completely fixing the configuration file casing layout.
>>>>
>>>> Once `OnlineCheckMode=none`, `OnlineCheckIPv4URL=`, and
>>>> `OnlineCheckIPv6URL=` are cleanly parsed by ConnMan, the infinite
>>>> runaway compounding leak is successfully contained. The daemon no longer
>>>> crawls endlessly toward 1024, confirming that the continuous crawl was
>>>> driven by the online probing subsystems reacting to the broken gateway
>>>> path.
>>>>
>>>> However, a structural "rest-leak" still occurs instantly upon the
>>>> WireGuard disconnect event over Wi-Fi. The descriptor count behaves as
>>>> follows:
>>>>
>>>> * VPN Connected baseline over Wi-Fi: 18 open files
>>>> * VPN Disconnected event (Immediate jump): 34 open files
>>>> * Idle sitting (Long-term tracking): Hard lock at 34 open files
>>>>
>>>> The daemon drops the continuous crawling behavior, but it permanently
>>>> holds onto exactly 16 leaked file descriptors/sockets from the dead
>>>> virtual interface. This confirms that while the online validation loop
>>>> causes the aggressive crash, ConnMan's core state engine still fails to
>>>> perform a clean garbage-collection teardown of the primary interface
>>>> handles when an out-of-band virtual tunnel drops over wireless layers.
>>>>
>>>> Best regards,
>>>> Doemela
>>>>
>>>> On 2026-08-12 16:41, kasuta@riseup.net wrote:
>>>>> Hello ConnMan Developers,
>>>>>
>>>>> Following up on my previous message regarding the WireGuard (wg0)
>>>>> disconnect trigger, I have conducted further exhaustive testing and
>>>>> isolated a massive behavioral divergence between Ethernet and Wi-Fi
>>>>> environments.
>>>>>
>>>>> ### New Discovery: Wi-Fi vs. Ethernet Behavior
>>>>> 1. Ethernet (Stable with Mitigations): When cycling the WireGuard tunnel
>>>>> while connected via hardwired Ethernet (eth0), resource tracking remains
>>>>> entirely flat and stable once the `--nodnsproxy` flag is applied.
>>>>>
>>>>> 2. Wi-Fi (Aggressive Persistent Leak): When operating over a wireless
>>>>> connection (wlan0), disconnecting the exact same WireGuard tunnel causes
>>>>> an immediate, massive surge in file descriptors within the Main Daemon
>>>>> (connmand). This leak persistently accumulates on Wi-Fi even with
>>>>> modifications active.
>>>>>
>>>>> Live Telemetry Capture (Wireless Environment):
>>>>> * VPN Connected (Baseline): Main Daemon (connmand) File Descriptors = 19
>>>>> * VPN Disconnected (Single Trigger Event): Main Daemon (connmand) File
>>>>> Descriptors = 91
>>>>>
>>>>> ConnMan instantly leaked 72 file descriptors in a single disconnect
>>>>> action on Wi-Fi. It appears that under wireless layers, ConnMan’s active
>>>>> link/BSSID scanning routines or netlink interface listeners (nl80211)
>>>>> completely fail to execute proper close() system calls when a concurrent
>>>>> virtual routing interface drops out-of-band.
>>>>>
>>>>> ### Deployed Workaround for Embedded Environments (LibreELEC)
>>>>> For users running this in read-only appliance platforms (tested on
>>>>> x86_64 and Raspberry Pi 3/4/5), I have implemented a systemd service
>>>>> override drop-in alongside a custom configuration to prevent total host
>>>>> networking freezes:
>>>>>
>>>>> /storage/.config/connman_main.conf:
>>>>> [General]
>>>>> OnlineCheckMode=none
>>>>>
>>>>> Systemd override block:
>>>>> [Service]
>>>>> ExecStart=
>>>>> ExecStart=/usr/sbin/connmand -n
>>>>> --config=/storage/.config/connman_main.conf --nodnsproxy
>>>>> LimitNOFILE=4096
>>>>> LogRateLimitIntervalSec=0
>>>>>
>>>>> Capping `LimitNOFILE=4096` ensures systemd will cleanly terminate and
>>>>> respawn connmand to reclaim leaked file descriptors before it exhausts
>>>>> the global operating system limit (1024).
>>>>>
>>>>> This strongly narrows down the root cause to missing resource garbage
>>>>> collection inside the wireless tracking state machines or netlink route
>>>>> management loops during out-of-band virtual interface teardowns.
>>>>>
>>>>> Best regards,
>>>>> Doemela
next prev parent reply other threads:[~2026-08-15 14:41 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-12 14:41 [BUG UPDATE] Further isolation: Massive Wi-Fi vs Ethernet divergence on WireGuard disconnect kasuta
2026-08-12 15:20 ` kasuta
2026-08-12 15:23 ` kasuta
2026-08-15 13:57 ` kasuta
2026-08-15 14:02 ` kasuta
2026-08-15 14:41 ` kasuta [this message]
2026-08-15 14:48 ` kasuta
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=19768b042fe429253e266be2b82e222a@riseup.net \
--to=kasuta@riseup.net \
--cc=connman@lists.linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox