* Re: ath12k_pci errors and loss of connectivity in 6.12.y branch
[not found] <CAPDiVH8gaBH6o_OY-zUWYpDbj5mhiqmofKGb71gLgHOi4vA=Vw@mail.gmail.com>
@ 2025-06-27 5:39 ` Baochen Qiang
2025-06-27 10:24 ` Robin Murphy
0 siblings, 1 reply; 10+ messages in thread
From: Baochen Qiang @ 2025-06-27 5:39 UTC (permalink / raw)
To: Matt Mower, Jeff Johnson, will, joro, Robin Murphy
Cc: linux-wireless, ath12k, 1107521, iommu
[+ IOMMU list]
On 6/27/2025 12:21 AM, Matt Mower wrote:
> Dear maintainer,
>
> I have been experiencing lost network connection with the ath12k_pci driver
> in the linux-6.12.y kernel branch. Often, when the issue occurs, the
> network does not recover until I reboot the computer. A full report of the
> errors I encounter, the symptoms that arise, and several dmesg attachments
> are in https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1107521 . I have
> attached a dmesg from 6.12.34 for convenience. The short summary is:
>
> 1. I started noticing log lines like the following soon after boot when I
> updated from 6.12.22 to 6.12.27. After these events occur, the network goes
> down and often does not come back up.
> ath12k_pci 0000:c2:00.0: AMD-Vi: Event logged [IO_PAGE_FAULT
> domain=0x0010 address=0xfea00000 flags=0x0020]
> 2. I was able to reproduce this issue very rarely in 6.12.12 and 6.12.22.
> The issue always occurs soon after boot in 6.12.27, 6.12.30, 6.12.33, and
> 6.12.34.
> 3. I have not reproduced the issue in 6.15.2 or 6.15.3.
> 4. In some cases, when shutting down the computer, a kernel bug caused my
> computer to hang. I haven't determined whether this is related to the issue
> above or an independent issue. Search the bug report
> for PXL_20250611_140820085.jpg to see a picture of the kernel bug on my
> laptop screen.
> 5. I have tested two firmware versions:
> a. fw_version 0x1108811c fw_build_timestamp 2025-05-17 00:21 fw_build_id
> QC_IMAGE_VERSION_STRING=WLAN.HMT.1.1.c5-00284.1-QCAHMTSWPL_V1.0_V2.0_SILICONZ-3
> b. fw_version 0x100301e1 fw_build_timestamp 2023-12-06 04:05 fw_build_id
> QC_IMAGE_VERSION_STRING=WLAN.HMT.1.0.c5-00481-QCAHMTSWPL_V1.0_V2.0_SILICONZ-3
>
> Thanks,
> Matt
>
I had a quick test with 6.12.27 kernel on both my Intel desktop and AMD RD but didn't hit
the issue. And I am using WLAN.HMT.1.1.c5-00284.1-QCAHMTSWPL_V1.0_V2.0_SILICONZ-3.
As mentioned in the Debian bug report, since reverting ath12k patches does not fix this
issue, maybe it comes from the IOMMU subsystem?
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: ath12k_pci errors and loss of connectivity in 6.12.y branch
2025-06-27 5:39 ` ath12k_pci errors and loss of connectivity in 6.12.y branch Baochen Qiang
@ 2025-06-27 10:24 ` Robin Murphy
2025-06-30 6:36 ` Vasant Hegde
0 siblings, 1 reply; 10+ messages in thread
From: Robin Murphy @ 2025-06-27 10:24 UTC (permalink / raw)
To: Baochen Qiang, Matt Mower, Jeff Johnson, will, joro, Hegde Vasant
Cc: linux-wireless, ath12k, 1107521, iommu
+Vasant
On 2025-06-27 6:39 am, Baochen Qiang wrote:
> [+ IOMMU list]
>
> On 6/27/2025 12:21 AM, Matt Mower wrote:
>> Dear maintainer,
>>
>> I have been experiencing lost network connection with the ath12k_pci driver
>> in the linux-6.12.y kernel branch. Often, when the issue occurs, the
>> network does not recover until I reboot the computer. A full report of the
>> errors I encounter, the symptoms that arise, and several dmesg attachments
>> are in https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1107521 . I have
>> attached a dmesg from 6.12.34 for convenience. The short summary is:
>>
>> 1. I started noticing log lines like the following soon after boot when I
>> updated from 6.12.22 to 6.12.27. After these events occur, the network goes
>> down and often does not come back up.
>> ath12k_pci 0000:c2:00.0: AMD-Vi: Event logged [IO_PAGE_FAULT
>> domain=0x0010 address=0xfea00000 flags=0x0020]
>> 2. I was able to reproduce this issue very rarely in 6.12.12 and 6.12.22.
>> The issue always occurs soon after boot in 6.12.27, 6.12.30, 6.12.33, and
>> 6.12.34.
>> 3. I have not reproduced the issue in 6.15.2 or 6.15.3.
>> 4. In some cases, when shutting down the computer, a kernel bug caused my
>> computer to hang. I haven't determined whether this is related to the issue
>> above or an independent issue. Search the bug report
>> for PXL_20250611_140820085.jpg to see a picture of the kernel bug on my
>> laptop screen.
>> 5. I have tested two firmware versions:
>> a. fw_version 0x1108811c fw_build_timestamp 2025-05-17 00:21 fw_build_id
>> QC_IMAGE_VERSION_STRING=WLAN.HMT.1.1.c5-00284.1-QCAHMTSWPL_V1.0_V2.0_SILICONZ-3
>> b. fw_version 0x100301e1 fw_build_timestamp 2023-12-06 04:05 fw_build_id
>> QC_IMAGE_VERSION_STRING=WLAN.HMT.1.0.c5-00481-QCAHMTSWPL_V1.0_V2.0_SILICONZ-3
>>
>> Thanks,
>> Matt
>>
>
> I had a quick test with 6.12.27 kernel on both my Intel desktop and AMD RD but didn't hit
> the issue. And I am using WLAN.HMT.1.1.c5-00284.1-QCAHMTSWPL_V1.0_V2.0_SILICONZ-3.
>
> As mentioned in the Debian bug report, since reverting ath12k patches does not fix this
> issue, maybe it comes from the IOMMU subsystem?
Faults are usually still indicative of the client driver/subsystem doing
something not quite right - racily performing dma_unmap before the
device has actually finished making accesses; mapping the wrong size
such that the device accesses off the end of the mapping (this can often
run into another valid mapping so not necessarily fault); mapping the
wrong DMA direction such that the device then tries to write to a
read-only page. However I suppose it's not impossible that some fix to
amd-iommu in that period might have changed its behaviour in a way that
exacerbates things - Vasant, does this strike a chord with anything
you're aware of?
A couple more things I'd try on the ath12k side: firstly, boot with
"iommu.strict=1" and see if that makes the faults any more
frequent/reproducible; if a fault is fairly easily reproducible, then
use the DMA API and/or IOMMU API tracepoints to compare the fault
address to prior DMA mapping activity - that can usually reveal the
nature of the bug enough to then know what to go looking for.
I wouldn't put much significance in whatever happens *after* the fault -
presumably the driver is assuming the blocked DMA write has completed,
so then goes on to read some incomplete descriptor as if it were valid,
and thus may fall over in all manner of entertaining ways on bogus data.
Thanks,
Robin.
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: ath12k_pci errors and loss of connectivity in 6.12.y branch
2025-06-27 10:24 ` Robin Murphy
@ 2025-06-30 6:36 ` Vasant Hegde
2025-07-02 5:17 ` Matt Mower
0 siblings, 1 reply; 10+ messages in thread
From: Vasant Hegde @ 2025-06-30 6:36 UTC (permalink / raw)
To: Robin Murphy, Baochen Qiang, Matt Mower, Jeff Johnson, will, joro
Cc: linux-wireless, ath12k, 1107521, iommu
Hi,
On 6/27/2025 3:54 PM, Robin Murphy wrote:
> +Vasant
>
> On 2025-06-27 6:39 am, Baochen Qiang wrote:
>> [+ IOMMU list]
>>
>> On 6/27/2025 12:21 AM, Matt Mower wrote:
>>> Dear maintainer,
>>>
>>> I have been experiencing lost network connection with the ath12k_pci driver
>>> in the linux-6.12.y kernel branch. Often, when the issue occurs, the
>>> network does not recover until I reboot the computer. A full report of the
>>> errors I encounter, the symptoms that arise, and several dmesg attachments
>>> are in https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1107521 . I have
>>> attached a dmesg from 6.12.34 for convenience. The short summary is:
>>>
>>> 1. I started noticing log lines like the following soon after boot when I
>>> updated from 6.12.22 to 6.12.27. After these events occur, the network goes
>>> down and often does not come back up.
>>> ath12k_pci 0000:c2:00.0: AMD-Vi: Event logged [IO_PAGE_FAULT
>>> domain=0x0010 address=0xfea00000 flags=0x0020]
>>> 2. I was able to reproduce this issue very rarely in 6.12.12 and 6.12.22.
>>> The issue always occurs soon after boot in 6.12.27, 6.12.30, 6.12.33, and
>>> 6.12.34.
>>> 3. I have not reproduced the issue in 6.15.2 or 6.15.3.
>>> 4. In some cases, when shutting down the computer, a kernel bug caused my
>>> computer to hang. I haven't determined whether this is related to the issue
>>> above or an independent issue. Search the bug report
>>> for PXL_20250611_140820085.jpg to see a picture of the kernel bug on my
>>> laptop screen.
>>> 5. I have tested two firmware versions:
>>> a. fw_version 0x1108811c fw_build_timestamp 2025-05-17 00:21 fw_build_id
>>> QC_IMAGE_VERSION_STRING=WLAN.HMT.1.1.c5-00284.1-QCAHMTSWPL_V1.0_V2.0_SILICONZ-3
>>> b. fw_version 0x100301e1 fw_build_timestamp 2023-12-06 04:05 fw_build_id
>>> QC_IMAGE_VERSION_STRING=WLAN.HMT.1.0.c5-00481-QCAHMTSWPL_V1.0_V2.0_SILICONZ-3
>>>
>>> Thanks,
>>> Matt
>>>
>>
>> I had a quick test with 6.12.27 kernel on both my Intel desktop and AMD RD but
>> didn't hit
>> the issue. And I am using WLAN.HMT.1.1.c5-00284.1-
>> QCAHMTSWPL_V1.0_V2.0_SILICONZ-3.
>>
>> As mentioned in the Debian bug report, since reverting ath12k patches does not
>> fix this
>> issue, maybe it comes from the IOMMU subsystem?
>
> Faults are usually still indicative of the client driver/subsystem doing
> something not quite right - racily performing dma_unmap before the device has
> actually finished making accesses; mapping the wrong size such that the device
> accesses off the end of the mapping (this can often run into another valid
> mapping so not necessarily fault); mapping the wrong DMA direction such that the
> device then tries to write to a read-only page. However I suppose it's not
> impossible that some fix to amd-iommu in that period might have changed its
> behaviour in a way that exacerbates things - Vasant, does this strike a chord
> with anything you're aware of?
I did look into kernel code and changes between v6.12.9..v6.12.22.. There are
only two changes in AMD iommu driver.
40c731472f41 iommu/amd: Expicitly enable CNTRL.EPHEn bit in resume path
-> This one was needed to fix the suspend/resume issue. This just adjusts
control bit after suspend. Its not touching page table.
6e1e451456e1 iommu/amd: Remove unused amd_iommu_domain_update()
- Code cleanup patch.
Looking into lspci output only `c2:00.0` is placed in group 15 and domain ID
0x10. I believe there is only one device in this domain.
Interpreting IO_PAGE_FAULT flags = 0x20 means It was a write request for the
page that was not present. So at this point I would still suspect on device
driver side than IOMMU side.
>
> A couple more things I'd try on the ath12k side: firstly, boot with
> "iommu.strict=1" and see if that makes the faults any more frequent/
> reproducible; if a fault is fairly easily reproducible, then use the DMA API
> and/or IOMMU API tracepoints to compare the fault address to prior DMA mapping
> activity - that can usually reveal the nature of the bug enough to then know
> what to go looking for.
>
> I wouldn't put much significance in whatever happens *after* the fault -
> presumably the driver is assuming the blocked DMA write has completed, so then
> goes on to read some incomplete descriptor as if it were valid, and thus may
> fall over in all manner of entertaining ways on bogus data.
Thanks Robin. I'd suggest to follow these suggestions.
-Vasant
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: ath12k_pci errors and loss of connectivity in 6.12.y branch
2025-06-30 6:36 ` Vasant Hegde
@ 2025-07-02 5:17 ` Matt Mower
2025-07-02 7:10 ` Baochen Qiang
0 siblings, 1 reply; 10+ messages in thread
From: Matt Mower @ 2025-07-02 5:17 UTC (permalink / raw)
To: Vasant Hegde
Cc: Robin Murphy, Baochen Qiang, Jeff Johnson, will, joro,
linux-wireless, ath12k, 1107521, iommu
> A couple more things I'd try on the ath12k side: firstly, boot with
> "iommu.strict=1" and see if that makes the faults any more
> frequent/reproducible;
The issue is easy enough to reproduce in 6.12.27 onward and I may be
mistaken about the rarity in 6.12.22; I reproduced it relatively
quickly in .22 today, so if this was the primary purpose for setting
iommu.strict=1, then testing with or without strict works. FWIW, I did
test iommu.strict=1 with 6.15.3 and still have not reproduced this
issue there.
> if a fault is fairly easily reproducible, then
> use the DMA API and/or IOMMU API tracepoints to compare the fault
> address to prior DMA mapping activity - that can usually reveal the
> nature of the bug enough to then know what to go looking for.
This is unfamiliar territory for me, so I hope the following is at
least close to what you requested. If not, happy to provide more test
results based on a set of instructions. Here's what I did:
1. Set CONFIG_DMA_API_DEBUG=y
2. Set kernel command line to: iommu.strict=1 log_buf_len=100M
dma_debug_driver=ath12k_pci trace_event=dma:*,iommu:*
3. Booted and waited for page fault, then cat'd
/sys/kernel/tracing/trace to a file.
Additionally, though I'm pretty sure this is irrelevant now, I added
logging after each dma_map_single() in the ath12k driver to print the
function name and resultant address to the kernel log.
Comparing the addresses of several io_page_fault lines in the trace
and in the kernel log, they line up. So, I'm hopeful this is on the
right track.
DMA/IOMMU trace: https://cmphys.com/ath12k/iommu_dma_trace-20250701.log
Kernel log with additional logging:
https://cmphys.com/ath12k/dmesg-6.12.35-20250701.log
Diff showing extra logging added to v6.12.35:
https://cmphys.com/ath12k/ath12k-extra-logging-6.12.35-20250701.diff
Thanks,
Matt
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: ath12k_pci errors and loss of connectivity in 6.12.y branch
2025-07-02 5:17 ` Matt Mower
@ 2025-07-02 7:10 ` Baochen Qiang
2025-07-02 14:53 ` Matt Mower
0 siblings, 1 reply; 10+ messages in thread
From: Baochen Qiang @ 2025-07-02 7:10 UTC (permalink / raw)
To: Matt Mower, Vasant Hegde
Cc: Robin Murphy, Jeff Johnson, will, joro, linux-wireless, ath12k,
1107521, iommu
On 7/2/2025 1:17 PM, Matt Mower wrote:
>> A couple more things I'd try on the ath12k side: firstly, boot with
>> "iommu.strict=1" and see if that makes the faults any more
>> frequent/reproducible;
>
> The issue is easy enough to reproduce in 6.12.27 onward and I may be
> mistaken about the rarity in 6.12.22; I reproduced it relatively
> quickly in .22 today, so if this was the primary purpose for setting
> iommu.strict=1, then testing with or without strict works. FWIW, I did
> test iommu.strict=1 with 6.15.3 and still have not reproduced this
> issue there.
>
>> if a fault is fairly easily reproducible, then
>> use the DMA API and/or IOMMU API tracepoints to compare the fault
>> address to prior DMA mapping activity - that can usually reveal the
>> nature of the bug enough to then know what to go looking for.
>
> This is unfamiliar territory for me, so I hope the following is at
> least close to what you requested. If not, happy to provide more test
> results based on a set of instructions. Here's what I did:
>
> 1. Set CONFIG_DMA_API_DEBUG=y
> 2. Set kernel command line to: iommu.strict=1 log_buf_len=100M
> dma_debug_driver=ath12k_pci trace_event=dma:*,iommu:*
> 3. Booted and waited for page fault, then cat'd
> /sys/kernel/tracing/trace to a file.
>
> Additionally, though I'm pretty sure this is irrelevant now, I added
> logging after each dma_map_single() in the ath12k driver to print the
> function name and resultant address to the kernel log.
>
> Comparing the addresses of several io_page_fault lines in the trace
> and in the kernel log, they line up. So, I'm hopeful this is on the
> right track.
>
> DMA/IOMMU trace: https://cmphys.com/ath12k/iommu_dma_trace-20250701.log
> Kernel log with additional logging:
> https://cmphys.com/ath12k/dmesg-6.12.35-20250701.log
> Diff showing extra logging added to v6.12.35:
> https://cmphys.com/ath12k/ath12k-extra-logging-6.12.35-20250701.diff
Thanks, the log is helpful.
So the whole thing is:
#1 ath12k allocates/maps it at a very early stage:
(udev-worker)-532 [010] ..... 4.878076: map: IOMMU: iova=0x00000000fe980000 -
0x00000000fea00000 paddr=0x000000010ec80000 size=524288
(udev-worker)-532 [010] ..... 4.878079: dma_alloc: 0000:c2:00.0
dma_addr=fe980000 size=524288 virt_addr=000000006cadbcb1 flags=GFP_KERNEL attrs=
#2 here it is unmapped/freed
kworker/u64:0-12 [011] ..... 327.747763: dma_free: 0000:c2:00.0
dma_addr=fe980000 size=524288 virt_addr=000000006cadbcb1 attrs=
kworker/u64:0-12 [011] ..... 327.747766: unmap: IOMMU: iova=0x00000000fe980000 -
0x00000000fea00000 size=524288 unmapped_size=524288
#3 then the page fault
irq/26-AMD-Vi-154 [006] ..... 327.753942: io_page_fault: IOMMU:ath12k_pci
0000:c2:00.0 iova=0x00000000fe980000 flags=0x0001
#4 here seems ath12k is recovering
[ 327.849022] mhi mhi0: Requested to power ON
This gives me the impression that the IOMMU page fault is caused by misbehaved firmware
which crashes. The sequence is, first firmware crashes, then host gets that event and
begins to recover, during which some DMA buffer is freed/unmapped. However the firmware
does not know that and continues to access it, and hence the page fault.
Matt, could you help enable verbose ath12k log to verify my guess?
modprobe ath12k debug_mask=0xffffffff
note this would make ath12k throws lots of logs. Here the purpose is to check whether
firmware crash happens before the page fault. You may monitor
ath12k_dbg(ab, ATH12K_DBG_BOOT, "reset starting\n");
in ath12k_core_reset(), which is the entry of ath12k recovery process.
And one more thing, the issue buffer is handled by dma_alloc/free_xxx API family, so
adding logs to dma_map/unmap_xxx API does not help here.
>
> Thanks,
> Matt
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: ath12k_pci errors and loss of connectivity in 6.12.y branch
2025-07-02 7:10 ` Baochen Qiang
@ 2025-07-02 14:53 ` Matt Mower
2025-07-03 2:19 ` Baochen Qiang
0 siblings, 1 reply; 10+ messages in thread
From: Matt Mower @ 2025-07-02 14:53 UTC (permalink / raw)
To: Baochen Qiang
Cc: Vasant Hegde, Robin Murphy, Jeff Johnson, will, joro,
linux-wireless, ath12k, 1107521, iommu
> Matt, could you help enable verbose ath12k log to verify my guess?
Here are kernel logs with ath12k debugging enabled:
1. WLAN.HMT.1.0.c5-00481-QCAHMTSWPL_V1.0_V2.0_SILICONZ-3
https://cmphys.com/ath12k/dmesg-6.12.35-ath12kdebug-fw0x100301e1-20250702.log
2. WLAN.HMT.1.1.c5-00284.1-QCAHMTSWPL_V1.0_V2.0_SILICONZ-3
https://cmphys.com/ath12k/dmesg-6.12.35-ath12kdebug-fw0x1108811c-20250702.log
I captured these after setting CONFIG_ATH12K_DEBUG=y and running "echo
0xffffffff > /sys/module/ath12k/parameters/debug_mask" during boot
(using @reboot in crontab).
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: ath12k_pci errors and loss of connectivity in 6.12.y branch
2025-07-02 14:53 ` Matt Mower
@ 2025-07-03 2:19 ` Baochen Qiang
2025-07-03 5:08 ` Matt Mower
0 siblings, 1 reply; 10+ messages in thread
From: Baochen Qiang @ 2025-07-03 2:19 UTC (permalink / raw)
To: Matt Mower
Cc: Vasant Hegde, Robin Murphy, Jeff Johnson, will, joro,
linux-wireless, ath12k, 1107521, iommu
On 7/2/2025 10:53 PM, Matt Mower wrote:
>> Matt, could you help enable verbose ath12k log to verify my guess?
>
> Here are kernel logs with ath12k debugging enabled:
Thanks Matt.
I see firmware crash before IOMMU fault in both logs, which verifies my guess.
> 1. WLAN.HMT.1.0.c5-00481-QCAHMTSWPL_V1.0_V2.0_SILICONZ-3
> https://cmphys.com/ath12k/dmesg-6.12.35-ath12kdebug-fw0x100301e1-20250702.log
[ 91.625809] ath12k_pci 0000:c2:00.0: mhi notify status reason MHI_CB_EE_RDDM
[ 91.625916] ath12k_pci 0000:c2:00.0: reset starting
[ 91.674375] ath12k_pci 0000:c2:00.0: waiting recovery start...
[ 91.679445] ath12k_pci 0000:c2:00.0: setting mhi state: POWER_OFF(3)
[ 91.680721] ath12k_pci 0000:c2:00.0: qmi wifi fw del server
[ 91.680754] ath12k_pci 0000:c2:00.0: setting mhi state: DEINIT(1)
[ 91.681842] ath12k_pci 0000:c2:00.0: cookie:0x0
[ 91.681858] ath12k_pci 0000:c2:00.0: WLAON_WARM_SW_ENTRY 0x14c4e54
[ 91.687109] ath12k_pci 0000:c2:00.0: AMD-Vi: Event logged [IO_PAGE_FAULT domain=0x0010
address=0xfe980000 flags=0x0020]
> 2. WLAN.HMT.1.1.c5-00284.1-QCAHMTSWPL_V1.0_V2.0_SILICONZ-3
> https://cmphys.com/ath12k/dmesg-6.12.35-ath12kdebug-fw0x1108811c-20250702.log
>
[ 113.621429] ath12k_pci 0000:c2:00.0: mhi notify status reason MHI_CB_EE_RDDM
[ 113.621794] ath12k_pci 0000:c2:00.0: reset starting
[ 113.670134] ath12k_pci 0000:c2:00.0: waiting recovery start...
[ 113.675177] ath12k_pci 0000:c2:00.0: setting mhi state: POWER_OFF(3)
[ 113.676331] ath12k_pci 0000:c2:00.0: setting mhi state: DEINIT(1)
[ 113.676581] ath12k_pci 0000:c2:00.0: qmi wifi fw del server
[ 113.676874] ath12k_pci 0000:c2:00.0: AMD-Vi: Event logged [IO_PAGE_FAULT domain=0x0010
address=0xfea50000 flags=0x0020]
> I captured these after setting CONFIG_ATH12K_DEBUG=y and running "echo
> 0xffffffff > /sys/module/ath12k/parameters/debug_mask" during boot
> (using @reboot in crontab).
Unfortunately I can not tell the root cause to the firmware crash from host log.
Internally I will try to repro this issue, in the meanwhile, Matt, could you help do some
more work to narrow down the problematic change?
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: ath12k_pci errors and loss of connectivity in 6.12.y branch
2025-07-03 2:19 ` Baochen Qiang
@ 2025-07-03 5:08 ` Matt Mower
2025-07-27 16:12 ` Matt Mower
0 siblings, 1 reply; 10+ messages in thread
From: Matt Mower @ 2025-07-03 5:08 UTC (permalink / raw)
To: Baochen Qiang
Cc: Vasant Hegde, Robin Murphy, Jeff Johnson, will, joro,
linux-wireless, ath12k, 1107521, iommu
> in the meanwhile, Matt, could you help do some
> more work to narrow down the problematic change?
I went the opposite direction and started cherry picking changes from
6.15.y, starting where ath12k diverged from 6.12.y in Sep 2024. I
found stability pretty quick and was able to bisect down to a single
commit where stability started. See branch
https://github.com/mdmower/linux/commits/mdm-6.12.35-ath12k-1/ for the
cherry pick history on top of v6.12.35. The commit where stability
started is:
wifi: ath12k: modify ath12k_mac_op_bss_info_changed() for MLO
https://git.kernel.org/pub/scm/linux/kernel/git/stable/linux.git/commit/drivers/net/wireless/ath/ath12k?h=linux-6.15.y&id=afbab6e4e88da68cca94cabfc1604d71db161d42
I know this isn't exactly what you needed (especially since the commit
is part of a patch series, not standalone), but maybe it'll help?
Logs with ath12k debugging enabled:
1. Last broken revision in my branch:
wifi: ath12k: modify ath12k_get_arvif_iter() for MLO
https://github.com/mdmower/linux/commit/841b7f4af08f9d80e6db862218d341402dfc9acf
https://cmphys.com/ath12k/dmesg-6.12.35-cherrypicks-841b7f4af08f-ath12kdebug.log
2. First fixed revision in my branch:
wifi: ath12k: modify ath12k_mac_op_bss_info_changed() for MLO
https://github.com/mdmower/linux/commit/931abb9e838e376a53a82c0f638fb63a9d31e737
https://cmphys.com/ath12k/dmesg-6.12.35-cherrypicks-931abb9e838e-ath12kdebug.log
I'll continue to test this branch to make sure it's not a false
positive, but so far it's looking good.
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: ath12k_pci errors and loss of connectivity in 6.12.y branch
2025-07-03 5:08 ` Matt Mower
@ 2025-07-27 16:12 ` Matt Mower
2025-11-23 22:44 ` Matt Mower
0 siblings, 1 reply; 10+ messages in thread
From: Matt Mower @ 2025-07-27 16:12 UTC (permalink / raw)
To: Baochen Qiang
Cc: Vasant Hegde, Robin Murphy, Jeff Johnson, will, joro,
linux-wireless, ath12k, 1107521, iommu
Baochen - aside from finding a commit where stability disappears (from
my last message, stability is achieved when I cherry pick "wifi:
ath12k: modify ath12k_mac_op_bss_info_changed() for MLO"), is there
any other debugging information I can provide? You mentioned that
firmware is a likely culprit; would it help to get a dump from the
ATH12K_COREDUMP feature (commit "wifi: ath12k: Add firmware coredump
collection support")? If so, could you give me a hint about how to
generate the dump and where to find it?
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: ath12k_pci errors and loss of connectivity in 6.12.y branch
2025-07-27 16:12 ` Matt Mower
@ 2025-11-23 22:44 ` Matt Mower
0 siblings, 0 replies; 10+ messages in thread
From: Matt Mower @ 2025-11-23 22:44 UTC (permalink / raw)
To: 1107521-done, linux-wireless, ath12k
Closing Debian bug report 1107521. I no longer experience these
crashes or loss of network connectivity with the following package
versions in Debian 13:
- firmware-atheros 20250808-1
- linux-image-6.12.48+deb13-amd64 6.12.48-1
- linux-image-6.12.57+deb13-amd64 6.12.57-1
^ permalink raw reply [flat|nested] 10+ messages in thread
end of thread, other threads:[~2025-11-23 22:44 UTC | newest]
Thread overview: 10+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
[not found] <CAPDiVH8gaBH6o_OY-zUWYpDbj5mhiqmofKGb71gLgHOi4vA=Vw@mail.gmail.com>
2025-06-27 5:39 ` ath12k_pci errors and loss of connectivity in 6.12.y branch Baochen Qiang
2025-06-27 10:24 ` Robin Murphy
2025-06-30 6:36 ` Vasant Hegde
2025-07-02 5:17 ` Matt Mower
2025-07-02 7:10 ` Baochen Qiang
2025-07-02 14:53 ` Matt Mower
2025-07-03 2:19 ` Baochen Qiang
2025-07-03 5:08 ` Matt Mower
2025-07-27 16:12 ` Matt Mower
2025-11-23 22:44 ` Matt Mower
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox