From: Laurence Oberman <loberman@redhat.com>
To: "Pluess, Tobias" <tpluess@ieee.org>
Cc: linux-scsi@vger.kernel.org, sathya.prakash@broadcom.com,
sreekanth.reddy@broadcom.com,
suganath-prabu.subramani@broadcom.com
Subject: Re:
Date: Wed, 30 Sep 2026 11:09:40 -0400 [thread overview]
Message-ID: <40901345e474c45f531a9bc48ff9658d7bfe02aa.camel@redhat.com> (raw)
In-Reply-To: <CAB4Wsm-HLX2haXvsq5CwnfpyxnzkCEnz__jGFh4_zr7RL78UZg@mail.gmail.com>
On Wed, 2026-09-30 at 10:09 +0200, Pluess, Tobias wrote:
> Hello Laurence,
> sorry for my late reply.
> Thanks for looking into this. I will setup my system accordingly to
> be
> able to boot older versions. It will take a few days until I have
> time
> but I will report ASAP my findings.
>
> From inspecting my kernel update logs,
>
> zgrep -E "status installed .*kernel-" /var/log/dpkg.log.*
>
> I see that around that time when the error started occuring was when
> updating from 6.14 to 6.17 and later. So I would guess the
> corresponding changes must have happened around that point. I'm not
> 100% sure, but I will try booting the older kernels and will report
> my
> findings.
>
> Thanks!
> greetings Tobias
>
> On Sat, Sep 26, 2026 at 2:30 PM Laurence Oberman
> <loberman@redhat.com> wrote:
> >
> > On Fri, 2026-09-25 at 21:00 +0200, Pluess, Tobias wrote:
> > > Hi,
> > >
> > > I am seeing a problem with the mpt3sas driver when using an
> > > LSI/Broadcom SAS9207-8e connected to an external tape drive (HP
> > > Ultrium 6650).
> > > This setup worked without issues with older kernels. I cannot
> > > 100%
> > > sure identify at which date the problem started to occur, but I
> > > would
> > > say it started to occur around April or May this year.
> > >
> > > What happens:
> > > a) the external tape drive is connected to the HBA using a SAS
> > > 8088
> > > cable.
> > > b) the tape drive is switched on and is detected normally and
> > > works
> > > as
> > > expected, backups can be made.
> > > c) after the backup finishes, the tape is ejected and the drive
> > > is
> > > idle, it is switched off.
> > >
> > > Here begins the interesting story. In the dmesg, I see the
> > > following
> > > errors:
> > >
> > > [Fri Sep 25 20:02:40 2026] mpt2sas_cm6: detecting:
> > > handle(0x0009),
> > > sas_address(0x50014380353cba70),
> > > phy(3)
> > > [Fri Sep 25 20:02:40 2026] mpt2sas_cm6: REPORT_LUNS:
> > > handle(0x0009),
> > > retries(0)
> > > [Fri Sep 25 20:02:40 2026] TEST_UNIT_READY: handle(0x0009) lun(0)
> > > [Fri Sep 25 20:02:41 2026] mpt2sas_cm6: detecting:
> > > handle(0x0009),
> > > sas_address(0x50014380353cba70),
> > > phy(3)
> > > [Fri Sep 25 20:02:41 2026] mpt2sas_cm6: REPORT_LUNS:
> > > handle(0x0009),
> > > retries(0)
> > > [Fri Sep 25 20:02:41 2026] TEST_UNIT_READY: handle(0x0009) lun(0)
> > > [Fri Sep 25 20:02:41 2026] START_UNIT: handle(0x0009), lun(0)
> > > [Fri Sep 25 20:02:45 2026] TEST_UNIT_READY: handle(0x0009) lun(0)
> > > [Fri Sep 25 20:02:46 2026] mpt2sas_cm6: detecting:
> > > handle(0x0009),
> > > sas_address(0x50014380353cba70),
> > > phy(3)
> > > [Fri Sep 25 20:02:46 2026] mpt2sas_cm6: REPORT_LUNS:
> > > handle(0x0009),
> > > retries(0)
> > > [Fri Sep 25 20:02:46 2026] TEST_UNIT_READY: handle(0x0009) lun(0)
> > > [Fri Sep 25 20:02:46 2026] START_UNIT: handle(0x0009), lun(0)
> > > [Fri Sep 25 20:02:49 2026] TEST_UNIT_READY: handle(0x0009) lun(0)
> > > [Fri Sep 25 20:02:52 2026] mpt2sas_cm6: log_info(0x31130000):
> > > originator(PL), code(0x13), sub_code(0x0000)
> > > [Fri Sep 25 20:02:53 2026] mpt2sas_cm6: handle(0x0009),
> > > ioc_status(0x0022) failure at
> > > drivers/scsi/mpt3sas/mpt3sas_transport.c:228/_transport_set_ident
> > > ify(
> > > )!
> > > [Fri Sep 25 20:02:53 2026] mpt2sas_cm6: failure at
> > > drivers/scsi/mpt3sas/mpt3sas_scsih.c:8407/_scsih_add_device()!
> > > [Fri Sep 25 20:02:54 2026] mpt2sas_cm6: handle(0x0009),
> > > ioc_status(0x0022) failure at
> > > drivers/scsi/mpt3sas/mpt3sas_transport.c:228/_transport_set_ident
> > > ify(
> > > )!
> > > [Fri Sep 25 20:02:54 2026] mpt2sas_cm6: failure at
> > > drivers/scsi/mpt3sas/mpt3sas_scsih.c:8407/_scsih_add_device()!
> > > [Fri Sep 25 20:02:55 2026] mpt2sas_cm6: handle(0x0009),
> > > ioc_status(0x0022) failure at
> > > drivers/scsi/mpt3sas/mpt3sas_transport.c:228/_transport_set_ident
> > > ify(
> > > )!
> > >
> > > Here, the last 2 lines are printed to dmesg approximately every
> > > second, ad infinitum until the system is rebooted; the error
> > > never
> > > stops. I am certain this did not occur with older versions of
> > > mpt3sas,
> > > but as I said, I cannot pinpoint exactly when the change
> > > happened.
> > >
> > > Powering up the tape drive again fixes the error, but as soon as
> > > the
> > > tape drive is switched off again, the error reappears.
> > >
> > > However, the error doesn't occur every time the drive is switched
> > > off.
> > > Sometimes the mpt3sas driver seems to be happy and no errors are
> > > reported. I think it has to do with some kind of timing, i.e. in
> > > which
> > > internal state the mpt3sas driver is when the external drive is
> > > disconnected.
> > >
> > > Hardware:
> > > * Serial Attached SCSI controller: Broadcom / LSI SAS2308 PCI-
> > > Express
> > > Fusion-MPT SAS-2 (rev 05)
> > > * external tape drive HP Ultrium 6650
> > > * Linux pve0 7.0.14-12-pve #1 SMP PREEMPT_DYNAMIC PMX 7.0.14-12
> > > (2026-08-11T11:05Z) x86_64 GNU/Linux
> > > * mpt3sas
> > >
> > >
> > > I wonder why this problem occurs and how it could be prevented. I
> > > have
> > > to say, my hot swap HDDs are connected to
> > >
> > > Serial Attached SCSI controller: Broadcom / LSI SAS3416 Fusion-
> > > MPT
> > > Tri-Mode I/O Controller Chip (IOC) (rev 01)
> > >
> > > and here, I never have the issue, it only appears with the 9207-
> > > 8e.
> > >
> > > I made my own "fix" to stop the driver from filling up my dmesg
> > > with
> > > errors; using a little bash script with these commands
> > >
> > > echo "0000:51:00.0" > /sys/bus/pci/drivers/mpt3sas/unbind
> > > echo "0000:51:00.0" > /sys/bus/pci/drivers/mpt3sas/bind
> > >
> > > resets the driver and it then stops printing errors. However I
> > > think
> > > this is not a true solution to the problem, it just fixes the
> > > symptom.
> > >
> > >
> > > Any ideas how to proceed?
> > > Unfortunately I don't quickly have another system handy where I
> > > could
> > > test older versions of the driver.
> > >
> > > Thanks,
> > > Greetings
> > > Tobias
> > >
> >
> > Hello
> > Any chance you can boot an earlier set of kernels to try narrow
> > down
> > when this started. I don't have the older hardware to reproduce in
> > our
> > lab here. I will do some code inspection in the meantime to try
> > figure
> > out what triggers this and come up with some ideas.
> >
> > Would be good to know which version older kernel was not seeing
> > this.
> > Thanks
> > Laurence
> >
Hi Tobias,
Thanks, that version range helps. From code inspection, one possibly
relevant change between 6.14 and 6.17 is:
37c4e72b0651 scsi: Fix sas_user_scan() to handle wildcard and multi-
channel scans
That changed SAS wildcard scan behavior so a scan may now cover more
channels than before. It may only be exposing an existing mpt3sas
stale-handle issue, but it could explain why the problem became visible
now.
Could you please test these if possible:
v6.14
v6.15
v6.16
v6.17
The key thing is to find the first version where powering off the tape
drive starts the repeated _transport_set_identify() /
_scsih_add_device() messages.
Also, while the messages are repeating, could you check whether
anything is repeatedly rescanning SCSI hosts?
grep -R "host.*/scan|/scan" /etc /usr/lib/systemd /lib/systemd
/etc/udev /lib/udev 2>/dev/null
If you are able to build/test a kernel with 37c4e72b0651 reverted, that
would also be very useful.
My current theory is that after the tape drive is powered off, the
SAS2308/mpt3sas path may still have a stale device handle, and a
wildcard scan keeps trying to add it again. The add then fails because
the firmware no longer returns valid identify data for that handle.
Thanks,
Laurence
next prev parent reply other threads:[~2026-09-30 15:09 UTC|newest]
Thread overview: 34+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-25 19:00 Pluess, Tobias
2026-09-26 12:30 ` Laurence Oberman
2026-09-30 8:09 ` Re: Pluess, Tobias
2026-09-30 15:09 ` Laurence Oberman [this message]
2026-09-30 15:52 ` Re: Laurence Oberman
-- strict thread matches above, loose matches on Subject: below --
2024-08-22 20:54 [PATCH 2/2] scsi: ufs: core: Fix the code for entering hibernation Bao D. Nguyen
2024-08-22 21:08 ` Bart Van Assche
2024-08-23 12:01 ` Manivannan Sadhasivam
2024-08-23 14:23 ` Bart Van Assche
2024-08-23 14:58 ` Manivannan Sadhasivam
2024-08-23 16:07 ` Bart Van Assche
2024-08-23 16:48 ` Manivannan Sadhasivam
2024-08-23 18:05 ` Bart Van Assche
2024-08-24 2:29 ` Manivannan Sadhasivam
2024-08-24 2:48 ` Bart Van Assche
2024-08-24 3:03 ` Manivannan Sadhasivam
2024-08-26 6:48 ` Can Guo
2022-11-21 11:11 Denis Arefev
2022-11-21 14:28 ` Jason Yan
[not found] <20211011231530.GA22856@t>
2021-10-12 1:23 ` James Bottomley
2021-10-12 2:30 ` Bart Van Assche
[not found] <5e7dc543.vYG3wru8B/me1sOV%chenanqing@oppo.com>
2020-03-27 15:53 ` Re: Lee Duncan
2017-11-13 14:55 Re: Amos Kalonzo
2017-05-03 6:23 Re: H.A
2017-02-23 15:09 Qin's Yanjun
2015-08-19 13:01 christain147
[not found] <132D0DB4B968F242BE373429794F35C22559D38329@NHS-PCLI-MBC011.AD1.NHS.NET>
2015-06-08 11:09 ` Practice Trinity (NHS SOUTH SEFTON CCG)
2014-11-14 20:50 salim
2014-11-14 18:56 milke
[not found] <6A286AB51AD8EC4180C4B2E9EF1D0A027AAD7EFF1E@exmb01.wrschool.net>
2014-09-08 17:36 ` Deborah Mayher
2014-07-24 8:37 Richard Wong
2014-06-16 7:10 Re: Angela D.Dawes
2014-06-15 20:36 Re: Angela D.Dawes
[not found] <blk-mq updates>
[not found] ` <1397464212-4454-1-git-send-email-hch@lst.de>
2014-04-15 20:16 ` Re: Jens Axboe
[not found] <B719EF0A9FB7A247B5147CD67A83E60E011FEB76D1@EXCH10-MB3.paterson.k12.nj.us>
2013-08-23 10:47 ` Ruiz, Irma
2012-05-20 22:20 Mr. Peter Wong
2011-03-06 21:28 Augusta Mubarak
2011-03-03 14:20 RE: Lukas Thompson
2011-03-03 13:34 RE: Lukas Thompson
2010-11-18 15:48 [PATCH v3] dm mpath: add feature flag to control call to blk_abort_queue Mike Snitzer
2010-11-18 19:16 ` (unknown), Mike Snitzer
2010-11-18 19:21 ` Mike Snitzer
2010-07-01 10:49 (unknown) FUJITA Tomonori
2010-07-01 12:29 ` Jens Axboe
[not found] <KC3KJ12CL8BAAEB7@vger.kernel.org>
2005-07-24 10:31 ` Re: wolman
[not found] <1KFJEFB27B4F3155@vger.kernel.org>
2005-05-30 2:49 ` Re: sevillar
2005-02-26 14:57 Yong Haynes
2004-03-17 22:03 Kendrick Logan
2003-06-03 23:51 (unknown) Justin T. Gibbs
2003-06-03 23:58 ` Marc-Christian Petersen
2002-11-08 5:35 Re: Randy.Dunlap
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=40901345e474c45f531a9bc48ff9658d7bfe02aa.camel@redhat.com \
--to=loberman@redhat.com \
--cc=linux-scsi@vger.kernel.org \
--cc=sathya.prakash@broadcom.com \
--cc=sreekanth.reddy@broadcom.com \
--cc=suganath-prabu.subramani@broadcom.com \
--cc=tpluess@ieee.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox