Linux SCSI subsystem development
 help / color / mirror / Atom feed
From: Laurence Oberman <loberman@redhat.com>
To: "Pluess, Tobias" <tpluess@ieee.org>
Cc: linux-scsi@vger.kernel.org, sathya.prakash@broadcom.com,
	 sreekanth.reddy@broadcom.com,
	suganath-prabu.subramani@broadcom.com
Subject: Re:
Date: Wed, 30 Sep 2026 11:09:40 -0400	[thread overview]
Message-ID: <40901345e474c45f531a9bc48ff9658d7bfe02aa.camel@redhat.com> (raw)
In-Reply-To: <CAB4Wsm-HLX2haXvsq5CwnfpyxnzkCEnz__jGFh4_zr7RL78UZg@mail.gmail.com>

On Wed, 2026-09-30 at 10:09 +0200, Pluess, Tobias wrote:
> Hello Laurence,
> sorry for my late reply.
> Thanks for looking into this. I will setup my system accordingly to
> be
> able to boot older versions. It will take a few days until I have
> time
> but I will report ASAP my findings.
> 
> From inspecting my kernel update logs,
> 
> zgrep -E "status installed .*kernel-" /var/log/dpkg.log.*
> 
> I see that around that time when the error started occuring was when
> updating from 6.14 to 6.17 and later. So I would guess the
> corresponding changes must have happened around that point. I'm not
> 100% sure, but I will try booting the older kernels and will report
> my
> findings.
> 
> Thanks!
> greetings Tobias
> 
> On Sat, Sep 26, 2026 at 2:30 PM Laurence Oberman
> <loberman@redhat.com> wrote:
> > 
> > On Fri, 2026-09-25 at 21:00 +0200, Pluess, Tobias wrote:
> > > Hi,
> > > 
> > > I am seeing a problem with the mpt3sas driver when using an
> > > LSI/Broadcom SAS9207-8e connected to an external tape drive (HP
> > > Ultrium 6650).
> > > This setup worked without issues with older kernels. I cannot
> > > 100%
> > > sure identify at which date the problem started to occur, but I
> > > would
> > > say it started to occur around April or May this year.
> > > 
> > > What happens:
> > > a) the external tape drive is connected to the HBA using a SAS
> > > 8088
> > > cable.
> > > b) the tape drive is switched on and is detected normally and
> > > works
> > > as
> > > expected, backups can be made.
> > > c) after the backup finishes, the tape is ejected and the drive
> > > is
> > > idle, it is switched off.
> > > 
> > > Here begins the interesting story. In the dmesg, I see the
> > > following
> > > errors:
> > > 
> > > [Fri Sep 25 20:02:40 2026] mpt2sas_cm6: detecting:
> > > handle(0x0009),
> > >                            sas_address(0x50014380353cba70),
> > > phy(3)
> > > [Fri Sep 25 20:02:40 2026] mpt2sas_cm6: REPORT_LUNS:
> > > handle(0x0009),
> > >                            retries(0)
> > > [Fri Sep 25 20:02:40 2026] TEST_UNIT_READY: handle(0x0009) lun(0)
> > > [Fri Sep 25 20:02:41 2026] mpt2sas_cm6: detecting:
> > > handle(0x0009),
> > >                            sas_address(0x50014380353cba70),
> > > phy(3)
> > > [Fri Sep 25 20:02:41 2026] mpt2sas_cm6: REPORT_LUNS:
> > > handle(0x0009),
> > >                            retries(0)
> > > [Fri Sep 25 20:02:41 2026] TEST_UNIT_READY: handle(0x0009) lun(0)
> > > [Fri Sep 25 20:02:41 2026] START_UNIT: handle(0x0009), lun(0)
> > > [Fri Sep 25 20:02:45 2026] TEST_UNIT_READY: handle(0x0009) lun(0)
> > > [Fri Sep 25 20:02:46 2026] mpt2sas_cm6: detecting:
> > > handle(0x0009),
> > >                            sas_address(0x50014380353cba70),
> > > phy(3)
> > > [Fri Sep 25 20:02:46 2026] mpt2sas_cm6: REPORT_LUNS:
> > > handle(0x0009),
> > >                            retries(0)
> > > [Fri Sep 25 20:02:46 2026] TEST_UNIT_READY: handle(0x0009) lun(0)
> > > [Fri Sep 25 20:02:46 2026] START_UNIT: handle(0x0009), lun(0)
> > > [Fri Sep 25 20:02:49 2026] TEST_UNIT_READY: handle(0x0009) lun(0)
> > > [Fri Sep 25 20:02:52 2026] mpt2sas_cm6: log_info(0x31130000):
> > > originator(PL), code(0x13), sub_code(0x0000)
> > > [Fri Sep 25 20:02:53 2026] mpt2sas_cm6: handle(0x0009),
> > > ioc_status(0x0022) failure at
> > > drivers/scsi/mpt3sas/mpt3sas_transport.c:228/_transport_set_ident
> > > ify(
> > > )!
> > > [Fri Sep 25 20:02:53 2026] mpt2sas_cm6: failure at
> > > drivers/scsi/mpt3sas/mpt3sas_scsih.c:8407/_scsih_add_device()!
> > > [Fri Sep 25 20:02:54 2026] mpt2sas_cm6: handle(0x0009),
> > > ioc_status(0x0022) failure at
> > > drivers/scsi/mpt3sas/mpt3sas_transport.c:228/_transport_set_ident
> > > ify(
> > > )!
> > > [Fri Sep 25 20:02:54 2026] mpt2sas_cm6: failure at
> > > drivers/scsi/mpt3sas/mpt3sas_scsih.c:8407/_scsih_add_device()!
> > > [Fri Sep 25 20:02:55 2026] mpt2sas_cm6: handle(0x0009),
> > > ioc_status(0x0022) failure at
> > > drivers/scsi/mpt3sas/mpt3sas_transport.c:228/_transport_set_ident
> > > ify(
> > > )!
> > > 
> > > Here, the last 2 lines are printed to dmesg approximately every
> > > second, ad infinitum until the system is rebooted; the error
> > > never
> > > stops. I am certain this did not occur with older versions of
> > > mpt3sas,
> > > but as I said, I cannot pinpoint exactly when the change
> > > happened.
> > > 
> > > Powering up the tape drive again fixes the error, but as soon as
> > > the
> > > tape drive is switched off again, the error reappears.
> > > 
> > > However, the error doesn't occur every time the drive is switched
> > > off.
> > > Sometimes the mpt3sas driver seems to be happy and no errors are
> > > reported. I think it has to do with some kind of timing, i.e. in
> > > which
> > > internal state the mpt3sas driver is when the external drive is
> > > disconnected.
> > > 
> > > Hardware:
> > > * Serial Attached SCSI controller: Broadcom / LSI SAS2308 PCI-
> > > Express
> > > Fusion-MPT SAS-2 (rev 05)
> > > * external tape drive HP Ultrium 6650
> > > * Linux pve0 7.0.14-12-pve #1 SMP PREEMPT_DYNAMIC PMX 7.0.14-12
> > > (2026-08-11T11:05Z) x86_64 GNU/Linux
> > > * mpt3sas
> > > 
> > > 
> > > I wonder why this problem occurs and how it could be prevented. I
> > > have
> > > to say, my hot swap HDDs are connected to
> > > 
> > > Serial Attached SCSI controller: Broadcom / LSI SAS3416 Fusion-
> > > MPT
> > > Tri-Mode I/O Controller Chip (IOC) (rev 01)
> > > 
> > > and here, I never have the issue, it only appears with the 9207-
> > > 8e.
> > > 
> > > I made my own "fix" to stop the driver from filling up my dmesg
> > > with
> > > errors; using a little bash script with these commands
> > > 
> > > echo "0000:51:00.0" > /sys/bus/pci/drivers/mpt3sas/unbind
> > > echo "0000:51:00.0" > /sys/bus/pci/drivers/mpt3sas/bind
> > > 
> > > resets the driver and it then stops printing errors. However I
> > > think
> > > this is not a true solution to the problem, it just fixes the
> > > symptom.
> > > 
> > > 
> > > Any ideas how to proceed?
> > > Unfortunately I don't quickly have another system handy where I
> > > could
> > > test older versions of the driver.
> > > 
> > > Thanks,
> > > Greetings
> > > Tobias
> > > 
> > 
> > Hello
> > Any chance you can boot an earlier set of kernels to try narrow
> > down
> > when this started. I don't have the older hardware to reproduce in
> > our
> > lab here. I will do some code inspection in the meantime to try
> > figure
> > out what triggers this and come up with some ideas.
> > 
> > Would be good to know which version older kernel was not seeing
> > this.
> > Thanks
> > Laurence
> > 
Hi Tobias,
Thanks, that version range helps. From code inspection, one possibly
relevant change between 6.14 and 6.17 is:
37c4e72b0651 scsi: Fix sas_user_scan() to handle wildcard and multi-
channel scans
That changed SAS wildcard scan behavior so a scan may now cover more
channels than before. It may only be exposing an existing mpt3sas
stale-handle issue, but it could explain why the problem became visible
now.

Could you please test these if possible:
v6.14
v6.15
v6.16
v6.17

The key thing is to find the first version where powering off the tape
drive starts the repeated _transport_set_identify() /
_scsih_add_device() messages.

Also, while the messages are repeating, could you check whether
anything is repeatedly rescanning SCSI hosts?
grep -R "host.*/scan|/scan" /etc /usr/lib/systemd /lib/systemd
/etc/udev /lib/udev 2>/dev/null

If you are able to build/test a kernel with 37c4e72b0651 reverted, that
would also be very useful.

My current theory is that after the tape drive is powered off, the
SAS2308/mpt3sas path may still have a stale device handle, and a
wildcard scan keeps trying to add it again. The add then fails because
the firmware no longer returns valid identify data for that handle.

Thanks,
Laurence


  reply	other threads:[~2026-09-30 15:09 UTC|newest]

Thread overview: 34+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-25 19:00 Pluess, Tobias
2026-09-26 12:30 ` Laurence Oberman
2026-09-30  8:09   ` Re: Pluess, Tobias
2026-09-30 15:09     ` Laurence Oberman [this message]
2026-09-30 15:52       ` Re: Laurence Oberman
  -- strict thread matches above, loose matches on Subject: below --
2024-08-22 20:54 [PATCH 2/2] scsi: ufs: core: Fix the code for entering hibernation Bao D. Nguyen
2024-08-22 21:08 ` Bart Van Assche
2024-08-23 12:01   ` Manivannan Sadhasivam
2024-08-23 14:23     ` Bart Van Assche
2024-08-23 14:58       ` Manivannan Sadhasivam
2024-08-23 16:07         ` Bart Van Assche
2024-08-23 16:48           ` Manivannan Sadhasivam
2024-08-23 18:05             ` Bart Van Assche
2024-08-24  2:29               ` Manivannan Sadhasivam
2024-08-24  2:48                 ` Bart Van Assche
2024-08-24  3:03                   ` Manivannan Sadhasivam
2024-08-26  6:48                     ` Can Guo
2022-11-21 11:11 Denis Arefev
2022-11-21 14:28 ` Jason Yan
     [not found] <20211011231530.GA22856@t>
2021-10-12  1:23 ` James Bottomley
2021-10-12  2:30   ` Bart Van Assche
     [not found] <5e7dc543.vYG3wru8B/me1sOV%chenanqing@oppo.com>
2020-03-27 15:53 ` Re: Lee Duncan
2017-11-13 14:55 Re: Amos Kalonzo
2017-05-03  6:23 Re: H.A
2017-02-23 15:09 Qin's Yanjun
2015-08-19 13:01 christain147
     [not found] <132D0DB4B968F242BE373429794F35C22559D38329@NHS-PCLI-MBC011.AD1.NHS.NET>
2015-06-08 11:09 ` Practice Trinity (NHS SOUTH SEFTON CCG)
2014-11-14 20:50 salim
2014-11-14 18:56 milke
     [not found] <6A286AB51AD8EC4180C4B2E9EF1D0A027AAD7EFF1E@exmb01.wrschool.net>
2014-09-08 17:36 ` Deborah Mayher
2014-07-24  8:37 Richard Wong
2014-06-16  7:10 Re: Angela D.Dawes
2014-06-15 20:36 Re: Angela D.Dawes
     [not found] <blk-mq updates>
     [not found] ` <1397464212-4454-1-git-send-email-hch@lst.de>
2014-04-15 20:16   ` Re: Jens Axboe
     [not found] <B719EF0A9FB7A247B5147CD67A83E60E011FEB76D1@EXCH10-MB3.paterson.k12.nj.us>
2013-08-23 10:47 ` Ruiz, Irma
2012-05-20 22:20 Mr. Peter Wong
2011-03-06 21:28 Augusta Mubarak
2011-03-03 14:20 RE: Lukas Thompson
2011-03-03 13:34 RE: Lukas Thompson
2010-11-18 15:48 [PATCH v3] dm mpath: add feature flag to control call to blk_abort_queue Mike Snitzer
2010-11-18 19:16 ` (unknown), Mike Snitzer
2010-11-18 19:21   ` Mike Snitzer
2010-07-01 10:49 (unknown) FUJITA Tomonori
2010-07-01 12:29 ` Jens Axboe
     [not found] <KC3KJ12CL8BAAEB7@vger.kernel.org>
2005-07-24 10:31 ` Re: wolman
     [not found] <1KFJEFB27B4F3155@vger.kernel.org>
2005-05-30  2:49 ` Re: sevillar
2005-02-26 14:57 Yong Haynes
2004-03-17 22:03 Kendrick Logan
2003-06-03 23:51 (unknown) Justin T. Gibbs
2003-06-03 23:58 ` Marc-Christian Petersen
2002-11-08  5:35 Re: Randy.Dunlap

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=40901345e474c45f531a9bc48ff9658d7bfe02aa.camel@redhat.com \
    --to=loberman@redhat.com \
    --cc=linux-scsi@vger.kernel.org \
    --cc=sathya.prakash@broadcom.com \
    --cc=sreekanth.reddy@broadcom.com \
    --cc=suganath-prabu.subramani@broadcom.com \
    --cc=tpluess@ieee.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox