From: Martin Steigerwald <martin@lichtvoll.de>
To: BTRFS <linux-btrfs@vger.kernel.org>, Qu Wenruo <quwenruo.btrfs@gmx.com>
Subject: Re: SSD overheating during scrub in laptop
Date: Tue, 16 Sep 2025 13:36:27 +0200 [thread overview]
Message-ID: <2031348.PYKUYFuaPT@laptop> (raw)
In-Reply-To: <45131321-5ed8-4abe-9edb-19b1936e83b4@gmx.com>
Thanks.
Qu Wenruo - 14.09.25, 23:33:01 CEST:
> 在 2025/9/15 06:58, Qu Wenruo 写道:
> > 在 2025/9/14 21:27, Martin Steigerwald 写道:
[…]
> >> Now I had these SSD goodbyes during regular use in times of heavy I/O
> >> and in the end it could not even scrub that /home partition anymore
> >> in one go.
> >> Linux hung and only way to recover was to forcefully power off the
> >> laptop.
> >
> > Can you setup netconsole and catch the dying message?
I suppose I could do that as I have several laptops at hand. However…
after cleansing out at least some of the dust it did not happen anymore.
And I do not quite feel like putting some dust back or otherwise provoke
the issue. I do not quite feel like trying to overheat the SSD again. :) I
looked at it more closely, it looks fine enough. Nothing melted
apparently, but what do I know? And I think it does not add to the
lifetime of the SSD to overheat it on purpose.
I might have made a photo of some earlier time, but I am not sure. I will
have a look and see whether I can find any. I remember the kernel wrote a
lot about NVME opcodes. However I do not recall the details. Unfortunately
at this recent occasion I did not make a photo.
In case it happens again by itself, I am at going for a netconsole log –
in case I can reproduce it then. Otherwise I would need to run netconsole
permanently to be sure to catch it on the first occurrence already. Not
sure whether that is a good idea.
> > I doubt if it's really the drive dying.
Why? The issue went away after removing at least some of the dust. Of
course that is a correlation and not necessarily a causation, but what
makes you think that something different is happening?
I checked the drive smart status. SMART status is passed. So everything
okay. There is 2% of spare area used, but that is still a very good value.
An indication that there is something to your suspicion is:
Media and Data Integrity Errors: 0
Error Information Log Entries: 0
Warning Comp. Temperature Time: 0
Critical Comp. Temperature Time: 0
So the drive itself recorded no critical composite temperature times. And
if it has fields for that, I suppose it could still record it on overheat
condition.
Oh by the way, it is a Samsung 990 Pro with a firmware version above that
firmware version that was known to be broken regarding SMART status
reporting. In the first I bought I made sure of that myself, but this one
came with a newer firmware version already.
> >> I opened the laptop and removed dust with high pressure air can while
> >> holding the fan still so it would not generate current. And with
> >> disabled laptop battery.
> >>
> >> This fixed the SSDs goodbye issue and I could even scrub that 2 TiB
> >> filesystem again. However, according to sensors command it still had
> >> about 80,8 °C composite temperature and even 100,8 °C for sensor 2
> >> for the NVME SSD at "nvme-pci-0300" shortly before the end of the
> >> scrub, with critical for composite at 84,8 °C, but no critical set
> >> for sensor 2. That is still quite high. Granted, it took about 7-8
> >> minutes of scrubbing at 3,5 to 4,2 GiB/s in one go for it to heat up
> >> like this. But on the other hand it is not summer anymore and room
> >> is not as hot as in summer.
> >
> > I have hit similar situation, but the symptom is very different, the
> > death come silently at boot, the drive is no longer recognized by the
> > BIOS thus no longer bootable, and Linux kernel from liveUSB will not
> > recognize it either.
Hmm, did you capture any logs?
> > That's why I'm asking if it's really dying caused by the heat.
Ah okay… You mentioned your physical solution being a thermal pad. So did
this symptom go away with it? Or was your physical solution a general
solution to reduce temperature, a solution unrelated to that symptom you
described?
> >> My solution to this will be to remove the dust inside laptop about
> >> every half year. However… I was a bit surprised that the NVME SSD
> >> would not throttle itself before saying goodbye. Or maybe it tried
> >> and it was not enough or to late? The laptop is a bit less than 15
> >> months old. So I conclude that it takes less than a year for the
> >> cooling system to become quite a bit less effective due to dust.
> >> Good old ThinkPad T520 took way longer for that. But it is way
> >> larger on the other hand with more space to distribute heat.
> >
> > You can refer to man page of btrfs-scrub, it provides the way to limit
> > the bandwidth of scrub using cgroup or even the btrfs sysfs interface.
>
> And I forgot my physical solution, with a thick thermal pad, you can
> connect the NVME to the back plate of the laptop to dissipate heat.
Well it has some violet colored pad below the NVME SSD. So a pad that is
between NVME SSD and motherboard. Not sure whether it is a cooling pad. It
might be more about isolating the motherboard from the heat of the SSD? It
has been there for the initial SSD Lenovo put into the laptop that I
replaced with the larger Samsung 990 Pro one.
But yeah, I may order a cooling plate to connection to the back plate.
Good idea.
> If there is enough space (I doubt for laptop though), you can add some
> thin heatsink for the drive.
Ha ha, no way. I would like to put a 8 TB SSD in there, but it would be
double layered and I doubt it would be a good idea to put a double layered
SSD into the laptop.
Best,
--
Martin
next prev parent reply other threads:[~2025-09-16 11:36 UTC|newest]
Thread overview: 5+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-09-14 11:57 SSD overheating during scrub in laptop Martin Steigerwald
2025-09-14 21:28 ` Qu Wenruo
2025-09-14 21:33 ` Qu Wenruo
2025-09-16 11:36 ` Martin Steigerwald [this message]
2025-09-16 21:17 ` Qu Wenruo
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=2031348.PYKUYFuaPT@laptop \
--to=martin@lichtvoll.de \
--cc=linux-btrfs@vger.kernel.org \
--cc=quwenruo.btrfs@gmx.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.