Netdev List
 help / color / mirror / Atom feed
From: Philipp Schuster <philipp.schuster@cyberus-technology.de>
To: Maciej Fijalkowski <maciej.fijalkowski@intel.com>
Cc: Philipp Schuster <philipp.schuster@cyberus-technology.de>,
	Petr Oros <poros@redhat.com>,
	Jacob Keller <jacob.e.keller@intel.com>,
	David Laight <david.laight.linux@gmail.com>,
	John Ousterhout <ouster@cs.stanford.edu>,
	stable@vger.kernel.org,
	"Anthony L . Nguyen" <anthony.l.nguyen@intel.com>,
	intel-wired-lan@lists.osuosl.org,
	Przemyslaw Kitszel <przemyslaw.kitszel@intel.com>,
	netdev@vger.kernel.org
Subject: Re: [Intel-wired-lan] [PATCH net v3] ice: fix packet corruption due to extraneous page flip
Date: Wed, 30 Sep 2026 17:21:00 +0200	[thread overview]
Message-ID: <20260930152101.1503079-1-philipp.schuster@cyberus-technology.de> (raw)
In-Reply-To: <ah2qcDLxMlGYNhgf@boxer>

Hi there,

We were hitting this bug pretty reliably in one of our setups and it
looks like this patch is fixing our problems. Is there any update on
that? For now, we've moved to a newer kernel (7.2.8) but it might still
be relevant to many folks out there.

I did not monitor memory growth during these tests.

A few details about my observation. Please note that I heavily used
an LLM to track down the issue and to create a reproducer. The LLM's
output and reproducer seemed valid to me:

I ran into this bug while benchmarking live migration with Cloud
Hypervisor between two hosts with Intel E810 NICs (ice, MTU 9000,
6.18.41). We migrate a VM with 120 vCPUs and 20 GiB of memory under a
write-heavy load, using TLS over 8 parallel TCP connections at about
8 GiB/s. On the receiver, up to 5 of 6 migrations failed because TLS
could not decrypt a record, and once even the plaintext record header
was broken. First I suspected our own code. But an in-process check
that decrypted 1.1 billion records again right after encryption found
no mismatch, and switching to another cipher did not help. So the data
was still correct when it left userspace. A plain-TCP test tool with
bulk traffic transferred more than 1.4 TB without any error, so the
trigger had to depend on the traffic pattern. At this point I found this
LKML thread. With the patched ice module on the receiver, my reproducer
passed all runs without a failure (432 GiB). Without the patch, 2 out of
3 runs had corrupted data.

The reproducer is a small C program (about 290 lines) that uses plain
TCP. To trigger the bug reliably, it does four things. First, it chooses
the write lengths so that frames exactly fill one or two 3 KiB Rx
buffers, and it uses TCP_NODELAY so that these frames are really sent
this way. Second, the sender is bursty and fast: 9 parallel connections,
plus 120 busy threads that simulate the vCPUs of a large VM with dirty
page tracking. Third, the receiver keeps many packets queued, because it
reads into scattered locations of a large buffer. Fourth, the stream
content is deterministic, so the receiver can check every byte. Without
the busy threads, 720 GiB went through without errors, and in single
test runs, leaving out one of the other conditions also made the
corruption disappear. Each corrupt block was at most one Rx buffer long
and contained data from the same connection, about 11 MiB later in the
stream. TCP itself reported no error.

Thanks,
Philipp

  reply	other threads:[~2026-09-30 15:21 UTC|newest]

Thread overview: 13+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-05-12 18:19 [PATCH net v3] ice: fix packet corruption due to extraneous page flip John Ousterhout
2026-05-13  9:07 ` David Laight
2026-05-13 16:28   ` John Ousterhout
2026-05-13 20:49     ` David Laight
2026-05-14  4:47       ` John Ousterhout
2026-05-14 10:01         ` David Laight
2026-05-14 16:43           ` Jacob Keller
2026-05-26 12:47             ` [Intel-wired-lan] " Petr Oros
2026-05-26 15:11               ` David Laight
2026-05-26 22:17               ` John Ousterhout
2026-06-01 15:51               ` Maciej Fijalkowski
2026-09-30 15:21                 ` Philipp Schuster [this message]
2026-05-14  9:08 ` Loktionov, Aleksandr

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260930152101.1503079-1-philipp.schuster@cyberus-technology.de \
    --to=philipp.schuster@cyberus-technology.de \
    --cc=anthony.l.nguyen@intel.com \
    --cc=david.laight.linux@gmail.com \
    --cc=intel-wired-lan@lists.osuosl.org \
    --cc=jacob.e.keller@intel.com \
    --cc=maciej.fijalkowski@intel.com \
    --cc=netdev@vger.kernel.org \
    --cc=ouster@cs.stanford.edu \
    --cc=poros@redhat.com \
    --cc=przemyslaw.kitszel@intel.com \
    --cc=stable@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox