All of lore.kernel.org
 help / color / mirror / Atom feed
From: David Laight <david.laight.linux@gmail.com>
To: Borislav Petkov <bp@alien8.de>
Cc: Li Zhe <lizhe.67@bytedance.com>,
	akpm@linux-foundation.org, apopple@nvidia.com, arnd@arndb.de,
	balbirs@nvidia.com, dave.hansen@linux.intel.com,
	david@kernel.org, kees@kernel.org, mingo@redhat.com,
	muchun.song@linux.dev, rppt@kernel.org, tglx@kernel.org,
	linux-arch@vger.kernel.org, linux-hardening@vger.kernel.org,
	linux-kernel@vger.kernel.org, linux-mm@kvack.org, x86@kernel.org
Subject: Re: [PATCH v8 7/9] x86/string: extend memcpy_flushcache() fixed-size fastpaths
Date: Thu, 30 Jul 2026 09:14:55 +0100	[thread overview]
Message-ID: <20260730091455.1242d01a@pumpkin> (raw)
In-Reply-To: <20260729234842.GFamqRWva8h7X59ccN@fat_crate.local>

On Wed, 29 Jul 2026 16:48:42 -0700
Borislav Petkov <bp@alien8.de> wrote:

...
> at least the code is making a lot more sense now.
> 
> The fact that you had to axe off so much cruft off of it tells me that you
> haven't really measured it right.

Especially since if you do actually measure the clock counts (non-trivial)
you'll find that loops are often completely free.
The out-of-order execution unit will (effectively) execute the loop control
instructions to generate a list of instructions that get executed at a
later time.
So provided the loop control doesn't use more clocks than the loop body
(and there are spare ALU units - usually true) loops really make little
difference.

This also means that unrolling loops often doesn't make things faster.
You do need to minimise the loop control instructions (and gcc doesn't
like the best loop that uses negative offsets from the end), and
intel cpu can't execute single clock loops (amd ones can).

Inlining also increases the code size, the I-cache reads are actually
likely to be significant.
You need to time a single 'cold-cache' call not just loops for long
enough that the result is also skewed by timer ticks (etc).

	David 

> 
> So why do I really want your patch?
> 
> Thx.
> 


  parent reply	other threads:[~2026-07-30  8:15 UTC|newest]

Thread overview: 15+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-27 12:34 [PATCH v8 0/9] mm: optimize zone-device memmap initialization Li Zhe
2026-07-27 12:34 ` [PATCH v8 1/9] mm: fix stale ZONE_DEVICE refcount comment Li Zhe
2026-07-27 12:34 ` [PATCH v8 2/9] mm: factor zone-device page init helpers out of __init_zone_device_page Li Zhe
2026-07-27 12:34 ` [PATCH v8 3/9] mm: add a set_page_section_from_pfn() helper Li Zhe
2026-07-27 12:34 ` [PATCH v8 4/9] mm: add a template-based fast path for zone-device page init Li Zhe
2026-07-27 12:34 ` [PATCH v8 5/9] mm: extend the template fast path to zone-device compound tails Li Zhe
2026-07-27 12:34 ` [PATCH v8 6/9] string: introduce memcpy_nontemporal() Li Zhe
2026-07-27 12:34 ` [PATCH v8 7/9] x86/string: extend memcpy_flushcache() fixed-size fastpaths Li Zhe
2026-07-29 23:48   ` Borislav Petkov
2026-07-30  8:05     ` Li Zhe
2026-07-30  8:14     ` David Laight [this message]
2026-07-27 12:34 ` [PATCH v8 8/9] mm: use memcpy_nontemporal() in zone-device template copies Li Zhe
2026-07-27 12:34 ` [PATCH v8 9/9] mm: always use the zone-device template init path Li Zhe
2026-07-27 20:57 ` [PATCH v8 0/9] mm: optimize zone-device memmap initialization Andrew Morton
2026-07-28  6:18   ` Li Zhe

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20260730091455.1242d01a@pumpkin \
    --to=david.laight.linux@gmail.com \
    --cc=akpm@linux-foundation.org \
    --cc=apopple@nvidia.com \
    --cc=arnd@arndb.de \
    --cc=balbirs@nvidia.com \
    --cc=bp@alien8.de \
    --cc=dave.hansen@linux.intel.com \
    --cc=david@kernel.org \
    --cc=kees@kernel.org \
    --cc=linux-arch@vger.kernel.org \
    --cc=linux-hardening@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-mm@kvack.org \
    --cc=lizhe.67@bytedance.com \
    --cc=mingo@redhat.com \
    --cc=muchun.song@linux.dev \
    --cc=rppt@kernel.org \
    --cc=tglx@kernel.org \
    --cc=x86@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.