From: Alison Schofield <alison.schofield@intel.com>
To: Robert Richter <rrichter@amd.com>
Cc: Davidlohr Bueso <dave@stgolabs.net>,
Jonathan Cameron <jic23@kernel.org>,
Dave Jiang <dave.jiang@intel.com>,
Vishal Verma <vishal.l.verma@intel.com>,
Ira Weiny <iweiny@kernel.org>, Li Ming <ming.li@zohomail.com>,
<linux-cxl@vger.kernel.org>
Subject: Re: [PATCH v5 1/7] Documentation/cxl: Describe mixed-granularity regions
Date: Fri, 18 Sep 2026 13:48:14 -0700 [thread overview]
Message-ID: <aq2jjiXfDKo0SwwQ@aschofie-mobl2.lan> (raw)
In-Reply-To: <aqqCpDux1vOwKugi@rric.localdomain>
On Wed, Sep 16, 2026 at 01:51:00PM +0200, Robert Richter wrote:
> On 03.09.26 16:23:43, Alison Schofield wrote:
> > Linux has required that a region's interleave granularity equal the
> > interleave granularity of its interleaving root decoder. A mixed-
> > granularity region lifts that restriction: the region granularity may be
> > finer than the root's, so the interleave refines from the root toward the
> > endpoints.
> >
> > Mixed-granularity region support introduces interleave relationships that
> > are not obvious from the existing CXL documentation. This is particularly
> > true for the 3-, 6-, and 12-way configurations and for understanding which
> > configurations permitted by the CXL Specification are supported by Linux.
> >
> > Document the mixed-granularity model and Linux's coarse-to-fine
> > restriction. Include the relevant configurations from CXL 4.0 Section
> > 9.13.1.1 and annotate their Linux support status.
> >
> > This intentionally repeats information from the CXL Specification. The
> > specification remains authoritative, but showing the Linux support policy
> > alongside the legal configurations ensures readers do not mistake an
> > unsupported Linux setup for a configuration unsupported by the CXL
> > Specification.
>
> This doc is too long. It should only contain the following:
Sure I will take another pass at trimming. I'll respond piece by piece
below.
>
> 1) Describe a brief description of how the specification defines
> interleaving along with references to it. Only add what is needed to
> describe the cxl driver specifics in 2).
>
> 2) Limitations and implementation specifics of the cxl driver compared
> to the specification.
>
> This document is ambiguous on what the specs defines and that the
> driver provides. This is blurred in this doc.
>
Let's go thru that point by point because I think that was the whole
point of the doc ;)
> >
> > Assisted-by: Claude:Opus-5
>
> I am fine with that, but we shouldn't use it to create tons of docs
> which we will need another AI assistant again to read. Though, I read
> it on my own. :-)
I find AI is particularly useful for the mechanical .rst work: building tables
and using correct formatting, while keeping the result readable both as rendered
HTML and plain-text patch review.
>
> > Signed-off-by: Alison Schofield <alison.schofield@intel.com>
> > ---
> > .../driver-api/cxl/linux/cxl-driver.rst | 153 ++++++++++++++++++
> > 1 file changed, 153 insertions(+)
> >
> > diff --git a/Documentation/driver-api/cxl/linux/cxl-driver.rst b/Documentation/driver-api/cxl/linux/cxl-driver.rst
> > index dd6dd17dc536..a61a20c8bb6d 100644
> > --- a/Documentation/driver-api/cxl/linux/cxl-driver.rst
> > +++ b/Documentation/driver-api/cxl/linux/cxl-driver.rst
> > @@ -602,6 +602,11 @@ derived from their upstream port connections. In `Cross-Link First` interleave
> > configurations, the :code:`interleave_granularity` of a decoder is equal to
> > :code:`parent_interleave_granularity * parent_interleave_ways`.
> >
> > +When the region granularity is finer than the granularity of an interleaving
> > +root decoder, the relation inverts: the :code:`interleave_granularity` of a
> > +decoder is equal to :code:`parent_interleave_granularity / interleave_ways`.
> > +See `Mixed Granularity`_.
>
> This introduces another unnecessary restriction: granularity must
> decrease from root down to the EP. E.g. The following config would not
> be supported:
Yes, those configurations are intentionally outside the supported subset.
Refer to the cover-letter requirements discussion.
>
> a)
>
> Level Ways Granularity
> ----- ---- -----------
> Root 2 4096
> Host bridge 2 1024
> Switch 2 2048
> Endpoint 8 1024
>
> Or:
>
> b)
>
> Level Ways Granularity
> ----- ---- -----------
> Root 3 4096
> Host bridge 2 1024
> Switch 2 2048
> Endpoint 12 1024
>
> In selector bits this is for the variants a/b:
>
> root: bits 12/-
> hb: bits 10
> switch: bits 11
>
> So the actual requirement is that the combined mask is consecutive.
> The lowest bit marks the EP granularity. The bit weight marks the
> 2-factor of the EP (total) ways.
Yes, that's the architectural requirement for the selector mask. The
restriction here is on the subset of those configs supported by Linux.
>
> This series should implement all cases as at some point the issue will
> pop up again.
Implementing all cases is not the goal of this set, and I don't think
that is just me being gunshy. We add support to the kernel as use cases
appear.
>
> > +
> > At Endpoint
> > ~~~~~~~~~~~
> > `Endpoint Decoders` are programmed similar to Host Bridge and Switch decoders,
> > @@ -619,6 +624,154 @@ from HPA to DPA. This is why they must be aware of the entire interleave set.
> > Linux does not support unbalanced interleave configurations. As a result, all
> > endpoints in an interleave set must have the same ways and granularity.
> >
> > +Mixed Granularity
> > +~~~~~~~~~~~~~~~~~
> > +Linux has required that a region's :code:`interleave_granularity` equal the
> > +:code:`interleave_granularity` of its interleaving root decoder. A
>
> This statement comments on the current kernel implementation which
> this series aims to change. It will be misleading to document that
> here. I rather would not document an implementation state that will be
> changed.
Agree. I'll remove the historical statement and describe the resulting
support only.
>
> > +*mixed-granularity* region lifts that restriction: the region granularity may
>
> This term is not a spec definition. It just describes a special case
> that may happen. I would avoid introducing it here. If the granularity
> implementation ware generic, there would be no need to describe that
> special case. So better describe the issue of current implementation
> and how it could be fixed.
Agree it's not CXL terminology, but it's useful terminology for the Linux
feature/subset this series implements
>
> > +be finer than the root's, so the interleave refines from the root toward the
> > +endpoints. The mix is between the root and the region, not granularity varying
> > +arbitrarily down the hierarchy. Linux requires granularity to change
> > +monotonically, either coarsening or refining, across the decoders that route
>
> Why that? It is not a requirement, but a limitation.
Yes, it is a Linux limitation. It is also a requirement for a configuration to be
supported by Linux, which is what this section is documenting.
>
> > +the request: the interleaving root, the host bridge, and any switches. The
> > +endpoint decoder is not part of that walk, it carries the region ways and
> > +granularity so that it can translate, as `At Endpoint`_ describes above. A
> > +mixed-granularity region requires an interleaving root decoder.
>
> Again, this is current. But it does not need to be documented here, as
> it is changed later.
I'll need to inspect what truly becomes obsolete. If this describes an
intermediate state superseded by the series, I'll trim it. If parts describe
the final Linux model, I'll keep those.
>
> > +
> > +Every decoder advances one target every multiple of its own granularity, and
> > +the decoders below it subdivide the address range their parent assigns to a
> > +single target. The `Cross-Link First` example above shows the other ordering,
> > +where the region granularity equals the granularity of the root decoder and
> > +granularity coarsens toward the endpoints. A mixed-granularity region refines
> > +instead, reaching the region granularity at the innermost interleaving decoder.
> > +
> > +Refining is not a preference. Once the region granularity is finer than the
> > +root's, Linux gives each level below the root a single granularity, its
> > +parent's divided by its own ways, so the ordering follows from the region and
> > +root settings rather than being chosen. The CXL Specification permits other
> > +orderings and does not require a monotonic one, see `Mod3 Interleave
> > +Configurations`_. Reaching a fine granularity as early as possible remains
> > +available through `Cross-Link First`, which this does not change.
> > +
> > +For an 8-way mixed-granularity region at 1024 below a 2-way interleaving root
> > +decoder at 4096, where each host bridge interleaves across two root ports and
> > +each root port hosts a 2-way switch, Linux programs::
> > +
> > + Level Ways Granularity
> > + ----- ---- -----------
> > + Root 2 4096
> > + Host bridge 2 2048
> > + Switch 2 1024
> > + Endpoint 8 1024
> > +
> > +A *region position* is the index of a region-granularity chunk within one full
> > +pass of the region interleave, and each endpoint of the region backs one
> > +position. Each decoder contributes to an endpoint's region position in
> > +proportion to its granularity::
> > +
> > + position += target_position *
> > + decoder_granularity / region_granularity
>
> The position is used to decode the DPA to SPA, and that is the actual
> definition of the position: The position bit mask are stripped off bye
> the endpoint to determine the DPA. Vice versa, it can be used to
> calculate the SPA by inserting the position bits (mask size depends on
> ways) at the granularity boundary into the DPA and adding the HPA
> offset.
>
> The above definition is just misleading and only reflects this special
> case and implementation.
I think I get you point, conceptually. I can define what position means
first using address translation. Then explain how this implementation
calcs it - if that needs documenting. I'll take a deeper look at this.
>
> It is also odd to use division when bit positions can be used. Same
> with the multiplier below.
Same result. For me, division better communicates the relationship being
documented.
>
> > +
> > +That multiplier is the *weight* of a level: how many region positions pass
> > +between successive advances of its target index. The root above selects a host
> > +bridge every 4096 bytes, so it advances one target every four region positions,
> > +weight 4, while the switch advances one target every position, weight 1.
> > +
> > +The two orderings weight the levels in opposite directions. Refining, a
> > +level's weight is the product of the ways of the interleaving levels below it,
> > +so the innermost interleaving level has weight 1. Coarsening, as in
> > +`Cross-Link First`, it is the product of the ways of the levels above it, so
> > +the root has weight 1 and advances one target per region position.
> > +
> > +The ways and granularity of a mixed-granularity region must describe the same
> > +interleave *span* as the root decoder, the address range in which the
> > +interleave pattern completes one full pass::
> > +
> > + root_ways * root_granularity == region_ways * region_granularity
>
> How is region_granularity defined?
I'll addxregion_granularity definition before this equation.
>
> In fact this describes, that the root selector bits must be the upper
> bits in the selector mask. Even that is an unnecessary limitation to
> implement support of "mixed" granularities.
Another scope issue.
>
> > +
> > +A region that does not describe that span either leaves part of the range
> > +unclaimed or reaches beyond it, and Linux rejects it. A same-granularity
> > +region below a power-of-two root decoder is exempt: it may be wider than the
> > +root and repeat the root targets, so its span is a multiple of the root's.
> > +
> > +Mod3 Interleave Configurations
> > +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
> > +A 3-way, 6-way, or 12-way interleave, known as a Mod3 interleave, selects its
> > +target with a factor-of-three selection rather than from binary interleave
> > +selector bits alone. CXL 4.0 Section 9.13.1.1 defines the 3-way selection as
> > +the address above the decoder granularity taken modulo 3. A 6-way selection
> > +claims one binary HPA bit at the decoder granularity and takes the modulo 3 of
> > +the address above that bit, and a 12-way selection claims two.
> > +
> > +The factor-of-three selection is not an additional binary selector bit, so a
> > +Mod3 interleave distributes across the hierarchy as::
> > +
> > + 3 = 3
> > + 6 = 3 * 2
> > + 12 = 3 * 4
> > +
> > +The cross-host bridge selection carries the factor of three, and the remaining
> > +x2 or x4 is binary interleave selection below it. A 6-way region at IGB across
> > +three host bridges is therefore::
> > +
> > + Device-level region: 6-way @ IGB
> > + Cross-host bridge: 3-way @ 2*IGB
> > + Below the root: 2-way @ IGB
> > +
> > +where both levels describe the same interleave span::
> > +
> > + 3 * (2 * IGB) == 6 * IGB
> > +
> > +The CXL Specification defines the legal Mod3 compositions and is normative.
> > +CXL 4.0 Section 9.13.1.1, "Legal Interleaving Configurations: 12-way, 6-way,
> > +and 3-way", Tables 9-6, 9-7, and 9-8 list them for a 12-way, 6-way, and 3-way
> > +device-level interleave at IGB. Those tables are summarized below, annotated
> > +with the subset Linux supports.
> > +
> > +CXL 4.0 Table 9-8, 3-way device-level interleave at IGB::
> > +
> > + Row Cross-host bridge Host bridge Switch Linux
> > + --- ----------------- ----------- ------ -----
> > + 1 3-way @ IGB none none supported
> > +
> > +CXL 4.0 Table 9-7, 6-way device-level interleave at IGB::
> > +
> > + Row Cross-host bridge Host bridge Switch Linux
> > + --- ----------------- ----------- ------ -----
> > + 1 6-way @ IGB none none supported
> > + 2 3-way @ 2*IGB 2-way @ IGB none supported
> > + 3 3-way @ 2*IGB none 2-way @ IGB supported
> > +
> > +CXL 4.0 Table 9-6, 12-way device-level interleave at IGB::
> > +
> > + Row Cross-host bridge Host bridge Switch Linux
> > + --- ----------------- ----------- ------ -----
> > + 1 12-way @ IGB none none supported
> > + 2 6-way @ 2*IGB 2-way @ IGB none supported
> > + 3 6-way @ 2*IGB none 2-way @ IGB supported
> > + 4 3-way @ 4*IGB 4-way @ IGB none supported
> > + 5 3-way @ 4*IGB none 4-way @ IGB supported
> > + 6 3-way @ 4*IGB 2-way @ IGB 2-way @ 2*IGB unsupported
> > + 7 3-way @ 4*IGB 2-way @ 2*IGB 2-way @ IGB supported
> > +
> > +Table 9-6 row 6 is legal per the CXL Specification but is unsupported by Linux.
>
> Exactly, unsupported config as described above.
>
> I am good with the 3-way section above, though it just reflects, what
> the spec describes and I wouldn't be that detailed here.
Will trim that.
>
> > +Walking it from the root toward the endpoints, granularity goes::
> > +
> > + 4*IGB -> IGB -> 2*IGB
> > +
> > +which refines and then coarsens. Row 7 interleaves the same 12 endpoints at
> > +the same granularity with those two levels exchanged::
> > +
> > + 4*IGB -> 2*IGB -> IGB
> > +
> > +which is monotonic. Linux programs row 7 for a user region and assembles an
> > +auto region whose decoders are programmed that way. An auto region matching
> > +row 6 is not assembled.
> > +
> > +Leaving row 6 unsupported does not prevent a 12-way device-level interleave.
> > +The specification defines six other legal compositions, all monotonic and
> > +supported by Linux.
>
> Again, we should (and can) avoid this limitation.
Scope again. Why? Use case? Let's keep that discussion in cover letter
thread.
>
> Suppose the following: This is a support cxl driver config:
>
> 4-way: 2-way @ IGB 2-way @ 2*IGB
>
> Root ports and switches are configured correctly. Now, you just want
> to enable 3-way which could be done without reprogramming the HDM
> decoder (accept for the endpoints). Those configs become unsupported
> and granularities need to be reprogrammed. This is very unexpected.
>
> Changing implementation to use selector bit logic will solve this
> limitation and make code (and doc) much easier.
v3 had selector-derived machinery. The simplification in v4 and v5 came
after narrowing the supported set and removing it.
>
> Thanks,
>
> -Robert
>
> > +
> > Example Configurations
> > ======================
> > .. toctree::
> > --
> > 2.37.3
> >
next prev parent reply other threads:[~2026-09-18 20:48 UTC|newest]
Thread overview: 32+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-03 23:23 [PATCH v5 0/7] cxl: Support mixed-granularity region interleaves Alison Schofield
2026-09-03 23:23 ` [PATCH v5 1/7] Documentation/cxl: Describe mixed-granularity regions Alison Schofield
2026-09-07 23:23 ` Jonathan Cameron
2026-09-16 11:51 ` Robert Richter
2026-09-18 20:48 ` Alison Schofield [this message]
2026-09-03 23:23 ` [PATCH v5 2/7] cxl/region: Warn on user region position mismatch Alison Schofield
2026-09-16 12:05 ` Robert Richter
2026-09-18 21:18 ` Alison Schofield
2026-09-25 0:30 ` Alison Schofield
2026-09-03 23:23 ` [PATCH v5 3/7] cxl/region: Generalize endpoint position mapping Alison Schofield
2026-09-07 23:42 ` Jonathan Cameron
2026-09-18 23:01 ` Alison Schofield
2026-09-16 14:15 ` Robert Richter
2026-09-18 23:20 ` Alison Schofield
2026-09-03 23:23 ` [PATCH v5 4/7] cxl/region: Name the interleave locals in cxl_port_setup_targets() Alison Schofield
2026-09-07 23:50 ` [PATCH v5 4/7] cxo_ol/region: " Jonathan Cameron
2026-09-19 0:42 ` Alison Schofield
2026-09-16 14:34 ` [PATCH v5 4/7] cxl/region: " Robert Richter
2026-09-16 17:45 ` Robert Richter
2026-09-19 0:58 ` Alison Schofield
2026-09-19 0:55 ` Alison Schofield
2026-09-03 23:23 ` [PATCH v5 5/7] cxl/region: Support mixed-granularity auto regions Alison Schofield
2026-09-08 0:02 ` Jonathan Cameron
2026-09-16 18:03 ` Robert Richter
2026-09-03 23:23 ` [PATCH v5 6/7] cxl/region: Support mixed-granularity user created regions Alison Schofield
2026-09-08 0:03 ` Jonathan Cameron
2026-09-16 18:13 ` Robert Richter
2026-09-03 23:23 ` [PATCH v5 7/7] cxl/test: Add a topology to test mixed-granularity regions Alison Schofield
2026-09-08 0:10 ` Jonathan Cameron
2026-09-16 18:19 ` [PATCH v5 0/7] cxl: Support mixed-granularity region interleaves Robert Richter
2026-09-18 19:49 ` Alison Schofield
2026-09-29 22:48 ` Dave Jiang
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aq2jjiXfDKo0SwwQ@aschofie-mobl2.lan \
--to=alison.schofield@intel.com \
--cc=dave.jiang@intel.com \
--cc=dave@stgolabs.net \
--cc=iweiny@kernel.org \
--cc=jic23@kernel.org \
--cc=linux-cxl@vger.kernel.org \
--cc=ming.li@zohomail.com \
--cc=rrichter@amd.com \
--cc=vishal.l.verma@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox