Netdev List
 help / color / mirror / Atom feed
* [RFC net-next 00/12] net: add the PON subsystem
@ 2026-10-08 14:32 John Crispin
  2026-10-08 14:32 ` [RFC net-next 01/12] net: pon: add the netlink specification of the pon family John Crispin
                   ` (13 more replies)
  0 siblings, 14 replies; 18+ messages in thread
From: John Crispin @ 2026-10-08 14:32 UTC (permalink / raw)
  To: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni
  Cc: Donald Hunter, Simon Horman, netdev, linux-kernel, Andrew Lunn,
	Christian Marangi, Jonathan Corbet, Shuah Khan, Randy Dunlap,
	linux-doc

This series adds net/pon, a subsystem for the ONU side of an ITU-T
passive optical network. It follows the thread "[RFC] net: towards a
generic PON framework" [1], which converged on a wiphy-like PON object
that is not a netdev, PLOAM in the MAC driver, service traffic on
ordinary netdevs and OMCI in userspace. This is that design as working
code.

The model is cfg80211 with a fullmac driver:

- struct pon_dev is the PON object. It is not a netdev. Userspace
  addresses it by a device id, as nl80211 addresses a wiphy.
- The MAC driver runs PLOAM activation and the key hierarchy. It
  reports the activation state through pon_dev_state_report(). The
  core validates the edge against ITU-T G.9807.1 Table C.12.4,
  publishes it and sets the carrier from it.
- The core owns the T-CONTs, the GEM ports and the upstream classifier
  rules (the 802.1p mapper of the thread). It stores them and dumps
  them. The driver offloads them.
- The core lends the driver a serialized context: an ordered workqueue
  per instance and the instance lock that the netlink handlers take.
- Service traffic uses a data netdev plus GEM netdevs (rtnetlink link
  kind "gem") where a GEM port needs a bridge port of its own. Frames
  ride the rings of an ethernet controller, the conduit, with the per
  frame information passed as arguments.
- OMCI travels over the pon genl family the way nl80211 carries
  management frames. One socket registers for the OMCI PDUs of a
  device, receives them as notifications and sends with omci-tx. No
  OMCI netdev and no new address family.
- Control is one generic netlink family with a YAML spec. The uapi
  header and the policy are generated.

This posting is the minimal core. The counters, alarms, events, the
ETS scheduler offload, a PON mode for the generic PHY framework, the PON
PCS object and the flow offload are further commits on top of it that
come in later series.

The scope is XGS-PON. The mode enum keeps gpon and xg-pon so that the
uapi order stays clean for later drivers. The core refuses a mode that
the device does not list.

The first driver is for the Airoha AN7581 PON MAC, with its PON PHY, an
optics driver and the conduit in airoha_eth. It is not part of this
posting, because it depends on the Airoha PCS series that is not
upstream yet. pon-tool [2] consumes the family through the libpon
library. It has a verb for every attribute of the family and runs in
OpenWrt today. The stack reaches O5, passes traffic and recovers from
fiber pulls against a production OLT with HGU, SFU and IP host services
at the same time.

[1] https://lore.kernel.org/netdev/20260926174601.1675-1-yhyxwgy@gmail.com/
[2] https://github.com/blogic/pon-tool

John Crispin (12):
  net: pon: add the netlink specification of the pon family
  net: add a PON device pointer to struct net_device
  net: pon: add the PLOAM vocabulary and message codec
  net: pon: add the device state and the lent context
  net: pon: add the conduit contract
  net: pon: add the GEM network devices
  net: pon: add the OMCI channel
  net: pon: add netlink support
  net: pon: add the device registration
  net: pon: build the PON subsystem
  Documentation: networking: describe the PON subsystem
  MAINTAINERS: add the PON subsystem

 Documentation/netlink/specs/pon.yaml     |  604 ++++++++
 Documentation/netlink/specs/rt-link.yaml |   14 +
 Documentation/networking/index.rst       |    1 +
 Documentation/networking/pon.rst         |  302 ++++
 MAINTAINERS                              |   11 +
 include/linux/netdevice.h                |    6 +
 include/net/pon.h                        |   14 +
 include/net/pon/functions.h              |  129 ++
 include/net/pon/ploam.h                  |  266 ++++
 include/net/pon/ploam_msg.h              |  306 ++++
 include/net/pon/types.h                  |  443 ++++++
 include/uapi/linux/if_link.h             |   10 +
 include/uapi/linux/pon.h                 |  166 +++
 net/Kconfig                              |    1 +
 net/Makefile                             |    1 +
 net/pon/Kconfig                          |   17 +
 net/pon/Makefile                         |    6 +
 net/pon/pon-nl-gen.c                     |  286 ++++
 net/pon/pon-nl-gen.h                     |   45 +
 net/pon/pon.h                            |  162 +++
 net/pon/pon_conduit.c                    |  614 ++++++++
 net/pon/pon_gem.c                        |  537 +++++++
 net/pon/pon_main.c                       |  454 ++++++
 net/pon/pon_nl.c                         | 1671 ++++++++++++++++++++++
 net/pon/pon_omci.c                       |  326 +++++
 net/pon/pon_ploam_msg.c                  |  312 ++++
 net/pon/pon_state.c                      |  408 ++++++
 net/pon/pon_work.c                       |  106 ++
 tools/net/ynl/Makefile.deps              |    1 +
 29 files changed, 7219 insertions(+)
 create mode 100644 Documentation/netlink/specs/pon.yaml
 create mode 100644 Documentation/networking/pon.rst
 create mode 100644 include/net/pon.h
 create mode 100644 include/net/pon/functions.h
 create mode 100644 include/net/pon/ploam.h
 create mode 100644 include/net/pon/ploam_msg.h
 create mode 100644 include/net/pon/types.h
 create mode 100644 include/uapi/linux/pon.h
 create mode 100644 net/pon/Kconfig
 create mode 100644 net/pon/Makefile
 create mode 100644 net/pon/pon-nl-gen.c
 create mode 100644 net/pon/pon-nl-gen.h
 create mode 100644 net/pon/pon.h
 create mode 100644 net/pon/pon_conduit.c
 create mode 100644 net/pon/pon_gem.c
 create mode 100644 net/pon/pon_main.c
 create mode 100644 net/pon/pon_nl.c
 create mode 100644 net/pon/pon_omci.c
 create mode 100644 net/pon/pon_ploam_msg.c
 create mode 100644 net/pon/pon_state.c
 create mode 100644 net/pon/pon_work.c


base-commit: 8df0638138d3e0344fd1fb36cf2d1ca1cf5028f0
-- 
2.34.1


^ permalink raw reply	[flat|nested] 18+ messages in thread

* [RFC net-next 01/12] net: pon: add the netlink specification of the pon family
  2026-10-08 14:32 [RFC net-next 00/12] net: add the PON subsystem John Crispin
@ 2026-10-08 14:32 ` John Crispin
  2026-10-08 14:32 ` [RFC net-next 02/12] net: add a PON device pointer to struct net_device John Crispin
                   ` (12 subsequent siblings)
  13 siblings, 0 replies; 18+ messages in thread
From: John Crispin @ 2026-10-08 14:32 UTC (permalink / raw)
  To: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni
  Cc: Donald Hunter, Simon Horman, netdev, linux-kernel, Andrew Lunn,
	Christian Marangi

The pon generic netlink family configures the ONU side of an ITU-T
passive optical network. The kernel owns activation (which the MAC
driver runs) and the datapath objects. Userspace owns OMCI (the
management plane) and expresses what the OLT asks for through this
family.

The family carries the identity and the activation state of a device,
the T-CONTs, the GEM ports, the upstream classifier rules and the OMCI
channel. Notifications are levels, so a listener that loses one reads
the state again. A write operation answers with the netlink ACK only.
No key material crosses the family.

The mode enum lists gpon, xg-pon and xgs-pon. No driver runs gpon or
xg-pon yet. The enum keeps them so that the uapi order stays clean for
the drivers that add them later. The core refuses them through the mode
capability of the device: dev-set refuses a mode the device does not
list, with an extack message and a line in the kernel log.

pon.yaml is the source. The uapi header is generated from it with
ynl-gen. The Makefile.deps of tools/net/ynl gets a line for the family,
so that the header of the tree wins over an older copy that a
distribution ships. The kernel side follows with the subsystem.

Assisted-by: LLM
Signed-off-by: John Crispin <john@phrozen.org>
---
 Documentation/netlink/specs/pon.yaml | 604 +++++++++++++++++++++++++++
 include/uapi/linux/pon.h             | 166 ++++++++
 tools/net/ynl/Makefile.deps          |   1 +
 3 files changed, 771 insertions(+)
 create mode 100644 Documentation/netlink/specs/pon.yaml
 create mode 100644 include/uapi/linux/pon.h

diff --git a/Documentation/netlink/specs/pon.yaml b/Documentation/netlink/specs/pon.yaml
new file mode 100644
index 000000000000..45780c179613
--- /dev/null
+++ b/Documentation/netlink/specs/pon.yaml
@@ -0,0 +1,604 @@
+# SPDX-License-Identifier: ((GPL-2.0 WITH Linux-syscall-note) OR BSD-3-Clause)
+---
+name: pon
+
+doc:
+  Passive Optical Network device configuration over generic netlink.
+
+  A PON device represents the ONU-side MAC of a GPON or XG(S)-PON link. The
+  kernel owns activation (the PLOAM state machine, ranging and the key
+  hierarchy run in hardware or in the driver, driven by MAC interrupts) and
+  the datapath objects (T-CONTs and GEM ports). Userspace owns the management
+  plane, OMCI, whose PDUs this family carries between the OMCC and the one
+  socket that registered for them. This family also carries identity,
+  datapath object configuration, state notifications and counters.
+
+  An operation is a set when it can change an object that already exists
+  and a new when it can only create one. A tcont-set rebinds a T-CONT, while
+  a gem-new refuses a GEM port that exists with other attributes.
+
+definitions:
+  -
+    type: enum
+    name: mode
+    doc: PON operating mode.
+    entries: [gpon, xg-pon, xgs-pon]
+  -
+    type: enum
+    name: ploam-state
+    doc: |
+      ITU-T G.9807.1 / G.984.3 ONU activation state. O5 is the only state in
+      which the OMCC carries traffic. G.9807.1 has one Serial Number state,
+      O2-3 (Table C.12.1), reported as o2. Only a G.984.3 device reports o3.
+    entries: [unknown, o1, o2, o3, o4, o5, o6, o7]
+  -
+    type: enum
+    name: gem-dir
+    doc: |
+      GEM port direction. The values are those of the direction attribute of
+      the GEM port network CTP managed entity of ITU-T G.988: UNI-to-ANI (1),
+      ANI-to-UNI (2) or bidirectional (3). An OMCI stack passes the value
+      through.
+    entries:
+      -
+        name: upstream
+        value: 1
+      -
+        name: downstream
+        value: 2
+      -
+        name: bidir
+        value: 3
+  -
+    type: enum
+    name: gem-key-ring
+    doc: |
+      Whether a GEM port is encrypted and with which keys, as the
+      encryption key ring attribute of the GEM port network CTP managed
+      entity of ITU-T G.988 clause 9.2.3 defines it.
+    entries:
+      -
+        name: none
+        doc: No encryption. Upstream is sent with key index 0.
+      -
+        name: unicast
+        doc: |
+          Unicast encryption in both directions, with the keys the ONU
+          generates.
+      -
+        name: broadcast
+        doc: |
+          Broadcast encryption, with the keys the OLT distributes over the
+          OMCI.
+      -
+        name: unicast-downstream
+        doc: |
+          Unicast encryption of the downstream only, with the keys the ONU
+          generates.
+  -
+    type: enum
+    name: gem-map-tag
+    doc: VLAN tag state an upstream classifier rule matches.
+    entries: [untagged, tagged]
+
+attribute-sets:
+  -
+    name: dev
+    attributes:
+      -
+        name: id
+        doc: PON device ID.
+        type: u32
+        checks:
+          min: 1
+      -
+        name: ifindex
+        doc: ifindex of the PON data network device (the PON MAC port).
+        type: u32
+      -
+        name: mode
+        doc: Active PON mode.
+        type: u32
+        enum: mode
+      -
+        name: modes-cap
+        doc: Bitmask of the modes the device supports.
+        type: u32
+        enum: mode
+        enum-as-flags: true
+      -
+        name: serial
+        doc: |
+          ONU serial number, 8 bytes: the 4 ASCII characters of the Vendor_ID
+          of ITU-T G.9807.1 clause C.11.2.6.1, then the 4 byte VSSN of clause
+          C.11.2.6.2. Eight 0x00 bytes address every ONU (Table C.11.23A) and
+          are refused.
+        type: binary
+        checks:
+          exact-len: 8
+      -
+        name: registration-id
+        doc: |
+          Registration_ID of the Registration message of ITU-T G.9807.1
+          (Table C.11.25), 1 to 36 bytes. An empty value is refused. The
+          kernel pads a shorter value with 0x00 bytes to 36, as the note of
+          the table recommends. Write only: it is a credential.
+        type: binary
+        checks:
+          max-len: 36
+      -
+        name: ploam-state
+        doc: Current activation state.
+        type: u32
+        enum: ploam-state
+      -
+        name: max-tconts
+        doc: Number of T-CONTs the device supports.
+        type: u32
+      -
+        name: max-gems
+        doc: |
+          Number of GEM ports gem-new can create, not counting the GEM port
+          of the OMCC. gem-new answers -ENOSPC beyond it.
+        type: u32
+      -
+        name: enable
+        doc: |
+          1 starts the ONU upstream link, 0 stops it. In a reply, whether the
+          link was last started.
+        type: u8
+        checks:
+          max: 1
+  -
+    name: tcont
+    attributes:
+      -
+        name: dev-id
+        doc: PON device ID.
+        type: u32
+        checks:
+          min: 1
+      -
+        name: index
+        doc: |
+          T-CONT index in the device, 0 based and below the max-tconts of
+          the device.
+        type: u32
+        checks:
+          max: 65534
+      -
+        name: alloc-id
+        doc: |
+          Alloc-ID bound to the T-CONT, as assigned by the OLT. ITU-T
+          G.9807.1 clause C.6.1.5.7 bounds it at 16383.
+        type: u32
+        checks:
+          max: 16383
+  -
+    name: gem
+    attributes:
+      -
+        name: dev-id
+        doc: PON device ID.
+        type: u32
+        checks:
+          min: 1
+      -
+        name: id
+        doc: |
+          GEM port ID, as the OLT assigns it over the OMCI. ITU-T G.9807.1
+          Table C.6.6 reserves 0 to 1020 for the default Port-ID (equal
+          to the ONU-ID, OMCC only) and 65535 for the idle Port-ID, so an
+          assigned GEM port is 1021 to 65534.
+        type: u32
+        checks:
+          min: 1021
+          max: 65534
+      -
+        name: dir
+        doc: Direction.
+        type: u32
+        enum: gem-dir
+      -
+        name: tcont-index
+        doc: |
+          T-CONT the upstream half rides. Required for an upstream or
+          bidirectional GEM port. A downstream GEM port rides no T-CONT and
+          is refused with this attribute.
+        type: u32
+        checks:
+          max: 65534
+      -
+        name: key-ring
+        doc: |
+          Encryption key ring of the GEM port. Absent means none. The
+          kernel refuses broadcast, since it takes no broadcast key.
+        type: u32
+        enum: gem-key-ring
+  -
+    name: gem-map
+    attributes:
+      -
+        name: dev-id
+        doc: PON device ID.
+        type: u32
+        checks:
+          min: 1
+      -
+        name: gem-id
+        doc: GEM port the matching frames map to, 1021 to 65534 as for gem id.
+        type: u32
+        checks:
+          min: 1021
+          max: 65534
+      -
+        name: tag
+        doc: Match the frame VLAN tag state. Absent matches either.
+        type: u32
+        enum: gem-map-tag
+      -
+        name: vid
+        doc: Match the outer VLAN ID. Absent matches any.
+        type: u32
+        checks:
+          max: 4094
+      -
+        name: pbit
+        doc: Match the outer VLAN priority. Absent matches any.
+        type: u32
+        checks:
+          max: 7
+      -
+        name: dscp
+        doc: Match the IP DSCP field. Absent matches any.
+        type: u32
+        checks:
+          max: 63
+  -
+    name: omci
+    attributes:
+      -
+        name: dev-id
+        doc: PON device ID.
+        type: u32
+        checks:
+          min: 1
+      -
+        name: pdu
+        doc: |
+          One OMCI PDU, from the transaction correlation id to the end of the
+          message contents: 44 bytes for the baseline format, the 10 byte
+          header and the contents for the extended format of ITU-T G.988
+          annex B. It carries no integrity field in either direction. The
+          MAC adds the MIC upstream and checks it downstream. Upstream the
+          device identifier selects the format and the length must match
+          it: 44 bytes for 0x0a, 10 bytes plus the message contents length
+          of the header for 0x0b. Any other PDU is refused.
+        type: binary
+        checks:
+          max-len: 1976
+
+operations:
+  list:
+    -
+      name: dev-get
+      doc: Get or dump PON devices.
+      attribute-set: dev
+      do:
+        request:
+          attributes:
+            - id
+        reply: &dev-all
+          attributes:
+            - id
+            - ifindex
+            - mode
+            - modes-cap
+            - serial
+            - ploam-state
+            - max-tconts
+            - max-gems
+            - enable
+        pre: pon-device-get-locked
+        post: pon-device-unlock
+      dump:
+        reply: *dev-all
+    -
+      name: dev-set
+      doc: |
+        Set the identity of a PON device, start or stop its upstream link, or
+        any combination. While the link is enabled, a mode or a serial number
+        other than the one the device holds is refused with EBUSY.
+        The settings are applied before the link moves.
+      attribute-set: dev
+      flags: [admin-perm]
+      do:
+        request:
+          attributes:
+            - id
+            - mode
+            - serial
+            - registration-id
+            - enable
+        pre: pon-device-get-locked
+        post: pon-device-unlock
+    -
+      name: dev-add-ntf
+      doc: Notification about a PON device appearing.
+      notify: dev-get
+      mcgrp: mgmt
+    -
+      name: dev-del-ntf
+      doc: |
+        Notification about a PON device disappearing. Every T-CONT, GEM port
+        and classifier rule of the device goes with it, without a
+        notification of its own.
+      notify: dev-get
+      mcgrp: mgmt
+    -
+      name: dev-change-ntf
+      doc: Notification about a PON device configuration change.
+      notify: dev-get
+      mcgrp: mgmt
+    -
+      name: ploam-ntf
+      doc: |
+        The activation state changed. Every transition is reported, so the
+        O5 edge, which is the OMCC becoming usable or unusable, is visible.
+      attribute-set: dev
+      mcgrp: state
+      event:
+        attributes:
+          - id
+          - ploam-state
+
+    -
+      name: tcont-get
+      doc: |
+        Get or dump T-CONTs. A dump with a device ID lists the T-CONTs of
+        that device only.
+      attribute-set: tcont
+      do:
+        request:
+          attributes:
+            - dev-id
+            - index
+        reply: &tcont-all
+          attributes:
+            - dev-id
+            - index
+            - alloc-id
+        pre: pon-device-get-locked
+        post: pon-device-unlock
+      dump:
+        request:
+          attributes:
+            - dev-id
+        reply: *tcont-all
+    -
+      name: tcont-set
+      doc: |
+        Bind an alloc-id to a T-CONT, or rebind it. A rebind moves the GEM
+        ports of the T-CONT too. If the driver refuses one, the old alloc-id
+        stays on the T-CONT and on its GEM ports and the request fails.
+      attribute-set: tcont
+      flags: [admin-perm]
+      do:
+        request:
+          attributes:
+            - dev-id
+            - index
+            - alloc-id
+        pre: pon-device-get-locked
+        post: pon-device-unlock
+    -
+      name: tcont-del
+      doc: |
+        Release a T-CONT. Fails with -EBUSY while a GEM port names the
+        T-CONT. Delete the GEM ports first.
+      attribute-set: tcont
+      flags: [admin-perm]
+      do:
+        request:
+          attributes:
+            - dev-id
+            - index
+        pre: pon-device-get-locked
+        post: pon-device-unlock
+    -
+      name: tcont-add-ntf
+      doc: A T-CONT object was created.
+      notify: tcont-get
+      mcgrp: mgmt
+    -
+      name: tcont-change-ntf
+      doc: A T-CONT object changed.
+      notify: tcont-get
+      mcgrp: mgmt
+    -
+      name: tcont-del-ntf
+      doc: A T-CONT object was removed.
+      notify: tcont-get
+      mcgrp: mgmt
+
+    -
+      name: gem-get
+      doc: |
+        Get or dump GEM ports. A dump with a device ID lists the GEM ports
+        of that device only.
+      attribute-set: gem
+      do:
+        request:
+          attributes:
+            - dev-id
+            - id
+        reply: &gem-all
+          attributes:
+            - dev-id
+            - id
+            - dir
+            - tcont-index
+            - key-ring
+        pre: pon-device-get-locked
+        post: pon-device-unlock
+      dump:
+        request:
+          attributes:
+            - dev-id
+        reply: *gem-all
+    -
+      name: gem-new
+      doc: |
+        Create a GEM port. A GEM port that exists with the same attributes
+        is accepted. One that exists with other attributes is refused.
+      attribute-set: gem
+      flags: [admin-perm]
+      do:
+        request:
+          attributes:
+            - dev-id
+            - id
+            - dir
+            - tcont-index
+            - key-ring
+        pre: pon-device-get-locked
+        post: pon-device-unlock
+    -
+      name: gem-del
+      doc: Destroy a GEM port.
+      attribute-set: gem
+      flags: [admin-perm]
+      do:
+        request:
+          attributes:
+            - dev-id
+            - id
+        pre: pon-device-get-locked
+        post: pon-device-unlock
+    -
+      name: gem-add-ntf
+      doc: A GEM port object was created.
+      notify: gem-get
+      mcgrp: mgmt
+    -
+      name: gem-del-ntf
+      doc: A GEM port object was removed.
+      notify: gem-get
+      mcgrp: mgmt
+
+    -
+      name: gem-map-get
+      doc: |
+        Dump the upstream classifier rules. The core keeps a copy of each
+        rule that the driver took. A dump with a device ID lists the rules of
+        that device only.
+      attribute-set: gem-map
+      dump:
+        request:
+          attributes:
+            - dev-id
+        reply:
+          attributes:
+            - dev-id
+            - gem-id
+            - tag
+            - vid
+            - pbit
+            - dscp
+    -
+      name: gem-map-new
+      doc: |
+        Add one upstream classifier rule mapping frames onto a GEM port.
+        A rule matches on the attributes it carries. A more specific rule
+        wins over a less specific one.
+      attribute-set: gem-map
+      flags: [admin-perm]
+      do:
+        request:
+          attributes:
+            - dev-id
+            - gem-id
+            - tag
+            - vid
+            - pbit
+            - dscp
+        pre: pon-device-get-locked
+        post: pon-device-unlock
+    -
+      name: gem-map-del
+      doc: Remove one upstream classifier rule.
+      attribute-set: gem-map
+      flags: [admin-perm]
+      do:
+        request:
+          attributes:
+            - dev-id
+            - gem-id
+            - tag
+            - vid
+            - pbit
+            - dscp
+        pre: pon-device-get-locked
+        post: pon-device-unlock
+    -
+      name: gem-map-add-ntf
+      doc: An upstream classifier rule was created.
+      notify: gem-map-get
+      mcgrp: mgmt
+    -
+      name: gem-map-del-ntf
+      doc: An upstream classifier rule was removed.
+      notify: gem-map-get
+      mcgrp: mgmt
+
+
+    -
+      name: omci-register
+      doc: |
+        Receive the OMCI PDUs of one device on this socket, as omci-ntf.
+        One socket at a time owns the OMCI channel of a device. Another
+        socket is refused with EBUSY until the owner closes. A registration
+        ends only when its socket closes.
+      attribute-set: omci
+      flags: [admin-perm]
+      do:
+        request:
+          attributes:
+            - dev-id
+        pre: pon-device-get-locked
+        post: pon-device-unlock
+    -
+      name: omci-tx
+      doc: |
+        Send one OMCI PDU to the OLT. Only the socket that owns the OMCI
+        channel of the device may send and the MAC refuses a PDU while the
+        ONU is not in O5.
+      attribute-set: omci
+      flags: [admin-perm]
+      do:
+        request:
+          attributes:
+            - dev-id
+            - pdu
+        pre: pon-device-get-locked
+        post: pon-device-unlock
+    -
+      name: omci-ntf
+      doc: |
+        One OMCI PDU from the OLT, sent only to the socket that owns the
+        OMCI channel of the device. A PDU that finds no room on that socket
+        is lost. The OLT retries.
+      attribute-set: omci
+      event:
+        attributes:
+          - dev-id
+          - pdu
+
+mcast-groups:
+  list:
+    -
+      name: mgmt
+    -
+      name: state
+
+...
diff --git a/include/uapi/linux/pon.h b/include/uapi/linux/pon.h
new file mode 100644
index 000000000000..1fedadb44305
--- /dev/null
+++ b/include/uapi/linux/pon.h
@@ -0,0 +1,166 @@
+/* SPDX-License-Identifier: ((GPL-2.0 WITH Linux-syscall-note) OR BSD-3-Clause) */
+/* Do not edit directly, auto-generated from: */
+/*	Documentation/netlink/specs/pon.yaml */
+/* YNL-GEN uapi header */
+/* To regenerate run: tools/net/ynl/ynl-regen.sh */
+
+#ifndef _UAPI_LINUX_PON_H
+#define _UAPI_LINUX_PON_H
+
+#define PON_FAMILY_NAME		"pon"
+#define PON_FAMILY_VERSION	1
+
+/*
+ * PON operating mode.
+ */
+enum pon_mode {
+	PON_MODE_GPON,
+	PON_MODE_XG_PON,
+	PON_MODE_XGS_PON,
+};
+
+/*
+ * ITU-T G.9807.1 / G.984.3 ONU activation state. O5 is the only state in which
+ * the OMCC carries traffic. G.9807.1 has one Serial Number state, O2-3 (Table
+ * C.12.1), reported as o2. Only a G.984.3 device reports o3.
+ */
+enum pon_ploam_state {
+	PON_PLOAM_STATE_UNKNOWN,
+	PON_PLOAM_STATE_O1,
+	PON_PLOAM_STATE_O2,
+	PON_PLOAM_STATE_O3,
+	PON_PLOAM_STATE_O4,
+	PON_PLOAM_STATE_O5,
+	PON_PLOAM_STATE_O6,
+	PON_PLOAM_STATE_O7,
+};
+
+/*
+ * GEM port direction. The values are those of the direction attribute of the
+ * GEM port network CTP managed entity of ITU-T G.988: UNI-to-ANI (1),
+ * ANI-to-UNI (2) or bidirectional (3). An OMCI stack passes the value through.
+ */
+enum pon_gem_dir {
+	PON_GEM_DIR_UPSTREAM = 1,
+	PON_GEM_DIR_DOWNSTREAM,
+	PON_GEM_DIR_BIDIR,
+};
+
+/**
+ * enum pon_gem_key_ring - Whether a GEM port is encrypted and with which keys,
+ *   as the encryption key ring attribute of the GEM port network CTP managed
+ *   entity of ITU-T G.988 clause 9.2.3 defines it.
+ * @PON_GEM_KEY_RING_NONE: No encryption. Upstream is sent with key index 0.
+ * @PON_GEM_KEY_RING_UNICAST: Unicast encryption in both directions, with the
+ *   keys the ONU generates.
+ * @PON_GEM_KEY_RING_BROADCAST: Broadcast encryption, with the keys the OLT
+ *   distributes over the OMCI.
+ * @PON_GEM_KEY_RING_UNICAST_DOWNSTREAM: Unicast encryption of the downstream
+ *   only, with the keys the ONU generates.
+ */
+enum pon_gem_key_ring {
+	PON_GEM_KEY_RING_NONE,
+	PON_GEM_KEY_RING_UNICAST,
+	PON_GEM_KEY_RING_BROADCAST,
+	PON_GEM_KEY_RING_UNICAST_DOWNSTREAM,
+};
+
+/*
+ * VLAN tag state an upstream classifier rule matches.
+ */
+enum pon_gem_map_tag {
+	PON_GEM_MAP_TAG_UNTAGGED,
+	PON_GEM_MAP_TAG_TAGGED,
+};
+
+enum {
+	PON_A_DEV_ID = 1,
+	PON_A_DEV_IFINDEX,
+	PON_A_DEV_MODE,
+	PON_A_DEV_MODES_CAP,
+	PON_A_DEV_SERIAL,
+	PON_A_DEV_REGISTRATION_ID,
+	PON_A_DEV_PLOAM_STATE,
+	PON_A_DEV_MAX_TCONTS,
+	PON_A_DEV_MAX_GEMS,
+	PON_A_DEV_ENABLE,
+
+	__PON_A_DEV_MAX,
+	PON_A_DEV_MAX = (__PON_A_DEV_MAX - 1)
+};
+
+enum {
+	PON_A_TCONT_DEV_ID = 1,
+	PON_A_TCONT_INDEX,
+	PON_A_TCONT_ALLOC_ID,
+
+	__PON_A_TCONT_MAX,
+	PON_A_TCONT_MAX = (__PON_A_TCONT_MAX - 1)
+};
+
+enum {
+	PON_A_GEM_DEV_ID = 1,
+	PON_A_GEM_ID,
+	PON_A_GEM_DIR,
+	PON_A_GEM_TCONT_INDEX,
+	PON_A_GEM_KEY_RING,
+
+	__PON_A_GEM_MAX,
+	PON_A_GEM_MAX = (__PON_A_GEM_MAX - 1)
+};
+
+enum {
+	PON_A_GEM_MAP_DEV_ID = 1,
+	PON_A_GEM_MAP_GEM_ID,
+	PON_A_GEM_MAP_TAG,
+	PON_A_GEM_MAP_VID,
+	PON_A_GEM_MAP_PBIT,
+	PON_A_GEM_MAP_DSCP,
+
+	__PON_A_GEM_MAP_MAX,
+	PON_A_GEM_MAP_MAX = (__PON_A_GEM_MAP_MAX - 1)
+};
+
+enum {
+	PON_A_OMCI_DEV_ID = 1,
+	PON_A_OMCI_PDU,
+
+	__PON_A_OMCI_MAX,
+	PON_A_OMCI_MAX = (__PON_A_OMCI_MAX - 1)
+};
+
+enum {
+	PON_CMD_DEV_GET = 1,
+	PON_CMD_DEV_SET,
+	PON_CMD_DEV_ADD_NTF,
+	PON_CMD_DEV_DEL_NTF,
+	PON_CMD_DEV_CHANGE_NTF,
+	PON_CMD_PLOAM_NTF,
+	PON_CMD_TCONT_GET,
+	PON_CMD_TCONT_SET,
+	PON_CMD_TCONT_DEL,
+	PON_CMD_TCONT_ADD_NTF,
+	PON_CMD_TCONT_CHANGE_NTF,
+	PON_CMD_TCONT_DEL_NTF,
+	PON_CMD_GEM_GET,
+	PON_CMD_GEM_NEW,
+	PON_CMD_GEM_DEL,
+	PON_CMD_GEM_ADD_NTF,
+	PON_CMD_GEM_DEL_NTF,
+	PON_CMD_GEM_MAP_GET,
+	PON_CMD_GEM_MAP_NEW,
+	PON_CMD_GEM_MAP_DEL,
+	PON_CMD_GEM_MAP_ADD_NTF,
+	PON_CMD_GEM_MAP_DEL_NTF,
+	PON_CMD_OMCI_REGISTER,
+	PON_CMD_OMCI_TX,
+	PON_CMD_OMCI_NTF,
+
+	__PON_CMD_MAX,
+	PON_CMD_MAX = (__PON_CMD_MAX - 1)
+};
+
+#define PON_MCGRP_MGMT	"mgmt"
+#define PON_MCGRP_STATE	"state"
+
+#endif /* _UAPI_LINUX_PON_H */
diff --git a/tools/net/ynl/Makefile.deps b/tools/net/ynl/Makefile.deps
index 1e746e25e2bc..fc84485b2c8b 100644
--- a/tools/net/ynl/Makefile.deps
+++ b/tools/net/ynl/Makefile.deps
@@ -38,6 +38,7 @@ CFLAGS_ovs_datapath:=$(call get_hdr_inc,__LINUX_OPENVSWITCH_H,openvswitch.h)
 CFLAGS_ovs_flow:=$(call get_hdr_inc,__LINUX_OPENVSWITCH_H,openvswitch.h)
 CFLAGS_ovs_packet:=$(call get_hdr_inc,__LINUX_OPENVSWITCH_H,openvswitch.h)
 CFLAGS_ovs_vport:=$(call get_hdr_inc,__LINUX_OPENVSWITCH_H,openvswitch.h)
+CFLAGS_pon:=$(call get_hdr_inc,_LINUX_PON_H,pon.h)
 CFLAGS_psp:=$(call get_hdr_inc,_LINUX_PSP_H,psp.h)
 CFLAGS_rt-addr:=$(call get_hdr_inc,__LINUX_RTNETLINK_H,rtnetlink.h) \
 	$(call get_hdr_inc,__LINUX_IF_ADDR_H,if_addr.h)
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 18+ messages in thread

* [RFC net-next 02/12] net: add a PON device pointer to struct net_device
  2026-10-08 14:32 [RFC net-next 00/12] net: add the PON subsystem John Crispin
  2026-10-08 14:32 ` [RFC net-next 01/12] net: pon: add the netlink specification of the pon family John Crispin
@ 2026-10-08 14:32 ` John Crispin
  2026-10-08 14:32 ` [RFC net-next 03/12] net: pon: add the PLOAM vocabulary and message codec John Crispin
                   ` (11 subsequent siblings)
  13 siblings, 0 replies; 18+ messages in thread
From: John Crispin @ 2026-10-08 14:32 UTC (permalink / raw)
  To: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni
  Cc: Andrew Lunn, Simon Horman, netdev, linux-kernel,
	Christian Marangi

A PON device attaches to three kinds of network device: the ethernet
device whose rings carry its frames (the conduit), its data network
device and the GEM network devices that stack on the data network
device. Give struct net_device the pointer, beside the other subsystem
pointers, as an RCU pointer under CONFIG_PON.

Nothing sets the pointer yet. The PON core and its accessor,
netdev_uses_pon(), follow.

Assisted-by: LLM
Signed-off-by: John Crispin <john@phrozen.org>
---
 include/linux/netdevice.h | 6 ++++++
 1 file changed, 6 insertions(+)

diff --git a/include/linux/netdevice.h b/include/linux/netdevice.h
index 8cae9b00211e..1b4ed6ae7e87 100644
--- a/include/linux/netdevice.h
+++ b/include/linux/netdevice.h
@@ -72,6 +72,7 @@ struct wireless_dev;
 /* 802.15.4 specific */
 struct wpan_dev;
 struct mpls_dev;
+struct pon_dev;
 /* UDP Tunnel offloads */
 struct udp_tunnel_info;
 struct udp_tunnel_nic_info;
@@ -1974,6 +1975,8 @@ enum netdev_reg_state {
  *	@mpls_ptr:	mpls_dev struct pointer
  *	@mctp_ptr:	MCTP specific data
  *	@psp_dev:	PSP crypto device registered for this netdev
+ *	@pon_dev:	PON device that uses this netdev as conduit, data
+ *			interface or GEM network device
  *
  *	@dev_addr:	Hw address (before bcast,
  *			because most packets are unicast)
@@ -2390,6 +2393,9 @@ struct net_device {
 #if IS_ENABLED(CONFIG_INET_PSP)
 	struct psp_dev __rcu	*psp_dev;
 #endif
+#if IS_ENABLED(CONFIG_PON)
+	struct pon_dev __rcu	*pon_dev;
+#endif
 
 /*
  * Cache lines mostly used on receive path (including eth_type_trans())
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 18+ messages in thread

* [RFC net-next 03/12] net: pon: add the PLOAM vocabulary and message codec
  2026-10-08 14:32 [RFC net-next 00/12] net: add the PON subsystem John Crispin
  2026-10-08 14:32 ` [RFC net-next 01/12] net: pon: add the netlink specification of the pon family John Crispin
  2026-10-08 14:32 ` [RFC net-next 02/12] net: add a PON device pointer to struct net_device John Crispin
@ 2026-10-08 14:32 ` John Crispin
  2026-10-08 14:32 ` [RFC net-next 04/12] net: pon: add the device state and the lent context John Crispin
                   ` (10 subsequent siblings)
  13 siblings, 0 replies; 18+ messages in thread
From: John Crispin @ 2026-10-08 14:32 UTC (permalink / raw)
  To: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, linux-kernel, netdev, Andrew Lunn,
	Christian Marangi

PLOAM is the activation protocol of ITU-T G.9807.1. ploam.h holds the
values that the Recommendation defines rather than what one MAC encodes:
the message identifiers, the completion codes, the key indexes of Table
C.11.12, the Alloc-ID and XGEM Port-ID ranges of Tables C.6.5 and C.6.6,
the ONU-ID range of Table C.6.4, the field lengths and the timers. Every
value is a G.9807.1 value. Every PON driver needs them, so they belong
to the subsystem.

Four values of the key exchange that every XG(S)-PON MAC needs are in
ploam.h as well:

- PON_PLOAM_KEY_NAME_CONSTANT, the ASCII form of the constant that the
  Key_Name hashes after the key (Table C.11.26), with its length
- PON_PLOAM_KEY_FRAGMENT_FIRST, the fragment number of a single
  fragment Key_Report (Table C.11.26)
- PON_PLOAM_DEFAULT_PLOAM_IK_BYTE, the byte of the default PLOAM_IK
  (clause C.15.3.3)
- PON_PLOAM_DEFAULT_MSK, the MSK of the well-known default
  Registration_ID (formula C.15-2), as an initializer. A MAC that
  derives an MSK only from the Registration_ID programmed into it needs
  the value as a constant.

The core does not use them. A PON MAC driver does.

The codec parses the downstream messages that the ONU handles into
structures and builds the upstream ones. It holds no state, no timer and
no policy. The builder refuses an ONU-ID above ten bits. It sends
Serial_Number_ONU with ONU-ID 0x3FF and sequence number 0, as Table
C.11.24 fixes both fields. The control member of the parsed Key_Control
carries the Control flag of Table C.11.12, generate or confirm.

The core itself never calls the codec. The PLOAM state machine runs in
the MAC driver, because its timing and its registers are the MAC's. The
codec is in the subsystem so that every PON MAC driver parses and builds
the messages with one implementation instead of its own. The core uses a
few of the ploam.h values itself: the netlink handlers validate the
Alloc-ID and XGEM Port-ID ranges with them. The other values serve the
drivers.

The core does not call two helpers either. A PON MAC driver needs
pon_ploam_alloc_is_assignable() to refuse an Assign_Alloc-ID outside the
assignable range of Table C.6.5. pon_ploam_profile_is_broadcast() tells
a driver that a Burst_Profile addresses every ONU of its upstream rate:
Table C.11.4 uses ONU-ID 0x3FE for the 9.95328 Gbit/s ONUs beside the
broadcast ONU-ID 0x3FF.

Nothing builds it yet. The build is wired at the end of the series.

Assisted-by: LLM
Signed-off-by: John Crispin <john@phrozen.org>
---
 include/net/pon/ploam.h     | 266 ++++++++++++++++++++++++++++++
 include/net/pon/ploam_msg.h | 306 +++++++++++++++++++++++++++++++++++
 net/pon/pon_ploam_msg.c     | 312 ++++++++++++++++++++++++++++++++++++
 3 files changed, 884 insertions(+)
 create mode 100644 include/net/pon/ploam.h
 create mode 100644 include/net/pon/ploam_msg.h
 create mode 100644 net/pon/pon_ploam_msg.c

diff --git a/include/net/pon/ploam.h b/include/net/pon/ploam.h
new file mode 100644
index 000000000000..2fc27be3f1cc
--- /dev/null
+++ b/include/net/pon/ploam.h
@@ -0,0 +1,266 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+/* Copyright (C) 2026 John Crispin <john@phrozen.org> */
+
+#ifndef __NET_PON_PLOAM_H
+#define __NET_PON_PLOAM_H
+
+#include <linux/bits.h>
+#include <linux/types.h>
+
+/**
+ * DOC: The PLOAM vocabulary
+ *
+ * What ITU-T G.9807.1 defines, rather than what any one MAC's registers
+ * encode. Every PON driver needs these, so they belong to the subsystem. The
+ * header holds no value of another ITU-T PON system, although the mode
+ * enumeration of the uapi names them.
+ *
+ * It holds the message identifiers, the completion codes, the Alloc-ID and
+ * XGEM Port-ID ranges of Tables C.6.5 and C.6.6, the ONU-ID range of
+ * Table C.6.4, the key indexes of Table C.11.12, the field lengths and the
+ * timers.
+ *
+ * This header is kernel internal. Only the handful of enumerations userspace
+ * reads are uapi, in <uapi/linux/pon.h>.
+ *
+ * A value that is a vendor's recovery policy or a register layout stays in
+ * that vendor's driver even when it sits next to a PLOAM field.
+ *
+ */
+
+/**
+ * enum pon_ploam_down_id - downstream PLOAM message identifiers
+ * @PON_PLOAM_DOWN_BURST_PROFILE: Burst_Profile
+ * @PON_PLOAM_DOWN_ASSIGN_ONU_ID: Assign_ONU-ID
+ * @PON_PLOAM_DOWN_RANGING_TIME: Ranging_Time
+ * @PON_PLOAM_DOWN_DEACTIVATE: Deactivate_ONU-ID
+ * @PON_PLOAM_DOWN_DISABLE_SN: Disable_Serial_Number
+ * @PON_PLOAM_DOWN_REQUEST_REG: Request_Registration
+ * @PON_PLOAM_DOWN_ASSIGN_ALLOC_ID: Assign_Alloc-ID
+ * @PON_PLOAM_DOWN_KEY_CONTROL: Key_Control
+ * @PON_PLOAM_DOWN_SLEEP_ALLOW: Sleep_Allow
+ * @PON_PLOAM_DOWN_REBOOT_ONU: Reboot_ONU
+ * @PON_PLOAM_DOWN_MAX: one above the highest identifier listed here, the
+ *	size of an array indexed by the identifier, not a wire value
+ *
+ * ITU-T G.9807.1 Table C.11.2.
+ */
+enum pon_ploam_down_id {
+	PON_PLOAM_DOWN_BURST_PROFILE	= 0x01,
+	PON_PLOAM_DOWN_ASSIGN_ONU_ID	= 0x03,
+	PON_PLOAM_DOWN_RANGING_TIME	= 0x04,
+	PON_PLOAM_DOWN_DEACTIVATE	= 0x05,
+	PON_PLOAM_DOWN_DISABLE_SN	= 0x06,
+	PON_PLOAM_DOWN_REQUEST_REG	= 0x09,
+	PON_PLOAM_DOWN_ASSIGN_ALLOC_ID	= 0x0a,
+	PON_PLOAM_DOWN_KEY_CONTROL	= 0x0d,
+	PON_PLOAM_DOWN_SLEEP_ALLOW	= 0x12,
+	PON_PLOAM_DOWN_REBOOT_ONU	= 0x1d,
+	PON_PLOAM_DOWN_MAX		= 0x1e,
+};
+
+/**
+ * enum pon_ploam_up_id - upstream PLOAM message identifiers
+ * @PON_PLOAM_UP_SERIAL_NUMBER: Serial_Number_ONU
+ * @PON_PLOAM_UP_REGISTRATION: Registration
+ * @PON_PLOAM_UP_KEY_REPORT: Key_Report
+ * @PON_PLOAM_UP_ACKNOWLEDGE: Acknowledgment
+ * @PON_PLOAM_UP_SLEEP_REQUEST: Sleep_Request
+ * @PON_PLOAM_UP_MAX: one above the highest identifier listed here, the size
+ *	of an array indexed by the identifier, not a wire value
+ *
+ * ITU-T G.9807.1 Table C.11.3.
+ */
+enum pon_ploam_up_id {
+	PON_PLOAM_UP_SERIAL_NUMBER	= 0x01,
+	PON_PLOAM_UP_REGISTRATION	= 0x02,
+	PON_PLOAM_UP_KEY_REPORT		= 0x05,
+	PON_PLOAM_UP_ACKNOWLEDGE	= 0x09,
+	PON_PLOAM_UP_SLEEP_REQUEST	= 0x10,
+	PON_PLOAM_UP_MAX		= 0x11,
+};
+
+/**
+ * enum pon_ploam_ack - completion codes of the upstream Acknowledgment message
+ * @PON_PLOAM_ACK_OK: OK
+ * @PON_PLOAM_ACK_NO_MESSAGE: no message to send
+ * @PON_PLOAM_ACK_BUSY: busy, preparing a response
+ * @PON_PLOAM_ACK_UNKNOWN_TYPE: unknown message type
+ * @PON_PLOAM_ACK_PARAM_ERR: parameter error
+ * @PON_PLOAM_ACK_PROCESS_ERR: processing error
+ *
+ * ITU-T G.9807.1 Table C.11.27, the Completion_code field.
+ */
+enum pon_ploam_ack {
+	PON_PLOAM_ACK_OK		= 0,
+	PON_PLOAM_ACK_NO_MESSAGE	= 1,
+	PON_PLOAM_ACK_BUSY		= 2,
+	PON_PLOAM_ACK_UNKNOWN_TYPE	= 3,
+	PON_PLOAM_ACK_PARAM_ERR		= 4,
+	PON_PLOAM_ACK_PROCESS_ERR	= 5,
+};
+
+/**
+ * enum pon_ploam_disable_mode - the Mode field of Disable_Serial_Number
+ * @PON_PLOAM_DISABLE_ALLOW_ONE: the ONU with this serial number is allowed
+ *	upstream access
+ * @PON_PLOAM_DISABLE_DENY_ALL: all ONUs are denied upstream access. The
+ *	serial number is ignored
+ * @PON_PLOAM_DISABLE_ALLOW_ALL: all ONUs are allowed upstream access
+ * @PON_PLOAM_DISABLE_DENY_ONE: the ONU with this serial number is denied
+ *	upstream access
+ *
+ * ITU-T G.9807.1 Table C.11.9, the Disable/enable field.
+ */
+enum pon_ploam_disable_mode {
+	PON_PLOAM_DISABLE_ALLOW_ONE	= 0x00,
+	PON_PLOAM_DISABLE_DENY_ALL	= 0x0f,
+	PON_PLOAM_DISABLE_ALLOW_ALL	= 0xf0,
+	PON_PLOAM_DISABLE_DENY_ONE	= 0xff,
+};
+
+/* The equalization delay encoding of Ranging_Time, G.9807.1 Table C.11.7. */
+#define PON_PLOAM_EQD_RELATIVE		0
+#define PON_PLOAM_EQD_ABSOLUTE		1
+#define PON_PLOAM_EQD_POSITIVE		0
+
+/* The Alloc-ID type field of Assign_Alloc-ID, G.9807.1 Table C.11.11. */
+#define PON_PLOAM_ALLOC_ASSIGN		0x01
+#define PON_PLOAM_ALLOC_DEALLOCATE	0xff
+
+/* Alloc-ID values, G.9807.1 Table C.6.5. The default alloc-id equals the
+ * ONU-ID, the three above the default range are serial number grants that
+ * are never assigned to an ONU and the rest are assignable. 1022 is the
+ * serial number grant for the 9.95328 Gbit/s upstream rate, 1023 the one
+ * for 2.48832 Gbit/s.
+ */
+#define PON_PLOAM_ALLOC_ID_DEFAULT_MAX	1020
+#define PON_PLOAM_ALLOC_ID_SN_GRANT_MIN	1021
+#define PON_PLOAM_ALLOC_ID_SN_GRANT_10G	1022
+#define PON_PLOAM_ALLOC_ID_SN_GRANT_2G5	1023
+#define PON_PLOAM_ALLOC_ID_MAX		16383
+
+/* XGEM Port-ID values, G.9807.1 Table C.6.6. 0 to 1020 is the default
+ * Port-ID, which equals the ONU-ID and carries only the OMCC. The OLT
+ * assigns every other GEM port of the ONU over the OMCC from 1021 to 65534
+ * and 65535 is the idle Port-ID.
+ */
+#define PON_GEM_PORT_ID_ASSIGNABLE_MIN	1021
+#define PON_GEM_PORT_ID_ASSIGNABLE_MAX	65534
+
+/* The Reboot_ONU fields, G.9807.1 Table C.11.23A. */
+#define PON_PLOAM_REBOOT_DEPTH_MAX		3
+#define PON_PLOAM_REBOOT_IMAGE_MAX		1
+#define PON_PLOAM_REBOOT_STATE_INACTIVE_ONLY	1
+#define PON_PLOAM_REBOOT_CALLS_MASK		0x3
+
+/* The Key_Length a Key_Control carries for the AES-128 cipher,
+ * G.9807.1 Table C.11.12.
+ */
+#define PON_PLOAM_KEY_LEN_AES128	16
+
+/* The Key index of Key_Control and Key_Report, the two low bits of the
+ * field. Values 00 and 11 are not defined. G.9807.1 Table C.11.12 and
+ * Table C.11.26.
+ */
+#define PON_PLOAM_KEY_INDEX_FIRST	1
+#define PON_PLOAM_KEY_INDEX_SECOND	2
+
+/* The Control flag of Key_Control (G.9807.1 Table C.11.12) and the Report
+ * type of Key_Report (Table C.11.26).
+ */
+#define PON_PLOAM_KEY_CONTROL_GENERATE		0
+#define PON_PLOAM_KEY_CONTROL_CONFIRM		1
+#define PON_PLOAM_KEY_REPORT_TYPE_NEW		0
+#define PON_PLOAM_KEY_REPORT_TYPE_EXISTING	1
+
+/* The Key_Name of a Key_Report on an existing key is AES-CMAC(KEK,
+ * encryption_key | 0x33313431353932363533353839373933, 128), G.9807.1
+ * Table C.11.26. The constant is the ASCII string below, without its NUL.
+ * The key fragment number of a single fragment report is 0.
+ */
+#define PON_PLOAM_KEY_NAME_CONSTANT		"3141592653589793"
+#define PON_PLOAM_KEY_NAME_CONSTANT_LEN		16
+#define PON_PLOAM_KEY_FRAGMENT_FIRST		0
+
+/* The default PLOAM_IK is this byte repeated sixteen times, G.9807.1
+ * clause C.15.3.3. The MSK of the well-known default Registration_ID
+ * (thirty-six zero bytes, Table C.11.25) is formula C.15-2 under that key,
+ * clause C.15.3.2. PON_PLOAM_DEFAULT_MSK holds it as an initializer. A MAC
+ * that derives an MSK only from the Registration_ID programmed into it
+ * needs the value as a constant.
+ */
+#define PON_PLOAM_DEFAULT_PLOAM_IK_BYTE		0x55
+#define PON_PLOAM_DEFAULT_MSK	{			\
+	0x24, 0x37, 0xbe, 0x54, 0xe9, 0x5e, 0x6e, 0xe3,	\
+	0x53, 0x8b, 0xb1, 0xb4, 0xb5, 0xd4, 0x32, 0xeb,	\
+}
+
+/* The delimiter and preamble lengths of the burst profile, in octets,
+ * G.9807.1 Table C.11.4.
+ */
+#define PON_PLOAM_BURST_DELIMITER_LEN_MAX	8
+#define PON_PLOAM_BURST_PREAMBLE_LEN_MIN	1
+#define PON_PLOAM_BURST_PREAMBLE_LEN_MAX	8
+
+/* The upstream line rate bit: R of the burst profile (G.9807.1
+ * Table C.11.4) and U of Assign_ONU-ID (Table C.11.6). Then the widths of the
+ * preamble repeat count of the burst profile, Table C.11.4.
+ */
+#define PON_PLOAM_LINE_RATE_XGPON	0
+#define PON_PLOAM_LINE_RATE_XGSPON	1
+#define PON_PLOAM_PREAMBLE_MASK_XGPON	0x1f
+#define PON_PLOAM_PREAMBLE_MASK_XGSPON	0xff
+
+/* The upstream line rate capability of Serial_Number_ONU, a bitmap of the
+ * form 0000 00HL, G.9807.1 Table C.11.24. H set: the ONU supports the
+ * 9.95328 Gbit/s upstream rate. L set: the ONU does not support the
+ * 2.48832 Gbit/s upstream rate.
+ */
+#define PON_PLOAM_SN_RATE_10G		BIT(1)
+#define PON_PLOAM_SN_RATE_NO_2G5	BIT(0)
+
+/* The ONU-ID is ten bits, G.9807.1 clause C.11.2.1. The OLT assigns 0 to
+ * 1020, Table C.6.4. 0x3ff is both the broadcast destination and the value
+ * an ONU carries before the OLT assigns it one. 0x3fe appears only in a
+ * Burst_Profile, as the broadcast destination of a profile for the
+ * 9.95328 Gbit/s upstream rate.
+ */
+#define PON_PLOAM_ONU_ID_MASK			GENMASK(9, 0)
+#define PON_PLOAM_ONU_ID_MAX			1020
+#define PON_PLOAM_ONU_ID_BROADCAST		0x3ff
+#define PON_PLOAM_ONU_ID_UNASSIGNED		0x3ff
+#define PON_PLOAM_ONU_ID_PROFILE_BCAST_10G	0x3fe
+
+/* Fields of the standard message, in bytes. The serial number is the
+ * Vendor_ID and the VSSN, G.9807.1 clauses C.11.2.6.1 and C.11.2.6.2. The
+ * Registration_ID is Table C.11.25 and the key fragment Table C.11.26.
+ */
+#define PON_PLOAM_SN_LEN		8
+#define PON_PLOAM_REG_ID_LEN		36
+#define PON_PLOAM_KEY_FRAGMENT_LEN	32
+
+/* The recommended initial value of TO1, in milliseconds, Table C.12.2. */
+#define PON_PLOAM_TO1_MS		10000
+
+/* The key exchange timers, in milliseconds, G.9807.1 clause C.15.5.3.3. */
+#define PON_PLOAM_TK4_MS		100
+#define PON_PLOAM_TK5_MS		20
+
+/**
+ * pon_ploam_alloc_is_assignable() - whether the OLT can assign an Alloc-ID
+ * @alloc_id: the fourteen bit Alloc-ID
+ *
+ * ITU-T G.9807.1 Table C.6.5. The core does not call it. A PON MAC driver
+ * needs it to refuse an Assign_Alloc-ID outside that range.
+ *
+ * Return: true for the values above PON_PLOAM_ALLOC_ID_SN_GRANT_2G5 up to
+ * PON_PLOAM_ALLOC_ID_MAX, 1024 to 16383.
+ */
+static inline bool pon_ploam_alloc_is_assignable(u16 alloc_id)
+{
+	return alloc_id > PON_PLOAM_ALLOC_ID_SN_GRANT_2G5 &&
+	       alloc_id <= PON_PLOAM_ALLOC_ID_MAX;
+}
+
+#endif /* __NET_PON_PLOAM_H */
diff --git a/include/net/pon/ploam_msg.h b/include/net/pon/ploam_msg.h
new file mode 100644
index 000000000000..fe25da8af12c
--- /dev/null
+++ b/include/net/pon/ploam_msg.h
@@ -0,0 +1,306 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+/* Copyright (C) 2026 John Crispin <john@phrozen.org> */
+
+#ifndef __NET_PON_PLOAM_MSG_H
+#define __NET_PON_PLOAM_MSG_H
+
+#include <linux/types.h>
+#include <net/pon/ploam.h>
+
+/**
+ * DOC: The PLOAM message codec
+ *
+ * Lays out and reads the standard message body. It is a pure function of an
+ * ITU-T wire format: no state, no timer, no register and no policy. A driver
+ * decides what a message means and when to send one. This decides where the
+ * bytes go.
+ *
+ * The body is what the standard defines and nothing else: the ONU-ID, the
+ * message identifier, the sequence number and 36 bytes of content. Whatever a
+ * MAC wraps around that, a prefix, a trailer, a FIFO word order or a message
+ * integrity check, stays in that MAC's driver.
+ */
+
+/* The generic PLOAM message structure, G.9807.1 Table C.11.1. */
+#define PON_PLOAM_CONTENT_LEN	36
+#define PON_PLOAM_BODY_LEN	(2 + 1 + 1 + PON_PLOAM_CONTENT_LEN)
+
+/**
+ * struct pon_ploam_sn - Serial_Number_ONU
+ * @sn: the eight byte serial number
+ * @random_delay: the response delay the ONU applies, in bit periods at
+ *	2.48832 Gbit/s whatever the upstream rate of the ONU
+ *
+ * ITU-T G.9807.1 Table C.11.24.
+ */
+struct pon_ploam_sn {
+	u8 sn[PON_PLOAM_SN_LEN];
+	u32 random_delay;
+};
+
+/**
+ * struct pon_ploam_registration - Registration
+ * @reg_id: the registration id
+ *
+ * ITU-T G.9807.1 Table C.11.25.
+ */
+struct pon_ploam_registration {
+	u8 reg_id[PON_PLOAM_REG_ID_LEN];
+};
+
+/**
+ * struct pon_ploam_key_report - Key_Report
+ * @key: the key fragment
+ * @len: its length, at most PON_PLOAM_KEY_FRAGMENT_LEN
+ * @type: new key or existing key, PON_PLOAM_KEY_REPORT_TYPE_*
+ * @index: the reported key index
+ * @num: the fragment number
+ *
+ * ITU-T G.9807.1 Table C.11.26.
+ */
+struct pon_ploam_key_report {
+	u8 key[PON_PLOAM_KEY_FRAGMENT_LEN];
+	u8 len;
+	u8 type;
+	u8 index;
+	u8 num;
+};
+
+/**
+ * struct pon_ploam_ack_msg - Acknowledgment
+ * @code: the completion code, enum pon_ploam_ack
+ *
+ * ITU-T G.9807.1 Table C.11.27.
+ */
+struct pon_ploam_ack_msg {
+	u8 code;
+};
+
+/**
+ * struct pon_ploam_up - one upstream message to lay out
+ * @onu_id: the ONU-ID to send it from, at most PON_PLOAM_ONU_ID_MASK
+ * @msg_id: which message, enum pon_ploam_up_id
+ * @seq_no: the downstream sequence number to echo, or 0
+ * @sn: Serial_Number_ONU content
+ * @registration: Registration content
+ * @key_report: Key_Report content
+ * @ack: Acknowledgment content
+ *
+ * ITU-T G.9807.1 Table C.11.3, the upstream message summary.
+ */
+struct pon_ploam_up {
+	u16 onu_id;
+	u8 msg_id;
+	u8 seq_no;
+	union {
+		struct pon_ploam_sn sn;
+		struct pon_ploam_registration registration;
+		struct pon_ploam_key_report key_report;
+		struct pon_ploam_ack_msg ack;
+	};
+};
+
+int pon_ploam_up_build(void *buf, size_t len, const struct pon_ploam_up *msg);
+
+/* The Burst_Profile field sizes, G.9807.1 Table C.11.4. */
+#define PON_PLOAM_BURST_PATTERN_LEN	8
+#define PON_PLOAM_PON_TAG_LEN		8
+
+/**
+ * struct pon_ploam_burst_profile - Burst_Profile
+ * @delimiter: the burst delimiter pattern
+ * @preamble: the burst preamble pattern
+ * @pon_tag: the PON tag the OLT assigns
+ * @index: which of the profiles this one is
+ * @version: the profile version
+ * @line_rate: the upstream line rate, PON_PLOAM_LINE_RATE_*
+ * @fec: upstream FEC is on
+ * @delimiter_len: significant bytes of @delimiter
+ * @preamble_len: significant bytes of @preamble
+ * @preamble_repeat: how many times the preamble repeats, masked to the width
+ *	that @line_rate gives the field: five bits for
+ *	PON_PLOAM_LINE_RATE_XGPON, eight for PON_PLOAM_LINE_RATE_XGSPON
+ *
+ * ITU-T G.9807.1 Table C.11.4.
+ */
+struct pon_ploam_burst_profile {
+	u8 delimiter[PON_PLOAM_BURST_PATTERN_LEN];
+	u8 preamble[PON_PLOAM_BURST_PATTERN_LEN];
+	u8 pon_tag[PON_PLOAM_PON_TAG_LEN];
+	u8 index;
+	u8 version;
+	u8 line_rate;
+	u8 fec;
+	u8 delimiter_len;
+	u8 preamble_len;
+	u8 preamble_repeat;
+};
+
+/**
+ * struct pon_ploam_assign_onu_id - Assign_ONU-ID
+ * @sn: the serial number the assignment is for
+ * @onu_id: the ONU-ID being assigned
+ * @line_rate: the upstream nominal line rate the OLT selects,
+ *	PON_PLOAM_LINE_RATE_*. It applies only to an ONU that supports both
+ *	upstream rates
+ *
+ * ITU-T G.9807.1 Table C.11.6.
+ */
+struct pon_ploam_assign_onu_id {
+	u8 sn[PON_PLOAM_SN_LEN];
+	u16 onu_id;
+	u8 line_rate;
+};
+
+/**
+ * struct pon_ploam_ranging_time - Ranging_Time
+ * @eqd: the equalization delay, in bit periods at 2.48832 Gbit/s whatever
+ *	the upstream rate of the ONU
+ * @absolute: @eqd replaces the current value rather than adjusting it
+ * @positive: an adjustment adds rather than subtracts
+ *
+ * ITU-T G.9807.1 Table C.11.7.
+ */
+struct pon_ploam_ranging_time {
+	u32 eqd;
+	bool absolute;
+	bool positive;
+};
+
+/**
+ * struct pon_ploam_disable_sn - Disable_Serial_Number
+ * @sn: the serial number the mode applies to
+ * @mode: enum pon_ploam_disable_mode
+ *
+ * ITU-T G.9807.1 Table C.11.9.
+ */
+struct pon_ploam_disable_sn {
+	u8 sn[PON_PLOAM_SN_LEN];
+	u8 mode;
+};
+
+/**
+ * struct pon_ploam_assign_alloc_id - Assign_Alloc-ID
+ * @alloc_id: the alloc-id
+ * @type: assign or deallocate, PON_PLOAM_ALLOC_*
+ *
+ * ITU-T G.9807.1 Table C.11.11.
+ */
+struct pon_ploam_assign_alloc_id {
+	u16 alloc_id;
+	u8 type;
+};
+
+/**
+ * struct pon_ploam_key_control - Key_Control
+ * @key_index: the key index to report
+ * @control: the Control flag, generate a new key or confirm a key,
+ *	PON_PLOAM_KEY_CONTROL_*
+ * @key_length: the key length the OLT asks for, in bytes, where 0 means 256
+ *
+ * ITU-T G.9807.1 Table C.11.12.
+ */
+struct pon_ploam_key_control {
+	u8 key_index;
+	u8 control;
+	u8 key_length;
+};
+
+/**
+ * struct pon_ploam_reboot - Reboot_ONU
+ * @sn: the serial number a broadcast message names, all zero for every ONU
+ * @depth: the reboot depth, where 0 is an OMCI MIB reset
+ * @image: which software image to load
+ * @state: the activation states the reboot applies in
+ * @flags: the call-in-progress conditions
+ *
+ * ITU-T G.9807.1 Table C.11.23A.
+ */
+struct pon_ploam_reboot {
+	u8 sn[PON_PLOAM_SN_LEN];
+	u8 depth;
+	u8 image;
+	u8 state;
+	u8 flags;
+};
+
+/**
+ * struct pon_ploam_down - one downstream message, parsed
+ * @onu_id: the destination, the ten bits the standard defines
+ * @msg_id: which message, enum pon_ploam_down_id
+ * @seq_no: the sequence number to echo in an acknowledgment
+ * @burst_profile: Burst_Profile content
+ * @assign_onu_id: Assign_ONU-ID content
+ * @ranging_time: Ranging_Time content
+ * @disable_sn: Disable_Serial_Number content
+ * @assign_alloc_id: Assign_Alloc-ID content
+ * @key_control: Key_Control content
+ * @reboot: Reboot_ONU content
+ *
+ * ITU-T G.9807.1 Table C.11.2, the downstream message summary.
+ */
+struct pon_ploam_down {
+	u16 onu_id;
+	u8 msg_id;
+	u8 seq_no;
+	union {
+		struct pon_ploam_burst_profile burst_profile;
+		struct pon_ploam_assign_onu_id assign_onu_id;
+		struct pon_ploam_ranging_time ranging_time;
+		struct pon_ploam_disable_sn disable_sn;
+		struct pon_ploam_assign_alloc_id assign_alloc_id;
+		struct pon_ploam_key_control key_control;
+		struct pon_ploam_reboot reboot;
+	};
+};
+
+int pon_ploam_down_parse(const void *buf, size_t len,
+			 struct pon_ploam_down *msg);
+
+/**
+ * pon_ploam_is_broadcast() - whether an ONU-ID is the broadcast destination
+ * @onu_id: the ten bit ONU-ID of a downstream message
+ *
+ * Only 0x3ff counts. 0x3fe, the broadcast of a Burst_Profile for the
+ * 9.95328 Gbit/s upstream rate, does not. ITU-T G.9807.1 clause C.11.2.1.
+ * A Burst_Profile tests its destination with
+ * pon_ploam_profile_is_broadcast().
+ *
+ * Return: true for PON_PLOAM_ONU_ID_BROADCAST.
+ */
+static inline bool pon_ploam_is_broadcast(u16 onu_id)
+{
+	return onu_id == PON_PLOAM_ONU_ID_BROADCAST;
+}
+
+/**
+ * pon_ploam_profile_is_broadcast() - whether a Burst_Profile is broadcast
+ * @onu_id: the ten bit ONU-ID of a downstream Burst_Profile
+ *
+ * A Burst_Profile goes to 0x3ff for every ONU or to 0x3fe for the ONUs that
+ * support the 9.95328 Gbit/s upstream rate. An XGS-PON ONU takes both alike.
+ * ITU-T G.9807.1 Table C.11.4 and Appendix II.
+ *
+ * Return: true for PON_PLOAM_ONU_ID_BROADCAST and
+ * PON_PLOAM_ONU_ID_PROFILE_BCAST_10G.
+ */
+static inline bool pon_ploam_profile_is_broadcast(u16 onu_id)
+{
+	return pon_ploam_is_broadcast(onu_id) ||
+	       onu_id == PON_PLOAM_ONU_ID_PROFILE_BCAST_10G;
+}
+
+/**
+ * pon_ploam_is_assignable() - whether the OLT can assign an ONU-ID
+ * @onu_id: the ten bit ONU-ID
+ *
+ * ITU-T G.9807.1 Table C.6.4.
+ *
+ * Return: true for 0 to PON_PLOAM_ONU_ID_MAX.
+ */
+static inline bool pon_ploam_is_assignable(u16 onu_id)
+{
+	return onu_id <= PON_PLOAM_ONU_ID_MAX;
+}
+
+#endif /* __NET_PON_PLOAM_MSG_H */
diff --git a/net/pon/pon_ploam_msg.c b/net/pon/pon_ploam_msg.c
new file mode 100644
index 000000000000..27b3285629f2
--- /dev/null
+++ b/net/pon/pon_ploam_msg.c
@@ -0,0 +1,312 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/* Copyright (C) 2026 John Crispin <john@phrozen.org> */
+
+#include <linux/errno.h>
+#include <linux/export.h>
+#include <linux/string.h>
+#include <linux/unaligned.h>
+#include <net/pon/ploam_msg.h>
+
+/* Offsets within the standard body. The ONU-ID is big endian and the content
+ * follows the four byte header. G.9807.1 Table C.11.1.
+ */
+#define PON_PLOAM_OFF_ONU_ID	0
+#define PON_PLOAM_OFF_MSG_ID	2
+#define PON_PLOAM_OFF_SEQ_NO	3
+#define PON_PLOAM_OFF_CONTENT	4
+
+/* Serial_Number_ONU content, G.9807.1 Table C.11.24. */
+#define PON_PLOAM_SN_OFF_SN		0
+#define PON_PLOAM_SN_OFF_DELAY		8
+#define PON_PLOAM_SN_OFF_CAPABILITY	32
+#define PON_PLOAM_SN_CAPABILITY		(PON_PLOAM_SN_RATE_10G | \
+					 PON_PLOAM_SN_RATE_NO_2G5)
+
+/* Registration content, G.9807.1 Table C.11.25. */
+#define PON_PLOAM_REG_OFF_ID		0
+
+/* Key_Report content, G.9807.1 Table C.11.26. */
+#define PON_PLOAM_KR_OFF_TYPE		0
+#define PON_PLOAM_KR_OFF_INDEX		1
+#define PON_PLOAM_KR_OFF_NUM		2
+#define PON_PLOAM_KR_OFF_FRAGMENT	4
+
+/* Acknowledgment content, G.9807.1 Table C.11.27. */
+#define PON_PLOAM_ACK_OFF_CODE		0
+
+/* Downstream content offsets, checked against the message tables of
+ * G.9807.1 clause C.11.3.3: Key_Control is Table C.11.12 and Reboot_ONU is
+ * Table C.11.23A.
+ */
+/* Burst_Profile, G.9807.1 Table C.11.4. */
+#define PON_PLOAM_BP_OFF_FLAGS		0
+#define PON_PLOAM_BP_OFF_FEC		1
+#define PON_PLOAM_BP_OFF_DELIM_LEN	2
+#define PON_PLOAM_BP_OFF_DELIM		3
+#define PON_PLOAM_BP_OFF_PRE_LEN	11
+#define PON_PLOAM_BP_OFF_PRE_REPEAT	12
+#define PON_PLOAM_BP_OFF_PRE		13
+#define PON_PLOAM_BP_OFF_PON_TAG	21
+
+/* Assign_ONU-ID, G.9807.1 Table C.11.6. */
+#define PON_PLOAM_AOI_OFF_ID		0
+#define PON_PLOAM_AOI_OFF_SN		2
+#define PON_PLOAM_AOI_OFF_RATE		10
+
+/* Ranging_Time, G.9807.1 Table C.11.7. */
+#define PON_PLOAM_RT_OFF_FLAGS		0
+#define PON_PLOAM_RT_OFF_EQD		1
+
+/* Disable_Serial_Number, G.9807.1 Table C.11.9. */
+#define PON_PLOAM_DSN_OFF_MODE		0
+#define PON_PLOAM_DSN_OFF_SN		1
+
+/* Assign_Alloc-ID, G.9807.1 Table C.11.11. */
+#define PON_PLOAM_AAI_OFF_ID		0
+#define PON_PLOAM_AAI_OFF_TYPE		2
+
+/* Key_Control, G.9807.1 Table C.11.12. */
+#define PON_PLOAM_KC_OFF_TYPE		1
+#define PON_PLOAM_KC_OFF_INDEX		2
+#define PON_PLOAM_KC_OFF_LENGTH		3
+
+/* Reboot_ONU, G.9807.1 Table C.11.23A. */
+#define PON_PLOAM_RB_OFF_SN		0
+#define PON_PLOAM_RB_OFF_DEPTH		8
+#define PON_PLOAM_RB_OFF_IMAGE		9
+#define PON_PLOAM_RB_OFF_STATE		10
+#define PON_PLOAM_RB_OFF_FLAGS		11
+
+/**
+ * pon_ploam_up_build() - lay out one upstream PLOAM message
+ * @buf:	where to write the standard body, at least PON_PLOAM_BODY_LEN
+ * @len:	the space available
+ * @msg:	what to send
+ *
+ * Writes the ONU-ID, the message identifier, the sequence number and the
+ * content. The caller owns everything around the body: any vendor prefix or
+ * trailer, the message integrity check and the order the bytes reach the
+ * hardware.
+ *
+ * The whole 36 byte content is cleared before the fields are written, so
+ * every octet a message does not use goes out as 0x00. The key report's type,
+ * index and fragment number are masked to their field widths.
+ *
+ * Serial_Number_ONU reports an ONU that supports the 9.95328 Gbit/s upstream
+ * rate only: PON_PLOAM_SN_RATE_10G and PON_PLOAM_SN_RATE_NO_2G5 are set. It
+ * always goes out with ONU-ID PON_PLOAM_ONU_ID_UNASSIGNED and sequence number
+ * 0, as Table C.11.24 fixes both. @msg->onu_id and @msg->seq_no are not used
+ * for it.
+ *
+ * The layouts are ITU-T G.9807.1 Table C.11.24 for Serial_Number_ONU,
+ * Table C.11.25 for Registration, Table C.11.26 for Key_Report and
+ * Table C.11.27 for Acknowledgment.
+ *
+ * Return: the number of bytes written, -EINVAL for a message identifier the
+ * codec does not build or an ONU-ID above PON_PLOAM_ONU_ID_MASK, -ENOSPC for
+ * a buffer that is too small, or -ERANGE for a key fragment longer than
+ * PON_PLOAM_KEY_FRAGMENT_LEN.
+ */
+int pon_ploam_up_build(void *buf, size_t len, const struct pon_ploam_up *msg)
+{
+	u8 *body = buf;
+	u8 *content = body + PON_PLOAM_OFF_CONTENT;
+	u16 onu_id = msg->onu_id;
+	u8 seq_no = msg->seq_no;
+
+	if (len < PON_PLOAM_BODY_LEN)
+		return -ENOSPC;
+
+	if (msg->onu_id > PON_PLOAM_ONU_ID_MASK)
+		return -EINVAL;
+
+	switch (msg->msg_id) {
+	case PON_PLOAM_UP_SERIAL_NUMBER:
+		onu_id = PON_PLOAM_ONU_ID_UNASSIGNED;
+		seq_no = 0;
+		break;
+	case PON_PLOAM_UP_REGISTRATION:
+	case PON_PLOAM_UP_ACKNOWLEDGE:
+		break;
+	case PON_PLOAM_UP_KEY_REPORT:
+		if (msg->key_report.len > PON_PLOAM_KEY_FRAGMENT_LEN)
+			return -ERANGE;
+		break;
+	default:
+		return -EINVAL;
+	}
+
+	put_unaligned_be16(onu_id, &body[PON_PLOAM_OFF_ONU_ID]);
+	body[PON_PLOAM_OFF_MSG_ID] = msg->msg_id;
+	body[PON_PLOAM_OFF_SEQ_NO] = seq_no;
+	memset(content, 0, PON_PLOAM_CONTENT_LEN);
+
+	switch (msg->msg_id) {
+	case PON_PLOAM_UP_SERIAL_NUMBER:
+		memcpy(&content[PON_PLOAM_SN_OFF_SN], msg->sn.sn,
+		       PON_PLOAM_SN_LEN);
+		put_unaligned_be32(msg->sn.random_delay,
+				   &content[PON_PLOAM_SN_OFF_DELAY]);
+		content[PON_PLOAM_SN_OFF_CAPABILITY] = PON_PLOAM_SN_CAPABILITY;
+		break;
+
+	case PON_PLOAM_UP_REGISTRATION:
+		memcpy(&content[PON_PLOAM_REG_OFF_ID], msg->registration.reg_id,
+		       PON_PLOAM_REG_ID_LEN);
+		break;
+
+	case PON_PLOAM_UP_KEY_REPORT:
+		content[PON_PLOAM_KR_OFF_TYPE] = msg->key_report.type & 1;
+		content[PON_PLOAM_KR_OFF_INDEX] = msg->key_report.index & 3;
+		content[PON_PLOAM_KR_OFF_NUM] = msg->key_report.num & 7;
+		memcpy(&content[PON_PLOAM_KR_OFF_FRAGMENT], msg->key_report.key,
+		       msg->key_report.len);
+		break;
+
+	case PON_PLOAM_UP_ACKNOWLEDGE:
+		content[PON_PLOAM_ACK_OFF_CODE] = msg->ack.code;
+		break;
+	}
+
+	return PON_PLOAM_BODY_LEN;
+}
+EXPORT_SYMBOL_GPL(pon_ploam_up_build);
+
+/**
+ * pon_ploam_down_parse() - read one downstream PLOAM message
+ * @buf:	the standard body, at least PON_PLOAM_BODY_LEN bytes
+ * @len:	how much is there
+ * @msg:	filled in on success
+ *
+ * Reads the header and, for the messages it knows, the content into the
+ * matching member. It decides nothing: whether the message is addressed to
+ * this ONU, whether the current state allows it and whether to acknowledge
+ * it are all the caller's.
+ *
+ * The layouts are ITU-T G.9807.1 Table C.11.4 for Burst_Profile, Table C.11.6
+ * for Assign_ONU-ID, Table C.11.7 for Ranging_Time, Table C.11.8 for
+ * Deactivate_ONU-ID, Table C.11.9 for Disable_Serial_Number, Table C.11.10 for
+ * Request_Registration, Table C.11.11 for Assign_Alloc-ID, Table C.11.12 for
+ * Key_Control and Table C.11.23A for Reboot_ONU.
+ *
+ * The codec checks one range only: the delimiter and preamble lengths of
+ * Burst_Profile. Every other field with a range in the Recommendation, such
+ * as the Reboot_ONU depth or the key index, is the caller's to check.
+ *
+ * Return: 0, -ENOSPC for a short buffer, -EINVAL for a Burst_Profile whose
+ * delimiter or preamble length is outside Table C.11.4, or -EOPNOTSUPP for a
+ * message the codec does not decode, which a caller may count and ignore.
+ * After -EINVAL and -EOPNOTSUPP, @msg->onu_id, @msg->msg_id and
+ * @msg->seq_no are valid, so that the caller can acknowledge the message.
+ * After -ENOSPC, @msg is not written.
+ */
+int pon_ploam_down_parse(const void *buf, size_t len,
+			 struct pon_ploam_down *msg)
+{
+	const u8 *body = buf;
+	const u8 *content = body + PON_PLOAM_OFF_CONTENT;
+
+	if (len < PON_PLOAM_BODY_LEN)
+		return -ENOSPC;
+
+	memset(msg, 0, sizeof(*msg));
+	msg->onu_id = get_unaligned_be16(&body[PON_PLOAM_OFF_ONU_ID]) &
+		      PON_PLOAM_ONU_ID_MASK;
+	msg->msg_id = body[PON_PLOAM_OFF_MSG_ID];
+	msg->seq_no = body[PON_PLOAM_OFF_SEQ_NO];
+
+	switch (msg->msg_id) {
+	case PON_PLOAM_DOWN_BURST_PROFILE: {
+		struct pon_ploam_burst_profile *profile = &msg->burst_profile;
+		u8 flags = content[PON_PLOAM_BP_OFF_FLAGS];
+		u8 repeat_mask;
+
+		profile->index = flags & 3;
+		profile->line_rate = (flags >> 2) & 1;
+		if (profile->line_rate == PON_PLOAM_LINE_RATE_XGSPON)
+			repeat_mask = PON_PLOAM_PREAMBLE_MASK_XGSPON;
+		else
+			repeat_mask = PON_PLOAM_PREAMBLE_MASK_XGPON;
+		profile->version = flags >> 4;
+		profile->fec = content[PON_PLOAM_BP_OFF_FEC] & 1;
+		profile->delimiter_len =
+			content[PON_PLOAM_BP_OFF_DELIM_LEN] & 0xf;
+		memcpy(profile->delimiter, &content[PON_PLOAM_BP_OFF_DELIM],
+		       PON_PLOAM_BURST_PATTERN_LEN);
+		profile->preamble_len = content[PON_PLOAM_BP_OFF_PRE_LEN] & 0xf;
+		profile->preamble_repeat =
+			content[PON_PLOAM_BP_OFF_PRE_REPEAT] & repeat_mask;
+		memcpy(profile->preamble, &content[PON_PLOAM_BP_OFF_PRE],
+		       PON_PLOAM_BURST_PATTERN_LEN);
+		memcpy(profile->pon_tag, &content[PON_PLOAM_BP_OFF_PON_TAG],
+		       PON_PLOAM_PON_TAG_LEN);
+		if (profile->delimiter_len > PON_PLOAM_BURST_DELIMITER_LEN_MAX)
+			return -EINVAL;
+		if (profile->preamble_len < PON_PLOAM_BURST_PREAMBLE_LEN_MIN ||
+		    profile->preamble_len > PON_PLOAM_BURST_PREAMBLE_LEN_MAX)
+			return -EINVAL;
+		return 0;
+	}
+
+	case PON_PLOAM_DOWN_ASSIGN_ONU_ID:
+		msg->assign_onu_id.onu_id =
+			get_unaligned_be16(&content[PON_PLOAM_AOI_OFF_ID]) &
+			PON_PLOAM_ONU_ID_MASK;
+		memcpy(msg->assign_onu_id.sn, &content[PON_PLOAM_AOI_OFF_SN],
+		       PON_PLOAM_SN_LEN);
+		msg->assign_onu_id.line_rate =
+			content[PON_PLOAM_AOI_OFF_RATE] & 1;
+		return 0;
+
+	case PON_PLOAM_DOWN_RANGING_TIME: {
+		u8 flags = content[PON_PLOAM_RT_OFF_FLAGS];
+
+		msg->ranging_time.absolute =
+			(flags & 1) == PON_PLOAM_EQD_ABSOLUTE;
+		msg->ranging_time.positive =
+			((flags >> 1) & 1) == PON_PLOAM_EQD_POSITIVE;
+		msg->ranging_time.eqd =
+			get_unaligned_be32(&content[PON_PLOAM_RT_OFF_EQD]);
+		return 0;
+	}
+
+	case PON_PLOAM_DOWN_DEACTIVATE:
+		return 0;
+
+	case PON_PLOAM_DOWN_DISABLE_SN:
+		msg->disable_sn.mode = content[PON_PLOAM_DSN_OFF_MODE];
+		memcpy(msg->disable_sn.sn, &content[PON_PLOAM_DSN_OFF_SN],
+		       PON_PLOAM_SN_LEN);
+		return 0;
+
+	case PON_PLOAM_DOWN_REQUEST_REG:
+		return 0;
+
+	case PON_PLOAM_DOWN_ASSIGN_ALLOC_ID:
+		msg->assign_alloc_id.alloc_id =
+			get_unaligned_be16(&content[PON_PLOAM_AAI_OFF_ID]) &
+			0x3fff;
+		msg->assign_alloc_id.type = content[PON_PLOAM_AAI_OFF_TYPE];
+		return 0;
+
+	case PON_PLOAM_DOWN_KEY_CONTROL:
+		msg->key_control.control = content[PON_PLOAM_KC_OFF_TYPE] & 1;
+		msg->key_control.key_index =
+			content[PON_PLOAM_KC_OFF_INDEX] & 3;
+		msg->key_control.key_length = content[PON_PLOAM_KC_OFF_LENGTH];
+		return 0;
+
+	case PON_PLOAM_DOWN_REBOOT_ONU:
+		memcpy(msg->reboot.sn, &content[PON_PLOAM_RB_OFF_SN],
+		       PON_PLOAM_SN_LEN);
+		msg->reboot.depth = content[PON_PLOAM_RB_OFF_DEPTH];
+		msg->reboot.image = content[PON_PLOAM_RB_OFF_IMAGE];
+		msg->reboot.state = content[PON_PLOAM_RB_OFF_STATE];
+		msg->reboot.flags = content[PON_PLOAM_RB_OFF_FLAGS];
+		return 0;
+
+	default:
+		return -EOPNOTSUPP;
+	}
+}
+EXPORT_SYMBOL_GPL(pon_ploam_down_parse);
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 18+ messages in thread

* [RFC net-next 04/12] net: pon: add the device state and the lent context
  2026-10-08 14:32 [RFC net-next 00/12] net: add the PON subsystem John Crispin
                   ` (2 preceding siblings ...)
  2026-10-08 14:32 ` [RFC net-next 03/12] net: pon: add the PLOAM vocabulary and message codec John Crispin
@ 2026-10-08 14:32 ` John Crispin
  2026-10-08 14:32 ` [RFC net-next 05/12] net: pon: add the conduit contract John Crispin
                   ` (9 subsequent siblings)
  13 siblings, 0 replies; 18+ messages in thread
From: John Crispin @ 2026-10-08 14:32 UTC (permalink / raw)
  To: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, linux-kernel, netdev, Andrew Lunn,
	Christian Marangi

The core lends a MAC driver its serialized context: one ordered
workqueue per instance and the instance lock, which the netlink handlers
take in their pre_doit. A driver queues a pon_work from any context. The
handler runs with the lock held, so a PLOAM transition never interleaves
with a netlink transaction.

The core owns the activation state. A driver reports the state it
reached through pon_dev_state_report(). The core validates the edge
against Table C.12.4 of ITU-T G.9807.1, publishes it and warns about an
illegal edge instead of hiding it. The PON network devices have a
carrier in O5 and O6 while a GEM port rides a bound alloc-id and the
conduit is paired and up. The driver reports the assignment and the
release of an alloc-id through pon_dev_event(), which is where the
carrier learns of them.

Every edge that pon_dev_state_report() publishes gets one line in the
kernel log at info level, for example "pon0: PLOAM state O5 -> O6". An
XGS-PON ONU is often a router whose developers watch the serial
console. The state reaches userspace through ploam-ntf only, so without
the line the console shows nothing when the ONU loses its link or comes
back. A repeated report changes nothing and logs nothing. The line uses
the names of the uapi states. O2 is named O2-3 in every mode but G-PON,
because ITU-T G.9807.1 Table C.12.1 has one Serial Number state that a
driver reports as o2.

A print to a serial console blocks for milliseconds, while the driver's
PLOAM work answers the OLT within a deadline. pon_dev_log() therefore
records a line in a ring of PON_LOG_LINES lines under the instance lock
and queues a work item on system_dfl_wq that prints it. The work holds
the lock only to take one line out, never across a print. The lines
keep their order. A full ring drops the oldest line and counts it. The
count is printed in its place. Nothing is recorded once unregister has
begun. Each line keeps its printk level. pon_dev_log() records at info
level. The core helper pon_dev_log_level() takes the level from its
caller.

pon_dev_log() is exported for a PON MAC driver: only the driver knows
why it leaves an activation state. It logs that from the same PLOAM
work, so it must not print there either. Through the same ring its line
keeps its place next to the edge that the core logs.

types.h and functions.h carry the whole contract between the core and a
driver, so the files that follow add code and no declarations.

The lent context exists for the MAC driver more than for the core. A PON
MAC signals PLOAM messages and PHY edges in hard interrupt context,
while its state machine sleeps and must not run against a netlink
transaction. The driver therefore queues a pon_work from its interrupt
handler. The handler runs under the instance lock. The core's own user
of this context is the OMCI receive path. A plain work item would serve
it. The mechanism exists for the driver.

The same holds for the reports. pon_dev_state_report() and
pon_dev_event() have no caller in the core: the driver runs the state
machine and receives the messages of the OLT. The core validates,
publishes and relays what the driver reports.

Assisted-by: LLM
Signed-off-by: John Crispin <john@phrozen.org>
---
 include/net/pon.h           |  14 ++
 include/net/pon/functions.h | 129 +++++++++++
 include/net/pon/types.h     | 443 ++++++++++++++++++++++++++++++++++++
 net/pon/pon.h               | 162 +++++++++++++
 net/pon/pon_state.c         | 408 +++++++++++++++++++++++++++++++++
 net/pon/pon_work.c          | 106 +++++++++
 6 files changed, 1262 insertions(+)
 create mode 100644 include/net/pon.h
 create mode 100644 include/net/pon/functions.h
 create mode 100644 include/net/pon/types.h
 create mode 100644 net/pon/pon.h
 create mode 100644 net/pon/pon_state.c
 create mode 100644 net/pon/pon_work.c

diff --git a/include/net/pon.h b/include/net/pon.h
new file mode 100644
index 000000000000..5c85f73f98f9
--- /dev/null
+++ b/include/net/pon.h
@@ -0,0 +1,14 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+/* Copyright (C) 2026 John Crispin <john@phrozen.org> */
+
+#ifndef __NET_PON_ALL_H
+#define __NET_PON_ALL_H
+
+#include <uapi/linux/pon.h>
+#include <net/pon/ploam.h>
+#include <net/pon/types.h>
+#include <net/pon/functions.h>
+
+/* Do not add any code here. Put it in the sub-headers instead. */
+
+#endif /* __NET_PON_ALL_H */
diff --git a/include/net/pon/functions.h b/include/net/pon/functions.h
new file mode 100644
index 000000000000..aa5110b47663
--- /dev/null
+++ b/include/net/pon/functions.h
@@ -0,0 +1,129 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+/* Copyright (C) 2026 John Crispin <john@phrozen.org> */
+
+#ifndef __NET_PON_FUNCTIONS_H
+#define __NET_PON_FUNCTIONS_H
+
+#include <linux/netdevice.h>
+#include <linux/rcupdate.h>
+#include <net/pon/types.h>
+
+struct pon_dev;
+struct pon_dev_caps;
+struct pon_dev_ops;
+struct sk_buff;
+
+#if IS_ENABLED(CONFIG_PON)
+
+void pon_work_queue(struct pon_dev *pdev, struct pon_work *work);
+
+/**
+ * pon_work_init() - prepare a work item
+ * @work: the item
+ * @func: the handler pon_work_queue() runs
+ *
+ * Call once before the item is first queued.
+ */
+static inline void pon_work_init(struct pon_work *work, pon_work_func_t func)
+{
+	INIT_LIST_HEAD(&work->entry);
+	work->func = func;
+}
+
+struct pon_dev *pon_dev_create(struct net_device *netdev,
+			       struct device *parent,
+			       const struct pon_dev_ops *ops,
+			       const struct pon_dev_caps *caps,
+			       enum pon_mode mode, void *priv_ptr);
+void pon_dev_unregister(struct pon_dev *pdev);
+void pon_dev_put(struct pon_dev *pdev);
+
+int pon_dev_state_report(struct pon_dev *pdev,
+			 enum pon_ploam_state state);
+__printf(2, 3) void pon_dev_log(struct pon_dev *pdev, const char *fmt, ...);
+void pon_dev_event(struct pon_dev *pdev, const struct pon_event *ev);
+
+int pon_conduit_register(struct net_device *conduit,
+			 const struct pon_conduit_ops *ops);
+void pon_conduit_unregister(struct net_device *conduit);
+int pon_conduit_rx(struct net_device *conduit, struct sk_buff *skb,
+		   const struct pon_rx_info *info);
+int pon_conduit_xmit(struct pon_dev *pdev, struct sk_buff *skb,
+		     const struct pon_tx_info *info);
+int pon_conduit_addr_set(struct pon_dev *pdev, const u8 *addr);
+int pon_conduit_mtu_set(struct pon_dev *pdev, const struct net_device *dev,
+			unsigned int mtu);
+/**
+ * netdev_uses_pon() - whether a network device belongs to a PON MAC
+ * @dev: the network device
+ *
+ * True for the PON data interface, for a GEM network device and for a
+ * conduit that is paired with a PON MAC.
+ *
+ * Return: true when @dev points at a PON instance.
+ */
+static inline bool netdev_uses_pon(const struct net_device *dev)
+{
+	return rcu_access_pointer(dev->pon_dev);
+}
+
+#else /* CONFIG_PON */
+
+/**
+ * pon_conduit_register() - offer a network device's rings to a PON MAC
+ * @conduit: the ethernet device
+ * @ops: what the driver does for the MAC
+ *
+ * The stub for a kernel without CONFIG_PON.
+ *
+ * Return: -ENOENT, there is no PON MAC to serve.
+ */
+static inline int pon_conduit_register(struct net_device *conduit,
+				       const struct pon_conduit_ops *ops)
+{
+	return -ENOENT;
+}
+
+/**
+ * pon_conduit_unregister() - take a network device's rings back
+ * @conduit: the ethernet device
+ *
+ * The stub for a kernel without CONFIG_PON. It does nothing.
+ */
+static inline void pon_conduit_unregister(struct net_device *conduit)
+{
+}
+
+/**
+ * pon_conduit_rx() - hand one received frame to the PON MAC that owns it
+ * @conduit: the ethernet device the frame arrived on
+ * @skb: the frame
+ * @info: what the receive descriptor said about it
+ *
+ * The stub for a kernel without CONFIG_PON. The frame stays the conduit's.
+ *
+ * Return: -ENODEV.
+ */
+static inline int pon_conduit_rx(struct net_device *conduit,
+				 struct sk_buff *skb,
+				 const struct pon_rx_info *info)
+{
+	return -ENODEV;
+}
+
+/**
+ * netdev_uses_pon() - whether a network device belongs to a PON MAC
+ * @dev: the network device
+ *
+ * The stub for a kernel without CONFIG_PON.
+ *
+ * Return: false.
+ */
+static inline bool netdev_uses_pon(const struct net_device *dev)
+{
+	return false;
+}
+
+#endif /* CONFIG_PON */
+
+#endif /* __NET_PON_FUNCTIONS_H */
diff --git a/include/net/pon/types.h b/include/net/pon/types.h
new file mode 100644
index 000000000000..0320c1e1dc10
--- /dev/null
+++ b/include/net/pon/types.h
@@ -0,0 +1,443 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+/* Copyright (C) 2026 John Crispin <john@phrozen.org> */
+
+#ifndef __NET_PON_TYPES_H
+#define __NET_PON_TYPES_H
+
+#include <linux/list.h>
+#include <linux/mutex.h>
+#include <linux/netdevice.h>
+#include <linux/refcount.h>
+#include <linux/skbuff.h>
+#include <linux/spinlock.h>
+#include <linux/workqueue.h>
+#include <linux/xarray.h>
+#include <net/pon/ploam.h>
+#include <uapi/linux/pon.h>
+
+struct net_device;
+struct netlink_ext_ack;
+struct pon_conduit;
+struct pon_dev;
+struct pon_work;
+struct sk_buff;
+
+#define PON_LOG_LINES		16
+#define PON_LOG_LINE_LEN	64
+
+typedef void (*pon_work_func_t)(struct pon_dev *pdev, struct pon_work *work);
+
+/**
+ * struct pon_work - a work item that runs in the instance's context
+ * @entry: link in the instance's work list
+ * @func: what to run
+ *
+ * The core lends a driver its serialized context. The handler runs on the
+ * instance's ordered workqueue with the instance lock held, so it may sleep
+ * and it cannot interleave with a netlink transaction.
+ */
+struct pon_work {
+	struct list_head entry;
+	pon_work_func_t func;
+};
+
+/* The wire defines these, so they come from the PLOAM vocabulary. */
+#define PON_SERIAL_LEN		PON_PLOAM_SN_LEN
+#define PON_REG_ID_LEN		PON_PLOAM_REG_ID_LEN
+
+/**
+ * struct pon_identity - what dev-set delivers, identity and the settings
+ *			 that travel with it
+ * @mode: requested operating mode, enum pon_mode
+ * @serial: serial number, 4 ASCII vendor id then 4 vendor specific bytes
+ * @reg_id: registration id used for XG(S)-PON authentication, padded with
+ *	    0x00 bytes to PON_REG_ID_LEN (ITU-T G.9807.1 Table C.11.25)
+ * @serial_set: @serial carries a value
+ * @mode_set: @mode carries a value
+ * @reg_id_len: PON_REG_ID_LEN when @reg_id carries a value, 0 when absent
+ */
+struct pon_identity {
+	enum pon_mode mode;
+	u8 serial[PON_SERIAL_LEN];
+	u8 reg_id[PON_REG_ID_LEN];
+	bool serial_set;
+	bool mode_set;
+	u8 reg_id_len;
+};
+
+/**
+ * struct pon_tcont_cfg - one T-CONT binding
+ * @index: T-CONT index in the device, 0 based
+ * @alloc_id: alloc-id assigned by the OLT
+ */
+struct pon_tcont_cfg {
+	u16 index;
+	u16 alloc_id;
+};
+
+/**
+ * struct pon_gem_cfg - one GEM port
+ * @id: GEM port id assigned by the OLT
+ * @dir: direction, enum pon_gem_dir
+ * @tcont_index: T-CONT the upstream half rides, valid when @tcont_valid
+ * @alloc_id: alloc-id of that T-CONT, resolved by the core, 0 when
+ *	@tcont_valid is false
+ * @key_ring: encryption key ring, enum pon_gem_key_ring (ITU-T G.988
+ *	clause 9.2.3), never broadcast
+ * @tcont_valid: @tcont_index carries a value. A downstream broadcast GEM
+ *	carries none
+ */
+struct pon_gem_cfg {
+	u16 id;
+	enum pon_gem_dir dir;
+	u16 tcont_index;
+	u16 alloc_id;
+	enum pon_gem_key_ring key_ring;
+	bool tcont_valid;
+};
+
+/**
+ * struct pon_gem_map_cfg - one upstream classifier rule
+ * @gem_id: GEM port the matching frames map to
+ * @vid: outer VLAN ID to match, valid when @vid_valid
+ * @tagged: tag state to match, valid when @tag_valid
+ * @pbit: outer VLAN priority to match, valid when @pbit_valid
+ * @dscp: IP DSCP to match, valid when @dscp_valid
+ * @tag_valid: @tagged carries a value
+ * @vid_valid: @vid carries a value
+ * @pbit_valid: @pbit carries a value
+ * @dscp_valid: @dscp carries a value
+ *
+ * A rule matches on the members marked valid. A more specific rule
+ * wins over a less specific one.
+ */
+struct pon_gem_map_cfg {
+	u16 gem_id;
+	u16 vid;
+	bool tagged;
+	u8 pbit;
+	u8 dscp;
+	bool tag_valid;
+	bool vid_valid;
+	bool pbit_valid;
+	bool dscp_valid;
+};
+
+/**
+ * enum pon_event_type - what a driver reports through pon_dev_event()
+ * @PON_EVENT_TYPE_TCONT_ALLOC: the OLT assigned an alloc-id with the
+ *	Assign_Alloc-ID message of ITU-T G.9807.1 clause C.11.3.3.7 and the
+ *	driver bound it to a transmit channel
+ * @PON_EVENT_TYPE_TCONT_DEALLOC: the OLT deallocated the alloc-id with the
+ *	same message and the driver released its channel. The GEM ports that
+ *	ride it carry no upstream traffic until the OLT assigns it again
+ */
+enum pon_event_type {
+	PON_EVENT_TYPE_TCONT_ALLOC,
+	PON_EVENT_TYPE_TCONT_DEALLOC,
+};
+
+/**
+ * struct pon_event - a discrete event a driver reports
+ * @type: what happened, enum pon_event_type
+ * @alloc_id: the alloc-id the OLT allocated or deallocated
+ */
+struct pon_event {
+	enum pon_event_type type;
+	u32 alloc_id;
+};
+
+/**
+ * struct pon_dev_caps - what the device supports
+ * @modes: bitmask of enum pon_mode the hardware can run
+ * @max_tconts: number of T-CONTs
+ * @max_gems: number of GEM ports userspace can create, not counting the
+ *	      GEM port of the OMCC, which the driver holds itself
+ */
+struct pon_dev_caps {
+	u32 modes;
+	u16 max_tconts;
+	u16 max_gems;
+};
+
+/**
+ * struct pon_tx_info - what the conduit needs to send one frame
+ * @gem: the GEM port id the frame goes out on
+ * @channel: the transmit channel the MAC driver bound the T-CONT to. The
+ *	     core carries it between the two drivers and does not interpret it
+ * @queue: the queue within that channel
+ * @mic_index: which OMCI integrity key signs the PDU, valid when @oam
+ * @oam: the frame is an OMCI PDU rather than an Ethernet frame
+ */
+struct pon_tx_info {
+	u16 gem;
+	u8 channel;
+	u8 queue;
+	u8 mic_index;
+	bool oam;
+};
+
+/**
+ * struct pon_rx_info - what the conduit learned from the receive descriptor
+ * @gem: the GEM port id the frame arrived on
+ * @oam: the frame is an OMCI PDU
+ * @mic_unchecked: the MAC did not verify the integrity of the PDU, valid when
+ *		   @oam
+ */
+struct pon_rx_info {
+	u16 gem;
+	bool oam;
+	bool mic_unchecked;
+};
+
+/**
+ * struct pon_conduit_ops - what the ethernet driver does for a PON MAC
+ * @xmit: queue one frame on the conduit's rings with a descriptor built from
+ *	  @info. Owns the skb from the call on (as ndo_start_xmit does) and
+ *	  returns NETDEV_TX_BUSY only when the frame was not taken. Called
+ *	  under rcu_read_lock() with bottom halves disabled, from every PON
+ *	  network device and the OMCI channel at once and without the
+ *	  conduit's own transmit lock, so the driver locks its ring itself.
+ */
+struct pon_conduit_ops {
+	netdev_tx_t (*xmit)(struct net_device *conduit, struct sk_buff *skb,
+			    const struct pon_tx_info *info);
+};
+
+/**
+ * struct pon_dev - one PON MAC
+ * @main_netdev: the PON data network device
+ * @conduit: the ethernet device whose rings carry the frames and what its
+ *	     driver does for the MAC, published as one pointer once the driver
+ *	     has registered it, read under RCU, written under pon_devs_lock
+ * @conduit_tracker: the reference held on the ethernet device of @conduit
+ * @parent: the MAC's device, which every interface the core creates
+ *	    parents onto and whose firmware node a conduit names
+ * @ops: driver callbacks, NULL once the driver has unregistered. The
+ *	 transmit paths read it once under RCU
+ * @caps: device capabilities
+ * @drv_priv: driver priv pointer
+ * @lock: instance lock and the driver's upcalls. It protects every field
+ *	  below it, with these exceptions. @conduit is @pon_devs_lock's. The
+ *	  OMCI fields name their own rules. @ploam and @going_away are
+ *	  written under the lock and read without it through READ_ONCE()
+ * @refcnt: reference count for the instance
+ * @id: instance id
+ * @mode: active mode, enum pon_mode
+ * @ploam: activation state as the driver last reported it
+ * @enabled: the upstream link was last started rather than stopped
+ * @omci_portid: netlink port id of the socket that owns the OMCI channel, 0
+ *		 when none does, changed with cmpxchg, because a socket that
+ *		 closes gives it up without the lock
+ * @omci_rxq: OMCI PDUs from the OLT in the order they arrived, waiting for
+ *	      the instance's context. Each one is marked when the MAC passed
+ *	      it up unchecked
+ * @omci_rx_work: has the driver verify the unchecked PDUs of @omci_rxq and
+ *		  hands every PDU to the owner of the OMCI channel
+ * @identity: the serial number as dev-set last delivered it, which dev-get
+ *	      reports. The registration id goes to the driver and is not kept
+ * @identity.serial: serial number, valid when @identity.serial_set
+ * @identity.serial_set: dev-set delivered a serial number
+ * @tconts: bound T-CONTs, struct pon_tcont
+ * @gems: GEM port objects, struct pon_gem
+ * @gem_maps: upstream classifier rules, struct pon_gem_map
+ * @gem_netdevs: GEM port id to its network device, read under RCU on receive
+ * @wq: ordered workqueue, the instance's one deferred context
+ * @work: the single worker that drains @work_list
+ * @work_list: work items waiting for the instance's context
+ * @work_lock: guards @work_list, taken from hard interrupt
+ * @log_work: prints the lines of @log_lines on system_dfl_wq, outside the
+ *	      instance's context
+ * @log_lines: the lines pon_dev_log() recorded and @log_work has not printed
+ *	       yet, a ring that starts at @log_head
+ * @log_levels: the printk level of each line of @log_lines, such as KERN_INFO
+ * @log_head: the oldest line of @log_lines
+ * @log_count: the lines held in @log_lines
+ * @log_dropped: the oldest lines a full @log_lines dropped since @log_work
+ *		 last printed the count
+ * @going_away: set on the unregister path, so a work item already past its
+ *		scheduling point does not reach a driver that is leaving
+ * @rcu: RCU head for freeing the structure
+ */
+struct pon_dev {
+	struct net_device *main_netdev;
+	struct pon_conduit __rcu *conduit;
+	netdevice_tracker conduit_tracker;
+	struct device *parent;
+
+	const struct pon_dev_ops *ops;
+	const struct pon_dev_caps *caps;
+	void *drv_priv;
+
+	/* guards every field below and the driver's upcalls */
+	struct mutex lock;
+	refcount_t refcnt;
+
+	u32 id;
+	enum pon_mode mode;
+	enum pon_ploam_state ploam;
+	bool enabled;
+	struct {
+		u8 serial[PON_SERIAL_LEN];
+		bool serial_set;
+	} identity;
+
+	struct list_head tconts;
+	struct list_head gems;
+	struct list_head gem_maps;
+	struct xarray gem_netdevs;
+
+	struct workqueue_struct *wq;
+	struct work_struct work;
+	struct list_head work_list;
+	/* guards @work_list, taken from hard interrupt */
+	spinlock_t work_lock;
+
+	u32 omci_portid;
+	struct sk_buff_head omci_rxq;
+	struct pon_work omci_rx_work;
+	struct work_struct log_work;
+	char log_lines[PON_LOG_LINES][PON_LOG_LINE_LEN];
+	const char *log_levels[PON_LOG_LINES];
+	unsigned int log_head;
+	unsigned int log_count;
+	unsigned int log_dropped;
+	bool going_away;
+
+	struct rcu_head rcu;
+};
+
+/**
+ * struct pon_dev_ops - netdev driver facing PON callbacks
+ *
+ * tcont_set, tcont_clear, gem_add, gem_del and omci_xmit are mandatory. The
+ * others may be NULL and the core then answers -EOPNOTSUPP.
+ *
+ * Every callback runs in process context and may sleep, except where its own
+ * description says otherwise. The ones that configure or read the device run
+ * with the instance lock held, from a netlink handler or from the instance's
+ * work, so a driver sees them one at a time and never interleaved with its
+ * own pon_work handlers. The datapath callbacks run without the lock.
+ */
+struct pon_dev_ops {
+	/**
+	 * @enable: start or stop the ONU upstream link
+	 * Instance lock held.
+	 */
+	int (*enable)(struct pon_dev *pdev, bool on,
+		      struct netlink_ext_ack *extack);
+
+	/**
+	 * @set_identity: set the ONU identity and the settings that travel
+	 *		  with it
+	 * Only the members whose _set flag or length is nonzero changed.
+	 * Instance lock held.
+	 */
+	int (*set_identity)(struct pon_dev *pdev,
+			    const struct pon_identity *id,
+			    struct netlink_ext_ack *extack);
+
+	/**
+	 * @tcont_set: bind an alloc-id to a T-CONT, or rebind it
+	 * Instance lock held.
+	 */
+	int (*tcont_set)(struct pon_dev *pdev,
+			 const struct pon_tcont_cfg *cfg,
+			 struct netlink_ext_ack *extack);
+
+	/**
+	 * @tcont_clear: release a T-CONT
+	 * The core passes the stored binding. Instance lock held.
+	 */
+	int (*tcont_clear)(struct pon_dev *pdev,
+			   const struct pon_tcont_cfg *cfg,
+			   struct netlink_ext_ack *extack);
+
+	/**
+	 * @tcont_channel: the conduit transmit channel a T-CONT is bound to,
+	 *		   optional. The core asks it for the carrier, which is
+	 *		   up while a GEM port rides a T-CONT with a channel.
+	 *		   Without it the carrier rises with the first GEM
+	 *		   port. Return the channel, -ENOLINK while the OLT has
+	 *		   not assigned the alloc-id and the T-CONT has no
+	 *		   channel yet, or another negative errno. The default
+	 *		   alloc-id has a channel once the ONU-ID is assigned.
+	 *		   Instance lock held. rtnl is held by some callers and
+	 *		   not by others, so the driver must not rely on it.
+	 */
+	int (*tcont_channel)(struct pon_dev *pdev,
+			     const struct pon_tcont_cfg *cfg);
+
+	/**
+	 * @gem_add: create a GEM port
+	 *
+	 * Must also succeed for a GEM port the driver already holds and
+	 * re-program it, which a T-CONT that moves to another alloc-id
+	 * needs. The driver keeps a GEM port across a loss of the link and
+	 * a new activation, as ITU-T G.9807.1 clause C.6.1.5.8 has the ONU
+	 * keep it. A GEM port whose alloc-id has no channel yet is held and
+	 * carries nothing until the OLT assigns the alloc-id. Instance lock
+	 * held.
+	 */
+	int (*gem_add)(struct pon_dev *pdev, const struct pon_gem_cfg *cfg,
+		       struct netlink_ext_ack *extack);
+
+	/**
+	 * @gem_del: destroy a GEM port
+	 * Instance lock held.
+	 */
+	int (*gem_del)(struct pon_dev *pdev, u16 gem_id,
+		       struct netlink_ext_ack *extack);
+
+	/**
+	 * @gem_xmit: send one frame on a GEM port, optional. Consumes the skb
+	 *	      whatever it returns. Without it a GEM network device
+	 *	      drops what it is given. Runs from ndo_start_xmit of the
+	 *	      GEM network device, without the instance lock. A
+	 *	      concurrent gem_del is the driver's to order.
+	 */
+	int (*gem_xmit)(struct pon_dev *pdev, u16 gem_id, struct sk_buff *skb);
+
+	/**
+	 * @omci_xmit: send one OMCI PDU to the OLT
+	 * The skb carries the bare PDU, validated by the core. The driver
+	 * consumes the skb on success and on failure. Runs from the omci-tx
+	 * netlink handler, with the instance lock held and bottom halves
+	 * disabled.
+	 */
+	int (*omci_xmit)(struct pon_dev *pdev, struct sk_buff *skb);
+
+	/**
+	 * @omci_verify: check the integrity of a received OMCI PDU that the
+	 *		 MAC passed up unchecked, optional. The skb is linear
+	 *		 and ends with the 4 byte MIC, which the driver checks
+	 *		 with its OMCI integrity key and strips on success
+	 *		 (ITU-T G.9807.1 clause C.15.7.2). The check is the
+	 *		 driver's because the key never reaches the core.
+	 *		 Return 0 to deliver it, a negative errno to drop it.
+	 *		 Without it such a PDU is dropped. Runs in the
+	 *		 instance's context with the lock held and may sleep.
+	 */
+	int (*omci_verify)(struct pon_dev *pdev, struct sk_buff *skb);
+
+	/**
+	 * @gem_map_set: add one upstream classifier rule. Idempotent: a rule
+	 *		 the driver already holds is success, not -EEXIST. A
+	 *		 rule outlives the GEM port it names.
+	 *		 Instance lock held.
+	 */
+	int (*gem_map_set)(struct pon_dev *pdev,
+			   const struct pon_gem_map_cfg *cfg,
+			   struct netlink_ext_ack *extack);
+
+	/**
+	 * @gem_map_del: remove one upstream classifier rule
+	 * Instance lock held.
+	 */
+	int (*gem_map_del)(struct pon_dev *pdev,
+			   const struct pon_gem_map_cfg *cfg,
+			   struct netlink_ext_ack *extack);
+
+};
+
+#endif /* __NET_PON_TYPES_H */
diff --git a/net/pon/pon.h b/net/pon/pon.h
new file mode 100644
index 000000000000..daf680790698
--- /dev/null
+++ b/net/pon/pon.h
@@ -0,0 +1,162 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+/* Copyright (C) 2026 John Crispin <john@phrozen.org> */
+
+#ifndef __PON_PON_H
+#define __PON_PON_H
+
+#include <linux/list.h>
+#include <linux/lockdep.h>
+#include <linux/mutex.h>
+#include <linux/workqueue.h>
+#include <linux/xarray.h>
+#include <net/pon.h>
+
+#define PON_GEM_KIND		"gem"
+
+extern struct xarray pon_devs;
+extern struct mutex pon_devs_lock;
+
+/**
+ * struct pon_tcont - one T-CONT of an instance
+ * @list:	entry in the instance's @tconts list
+ * @cfg:	the configuration tcont-set last stored
+ */
+struct pon_tcont {
+	struct list_head list;
+	struct pon_tcont_cfg cfg;
+};
+
+/**
+ * struct pon_gem - one GEM port of an instance
+ * @list:	entry in the instance's @gems list
+ * @cfg:	the configuration gem-new stored
+ * @gem_netdev:	the GEM port's network device, or NULL when it has none
+ */
+struct pon_gem {
+	struct list_head list;
+	struct pon_gem_cfg cfg;
+	struct net_device *gem_netdev;
+};
+
+/**
+ * struct pon_gem_map - one upstream classifier rule
+ * @list:	entry in the instance's @gem_maps list
+ * @cfg:	what the rule matches on and the GEM port it selects
+ *
+ * One upstream classifier rule, kept for the dumps and the notifications.
+ */
+struct pon_gem_map {
+	struct list_head list;
+	struct pon_gem_map_cfg cfg;
+};
+
+struct pon_tcont *pon_tcont_find(struct pon_dev *pdev, u16 index);
+struct pon_gem *pon_gem_find(struct pon_dev *pdev, u16 gem_id);
+int pon_gem_channel(struct pon_dev *pdev, const struct pon_gem *gem);
+struct pon_gem_map *pon_gem_map_find(struct pon_dev *pdev,
+				     const struct pon_gem_map_cfg *cfg);
+bool pon_tcont_in_use(struct pon_dev *pdev, u16 index);
+bool pon_gems_full(struct pon_dev *pdev);
+bool pon_tcont_alloc_taken(struct pon_dev *pdev, u16 index, u16 alloc_id);
+int pon_tcont_gems_rebind(struct pon_dev *pdev, u16 index, u16 old_alloc_id,
+			  u16 new_alloc_id, struct netlink_ext_ack *extack);
+
+void pon_work_worker(struct work_struct *work);
+void pon_work_drain(struct pon_dev *pdev);
+
+void pon_nl_obj_gen_inc(void);
+void pon_nl_notify_dev(struct pon_dev *pdev, u32 cmd);
+void pon_nl_notify_ploam(struct pon_dev *pdev);
+void pon_nl_notify_tcont(struct pon_dev *pdev, struct pon_tcont *tcont,
+			 u32 cmd);
+void pon_nl_notify_gem(struct pon_dev *pdev, struct pon_gem *gem, u32 cmd);
+void pon_nl_notify_gem_map(struct pon_dev *pdev, struct pon_gem_map *map,
+			   u32 cmd);
+
+void pon_dev_carrier_update(struct pon_dev *pdev);
+
+/* ITU-T G.988 clause 11.2.5 and Table 11.2-2: header and length are 10
+ * bytes and a PDU is at most 1980 bytes including the 4 byte MIC.
+ */
+#define PON_OMCI_MIN_LEN	10
+#define PON_OMCI_MAX_LEN	1976
+#define PON_OMCI_MIC_LEN	4
+#define PON_OMCI_QUEUE_MAX	512
+
+void pon_omci_init(struct pon_dev *pdev);
+void pon_omci_destroy(struct pon_dev *pdev);
+int pon_omci_conduit_rx(struct pon_dev *pdev, struct sk_buff *skb,
+			bool unverified);
+int pon_omci_xmit(struct pon_dev *pdev, const void *pdu, unsigned int len,
+		  struct netlink_ext_ack *extack);
+int pon_omci_register(struct pon_dev *pdev, u32 portid);
+int pon_omci_notifier_register(void);
+void pon_omci_notifier_unregister(void);
+int pon_nl_omci_ntf(struct pon_dev *pdev, const struct sk_buff *skb);
+
+bool pon_conduit_attach(struct pon_dev *pdev);
+void pon_conduit_detach(struct pon_dev *pdev);
+void pon_conduit_sync(struct pon_dev *pdev);
+bool pon_conduit_running(struct pon_dev *pdev);
+void pon_log_init(struct pon_dev *pdev);
+__printf(3, 4) void pon_dev_log_level(struct pon_dev *pdev, const char *level,
+				      const char *fmt, ...);
+int pon_conduit_notifier_register(void);
+void pon_conduit_notifier_unregister(void);
+struct net_device *pon_conduit_hold(struct pon_dev *pdev,
+				    const struct pon_conduit_ops **ops,
+				    netdevice_tracker *tracker);
+
+int pon_gem_link_register(void);
+void pon_gem_link_unregister(void);
+void pon_gem_netdevs_unregister(struct pon_dev *pdev);
+int pon_gem_netdev_id(const struct net_device *dev, u16 *gem_id);
+
+/**
+ * pon_dev_get() - take a reference to a PON device
+ * @pdev:	PON device structure, on which the caller already holds a
+ *		reference or which it found under pon_devs_lock
+ *
+ * The reference is dropped with pon_dev_put().
+ *
+ * Context: Any context.
+ */
+static inline void pon_dev_get(struct pon_dev *pdev)
+{
+	refcount_inc(&pdev->refcnt);
+}
+
+/**
+ * pon_dev_tryget() - take a reference to a PON device that may be going
+ * @pdev:	PON device structure, found under RCU
+ *
+ * Fails once the last reference is gone and the device waits to be freed.
+ *
+ * Context: Any context.
+ * Return: true when the reference was taken, false otherwise.
+ */
+static inline bool pon_dev_tryget(struct pon_dev *pdev)
+{
+	return refcount_inc_not_zero(&pdev->refcnt);
+}
+
+/**
+ * pon_dev_is_registered() - test that a PON device is still registered
+ * @pdev:	PON device structure
+ *
+ * @pdev->ops survives until the end of pon_dev_unregister(), which frees the
+ * instance's objects long before it. @pdev->going_away is set first, so both
+ * are tested: a caller that only asked about @pdev->ops could act on a T-CONT
+ * or a GEM port that has already been freed.
+ *
+ * Context: Called with @pdev->lock held.
+ * Return: true while the driver is registered and pon_dev_unregister() has
+ * not begun, false otherwise.
+ */
+static inline bool pon_dev_is_registered(struct pon_dev *pdev)
+{
+	lockdep_assert_held(&pdev->lock);
+	return pdev->ops && !pdev->going_away;
+}
+
+#endif /* __PON_PON_H */
diff --git a/net/pon/pon_state.c b/net/pon/pon_state.c
new file mode 100644
index 000000000000..5e06df5b072c
--- /dev/null
+++ b/net/pon/pon_state.c
@@ -0,0 +1,408 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/* Copyright (C) 2026 John Crispin <john@phrozen.org> */
+
+#include <linux/bits.h>
+#include <linux/netdevice.h>
+#include <linux/sprintf.h>
+#include <linux/stdarg.h>
+#include <linux/string.h>
+#include <linux/workqueue.h>
+#include <net/pon.h>
+
+#include "pon.h"
+
+#define PON_S(x)	BIT(PON_PLOAM_STATE_##x)
+
+/* The activation edges the ITU-T state machine permits, indexed by the state
+ * being left.
+ *
+ * One table serves every mode. G.984.3 ranges in O3 where G.9807.1 merges O2
+ * and O3 into one state and a MAC has one register code for the merged pair,
+ * so a driver reports O2 and then O4. Both edges are legal here because what
+ * is being checked is whether a driver reported something impossible, not
+ * whether it conforms to one mode's text. A per mode split is worth adding
+ * the day a second MAC needs it.
+ *
+ * O1 is reachable from every state: a loss of signal, a deactivation or the
+ * end of TO2 returns the ONU to it. O7 is reachable from O1 to O5, where the
+ * OLT's Disable_Serial_Number stops the ONU, but not from O6, which leaves
+ * only for O5 or O1. O7 itself leaves only for O1, when the OLT enables the
+ * ONU again (G.9807.1 Table C.12.4).
+ */
+static const u32 pon_state_legal[] = {
+	[PON_PLOAM_STATE_UNKNOWN] = PON_S(O1) | PON_S(O2) | PON_S(O3) |
+				    PON_S(O4) | PON_S(O5) | PON_S(O6) |
+				    PON_S(O7),
+	[PON_PLOAM_STATE_O1]	  = PON_S(O2) | PON_S(O7),
+	[PON_PLOAM_STATE_O2]	  = PON_S(O1) | PON_S(O3) | PON_S(O4) |
+				    PON_S(O7),
+	[PON_PLOAM_STATE_O3]	  = PON_S(O1) | PON_S(O4) | PON_S(O7),
+	[PON_PLOAM_STATE_O4]	  = PON_S(O1) | PON_S(O2) | PON_S(O5) |
+				    PON_S(O7),
+	[PON_PLOAM_STATE_O5]	  = PON_S(O1) | PON_S(O6) | PON_S(O7),
+	[PON_PLOAM_STATE_O6]	  = PON_S(O1) | PON_S(O5),
+	[PON_PLOAM_STATE_O7]	  = PON_S(O1),
+};
+
+static const char *const pon_state_names[] = {
+	[PON_PLOAM_STATE_UNKNOWN] = "unknown",
+	[PON_PLOAM_STATE_O1]	  = "O1",
+	[PON_PLOAM_STATE_O2]	  = "O2",
+	[PON_PLOAM_STATE_O3]	  = "O3",
+	[PON_PLOAM_STATE_O4]	  = "O4",
+	[PON_PLOAM_STATE_O5]	  = "O5",
+	[PON_PLOAM_STATE_O6]	  = "O6",
+	[PON_PLOAM_STATE_O7]	  = "O7",
+};
+
+/**
+ * pon_state_edge_legal() - test an activation edge against the state machine
+ * @from:	the state being left, enum pon_ploam_state
+ * @to:		the state being entered, enum pon_ploam_state
+ *
+ * The edges follow the ONU activation cycle state transition table, ITU-T
+ * G.9807.1 clause C.12.1.4.3, Table C.12.4, over the states of Table C.12.1.
+ *
+ * Return: true when pon_state_legal permits the edge, false otherwise or
+ * when @from is out of range.
+ */
+static bool pon_state_edge_legal(u32 from, u32 to)
+{
+	if (from >= ARRAY_SIZE(pon_state_legal))
+		return false;
+
+	return !!(pon_state_legal[from] & BIT(to));
+}
+
+/**
+ * pon_dev_state_name() - the name of an activation state in the device's mode
+ * @pdev:	PON device structure
+ * @state:	the state, enum pon_ploam_state, in range
+ *
+ * ITU-T G.9807.1 Table C.12.1 has one Serial Number state, O2-3, which a
+ * driver reports as O2. Only a G-PON device has an O2 of its own.
+ *
+ * Context: Called with @pdev->lock held.
+ * Return: "O2-3" for O2 in a mode other than G-PON, otherwise the name in
+ * pon_state_names.
+ */
+static const char *pon_dev_state_name(const struct pon_dev *pdev, u32 state)
+{
+	if (state == PON_PLOAM_STATE_O2 && pdev->mode != PON_MODE_GPON)
+		return "O2-3";
+
+	return pon_state_names[state];
+}
+
+/**
+ * pon_netdev_carrier_set() - set the carrier of one network device
+ * @dev:	the network device
+ * @up:		true for carrier on, false for carrier off
+ */
+static void pon_netdev_carrier_set(struct net_device *dev, bool up)
+{
+	if (up)
+		netif_carrier_on(dev);
+	else
+		netif_carrier_off(dev);
+}
+
+/**
+ * pon_dev_gem_bound() - test whether a GEM port can carry traffic
+ * @pdev:	PON device structure
+ *
+ * Without the driver's tcont_channel callback, any GEM port is enough.
+ *
+ * Context: Called with @pdev->lock held.
+ * Return: true when a GEM port rides a T-CONT that has a transmit channel,
+ * false otherwise.
+ */
+static bool pon_dev_gem_bound(struct pon_dev *pdev)
+{
+	struct pon_gem *gem;
+
+	if (list_empty(&pdev->gems))
+		return false;
+
+	if (!pdev->ops->tcont_channel)
+		return true;
+
+	list_for_each_entry(gem, &pdev->gems, list)
+		if (pon_gem_channel(pdev, gem) >= 0)
+			return true;
+
+	return false;
+}
+
+/**
+ * pon_dev_carrier_update() - set the carrier of the PON netdevs from the state
+ * @pdev:	PON device structure
+ *
+ * The carrier is up in O5 or O6 while a conduit that is up is paired with
+ * the instance and at least one GEM port rides a T-CONT that has a transmit
+ * channel, as the driver's tcont_channel callback answers. That includes
+ * the default alloc-id, which carries data too and which the driver binds
+ * without an alloc-id event. Without the callback a GEM port is enough.
+ * The GEM network devices follow the data netdev.
+ *
+ * Called on every change that can alter the answer: an activation edge, an
+ * alloc-id event, a T-CONT or a GEM port that is set or deleted and a
+ * conduit that pairs, goes up, goes down or leaves.
+ */
+void pon_dev_carrier_update(struct pon_dev *pdev)
+{
+	struct pon_gem *gem;
+	bool up;
+
+	lockdep_assert_held(&pdev->lock);
+
+	up = (pdev->ploam == PON_PLOAM_STATE_O5 ||
+	      pdev->ploam == PON_PLOAM_STATE_O6) && pon_dev_gem_bound(pdev) &&
+	     pon_conduit_running(pdev);
+
+	pon_netdev_carrier_set(pdev->main_netdev, up);
+	list_for_each_entry(gem, &pdev->gems, list)
+		if (gem->gem_netdev)
+			pon_netdev_carrier_set(gem->gem_netdev, up);
+}
+
+/**
+ * pon_dev_log_pop() - take the oldest recorded line out of the ring
+ * @pdev:	PON device structure
+ * @line:	filled with the line, PON_LOG_LINE_LEN bytes
+ * @level:	set to the printk level of the line
+ * @dropped:	set to the lines a full ring dropped before this one
+ *
+ * Context: Process context. Takes @pdev->lock.
+ * Return: true when @line holds a line, false when the ring is empty.
+ */
+static bool pon_dev_log_pop(struct pon_dev *pdev, char *line,
+			    const char **level, unsigned int *dropped)
+{
+	bool found;
+
+	mutex_lock(&pdev->lock);
+	*dropped = pdev->log_dropped;
+	pdev->log_dropped = 0;
+	found = pdev->log_count;
+	if (found) {
+		strscpy(line, pdev->log_lines[pdev->log_head],
+			PON_LOG_LINE_LEN);
+		*level = pdev->log_levels[pdev->log_head];
+		pdev->log_head = (pdev->log_head + 1) % PON_LOG_LINES;
+		pdev->log_count--;
+	}
+	mutex_unlock(&pdev->lock);
+
+	return found;
+}
+
+/**
+ * pon_dev_log_work() - print the lines pon_dev_log() recorded
+ * @work:	the instance's @log_work
+ *
+ * Prints the lines in the order they were recorded, each at its own level
+ * with the name of the data network device. Where a full ring dropped lines,
+ * one line at info level with their count comes first, because they were
+ * older. The instance lock is held only to take a line out, never across a
+ * print.
+ *
+ * Context: Process context, on system_dfl_wq. Takes @pdev->lock.
+ */
+static void pon_dev_log_work(struct work_struct *work)
+{
+	struct pon_dev *pdev = container_of(work, struct pon_dev, log_work);
+	char line[PON_LOG_LINE_LEN];
+	unsigned int dropped;
+	const char *level;
+	bool found;
+
+	do {
+		found = pon_dev_log_pop(pdev, line, &level, &dropped);
+		if (dropped)
+			netdev_info(pdev->main_netdev,
+				    "%u log lines dropped\n", dropped);
+		if (found)
+			netdev_printk(level, pdev->main_netdev, "%s\n", line);
+	} while (found);
+}
+
+/**
+ * pon_log_init() - prepare the log of a new PON device
+ * @pdev:	PON device structure
+ *
+ * Context: From pon_dev_create(), before the device is published.
+ */
+void pon_log_init(struct pon_dev *pdev)
+{
+	INIT_WORK(&pdev->log_work, pon_dev_log_work);
+}
+
+/**
+ * pon_dev_log_record() - record one line for the log work
+ * @pdev:	PON device structure
+ * @level:	printk level of the line, such as KERN_INFO
+ * @fmt:	printf format of the line, without a newline
+ * @args:	the arguments of @fmt
+ *
+ * Context: Called with @pdev->lock held.
+ */
+static __printf(3, 0) void pon_dev_log_record(struct pon_dev *pdev,
+					      const char *level,
+					      const char *fmt, va_list args)
+{
+	unsigned int slot;
+
+	lockdep_assert_held(&pdev->lock);
+
+	if (pdev->going_away)
+		return;
+
+	if (pdev->log_count == PON_LOG_LINES) {
+		pdev->log_head = (pdev->log_head + 1) % PON_LOG_LINES;
+		pdev->log_count--;
+		pdev->log_dropped++;
+	}
+
+	slot = (pdev->log_head + pdev->log_count) % PON_LOG_LINES;
+	vsnprintf(pdev->log_lines[slot], PON_LOG_LINE_LEN, fmt, args);
+	pdev->log_levels[slot] = level;
+	pdev->log_count++;
+
+	queue_work(system_dfl_wq, &pdev->log_work);
+}
+
+/**
+ * pon_dev_log_level() - log one line at a given level from the core
+ * @pdev:	PON device structure
+ * @level:	printk level of the line, such as KERN_ERR
+ * @fmt:	printf format of the line, without a newline
+ *
+ * As pon_dev_log(), with the level of the line chosen by the caller.
+ *
+ * Context: Called with @pdev->lock held.
+ */
+void pon_dev_log_level(struct pon_dev *pdev, const char *level,
+		       const char *fmt, ...)
+{
+	va_list args;
+
+	va_start(args, fmt);
+	pon_dev_log_record(pdev, level, fmt, args);
+	va_end(args);
+}
+
+/**
+ * pon_dev_log() - log one line from the instance's context
+ * @pdev:	PON device structure
+ * @fmt:	printf format of the line, without a newline
+ *
+ * Records the line and leaves the print to a work item on system_dfl_wq. A
+ * print to a serial console blocks for milliseconds, while the instance's
+ * context answers the OLT within a deadline. The caller never waits for it.
+ * The lines are printed at info level with the name of the data network
+ * device, in the order they were recorded, the activation edges of
+ * pon_dev_state_report() among them. The ring holds PON_LOG_LINES lines.
+ * When it is full the oldest line is dropped and counted. The count is
+ * printed in its place. A line is cut to PON_LOG_LINE_LEN - 1 characters.
+ * Nothing is recorded once pon_dev_unregister() has begun.
+ *
+ * Context: Called with @pdev->lock held.
+ */
+void pon_dev_log(struct pon_dev *pdev, const char *fmt, ...)
+{
+	va_list args;
+
+	va_start(args, fmt);
+	pon_dev_log_record(pdev, KERN_INFO, fmt, args);
+	va_end(args);
+}
+EXPORT_SYMBOL_GPL(pon_dev_log);
+
+/**
+ * pon_dev_state_report() - report the activation state the MAC reached
+ * @pdev:	PON device structure
+ * @state:	the new state, enum pon_ploam_state
+ *
+ * The driver computes the state. The core owns, validates and publishes it.
+ * Call with @pdev->lock held, from the instance's work or from a pon_dev_ops
+ * handler.
+ *
+ * A repeat of the state already published is a no-op: this is a level, not an
+ * edge. An edge the standard does not permit is published anyway, because the
+ * hardware is the truth and a core that refused it would publish a state the
+ * ONU is not in. It is warned about instead. Every edge that is published
+ * gets one line in the kernel log at info level through pon_dev_log().
+ *
+ * Return: 0, -ENODEV once pon_dev_unregister() has begun, or -EINVAL for a
+ * value out of range or an illegal edge. An illegal edge is published anyway
+ * and the return value is for the driver author and for a selftest. A value
+ * out of range is not a state and is refused before anything is published.
+ */
+int pon_dev_state_report(struct pon_dev *pdev, enum pon_ploam_state state)
+{
+	u32 old = pdev->ploam;
+	int err = 0;
+
+	lockdep_assert_held(&pdev->lock);
+
+	if (pdev->going_away)
+		return -ENODEV;
+
+	if (state > PON_PLOAM_STATE_O7) {
+		netdev_warn(pdev->main_netdev,
+			    "activation state %u is not a state\n", state);
+		return -EINVAL;
+	}
+
+	if (state == old)
+		return 0;
+
+	if (!pon_state_edge_legal(old, state)) {
+		net_warn_ratelimited("%s: activation went O%u to O%u, which the standard does not permit\n",
+				     netdev_name(pdev->main_netdev), old,
+				     state);
+		err = -EINVAL;
+	}
+
+	WRITE_ONCE(pdev->ploam, state);
+	pon_dev_log(pdev, "PLOAM state %s -> %s",
+		    pon_dev_state_name(pdev, old),
+		    pon_dev_state_name(pdev, state));
+
+	/* Carrier before the notification, so a daemon that reads both sees a
+	 * netdev that agrees with the state. rtnetlink and generic netlink
+	 * have no ordering between them, so this is best effort.
+	 */
+	pon_dev_carrier_update(pdev);
+	pon_nl_notify_ploam(pdev);
+
+	return err;
+}
+EXPORT_SYMBOL_GPL(pon_dev_state_report);
+
+/**
+ * pon_dev_event() - report a discrete event
+ * @pdev:	PON device structure
+ * @ev:		what happened and the alloc-id it names
+ *
+ * Runs in the instance's context, with its lock held.
+ *
+ * A PON_EVENT_TYPE_TCONT_ALLOC event means the driver bound the alloc-id to
+ * a channel, so the GEM ports that ride it carry traffic from then on: the
+ * carrier may rise. A PON_EVENT_TYPE_TCONT_DEALLOC event means the driver
+ * released the channel: the carrier may fall. An event that arrives once
+ * pon_dev_unregister() has begun is dropped.
+ */
+void pon_dev_event(struct pon_dev *pdev, const struct pon_event *ev)
+{
+	lockdep_assert_held(&pdev->lock);
+
+	if (pdev->going_away)
+		return;
+
+	if (ev->type == PON_EVENT_TYPE_TCONT_ALLOC ||
+	    ev->type == PON_EVENT_TYPE_TCONT_DEALLOC)
+		pon_dev_carrier_update(pdev);
+}
+EXPORT_SYMBOL_GPL(pon_dev_event);
diff --git a/net/pon/pon_work.c b/net/pon/pon_work.c
new file mode 100644
index 000000000000..59264cec2b41
--- /dev/null
+++ b/net/pon/pon_work.c
@@ -0,0 +1,106 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/* Copyright (C) 2026 John Crispin <john@phrozen.org> */
+
+#include <linux/list.h>
+#include <linux/spinlock.h>
+#include <linux/workqueue.h>
+#include <net/pon.h>
+
+#include "pon.h"
+
+/**
+ * DOC: The lent context
+ *
+ * A PON MAC runs the ITU-T activation state machine itself, but it must not
+ * run it in hard interrupt context and it must not run it against a half
+ * finished netlink transaction. The core therefore lends the driver its own
+ * serialized context: one ordered workqueue per instance and one lock that
+ * the netlink handlers take in their pre_doit.
+ *
+ * A driver queues a pon_work from any context, hard interrupt included. The
+ * handler runs with pdev->lock held, so it may sleep.
+ *
+ * One worker drains the list one item at a time and re-queues itself while
+ * the list is not empty, so the ordering a driver sees is the order it
+ * queued in.
+ */
+
+void pon_work_worker(struct work_struct *work)
+{
+	struct pon_dev *pdev = container_of(work, struct pon_dev, work);
+	struct pon_work *item;
+
+	mutex_lock(&pdev->lock);
+
+	if (pdev->going_away) {
+		mutex_unlock(&pdev->lock);
+		return;
+	}
+
+	spin_lock_irq(&pdev->work_lock);
+	item = list_first_entry_or_null(&pdev->work_list, struct pon_work,
+					entry);
+	if (item) {
+		list_del_init(&item->entry);
+		if (!list_empty(&pdev->work_list))
+			queue_work(pdev->wq, work);
+	}
+	spin_unlock_irq(&pdev->work_lock);
+
+	if (item)
+		item->func(pdev, item);
+
+	mutex_unlock(&pdev->lock);
+}
+
+/**
+ * pon_work_queue() - run a work item in the instance's context
+ * @pdev: PON device structure
+ * @work: the item, initialized with pon_work_init()
+ *
+ * Safe from any context, hard interrupt included. Queueing an item that is
+ * already queued and has not yet run does nothing, so a burst of reports
+ * costs one run.
+ */
+void pon_work_queue(struct pon_dev *pdev, struct pon_work *work)
+{
+	unsigned long flags;
+
+	/*
+	 * A driver that reports after pon_dev_unregister() canceled the work
+	 * would queue onto a workqueue that is being drained, which warns. The
+	 * instance is going: drop the item instead.
+	 */
+	if (READ_ONCE(pdev->going_away))
+		return;
+
+	spin_lock_irqsave(&pdev->work_lock, flags);
+	if (list_empty(&work->entry))
+		list_add_tail(&work->entry, &pdev->work_list);
+	spin_unlock_irqrestore(&pdev->work_lock, flags);
+
+	queue_work(pdev->wq, &pdev->work);
+}
+EXPORT_SYMBOL_GPL(pon_work_queue);
+
+/**
+ * pon_work_drain() - drop every queued item
+ * @pdev: PON device structure
+ *
+ * For the unregister path, after @going_away is set. Items still on the list
+ * are dropped rather than run, because the driver that owns them is leaving.
+ */
+void pon_work_drain(struct pon_dev *pdev)
+{
+	unsigned long flags;
+
+	spin_lock_irqsave(&pdev->work_lock, flags);
+	while (!list_empty(&pdev->work_list)) {
+		struct pon_work *item;
+
+		item = list_first_entry(&pdev->work_list, struct pon_work,
+					entry);
+		list_del_init(&item->entry);
+	}
+	spin_unlock_irqrestore(&pdev->work_lock, flags);
+}
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 18+ messages in thread

* [RFC net-next 05/12] net: pon: add the conduit contract
  2026-10-08 14:32 [RFC net-next 00/12] net: add the PON subsystem John Crispin
                   ` (3 preceding siblings ...)
  2026-10-08 14:32 ` [RFC net-next 04/12] net: pon: add the device state and the lent context John Crispin
@ 2026-10-08 14:32 ` John Crispin
  2026-10-08 14:32 ` [RFC net-next 06/12] net: pon: add the GEM network devices John Crispin
                   ` (8 subsequent siblings)
  13 siblings, 0 replies; 18+ messages in thread
From: John Crispin @ 2026-10-08 14:32 UTC (permalink / raw)
  To: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, linux-kernel, netdev, Andrew Lunn,
	Christian Marangi

A PON MAC on a system on chip has no DMA of its own, so its frames ride
the rings of an ethernet controller, the conduit. The driver of the
conduit registers it with pon_conduit_register() when its firmware node
names the MAC through a pon-handle. The core pairs the two in whichever
order they probe. Neither driver includes the headers of the other.

Per frame information crosses as explicit arguments: the GEM port, the
T-CONT channel and the queue on transmit, the GEM port and the OMCI
flags on receive. pon_conduit_rx() delivers a frame to the OMCI channel,
to the network device of its GEM port, or to the PON data interface. A
frame for a PON network device that is down is dropped and counted.

The conduit keeps the address of the data interface and the largest MTU
among the PON network devices, never less than an extended OMCI PDU with
its integrity field. A conduit that refuses the address or the MTU when
it pairs is warned about and stays paired.

A netdev notifier follows the conduit up and down. The PON network
devices keep their carrier only while the conduit is paired and up, so
the carrier falls before the conduit closes or goes.

pon_conduit_register(), pon_conduit_unregister() and pon_conduit_rx()
have no caller in the core. The driver of the conduit calls them.
pon_conduit_register() takes rtnl, so the driver calls it without rtnl.
pon_conduit_xmit() has no caller in the core either: the MAC driver
calls it from its transmit paths, because only the MAC driver can fill
struct pon_tx_info. The core carries that structure from one driver to
the other without reading it.

Assisted-by: LLM
Signed-off-by: John Crispin <john@phrozen.org>
---
 net/pon/pon_conduit.c | 614 ++++++++++++++++++++++++++++++++++++++++++
 1 file changed, 614 insertions(+)
 create mode 100644 net/pon/pon_conduit.c

diff --git a/net/pon/pon_conduit.c b/net/pon/pon_conduit.c
new file mode 100644
index 000000000000..ed0e0dff8cd3
--- /dev/null
+++ b/net/pon/pon_conduit.c
@@ -0,0 +1,614 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/* Copyright (C) 2026 John Crispin <john@phrozen.org> */
+
+#include <linux/etherdevice.h>
+#include <linux/netdevice.h>
+#include <linux/notifier.h>
+#include <linux/property.h>
+#include <linux/rtnetlink.h>
+#include <linux/slab.h>
+#include <net/pon.h>
+
+#include "pon.h"
+
+/* One registered conduit. It stays on the list whether or not a PON MAC has
+ * claimed it, so that a MAC which registers later, or registers again, finds
+ * it.
+ */
+struct pon_conduit {
+	struct list_head list;
+	struct net_device *netdev;
+	struct fwnode_handle *mac;
+	const struct pon_conduit_ops *ops;
+	struct pon_dev *pdev;
+	bool up;
+};
+
+/* guarded by pon_devs_lock */
+static LIST_HEAD(pon_conduits);
+
+/**
+ * pon_conduit_find() - look up the conduit of a network device
+ * @netdev:	the ethernet device a driver registered as a conduit
+ *
+ * Context: Called with pon_devs_lock held.
+ * Return: the conduit, or NULL when @netdev is not a registered conduit.
+ */
+static struct pon_conduit *pon_conduit_find(const struct net_device *netdev)
+{
+	struct pon_conduit *entry;
+
+	lockdep_assert_held(&pon_devs_lock);
+	list_for_each_entry(entry, &pon_conduits, list)
+		if (entry->netdev == netdev)
+			return entry;
+	return NULL;
+}
+
+/**
+ * pon_conduit_find_free() - look up an unclaimed conduit that names a MAC
+ * @pdev:	PON device structure
+ *
+ * A conduit names its MAC through the firmware node of @pdev's parent.
+ *
+ * Context: Called with pon_devs_lock held.
+ * Return: a conduit that names the MAC of @pdev and that no instance has
+ * claimed, or NULL when there is none.
+ */
+static struct pon_conduit *pon_conduit_find_free(const struct pon_dev *pdev)
+{
+	struct fwnode_handle *mac = dev_fwnode(pdev->parent);
+	struct pon_conduit *entry;
+
+	lockdep_assert_held(&pon_devs_lock);
+	list_for_each_entry(entry, &pon_conduits, list)
+		if (entry->mac == mac && !entry->pdev)
+			return entry;
+	return NULL;
+}
+
+/**
+ * pon_conduit_pair() - pair a conduit with an instance
+ * @entry:	the conduit
+ * @pdev:	PON device structure
+ *
+ * Takes a reference on the conduit's network device and publishes the
+ * pointer from each side to the other for the receive and transmit paths.
+ *
+ * Context: Called with pon_devs_lock held.
+ */
+static void pon_conduit_pair(struct pon_conduit *entry, struct pon_dev *pdev)
+{
+	lockdep_assert_held(&pon_devs_lock);
+
+	entry->pdev = pdev;
+	netdev_hold(entry->netdev, &pdev->conduit_tracker, GFP_KERNEL);
+	rcu_assign_pointer(pdev->conduit, entry);
+	rcu_assign_pointer(entry->netdev->pon_dev, pdev);
+}
+
+/**
+ * pon_conduit_unpair() - break the pairing of a conduit and its instance
+ * @entry:	the conduit, paired with an instance
+ *
+ * The receive path finds the instance through the conduit and the transmit
+ * path finds the conduit through the instance, both under RCU, so both
+ * pointers go before the grace period and the reference after it.
+ *
+ * Context: Called with pon_devs_lock held. Sleeps.
+ */
+static void pon_conduit_unpair(struct pon_conduit *entry)
+{
+	struct pon_dev *pdev = entry->pdev;
+
+	lockdep_assert_held(&pon_devs_lock);
+
+	rcu_assign_pointer(entry->netdev->pon_dev, NULL);
+	rcu_assign_pointer(pdev->conduit, NULL);
+	synchronize_net();
+	netdev_put(entry->netdev, &pdev->conduit_tracker);
+	entry->pdev = NULL;
+}
+
+/**
+ * pon_conduit_hold() - hold the conduit across a call that may sleep
+ * @pdev:	PON device structure
+ * @ops:	where to store what the conduit's driver does, or NULL
+ * @tracker:	the tracker of the reference
+ *
+ * Return: the conduit's ethernet device with a reference held, or NULL when
+ * no driver has claimed the instance.
+ */
+struct net_device *pon_conduit_hold(struct pon_dev *pdev,
+				    const struct pon_conduit_ops **ops,
+				    netdevice_tracker *tracker)
+{
+	struct net_device *conduit = NULL;
+	struct pon_conduit *entry;
+
+	rcu_read_lock();
+	entry = rcu_dereference(pdev->conduit);
+	if (entry) {
+		conduit = entry->netdev;
+		if (ops)
+			*ops = entry->ops;
+		netdev_hold(conduit, tracker, GFP_ATOMIC);
+	}
+	rcu_read_unlock();
+
+	return conduit;
+}
+
+/**
+ * pon_conduit_running() - whether the instance has a conduit that is up
+ * @pdev:	PON device structure
+ *
+ * A conduit counts as up from NETDEV_UP until NETDEV_GOING_DOWN, so the PON
+ * network devices lose their carrier before the conduit closes.
+ *
+ * Context: Any context.
+ * Return: true when a conduit is paired with @pdev and is up.
+ */
+bool pon_conduit_running(struct pon_dev *pdev)
+{
+	struct pon_conduit *entry;
+	bool up;
+
+	rcu_read_lock();
+	entry = rcu_dereference(pdev->conduit);
+	up = entry && READ_ONCE(entry->up);
+	rcu_read_unlock();
+
+	return up;
+}
+
+/**
+ * pon_conduit_carrier_update() - set the carrier after a conduit change
+ * @pdev:	PON device structure
+ *
+ * Context: Takes @pdev->lock. Called without it.
+ */
+static void pon_conduit_carrier_update(struct pon_dev *pdev)
+{
+	mutex_lock(&pdev->lock);
+	if (pon_dev_is_registered(pdev))
+		pon_dev_carrier_update(pdev);
+	mutex_unlock(&pdev->lock);
+}
+
+/**
+ * pon_conduit_addr_set() - give the conduit the PON data interface's address
+ * @pdev:	PON device structure
+ * @addr:	the address the PON network devices carry
+ *
+ * The frame engine behind a conduit recognizes the frames addressed to the
+ * ONU by the conduit's own address, so the two are kept the same. Called with
+ * rtnl held.
+ *
+ * Return: 0, or what the conduit's driver answered.
+ */
+int pon_conduit_addr_set(struct pon_dev *pdev, const u8 *addr)
+{
+	struct sockaddr_storage ss = {};
+	struct net_device *conduit;
+	netdevice_tracker tracker;
+	int err = 0;
+
+	ASSERT_RTNL();
+
+	conduit = pon_conduit_hold(pdev, NULL, &tracker);
+	if (!conduit)
+		return 0;
+
+	if (!ether_addr_equal(conduit->dev_addr, addr)) {
+		ss.ss_family = conduit->type;
+		memcpy(ss.__data, addr, ETH_ALEN);
+		err = dev_set_mac_address(conduit, &ss, NULL);
+	}
+
+	netdev_put(conduit, &tracker);
+	return err;
+}
+EXPORT_SYMBOL_GPL(pon_conduit_addr_set);
+
+/**
+ * pon_conduit_sync() - give a newly paired conduit the address and the MTU
+ * @pdev:	PON device structure
+ *
+ * After a pairing, outside every lock: the conduit's driver programs its
+ * address and its MTU under rtnl, which is never taken while an instance lock
+ * is held. The carrier follows the new conduit. A conduit that refuses the
+ * address or the MTU is warned about and stays paired.
+ *
+ * Does nothing once pon_dev_unregister() has begun. It tests that under rtnl,
+ * which pon_dev_unregister() takes before it lets the driver free the data
+ * network device, so the device read here is alive.
+ *
+ * Context: Takes rtnl and @pdev->lock. Called without rtnl, pon_devs_lock or
+ * @pdev->lock held.
+ */
+void pon_conduit_sync(struct pon_dev *pdev)
+{
+	int err;
+
+	rtnl_lock();
+	if (READ_ONCE(pdev->going_away)) {
+		rtnl_unlock();
+		return;
+	}
+	err = pon_conduit_addr_set(pdev, pdev->main_netdev->dev_addr);
+	if (err)
+		netdev_warn(pdev->main_netdev,
+			    "the conduit refused the address: %pe\n",
+			    ERR_PTR(err));
+	err = pon_conduit_mtu_set(pdev, NULL, 0);
+	if (err)
+		netdev_warn(pdev->main_netdev,
+			    "the conduit refused the MTU: %pe\n", ERR_PTR(err));
+	rtnl_unlock();
+
+	pon_conduit_carrier_update(pdev);
+}
+
+/**
+ * pon_conduit_register() - offer a network device's rings to a PON MAC
+ * @netdev:	the registered ethernet device, whose firmware node names the
+ *		MAC through a pon-handle reference
+ * @ops:	what the driver does for the MAC
+ *
+ * Neither side waits for the other. A MAC that registers later, or again,
+ * finds the conduit. Until then its frames are dropped and the conduit's own
+ * receive path keeps what arrives.
+ *
+ * Context: Takes rtnl. Called without it.
+ * Return: 0 once registered, -ENOENT when the firmware node names no PON MAC,
+ * -EBUSY when the network device is a conduit already, or a negative errno.
+ */
+int pon_conduit_register(struct net_device *netdev,
+			 const struct pon_conduit_ops *ops)
+{
+	struct pon_dev *pdev = NULL, *iter;
+	struct fwnode_handle *mac;
+	struct pon_conduit *entry;
+	unsigned long id;
+
+	mac = fwnode_find_reference(dev_fwnode(&netdev->dev), "pon-handle", 0);
+	if (IS_ERR(mac))
+		return PTR_ERR(mac);
+
+	entry = kzalloc_obj(*entry, GFP_KERNEL);
+	if (!entry) {
+		fwnode_handle_put(mac);
+		return -ENOMEM;
+	}
+	entry->netdev = netdev;
+	entry->mac = mac;
+	entry->ops = ops;
+
+	rtnl_lock();
+	mutex_lock(&pon_devs_lock);
+
+	if (pon_conduit_find(netdev)) {
+		mutex_unlock(&pon_devs_lock);
+		rtnl_unlock();
+		fwnode_handle_put(mac);
+		kfree(entry);
+		return -EBUSY;
+	}
+	entry->up = netif_running(netdev);
+	list_add_tail(&entry->list, &pon_conduits);
+
+	xa_for_each(&pon_devs, id, iter) {
+		if (dev_fwnode(iter->parent) != mac ||
+		    rcu_access_pointer(iter->conduit))
+			continue;
+		pon_conduit_pair(entry, iter);
+		pon_dev_get(iter);
+		pdev = iter;
+		break;
+	}
+
+	mutex_unlock(&pon_devs_lock);
+	rtnl_unlock();
+
+	if (pdev) {
+		pon_conduit_sync(pdev);
+		pon_dev_put(pdev);
+	}
+
+	return 0;
+}
+EXPORT_SYMBOL_GPL(pon_conduit_register);
+
+/**
+ * pon_conduit_unregister() - take a network device's rings back
+ * @netdev:	the device pon_conduit_register() was given
+ *
+ * Call before unregistering the device. A MAC the device served keeps
+ * running without a conduit and its PON network devices lose their carrier.
+ */
+void pon_conduit_unregister(struct net_device *netdev)
+{
+	struct pon_conduit *entry;
+	struct pon_dev *pdev;
+
+	mutex_lock(&pon_devs_lock);
+	entry = pon_conduit_find(netdev);
+	if (entry) {
+		pdev = entry->pdev;
+		if (pdev) {
+			pon_conduit_unpair(entry);
+			pon_conduit_carrier_update(pdev);
+		}
+		list_del(&entry->list);
+	}
+	mutex_unlock(&pon_devs_lock);
+
+	if (!entry)
+		return;
+
+	fwnode_handle_put(entry->mac);
+	kfree(entry);
+}
+EXPORT_SYMBOL_GPL(pon_conduit_unregister);
+
+/**
+ * pon_conduit_netdev_event() - follow a conduit that goes up or down
+ * @nb:		the notifier block
+ * @event:	what happened to the network device
+ * @ptr:	struct netdev_notifier_info of the network device
+ *
+ * A registered conduit counts as up from NETDEV_UP until NETDEV_GOING_DOWN.
+ * The PON network devices of the instance it serves get their carrier
+ * again, so they never carry traffic into a conduit that is down.
+ *
+ * Context: Called with rtnl held. Takes pon_devs_lock and the instance lock.
+ * Return: NOTIFY_DONE.
+ */
+static int pon_conduit_netdev_event(struct notifier_block *nb,
+				    unsigned long event, void *ptr)
+{
+	struct net_device *netdev = netdev_notifier_info_to_dev(ptr);
+	struct pon_conduit *entry;
+
+	if (event != NETDEV_UP && event != NETDEV_GOING_DOWN)
+		return NOTIFY_DONE;
+
+	mutex_lock(&pon_devs_lock);
+	entry = pon_conduit_find(netdev);
+	if (entry) {
+		WRITE_ONCE(entry->up, event == NETDEV_UP);
+		if (entry->pdev)
+			pon_conduit_carrier_update(entry->pdev);
+	}
+	mutex_unlock(&pon_devs_lock);
+
+	return NOTIFY_DONE;
+}
+
+static struct notifier_block pon_conduit_netdev_notifier = {
+	.notifier_call = pon_conduit_netdev_event,
+};
+
+/**
+ * pon_conduit_notifier_register() - watch the conduits go up and down
+ *
+ * Return: 0, or the errno of register_netdevice_notifier().
+ */
+int pon_conduit_notifier_register(void)
+{
+	return register_netdevice_notifier(&pon_conduit_netdev_notifier);
+}
+
+/**
+ * pon_conduit_notifier_unregister() - stop the watch of the conduits
+ */
+void pon_conduit_notifier_unregister(void)
+{
+	unregister_netdevice_notifier(&pon_conduit_netdev_notifier);
+}
+
+/**
+ * pon_conduit_attach() - claim a waiting conduit for a new instance
+ * @pdev:	PON device structure
+ *
+ * A new instance claims the conduit that names its MAC, if one is waiting.
+ *
+ * Context: Called with pon_devs_lock held.
+ * Return: true when a conduit was paired and there is an address to sync
+ * with pon_conduit_sync() once the locks are dropped, false otherwise.
+ */
+bool pon_conduit_attach(struct pon_dev *pdev)
+{
+	struct pon_conduit *entry;
+
+	entry = pon_conduit_find_free(pdev);
+	if (!entry)
+		return false;
+
+	pon_conduit_pair(entry, pdev);
+	return true;
+}
+
+/**
+ * pon_conduit_detach() - release the conduit an instance has claimed
+ * @pdev:	PON device structure
+ *
+ * The conduit stays registered and waits for the next instance of its MAC.
+ * Once this returns, no frame of the conduit reaches @pdev.
+ *
+ * Context: Takes pon_devs_lock. Sleeps.
+ */
+void pon_conduit_detach(struct pon_dev *pdev)
+{
+	struct pon_conduit *entry;
+
+	mutex_lock(&pon_devs_lock);
+	entry = rcu_dereference_protected(pdev->conduit,
+					  lockdep_is_held(&pon_devs_lock));
+	if (entry)
+		pon_conduit_unpair(entry);
+	mutex_unlock(&pon_devs_lock);
+}
+
+/**
+ * pon_conduit_mtu_largest() - the MTU the conduit needs
+ * @pdev:	PON device structure
+ * @dev:	the PON network device whose MTU changes, or NULL when none does
+ * @mtu:	the MTU @dev changes to
+ *
+ * Takes @mtu for @dev and the current MTU of every other PON network device
+ * of @pdev. The floor is the extended OMCI PDU with its MIC, 1980 bytes,
+ * ITU-T G.988 clause 11.2.5.
+ *
+ * Return: the largest of these MTUs.
+ */
+static unsigned int pon_conduit_mtu_largest(struct pon_dev *pdev,
+					    const struct net_device *dev,
+					    unsigned int mtu)
+{
+	struct net_device *gem;
+	unsigned long id;
+
+	mtu = max(mtu, PON_OMCI_MAX_LEN + PON_OMCI_MIC_LEN);
+
+	if (dev != pdev->main_netdev)
+		mtu = max(mtu, READ_ONCE(pdev->main_netdev->mtu));
+
+	xa_for_each(&pdev->gem_netdevs, id, gem)
+		if (dev != gem)
+			mtu = max(mtu, READ_ONCE(gem->mtu));
+
+	return mtu;
+}
+
+/**
+ * pon_conduit_mtu_set() - make room on the conduit for a PON network device
+ * @pdev:	PON device structure
+ * @dev:	the PON network device whose MTU changes, or NULL when none does
+ * @mtu:	the MTU @dev changes to
+ *
+ * The frames of every PON network device and the OMCI PDUs ride the conduit,
+ * so the conduit takes the largest MTU among them and never less than the
+ * longest OMCI PDU with its integrity field. Called with rtnl held.
+ *
+ * Return: 0, or what the conduit's driver answered.
+ */
+int pon_conduit_mtu_set(struct pon_dev *pdev, const struct net_device *dev,
+			unsigned int mtu)
+{
+	const struct pon_conduit_ops *ops;
+	struct net_device *conduit;
+	netdevice_tracker tracker;
+	int err = 0;
+
+	ASSERT_RTNL();
+
+	conduit = pon_conduit_hold(pdev, &ops, &tracker);
+	if (!conduit)
+		return 0;
+
+	mtu = pon_conduit_mtu_largest(pdev, dev, mtu);
+	if (mtu != READ_ONCE(conduit->mtu))
+		err = dev_set_mtu(conduit, mtu);
+
+	netdev_put(conduit, &tracker);
+
+	return err;
+}
+EXPORT_SYMBOL_GPL(pon_conduit_mtu_set);
+
+/**
+ * pon_conduit_xmit() - send one frame through the conduit
+ * @pdev:	PON device structure
+ * @skb:	the frame, consumed whatever this returns
+ * @info:	what the conduit needs to build the descriptor
+ *
+ * Return: 0 once the conduit has the frame, -ENETDOWN without a conduit,
+ * -EBUSY when its ring is full.
+ */
+int pon_conduit_xmit(struct pon_dev *pdev, struct sk_buff *skb,
+		     const struct pon_tx_info *info)
+{
+	struct pon_conduit *entry;
+	netdev_tx_t ret;
+
+	rcu_read_lock();
+	entry = rcu_dereference(pdev->conduit);
+	if (!entry) {
+		rcu_read_unlock();
+		dev_kfree_skb_any(skb);
+		return -ENETDOWN;
+	}
+	ret = entry->ops->xmit(entry->netdev, skb, info);
+	rcu_read_unlock();
+
+	if (ret == NETDEV_TX_BUSY) {
+		dev_kfree_skb_any(skb);
+		return -EBUSY;
+	}
+
+	return 0;
+}
+EXPORT_SYMBOL_GPL(pon_conduit_xmit);
+
+/**
+ * pon_conduit_rx() - hand one received frame to the PON MAC that owns it
+ * @conduit:	the ethernet device the frame arrived on
+ * @skb:	the frame, starting at its first byte: the MAC header of an
+ *		Ethernet frame, the transaction id of an OMCI PDU
+ * @info:	what the receive descriptor said about it
+ *
+ * Runs in the conduit's NAPI context. An OMCI PDU goes to the socket that
+ * owns the OMCI channel and is consumed here, a frame on a GEM port that
+ * carries a network device of its own goes there and everything else goes
+ * to the PON data interface. A frame for a network device that is down is
+ * dropped and counted in its rx_dropped.
+ *
+ * Return: 1 when the frame belongs to a PON network device (now skb->dev)
+ * and the conduit delivers it, 0 when it was consumed here, or -ENODEV when
+ * no PON MAC owns the conduit and the frame stays the conduit's.
+ */
+int pon_conduit_rx(struct net_device *conduit, struct sk_buff *skb,
+		   const struct pon_rx_info *info)
+{
+	struct net_device *dev;
+	struct pon_dev *pdev;
+	unsigned int len;
+	int ret;
+
+	rcu_read_lock();
+
+	pdev = rcu_dereference(conduit->pon_dev);
+	if (!pdev) {
+		ret = -ENODEV;
+		goto out;
+	}
+
+	if (info->oam) {
+		ret = pon_omci_conduit_rx(pdev, skb, info->mic_unchecked);
+		goto out;
+	}
+
+	dev = xa_load(&pdev->gem_netdevs, info->gem);
+	if (!dev)
+		dev = pdev->main_netdev;
+
+	if (!(READ_ONCE(dev->flags) & IFF_UP)) {
+		dev_core_stats_rx_dropped_inc(dev);
+		kfree_skb(skb);
+		ret = 0;
+		goto out;
+	}
+
+	len = skb->len;
+	skb->dev = dev;
+	skb->protocol = eth_type_trans(skb, dev);
+	dev_sw_netstats_rx_add(dev, len);
+	ret = 1;
+
+out:
+	rcu_read_unlock();
+	return ret;
+}
+EXPORT_SYMBOL_GPL(pon_conduit_rx);
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 18+ messages in thread

* [RFC net-next 06/12] net: pon: add the GEM network devices
  2026-10-08 14:32 [RFC net-next 00/12] net: add the PON subsystem John Crispin
                   ` (4 preceding siblings ...)
  2026-10-08 14:32 ` [RFC net-next 05/12] net: pon: add the conduit contract John Crispin
@ 2026-10-08 14:32 ` John Crispin
  2026-10-08 14:32 ` [RFC net-next 07/12] net: pon: add the OMCI channel John Crispin
                   ` (7 subsequent siblings)
  13 siblings, 0 replies; 18+ messages in thread
From: John Crispin @ 2026-10-08 14:32 UTC (permalink / raw)
  To: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni
  Cc: Donald Hunter, Andrew Lunn, Simon Horman, netdev, linux-kernel,
	Christian Marangi

A service that wants a network device of its own gets one per GEM port,
created with the rtnetlink link kind "gem" on top of the PON data
interface. The one attribute of the link kind is the GEM port id
(IFLA_GEM_ID in uapi/linux/if_link.h, listed in rt-link.yaml for ynl).
It must lie in the assignable range of ITU-T G.9807.1 Table C.6.6.

The device is registered outside the instance lock and attached to its
GEM port under it. The receive path finds it under RCU. It transmits
through the gem_xmit callback of the driver and takes the MTU range of
the data interface. A frame that the callback refuses is counted as
dropped, as the data interface counts it.

As a VLAN device does, the device names the data interface as its link
and transmits without a lock of its own (lltx). A module alias loads
the core for the link kind.

The gem_xmit callback belongs to the MAC driver, because only the MAC
driver knows the transmit channel and the queue of a GEM port. It fills
the descriptor information and hands the frame to the conduit.

Assisted-by: LLM
Signed-off-by: John Crispin <john@phrozen.org>
---
 Documentation/netlink/specs/rt-link.yaml |  14 +
 include/uapi/linux/if_link.h             |  10 +
 net/pon/pon_gem.c                        | 537 +++++++++++++++++++++++
 3 files changed, 561 insertions(+)
 create mode 100644 net/pon/pon_gem.c

diff --git a/Documentation/netlink/specs/rt-link.yaml b/Documentation/netlink/specs/rt-link.yaml
index 81b58633911d..ee2cb7d12c28 100644
--- a/Documentation/netlink/specs/rt-link.yaml
+++ b/Documentation/netlink/specs/rt-link.yaml
@@ -2368,6 +2368,17 @@ attribute-sets:
         name: mode
         type: u8
         enum: ovpn-mode
+  -
+    name: linkinfo-gem-attrs
+    name-prefix: ifla-gem-
+    attributes:
+      -
+        name: id
+        doc: GEM port ID of the PON device the GEM network device rides.
+        type: u32
+        checks:
+          min: 1021
+          max: 65534
 
 sub-messages:
   -
@@ -2427,6 +2438,9 @@ sub-messages:
       -
         value: ovpn
         attribute-set: linkinfo-ovpn-attrs
+      -
+        value: gem
+        attribute-set: linkinfo-gem-attrs
   -
     name: linkinfo-member-data-msg
     formats:
diff --git a/include/uapi/linux/if_link.h b/include/uapi/linux/if_link.h
index 245b36204525..6d4d149cf4ab 100644
--- a/include/uapi/linux/if_link.h
+++ b/include/uapi/linux/if_link.h
@@ -2075,4 +2075,14 @@ enum {
 
 #define IFLA_OVPN_MAX	(__IFLA_OVPN_MAX - 1)
 
+/* GEM section */
+
+enum {
+	IFLA_GEM_UNSPEC,
+	IFLA_GEM_ID,
+	__IFLA_GEM_MAX,
+};
+
+#define IFLA_GEM_MAX	(__IFLA_GEM_MAX - 1)
+
 #endif /* _UAPI_LINUX_IF_LINK_H */
diff --git a/net/pon/pon_gem.c b/net/pon/pon_gem.c
new file mode 100644
index 000000000000..ddb5b4f0a373
--- /dev/null
+++ b/net/pon/pon_gem.c
@@ -0,0 +1,537 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/* Copyright (C) 2026 John Crispin <john@phrozen.org> */
+
+#include <linux/etherdevice.h>
+#include <linux/if_link.h>
+#include <linux/module.h>
+#include <linux/netdevice.h>
+#include <net/rtnetlink.h>
+#include <net/pon.h>
+
+#include "pon.h"
+
+struct pon_gem_priv {
+	struct pon_dev *pdev;
+	/* Fixed for the life of the interface: a GEM port that carries one
+	 * cannot be removed, so the transmit path reads it without the lock.
+	 */
+	u16 gem_id;
+};
+
+/* A policy carries its bounds in an s16, so anything above 32767 needs the
+ * full range form. NLA_POLICY_MAX() would store a negative maximum and warn
+ * on every message.
+ */
+static const struct netlink_range_validation pon_gem_id_range = {
+	.min = PON_GEM_PORT_ID_ASSIGNABLE_MIN,
+	.max = PON_GEM_PORT_ID_ASSIGNABLE_MAX,
+};
+
+static const struct nla_policy pon_gem_link_policy[IFLA_GEM_MAX + 1] = {
+	[IFLA_GEM_ID] = NLA_POLICY_FULL_RANGE(NLA_U32, &pon_gem_id_range),
+};
+
+/**
+ * pon_gem_xmit() - send one frame on the GEM port of a GEM interface
+ * @skb:	the frame
+ * @dev:	the GEM network device
+ *
+ * Implements ndo_start_xmit. The driver's gem_xmit callback stamps the GEM
+ * port into the descriptor. Without that callback, or once the driver has
+ * unregistered, the frame is dropped and counted. A frame the callback
+ * refuses is counted as dropped too: the link is down, the GEM port has no
+ * channel or the ring is full, which the data interface counts the same way.
+ *
+ * Return: NETDEV_TX_OK, always.
+ */
+static netdev_tx_t pon_gem_xmit(struct sk_buff *skb, struct net_device *dev)
+{
+	struct pon_gem_priv *priv = netdev_priv(dev);
+	struct pon_dev *pdev = priv->pdev;
+	const struct pon_dev_ops *ops;
+	unsigned int len = skb->len;
+	int err;
+
+	/* A driver that cannot stamp a GEM port into its descriptor gives the
+	 * interface a name, a place in the link hierarchy and counters and
+	 * nothing else. A driver that has unregistered has no ops at all.
+	 */
+	rcu_read_lock();
+	ops = READ_ONCE(pdev->ops);
+	if (!ops || !ops->gem_xmit) {
+		rcu_read_unlock();
+		DEV_STATS_INC(dev, tx_dropped);
+		kfree_skb(skb);
+		return NETDEV_TX_OK;
+	}
+	err = ops->gem_xmit(pdev, priv->gem_id, skb);
+	rcu_read_unlock();
+
+	if (err) {
+		DEV_STATS_INC(dev, tx_dropped);
+		return NETDEV_TX_OK;
+	}
+
+	dev_sw_netstats_tx_add(dev, 1, len);
+	return NETDEV_TX_OK;
+}
+
+/**
+ * pon_gem_get_iflink() - the link of a GEM interface
+ * @dev:	the GEM network device
+ *
+ * Implements ndo_get_iflink. A GEM interface is created on the PON data
+ * interface, so a dump names that one as its link.
+ *
+ * Return: the ifindex of the data interface, or 0 before the interface is
+ * attached.
+ */
+static int pon_gem_get_iflink(const struct net_device *dev)
+{
+	const struct pon_gem_priv *priv = netdev_priv(dev);
+
+	if (!priv->pdev)
+		return 0;
+
+	return READ_ONCE(priv->pdev->main_netdev->ifindex);
+}
+
+/**
+ * pon_gem_change_mtu() - change the MTU of a GEM interface
+ * @dev:	the GEM network device
+ * @mtu:	the new MTU
+ *
+ * Implements ndo_change_mtu. The conduit carries the frames of every PON
+ * network device, so it must take the new MTU first.
+ *
+ * Context: Called with rtnl held.
+ * Return: 0, or what pon_conduit_mtu_set() answered.
+ */
+static int pon_gem_change_mtu(struct net_device *dev, int mtu)
+{
+	struct pon_gem_priv *priv = netdev_priv(dev);
+	int err;
+
+	if (priv->pdev) {
+		err = pon_conduit_mtu_set(priv->pdev, dev, mtu);
+		if (err)
+			return err;
+	}
+
+	WRITE_ONCE(dev->mtu, mtu);
+
+	return 0;
+}
+
+static const struct net_device_ops pon_gem_netdev_ops = {
+	.ndo_start_xmit		= pon_gem_xmit,
+	.ndo_change_mtu		= pon_gem_change_mtu,
+	.ndo_set_mac_address	= eth_mac_addr,
+	.ndo_validate_addr	= eth_validate_addr,
+	.ndo_get_iflink		= pon_gem_get_iflink,
+};
+
+/**
+ * pon_gem_netdev_id() - the GEM port a GEM network device names
+ * @dev:	a network device
+ * @gem_id:	filled in with the GEM port id on success
+ *
+ * The GEM port id is fixed for the life of the interface, so a caller that
+ * only wants to know which GEM port a device names needs no lock.
+ *
+ * Return: 0, or -ENODEV when @dev is not a GEM network device.
+ */
+int pon_gem_netdev_id(const struct net_device *dev, u16 *gem_id)
+{
+	const struct pon_gem_priv *priv;
+
+	if (dev->netdev_ops != &pon_gem_netdev_ops)
+		return -ENODEV;
+
+	priv = netdev_priv(dev);
+	*gem_id = priv->gem_id;
+
+	return 0;
+}
+
+/**
+ * pon_gem_dev_destructor() - release what a GEM interface holds
+ * @dev:	the GEM network device
+ *
+ * The priv_destructor of a GEM network device. Drops the reference on the
+ * instance that pon_gem_newlink() took, if the interface was attached.
+ */
+static void pon_gem_dev_destructor(struct net_device *dev)
+{
+	struct pon_gem_priv *priv = netdev_priv(dev);
+
+	if (priv->pdev)
+		pon_dev_put(priv->pdev);
+}
+
+/**
+ * pon_gem_dev_setup() - initialize a new GEM network device
+ * @dev:	the GEM network device
+ *
+ * Implements the setup op of the "gem" rtnl_link_ops.
+ */
+static void pon_gem_dev_setup(struct net_device *dev)
+{
+	ether_setup(dev);
+	dev->netdev_ops = &pon_gem_netdev_ops;
+	dev->pcpu_stat_type = NETDEV_PCPU_STAT_TSTATS;
+	dev->needs_free_netdev = true;
+	dev->priv_destructor = pon_gem_dev_destructor;
+	dev->priv_flags |= IFF_NO_QUEUE;
+	dev->lltx = true;
+	dev->netns_immutable = true;
+	dev->max_mtu = ETH_MAX_MTU;
+}
+
+/**
+ * pon_gem_link_target() - the GEM object a new GEM interface is created for
+ * @pdev:	PON device structure
+ * @dev:	the new GEM network device
+ * @lower:	the lower device the request named
+ * @gem_id:	the GEM port id the request named
+ * @extack:	netlink extended ack for the reason of a refusal
+ *
+ * Context: Called with @pdev->lock held.
+ * Return: the GEM object, or ERR_PTR() of -ENODEV when @pdev is not
+ * registered, -EINVAL when @lower is not the data interface or lives in
+ * another network namespace than @dev, -ENOENT when no GEM object has
+ * @gem_id, or -EBUSY when the GEM port has a network device already.
+ */
+static struct pon_gem *pon_gem_link_target(struct pon_dev *pdev,
+					   const struct net_device *dev,
+					   const struct net_device *lower,
+					   u16 gem_id,
+					   struct netlink_ext_ack *extack)
+{
+	struct pon_gem *gem;
+
+	lockdep_assert_held(&pdev->lock);
+
+	if (!pon_dev_is_registered(pdev))
+		return ERR_PTR(-ENODEV);
+
+	/* The conduit carries the same pointer and is not a PON interface. */
+	if (lower != pdev->main_netdev) {
+		NL_SET_ERR_MSG(extack,
+			       "lower device is not the PON data interface");
+		return ERR_PTR(-EINVAL);
+	}
+	if (!net_eq(dev_net(dev), dev_net(lower))) {
+		NL_SET_ERR_MSG(extack,
+			       "a GEM device lives in its PON device's namespace");
+		return ERR_PTR(-EINVAL);
+	}
+
+	gem = pon_gem_find(pdev, gem_id);
+	if (!gem) {
+		NL_SET_ERR_MSG(extack, "no GEM object with this id");
+		return ERR_PTR(-ENOENT);
+	}
+	if (gem->gem_netdev) {
+		NL_SET_ERR_MSG(extack,
+			       "the GEM port already has a network device");
+		return ERR_PTR(-EBUSY);
+	}
+
+	return gem;
+}
+
+/**
+ * pon_gem_link_attach() - attach a registered GEM interface to its GEM port
+ * @pdev:	PON device structure
+ * @dev:	the GEM network device, registered
+ * @lower:	the lower device the request named
+ * @gem_id:	the GEM port id the request named
+ * @extack:	netlink extended ack for the reason of a refusal
+ *
+ * Checks the target again, because the instance lock was dropped while
+ * @dev was registered. Then it links @dev and the GEM object and lets the
+ * receive path find @dev.
+ *
+ * Context: Called with rtnl and @pdev->lock held.
+ * Return: 0, or a negative errno from pon_gem_link_target() or the xarray.
+ */
+static int pon_gem_link_attach(struct pon_dev *pdev, struct net_device *dev,
+			       const struct net_device *lower, u16 gem_id,
+			       struct netlink_ext_ack *extack)
+{
+	struct pon_gem_priv *priv = netdev_priv(dev);
+	struct pon_gem *gem;
+	int err;
+
+	lockdep_assert_held(&pdev->lock);
+
+	gem = pon_gem_link_target(pdev, dev, lower, gem_id, extack);
+	if (IS_ERR(gem))
+		return PTR_ERR(gem);
+
+	err = xa_err(xa_store(&pdev->gem_netdevs, gem_id, dev, GFP_KERNEL));
+	if (err)
+		return err;
+
+	priv->pdev = pdev;
+	gem->gem_netdev = dev;
+	rcu_assign_pointer(dev->pon_dev, pdev);
+	pon_dev_carrier_update(pdev);
+
+	return 0;
+}
+
+/**
+ * pon_gem_newlink() - create a GEM interface
+ * @dev:	the new GEM network device
+ * @params:	the attributes of the request
+ * @extack:	netlink extended ack for the reason of a refusal
+ *
+ * Implements the newlink op of the "gem" rtnl_link_ops. The request names
+ * the PON data interface as its link and a GEM port id, which the policy
+ * limits to the assignable XGEM Port-IDs of ITU-T G.9807.1 Table C.6.6.
+ * The interface takes its MTU range from the data interface. On success
+ * it holds a reference on the instance until pon_gem_dev_destructor().
+ *
+ * Context: Called with rtnl held.
+ * Return: 0, -EINVAL for a missing link, a missing GEM port id or an MTU
+ * out of range, -ENODEV when the link does not exist, -EOPNOTSUPP when the
+ * link is not a PON device, or a negative errno.
+ */
+static int pon_gem_newlink(struct net_device *dev,
+			   struct rtnl_newlink_params *params,
+			   struct netlink_ext_ack *extack)
+{
+	struct nlattr **data = params->data;
+	struct nlattr **tb = params->tb;
+	struct pon_gem_priv *priv;
+	struct net_device *lower;
+	struct pon_dev *pdev;
+	struct pon_gem *gem;
+	u16 gem_id;
+	int err;
+
+	if (!tb[IFLA_LINK]) {
+		NL_SET_ERR_MSG(extack, "a lower PON device is required");
+		return -EINVAL;
+	}
+	if (!data || !data[IFLA_GEM_ID]) {
+		NL_SET_ERR_MSG(extack, "the GEM port id is required");
+		return -EINVAL;
+	}
+	gem_id = nla_get_u32(data[IFLA_GEM_ID]);
+
+	lower = __dev_get_by_index(rtnl_newlink_link_net(params),
+				   nla_get_u32(tb[IFLA_LINK]));
+	if (!lower)
+		return -ENODEV;
+
+	rcu_read_lock();
+	pdev = rcu_dereference(lower->pon_dev);
+	if (pdev && !pon_dev_tryget(pdev))
+		pdev = NULL;
+	rcu_read_unlock();
+
+	if (!pdev) {
+		NL_SET_ERR_MSG(extack, "lower device is not a PON device");
+		return -EOPNOTSUPP;
+	}
+
+	mutex_lock(&pdev->lock);
+	gem = pon_gem_link_target(pdev, dev, lower, gem_id, extack);
+	mutex_unlock(&pdev->lock);
+	if (IS_ERR(gem)) {
+		err = PTR_ERR(gem);
+		goto err_put;
+	}
+
+	if (!tb[IFLA_MTU]) {
+		dev->mtu = lower->mtu;
+	} else if (dev->mtu < lower->min_mtu || dev->mtu > lower->max_mtu) {
+		NL_SET_ERR_MSG(extack,
+			       "the MTU is outside the range of the PON data interface");
+		err = -EINVAL;
+		goto err_put;
+	}
+	dev->min_mtu = lower->min_mtu;
+	dev->max_mtu = lower->max_mtu;
+
+	err = pon_conduit_mtu_set(pdev, dev, dev->mtu);
+	if (err) {
+		NL_SET_ERR_MSG(extack, "the conduit refused the MTU");
+		goto err_put;
+	}
+
+	if (!tb[IFLA_ADDRESS])
+		eth_hw_addr_inherit(dev, lower);
+
+	SET_NETDEV_DEV(dev, pdev->parent);
+	priv = netdev_priv(dev);
+	priv->gem_id = gem_id;
+
+	err = register_netdevice(dev);
+	if (err)
+		goto err_put;
+
+	mutex_lock(&pdev->lock);
+	err = pon_gem_link_attach(pdev, dev, lower, gem_id, extack);
+	mutex_unlock(&pdev->lock);
+	if (err) {
+		unregister_netdevice(dev);
+		goto err_put;
+	}
+
+	return 0;
+
+err_put:
+	pon_dev_put(pdev);
+	return err;
+}
+
+/**
+ * pon_gem_link_detach() - detach a GEM interface from its GEM port
+ * @pdev:	PON device structure
+ * @dev:	the GEM network device
+ *
+ * Unlinks @dev and the GEM object and stops the receive path from finding
+ * @dev. Does nothing for an interface that is detached already.
+ *
+ * Context: Called with rtnl and @pdev->lock held.
+ */
+static void pon_gem_link_detach(struct pon_dev *pdev, struct net_device *dev)
+{
+	struct pon_gem_priv *priv = netdev_priv(dev);
+	struct pon_gem *gem;
+
+	lockdep_assert_held(&pdev->lock);
+
+	if (xa_load(&pdev->gem_netdevs, priv->gem_id) != dev)
+		return;
+
+	rcu_assign_pointer(dev->pon_dev, NULL);
+	xa_erase(&pdev->gem_netdevs, priv->gem_id);
+	gem = pon_gem_find(pdev, priv->gem_id);
+	if (gem && gem->gem_netdev == dev)
+		gem->gem_netdev = NULL;
+}
+
+/**
+ * pon_gem_dellink() - delete a GEM interface
+ * @dev:	the GEM network device
+ * @head:	the list to queue @dev on for unregistration
+ *
+ * Implements the dellink op of the "gem" rtnl_link_ops. Detaches @dev from
+ * its GEM port (if it is still attached) and queues it for unregistration.
+ *
+ * Context: Called with rtnl held.
+ */
+static void pon_gem_dellink(struct net_device *dev, struct list_head *head)
+{
+	struct pon_gem_priv *priv = netdev_priv(dev);
+	struct pon_dev *pdev = priv->pdev;
+
+	if (pdev) {
+		mutex_lock(&pdev->lock);
+		pon_gem_link_detach(pdev, dev);
+		mutex_unlock(&pdev->lock);
+	}
+
+	unregister_netdevice_queue(dev, head);
+}
+
+/**
+ * pon_gem_get_size() - the size of the link info of a GEM interface
+ * @dev:	the GEM network device
+ *
+ * Implements the get_size op of the "gem" rtnl_link_ops.
+ *
+ * Return: the room pon_gem_fill_info() needs.
+ */
+static size_t pon_gem_get_size(const struct net_device *dev)
+{
+	return nla_total_size(sizeof(u32));
+}
+
+/**
+ * pon_gem_fill_info() - put the link info of a GEM interface
+ * @skb:	the netlink message
+ * @dev:	the GEM network device
+ *
+ * Implements the fill_info op of the "gem" rtnl_link_ops. Puts the GEM
+ * port id as IFLA_GEM_ID.
+ *
+ * Return: 0, or -EMSGSIZE when @skb has no room.
+ */
+static int pon_gem_fill_info(struct sk_buff *skb, const struct net_device *dev)
+{
+	const struct pon_gem_priv *priv = netdev_priv(dev);
+
+	if (nla_put_u32(skb, IFLA_GEM_ID, priv->gem_id))
+		return -EMSGSIZE;
+
+	return 0;
+}
+
+static struct rtnl_link_ops pon_gem_link_ops = {
+	.kind		= PON_GEM_KIND,
+	.priv_size	= sizeof(struct pon_gem_priv),
+	.setup		= pon_gem_dev_setup,
+	.maxtype	= IFLA_GEM_MAX,
+	.policy		= pon_gem_link_policy,
+	.newlink	= pon_gem_newlink,
+	.dellink	= pon_gem_dellink,
+	.get_size	= pon_gem_get_size,
+	.fill_info	= pon_gem_fill_info,
+};
+MODULE_ALIAS_RTNL_LINK(PON_GEM_KIND);
+
+/**
+ * pon_gem_netdevs_unregister() - drop every GEM network device of a device
+ * @pdev:	PON device structure
+ *
+ * Called on the unregister path without rtnl or the instance lock held. It
+ * takes rtnl and then the instance lock, detaches every GEM network device
+ * and unregisters them all before it drops rtnl. pon_gem_dellink() runs
+ * under rtnl too, so each device is detached and unregistered exactly once,
+ * by whichever of the two takes rtnl first.
+ */
+void pon_gem_netdevs_unregister(struct pon_dev *pdev)
+{
+	struct net_device *ndev;
+	struct pon_gem *gem;
+	LIST_HEAD(list_kill);
+
+	rtnl_lock();
+	mutex_lock(&pdev->lock);
+	list_for_each_entry(gem, &pdev->gems, list) {
+		ndev = gem->gem_netdev;
+		if (!ndev)
+			continue;
+
+		pon_gem_link_detach(pdev, ndev);
+		unregister_netdevice_queue(ndev, &list_kill);
+	}
+	mutex_unlock(&pdev->lock);
+
+	unregister_netdevice_many(&list_kill);
+	rtnl_unlock();
+}
+
+/**
+ * pon_gem_link_register() - register the "gem" link type
+ *
+ * Return: 0, or the negative errno of rtnl_link_register().
+ */
+int pon_gem_link_register(void)
+{
+	return rtnl_link_register(&pon_gem_link_ops);
+}
+
+/**
+ * pon_gem_link_unregister() - unregister the "gem" link type
+ */
+void pon_gem_link_unregister(void)
+{
+	rtnl_link_unregister(&pon_gem_link_ops);
+}
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 18+ messages in thread

* [RFC net-next 07/12] net: pon: add the OMCI channel
  2026-10-08 14:32 [RFC net-next 00/12] net: add the PON subsystem John Crispin
                   ` (5 preceding siblings ...)
  2026-10-08 14:32 ` [RFC net-next 06/12] net: pon: add the GEM network devices John Crispin
@ 2026-10-08 14:32 ` John Crispin
  2026-10-08 14:32 ` [RFC net-next 08/12] net: pon: add netlink support John Crispin
                   ` (6 subsequent siblings)
  13 siblings, 0 replies; 18+ messages in thread
From: John Crispin @ 2026-10-08 14:32 UTC (permalink / raw)
  To: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, linux-kernel, netdev, Andrew Lunn,
	Christian Marangi

OMCI is the management protocol of ITU-T G.988. It is the job of
userspace. The kernel carries it over the pon netlink family, as nl80211
carries management frames. One socket registers for the OMCI PDUs of a
device, receives them as notifications and sends with omci-tx. The
registration ends when the socket closes.

A PDU crosses the boundary bare, without its integrity field. The core
checks the framing of a PDU to send against G.988 clause 11.2.3: 44
bytes in the baseline format, the header and the stated contents length
in the extended format. So every driver gets the same PDUs. A PDU that
the MAC passed up unchecked is verified by the driver, which holds the
OMCI integrity key. Received PDUs wait for the context of the instance
in one queue bounded at 512, in the order they arrived, checked or not.
A PDU that finds no owner or no room is dropped and counted.

The omci_xmit callback runs with bottom halves disabled, because it ends
in the transmit path of the conduit.

A MAC can check the integrity of a downstream OMCI PDU inline. A PDU
that it passes up unchecked goes to the omci_verify callback of the
driver, which checks it with the OMCI integrity key. The key stays in
the driver and the check can sleep, so the core defers the call to the
lent context. The core has no key and never checks a MIC itself.
Without omci_verify it drops an unchecked PDU.

Assisted-by: LLM
Signed-off-by: John Crispin <john@phrozen.org>
---
 net/pon/pon_omci.c | 326 +++++++++++++++++++++++++++++++++++++++++++++
 1 file changed, 326 insertions(+)
 create mode 100644 net/pon/pon_omci.c

diff --git a/net/pon/pon_omci.c b/net/pon/pon_omci.c
new file mode 100644
index 000000000000..c8093a9c3749
--- /dev/null
+++ b/net/pon/pon_omci.c
@@ -0,0 +1,326 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/* Copyright (C) 2026 John Crispin <john@phrozen.org> */
+
+#include <linux/atomic.h>
+#include <linux/bottom_half.h>
+#include <linux/netlink.h>
+#include <linux/notifier.h>
+#include <linux/skbuff.h>
+#include <linux/xarray.h>
+#include <linux/unaligned.h>
+#include <net/net_namespace.h>
+#include <net/pon.h>
+
+#include "pon.h"
+
+#define PON_OMCI_DEVID_OFF		3
+#define PON_OMCI_DEVID_BASELINE		0x0a
+#define PON_OMCI_DEVID_EXTENDED		0x0b
+#define PON_OMCI_CONTENTS_LEN_OFF	8
+#define PON_OMCI_EXTENDED_HDR_LEN	10
+#define PON_OMCI_BASELINE_LEN		44
+#define PON_OMCI_CB(skb)		((struct pon_omci_cb *)(skb)->cb)
+
+/**
+ * struct pon_omci_cb - what the receive queue keeps with a PDU
+ * @unverified: the MAC did not check the integrity of the PDU, which still
+ *		ends with its MIC
+ */
+struct pon_omci_cb {
+	bool unverified;
+};
+
+/**
+ * pon_omci_len_valid() - check the length of an OMCI PDU
+ * @len:	the length of the PDU in bytes
+ * @trailer:	the bytes the PDU carries after the message, PON_OMCI_MIC_LEN
+ *		when it still ends with its MIC, 0 otherwise
+ *
+ * The limits are those of the extended message format, ITU-T G.988 clause
+ * 11.2.5 and Table 11.2-2.
+ *
+ * Return: true when @len lies between PON_OMCI_MIN_LEN and PON_OMCI_MAX_LEN,
+ * both plus @trailer, false otherwise.
+ */
+static bool pon_omci_len_valid(unsigned int len, unsigned int trailer)
+{
+	return len >= PON_OMCI_MIN_LEN + trailer &&
+	       len <= PON_OMCI_MAX_LEN + trailer;
+}
+
+/**
+ * pon_omci_framing_valid() - check the framing of an OMCI PDU to send
+ * @pdu:	the PDU, without its MIC
+ * @len:	the length of @pdu in bytes, at least PON_OMCI_MIN_LEN
+ *
+ * The device identifier of ITU-T G.988 clause 11.2.3 selects the format. A
+ * baseline PDU (Table 11.2-1) is 44 bytes without the MIC. An extended PDU
+ * (Table 11.2-2) is the 10 byte header and as many bytes of contents as its
+ * message contents length states.
+ *
+ * Return: true when @len is the length the header of @pdu calls for, false
+ * otherwise.
+ */
+static bool pon_omci_framing_valid(const u8 *pdu, unsigned int len)
+{
+	unsigned int contents;
+
+	switch (pdu[PON_OMCI_DEVID_OFF]) {
+	case PON_OMCI_DEVID_BASELINE:
+		return len == PON_OMCI_BASELINE_LEN;
+	case PON_OMCI_DEVID_EXTENDED:
+		contents = get_unaligned_be16(pdu + PON_OMCI_CONTENTS_LEN_OFF);
+		return len == PON_OMCI_EXTENDED_HDR_LEN + contents;
+	default:
+		return false;
+	}
+}
+
+/**
+ * pon_omci_deliver() - hand one OMCI PDU to the owner of the OMCI channel
+ * @pdev:	PON device structure
+ * @skb:	the verified PDU, which is consumed
+ *
+ * A PDU that cannot be sent to the owner, or that arrives while no socket
+ * owns the channel, is dropped.
+ *
+ * Context: The instance's context, with @pdev->lock held.
+ */
+static void pon_omci_deliver(struct pon_dev *pdev, struct sk_buff *skb)
+{
+	if (pon_nl_omci_ntf(pdev, skb)) {
+		kfree_skb(skb);
+		return;
+	}
+
+	consume_skb(skb);
+}
+
+/**
+ * pon_omci_rx_verify() - check a received OMCI PDU through the driver
+ * @pdev:	PON device structure
+ * @skb:	the PDU the MAC passed up unchecked, ending with its MIC
+ *
+ * The driver's omci_verify callback checks and strips the MIC (ITU-T
+ * G.9807.1 clause C.15.7.2). A PDU that fails is freed. A PDU that cannot be
+ * made linear is freed too, because its integrity was never checked.
+ *
+ * Context: The instance's context, with @pdev->lock held.
+ * Return: true when the PDU passed and is now bare, false when it is gone.
+ */
+static bool pon_omci_rx_verify(struct pon_dev *pdev, struct sk_buff *skb)
+{
+	if (skb_linearize(skb) || pdev->ops->omci_verify(pdev, skb)) {
+		kfree_skb(skb);
+		return false;
+	}
+
+	return true;
+}
+
+/**
+ * pon_omci_rx_work() - verify and deliver the queued OMCI PDUs
+ * @pdev:	PON device structure
+ * @work:	the instance's @omci_rx_work
+ *
+ * Takes the PDUs in the order they arrived. One the MAC passed up unchecked
+ * goes through pon_omci_rx_verify() first. One the MAC verified goes to the
+ * owner as it is.
+ *
+ * Context: The instance's context, with @pdev->lock held.
+ */
+static void pon_omci_rx_work(struct pon_dev *pdev, struct pon_work *work)
+{
+	struct sk_buff *skb;
+
+	while ((skb = skb_dequeue(&pdev->omci_rxq))) {
+		if (PON_OMCI_CB(skb)->unverified &&
+		    !pon_omci_rx_verify(pdev, skb))
+			continue;
+
+		pon_omci_deliver(pdev, skb);
+	}
+}
+
+/**
+ * pon_omci_conduit_rx() - take one OMCI PDU from the conduit
+ * @pdev:	PON device structure
+ * @skb:	the PDU, starting at the transaction correlation id and ending
+ *		with its 4 byte MIC when @unverified
+ * @unverified:	the MAC did not check the integrity of the PDU
+ *
+ * Runs in the conduit's NAPI context. The PDU waits for the instance's
+ * context, which verifies it through the driver when @unverified and hands
+ * it to the owner of the OMCI channel. Checked and unchecked PDUs share one
+ * queue, so the owner gets them in the order they arrived. A PDU is at most
+ * PON_OMCI_MAX_LEN bytes and PON_OMCI_MIC_LEN more when @unverified, because
+ * the driver strips the MIC only once it has checked it (ITU-T G.988 clause
+ * 11.2.5). A PDU that arrives while no socket owns the channel is dropped
+ * here.
+ *
+ * Return: 0, the skb is consumed.
+ */
+int pon_omci_conduit_rx(struct pon_dev *pdev, struct sk_buff *skb,
+			bool unverified)
+{
+	if (!pon_omci_len_valid(skb->len, unverified ? PON_OMCI_MIC_LEN : 0) ||
+	    !READ_ONCE(pdev->omci_portid) ||
+	    (unverified && !pdev->ops->omci_verify) ||
+	    skb_queue_len_lockless(&pdev->omci_rxq) >= PON_OMCI_QUEUE_MAX) {
+		kfree_skb(skb);
+		return 0;
+	}
+
+	PON_OMCI_CB(skb)->unverified = unverified;
+	skb_queue_tail(&pdev->omci_rxq, skb);
+	pon_work_queue(pdev, &pdev->omci_rx_work);
+
+	return 0;
+}
+
+/**
+ * pon_omci_xmit() - send one OMCI PDU to the OLT
+ * @pdev:	PON device structure
+ * @pdu:	the PDU, without its MIC
+ * @len:	the length of @pdu in bytes
+ * @extack:	netlink extended ack for the reason the core refuses the PDU,
+ *		or NULL
+ *
+ * Copies the PDU into an skb and hands it to the driver's omci_xmit
+ * callback with bottom halves disabled. The length must fit the extended
+ * message format, ITU-T G.988 clause 11.2.5 and Table 11.2-2, less the 4
+ * byte MIC, which @pdu does not carry. It must also be the length that the
+ * header of the PDU calls for, as pon_omci_framing_valid() checks, so the
+ * driver receives a bare PDU of a known format. Only that refusal sets
+ * @extack. An error of the driver leaves it alone.
+ *
+ * Context: Called with @pdev->lock held, from the omci-tx netlink handler.
+ * Return: 0, -EINVAL for a length out of range or a length that does not
+ * match the format, -ENOMEM, or the driver's errno.
+ */
+int pon_omci_xmit(struct pon_dev *pdev, const void *pdu, unsigned int len,
+		  struct netlink_ext_ack *extack)
+{
+	struct sk_buff *skb;
+	int err;
+
+	lockdep_assert_held(&pdev->lock);
+
+	if (!pon_omci_len_valid(len, 0) || !pon_omci_framing_valid(pdu, len)) {
+		NL_SET_ERR_MSG(extack,
+			       "OMCI PDU length does not match its format");
+		return -EINVAL;
+	}
+
+	skb = alloc_skb(len, GFP_KERNEL);
+	if (!skb)
+		return -ENOMEM;
+	skb_put_data(skb, pdu, len);
+
+	local_bh_disable();
+	err = pdev->ops->omci_xmit(pdev, skb);
+	local_bh_enable();
+
+	return err;
+}
+
+/**
+ * pon_omci_register() - claim the OMCI channel of a PON device for a socket
+ * @pdev:	PON device structure
+ * @portid:	the netlink port id of the socket
+ *
+ * The first socket to claim the channel owns it until it closes, when
+ * pon_omci_netlink_notify() releases it. A second claim by the owner
+ * succeeds.
+ *
+ * Context: Called with @pdev->lock held.
+ * Return: 0 when @portid owns the channel, -EBUSY when another socket does.
+ */
+int pon_omci_register(struct pon_dev *pdev, u32 portid)
+{
+	u32 owner;
+
+	lockdep_assert_held(&pdev->lock);
+
+	owner = cmpxchg(&pdev->omci_portid, 0, portid);
+
+	return owner && owner != portid ? -EBUSY : 0;
+}
+
+/**
+ * pon_omci_netlink_notify() - release the OMCI channels of a closed socket
+ * @nb:		the notifier block
+ * @state:	what happened to the socket
+ * @data:	struct netlink_notify of the socket
+ *
+ * Runs when a netlink socket is released. Every device in the netns of the
+ * socket that it owned the OMCI channel of loses its owner.
+ *
+ * Return: NOTIFY_DONE.
+ */
+static int pon_omci_netlink_notify(struct notifier_block *nb,
+				   unsigned long state, void *data)
+{
+	struct netlink_notify *notify = data;
+	struct pon_dev *pdev;
+	unsigned long id;
+
+	if (state != NETLINK_URELEASE || notify->protocol != NETLINK_GENERIC)
+		return NOTIFY_DONE;
+
+	rcu_read_lock();
+	xa_for_each(&pon_devs, id, pdev) {
+		if (!net_eq(dev_net(pdev->main_netdev), notify->net))
+			continue;
+		cmpxchg(&pdev->omci_portid, notify->portid, 0);
+	}
+	rcu_read_unlock();
+
+	return NOTIFY_DONE;
+}
+
+static struct notifier_block pon_omci_netlink_notifier = {
+	.notifier_call = pon_omci_netlink_notify,
+};
+
+/**
+ * pon_omci_notifier_register() - watch for closed netlink sockets
+ *
+ * Return: 0, or the errno of netlink_register_notifier().
+ */
+int pon_omci_notifier_register(void)
+{
+	return netlink_register_notifier(&pon_omci_netlink_notifier);
+}
+
+/**
+ * pon_omci_notifier_unregister() - stop the watch for closed netlink sockets
+ */
+void pon_omci_notifier_unregister(void)
+{
+	netlink_unregister_notifier(&pon_omci_netlink_notifier);
+}
+
+/**
+ * pon_omci_init() - prepare the OMCI receive path of a new PON device
+ * @pdev:	PON device structure
+ *
+ * Context: From pon_dev_create(), before the device is published.
+ */
+void pon_omci_init(struct pon_dev *pdev)
+{
+	skb_queue_head_init(&pdev->omci_rxq);
+	pon_work_init(&pdev->omci_rx_work, pon_omci_rx_work);
+}
+
+/**
+ * pon_omci_destroy() - free the OMCI PDUs still queued
+ * @pdev:	PON device structure
+ *
+ * Context: From pon_dev_unregister(), after the conduit is detached and the
+ * instance's work is canceled.
+ */
+void pon_omci_destroy(struct pon_dev *pdev)
+{
+	skb_queue_purge(&pdev->omci_rxq);
+}
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 18+ messages in thread

* [RFC net-next 08/12] net: pon: add netlink support
  2026-10-08 14:32 [RFC net-next 00/12] net: add the PON subsystem John Crispin
                   ` (6 preceding siblings ...)
  2026-10-08 14:32 ` [RFC net-next 07/12] net: pon: add the OMCI channel John Crispin
@ 2026-10-08 14:32 ` John Crispin
  2026-10-08 14:32 ` [RFC net-next 09/12] net: pon: add the device registration John Crispin
                   ` (5 subsequent siblings)
  13 siblings, 0 replies; 18+ messages in thread
From: John Crispin @ 2026-10-08 14:32 UTC (permalink / raw)
  To: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, linux-kernel, netdev, Andrew Lunn,
	Christian Marangi

The handlers of the pon family, whose spec came first: the identity and
the link, the T-CONTs, the GEM ports and the upstream classifier rules
and the OMCI channel. pon-nl-gen.c and pon-nl-gen.h are generated from
pon.yaml with ynl-gen.

Every handler that names a device takes the instance lock in its
pre_doit. A write handler sends no message of its own. The netlink ACK
carries its result. Object notifications share the reply format of the
matching get. An object dump takes an optional device id and is marked
interrupted when a list changed under it, also when the part that
follows the change is empty.

A T-CONT and its alloc-id are one to one, as ITU-T G.988 clause 9.2.2
defines. A rebind moves the GEM ports of the T-CONT and restores them
when the driver refuses one. A GEM port that exists with other
attributes is refused. An upstream or bidirectional GEM port needs a
T-CONT. A downstream one takes none. No key material crosses the
family: a GEM port selects its key ring as clause 9.2.3 defines it. The
broadcast ring is refused.

The set_identity callback gets the registration id padded with 0x00 to
the 36 octets of ITU-T G.9807.1 Table C.11.25, so every driver pads the
same way. An empty registration id is refused. The core keeps the serial
number only. The serial number and the mode do not change while the
link is enabled, because the OLT addresses the ONU by its serial number.
The zero serial number, which addresses every ONU, is refused.

A mode the device does not list in its capabilities is refused with an
extack message and one line at error level through pon_dev_log_level()
that names the mode, for example "mode gpon not supported by the
device". The log work prints it, so no print runs under the instance
lock. A refused request is no failure of the system, so it raises no
WARN.

Assisted-by: LLM
Signed-off-by: John Crispin <john@phrozen.org>
---
 net/pon/pon-nl-gen.c |  286 ++++++++
 net/pon/pon-nl-gen.h |   45 ++
 net/pon/pon_nl.c     | 1671 ++++++++++++++++++++++++++++++++++++++++++
 3 files changed, 2002 insertions(+)
 create mode 100644 net/pon/pon-nl-gen.c
 create mode 100644 net/pon/pon-nl-gen.h
 create mode 100644 net/pon/pon_nl.c

diff --git a/net/pon/pon-nl-gen.c b/net/pon/pon-nl-gen.c
new file mode 100644
index 000000000000..e9898952c9fa
--- /dev/null
+++ b/net/pon/pon-nl-gen.c
@@ -0,0 +1,286 @@
+// SPDX-License-Identifier: ((GPL-2.0 WITH Linux-syscall-note) OR BSD-3-Clause)
+/* Do not edit directly, auto-generated from: */
+/*	Documentation/netlink/specs/pon.yaml */
+/* YNL-GEN kernel source */
+/* To regenerate run: tools/net/ynl/ynl-regen.sh */
+
+#include <net/netlink.h>
+#include <net/genetlink.h>
+
+#include "pon-nl-gen.h"
+
+#include <uapi/linux/pon.h>
+
+/* Integer value ranges */
+static const struct netlink_range_validation pon_a_tcont_index_range = {
+	.max	= 65534ULL,
+};
+
+static const struct netlink_range_validation pon_a_gem_id_range = {
+	.min	= 1021ULL,
+	.max	= 65534ULL,
+};
+
+static const struct netlink_range_validation pon_a_gem_tcont_index_range = {
+	.max	= 65534ULL,
+};
+
+static const struct netlink_range_validation pon_a_gem_map_gem_id_range = {
+	.min	= 1021ULL,
+	.max	= 65534ULL,
+};
+
+/* PON_CMD_DEV_GET - do */
+static const struct nla_policy pon_dev_get_nl_policy[PON_A_DEV_ID + 1] = {
+	[PON_A_DEV_ID] = NLA_POLICY_MIN(NLA_U32, 1),
+};
+
+/* PON_CMD_DEV_SET - do */
+static const struct nla_policy pon_dev_set_nl_policy[PON_A_DEV_ENABLE + 1] = {
+	[PON_A_DEV_ID] = NLA_POLICY_MIN(NLA_U32, 1),
+	[PON_A_DEV_MODE] = NLA_POLICY_MAX(NLA_U32, 2),
+	[PON_A_DEV_SERIAL] = NLA_POLICY_EXACT_LEN(8),
+	[PON_A_DEV_REGISTRATION_ID] = NLA_POLICY_MAX_LEN(36),
+	[PON_A_DEV_ENABLE] = NLA_POLICY_MAX(NLA_U8, 1),
+};
+
+/* PON_CMD_TCONT_GET - do */
+static const struct nla_policy pon_tcont_get_do_nl_policy[PON_A_TCONT_INDEX + 1] = {
+	[PON_A_TCONT_DEV_ID] = NLA_POLICY_MIN(NLA_U32, 1),
+	[PON_A_TCONT_INDEX] = NLA_POLICY_FULL_RANGE(NLA_U32, &pon_a_tcont_index_range),
+};
+
+/* PON_CMD_TCONT_GET - dump */
+static const struct nla_policy pon_tcont_get_dump_nl_policy[PON_A_TCONT_DEV_ID + 1] = {
+	[PON_A_TCONT_DEV_ID] = NLA_POLICY_MIN(NLA_U32, 1),
+};
+
+/* PON_CMD_TCONT_SET - do */
+static const struct nla_policy pon_tcont_set_nl_policy[PON_A_TCONT_ALLOC_ID + 1] = {
+	[PON_A_TCONT_DEV_ID] = NLA_POLICY_MIN(NLA_U32, 1),
+	[PON_A_TCONT_INDEX] = NLA_POLICY_FULL_RANGE(NLA_U32, &pon_a_tcont_index_range),
+	[PON_A_TCONT_ALLOC_ID] = NLA_POLICY_MAX(NLA_U32, 16383),
+};
+
+/* PON_CMD_TCONT_DEL - do */
+static const struct nla_policy pon_tcont_del_nl_policy[PON_A_TCONT_INDEX + 1] = {
+	[PON_A_TCONT_DEV_ID] = NLA_POLICY_MIN(NLA_U32, 1),
+	[PON_A_TCONT_INDEX] = NLA_POLICY_FULL_RANGE(NLA_U32, &pon_a_tcont_index_range),
+};
+
+/* PON_CMD_GEM_GET - do */
+static const struct nla_policy pon_gem_get_do_nl_policy[PON_A_GEM_ID + 1] = {
+	[PON_A_GEM_DEV_ID] = NLA_POLICY_MIN(NLA_U32, 1),
+	[PON_A_GEM_ID] = NLA_POLICY_FULL_RANGE(NLA_U32, &pon_a_gem_id_range),
+};
+
+/* PON_CMD_GEM_GET - dump */
+static const struct nla_policy pon_gem_get_dump_nl_policy[PON_A_GEM_DEV_ID + 1] = {
+	[PON_A_GEM_DEV_ID] = NLA_POLICY_MIN(NLA_U32, 1),
+};
+
+/* PON_CMD_GEM_NEW - do */
+static const struct nla_policy pon_gem_new_nl_policy[PON_A_GEM_KEY_RING + 1] = {
+	[PON_A_GEM_DEV_ID] = NLA_POLICY_MIN(NLA_U32, 1),
+	[PON_A_GEM_ID] = NLA_POLICY_FULL_RANGE(NLA_U32, &pon_a_gem_id_range),
+	[PON_A_GEM_DIR] = NLA_POLICY_RANGE(NLA_U32, 1, 3),
+	[PON_A_GEM_TCONT_INDEX] = NLA_POLICY_FULL_RANGE(NLA_U32, &pon_a_gem_tcont_index_range),
+	[PON_A_GEM_KEY_RING] = NLA_POLICY_MAX(NLA_U32, 3),
+};
+
+/* PON_CMD_GEM_DEL - do */
+static const struct nla_policy pon_gem_del_nl_policy[PON_A_GEM_ID + 1] = {
+	[PON_A_GEM_DEV_ID] = NLA_POLICY_MIN(NLA_U32, 1),
+	[PON_A_GEM_ID] = NLA_POLICY_FULL_RANGE(NLA_U32, &pon_a_gem_id_range),
+};
+
+/* PON_CMD_GEM_MAP_GET - dump */
+static const struct nla_policy pon_gem_map_get_nl_policy[PON_A_GEM_MAP_DEV_ID + 1] = {
+	[PON_A_GEM_MAP_DEV_ID] = NLA_POLICY_MIN(NLA_U32, 1),
+};
+
+/* PON_CMD_GEM_MAP_NEW - do */
+static const struct nla_policy pon_gem_map_new_nl_policy[PON_A_GEM_MAP_DSCP + 1] = {
+	[PON_A_GEM_MAP_DEV_ID] = NLA_POLICY_MIN(NLA_U32, 1),
+	[PON_A_GEM_MAP_GEM_ID] = NLA_POLICY_FULL_RANGE(NLA_U32, &pon_a_gem_map_gem_id_range),
+	[PON_A_GEM_MAP_TAG] = NLA_POLICY_MAX(NLA_U32, 1),
+	[PON_A_GEM_MAP_VID] = NLA_POLICY_MAX(NLA_U32, 4094),
+	[PON_A_GEM_MAP_PBIT] = NLA_POLICY_MAX(NLA_U32, 7),
+	[PON_A_GEM_MAP_DSCP] = NLA_POLICY_MAX(NLA_U32, 63),
+};
+
+/* PON_CMD_GEM_MAP_DEL - do */
+static const struct nla_policy pon_gem_map_del_nl_policy[PON_A_GEM_MAP_DSCP + 1] = {
+	[PON_A_GEM_MAP_DEV_ID] = NLA_POLICY_MIN(NLA_U32, 1),
+	[PON_A_GEM_MAP_GEM_ID] = NLA_POLICY_FULL_RANGE(NLA_U32, &pon_a_gem_map_gem_id_range),
+	[PON_A_GEM_MAP_TAG] = NLA_POLICY_MAX(NLA_U32, 1),
+	[PON_A_GEM_MAP_VID] = NLA_POLICY_MAX(NLA_U32, 4094),
+	[PON_A_GEM_MAP_PBIT] = NLA_POLICY_MAX(NLA_U32, 7),
+	[PON_A_GEM_MAP_DSCP] = NLA_POLICY_MAX(NLA_U32, 63),
+};
+
+/* PON_CMD_OMCI_REGISTER - do */
+static const struct nla_policy pon_omci_register_nl_policy[PON_A_OMCI_DEV_ID + 1] = {
+	[PON_A_OMCI_DEV_ID] = NLA_POLICY_MIN(NLA_U32, 1),
+};
+
+/* PON_CMD_OMCI_TX - do */
+static const struct nla_policy pon_omci_tx_nl_policy[PON_A_OMCI_PDU + 1] = {
+	[PON_A_OMCI_DEV_ID] = NLA_POLICY_MIN(NLA_U32, 1),
+	[PON_A_OMCI_PDU] = NLA_POLICY_MAX_LEN(1976),
+};
+
+/* Ops table for pon */
+static const struct genl_split_ops pon_nl_ops[] = {
+	{
+		.cmd		= PON_CMD_DEV_GET,
+		.pre_doit	= pon_device_get_locked,
+		.doit		= pon_nl_dev_get_doit,
+		.post_doit	= pon_device_unlock,
+		.policy		= pon_dev_get_nl_policy,
+		.maxattr	= PON_A_DEV_ID,
+		.flags		= GENL_CMD_CAP_DO,
+	},
+	{
+		.cmd	= PON_CMD_DEV_GET,
+		.dumpit	= pon_nl_dev_get_dumpit,
+		.flags	= GENL_CMD_CAP_DUMP,
+	},
+	{
+		.cmd		= PON_CMD_DEV_SET,
+		.pre_doit	= pon_device_get_locked,
+		.doit		= pon_nl_dev_set_doit,
+		.post_doit	= pon_device_unlock,
+		.policy		= pon_dev_set_nl_policy,
+		.maxattr	= PON_A_DEV_ENABLE,
+		.flags		= GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
+	},
+	{
+		.cmd		= PON_CMD_TCONT_GET,
+		.pre_doit	= pon_device_get_locked,
+		.doit		= pon_nl_tcont_get_doit,
+		.post_doit	= pon_device_unlock,
+		.policy		= pon_tcont_get_do_nl_policy,
+		.maxattr	= PON_A_TCONT_INDEX,
+		.flags		= GENL_CMD_CAP_DO,
+	},
+	{
+		.cmd		= PON_CMD_TCONT_GET,
+		.dumpit		= pon_nl_tcont_get_dumpit,
+		.policy		= pon_tcont_get_dump_nl_policy,
+		.maxattr	= PON_A_TCONT_DEV_ID,
+		.flags		= GENL_CMD_CAP_DUMP,
+	},
+	{
+		.cmd		= PON_CMD_TCONT_SET,
+		.pre_doit	= pon_device_get_locked,
+		.doit		= pon_nl_tcont_set_doit,
+		.post_doit	= pon_device_unlock,
+		.policy		= pon_tcont_set_nl_policy,
+		.maxattr	= PON_A_TCONT_ALLOC_ID,
+		.flags		= GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
+	},
+	{
+		.cmd		= PON_CMD_TCONT_DEL,
+		.pre_doit	= pon_device_get_locked,
+		.doit		= pon_nl_tcont_del_doit,
+		.post_doit	= pon_device_unlock,
+		.policy		= pon_tcont_del_nl_policy,
+		.maxattr	= PON_A_TCONT_INDEX,
+		.flags		= GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
+	},
+	{
+		.cmd		= PON_CMD_GEM_GET,
+		.pre_doit	= pon_device_get_locked,
+		.doit		= pon_nl_gem_get_doit,
+		.post_doit	= pon_device_unlock,
+		.policy		= pon_gem_get_do_nl_policy,
+		.maxattr	= PON_A_GEM_ID,
+		.flags		= GENL_CMD_CAP_DO,
+	},
+	{
+		.cmd		= PON_CMD_GEM_GET,
+		.dumpit		= pon_nl_gem_get_dumpit,
+		.policy		= pon_gem_get_dump_nl_policy,
+		.maxattr	= PON_A_GEM_DEV_ID,
+		.flags		= GENL_CMD_CAP_DUMP,
+	},
+	{
+		.cmd		= PON_CMD_GEM_NEW,
+		.pre_doit	= pon_device_get_locked,
+		.doit		= pon_nl_gem_new_doit,
+		.post_doit	= pon_device_unlock,
+		.policy		= pon_gem_new_nl_policy,
+		.maxattr	= PON_A_GEM_KEY_RING,
+		.flags		= GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
+	},
+	{
+		.cmd		= PON_CMD_GEM_DEL,
+		.pre_doit	= pon_device_get_locked,
+		.doit		= pon_nl_gem_del_doit,
+		.post_doit	= pon_device_unlock,
+		.policy		= pon_gem_del_nl_policy,
+		.maxattr	= PON_A_GEM_ID,
+		.flags		= GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
+	},
+	{
+		.cmd		= PON_CMD_GEM_MAP_GET,
+		.dumpit		= pon_nl_gem_map_get_dumpit,
+		.policy		= pon_gem_map_get_nl_policy,
+		.maxattr	= PON_A_GEM_MAP_DEV_ID,
+		.flags		= GENL_CMD_CAP_DUMP,
+	},
+	{
+		.cmd		= PON_CMD_GEM_MAP_NEW,
+		.pre_doit	= pon_device_get_locked,
+		.doit		= pon_nl_gem_map_new_doit,
+		.post_doit	= pon_device_unlock,
+		.policy		= pon_gem_map_new_nl_policy,
+		.maxattr	= PON_A_GEM_MAP_DSCP,
+		.flags		= GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
+	},
+	{
+		.cmd		= PON_CMD_GEM_MAP_DEL,
+		.pre_doit	= pon_device_get_locked,
+		.doit		= pon_nl_gem_map_del_doit,
+		.post_doit	= pon_device_unlock,
+		.policy		= pon_gem_map_del_nl_policy,
+		.maxattr	= PON_A_GEM_MAP_DSCP,
+		.flags		= GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
+	},
+	{
+		.cmd		= PON_CMD_OMCI_REGISTER,
+		.pre_doit	= pon_device_get_locked,
+		.doit		= pon_nl_omci_register_doit,
+		.post_doit	= pon_device_unlock,
+		.policy		= pon_omci_register_nl_policy,
+		.maxattr	= PON_A_OMCI_DEV_ID,
+		.flags		= GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
+	},
+	{
+		.cmd		= PON_CMD_OMCI_TX,
+		.pre_doit	= pon_device_get_locked,
+		.doit		= pon_nl_omci_tx_doit,
+		.post_doit	= pon_device_unlock,
+		.policy		= pon_omci_tx_nl_policy,
+		.maxattr	= PON_A_OMCI_PDU,
+		.flags		= GENL_ADMIN_PERM | GENL_CMD_CAP_DO,
+	},
+};
+
+static const struct genl_multicast_group pon_nl_mcgrps[] = {
+	[PON_NLGRP_MGMT] = { "mgmt", },
+	[PON_NLGRP_STATE] = { "state", },
+};
+
+struct genl_family pon_nl_family __ro_after_init = {
+	.name		= PON_FAMILY_NAME,
+	.version	= PON_FAMILY_VERSION,
+	.netnsok	= true,
+	.parallel_ops	= true,
+	.module		= THIS_MODULE,
+	.split_ops	= pon_nl_ops,
+	.n_split_ops	= ARRAY_SIZE(pon_nl_ops),
+	.mcgrps		= pon_nl_mcgrps,
+	.n_mcgrps	= ARRAY_SIZE(pon_nl_mcgrps),
+};
diff --git a/net/pon/pon-nl-gen.h b/net/pon/pon-nl-gen.h
new file mode 100644
index 000000000000..24d773ee2651
--- /dev/null
+++ b/net/pon/pon-nl-gen.h
@@ -0,0 +1,45 @@
+/* SPDX-License-Identifier: ((GPL-2.0 WITH Linux-syscall-note) OR BSD-3-Clause) */
+/* Do not edit directly, auto-generated from: */
+/*	Documentation/netlink/specs/pon.yaml */
+/* YNL-GEN kernel header */
+/* To regenerate run: tools/net/ynl/ynl-regen.sh */
+
+#ifndef _LINUX_PON_GEN_H
+#define _LINUX_PON_GEN_H
+
+#include <net/netlink.h>
+#include <net/genetlink.h>
+
+#include <uapi/linux/pon.h>
+
+int pon_device_get_locked(const struct genl_split_ops *ops,
+			  struct sk_buff *skb, struct genl_info *info);
+void
+pon_device_unlock(const struct genl_split_ops *ops, struct sk_buff *skb,
+		  struct genl_info *info);
+
+int pon_nl_dev_get_doit(struct sk_buff *skb, struct genl_info *info);
+int pon_nl_dev_get_dumpit(struct sk_buff *skb, struct netlink_callback *cb);
+int pon_nl_dev_set_doit(struct sk_buff *skb, struct genl_info *info);
+int pon_nl_tcont_get_doit(struct sk_buff *skb, struct genl_info *info);
+int pon_nl_tcont_get_dumpit(struct sk_buff *skb, struct netlink_callback *cb);
+int pon_nl_tcont_set_doit(struct sk_buff *skb, struct genl_info *info);
+int pon_nl_tcont_del_doit(struct sk_buff *skb, struct genl_info *info);
+int pon_nl_gem_get_doit(struct sk_buff *skb, struct genl_info *info);
+int pon_nl_gem_get_dumpit(struct sk_buff *skb, struct netlink_callback *cb);
+int pon_nl_gem_new_doit(struct sk_buff *skb, struct genl_info *info);
+int pon_nl_gem_del_doit(struct sk_buff *skb, struct genl_info *info);
+int pon_nl_gem_map_get_dumpit(struct sk_buff *skb, struct netlink_callback *cb);
+int pon_nl_gem_map_new_doit(struct sk_buff *skb, struct genl_info *info);
+int pon_nl_gem_map_del_doit(struct sk_buff *skb, struct genl_info *info);
+int pon_nl_omci_register_doit(struct sk_buff *skb, struct genl_info *info);
+int pon_nl_omci_tx_doit(struct sk_buff *skb, struct genl_info *info);
+
+enum {
+	PON_NLGRP_MGMT,
+	PON_NLGRP_STATE,
+};
+
+extern struct genl_family pon_nl_family;
+
+#endif /* _LINUX_PON_GEN_H */
diff --git a/net/pon/pon_nl.c b/net/pon/pon_nl.c
new file mode 100644
index 000000000000..4ca55c992063
--- /dev/null
+++ b/net/pon/pon_nl.c
@@ -0,0 +1,1671 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/* Copyright (C) 2026 John Crispin <john@phrozen.org> */
+
+#include <linux/atomic.h>
+#include <linux/limits.h>
+#include <linux/skbuff.h>
+#include <linux/slab.h>
+#include <linux/xarray.h>
+#include <net/genetlink.h>
+#include <net/pon.h>
+#include <net/sock.h>
+
+#include "pon-nl-gen.h"
+#include "pon.h"
+
+static const char *const pon_nl_mode_names[] = {
+	[PON_MODE_GPON]		= "gpon",
+	[PON_MODE_XG_PON]	= "xg-pon",
+	[PON_MODE_XGS_PON]	= "xgs-pon",
+};
+
+/* The object dumps walk every device the caller's namespace can see, or the
+ * one device the request names and within a device one list. A dump that
+ * fills its skb stops at the object it could not fit and resumes there:
+ * args[0] is the device, args[1] the objects of it already sent. The count
+ * is only good for the list it was taken from, so every object that joins or
+ * leaves a list moves the generation and a dump that resumes on another
+ * generation is marked as interrupted. The generation is read under the lock
+ * that the walk holds, so no change falls between the two.
+ */
+typedef int (*pon_nl_obj_fill_t)(struct pon_dev *pdev, struct list_head *pos,
+				 struct sk_buff *rsp,
+				 const struct genl_info *info);
+
+typedef int (*pon_nl_dev_fill_t)(struct pon_dev *pdev, struct sk_buff *rsp,
+				 const struct genl_info *info);
+
+static atomic_t pon_nl_obj_gen = ATOMIC_INIT(1);
+
+/* Netlink helpers */
+
+/**
+ * pon_nl_dev_reply() - answer a request with one message about the device
+ * @info: the request info, user_ptr[0] holds the device
+ * @fill: puts the message into the reply
+ *
+ * Context: Called with the device lock held by pon_device_get_locked().
+ * Return: 0, -ENOMEM, the error of @fill, or the error of genlmsg_reply().
+ */
+static int pon_nl_dev_reply(struct genl_info *info, pon_nl_dev_fill_t fill)
+{
+	struct pon_dev *pdev = info->user_ptr[0];
+	struct sk_buff *rsp;
+	int err;
+
+	rsp = genlmsg_new(GENLMSG_DEFAULT_SIZE, GFP_KERNEL);
+	if (!rsp)
+		return -ENOMEM;
+
+	err = fill(pdev, rsp, info);
+	if (err) {
+		nlmsg_free(rsp);
+		return err;
+	}
+
+	return genlmsg_reply(rsp, info);
+}
+
+/**
+ * pon_nl_obj_reply() - answer a request with one message about an object
+ * @info: the request info, user_ptr[0] holds the device
+ * @pos: the list entry of the object
+ * @fill: puts the message into the reply
+ *
+ * Context: Called with the device lock held by pon_device_get_locked().
+ * Return: 0, -ENOMEM, the error of @fill, or the error of genlmsg_reply().
+ */
+static int pon_nl_obj_reply(struct genl_info *info, struct list_head *pos,
+			    pon_nl_obj_fill_t fill)
+{
+	struct pon_dev *pdev = info->user_ptr[0];
+	struct sk_buff *rsp;
+	int err;
+
+	rsp = genlmsg_new(GENLMSG_DEFAULT_SIZE, GFP_KERNEL);
+	if (!rsp)
+		return -ENOMEM;
+
+	err = fill(pdev, pos, rsp, info);
+	if (err) {
+		nlmsg_free(rsp);
+		return err;
+	}
+
+	return genlmsg_reply(rsp, info);
+}
+
+/**
+ * pon_nl_notify_obj() - send an object notification to the mgmt group
+ * @pdev: the PON device
+ * @pos: the list entry of the object
+ * @cmd: the notification command
+ * @fill: puts the object into the notification
+ *
+ * Object notifications share the GET reply format, so a listener parses one
+ * message shape whether it asked for the object or was told about it. Does
+ * nothing when the group has no listener in the device's namespace.
+ *
+ * Context: Called with @pdev->lock held. May sleep.
+ */
+static void pon_nl_notify_obj(struct pon_dev *pdev, struct list_head *pos,
+			      u32 cmd, pon_nl_obj_fill_t fill)
+{
+	struct net *net = dev_net(pdev->main_netdev);
+	struct genl_info info;
+	struct sk_buff *ntf;
+
+	lockdep_assert_held(&pdev->lock);
+
+	if (!genl_has_listeners(&pon_nl_family, net, PON_NLGRP_MGMT))
+		return;
+
+	ntf = genlmsg_new(GENLMSG_DEFAULT_SIZE, GFP_KERNEL);
+	if (!ntf)
+		return;
+
+	genl_info_init_ntf(&info, &pon_nl_family, cmd);
+	genl_info_net_set(&info, net);
+	if (fill(pdev, pos, ntf, &info)) {
+		nlmsg_free(ntf);
+		return;
+	}
+
+	genlmsg_multicast_netns(&pon_nl_family, net, ntf, 0, PON_NLGRP_MGMT,
+				GFP_KERNEL);
+}
+
+/* Device lookup and locking */
+
+/**
+ * pon_device_get_and_lock() - look up a PON device and take its lock
+ * @net: the network namespace of the request
+ * @dev_id: the device id attribute
+ *
+ * Takes pon_devs_lock for the lookup and holds it until @pdev->lock is
+ * taken, so the device cannot go away in between. A device in another
+ * namespace is not found.
+ *
+ * Return: the device with its lock held, or ERR_PTR(-ENODEV).
+ */
+static struct pon_dev *
+pon_device_get_and_lock(struct net *net, struct nlattr *dev_id)
+{
+	struct pon_dev *pdev;
+
+	mutex_lock(&pon_devs_lock);
+	pdev = xa_load(&pon_devs, nla_get_u32(dev_id));
+	if (!pdev) {
+		mutex_unlock(&pon_devs_lock);
+		return ERR_PTR(-ENODEV);
+	}
+
+	mutex_lock(&pdev->lock);
+	mutex_unlock(&pon_devs_lock);
+
+	if (dev_net(pdev->main_netdev) != net) {
+		mutex_unlock(&pdev->lock);
+		return ERR_PTR(-ENODEV);
+	}
+
+	return pdev;
+}
+
+/**
+ * pon_device_get_locked() - genl pre_doit that looks up and locks the device
+ * @ops: the operation of the request
+ * @skb: the request
+ * @info: the request info, user_ptr[0] receives the device
+ *
+ * Every attribute set carries the device id as attribute 1, so one pre_doit
+ * serves all commands that name a device.
+ *
+ * Context: Runs as the genl pre_doit. On success the device lock stays held
+ * until pon_device_unlock().
+ * Return: 0, -EINVAL without a device id, or -ENODEV.
+ */
+int pon_device_get_locked(const struct genl_split_ops *ops,
+			  struct sk_buff *skb, struct genl_info *info)
+{
+	struct nlattr *id = info->attrs[PON_A_DEV_ID];
+
+	/* Every attribute set carries the device id as attribute 1. */
+	BUILD_BUG_ON((int)PON_A_DEV_ID != (int)PON_A_TCONT_DEV_ID ||
+		     (int)PON_A_DEV_ID != (int)PON_A_GEM_DEV_ID ||
+		     (int)PON_A_DEV_ID != (int)PON_A_GEM_MAP_DEV_ID ||
+		     (int)PON_A_DEV_ID != (int)PON_A_OMCI_DEV_ID);
+
+	if (!id) {
+		NL_SET_ERR_MSG(info->extack, "device id is missing");
+		return -EINVAL;
+	}
+
+	info->user_ptr[0] = pon_device_get_and_lock(genl_info_net(info), id);
+	if (IS_ERR(info->user_ptr[0])) {
+		NL_SET_ERR_MSG_ATTR(info->extack, id, "no such device");
+		return PTR_ERR(info->user_ptr[0]);
+	}
+
+	return 0;
+}
+
+/**
+ * pon_device_unlock() - genl post_doit that drops the device lock
+ * @ops: the operation of the request
+ * @skb: the request
+ * @info: the request info, user_ptr[0] holds the device
+ *
+ * Context: Runs as the genl post_doit, with the lock that
+ * pon_device_get_locked() took.
+ */
+void
+pon_device_unlock(const struct genl_split_ops *ops, struct sk_buff *skb,
+		  struct genl_info *info)
+{
+	struct pon_dev *pdev = info->user_ptr[0];
+
+	mutex_unlock(&pdev->lock);
+}
+
+/* Device */
+
+/**
+ * pon_nl_dev_fill() - put one device message into an skb
+ * @pdev: the PON device
+ * @rsp: the skb to fill
+ * @info: the request, or the notification info
+ *
+ * Context: Called with @pdev->lock held.
+ * Return: 0, or -EMSGSIZE when the message does not fit.
+ */
+static int
+pon_nl_dev_fill(struct pon_dev *pdev, struct sk_buff *rsp,
+		const struct genl_info *info)
+{
+	void *hdr;
+
+	hdr = genlmsg_iput(rsp, info);
+	if (!hdr)
+		return -EMSGSIZE;
+
+	if (nla_put_u32(rsp, PON_A_DEV_ID, pdev->id) ||
+	    nla_put_u32(rsp, PON_A_DEV_IFINDEX, pdev->main_netdev->ifindex) ||
+	    nla_put_u32(rsp, PON_A_DEV_MODE, pdev->mode) ||
+	    nla_put_u32(rsp, PON_A_DEV_MODES_CAP, pdev->caps->modes) ||
+	    nla_put_u32(rsp, PON_A_DEV_PLOAM_STATE, READ_ONCE(pdev->ploam)) ||
+	    nla_put_u32(rsp, PON_A_DEV_MAX_TCONTS, pdev->caps->max_tconts) ||
+	    nla_put_u32(rsp, PON_A_DEV_MAX_GEMS, pdev->caps->max_gems))
+		goto err_cancel_msg;
+
+	if (pdev->identity.serial_set &&
+	    nla_put(rsp, PON_A_DEV_SERIAL, PON_SERIAL_LEN,
+		    pdev->identity.serial))
+		goto err_cancel_msg;
+
+	if (nla_put_u8(rsp, PON_A_DEV_ENABLE, pdev->enabled))
+		goto err_cancel_msg;
+
+	genlmsg_end(rsp, hdr);
+	return 0;
+
+err_cancel_msg:
+	genlmsg_cancel(rsp, hdr);
+	return -EMSGSIZE;
+}
+
+/**
+ * pon_nl_notify_dev() - send a device notification to the mgmt group
+ * @pdev: the PON device
+ * @cmd: the notification command, PON_CMD_DEV_ADD_NTF,
+ *       PON_CMD_DEV_CHANGE_NTF or PON_CMD_DEV_DEL_NTF
+ *
+ * Does nothing when the group has no listener in the device's namespace.
+ *
+ * Context: Called with @pdev->lock held. May sleep.
+ */
+void pon_nl_notify_dev(struct pon_dev *pdev, u32 cmd)
+{
+	struct net *net = dev_net(pdev->main_netdev);
+	struct genl_info info;
+	struct sk_buff *ntf;
+
+	lockdep_assert_held(&pdev->lock);
+
+	if (!genl_has_listeners(&pon_nl_family, net, PON_NLGRP_MGMT))
+		return;
+
+	ntf = genlmsg_new(GENLMSG_DEFAULT_SIZE, GFP_KERNEL);
+	if (!ntf)
+		return;
+
+	genl_info_init_ntf(&info, &pon_nl_family, cmd);
+	genl_info_net_set(&info, net);
+	if (pon_nl_dev_fill(pdev, ntf, &info)) {
+		nlmsg_free(ntf);
+		return;
+	}
+
+	genlmsg_multicast_netns(&pon_nl_family, net, ntf, 0, PON_NLGRP_MGMT,
+				GFP_KERNEL);
+}
+
+/**
+ * pon_nl_dev_get_doit() - handle PON_CMD_DEV_GET for one device
+ * @req: the request
+ * @info: the request info, user_ptr[0] holds the device
+ *
+ * Context: Called with the device lock held by pon_device_get_locked().
+ * Return: 0, or a negative errno.
+ */
+int pon_nl_dev_get_doit(struct sk_buff *req, struct genl_info *info)
+{
+	return pon_nl_dev_reply(info, pon_nl_dev_fill);
+}
+
+/**
+ * pon_nl_dev_get_dumpit() - handle the PON_CMD_DEV_GET dump
+ * @rsp: the skb to fill
+ * @cb: the dump state, args[0] is the next device id
+ *
+ * Dumps every device in the namespace of the requesting socket.
+ *
+ * Context: Takes pon_devs_lock and each device lock in turn.
+ * Return: 0, or -EMSGSIZE when the skb is full and the dump resumes.
+ */
+int pon_nl_dev_get_dumpit(struct sk_buff *rsp, struct netlink_callback *cb)
+{
+	struct pon_dev *pdev;
+	int err = 0;
+
+	mutex_lock(&pon_devs_lock);
+	xa_for_each_start(&pon_devs, cb->args[0], pdev, cb->args[0]) {
+		mutex_lock(&pdev->lock);
+		if (dev_net(pdev->main_netdev) == sock_net(rsp->sk))
+			err = pon_nl_dev_fill(pdev, rsp, genl_info_dump(cb));
+		mutex_unlock(&pdev->lock);
+		if (err)
+			break;
+	}
+	mutex_unlock(&pon_devs_lock);
+
+	return err;
+}
+
+/**
+ * pon_nl_identity_commit() - copy accepted identity settings into the device
+ * @pdev: the PON device
+ * @id: the settings the driver took
+ *
+ * The driver took the settings. The core keeps the mode and the serial
+ * number, which dev-get reports. The registration id stays with the driver.
+ *
+ * Context: Called with @pdev->lock held.
+ */
+static void pon_nl_identity_commit(struct pon_dev *pdev,
+				   const struct pon_identity *id)
+{
+	if (id->mode_set)
+		pdev->mode = id->mode;
+	if (id->serial_set) {
+		memcpy(pdev->identity.serial, id->serial, PON_SERIAL_LEN);
+		pdev->identity.serial_set = true;
+	}
+}
+
+/**
+ * pon_nl_serial_same() - compare a requested serial number with the stored one
+ * @pdev: the PON device
+ * @id: the settings of the request, with a serial number
+ *
+ * Context: Called with @pdev->lock held.
+ * Return: true when the device holds a serial number equal to that of @id.
+ */
+static bool pon_nl_serial_same(const struct pon_dev *pdev,
+			       const struct pon_identity *id)
+{
+	return pdev->identity.serial_set &&
+	       !memcmp(pdev->identity.serial, id->serial, PON_SERIAL_LEN);
+}
+
+/**
+ * pon_nl_dev_set() - apply a PON_CMD_DEV_SET request
+ * @info: the request info, user_ptr[0] holds the device
+ * @id: zeroed space for the identity settings, which the caller clears
+ *	after the call because it holds credentials
+ *
+ * Passes the identity settings to the set_identity op, then the enable
+ * setting to the enable op. A registration id shorter than PON_REG_ID_LEN
+ * reaches the op padded with 0x00 bytes, as the note of ITU-T G.9807.1 Table
+ * C.11.25 recommends. An empty one is refused. A mode the device does not
+ * list in its capabilities is refused with an error line through
+ * pon_dev_log_level(). While the link is enabled, a mode or a serial number
+ * other than the one the device holds is refused: the OLT addresses the ONU
+ * by its serial number (G.9807.1 clause C.11.2.6.1). The serial number of
+ * eight 0x00 bytes is refused, since G.9807.1 Table C.11.23A uses it to
+ * address every ONU. A PON_CMD_DEV_CHANGE_NTF follows once the
+ * identity changed, even when the enable op then fails.
+ *
+ * Context: Called with the device lock held by pon_device_get_locked().
+ * Return: 0, or a negative errno.
+ */
+static int pon_nl_dev_set(struct genl_info *info, struct pon_identity *id)
+{
+	struct pon_dev *pdev = info->user_ptr[0];
+	bool identity = false;
+	int err;
+
+	if (info->attrs[PON_A_DEV_MODE]) {
+		id->mode = nla_get_u32(info->attrs[PON_A_DEV_MODE]);
+		id->mode_set = true;
+		if (!(pdev->caps->modes & BIT(id->mode))) {
+			NL_SET_ERR_MSG(info->extack,
+				       "mode not supported by the device");
+			pon_dev_log_level(pdev, KERN_ERR,
+					  "mode %s not supported by the device",
+					  pon_nl_mode_names[id->mode]);
+			return -EINVAL;
+		}
+		if (pdev->enabled && id->mode != pdev->mode) {
+			NL_SET_ERR_MSG_ATTR(info->extack,
+					    info->attrs[PON_A_DEV_MODE],
+					    "the mode cannot change while the link is enabled");
+			return -EBUSY;
+		}
+	}
+	if (info->attrs[PON_A_DEV_SERIAL]) {
+		nla_memcpy(id->serial, info->attrs[PON_A_DEV_SERIAL],
+			   PON_SERIAL_LEN);
+		id->serial_set = true;
+		if (!memchr_inv(id->serial, 0, PON_SERIAL_LEN)) {
+			NL_SET_ERR_MSG_ATTR(info->extack,
+					    info->attrs[PON_A_DEV_SERIAL],
+					    "the all-zero serial number addresses every ONU");
+			return -EINVAL;
+		}
+		if (pdev->enabled && !pon_nl_serial_same(pdev, id)) {
+			NL_SET_ERR_MSG_ATTR(info->extack,
+					    info->attrs[PON_A_DEV_SERIAL],
+					    "the serial number cannot change while the link is enabled");
+			return -EBUSY;
+		}
+	}
+	if (info->attrs[PON_A_DEV_REGISTRATION_ID]) {
+		struct nlattr *reg_id = info->attrs[PON_A_DEV_REGISTRATION_ID];
+
+		if (!nla_len(reg_id)) {
+			NL_SET_ERR_MSG_ATTR(info->extack, reg_id,
+					    "the registration id is empty");
+			return -EINVAL;
+		}
+		nla_memcpy(id->reg_id, reg_id, PON_REG_ID_LEN);
+		id->reg_id_len = PON_REG_ID_LEN;
+	}
+
+	identity = id->mode_set || id->serial_set || id->reg_id_len;
+
+	if (!identity && !info->attrs[PON_A_DEV_ENABLE]) {
+		NL_SET_ERR_MSG(info->extack, "no settings present");
+		return -EINVAL;
+	}
+
+	if (identity && !pdev->ops->set_identity)
+		return -EOPNOTSUPP;
+	if (info->attrs[PON_A_DEV_ENABLE] && !pdev->ops->enable)
+		return -EOPNOTSUPP;
+
+	if (identity) {
+		err = pdev->ops->set_identity(pdev, id, info->extack);
+		if (err)
+			return err;
+		pon_nl_identity_commit(pdev, id);
+	}
+
+	if (info->attrs[PON_A_DEV_ENABLE]) {
+		bool on = nla_get_u8(info->attrs[PON_A_DEV_ENABLE]);
+
+		err = pdev->ops->enable(pdev, on, info->extack);
+		if (err) {
+			if (identity)
+				pon_nl_notify_dev(pdev, PON_CMD_DEV_CHANGE_NTF);
+			return err;
+		}
+		pdev->enabled = on;
+	}
+
+	pon_nl_notify_dev(pdev, PON_CMD_DEV_CHANGE_NTF);
+
+	return 0;
+}
+
+/**
+ * pon_nl_dev_set_doit() - handle PON_CMD_DEV_SET
+ * @skb: the request
+ * @info: the request info, user_ptr[0] holds the device
+ *
+ * Applies the request through pon_nl_dev_set() and clears the copy of the
+ * identity settings on every exit, since it holds the registration id.
+ *
+ * Context: Called with the device lock held by pon_device_get_locked().
+ * Return: 0, or a negative errno.
+ */
+int pon_nl_dev_set_doit(struct sk_buff *skb, struct genl_info *info)
+{
+	struct pon_identity id = {};
+	int err;
+
+	err = pon_nl_dev_set(info, &id);
+	memzero_explicit(&id, sizeof(id));
+
+	return err;
+}
+
+/* State notifications. The instance lock is held and they may sleep. */
+
+/**
+ * pon_nl_notify_ploam() - send the activation state notification
+ * @pdev: the PON device
+ *
+ * Sends PON_CMD_PLOAM_NTF with the ONU activation state of ITU-T G.9807.1
+ * Table C.12.1 to the state group. pon_dev_state_report() owns the state
+ * and calls this once it has validated and published it, so this only
+ * formats the message.
+ *
+ * Context: Called with @pdev->lock held. May sleep.
+ */
+void pon_nl_notify_ploam(struct pon_dev *pdev)
+{
+	struct net *net = dev_net(pdev->main_netdev);
+	struct sk_buff *ntf;
+	void *hdr;
+
+	lockdep_assert_held(&pdev->lock);
+
+	if (!genl_has_listeners(&pon_nl_family, net, PON_NLGRP_STATE))
+		return;
+
+	ntf = genlmsg_new(GENLMSG_DEFAULT_SIZE, GFP_KERNEL);
+	if (!ntf)
+		return;
+
+	hdr = genlmsg_put(ntf, 0, 0, &pon_nl_family, 0, PON_CMD_PLOAM_NTF);
+	if (!hdr)
+		goto err_free;
+
+	if (nla_put_u32(ntf, PON_A_DEV_ID, pdev->id) ||
+	    nla_put_u32(ntf, PON_A_DEV_PLOAM_STATE, pdev->ploam))
+		goto err_free;
+
+	genlmsg_end(ntf, hdr);
+	genlmsg_multicast_netns(&pon_nl_family, net, ntf, 0, PON_NLGRP_STATE,
+				GFP_KERNEL);
+	return;
+
+err_free:
+	nlmsg_free(ntf);
+}
+
+/**
+ * pon_nl_obj_gen_inc() - move the object generation
+ *
+ * Called whenever an object joins or leaves a list, so that an object dump
+ * that resumes on a new generation is marked as interrupted. The value
+ * wraps from INT_MAX to 1 and so never becomes 0.
+ */
+void pon_nl_obj_gen_inc(void)
+{
+	int old = atomic_read(&pon_nl_obj_gen);
+	int new;
+
+	do {
+		new = old == INT_MAX ? 1 : old + 1;
+	} while (!atomic_try_cmpxchg(&pon_nl_obj_gen, &old, new));
+}
+
+/**
+ * pon_nl_obj_dev_dump() - dump one object list of one device
+ * @pdev: the PON device
+ * @rsp: the skb to fill
+ * @cb: the dump state, args[1] counts the objects sent
+ * @head: the object list
+ * @skip: the number of objects sent before this call
+ * @fill: puts one object into @rsp
+ *
+ * Context: Called with @pdev->lock held.
+ * Return: 0 once the list is done, or the error of @fill.
+ */
+static int pon_nl_obj_dev_dump(struct pon_dev *pdev, struct sk_buff *rsp,
+			       struct netlink_callback *cb,
+			       struct list_head *head, unsigned long skip,
+			       pon_nl_obj_fill_t fill)
+{
+	struct list_head *pos;
+	int err;
+
+	lockdep_assert_held(&pdev->lock);
+
+	cb->seq = atomic_read(&pon_nl_obj_gen);
+
+	list_for_each(pos, head) {
+		if (skip) {
+			skip--;
+			continue;
+		}
+		err = fill(pdev, pos, rsp, genl_info_dump(cb));
+		if (err)
+			return err;
+		cb->args[1]++;
+	}
+
+	cb->args[1] = 0;
+	return 0;
+}
+
+/**
+ * pon_nl_obj_dumpit() - dump one object list of each visible device
+ * @rsp: the skb to fill
+ * @cb: the dump state, args[0] is the device, args[1] the objects sent
+ * @list_offset: the offset of the list head in struct pon_dev
+ * @fill: puts one object into @rsp
+ *
+ * Walks every device in the namespace of the requesting socket, or only
+ * the device the request names. A part that put no message leaves the
+ * consistency check to the NLMSG_DONE that follows, which then carries
+ * NLM_F_DUMP_INTR when the generation moved since the previous part.
+ *
+ * Context: Takes pon_devs_lock and each device lock in turn.
+ * Return: 0, the error of @fill, or -ENODEV when the named device is not
+ * found.
+ */
+static int pon_nl_obj_dumpit(struct sk_buff *rsp, struct netlink_callback *cb,
+			     size_t list_offset, pon_nl_obj_fill_t fill)
+{
+	const struct genl_info *info = genl_info_dump(cb);
+	unsigned long resumed = cb->args[0];
+	unsigned long last = ULONG_MAX;
+	struct pon_dev *pdev;
+	bool found = false;
+	int err = 0;
+
+	if (info->attrs[PON_A_DEV_ID]) {
+		last = nla_get_u32(info->attrs[PON_A_DEV_ID]);
+		cb->args[0] = last;
+	}
+
+	mutex_lock(&pon_devs_lock);
+	cb->seq = atomic_read(&pon_nl_obj_gen);
+	xa_for_each_range(&pon_devs, cb->args[0], pdev, cb->args[0], last) {
+		struct list_head *head = (void *)pdev + list_offset;
+		unsigned long skip = pdev->id == resumed ? cb->args[1] : 0;
+
+		mutex_lock(&pdev->lock);
+		if (dev_net(pdev->main_netdev) == sock_net(rsp->sk)) {
+			found = true;
+			err = pon_nl_obj_dev_dump(pdev, rsp, cb, head, skip,
+						  fill);
+		}
+		mutex_unlock(&pdev->lock);
+		if (err)
+			break;
+	}
+	mutex_unlock(&pon_devs_lock);
+
+	if (!found && last != ULONG_MAX) {
+		NL_SET_ERR_MSG(cb->extack, "no such device");
+		return -ENODEV;
+	}
+
+	if (rsp->len)
+		nl_dump_check_consistent(cb, nlmsg_hdr(rsp));
+
+	return err;
+}
+
+/* T-CONT */
+
+/**
+ * pon_gem_on_tcont() - check whether a GEM port is bound to a T-CONT
+ * @gem: the GEM port
+ * @index: the T-CONT index
+ *
+ * Return: true when @gem names the T-CONT @index.
+ */
+static bool pon_gem_on_tcont(const struct pon_gem *gem, u16 index)
+{
+	return gem->cfg.tcont_valid && gem->cfg.tcont_index == index;
+}
+
+/**
+ * pon_tcont_in_use() - check whether a GEM port names a T-CONT
+ * @pdev: the PON device
+ * @index: the T-CONT index
+ *
+ * Return: true while a GEM port object holds @index as its T-CONT.
+ */
+bool pon_tcont_in_use(struct pon_dev *pdev, u16 index)
+{
+	struct pon_gem *gem;
+
+	lockdep_assert_held(&pdev->lock);
+	list_for_each_entry(gem, &pdev->gems, list)
+		if (pon_gem_on_tcont(gem, index))
+			return true;
+	return false;
+}
+
+/**
+ * pon_tcont_alloc_taken() - check whether another T-CONT holds an alloc-id
+ * @pdev: the PON device
+ * @index: the T-CONT that asks for @alloc_id
+ * @alloc_id: the alloc-id
+ *
+ * ITU-T G.988 clause 9.2.2 relates alloc-ids and T-CONTs one to one and
+ * leaves several T-CONTs on one alloc-id undefined.
+ *
+ * Return: true while a T-CONT other than @index holds @alloc_id.
+ */
+bool pon_tcont_alloc_taken(struct pon_dev *pdev, u16 index, u16 alloc_id)
+{
+	struct pon_tcont *tcont;
+
+	lockdep_assert_held(&pdev->lock);
+	list_for_each_entry(tcont, &pdev->tconts, list)
+		if (tcont->cfg.index != index &&
+		    tcont->cfg.alloc_id == alloc_id)
+			return true;
+	return false;
+}
+
+/**
+ * pon_tcont_gems_rebind() - move the GEM ports of a T-CONT to another alloc-id
+ * @pdev: the PON device
+ * @index: the T-CONT index
+ * @old_alloc_id: the alloc-id the GEM ports are bound to
+ * @new_alloc_id: the alloc-id to bind them to
+ * @extack: netlink extended ack
+ *
+ * Programs each GEM port of the T-CONT again through the gem_add op. If the
+ * driver refuses one, the refused GEM port and the ones already moved are
+ * programmed again with @old_alloc_id, so that the core and the driver agree.
+ *
+ * Return: 0, or the error of the refused GEM port.
+ */
+int pon_tcont_gems_rebind(struct pon_dev *pdev, u16 index, u16 old_alloc_id,
+			  u16 new_alloc_id, struct netlink_ext_ack *extack)
+{
+	struct pon_gem *gem;
+	int err;
+
+	lockdep_assert_held(&pdev->lock);
+
+	list_for_each_entry(gem, &pdev->gems, list) {
+		if (!pon_gem_on_tcont(gem, index) ||
+		    gem->cfg.alloc_id == new_alloc_id)
+			continue;
+
+		gem->cfg.alloc_id = new_alloc_id;
+		err = pdev->ops->gem_add(pdev, &gem->cfg, extack);
+		if (err)
+			goto err_restore;
+	}
+
+	return 0;
+
+err_restore:
+	NL_SET_ERR_MSG_WEAK(extack,
+			    "a GEM port of the T-CONT cannot move to the alloc-id");
+
+	list_for_each_entry_from_reverse(gem, &pdev->gems, list) {
+		if (!pon_gem_on_tcont(gem, index) ||
+		    gem->cfg.alloc_id != new_alloc_id)
+			continue;
+
+		gem->cfg.alloc_id = old_alloc_id;
+		if (pdev->ops->gem_add(pdev, &gem->cfg, NULL))
+			netdev_warn(pdev->main_netdev,
+				    "restoring GEM %u to alloc-id %u failed\n",
+				    gem->cfg.id, old_alloc_id);
+	}
+
+	return err;
+}
+
+/**
+ * pon_nl_tcont_set_doit() - handle PON_CMD_TCONT_SET
+ * @skb: the request
+ * @info: the request info, user_ptr[0] holds the device
+ *
+ * Creates the T-CONT, or binds an existing one to another alloc-id and
+ * moves its GEM ports with it. The alloc-id must lie in the ranges of
+ * ITU-T G.9807.1 Table C.6.5 and must not be a serial number grant. Alloc-ids
+ * and T-CONTs relate one to one, as ITU-T G.988 clause 9.2.2 defines.
+ *
+ * Context: Called with the device lock held by pon_device_get_locked().
+ * Return: 0, or a negative errno.
+ */
+int pon_nl_tcont_set_doit(struct sk_buff *skb, struct genl_info *info)
+{
+	struct pon_dev *pdev = info->user_ptr[0];
+	struct pon_tcont_cfg cfg = {};
+	struct pon_tcont *tcont;
+	bool changed = false;
+	bool is_new;
+	int err;
+
+	if (GENL_REQ_ATTR_CHECK(info, PON_A_TCONT_INDEX) ||
+	    GENL_REQ_ATTR_CHECK(info, PON_A_TCONT_ALLOC_ID))
+		return -EINVAL;
+
+	cfg.index = nla_get_u32(info->attrs[PON_A_TCONT_INDEX]);
+	cfg.alloc_id = nla_get_u32(info->attrs[PON_A_TCONT_ALLOC_ID]);
+
+	if (cfg.index >= pdev->caps->max_tconts) {
+		NL_SET_ERR_MSG_ATTR(info->extack,
+				    info->attrs[PON_A_TCONT_INDEX],
+				    "T-CONT index out of range");
+		return -ERANGE;
+	}
+	if (cfg.alloc_id >= PON_PLOAM_ALLOC_ID_SN_GRANT_MIN &&
+	    cfg.alloc_id <= PON_PLOAM_ALLOC_ID_SN_GRANT_2G5) {
+		NL_SET_ERR_MSG_ATTR(info->extack,
+				    info->attrs[PON_A_TCONT_ALLOC_ID],
+				    "the alloc-id is a serial number grant");
+		return -EINVAL;
+	}
+	if (pon_tcont_alloc_taken(pdev, cfg.index, cfg.alloc_id)) {
+		NL_SET_ERR_MSG_ATTR(info->extack,
+				    info->attrs[PON_A_TCONT_ALLOC_ID],
+				    "another T-CONT holds the alloc-id");
+		return -EEXIST;
+	}
+
+	tcont = pon_tcont_find(pdev, cfg.index);
+	is_new = !tcont;
+	if (is_new) {
+		tcont = kzalloc_obj(*tcont, GFP_KERNEL);
+		if (!tcont)
+			return -ENOMEM;
+	} else {
+		changed = tcont->cfg.alloc_id != cfg.alloc_id;
+	}
+
+	err = pdev->ops->tcont_set(pdev, &cfg, info->extack);
+	if (err)
+		goto err_free_tcont;
+
+	if (changed) {
+		err = pon_tcont_gems_rebind(pdev, cfg.index,
+					    tcont->cfg.alloc_id, cfg.alloc_id,
+					    info->extack);
+		if (err)
+			goto err_restore_tcont;
+	}
+
+	tcont->cfg = cfg;
+	if (is_new) {
+		list_add_tail(&tcont->list, &pdev->tconts);
+		pon_nl_obj_gen_inc();
+	}
+
+	pon_dev_carrier_update(pdev);
+
+	if (is_new)
+		pon_nl_notify_tcont(pdev, tcont, PON_CMD_TCONT_ADD_NTF);
+	else if (changed)
+		pon_nl_notify_tcont(pdev, tcont, PON_CMD_TCONT_CHANGE_NTF);
+
+	return 0;
+
+err_restore_tcont:
+	if (pdev->ops->tcont_set(pdev, &tcont->cfg, NULL))
+		netdev_warn(pdev->main_netdev,
+			    "restoring T-CONT %u to alloc-id %u failed\n",
+			    tcont->cfg.index, tcont->cfg.alloc_id);
+err_free_tcont:
+	if (is_new)
+		kfree(tcont);
+	return err;
+}
+
+/**
+ * pon_nl_tcont_fill() - put one T-CONT message into an skb
+ * @pdev: the PON device
+ * @tcont: the T-CONT
+ * @rsp: the skb to fill
+ * @info: the request, or the notification info
+ *
+ * Context: Called with @pdev->lock held.
+ * Return: 0, or -EMSGSIZE when the message does not fit.
+ */
+static int
+pon_nl_tcont_fill(struct pon_dev *pdev, struct pon_tcont *tcont,
+		  struct sk_buff *rsp, const struct genl_info *info)
+{
+	struct pon_tcont_cfg *cfg = &tcont->cfg;
+	void *hdr;
+
+	hdr = genlmsg_iput(rsp, info);
+	if (!hdr)
+		return -EMSGSIZE;
+
+	if (nla_put_u32(rsp, PON_A_TCONT_DEV_ID, pdev->id) ||
+	    nla_put_u32(rsp, PON_A_TCONT_INDEX, cfg->index) ||
+	    nla_put_u32(rsp, PON_A_TCONT_ALLOC_ID, cfg->alloc_id))
+		goto err_cancel_msg;
+
+	genlmsg_end(rsp, hdr);
+	return 0;
+
+err_cancel_msg:
+	genlmsg_cancel(rsp, hdr);
+	return -EMSGSIZE;
+}
+
+/**
+ * pon_nl_tcont_fill_pos() - pon_nl_obj_fill_t for the T-CONT list
+ * @pdev: the PON device
+ * @pos: the list entry of the T-CONT
+ * @rsp: the skb to fill
+ * @info: the dump info
+ *
+ * Return: the result of pon_nl_tcont_fill().
+ */
+static int pon_nl_tcont_fill_pos(struct pon_dev *pdev, struct list_head *pos,
+				 struct sk_buff *rsp,
+				 const struct genl_info *info)
+{
+	return pon_nl_tcont_fill(pdev, list_entry(pos, struct pon_tcont, list),
+				 rsp, info);
+}
+
+/**
+ * pon_nl_notify_tcont() - send a T-CONT notification to the mgmt group
+ * @pdev: the PON device
+ * @tcont: the T-CONT
+ * @cmd: the notification command, PON_CMD_TCONT_ADD_NTF,
+ *       PON_CMD_TCONT_CHANGE_NTF or PON_CMD_TCONT_DEL_NTF
+ *
+ * Context: Called with @pdev->lock held. May sleep.
+ */
+void pon_nl_notify_tcont(struct pon_dev *pdev, struct pon_tcont *tcont, u32 cmd)
+{
+	pon_nl_notify_obj(pdev, &tcont->list, cmd, pon_nl_tcont_fill_pos);
+}
+
+/**
+ * pon_nl_tcont_get_doit() - handle PON_CMD_TCONT_GET for one T-CONT
+ * @skb: the request
+ * @info: the request info, user_ptr[0] holds the device
+ *
+ * Context: Called with the device lock held by pon_device_get_locked().
+ * Return: 0, -ENOENT when the T-CONT does not exist, or a negative errno.
+ */
+int pon_nl_tcont_get_doit(struct sk_buff *skb, struct genl_info *info)
+{
+	struct pon_dev *pdev = info->user_ptr[0];
+	struct pon_tcont *tcont;
+
+	if (GENL_REQ_ATTR_CHECK(info, PON_A_TCONT_INDEX))
+		return -EINVAL;
+
+	tcont = pon_tcont_find(pdev,
+			       nla_get_u32(info->attrs[PON_A_TCONT_INDEX]));
+	if (!tcont) {
+		NL_SET_ERR_MSG(info->extack, "no such T-CONT");
+		return -ENOENT;
+	}
+
+	return pon_nl_obj_reply(info, &tcont->list, pon_nl_tcont_fill_pos);
+}
+
+/**
+ * pon_nl_tcont_get_dumpit() - handle the PON_CMD_TCONT_GET dump
+ * @rsp: the skb to fill
+ * @cb: the dump state
+ *
+ * Return: the result of pon_nl_obj_dumpit().
+ */
+int pon_nl_tcont_get_dumpit(struct sk_buff *rsp, struct netlink_callback *cb)
+{
+	return pon_nl_obj_dumpit(rsp, cb, offsetof(struct pon_dev, tconts),
+				 pon_nl_tcont_fill_pos);
+}
+
+/**
+ * pon_nl_tcont_del_doit() - handle PON_CMD_TCONT_DEL
+ * @skb: the request
+ * @info: the request info, user_ptr[0] holds the device
+ *
+ * Refuses while a GEM port names the T-CONT.
+ *
+ * Context: Called with the device lock held by pon_device_get_locked().
+ * Return: 0, -ENOENT when the T-CONT does not exist, -EBUSY while a GEM port
+ * names it, or another negative errno.
+ */
+int pon_nl_tcont_del_doit(struct sk_buff *skb, struct genl_info *info)
+{
+	struct pon_dev *pdev = info->user_ptr[0];
+	struct pon_tcont *tcont;
+	u16 index;
+	int err;
+
+	if (GENL_REQ_ATTR_CHECK(info, PON_A_TCONT_INDEX))
+		return -EINVAL;
+
+	index = nla_get_u32(info->attrs[PON_A_TCONT_INDEX]);
+	tcont = pon_tcont_find(pdev, index);
+	if (!tcont) {
+		NL_SET_ERR_MSG(info->extack, "no such T-CONT");
+		return -ENOENT;
+	}
+	if (pon_tcont_in_use(pdev, index)) {
+		NL_SET_ERR_MSG(info->extack, "a GEM port names the T-CONT");
+		return -EBUSY;
+	}
+
+	err = pdev->ops->tcont_clear(pdev, &tcont->cfg, info->extack);
+	if (err)
+		return err;
+
+	pon_nl_notify_tcont(pdev, tcont, PON_CMD_TCONT_DEL_NTF);
+
+	list_del(&tcont->list);
+	kfree(tcont);
+	pon_nl_obj_gen_inc();
+
+	pon_dev_carrier_update(pdev);
+
+	return 0;
+}
+
+/* GEM ports */
+
+/**
+ * pon_gems_full() - check whether the device holds all the GEM ports it can
+ * @pdev: the PON device
+ *
+ * Return: true when the GEM port objects number the device's max_gems.
+ */
+bool pon_gems_full(struct pon_dev *pdev)
+{
+	struct pon_gem *gem;
+	unsigned int n = 0;
+
+	lockdep_assert_held(&pdev->lock);
+	list_for_each_entry(gem, &pdev->gems, list)
+		n++;
+	return n >= pdev->caps->max_gems;
+}
+
+/**
+ * pon_gem_cfg_same() - compare two GEM port configurations
+ * @existing: the configuration of the GEM port the instance holds
+ * @requested: the configuration a request names
+ *
+ * Return: true when every member that a request sets is equal.
+ */
+static bool pon_gem_cfg_same(const struct pon_gem_cfg *existing,
+			     const struct pon_gem_cfg *requested)
+{
+	return existing->id == requested->id &&
+	       existing->dir == requested->dir &&
+	       existing->tcont_valid == requested->tcont_valid &&
+	       existing->tcont_index == requested->tcont_index &&
+	       existing->alloc_id == requested->alloc_id &&
+	       existing->key_ring == requested->key_ring;
+}
+
+/**
+ * pon_nl_gem_new_doit() - handle PON_CMD_GEM_NEW
+ * @skb: the request
+ * @info: the request info, user_ptr[0] holds the device
+ *
+ * Creates the GEM port of the GEM port network CTP of ITU-T G.988 clause
+ * 9.2.3. The policy holds the GEM port ID to the assignable range of ITU-T
+ * G.9807.1 Table C.6.6. The broadcast key ring of clause 9.2.3 is refused.
+ * A GEM port with an upstream half names its T-CONT and takes the alloc-id of
+ * it. A downstream GEM port names none. A request for a GEM port that exists
+ * with the same attributes succeeds and changes nothing.
+ *
+ * Context: Called with the device lock held by pon_device_get_locked().
+ * Return: 0, or a negative errno.
+ */
+int pon_nl_gem_new_doit(struct sk_buff *skb, struct genl_info *info)
+{
+	struct pon_dev *pdev = info->user_ptr[0];
+	struct pon_gem_cfg cfg = {};
+	struct pon_gem *gem;
+	int err;
+
+	if (GENL_REQ_ATTR_CHECK(info, PON_A_GEM_ID) ||
+	    GENL_REQ_ATTR_CHECK(info, PON_A_GEM_DIR))
+		return -EINVAL;
+
+	cfg.id = nla_get_u32(info->attrs[PON_A_GEM_ID]);
+	cfg.dir = nla_get_u32(info->attrs[PON_A_GEM_DIR]);
+	if (cfg.dir == PON_GEM_DIR_DOWNSTREAM) {
+		if (info->attrs[PON_A_GEM_TCONT_INDEX]) {
+			NL_SET_ERR_MSG_ATTR(info->extack,
+					    info->attrs[PON_A_GEM_TCONT_INDEX],
+					    "a downstream GEM port rides no T-CONT");
+			return -EINVAL;
+		}
+	} else if (GENL_REQ_ATTR_CHECK(info, PON_A_GEM_TCONT_INDEX)) {
+		NL_SET_ERR_MSG(info->extack,
+			       "an upstream GEM port needs a T-CONT");
+		return -EINVAL;
+	}
+	if (info->attrs[PON_A_GEM_TCONT_INDEX]) {
+		cfg.tcont_index =
+			nla_get_u32(info->attrs[PON_A_GEM_TCONT_INDEX]);
+		cfg.tcont_valid = true;
+	}
+	if (info->attrs[PON_A_GEM_KEY_RING])
+		cfg.key_ring = nla_get_u32(info->attrs[PON_A_GEM_KEY_RING]);
+	if (cfg.key_ring == PON_GEM_KEY_RING_BROADCAST) {
+		NL_SET_ERR_MSG_ATTR(info->extack,
+				    info->attrs[PON_A_GEM_KEY_RING],
+				    "broadcast keys are not supported");
+		return -EOPNOTSUPP;
+	}
+
+	if (cfg.tcont_valid) {
+		struct pon_tcont *tcont;
+
+		tcont = pon_tcont_find(pdev, cfg.tcont_index);
+		if (!tcont) {
+			NL_SET_ERR_MSG_ATTR(info->extack,
+					    info->attrs[PON_A_GEM_TCONT_INDEX],
+					    "no such T-CONT");
+			return -ENOENT;
+		}
+		cfg.alloc_id = tcont->cfg.alloc_id;
+	}
+
+	gem = pon_gem_find(pdev, cfg.id);
+	if (gem) {
+		if (!pon_gem_cfg_same(&gem->cfg, &cfg)) {
+			NL_SET_ERR_MSG(info->extack,
+				       "the GEM port exists with other attributes");
+			return -EEXIST;
+		}
+
+		return 0;
+	}
+
+	if (pon_gems_full(pdev)) {
+		NL_SET_ERR_MSG(info->extack,
+			       "the device holds no more GEM ports");
+		return -ENOSPC;
+	}
+
+	gem = kzalloc_obj(*gem, GFP_KERNEL);
+	if (!gem)
+		return -ENOMEM;
+
+	err = pdev->ops->gem_add(pdev, &cfg, info->extack);
+	if (err)
+		goto err_free_gem;
+
+	gem->cfg = cfg;
+	list_add_tail(&gem->list, &pdev->gems);
+	pon_nl_obj_gen_inc();
+	pon_dev_carrier_update(pdev);
+	pon_nl_notify_gem(pdev, gem, PON_CMD_GEM_ADD_NTF);
+
+	return 0;
+
+err_free_gem:
+	kfree(gem);
+	return err;
+}
+
+/**
+ * pon_nl_gem_del_doit() - handle PON_CMD_GEM_DEL
+ * @skb: the request
+ * @info: the request info, user_ptr[0] holds the device
+ *
+ * Refuses while a GEM network device is attached to the GEM port.
+ *
+ * Context: Called with the device lock held by pon_device_get_locked().
+ * Return: 0, -ENOENT when the GEM port does not exist, -EBUSY while a GEM
+ * network device is attached, or another negative errno.
+ */
+int pon_nl_gem_del_doit(struct sk_buff *skb, struct genl_info *info)
+{
+	struct pon_dev *pdev = info->user_ptr[0];
+	struct pon_gem *gem;
+	u16 gem_id;
+	int err;
+
+	if (GENL_REQ_ATTR_CHECK(info, PON_A_GEM_ID))
+		return -EINVAL;
+
+	gem_id = nla_get_u32(info->attrs[PON_A_GEM_ID]);
+	gem = pon_gem_find(pdev, gem_id);
+	if (!gem) {
+		NL_SET_ERR_MSG(info->extack, "no such GEM port");
+		return -ENOENT;
+	}
+	if (gem->gem_netdev) {
+		NL_SET_ERR_MSG(info->extack,
+			       "a GEM network device is attached");
+		return -EBUSY;
+	}
+
+	err = pdev->ops->gem_del(pdev, gem_id, info->extack);
+	if (err)
+		return err;
+
+	pon_nl_notify_gem(pdev, gem, PON_CMD_GEM_DEL_NTF);
+
+	list_del(&gem->list);
+	kfree(gem);
+	pon_nl_obj_gen_inc();
+
+	pon_dev_carrier_update(pdev);
+
+	return 0;
+}
+
+/**
+ * pon_nl_gem_fill() - put one GEM port message into an skb
+ * @pdev: the PON device
+ * @gem: the GEM port
+ * @rsp: the skb to fill
+ * @info: the request, or the notification info
+ *
+ * Context: Called with @pdev->lock held.
+ * Return: 0, or -EMSGSIZE when the message does not fit.
+ */
+static int
+pon_nl_gem_fill(struct pon_dev *pdev, struct pon_gem *gem,
+		struct sk_buff *rsp, const struct genl_info *info)
+{
+	struct pon_gem_cfg *cfg = &gem->cfg;
+	void *hdr;
+
+	hdr = genlmsg_iput(rsp, info);
+	if (!hdr)
+		return -EMSGSIZE;
+
+	if (nla_put_u32(rsp, PON_A_GEM_DEV_ID, pdev->id) ||
+	    nla_put_u32(rsp, PON_A_GEM_ID, cfg->id) ||
+	    nla_put_u32(rsp, PON_A_GEM_DIR, cfg->dir) ||
+	    nla_put_u32(rsp, PON_A_GEM_KEY_RING, cfg->key_ring))
+		goto err_cancel_msg;
+
+	if (cfg->tcont_valid &&
+	    nla_put_u32(rsp, PON_A_GEM_TCONT_INDEX, cfg->tcont_index))
+		goto err_cancel_msg;
+
+	genlmsg_end(rsp, hdr);
+	return 0;
+
+err_cancel_msg:
+	genlmsg_cancel(rsp, hdr);
+	return -EMSGSIZE;
+}
+
+/**
+ * pon_nl_gem_fill_pos() - pon_nl_obj_fill_t for the GEM port list
+ * @pdev: the PON device
+ * @pos: the list entry of the GEM port
+ * @rsp: the skb to fill
+ * @info: the dump info
+ *
+ * Return: the result of pon_nl_gem_fill().
+ */
+static int pon_nl_gem_fill_pos(struct pon_dev *pdev, struct list_head *pos,
+			       struct sk_buff *rsp,
+			       const struct genl_info *info)
+{
+	return pon_nl_gem_fill(pdev, list_entry(pos, struct pon_gem, list), rsp,
+			       info);
+}
+
+/**
+ * pon_nl_notify_gem() - send a GEM port notification to the mgmt group
+ * @pdev: the PON device
+ * @gem: the GEM port
+ * @cmd: the notification command, PON_CMD_GEM_ADD_NTF or
+ *       PON_CMD_GEM_DEL_NTF
+ *
+ * Context: Called with @pdev->lock held. May sleep.
+ */
+void pon_nl_notify_gem(struct pon_dev *pdev, struct pon_gem *gem, u32 cmd)
+{
+	pon_nl_notify_obj(pdev, &gem->list, cmd, pon_nl_gem_fill_pos);
+}
+
+/**
+ * pon_nl_gem_get_doit() - handle PON_CMD_GEM_GET for one GEM port
+ * @skb: the request
+ * @info: the request info, user_ptr[0] holds the device
+ *
+ * Context: Called with the device lock held by pon_device_get_locked().
+ * Return: 0, -ENOENT when the GEM port does not exist, or a negative errno.
+ */
+int pon_nl_gem_get_doit(struct sk_buff *skb, struct genl_info *info)
+{
+	struct pon_dev *pdev = info->user_ptr[0];
+	struct pon_gem *gem;
+
+	if (GENL_REQ_ATTR_CHECK(info, PON_A_GEM_ID))
+		return -EINVAL;
+
+	gem = pon_gem_find(pdev, nla_get_u32(info->attrs[PON_A_GEM_ID]));
+	if (!gem) {
+		NL_SET_ERR_MSG(info->extack, "no such GEM port");
+		return -ENOENT;
+	}
+
+	return pon_nl_obj_reply(info, &gem->list, pon_nl_gem_fill_pos);
+}
+
+/**
+ * pon_nl_gem_get_dumpit() - handle the PON_CMD_GEM_GET dump
+ * @rsp: the skb to fill
+ * @cb: the dump state
+ *
+ * Return: the result of pon_nl_obj_dumpit().
+ */
+int pon_nl_gem_get_dumpit(struct sk_buff *rsp, struct netlink_callback *cb)
+{
+	return pon_nl_obj_dumpit(rsp, cb, offsetof(struct pon_dev, gems),
+				 pon_nl_gem_fill_pos);
+}
+
+/* GEM upstream classifier */
+
+/**
+ * pon_nl_gem_map_parse() - read a classifier rule from a request
+ * @info: the request info
+ * @cfg: receives the rule, members without an attribute stay unset
+ */
+static void pon_nl_gem_map_parse(struct genl_info *info,
+				 struct pon_gem_map_cfg *cfg)
+{
+	cfg->gem_id = nla_get_u32(info->attrs[PON_A_GEM_MAP_GEM_ID]);
+	if (info->attrs[PON_A_GEM_MAP_TAG]) {
+		cfg->tag_valid = true;
+		cfg->tagged = nla_get_u32(info->attrs[PON_A_GEM_MAP_TAG]) ==
+			      PON_GEM_MAP_TAG_TAGGED;
+	}
+	if (info->attrs[PON_A_GEM_MAP_VID]) {
+		cfg->vid_valid = true;
+		cfg->vid = nla_get_u32(info->attrs[PON_A_GEM_MAP_VID]);
+	}
+	if (info->attrs[PON_A_GEM_MAP_PBIT]) {
+		cfg->pbit_valid = true;
+		cfg->pbit = nla_get_u32(info->attrs[PON_A_GEM_MAP_PBIT]);
+	}
+	if (info->attrs[PON_A_GEM_MAP_DSCP]) {
+		cfg->dscp_valid = true;
+		cfg->dscp = nla_get_u32(info->attrs[PON_A_GEM_MAP_DSCP]);
+	}
+}
+
+/**
+ * pon_nl_gem_map_new_doit() - handle PON_CMD_GEM_MAP_NEW
+ * @skb: the request
+ * @info: the request info, user_ptr[0] holds the device
+ *
+ * Programs an upstream classifier rule through the gem_map_set op. The rule
+ * joins the list only once the driver has taken it. A rule that exists is
+ * programmed again and sends no notification.
+ *
+ * Context: Called with the device lock held by pon_device_get_locked().
+ * Return: 0, -EOPNOTSUPP without a classifier in the driver, or a negative
+ * errno.
+ */
+int pon_nl_gem_map_new_doit(struct sk_buff *skb, struct genl_info *info)
+{
+	struct pon_dev *pdev = info->user_ptr[0];
+	struct pon_gem_map_cfg cfg = {};
+	struct pon_gem_map *map;
+	bool is_new;
+	int err;
+
+	if (GENL_REQ_ATTR_CHECK(info, PON_A_GEM_MAP_GEM_ID))
+		return -EINVAL;
+
+	if (!pdev->ops->gem_map_set) {
+		NL_SET_ERR_MSG(info->extack,
+			       "the driver carries no GEM classifier");
+		return -EOPNOTSUPP;
+	}
+
+	pon_nl_gem_map_parse(info, &cfg);
+
+	map = pon_gem_map_find(pdev, &cfg);
+	is_new = !map;
+	if (is_new) {
+		map = kzalloc_obj(*map, GFP_KERNEL);
+		if (!map)
+			return -ENOMEM;
+	}
+
+	err = pdev->ops->gem_map_set(pdev, &cfg, info->extack);
+	if (err)
+		goto err_free_map;
+
+	map->cfg = cfg;
+	if (is_new) {
+		list_add_tail(&map->list, &pdev->gem_maps);
+		pon_nl_obj_gen_inc();
+		pon_nl_notify_gem_map(pdev, map, PON_CMD_GEM_MAP_ADD_NTF);
+	}
+
+	return 0;
+
+err_free_map:
+	if (is_new)
+		kfree(map);
+	return err;
+}
+
+/**
+ * pon_nl_gem_map_del_doit() - handle PON_CMD_GEM_MAP_DEL
+ * @skb: the request
+ * @info: the request info, user_ptr[0] holds the device
+ *
+ * Context: Called with the device lock held by pon_device_get_locked().
+ * Return: 0, -EOPNOTSUPP without a classifier in the driver, -ENOENT when
+ * the rule does not exist, or another negative errno.
+ */
+int pon_nl_gem_map_del_doit(struct sk_buff *skb, struct genl_info *info)
+{
+	struct pon_dev *pdev = info->user_ptr[0];
+	struct pon_gem_map_cfg cfg = {};
+	struct pon_gem_map *map;
+	int err;
+
+	if (GENL_REQ_ATTR_CHECK(info, PON_A_GEM_MAP_GEM_ID))
+		return -EINVAL;
+
+	if (!pdev->ops->gem_map_del) {
+		NL_SET_ERR_MSG(info->extack,
+			       "the driver carries no GEM classifier");
+		return -EOPNOTSUPP;
+	}
+
+	pon_nl_gem_map_parse(info, &cfg);
+
+	map = pon_gem_map_find(pdev, &cfg);
+	if (!map) {
+		NL_SET_ERR_MSG(info->extack, "no such classifier rule");
+		return -ENOENT;
+	}
+
+	err = pdev->ops->gem_map_del(pdev, &cfg, info->extack);
+	if (err)
+		return err;
+
+	pon_nl_notify_gem_map(pdev, map, PON_CMD_GEM_MAP_DEL_NTF);
+	list_del(&map->list);
+	kfree(map);
+	pon_nl_obj_gen_inc();
+
+	return 0;
+}
+
+/**
+ * pon_nl_omci_register_doit() - handle PON_CMD_OMCI_REGISTER
+ * @skb: the request
+ * @info: the request info, user_ptr[0] holds the device
+ *
+ * Makes the requesting socket the owner of the OMCI channel of the device.
+ * The received OMCI PDUs go to the owner as PON_CMD_OMCI_NTF. A second
+ * request from the owner succeeds.
+ *
+ * Context: Called with the device lock held by pon_device_get_locked().
+ * Return: 0, or -EBUSY while another socket owns the channel.
+ */
+int pon_nl_omci_register_doit(struct sk_buff *skb, struct genl_info *info)
+{
+	struct pon_dev *pdev = info->user_ptr[0];
+	int err;
+
+	err = pon_omci_register(pdev, info->snd_portid);
+	if (err)
+		NL_SET_ERR_MSG(info->extack,
+			       "another socket owns the OMCI channel");
+
+	return err;
+}
+
+/**
+ * pon_nl_omci_tx_doit() - handle PON_CMD_OMCI_TX
+ * @skb: the request
+ * @info: the request info, user_ptr[0] holds the device
+ *
+ * Sends one OMCI PDU without its MIC. Only the owner of the OMCI channel
+ * sends. The PDU holds at least the 10 byte header of the extended message
+ * format of ITU-T G.988 Table 11.2-2. The policy limits it to the 1980 byte
+ * PDU of G.988 clause 11.2.5 less the 4 byte MIC. pon_omci_xmit() refuses a
+ * PDU whose length does not match its format and sets the extended ack text
+ * for that refusal only.
+ *
+ * Context: Called with the device lock held by pon_device_get_locked().
+ * Return: 0, -EPERM when the socket does not own the channel, -EINVAL for
+ * a PDU that is too short or does not match its format, or the error of
+ * pon_omci_xmit().
+ */
+int pon_nl_omci_tx_doit(struct sk_buff *skb, struct genl_info *info)
+{
+	struct pon_dev *pdev = info->user_ptr[0];
+	struct nlattr *pdu;
+
+	if (GENL_REQ_ATTR_CHECK(info, PON_A_OMCI_PDU))
+		return -EINVAL;
+
+	if (READ_ONCE(pdev->omci_portid) != info->snd_portid) {
+		NL_SET_ERR_MSG(info->extack,
+			       "the socket does not own the OMCI channel");
+		return -EPERM;
+	}
+
+	pdu = info->attrs[PON_A_OMCI_PDU];
+
+	return pon_omci_xmit(pdev, nla_data(pdu), nla_len(pdu), info->extack);
+}
+
+/**
+ * pon_nl_omci_ntf() - pass a received OMCI PDU to the channel owner
+ * @pdev: the PON device
+ * @skb: the verified PDU
+ *
+ * Sends PON_CMD_OMCI_NTF as a unicast to the socket that owns the OMCI
+ * channel. The caller keeps @skb.
+ *
+ * Context: Called with @pdev->lock held. May sleep.
+ * Return: 0, -ENOTCONN without an owner, -ENOMEM, -EMSGSIZE, or the error
+ * of genlmsg_unicast().
+ */
+int pon_nl_omci_ntf(struct pon_dev *pdev, const struct sk_buff *skb)
+{
+	u32 portid = READ_ONCE(pdev->omci_portid);
+	struct sk_buff *ntf;
+	struct nlattr *pdu;
+	void *hdr;
+
+	lockdep_assert_held(&pdev->lock);
+
+	if (!portid)
+		return -ENOTCONN;
+
+	ntf = genlmsg_new(nla_total_size(sizeof(u32)) +
+			  nla_total_size(skb->len), GFP_KERNEL);
+	if (!ntf)
+		return -ENOMEM;
+
+	hdr = genlmsg_put(ntf, 0, 0, &pon_nl_family, 0, PON_CMD_OMCI_NTF);
+	if (!hdr)
+		goto err_free;
+
+	if (nla_put_u32(ntf, PON_A_OMCI_DEV_ID, pdev->id))
+		goto err_free;
+
+	pdu = nla_reserve(ntf, PON_A_OMCI_PDU, skb->len);
+	if (!pdu || skb_copy_bits(skb, 0, nla_data(pdu), skb->len))
+		goto err_free;
+
+	genlmsg_end(ntf, hdr);
+
+	return genlmsg_unicast(dev_net(pdev->main_netdev), ntf, portid);
+
+err_free:
+	nlmsg_free(ntf);
+	return -EMSGSIZE;
+}
+
+/* Upstream classifier rules */
+
+/**
+ * pon_nl_gem_map_fill() - put one classifier rule message into an skb
+ * @pdev: the PON device
+ * @map: the rule
+ * @rsp: the skb to fill
+ * @info: the request, or the notification info
+ *
+ * Context: Called with @pdev->lock held.
+ * Return: 0, or -EMSGSIZE when the message does not fit.
+ */
+static int pon_nl_gem_map_fill(struct pon_dev *pdev, struct pon_gem_map *map,
+			       struct sk_buff *rsp,
+			       const struct genl_info *info)
+{
+	struct pon_gem_map_cfg *cfg = &map->cfg;
+	void *hdr;
+
+	hdr = genlmsg_iput(rsp, info);
+	if (!hdr)
+		return -EMSGSIZE;
+
+	if (nla_put_u32(rsp, PON_A_GEM_MAP_DEV_ID, pdev->id) ||
+	    nla_put_u32(rsp, PON_A_GEM_MAP_GEM_ID, cfg->gem_id))
+		goto err_cancel_msg;
+
+	if (cfg->tag_valid &&
+	    nla_put_u32(rsp, PON_A_GEM_MAP_TAG, cfg->tagged))
+		goto err_cancel_msg;
+	if (cfg->vid_valid && nla_put_u32(rsp, PON_A_GEM_MAP_VID, cfg->vid))
+		goto err_cancel_msg;
+	if (cfg->pbit_valid && nla_put_u32(rsp, PON_A_GEM_MAP_PBIT, cfg->pbit))
+		goto err_cancel_msg;
+	if (cfg->dscp_valid && nla_put_u32(rsp, PON_A_GEM_MAP_DSCP, cfg->dscp))
+		goto err_cancel_msg;
+
+	genlmsg_end(rsp, hdr);
+	return 0;
+
+err_cancel_msg:
+	genlmsg_cancel(rsp, hdr);
+	return -EMSGSIZE;
+}
+
+/**
+ * pon_nl_gem_map_fill_pos() - pon_nl_obj_fill_t for the classifier rules
+ * @pdev: the PON device
+ * @pos: the list entry of the rule
+ * @rsp: the skb to fill
+ * @info: the dump info
+ *
+ * Return: the result of pon_nl_gem_map_fill().
+ */
+static int pon_nl_gem_map_fill_pos(struct pon_dev *pdev,
+				   struct list_head *pos, struct sk_buff *rsp,
+				   const struct genl_info *info)
+{
+	return pon_nl_gem_map_fill(pdev,
+				   list_entry(pos, struct pon_gem_map, list),
+				   rsp, info);
+}
+
+/**
+ * pon_nl_notify_gem_map() - send a classifier rule notification
+ * @pdev: the PON device
+ * @map: the rule
+ * @cmd: the notification command, PON_CMD_GEM_MAP_ADD_NTF or
+ *       PON_CMD_GEM_MAP_DEL_NTF
+ *
+ * Context: Called with @pdev->lock held. May sleep.
+ */
+void pon_nl_notify_gem_map(struct pon_dev *pdev, struct pon_gem_map *map,
+			   u32 cmd)
+{
+	pon_nl_notify_obj(pdev, &map->list, cmd, pon_nl_gem_map_fill_pos);
+}
+
+/**
+ * pon_nl_gem_map_get_dumpit() - handle the PON_CMD_GEM_MAP_GET dump
+ * @rsp: the skb to fill
+ * @cb: the dump state
+ *
+ * Return: the result of pon_nl_obj_dumpit().
+ */
+int pon_nl_gem_map_get_dumpit(struct sk_buff *rsp, struct netlink_callback *cb)
+{
+	return pon_nl_obj_dumpit(rsp, cb, offsetof(struct pon_dev, gem_maps),
+				 pon_nl_gem_map_fill_pos);
+}
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 18+ messages in thread

* [RFC net-next 09/12] net: pon: add the device registration
  2026-10-08 14:32 [RFC net-next 00/12] net: add the PON subsystem John Crispin
                   ` (7 preceding siblings ...)
  2026-10-08 14:32 ` [RFC net-next 08/12] net: pon: add netlink support John Crispin
@ 2026-10-08 14:32 ` John Crispin
  2026-10-08 14:32 ` [RFC net-next 10/12] net: pon: build the PON subsystem John Crispin
                   ` (4 subsequent siblings)
  13 siblings, 0 replies; 18+ messages in thread
From: John Crispin @ 2026-10-08 14:32 UTC (permalink / raw)
  To: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, linux-kernel, netdev, Andrew Lunn,
	Christian Marangi

pon_dev_create() ties the components to one instance. It takes the data
network device, the device and the ops of the MAC, pairs a waiting
conduit and publishes the instance. The data network device must be
bound to its network namespace, because the core looks a device up by
that namespace.

pon_dev_unregister() stops the upstream link and withdraws the instance:
its conduit, its work, its GEM network devices and its objects. It
disables the work item, so that a late queue runs nothing. It prints
the lines that pon_dev_log() recorded until then, then disables the log
work. The structure stays until the driver drops it with pon_dev_put(),
so the interrupts and timers of a driver never name a freed device.

The module registers the netlink family, the gem link kind, the notifier
that releases the OMCI channels and the notifier that follows the
conduits. A module alias loads it for the family. DOC: PON locking
documents the lock order.

The workqueue of an instance has high priority because the MAC driver
runs the handlers of its activation timers on it. drv_priv is cleared
only after synchronize_net(): a GEM transmit that read the ops under RCU
can still reach the driver, which finds itself through drv_priv.

Assisted-by: LLM
Signed-off-by: John Crispin <john@phrozen.org>
---
 net/pon/pon_main.c | 454 +++++++++++++++++++++++++++++++++++++++++++++
 1 file changed, 454 insertions(+)
 create mode 100644 net/pon/pon_main.c

diff --git a/net/pon/pon_main.c b/net/pon/pon_main.c
new file mode 100644
index 000000000000..977b2635a8ea
--- /dev/null
+++ b/net/pon/pon_main.c
@@ -0,0 +1,454 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/* Copyright (C) 2026 John Crispin <john@phrozen.org> */
+
+#include <linux/list.h>
+#include <linux/module.h>
+#include <linux/netdevice.h>
+#include <linux/slab.h>
+#include <linux/spinlock.h>
+#include <linux/string.h>
+#include <linux/workqueue.h>
+#include <linux/xarray.h>
+#include <net/pon.h>
+
+#include "pon.h"
+#include "pon-nl-gen.h"
+
+DEFINE_XARRAY_ALLOC1(pon_devs);
+/* guards the pon_devs xarray, taken before any instance lock */
+DEFINE_MUTEX(pon_devs_lock);
+
+/**
+ * DOC: PON locking
+ *
+ * pon_devs_lock protects the pon_devs xarray, the conduit list and the
+ * pairing of a conduit with an instance.
+ * Ordering is take the pon_devs_lock and then the instance lock.
+ * rtnl is taken before them and never while an instance lock or pon_devs_lock
+ * is held: the GEM netdevs are registered and unregistered outside them and
+ * the conduit's address and MTU are set after a pairing once both are
+ * dropped.
+ * Each instance is protected by RCU and has a refcount.
+ * When the driver unregisters, the instance gets flushed, but the struct
+ * sticks around.
+ *
+ * The instance lock is also the activation lock. Every deferred report
+ * reaches the state machine through pdev->wq, whose handler takes it. The
+ * netlink handlers take the same lock in their pre_doit, so a userspace
+ * transaction cannot interleave with a half run transition. pdev->work_lock
+ * is the one lock below it, held only to add and remove work items from
+ * hard interrupt context.
+ */
+
+/**
+ * pon_tcont_find() - look up a T-CONT by its index
+ * @pdev:	PON device structure
+ * @index:	the T-CONT index
+ *
+ * Context: Called with @pdev->lock held.
+ * Return: the T-CONT, or NULL when no T-CONT has @index.
+ */
+struct pon_tcont *pon_tcont_find(struct pon_dev *pdev, u16 index)
+{
+	struct pon_tcont *tcont;
+
+	lockdep_assert_held(&pdev->lock);
+	list_for_each_entry(tcont, &pdev->tconts, list)
+		if (tcont->cfg.index == index)
+			return tcont;
+	return NULL;
+}
+
+/**
+ * pon_gem_find() - look up a GEM port by its id
+ * @pdev:	PON device structure
+ * @gem_id:	the GEM port id
+ *
+ * Context: Called with @pdev->lock held.
+ * Return: the GEM port, or NULL when no GEM port has @gem_id.
+ */
+struct pon_gem *pon_gem_find(struct pon_dev *pdev, u16 gem_id)
+{
+	struct pon_gem *gem;
+
+	lockdep_assert_held(&pdev->lock);
+	list_for_each_entry(gem, &pdev->gems, list)
+		if (gem->cfg.id == gem_id)
+			return gem;
+	return NULL;
+}
+
+/**
+ * pon_gem_channel() - the conduit transmit channel of a GEM port
+ * @pdev:	PON device structure
+ * @gem:	the GEM port
+ *
+ * Asks the driver's tcont_channel callback for the channel of the T-CONT
+ * that @gem rides.
+ *
+ * Context: Called with @pdev->lock held.
+ * Return: the channel, -ENOLINK when @gem rides no T-CONT, the T-CONT does
+ * not exist, the driver has no tcont_channel callback or the T-CONT has no
+ * channel yet, or another negative errno from the driver.
+ */
+int pon_gem_channel(struct pon_dev *pdev, const struct pon_gem *gem)
+{
+	struct pon_tcont *tcont;
+
+	if (!gem->cfg.tcont_valid || !pdev->ops->tcont_channel)
+		return -ENOLINK;
+
+	tcont = pon_tcont_find(pdev, gem->cfg.tcont_index);
+	if (!tcont)
+		return -ENOLINK;
+
+	return pdev->ops->tcont_channel(pdev, &tcont->cfg);
+}
+
+/**
+ * pon_gem_map_same() - test whether two classifier rules are the same rule
+ * @existing:	the rule the instance holds
+ * @requested:	the rule a request names
+ *
+ * A rule is identified by everything it matches on, because two rules for one
+ * GEM that differ only in their VLAN are different rules.
+ *
+ * Return: true when @existing and @requested match on the same fields for the
+ * same GEM port.
+ */
+static bool pon_gem_map_same(const struct pon_gem_map_cfg *existing,
+			     const struct pon_gem_map_cfg *requested)
+{
+	return existing->gem_id == requested->gem_id &&
+	       existing->tag_valid == requested->tag_valid &&
+	       existing->tagged == requested->tagged &&
+	       existing->vid_valid == requested->vid_valid &&
+	       existing->vid == requested->vid &&
+	       existing->pbit_valid == requested->pbit_valid &&
+	       existing->pbit == requested->pbit &&
+	       existing->dscp_valid == requested->dscp_valid &&
+	       existing->dscp == requested->dscp;
+}
+
+/**
+ * pon_gem_map_find() - look up an upstream classifier rule
+ * @pdev:	PON device structure
+ * @cfg:	the rule to look for, compared with pon_gem_map_same()
+ *
+ * Context: Called with @pdev->lock held.
+ * Return: the stored rule, or NULL when the instance holds no such rule.
+ */
+struct pon_gem_map *pon_gem_map_find(struct pon_dev *pdev,
+				     const struct pon_gem_map_cfg *cfg)
+{
+	struct pon_gem_map *map;
+
+	lockdep_assert_held(&pdev->lock);
+	list_for_each_entry(map, &pdev->gem_maps, list)
+		if (pon_gem_map_same(&map->cfg, cfg))
+			return map;
+	return NULL;
+}
+
+/**
+ * pon_dev_create() - create and register a PON device
+ * @netdev:	the PON data network device, already registered, with
+ *		per-CPU transmit and receive statistics
+ *		(NETDEV_PCPU_STAT_TSTATS) for the conduit's receive path to
+ *		count into and bound to its network namespace
+ *		(netns_immutable) since before it was registered
+ * @parent:	the MAC's device, which the interfaces the core creates
+ *		parent onto and whose firmware node a conduit names to be
+ *		paired with this instance
+ * @ops:	driver callbacks
+ * @caps:	device capabilities
+ * @mode:	the mode the MAC is configured for, enum pon_mode
+ * @priv_ptr:	back-pointer to driver private data
+ *
+ * Syncs the conduit when one is already registered, so it takes rtnl and must
+ * not be called with rtnl held. The activation state starts unknown. The driver
+ * reports it through pon_dev_state_report() from its own work.
+ *
+ * Return: pointer to the allocated PON device, or ERR_PTR.
+ */
+struct pon_dev *pon_dev_create(struct net_device *netdev,
+			       struct device *parent,
+			       const struct pon_dev_ops *ops,
+			       const struct pon_dev_caps *caps,
+			       enum pon_mode mode, void *priv_ptr)
+{
+	static u32 last_id;
+	struct pon_dev *pdev;
+	bool paired;
+	int err;
+
+	if (WARN_ON(!netdev || !parent || !ops || !caps ||
+		    netdev->pcpu_stat_type != NETDEV_PCPU_STAT_TSTATS ||
+		    !netdev->netns_immutable ||
+		    !ops->tcont_set ||
+		    !ops->tcont_clear ||
+		    !ops->gem_add ||
+		    !ops->gem_del ||
+		    !ops->omci_xmit))
+		return ERR_PTR(-EINVAL);
+
+	pdev = kzalloc_obj(*pdev, GFP_KERNEL);
+	if (!pdev)
+		return ERR_PTR(-ENOMEM);
+
+	pdev->main_netdev = netdev;
+	pdev->parent = parent;
+	pdev->ops = ops;
+	pdev->caps = caps;
+	pdev->mode = mode;
+	pdev->drv_priv = priv_ptr;
+	pdev->ploam = PON_PLOAM_STATE_UNKNOWN;
+
+	mutex_init(&pdev->lock);
+	spin_lock_init(&pdev->work_lock);
+	INIT_LIST_HEAD(&pdev->tconts);
+	INIT_LIST_HEAD(&pdev->gems);
+	INIT_LIST_HEAD(&pdev->gem_maps);
+	xa_init(&pdev->gem_netdevs);
+	INIT_LIST_HEAD(&pdev->work_list);
+	INIT_WORK(&pdev->work, pon_work_worker);
+	pon_log_init(pdev);
+	refcount_set(&pdev->refcnt, 1);
+
+	/* Ordered, so that the work items run one at a time and in the order
+	 * they were queued. High priority, because the activation and key
+	 * exchange timers whose handlers a driver runs on it are short.
+	 */
+	pdev->wq = alloc_ordered_workqueue("pon-%s", WQ_HIGHPRI, netdev->name);
+	if (!pdev->wq) {
+		err = -ENOMEM;
+		goto err_free;
+	}
+
+	pon_omci_init(pdev);
+
+	mutex_lock(&pon_devs_lock);
+	err = xa_alloc_cyclic(&pon_devs, &pdev->id, pdev, xa_limit_16b,
+			      &last_id, GFP_KERNEL);
+	if (err) {
+		mutex_unlock(&pon_devs_lock);
+		goto err_wq;
+	}
+	paired = pon_conduit_attach(pdev);
+	mutex_lock(&pdev->lock);
+	mutex_unlock(&pon_devs_lock);
+
+	pon_nl_notify_dev(pdev, PON_CMD_DEV_ADD_NTF);
+
+	rcu_assign_pointer(netdev->pon_dev, pdev);
+
+	mutex_unlock(&pdev->lock);
+
+	if (paired)
+		pon_conduit_sync(pdev);
+
+	return pdev;
+
+err_wq:
+	destroy_workqueue(pdev->wq);
+err_free:
+	mutex_destroy(&pdev->lock);
+	kfree(pdev);
+	return ERR_PTR(err);
+}
+EXPORT_SYMBOL_GPL(pon_dev_create);
+
+/**
+ * pon_dev_free() - free a PON device once its last reference is gone
+ * @pdev:	PON device structure
+ *
+ * Releases the device's id, destroys its workqueue and frees the structure
+ * after an RCU grace period.
+ *
+ * Context: Process context. Takes pon_devs_lock and may sleep.
+ */
+static void pon_dev_free(struct pon_dev *pdev)
+{
+	mutex_lock(&pon_devs_lock);
+	xa_erase(&pon_devs, pdev->id);
+	mutex_unlock(&pon_devs_lock);
+
+	destroy_workqueue(pdev->wq);
+	xa_destroy(&pdev->gem_netdevs);
+	mutex_destroy(&pdev->lock);
+	kfree_rcu(pdev, rcu);
+}
+
+/**
+ * pon_dev_put() - drop a reference to a PON device
+ * @pdev:	PON device structure
+ *
+ * A driver drops the reference pon_dev_create() returned, once, after
+ * pon_dev_unregister() and after it has released everything of its own that
+ * names the device: its interrupts, its timers and the data network device.
+ */
+void pon_dev_put(struct pon_dev *pdev)
+{
+	if (refcount_dec_and_test(&pdev->refcnt))
+		pon_dev_free(pdev);
+}
+EXPORT_SYMBOL_GPL(pon_dev_put);
+
+/**
+ * pon_dev_unregister() - unregister a PON device
+ * @pdev:	PON device structure
+ *
+ * Stops the upstream link through the driver's enable callback when it is
+ * enabled, then withdraws the device. Once this returns no netlink request,
+ * no work item and no network device of the core reaches the driver. The
+ * lines that pon_dev_log() recorded until then are printed before it
+ * returns.
+ *
+ * The device itself stays until the driver calls pon_dev_put(). What a
+ * driver still reports to it in between, from an interrupt or from the data
+ * network device, is refused.
+ *
+ * Unregisters network devices, so it takes rtnl and must not be called with
+ * rtnl held.
+ */
+void pon_dev_unregister(struct pon_dev *pdev)
+{
+	struct pon_gem_map *map, *map_next;
+	struct pon_tcont *tcont, *tcont_next;
+	struct pon_gem *gem, *gem_next;
+
+	mutex_lock(&pon_devs_lock);
+	mutex_lock(&pdev->lock);
+
+	/* Wait until pon_dev_free() to call xa_erase() so this id cannot be
+	 * reused while references are still held. Storing NULL makes every
+	 * netlink lookup fail from here on.
+	 */
+	xa_store(&pon_devs, pdev->id, NULL, GFP_KERNEL);
+	pon_nl_obj_gen_inc();
+	mutex_unlock(&pon_devs_lock);
+
+	if (pdev->enabled) {
+		pdev->ops->enable(pdev, false, NULL);
+		pdev->enabled = false;
+	}
+
+	pon_nl_notify_dev(pdev, PON_CMD_DEV_DEL_NTF);
+
+	WRITE_ONCE(pdev->going_away, true);
+	mutex_unlock(&pdev->lock);
+
+	/* First, so that no frame reaches the instance once its objects go. */
+	pon_conduit_detach(pdev);
+
+	/* Drop what is queued before stopping the worker, or an item requeued
+	 * by the one still running would outlive the cancel.
+	 */
+	pon_work_drain(pdev);
+
+	/* Each takes the instance lock, so none is waited for with it held. */
+	disable_work_sync(&pdev->work);
+	flush_work(&pdev->log_work);
+	disable_work_sync(&pdev->log_work);
+
+	pon_gem_netdevs_unregister(pdev);
+
+	pon_omci_destroy(pdev);
+
+	mutex_lock(&pdev->lock);
+
+	list_for_each_entry_safe(tcont, tcont_next, &pdev->tconts, list) {
+		list_del(&tcont->list);
+		kfree(tcont);
+	}
+	list_for_each_entry_safe(gem, gem_next, &pdev->gems, list) {
+		list_del(&gem->list);
+		kfree(gem);
+	}
+	list_for_each_entry_safe(map, map_next, &pdev->gem_maps, list) {
+		list_del(&map->list);
+		kfree(map);
+	}
+
+	rcu_assign_pointer(pdev->main_netdev->pon_dev, NULL);
+
+	WRITE_ONCE(pdev->ops, NULL);
+	memzero_explicit(&pdev->identity, sizeof(pdev->identity));
+
+	mutex_unlock(&pdev->lock);
+
+	/* The GEM transmit path reads @ops under RCU. */
+	synchronize_net();
+
+	/*
+	 * Only now: a transmit that loaded @ops before the store above is
+	 * still inside its read side section and reaches the driver, which
+	 * reads @drv_priv to find itself.
+	 */
+	pdev->drv_priv = NULL;
+}
+EXPORT_SYMBOL_GPL(pon_dev_unregister);
+
+/**
+ * pon_init() - register the PON core
+ *
+ * Registers the generic netlink family, the GEM link type, the netlink
+ * notifier that releases the OMCI channels and the netdev notifier that
+ * follows the conduits.
+ *
+ * Runs before the device drivers when built in, because an ethernet driver
+ * registers its conduit from its own probe. A module carries the same
+ * ordering through its dependencies.
+ *
+ * Return: 0, or a negative errno when a registration fails.
+ */
+static int __init pon_init(void)
+{
+	int err;
+
+	err = genl_register_family(&pon_nl_family);
+	if (err)
+		return err;
+
+	err = pon_gem_link_register();
+	if (err)
+		goto err_family;
+
+	err = pon_omci_notifier_register();
+	if (err)
+		goto err_link;
+
+	err = pon_conduit_notifier_register();
+	if (err)
+		goto err_omci;
+
+	return 0;
+
+err_omci:
+	pon_omci_notifier_unregister();
+err_link:
+	pon_gem_link_unregister();
+err_family:
+	genl_unregister_family(&pon_nl_family);
+	return err;
+}
+subsys_initcall(pon_init);
+
+/**
+ * pon_exit() - unregister the PON core
+ *
+ * Undoes pon_init() in the reverse order.
+ */
+static void __exit pon_exit(void)
+{
+	pon_conduit_notifier_unregister();
+	pon_omci_notifier_unregister();
+	pon_gem_link_unregister();
+	genl_unregister_family(&pon_nl_family);
+}
+module_exit(pon_exit);
+
+MODULE_ALIAS_GENL_FAMILY(PON_FAMILY_NAME);
+MODULE_AUTHOR("John Crispin <john@phrozen.org>");
+MODULE_DESCRIPTION("Passive optical network ONU support");
+MODULE_LICENSE("GPL");
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 18+ messages in thread

* [RFC net-next 10/12] net: pon: build the PON subsystem
  2026-10-08 14:32 [RFC net-next 00/12] net: add the PON subsystem John Crispin
                   ` (8 preceding siblings ...)
  2026-10-08 14:32 ` [RFC net-next 09/12] net: pon: add the device registration John Crispin
@ 2026-10-08 14:32 ` John Crispin
  2026-10-08 14:32 ` [RFC net-next 11/12] Documentation: networking: describe " John Crispin
                   ` (3 subsequent siblings)
  13 siblings, 0 replies; 18+ messages in thread
From: John Crispin @ 2026-10-08 14:32 UTC (permalink / raw)
  To: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni
  Cc: Simon Horman, linux-kernel, netdev, Andrew Lunn,
	Christian Marangi

Add CONFIG_PON (a tristate) and the Makefile of the core. Hook net/pon
into net/Kconfig and net/Makefile.

The build is wired last, so every commit before it builds and the
series bisects.

Assisted-by: LLM
Signed-off-by: John Crispin <john@phrozen.org>
---
 net/Kconfig      |  1 +
 net/Makefile     |  1 +
 net/pon/Kconfig  | 17 +++++++++++++++++
 net/pon/Makefile |  6 ++++++
 4 files changed, 25 insertions(+)
 create mode 100644 net/pon/Kconfig
 create mode 100644 net/pon/Makefile

diff --git a/net/Kconfig b/net/Kconfig
index 76ab44aa439a..117845100a06 100644
--- a/net/Kconfig
+++ b/net/Kconfig
@@ -82,6 +82,7 @@ menu "Networking options"
 
 source "net/packet/Kconfig"
 source "net/psp/Kconfig"
+source "net/pon/Kconfig"
 source "net/unix/Kconfig"
 source "net/tls/Kconfig"
 source "net/xfrm/Kconfig"
diff --git a/net/Makefile b/net/Makefile
index 5b2dd7f07a85..757650c5428d 100644
--- a/net/Makefile
+++ b/net/Makefile
@@ -19,6 +19,7 @@ obj-$(CONFIG_TLS)		+= tls/
 obj-$(CONFIG_XFRM)		+= xfrm/
 obj-$(CONFIG_UNIX)		+= unix/
 obj-$(CONFIG_INET_PSP)		+= psp/
+obj-$(CONFIG_PON)		+= pon/
 obj-y				+= ipv6/
 obj-$(CONFIG_PACKET)		+= packet/
 obj-$(CONFIG_NET_KEY)		+= key/
diff --git a/net/pon/Kconfig b/net/pon/Kconfig
new file mode 100644
index 000000000000..af40d4ebd938
--- /dev/null
+++ b/net/pon/Kconfig
@@ -0,0 +1,17 @@
+# SPDX-License-Identifier: GPL-2.0-only
+#
+# PON subsystem configuration
+#
+config PON
+	tristate "Passive Optical Network (PON) device support"
+	depends on NET
+	help
+	  Enable kernel support for PON (GPON / XG(S)-PON) ONU devices: the
+	  netlink configuration family (which also carries the OMCI
+	  management channel to a userspace OMCI stack) and per GEM port
+	  virtual devices.
+
+	  To compile this as a module, choose M here: the module will be
+	  called pon.
+
+	  If unsure, say N.
diff --git a/net/pon/Makefile b/net/pon/Makefile
new file mode 100644
index 000000000000..4b23d2265be9
--- /dev/null
+++ b/net/pon/Makefile
@@ -0,0 +1,6 @@
+# SPDX-License-Identifier: GPL-2.0-only
+
+obj-$(CONFIG_PON) += pon.o
+
+pon-y := pon_main.o pon_nl.o pon_omci.o pon_gem.o pon_conduit.o \
+	 pon_work.o pon_state.o pon_ploam_msg.o pon-nl-gen.o
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 18+ messages in thread

* [RFC net-next 11/12] Documentation: networking: describe the PON subsystem
  2026-10-08 14:32 [RFC net-next 00/12] net: add the PON subsystem John Crispin
                   ` (9 preceding siblings ...)
  2026-10-08 14:32 ` [RFC net-next 10/12] net: pon: build the PON subsystem John Crispin
@ 2026-10-08 14:32 ` John Crispin
  2026-10-08 14:32 ` [RFC net-next 12/12] MAINTAINERS: add " John Crispin
                   ` (2 subsequent siblings)
  13 siblings, 0 replies; 18+ messages in thread
From: John Crispin @ 2026-10-08 14:32 UTC (permalink / raw)
  To: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni
  Cc: Jonathan Corbet, Simon Horman, Shuah Khan, Randy Dunlap, netdev,
	linux-doc, linux-kernel, Andrew Lunn, Christian Marangi

pon.rst describes the model, activation and why its state machine runs
in the driver, the activation log, the objects across a loss of the
link, the lent context, the OMCI channel, the conduit and the netlink
interface. It says which modes the uapi reserves and that the core
refuses a mode the device does not list.

Assisted-by: LLM
Signed-off-by: John Crispin <john@phrozen.org>
---
 Documentation/networking/index.rst |   1 +
 Documentation/networking/pon.rst   | 302 +++++++++++++++++++++++++++++
 2 files changed, 303 insertions(+)
 create mode 100644 Documentation/networking/pon.rst

diff --git a/Documentation/networking/index.rst b/Documentation/networking/index.rst
index 44a422ad3b05..05eaac7212f3 100644
--- a/Documentation/networking/index.rst
+++ b/Documentation/networking/index.rst
@@ -96,6 +96,7 @@ Contents:
    phy-port
    pktgen
    plip
+   pon
    ppp_generic
    proc_net_tcp
    pse-pd/index
diff --git a/Documentation/networking/pon.rst b/Documentation/networking/pon.rst
new file mode 100644
index 000000000000..732e6449d522
--- /dev/null
+++ b/Documentation/networking/pon.rst
@@ -0,0 +1,302 @@
+.. SPDX-License-Identifier: GPL-2.0
+
+===========================
+Passive optical network ONU
+===========================
+
+Overview
+========
+
+``net/pon`` is the kernel side of an optical network unit, the subscriber end
+of a passive optical network. One optical line terminal in the exchange serves
+many ONUs over a shared fiber. Downstream is a broadcast the ONU filters.
+Upstream is time division multiple access and an ONU may only transmit inside
+the windows the OLT grants it.
+
+The subsystem implements ITU-T G.9807.1 (XGS-PON): the PLOAM codec and the
+identifier ranges are those of G.9807.1. The uapi also
+reserves the modes of G.984 (GPON) and G.987 (XG-PON) for later drivers, as
+well as the activation state o3 of G.984.3. The core implements none of those
+systems. A device lists the modes it supports in its capabilities. The core
+refuses any other mode. The subsystem does not cover EPON, which is IEEE
+802.3ah and a different access method.
+
+The model
+=========
+
+Three objects carry the datapath and they nest:
+
+Device
+  One PON MAC. It has an identity the OLT authenticates, a mode and an
+  activation state.
+
+T-CONT
+  A transmission container: the entity the OLT grants upstream windows to,
+  named by an alloc-id the OLT assigns over PLOAM. Upstream bandwidth is
+  arbitrated between T-CONTs by the OLT and within one T-CONT by the ONU.
+
+GEM port
+  A flow, named by a GEM port id the OLT assigns over OMCI. Several GEM ports
+  ride one T-CONT. A GEM port id is a label in the frame header, not a
+  scheduling entity. G.9807.1 Table C.6.6 reserves 0 to 1020 for the default
+  GEM port (equal to the ONU-ID, OMCI only) and 65535 for the idle GEM port,
+  so the uapi takes only 1021 to 65534.
+
+One more object supports them:
+
+Classifier rule
+  What upstream traffic maps onto which GEM port, matched on VLAN tag state,
+  VLAN id, priority or DSCP.
+
+Activation
+==========
+
+An ONU reaches the operational state O5 through the PLOAM state machine of
+G.984.3 or G.9807.1: serial number exchange, ONU-ID assignment, ranging and
+finally O5, which is the only state in which the OMCI management channel
+carries traffic.
+
+**The state machine runs in the driver, not in the core.** That is a
+deliberate choice and the reason this subsystem has no software MAC layer.
+On the hardware it was written against, the key hierarchy of G.9807.1 clause
+C.15.3 runs inside the MAC's key generator and the OMCI integrity key never
+leaves it, so no generic layer can compute the OMCI message integrity check.
+A core that owned the state machine would be a core that could not finish the
+job.
+
+What the core owns instead is the vocabulary and the object model:
+
+* the ITU-T constants and message layouts, in ``include/net/pon/ploam.h``
+* the PLOAM message codec, which builds and parses standard message bodies and
+  holds no state, no timer and no policy
+* the objects above, which it stores
+* the activation state, which the driver reports and the core validates and
+  publishes
+* the network devices, their carrier and the OMCI channel
+* the netlink uapi
+
+The driver owns the state machine, its timers, the key hierarchy and every
+register.
+
+Why the core validates the state
+--------------------------------
+
+``ploam-ntf`` and the activation state in the netlink reply are uapi. If each
+driver published its own state, each vendor would invent its own notification
+timing and an OMCI daemon would behave differently per MAC. So the driver
+reports through ``pon_dev_state_report()`` and the core checks the transition
+against the state machine of the standard, publishes it and updates the
+carrier. An out of range value is refused. An illegal edge is published
+anyway and warned about: the hardware is the truth and a driver that reports
+something the standard does not permit is a bug to find, not a state to hide.
+
+Everything the uapi carries follows the ITU-T recommendations. Where a MAC
+deviates from them, its driver translates the hardware to the standard before
+it reports, as a quirk of that driver. The core and the uapi never carry a
+vendor state or a vendor meaning. The states are the union of G.984.3 and
+G.9807.1: G.9807.1 has one Serial Number state, O2-3 (Table C.12.1), which a
+driver reports as ``o2``. Only a G.984.3 device reports ``o3``.
+
+The core logs every edge it publishes at info level, one line per edge, for
+example ``pon0: PLOAM state O5 -> O6``. The log names the Serial Number state
+``O2-3`` in every mode but G-PON. A repeated report logs nothing. A work
+item outside the instance's context prints the line, so a slow console never
+delays the PLOAM exchange. A driver logs its own lines the same way through
+``pon_dev_log()``. All lines keep their order.
+
+Objects outlive the link
+========================
+
+A GEM port is OMCI configuration and the ONU keeps it when the link goes and
+a new activation starts: G.9807.1 clause C.6.1.5.8 has the ONU retain every
+XGEM port id the OMCI assigned when it enters O1 and discard only the default
+one that carries the OMCC. So the driver keeps its GEM ports across a loss of
+the link and the core keeps the objects, exactly as an address on an ethernet
+device outlives a cable pull. Nothing is handed back when the link returns,
+because nothing was taken away.
+
+What the ONU does discard are the alloc-ids (clause C.6.1.5.7). The OLT assigns
+them again over PLOAM once the ONU is back in O5 and the driver binds each one
+to a transmit channel and reports that to the core through
+``pon_dev_event()``. The default alloc-id, equal to the ONU-ID, is not assigned
+by a message. The driver binds it with the ONU-ID and it may carry user traffic
+as well as the OMCC. A GEM port whose
+alloc-id has no channel yet is held and carries no traffic. The carrier of the
+PON interfaces is up in O5 and O6 while at least one GEM port rides an alloc-id
+that has a channel and a conduit (see below) is paired with the device and up.
+The core follows the conduit with a netdev notifier: the carrier falls when the
+conduit starts to go down and when its driver takes it back.
+
+Classifier rules are objects of their own. In G.988 they come from their own
+managed entities, which reference a GEM port rather than belonging to it, so a
+rule outlives the GEM port it names, in the core and in the driver.
+
+The lent context
+================
+
+A PON MAC runs the activation state machine itself, but it must not run it in
+hard interrupt context and it must not run it against a half finished netlink
+transaction. The core therefore lends the driver its own serialized context:
+one ordered workqueue per instance and one lock that the netlink handlers
+take in their ``pre_doit``.
+
+A driver queues a ``pon_work`` from any context, hard interrupt included. The
+handler runs with the instance lock held, so it may sleep. One worker drains
+the list one item at a time, so the ordering a driver sees is the order it
+queued in.
+
+The OMCI channel
+================
+
+OMCI is the management protocol of G.988 and it is userspace's job. The kernel
+carries it over the netlink family, the way nl80211 carries management frames
+to hostapd. A daemon sends ``omci-register`` for a device and from then on
+receives every OMCI PDU of that device as an ``omci-ntf`` message, sent to
+its socket alone. It sends a PDU with ``omci-tx``. One socket owns the OMCI
+channel of a device at a time: another socket is refused with ``EBUSY`` and
+only the owner may send. The registration ends when the socket closes, so a
+daemon that restarts registers again and nothing is left behind.
+
+A PDU crosses the boundary bare, from the transaction correlation id to the
+end of the message contents, in both directions and in both the baseline and
+the extended format. It carries no integrity field, because the MAC computes
+the upstream one and checks the downstream one. A PDU is therefore at most
+1976 bytes: G.988 clause 11.2.5 limits an extended message to 1980 bytes
+and 4 of them are the integrity field.
+
+The downstream check works like a receive checksum offload. The MAC that
+checks the integrity field says so in its receive descriptor, per PDU and
+the core takes such a PDU as verified and stripped. A PDU the MAC passed up
+unchecked still ends with its integrity field, so the core accepts it up to
+1980 bytes and hands it to the driver's ``omci_verify``, which checks and
+strips the field. That check cannot be a generic helper of the core: the
+field is an AES-CMAC with the OMCI integrity key (G.9807.1 clause C.15.7.2) and
+the key belongs to the key hierarchy the driver runs. On some MACs it never
+leaves the key generator.
+
+No network device and no
+EtherType is involved: the OMCC carries the PDU on its own GEM port with no
+Ethernet header and the core knows that port, so there is nothing to
+classify.
+
+A received PDU waits in the instance's context until it reaches the owner. A
+PDU that finds no owner, a full queue or a full socket is dropped. The OLT
+retries. The exchange can be captured on an ``nlmon`` device like any other
+netlink traffic.
+
+No OMCI managed entity is modeled in the kernel. The daemon reads the OLT's
+intent from the MIB and expresses it through the netlink family below.
+
+The conduit
+===========
+
+A PON MAC on a system on chip has no DMA of its own: its frames ride the rings
+of an ethernet port of the same chip. That port is the conduit, in the sense
+DSA gives the word and its driver registers it with ``pon_conduit_register()``
+when the port's firmware node names the MAC through a ``pon-handle``
+reference. The core pairs the two by that node, in whichever order they
+probe and neither waits for the other.
+
+Per frame information crosses the boundary as explicit arguments and never in
+``skb->cb`` or in a frame tag. On transmit the MAC driver names the GEM port,
+the T-CONT channel, the queue and, for an OMCI PDU, the integrity key index
+and the conduit's driver builds its descriptor from them. On receive the
+conduit's driver reads the GEM port and the OMCI and integrity flags from its
+descriptor and hands the frame to ``pon_conduit_rx()``, which delivers it to
+the OMCI channel, to the GEM port's own network device when it has one, or to
+the PON data interface. The MAC driver verifies an OMCI PDU the hardware
+passed up unchecked through ``omci_verify``. Without that callback such a PDU
+is dropped. Such a PDU still carries its integrity field, so the core accepts
+it at up to 1980 bytes and the conduit's MTU never drops below that.
+
+The conduit keeps the address of the PON data interface, because the frame
+engine behind it recognizes the ONU's frames by the conduit's own address.
+The core sets it when the two are paired. When the data interface's address
+changes afterwards, the MAC driver calls ``pon_conduit_addr_set()`` from its
+``ndo_set_mac_address``.
+
+The conduit also takes the largest MTU among the PON network devices, because
+the frames of all of them ride its rings. The core sets it when the two are
+paired and when a GEM port's network device is created or changes its MTU. The
+MAC driver calls ``pon_conduit_mtu_set()`` from the ``ndo_change_mtu`` of the
+data interface. A change that the conduit's driver refuses is refused.
+
+Netlink interface
+=================
+
+The generic netlink family ``pon`` is specified in
+``Documentation/netlink/specs/pon.yaml`` and the uapi header is generated from
+it. Every object can be listed, watched and read back:
+
+=================  ==============================================
+``dev-get``        devices, with a dump
+``tcont-get``      T-CONTs, with a dump
+``gem-get``        GEM ports, with a dump
+``gem-map-get``    classifier rules, a dump
+=================  ==============================================
+
+Notifications share the reply format of the matching get, so a listener parses
+one message shape whether it asked for an object or was told about it. There
+is no sequence number: a listener that overruns its socket re-reads.
+``dev-del-ntf`` stands for the removal of every object of that device. The
+objects go without a notification of their own.
+
+The dumps of the T-CONTs, the GEM ports and the classifier rules take an
+optional device ID and then list the objects of that device only. A
+device ID that names no device ends the dump with ``-ENODEV``. A dump that
+does not fit one message resumes by count, so the kernel sets
+``NLM_F_DUMP_INTR`` when an object joined or left a list in between and the
+reader repeats the dump.
+
+``omci-register``, ``omci-tx`` and ``omci-ntf`` are the OMCI channel, described
+above. ``omci-ntf`` goes to the owner of the channel alone and to no
+multicast group.
+
+``ploam-ntf`` reports every activation transition.
+
+An operation is a ``set`` when it can change an object that already exists
+and a ``new`` when it can only create one. ``tcont-set`` rebinds a T-CONT,
+while ``gem-new`` and ``gem-map-new`` refuse an object that exists with other
+attributes.
+
+A GEM port reaches an alloc-id only through its T-CONT. ``tcont-set`` that
+changes the alloc-id moves the GEM ports of the T-CONT too. If the driver
+refuses one, the core restores the old alloc-id on the T-CONT and on its GEM
+ports and returns the error. ``tcont-del`` answers ``-EBUSY`` while a GEM port
+names the T-CONT, so userspace deletes the GEM ports first. ``tcont-set``
+answers ``-EEXIST`` for an alloc-id that another T-CONT holds: G.988 clause
+9.2.2 relates alloc-ids and T-CONTs one to one, so userspace that moves an
+alloc-id releases it first.
+
+``dev-set`` refuses to start the link until userspace gave the device a serial
+number. The driver carries no serial of its own, so an ONU that has not been
+provisioned stays silent rather than range under a vendor default.
+
+No key material crosses this interface in either direction. ``gem-new`` selects
+the key ring of each GEM port, as G.988 clause 9.2.3 defines it. The key
+ring alone decides whether a GEM port is encrypted and in which direction. The
+keys themselves are derived and rotated below. The broadcast key ring is
+refused, because the OLT distributes those keys over the OMCI and the kernel
+takes none.
+
+Internals
+=========
+
+The kernel-internal interfaces, pulled from the source.
+
+.. kernel-doc:: net/pon/pon_main.c
+   :doc: PON locking
+
+.. kernel-doc:: net/pon/pon_work.c
+   :doc: The lent context
+
+.. kernel-doc:: include/net/pon/ploam.h
+   :doc: The PLOAM vocabulary
+
+.. kernel-doc:: include/net/pon/ploam_msg.h
+   :doc: The PLOAM message codec
+
+.. kernel-doc:: include/net/pon/types.h
+   :internal:
+
+.. kernel-doc:: include/net/pon/functions.h
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 18+ messages in thread

* [RFC net-next 12/12] MAINTAINERS: add the PON subsystem
  2026-10-08 14:32 [RFC net-next 00/12] net: add the PON subsystem John Crispin
                   ` (10 preceding siblings ...)
  2026-10-08 14:32 ` [RFC net-next 11/12] Documentation: networking: describe " John Crispin
@ 2026-10-08 14:32 ` John Crispin
  2026-10-09  2:27 ` [RFC net-next 00/12] net: " Gaoyang Wei
  2026-10-09 11:04 ` Andre Ziviani
  13 siblings, 0 replies; 18+ messages in thread
From: John Crispin @ 2026-10-08 14:32 UTC (permalink / raw)
  To: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni
  Cc: linux-kernel, netdev, Andrew Lunn, Christian Marangi

Add an entry for the PON subsystem: the core under net/pon, its headers,
the uapi header, the netlink specification and the documentation.

Assisted-by: LLM
Signed-off-by: John Crispin <john@phrozen.org>
---
 MAINTAINERS | 11 +++++++++++
 1 file changed, 11 insertions(+)

diff --git a/MAINTAINERS b/MAINTAINERS
index 72ca3aab2106..0513cd25a039 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -20774,6 +20774,17 @@ F:	include/linux/parman.h
 F:	lib/parman.c
 F:	lib/test_parman.c
 
+PASSIVE OPTICAL NETWORK (PON) SUBSYSTEM
+M:	John Crispin <john@phrozen.org>
+L:	netdev@vger.kernel.org
+S:	Maintained
+F:	Documentation/netlink/specs/pon.yaml
+F:	Documentation/networking/pon.rst
+F:	include/net/pon.h
+F:	include/net/pon/
+F:	include/uapi/linux/pon.h
+F:	net/pon/
+
 PC ENGINES APU BOARD DRIVER
 M:	Enrico Weigelt, metux IT consult <info@metux.net>
 S:	Maintained
-- 
2.34.1


^ permalink raw reply related	[flat|nested] 18+ messages in thread

* Re: [RFC net-next 00/12] net: add the PON subsystem
  2026-10-08 14:32 [RFC net-next 00/12] net: add the PON subsystem John Crispin
                   ` (11 preceding siblings ...)
  2026-10-08 14:32 ` [RFC net-next 12/12] MAINTAINERS: add " John Crispin
@ 2026-10-09  2:27 ` Gaoyang Wei
  2026-10-09 11:04 ` Andre Ziviani
  13 siblings, 0 replies; 18+ messages in thread
From: Gaoyang Wei @ 2026-10-09  2:27 UTC (permalink / raw)
  To: John Crispin, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni
  Cc: Donald Hunter, Simon Horman, netdev, linux-kernel, Andrew Lunn,
	Christian Marangi, Jonathan Corbet, Shuah Khan, Randy Dunlap,
	linux-doc, Gaoyang Wei, Ziyou Xu

Hi John,

Your cover letter explicitly presents this series as following the
framework discussion I started. Yet neither Ziyou nor I was copied on
the cover letter or any of the twelve patches.

Given our participation in the original discussion and the cross-vendor
follow-up, both of us should have been copied. We find this omission
inappropriate.

I've added both of us to Cc. Please keep us copied on future revisions
and related framework discussions.

Gaoyang

^ permalink raw reply	[flat|nested] 18+ messages in thread

* Re: [RFC net-next 00/12] net: add the PON subsystem
  2026-10-08 14:32 [RFC net-next 00/12] net: add the PON subsystem John Crispin
                   ` (12 preceding siblings ...)
  2026-10-09  2:27 ` [RFC net-next 00/12] net: " Gaoyang Wei
@ 2026-10-09 11:04 ` Andre Ziviani
  2026-10-09 12:11   ` John Crispin
  13 siblings, 1 reply; 18+ messages in thread
From: Andre Ziviani @ 2026-10-09 11:04 UTC (permalink / raw)
  To: John Crispin
  Cc: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Donald Hunter, Simon Horman, netdev, linux-kernel, Andrew Lunn,
	Christian Marangi, Jonathan Corbet, Shuah Khan, Randy Dunlap,
	linux-doc, Gaoyang Wei, Ziyou Xu, Benjamin Larsson

Hi John,

On Thu, Oct 08, 2026 at 04:32:37PM +0200, John Crispin wrote:
> The scope is XGS-PON. The mode enum keeps gpon and xg-pon so that the
> uapi order stays clean for later drivers.

A G-PON data point, since the framework thread mentioned Realtek only
as work with proprietary components. odi-oss [1] is open firmware for a
GPON SFP ONU stick on the Realtek RTL9602C. It runs mainline 6.18 with
GPL drivers of our own for the GPON MAC, the PLOAM state machine of
G.984.3, the CPU NIC and the switch, and with a userspace OMCI daemon.
Apart from the stock bootloader, the image has no vendor code or
binaries. It runs in production on two ISPs, and a user has it working
on a third.

I should be upfront: I am not a PON or kernel expert. The drivers were
written with an AI coding assistant, working from the ITU-T
recommendations and from register traces of the stock firmware, and
checked by trial and error on my own sticks against live OLTs. So take
what follows as field data from one G-PON implementation, not as
review from someone who knows the standards well. Corrections are very
welcome.

Our split is already the one this series takes: PLOAM in the MAC driver,
OMCI in userspace over netlink (a private family for now). So our G-PON
driver should be able to sit on net/pon, as a second MAC and a first
G-PON one, at least out of tree (the RTL9602C platform itself is not
upstream). Reading the series with that in mind, four things would get
in the way:

1. The activation edges. pon_state_legal[] is one table for every mode,
   and the comment says a per-mode split can wait for a second MAC. Our
   state machine follows G.984.3 Table 10-1, and five of its edges are
   not in the table:

     O3 -> O2  TO1 expires in Serial_Number
     O5 -> O2  Deactivate_ONU-ID in Operation (O6 -> O2 too)
     O6 -> O4  broadcast POPUP
     O6 -> O7  Disable_Serial_Number in POPUP
     O7 -> O2  Disable_Serial_Number "enable"; G.9807.1 goes to O1

   Each would be published with a warning on every occurrence. A table
   per mode, selected by the device mode, would fix it.

2. GEM port ids. The spec takes 1021 to 65534, the XGEM Port-ID range. A
   G-PON Port-ID is 12 bits. The operator with six GEM ports mentioned
   above assigns 269 to 909, and ours on the other ISP 1434 to 1946.
   The range would need to depend on the mode too.

3. The datapath of an SFP ONU. Service traffic never reaches the CPU on
   this stick. The switch bridges the host SerDes and the PON in
   hardware, and our OMCI daemon programs it from the MIB: GEM flows,
   upstream queues and VLAN treatment from the Extended VLAN Tagging ME.
   Only OMCI goes through the CPU NIC: received frames carry a trap
   reason in the RX descriptor, and we send with per-frame descriptor
   words (port and GEM stream). So OMCI fits the conduit contract well,
   but there is no data netdev to pair and nothing to bridge in Linux.
   Can a driver own the T-CONT, GEM and gem-map objects and offload
   them without a data netdev or GEM netdevs? The VLAN rules we need
   also go beyond tag/vid/pbit/dscp (double-tag rules, TPID, the
   treatment side), but that may belong to switchdev rather than here.

4. Observing OMCI. omci-ntf goes to the one registered socket. Being able
   to watch the exchange without taking over the channel, as Benjamin and
   Ziyou asked in the earlier thread, is what we use most when an OLT
   behaves oddly. A read-only multicast copy of the PDUs in both
   directions would cover it.

One small note for the docs: baseline G-PON OMCI ends with a CRC-32
trailer, not a MIC. On the RTL9602C neither direction is checked or
added by the MAC, so our driver would compute it in software. That fits
the current pdu attribute; the docs could just name the G-PON trailer
alongside the MIC.

If a per-mode state table and Port-ID range are acceptable, I can port
our driver onto the series and report back. I can also share PLOAM
traces from both ISPs.

[1] https://github.com/AndreZiviani/odi-oss

Thanks,
Andre Ziviani

^ permalink raw reply	[flat|nested] 18+ messages in thread

* Re: [RFC net-next 00/12] net: add the PON subsystem
  2026-10-09 11:04 ` Andre Ziviani
@ 2026-10-09 12:11   ` John Crispin
  0 siblings, 0 replies; 18+ messages in thread
From: John Crispin @ 2026-10-09 12:11 UTC (permalink / raw)
  To: Andre Ziviani
  Cc: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
	Donald Hunter, Simon Horman, netdev, linux-kernel, Andrew Lunn,
	Christian Marangi, Jonathan Corbet, Shuah Khan, Randy Dunlap,
	linux-doc, Gaoyang Wei, Ziyou Xu, Benjamin Larsson



On 09.10.26 13:04, Andre Ziviani wrote:
> Hi John,
> 
> On Thu, Oct 08, 2026 at 04:32:37PM +0200, John Crispin wrote:
>> The scope is XGS-PON. The mode enum keeps gpon and xg-pon so that the
>> uapi order stays clean for later drivers.
> 
> A G-PON data point, since the framework thread mentioned Realtek only
> as work with proprietary components. odi-oss [1] is open firmware for a
> GPON SFP ONU stick on the Realtek RTL9602C. It runs mainline 6.18 with
> GPL drivers of our own for the GPON MAC, the PLOAM state machine of
> G.984.3, the CPU NIC and the switch, and with a userspace OMCI daemon.
> Apart from the stock bootloader, the image has no vendor code or
> binaries. It runs in production on two ISPs, and a user has it working
> on a third.
> 
[...]
> traces from both ISPs.
> 
> [1] https://github.com/AndreZiviani/odi-oss
> 
> Thanks,
> Andre Ziviani
> 


and those dongle can be bought on fs.com for ~12euro

I'll grab a couple. Thanks for the info !


^ permalink raw reply	[flat|nested] 18+ messages in thread

* [RFC net-next 00/12] net: add the PON subsystem
@ 2026-10-09 13:39 Stephan Pruecklmayer
  0 siblings, 0 replies; 18+ messages in thread
From: Stephan Pruecklmayer @ 2026-10-09 13:39 UTC (permalink / raw)
  To: John Crispin, David S. Miller, Eric Dumazet, Jakub Kicinski,
	Paolo Abeni, 20261008143249.3439762-1-john@phrozen.org
  Cc: Donald Hunter, Simon Horman, netdev@vger.kernel.org,
	linux-kernel@vger.kernel.org, Andrew Lunn, Christian Marangi,
	Jonathan Corbet, Shuah Khan, Randy Dunlap,
	linux-doc@vger.kernel.org, Gaoyang Wei, Ziyou Xu

Hi all,

I appreciate to get a common structure for PON gateways, 
HGUs and SFUs in place. 
As mentioned in the other thread on this topic 
([RFC] net: towards a generic PON framework)
https://lore.kernel.org/all/785bffbc-b028-46dd-b3de-b95a28122258@gmail.com/

we from MaxLinear have also discussed and implemented an abstraction at 
the Linux kernel boundary layer to abstract hardware devices and the OMCI stack 
accordingly. For that reason we definitely endorse such endevor in principle
and we will further comment and reply on this mailing list.

Having said that, such implementation shall be an implementation for all 
people using Linux and PON and not only one specific company implementation
that is being pushed here. The required abstraction needs to support multiple
chipsets and their underlying architectures. Different chips may do certain things
already in hardware (example PLOAM) while others do it in software.
Or how the data is being transferred between driver and OMCI (Ethernet or 
something else).

Before nailing down such interface, this homework needs to be done. 
From MaxLinear side we will look into this and compare to whatever public
information on the interface for the proposed Airoha chip is available, but from
a very first glance it cannot  and will not work as is, but needs abstractions for 
several building blocks. 
Our experts will reach out in this mailing list concerning specific implementations 
and questions on this current proposal.


Stephan Pruecklmayer
MaxLinear
Senior Technical Director Strategic Marketing and Standards




^ permalink raw reply	[flat|nested] 18+ messages in thread

* Re: [RFC net-next 00/12] net: add the PON subsystem
@ 2026-10-09 13:45 Stephan Pruecklmayer
  0 siblings, 0 replies; 18+ messages in thread
From: Stephan Pruecklmayer @ 2026-10-09 13:45 UTC (permalink / raw)
  To: Stephan Pruecklmayer, John Crispin
  Cc: andrew+netdev@lunn.ch, ansuelsmth@gmail.com, corbet@lwn.net,
	davem@davemloft.net, donald.hunter@gmail.com, edumazet@kernel.org,
	horms@kernel.org, kuba@kernel.org, linux-doc@vger.kernel.org,
	linux-kernel@vger.kernel.org, netdev@vger.kernel.org,
	pabeni@redhat.com, rdunlap@infradead.org,
	skhan@linuxfoundation.org, xuziyougm@gmail.com, yhyxwgy@gmail.com,
	Thomas Langer, Zahari Doychev

Apologize for the add-on - looping in my colleagues to the thread
tlanger@maxlinear.com
zdoychev@maxlinear.com


^ permalink raw reply	[flat|nested] 18+ messages in thread

end of thread, other threads:[~2026-10-09 13:45 UTC | newest]

Thread overview: 18+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-08 14:32 [RFC net-next 00/12] net: add the PON subsystem John Crispin
2026-10-08 14:32 ` [RFC net-next 01/12] net: pon: add the netlink specification of the pon family John Crispin
2026-10-08 14:32 ` [RFC net-next 02/12] net: add a PON device pointer to struct net_device John Crispin
2026-10-08 14:32 ` [RFC net-next 03/12] net: pon: add the PLOAM vocabulary and message codec John Crispin
2026-10-08 14:32 ` [RFC net-next 04/12] net: pon: add the device state and the lent context John Crispin
2026-10-08 14:32 ` [RFC net-next 05/12] net: pon: add the conduit contract John Crispin
2026-10-08 14:32 ` [RFC net-next 06/12] net: pon: add the GEM network devices John Crispin
2026-10-08 14:32 ` [RFC net-next 07/12] net: pon: add the OMCI channel John Crispin
2026-10-08 14:32 ` [RFC net-next 08/12] net: pon: add netlink support John Crispin
2026-10-08 14:32 ` [RFC net-next 09/12] net: pon: add the device registration John Crispin
2026-10-08 14:32 ` [RFC net-next 10/12] net: pon: build the PON subsystem John Crispin
2026-10-08 14:32 ` [RFC net-next 11/12] Documentation: networking: describe " John Crispin
2026-10-08 14:32 ` [RFC net-next 12/12] MAINTAINERS: add " John Crispin
2026-10-09  2:27 ` [RFC net-next 00/12] net: " Gaoyang Wei
2026-10-09 11:04 ` Andre Ziviani
2026-10-09 12:11   ` John Crispin
  -- strict thread matches above, loose matches on Subject: below --
2026-10-09 13:39 Stephan Pruecklmayer
2026-10-09 13:45 Stephan Pruecklmayer

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox