* Re: [PATCH v2 net-next] net: lpc_eth: fix trivial comment typo
From: David Miller @ 2018-11-22 0:18 UTC (permalink / raw)
To: aclaudi; +Cc: netdev, vz
In-Reply-To: <b78920afc46c38a6a7d2cd843615bda06574c4b0.1542734728.git.aclaudi@redhat.com>
From: Andrea Claudi <aclaudi@redhat.com>
Date: Tue, 20 Nov 2018 18:30:30 +0100
> Fix comment typo rxfliterctrl -> rxfilterctrl
>
> Signed-off-by: Andrea Claudi <aclaudi@redhat.com>
Applied.
^ permalink raw reply
* Re: [PATCH net V4 0/5] net/smc: fixes 2018-11-12
From: David Miller @ 2018-11-22 0:15 UTC (permalink / raw)
To: ubraun; +Cc: netdev, linux-s390, schwidefsky, heiko.carstens, raspl
In-Reply-To: <20181120154643.24111-1-ubraun@linux.ibm.com>
From: Ursula Braun <ubraun@linux.ibm.com>
Date: Tue, 20 Nov 2018 16:46:38 +0100
> here is V4 of some net/smc fixes in different areas for the net tree.
>
> v1->v2:
> do not define 8-byte alignment for union smcd_cdc_cursor in
> patch 4/5 "net/smc: atomic SMCD cursor handling"
> v2->v3:
> stay with 8-byte alignment for union smcd_cdc_cursor in
> patch 4/5 "net/smc: atomic SMCD cursor handling", but get rid of
> __packed for struct smcd_cdc_msg
> v3->v4:
> get rid of another __packed for struct smc_cdc_msg in
> patch 4/5 "net/smc: atomic SMCD cursor handling"
Series applied, thanks.
^ permalink raw reply
* Re: [PATCH net-next 3/3] tcp: implement head drops in backlog queue
From: Eric Dumazet @ 2018-11-21 23:52 UTC (permalink / raw)
To: Yuchung Cheng
Cc: Eric Dumazet, David Miller, netdev, Jean-Louis Dupond,
Neal Cardwell
In-Reply-To: <CAK6E8=e+o06AjzuOT5tZLS7uUp_h0gRV2K5KUbDX7mxVBQNj2g@mail.gmail.com>
On Wed, Nov 21, 2018 at 3:47 PM Yuchung Cheng <ycheng@google.com> wrote:
>
> On Wed, Nov 21, 2018 at 2:47 PM, Eric Dumazet <eric.dumazet@gmail.com> wrote:
> >
> >
> > On 11/21/2018 02:40 PM, Yuchung Cheng wrote:
> >> On Wed, Nov 21, 2018 at 9:52 AM, Eric Dumazet <edumazet@google.com> wrote:
> >>> Under high stress, and if GRO or coalescing does not help,
> >>> we better make room in backlog queue to be able to keep latest
> >>> packet coming.
> >>>
> >>> This generally helps fast recovery, given that we often receive
> >>> packets in order.
> >>
> >> I like the benefit of fast recovery but I am a bit leery about head
> >> drop causing HoLB on large read, while tail drops can be repaired by
> >> RACK and TLP already. Hmm -
> >
> > This is very different pattern here.
> >
> > We have a train of packets coming, the last packet is not a TLP probe...
> >
> > Consider this train coming from an old stack without burst control nor pacing.
> >
> > This patch guarantees last packet will be processed, and either :
> >
> > 1) We are a receiver, we will send a SACK. Sender will typically start recovery
> >
> > 2) We are a sender, we will process the most recent ACK sent by the receiver.
> >
>
> Sure on the sender it's universally good.
>
> On the receiver side my scenario was not the last packet being TLP.
> AFAIU the new design will first try coalesce the incoming skb to the
> tail one then exit. Otherwise it's queued to the back with an
> additional 64KB space credit. This patch checks the space w/o the
> extra credit and drops the head skb. If the head skb has been
> coalesced, we might end dropping the first big chunk that may need a
> few rounds of fast recovery to repair. But I am likely to
> misunderstand the patch :-)
This situation happens only when we are largely over committing
the memory, due to bad skb->len/skb->truesize ratio.
Think about MTU=9000 and some drivers allocating 16 KB of memory per frame,
even if the packet has 1500 bytes in it.
>
> Would it make sense to check the space first before the coalesce
> operation, and drop just enough bytes of the head to make room?
This is basically what the patch does, the while loop breaks when we have freed
just enough skbs.
^ permalink raw reply
* [v5, PATCH 2/2] dt-binding: mediatek-dwmac: add binding document for MediaTek MT2712 DWMAC
From: Biao Huang @ 2018-11-22 10:28 UTC (permalink / raw)
To: davem, robh+dt
Cc: honghui.zhang, yt.shen, liguo.zhang, mark.rutland, nelson.chang,
matthias.bgg, biao.huang, netdev, devicetree, linux-kernel,
linux-arm-kernel, linux-mediatek, joabreu, andrew
In-Reply-To: <1542882521-18874-1-git-send-email-biao.huang@mediatek.com>
The commit adds the device tree binding documentation for the MediaTek DWMAC
found on MediaTek MT2712.
Signed-off-by: Biao Huang <biao.huang@mediatek.com>
---
.../devicetree/bindings/net/mediatek-dwmac.txt | 78 ++++++++++++++++++++
1 file changed, 78 insertions(+)
create mode 100644 Documentation/devicetree/bindings/net/mediatek-dwmac.txt
diff --git a/Documentation/devicetree/bindings/net/mediatek-dwmac.txt b/Documentation/devicetree/bindings/net/mediatek-dwmac.txt
new file mode 100644
index 0000000..0f8a915
--- /dev/null
+++ b/Documentation/devicetree/bindings/net/mediatek-dwmac.txt
@@ -0,0 +1,78 @@
+MediaTek DWMAC glue layer controller
+
+This file documents platform glue layer for stmmac.
+Please see stmmac.txt for the other unchanged properties.
+
+The device node has following properties.
+
+Required properties:
+- compatible: Should be "mediatek,mt2712-gmac" for MT2712 SoC
+- reg: Address and length of the register set for the device
+- interrupts: Should contain the MAC interrupts
+- interrupt-names: Should contain a list of interrupt names corresponding to
+ the interrupts in the interrupts property, if available.
+ Should be "macirq" for the main MAC IRQ
+- clocks: Must contain a phandle for each entry in clock-names.
+- clock-names: The name of the clock listed in the clocks property. These are
+ "axi", "apb", "mac_main", "ptp_ref" for MT2712 SoC
+- mac-address: See ethernet.txt in the same directory
+- phy-mode: See ethernet.txt in the same directory
+
+Optional properties:
+- mediatek,tx-delay: TX clock delay macro value. Range is 0~31. Default is 0.
+ It should be defined for rgmii/rgmii-rxid/mii interface.
+- mediatek,rx-delay: RX clock delay macro value. Range is 0~31. Default is 0.
+ It should be defined for rgmii/rgmii-txid/mii/rmii interface.
+- mediatek,fine-tune: boolean property, if present indicates that fine delay
+ is selected for rgmii interface.
+ If present, tx-delay/rx-delay is 170+/-50ps per stage.
+ Else tx-delay/rx-delay of coarse delay macro is 0.55+/-0.2ns per stage.
+ This property do not apply to non-rgmii PHYs.
+ Only coarse-tune delay is supported for mii/rmii PHYs.
+- mediatek,rmii-rxc: boolean property, if present indicates that the rmii
+ reference clock, which is from external PHYs, is connected to RXC pin
+ on MT2712 SoC.
+ Otherwise, is connected to TXC pin.
+- mediatek,txc-inverse: boolean property, if present indicates that
+ 1. tx clock will be inversed in mii/rgmii case,
+ 2. tx clock inside MAC will be inversed relative to reference clock
+ which is from external PHYs in rmii case, and it rarely happen.
+- mediatek,rxc-inverse: boolean property, if present indicates that
+ 1. rx clock will be inversed in mii/rgmii case.
+ 2. reference clock will be inversed when arrived at MAC in rmii case.
+- assigned-clocks: mac_main and ptp_ref clocks
+- assigned-clock-parents: parent clocks of the assigned clocks
+
+Example:
+ eth: ethernet@1101c000 {
+ compatible = "mediatek,mt2712-gmac";
+ reg = <0 0x1101c000 0 0x1300>;
+ interrupts = <GIC_SPI 237 IRQ_TYPE_LEVEL_LOW>;
+ interrupt-names = "macirq";
+ phy-mode ="rgmii-id";
+ mac-address = [00 55 7b b5 7d f7];
+ clock-names = "axi",
+ "apb",
+ "mac_main",
+ "ptp_ref",
+ "ptp_top";
+ clocks = <&pericfg CLK_PERI_GMAC>,
+ <&pericfg CLK_PERI_GMAC_PCLK>,
+ <&topckgen CLK_TOP_ETHER_125M_SEL>,
+ <&topckgen CLK_TOP_ETHER_50M_SEL>;
+ assigned-clocks = <&topckgen CLK_TOP_ETHER_125M_SEL>,
+ <&topckgen CLK_TOP_ETHER_50M_SEL>;
+ assigned-clock-parents = <&topckgen CLK_TOP_ETHERPLL_125M>,
+ <&topckgen CLK_TOP_APLL1_D3>;
+ mediatek,pericfg = <&pericfg>;
+ mediatek,tx-delay = <9>;
+ mediatek,rx-delay = <9>;
+ mediatek,fine-tune;
+ mediatek,rmii-rxc;
+ mediatek,txc-inverse;
+ mediatek,rxc-inverse;
+ snps,txpbl = <32>;
+ snps,rxpbl = <32>;
+ snps,reset-gpio = <&pio 87 GPIO_ACTIVE_LOW>;
+ snps,reset-active-low;
+ };
--
1.7.9.5
^ permalink raw reply related
* [v5, PATCH 1/2] net:stmmac: dwmac-mediatek: add support for mt2712
From: Biao Huang @ 2018-11-22 10:28 UTC (permalink / raw)
To: davem, robh+dt
Cc: honghui.zhang, yt.shen, liguo.zhang, mark.rutland, nelson.chang,
matthias.bgg, biao.huang, netdev, devicetree, linux-kernel,
linux-arm-kernel, linux-mediatek, joabreu, andrew
In-Reply-To: <1542882521-18874-1-git-send-email-biao.huang@mediatek.com>
Add Ethernet support for MediaTek SoCs from the mt2712 family
Signed-off-by: Biao Huang <biao.huang@mediatek.com>
---
drivers/net/ethernet/stmicro/stmmac/Kconfig | 8 +
drivers/net/ethernet/stmicro/stmmac/Makefile | 1 +
.../net/ethernet/stmicro/stmmac/dwmac-mediatek.c | 364 ++++++++++++++++++++
3 files changed, 373 insertions(+)
create mode 100644 drivers/net/ethernet/stmicro/stmmac/dwmac-mediatek.c
diff --git a/drivers/net/ethernet/stmicro/stmmac/Kconfig b/drivers/net/ethernet/stmicro/stmmac/Kconfig
index 324049e..6209cc1 100644
--- a/drivers/net/ethernet/stmicro/stmmac/Kconfig
+++ b/drivers/net/ethernet/stmicro/stmmac/Kconfig
@@ -75,6 +75,14 @@ config DWMAC_LPC18XX
---help---
Support for NXP LPC18xx/43xx DWMAC Ethernet.
+config DWMAC_MEDIATEK
+ tristate "MediaTek MT27xx GMAC support"
+ depends on OF && (ARCH_MEDIATEK || COMPILE_TEST)
+ help
+ Support for MediaTek GMAC Ethernet controller.
+
+ This selects the MT2712 SoC support for the stmmac driver.
+
config DWMAC_MESON
tristate "Amlogic Meson dwmac support"
default ARCH_MESON
diff --git a/drivers/net/ethernet/stmicro/stmmac/Makefile b/drivers/net/ethernet/stmicro/stmmac/Makefile
index 99967a8..bf09701 100644
--- a/drivers/net/ethernet/stmicro/stmmac/Makefile
+++ b/drivers/net/ethernet/stmicro/stmmac/Makefile
@@ -13,6 +13,7 @@ obj-$(CONFIG_STMMAC_PLATFORM) += stmmac-platform.o
obj-$(CONFIG_DWMAC_ANARION) += dwmac-anarion.o
obj-$(CONFIG_DWMAC_IPQ806X) += dwmac-ipq806x.o
obj-$(CONFIG_DWMAC_LPC18XX) += dwmac-lpc18xx.o
+obj-$(CONFIG_DWMAC_MEDIATEK) += dwmac-mediatek.o
obj-$(CONFIG_DWMAC_MESON) += dwmac-meson.o dwmac-meson8b.o
obj-$(CONFIG_DWMAC_OXNAS) += dwmac-oxnas.o
obj-$(CONFIG_DWMAC_ROCKCHIP) += dwmac-rk.o
diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac-mediatek.c b/drivers/net/ethernet/stmicro/stmmac/dwmac-mediatek.c
new file mode 100644
index 0000000..dd8d4cc
--- /dev/null
+++ b/drivers/net/ethernet/stmicro/stmmac/dwmac-mediatek.c
@@ -0,0 +1,364 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2018 MediaTek Inc.
+ */
+#include <linux/io.h>
+#include <linux/mfd/syscon.h>
+#include <linux/of.h>
+#include <linux/of_device.h>
+#include <linux/of_net.h>
+#include <linux/regmap.h>
+#include <linux/stmmac.h>
+
+#include "stmmac.h"
+#include "stmmac_platform.h"
+
+/* Peri Configuration register for mt2712 */
+#define PERI_ETH_PHY_INTF_SEL 0x418
+#define PHY_INTF_MII_GMII 0
+#define PHY_INTF_RGMII 1
+#define PHY_INTF_RMII 4
+#define RMII_CLK_SRC_RXC BIT(4)
+#define RMII_CLK_SRC_INTERNAL BIT(5)
+
+#define PERI_ETH_PHY_DLY 0x428
+#define PHY_DLY_GTXC_INV BIT(6)
+#define PHY_DLY_GTXC_ENABLE BIT(5)
+#define PHY_DLY_GTXC_STAGES GENMASK(4, 0)
+#define PHY_DLY_TXC_INV BIT(20)
+#define PHY_DLY_TXC_ENABLE BIT(19)
+#define PHY_DLY_TXC_STAGES GENMASK(18, 14)
+#define PHY_DLY_TXC_SHIFT 14
+#define PHY_DLY_RXC_INV BIT(13)
+#define PHY_DLY_RXC_ENABLE BIT(12)
+#define PHY_DLY_RXC_STAGES GENMASK(11, 7)
+#define PHY_DLY_RXC_SHIFT 7
+
+#define PERI_ETH_DLY_FINE 0x800
+#define ETH_RMII_DLY_TX_INV BIT(2)
+#define ETH_FINE_DLY_GTXC BIT(1)
+#define ETH_FINE_DLY_RXC BIT(0)
+
+struct mac_delay_struct {
+ u32 tx_delay;
+ u32 rx_delay;
+ u32 tx_inv;
+ u32 rx_inv;
+};
+
+struct mediatek_dwmac_plat_data {
+ const struct mediatek_dwmac_variant *variant;
+ struct mac_delay_struct mac_delay;
+ struct clk_bulk_data *clks;
+ struct device_node *np;
+ struct regmap *peri_regmap;
+ struct device *dev;
+ int fine_tune;
+ int phy_mode;
+ int rmii_rxc;
+};
+
+struct mediatek_dwmac_variant {
+ int (*dwmac_set_phy_interface)(struct mediatek_dwmac_plat_data *plat);
+ int (*dwmac_set_delay)(struct mediatek_dwmac_plat_data *plat);
+
+ /* clock ids to be requested */
+ const char * const *clk_list;
+ int num_clks;
+
+ u32 dma_bit_mask;
+ u32 rx_delay_max;
+ u32 tx_delay_max;
+};
+
+/* list of clocks required for mac */
+static const char * const mt2712_dwmac_clk_l[] = {
+ "axi", "apb", "mac_main", "ptp_ref"
+};
+
+static int mt2712_set_interface(struct mediatek_dwmac_plat_data *plat)
+{
+ int rmii_rxc = plat->rmii_rxc ? RMII_CLK_SRC_RXC : 0;
+ u32 intf_val = 0;
+
+ /* select phy interface in top control domain */
+ switch (plat->phy_mode) {
+ case PHY_INTERFACE_MODE_MII:
+ intf_val |= PHY_INTF_MII_GMII;
+ break;
+ case PHY_INTERFACE_MODE_RMII:
+ intf_val |= PHY_INTF_RMII;
+ intf_val |= rmii_rxc;
+ break;
+ case PHY_INTERFACE_MODE_RGMII:
+ case PHY_INTERFACE_MODE_RGMII_TXID:
+ case PHY_INTERFACE_MODE_RGMII_RXID:
+ case PHY_INTERFACE_MODE_RGMII_ID:
+ intf_val |= PHY_INTF_RGMII;
+ break;
+ default:
+ dev_err(plat->dev, "phy interface not supported\n");
+ return -EINVAL;
+ }
+
+ regmap_write(plat->peri_regmap, PERI_ETH_PHY_INTF_SEL, intf_val);
+
+ return 0;
+}
+
+static int mt2712_set_delay(struct mediatek_dwmac_plat_data *plat)
+{
+ struct mac_delay_struct *mac_delay = &plat->mac_delay;
+ u32 delay_val = 0;
+ u32 fine_val = 0;
+
+ switch (plat->phy_mode) {
+ case PHY_INTERFACE_MODE_MII:
+ delay_val |= mac_delay->tx_delay ? PHY_DLY_TXC_ENABLE : 0;
+ delay_val |= (mac_delay->tx_delay << PHY_DLY_TXC_SHIFT) &
+ PHY_DLY_TXC_STAGES;
+ delay_val |= mac_delay->tx_inv ? PHY_DLY_TXC_INV : 0;
+ delay_val |= mac_delay->rx_delay ? PHY_DLY_RXC_ENABLE : 0;
+ delay_val |= (mac_delay->rx_delay << PHY_DLY_RXC_SHIFT) &
+ PHY_DLY_RXC_STAGES;
+ delay_val |= mac_delay->rx_inv ? PHY_DLY_RXC_INV : 0;
+ break;
+ case PHY_INTERFACE_MODE_RMII:
+ if (plat->rmii_rxc) {
+ delay_val |= mac_delay->rx_delay ?
+ PHY_DLY_RXC_ENABLE : 0;
+ delay_val |= (mac_delay->rx_delay <<
+ PHY_DLY_RXC_SHIFT) & PHY_DLY_RXC_STAGES;
+ delay_val |= mac_delay->rx_inv ? PHY_DLY_RXC_INV : 0;
+ fine_val |= mac_delay->tx_inv ?
+ ETH_RMII_DLY_TX_INV : 0;
+ } else {
+ delay_val |= mac_delay->rx_delay ?
+ PHY_DLY_TXC_ENABLE : 0;
+ delay_val |= (mac_delay->rx_delay <<
+ PHY_DLY_TXC_SHIFT) & PHY_DLY_TXC_STAGES;
+ delay_val |= mac_delay->rx_inv ? PHY_DLY_TXC_INV : 0;
+ fine_val |= mac_delay->tx_inv ?
+ ETH_RMII_DLY_TX_INV : 0;
+ }
+ break;
+ case PHY_INTERFACE_MODE_RGMII:
+ fine_val = plat->fine_tune ?
+ (ETH_FINE_DLY_GTXC | ETH_FINE_DLY_RXC) : 0;
+ delay_val |= mac_delay->tx_delay ? PHY_DLY_GTXC_ENABLE : 0;
+ delay_val |= mac_delay->tx_delay & PHY_DLY_GTXC_STAGES;
+ delay_val |= mac_delay->tx_inv ? PHY_DLY_GTXC_INV : 0;
+ delay_val |= mac_delay->rx_delay ? PHY_DLY_RXC_ENABLE : 0;
+ delay_val |= (mac_delay->rx_delay << PHY_DLY_RXC_SHIFT) &
+ PHY_DLY_RXC_STAGES;
+ delay_val |= mac_delay->rx_inv ? PHY_DLY_RXC_INV : 0;
+ break;
+ case PHY_INTERFACE_MODE_RGMII_TXID:
+ fine_val = plat->fine_tune ? ETH_FINE_DLY_RXC : 0;
+ delay_val |= mac_delay->rx_delay ? PHY_DLY_RXC_ENABLE : 0;
+ delay_val |= (mac_delay->rx_delay << PHY_DLY_RXC_SHIFT) &
+ PHY_DLY_RXC_STAGES;
+ delay_val |= mac_delay->rx_inv ? PHY_DLY_RXC_INV : 0;
+ break;
+ case PHY_INTERFACE_MODE_RGMII_RXID:
+ fine_val = plat->fine_tune ? ETH_FINE_DLY_GTXC : 0;
+ delay_val |= mac_delay->tx_delay ? PHY_DLY_GTXC_ENABLE : 0;
+ delay_val |= mac_delay->tx_delay & PHY_DLY_GTXC_STAGES;
+ delay_val |= mac_delay->tx_inv ? PHY_DLY_GTXC_INV : 0;
+ break;
+ case PHY_INTERFACE_MODE_RGMII_ID:
+ break;
+ default:
+ dev_err(plat->dev, "phy interface not supported\n");
+ return -EINVAL;
+ }
+ regmap_write(plat->peri_regmap, PERI_ETH_PHY_DLY, delay_val);
+ regmap_write(plat->peri_regmap, PERI_ETH_DLY_FINE, fine_val);
+
+ return 0;
+}
+
+static const struct mediatek_dwmac_variant mt2712_gmac_variant = {
+ .dwmac_set_phy_interface = mt2712_set_interface,
+ .dwmac_set_delay = mt2712_set_delay,
+ .clk_list = mt2712_dwmac_clk_l,
+ .num_clks = ARRAY_SIZE(mt2712_dwmac_clk_l),
+ .dma_bit_mask = 33,
+ .rx_delay_max = 32,
+ .tx_delay_max = 32,
+};
+
+static int mediatek_dwmac_config_dt(struct mediatek_dwmac_plat_data *plat)
+{
+ u32 tx_delay, rx_delay;
+
+ plat->peri_regmap = syscon_regmap_lookup_by_phandle(plat->np, "mediatek,pericfg");
+ if (IS_ERR(plat->peri_regmap)) {
+ dev_err(plat->dev, "Failed to get pericfg syscon\n");
+ return PTR_ERR(plat->peri_regmap);
+ }
+
+ plat->phy_mode = of_get_phy_mode(plat->np);
+ if (plat->phy_mode < 0) {
+ dev_err(plat->dev, "not find phy-mode\n");
+ return -EINVAL;
+ }
+
+ if (!of_property_read_u32(plat->np, "mediatek,tx-delay", &tx_delay)) {
+ if (tx_delay < plat->variant->tx_delay_max) {
+ plat->mac_delay.tx_delay = tx_delay;
+ } else {
+ dev_err(plat->dev, "Invalid TX clock delay: %d\n", tx_delay);
+ return -EINVAL;
+ }
+ }
+
+ if (!of_property_read_u32(plat->np, "mediatek,rx-delay", &rx_delay)) {
+ if (rx_delay < plat->variant->rx_delay_max) {
+ plat->mac_delay.rx_delay = rx_delay;
+ } else {
+ dev_err(plat->dev, "Invalid RX clock delay: %d\n", rx_delay);
+ return -EINVAL;
+ }
+ }
+
+ plat->mac_delay.tx_inv = of_property_read_bool(plat->np, "mediatek,txc-inverse");
+ plat->mac_delay.rx_inv = of_property_read_bool(plat->np, "mediatek,rxc-inverse");
+ plat->fine_tune = of_property_read_bool(plat->np, "mediatek,fine-tune");
+ plat->rmii_rxc = of_property_read_bool(plat->np, "mediatek,rmii-rxc");
+
+ return 0;
+}
+
+static int mediatek_dwmac_clk_init(struct mediatek_dwmac_plat_data *plat)
+{
+ const struct mediatek_dwmac_variant *variant = plat->variant;
+ int num = variant->num_clks;
+ int i;
+
+ plat->clks = devm_kcalloc(plat->dev, num, sizeof(*plat->clks), GFP_KERNEL);
+ if (!plat->clks)
+ return -ENOMEM;
+
+ for (i = 0; i < num; i++)
+ plat->clks[i].id = variant->clk_list[i];
+
+ return devm_clk_bulk_get(plat->dev, num, plat->clks);
+}
+
+static int mediatek_dwmac_init(struct platform_device *pdev, void *priv)
+{
+ struct mediatek_dwmac_plat_data *plat = priv;
+ const struct mediatek_dwmac_variant *variant = plat->variant;
+ int ret = 0;
+
+ ret = dma_set_mask_and_coherent(plat->dev, DMA_BIT_MASK(variant->dma_bit_mask));
+ if (ret) {
+ dev_err(plat->dev, "No suitable DMA available, err = %d\n", ret);
+ return ret;
+ }
+
+ ret = variant->dwmac_set_phy_interface(plat);
+ if (ret) {
+ dev_err(plat->dev, "failed to set phy interface, err = %d\n", ret);
+ return ret;
+ }
+
+ ret = variant->dwmac_set_delay(plat);
+ if (ret) {
+ dev_err(plat->dev, "failed to set delay value, err = %d\n", ret);
+ return ret;
+ }
+
+ ret = clk_bulk_prepare_enable(variant->num_clks, plat->clks);
+ if (ret) {
+ dev_err(plat->dev, "failed to enable clks, err = %d\n", ret);
+ return ret;
+ }
+
+ return 0;
+}
+
+static void mediatek_dwmac_exit(struct platform_device *pdev, void *priv)
+{
+ struct mediatek_dwmac_plat_data *plat = priv;
+ const struct mediatek_dwmac_variant *variant = plat->variant;
+
+ clk_bulk_disable_unprepare(variant->num_clks, plat->clks);
+}
+
+static int mediatek_dwmac_probe(struct platform_device *pdev)
+{
+ struct mediatek_dwmac_plat_data *priv_plat;
+ struct plat_stmmacenet_data *plat_dat;
+ struct stmmac_resources stmmac_res;
+ int ret = 0;
+
+ priv_plat = devm_kzalloc(&pdev->dev, sizeof(*priv_plat), GFP_KERNEL);
+ if (!priv_plat)
+ return -ENOMEM;
+
+ priv_plat->variant = of_device_get_match_data(&pdev->dev);
+ if (!priv_plat->variant) {
+ dev_err(&pdev->dev, "Missing dwmac-mediatek variant\n");
+ return -EINVAL;
+ }
+
+ priv_plat->dev = &pdev->dev;
+ priv_plat->np = pdev->dev.of_node;
+
+ ret = mediatek_dwmac_config_dt(priv_plat);
+ if (ret)
+ return ret;
+
+ ret = mediatek_dwmac_clk_init(priv_plat);
+ if (ret)
+ return ret;
+
+ ret = stmmac_get_platform_resources(pdev, &stmmac_res);
+ if (ret)
+ return ret;
+
+ plat_dat = stmmac_probe_config_dt(pdev, &stmmac_res.mac);
+ if (IS_ERR(plat_dat))
+ return PTR_ERR(plat_dat);
+
+ plat_dat->interface = priv_plat->phy_mode;
+ /* clk_csr_i = 250-300MHz & MDC = clk_csr_i/124 */
+ plat_dat->clk_csr = 5;
+ plat_dat->has_gmac4 = 1;
+ plat_dat->has_gmac = 0;
+ plat_dat->pmt = 0;
+ plat_dat->maxmtu = ETH_DATA_LEN;
+ plat_dat->bsp_priv = priv_plat;
+ plat_dat->init = mediatek_dwmac_init;
+ plat_dat->exit = mediatek_dwmac_exit;
+ mediatek_dwmac_init(pdev, priv_plat);
+
+ ret = stmmac_dvr_probe(&pdev->dev, plat_dat, &stmmac_res);
+ if (ret) {
+ stmmac_remove_config_dt(pdev, plat_dat);
+ return ret;
+ }
+
+ return 0;
+}
+
+static const struct of_device_id mediatek_dwmac_match[] = {
+ { .compatible = "mediatek,mt2712-gmac",
+ .data = &mt2712_gmac_variant },
+ { }
+};
+
+MODULE_DEVICE_TABLE(of, mediatek_dwmac_match);
+
+static struct platform_driver mediatek_dwmac_driver = {
+ .probe = mediatek_dwmac_probe,
+ .remove = stmmac_pltfr_remove,
+ .driver = {
+ .name = "dwmac-mediatek",
+ .pm = &stmmac_pltfr_pm_ops,
+ .of_match_table = mediatek_dwmac_match,
+ },
+};
+module_platform_driver(mediatek_dwmac_driver);
--
1.7.9.5
^ permalink raw reply related
* [v5, PATCH 0/2] add Ethernet driver support for mt2712
From: Biao Huang @ 2018-11-22 10:28 UTC (permalink / raw)
To: davem, robh+dt
Cc: honghui.zhang, yt.shen, liguo.zhang, mark.rutland, nelson.chang,
matthias.bgg, biao.huang, netdev, devicetree, linux-kernel,
linux-arm-kernel, linux-mediatek, joabreu, andrew
Changes in v5:
1. fix dma bit mask setting.
2. remove one training whitespace.
^ permalink raw reply
* Re: [PATCH net] tcp: defer SACK compression after DupThresh
From: David Miller @ 2018-11-21 23:50 UTC (permalink / raw)
To: edumazet; +Cc: netdev, jean-louis, ncardwell, ycheng, eric.dumazet
In-Reply-To: <20181120135359.7539-1-edumazet@google.com>
From: Eric Dumazet <edumazet@google.com>
Date: Tue, 20 Nov 2018 05:53:59 -0800
> Jean-Louis reported a TCP regression and bisected to recent SACK
> compression.
>
> After a loss episode (receiver not able to keep up and dropping
> packets because its backlog is full), linux TCP stack is sending
> a single SACK (DUPACK).
>
> Sender waits a full RTO timer before recovering losses.
>
> While RFC 6675 says in section 5, "Algorithm Details",
>
> (2) If DupAcks < DupThresh but IsLost (HighACK + 1) returns true --
> indicating at least three segments have arrived above the current
> cumulative acknowledgment point, which is taken to indicate loss
> -- go to step (4).
> ...
> (4) Invoke fast retransmit and enter loss recovery as follows:
>
> there are old TCP stacks not implementing this strategy, and
> still counting the dupacks before starting fast retransmit.
>
> While these stacks probably perform poorly when receivers implement
> LRO/GRO, we should be a little more gentle to them.
>
> This patch makes sure we do not enable SACK compression unless
> 3 dupacks have been sent since last rcv_nxt update.
>
> Ideally we should even rearm the timer to send one or two
> more DUPACK if no more packets are coming, but that will
> be work aiming for linux-4.21.
>
> Many thanks to Jean-Louis for bisecting the issue, providing
> packet captures and testing this patch.
>
> Fixes: 5d9f4262b7ea ("tcp: add SACK compression")
> Reported-by: Jean-Louis Dupond <jean-louis@dupond.be>
> Tested-by: Jean-Louis Dupond <jean-louis@dupond.be>
> Signed-off-by: Eric Dumazet <edumazet@google.com>
> Acked-by: Neal Cardwell <ncardwell@google.com>
Applied and queued up for -stable.
Thanks Eric.
^ permalink raw reply
* Re: [PATCH net-next,v3 04/12] cls_api: add translator to flow_action representation
From: Pablo Neira Ayuso @ 2018-11-21 23:48 UTC (permalink / raw)
To: Marcelo Ricardo Leitner
Cc: netdev, davem, thomas.lendacky, f.fainelli, ariel.elior,
michael.chan, santosh, madalin.bucur, yisen.zhuang, salil.mehta,
jeffrey.t.kirsher, tariqt, saeedm, jiri, idosch, jakub.kicinski,
peppe.cavallaro, grygorii.strashko, andrew, vivien.didelot,
alexandre.torgue, joabreu, linux-net-drivers, ganeshgr, ogerlitz,
Manish.Chopra
In-Reply-To: <20181121211541.GA8353@localhost.localdomain>
On Wed, Nov 21, 2018 at 07:15:41PM -0200, Marcelo Ricardo Leitner wrote:
> On Wed, Nov 21, 2018 at 03:51:24AM +0100, Pablo Neira Ayuso wrote:
[...]
> > diff --git a/net/sched/cls_flower.c b/net/sched/cls_flower.c
> > index d2971fbfc3d9..8898943b8ee6 100644
> > --- a/net/sched/cls_flower.c
> > +++ b/net/sched/cls_flower.c
> > @@ -382,7 +382,7 @@ static int fl_hw_replace_filter(struct tcf_proto *tp,
> > bool skip_sw = tc_skip_sw(f->flags);
> > int err;
> >
> > - cls_flower.rule = flow_rule_alloc();
> > + cls_flower.rule = flow_rule_alloc(tcf_exts_num_actions(&f->exts));
>
> As previous patch did:
> -struct flow_rule *flow_rule_alloc(void);
> +struct flow_rule *flow_rule_alloc(unsigned int num_actions);
> the build is broken without this change (bisect-ability).
> (applies to similar lines too)
Sorry about this, I'll fix this, thanks Marcelo.
^ permalink raw reply
* Re: [PATCH bpf] tools: bpftool: fix potential NULL pointer dereference in do_load
From: Daniel Borkmann @ 2018-11-21 23:47 UTC (permalink / raw)
To: Jakub Kicinski, alexei.starovoitov
Cc: oss-drivers, netdev, Wen Yang, Julia Lawall
In-Reply-To: <20181121215317.14887-1-jakub.kicinski@netronome.com>
On 11/21/2018 10:53 PM, Jakub Kicinski wrote:
> This patch fixes a possible null pointer dereference in
> do_load, detected by the semantic patch deref_null.cocci,
> with the following warning:
>
> ./tools/bpf/bpftool/prog.c:1021:23-25: ERROR: map_replace is NULL but dereferenced.
>
> The following code has potential null pointer references:
> 881 map_replace = reallocarray(map_replace, old_map_fds + 1,
> 882 sizeof(*map_replace));
> 883 if (!map_replace) {
> 884 p_err("mem alloc failed");
> 885 goto err_free_reuse_maps;
> 886 }
>
> ...
> 1019 err_free_reuse_maps:
> 1020 for (i = 0; i < old_map_fds; i++)
> 1021 close(map_replace[i].fd);
> 1022 free(map_replace);
>
> Fixes: 3ff5a4dc5d89 ("tools: bpftool: allow reuse of maps with bpftool prog load")
> Co-developed-by: Wen Yang <wen.yang99@zte.com.cn>
> Signed-off-by: Wen Yang <wen.yang99@zte.com.cn>
> Signed-off-by: Jakub Kicinski <jakub.kicinski@netronome.com>
Applied to bpf, thanks everyone!
^ permalink raw reply
* Re: [PATCH net-next 3/3] tcp: implement head drops in backlog queue
From: Yuchung Cheng @ 2018-11-21 23:46 UTC (permalink / raw)
To: Eric Dumazet
Cc: Eric Dumazet, David S . Miller, netdev, Jean-Louis Dupond,
Neal Cardwell
In-Reply-To: <f86128c2-1109-a397-3a70-098556c21cea@gmail.com>
On Wed, Nov 21, 2018 at 2:47 PM, Eric Dumazet <eric.dumazet@gmail.com> wrote:
>
>
> On 11/21/2018 02:40 PM, Yuchung Cheng wrote:
>> On Wed, Nov 21, 2018 at 9:52 AM, Eric Dumazet <edumazet@google.com> wrote:
>>> Under high stress, and if GRO or coalescing does not help,
>>> we better make room in backlog queue to be able to keep latest
>>> packet coming.
>>>
>>> This generally helps fast recovery, given that we often receive
>>> packets in order.
>>
>> I like the benefit of fast recovery but I am a bit leery about head
>> drop causing HoLB on large read, while tail drops can be repaired by
>> RACK and TLP already. Hmm -
>
> This is very different pattern here.
>
> We have a train of packets coming, the last packet is not a TLP probe...
>
> Consider this train coming from an old stack without burst control nor pacing.
>
> This patch guarantees last packet will be processed, and either :
>
> 1) We are a receiver, we will send a SACK. Sender will typically start recovery
>
> 2) We are a sender, we will process the most recent ACK sent by the receiver.
>
Sure on the sender it's universally good.
On the receiver side my scenario was not the last packet being TLP.
AFAIU the new design will first try coalesce the incoming skb to the
tail one then exit. Otherwise it's queued to the back with an
additional 64KB space credit. This patch checks the space w/o the
extra credit and drops the head skb. If the head skb has been
coalesced, we might end dropping the first big chunk that may need a
few rounds of fast recovery to repair. But I am likely to
misunderstand the patch :-)
Would it make sense to check the space first before the coalesce
operation, and drop just enough bytes of the head to make room?
^ permalink raw reply
* Re: [PATCH net-next 0/4] VLAN tag handling cleanup
From: David Miller @ 2018-11-21 23:41 UTC (permalink / raw)
To: mirq-linux
Cc: netdev, ajit.khaparde, devel, haiyangz, kys, leon, linux-rdma,
saeedm, sathya.perla, somnath.kotur, sriharsha.basavapatna,
sthemmin
In-Reply-To: <cover.1542716156.git.mirq-linux@rere.qmqm.pl>
From: Michał Mirosław <mirq-linux@rere.qmqm.pl>
Date: Tue, 20 Nov 2018 13:20:30 +0100
> This is a cleanup set after VLAN_TAG_PRESENT removal. The CFI bit
> handling is made similar to how other tag fields are used.
Series applied, thanks.
^ permalink raw reply
* Re: [PATCH net] net: skb_scrub_packet(): Scrub offload_fwd_mark
From: David Miller @ 2018-11-21 23:40 UTC (permalink / raw)
To: petrm; +Cc: netdev, idosch
In-Reply-To: <1bc608bd028b41a21c99c982f459b7434f0948ed.1542712365.git.petrm@mellanox.com>
From: Petr Machata <petrm@mellanox.com>
Date: Tue, 20 Nov 2018 11:39:56 +0000
> When a packet is trapped and the corresponding SKB marked as
> already-forwarded, it retains this marking even after it is forwarded
> across veth links into another bridge. There, since it ingresses the
> bridge over veth, which doesn't have offload_fwd_mark, it triggers a
> warning in nbp_switchdev_frame_mark().
>
> Then nbp_switchdev_allowed_egress() decides not to allow egress from
> this bridge through another veth, because the SKB is already marked, and
> the mark (of 0) of course matches. Thus the packet is incorrectly
> blocked.
>
> Solve by resetting offload_fwd_mark() in skb_scrub_packet(). That
> function is called from tunnels and also from veth, and thus catches the
> cases where traffic is forwarded between bridges and transformed in a
> way that invalidates the marking.
>
> Fixes: 6bc506b4fb06 ("bridge: switchdev: Add forward mark support for stacked devices")
> Fixes: abf4bb6b63d0 ("skbuff: Add the offload_mr_fwd_mark field")
> Signed-off-by: Petr Machata <petrm@mellanox.com>
> Suggested-by: Ido Schimmel <idosch@mellanox.com>
Applied and queued up for -stable, thanks.
^ permalink raw reply
* [PATCH][next] tools/bpf: fix spelling mistake "memeory" -> "memory"
From: Colin King @ 2018-11-22 10:13 UTC (permalink / raw)
To: Alexei Starovoitov, Daniel Borkmann, Shuah Khan, netdev,
linux-kselftest
Cc: kernel-janitors, linux-kernel
From: Colin Ian King <colin.king@canonical.com>
The CHECK message contains a spelling mistake, fix it.
Signed-off-by: Colin Ian King <colin.king@canonical.com>
---
tools/testing/selftests/bpf/test_btf.c | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
diff --git a/tools/testing/selftests/bpf/test_btf.c b/tools/testing/selftests/bpf/test_btf.c
index 7b1b160d6e67..bcbda7037840 100644
--- a/tools/testing/selftests/bpf/test_btf.c
+++ b/tools/testing/selftests/bpf/test_btf.c
@@ -2573,7 +2573,7 @@ static int do_test_file(unsigned int test_num)
}
func_info = malloc(info.func_info_cnt * rec_size);
- if (CHECK(!func_info, "out of memeory")) {
+ if (CHECK(!func_info, "out of memory")) {
err = -1;
goto done;
}
@@ -3299,7 +3299,7 @@ static int do_test_func_type(int test_num)
}
func_info = malloc(info.func_info_cnt * rec_size);
- if (CHECK(!func_info, "out of memeory")) {
+ if (CHECK(!func_info, "out of memory")) {
err = -1;
goto done;
}
--
2.19.1
^ permalink raw reply related
* [PATCH][V2] net: hinic: fix null pointer dereference on pointer hwdev
From: Colin King @ 2018-11-22 10:05 UTC (permalink / raw)
To: Aviad Krawczyk, David S . Miller, netdev
Cc: kernel-janitors, sergei.shtylyov, linux-kernel
From: Colin Ian King <colin.king@canonical.com>
Pointer hwdev is being dereferenced when declaring hwif , however, later
on hwdev is being null checked, hence we have dereference before null
check error. Fix this by assigning hwif and pdef only once hwdev has
been null checked.
Detected by CoverityScan, CID#1485581 ("Dereference before null check")
Fixes: 4a61abb100c8 ("net-next/hinic:add rx checksum offload for HiNIC")
Signed-off-by: Colin Ian King <colin.king@canonical.com>
---
V2: fix spelling mistake in commit message, remove a blank line in
my original commit.
---
drivers/net/ethernet/huawei/hinic/hinic_port.c | 6 ++++--
1 file changed, 4 insertions(+), 2 deletions(-)
diff --git a/drivers/net/ethernet/huawei/hinic/hinic_port.c b/drivers/net/ethernet/huawei/hinic/hinic_port.c
index e9f76e904610..122c93597268 100644
--- a/drivers/net/ethernet/huawei/hinic/hinic_port.c
+++ b/drivers/net/ethernet/huawei/hinic/hinic_port.c
@@ -414,14 +414,16 @@ int hinic_set_rx_csum_offload(struct hinic_dev *nic_dev, u32 en)
{
struct hinic_checksum_offload rx_csum_cfg = {0};
struct hinic_hwdev *hwdev = nic_dev->hwdev;
- struct hinic_hwif *hwif = hwdev->hwif;
- struct pci_dev *pdev = hwif->pdev;
+ struct hinic_hwif *hwif;
+ struct pci_dev *pdev;
u16 out_size;
int err;
if (!hwdev)
return -EINVAL;
+ hwif = hwdev->hwif;
+ pdev = hwif->pdev;
rx_csum_cfg.func_id = HINIC_HWIF_FUNC_IDX(hwif);
rx_csum_cfg.rx_csum_offload = en;
--
2.19.1
^ permalink raw reply related
* Re: [PATCH bpf-next] bpf: add read/write access to skb->tstamp from tc clsact progs
From: Daniel Borkmann @ 2018-11-21 22:57 UTC (permalink / raw)
To: Vlad Dumitrescu, eric.dumazet, alexei.starovoitov
Cc: Vlad Dumitrescu, ast, netdev, edumazet, willemb
In-Reply-To: <CALpBo+URiVUOkATDFtuyTnbX-FQVofy05Lh4tUb88+bihR+AXw@mail.gmail.com>
On 11/21/2018 07:48 PM, Vlad Dumitrescu wrote:
> On Wed, Nov 21, 2018 at 5:08 AM Eric Dumazet <eric.dumazet@gmail.com> wrote:
>> On 11/20/2018 06:40 PM, Alexei Starovoitov wrote:
>>>
>>> looks good to me.
>>>
>>> Any particular reason you decided to disable it for cg_skb ?
>>> It seems to me the same EDT approach will work from
>>> cgroup-bpf skb hooks just as well and then we can have neat
>>> way of controlling traffic per-container instead of tc-clsbpf global.
>>> If you're already on cgroup v2 it will save you a lot of classifier
>>> cycles, since you'd be able to group apps by cgroup
>>> instead of relying on ip only.
>>
>> Vlad first wrote a complete version, but we felt explaining the _why_
>> was probably harder.
>>
>> No particular reason, other than having to write more tests perhaps.
>
> This sounds reasonable to me. I can prepare a v2.
>
> Any concerns regarding capabilities? For example data and data_end are
> only available to CAP_SYS_ADMIN. Note that enforcement of this would
> be done by a global component later in the pipeline (e.g., FQ qdisc).
cg_skb_is_valid_access() has the CAP_SYS_ADMIN enforcement for direct
packet access since cg_skb can also run from unprivileged. Makes sense
to do the same for skb->tstamp for the STX_MEM part at least.
> Any opinions on sk_filter, lwt, and sk_skb before I send v2?
I'd probably leave that out for the time being if there is no concrete
use at this point.
Thanks,
Daniel
^ permalink raw reply
* Re: [PATCH perf,bpf 0/5] reveal invisible bpf programs
From: Peter Zijlstra @ 2018-11-22 9:32 UTC (permalink / raw)
To: Song Liu; +Cc: netdev, linux-kernel, ast, daniel, acme, kernel-team
In-Reply-To: <20181121195502.3259930-1-songliubraving@fb.com>
On Wed, Nov 21, 2018 at 11:54:57AM -0800, Song Liu wrote:
> Changes RFC -> PATCH v1:
>
> 1. In perf-record, poll vip events in a separate thread;
> 2. Add tag to bpf prog name;
> 3. Small refactorings.
>
> Original cover letter (with minor revisions):
>
> This is to follow up Alexei's early effort to show bpf programs
>
> https://www.spinics.net/lists/netdev/msg524232.html
>
> In this version, PERF_RECORD_BPF_EVENT is introduced to send real time BPF
> load/unload events to user space. In user space, perf-record is modified
> to listen to these events (through a dedicated ring buffer) and generate
> detailed information about the program (struct bpf_prog_info_event). Then,
> perf-report translates these events into proper symbols.
>
> With this set, perf-report will show bpf program as:
>
> 18.49% 0.16% test [kernel.vmlinux] [k] ksys_write
> 18.01% 0.47% test [kernel.vmlinux] [k] vfs_write
> 17.02% 0.40% test bpf_prog [k] bpf_prog_07367f7ba80df72b_
> 16.97% 0.10% test [kernel.vmlinux] [k] __vfs_write
> 16.86% 0.12% test [kernel.vmlinux] [k] comm_write
> 16.67% 0.39% test [kernel.vmlinux] [k] bpf_probe_read
>
> Note that, the program name is still work in progress, it will be cleaner
> with function types in BTF.
>
> Please share your comments on this.
So I see:
kernel/bpf/core.c:void bpf_prog_kallsyms_add(struct bpf_prog *fp)
which should already provide basic symbol information for extant eBPF
programs, right?
And (AFAIK) perf uses /proc/kcore for annotate on the current running
kernel (if not, it really should, given alternatives, jump_labels and
all other other self-modifying code).
So this fancy new stuff is only for the case where your profile spans
eBPF load/unload events (which should be relatively rare in the normal
case, right), or when you want source annotated asm output (I normally
don't bother with that).
That is; I would really like this fancy stuff to be an optional extra
that is typically not needed.
Does that make sense?
^ permalink raw reply
* Re: [PATCH net-next 3/3] tcp: implement head drops in backlog queue
From: Eric Dumazet @ 2018-11-21 22:47 UTC (permalink / raw)
To: Yuchung Cheng, Eric Dumazet
Cc: David S . Miller, netdev, Jean-Louis Dupond, Neal Cardwell
In-Reply-To: <CAK6E8=d1CxqkYy+Ad9Vw4NpP_6UKir5Rik=EuN=zGD8o-4USWQ@mail.gmail.com>
On 11/21/2018 02:40 PM, Yuchung Cheng wrote:
> On Wed, Nov 21, 2018 at 9:52 AM, Eric Dumazet <edumazet@google.com> wrote:
>> Under high stress, and if GRO or coalescing does not help,
>> we better make room in backlog queue to be able to keep latest
>> packet coming.
>>
>> This generally helps fast recovery, given that we often receive
>> packets in order.
>
> I like the benefit of fast recovery but I am a bit leery about head
> drop causing HoLB on large read, while tail drops can be repaired by
> RACK and TLP already. Hmm -
This is very different pattern here.
We have a train of packets coming, the last packet is not a TLP probe...
Consider this train coming from an old stack without burst control nor pacing.
This patch guarantees last packet will be processed, and either :
1) We are a receiver, we will send a SACK. Sender will typically start recovery
2) We are a sender, we will process the most recent ACK sent by the receiver.
^ permalink raw reply
* Re: [PATCH bpf-next] bpf: add read/write access to skb->tstamp from tc clsact progs
From: Alexei Starovoitov @ 2018-11-21 22:46 UTC (permalink / raw)
To: Vlad Dumitrescu
Cc: eric.dumazet, Vlad Dumitrescu, ast, Daniel Borkmann, netdev,
edumazet, willemb
In-Reply-To: <CALpBo+URiVUOkATDFtuyTnbX-FQVofy05Lh4tUb88+bihR+AXw@mail.gmail.com>
On Wed, Nov 21, 2018 at 10:48:21AM -0800, Vlad Dumitrescu wrote:
> On Wed, Nov 21, 2018 at 5:08 AM Eric Dumazet <eric.dumazet@gmail.com> wrote:
> >
> >
> >
> > On 11/20/2018 06:40 PM, Alexei Starovoitov wrote:
> >
> > >
> > > looks good to me.
> > >
> > > Any particular reason you decided to disable it for cg_skb ?
> > > It seems to me the same EDT approach will work from
> > > cgroup-bpf skb hooks just as well and then we can have neat
> > > way of controlling traffic per-container instead of tc-clsbpf global.
> > > If you're already on cgroup v2 it will save you a lot of classifier
> > > cycles, since you'd be able to group apps by cgroup
> > > instead of relying on ip only.
> >
> > Vlad first wrote a complete version, but we felt explaining the _why_
> > was probably harder.
> >
> > No particular reason, other than having to write more tests perhaps.
>
> This sounds reasonable to me. I can prepare a v2.
thank you
> Any concerns regarding capabilities? For example data and data_end are
> only available to CAP_SYS_ADMIN. Note that enforcement of this would
> be done by a global component later in the pipeline (e.g., FQ qdisc).
I'd do cap_sys_admin for now, since i'm not sure whether any tstamp
values will be acceptable to fq.
> Any opinions on sk_filter, lwt, and sk_skb before I send v2?
sk_filter not appealing, since it's too late in the stack.
lwt could be interesting, but I'd wait until first user appears.
sk_skb - useful, but it requires more work.
We'll follow up to that sk_skb with our own patches.
Thanks!
^ permalink raw reply
* Re: [PATCH][net-next] net: hinic: fix null pointer dereference on pointer hwdev
From: Sergei Shtylyov @ 2018-11-22 9:22 UTC (permalink / raw)
To: Colin King, Aviad Krawczyk, David S . Miller, netdev
Cc: kernel-janitors, linux-kernel
In-Reply-To: <20181121195956.2730-1-colin.king@canonical.com>
Hello!
On 21.11.2018 22:59, Colin King wrote:
> From: Colin Ian King <colin.king@canonical.com>
>
> Pointer hwdev is being dereferenced when declaring hwif , however, later
> on hwdev is being null checked, hence we have dereference before null
> check error. Fix this by assigning hwiw and pdef only once hwdev has
It's hwif, right?
> been null checked.
>
> Detected by CoverityScan, CID#1485581 ("Dereference before null check")
>
> Fixes: 4a61abb100c8 ("net-next/hinic:add rx checksum offload for HiNIC")
>
No need for this empty line.
> Signed-off-by: Colin Ian King <colin.king@canonical.com>
[...]
MBR, Sergei
^ permalink raw reply
* Re: [PATCH net-next 1/3] tcp: remove hdrlen argument from tcp_queue_rcv()
From: Yuchung Cheng @ 2018-11-21 22:41 UTC (permalink / raw)
To: Eric Dumazet
Cc: David S . Miller, netdev, Jean-Louis Dupond, Neal Cardwell,
Eric Dumazet
In-Reply-To: <20181121175240.6075-2-edumazet@google.com>
On Wed, Nov 21, 2018 at 9:52 AM, Eric Dumazet <edumazet@google.com> wrote:
> Only one caller needs to pull TCP headers, so lets
> move __skb_pull() to the caller side.
>
> Signed-off-by: Eric Dumazet <edumazet@google.com>
> ---
Acked-by: Yuchung Cheng <ycheng@google.com>
> net/ipv4/tcp_input.c | 13 ++++++-------
> 1 file changed, 6 insertions(+), 7 deletions(-)
>
> diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c
> index edaaebfbcd4693ef6689f3f4c71733b3888c7c2c..e0ad7d3825b5945049c099171ce36e5c8bb9ba99 100644
> --- a/net/ipv4/tcp_input.c
> +++ b/net/ipv4/tcp_input.c
> @@ -4602,13 +4602,12 @@ static void tcp_data_queue_ofo(struct sock *sk, struct sk_buff *skb)
> }
> }
>
> -static int __must_check tcp_queue_rcv(struct sock *sk, struct sk_buff *skb, int hdrlen,
> - bool *fragstolen)
> +static int __must_check tcp_queue_rcv(struct sock *sk, struct sk_buff *skb,
> + bool *fragstolen)
> {
> int eaten;
> struct sk_buff *tail = skb_peek_tail(&sk->sk_receive_queue);
>
> - __skb_pull(skb, hdrlen);
> eaten = (tail &&
> tcp_try_coalesce(sk, tail,
> skb, fragstolen)) ? 1 : 0;
> @@ -4659,7 +4658,7 @@ int tcp_send_rcvq(struct sock *sk, struct msghdr *msg, size_t size)
> TCP_SKB_CB(skb)->end_seq = TCP_SKB_CB(skb)->seq + size;
> TCP_SKB_CB(skb)->ack_seq = tcp_sk(sk)->snd_una - 1;
>
> - if (tcp_queue_rcv(sk, skb, 0, &fragstolen)) {
> + if (tcp_queue_rcv(sk, skb, &fragstolen)) {
> WARN_ON_ONCE(fragstolen); /* should not happen */
> __kfree_skb(skb);
> }
> @@ -4719,7 +4718,7 @@ static void tcp_data_queue(struct sock *sk, struct sk_buff *skb)
> goto drop;
> }
>
> - eaten = tcp_queue_rcv(sk, skb, 0, &fragstolen);
> + eaten = tcp_queue_rcv(sk, skb, &fragstolen);
> if (skb->len)
> tcp_event_data_recv(sk, skb);
> if (TCP_SKB_CB(skb)->tcp_flags & TCPHDR_FIN)
> @@ -5585,8 +5584,8 @@ void tcp_rcv_established(struct sock *sk, struct sk_buff *skb)
> NET_INC_STATS(sock_net(sk), LINUX_MIB_TCPHPHITS);
>
> /* Bulk data transfer: receiver */
> - eaten = tcp_queue_rcv(sk, skb, tcp_header_len,
> - &fragstolen);
> + __skb_pull(skb, tcp_header_len);
> + eaten = tcp_queue_rcv(sk, skb, &fragstolen);
>
> tcp_event_data_recv(sk, skb);
>
> --
> 2.19.1.1215.g8438c0b245-goog
>
^ permalink raw reply
* Re: [PATCH net-next 3/3] tcp: implement head drops in backlog queue
From: Yuchung Cheng @ 2018-11-21 22:40 UTC (permalink / raw)
To: Eric Dumazet
Cc: David S . Miller, netdev, Jean-Louis Dupond, Neal Cardwell,
Eric Dumazet
In-Reply-To: <20181121175240.6075-4-edumazet@google.com>
On Wed, Nov 21, 2018 at 9:52 AM, Eric Dumazet <edumazet@google.com> wrote:
> Under high stress, and if GRO or coalescing does not help,
> we better make room in backlog queue to be able to keep latest
> packet coming.
>
> This generally helps fast recovery, given that we often receive
> packets in order.
I like the benefit of fast recovery but I am a bit leery about head
drop causing HoLB on large read, while tail drops can be repaired by
RACK and TLP already. Hmm -
>
> Signed-off-by: Eric Dumazet <edumazet@google.com>
> Tested-by: Jean-Louis Dupond <jean-louis@dupond.be>
> Cc: Neal Cardwell <ncardwell@google.com>
> Cc: Yuchung Cheng <ycheng@google.com>
> ---
> net/ipv4/tcp_ipv4.c | 14 ++++++++++++++
> 1 file changed, 14 insertions(+)
>
> diff --git a/net/ipv4/tcp_ipv4.c b/net/ipv4/tcp_ipv4.c
> index 401e1d1cb904a4c7963d8baa419cfbf178593344..36c9d715bf2aa7eb7bf58b045bfeb85a2ec1a696 100644
> --- a/net/ipv4/tcp_ipv4.c
> +++ b/net/ipv4/tcp_ipv4.c
> @@ -1693,6 +1693,20 @@ bool tcp_add_backlog(struct sock *sk, struct sk_buff *skb)
> __skb_push(skb, hdrlen);
> }
>
> + while (sk_rcvqueues_full(sk, limit)) {
> + struct sk_buff *head;
> +
> + head = sk->sk_backlog.head;
> + if (!head)
> + break;
> + sk->sk_backlog.head = head->next;
> + if (!head->next)
> + sk->sk_backlog.tail = NULL;
> + skb_mark_not_on_list(head);
> + sk->sk_backlog.len -= head->truesize;
> + kfree_skb(head);
> + __NET_INC_STATS(sock_net(sk), LINUX_MIB_TCPBACKLOGDROP);
> + }
> /* Only socket owner can try to collapse/prune rx queues
> * to reduce memory overhead, so add a little headroom here.
> * Few sockets backlog are possibly concurrently non empty.
> --
> 2.19.1.1215.g8438c0b245-goog
>
^ permalink raw reply
* Re: [PATCH net-next 2/3] tcp: implement coalescing on backlog queue
From: Eric Dumazet @ 2018-11-21 22:40 UTC (permalink / raw)
To: Yuchung Cheng, Eric Dumazet
Cc: David S . Miller, netdev, Jean-Louis Dupond, Neal Cardwell
In-Reply-To: <CAK6E8=en_sPbKnb60YnPYrVBYKY1doG_V-FFgzrbAv4gnQwy5w@mail.gmail.com>
On 11/21/2018 02:31 PM, Yuchung Cheng wrote:
> On Wed, Nov 21, 2018 at 9:52 AM, Eric Dumazet <edumazet@google.com> wrote:
>> +
> Really nice! would it make sense to re-use (some of) the similar
> tcp_try_coalesce()?
>
Maybe, but it is a bit complex, since skbs in receive queues (regular or out of order)
are accounted differently (they have skb->destructor set)
Also they had the TCP header pulled already, while the backlog coalescing also has
to make sure TCP options match.
Not sure if we want to add extra parameters and conditional checks...
^ permalink raw reply
* Re: [PATCH v5 bpf-next 0/2] bpf: adding support for mapinmap in libbpf
From: Daniel Borkmann @ 2018-11-21 22:38 UTC (permalink / raw)
To: Nikita V. Shirokov, Alexei Starovoitov, Jakub Kicinski; +Cc: netdev
In-Reply-To: <20181121045557.17924-1-tehnerd@tehnerd.com>
On 11/21/2018 05:55 AM, Nikita V. Shirokov wrote:
> in this patch series i'm adding a helper for libbpf which would allow
> it to load map-in-map(BPF_MAP_TYPE_ARRAY_OF_MAPS and
> BPF_MAP_TYPE_HASH_OF_MAPS).
> first patch contains new helper + explains proposed workflow
> second patch contains tests which also could be used as example of usage
>
> v4->v5:
> - naming: renamed everything to map_in_map instead of mapinmap
> - start to return nonzero val if set_inner_map_fd failed
>
> v3->v4:
> - renamed helper to set_inner_map_fd
> - now we set this value only if it haven't
> been set before and only for (array|hash) of maps
>
> v2->v3:
> - fixing typo in patch description
> - initializing inner_map_fd to -1 by default
>
> v1->v2:
> - addressing nits
> - removing const identifier from fd in new helper
> - starting to check return val for bpf_map_update_elem
>
> Nikita V. Shirokov (2):
> bpf: adding support for map in map in libbpf
> bpf: adding tests for mapinmap helpber in libbpf
>
> tools/lib/bpf/libbpf.c | 40 ++++++++++--
> tools/lib/bpf/libbpf.h | 2 +
> tools/testing/selftests/bpf/Makefile | 3 +-
> tools/testing/selftests/bpf/test_map_in_map.c | 49 +++++++++++++++
> tools/testing/selftests/bpf/test_maps.c | 90 +++++++++++++++++++++++++++
> 5 files changed, 177 insertions(+), 7 deletions(-)
> create mode 100644 tools/testing/selftests/bpf/test_map_in_map.c
>
Applied, thanks!
^ permalink raw reply
* Re: [PATCH net-next 2/3] tcp: implement coalescing on backlog queue
From: Yuchung Cheng @ 2018-11-21 22:31 UTC (permalink / raw)
To: Eric Dumazet
Cc: David S . Miller, netdev, Jean-Louis Dupond, Neal Cardwell,
Eric Dumazet
In-Reply-To: <20181121175240.6075-3-edumazet@google.com>
On Wed, Nov 21, 2018 at 9:52 AM, Eric Dumazet <edumazet@google.com> wrote:
>
> In case GRO is not as efficient as it should be or disabled,
> we might have a user thread trapped in __release_sock() while
> softirq handler flood packets up to the point we have to drop.
>
> This patch balances work done from user thread and softirq,
> to give more chances to __release_sock() to complete its work.
>
> This also helps if we receive many ACK packets, since GRO
> does not aggregate them.
>
> Signed-off-by: Eric Dumazet <edumazet@google.com>
> Tested-by: Jean-Louis Dupond <jean-louis@dupond.be>
> Cc: Neal Cardwell <ncardwell@google.com>
> Cc: Yuchung Cheng <ycheng@google.com>
> ---
> include/uapi/linux/snmp.h | 1 +
> net/ipv4/proc.c | 1 +
> net/ipv4/tcp_ipv4.c | 75 +++++++++++++++++++++++++++++++++++----
> 3 files changed, 71 insertions(+), 6 deletions(-)
>
> diff --git a/include/uapi/linux/snmp.h b/include/uapi/linux/snmp.h
> index f80135e5feaa886000009db6dff75b2bc2d637b2..86dc24a96c90ab047d5173d625450facd6c6dd79 100644
> --- a/include/uapi/linux/snmp.h
> +++ b/include/uapi/linux/snmp.h
> @@ -243,6 +243,7 @@ enum
> LINUX_MIB_TCPREQQFULLDROP, /* TCPReqQFullDrop */
> LINUX_MIB_TCPRETRANSFAIL, /* TCPRetransFail */
> LINUX_MIB_TCPRCVCOALESCE, /* TCPRcvCoalesce */
> + LINUX_MIB_TCPBACKLOGCOALESCE, /* TCPBacklogCoalesce */
> LINUX_MIB_TCPOFOQUEUE, /* TCPOFOQueue */
> LINUX_MIB_TCPOFODROP, /* TCPOFODrop */
> LINUX_MIB_TCPOFOMERGE, /* TCPOFOMerge */
> diff --git a/net/ipv4/proc.c b/net/ipv4/proc.c
> index 70289682a6701438aed99a00a9705c39fa4394d3..c3610b37bb4ce665b1976d8cc907b6dd0de42ab9 100644
> --- a/net/ipv4/proc.c
> +++ b/net/ipv4/proc.c
> @@ -219,6 +219,7 @@ static const struct snmp_mib snmp4_net_list[] = {
> SNMP_MIB_ITEM("TCPRenoRecoveryFail", LINUX_MIB_TCPRENORECOVERYFAIL),
> SNMP_MIB_ITEM("TCPSackRecoveryFail", LINUX_MIB_TCPSACKRECOVERYFAIL),
> SNMP_MIB_ITEM("TCPRcvCollapsed", LINUX_MIB_TCPRCVCOLLAPSED),
> + SNMP_MIB_ITEM("TCPBacklogCoalesce", LINUX_MIB_TCPBACKLOGCOALESCE),
> SNMP_MIB_ITEM("TCPDSACKOldSent", LINUX_MIB_TCPDSACKOLDSENT),
> SNMP_MIB_ITEM("TCPDSACKOfoSent", LINUX_MIB_TCPDSACKOFOSENT),
> SNMP_MIB_ITEM("TCPDSACKRecv", LINUX_MIB_TCPDSACKRECV),
> diff --git a/net/ipv4/tcp_ipv4.c b/net/ipv4/tcp_ipv4.c
> index 795605a2327504b8a025405826e7e0ca8dc8501d..401e1d1cb904a4c7963d8baa419cfbf178593344 100644
> --- a/net/ipv4/tcp_ipv4.c
> +++ b/net/ipv4/tcp_ipv4.c
> @@ -1619,12 +1619,10 @@ int tcp_v4_early_demux(struct sk_buff *skb)
> bool tcp_add_backlog(struct sock *sk, struct sk_buff *skb)
> {
> u32 limit = sk->sk_rcvbuf + sk->sk_sndbuf;
> -
> - /* Only socket owner can try to collapse/prune rx queues
> - * to reduce memory overhead, so add a little headroom here.
> - * Few sockets backlog are possibly concurrently non empty.
> - */
> - limit += 64*1024;
> + struct skb_shared_info *shinfo;
> + const struct tcphdr *th;
> + struct sk_buff *tail;
> + unsigned int hdrlen;
>
> /* In case all data was pulled from skb frags (in __pskb_pull_tail()),
> * we can fix skb->truesize to its real value to avoid future drops.
> @@ -1636,6 +1634,71 @@ bool tcp_add_backlog(struct sock *sk, struct sk_buff *skb)
>
> skb_dst_drop(skb);
>
> + if (unlikely(tcp_checksum_complete(skb))) {
> + bh_unlock_sock(sk);
> + __TCP_INC_STATS(sock_net(sk), TCP_MIB_CSUMERRORS);
> + __TCP_INC_STATS(sock_net(sk), TCP_MIB_INERRS);
> + return true;
> + }
> +
> + /* Attempt coalescing to last skb in backlog, even if we are
> + * above the limits.
> + * This is okay because skb capacity is limited to MAX_SKB_FRAGS.
> + */
> + th = (const struct tcphdr *)skb->data;
> + hdrlen = th->doff * 4;
> + shinfo = skb_shinfo(skb);
> +
> + if (!shinfo->gso_size)
> + shinfo->gso_size = skb->len - hdrlen;
> +
> + if (!shinfo->gso_segs)
> + shinfo->gso_segs = 1;
> +
> + tail = sk->sk_backlog.tail;
> + if (tail &&
> + TCP_SKB_CB(tail)->end_seq == TCP_SKB_CB(skb)->seq &&
> +#ifdef CONFIG_TLS_DEVICE
> + tail->decrypted == skb->decrypted &&
> +#endif
> + !memcmp(tail->data + sizeof(*th), skb->data + sizeof(*th),
> + hdrlen - sizeof(*th))) {
> + bool fragstolen;
> + int delta;
> +
> + __skb_pull(skb, hdrlen);
> + if (skb_try_coalesce(tail, skb, &fragstolen, &delta)) {
> + TCP_SKB_CB(tail)->end_seq = TCP_SKB_CB(skb)->end_seq;
> + TCP_SKB_CB(tail)->ack_seq = TCP_SKB_CB(skb)->ack_seq;
> + TCP_SKB_CB(tail)->tcp_flags |= TCP_SKB_CB(skb)->tcp_flags;
> +
> + if (TCP_SKB_CB(skb)->has_rxtstamp) {
> + TCP_SKB_CB(tail)->has_rxtstamp = true;
> + tail->tstamp = skb->tstamp;
> + skb_hwtstamps(tail)->hwtstamp = skb_hwtstamps(skb)->hwtstamp;
> + }
> +
Really nice! would it make sense to re-use (some of) the similar
tcp_try_coalesce()?
> + /* Not as strict as GRO. We only need to carry mss max value */
> + skb_shinfo(tail)->gso_size = max(shinfo->gso_size,
> + skb_shinfo(tail)->gso_size);
> +
> + skb_shinfo(tail)->gso_segs += shinfo->gso_segs;
> +
> + sk->sk_backlog.len += delta;
> + __NET_INC_STATS(sock_net(sk),
> + LINUX_MIB_TCPBACKLOGCOALESCE);
> + kfree_skb_partial(skb, fragstolen);
> + return false;
> + }
> + __skb_push(skb, hdrlen);
> + }
> +
> + /* Only socket owner can try to collapse/prune rx queues
> + * to reduce memory overhead, so add a little headroom here.
> + * Few sockets backlog are possibly concurrently non empty.
> + */
> + limit += 64*1024;
> +
> if (unlikely(sk_add_backlog(sk, skb, limit))) {
> bh_unlock_sock(sk);
> __NET_INC_STATS(sock_net(sk), LINUX_MIB_TCPBACKLOGDROP);
> --
> 2.19.1.1215.g8438c0b245-goog
>
^ permalink raw reply
* Re: [PATCH bpf-next v3 1/3] bpf, libbpf: introduce bpf_object__probe_caps to test BPF capabilities
From: Daniel Borkmann @ 2018-11-21 22:30 UTC (permalink / raw)
To: Stanislav Fomichev, netdev, ast
In-Reply-To: <20181121011121.159355-1-sdf@google.com>
On 11/21/2018 02:11 AM, Stanislav Fomichev wrote:
> It currently only checks whether kernel supports map/prog names.
> This capability check will be used in the next two commits to skip setting
> prog/map names.
>
> Suggested-by: Daniel Borkmann <daniel@iogearbox.net>
> Signed-off-by: Stanislav Fomichev <sdf@google.com>
Looks great, thanks for following through. Applied, thanks!
^ permalink raw reply
page: next (older) | prev (newer) | latest
- recent:[subjects (threaded)|topics (new)|topics (active)]
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox