* Re: linux-next: manual merge of the net-next tree with the rdma tree
From: Jason Gunthorpe @ 2018-07-27 2:48 UTC (permalink / raw)
To: Stephen Rothwell
Cc: David Miller, Networking, Doug Ledford, Linux-Next Mailing List,
Linux Kernel Mailing List, Parav Pandit, Ursula Braun,
Leon Romanovsky, linux-rdma
In-Reply-To: <20180727123301.3ac97ddc@canb.auug.org.au>
On Fri, Jul 27, 2018 at 12:33:01PM +1000, Stephen Rothwell wrote:
> I fixed it up (I wasn't sure how to fix this up as so much has changed
> in the net-next tree and both modified functions had been (re)moved,
> so I effectively reverted the rdma tree commit) and can carry the fix
> as necessary. Please come to some arrangement about this.
How does that still compile? We removed ib_query_gid() from the rdma
tree and replaced it with rdma_get_gid_attr()..
I think the merge resolution is going to be a bit nasty to absorb
that much changing..
Perhaps we should add a compatability ib_query_gid back to the RDMA
tree and then send DaveM a commit to fix SMC and remove it during the
next cycle? Linus can resolve smc_ib.c by using the net version
Does someone else have a better idea?
Thanks,
Jason
^ permalink raw reply
* Re: [PATCH V3 bpf] xdp: add NULL pointer check in __xdp_return()
From: Daniel Borkmann @ 2018-07-27 1:44 UTC (permalink / raw)
To: Taehee Yoo, ast; +Cc: netdev, bjorn.topel, brouer, kafai, jakub.kicinski
In-Reply-To: <20180726141703.6236-1-ap420073@gmail.com>
On 07/26/2018 04:17 PM, Taehee Yoo wrote:
> rhashtable_lookup() can return NULL. so that NULL pointer
> check routine should be added.
>
> Fixes: 02b55e5657c3 ("xdp: add MEM_TYPE_ZERO_COPY")
> Acked-by: Martin KaFai Lau <kafai@fb.com>
> Signed-off-by: Taehee Yoo <ap420073@gmail.com>
Applied to bpf, thanks Taehee!
^ permalink raw reply
* Re: [PATCH bpf] bpf: btf: Use exact btf value_size match in map_check_btf()
From: Daniel Borkmann @ 2018-07-27 1:46 UTC (permalink / raw)
To: Martin KaFai Lau, netdev; +Cc: Alexei Starovoitov, kernel-team
In-Reply-To: <20180726165759.3788287-1-kafai@fb.com>
On 07/26/2018 06:57 PM, Martin KaFai Lau wrote:
> The current map_check_btf() in BPF_MAP_TYPE_ARRAY rejects
> '> map->value_size' to ensure map_seq_show_elem() will not
> access things beyond an array element.
>
> Yonghong suggested that using '!=' is a more correct
> check. The 8 bytes round_up on value_size is stored
> in array->elem_size. Hence, using '!=' on map->value_size
> is a proper check.
>
> This patch also adds new tests to check the btf array
> key type and value type. Two of these new tests verify
> the btf's value_size (the change in this patch).
>
> It also fixes two existing tests that wrongly encoded
> a btf's type size (pprint_test) and the value_type_id (in one
> of the raw_tests[]). However, that do not affect these two
> BTF verification tests before or after this test changes.
> These two tests mainly failed at array creation time after
> this patch.
>
> Fixes: a26ca7c982cb ("bpf: btf: Add pretty print support to the basic arraymap")
> Suggested-by: Yonghong Song <yhs@fb.com>
> Acked-by: Yonghong Song <yhs@fb.com>
> Signed-off-by: Martin KaFai Lau <kafai@fb.com>
Applied to bpf, thanks Martin!
^ permalink raw reply
* Re: linux-next: manual merge of the net-next tree with the rdma tree
From: Stephen Rothwell @ 2018-07-27 3:28 UTC (permalink / raw)
To: Jason Gunthorpe
Cc: David Miller, Networking, Doug Ledford, Linux-Next Mailing List,
Linux Kernel Mailing List, Parav Pandit, Ursula Braun,
Leon Romanovsky, linux-rdma
In-Reply-To: <20180727024832.GA30023@mellanox.com>
[-- Attachment #1: Type: text/plain, Size: 3940 bytes --]
Hi Jason,
On Thu, 26 Jul 2018 20:48:32 -0600 Jason Gunthorpe <jgg@mellanox.com> wrote:
>
> On Fri, Jul 27, 2018 at 12:33:01PM +1000, Stephen Rothwell wrote:
>
> > I fixed it up (I wasn't sure how to fix this up as so much has changed
> > in the net-next tree and both modified functions had been (re)moved,
> > so I effectively reverted the rdma tree commit) and can carry the fix
> > as necessary. Please come to some arrangement about this.
>
> How does that still compile? We removed ib_query_gid() from the rdma
> tree and replaced it with rdma_get_gid_attr()..
Yeah, it doesn't :-(
> I think the merge resolution is going to be a bit nasty to absorb
> that much changing..
>
> Perhaps we should add a compatability ib_query_gid back to the RDMA
> tree and then send DaveM a commit to fix SMC and remove it during the
> next cycle? Linus can resolve smc_ib.c by using the net version
>
> Does someone else have a better idea?
I applied this merge fix patch:
From: Stephen Rothwell <sfr@canb.auug.org.au>
Date: Fri, 27 Jul 2018 13:19:31 +1000
Subject: [PATCH] net/smc: fixups for ip_query_gid API removal
Signed-off-by: Stephen Rothwell <sfr@canb.auug.org.au>
---
net/smc/smc_ib.c | 47 +++++++++++++++++++++++++----------------------
1 file changed, 25 insertions(+), 22 deletions(-)
diff --git a/net/smc/smc_ib.c b/net/smc/smc_ib.c
index 2cc64bc8ae20..debc6e44f738 100644
--- a/net/smc/smc_ib.c
+++ b/net/smc/smc_ib.c
@@ -16,6 +16,7 @@
#include <linux/workqueue.h>
#include <linux/scatterlist.h>
#include <rdma/ib_verbs.h>
+#include <rdma/ib_cache.h>
#include "smc_pnet.h"
#include "smc_ib.h"
@@ -144,17 +145,21 @@ int smc_ib_ready_link(struct smc_link *lnk)
static int smc_ib_fill_mac(struct smc_ib_device *smcibdev, u8 ibport)
{
- struct ib_gid_attr gattr;
- union ib_gid gid;
- int rc;
+ const struct ib_gid_attr *gattr;
+ int rc = 0;
- rc = ib_query_gid(smcibdev->ibdev, ibport, 0, &gid, &gattr);
- if (rc || !gattr.ndev)
- return -ENODEV;
+ gattr = rdma_get_gid_attr(smcibdev->ibdev, ibport, 0);
+ if (IS_ERR(gattr))
+ return PTR_ERR(gattr);
+ if (!gattr->ndev) {
+ rc = -ENODEV;
+ goto done;
+ }
- memcpy(smcibdev->mac[ibport - 1], gattr.ndev->dev_addr, ETH_ALEN);
- dev_put(gattr.ndev);
- return 0;
+ memcpy(smcibdev->mac[ibport - 1], gattr->ndev->dev_addr, ETH_ALEN);
+done:
+ rdma_put_gid_attr(gattr);
+ return rc;
}
/* Create an identifier unique for this instance of SMC-R.
@@ -179,29 +184,27 @@ bool smc_ib_port_active(struct smc_ib_device *smcibdev, u8 ibport)
int smc_ib_determine_gid(struct smc_ib_device *smcibdev, u8 ibport,
unsigned short vlan_id, u8 gid[], u8 *sgid_index)
{
- struct ib_gid_attr gattr;
- union ib_gid _gid;
+ const struct ib_gid_attr *gattr;
int i;
for (i = 0; i < smcibdev->pattr[ibport - 1].gid_tbl_len; i++) {
- memset(&_gid, 0, SMC_GID_SIZE);
- memset(&gattr, 0, sizeof(gattr));
- if (ib_query_gid(smcibdev->ibdev, ibport, i, &_gid, &gattr))
+ gattr = rdma_get_gid_attr(smcibdev->ibdev, ibport, i);
+ if (IS_ERR(gattr))
continue;
- if (!gattr.ndev)
+ if (!gattr->ndev)
continue;
- if (((!vlan_id && !is_vlan_dev(gattr.ndev)) ||
- (vlan_id && is_vlan_dev(gattr.ndev) &&
- vlan_dev_vlan_id(gattr.ndev) == vlan_id)) &&
- gattr.gid_type == IB_GID_TYPE_IB) {
+ if (((!vlan_id && !is_vlan_dev(gattr->ndev)) ||
+ (vlan_id && is_vlan_dev(gattr->ndev) &&
+ vlan_dev_vlan_id(gattr->ndev) == vlan_id)) &&
+ gattr->gid_type == IB_GID_TYPE_IB) {
if (gid)
- memcpy(gid, &_gid, SMC_GID_SIZE);
+ memcpy(gid, &gattr->gid, SMC_GID_SIZE);
if (sgid_index)
*sgid_index = i;
- dev_put(gattr.ndev);
+ rdma_put_gid_attr(gattr);
return 0;
}
- dev_put(gattr.ndev);
+ rdma_put_gid_attr(gattr);
}
return -ENODEV;
}
--
2.18.0
--
Cheers,
Stephen Rothwell
[-- Attachment #2: OpenPGP digital signature --]
[-- Type: application/pgp-signature, Size: 488 bytes --]
^ permalink raw reply related
* Re: linux-next: manual merge of the net-next tree with the rdma tree
From: Stephen Rothwell @ 2018-07-27 3:45 UTC (permalink / raw)
To: Jason Gunthorpe
Cc: David Miller, Networking, Doug Ledford, Linux-Next Mailing List,
Linux Kernel Mailing List, Parav Pandit, Ursula Braun,
Leon Romanovsky, linux-rdma
In-Reply-To: <20180727132847.2be8563b@canb.auug.org.au>
[-- Attachment #1: Type: text/plain, Size: 4213 bytes --]
Hi all,
On Fri, 27 Jul 2018 13:28:47 +1000 Stephen Rothwell <sfr@canb.auug.org.au> wrote:
>
> I applied this merge fix patch:
The final conflict resolution actually looks like this:
(the rdma tree changes to net/smc/smc_core.c are dropped)
c1d4bb2af93573ee4a21538a1a97b568a2344499
diff --cc net/smc/smc_ib.c
index 74f29f814ec1,2cc64bc8ae20..debc6e44f738
--- a/net/smc/smc_ib.c
+++ b/net/smc/smc_ib.c
@@@ -144,6 -142,93 +143,95 @@@ out
return rc;
}
+ static int smc_ib_fill_mac(struct smc_ib_device *smcibdev, u8 ibport)
+ {
- struct ib_gid_attr gattr;
- union ib_gid gid;
- int rc;
++ const struct ib_gid_attr *gattr;
++ int rc = 0;
+
- rc = ib_query_gid(smcibdev->ibdev, ibport, 0, &gid, &gattr);
- if (rc || !gattr.ndev)
- return -ENODEV;
++ gattr = rdma_get_gid_attr(smcibdev->ibdev, ibport, 0);
++ if (IS_ERR(gattr))
++ return PTR_ERR(gattr);
++ if (!gattr->ndev) {
++ rc = -ENODEV;
++ goto done;
++ }
+
- memcpy(smcibdev->mac[ibport - 1], gattr.ndev->dev_addr, ETH_ALEN);
- dev_put(gattr.ndev);
- return 0;
++ memcpy(smcibdev->mac[ibport - 1], gattr->ndev->dev_addr, ETH_ALEN);
++done:
++ rdma_put_gid_attr(gattr);
++ return rc;
+ }
+
+ /* Create an identifier unique for this instance of SMC-R.
+ * The MAC-address of the first active registered IB device
+ * plus a random 2-byte number is used to create this identifier.
+ * This name is delivered to the peer during connection initialization.
+ */
+ static inline void smc_ib_define_local_systemid(struct smc_ib_device *smcibdev,
+ u8 ibport)
+ {
+ memcpy(&local_systemid[2], &smcibdev->mac[ibport - 1],
+ sizeof(smcibdev->mac[ibport - 1]));
+ get_random_bytes(&local_systemid[0], 2);
+ }
+
+ bool smc_ib_port_active(struct smc_ib_device *smcibdev, u8 ibport)
+ {
+ return smcibdev->pattr[ibport - 1].state == IB_PORT_ACTIVE;
+ }
+
+ /* determine the gid for an ib-device port and vlan id */
+ int smc_ib_determine_gid(struct smc_ib_device *smcibdev, u8 ibport,
+ unsigned short vlan_id, u8 gid[], u8 *sgid_index)
+ {
- struct ib_gid_attr gattr;
- union ib_gid _gid;
++ const struct ib_gid_attr *gattr;
+ int i;
+
+ for (i = 0; i < smcibdev->pattr[ibport - 1].gid_tbl_len; i++) {
- memset(&_gid, 0, SMC_GID_SIZE);
- memset(&gattr, 0, sizeof(gattr));
- if (ib_query_gid(smcibdev->ibdev, ibport, i, &_gid, &gattr))
++ gattr = rdma_get_gid_attr(smcibdev->ibdev, ibport, i);
++ if (IS_ERR(gattr))
+ continue;
- if (!gattr.ndev)
++ if (!gattr->ndev)
+ continue;
- if (((!vlan_id && !is_vlan_dev(gattr.ndev)) ||
- (vlan_id && is_vlan_dev(gattr.ndev) &&
- vlan_dev_vlan_id(gattr.ndev) == vlan_id)) &&
- gattr.gid_type == IB_GID_TYPE_IB) {
++ if (((!vlan_id && !is_vlan_dev(gattr->ndev)) ||
++ (vlan_id && is_vlan_dev(gattr->ndev) &&
++ vlan_dev_vlan_id(gattr->ndev) == vlan_id)) &&
++ gattr->gid_type == IB_GID_TYPE_IB) {
+ if (gid)
- memcpy(gid, &_gid, SMC_GID_SIZE);
++ memcpy(gid, &gattr->gid, SMC_GID_SIZE);
+ if (sgid_index)
+ *sgid_index = i;
- dev_put(gattr.ndev);
++ rdma_put_gid_attr(gattr);
+ return 0;
+ }
- dev_put(gattr.ndev);
++ rdma_put_gid_attr(gattr);
+ }
+ return -ENODEV;
+ }
+
+ static int smc_ib_remember_port_attr(struct smc_ib_device *smcibdev, u8 ibport)
+ {
+ int rc;
+
+ memset(&smcibdev->pattr[ibport - 1], 0,
+ sizeof(smcibdev->pattr[ibport - 1]));
+ rc = ib_query_port(smcibdev->ibdev, ibport,
+ &smcibdev->pattr[ibport - 1]);
+ if (rc)
+ goto out;
+ /* the SMC protocol requires specification of the RoCE MAC address */
+ rc = smc_ib_fill_mac(smcibdev, ibport);
+ if (rc)
+ goto out;
+ if (!strncmp(local_systemid, SMC_LOCAL_SYSTEMID_RESET,
+ sizeof(local_systemid)) &&
+ smc_ib_port_active(smcibdev, ibport))
+ /* create unique system identifier */
+ smc_ib_define_local_systemid(smcibdev, ibport);
+ out:
+ return rc;
+ }
+
/* process context wrapper for might_sleep smc_ib_remember_port_attr */
static void smc_ib_port_event_work(struct work_struct *work)
{
--
Cheers,
Stephen Rothwell
[-- Attachment #2: OpenPGP digital signature --]
[-- Type: application/pgp-signature, Size: 488 bytes --]
^ permalink raw reply
* [PATCH] net: adaptec: Replace mdelay() with msleep() in starfire_init_one()
From: Jia-Ju Bai @ 2018-07-27 3:51 UTC (permalink / raw)
To: ionut; +Cc: netdev, linux-kernel, Jia-Ju Bai
starfire_init_one() is never called in atomic context.
It calls mdelay() to busily wait, which is not necessary.
mdelay() can be replaced with msleep().
This is found by a static analysis tool named DCNS written by myself.
Signed-off-by: Jia-Ju Bai <baijiaju1990@gmail.com>
---
drivers/net/ethernet/adaptec/starfire.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/drivers/net/ethernet/adaptec/starfire.c b/drivers/net/ethernet/adaptec/starfire.c
index 3872ab96b80a..097467f44b0d 100644
--- a/drivers/net/ethernet/adaptec/starfire.c
+++ b/drivers/net/ethernet/adaptec/starfire.c
@@ -802,7 +802,7 @@ static int starfire_init_one(struct pci_dev *pdev,
int mii_status;
for (phy = 0; phy < 32 && phy_idx < PHY_CNT; phy++) {
mdio_write(dev, phy, MII_BMCR, BMCR_RESET);
- mdelay(100);
+ msleep(100);
boguscnt = 1000;
while (--boguscnt > 0)
if ((mdio_read(dev, phy, MII_BMCR) & BMCR_RESET) == 0)
--
2.17.0
^ permalink raw reply related
* [PATCH net-next] sch_cake: Make gso-splitting configurable from userspace
From: Dave Taht @ 2018-07-27 2:45 UTC (permalink / raw)
To: netdev, cake; +Cc: Dave Taht
I expect the first part of this patch to generate no controversy,
as being able to enable configure gso-splitting on or off in all
use cases of cake is a goodness.
But: I expect the single line re-enabling cake's fielded default of
always splitting gro and gso packets, in shaped or unshaped mode, back
into packets, to reduce my email systems' hard disk inbox to a barren,
burnt cylinder, even if it is made easy to override thusly:
tc qdisc replace dev whatever root cake no-split-gso
While I agree that gro/gso is needed at 10gigE+ speeds, I feel
offering an option to disable splitting to those users trying to run
cake at those speeds is better than the alternative of forcing users
running at 1Gbit, 100mbit, 10mbit and below, with and without pause
frames, shaped or unshaped, to remember to split-gso.
While I have assembled tons of data in use cases ranging from nearly 0
to a gbit, the first, and most compelling argument I can make is
made in the commit that follows, where allowing GSO/GRO superpackets
triples the size of the underlying BQL when running at a gbit.
Dave Taht (1):
sch_cake: Make gso-splitting configurable
net/sched/sch_cake.c | 13 +++++++------
1 file changed, 7 insertions(+), 6 deletions(-)
--
2.7.4
^ permalink raw reply
* [PATCH net-next] sch_cake: Make gso-splitting configurable from userspace
From: Dave Taht @ 2018-07-27 2:45 UTC (permalink / raw)
To: netdev, cake; +Cc: Dave Taht
In-Reply-To: <1532659510-17385-1-git-send-email-dave.taht@gmail.com>
This patch restores cake's deployed behavior at line rate to always
split gso, and makes gso splitting configurable from userspace.
running cake unlimited (unshaped) at 1gigE, local traffic:
no-split-gso bql limit: 131966
split-gso bql limit: ~42392-45420
On this 4 stream test splitting gso apart results in halving the
observed interpacket latency at no loss in throughput.
Summary of tcp_nup test run 'gso-split' (at 2018-07-26 16:03:51.824728):
Ping (ms) ICMP : 0.83 0.81 ms 341
TCP upload avg : 235.43 235.39 Mbits/s 301
TCP upload sum : 941.71 941.56 Mbits/s 301
TCP upload::1 : 235.45 235.43 Mbits/s 271
TCP upload::2 : 235.45 235.41 Mbits/s 289
TCP upload::3 : 235.40 235.40 Mbits/s 288
TCP upload::4 : 235.41 235.40 Mbits/s 291
verses
Summary of tcp_nup test run 'no-split-gso' (at 2018-07-26 16:37:23.563960):
avg median # data pts
Ping (ms) ICMP : 1.67 1.73 ms 348
TCP upload avg : 234.56 235.37 Mbits/s 301
TCP upload sum : 938.24 941.49 Mbits/s 301
TCP upload::1 : 234.55 235.38 Mbits/s 285
TCP upload::2 : 234.57 235.37 Mbits/s 286
TCP upload::3 : 234.58 235.37 Mbits/s 274
TCP upload::4 : 234.54 235.42 Mbits/s 288
---
net/sched/sch_cake.c | 13 +++++++------
1 file changed, 7 insertions(+), 6 deletions(-)
diff --git a/net/sched/sch_cake.c b/net/sched/sch_cake.c
index 539c949..35fc725 100644
--- a/net/sched/sch_cake.c
+++ b/net/sched/sch_cake.c
@@ -80,7 +80,6 @@
#define CAKE_QUEUES (1024)
#define CAKE_FLOW_MASK 63
#define CAKE_FLOW_NAT_FLAG 64
-#define CAKE_SPLIT_GSO_THRESHOLD (125000000) /* 1Gbps */
/* struct cobalt_params - contains codel and blue parameters
* @interval: codel initial drop rate
@@ -2569,10 +2568,12 @@ static int cake_change(struct Qdisc *sch, struct nlattr *opt,
if (tb[TCA_CAKE_MEMORY])
q->buffer_config_limit = nla_get_u32(tb[TCA_CAKE_MEMORY]);
- if (q->rate_bps && q->rate_bps <= CAKE_SPLIT_GSO_THRESHOLD)
- q->rate_flags |= CAKE_FLAG_SPLIT_GSO;
- else
- q->rate_flags &= ~CAKE_FLAG_SPLIT_GSO;
+ if (tb[TCA_CAKE_SPLIT_GSO]) {
+ if (!!nla_get_u32(tb[TCA_CAKE_SPLIT_GSO]))
+ q->rate_flags |= CAKE_FLAG_SPLIT_GSO;
+ else
+ q->rate_flags &= ~CAKE_FLAG_SPLIT_GSO;
+ }
if (q->tins) {
sch_tree_lock(sch);
@@ -2608,7 +2609,7 @@ static int cake_init(struct Qdisc *sch, struct nlattr *opt,
q->target = 5000; /* 5ms: codel RFC argues
* for 5 to 10% of interval
*/
-
+ q->rate_flags |= CAKE_FLAG_SPLIT_GSO;
q->cur_tin = 0;
q->cur_flow = 0;
--
2.7.4
^ permalink raw reply related
* Re: [PATCH net-next v3 1/5] tc/act: user space can't use TC_ACT_REDIRECT directly
From: Daniel Borkmann @ 2018-07-27 2:48 UTC (permalink / raw)
To: Jiri Pirko
Cc: Paolo Abeni, Jamal Hadi Salim, netdev, Cong Wang,
Marcelo Ricardo Leitner, Eyal Birger, David S. Miller,
alexei.starovoitov
In-Reply-To: <20180726074300.GE2222@nanopsycho>
On 07/26/2018 09:43 AM, Jiri Pirko wrote:
> Wed, Jul 25, 2018 at 06:29:54PM CEST, daniel@iogearbox.net wrote:
>> On 07/25/2018 05:48 PM, Paolo Abeni wrote:
>>> On Wed, 2018-07-25 at 15:03 +0200, Jiri Pirko wrote:
>>>> Wed, Jul 25, 2018 at 02:54:04PM CEST, pabeni@redhat.com wrote:
>>>>> On Wed, 2018-07-25 at 13:56 +0200, Jiri Pirko wrote:
>>>>>> Tue, Jul 24, 2018 at 10:06:39PM CEST, pabeni@redhat.com wrote:
>>>>>>> Only cls_bpf and act_bpf can safely use such value. If a generic
>>>>>>> action is configured by user space to return TC_ACT_REDIRECT,
>>>>>>> the usually visible behavior is passing the skb up the stack - as
>>>>>>> for unknown action, but, with complex configuration, more random
>>>>>>> results can be obtained.
>>>>>>>
>>>>>>> This patch forcefully converts TC_ACT_REDIRECT to TC_ACT_UNSPEC
>>>>>>> at action init time, making the kernel behavior more consistent.
>>>>>>>
>>>>>>> v1 -> v3: use TC_ACT_UNSPEC instead of a newly definied act value
>>>>>>>
>>>>>>> Signed-off-by: Paolo Abeni <pabeni@redhat.com>
>>>>>>> ---
>>>>>>> net/sched/act_api.c | 5 +++++
>>>>>>> 1 file changed, 5 insertions(+)
>>>>>>>
>>>>>>> diff --git a/net/sched/act_api.c b/net/sched/act_api.c
>>>>>>> index 148a89ab789b..24b5534967fe 100644
>>>>>>> --- a/net/sched/act_api.c
>>>>>>> +++ b/net/sched/act_api.c
>>>>>>> @@ -895,6 +895,11 @@ struct tc_action *tcf_action_init_1(struct net *net, struct tcf_proto *tp,
>>>>>>> }
>>>>>>> }
>>>>>>>
>>>>>>> + if (a->tcfa_action == TC_ACT_REDIRECT) {
>>>>>>> + net_warn_ratelimited("TC_ACT_REDIRECT can't be used directly");
>>>>>>
>>>>>> Can't you push this warning through extack?
>>>>>>
>>>>>> But, wouldn't it be more appropriate to fail here? User is passing
>>>>>> invalid configuration....
>>>>>
>>>>> Jiri, Jamal, thank you for the feedback.
>>>>>
>>>>> Please allow me to answer both of you here, since you raised similar
>>>>> concers.
>>>>>
>>>>> I thought about rejecting the action, but that change of behavior could
>>>>> break some users, as currently most kind of invalid tcfa_action values
>>>>> are simply accepted.
>>>>>
>>>>> If there is consensus about it, I can simply fail.
>>>>
>>>> Well it was obviously wrong to expose TC_ACT_REDIRECT to uapi and it
>>>> really has no meaning for anyone to use it throughout its whole history.
>>
>> That claim is completely wrong.
>
> Why? Does addition of TC_ACT_REDIRECT to uapi have any meaning?
BPF programs return TC_ACT_* as a verdict which is then further processed,
including TC_ACT_REDIRECT. Hence, it's a contract for them similarly as e.g.
helper enums (BPF_FUNC_*) and other things they make use of.
Cheers,
Daniel
^ permalink raw reply
* Re: [PATCH net-next] net: hns: make hns_dsaf_roce_reset non static
From: David Miller @ 2018-07-27 4:19 UTC (permalink / raw)
To: yuehaibing
Cc: yisen.zhuang, salil.mehta, linux-kernel, netdev, linyunsheng,
matthias.bgg, lipeng321
In-Reply-To: <20180727015312.19256-1-yuehaibing@huawei.com>
From: YueHaibing <yuehaibing@huawei.com>
Date: Fri, 27 Jul 2018 09:53:12 +0800
> hns_dsaf_roce_reset is exported and used in hns_roce_hw_v1.c
> In commit 336a443bd9dd ("net: hns: Make many functions static") I make
> it static wrongly.
>
> drivers/infiniband/hw/hns/hns_roce_hw_v1.o: In function `hns_roce_v1_reset':
> hns_roce_hw_v1.c:(.text+0x37ac): undefined reference to `hns_dsaf_roce_reset'
> hns_roce_hw_v1.c:(.text+0x37cc): undefined reference to `hns_dsaf_roce_reset'
>
> Fixes: 336a443bd9dd ("net: hns: Make many functions static")
> Signed-off-by: YueHaibing <yuehaibing@huawei.com>
Oops, applied.
^ permalink raw reply
* Re: [PATCH v3 bpf-next 04/14] bpf: allocate cgroup storage entries on attaching bpf programs
From: Daniel Borkmann @ 2018-07-27 4:21 UTC (permalink / raw)
To: Roman Gushchin, netdev; +Cc: linux-kernel, kernel-team, Alexei Starovoitov
In-Reply-To: <20180720174558.5829-5-guro@fb.com>
On 07/20/2018 07:45 PM, Roman Gushchin wrote:
> If a bpf program is using cgroup local storage, allocate
> a bpf_cgroup_storage structure automatically on attaching the program
> to a cgroup and save the pointer into the corresponding bpf_prog_list
> entry.
> Analogically, release the cgroup local storage on detaching
> of the bpf program.
>
> Signed-off-by: Roman Gushchin <guro@fb.com>
> Cc: Alexei Starovoitov <ast@kernel.org>
> Cc: Daniel Borkmann <daniel@iogearbox.net>
> Acked-by: Martin KaFai Lau <kafai@fb.com>
> ---
> include/linux/bpf-cgroup.h | 1 +
> kernel/bpf/cgroup.c | 28 ++++++++++++++++++++++++++--
> 2 files changed, 27 insertions(+), 2 deletions(-)
>
> diff --git a/include/linux/bpf-cgroup.h b/include/linux/bpf-cgroup.h
> index 1b1b4e94d77d..f37347331fdb 100644
> --- a/include/linux/bpf-cgroup.h
> +++ b/include/linux/bpf-cgroup.h
> @@ -42,6 +42,7 @@ struct bpf_cgroup_storage {
> struct bpf_prog_list {
> struct list_head node;
> struct bpf_prog *prog;
> + struct bpf_cgroup_storage *storage;
> };
>
> struct bpf_prog_array;
> diff --git a/kernel/bpf/cgroup.c b/kernel/bpf/cgroup.c
> index badabb0b435c..986ff18ef92e 100644
> --- a/kernel/bpf/cgroup.c
> +++ b/kernel/bpf/cgroup.c
> @@ -34,6 +34,8 @@ void cgroup_bpf_put(struct cgroup *cgrp)
> list_for_each_entry_safe(pl, tmp, progs, node) {
> list_del(&pl->node);
> bpf_prog_put(pl->prog);
> + bpf_cgroup_storage_unlink(pl->storage);
> + bpf_cgroup_storage_free(pl->storage);
> kfree(pl);
> static_branch_dec(&cgroup_bpf_enabled_key);
> }
> @@ -188,6 +190,7 @@ int __cgroup_bpf_attach(struct cgroup *cgrp, struct bpf_prog *prog,
> {
> struct list_head *progs = &cgrp->bpf.progs[type];
> struct bpf_prog *old_prog = NULL;
> + struct bpf_cgroup_storage *storage, *old_storage = NULL;
> struct cgroup_subsys_state *css;
> struct bpf_prog_list *pl;
> bool pl_was_allocated;
> @@ -210,6 +213,10 @@ int __cgroup_bpf_attach(struct cgroup *cgrp, struct bpf_prog *prog,
> if (prog_list_length(progs) >= BPF_CGROUP_MAX_PROGS)
> return -E2BIG;
>
> + storage = bpf_cgroup_storage_alloc(prog);
> + if (IS_ERR(storage))
> + return -ENOMEM;
> +
> if (flags & BPF_F_ALLOW_MULTI) {
> list_for_each_entry(pl, progs, node)
> if (pl->prog == prog)
> @@ -217,24 +224,33 @@ int __cgroup_bpf_attach(struct cgroup *cgrp, struct bpf_prog *prog,
> return -EINVAL;
>
> pl = kmalloc(sizeof(*pl), GFP_KERNEL);
> - if (!pl)
> + if (!pl) {
> + bpf_cgroup_storage_free(storage);
> return -ENOMEM;
> + }
Above code is:
storage = bpf_cgroup_storage_alloc(prog);
if (IS_ERR(storage))
return -ENOMEM;
if (flags & BPF_F_ALLOW_MULTI) {
list_for_each_entry(pl, progs, node)
if (pl->prog == prog)
/* disallow attaching the same prog twice */
return -EINVAL;
pl = kmalloc(sizeof(*pl), GFP_KERNEL);
if (!pl) {
bpf_cgroup_storage_free(storage);
return -ENOMEM;
}
Given bpf_cgroup_storage_alloc() only changes the mem and prepares (kmallocs)
the local storage buffer, we would leak it above where we attach the same
prog twice, no?
> +
> pl_was_allocated = true;
> pl->prog = prog;
> + pl->storage = storage;
> list_add_tail(&pl->node, progs);
> } else {
> if (list_empty(progs)) {
> pl = kmalloc(sizeof(*pl), GFP_KERNEL);
> - if (!pl)
> + if (!pl) {
> + bpf_cgroup_storage_free(storage);
> return -ENOMEM;
> + }
> pl_was_allocated = true;
> list_add_tail(&pl->node, progs);
> } else {
> pl = list_first_entry(progs, typeof(*pl), node);
> old_prog = pl->prog;
> + old_storage = pl->storage;
> + bpf_cgroup_storage_unlink(old_storage);
> pl_was_allocated = false;
> }
> pl->prog = prog;
> + pl->storage = storage;
> }
>
> cgrp->bpf.flags[type] = flags;
> @@ -257,10 +273,13 @@ int __cgroup_bpf_attach(struct cgroup *cgrp, struct bpf_prog *prog,
> }
>
> static_branch_inc(&cgroup_bpf_enabled_key);
> + if (old_storage)
> + bpf_cgroup_storage_free(old_storage);
> if (old_prog) {
> bpf_prog_put(old_prog);
> static_branch_dec(&cgroup_bpf_enabled_key);
> }
> + bpf_cgroup_storage_link(storage, cgrp, type);
> return 0;
>
> cleanup:
> @@ -276,6 +295,9 @@ int __cgroup_bpf_attach(struct cgroup *cgrp, struct bpf_prog *prog,
>
> /* and cleanup the prog list */
> pl->prog = old_prog;
> + bpf_cgroup_storage_free(pl->storage);
> + pl->storage = old_storage;
> + bpf_cgroup_storage_link(old_storage, cgrp, type);
> if (pl_was_allocated) {
> list_del(&pl->node);
> kfree(pl);
> @@ -356,6 +378,8 @@ int __cgroup_bpf_detach(struct cgroup *cgrp, struct bpf_prog *prog,
>
> /* now can actually delete it from this cgroup list */
> list_del(&pl->node);
> + bpf_cgroup_storage_unlink(pl->storage);
> + bpf_cgroup_storage_free(pl->storage);
> kfree(pl);
> if (list_empty(progs))
> /* last program was detached, reset flags to zero */
>
^ permalink raw reply
* Re: [PATCH v5 bpf-next 2/9] veth: Add driver XDP
From: John Fastabend @ 2018-07-27 3:02 UTC (permalink / raw)
To: Toshiaki Makita, netdev, Alexei Starovoitov, Daniel Borkmann
Cc: Toshiaki Makita, Jesper Dangaard Brouer, Jakub Kicinski
In-Reply-To: <20180726144032.2116-3-toshiaki.makita1@gmail.com>
On 07/26/2018 07:40 AM, Toshiaki Makita wrote:
> From: Toshiaki Makita <makita.toshiaki@lab.ntt.co.jp>
>
> This is the basic implementation of veth driver XDP.
>
> Incoming packets are sent from the peer veth device in the form of skb,
> so this is generally doing the same thing as generic XDP.
>
> This itself is not so useful, but a starting point to implement other
> useful veth XDP features like TX and REDIRECT.
>
> This introduces NAPI when XDP is enabled, because XDP is now heavily
> relies on NAPI context. Use ptr_ring to emulate NIC ring. Tx function
> enqueues packets to the ring and peer NAPI handler drains the ring.
>
> Currently only one ring is allocated for each veth device, so it does
> not scale on multiqueue env. This can be resolved by allocating rings
> on the per-queue basis later.
>
> Note that NAPI is not used but netif_rx is used when XDP is not loaded,
> so this does not change the default behaviour.
>
> v3:
> - Fix race on closing the device.
> - Add extack messages in ndo_bpf.
>
> v2:
> - Squashed with the patch adding NAPI.
> - Implement adjust_tail.
> - Don't acquire consumer lock because it is guarded by NAPI.
> - Make poll_controller noop since it is unnecessary.
> - Register rxq_info on enabling XDP rather than on opening the device.
>
> Signed-off-by: Toshiaki Makita <makita.toshiaki@lab.ntt.co.jp>
> ---
[...]
One nit and one question.
> +
> +static struct sk_buff *veth_xdp_rcv_skb(struct veth_priv *priv,
> + struct sk_buff *skb)
> +{
> + u32 pktlen, headroom, act, metalen;
> + void *orig_data, *orig_data_end;
> + int size, mac_len, delta, off;
> + struct bpf_prog *xdp_prog;
> + struct xdp_buff xdp;
> +
> + rcu_read_lock();
> + xdp_prog = rcu_dereference(priv->xdp_prog);
> + if (unlikely(!xdp_prog)) {
> + rcu_read_unlock();
> + goto out;
> + }
> +
> + mac_len = skb->data - skb_mac_header(skb);
> + pktlen = skb->len + mac_len;
> + size = SKB_DATA_ALIGN(VETH_XDP_HEADROOM + pktlen) +
> + SKB_DATA_ALIGN(sizeof(struct skb_shared_info));
> + if (size > PAGE_SIZE)
> + goto drop;
I'm not sure why it matters if size > PAGE_SIZE here. Why not
just consume it and use the correct page order in alloc_page if
its not linear.
> +
> + headroom = skb_headroom(skb) - mac_len;
> + if (skb_shared(skb) || skb_head_is_locked(skb) ||
> + skb_is_nonlinear(skb) || headroom < XDP_PACKET_HEADROOM) {
> + struct sk_buff *nskb;
> + void *head, *start;
> + struct page *page;
> + int head_off;
> +
> + page = alloc_page(GFP_ATOMIC);
Should also have __NO_WARN here as well this can be triggered by
external events so we don't want DDOS here to flood system logs.
> + if (!page)
> + goto drop;
> +
> + head = page_address(page);
> + start = head + VETH_XDP_HEADROOM;
> + if (skb_copy_bits(skb, -mac_len, start, pktlen)) {
> + page_frag_free(head);
> + goto drop;
> + }
> +
> + nskb = veth_build_skb(head,
> + VETH_XDP_HEADROOM + mac_len, skb->len,
> + PAGE_SIZE);
> + if (!nskb) {
> + page_frag_free(head);
> + goto drop;
> + }
> +
> + skb_copy_header(nskb, skb);
> + head_off = skb_headroom(nskb) - skb_headroom(skb);
> + skb_headers_offset_update(nskb, head_off);
> + if (skb->sk)
> + skb_set_owner_w(nskb, skb->sk);
> + consume_skb(skb);
> + skb = nskb;
> + }
> +
> + xdp.data_hard_start = skb->head;
> + xdp.data = skb_mac_header(skb);
> + xdp.data_end = xdp.data + pktlen;
> + xdp.data_meta = xdp.data;
> + xdp.rxq = &priv->xdp_rxq;
> + orig_data = xdp.data;
> + orig_data_end = xdp.data_end;
> +
> + act = bpf_prog_run_xdp(xdp_prog, &xdp);
> +
> + switch (act) {
> + case XDP_PASS:
> + break;
> + default:
> + bpf_warn_invalid_xdp_action(act);
> + case XDP_ABORTED:
> + trace_xdp_exception(priv->dev, xdp_prog, act);
> + case XDP_DROP:
> + goto drop;
> + }
> + rcu_read_unlock();
> +
> + delta = orig_data - xdp.data;
> + off = mac_len + delta;
> + if (off > 0)
> + __skb_push(skb, off);
> + else if (off < 0)
> + __skb_pull(skb, -off);
> + skb->mac_header -= delta;
> + off = xdp.data_end - orig_data_end;
> + if (off != 0)
> + __skb_put(skb, off);
> + skb->protocol = eth_type_trans(skb, priv->dev);
> +
> + metalen = xdp.data - xdp.data_meta;
> + if (metalen)
> + skb_metadata_set(skb, metalen);
> +out:
> + return skb;
> +drop:
> + rcu_read_unlock();
> + kfree_skb(skb);
> + return NULL;
> +}
> +
Thanks,
John
^ permalink raw reply
* Re: [PATCH] isdn: mISDN: hfcpci: Replace GFP_ATOMIC with GFP_KERNEL in hfc_probe()
From: David Miller @ 2018-07-27 4:24 UTC (permalink / raw)
To: baijiaju1990; +Cc: isdn, keescook, netdev, linux-kernel
In-Reply-To: <20180727023906.1590-1-baijiaju1990@gmail.com>
From: Jia-Ju Bai <baijiaju1990@gmail.com>
Date: Fri, 27 Jul 2018 10:39:06 +0800
> hfc_probe() is never called in atomic context.
> It calls kzalloc() with GFP_ATOMIC, which is not necessary.
> GFP_ATOMIC can be replaced with GFP_KERNEL.
>
> This is found by a static analysis tool named DCNS written by myself.
>
> Signed-off-by: Jia-Ju Bai <baijiaju1990@gmail.com>
Applied.
^ permalink raw reply
* Re: [PATCH] isdn: mISDN: netjet: Replace GFP_ATOMIC with GFP_KERNEL in nj_probe()
From: David Miller @ 2018-07-27 4:24 UTC (permalink / raw)
To: baijiaju1990; +Cc: isdn, keescook, jeyu, netdev, linux-kernel
In-Reply-To: <20180727024109.1702-1-baijiaju1990@gmail.com>
From: Jia-Ju Bai <baijiaju1990@gmail.com>
Date: Fri, 27 Jul 2018 10:41:09 +0800
> nj_probe() is never called in atomic context.
> It calls kzalloc() with GFP_ATOMIC, which is not necessary.
> GFP_ATOMIC can be replaced with GFP_KERNEL.
>
> This is found by a static analysis tool named DCNS written by myself.
>
> Signed-off-by: Jia-Ju Bai <baijiaju1990@gmail.com>
Applied.
^ permalink raw reply
* Re: [PATCH] isdn: hisax: config: Replace GFP_ATOMIC with GFP_KERNEL
From: David Miller @ 2018-07-27 4:25 UTC (permalink / raw)
To: baijiaju1990; +Cc: isdn, netdev, linux-kernel
In-Reply-To: <20180727024828.2107-1-baijiaju1990@gmail.com>
From: Jia-Ju Bai <baijiaju1990@gmail.com>
Date: Fri, 27 Jul 2018 10:48:28 +0800
> hisax_cs_new() and hisax_cs_setup() are never called in atomic context.
> They call kmalloc() and kzalloc() with GFP_ATOMIC, which is not necessary.
> GFP_ATOMIC can be replaced with GFP_KERNEL.
>
> This is found by a static analysis tool named DCNS written by myself.
>
> Signed-off-by: Jia-Ju Bai <baijiaju1990@gmail.com>
Applied.
^ permalink raw reply
* Re: [PATCH] net: adaptec: Replace mdelay() with msleep() in starfire_init_one()
From: David Miller @ 2018-07-27 4:25 UTC (permalink / raw)
To: baijiaju1990; +Cc: ionut, netdev, linux-kernel
In-Reply-To: <20180727035106.5185-1-baijiaju1990@gmail.com>
From: Jia-Ju Bai <baijiaju1990@gmail.com>
Date: Fri, 27 Jul 2018 11:51:06 +0800
> starfire_init_one() is never called in atomic context.
> It calls mdelay() to busily wait, which is not necessary.
> mdelay() can be replaced with msleep().
>
> This is found by a static analysis tool named DCNS written by myself.
>
> Signed-off-by: Jia-Ju Bai <baijiaju1990@gmail.com>
Applied.
^ permalink raw reply
* Re: [PATCH v2 net-next 0/3] docs: net: Convert netdev-FAQ to RST
From: David Miller @ 2018-07-27 4:28 UTC (permalink / raw)
To: me; +Cc: corbet, ecree, ast, daniel, linux-doc, netdev, linux-kernel
In-Reply-To: <20180726050031.1697-1-me@tobin.cc>
From: "Tobin C. Harding" <me@tobin.cc>
Date: Thu, 26 Jul 2018 15:00:28 +1000
> Jon answered all the tree questions on v1 so if you will please take
> this through your tree that would be awesome.
>
> v2:
> - Fix typo 'canonical_path_format' (thanks Edward)
> - Add patch fixing references netdev-FAQ
Series applied.
^ permalink raw reply
* Re: [PATCH 2/5] rhashtable: don't hold lock on first table throughout insertion.
From: Paul E. McKenney @ 2018-07-27 3:18 UTC (permalink / raw)
To: NeilBrown; +Cc: Herbert Xu, Thomas Graf, netdev, linux-kernel
In-Reply-To: <87r2jpmqu2.fsf@notabene.neil.brown.name>
On Fri, Jul 27, 2018 at 11:04:37AM +1000, NeilBrown wrote:
> On Wed, Jul 25 2018, Paul E. McKenney wrote:
> >>
> >> Looks good ... except ... naming is hard.
> >>
> >> is_after_call_rcu_init() asserts where in the lifecycle we are,
> >> is_after_call_rcu() tests where in the lifecycle we are.
> >>
> >> The names are similar but the purpose is quite different.
> >> Maybe s/is_after_call_rcu_init/call_rcu_init/ ??
> >
> > How about rcu_head_init() and rcu_head_after_call_rcu()?
Very well, I will pull this change in on my next rebase.
Thanx, Paul
^ permalink raw reply
* RE: linux-next: manual merge of the net-next tree with the rdma tree
From: Parav Pandit @ 2018-07-27 4:57 UTC (permalink / raw)
To: Stephen Rothwell, Jason Gunthorpe
Cc: David Miller, Networking, Doug Ledford, Linux-Next Mailing List,
Linux Kernel Mailing List, Ursula Braun, Leon Romanovsky,
linux-rdma@vger.kernel.org
In-Reply-To: <20180727132847.2be8563b@canb.auug.org.au>
> -----Original Message-----
> From: linux-rdma-owner@vger.kernel.org <linux-rdma-owner@vger.kernel.org>
> On Behalf Of Stephen Rothwell
> Sent: Thursday, July 26, 2018 10:29 PM
> To: Jason Gunthorpe <jgg@mellanox.com>
> Cc: David Miller <davem@davemloft.net>; Networking
> <netdev@vger.kernel.org>; Doug Ledford <dledford@redhat.com>; Linux-Next
> Mailing List <linux-next@vger.kernel.org>; Linux Kernel Mailing List <linux-
> kernel@vger.kernel.org>; Parav Pandit <parav@mellanox.com>; Ursula Braun
> <ubraun@linux.ibm.com>; Leon Romanovsky <leonro@mellanox.com>; linux-
> rdma@vger.kernel.org
> Subject: Re: linux-next: manual merge of the net-next tree with the rdma tree
>
> Hi Jason,
>
> On Thu, 26 Jul 2018 20:48:32 -0600 Jason Gunthorpe <jgg@mellanox.com>
> wrote:
> >
> > On Fri, Jul 27, 2018 at 12:33:01PM +1000, Stephen Rothwell wrote:
> >
> > > I fixed it up (I wasn't sure how to fix this up as so much has
> > > changed in the net-next tree and both modified functions had been
> > > (re)moved, so I effectively reverted the rdma tree commit) and can
> > > carry the fix as necessary. Please come to some arrangement about this.
> >
> > How does that still compile? We removed ib_query_gid() from the rdma
> > tree and replaced it with rdma_get_gid_attr()..
>
> Yeah, it doesn't :-(
>
> > I think the merge resolution is going to be a bit nasty to absorb that
> > much changing..
> >
> > Perhaps we should add a compatability ib_query_gid back to the RDMA
> > tree and then send DaveM a commit to fix SMC and remove it during the
> > next cycle? Linus can resolve smc_ib.c by using the net version
> >
> > Does someone else have a better idea?
>
> I applied this merge fix patch:
>
> From: Stephen Rothwell <sfr@canb.auug.org.au>
> Date: Fri, 27 Jul 2018 13:19:31 +1000
> Subject: [PATCH] net/smc: fixups for ip_query_gid API removal
>
> Signed-off-by: Stephen Rothwell <sfr@canb.auug.org.au>
> ---
> net/smc/smc_ib.c | 47 +++++++++++++++++++++++++----------------------
> 1 file changed, 25 insertions(+), 22 deletions(-)
>
> diff --git a/net/smc/smc_ib.c b/net/smc/smc_ib.c index
> 2cc64bc8ae20..debc6e44f738 100644
> --- a/net/smc/smc_ib.c
> +++ b/net/smc/smc_ib.c
> @@ -16,6 +16,7 @@
> #include <linux/workqueue.h>
> #include <linux/scatterlist.h>
> #include <rdma/ib_verbs.h>
> +#include <rdma/ib_cache.h>
>
> #include "smc_pnet.h"
> #include "smc_ib.h"
> @@ -144,17 +145,21 @@ int smc_ib_ready_link(struct smc_link *lnk)
>
> static int smc_ib_fill_mac(struct smc_ib_device *smcibdev, u8 ibport) {
> - struct ib_gid_attr gattr;
> - union ib_gid gid;
> - int rc;
> + const struct ib_gid_attr *gattr;
> + int rc = 0;
>
> - rc = ib_query_gid(smcibdev->ibdev, ibport, 0, &gid, &gattr);
> - if (rc || !gattr.ndev)
> - return -ENODEV;
> + gattr = rdma_get_gid_attr(smcibdev->ibdev, ibport, 0);
> + if (IS_ERR(gattr))
> + return PTR_ERR(gattr);
> + if (!gattr->ndev) {
> + rc = -ENODEV;
> + goto done;
> + }
>
> - memcpy(smcibdev->mac[ibport - 1], gattr.ndev->dev_addr, ETH_ALEN);
> - dev_put(gattr.ndev);
> - return 0;
> + memcpy(smcibdev->mac[ibport - 1], gattr->ndev->dev_addr,
> ETH_ALEN);
> +done:
> + rdma_put_gid_attr(gattr);
> + return rc;
> }
>
> /* Create an identifier unique for this instance of SMC-R.
> @@ -179,29 +184,27 @@ bool smc_ib_port_active(struct smc_ib_device
> *smcibdev, u8 ibport) int smc_ib_determine_gid(struct smc_ib_device
> *smcibdev, u8 ibport,
> unsigned short vlan_id, u8 gid[], u8 *sgid_index) {
> - struct ib_gid_attr gattr;
> - union ib_gid _gid;
> + const struct ib_gid_attr *gattr;
> int i;
>
> for (i = 0; i < smcibdev->pattr[ibport - 1].gid_tbl_len; i++) {
> - memset(&_gid, 0, SMC_GID_SIZE);
> - memset(&gattr, 0, sizeof(gattr));
> - if (ib_query_gid(smcibdev->ibdev, ibport, i, &_gid, &gattr))
> + gattr = rdma_get_gid_attr(smcibdev->ibdev, ibport, i);
> + if (IS_ERR(gattr))
> continue;
> - if (!gattr.ndev)
> + if (!gattr->ndev)
> continue;
This requires a small fix.
If (!gattr->ndev {
rdma_put_gid_attr(gattr);
continue;
}
> - if (((!vlan_id && !is_vlan_dev(gattr.ndev)) ||
> - (vlan_id && is_vlan_dev(gattr.ndev) &&
> - vlan_dev_vlan_id(gattr.ndev) == vlan_id)) &&
> - gattr.gid_type == IB_GID_TYPE_IB) {
> + if (((!vlan_id && !is_vlan_dev(gattr->ndev)) ||
> + (vlan_id && is_vlan_dev(gattr->ndev) &&
> + vlan_dev_vlan_id(gattr->ndev) == vlan_id)) &&
> + gattr->gid_type == IB_GID_TYPE_IB) {
> if (gid)
> - memcpy(gid, &_gid, SMC_GID_SIZE);
> + memcpy(gid, &gattr->gid, SMC_GID_SIZE);
> if (sgid_index)
> *sgid_index = i;
> - dev_put(gattr.ndev);
> + rdma_put_gid_attr(gattr);
> return 0;
> }
> - dev_put(gattr.ndev);
> + rdma_put_gid_attr(gattr);
> }
> return -ENODEV;
> }
> --
> 2.18.0
>
> --
> Cheers,
> Stephen Rothwell
^ permalink raw reply
* Re: [PATCH v5 bpf-next 4/9] veth: Handle xdp_frames in xdp napi ring
From: John Fastabend @ 2018-07-27 3:40 UTC (permalink / raw)
To: Toshiaki Makita, netdev, Alexei Starovoitov, Daniel Borkmann
Cc: Toshiaki Makita, Jesper Dangaard Brouer, Jakub Kicinski
In-Reply-To: <20180726144032.2116-5-toshiaki.makita1@gmail.com>
On 07/26/2018 07:40 AM, Toshiaki Makita wrote:
> From: Toshiaki Makita <makita.toshiaki@lab.ntt.co.jp>
>
> This is preparation for XDP TX and ndo_xdp_xmit.
> This allows napi handler to handle xdp_frames through xdp ring as well
> as sk_buff.
>
> v3:
> - Revert v2 change around rings and use a flag to differentiate skb and
> xdp_frame, since bulk skb xmit makes little performance difference
> for now.
>
> v2:
> - Use another ring instead of using flag to differentiate skb and
> xdp_frame. This approach makes bulk skb transmit possible in
> veth_xmit later.
> - Clear xdp_frame feilds in skb->head.
> - Implement adjust_tail.
>
> Signed-off-by: Toshiaki Makita <makita.toshiaki@lab.ntt.co.jp>
> ---
> drivers/net/veth.c | 87 ++++++++++++++++++++++++++++++++++++++++++++++++++----
> 1 file changed, 82 insertions(+), 5 deletions(-)
>
Acked-by: John Fastabend <john.fastabend@gmail.com>
^ permalink raw reply
* RE: linux-next: manual merge of the net-next tree with the rdma tree
From: Parav Pandit @ 2018-07-27 5:03 UTC (permalink / raw)
To: Stephen Rothwell, Jason Gunthorpe
Cc: David Miller, Networking, Doug Ledford, Linux-Next Mailing List,
Linux Kernel Mailing List, Ursula Braun, Leon Romanovsky,
linux-rdma@vger.kernel.org
In-Reply-To: <20180727134524.366e9425@canb.auug.org.au>
> -----Original Message-----
> From: linux-rdma-owner@vger.kernel.org <linux-rdma-owner@vger.kernel.org>
> On Behalf Of Stephen Rothwell
> Sent: Thursday, July 26, 2018 10:45 PM
> To: Jason Gunthorpe <jgg@mellanox.com>
> Cc: David Miller <davem@davemloft.net>; Networking
> <netdev@vger.kernel.org>; Doug Ledford <dledford@redhat.com>; Linux-Next
> Mailing List <linux-next@vger.kernel.org>; Linux Kernel Mailing List <linux-
> kernel@vger.kernel.org>; Parav Pandit <parav@mellanox.com>; Ursula Braun
> <ubraun@linux.ibm.com>; Leon Romanovsky <leonro@mellanox.com>; linux-
> rdma@vger.kernel.org
> Subject: Re: linux-next: manual merge of the net-next tree with the rdma tree
>
> Hi all,
>
> On Fri, 27 Jul 2018 13:28:47 +1000 Stephen Rothwell <sfr@canb.auug.org.au>
> wrote:
> >
> > I applied this merge fix patch:
>
> The final conflict resolution actually looks like this:
>
> (the rdma tree changes to net/smc/smc_core.c are dropped)
>
> c1d4bb2af93573ee4a21538a1a97b568a2344499
> diff --cc net/smc/smc_ib.c
> index 74f29f814ec1,2cc64bc8ae20..debc6e44f738
> --- a/net/smc/smc_ib.c
> +++ b/net/smc/smc_ib.c
> @@@ -144,6 -142,93 +143,95 @@@ out
> return rc;
> }
>
> + static int smc_ib_fill_mac(struct smc_ib_device *smcibdev, u8 ibport)
> + {
> - struct ib_gid_attr gattr;
> - union ib_gid gid;
> - int rc;
> ++ const struct ib_gid_attr *gattr;
> ++ int rc = 0;
> +
> - rc = ib_query_gid(smcibdev->ibdev, ibport, 0, &gid, &gattr);
> - if (rc || !gattr.ndev)
> - return -ENODEV;
> ++ gattr = rdma_get_gid_attr(smcibdev->ibdev, ibport, 0);
> ++ if (IS_ERR(gattr))
> ++ return PTR_ERR(gattr);
> ++ if (!gattr->ndev) {
> ++ rc = -ENODEV;
> ++ goto done;
> ++ }
> +
> - memcpy(smcibdev->mac[ibport - 1], gattr.ndev->dev_addr, ETH_ALEN);
> - dev_put(gattr.ndev);
> - return 0;
> ++ memcpy(smcibdev->mac[ibport - 1], gattr->ndev->dev_addr,
> ETH_ALEN);
> ++done:
> ++ rdma_put_gid_attr(gattr);
> ++ return rc;
> + }
> +
> + /* Create an identifier unique for this instance of SMC-R.
> + * The MAC-address of the first active registered IB device
> + * plus a random 2-byte number is used to create this identifier.
> + * This name is delivered to the peer during connection initialization.
> + */
> + static inline void smc_ib_define_local_systemid(struct smc_ib_device
> *smcibdev,
> + u8 ibport)
> + {
> + memcpy(&local_systemid[2], &smcibdev->mac[ibport - 1],
> + sizeof(smcibdev->mac[ibport - 1]));
> + get_random_bytes(&local_systemid[0], 2); }
> +
> + bool smc_ib_port_active(struct smc_ib_device *smcibdev, u8 ibport) {
> + return smcibdev->pattr[ibport - 1].state == IB_PORT_ACTIVE; }
> +
> + /* determine the gid for an ib-device port and vlan id */ int
> + smc_ib_determine_gid(struct smc_ib_device *smcibdev, u8 ibport,
> + unsigned short vlan_id, u8 gid[], u8 *sgid_index) {
> - struct ib_gid_attr gattr;
> - union ib_gid _gid;
> ++ const struct ib_gid_attr *gattr;
> + int i;
> +
> + for (i = 0; i < smcibdev->pattr[ibport - 1].gid_tbl_len; i++) {
> - memset(&_gid, 0, SMC_GID_SIZE);
> - memset(&gattr, 0, sizeof(gattr));
> - if (ib_query_gid(smcibdev->ibdev, ibport, i, &_gid, &gattr))
> ++ gattr = rdma_get_gid_attr(smcibdev->ibdev, ibport, i);
> ++ if (IS_ERR(gattr))
> + continue;
> - if (!gattr.ndev)
> ++ if (!gattr->ndev)
> + continue;
Seeing this updated patch, so for completeness same reply as the previous email.
If (!gattr->ndev) {
rdma_put_gid_attr(gattr);
continue;
}
Rest changes above and below looks fine to me.
Thanks for doing it, I am not part of netdev mailing list so didn't see the compile error until this patch came up.
> - if (((!vlan_id && !is_vlan_dev(gattr.ndev)) ||
> - (vlan_id && is_vlan_dev(gattr.ndev) &&
> - vlan_dev_vlan_id(gattr.ndev) == vlan_id)) &&
> - gattr.gid_type == IB_GID_TYPE_IB) {
> ++ if (((!vlan_id && !is_vlan_dev(gattr->ndev)) ||
> ++ (vlan_id && is_vlan_dev(gattr->ndev) &&
> ++ vlan_dev_vlan_id(gattr->ndev) == vlan_id)) &&
> ++ gattr->gid_type == IB_GID_TYPE_IB) {
> + if (gid)
> - memcpy(gid, &_gid, SMC_GID_SIZE);
> ++ memcpy(gid, &gattr->gid, SMC_GID_SIZE);
> + if (sgid_index)
> + *sgid_index = i;
> - dev_put(gattr.ndev);
> ++ rdma_put_gid_attr(gattr);
> + return 0;
> + }
> - dev_put(gattr.ndev);
> ++ rdma_put_gid_attr(gattr);
> + }
> + return -ENODEV;
> + }
> +
> + static int smc_ib_remember_port_attr(struct smc_ib_device *smcibdev,
> + u8 ibport) {
> + int rc;
> +
> + memset(&smcibdev->pattr[ibport - 1], 0,
> + sizeof(smcibdev->pattr[ibport - 1]));
> + rc = ib_query_port(smcibdev->ibdev, ibport,
> + &smcibdev->pattr[ibport - 1]);
> + if (rc)
> + goto out;
> + /* the SMC protocol requires specification of the RoCE MAC address */
> + rc = smc_ib_fill_mac(smcibdev, ibport);
> + if (rc)
> + goto out;
> + if (!strncmp(local_systemid, SMC_LOCAL_SYSTEMID_RESET,
> + sizeof(local_systemid)) &&
> + smc_ib_port_active(smcibdev, ibport))
> + /* create unique system identifier */
> + smc_ib_define_local_systemid(smcibdev, ibport);
> + out:
> + return rc;
> + }
> +
> /* process context wrapper for might_sleep smc_ib_remember_port_attr */
> static void smc_ib_port_event_work(struct work_struct *work)
> {
>
> --
> Cheers,
> Stephen Rothwell
^ permalink raw reply
* Re: linux-next: manual merge of the net-next tree with the rdma tree
From: Stephen Rothwell @ 2018-07-27 5:09 UTC (permalink / raw)
To: Parav Pandit
Cc: Jason Gunthorpe, David Miller, Networking, Doug Ledford,
Linux-Next Mailing List, Linux Kernel Mailing List, Ursula Braun,
Leon Romanovsky, linux-rdma@vger.kernel.org
In-Reply-To: <VI1PR0502MB3008702AFA64F44DE57DC9F5D12A0@VI1PR0502MB3008.eurprd05.prod.outlook.com>
[-- Attachment #1: Type: text/plain, Size: 673 bytes --]
Hi Parav,
On Fri, 27 Jul 2018 04:57:41 +0000 Parav Pandit <parav@mellanox.com> wrote:
>
> > for (i = 0; i < smcibdev->pattr[ibport - 1].gid_tbl_len; i++) {
> > - memset(&_gid, 0, SMC_GID_SIZE);
> > - memset(&gattr, 0, sizeof(gattr));
> > - if (ib_query_gid(smcibdev->ibdev, ibport, i, &_gid, &gattr))
> > + gattr = rdma_get_gid_attr(smcibdev->ibdev, ibport, i);
> > + if (IS_ERR(gattr))
> > continue;
> > - if (!gattr.ndev)
> > + if (!gattr->ndev)
> > continue;
> This requires a small fix.
> If (!gattr->ndev {
> rdma_put_gid_attr(gattr);
> continue;
> }
Thanks, I have fixed this up for Monday.
--
Cheers,
Stephen Rothwell
[-- Attachment #2: OpenPGP digital signature --]
[-- Type: application/pgp-signature, Size: 488 bytes --]
^ permalink raw reply
* Re: [PATCH v5 bpf-next 5/9] veth: Add ndo_xdp_xmit
From: John Fastabend @ 2018-07-27 3:54 UTC (permalink / raw)
To: Toshiaki Makita, netdev, Alexei Starovoitov, Daniel Borkmann
Cc: Toshiaki Makita, Jesper Dangaard Brouer, Jakub Kicinski
In-Reply-To: <20180726144032.2116-6-toshiaki.makita1@gmail.com>
On 07/26/2018 07:40 AM, Toshiaki Makita wrote:
> From: Toshiaki Makita <makita.toshiaki@lab.ntt.co.jp>
>
> This allows NIC's XDP to redirect packets to veth. The destination veth
> device enqueues redirected packets to the napi ring of its peer, then
> they are processed by XDP on its peer veth device.
> This can be thought as calling another XDP program by XDP program using
> REDIRECT, when the peer enables driver XDP.
>
> Note that when the peer veth device does not set driver xdp, redirected
> packets will be dropped because the peer is not ready for NAPI.
>
> v4:
> - Don't use xdp_ok_fwd_dev() because checking IFF_UP is not necessary.
> Add comments about it and check only MTU.
>
> v2:
> - Drop the part converting xdp_frame into skb when XDP is not enabled.
> - Implement bulk interface of ndo_xdp_xmit.
> - Implement XDP_XMIT_FLUSH bit and drop ndo_xdp_flush.
>
> Signed-off-by: Toshiaki Makita <makita.toshiaki@lab.ntt.co.jp>
> ---
> drivers/net/veth.c | 51 +++++++++++++++++++++++++++++++++++++++++++++++++++
> 1 file changed, 51 insertions(+)
>
Acked-by: John Fastabend <john.fastabend@gmail.com>
^ permalink raw reply
* Re: [PATCH v2 0/4] docs: bpf: Fix RST conversion
From: Daniel Borkmann @ 2018-07-27 5:25 UTC (permalink / raw)
To: Tobin C. Harding, David S. Miller
Cc: Jonathan Corbet, Sergei Shtylyov, Alexei Starovoitov, linux-doc,
netdev, linux-kernel
In-Reply-To: <20180726050305.2147-1-me@tobin.cc>
On 07/26/2018 07:03 AM, Tobin C. Harding wrote:
> Hi Dave,
>
> I'm sending v2 of this to you instead of to Jon. Rationale: BPF is to
> do with networking anyways and there is a broken link that requires
> conversion of Documentation/networking/filter.txt to fix and that will
> go to you. FTR there are no merge conflicts between this set and the
> other set I just sent (either can be applied on top of the other)
>
> [PATCH v2 net-next 0/3] docs: net: Convert netdev-FAQ to RST
>
>
> Recently BPF docs were converted to RST format. A couple of things were
> missed.
>
> - Use 'index.rst' instead of 'README.rst'. Although README.rst will
> work just fine it is more typical to keep the subdirectory indices
> in a file called 'index.rst'.
>
> - Integrate files Documentation/bpf/*.rst into build system using
> toctree in Documentation/bpf/index.rst
>
> - Include bpf/index in top level toctree so bpf is indexed in the main
> kernel docs.
>
> - Make anal change to heading format (inline with rest of Documentation/).
>
>
> thanks,
> Tobin.
Applied to bpf-next, thanks Tobin!
^ permalink raw reply
* Re: [PATCH v3 bpf-next 02/14] bpf: introduce cgroup storage maps
From: Daniel Borkmann @ 2018-07-27 4:11 UTC (permalink / raw)
To: Roman Gushchin, netdev; +Cc: linux-kernel, kernel-team, Alexei Starovoitov
In-Reply-To: <20180720174558.5829-3-guro@fb.com>
On 07/20/2018 07:45 PM, Roman Gushchin wrote:
> This commit introduces BPF_MAP_TYPE_CGROUP_STORAGE maps:
> a special type of maps which are implementing the cgroup storage.
>
> From the userspace point of view it's almost a generic
> hash map with the (cgroup inode id, attachment type) pair
> used as a key.
>
> The only difference is that some operations are restricted:
> 1) a user can't create new entries,
> 2) a user can't remove existing entries.
>
> The lookup from userspace is o(log(n)).
>
> Signed-off-by: Roman Gushchin <guro@fb.com>
> Cc: Alexei Starovoitov <ast@kernel.org>
> Cc: Daniel Borkmann <daniel@iogearbox.net>
> Acked-by: Martin KaFai Lau <kafai@fb.com>
(First of all sorry for the late review, only limited availability this week
on my side.)
> ---
> include/linux/bpf-cgroup.h | 38 +++++
> include/linux/bpf.h | 1 +
> include/linux/bpf_types.h | 3 +
> include/uapi/linux/bpf.h | 6 +
> kernel/bpf/Makefile | 1 +
> kernel/bpf/local_storage.c | 367 +++++++++++++++++++++++++++++++++++++++++++++
> kernel/bpf/verifier.c | 12 ++
> 7 files changed, 428 insertions(+)
> create mode 100644 kernel/bpf/local_storage.c
>
> diff --git a/include/linux/bpf-cgroup.h b/include/linux/bpf-cgroup.h
> index 79795c5fa7c3..6b0e7bd4b154 100644
> --- a/include/linux/bpf-cgroup.h
> +++ b/include/linux/bpf-cgroup.h
> @@ -3,19 +3,39 @@
> #define _BPF_CGROUP_H
>
> #include <linux/jump_label.h>
> +#include <linux/rbtree.h>
> #include <uapi/linux/bpf.h>
>
[...]
> diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c
> index 15d69b278277..0b089ba4595d 100644
> --- a/kernel/bpf/verifier.c
> +++ b/kernel/bpf/verifier.c
> @@ -5140,6 +5140,14 @@ static int replace_map_fd_with_map_ptr(struct bpf_verifier_env *env)
> return -E2BIG;
> }
>
> + if (map->map_type == BPF_MAP_TYPE_CGROUP_STORAGE &&
> + bpf_cgroup_storage_assign(env->prog, map)) {
> + verbose(env,
> + "only one cgroup storage is allowed\n");
> + fdput(f);
> + return -EBUSY;
> + }
> +
> /* hold the map. If the program is rejected by verifier,
> * the map will be released by release_maps() or it
> * will be used by the valid program until it's unloaded
> @@ -5148,6 +5156,10 @@ static int replace_map_fd_with_map_ptr(struct bpf_verifier_env *env)
> map = bpf_map_inc(map, false);
> if (IS_ERR(map)) {
> fdput(f);
> + if (map->map_type ==
> + BPF_MAP_TYPE_CGROUP_STORAGE)
> + bpf_cgroup_storage_release(env->prog,
> + map);
I think this behavior is a bit strange, meaning that we would reset the map via
bpf_cgroup_storage_release() in this case, but if we error out and exit in any
later instruction the prior bpf_cgroup_storage_assign() is not undone, meaning
at this point we have no other choice but to destroy the map since any later
BPF prog load with bpf_cgroup_storage_assign() attempt would fail with -EBUSY
even though it's not assigned anywhere, is that correct? Same also on any other
errors along the prog load path. E.g. say, as one example, your verifier buffer
is too small, so any retry with the very same program from loader side would fail
above due to different prog pointers?
> return PTR_ERR(map);
> }
> env->used_maps[env->used_map_cnt++] = map;
>
^ permalink raw reply
page: next (older) | prev (newer) | latest
- recent:[subjects (threaded)|topics (new)|topics (active)]
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox