Netdev List
 help / color / mirror / Atom feed
* Re: [PATCH net 2/2] sh_eth: Fix ethtool operation crash when net device is down
From: Ben Hutchings @ 2015-01-19 10:41 UTC (permalink / raw)
  To: Florian Fainelli
  Cc: ct178-internal, David S. Miller, netdev, linux-kernel,
	Nobuhiro Iwamatsu, Mitsuhiro Kimura, Hisashi Nakamura,
	Yoshihiro Kaneko
In-Reply-To: <CAGVrzcbj2JANTECtxgDUrb7c_20nyoW4n5ARndHGKYuQu-0kQg@mail.gmail.com>

On Fri, 2015-01-16 at 10:45 -0800, Florian Fainelli wrote:
> 2015-01-16 9:51 GMT-08:00 Ben Hutchings <ben.hutchings@codethink.co.uk>:
[...]
> > --- a/drivers/net/ethernet/renesas/sh_eth.c
> > +++ b/drivers/net/ethernet/renesas/sh_eth.c
> > @@ -1827,6 +1827,9 @@ static int sh_eth_get_settings(struct net_device *ndev,
> >         unsigned long flags;
> >         int ret;
> >
> > +       if (!mdp->phydev)
> > +               return -ENODEV;
> 
> Since the PHY is disconnected, would not checking for netif_running()
> make sense here, unless there is a good reason to still allow
> phy_ethtool_gset() to be called?
[...]

I think those two conditions will be equivalent, won't they?  Writing
the condition like this will also work if the driver later supports
PHY-less configurations.

Ben.

^ permalink raw reply

* Re: [PATCH net 1/2] sh_eth: Fix promiscuous mode on chips without TSU
From: Ben Hutchings @ 2015-01-19 10:45 UTC (permalink / raw)
  To: Sergei Shtylyov
  Cc: ct178-internal, David S. Miller, netdev, linux-kernel,
	Nobuhiro Iwamatsu, Mitsuhiro Kimura, Hisashi Nakamura,
	Yoshihiro Kaneko
In-Reply-To: <54B9663D.2040008@cogentembedded.com>

On Fri, 2015-01-16 at 22:27 +0300, Sergei Shtylyov wrote:
> Hello.
> 
> On 01/16/2015 08:51 PM, Ben Hutchings wrote:
> 
> > Currently net_device_ops::set_rx_mode is only implemented for
> > chips with a TSU (multiple address table).  However we do need
> > to turn the PRM (promiscuous) flag on and off for other chips.
> 
> > - Remove the unlikely() from the TSU functions that we may safely
> >    call for chips without a TSU
> 
>     This is just optimization, worth pushing thru net-next instead.

This patch introduces calls to those functions for chips without a TSU,
and that makes the branch hint no longer correct.  (Not that these
functions are speed-critical, so it doesn't matter that much.)

> > - Make setting of the MCT flag conditional on the tsu capability flag
> > - Rename sh_eth_set_multicast_list() to sh_eth_set_rx_mode() and plumb
> >    it into both net_device_ops structures
> > - Remove the previously-unreachable branch in sh_eth_rx_mode() that
> >    would otherwise reset the flags to defaults for non-TSU chips
> 
>     It couldn't be default for non-TSU chips, they don't seem to have 
> ECMR.MCT. So it was just wrong.
> 
>     It would have been better if you did one thing per patch or at least 
> didn't mix fixes with clean-ups...
[...]

I think this is one logical change.

Ben.

^ permalink raw reply

* [PATCH] phonet netlink: allow multiple messages per skb in route dump
From: Johannes Berg @ 2015-01-19 11:15 UTC (permalink / raw)
  To: netdev; +Cc: Sakari Ailus, Remi Denis-Courmont, Johannes Berg

From: Johannes Berg <johannes.berg@intel.com>

My previous patch to this file changed the code to be bug-compatible
towards userspace. Unless userspace (which I wasn't able to find)
implements the dump reader by hand in a wrong way, this isn't needed.
If it uses libnl or similar code putting multiple messages into a
single SKB is far more efficient.

Change the code to do this. While at it, also clean it up and don't
use so many variables - just store the address in the callback args
directly.

Signed-off-by: Johannes Berg <johannes.berg@intel.com>
---
 net/phonet/pn_netlink.c | 22 +++++++---------------
 1 file changed, 7 insertions(+), 15 deletions(-)

diff --git a/net/phonet/pn_netlink.c b/net/phonet/pn_netlink.c
index 54d766842c2b..bc5ee5fbe6ae 100644
--- a/net/phonet/pn_netlink.c
+++ b/net/phonet/pn_netlink.c
@@ -272,31 +272,23 @@ static int route_doit(struct sk_buff *skb, struct nlmsghdr *nlh)
 static int route_dumpit(struct sk_buff *skb, struct netlink_callback *cb)
 {
 	struct net *net = sock_net(skb->sk);
-	u8 addr, addr_idx = 0, addr_start_idx = cb->args[0];
+	u8 addr;
 
 	rcu_read_lock();
-	for (addr = 0; addr < 64; addr++) {
-		struct net_device *dev;
+	for (addr = cb->args[0]; addr < 64; addr++) {
+		struct net_device *dev = phonet_route_get_rcu(net, addr << 2);
 
-		dev = phonet_route_get_rcu(net, addr << 2);
 		if (!dev)
 			continue;
 
-		if (addr_idx++ < addr_start_idx)
-			continue;
-		fill_route(skb, dev, addr << 2, NETLINK_CB(cb->skb).portid,
-			   cb->nlh->nlmsg_seq, RTM_NEWROUTE);
-		/* fill_route() used to return > 0 (or negative errors) but
-		 * never 0 - ignore the return value and just go out to
-		 * call dumpit again from outside to preserve the behavior
-		 */
-		goto out;
+		if (fill_route(skb, dev, addr << 2, NETLINK_CB(cb->skb).portid,
+			       cb->nlh->nlmsg_seq, RTM_NEWROUTE) < 0)
+			goto out;
 	}
 
 out:
 	rcu_read_unlock();
-	cb->args[0] = addr_idx;
-	cb->args[1] = 0;
+	cb->args[0] = addr;
 
 	return skb->len;
 }
-- 
2.1.4

^ permalink raw reply related

* GOOD DAY
From: John Watson @ 2015-01-19 11:53 UTC (permalink / raw)


Watson Investment Group.
22 Pasteur Drive,Leegomery Telford
United Kingdom ,TF1 6PQ
Tel: 447053834136
Fax: +44 8435 622250
E-mail:jw25743@gmail.com


We are a private financial company based in the united Kingdom. We
offer funding for investments across the world. We have funded and
invested in businesses and given personal,multi-purpose loans to
individuals and groups. We have transformed DREAMS into REALITIES.

We will like to invest in your company or project using our
BG/Monetization Program, please send more details about your company
to provide you with our procedure for the loan.


John Watson
CEO Watson Investment Group

^ permalink raw reply

* Re: [PATCH 7/9] rhashtable: Per bucket locks & deferred expansion/shrinking
From: Thomas Graf @ 2015-01-19 12:58 UTC (permalink / raw)
  To: Patrick McHardy
  Cc: David Laight, davem@davemloft.net, netdev@vger.kernel.org,
	herbert@gondor.apana.org.au, paulmck@linux.vnet.ibm.com,
	edumazet@google.com, john.r.fastabend@intel.com,
	josh@joshtriplett.org, netfilter-devel@vger.kernel.org
In-Reply-To: <20150117080255.GA3968@acer.localdomain>

On 01/17/15 at 08:02am, Patrick McHardy wrote:
> On 16.01, Thomas Graf wrote:
> > Resize operations should be *really* rare as well unless you start
> > with really small hash table sizes and constantly add/remove at the
> > watermark.
> 
> Which are far enough from each other that this should only happen
> in really unlucky cases.
> 
> > Re-dumping on insert/remove is a different story of course. Do you
> > care about missed insert/removals for dumps? If not we can do the
> > sequence number consistency checking for resizing only.
> 
> No, that has always been undeterministic with netlink. We want to
> dump everything that was present when the dump was started and is
> still present when it finishes. Anything else can be handled using
> notifications.

It looks like we want to provide two ways to resolve this:

1) Walker holds ht->mutex the entire time to block out resizes.
   Optionally the walker can acquire all bucket locks. Such
   scenarios would seem to benefit from either a single or a very
   small number of bucket locks.

2) Walker holds ht->mutex during individual Netlink message
   construction periods and relases it while user space reads the
   message. rhashtable provides a hook which is called when a
   resize operation is scheduled allowing for the walker code to
   bump a sequence number and notify user space that the dump is
   inconsistent, causing it to request a new dump.

I'll provide an API to achieve (2). (1) is already achieveable with
the current API.

^ permalink raw reply

* Re: [3.19.0-rc4+] rhashtable: BUG kmalloc-2048 (Not tainted): Poison overwritten
From: Thomas Graf @ 2015-01-19 12:59 UTC (permalink / raw)
  To: Ying Xue
  Cc: richard.alpe@ericsson.com >> Richard Alpe, Netdev,
	tipc-discussion
In-Reply-To: <54BCBA35.2080103@windriver.com>

On 01/19/15 at 04:03pm, Ying Xue wrote:
> On 3.19.0-rc4+, I encountered below error with attached test
> case(bind_netlink.c). Please execute the following commands to reproduce
> the error:
> 
> gcc -Wall -o bind_netlink bind_netlink.c
> ./bind_netlink 1000
> 
> By the way, if we run another test case(bind_tipc.c), the similar issue
> will happen on TIPC socket.
> 
> Therefore, it seems that the issue is closely associated with rhashtable
> instead of specific stacks like netlink or tipc.

Looks like a RCU read side critical section was missed. Does the TIPC
poision warning look the same? offset 2048?

^ permalink raw reply

* Re: tcp: Do not apply TSO segment limit to non-TSO packets
From: Thomas Jarosch @ 2015-01-19 13:39 UTC (permalink / raw)
  To: Herbert Xu
  Cc: netdev, edumazet, Steffen Klassert, Ben Hutchings,
	David S. Miller
In-Reply-To: <20141231133923.GA30248@gondor.apana.org.au>

Hi Herbert,

On Thursday, 1. January 2015 00:39:23 Herbert Xu wrote:
> The problem is that when the MSS goes down, existing queued packet
> on the TX queue that have not been transmitted yet all look like
> TSO packets and get treated as such.
> 
> This then triggers a bug where tcp_mss_split_point tells us to
> generate a zero-sized packet on the TX queue.  Once that happens
> we're screwed because the zero-sized packet can never be removed
> by ACKs.

picking up this one again: Is there any valid use case to have
zero-sized packets in the TX queue? If not, may be a WARN_ON() could
be added to the processing of the TX queue. That would help to
prevent future issues like this.

Cheers,
Thomas

^ permalink raw reply

* Re: [PATCH net] ipv6: stop sending PTB packets for MTU < 1280
From: Hannes Frederic Sowa @ 2015-01-19 13:55 UTC (permalink / raw)
  To: Hagen Paul Pfeifer; +Cc: netdev, stable, Fernando Gont
In-Reply-To: <1421357665-2804-1-git-send-email-hagen@jauu.net>

On Do, 2015-01-15 at 22:34 +0100, Hagen Paul Pfeifer wrote:
> Reduce the attack vector and stop generating IPv6 Fragment Header for
> paths with an MTU smaller than the minimum required IPv6 MTU
> size (1280 byte) - called atomic fragments.
> 
> See IETF I-D "Deprecating the Generation of IPv6 Atomic Fragments" [1]
> for more information and how this "feature" can be misused.
> 
> [1] https://tools.ietf.org/html/draft-ietf-6man-deprecate-atomfrag-generation-00
> 
> Cc: stable@vger.kernel.org
> Signed-off-by: Fernando Gont <fgont@si6networks.com>
> Signed-off-by: Hagen Paul Pfeifer <hagen@jauu.net>

Acked-by: Hannes Frederic Sowa <hannes@stressinduktion.org>

I think this is the correct way forward on how to deal with atomic
fragments.

Hagen, do you submit patches to remove dst_allfrag/RTAX_FEATURE_ALLFRAG,
IPCORK_ALLFRAG, etc. for net-next, too?

Thanks,
Hannes

^ permalink raw reply

* Re: [PATCH iproute2 3/3] ss: Unify tcp stats output
From: Hagen Paul Pfeifer @ 2015-01-19 13:57 UTC (permalink / raw)
  To: Vadim Kochan; +Cc: netdev
In-Reply-To: <1421613815-6635-4-git-send-email-vadim4j@gmail.com>

On 18 January 2015 at 21:43, Vadim Kochan <vadim4j@gmail.com> wrote:

> Signed-off-by: Vadim Kochan <vadim4j@gmail.com>
> ---
>  misc/ss.c | 362 +++++++++++++++++++++++++++++++++++++++-----------------------
>  1 file changed, 231 insertions(+), 131 deletions(-)

Hey Vadin,

your patch do *not* change the output format, right? ss output is
parsed by scripts and tools.

hgn

^ permalink raw reply

* Re: [PATCH net] ipv6: stop sending PTB packets for MTU < 1280
From: Hagen Paul Pfeifer @ 2015-01-19 14:00 UTC (permalink / raw)
  To: Hannes Frederic Sowa; +Cc: netdev, stable, Fernando Gont
In-Reply-To: <1421675722.32277.3.camel@stressinduktion.org>

On 19 January 2015 at 14:55, Hannes Frederic Sowa
<hannes@stressinduktion.org> wrote:
> Acked-by: Hannes Frederic Sowa <hannes@stressinduktion.org>
>
> I think this is the correct way forward on how to deal with atomic
> fragments.
>
> Hagen, do you submit patches to remove dst_allfrag/RTAX_FEATURE_ALLFRAG,
> IPCORK_ALLFRAG, etc. for net-next, too?

Yes, patch sits already in the pipe. I wanted to wait for davem's pull.

Hagen

^ permalink raw reply

* Re: [PATCH iproute2 3/3] ss: Unify tcp stats output
From: Vadim Kochan @ 2015-01-19 14:04 UTC (permalink / raw)
  To: Hagen Paul Pfeifer; +Cc: Vadim Kochan, netdev
In-Reply-To: <CAPh34mdJ7mjzueaX2qakoT=ODbHDLm4k-K_3Eg__6+s7ydf4Gg@mail.gmail.com>

On Mon, Jan 19, 2015 at 02:57:50PM +0100, Hagen Paul Pfeifer wrote:
> On 18 January 2015 at 21:43, Vadim Kochan <vadim4j@gmail.com> wrote:
> 
> > Signed-off-by: Vadim Kochan <vadim4j@gmail.com>
> > ---
> >  misc/ss.c | 362 +++++++++++++++++++++++++++++++++++++++-----------------------
> >  1 file changed, 231 insertions(+), 131 deletions(-)
> 
> Hey Vadin,
> 
> your patch do *not* change the output format, right? ss output is
> parsed by scripts and tools.
> 
> hgn

Hi,

It should not for tcp but does for memory info in the 1st patch where
the memeinfo param names were changed.

Regarding parsing of ss by scripts and tools, thats really painful
for me to see how ss outputs additional info, it is not human readable,
actually I'd like to change the output layout in the future, and add
something like online output option '-O'.

May be you can test these patches if they breaks output parsing ?

Anyway I will give up with ss output changes if we really carrying about
to keep the same output for scripts/tools.

Here are some comments from Stephen:
    http://marc.info/?l=linux-netdev&m=142129033800881&w=2

Regards,

^ permalink raw reply

* Re: [PATCH iproute2 3/3] ss: Unify tcp stats output
From: Hagen Paul Pfeifer @ 2015-01-19 14:28 UTC (permalink / raw)
  To: Vadim Kochan; +Cc: netdev
In-Reply-To: <20150119140421.GA1786@angus-think.wlc.globallogic.com>

Hey Vadin,

to make this short. We already discussed about changing the layout and
Stephen nacked this. My proposal was to key:value the output. Because
nearly all outputed data is already in this format - except the
congestion control algo, ts, sack and tx'ed data. Where I proposed
cc:<algo>. The key:value format has the advantages that the ordering
do not mather anymore, An python parser would be something like split
for whitespaces and later split for colon. Currently parsing this is a
mess, see [1].

Anyway, the more clever idea is to add an json outputer like already
supported by some ss modules and get rid of this mess.

Hagen


[1] https://github.com/hgn/captcp/blob/master/captcp.py#L4861

^ permalink raw reply

* Re: [PATCH iproute2 3/3] ss: Unify tcp stats output
From: Vadim Kochan @ 2015-01-19 14:28 UTC (permalink / raw)
  To: Hagen Paul Pfeifer; +Cc: Vadim Kochan, netdev
In-Reply-To: <CAPh34me7KFWnwgOaJmWZg3NxtzOZVce3VYShSy5PfeUkYkJcag@mail.gmail.com>

On Mon, Jan 19, 2015 at 03:28:21PM +0100, Hagen Paul Pfeifer wrote:
> Hey Vadin,
> 
> to make this short. We already discussed about changing the layout and
> Stephen nacked this. My proposal was to key:value the output. Because
> nearly all outputed data is already in this format - except the
> congestion control algo, ts, sack and tx'ed data. Where I proposed
> cc:<algo>. The key:value format has the advantages that the ordering
> do not mather anymore, An python parser would be something like split
> for whitespaces and later split for colon. Currently parsing this is a
> mess, see [1].
> 
> Anyway, the more clever idea is to add an json outputer like already
> supported by some ss modules and get rid of this mess.
> 
> Hagen
> 
> 
> [1] https://github.com/hgn/captcp/blob/master/captcp.py#L4861

Seems these patches will break your script at least for memory info
output.

I am thinking may be 1st of all it is better to make output in json
format and after this trying to make changes with human readability.

What do you think ?

Regards,
Vadim Kochan

^ permalink raw reply

* re: ipvlan: Initial check-in of the IPVLAN driver.
From: Dan Carpenter @ 2015-01-19 14:40 UTC (permalink / raw)
  To: maheshb; +Cc: netdev

Hello Mahesh Bandewar,

The patch 2ad7bf363841: "ipvlan: Initial check-in of the IPVLAN
driver." from Nov 23, 2014, leads to the following static checker
warning:

	drivers/net/ipvlan/ipvlan_core.c:380 ipvlan_process_v6_outbound()
	warn: 'dst' isn't an ERR_PTR

drivers/net/ipvlan/ipvlan_core.c
   378  
   379          dst = ip6_route_output(dev_net(dev), NULL, &fl6);
   380          if (IS_ERR(dst))
   381                  goto err;

The ip6_route_output() function is not documented but it always returns
a valid pointer.  I believe you are supposed to check something like:

		if (dst->error) {
			ret = dst->error;
			goto error;
		}

   382  
   383          skb_dst_drop(skb);
   384          skb_dst_set(skb, dst);

regards,
dan carpenter

^ permalink raw reply

* Re: [patch-net-next v3 2/2] net: ethernet: cpsw: don't requests IRQs we don't use
From: Felipe Balbi @ 2015-01-19 14:40 UTC (permalink / raw)
  To: David Miller; +Cc: balbi, tony, linux-omap, mugunthanvnm, netdev
In-Reply-To: <20150118.010750.480741880077015933.davem@davemloft.net>

[-- Attachment #1: Type: text/plain, Size: 663 bytes --]

On Sun, Jan 18, 2015 at 01:07:50AM -0500, David Miller wrote:
> From: Felipe Balbi <balbi@ti.com>
> Date: Fri, 16 Jan 2015 10:11:12 -0600
> 
> > CPSW never uses RX_THRESHOLD or MISC interrupts. In
> > fact, they are always kept masked in their appropriate
> > IRQ Enable register.
> > 
> > Instead of allocating an IRQ that never fires, it's best
> > to remove that code altogether and let future patches
> > implement it if anybody needs those.
> > 
> > Signed-off-by: Felipe Balbi <balbi@ti.com>
> 
> Applied.

looks like randconfig caught a build break. Do you want an incremental
patch or this patch again with the fix in it ?

-- 
balbi

[-- Attachment #2: Digital signature --]
[-- Type: application/pgp-signature, Size: 819 bytes --]

^ permalink raw reply

* Re: [PATCH mac80211-next] nl80211: Allow set network namespace by fd
From: Vadim Kochan @ 2015-01-19 14:34 UTC (permalink / raw)
  To: Johannes Berg; +Cc: Vadim Kochan, Eric W. Biederman, linux-wireless, netdev
In-Reply-To: <1421225246.1950.22.camel@sipsolutions.net>

On Wed, Jan 14, 2015 at 09:47:26AM +0100, Johannes Berg wrote:
> On Mon, 2015-01-12 at 16:34 +0200, Vadim Kochan wrote:
> 
> > --- a/net/core/net_namespace.c
> > +++ b/net/core/net_namespace.c
> > @@ -361,6 +361,7 @@ struct net *get_net_ns_by_fd(int fd)
> >  	return ERR_PTR(-EINVAL);
> >  }
> >  #endif
> > +EXPORT_SYMBOL_GPL(get_net_ns_by_fd);
> 
> Does this seem OK? Vadim is adding support for using the ns-by-fd in
> nl80211, which can be a module as part of cfg80211.
> 
> johannes
> 
PING ... in case if this email was missed ...

Thanks,

^ permalink raw reply

* Re: [PATCH iproute2 3/3] ss: Unify tcp stats output
From: Daniel Borkmann @ 2015-01-19 14:48 UTC (permalink / raw)
  To: Hagen Paul Pfeifer; +Cc: Vadim Kochan, netdev, stephen
In-Reply-To: <CAPh34me7KFWnwgOaJmWZg3NxtzOZVce3VYShSy5PfeUkYkJcag@mail.gmail.com>

On 01/19/2015 03:28 PM, Hagen Paul Pfeifer wrote:
> Hey Vadin,
>
> to make this short. We already discussed about changing the layout and
> Stephen nacked this. My proposal was to key:value the output. Because
> nearly all outputed data is already in this format - except the
> congestion control algo, ts, sack and tx'ed data. Where I proposed
> cc:<algo>. The key:value format has the advantages that the ordering
> do not mather anymore, An python parser would be something like split
> for whitespaces and later split for colon. Currently parsing this is a
> mess, see [1].
>
> Anyway, the more clever idea is to add an json outputer like already
> supported by some ss modules and get rid of this mess.

+1

I was also thinking in addition to json, that it might be useful to have
an optional ncurses top-like mode in ss. The level of detail could be
unfolded for a specific entry on demand, etc. I would not add it as a
hard library requirement, but in case ncurses headers are detected by
the configure script, it could be compiled in then. It's also easily
changeable since there's no such requirement that the way data is being
displayed needs to be stable for scripts.

> Hagen
>
> [1] https://github.com/hgn/captcp/blob/master/captcp.py#L4861

^ permalink raw reply

* Re: [PATCH iproute2 3/3] ss: Unify tcp stats output
From: Vadim Kochan @ 2015-01-19 14:50 UTC (permalink / raw)
  To: Daniel Borkmann; +Cc: Hagen Paul Pfeifer, Vadim Kochan, netdev, stephen
In-Reply-To: <54BD1958.6030403@redhat.com>

On Mon, Jan 19, 2015 at 03:48:56PM +0100, Daniel Borkmann wrote:
> On 01/19/2015 03:28 PM, Hagen Paul Pfeifer wrote:
> >Hey Vadin,
> >
> >to make this short. We already discussed about changing the layout and
> >Stephen nacked this. My proposal was to key:value the output. Because
> >nearly all outputed data is already in this format - except the
> >congestion control algo, ts, sack and tx'ed data. Where I proposed
> >cc:<algo>. The key:value format has the advantages that the ordering
> >do not mather anymore, An python parser would be something like split
> >for whitespaces and later split for colon. Currently parsing this is a
> >mess, see [1].
> >
> >Anyway, the more clever idea is to add an json outputer like already
> >supported by some ss modules and get rid of this mess.
> 
> +1
> 
> I was also thinking in addition to json, that it might be useful to have
> an optional ncurses top-like mode in ss. The level of detail could be
> unfolded for a specific entry on demand, etc. I would not add it as a
> hard library requirement, but in case ncurses headers are detected by
> the configure script, it could be compiled in then. It's also easily
> changeable since there's no such requirement that the way data is being
> displayed needs to be stable for scripts.
> 
> >Hagen
> >
> >[1] https://github.com/hgn/captcp/blob/master/captcp.py#L4861

OK I will re-work series to do only refactoring/cleanups. And in future
I will keep in mind about any surprises for ss parsers.

Regards,

^ permalink raw reply

* Re: [PATCH iproute2 3/3] ss: Unify tcp stats output
From: Hagen Paul Pfeifer @ 2015-01-19 15:01 UTC (permalink / raw)
  To: Vadim Kochan; +Cc: netdev
In-Reply-To: <20150119142803.GA5549@angus-think.wlc.globallogic.com>

On 19 January 2015 at 15:28, Vadim Kochan <vadim4j@gmail.com> wrote:

> I am thinking may be 1st of all it is better to make output in json
> format and after this trying to make changes with human readability.
>
> What do you think ?

Mhh, if an JSON outputer is also provided a prioi *I* had no problems
with an incompatible change. It is awful to parse the current output.
*But* I am not sure if this opinion is shared by Stephen too. We
probably break several scripts/applications with this change. Breaking
the output *and* do not provide an stable alternative (JSON) is bad.

+1 for JSON (which is not that hard to implement)

hgn

^ permalink raw reply

* Re: [RFC PATCH] net: ipv6: Make address flushing on ifdown optional
From: Hannes Frederic Sowa @ 2015-01-19 15:02 UTC (permalink / raw)
  To: David Ahern; +Cc: netdev
In-Reply-To: <1421263039-96198-1-git-send-email-dsahern@gmail.com>

On Mi, 2015-01-14 at 12:17 -0700, David Ahern wrote:
> Currently, ipv6 addresses are flushed when the interface is configured down:
> 
> [root@f20 ~]# ip -6 addr add dev eth1 2000:11:1:1::1/64
> [root@f20 ~]# ip addr show dev eth1
> 3: eth1: <BROADCAST,MULTICAST> mtu 1500 qdisc noop state DOWN group default qlen 1000
>     link/ether 02:04:11:22:33:01 brd ff:ff:ff:ff:ff:ff
>     inet6 2000:11:1:1::1/64 scope global tentative
>        valid_lft forever preferred_lft forever
> [root@f20 ~]# ip link set dev eth1 up
> [root@f20 ~]# ip link set dev eth1 down
> [root@f20 ~]# ip addr show dev eth1
> 3: eth1: <BROADCAST,MULTICAST> mtu 1500 qdisc pfifo_fast state DOWN group default qlen 1000
>     link/ether 02:04:11:22:33:01 brd ff:ff:ff:ff:ff:ff
> 
> Add a new sysctl to make this behavior optional. Setting defaults to flush
> addresses to maintain backwards compatibility. When reset flushing is bypassed:
> 
> [root@f20 ~]# echo 0 > /proc/sys/net/ipv6/conf/eth1/flush_addr_on_down
> [root@f20 ~]# ip -6 addr add dev eth1 2000:11:1:1::1/64
> [root@f20 ~]# ip addr show dev eth1
> 3: eth1: <BROADCAST,MULTICAST> mtu 1500 qdisc pfifo_fast state DOWN group default qlen 1000
>     link/ether 02:04:11:22:33:01 brd ff:ff:ff:ff:ff:ff
>     inet6 2000:11:1:1::1/64 scope global tentative
>        valid_lft forever preferred_lft forever
> [root@f20 ~]#  ip link set dev eth1 up
> [root@f20 ~]#  ip link set dev eth1 down
> [root@f20 ~]# ip addr show dev eth1
> 3: eth1: <BROADCAST,MULTICAST> mtu 1500 qdisc pfifo_fast state DOWN group default qlen 1000
>     link/ether 02:04:11:22:33:01 brd ff:ff:ff:ff:ff:ff
>     inet6 2000:11:1:1::1/64 scope global
>        valid_lft forever preferred_lft forever
>     inet6 fe80::4:11ff:fe22:3301/64 scope link
>        valid_lft forever preferred_lft forever
> 
> Suggested-by: Hannes Frederic Sowa <hannes@redhat.com>
> Signed-off-by: David Ahern <dsahern@gmail.com>
> Cc: Hannes Frederic Sowa <hannes@redhat.com>
> ---
>  include/linux/ipv6.h      |  1 +
>  include/uapi/linux/ipv6.h |  1 +
>  net/ipv6/addrconf.c       | 15 +++++++++++++++
>  3 files changed, 17 insertions(+)
> 
> diff --git a/include/linux/ipv6.h b/include/linux/ipv6.h
> index c694e7baa621..1d726e39f09f 100644
> --- a/include/linux/ipv6.h
> +++ b/include/linux/ipv6.h
> @@ -52,6 +52,7 @@ struct ipv6_devconf {
>  	__s32		force_tllao;
>  	__s32           ndisc_notify;
>  	__s32		suppress_frag_ndisc;
> +	__s32		flush_addr_on_down;
>  	void		*sysctl;
>  };
>  
> diff --git a/include/uapi/linux/ipv6.h b/include/uapi/linux/ipv6.h
> index e863d088b9a5..c7cb79e0f0fe 100644
> --- a/include/uapi/linux/ipv6.h
> +++ b/include/uapi/linux/ipv6.h
> @@ -165,6 +165,7 @@ enum {
>  	DEVCONF_SUPPRESS_FRAG_NDISC,
>  	DEVCONF_ACCEPT_RA_FROM_LOCAL,
>  	DEVCONF_USE_OPTIMISTIC,
> +	DEVCONF_FLUSH_ON_DOWN,
>  	DEVCONF_MAX
>  };
>  
> diff --git a/net/ipv6/addrconf.c b/net/ipv6/addrconf.c
> index f7c8bbeb27b7..5c0d49073cb1 100644
> --- a/net/ipv6/addrconf.c
> +++ b/net/ipv6/addrconf.c
> @@ -201,6 +201,7 @@ static struct ipv6_devconf ipv6_devconf __read_mostly = {
>  	.disable_ipv6		= 0,
>  	.accept_dad		= 1,
>  	.suppress_frag_ndisc	= 1,
> +	.flush_addr_on_down	= 1,
>  };
>  
>  static struct ipv6_devconf ipv6_devconf_dflt __read_mostly = {
> @@ -238,6 +239,7 @@ static struct ipv6_devconf ipv6_devconf_dflt __read_mostly = {
>  	.disable_ipv6		= 0,
>  	.accept_dad		= 1,
>  	.suppress_frag_ndisc	= 1,
> +	.flush_addr_on_down	= 1,
>  };
>  
>  /* Check if a valid qdisc is available */
> @@ -3083,6 +3085,9 @@ static int addrconf_ifdown(struct net_device *dev, int how)
>  	if (how && del_timer(&idev->regen_timer))
>  		in6_dev_put(idev);
>  
> +	if (!how && !idev->cnf.flush_addr_on_down)
> +		goto unlock;

I would still prefer that we flush automatically generated addresses and
only keep the static and permanent ones.

What do you think?

Bye,
Hannes

^ permalink raw reply

* NetDev 0.1 conference new proposals accepted + misc updates
From: Jamal Hadi Salim @ 2015-01-19 15:28 UTC (permalink / raw)
  To: netdev-u79uwXL29TY76Z2rM5mHXA,
	linux-wireless-u79uwXL29TY76Z2rM5mHXA, lwn-T1hC0tSOHrs,
	netdev01-wool9L35kiczKOhml7GhPkB+6BGkLq7r,
	lartc-u79uwXL29TY76Z2rM5mHXA, netfilter-u79uwXL29TY76Z2rM5mHXA,
	netfilter-devel-u79uwXL29TY76Z2rM5mHXA
  Cc: Richard Guy Briggs, info-4R/QDJtgm2dg9hUCZPvPmw

Fellow netheads:

On behalf of rgb, yours truly is sending out the weekly announcement.

This conference is turning out to be _the best ever_ Linux networking
content put together under one roof. You will be committing a netdev
crime if you dont show up!
Please spread the news. Forward this email appropriately.

First things first.

A reminder that the Westin Hotel is still holding a block of rooms for
Netdev01 at a guaranteed rate of $159.00 or $179.00 (depending on the 
type of room required) and there are still rooms available.
That guarantee expires on January 23, so please dont procastinate
and book now. The rates *WILL* go up.
Reservations: 
https://www.starwoodmeeting.com/StarGroupsWeb/res?id=1412035802&key=1AC9C1F8

Please register soon to help us plan this event better.
Registration https://onlineregistrations.ca/netdev01/
$100/day, or $350 for 4 days (Cdn dollars, reel cheep).
(online reg closes Feb 12th). Again:
Registering helps us plan properly for numbers of attendees,
ensuring venue sizes and supplies are appropriate without
wasting resources.

NetDev 0.1 would like to gratefully acknowledge our sponsors: 
https://netdev01.org/sponsors
Google https://www.google.ca
Qualcomm  https://www.qualcomm.com/
Verizon http://www.verizon.com/
Cumulus Networks http://cumulusnetworks.com/
Mojatatu Networks http://mojatatu.com/

There are a few more talks and tutorials and BOFs/workshops
in the pipeline (we will close acceptance at the end of the week,
so if you want to submit, please do it NOW).

Here is an update on new proposals accepted this past week for NetDev
0.1 that you may have missed if you aren't following the RSS feed or
twitter:

New Sponsors (All listed at: https://www.netdev01.org/sponsors)
=============
Qualcomm - Without sponsors like Qualcomm netdev 0.1 couldnt have
have happened. Thanks Qualcomm for your support of the open source
community.
Google - Google is a giant in the open source world. We thank them
for their sponsorship.


Accepted talks: (All listed: https://www.netdev01.org/sessions )
===============
1) Breaking Open Linux Switching Drivers by Andy Gospodarek
Andy will talk about taking an evolutionary path of a compromise
to integrate between vendor SDKs and the Linux kernel for hardware
offload. Graphics cards all over again? Come and find out.
https://www.netdev01.org/sessions/19

2) Weighing Three Methods of Scaling UDP Applications by Alan Dekok
Alan, the maintainer of FreeRadius, has had to battle UDP for years.
In this talk he will describe his experiences on 3 different approaches
he has recently taken to try and scale UDP applications on Linux
https://www.netdev01.org/sessions/20

3) TC Classifier Action Subsystem by Jamal Hadi Salim
Jamal will describe the tc classifier-action subsystem in details (A 
decade too late?). He claims to be inspired by two talks at netdev01
which refer to this subsystem. He also claims it is a better
architecture than OF and P4. Come beat up on him.
https://www.netdev01.org/sessions/21

4)  MLAG (or Multi-Chassis Link aggregation Group) integration with 
Linux by Matty Kadosh
Mellanox MLAG implementation at  https://github.com/open-ethernet/MLAG
will be discussed in this talk. Matty will talk about experiences and
will try to convince us that it is time to mainstream MLAG.
https://www.netdev01.org/sessions/22

5) MLAG on Linux - Lessons Learned by Scott Emery
Cumulus Networks has extensive MLAG deployment on Linux using hardware 
offloads on high speed hardware.
They have had to battle dragons in many dungeons and slay
many monsters along the way. Come hear all that experience in one talk.
https://www.netdev01.org/sessions/23

Accepted Tutorials (All listed: https://www.netdev01.org/sessions )
===================
1)Hardware Accelerating Linux Network Functions by
  Roopa Prabhu and Toshiaki Makita

Everything you ever wanted to know about using Linux bridging
variations for virtualization, data centre & enterprise, L3, ACLs and 
more. How well do you know how to use bridge/brctl/ip/ethtool and
their friends to setup a desired network(bridge, SRIOV, Macvlan etc)?
Roopa and Toshiaki will try to guide you through this vast ocean of 
knowledge. You will get to learn about how to learn the theory
and practise and beauty of using these toolsets consistently
whether on a NIC, offloaded-NIC or multi-terabit capacity switch
hardware.
https://www.netdev01.org/sessions/24


More coming in future updates.
If youve got this far reading - please go and register.

cheers,
jamal (on behalf of Richard Guy Briggs)
--
To unsubscribe from this list: send the line "unsubscribe linux-wireless" in
the body of a message to majordomo-u79uwXL29TY76Z2rM5mHXA@public.gmane.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html

^ permalink raw reply

* [patch iproute2 repost 2/2] tc: add support for BPF based actions
From: Jiri Pirko @ 2015-01-19 15:56 UTC (permalink / raw)
  To: netdev; +Cc: jhs, stephen
In-Reply-To: <1421682990-11072-1-git-send-email-jiri@resnulli.us>

Signed-off-by: Jiri Pirko <jiri@resnulli.us>
---
 include/linux/tc_act/tc_bpf.h |  31 +++++++
 tc/Makefile                   |   1 +
 tc/m_bpf.c                    | 183 ++++++++++++++++++++++++++++++++++++++++++
 3 files changed, 215 insertions(+)
 create mode 100644 include/linux/tc_act/tc_bpf.h
 create mode 100644 tc/m_bpf.c

diff --git a/include/linux/tc_act/tc_bpf.h b/include/linux/tc_act/tc_bpf.h
new file mode 100644
index 0000000..5288bd7
--- /dev/null
+++ b/include/linux/tc_act/tc_bpf.h
@@ -0,0 +1,31 @@
+/*
+ * Copyright (c) 2015 Jiri Pirko <jiri@resnulli.us>
+ *
+ * This program is free software; you can redistribute it and/or modify
+ * it under the terms of the GNU General Public License as published by
+ * the Free Software Foundation; either version 2 of the License, or
+ * (at your option) any later version.
+ */
+
+#ifndef __LINUX_TC_BPF_H
+#define __LINUX_TC_BPF_H
+
+#include <linux/pkt_cls.h>
+
+#define TCA_ACT_BPF 13
+
+struct tc_act_bpf {
+	tc_gen;
+};
+
+enum {
+	TCA_ACT_BPF_UNSPEC,
+	TCA_ACT_BPF_TM,
+	TCA_ACT_BPF_PARMS,
+	TCA_ACT_BPF_OPS_LEN,
+	TCA_ACT_BPF_OPS,
+	__TCA_ACT_BPF_MAX,
+};
+#define TCA_ACT_BPF_MAX (__TCA_ACT_BPF_MAX - 1)
+
+#endif
diff --git a/tc/Makefile b/tc/Makefile
index 15f68ce..d831a15 100644
--- a/tc/Makefile
+++ b/tc/Makefile
@@ -46,6 +46,7 @@ TCMODULES += m_skbedit.o
 TCMODULES += m_csum.o
 TCMODULES += m_simple.o
 TCMODULES += m_vlan.o
+TCMODULES += m_bpf.o
 TCMODULES += p_ip.o
 TCMODULES += p_icmp.o
 TCMODULES += p_tcp.o
diff --git a/tc/m_bpf.c b/tc/m_bpf.c
new file mode 100644
index 0000000..611135e
--- /dev/null
+++ b/tc/m_bpf.c
@@ -0,0 +1,183 @@
+/*
+ * m_bpf.c	BFP based action module
+ *
+ *              This program is free software; you can redistribute it and/or
+ *              modify it under the terms of the GNU General Public License
+ *              as published by the Free Software Foundation; either version
+ *              2 of the License, or (at your option) any later version.
+ *
+ * Authors:     Jiri Pirko <jiri@resnulli.us>
+ */
+
+#include <stdio.h>
+#include <stdlib.h>
+#include <unistd.h>
+#include <string.h>
+#include <stdbool.h>
+#include <linux/tc_act/tc_bpf.h>
+
+#include "utils.h"
+#include "rt_names.h"
+#include "tc_util.h"
+#include "tc_bpf.h"
+
+static void explain(void)
+{
+	fprintf(stderr, "Usage: ... bpf ...\n");
+	fprintf(stderr, "\n");
+	fprintf(stderr, " [inline]:     run bytecode BPF_BYTECODE\n");
+	fprintf(stderr, " [from file]:  run bytecode-file FILE\n");
+	fprintf(stderr, "\n");
+	fprintf(stderr, "Where BPF_BYTECODE := \'s,c t f k,c t f k,c t f k,...\'\n");
+	fprintf(stderr, "      c,t,f,k and s are decimals; s denotes number of 4-tuples\n");
+	fprintf(stderr, "Where FILE points to a file containing the BPF_BYTECODE string\n");
+	fprintf(stderr, "\nACTION_SPEC := ... look at individual actions\n");
+	fprintf(stderr, "NOTE: CLASSID is parsed as hexadecimal input.\n");
+}
+
+static void usage(void)
+{
+	explain();
+	exit(-1);
+}
+
+static int parse_bpf(struct action_util *a, int *argc_p, char ***argv_p,
+		     int tca_id, struct nlmsghdr *n)
+{
+	int argc = *argc_p;
+	char **argv = *argv_p;
+	struct rtattr *tail;
+	struct tc_act_bpf parm = { 0 };
+	struct sock_filter bpf_ops[BPF_MAXINSNS];
+	__u16 bpf_len = 0;
+
+	if (matches(*argv, "bpf") != 0)
+		return -1;
+
+	NEXT_ARG();
+
+	while (argc > 0) {
+		if (matches(*argv, "run") == 0) {
+			bool from_file;
+			int ret;
+
+			NEXT_ARG();
+			if (strcmp(*argv, "bytecode-file") == 0) {
+				from_file = true;
+			} else if (strcmp(*argv, "bytecode") == 0) {
+				from_file = false;
+			} else {
+				fprintf(stderr, "unexpected \"%s\"\n", *argv);
+				explain();
+				return -1;
+			}
+			NEXT_ARG();
+			ret = bpf_parse_ops(argc, argv, bpf_ops, from_file);
+			if (ret < 0) {
+				fprintf(stderr, "Illegal \"bytecode\"\n");
+				return -1;
+			}
+			bpf_len = ret;
+		} else if (matches(*argv, "help") == 0) {
+			usage();
+		} else {
+			break;
+		}
+		argc--;
+		argv++;
+	}
+
+	parm.action = TC_ACT_PIPE;
+	if (argc) {
+		if (matches(*argv, "reclassify") == 0) {
+			parm.action = TC_ACT_RECLASSIFY;
+			NEXT_ARG();
+		} else if (matches(*argv, "pipe") == 0) {
+			parm.action = TC_ACT_PIPE;
+			NEXT_ARG();
+		} else if (matches(*argv, "drop") == 0 ||
+			   matches(*argv, "shot") == 0) {
+			parm.action = TC_ACT_SHOT;
+			NEXT_ARG();
+		} else if (matches(*argv, "continue") == 0) {
+			parm.action = TC_ACT_UNSPEC;
+			NEXT_ARG();
+		} else if (matches(*argv, "pass") == 0) {
+			parm.action = TC_ACT_OK;
+			NEXT_ARG();
+		}
+	}
+
+	if (argc) {
+		if (matches(*argv, "index") == 0) {
+			NEXT_ARG();
+			if (get_u32(&parm.index, *argv, 10)) {
+				fprintf(stderr, "bpf: Illegal \"index\"\n");
+				return -1;
+			}
+			argc--;
+			argv++;
+		}
+	}
+
+	if (!bpf_len) {
+		fprintf(stderr, "bpf: Bytecode needs to be passed\n");
+		explain();
+		return -1;
+	}
+
+	tail = NLMSG_TAIL(n);
+	addattr_l(n, MAX_MSG, tca_id, NULL, 0);
+	addattr_l(n, MAX_MSG, TCA_ACT_BPF_PARMS, &parm, sizeof(parm));
+	addattr16(n, MAX_MSG, TCA_ACT_BPF_OPS_LEN, bpf_len);
+	addattr_l(n, MAX_MSG, TCA_ACT_BPF_OPS, &bpf_ops,
+		  bpf_len * sizeof(struct sock_filter));
+	tail->rta_len = (char *)NLMSG_TAIL(n) - (char *)tail;
+
+	*argc_p = argc;
+	*argv_p = argv;
+	return 0;
+}
+
+static int print_bpf(struct action_util *au, FILE *f, struct rtattr *arg)
+{
+	struct rtattr *tb[TCA_ACT_BPF_MAX + 1];
+	struct tc_act_bpf *parm;
+
+	if (arg == NULL)
+		return -1;
+
+	parse_rtattr_nested(tb, TCA_ACT_BPF_MAX, arg);
+
+	if (!tb[TCA_ACT_BPF_PARMS]) {
+		fprintf(f, "[NULL bpf parameters]");
+		return -1;
+	}
+	parm = RTA_DATA(tb[TCA_ACT_BPF_PARMS]);
+
+	fprintf(f, " bpf ");
+
+	if (tb[TCA_ACT_BPF_OPS] && tb[TCA_ACT_BPF_OPS_LEN])
+		bpf_print_ops(f, tb[TCA_ACT_BPF_OPS],
+			      rta_getattr_u16(tb[TCA_ACT_BPF_OPS_LEN]));
+
+	fprintf(f, "\n\tindex %d ref %d bind %d", parm->index, parm->refcnt,
+		parm->bindcnt);
+
+	if (show_stats) {
+		if (tb[TCA_ACT_BPF_TM]) {
+			struct tcf_t *tm = RTA_DATA(tb[TCA_ACT_BPF_TM]);
+			print_tm(f, tm);
+		}
+	}
+
+	fprintf(f, "\n ");
+
+	return 0;
+}
+
+struct action_util bpf_action_util = {
+	.id = "bpf",
+	.parse_aopt = parse_bpf,
+	.print_aopt = print_bpf,
+};
-- 
1.9.3

^ permalink raw reply related

* [patch iproute2 repost 1/2] tc: push bpf common code into separate file
From: Jiri Pirko @ 2015-01-19 15:56 UTC (permalink / raw)
  To: netdev; +Cc: jhs, stephen

Signed-off-by: Jiri Pirko <jiri@resnulli.us>
---
 tc/Makefile |   2 +-
 tc/f_bpf.c  | 136 +++++--------------------------------------------------
 tc/tc_bpf.c | 146 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
 tc/tc_bpf.h |  28 ++++++++++++
 4 files changed, 186 insertions(+), 126 deletions(-)
 create mode 100644 tc/tc_bpf.c
 create mode 100644 tc/tc_bpf.h

diff --git a/tc/Makefile b/tc/Makefile
index 9412094..15f68ce 100644
--- a/tc/Makefile
+++ b/tc/Makefile
@@ -1,5 +1,5 @@
 TCOBJ= tc.o tc_qdisc.o tc_class.o tc_filter.o tc_util.o \
-       tc_monitor.o m_police.o m_estimator.o m_action.o \
+       tc_monitor.o tc_bpf.o m_police.o m_estimator.o m_action.o \
        m_ematch.o emp_ematch.yacc.o emp_ematch.lex.o
 
 include ../Config
diff --git a/tc/f_bpf.c b/tc/f_bpf.c
index 48635a7..e2af94e 100644
--- a/tc/f_bpf.c
+++ b/tc/f_bpf.c
@@ -26,6 +26,7 @@
 
 #include "utils.h"
 #include "tc_util.h"
+#include "tc_bpf.h"
 
 static void explain(void)
 {
@@ -44,130 +45,6 @@ static void explain(void)
 	fprintf(stderr, "NOTE: CLASSID is parsed as hexadecimal input.\n");
 }
 
-static int bpf_parse_string(char *arg, bool from_file, __u16 *bpf_len,
-			    char **bpf_string, bool *need_release,
-			    const char separator)
-{
-	char sp;
-
-	if (from_file) {
-		size_t tmp_len, op_len = sizeof("65535 255 255 4294967295,");
-		char *tmp_string;
-		FILE *fp;
-
-		tmp_len = sizeof("4096,") + BPF_MAXINSNS * op_len;
-		tmp_string = malloc(tmp_len);
-		if (tmp_string == NULL)
-			return -ENOMEM;
-
-		memset(tmp_string, 0, tmp_len);
-
-		fp = fopen(arg, "r");
-		if (fp == NULL) {
-			perror("Cannot fopen");
-			free(tmp_string);
-			return -ENOENT;
-		}
-
-		if (!fgets(tmp_string, tmp_len, fp)) {
-			free(tmp_string);
-			fclose(fp);
-			return -EIO;
-		}
-
-		fclose(fp);
-
-		*need_release = true;
-		*bpf_string = tmp_string;
-	} else {
-		*need_release = false;
-		*bpf_string = arg;
-	}
-
-	if (sscanf(*bpf_string, "%hu%c", bpf_len, &sp) != 2 ||
-	    sp != separator) {
-		if (*need_release)
-			free(*bpf_string);
-		return -EINVAL;
-	}
-
-	return 0;
-}
-
-static int bpf_parse_ops(int argc, char **argv, struct nlmsghdr *n,
-			 bool from_file)
-{
-	char *bpf_string, *token, separator = ',';
-	struct sock_filter bpf_ops[BPF_MAXINSNS];
-	int ret = 0, i = 0;
-	bool need_release;
-	__u16 bpf_len = 0;
-
-	if (argc < 1)
-		return -EINVAL;
-	if (bpf_parse_string(argv[0], from_file, &bpf_len, &bpf_string,
-			     &need_release, separator))
-		return -EINVAL;
-	if (bpf_len == 0 || bpf_len > BPF_MAXINSNS) {
-		ret = -EINVAL;
-		goto out;
-	}
-
-	token = bpf_string;
-	while ((token = strchr(token, separator)) && (++token)[0]) {
-		if (i >= bpf_len) {
-			fprintf(stderr, "Real program length exceeds encoded "
-				"length parameter!\n");
-			ret = -EINVAL;
-			goto out;
-		}
-
-		if (sscanf(token, "%hu %hhu %hhu %u,",
-			   &bpf_ops[i].code, &bpf_ops[i].jt,
-			   &bpf_ops[i].jf, &bpf_ops[i].k) != 4) {
-			fprintf(stderr, "Error at instruction %d!\n", i);
-			ret = -EINVAL;
-			goto out;
-		}
-
-		i++;
-	}
-
-	if (i != bpf_len) {
-		fprintf(stderr, "Parsed program length is less than encoded"
-			"length parameter!\n");
-		ret = -EINVAL;
-		goto out;
-	}
-
-	addattr_l(n, MAX_MSG, TCA_BPF_OPS_LEN, &bpf_len, sizeof(bpf_len));
-	addattr_l(n, MAX_MSG, TCA_BPF_OPS, &bpf_ops,
-		  bpf_len * sizeof(struct sock_filter));
-out:
-	if (need_release)
-		free(bpf_string);
-
-	return ret;
-}
-
-static void bpf_print_ops(FILE *f, struct rtattr *bpf_ops, __u16 len)
-{
-	struct sock_filter *ops = (struct sock_filter *) RTA_DATA(bpf_ops);
-	int i;
-
-	if (len == 0)
-		return;
-
-	fprintf(f, "bytecode \'%u,", len);
-
-	for (i = 0; i < len - 1; i++)
-		fprintf(f, "%hu %hhu %hhu %u,", ops[i].code, ops[i].jt,
-			ops[i].jf, ops[i].k);
-
-	fprintf(f, "%hu %hhu %hhu %u\'\n", ops[i].code, ops[i].jt,
-		ops[i].jf, ops[i].k);
-}
-
 static int bpf_parse_opt(struct filter_util *qu, char *handle,
 			 int argc, char **argv, struct nlmsghdr *n)
 {
@@ -195,6 +72,10 @@ static int bpf_parse_opt(struct filter_util *qu, char *handle,
 	while (argc > 0) {
 		if (matches(*argv, "run") == 0) {
 			bool from_file;
+			struct sock_filter bpf_ops[BPF_MAXINSNS];
+			__u16 bpf_len;
+			int ret;
+
 			NEXT_ARG();
 			if (strcmp(*argv, "bytecode-file") == 0) {
 				from_file = true;
@@ -206,10 +87,15 @@ static int bpf_parse_opt(struct filter_util *qu, char *handle,
 				return -1;
 			}
 			NEXT_ARG();
-			if (bpf_parse_ops(argc, argv, n, from_file)) {
+			ret = bpf_parse_ops(argc, argv, bpf_ops, from_file);
+			if (ret < 0) {
 				fprintf(stderr, "Illegal \"bytecode\"\n");
 				return -1;
 			}
+			bpf_len = ret;
+			addattr16(n, MAX_MSG, TCA_BPF_OPS_LEN, bpf_len);
+			addattr_l(n, MAX_MSG, TCA_BPF_OPS, &bpf_ops,
+				  bpf_len * sizeof(struct sock_filter));
 		} else if (matches(*argv, "classid") == 0 ||
 			   strcmp(*argv, "flowid") == 0) {
 			unsigned handle;
diff --git a/tc/tc_bpf.c b/tc/tc_bpf.c
new file mode 100644
index 0000000..c6901d6
--- /dev/null
+++ b/tc/tc_bpf.c
@@ -0,0 +1,146 @@
+/*
+ * tc_bpf.c	BPF common code
+ *
+ *		This program is free software; you can distribute it and/or
+ *		modify it under the terms of the GNU General Public License
+ *		as published by the Free Software Foundation; either version
+ *		2 of the License, or (at your option) any later version.
+ *
+ * Authors:	Daniel Borkmann <dborkman@redhat.com>
+ *		Jiri Pirko <jiri@resnulli.us>
+ */
+
+#include <stdio.h>
+#include <stdlib.h>
+#include <unistd.h>
+#include <string.h>
+#include <stdbool.h>
+#include <errno.h>
+#include <linux/filter.h>
+#include <linux/netlink.h>
+#include <linux/rtnetlink.h>
+
+#include "utils.h"
+#include "tc_util.h"
+#include "tc_bpf.h"
+
+int bpf_parse_string(char *arg, bool from_file, __u16 *bpf_len,
+		     char **bpf_string, bool *need_release,
+		     const char separator)
+{
+	char sp;
+
+	if (from_file) {
+		size_t tmp_len, op_len = sizeof("65535 255 255 4294967295,");
+		char *tmp_string;
+		FILE *fp;
+
+		tmp_len = sizeof("4096,") + BPF_MAXINSNS * op_len;
+		tmp_string = malloc(tmp_len);
+		if (tmp_string == NULL)
+			return -ENOMEM;
+
+		memset(tmp_string, 0, tmp_len);
+
+		fp = fopen(arg, "r");
+		if (fp == NULL) {
+			perror("Cannot fopen");
+			free(tmp_string);
+			return -ENOENT;
+		}
+
+		if (!fgets(tmp_string, tmp_len, fp)) {
+			free(tmp_string);
+			fclose(fp);
+			return -EIO;
+		}
+
+		fclose(fp);
+
+		*need_release = true;
+		*bpf_string = tmp_string;
+	} else {
+		*need_release = false;
+		*bpf_string = arg;
+	}
+
+	if (sscanf(*bpf_string, "%hu%c", bpf_len, &sp) != 2 ||
+	    sp != separator) {
+		if (*need_release)
+			free(*bpf_string);
+		return -EINVAL;
+	}
+
+	return 0;
+}
+
+int bpf_parse_ops(int argc, char **argv, struct sock_filter *bpf_ops,
+		  bool from_file)
+{
+	char *bpf_string, *token, separator = ',';
+	int ret = 0, i = 0;
+	bool need_release;
+	__u16 bpf_len = 0;
+
+	if (argc < 1)
+		return -EINVAL;
+	if (bpf_parse_string(argv[0], from_file, &bpf_len, &bpf_string,
+			     &need_release, separator))
+		return -EINVAL;
+	if (bpf_len == 0 || bpf_len > BPF_MAXINSNS) {
+		ret = -EINVAL;
+		goto out;
+	}
+
+	token = bpf_string;
+	while ((token = strchr(token, separator)) && (++token)[0]) {
+		if (i >= bpf_len) {
+			fprintf(stderr, "Real program length exceeds encoded "
+				"length parameter!\n");
+			ret = -EINVAL;
+			goto out;
+		}
+
+		if (sscanf(token, "%hu %hhu %hhu %u,",
+			   &bpf_ops[i].code, &bpf_ops[i].jt,
+			   &bpf_ops[i].jf, &bpf_ops[i].k) != 4) {
+			fprintf(stderr, "Error at instruction %d!\n", i);
+			ret = -EINVAL;
+			goto out;
+		}
+
+		i++;
+	}
+
+	if (i != bpf_len) {
+		fprintf(stderr, "Parsed program length is less than encoded"
+			"length parameter!\n");
+		ret = -EINVAL;
+		goto out;
+	}
+	ret = bpf_len;
+
+out:
+	if (need_release)
+		free(bpf_string);
+
+	return ret;
+}
+
+void bpf_print_ops(FILE *f, struct rtattr *bpf_ops, __u16 len)
+{
+	struct sock_filter *ops = (struct sock_filter *) RTA_DATA(bpf_ops);
+	int i;
+
+	if (len == 0)
+		return;
+
+	fprintf(f, "bytecode \'%u,", len);
+
+	for (i = 0; i < len - 1; i++)
+		fprintf(f, "%hu %hhu %hhu %u,", ops[i].code, ops[i].jt,
+			ops[i].jf, ops[i].k);
+
+	fprintf(f, "%hu %hhu %hhu %u\'\n", ops[i].code, ops[i].jt,
+		ops[i].jf, ops[i].k);
+}
diff --git a/tc/tc_bpf.h b/tc/tc_bpf.h
new file mode 100644
index 0000000..08cca92
--- /dev/null
+++ b/tc/tc_bpf.h
@@ -0,0 +1,28 @@
+/*
+ * tc_bpf.h	BPF common code
+ *
+ *		This program is free software; you can distribute it and/or
+ *		modify it under the terms of the GNU General Public License
+ *		as published by the Free Software Foundation; either version
+ *		2 of the License, or (at your option) any later version.
+ *
+ * Authors:	Daniel Borkmann <dborkman@redhat.com>
+ *		Jiri Pirko <jiri@resnulli.us>
+ */
+
+#ifndef _TC_BPF_H_
+#define _TC_BPF_H_ 1
+
+#include <stdio.h>
+#include <linux/filter.h>
+#include <linux/netlink.h>
+#include <linux/rtnetlink.h>
+
+int bpf_parse_string(char *arg, bool from_file, __u16 *bpf_len,
+		     char **bpf_string, bool *need_release,
+		     const char separator);
+int bpf_parse_ops(int argc, char **argv, struct sock_filter *bpf_ops,
+		  bool from_file);
+void bpf_print_ops(FILE *f, struct rtattr *bpf_ops, __u16 len);
+
+#endif
-- 
1.9.3

^ permalink raw reply related

* Re: [PATCH net-next] tipc: ratelimit network event traces
From: Jon Maloy @ 2015-01-19 16:07 UTC (permalink / raw)
  To: Erik Hugne, Richard Alpe, ying.xue@windriver.com,
	netdev@vger.kernel.org
  Cc: tipc-discussion@lists.sourceforge.net
In-Reply-To: <1421658164-26185-1-git-send-email-erik.hugne@ericsson.com>



> -----Original Message-----
> From: Erik Hugne
> Sent: January-19-15 4:03 AM
> To: Richard Alpe; ying.xue@windriver.com; Jon Maloy;
> netdev@vger.kernel.org
> Cc: tipc-discussion@lists.sourceforge.net; Erik Hugne
> Subject: [PATCH net-next] tipc: ratelimit network event traces
> 
> From: Erik Hugne <erik.hugne@ericsson.com>
> 
> If a large number of namespaces is spawned on a node and TIPC is enabled in
> each of these, the excessive printk tracing of network events will cause the
> system to grind down to a near halt.


The patch is ok, but I don't quite understand how this can happen. Are you connecting
dozens of nodes in dozens of namespaces?  How many "Established link''  printouts
are there? Even if there are hundreds of them in a burst, I don't quite understand how
 this can kill the whole system.  Just curious.

///jon



> We fix this by adding ratelimiting to the info/warning logs regarding link state
> and node availability.
> 
> Signed-off-by: Erik Hugne <erik.hugne@ericsson.com>
> Reviewed-by: Ying Xue <ying.xue@windriver.com>
> ---
>  net/tipc/link.c | 21 +++++++++++----------  net/tipc/node.c | 24
> +++++++++++++-----------
>  2 files changed, 24 insertions(+), 21 deletions(-)
> 
> diff --git a/net/tipc/link.c b/net/tipc/link.c index 193bc15..bedb590 100644
> --- a/net/tipc/link.c
> +++ b/net/tipc/link.c
> @@ -538,8 +538,8 @@ static void link_state_event(struct tipc_link *l_ptr,
> unsigned int event)
>  			link_set_timer(l_ptr, cont_intv / 4);
>  			break;
>  		case RESET_MSG:
> -			pr_info("%s<%s>, requested by peer\n",
> link_rst_msg,
> -				l_ptr->name);
> +			pr_info_ratelimited("%s<%s>, requested by peer\n",
> +					    link_rst_msg, l_ptr->name);
>  			tipc_link_reset(l_ptr);
>  			l_ptr->state = RESET_RESET;
>  			l_ptr->fsm_msg_cnt = 0;
> @@ -549,7 +549,8 @@ static void link_state_event(struct tipc_link *l_ptr,
> unsigned int event)
>  			link_set_timer(l_ptr, cont_intv);
>  			break;
>  		default:
> -			pr_err("%s%u in WW state\n", link_unk_evt, event);
> +			pr_err_ratelimited("%s%u in WW state\n",
> link_unk_evt,
> +					   event);
>  		}
>  		break;
>  	case WORKING_UNKNOWN:
> @@ -561,8 +562,8 @@ static void link_state_event(struct tipc_link *l_ptr,
> unsigned int event)
>  			link_set_timer(l_ptr, cont_intv);
>  			break;
>  		case RESET_MSG:
> -			pr_info("%s<%s>, requested by peer while
> probing\n",
> -				link_rst_msg, l_ptr->name);
> +			pr_info_ratelimited("%s<%s>, requested by peer
> while probing\n",
> +					    link_rst_msg, l_ptr->name);
>  			tipc_link_reset(l_ptr);
>  			l_ptr->state = RESET_RESET;
>  			l_ptr->fsm_msg_cnt = 0;
> @@ -588,8 +589,8 @@ static void link_state_event(struct tipc_link *l_ptr,
> unsigned int event)
>  				l_ptr->fsm_msg_cnt++;
>  				link_set_timer(l_ptr, cont_intv / 4);
>  			} else {	/* Link has failed */
> -				pr_warn("%s<%s>, peer not responding\n",
> -					link_rst_msg, l_ptr->name);
> +				pr_warn_ratelimited("%s<%s>, peer not
> responding\n",
> +						    link_rst_msg, l_ptr-
> >name);
>  				tipc_link_reset(l_ptr);
>  				l_ptr->state = RESET_UNKNOWN;
>  				l_ptr->fsm_msg_cnt = 0;
> @@ -1568,9 +1569,9 @@ static void tipc_link_proto_rcv(struct net *net,
> struct tipc_link *l_ptr,
> 
>  		if (msg_linkprio(msg) &&
>  		    (msg_linkprio(msg) != l_ptr->priority)) {
> -			pr_warn("%s<%s>, priority change %u->%u\n",
> -				link_rst_msg, l_ptr->name, l_ptr->priority,
> -				msg_linkprio(msg));
> +			pr_warn_ratelimited("%s<%s>, priority change %u-
> >%u\n",
> +					    link_rst_msg, l_ptr->name,
> +					    l_ptr->priority, msg_linkprio(msg));
>  			l_ptr->priority = msg_linkprio(msg);
>  			tipc_link_reset(l_ptr); /* Enforce change to take
> effect */
>  			break;
> diff --git a/net/tipc/node.c b/net/tipc/node.c index b1eb092..01a03a7 100644
> --- a/net/tipc/node.c
> +++ b/net/tipc/node.c
> @@ -230,8 +230,8 @@ void tipc_node_link_up(struct tipc_node *n_ptr,
> struct tipc_link *l_ptr)
>  	n_ptr->action_flags |= TIPC_NOTIFY_LINK_UP;
>  	n_ptr->link_id = l_ptr->peer_bearer_id << 16 | l_ptr->bearer_id;
> 
> -	pr_info("Established link <%s> on network plane %c\n",
> -		l_ptr->name, l_ptr->net_plane);
> +	pr_info_ratelimited("Established link <%s> on network plane %c\n",
> +			    l_ptr->name, l_ptr->net_plane);
> 
>  	if (!active[0]) {
>  		active[0] = active[1] = l_ptr;
> @@ -239,7 +239,8 @@ void tipc_node_link_up(struct tipc_node *n_ptr,
> struct tipc_link *l_ptr)
>  		goto exit;
>  	}
>  	if (l_ptr->priority < active[0]->priority) {
> -		pr_info("New link <%s> becomes standby\n", l_ptr->name);
> +		pr_info_ratelimited("New link <%s> becomes standby\n",
> +				    l_ptr->name);
>  		goto exit;
>  	}
>  	tipc_link_dup_queue_xmit(active[0], l_ptr); @@ -247,9 +248,10 @@
> void tipc_node_link_up(struct tipc_node *n_ptr, struct tipc_link *l_ptr)
>  		active[0] = l_ptr;
>  		goto exit;
>  	}
> -	pr_info("Old link <%s> becomes standby\n", active[0]->name);
> +	pr_info_ratelimited("Old link <%s> becomes standby\n",
> +active[0]->name);
>  	if (active[1] != active[0])
> -		pr_info("Old link <%s> becomes standby\n", active[1]-
> >name);
> +		pr_info_ratelimited("Old link <%s> becomes standby\n",
> +				    active[1]->name);
>  	active[0] = active[1] = l_ptr;
>  exit:
>  	/* Leave room for changeover header when returning 'mtu' to users:
> */ @@ -297,12 +299,12 @@ void tipc_node_link_down(struct tipc_node
> *n_ptr, struct tipc_link *l_ptr)
>  	n_ptr->link_id = l_ptr->peer_bearer_id << 16 | l_ptr->bearer_id;
> 
>  	if (!tipc_link_is_active(l_ptr)) {
> -		pr_info("Lost standby link <%s> on network plane %c\n",
> -			l_ptr->name, l_ptr->net_plane);
> +		pr_info_ratelimited("Lost standby link <%s> on network
> plane %c\n",
> +				    l_ptr->name, l_ptr->net_plane);
>  		return;
>  	}
> -	pr_info("Lost link <%s> on network plane %c\n",
> -		l_ptr->name, l_ptr->net_plane);
> +	pr_info_ratelimited("Lost link <%s> on network plane %c\n",
> +			    l_ptr->name, l_ptr->net_plane);
> 
>  	active = &n_ptr->active_links[0];
>  	if (active[0] == l_ptr)
> @@ -380,8 +382,8 @@ static void node_lost_contact(struct tipc_node
> *n_ptr)
>  	char addr_string[16];
>  	u32 i;
> 
> -	pr_info("Lost contact with %s\n",
> -		tipc_addr_string_fill(addr_string, n_ptr->addr));
> +	pr_info_ratelimited("Lost contact with %s\n",
> +			    tipc_addr_string_fill(addr_string, n_ptr->addr));
> 
>  	/* Flush broadcast link info associated with lost node */
>  	if (n_ptr->bclink.recv_permitted) {
> --
> 2.1.3


------------------------------------------------------------------------------
New Year. New Location. New Benefits. New Data Center in Ashburn, VA.
GigeNET is offering a free month of service with a new server in Ashburn.
Choose from 2 high performing configs, both with 100TB of bandwidth.
Higher redundancy.Lower latency.Increased capacity.Completely compliant.
http://p.sf.net/sfu/gigenet

^ permalink raw reply

* Re: [net-next PATCH v2 01/12] net: flow_table: create interface for hw match/action tables
From: John Fastabend @ 2015-01-19 16:11 UTC (permalink / raw)
  To: Simon Horman; +Cc: tgraf, sfeldma, netdev, gerlitz.or, jhs, andy, davem
In-Reply-To: <20150119050937.GD5612@vergenet.net>

On 01/18/2015 09:09 PM, Simon Horman wrote:
> [snip]
>
>> --- /dev/null
>> +++ b/net/core/flow_table.c
>> @@ -0,0 +1,942 @@
>> +/*
>> + * include/uapi/linux/if_flow.h - Flow table interface for Switch devices
>
> Hi John,
>
> its seems that the path in the above comment should be
> net/core/flow_table.c.
>
> [snip]
>

Thanks Simon, I've merged all three comments.

-- 
John Fastabend         Intel Corporation

^ permalink raw reply


This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox