Netdev List
 help / color / mirror / Atom feed
* Re: linux-next: Tree for Aug 6 [ wireless | iwlwifi | mac80211 ? ]
From: Sedat Dilek @ 2013-08-06 19:14 UTC (permalink / raw)
  To: Johannes Berg; +Cc: David Miller, Stephen Rothwell, wireless, netdev
In-Reply-To: <1375816128.8219.28.camel@jlt4.sipsolutions.net>

On Tue, Aug 6, 2013 at 9:08 PM, Johannes Berg <johannes@sipsolutions.net> wrote:
> On Tue, 2013-08-06 at 20:35 +0200, Sedat Dilek wrote:
>
>> Attached is a diff comparing all new commits in next-20130805.
>> If one of the commits smells bad to you, please let me know.
>
> Out of that list, only the af_packet changes would seem to have any
> impact on wireless at all.
>

git-bisecting... 2 steps to go...

This one is bad... "af_packet: simplify VLAN frame check in packet_snd"

http://git.kernel.org/cgit/linux/kernel/git/davem/net-next.git/commit/?id=c483e02614551e44ced3fe6eedda8e36d3277ccc

- Sedat -

^ permalink raw reply

* Re: linux-next: Tree for Aug 6 [ wireless | iwlwifi | mac80211 ? ]
From: Johannes Berg @ 2013-08-06 19:08 UTC (permalink / raw)
  To: sedat.dilek; +Cc: David Miller, Stephen Rothwell, wireless, netdev
In-Reply-To: <CA+icZUXpZaHu-FjR6FGnbkG0XEkiukEgYYbDmrSjdqNz4rWCNg@mail.gmail.com>

On Tue, 2013-08-06 at 20:35 +0200, Sedat Dilek wrote:

> Attached is a diff comparing all new commits in next-20130805.
> If one of the commits smells bad to you, please let me know.

Out of that list, only the af_packet changes would seem to have any
impact on wireless at all.

johannes

^ permalink raw reply

* Re: linux-next: Tree for Aug 6 [ wireless | iwlwifi | mac80211 ? ]
From: Sedat Dilek @ 2013-08-06 18:35 UTC (permalink / raw)
  To: Johannes Berg, David Miller
  Cc: Stephen Rothwell, wireless, netdev-u79uwXL29TY76Z2rM5mHXA
In-Reply-To: <CA+icZUV++2+OG=P1SVbNrpHPLyzrfFX0eu8NHrd0nPjkPDUOtQ-JsoAwUIsXosN+BqQ9rBEUg@public.gmane.org>

[-- Attachment #1: Type: text/plain, Size: 2436 bytes --]

On Tue, Aug 6, 2013 at 8:26 PM, Sedat Dilek <sedat.dilek-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org> wrote:
> On Tue, Aug 6, 2013 at 6:51 PM, Sedat Dilek <sedat.dilek-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org> wrote:
>> On Tue, Aug 6, 2013 at 6:07 PM, Sedat Dilek <sedat.dilek-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org> wrote:
>>> On Tue, Aug 6, 2013 at 6:03 PM, Johannes Berg <johannes-cdvu00un1VgdHxzADdlk8Q@public.gmane.org> wrote:
>>>> [reducing mailing lists - Stephen if you want off let us know]
>>>>
>>>> On Tue, 2013-08-06 at 17:58 +0200, Sedat Dilek wrote:
>>>>
>>>>> $ cat /proc/version
>>>>> Linux version 3.11.0-rc4-1-wl-20130805 (sedat.dilek-Re5JQEeQqe8AvxtiuMwx3w@public.gmane.org@fambox)
>>>>> (gcc version 4.6.3 (Ubuntu/Linaro 4.6.3-1ubuntu5) ) #1 SMP Tue Aug 6
>>>>> 17:52:45 CEST 2013
>>>>>
>>>>> This kernel boots fine and has no WiFi problems.
>>>>>
>>>>> > You assume a breakage in a different area than wireless or network in general?
>>>>> >
>>>>>
>>>>> Any wild guess?
>>>>
>>>> Not really, but maybe you can pull the other trees on top again and see
>>>> if those broke it? I don't really think/hope so but it's possible that
>>>> one of my pending pull requests broke it, I guess?
>>>>
>>>
>>> Hmm, I pulled in latest for-john of mac80211(-next) and
>>> iwlwifi-(fixes|next) which did not help.
>>>
>>> I am testing next-20130805 and next-20130802 to identify which
>>> Linux-next release is BROKEN.
>>>
>>
>> GOOD: next-20130801 and next-20130802
>>
>> BROKEN: next-20130805 and next-20130806
>>
>> Comparing diffs between next-20130802 and next-20130805.
>>
>
> [ To Dave and CC netdev ]
>
> [ Full thread see [1] ]
>
> The diff was not very helpful.
>
> $ cd linux-next-20130805
>
> $ git log --oneline v3.11-rc4.. | grep "Merge remote-tracking" | egrep
> 'net-next|slave-dma|wireless'
> 7f84622 Merge remote-tracking branch 'net-next/master' <--- BAD
> (before slave-dma/next)
> 3580d24 Merge remote-tracking branch 'slave-dma/next' <--- GOOD (after
> net-next/master)
> 88825c7 Merge remote-tracking branch 'wireless/master' <--- GOOD
>
> So, the culprit commit seems to be in the merge of 'net-next/master'.
>
> ( BTW, disabling IPv6 in NM/wifi-setup did not change behaviour -
> still disconnected to AP. )
>
> - Sedat -
>
> [1] http://marc.info/?t=137579712800008&r=1&w=2

Attached is a diff comparing all new commits in next-20130805.
If one of the commits smells bad to you, please let me know.

- Sedat -

[-- Attachment #2: list-of-commits-from-net-next_next-20130802-VS-next-20130805.diff --]
[-- Type: application/octet-stream, Size: 1855 bytes --]

--- list-of-commits-from-net-next-in-next-20130802.txt	2013-08-06 20:29:45.312395359 +0200
+++ list-of-commits-from-next-next-in-next-20130805.txt	2013-08-06 20:26:03.829704013 +0200
@@ -1,3 +1,26 @@
+f270701 ax88179_178a: avoid copy of tx tcp packets
+fba3679 fib_rules: reorder struct fib_rules fields
+73f5698 fib_rules: fix suppressor names and default values
+0c0667a vlan: cleanup the usage of vlan_dev_priv(dev)
+9e9402e cnic, bnx2i: Fix bug on some bnx2x devices that don't support iSCSI
+9918d5b bonding: modify only neigh_parms owned by us
+6313480 neighbour: populate neigh_parms on alloc before calling ndo_neigh_setup
+8a849bb net: netlink: minor: remove unused pointer in alloc_pg_vec
+6ef94cf fib_rules: add route suppression based on ifgroup
+d1c53c8 icmpv6_filter: allow ICMPv6 messages with bodies < 4 bytes
+9cc08af icmpv6_filter: fix "_hdr" incorrectly being a pointer
+c483e02 af_packet: simplify VLAN frame check in packet_snd
+cbd89ac af_packet: fix for sending VLAN frames via packet_mmap
+0f75b09 af_packet: when sending ethernet frames, parse header for skb->protocol
+d27fc78 sctp: Don't lookup dst if transport dst is still valid
+1409a93 ethernet: Convert mac address uses of 6 to ETH_ALEN
+574e2af include: Convert ethernet mac address declarations to use ETH_ALEN
+e216975 uapi: Convert some uses of 6 to ETH_ALEN
+3753456 qlcnic: Update version to 5.2.45
+79da4d0 qlcnic: Enable mailbox interface in poll mode when interrupts are not available
+068a8d1 qlcnic: Replace poll mode mailbox interface with interrupt based mailbox interface
+e5c4e6c qlcnic: Interrupt based driver firmware mailbox mechanism
+b9c1198 qlcnic: Enhance diagnostic loopback error codes.
 278b208 bonding: initial RCU conversion
 1507722 bonding: factor out slave id tx code and simplify xmit paths
 78a646c bonding: simplify broadcast_xmit function

^ permalink raw reply

* Re: linux-next: Tree for Aug 6 [ wireless | iwlwifi | mac80211 ? ]
From: Sedat Dilek @ 2013-08-06 18:26 UTC (permalink / raw)
  To: Johannes Berg, David Miller; +Cc: Stephen Rothwell, wireless, netdev
In-Reply-To: <CA+icZUUmQbovkwXcdEp6eu=WJkiuiyO8nDVwCHFUuPDbJqU7Kg@mail.gmail.com>

[-- Attachment #1: Type: text/plain, Size: 2003 bytes --]

On Tue, Aug 6, 2013 at 6:51 PM, Sedat Dilek <sedat.dilek@gmail.com> wrote:
> On Tue, Aug 6, 2013 at 6:07 PM, Sedat Dilek <sedat.dilek@gmail.com> wrote:
>> On Tue, Aug 6, 2013 at 6:03 PM, Johannes Berg <johannes@sipsolutions.net> wrote:
>>> [reducing mailing lists - Stephen if you want off let us know]
>>>
>>> On Tue, 2013-08-06 at 17:58 +0200, Sedat Dilek wrote:
>>>
>>>> $ cat /proc/version
>>>> Linux version 3.11.0-rc4-1-wl-20130805 (sedat.dilek@gmail.com@fambox)
>>>> (gcc version 4.6.3 (Ubuntu/Linaro 4.6.3-1ubuntu5) ) #1 SMP Tue Aug 6
>>>> 17:52:45 CEST 2013
>>>>
>>>> This kernel boots fine and has no WiFi problems.
>>>>
>>>> > You assume a breakage in a different area than wireless or network in general?
>>>> >
>>>>
>>>> Any wild guess?
>>>
>>> Not really, but maybe you can pull the other trees on top again and see
>>> if those broke it? I don't really think/hope so but it's possible that
>>> one of my pending pull requests broke it, I guess?
>>>
>>
>> Hmm, I pulled in latest for-john of mac80211(-next) and
>> iwlwifi-(fixes|next) which did not help.
>>
>> I am testing next-20130805 and next-20130802 to identify which
>> Linux-next release is BROKEN.
>>
>
> GOOD: next-20130801 and next-20130802
>
> BROKEN: next-20130805 and next-20130806
>
> Comparing diffs between next-20130802 and next-20130805.
>

[ To Dave and CC netdev ]

[ Full thread see [1] ]

The diff was not very helpful.

$ cd linux-next-20130805

$ git log --oneline v3.11-rc4.. | grep "Merge remote-tracking" | egrep
'net-next|slave-dma|wireless'
7f84622 Merge remote-tracking branch 'net-next/master' <--- BAD
(before slave-dma/next)
3580d24 Merge remote-tracking branch 'slave-dma/next' <--- GOOD (after
net-next/master)
88825c7 Merge remote-tracking branch 'wireless/master' <--- GOOD

So, the culprit commit seems to be in the merge of 'net-next/master'.

( BTW, disabling IPv6 in NM/wifi-setup did not change behaviour -
still disconnected to AP. )

- Sedat -

[1] http://marc.info/?t=137579712800008&r=1&w=2

[-- Attachment #2: list-of-commits-from-next-next-in-next-20130805.txt --]
[-- Type: text/plain, Size: 10653 bytes --]

f270701 ax88179_178a: avoid copy of tx tcp packets
fba3679 fib_rules: reorder struct fib_rules fields
73f5698 fib_rules: fix suppressor names and default values
0c0667a vlan: cleanup the usage of vlan_dev_priv(dev)
9e9402e cnic, bnx2i: Fix bug on some bnx2x devices that don't support iSCSI
9918d5b bonding: modify only neigh_parms owned by us
6313480 neighbour: populate neigh_parms on alloc before calling ndo_neigh_setup
8a849bb net: netlink: minor: remove unused pointer in alloc_pg_vec
6ef94cf fib_rules: add route suppression based on ifgroup
d1c53c8 icmpv6_filter: allow ICMPv6 messages with bodies < 4 bytes
9cc08af icmpv6_filter: fix "_hdr" incorrectly being a pointer
c483e02 af_packet: simplify VLAN frame check in packet_snd
cbd89ac af_packet: fix for sending VLAN frames via packet_mmap
0f75b09 af_packet: when sending ethernet frames, parse header for skb->protocol
d27fc78 sctp: Don't lookup dst if transport dst is still valid
1409a93 ethernet: Convert mac address uses of 6 to ETH_ALEN
574e2af include: Convert ethernet mac address declarations to use ETH_ALEN
e216975 uapi: Convert some uses of 6 to ETH_ALEN
3753456 qlcnic: Update version to 5.2.45
79da4d0 qlcnic: Enable mailbox interface in poll mode when interrupts are not available
068a8d1 qlcnic: Replace poll mode mailbox interface with interrupt based mailbox interface
e5c4e6c qlcnic: Interrupt based driver firmware mailbox mechanism
b9c1198 qlcnic: Enhance diagnostic loopback error codes.
278b208 bonding: initial RCU conversion
1507722 bonding: factor out slave id tx code and simplify xmit paths
78a646c bonding: simplify broadcast_xmit function
71bc3b2 bonding: remove unnecessary read_locks of curr_slave_lock
dec1e90 bonding: convert to list API and replace bond's custom list
439677d ipv6: bump genid when delete/add address
8b09be5 bnx2x: Revising locking scheme for MAC configuration
4beac02 bonding: fix system hang due to fast igmp timer rescheduling
9ab5ec5 tile: support PTP using the tilegx mPIPE (IEEE 1588)
84e181b tile: remove deprecated NETIF_F_LLTX flag from tile drivers
4aa0264 tile: make "tile_net.custom" a proper bool module parameter
2c7d04a tile: support TSO for IPv6 in tilegx network driver
f3286a3 tile: support multiple mPIPE shims in tilegx network driver
6ab4ae9 tile: enable GRO in the tilegx network driver
5e7a54a tile: fix panic bug in napi support for tilegx network driver
ad01818 tile: update dev->stats directly in tilegx network driver
2628e8a tile: support jumbo frames in the tilegx network driver
48f2a4e tile: remove dead is_dup_ack() function from tilepro net driver
815d3ba tile: avoid bug in tilepro net driver built with old hypervisor
439a93a tile: support rx_dropped/rx_errors in tilepro net driver
a8eaed5 tile: set hw_features and vlan_features in setup
84915c6 gianfar: Remove unused field grp_id from gfar_priv_grp
376c731 net: add a temporary sanity check in skb_orphan()
aa10181 can: flexcan: Check the return value from clk_prepare_enable()
933e4af can: flexcan: Use devm_ioremap_resource()
46b3a42 ipv6: fib6_rules should return exact return value
3783072 cls_cgroup.h netprio_cgroup.h: Remove extern from function prototypes
4fc7074 checksum: Remove extern from function prototypes
10dd9b7 cfg80211.h/mac80211.h: Remove extern from function prototypes
c1d8f80 ax25.h: Remove extern from function prototypes
90972b2 arp/neighbour.h: Remove extern from function prototypes
cd2cf63 af_rxrpc.h: Remove extern from function prototypes
b60a828 af_unix.h: Remove extern from function prototypes
e8e54d3 addrconf.h: Remove extern function prototypes
49dfe76 Documentation: add networking/netdev-FAQ.txt
7764a45 fib_rules: add .suppress operation
5c15257 net: Remove extern from include/net/ scheduling prototypes
c34a761 net: skb_orphan() changes
f2f872f netem: Introduce skb_orphan_partial() helper
ca4c3fc net: split rt_genid for ipv4 and ipv6
ba361cb sh_eth: r8a7790: Handle the RFE (Receive FIFO overflow Error) interrupt
c0155b2 tcp: Remove unused tcpct declarations and comments
8f58332 ixgbe: add support for quad-port x520 adapter
674c18b ixgbe: clear semaphore bits on timeouts
9c432ad ixgbe: rename LL_EXTENDED_STATS to use queue instead of q
8fecf67 ixgbe: fix lockdep annotation issue for ptp's work item
e027d1a ixgbe: call pcie_get_mimimum_link to check if device has enough bandwidth
81377c8 PCI: Add function to obtain minimum link width and speed
4299c8a net: remove an unneeded check
b438f94 flow_dissector: add support for IPPROTO_IPV6
fca4189 flow_dissector: clean up IPIP case
59da381 PCI: move enum pcie_link_width into pci.h
343e51a PCI: expose pcie_link_speed and pcix_bus_speed arrays
a4b6fc6 ixgbe: fix SFF data dumps of SFP+ modules
3dcc2f4 ixgbe: fix semaphore lock for I2C read/writes on 82598
93ac03b ixgbe: bump version number
4e8e1bc ixgbe: add new media type.
ff80e51 net: export physical port id via sysfs
66cae9e rtnl: export physical port id via RT netlink
66b52b0 net: add ndo to get id of physical port of the device
73d8095 ixgbe: fix fc autoneg ethtool reporting.
e507d0c ixgbe: Use pci_vfs_assigned instead of ixgbe_vfs_are_assigned
670224f ixgbe: Retain VLAN filtering in promiscuous + VT mode
9ad8fef net: mvneta: support big endian
6083ed4 net: mvneta: move the RX and TX desc macros outside of the structs
ffd756b pktgen: Require CONFIG_INET due to use of IPv4 checksum function
5ad37d5 tcp: add tcp_syncookies mode to allow unconditionally generation of syncookies
dcfd8d5 drivers: net: cpsw: Add support for set MAC address
d68e2d3 tile: handle 64-bit statistics in tilepro network driver
60ff779 9p: client: remove unused code and any reference to "cancelled" function
4ce1fd6 be2net: don't use dev_err when AER enabling fails
cd77b2e tg3: Update version to 3.133
378b72c tg3: Fix UDP fragments treated as RMCP
92e6457 tg3: Enable support for timesync gpio output
4c305fa tg3: Implement the shutdown handler
5137a2e tg3: Allow NVRAM programming when interface is down
c145935 tg3: Remove incorrect switch to aux power
ca67a3c cnic: Update version to 2.5.17 and copyright year.
28e3a8f cnic: Add missing error checking for RAMROD_CMD_ID_CLOSE
b3bd2d6 cnic: Update TCP options setup for iSCSI.
6cdcdbb cnic: Reset tcp_flags during cnic_cm_create().
b54345e cnic: Simplify cnic_release().
415fb87 cnic: Simplify netdev events handling.
fe6f700 net/mlx4_core: Respond to operation request by firmware
2d4b646 net/mlx4_en: Fix BlueFlame race
73d94e9 pktgen: add needed include file
9d4a031 ipv4, ipv6: send igmpv3/mld packets with TC_PRIO_CONTROL
16b095a e1000e: fix I217/I218 PHY initialization flow
97390ab e1000e: do not resume device from RPM suspend to read PHY status registers
91a3d82 e1000e: enable support for new device IDs
3ef672a e1000e: ethtool unnecessarily takes device out of RPM suspend
e0236ad e1000e: Tx hang on I218 when linked at 100Half and slow response at 10Mbps
ce345e0 e1000e: low throughput using 4K jumbos on I218
da1e204 e1000e: iAMT connections drop on driver unload when jumbo frames enabled
b43e867 e1000e: disable ASPM L1 on 82583
c96ddb0 e1000e: Use marco instead of digit for defining e1000_rx_desc_packet_split
2592881 e1000e: Remove duplicate assignment of default rx/tx ring size
24b41c9 e1000e: restore call to pci_clear_master()
ab90695 e100: dump small buffers via %*ph
dcfe804 bonding: remove bond_resend_igmp_join_requests read_unlock leftover
03c633e pktgen: Use ip_send_check() to compute checksum
c26bf4a pktgen: Add UDPCSUM flag to support UDP checksums
82a54d0 VSOCK: Move af_vsock.h and vsock_addr.h to include/net
452c447 USBNET: increase max rx/tx qlen for improving USB3 thoughtput
a88c32a USBNET: centralize computing of max rx/tx qlen
6680ec6 tuntap: hardware vlan tx support
024ec3d net/sctp: Refactor SCTP skb checksum computation
e7428e9 virtio-net: put virtio net header inline with data
10eccb4 bond: cleanup netpoll code
0fb52a2 team: cleanup netpoll clode
93d8bf9 bridge: cleanup netpoll code
f528094 bonding: use pre-defined macro in bond_mode_name instead of magic number 0
f1a26fd pch_gbe: Add MinnowBoard support
9025c8e drivers/net/ethernet/stmicro/stmmac: don't check resource with devm_ioremap_resource
b04d68e pch_gbe: Use PCH_GBE_PHY_REGS_LEN instead of 32
18afa4b net: Make devnet_rename_seq static
c9bee3b tcp: TCP_NOTSENT_LOWAT socket option
64dc613 net: add sk_stream_is_writeable() helper
4d58c02 net: sctp: trivial: add uapi/linux/sctp.h into maintainers
91705c6 net: sctp: trivial: update mailing list address
d971854 drivers: net: cpsw: add support to show hw stats via ethtool
b07ea07 bonding: Fixed up a error "do not initialise statics to 0 or NULL" in bond_main.c
9402b74 bonding: add rtnl protection for bonding_store_fail_over_mac
38c4916 bonding: bond_sysfs.c checkpatch cleanup
c4cdef9 bonding: don't call slave_xxx_netpoll under spinlocks
f13bbc2 drivers/net: enic: Move ethtool code to a separate file
59ea52d net: trans_rdma: remove unused function
2d17f40 be2net: delete primary MAC address while unloading
3175d8c be2net: use SET/GET_MAC_LIST for SH-R
95046b9 be2net: refactor MAC-addr setup code
b5bb977 be2net: fix pmac_id for BE3 VFs
04a0602 be2net: allow VFs to program MAC and VLAN filters
5a712c1 be2net: fix MAC address modification for VF
e18dbf7 sh_eth: Add support for r8a7790 SoC
55754f1 sh_eth: add support for RMIIMODE register
9225b23 net: ipv6 eliminate parameter "int addrlen" in function fib6_add_1
86a37de net ipv6: Remove rebundant rt6i_nsiblings initialization
b9959fd vti: switch to new ip tunnel code
2b52c3a ip6mr: change the prototype of ip6_mr_forward().
c4854ec ipmr: change the prototype of ip_mr_forward().
492b200 team: add support for sending multicast rejoins
4aa5dee net: convert resend IGMP to notifier event
fc423ff team: add peer notification
ab2cfbb macvlan fdb replace support
906dc18 vxlan fdb replace an existing entry
ed08495 tcp: use RTT from SACK for RTO
59c9af4 tcp: measure RTT from new SACK
5b08e47 tcp: prefer packet timing to TS-ECR for RTT
375fe02 tcp: consolidate SYNACK RTT sampling
0d9b2ab fec: Use devm_request_irq()
399db75 fec: Remove unneeded check in platform_get_resource()
13a097b fec: Check the return value from clk_prepare_enable()
79820e7 fec: Enable/disable clk_ptp in suspend/resume
d265cf4 fec: Fix the order for enabling/disabling the clocks
9514fe7 fec: Do not enable/disable optional clocks unconditionally
eda2977 tun: Support software transmit time stamping.
cb820f8 net: Provide a generic socket error queue delivery method for Tx time stamps.
0887a57 net/velocity: add poll controller function for velocity nic
6e3d677 mISDN: replace sum of bitmasks with OR operation.
aafee33 net/irda: fixed style issues in irttp

^ permalink raw reply

* Re: low latency/busy poll feedback and bugs
From: Eliezer Tamir @ 2013-08-06 18:25 UTC (permalink / raw)
  To: Shawn Bohrer; +Cc: Amir Vadai, netdev
In-Reply-To: <20130806180806.GA8993@sbohrermbp13-local.rgmadvisors.com>

On 06/08/2013 21:08, Shawn Bohrer wrote:
> On Tue, Aug 06, 2013 at 10:41:48AM +0300, Eliezer Tamir wrote:
>> For multicast, it is possible that incoming packets to come from more
>> than one port (and therefore more than one queue).
>> I'm not sure how we could handle that, but what we have today won't do
>> well for that use-case.
>  
> It is unclear to me exactly what happens in this case.  With my simple
> patch I'm assuming it will spin on the receive queue that received the
> last packet for that socket.  What happens when a packet arrives on a
> different receive queue than the one we were spinning on? I assume it
> is still delivered but perhaps the spinning process won't get it until
> the spinning time expires?  I'm just guessing and haven't attempted to
> figure it out from looking through the code.

What will happen is that the current code will only busy poll on one
queue, sometimes on this one, sometimes on that one.

packets arriving on the other queue will still be serviced but will
suffer the latency of waiting for NAPI to schedule.

So your avg will be better, but your std. dev. much worse and it's
probably not worth it if you really expect two devices to receive
data at the same time.

^ permalink raw reply

* Re: [PATCH net] net: rename and move busy poll mib counter
From: Shawn Bohrer @ 2013-08-06 18:18 UTC (permalink / raw)
  To: Eric Dumazet
  Cc: Eliezer Tamir, David Miller, linux-kernel, netdev, Eliezer Tamir
In-Reply-To: <1375784088.4457.74.camel@edumazet-glaptop>

On Tue, Aug 06, 2013 at 03:14:48AM -0700, Eric Dumazet wrote:
> On Tue, 2013-08-06 at 12:52 +0300, Eliezer Tamir wrote:
> > Move the low latency mib counter to the ip section.
> > Rename it from low latency to busy poll.
> > 
> > Reported-by: Shawn Bohrer <sbohrer@rgmadvisors.com>
> > Signed-off-by: Eliezer Tamir <eliezer.tamir@linux.intel.com>
> > ---
> 
> Well, it should not be part of IP mib, but a socket one (not existing so
> far)
> 
> Linux MIB already contains few non TCP counters :
> 
> LINUX_MIB_ARPFILTER
> LINUX_MIB_IPRPFILTER

Doesn't mean they are in the correct place either, but perhaps it's too
late for them.

> Its mostly populated by TCP counters, sure.

See, on the kernel side these are called "LINUX_MIB*" which seems
perfectly sane and I wouldn't even think the statistic is out of
place.  On the user-mode side these are all reported in
/proc/net/netstat as TcpExt statistics.  I can tell you that I don't
look at TCP statistics when I'm debugging/testing UDP issues
(apparently I should).

--
Shawn

-- 

---------------------------------------------------------------
This email, along with any attachments, is confidential. If you 
believe you received this message in error, please contact the 
sender immediately and delete all copies of the message.  
Thank you.

^ permalink raw reply

* Re: low latency/busy poll feedback and bugs
From: Shawn Bohrer @ 2013-08-06 18:08 UTC (permalink / raw)
  To: Eliezer Tamir; +Cc: Amir Vadai, netdev
In-Reply-To: <5200A8BC.4010402@linux.intel.com>

On Tue, Aug 06, 2013 at 10:41:48AM +0300, Eliezer Tamir wrote:
> On 06/08/2013 00:22, Shawn Bohrer wrote:
> > 3) I don't know if this was intentional, an oversight, or simply a
> > missing feature but UDP multicast currently is not supported.  In
> > order to add support I believe you would need to call
> > sk_mark_napi_id() in __udp4_lib_mcast_deliver().  Assuming there isn't
> > some intentional reason this wasn't done I'd be happy to test this and
> > send a patch.
> 
> This is still WIP, so our goal was to make it easy to extend for new
> cases and protocols.
> 
> For multicast, it is possible that incoming packets to come from more
> than one port (and therefore more than one queue).
> I'm not sure how we could handle that, but what we have today won't do
> well for that use-case.
 
It is unclear to me exactly what happens in this case.  With my simple
patch I'm assuming it will spin on the receive queue that received the
last packet for that socket.  What happens when a packet arrives on a
different receive queue than the one we were spinning on? I assume it
is still delivered but perhaps the spinning process won't get it until
the spinning time expires?  I'm just guessing and haven't attempted to
figure it out from looking through the code.

I put together a small test case with two senders and a single
receiver, and visually (by watching /proc/interrups) verified that
their traffic went to two different queues.  The receiver received all
of the packets with busy_read enabled so it appears that it at least
superficially works.  I did not verify the effect on latency.

> What do you use for testing?

In fio 2.1.2 [1] I added support for UDP multicast.  It's not quite as
flexible as I would like but you can still test a number of scenarios
like the one above or do a basic pingpong test.

Here are my fio job files for a pingpong test:

$ cat mcudp_rr_receive
[global]
ioengine=net
protocol=udp
bs=64
size=100m
# IP address of interface to receive packets
#interface=10.8.16.21
rw=read

[pingpong]
pingpong=1
port=10000
hostname=239.0.0.0

$ cat mcudp_rr_send
[global]
ioengine=net
protocol=udp
bs=64
size=100m
# IP address of interface to send packets
#interface=10.8.16.22
rw=write

[pingpong]
pingpong=1
port=10000
hostname=239.0.0.0

Just start the receiver on one host then start the sender on a second
host:

[host1] $ fio mcudp_rr_receive

[host2] $ fio mcudp_rr_send

[1] http://brick.kernel.dk/snaps/fio-2.1.2.tar.bz2

--
Shawn

-- 

---------------------------------------------------------------
This email, along with any attachments, is confidential. If you 
believe you received this message in error, please contact the 
sender immediately and delete all copies of the message.  
Thank you.

^ permalink raw reply

* Re: [PATCH 2/2 net-next] ip_tunnel: operstate support and link state transfer
From: Pravin Shelar @ 2013-08-06 18:07 UTC (permalink / raw)
  To: Stephen Hemminger; +Cc: David Miller, netdev
In-Reply-To: <20130806104942.25bdcfa1@nehalam.linuxnetplumber.net>

On Tue, Aug 6, 2013 at 10:49 AM, Stephen Hemminger
<stephen@networkplumber.org> wrote:
> On Tue, 6 Aug 2013 10:42:53 -0700
> Pravin Shelar <pshelar@nicira.com> wrote:
>
>> On Mon, Aug 5, 2013 at 10:53 PM, Stephen Hemminger
>> <stephen@networkplumber.org> wrote:
>> > Tunnel devices should reflect the carrier state of the lower device.
>> > I.e if carrier goes down on the lower (ethernet) device, it should
>> > change on the tunnel as well.
>> >
>> > This patch also adds full RFC2863 compatible state so that the
>> > tunnel state can be controlled from user space as described in
>> > Documentation/networking/operstats.txt
>> >
>> > Example of usage:
>> > ip li add tnl1 mode dormant \
>> >   type gretap remote 172.19.20.21 local 172.16.17.18 dev eth1
>> > ip li set dev tnl1 up
>> > ip li set dev tnl1 state UP
>> >
>> > In real life, this would be managed by tunnel broker, not
>> > iproute2 shell commands.
>> >
>> >
>> I sent out similar patch which try to add this feature at ip_tunnel
>> generic layer rather than in tunnel implementation.  This way we can
>> share single notifier for all tunneling protocols.
>> Can you have something similar?
>> http://marc.info/?l=linux-netdev&m=135761231222711&w=2
>
> How does that work well for case of GRE where gre and gretap
> have different net namespace id's. If it can handle that, then
> this is better.
>
ip_tunnel registers its own net namespace struct to keep track of all
tunnels from different tunneling protocols.

> Also link_map could be array (not allocated), and avoid
> another layer of indirection
>
> Also, the rfc2863 policy needs to handle CHANGE (for carrier).

that event is not handled in the patch.

^ permalink raw reply

* [PATCH v4 3/3] net: igmp: Allow user-space configuration of igmp unsolicited report interval
From: William Manley @ 2013-08-06 18:03 UTC (permalink / raw)
  To: william.manley, netdev, bcrl, luky-37, sergei.shtylyov,
	bhutchings, davem, hannes
In-Reply-To: <1375812195-6575-1-git-send-email-william.manley@youview.com>

Adds the new procfs knobs:

    /proc/sys/net/ipv4/conf/*/igmpv2_unsolicited_report_interval
    /proc/sys/net/ipv4/conf/*/igmpv3_unsolicited_report_interval

Which will allow userspace configuration of the IGMP unsolicited report
interval (see below) in milliseconds.  The defaults are 10000ms for IGMPv2
and 1000ms for IGMPv3 in accordance with RFC2236 and RFC3376.

Background:

If an IGMP join packet is lost you will not receive data sent to the
multicast group so if no data arrives from that multicast group in a
period of time after the IGMP join a second IGMP join will be sent.  The
delay between joins is the "IGMP Unsolicited Report Interval".

Prior to this patch this value was hard coded in the kernel to 10s for
IGMPv2 and 1s for IGMPv3.  10s is unsuitable for some use-cases, such as
IPTV as it can cause channel change to be slow in the presence of packet
loss.

This patch allows the value to be overridden from userspace for both
IGMPv2 and IGMPv3 such that it can be tuned accoding to the network.

Tested with Wireshark and a simple program to join a (non-existent)
multicast group.  The distribution of timings for the second join differ
based upon setting the procfs knobs.

igmpvX_unsolicited_report_interval is intended to follow the pattern
established by force_igmp_version, and while a procfs entry has been added
a corresponding sysctl knob has not as it is my understanding that sysctl
is deprecated[1].

[1]: http://lwn.net/Articles/247243/

Signed-off-by: William Manley <william.manley@youview.com>
---
 include/linux/inetdevice.h |    2 ++
 net/ipv4/devinet.c         |    8 ++++++++
 net/ipv4/igmp.c            |   19 +++++++++++++++++--
 3 files changed, 27 insertions(+), 2 deletions(-)

diff --git a/include/linux/inetdevice.h b/include/linux/inetdevice.h
index d4c56fc..43b3c72 100644
--- a/include/linux/inetdevice.h
+++ b/include/linux/inetdevice.h
@@ -28,6 +28,8 @@ enum
 	IPV4_DEVCONF_ARPFILTER,
 	IPV4_DEVCONF_MEDIUM_ID,
 	IPV4_DEVCONF_FORCE_IGMP_VERSION,
+	IPV4_DEVCONF_IGMPV2_UNSOLICITED_REPORT_INTERVAL,
+	IPV4_DEVCONF_IGMPV3_UNSOLICITED_REPORT_INTERVAL,
 	IPV4_DEVCONF_NOXFRM,
 	IPV4_DEVCONF_NOPOLICY,
 	IPV4_DEVCONF_ARP_ANNOUNCE,
diff --git a/net/ipv4/devinet.c b/net/ipv4/devinet.c
index a9561c4..56ecb1b 100644
--- a/net/ipv4/devinet.c
+++ b/net/ipv4/devinet.c
@@ -73,6 +73,8 @@ static struct ipv4_devconf ipv4_devconf = {
 		[IPV4_DEVCONF_SEND_REDIRECTS - 1] = 1,
 		[IPV4_DEVCONF_SECURE_REDIRECTS - 1] = 1,
 		[IPV4_DEVCONF_SHARED_MEDIA - 1] = 1,
+		[IPV4_DEVCONF_IGMPV2_UNSOLICITED_REPORT_INTERVAL - 1] = 10000 /*ms*/,
+		[IPV4_DEVCONF_IGMPV3_UNSOLICITED_REPORT_INTERVAL - 1] =  1000 /*ms*/,
 	},
 };
 
@@ -83,6 +85,8 @@ static struct ipv4_devconf ipv4_devconf_dflt = {
 		[IPV4_DEVCONF_SECURE_REDIRECTS - 1] = 1,
 		[IPV4_DEVCONF_SHARED_MEDIA - 1] = 1,
 		[IPV4_DEVCONF_ACCEPT_SOURCE_ROUTE - 1] = 1,
+		[IPV4_DEVCONF_IGMPV2_UNSOLICITED_REPORT_INTERVAL - 1] = 10000 /*ms*/,
+		[IPV4_DEVCONF_IGMPV3_UNSOLICITED_REPORT_INTERVAL - 1] =  1000 /*ms*/,
 	},
 };
 
@@ -2096,6 +2100,10 @@ static struct devinet_sysctl_table {
 		DEVINET_SYSCTL_RW_ENTRY(PROXY_ARP_PVLAN, "proxy_arp_pvlan"),
 		DEVINET_SYSCTL_RW_ENTRY(FORCE_IGMP_VERSION,
 					"force_igmp_version"),
+		DEVINET_SYSCTL_RW_ENTRY(IGMPV2_UNSOLICITED_REPORT_INTERVAL,
+					"igmpv2_unsolicited_report_interval"),
+		DEVINET_SYSCTL_RW_ENTRY(IGMPV3_UNSOLICITED_REPORT_INTERVAL,
+					"igmpv3_unsolicited_report_interval"),
 
 		DEVINET_SYSCTL_FLUSHING_ENTRY(NOXFRM, "disable_xfrm"),
 		DEVINET_SYSCTL_FLUSHING_ENTRY(NOPOLICY, "disable_policy"),
diff --git a/net/ipv4/igmp.c b/net/ipv4/igmp.c
index 9f0aaea..c5541da 100644
--- a/net/ipv4/igmp.c
+++ b/net/ipv4/igmp.c
@@ -141,10 +141,25 @@
 
 static int unsolicited_report_interval(struct in_device *in_dev)
 {
+	int interval_ms, interval_jiffies;
+
 	if (IGMP_V1_SEEN(in_dev) || IGMP_V2_SEEN(in_dev))
-		return IGMP_V2_Unsolicited_Report_Interval;
+		interval_ms = IN_DEV_CONF_GET(
+			in_dev,
+			IGMPV2_UNSOLICITED_REPORT_INTERVAL);
 	else /* v3 */
-		return IGMP_V3_Unsolicited_Report_Interval;
+		interval_ms = IN_DEV_CONF_GET(
+			in_dev,
+			IGMPV3_UNSOLICITED_REPORT_INTERVAL);
+
+	interval_jiffies = msecs_to_jiffies(interval_ms);
+
+	/* _timer functions can't handle a delay of 0 jiffies so ensure
+	 *  we always return a positive value.
+	 */
+	if (interval_jiffies <= 0)
+		interval_jiffies = 1;
+	return interval_jiffies;
 }
 
 static void igmpv3_add_delrec(struct in_device *in_dev, struct ip_mc_list *im);
-- 
1.7.10.4

^ permalink raw reply related

* [PATCH v4 2/3] net: igmp: Don't flush routing cache when force_igmp_version is modified
From: William Manley @ 2013-08-06 18:03 UTC (permalink / raw)
  To: william.manley, netdev, bcrl, luky-37, sergei.shtylyov,
	bhutchings, davem, hannes
In-Reply-To: <1375812195-6575-1-git-send-email-william.manley@youview.com>

The procfs knob /proc/sys/net/ipv4/conf/*/force_igmp_version allows the
IGMP protocol version to use to be explicitly set.  As a side effect this
caused the routing cache to be flushed as it was declared as a
DEVINET_SYSCTL_FLUSHING_ENTRY.  Flushing is unnecessary and this patch
makes it so flushing does not occur.

Requested by Hannes Frederic Sowa as he was reviewing other patches
adding procfs entries.

Suggested-by: Hannes Frederic Sowa <hannes@stressinduktion.org>
Signed-off-by: William Manley <william.manley@youview.com>
---
 include/linux/inetdevice.h |    2 +-
 net/ipv4/devinet.c         |    4 ++--
 2 files changed, 3 insertions(+), 3 deletions(-)

diff --git a/include/linux/inetdevice.h b/include/linux/inetdevice.h
index ea1e3b8..d4c56fc 100644
--- a/include/linux/inetdevice.h
+++ b/include/linux/inetdevice.h
@@ -27,9 +27,9 @@ enum
 	IPV4_DEVCONF_TAG,
 	IPV4_DEVCONF_ARPFILTER,
 	IPV4_DEVCONF_MEDIUM_ID,
+	IPV4_DEVCONF_FORCE_IGMP_VERSION,
 	IPV4_DEVCONF_NOXFRM,
 	IPV4_DEVCONF_NOPOLICY,
-	IPV4_DEVCONF_FORCE_IGMP_VERSION,
 	IPV4_DEVCONF_ARP_ANNOUNCE,
 	IPV4_DEVCONF_ARP_IGNORE,
 	IPV4_DEVCONF_PROMOTE_SECONDARIES,
diff --git a/net/ipv4/devinet.c b/net/ipv4/devinet.c
index dfc39d4..a9561c4 100644
--- a/net/ipv4/devinet.c
+++ b/net/ipv4/devinet.c
@@ -2094,11 +2094,11 @@ static struct devinet_sysctl_table {
 		DEVINET_SYSCTL_RW_ENTRY(ARP_ACCEPT, "arp_accept"),
 		DEVINET_SYSCTL_RW_ENTRY(ARP_NOTIFY, "arp_notify"),
 		DEVINET_SYSCTL_RW_ENTRY(PROXY_ARP_PVLAN, "proxy_arp_pvlan"),
+		DEVINET_SYSCTL_RW_ENTRY(FORCE_IGMP_VERSION,
+					"force_igmp_version"),
 
 		DEVINET_SYSCTL_FLUSHING_ENTRY(NOXFRM, "disable_xfrm"),
 		DEVINET_SYSCTL_FLUSHING_ENTRY(NOPOLICY, "disable_policy"),
-		DEVINET_SYSCTL_FLUSHING_ENTRY(FORCE_IGMP_VERSION,
-					      "force_igmp_version"),
 		DEVINET_SYSCTL_FLUSHING_ENTRY(PROMOTE_SECONDARIES,
 					      "promote_secondaries"),
 		DEVINET_SYSCTL_FLUSHING_ENTRY(ROUTE_LOCALNET,
-- 
1.7.10.4

^ permalink raw reply related

* [PATCH v4 1/3] net: igmp: Reduce Unsolicited report interval to 1s when using IGMPv3
From: William Manley @ 2013-08-06 18:03 UTC (permalink / raw)
  To: william.manley, netdev, bcrl, luky-37, sergei.shtylyov,
	bhutchings, davem, hannes
In-Reply-To: <1375812195-6575-1-git-send-email-william.manley@youview.com>

If an IGMP join packet is lost you will not receive data sent to the
multicast group so if no data arrives from that multicast group in a
period of time after the IGMP join a second IGMP join will be sent.  The
delay between joins is the "IGMP Unsolicited Report Interval".

Previously this value was hard coded to be chosen randomly between 0-10s.
This can be too long for some use-cases, such as IPTV as it can cause
channel change to be slow in the presence of packet loss.

The value 10s has come from IGMPv2 RFC2236, which was reduced to 1s in
IGMPv3 RFC3376.  This patch makes the kernel use the 1s value from the
later RFC if we are operating in IGMPv3 mode.  IGMPv2 behaviour is
unaffected.

Tested with Wireshark and a simple program to join a (non-existent)
multicast group.  The distribution of timings for the second join differ
based upon setting /proc/sys/net/ipv4/conf/eth0/force_igmp_version.

Signed-off-by: William Manley <william.manley@youview.com>
---
 net/ipv4/igmp.c |   16 +++++++++++++---
 1 file changed, 13 insertions(+), 3 deletions(-)

diff --git a/net/ipv4/igmp.c b/net/ipv4/igmp.c
index d8c2327..9f0aaea 100644
--- a/net/ipv4/igmp.c
+++ b/net/ipv4/igmp.c
@@ -113,7 +113,8 @@
 
 #define IGMP_V1_Router_Present_Timeout		(400*HZ)
 #define IGMP_V2_Router_Present_Timeout		(400*HZ)
-#define IGMP_Unsolicited_Report_Interval	(10*HZ)
+#define IGMP_V2_Unsolicited_Report_Interval	(10*HZ)
+#define IGMP_V3_Unsolicited_Report_Interval	(1*HZ)
 #define IGMP_Query_Response_Interval		(10*HZ)
 #define IGMP_Unsolicited_Report_Count		2
 
@@ -138,6 +139,14 @@
 	 ((in_dev)->mr_v2_seen && \
 	  time_before(jiffies, (in_dev)->mr_v2_seen)))
 
+static int unsolicited_report_interval(struct in_device *in_dev)
+{
+	if (IGMP_V1_SEEN(in_dev) || IGMP_V2_SEEN(in_dev))
+		return IGMP_V2_Unsolicited_Report_Interval;
+	else /* v3 */
+		return IGMP_V3_Unsolicited_Report_Interval;
+}
+
 static void igmpv3_add_delrec(struct in_device *in_dev, struct ip_mc_list *im);
 static void igmpv3_del_delrec(struct in_device *in_dev, __be32 multiaddr);
 static void igmpv3_clear_delrec(struct in_device *in_dev);
@@ -719,7 +728,8 @@ static void igmp_ifc_timer_expire(unsigned long data)
 	igmpv3_send_cr(in_dev);
 	if (in_dev->mr_ifc_count) {
 		in_dev->mr_ifc_count--;
-		igmp_ifc_start_timer(in_dev, IGMP_Unsolicited_Report_Interval);
+		igmp_ifc_start_timer(in_dev,
+				     unsolicited_report_interval(in_dev));
 	}
 	__in_dev_put(in_dev);
 }
@@ -744,7 +754,7 @@ static void igmp_timer_expire(unsigned long data)
 
 	if (im->unsolicit_count) {
 		im->unsolicit_count--;
-		igmp_start_timer(im, IGMP_Unsolicited_Report_Interval);
+		igmp_start_timer(im, unsolicited_report_interval(in_dev));
 	}
 	im->reporter = 1;
 	spin_unlock(&im->lock);
-- 
1.7.10.4

^ permalink raw reply related

* IGMP Unsolicited report interval patches
From: William Manley @ 2013-08-06 18:03 UTC (permalink / raw)
  To: william.manley, netdev, bcrl, luky-37, sergei.shtylyov,
	bhutchings, davem, hannes
In-Reply-To: <20130731063442.GA10498@order.stressinduktion.org>

4th version of the patches.

The significant changes since last review are: 

1. there is a new patch (2/3) as requested by Hannes.
2. the third patch now uses IN_DEV_CONF_GET in place of
   IPV4_DEVCONF_ALL.  This means that the unsolicited report interval can
   now be configured on an interface-by-interface basis as I'd originally
   intended but messed up in the implementation.  One concern I have now
   is that with this latest patch-set is that while
   /proc/sys/net/ipv4/conf/eth0/igmp... will now have an effect 
   /proc/sys/net/ipv4/conf/all/igmp... will not.  I'm not sure how to
   resolve this.

   One option would be to have a special value of -1 to mean use the
   default so I could implement fall-back semantics.  A down-side of this
   approach is that it makes the meaning of the knobs less clear for
   someone browsing through the filesystem.  Another option would be to
   remove the knob from all/ entirely, although I'm not sure how to do
   this.  Suggestions are very much welcome :)

Thanks 

Will

^ permalink raw reply

* [RFC PATCH net-next 1/6] bonding: add vlan_uses_dev_rcu() and make bond_vlan_used() use it
From: Veaceslav Falico @ 2013-08-06 17:58 UTC (permalink / raw)
  To: netdev
  Cc: Veaceslav Falico, Jay Vosburgh, Andy Gospodarek, Patrick McHardy,
	David S. Miller, Nikolay Aleksandrov
In-Reply-To: <1375811944-31942-1-git-send-email-vfalico@redhat.com>

Currently, bond_vlan_used() looks for any vlan, including the pseudo-vlan
id 0, and always returns true if 8021q is loaded. This creates several bad
situations - some warnings in __bond_release_one() because it thinks that
we still have vlans while removing, sending LB packets with vlan id 0 and,
possibly, other caused by vlan id 0.

Fix it by adding a new call, vlan_uses_dev_rcu(), which is the same as
vlan_uses_dev(), but uses rcu_dereference() instead of rtnl, and thus we
can use it in bond_vlan_used() wrapped in rcu_read_lock().

By the time vlan_uses_dev_rcu() returns we cannot be sure if there were no
vlans added/removed, so it's basicly an optimization function.

Also, use the pure vlan_uses_dev() in __bond_release_one() cause the rtnl
lock is held there.

For this call to be visible in bonding.h, add include <linux/if_vlan.h>,
and also remove it from any other bonding file, cause they all include
bonding.h, and thus linux/if_vlan.h.

CC: Jay Vosburgh <fubar@us.ibm.com>
CC: Andy Gospodarek <andy@greyhouse.net>
CC: Patrick McHardy <kaber@trash.net>
CC: "David S. Miller" <davem@davemloft.net>
CC: Nikolay Aleksandrov <nikolay@redhat.com>
Signed-off-by: Veaceslav Falico <vfalico@redhat.com>
---
 drivers/net/bonding/bond_alb.c  |    1 -
 drivers/net/bonding/bond_main.c |    3 +--
 drivers/net/bonding/bonding.h   |   10 +++++++++-
 include/linux/if_vlan.h         |    6 ++++++
 net/8021q/vlan_core.c           |   11 +++++++++++
 5 files changed, 27 insertions(+), 4 deletions(-)

diff --git a/drivers/net/bonding/bond_alb.c b/drivers/net/bonding/bond_alb.c
index 3a5db7b..2684329 100644
--- a/drivers/net/bonding/bond_alb.c
+++ b/drivers/net/bonding/bond_alb.c
@@ -34,7 +34,6 @@
 #include <linux/if_arp.h>
 #include <linux/if_ether.h>
 #include <linux/if_bonding.h>
-#include <linux/if_vlan.h>
 #include <linux/in.h>
 #include <net/ipx.h>
 #include <net/arp.h>
diff --git a/drivers/net/bonding/bond_main.c b/drivers/net/bonding/bond_main.c
index 4264a76..9d1045d 100644
--- a/drivers/net/bonding/bond_main.c
+++ b/drivers/net/bonding/bond_main.c
@@ -69,7 +69,6 @@
 #include <net/arp.h>
 #include <linux/mii.h>
 #include <linux/ethtool.h>
-#include <linux/if_vlan.h>
 #include <linux/if_bonding.h>
 #include <linux/jiffies.h>
 #include <linux/preempt.h>
@@ -1953,7 +1952,7 @@ static int __bond_release_one(struct net_device *bond_dev,
 		bond_set_carrier(bond);
 		eth_hw_addr_random(bond_dev);
 
-		if (bond_vlan_used(bond)) {
+		if (vlan_uses_dev(bond_dev)) {
 			pr_warning("%s: Warning: clearing HW address of %s while it still has VLANs.\n",
 				   bond_dev->name, bond_dev->name);
 			pr_warning("%s: When re-adding slaves, make sure the bond's HW address matches its VLANs'.\n",
diff --git a/drivers/net/bonding/bonding.h b/drivers/net/bonding/bonding.h
index 4bf52d5..cb49313 100644
--- a/drivers/net/bonding/bonding.h
+++ b/drivers/net/bonding/bonding.h
@@ -23,6 +23,7 @@
 #include <linux/netpoll.h>
 #include <linux/inetdevice.h>
 #include <linux/etherdevice.h>
+#include <linux/if_vlan.h>
 #include "bond_3ad.h"
 #include "bond_alb.h"
 
@@ -267,9 +268,16 @@ struct bonding {
 #endif /* CONFIG_DEBUG_FS */
 };
 
+/* use vlan_uses_dev() if under rtnl */
 static inline bool bond_vlan_used(struct bonding *bond)
 {
-	return !list_empty(&bond->vlan_list);
+	bool ret;
+
+	rcu_read_lock();
+	ret = vlan_uses_dev_rcu(bond->dev);
+	rcu_read_unlock();
+
+	return ret;
 }
 
 #define bond_slave_get_rcu(dev) \
diff --git a/include/linux/if_vlan.h b/include/linux/if_vlan.h
index 715c343..1fcea36 100644
--- a/include/linux/if_vlan.h
+++ b/include/linux/if_vlan.h
@@ -101,6 +101,7 @@ extern void vlan_vids_del_by_dev(struct net_device *dev,
 				 const struct net_device *by_dev);
 
 extern bool vlan_uses_dev(const struct net_device *dev);
+extern bool vlan_uses_dev_rcu(const struct net_device *dev);
 #else
 static inline struct net_device *
 __vlan_find_dev_deep(struct net_device *real_dev,
@@ -155,6 +156,11 @@ static inline bool vlan_uses_dev(const struct net_device *dev)
 {
 	return false;
 }
+
+static inline bool vlan_uses_dev_rcu(const struct net_device *dev)
+{
+	return false;
+}
 #endif
 
 static inline bool vlan_hw_offload_capable(netdev_features_t features,
diff --git a/net/8021q/vlan_core.c b/net/8021q/vlan_core.c
index 4a78c4d..52e3fb3 100644
--- a/net/8021q/vlan_core.c
+++ b/net/8021q/vlan_core.c
@@ -399,3 +399,14 @@ bool vlan_uses_dev(const struct net_device *dev)
 	return vlan_info->grp.nr_vlan_devs ? true : false;
 }
 EXPORT_SYMBOL(vlan_uses_dev);
+
+bool vlan_uses_dev_rcu(const struct net_device *dev)
+{
+	struct vlan_info *vlan_info;
+
+	vlan_info = rcu_dereference(dev->vlan_info);
+	if (!vlan_info)
+		return false;
+	return vlan_info->grp.nr_vlan_devs ? true : false;
+}
+EXPORT_SYMBOL(vlan_uses_dev_rcu);
-- 
1.7.1

^ permalink raw reply related

* [RFC PATCH net-next 6/6] bonding: remove unused bond->vlan_list
From: Veaceslav Falico @ 2013-08-06 17:59 UTC (permalink / raw)
  To: netdev; +Cc: Veaceslav Falico, Jay Vosburgh, Andy Gospodarek
In-Reply-To: <1375811944-31942-1-git-send-email-vfalico@redhat.com>

There are currently no users of bond->vlan_list (other than the maintaining
functions add/remove) - so remove it and every unneeded helper.

In this patch we remove:
vlan_list from struct bonding
bond_next_vlan - we don't need it anymore
struct vlan_entry - it was a helper struct for bond->vlan_list
some bits from bond_vlan_rx_add/kill_vid() - which were related to
	bond->vlan_list
(de)initialization of vlan_list from bond_setup/uninit
bond_add_vlan - its only scope was to maintain bond->vlan_list

We don't fully remove bond_del_vlan() - bond_alb_clear_vlan() still needs
to be called when a vlan disappears. And we make bond_del_vlan() to not
return anything cause it cannot fail already.

CC: Jay Vosburgh <fubar@us.ibm.com>
CC: Andy Gospodarek <andy@greyhouse.net>
Signed-off-by: Veaceslav Falico <vfalico@redhat.com>
---
 drivers/net/bonding/bond_main.c |  117 +-------------------------------------
 drivers/net/bonding/bonding.h   |    6 --
 2 files changed, 4 insertions(+), 119 deletions(-)

diff --git a/drivers/net/bonding/bond_main.c b/drivers/net/bonding/bond_main.c
index b85c96f..69067d4 100644
--- a/drivers/net/bonding/bond_main.c
+++ b/drivers/net/bonding/bond_main.c
@@ -282,113 +282,24 @@ const char *bond_mode_name(int mode)
 /*---------------------------------- VLAN -----------------------------------*/
 
 /**
- * bond_add_vlan - add a new vlan id on bond
- * @bond: bond that got the notification
- * @vlan_id: the vlan id to add
- *
- * Returns -ENOMEM if allocation failed.
- */
-static int bond_add_vlan(struct bonding *bond, unsigned short vlan_id)
-{
-	struct vlan_entry *vlan;
-
-	pr_debug("bond: %s, vlan id %d\n",
-		 (bond ? bond->dev->name : "None"), vlan_id);
-
-	vlan = kzalloc(sizeof(struct vlan_entry), GFP_KERNEL);
-	if (!vlan)
-		return -ENOMEM;
-
-	INIT_LIST_HEAD(&vlan->vlan_list);
-	vlan->vlan_id = vlan_id;
-
-	write_lock_bh(&bond->lock);
-
-	list_add_tail(&vlan->vlan_list, &bond->vlan_list);
-
-	write_unlock_bh(&bond->lock);
-
-	pr_debug("added VLAN ID %d on bond %s\n", vlan_id, bond->dev->name);
-
-	return 0;
-}
-
-/**
  * bond_del_vlan - delete a vlan id from bond
  * @bond: bond that got the notification
  * @vlan_id: the vlan id to delete
  *
  * returns -ENODEV if @vlan_id was not found in @bond.
  */
-static int bond_del_vlan(struct bonding *bond, unsigned short vlan_id)
+static void bond_del_vlan(struct bonding *bond, unsigned short vlan_id)
 {
-	struct vlan_entry *vlan;
-	int res = -ENODEV;
-
 	pr_debug("bond: %s, vlan id %d\n", bond->dev->name, vlan_id);
 
 	block_netpoll_tx();
 	write_lock_bh(&bond->lock);
 
-	list_for_each_entry(vlan, &bond->vlan_list, vlan_list) {
-		if (vlan->vlan_id == vlan_id) {
-			list_del(&vlan->vlan_list);
-
-			if (bond_is_lb(bond))
-				bond_alb_clear_vlan(bond, vlan_id);
-
-			pr_debug("removed VLAN ID %d from bond %s\n",
-				 vlan_id, bond->dev->name);
+	if (bond_is_lb(bond))
+		bond_alb_clear_vlan(bond, vlan_id);
 
-			kfree(vlan);
-
-			res = 0;
-			goto out;
-		}
-	}
-
-	pr_debug("couldn't find VLAN ID %d in bond %s\n",
-		 vlan_id, bond->dev->name);
-
-out:
 	write_unlock_bh(&bond->lock);
 	unblock_netpoll_tx();
-	return res;
-}
-
-/**
- * bond_next_vlan - safely skip to the next item in the vlans list.
- * @bond: the bond we're working on
- * @curr: item we're advancing from
- *
- * Returns %NULL if list is empty, bond->next_vlan if @curr is %NULL,
- * or @curr->next otherwise (even if it is @curr itself again).
- *
- * Caller must hold bond->lock
- */
-struct vlan_entry *bond_next_vlan(struct bonding *bond, struct vlan_entry *curr)
-{
-	struct vlan_entry *next, *last;
-
-	if (list_empty(&bond->vlan_list))
-		return NULL;
-
-	if (!curr) {
-		next = list_entry(bond->vlan_list.next,
-				  struct vlan_entry, vlan_list);
-	} else {
-		last = list_entry(bond->vlan_list.prev,
-				  struct vlan_entry, vlan_list);
-		if (last == curr) {
-			next = list_entry(bond->vlan_list.next,
-					  struct vlan_entry, vlan_list);
-		} else {
-			next = list_entry(curr->vlan_list.next,
-					  struct vlan_entry, vlan_list);
-		}
-	}
-
-	return next;
 }
 
 /**
@@ -450,13 +361,6 @@ static int bond_vlan_rx_add_vid(struct net_device *bond_dev,
 			goto unwind;
 	}
 
-	res = bond_add_vlan(bond, vid);
-	if (res) {
-		pr_err("%s: Error: Failed to add vlan id %d\n",
-		       bond_dev->name, vid);
-		goto unwind;
-	}
-
 	return 0;
 
 unwind:
@@ -477,17 +381,11 @@ static int bond_vlan_rx_kill_vid(struct net_device *bond_dev,
 {
 	struct bonding *bond = netdev_priv(bond_dev);
 	struct slave *slave;
-	int res;
 
 	bond_for_each_slave(bond, slave)
 		vlan_vid_del(slave->dev, proto, vid);
 
-	res = bond_del_vlan(bond, vid);
-	if (res) {
-		pr_err("%s: Error: Failed to remove vlan id %d\n",
-		       bond_dev->name, vid);
-		return res;
-	}
+	bond_del_vlan(bond, vid);
 
 	return 0;
 }
@@ -4141,7 +4039,6 @@ static void bond_setup(struct net_device *bond_dev)
 
 	/* Initialize pointers */
 	bond->dev = bond_dev;
-	INIT_LIST_HEAD(&bond->vlan_list);
 
 	/* Initialize the device entry points */
 	ether_setup(bond_dev);
@@ -4194,7 +4091,6 @@ static void bond_uninit(struct net_device *bond_dev)
 {
 	struct bonding *bond = netdev_priv(bond_dev);
 	struct slave *slave, *tmp_slave;
-	struct vlan_entry *vlan, *tmp;
 
 	bond_netpoll_cleanup(bond_dev);
 
@@ -4206,11 +4102,6 @@ static void bond_uninit(struct net_device *bond_dev)
 	list_del(&bond->bond_list);
 
 	bond_debug_unregister(bond);
-
-	list_for_each_entry_safe(vlan, tmp, &bond->vlan_list, vlan_list) {
-		list_del(&vlan->vlan_list);
-		kfree(vlan);
-	}
 }
 
 /*------------------------- Module initialization ---------------------------*/
diff --git a/drivers/net/bonding/bonding.h b/drivers/net/bonding/bonding.h
index cb49313..da27e3d 100644
--- a/drivers/net/bonding/bonding.h
+++ b/drivers/net/bonding/bonding.h
@@ -186,11 +186,6 @@ struct bond_parm_tbl {
 
 #define BOND_MAX_MODENAME_LEN 20
 
-struct vlan_entry {
-	struct list_head vlan_list;
-	unsigned short vlan_id;
-};
-
 struct slave {
 	struct net_device *dev; /* first - useful for panic debug */
 	struct list_head list;
@@ -255,7 +250,6 @@ struct bonding {
 	struct   ad_bond_info ad_info;
 	struct   alb_bond_info alb_info;
 	struct   bond_params params;
-	struct   list_head vlan_list;
 	struct   workqueue_struct *wq;
 	struct   delayed_work mii_work;
 	struct   delayed_work arp_work;
-- 
1.7.1

^ permalink raw reply related

* [RFC PATCH net-next 5/6] bonding: convert bond_arp_send_all to use bond->dev->vlan_info
From: Veaceslav Falico @ 2013-08-06 17:59 UTC (permalink / raw)
  To: netdev; +Cc: Veaceslav Falico, Jay Vosburgh, Andy Gospodarek
In-Reply-To: <1375811944-31942-1-git-send-email-vfalico@redhat.com>

Instead of looping through bond->vlan_list, loop through
bond->dev->vlan_info via __vlan_find_dev_next_id() and dereference the
vlan_dev through __vlan_find_dev_deep(), all under rcu_read_lock().

CC: Jay Vosburgh <fubar@us.ibm.com>
CC: Andy Gospodarek <andy@greyhouse.net>
Signed-off-by: Veaceslav Falico <vfalico@redhat.com>
---
 drivers/net/bonding/bond_main.c |   10 ++++------
 1 files changed, 4 insertions(+), 6 deletions(-)

diff --git a/drivers/net/bonding/bond_main.c b/drivers/net/bonding/bond_main.c
index 00f6011..b85c96f 100644
--- a/drivers/net/bonding/bond_main.c
+++ b/drivers/net/bonding/bond_main.c
@@ -2447,7 +2447,6 @@ static void bond_arp_send_all(struct bonding *bond, struct slave *slave)
 {
 	int i, vlan_id;
 	__be32 *targets = bond->params.arp_targets;
-	struct vlan_entry *vlan;
 	struct net_device *vlan_dev = NULL;
 	struct rtable *rt;
 
@@ -2492,19 +2491,18 @@ static void bond_arp_send_all(struct bonding *bond, struct slave *slave)
 		}
 
 		vlan_id = 0;
-		list_for_each_entry(vlan, &bond->vlan_list, vlan_list) {
-			rcu_read_lock();
+		rcu_read_lock();
+		while ((vlan_id = __vlan_find_dev_next_id(bond->dev, vlan_id+1))) {
 			vlan_dev = __vlan_find_dev_deep(bond->dev,
 							htons(ETH_P_8021Q),
-							vlan->vlan_id);
-			rcu_read_unlock();
+							vlan_id);
 			if (vlan_dev == rt->dst.dev) {
-				vlan_id = vlan->vlan_id;
 				pr_debug("basa: vlan match on %s %d\n",
 				       vlan_dev->name, vlan_id);
 				break;
 			}
 		}
+		rcu_read_unlock();
 
 		if (vlan_id && vlan_dev) {
 			ip_rt_put(rt);
-- 
1.7.1

^ permalink raw reply related

* [RFC PATCH net-next 4/6] bonding: convert bond_has_this_ip to use bond->dev->vlan_info
From: Veaceslav Falico @ 2013-08-06 17:59 UTC (permalink / raw)
  To: netdev; +Cc: Veaceslav Falico, Jay Vosburgh, Andy Gospodarek
In-Reply-To: <1375811944-31942-1-git-send-email-vfalico@redhat.com>

Use __vlan_find_dev_next_id() to loop through __vlan_find_dev_next_id() and
__vlan_find_dev_deep() to dereference the actual vlan device and verify if
it has the ip requested.

CC: Jay Vosburgh <fubar@us.ibm.com>
CC: Andy Gospodarek <andy@greyhouse.net>
Signed-off-by: Veaceslav Falico <vfalico@redhat.com>
---
 drivers/net/bonding/bond_main.c |   15 +++++++++------
 1 files changed, 9 insertions(+), 6 deletions(-)

diff --git a/drivers/net/bonding/bond_main.c b/drivers/net/bonding/bond_main.c
index 9d1045d..00f6011 100644
--- a/drivers/net/bonding/bond_main.c
+++ b/drivers/net/bonding/bond_main.c
@@ -2392,20 +2392,23 @@ re_arm:
 
 static int bond_has_this_ip(struct bonding *bond, __be32 ip)
 {
-	struct vlan_entry *vlan;
 	struct net_device *vlan_dev;
+	u16 vlan_id = 0;
 
 	if (ip == bond_confirm_addr(bond->dev, 0, ip))
 		return 1;
 
-	list_for_each_entry(vlan, &bond->vlan_list, vlan_list) {
-		rcu_read_lock();
+	rcu_read_lock();
+	while ((vlan_id = __vlan_find_dev_next_id(bond->dev, vlan_id+1))) {
 		vlan_dev = __vlan_find_dev_deep(bond->dev, htons(ETH_P_8021Q),
-						vlan->vlan_id);
-		rcu_read_unlock();
-		if (vlan_dev && ip == bond_confirm_addr(vlan_dev, 0, ip))
+						vlan_id);
+
+		if (vlan_dev && ip == bond_confirm_addr(vlan_dev, 0, ip)) {
+			rcu_read_unlock();
 			return 1;
+		}
 	}
+	rcu_read_unlock();
 
 	return 0;
 }
-- 
1.7.1

^ permalink raw reply related

* [RFC PATCH net-next 3/6] bonding: make bond_alb use 8021q's dev->vlan_info instead of vlan_list
From: Veaceslav Falico @ 2013-08-06 17:59 UTC (permalink / raw)
  To: netdev; +Cc: Veaceslav Falico, Jay Vosburgh, Andy Gospodarek
In-Reply-To: <1375811944-31942-1-git-send-email-vfalico@redhat.com>

In alb mode, we only need each vlan's id (that is on top of bond) to tag
learning packets, so get them via __vlan_find_dev_next_id(bond->dev,
last_id).

We must also find *any* vlan (including last id stored in current_alb_vlan)
if we can't find anything >= current_alb_vlan id.

For that, we verify if bond has any vlans at all, and if yes - find the
next vlan id after current_alb_vlan id. So, if vlan id is not 0, we tag the
skb.

CC: Jay Vosburgh <fubar@us.ibm.com>
CC: Andy Gospodarek <andy@greyhouse.net>
Signed-off-by: Veaceslav Falico <vfalico@redhat.com>
---
 drivers/net/bonding/bond_alb.c |   28 ++++++++++++++++++----------
 drivers/net/bonding/bond_alb.h |    2 +-
 2 files changed, 19 insertions(+), 11 deletions(-)

diff --git a/drivers/net/bonding/bond_alb.c b/drivers/net/bonding/bond_alb.c
index 2684329..5731540 100644
--- a/drivers/net/bonding/bond_alb.c
+++ b/drivers/net/bonding/bond_alb.c
@@ -976,6 +976,7 @@ static void alb_send_learning_packets(struct slave *slave, u8 mac_addr[])
 	struct learning_pkt pkt;
 	int size = sizeof(struct learning_pkt);
 	int i;
+	u16 vlan_id = 0;
 
 	memset(&pkt, 0, size);
 	memcpy(pkt.mac_dst, mac_addr, ETH_ALEN);
@@ -1000,22 +1001,29 @@ static void alb_send_learning_packets(struct slave *slave, u8 mac_addr[])
 		skb->priority = TC_PRIO_CONTROL;
 		skb->dev = slave->dev;
 
+		rcu_read_lock();
 		if (bond_vlan_used(bond)) {
-			struct vlan_entry *vlan;
+			vlan_id = bond->alb_info.current_alb_vlan+1;
 
-			vlan = bond_next_vlan(bond,
-					      bond->alb_info.current_alb_vlan);
+			vlan_id = __vlan_find_dev_next_id(bond->dev, vlan_id);
+			if (!vlan_id)
+				vlan_id = __vlan_find_dev_next_id(bond->dev, 1);
 
-			bond->alb_info.current_alb_vlan = vlan;
-			if (!vlan) {
+			bond->alb_info.current_alb_vlan = vlan_id;
+
+			if (!vlan_id) {
+				rcu_read_unlock();
 				kfree_skb(skb);
 				continue;
 			}
+		}
+		rcu_read_unlock();
 
-			skb = vlan_put_tag(skb, htons(ETH_P_8021Q), vlan->vlan_id);
+		if (vlan_id) {
+			skb = vlan_put_tag(skb, htons(ETH_P_8021Q), vlan_id);
 			if (!skb) {
-				pr_err("%s: Error: failed to insert VLAN tag\n",
-				       bond->dev->name);
+				pr_err("%s: Error: failed to insert VLAN tag %d\n",
+				       bond->dev->name, vlan_id);
 				continue;
 			}
 		}
@@ -1759,8 +1767,8 @@ int bond_alb_set_mac_address(struct net_device *bond_dev, void *addr)
 void bond_alb_clear_vlan(struct bonding *bond, unsigned short vlan_id)
 {
 	if (bond->alb_info.current_alb_vlan &&
-	    (bond->alb_info.current_alb_vlan->vlan_id == vlan_id)) {
-		bond->alb_info.current_alb_vlan = NULL;
+	    (bond->alb_info.current_alb_vlan == vlan_id)) {
+		bond->alb_info.current_alb_vlan = 0;
 	}
 
 	if (bond->alb_info.rlb_enabled) {
diff --git a/drivers/net/bonding/bond_alb.h b/drivers/net/bonding/bond_alb.h
index e7a5b8b..b2452aa 100644
--- a/drivers/net/bonding/bond_alb.h
+++ b/drivers/net/bonding/bond_alb.h
@@ -170,7 +170,7 @@ struct alb_bond_info {
 						 * rx traffic should be
 						 * rebalanced
 						 */
-	struct vlan_entry	*current_alb_vlan;
+	u16			current_alb_vlan;
 };
 
 int bond_alb_initialize(struct bonding *bond, int rlb_enabled);
-- 
1.7.1

^ permalink raw reply related

* [RFC PATCH net-next] bonding: remove bond->vlan_list
From: Veaceslav Falico @ 2013-08-06 17:58 UTC (permalink / raw)
  To: netdev
  Cc: Jay Vosburgh, Andy Gospodarek, Patrick McHardy, David S. Miller,
	Nikolay Aleksandrov, Veaceslav Falico

The aim of this patchset is to remove bond->vlan_list completely, and use
8021q's standard functions instead of it.

The patchset is on top of Nik's latest two patches:
[net-next,v2,1/2] bonding: change the bond's vlan syncing functions with
			   the standard ones
[net-next,v2,2/2] bonding: unwind on bond_add_vlan failure


First two patches add two helper functions to vlan code:

bonding: add vlan_uses_dev_rcu() and make bond_vlan_used() use it

	Here we add a copy of vlan_uses_dev(), only rcu'd instead of rtnl,
	and make bond_vlan_used() use it under rcu_read_lock().

vlan: add __vlan_find_dev_next_id()

	This function takes dev and vlan_id and returns the next vlan id
	that uses this dev. It can be used to cycle through the vlan list,
	and not only by bonding - but by any network driver that uses its
	private vlan list.

Next four patches actually convert bonding to use the new
functions/approach and remove the vlan_list completely.

This patchset solves several issues with bonding, simplify it overall,
RCUify further and add infrastructure to anyone else who'd like to use
8021q standard functions instead of their own vlan_list, which is quite
common amongst network drivers currently.

I'm testing it continuously currently, no issues found, will update on
anything.

CC: Jay Vosburgh <fubar@us.ibm.com>
CC: Andy Gospodarek <andy@greyhouse.net>
CC: Patrick McHardy <kaber@trash.net>
CC: "David S. Miller" <davem@davemloft.net>
CC: Nikolay Aleksandrov <nikolay@redhat.com>
Signed-off-by: Veaceslav Falico <vfalico@redhat.com>

---
 drivers/net/bonding/bond_alb.c  |   29 +++++---
 drivers/net/bonding/bond_alb.h  |    2 +-
 drivers/net/bonding/bond_main.c |  145 +++++----------------------------------
 drivers/net/bonding/bonding.h   |   16 +++--
 include/linux/if_vlan.h         |   12 +++
 net/8021q/vlan_core.c           |   47 +++++++++++++
 6 files changed, 105 insertions(+), 146 deletions(-)

^ permalink raw reply

* [RFC PATCH net-next 2/6] vlan: add __vlan_find_dev_next_id()
From: Veaceslav Falico @ 2013-08-06 17:59 UTC (permalink / raw)
  To: netdev; +Cc: Veaceslav Falico, Patrick McHardy, David S. Miller
In-Reply-To: <1375811944-31942-1-git-send-email-vfalico@redhat.com>

Add a new exported function __vlan_find_dev_next_id(dev, vlan_id), which
returns the a vlan id that is used by the dev and is greater or equal to
vlan_id.

This function must be under rcu_read_lock(), is aware of master devices and
doesn't guarantee that, once it returns, the vlan id will still be used.

It's basically a helper for "for_each_vlan_in_dev(dev, vlan_dev)" logic,
and is supposed to be used like this:

vlan_id = 0;

while ((vlan_id = __vlan_find_dev_next_id(dev, vlan_id+1))) {
	vlan_dev = __vlan_find_dev_deep(dev, htons(ETH_P_8021Q), vlan_id);
	if (!vlan_dev)
		continue;

	do_work(vlan_dev);
}

In that case we're sure that vlan_dev at least was used as a vlan, after
__vlan_find_dev_deep() returns, and won't go away while we're holding
rcu_read_lock().

However, if we don't hold rtnl_lock(), we can't be sure that that vlan_dev
is still in dev's vlan_list.

CC: Patrick McHardy <kaber@trash.net>
CC: "David S. Miller" <davem@davemloft.net>
Signed-off-by: Veaceslav Falico <vfalico@redhat.com>
---
 include/linux/if_vlan.h |    6 ++++++
 net/8021q/vlan_core.c   |   36 ++++++++++++++++++++++++++++++++++++
 2 files changed, 42 insertions(+), 0 deletions(-)

diff --git a/include/linux/if_vlan.h b/include/linux/if_vlan.h
index 1fcea36..a69219b 100644
--- a/include/linux/if_vlan.h
+++ b/include/linux/if_vlan.h
@@ -86,6 +86,7 @@ static inline int is_vlan_dev(struct net_device *dev)
 
 extern struct net_device *__vlan_find_dev_deep(struct net_device *real_dev,
 					       __be16 vlan_proto, u16 vlan_id);
+extern u16 __vlan_find_dev_next_id(struct net_device *dev, u16 vlan_id);
 extern struct net_device *vlan_dev_real_dev(const struct net_device *dev);
 extern u16 vlan_dev_vlan_id(const struct net_device *dev);
 
@@ -110,6 +111,11 @@ __vlan_find_dev_deep(struct net_device *real_dev,
 	return NULL;
 }
 
+static inline u16 __vlan_find_dev_next_id(struct net_device *dev, u16 vlan_id)
+{
+	return 0;
+}
+
 static inline struct net_device *vlan_dev_real_dev(const struct net_device *dev)
 {
 	BUG();
diff --git a/net/8021q/vlan_core.c b/net/8021q/vlan_core.c
index 52e3fb3..4e9ff3f 100644
--- a/net/8021q/vlan_core.c
+++ b/net/8021q/vlan_core.c
@@ -89,6 +89,42 @@ struct net_device *__vlan_find_dev_deep(struct net_device *dev,
 }
 EXPORT_SYMBOL(__vlan_find_dev_deep);
 
+#define vlan_group_for_each_from(grp, i, from) \
+	for ((i) = from; i < VLAN_PROTO_NUM * VLAN_N_VID; i++) \
+		if (__vlan_group_get_device((grp), (i) / VLAN_N_VID, \
+						   (i) % VLAN_N_VID))
+
+/* Must be called under rcu_read_lock(), returns next vlan_id used by a
+ * vlan device on dev, starting with vlan_id, or 0 if no vlan_id found.
+ */
+u16 __vlan_find_dev_next_id(struct net_device *dev, u16 vlan_id)
+{
+	struct net_device *real_dev = dev, *master_dev;
+	struct vlan_info *vlan_info;
+	struct vlan_group *grp;
+	u16 i;
+
+	BUG_ON(vlan_id >= VLAN_PROTO_NUM * VLAN_N_VID);
+
+	master_dev = netdev_master_upper_dev_get_rcu(real_dev);
+
+	if (master_dev)
+		real_dev = master_dev;
+
+	vlan_info = rcu_dereference(real_dev->vlan_info);
+
+	if (!vlan_info)
+		return 0;
+
+	grp = &vlan_info->grp;
+
+	vlan_group_for_each_from(grp, i, vlan_id)
+		return i;
+
+	return 0;
+}
+EXPORT_SYMBOL(__vlan_find_dev_next_id);
+
 struct net_device *vlan_dev_real_dev(const struct net_device *dev)
 {
 	return vlan_dev_priv(dev)->real_dev;
-- 
1.7.1

^ permalink raw reply related

* Re: [PATCH 2/2 net-next] ip_tunnel: operstate support and link state transfer
From: Stephen Hemminger @ 2013-08-06 17:49 UTC (permalink / raw)
  To: Pravin Shelar; +Cc: David Miller, netdev
In-Reply-To: <CALnjE+qQWampgB5OOR54XkbuBe+Hb6p9d15aPvSEAvjxgXqhwQ@mail.gmail.com>

On Tue, 6 Aug 2013 10:42:53 -0700
Pravin Shelar <pshelar@nicira.com> wrote:

> On Mon, Aug 5, 2013 at 10:53 PM, Stephen Hemminger
> <stephen@networkplumber.org> wrote:
> > Tunnel devices should reflect the carrier state of the lower device.
> > I.e if carrier goes down on the lower (ethernet) device, it should
> > change on the tunnel as well.
> >
> > This patch also adds full RFC2863 compatible state so that the
> > tunnel state can be controlled from user space as described in
> > Documentation/networking/operstats.txt
> >
> > Example of usage:
> > ip li add tnl1 mode dormant \
> >   type gretap remote 172.19.20.21 local 172.16.17.18 dev eth1
> > ip li set dev tnl1 up
> > ip li set dev tnl1 state UP
> >
> > In real life, this would be managed by tunnel broker, not
> > iproute2 shell commands.
> >
> >
> I sent out similar patch which try to add this feature at ip_tunnel
> generic layer rather than in tunnel implementation.  This way we can
> share single notifier for all tunneling protocols.
> Can you have something similar?
> http://marc.info/?l=linux-netdev&m=135761231222711&w=2

How does that work well for case of GRE where gre and gretap
have different net namespace id's. If it can handle that, then
this is better.

Also link_map could be array (not allocated), and avoid
another layer of indirection

Also, the rfc2863 policy needs to handle CHANGE (for carrier).

^ permalink raw reply

* Re: [PATCH 2/2 net-next] ip_tunnel: operstate support and link state transfer
From: Pravin Shelar @ 2013-08-06 17:42 UTC (permalink / raw)
  To: Stephen Hemminger; +Cc: David Miller, netdev
In-Reply-To: <20130805225359.7ba89042@nehalam.linuxnetplumber.net>

On Mon, Aug 5, 2013 at 10:53 PM, Stephen Hemminger
<stephen@networkplumber.org> wrote:
> Tunnel devices should reflect the carrier state of the lower device.
> I.e if carrier goes down on the lower (ethernet) device, it should
> change on the tunnel as well.
>
> This patch also adds full RFC2863 compatible state so that the
> tunnel state can be controlled from user space as described in
> Documentation/networking/operstats.txt
>
> Example of usage:
> ip li add tnl1 mode dormant \
>   type gretap remote 172.19.20.21 local 172.16.17.18 dev eth1
> ip li set dev tnl1 up
> ip li set dev tnl1 state UP
>
> In real life, this would be managed by tunnel broker, not
> iproute2 shell commands.
>
>
I sent out similar patch which try to add this feature at ip_tunnel
generic layer rather than in tunnel implementation.  This way we can
share single notifier for all tunneling protocols.
Can you have something similar?
http://marc.info/?l=linux-netdev&m=135761231222711&w=2

^ permalink raw reply

* Re: [PATCH net-next 1/2] ip_tunnel: embed hash list head
From: Pravin Shelar @ 2013-08-06 17:41 UTC (permalink / raw)
  To: Stephen Hemminger; +Cc: David Miller, netdev
In-Reply-To: <20130805225137.6c934d7f@nehalam.linuxnetplumber.net>

On Mon, Aug 5, 2013 at 10:51 PM, Stephen Hemminger
<stephen@networkplumber.org> wrote:
> The IP tunnel hash heads can be embedded in the per-net structure
> since it is a fixed size. Reduce the size so that the total structure
> fits in a page size. The original size was overly large, even NETDEV_HASHBITS
> is only 8 bits!
>
> Also, add some white space for readability.
>
> Signed-off-by: Stephen Hemminger <stephen@networkplumber.org>
>
>
Looks good.
Acked-by: Pravin B Shelar <pshelar@nicira.com>.

^ permalink raw reply

* GOOD....NEWS......GOOD....NEWS
From: mrs.awa.sule @ 2013-08-06 17:27 UTC (permalink / raw)






With All Due Respect


My name is MRS AWA SULE, am the manager of  auditing and accounting department Bank of  African, I need your urgent assistance in  transferring the sum of $17, 200, 000.00  (Seventeen Million Two Hundred Thousand United  States Dollars Only) immediately to your  account.Meanwhile, it was very fortunate for me to  come across the deceased file, when I was  arranging the old and abandoned customers files in  other to sign and submit to the entire bank  management, as official re-documentation.


However, it is not authorized by the rules guiding  our bank for a citizen of Burkina Faso to make the  claim of the fund unless you are a foreigner, no  matter the country you come from, that’s the  reason I contacts you as a foreigner to apply for  the claim and transfer of the fund smoothly into  your reliable bank account as the next of kin to  the deceased, and I assure you that this  transaction is 100% risks free.If you are really  sure of your, Trust worthiness, Accountability and  confidentiality on this transaction contact me and  accept not to change your mind to cheat, or  disappoint me when the deposited funds are  released to you by our bank; 


FullName:........ ....................................
Address:.......... . .................................
Occupation........ ..................................
Age.............. ......................................
Sex...............  ....................................
Telephone Number. ...............................
Country..... .. ... ....................................


Please contact me back immediately. so that I will let you know the next steps and procedures to  follow in order to finalize this transaction  immediately ok.
I am waiting for your urgent response!!! 
Thanks and remain blessed, 
MRS AWA SULE,

^ permalink raw reply

* Re: [PATCH v3 4/4] USBNET: ax88179_178a: enable tso if usb host supports sg dma
From: Grant Grundler @ 2013-08-06 17:09 UTC (permalink / raw)
  To: Eric Dumazet
  Cc: Ming Lei, David S. Miller, Greg Kroah-Hartman, Oliver Neukum,
	Sarah Sharp, netdev, linux-usb, Ben Hutchings, Alan Stern,
	Freddy Xin
In-Reply-To: <1375791737.4457.98.camel@edumazet-glaptop>

On Tue, Aug 6, 2013 at 5:22 AM, Eric Dumazet <eric.dumazet@gmail.com> wrote:
...
>> @@ -1310,6 +1318,10 @@ static int ax88179_reset(struct usbnet *dev)
>>
>>       dev->net->hw_features |= NETIF_F_IP_CSUM | NETIF_F_IPV6_CSUM |
>>                                NETIF_F_RXCSUM;
>> +     if (dev->can_dma_sg) {
>> +             dev->net->features |= NETIF_F_SG | NETIF_F_TSO;
>> +             dev->net->hw_features |= NETIF_F_SG | NETIF_F_TSO;
>> +     }
>>
>
> My concern with setting TSO on reset() is the following :
>
> Admin can disable TSO with
>
> ethtool -K ethX tso off
>
>
> Then, one hour later, or one month later, a reset happens, and this code
> magically re-enables TSO
>
> So, I really think this part should be removed from your patch.

Following that logic, shouldn't all the features/hw_features settings
be removed from reset code path?

hw_features shouldn't change since power up.

FWIW, I do agree with you.

I'll note that any "hiccup" in the USB side that causes the device to
get dropped and re-probed will cause the same symptom. There is
nothing the driver can do about it in this case. Perhaps add some udev
rules to preserve ethtool settings the same way I've seen udev rules
to record MAC address to enumerate devices (eth0, eth1, etc.)

cheers,
grant

^ permalink raw reply

* [ANNOUNCE] conntrack-tools 1.4.2 release
From: Pablo Neira Ayuso @ 2013-08-06 16:55 UTC (permalink / raw)
  To: netfilter-devel; +Cc: netdev, netfilter, netfilter-announce, lwn

[-- Attachment #1: Type: text/plain, Size: 876 bytes --]

Hi!

The Netfilter project proudly presents:

        conntrack-tools 1.4.2

The conntrack-tools are the userspace command line interface
`conntrack' and the userspace daemon `conntrackd'. The conntrack
utility replaces the old /proc/net/nf_conntrack interface. With
conntrack, you can dump, modify and delete entries from the connection
tracking state table from userspace. On the other hand, conntrackd
allows you to deploy highly available stateful firewall clusters and
to run connection tracking helpers from user-space.

More information in the official manual at:
http://conntrack-tools.netfilter.org/manual.html

This release includes bugfixes and the connlabel support. See ChangeLog that
comes attached to this email for more details.

You can download it from:

http://www.netfilter.org/projects/nfacct/downloads.html
ftp://ftp.netfilter.org/pub/nfacct/

Have fun!

[-- Attachment #2: changes-conntrack-tools-1.4.2.txt --]
[-- Type: text/plain, Size: 1124 bytes --]

Clemence Faure (2):
      conntrack: introduce -l option to filter by labels
      conntrack: fix reporting of unknown arguments

Florian Westphal (5):
      conntrackd: fix compiler warnings
      include: kill unused PLD_* macros
      conntrack: add connlabel format attribute
      conntrackd: support replication of connlabels
      conntrack: fix -L format output

James Guthrie (1):
      conntrackd: fix parsing of non-abbreviated IPv6 address in config file

Pablo Neira Ayuso (11):
      build: requires libnetfilter_conntrack >= 1.0.3
      conntrack: fix timestamps when microseconds are less than 100000
      tests: cthelper: remove test infrastructure from this tree
      cthelper: add IPv6 support
      cthelper: helpers may not use private information area
      conntrackd: cache: fix hashing based on IPv6 address
      conntrackd: deprecate `Family' in configuration file
      conntrackd: fix crash with IPv6 expectation in the filtering code
      conntrackd: simplify expectation filtering
      cthelper: fix IPv6 address and mask in newly created expectations
      conntrack-tools 1.4.2 release


^ permalink raw reply


This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox