* Re: 2.6.23-rc8-mm2 BUG: register_netdevice() issue as (ab)used by ISDN
From: Jeff Garzik @ 2007-10-07 12:24 UTC (permalink / raw)
To: Andreas Mohr; +Cc: isdn4linux, netdev, akpm
In-Reply-To: <20071007120653.GA16184@rhlx01.hs-esslingen.de>
Andreas Mohr wrote:
> I intend to still try to get it up and running with 2.6.23-rc8-mm2 today
> (with some workarounds hopefully, maybe even disabling ISDN completely)...
>
> The last running kernel (I didn't have newer ones in between), up for some 110
> days was 2.6.19-cks2 (IOW, I cannot quite say that
> "this is an important regression, it has been broken very recently").
It definitely looks like ISDN is somehow to blame. Any chance you could
try git://git.kernel.org/pub/scm/linux/kernel/git/davem/net-2.6.24.git ?
That would help narrow down the problem to a very-large set of
networking changes pending for 2.6.24.
If you can isolate it to net-2.6.24, it should be easily bisect-able
from there, if you have the patience :)
Jeff
^ permalink raw reply
* 2.6.23-rc8-mm2 BUG: register_netdevice() issue as (ab)used by ISDN
From: Andreas Mohr @ 2007-10-07 12:06 UTC (permalink / raw)
To: isdn4linux; +Cc: netdev, akpm
[not necessarily a very recent regression, used 2.6.19 kernels before...]
Hi all,
wondered why my main internet server (headless!) didn't come up properly
on a new 2.6.23-rc8-mm2
(connections almost completely refused: firewalling not executed due to
earlier OOPS?).
Upon LILO emergency fallback into an older version (2.6.16...) saw this
in /var/log/messages:
Oct 7 13:07:34 gate kernel: e100: Intel(R) PRO/100 Network Driver, 3.5.23-k4-NAPI
Oct 7 13:07:34 gate kernel: e100: Copyright(c) 1999-2006 Intel Corporation
Oct 7 13:07:34 gate kernel: Atmel at76x USB Wireless LAN Driver 0.16 loading
Oct 7 13:07:34 gate kernel: eth2: MAC address 00:05:5d:95:ab:f0
Oct 7 13:07:34 gate kernel: eth2: firmware version 1.101.5-84
Oct 7 13:07:34 gate kernel: eth2: regulatory domain 0x30: ETSI (most of Europe)
Oct 7 13:07:34 gate kernel: usbcore: registered new interface driver at76_usb
Oct 7 13:07:34 gate kernel: Netfilter messages via NETLINK v0.30.
Oct 7 13:07:34 gate kernel: nf_conntrack version 0.5.0 (2048 buckets, 8192 max)
Oct 7 13:07:34 gate kernel: nf_conntrack_ipv4: Unknown parameter `hashsize'
Oct 7 13:07:34 gate kernel: ip_tables: (C) 2000-2006 Netfilter Core Team
Oct 7 13:07:34 gate kernel: process `syslogd' is using obsolete setsockopt SO_BSD
COMPAT
Oct 7 13:07:40 gate kernel: ------------[ cut here ]------------
Oct 7 13:07:40 gate kernel: kernel BUG at net/core/dev.c:3485!
Oct 7 13:07:40 gate kernel: invalid opcode: 0000 [#1]
Oct 7 13:07:40 gate kernel: last sysfs file: /devices/platform/sis5595.656/temp1_
max
Oct 7 13:07:40 gate kernel: Modules linked in: xt_state xt_limit xt_tcpudp xt_mul
tiport iptable_mangle iptable_nat nf_conntrack_ipv4 iptable_filter ip_tables nf_na
t_tftp nf_conntrack_tftp nf_nat_h323 nf_conntrack_h323 nf_nat_irc nf_nat_ftp nf_co
nntrack_irc nf_conntrack_ftp ipt_MASQUERADE nf_nat nf_conntrack nfnetlink ipt_REJE
CT ipt_LOG x_tables at76_usb firmware_class e100 ohci_hcd usbcore i2c_sis630 i2c_s
is5595 i2c_core sis5595 hisax isdn eepro100
Oct 7 13:07:40 gate kernel:
Oct 7 13:07:40 gate kernel: Pid: 2235, comm: isdnctrl Not tainted (2.6.23-rc8-mm2
-gate #1)
Oct 7 13:07:40 gate kernel: EIP: 0060:[<c029ad8a>] EFLAGS: 00010246 CPU: 0
Oct 7 13:07:40 gate kernel: EIP is at register_netdevice+0x6d/0x2c1
Oct 7 13:07:40 gate kernel: EAX: 00000000 EBX: c786501c ECX: 00000000 EDX: 00000d
99
Oct 7 13:07:40 gate kernel: ESI: c786501c EDI: 00000000 EBP: c1657d08 ESP: c1657c
ec
Oct 7 13:07:40 gate kernel: DS: 007b ES: 007b FS: 0000 GS: 0033 SS: 0068
Oct 7 13:07:40 gate kernel: Process isdnctrl (pid: 2235, ti=c1656000 task=c14c306
0 task.ti=c1656000)
Oct 7 13:07:40 gate kernel: Stack: c786501c 00000000 c1657d00 c0309ec6 c1657d68 c
786501c 00000000 c1657d18
Oct 7 13:07:40 gate kernel: c029b010 c1657d68 c7865000 c1657d38 c885a0fa c
884e39c c8864598 c786501c
Oct 7 13:07:40 gate kernel: c1657d68 00000000 bfbf779d c1657f60 c886484d c
212ef80 c1561980 c1805ab0
Oct 7 13:07:40 gate kernel: Call Trace:
Oct 7 13:07:40 gate kernel: [<c029b010>] register_netdev+0x32/0x3f
Oct 7 13:07:40 gate kernel: [<c885a0fa>] isdn_net_new+0x111/0x2ca [isdn]
Oct 7 13:07:40 gate kernel: [<c886484d>] isdn_ioctl+0x2b5/0xb46 [isdn]
Oct 7 13:07:40 gate kernel: [<c015f53c>] do_ioctl+0x40/0x50
Oct 7 13:07:40 gate kernel: [<c015f738>] vfs_ioctl+0x1ec/0x203
Oct 7 13:07:40 gate kernel: [<c015f780>] sys_ioctl+0x31/0x49
Oct 7 13:07:40 gate kernel: [<c0103e12>] syscall_call+0x7/0xb
Oct 7 13:07:40 gate kernel: [<b7f54de4>] 0xb7f54de4
Oct 7 13:07:40 gate kernel: =======================
Oct 7 13:07:40 gate kernel: Code: e8 3f a5 e6 ff ba 99 0d 00 00 b8 b4 3c 38 c0 e8
96 82 e7 ff 83 bb 20 02 00 00 00 74 04 0f 0b eb fe 8b 83 80 02 00 00 85 c0 75 04
<0f> 0b eb fe 89 45 f0 b9 4c 2e 48 c0 ba 8e 3e 38 c0 8d 83 78 01
Oct 7 13:07:40 gate kernel: EIP: [<c029ad8a>] register_netdevice+0x6d/0x2c1 SS:ES
P 0068:c1657cec
Oct 7 13:18:17 gate syslogd 1.4.1#20: restart.
Oct 7 13:18:17 gate kernel: klogd 1.4.1#20, log source = /proc/kmsg started.
Oct 7 13:18:17 gate kernel: Linux version 2.6.16-cks11-gate (root@gate) (gcc vers
ion 4.0.4 20060507 (prerelease) (Debian 4.0.3-3)) #1 Sat May 27 14:45:18 CEST 2006
[logging was then back into 2.6.16 above]
2.6.23-rc8-mm2/net/core/dev.c/register_netdevice():
int register_netdevice(struct net_device *dev)
{
struct hlist_head *head;
struct hlist_node *p;
int ret;
struct net *net;
BUG_ON(dev_boot_phase);
ASSERT_RTNL();
might_sleep();
/* When net_device's are persistent, this will be fatal. */
BUG_ON(dev->reg_state != NETREG_UNINITIALIZED);
BUG_ON(!dev->nd_net);
net = dev->nd_net;
spin_lock_init(&dev->queue_lock);
spin_lock_init(&dev->_xmit_lock);
netdev_set_lockdep_class(&dev->_xmit_lock, dev->type);
dev->xmit_lock_owner = -1;
spin_lock_init(&dev->ingress_lock);
dev->iflink = -1;
/* Init, if this function is available */
if (dev->init) {
ret = dev->init(dev);
if (ret) {
if (ret > 0)
ret = -EIO;
goto out;
}
}
objdump -D linux-2.6.23-rc8-mm2/vmlinux|less :
c029ad1d <register_netdevice>:
c029ad1d: 55 push %ebp
c029ad1e: 89 e5 mov %esp,%ebp
c029ad20: 57 push %edi
c029ad21: 56 push %esi
c029ad22: 53 push %ebx
c029ad23: 89 c3 mov %eax,%ebx
c029ad25: 83 ec 10 sub $0x10,%esp
c029ad28: 83 3d 4c fa 41 c0 00 cmpl $0x0,0xc041fa4c
c029ad2f: 74 04 je c029ad35 <register_netdevice+0x18>
c029ad31: 0f 0b ud2a
c029ad33: eb fe jmp c029ad33 <register_netdevice+0x16>
c029ad35: e8 7e 7f 00 00 call c02a2cb8 <rtnl_trylock>
c029ad3a: 85 c0 test %eax,%eax
c029ad3c: 74 26 je c029ad64 <register_netdevice+0x47>
c029ad3e: e8 79 78 00 00 call c02a25bc <rtnl_unlock>
c029ad43: c7 44 24 08 97 0d 00 movl $0xd97,0x8(%esp)
c029ad4a: 00
c029ad4b: c7 44 24 04 b4 3c 38 movl $0xc0383cb4,0x4(%esp)
c029ad52: c0
c029ad53: c7 04 24 c3 3c 38 c0 movl $0xc0383cc3,(%esp)
c029ad5a: e8 13 e3 e7 ff call c0119072 <printk>
c029ad5f: e8 3f a5 e6 ff call c01052a3 <dump_stack>
c029ad64: ba 99 0d 00 00 mov $0xd99,%edx
c029ad69: b8 b4 3c 38 c0 mov $0xc0383cb4,%eax
c029ad6e: e8 96 82 e7 ff call c0113009 <__might_sleep>
c029ad73: 83 bb 20 02 00 00 00 cmpl $0x0,0x220(%ebx)
c029ad7a: 74 04 je c029ad80 <register_netdevice+0x63>
c029ad7c: 0f 0b ud2a
c029ad7e: eb fe jmp c029ad7e <register_netdevice+0x61>
c029ad80: 8b 83 80 02 00 00 mov 0x280(%ebx),%eax
c029ad86: 85 c0 test %eax,%eax
c029ad88: 75 04 jne c029ad8e <register_netdevice+0x71>
c029ad8a: 0f 0b ud2a
c029ad8c: eb fe jmp c029ad8c <register_netdevice+0x6f>
c029ad8e: 89 45 f0 mov %eax,0xfffffff0(%ebp)
c029ad91: b9 4c 2e 48 c0 mov $0xc0482e4c,%ecx
c029ad96: ba 8e 3e 38 c0 mov $0xc0383e8e,%edx
c029ad9b: 8d 83 78 01 00 00 lea 0x178(%ebx),%eax
c029ada1: e8 8e d9 f7 ff call c0218734 <__spin_lock_init>
c029ada6: 8d 83 b4 01 00 00 lea 0x1b4(%ebx),%eax
c029adac: b9 4c 2e 48 c0 mov $0xc0482e4c,%ecx
c029adb1: ba 9f 3e 38 c0 mov $0xc0383e9f,%edx
c029adb6: e8 79 d9 f7 ff call c0218734 <__spin_lock_init>
c029adbb: ba b0 3e 38 c0 mov $0xc0383eb0,%edx
c029adc0: b9 4c 2e 48 c0 mov $0xc0482e4c,%ecx
c029adc5: c7 83 c4 01 00 00 ff movl $0xffffffff,0x1c4(%ebx)
c029adcc: ff ff ff
c029adcf: 8d 83 a0 01 00 00 lea 0x1a0(%ebx),%eax
c029add5: e8 5a d9 f7 ff call c0218734 <__spin_lock_init>
c029adda: 8b 53 40 mov 0x40(%ebx),%edx
c029addd: c7 43 50 ff ff ff ff movl $0xffffffff,0x50(%ebx)
c029ade4: 85 d2 test %edx,%edx
c029ade6: 74 1b je c029ae03 <register_netdevice+0xe6>
Since EIP is c029ad8a, following the code flow it clearly looks to me
as if it's the
BUG_ON(!dev->nd_net);
check which caused the BUG message.
This is a K6-3/450@150 running Debian stable.
lsmod (NOTE: this is on 2.6.16!!):
Module Size Used by
sch_ingress 4356 1
cls_u32 7044 5
sch_tbf 6272 1
sch_sfq 5632 4
sch_htb 15616 1
ppp_async 11648 1
crc_ccitt 2176 1 ppp_async
ipt_ULOG 7712 1
xt_state 2176 4
xt_limit 2688 11
xt_tcpudp 3328 48
ipt_multiport 2560 10
iptable_mangle 2816 0
iptable_nat 7684 1
iptable_filter 3072 1
ip_tables 11736 3 iptable_mangle,iptable_nat,iptable_filter
ip_nat_tftp 1920 0
ip_conntrack_tftp 4244 1 ip_nat_tftp
ip_nat_irc 2688 0
ip_nat_ftp 3200 0
ip_conntrack_irc 6680 1 ip_nat_irc
ip_conntrack_ftp 7324 1 ip_nat_ftp
ipt_MASQUERADE 3712 1
ip_nat 15764 5 iptable_nat,ip_nat_tftp,ip_nat_irc,ip_nat_ftp,ipt_MASQUERADE
ip_conntrack 44196 10 xt_state,iptable_nat,ip_nat_tftp,ip_conntrack_tftp,ip_nat_irc,ip_nat_ftp,ip_conntrack_irc,ip_conntrack_ftp,ipt_MASQUERADE,ip_nat
ipt_REJECT 5248 0
ipt_LOG 6144 11
x_tables 12292 10 ipt_ULOG,xt_state,xt_limit,xt_tcpudp,ipt_multiport,iptable_nat,ip_tables,ipt_MASQUERADE,ipt_REJECT,ipt_LOG
e100 33284 0
at76c503_rfmd 5260 0
firmware_class 10112 1 at76c503_rfmd
at76c503 79840 1 at76c503_rfmd
at76_usbdfu 4996 1 at76c503
ohci_hcd 29700 0
usbcore 126240 5 at76c503_rfmd,at76c503,at76_usbdfu,ohci_hcd
i2c_sis630 8844 0
i2c_sis5595 7940 0
sis5595 14216 0
i2c_isa 5888 1 sis5595
i2c_core 22144 4 i2c_sis630,i2c_sis5595,sis5595,i2c_isa
hisax 175100 2
isdn 108864 5 hisax
eepro100 29456 0
root@gate:/usr/src/linux-2.6.23-rc8-mm2/net# dpkg -l|grep isdn
ii isdnlog 3.9.20060704-3 ISDN connection logger
ii isdnlog-data 3.9.20060704-3 data for isdnlog users
ii isdnutils 3.9.20060704-3 Most important ISDN-related packages and uti
ii isdnutils-base 3.9.20060704-3 ISDN utilities, the basic (minimal) set
ii isdnutils-xtools 3.9.20060704-3 ISDN utilities that use X
ii isdnvboxclient 3.9.20060704-3 ISDN answering machine, client
ii isdnvboxserver 3.9.20060704-3 ISDN answering machine, server
CONFIG_ISDN=m
CONFIG_ISDN_I4L=m
CONFIG_ISDN_PPP=y
CONFIG_ISDN_PPP_VJ=y
CONFIG_ISDN_MPP=y
# CONFIG_IPPP_FILTER is not set
CONFIG_ISDN_PPP_BSDCOMP=m
CONFIG_ISDN_AUDIO=y
# CONFIG_ISDN_TTY_FAX is not set
#
# ISDN feature submodules
#
CONFIG_ISDN_DRV_LOOP=m
# CONFIG_ISDN_DIVERSION is not set
#
# ISDN4Linux hardware drivers
#
#
# Passive cards
#
CONFIG_ISDN_DRV_HISAX=m
#
# D-channel protocol features
#
CONFIG_HISAX_EURO=y
CONFIG_DE_AOC=y
# CONFIG_HISAX_NO_SENDCOMPLETE is not set
# CONFIG_HISAX_NO_LLC is not set
# CONFIG_HISAX_NO_KEYPAD is not set
# CONFIG_HISAX_1TR6 is not set
# CONFIG_HISAX_NI1 is not set
CONFIG_HISAX_MAX_CARDS=8
#
# HiSax supported cards
#
# CONFIG_HISAX_16_0 is not set
CONFIG_HISAX_16_3=y
CONFIG_HISAX_TELESPCI=y
# CONFIG_HISAX_S0BOX is not set
# CONFIG_HISAX_AVM_A1 is not set
CONFIG_HISAX_FRITZPCI=y
# CONFIG_HISAX_AVM_A1_PCMCIA is not set
# CONFIG_HISAX_ELSA is not set
# CONFIG_HISAX_IX1MICROR2 is not set
# CONFIG_HISAX_DIEHLDIVA is not set
# CONFIG_HISAX_ASUSCOM is not set
# CONFIG_HISAX_TELEINT is not set
# CONFIG_HISAX_HFCS is not set
# CONFIG_HISAX_SEDLBAUER is not set
CONFIG_HISAX_SPORTSTER=y
# CONFIG_HISAX_MIC is not set
# CONFIG_HISAX_NETJET is not set
# CONFIG_HISAX_NETJET_U is not set
# CONFIG_HISAX_NICCY is not set
# CONFIG_HISAX_ISURF is not set
# CONFIG_HISAX_HSTSAPHIR is not set
# CONFIG_HISAX_BKM_A4T is not set
# CONFIG_HISAX_SCT_QUADRO is not set
# CONFIG_HISAX_GAZEL is not set
CONFIG_HISAX_HFC_PCI=y
CONFIG_HISAX_W6692=y
# CONFIG_HISAX_HFC_SX is not set
# CONFIG_HISAX_DEBUG is not set
#
# HiSax PCMCIA card service modules
#
#
# HiSax sub driver modules
#
# CONFIG_HISAX_ST5481 is not set
# CONFIG_HISAX_HFCUSB is not set
# CONFIG_HISAX_HFC4S8S is not set
CONFIG_HISAX_FRITZ_PCIPNP=m
#
# Active cards
#
# CONFIG_ISDN_DRV_ICN is not set
# CONFIG_ISDN_DRV_PCBIT is not set
# CONFIG_ISDN_DRV_SC is not set
# CONFIG_ISDN_DRV_ACT2000 is not set
# CONFIG_HYSDN is not set
# CONFIG_ISDN_DRV_GIGASET is not set
# CONFIG_ISDN_CAPI is not set
# CONFIG_PHONE is not set
lspci (NOTE: on 2.6.16!!):
00:09.0 Network controller: Cologne Chip Designs GmbH ISDN network controller [HFC-PCI] (rev 02)
I used to have an ISA-based card (Teles?) in there, replaced by the PCI one
about 2 years ago.
grep isdnctrl /etc/isdn/*:
/etc/isdn/device.ippp0:# Read the isdnctrl manpage for more info.
/etc/isdn/device.ippp0: isdnctrl addif ${device}
/etc/isdn/device.ippp0: isdnctrl eaz ${device} $LOCALMSN
/etc/isdn/device.ippp0: # "name". More than one number can be set by calling isdnc
trl addphone
/etc/isdn/device.ippp0: isdnctrl addphone ${device} out $LEADINGZE
RO$MSN
/etc/isdn/device.ippp0: # disabled. More than one number can be set by calling isd
nctrl addphone
/etc/isdn/device.ippp0: # isdnctrl addphone ${device} in $MSN
/etc/isdn/device.ippp0: # been added to the access list with isdnctrl addphone nam
e in.
/etc/isdn/device.ippp0: isdnctrl secure ${device} on
/etc/isdn/device.ippp0: isdnctrl huptimeout ${device} 180 # XXX_
/etc/isdn/device.ippp0: #isdnctrl dialmax ${device} 3
/etc/isdn/device.ippp0: #isdnctrl ihup ${device} on
/etc/isdn/device.ippp0: isdnctrl encap ${device} $ENCAP
/etc/isdn/device.ippp0: isdnctrl l2_prot ${device} hdlc
/etc/isdn/device.ippp0: isdnctrl l3_prot ${device} trans
/etc/isdn/device.ippp0: isdnctrl verbose 2
/etc/isdn/device.ippp0: #isdnctrl chargehup ${device} on
/etc/isdn/device.ippp0: #isdnctrl chargeint ${device} NUM
/etc/isdn/device.ippp0: #isdnctrl callback ${device} MODE
/etc/isdn/device.ippp0: #isdnctrl cbdelay ${device} SECONDS
/etc/isdn/device.ippp0: #isdnctrl cbhup ${device} MODE
/etc/isdn/device.ippp0: # See also : isdnctrl(8), isdnctrl help text
/etc/isdn/device.ippp0: isdnctrl pppbind ${device} $bindnum
/etc/isdn/device.ippp0: isdnctrl dialmode $device $DIALMODE >/dev/null 2>&1
/etc/isdn/device.ippp0: isdnctrl dialmode $device off >/dev/null 2>&1
/etc/isdn/device.ippp0: isdnctrl delif $device 2> /dev/null
/etc/isdn/init.d.functions: # can't count on "isdnctrl status all" working yet,
unfortunately...
/etc/isdn/init.d.functions: DEVS=`/usr/sbin/isdnctrl list all | grep 'Current
setup' | cut -f2 -d"'" | sort`
/etc/isdn/init.d.functions: if [ ! -e /dev/isdnctrl ]; then
/etc/isdn/init.d.functions: cd /dev && ln -s isdnctrl0 isdnctrl
/etc/isdn/init.d.functions: cardnum=0 # counts in channels, just like /dev/is
dnctrlX
/etc/isdn/init.d.functions: for optionfile in /etc/isdn/isdnlog.isdnctrl[02468]
; do
/etc/isdn/init.d.functions: devicenum=${device#isdnctrl}
/etc/isdn/init.d.functions: # test for isdnctrl dualmode. With dualmode, on
e isdnlog listens to
/etc/isdn/init.d.functions: for optionfile in /etc/isdn/isdnlog.isdnctrl?; do
/etc/isdn/init.d.functions: for optionfile in /etc/isdn/isdnlog.isdnctrl?; do
/etc/isdn/init.d.functions: /usr/sbin/isdnctrl delif $device >/
dev/null 2>&1 || true
/etc/isdn/init.d.functions: /usr/sbin/isdnctrl delif $device >/dev/null
2>&1 || true
/etc/isdn/netdown.old:eval `grep '^ isdnctrl addphone' /etc/isdn/device.ippp0
| sed 's,addphone,delphone,'`
/etc/isdn/netdown.old:/sbin/isdnctrl hangup ippp0
/etc/isdn/netup.old:eval `grep '^ isdnctrl addphone' /etc/isdn/device.ippp0`
/etc/isdn/stop: /usr/sbin/isdnctrl system off
/etc/isdn/xisdnload-netdown:# script again). So, putting "isdnctrl dialmode all of
f" here is not that
/etc/isdn/xisdnload-netdown:# useful, as you have to do "isdnctrl dialmode all aut
o" manually...
/etc/isdn/xisdnload-netdown:/usr/sbin/isdnctrl hangup ippp0 > /dev/null
/etc/isdn/xmonisdn-netdown:/usr/sbin/isdnctrl dialmode all off
/etc/isdn/xmonisdn-netup:/usr/sbin/isdnctrl dialmode all auto
I intend to still try to get it up and running with 2.6.23-rc8-mm2 today
(with some workarounds hopefully, maybe even disabling ISDN completely)...
The last running kernel (I didn't have newer ones in between), up for some 110
days was 2.6.19-cks2 (IOW, I cannot quite say that
"this is an important regression, it has been broken very recently").
Thanks,
Andreas Mohr
^ permalink raw reply
* Re: [PATCH 2/5] forcedeth: interrupt handling cleanup
From: Jeff Garzik @ 2007-10-07 11:40 UTC (permalink / raw)
To: Yinghai Lu; +Cc: netdev, Ayaz Abdulla, LKML, Andrew Morton
In-Reply-To: <86802c440710062143x3bb801c3obb91292208073588@mail.gmail.com>
Yinghai Lu wrote:
> On 10/6/07, Jeff Garzik <jeff@garzik.org> wrote:
>> commit a606d2a111cdf948da5d69eb1de5526c5c2dafef
>> Author: Jeff Garzik <jeff@garzik.org>
>> Date: Fri Oct 5 22:56:05 2007 -0400
>>
>> [netdrvr] forcedeth: interrupt handling cleanup
>>
>> * nv_nic_irq_optimized() and nv_nic_irq_other() were complete duplicates
>> of nv_nic_irq(), with the exception of one function call. Consolidate
>> all three into a single interrupt handler, deleting a lot of redundant
>> code.
>>
>> * greatly simplify irq handler locking.
>>
>> Prior to this change, the irq handler(s) would acquire and release
>> np->lock for each action (RX, TX, other events).
>>
>> For the common case -- RX or TX work -- the lock is always acquired,
>> making all successive acquire/release combinations largely redundant.
>>
>> Acquire the lock at the beginning of the irq handler, and release it at
>> the end of the irq handler. This is simple, easy, and obvious.
>>
>> * remove irq handler work loop.
>>
>> All interesting events emanating from the irq handler either have
>> their own work loops, or they poke a timer into action.
>>
>> Therefore, delete the pointless master interrupt handler work loop.
>>
>> Signed-off-by: Jeff Garzik <jgarzik@redhat.com>
>>
>> drivers/net/forcedeth.c | 325 +++++++++++-------------------------------------
>> 1 file changed, 77 insertions(+), 248 deletions(-)
>>
> any chance to create three verion irq handlers for ioapic, msi, msi-x...?
>
> MACRO or inline function?
MSI-X already has its own separate interrupt handlers. MSI and INTx
call the same interrupt handling code, like the unmodified driver goes.
Creating an MSI-specific irq handler would not save very much AFAICS,
but I might be missing something.
Do you have ideas/suggestions for a different method?
Jeff
^ permalink raw reply
* Re: [PATCH 0/5] forcedeth: several proposed updates for testing
From: Jeff Garzik @ 2007-10-07 11:34 UTC (permalink / raw)
To: Ingo Molnar; +Cc: netdev, Ayaz Abdulla, LKML, Andrew Morton
In-Reply-To: <20071007090808.GB733@elte.hu>
Ingo Molnar wrote:
> * Jeff Garzik <jeff@garzik.org> wrote:
>
>> * I feel TX NAPI is a useful tool, because it provides an independent TX
>> process control point and system load feedback point.
>> Thus I felt this was slightly superior to tasklets.
>
> /me agrees violently
>
> btw., when i played with this tunable under -rt:
>
> enum {
> NV_OPTIMIZATION_MODE_THROUGHPUT,
> NV_OPTIMIZATION_MODE_CPU
> };
> static int optimization_mode = NV_OPTIMIZATION_MODE_THROUGHPUT;
>
> the MODE_CPU one gave (much) _higher_ bandwidth. The queueing model in
> forcedeth seemed to be not that robust and i think a single queueing
> model should be adopted instead of this tunable. (which i think just hid
> some bug/dependency) But i never got to the bottom of it so it's just
> the impression i got.
That's interesting. It will be informative to narrow down the variables
affected by this. My changes stirred the pot quite a bit :)
* 'throughput' mode enables MSI-X, and separate interrupt vectors for RX
and TX. so, NVIDIA's MSI-X implementation, our generic MSI-X support,
or "Known bugs" (see top of file) may be a factor here.
* 'throughput' mode also changes the NIC's timer interrupt frequency
* do you recall if you were running in NAPI mode? It defaulted to off
in Kconfig, but I turned it on unconditionally.
* I think TX NAPI has the potential to make the optimization_mode
irrelevant (along with the other changes, most notably the interrupt
handling change)
* and overall, yes, if we can have a single queueing model /
optimization mode I am strongly in favor of that.
Testing welcome ;-) Though these patches are raw and "hot off the
presses", so unrelated bugs are practically a certainty. And I am
worrying about the "Known bugs" note at the top. My gut feeling is that
this was, in part, misunderstanding on the part of reverse-engineers,
since corrected when NVIDIA started contributing to the driver.
Jeff
^ permalink raw reply
* [PATCH 6/n] forcedeth: protect slow path with mutex
From: Jeff Garzik @ 2007-10-07 11:23 UTC (permalink / raw)
To: netdev, Ayaz Abdulla; +Cc: LKML, Andrew Morton
In-Reply-To: <20071006151250.GA17020@havoc.gtf.org>
commit abca163a14b28c234df9bf38034bc967ff81c3aa
Author: Jeff Garzik <jeff@garzik.org>
Date: Sun Oct 7 07:22:14 2007 -0400
[netdrvr] forcedeth: wrap slow path hw manipulation inside hw_mutex
* This makes sure everybody who wants to start/stop the RX and TX engines
first acquires this mutex.
* tx_timeout code was deleted, replaced by scheduling reset_task.
* linkchange moved to a workqueue (always inside hw_mutex)
* simplified irq handling a bit
* make sure to disable workqueues before NAPI
Signed-off-by: Jeff Garzik <jgarzik@redhat.com>
drivers/net/forcedeth.c | 272 ++++++++++++++++++++++++++++++------------------
1 file changed, 175 insertions(+), 97 deletions(-)
abca163a14b28c234df9bf38034bc967ff81c3aa
diff --git a/drivers/net/forcedeth.c b/drivers/net/forcedeth.c
index a037f49..d1c1efa 100644
--- a/drivers/net/forcedeth.c
+++ b/drivers/net/forcedeth.c
@@ -63,6 +63,7 @@
#include <linux/if_vlan.h>
#include <linux/dma-mapping.h>
#include <linux/workqueue.h>
+#include <linux/mutex.h>
#include <asm/irq.h>
#include <asm/io.h>
@@ -647,6 +648,12 @@ struct nv_skb_map {
unsigned int dma_len;
};
+struct nv_mc_info {
+ u32 addr[2];
+ u32 mask[2];
+ u32 pff;
+};
+
/*
* SMP locking:
* All hardware access under dev->priv->lock, except the performance
@@ -709,6 +716,8 @@ struct fe_priv {
unsigned int pkt_limit;
struct timer_list oom_kick;
struct work_struct reset_task;
+ struct work_struct linkchange_task;
+ struct work_struct mcast_task;
struct delayed_work stats_task;
u32 reset_task_irq;
int rx_ring_size;
@@ -731,14 +740,18 @@ struct fe_priv {
int tx_ring_size;
/* vlan fields */
- struct vlan_group *vlangrp;
+ struct vlan_group *vlangrp;
/* msi/msi-x fields */
- u32 msi_flags;
- struct msix_entry msi_x_entry[NV_MSI_X_MAX_VECTORS];
+ u32 msi_flags;
+ struct msix_entry msi_x_entry[NV_MSI_X_MAX_VECTORS];
/* flow control */
- u32 pause_flags;
+ u32 pause_flags;
+
+ struct mutex hw_mutex;
+
+ struct nv_mc_info mci;
};
/*
@@ -2120,27 +2133,9 @@ static void nv_tx_timeout(struct net_device *dev)
spin_lock_irq(&np->lock);
- /* 1) stop tx engine */
- nv_stop_tx(dev);
-
- /* 2) process all pending tx completions */
- if (!nv_optimized(np))
- nv_tx_done(dev, np->tx_ring_size);
- else
- nv_tx_done_optimized(dev, np->tx_ring_size);
+ np->reset_task_irq = np->irqmask;
+ schedule_work(&np->reset_task);
- /* 3) if there are dead entries: clear everything */
- if (np->get_tx_ctx != np->put_tx_ctx) {
- printk(KERN_DEBUG "%s: tx_timeout: dead entries!\n", dev->name);
- nv_drain_tx(dev);
- nv_init_tx(dev);
- setup_hw_rings(dev, NV_SETUP_TX_RING);
- }
-
- netif_wake_queue(dev);
-
- /* 4) restart tx engine */
- nv_start_tx(dev);
spin_unlock_irq(&np->lock);
}
@@ -2476,6 +2471,7 @@ static int nv_change_mtu(struct net_device *dev, int new_mtu)
* guessed, there is probably a simpler approach.
* Changing the MTU is a rare event, it shouldn't matter.
*/
+ mutex_lock(&np->hw_mutex);
nv_disable_irq(dev);
netif_tx_lock_bh(dev);
spin_lock(&np->lock);
@@ -2503,6 +2499,7 @@ static int nv_change_mtu(struct net_device *dev, int new_mtu)
spin_unlock(&np->lock);
netif_tx_unlock_bh(dev);
nv_enable_irq(dev);
+ mutex_unlock(&np->hw_mutex);
}
return 0;
}
@@ -2535,6 +2532,8 @@ static int nv_set_mac_address(struct net_device *dev, void *addr)
/* synchronized against open : rtnl_lock() held by caller */
memcpy(dev->dev_addr, macaddr->sa_data, ETH_ALEN);
+ mutex_lock(&np->hw_mutex);
+
if (netif_running(dev)) {
netif_tx_lock_bh(dev);
spin_lock_irq(&np->lock);
@@ -2552,6 +2551,8 @@ static int nv_set_mac_address(struct net_device *dev, void *addr)
} else {
nv_copy_mac_to_hw(dev);
}
+
+ mutex_unlock(&np->hw_mutex);
return 0;
}
@@ -2605,17 +2606,61 @@ static void nv_set_multicast(struct net_device *dev)
}
addr[0] |= NVREG_MCASTADDRA_FORCE;
pff |= NVREG_PFF_ALWAYS;
+
+ if (mutex_trylock(&np->hw_mutex)) {
+ spin_lock_irq(&np->lock);
+
+ nv_stop_rx(dev);
+
+ writel(addr[0], base + NvRegMulticastAddrA);
+ writel(addr[1], base + NvRegMulticastAddrB);
+ writel(mask[0], base + NvRegMulticastMaskA);
+ writel(mask[1], base + NvRegMulticastMaskB);
+ writel(pff, base + NvRegPacketFilterFlags);
+ dprintk(KERN_INFO "%s: reconfiguration for multicast lists.\n",
+ dev->name);
+
+ nv_start_rx(dev);
+
+ spin_unlock_irq(&np->lock);
+ } else {
+ spin_lock_irq(&np->lock);
+ np->mci.addr[0] = addr[0];
+ np->mci.addr[1] = addr[1];
+ np->mci.mask[0] = mask[0];
+ np->mci.mask[1] = mask[1];
+ np->mci.pff = pff;
+ spin_unlock_irq(&np->lock);
+
+ schedule_work(&np->mcast_task);
+ }
+}
+
+static void nv_mcast_task(struct work_struct *work)
+{
+ struct fe_priv *np = container_of(work, struct fe_priv, mcast_task);
+ struct net_device *dev = np->dev;
+ u8 __iomem *base = get_hwbase(dev);
+
+ mutex_lock(&np->hw_mutex);
+
spin_lock_irq(&np->lock);
+
nv_stop_rx(dev);
- writel(addr[0], base + NvRegMulticastAddrA);
- writel(addr[1], base + NvRegMulticastAddrB);
- writel(mask[0], base + NvRegMulticastMaskA);
- writel(mask[1], base + NvRegMulticastMaskB);
- writel(pff, base + NvRegPacketFilterFlags);
+
+ writel(np->mci.addr[0], base + NvRegMulticastAddrA);
+ writel(np->mci.addr[1], base + NvRegMulticastAddrB);
+ writel(np->mci.mask[0], base + NvRegMulticastMaskA);
+ writel(np->mci.mask[1], base + NvRegMulticastMaskB);
+ writel(np->mci.pff, base + NvRegPacketFilterFlags);
dprintk(KERN_INFO "%s: reconfiguration for multicast lists.\n",
dev->name);
+
nv_start_rx(dev);
+
spin_unlock_irq(&np->lock);
+
+ mutex_unlock(&np->hw_mutex);
}
static void nv_update_pause(struct net_device *dev, u32 pause_flags)
@@ -2873,6 +2918,15 @@ static void nv_linkchange(struct net_device *dev)
}
}
+static void nv_linkchange_task(struct work_struct *work)
+{
+ struct fe_priv *np = container_of(work, struct fe_priv, linkchange_task);
+
+ mutex_lock(&np->hw_mutex);
+ nv_linkchange(np->dev);
+ mutex_unlock(&np->hw_mutex);
+}
+
static void nv_link_irq(struct net_device *dev)
{
u8 __iomem *base = get_hwbase(dev);
@@ -2883,7 +2937,7 @@ static void nv_link_irq(struct net_device *dev)
dprintk(KERN_INFO "%s: link change irq, status 0x%x.\n", dev->name, miistat);
if (miistat & (NVREG_MIISTAT_LINKCHANGE))
- nv_linkchange(dev);
+ schedule_work(&np->linkchange_task);
dprintk(KERN_DEBUG "%s: link change notification done.\n", dev->name);
}
@@ -2894,34 +2948,39 @@ static irqreturn_t __nv_nic_irq(struct net_device *dev, bool optimized)
u32 events;
int handled = 0;
u32 upd_mask = 0;
+ bool msix = (np->msi_flags & NV_MSI_X_ENABLED);
dprintk(KERN_DEBUG "%s: nv_nic_irq%s\n", dev->name,
optimized ? "_optimized" : "");
- spin_lock(&np->lock);
-
- if (!(np->msi_flags & NV_MSI_X_ENABLED)) {
+ if (!msix)
events = readl(base + NvRegIrqStatus) & NVREG_IRQSTAT_MASK;
- writel(NVREG_IRQSTAT_MASK, base + NvRegIrqStatus);
- } else {
+ else
events = readl(base + NvRegMSIXIrqStatus) & NVREG_IRQSTAT_MASK;
- writel(NVREG_IRQSTAT_MASK, base + NvRegMSIXIrqStatus);
- }
dprintk(KERN_DEBUG "%s: irq: %08x\n", dev->name, events);
+ spin_lock(&np->lock);
+
if (!(events & np->irqmask))
goto out;
- if (events & NVREG_IRQ_RX_ALL) {
- netif_rx_schedule(dev, &np->napi);
+ if (!msix)
+ writel(NVREG_IRQSTAT_MASK, base + NvRegIrqStatus);
+ else
+ writel(NVREG_IRQSTAT_MASK, base + NvRegMSIXIrqStatus);
+
+ if ((events & NVREG_IRQ_RX_ALL) &&
+ (netif_rx_schedule_prep(dev, &np->tx_napi))) {
+ __netif_rx_schedule(dev, &np->napi);
/* Disable furthur receive irq's */
upd_mask |= NVREG_IRQ_RX_ALL;
}
- if (events & NVREG_IRQ_TX_ALL) {
- netif_rx_schedule(dev, &np->tx_napi);
+ if ((events & NVREG_IRQ_TX_ALL) &&
+ (netif_rx_schedule_prep(dev, &np->tx_napi))) {
+ __netif_rx_schedule(dev, &np->tx_napi);
/* Disable furthur xmit irq's */
upd_mask |= NVREG_IRQ_TX_ALL;
@@ -2930,7 +2989,7 @@ static irqreturn_t __nv_nic_irq(struct net_device *dev, bool optimized)
if (upd_mask) {
np->irqmask &= ~upd_mask;
- if (np->msi_flags & NV_MSI_X_ENABLED)
+ if (msix)
writel(upd_mask, base + NvRegIrqMask);
else
writel(np->irqmask, base + NvRegIrqMask);
@@ -2940,7 +2999,7 @@ static irqreturn_t __nv_nic_irq(struct net_device *dev, bool optimized)
nv_link_irq(dev);
if (unlikely(np->need_linktimer && time_after(jiffies, np->link_timeout))) {
- nv_linkchange(dev);
+ schedule_work(&np->linkchange_task);
np->link_timeout = jiffies + LINK_TIMEOUT;
}
@@ -2958,7 +3017,7 @@ static irqreturn_t __nv_nic_irq(struct net_device *dev, bool optimized)
if (unlikely(events & NVREG_IRQ_RECOVER_ERROR)) {
/* disable interrupts on the nic */
- if (!(np->msi_flags & NV_MSI_X_ENABLED))
+ if (!msix)
writel(0, base + NvRegIrqMask);
else
writel(np->irqmask, base + NvRegIrqMask);
@@ -3001,20 +3060,15 @@ static irqreturn_t nv_nic_irq_other(int foo, void *data)
static irqreturn_t nv_nic_irq_tx(int foo, void *data)
{
- struct net_device *dev = (struct net_device *) data;
+ struct net_device *dev = data;
struct fe_priv *np = netdev_priv(dev);
u8 __iomem *base = get_hwbase(dev);
- u32 events;
- events = readl(base + NvRegMSIXIrqStatus) & NVREG_IRQ_TX_ALL;
- writel(NVREG_IRQ_TX_ALL, base + NvRegMSIXIrqStatus);
+ writel(NVREG_IRQ_TX_ALL, base + NvRegMSIXIrqStatus); /* ack ints */
+ writel(NVREG_IRQ_TX_ALL, base + NvRegIrqMask); /* disable ints */
+
+ netif_rx_schedule(dev, &np->tx_napi);
- if (events) {
- netif_rx_schedule(dev, &np->tx_napi);
- /* disable receive interrupts on the nic */
- writel(NVREG_IRQ_TX_ALL, base + NvRegIrqMask);
- pci_push(base);
- }
return IRQ_HANDLED;
}
@@ -3090,26 +3144,21 @@ static int nv_napi_poll(struct napi_struct *napi, int budget)
static irqreturn_t nv_nic_irq_rx(int foo, void *data)
{
- struct net_device *dev = (struct net_device *) data;
+ struct net_device *dev = data;
struct fe_priv *np = netdev_priv(dev);
u8 __iomem *base = get_hwbase(dev);
- u32 events;
- events = readl(base + NvRegMSIXIrqStatus) & NVREG_IRQ_RX_ALL;
- writel(NVREG_IRQ_RX_ALL, base + NvRegMSIXIrqStatus);
+ writel(NVREG_IRQ_RX_ALL, base + NvRegMSIXIrqStatus); /* ack ints */
+ writel(NVREG_IRQ_RX_ALL, base + NvRegIrqMask); /* disable ints */
+
+ netif_rx_schedule(dev, &np->napi);
- if (events) {
- netif_rx_schedule(dev, &np->napi);
- /* disable receive interrupts on the nic */
- writel(NVREG_IRQ_RX_ALL, base + NvRegIrqMask);
- pci_push(base);
- }
return IRQ_HANDLED;
}
static irqreturn_t nv_nic_irq_test(int foo, void *data)
{
- struct net_device *dev = (struct net_device *) data;
+ struct net_device *dev = data;
struct fe_priv *np = netdev_priv(dev);
u8 __iomem *base = get_hwbase(dev);
u32 events;
@@ -3287,12 +3336,17 @@ static void nv_reset_task(struct work_struct *work)
struct net_device *dev = np->dev;
u8 __iomem *base = get_hwbase(dev);
u32 mask;
+ unsigned long flags;
+
+ mutex_lock(&np->hw_mutex);
/*
* First disable irq(s) and then
* reenable interrupts on the nic, we have to do this before calling
* nv_nic_irq because that may decide to do otherwise
*/
+ netif_tx_lock_bh(dev);
+ spin_lock_irqsave(&np->lock, flags);
if (!using_multi_irqs(dev)) {
mask = np->irqmask;
@@ -3308,11 +3362,7 @@ static void nv_reset_task(struct work_struct *work)
np->reset_task_irq = 0;
printk(KERN_INFO "forcedeth: MAC in recoverable error state\n");
- if (!netif_running(dev))
- goto out;
- netif_tx_lock_bh(dev);
- spin_lock(&np->lock);
/* stop engines */
nv_stop_txrx(dev);
nv_txrx_reset(dev);
@@ -3334,12 +3384,14 @@ static void nv_reset_task(struct work_struct *work)
/* restart rx engine */
nv_start_txrx(dev);
- spin_unlock(&np->lock);
- netif_tx_unlock_bh(dev);
-out:
writel(mask, base + NvRegIrqMask);
pci_push(base);
+
+ spin_unlock_irqrestore(&np->lock, flags);
+ netif_tx_unlock_bh(dev);
+
+ mutex_unlock(&np->hw_mutex);
}
#ifdef CONFIG_NET_POLL_CONTROLLER
@@ -4321,29 +4373,45 @@ static void nv_get_strings(struct net_device *dev, u32 stringset, u8 *buffer)
}
}
+static int nv_ethtool_begin (struct net_device *dev)
+{
+ struct fe_priv *np = get_nvpriv(dev);
+
+ mutex_lock(&np->hw_mutex);
+}
+
+static void nv_ethtool_complete (struct net_device *dev)
+{
+ struct fe_priv *np = get_nvpriv(dev);
+
+ mutex_unlock(&np->hw_mutex);
+}
+
static const struct ethtool_ops ops = {
- .get_drvinfo = nv_get_drvinfo,
- .get_link = ethtool_op_get_link,
- .get_wol = nv_get_wol,
- .set_wol = nv_set_wol,
- .get_settings = nv_get_settings,
- .set_settings = nv_set_settings,
- .get_regs_len = nv_get_regs_len,
- .get_regs = nv_get_regs,
- .nway_reset = nv_nway_reset,
- .set_tso = nv_set_tso,
- .get_ringparam = nv_get_ringparam,
- .set_ringparam = nv_set_ringparam,
- .get_pauseparam = nv_get_pauseparam,
- .set_pauseparam = nv_set_pauseparam,
- .get_rx_csum = nv_get_rx_csum,
- .set_rx_csum = nv_set_rx_csum,
- .set_tx_csum = nv_set_tx_csum,
- .set_sg = nv_set_sg,
- .get_strings = nv_get_strings,
- .get_ethtool_stats = nv_get_ethtool_stats,
- .get_sset_count = nv_get_sset_count,
- .self_test = nv_self_test,
+ .begin = nv_ethtool_begin,
+ .complete = nv_ethtool_complete,
+ .get_drvinfo = nv_get_drvinfo,
+ .get_link = ethtool_op_get_link,
+ .get_wol = nv_get_wol,
+ .set_wol = nv_set_wol,
+ .get_settings = nv_get_settings,
+ .set_settings = nv_set_settings,
+ .get_regs_len = nv_get_regs_len,
+ .get_regs = nv_get_regs,
+ .nway_reset = nv_nway_reset,
+ .set_tso = nv_set_tso,
+ .get_ringparam = nv_get_ringparam,
+ .set_ringparam = nv_set_ringparam,
+ .get_pauseparam = nv_get_pauseparam,
+ .set_pauseparam = nv_set_pauseparam,
+ .get_rx_csum = nv_get_rx_csum,
+ .set_rx_csum = nv_set_rx_csum,
+ .set_tx_csum = nv_set_tx_csum,
+ .set_sg = nv_set_sg,
+ .get_strings = nv_get_strings,
+ .get_ethtool_stats = nv_get_ethtool_stats,
+ .get_sset_count = nv_get_sset_count,
+ .self_test = nv_self_test,
};
static void nv_vlan_rx_register(struct net_device *dev, struct vlan_group *grp)
@@ -4524,9 +4592,13 @@ static int nv_open(struct net_device *dev)
}
/* set linkspeed to invalid value, thus force nv_update_linkspeed
* to init hw */
+ mutex_lock(&np->hw_mutex);
np->linkspeed = 0;
ret = nv_update_linkspeed(dev);
+ mutex_unlock(&np->hw_mutex);
+
nv_start_txrx(dev);
+
napi_enable(&np->napi);
napi_enable(&np->tx_napi);
netif_start_queue(dev);
@@ -4558,13 +4630,17 @@ static int nv_close(struct net_device *dev)
u8 __iomem *base;
netif_stop_queue(dev);
- napi_disable(&np->napi);
- napi_disable(&np->tx_napi);
- synchronize_irq(dev->irq);
del_timer_sync(&np->oom_kick);
cancel_rearming_delayed_work(&np->stats_task);
cancel_work_sync(&np->reset_task);
+ cancel_work_sync(&np->linkchange_task);
+ cancel_work_sync(&np->mcast_task);
+
+ napi_disable(&np->napi);
+ napi_disable(&np->tx_napi);
+
+ synchronize_irq(dev->irq);
spin_lock_irq(&np->lock);
nv_stop_txrx(dev);
@@ -4619,6 +4695,8 @@ static int __devinit nv_probe(struct pci_dev *pci_dev, const struct pci_device_i
np->oom_kick.data = (unsigned long) dev;
np->oom_kick.function = &nv_do_rx_refill; /* timer handler */
INIT_WORK(&np->reset_task, nv_reset_task);
+ INIT_WORK(&np->linkchange_task, nv_linkchange_task);
+ INIT_WORK(&np->mcast_task, nv_mcast_task);
INIT_DELAYED_WORK(&np->stats_task, nv_stats_task);
err = pci_enable_device(pci_dev);
^ permalink raw reply related
* libertas and endianness
From: Geert Uytterhoeven @ 2007-10-07 10:15 UTC (permalink / raw)
To: Jeff Garzik, netdev, linux-wireless
Cc: Andrew Morton, Linux Kernel Development
Somehow (haven't found out why it suddenly got compiled, no .config
changes) this showed up in the list of warnings in 2.6.23-rc9 compared
to -rc8 on one of my m68k builds:
| drivers/net/wireless/libertas/cmd.c:189: warning: large integer implicitly truncated to unsigned type
| drivers/net/wireless/libertas/cmd.c:195: warning: large integer implicitly truncated to unsigned type
The offending lines are:
| wep->keytype[i] = cpu_to_le16(cmd_type_wep_40_bit);
| wep->keytype[i] = cpu_to_le16(cmd_type_wep_104_bit);
I.e. it tries to store 0x0100 resp. 0x0200 into keytype[i], which is is u8.
Gr{oetje,eeting}s,
Geert
--
Geert Uytterhoeven -- There's lots of Linux beyond ia32 -- geert@linux-m68k.org
In personal conversations with technical people, I call myself a hacker. But
when I'm talking to journalists I just say "programmer" or something like that.
-- Linus Torvalds
^ permalink raw reply
* Re: [RFC/PATCH 2/4] UDP memory usage accounting (take 4): accounting unit and variable
From: Evgeniy Polyakov @ 2007-10-07 10:09 UTC (permalink / raw)
To: Satoshi OSHIMA
Cc: Andi Kleen, David Miller, Herbert Xu, netdev, ?? ??,
Yumiko SUGITA, ??@RedHat
In-Reply-To: <470651B3.6060001@hitachi.com>
Hi.
On Sat, Oct 06, 2007 at 12:01:07AM +0900, Satoshi OSHIMA (satoshi.oshima.fk@hitachi.com) wrote:
> --- 2.6.23-rc3-udp_limit.orig/net/ipv4/udp.c
> +++ 2.6.23-rc3-udp_limit/net/ipv4/udp.c
> @@ -113,6 +113,10 @@ DEFINE_SNMP_STAT(struct udp_mib, udp_sta
> struct hlist_head udp_hash[UDP_HTABLE_SIZE];
> DEFINE_RWLOCK(udp_hash_lock);
>
> +atomic_t udp_memory_allocated;
> +
> +EXPORT_SYMBOL(udp_memory_allocated);
> +
Why do you export this variable?
It is not accessed from modules in your patchset.
--
Evgeniy Polyakov
^ permalink raw reply
* Re: [PATCH 0/5] forcedeth: several proposed updates for testing
From: Ingo Molnar @ 2007-10-07 9:08 UTC (permalink / raw)
To: Jeff Garzik; +Cc: netdev, Ayaz Abdulla, LKML, Andrew Morton
In-Reply-To: <20071006151250.GA17020@havoc.gtf.org>
* Jeff Garzik <jeff@garzik.org> wrote:
> * I feel TX NAPI is a useful tool, because it provides an independent TX
> process control point and system load feedback point.
> Thus I felt this was slightly superior to tasklets.
/me agrees violently
btw., when i played with this tunable under -rt:
enum {
NV_OPTIMIZATION_MODE_THROUGHPUT,
NV_OPTIMIZATION_MODE_CPU
};
static int optimization_mode = NV_OPTIMIZATION_MODE_THROUGHPUT;
the MODE_CPU one gave (much) _higher_ bandwidth. The queueing model in
forcedeth seemed to be not that robust and i think a single queueing
model should be adopted instead of this tunable. (which i think just hid
some bug/dependency) But i never got to the bottom of it so it's just
the impression i got.
Ingo
^ permalink raw reply
* Re: [PATCH 2/5] forcedeth: interrupt handling cleanup
From: Ingo Molnar @ 2007-10-07 9:03 UTC (permalink / raw)
To: Jeff Garzik; +Cc: netdev, Ayaz Abdulla, LKML, Andrew Morton
In-Reply-To: <20071006151400.GB17488@havoc.gtf.org>
* Jeff Garzik <jeff@garzik.org> wrote:
> - spin_unlock(&np->lock);
> - printk(KERN_DEBUG "%s: too many iterations (%d) in nv_nic_irq.\n", dev->name, i);
> - break;
i like that! One forcedeth annoyance that triggers frequently on one of
my testboxes is:
[ 120.955202] eth0: too many iterations (6) in nv_nic_irq.
[ 121.233865] eth0: too many iterations (6) in nv_nic_irq.
[ 129.215450] eth0: too many iterations (6) in nv_nic_irq.
[ 139.734408] eth0: too many iterations (6) in nv_nic_irq.
[ 144.546811] eth0: too many iterations (6) in nv_nic_irq.
[ 153.811005] eth0: too many iterations (6) in nv_nic_irq.
[ 154.695879] eth0: too many iterations (6) in nv_nic_irq.
[ 155.455078] eth0: too many iterations (6) in nv_nic_irq.
[ 173.912162] eth0: too many iterations (6) in nv_nic_irq.
Ingo
^ permalink raw reply
* Re: [RFC][PATCH 1/2] TCP: fix lost retransmit detection
From: TAKANO Ryousei @ 2007-10-07 8:42 UTC (permalink / raw)
To: davem; +Cc: ilpo.jarvinen, netdev, y-kodama
In-Reply-To: <20071006.231714.41173786.davem@davemloft.net>
From: David Miller <davem@davemloft.net>
Subject: Re: [RFC][PATCH 1/2] TCP: fix lost retransmit detection
Date: Sat, 06 Oct 2007 23:17:14 -0700 (PDT)
> From: TAKANO Ryousei <takano@axe-inc.co.jp>
> Date: Sun, 07 Oct 2007 14:51:00 +0900 (JST)
>
> > BTW, what is difference among netdev-2.6, net-2.6 (net-2.6.24), and
> > tcp-2.6? I am not familiar with linux kernel development process.
>
> net-2.6.24 tree is current new development
>
> net-2.6 tree is only bug fixes
>
> tcp-2.6 is an old abandonded tree that holds lots of old
> TCP work, most of which is merged already, but one part
> (my RB-Tree SACK patches) are not integrated yet.
Thanks, and I got it.
I need to check the related recent posts (SACK block validation, sacktag
cache usage, and RT-tree SACK) on this list.
Ryousei Takano
^ permalink raw reply
* Re: [RFC][PATCH 1/2] TCP: fix lost retransmit detection
From: David Miller @ 2007-10-07 6:17 UTC (permalink / raw)
To: takano; +Cc: ilpo.jarvinen, netdev, y-kodama
In-Reply-To: <20071007.145100.00679926.takano@axe-inc.co.jp>
From: TAKANO Ryousei <takano@axe-inc.co.jp>
Date: Sun, 07 Oct 2007 14:51:00 +0900 (JST)
> BTW, what is difference among netdev-2.6, net-2.6 (net-2.6.24), and
> tcp-2.6? I am not familiar with linux kernel development process.
net-2.6.24 tree is current new development
net-2.6 tree is only bug fixes
tcp-2.6 is an old abandonded tree that holds lots of old
TCP work, most of which is merged already, but one part
(my RB-Tree SACK patches) are not integrated yet.
^ permalink raw reply
* Re: [RFC][PATCH 2/2] TCP: skip processing cached SACK blocks
From: TAKANO Ryousei @ 2007-10-07 6:11 UTC (permalink / raw)
To: ilpo.jarvinen; +Cc: netdev, y-kodama
In-Reply-To: <Pine.LNX.4.64.0710041516170.31129@kivilampi-30.cs.helsinki.fi>
Hi Ilpo,
Thanks for your reply.
Most of my response are in the reply to patch 1/2.
From: "Ilpo Järvinen" <ilpo.jarvinen@helsinki.fi>
Subject: Re: [RFC][PATCH 2/2] TCP: skip processing cached SACK blocks
Date: Fri, 5 Oct 2007 13:37:21 +0300 (EEST)
> On Thu, 4 Oct 2007, TAKANO Ryousei wrote:
>
> > This patch allows to process only newly reported SACK blocks at the
> > sender side. An ACK packet contains up to three SACK blocks, and some
>
> "A SACK option that specifies n blocks will have a length of 8*n+2
> bytes, so the 40 bytes available for TCP options can specify a
> maximum of 4 blocks. It is expected that SACK will often be used in
> conjunction with the Timestamp option used for RTTM [Jacobson92],
> which takes an additional 10 bytes (plus two bytes of padding); thus
> a maximum of 3 SACK blocks will be allowed in this case." [RFC2018]
>
> :-)
>
Yes, indeed:-)
> > of them may be already reported and processed blocks. This patch
> > prevents processing of such already processed SACK blocks.
> >
> > Signed-off-by: Ryousei Takano <takano-ryousei@aist.go.jp>
> > Signed-off-by: Yuetsu Kodama <y-kodama@aist.go.jp>
> > ---
> > net/ipv4/tcp_input.c | 24 ++++++++++++++++++++++++
> > 1 files changed, 24 insertions(+), 0 deletions(-)
> >
> > diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c
> > index bbad2cd..9615fc9 100644
> > --- a/net/ipv4/tcp_input.c
> > +++ b/net/ipv4/tcp_input.c
> > @@ -978,6 +978,7 @@ tcp_sacktag_write_queue(struct sock *sk, struct sk_buff *ack_skb, u32 prior_snd_
> > int cached_fack_count;
> > int i;
> > int first_sack_index;
> > + u8 sack_block_skip[4] = {0,0,0,0};
> >
> > if (!tp->sacked_out)
> > tp->fackets_out = 0;
> > @@ -1012,6 +1013,21 @@ tcp_sacktag_write_queue(struct sock *sk, struct sk_buff *ack_skb, u32 prior_snd_
> > if (before(TCP_SKB_CB(ack_skb)->ack_seq, prior_snd_una - tp->max_window))
> > return 0;
> >
> > + /* Skip processing cached SACK blocks. */
> > + for (i = 0; i < num_sacks; i++) {
> > + __be32 start_seq = sp[i].start_seq;
> > + __be32 end_seq = sp[i].end_seq;
> > + int j;
> > +
> > + for (j = 0; j < ARRAY_SIZE(tp->recv_sack_cache); j++) {
> > + if ((tp->recv_sack_cache[j].start_seq == start_seq) &&
> > + (tp->recv_sack_cache[j].end_seq == end_seq)) {
> > + sack_block_skip[i] = 1;
> > + break;
> > + }
> > + }
> > + }
> > +
>
> I'm somewhat against adding more and more special cases to sacktag,
> there's still need for more special cases after this one to avoid very
> expensive processing (I guess they just won't occur in your scenario)!
> ...I would rather remove whole special case mess of the fastpath and
> have a more generic solution (see the patch I point into in the reply
> to patch 1/2)...
>
> > /* SACK fastpath:
> > * if the only SACK change is the increase of the end_seq of
> > * the first block then only apply that SACK block
> > @@ -1051,11 +1067,16 @@ tcp_sacktag_write_queue(struct sock *sk, struct sk_buff *ack_skb, u32 prior_snd_
> > if (after(ntohl(sp[j].start_seq),
> > ntohl(sp[j+1].start_seq))){
> > struct tcp_sack_block_wire tmp;
> > + u8 sbtmp;
> >
> > tmp = sp[j];
> > sp[j] = sp[j+1];
> > sp[j+1] = tmp;
> >
> > + sbtmp = sack_block_skip[j];
> > + sack_block_skip[j] = sack_block_skip[j+1];
> > + sack_block_skip[j+1] = sbtmp;
> > +
> > /* Track where the first SACK block goes to */
> > if (j == first_sack_index)
> > first_sack_index = j+1;
> > @@ -1083,6 +1104,9 @@ tcp_sacktag_write_queue(struct sock *sk, struct sk_buff *ack_skb, u32 prior_snd_
> > int fack_count;
> > int dup_sack = (found_dup_sack && (i == first_sack_index));
> >
> > + if (sack_block_skip[i])
>
> DSACKs must always be processed, so please add:
>
> && !dup_sack
>
I did not notice DSACKs. Thanks.
> > + continue;
>
> By doing this skipping here, you actually end up crippling lost_retrans
> detection even more than it was broken before. ...You probably didn't just
> notice that during tests because of unrelated suboptimal behavior (in
> fastpath_skb_hint handling). ...Anyway, correctness of this should be
> evaluated against the fixed lost_retrans, rather than the already
> broken one.
>
You can find the result that the average goodput slightly improves
against the only PATCH #1 (fixed lost_retrans) applied kernel.
Sorry, our web page is down this weekend for the power outage.
http://projects.gtrc.aist.go.jp/gnet/sack-bug.html
> > +
> > skb = cached_skb;
> > fack_count = cached_fack_count;
>
> Other than what's noted above:
>
> Acked-by: Ilpo Järvinen <ilpo.jarvinen@helsinki.fi>
>
>
> --
> i.
Regards,
Ryousei Takano
^ permalink raw reply
* Re: [RFC][PATCH 1/2] TCP: fix lost retransmit detection
From: TAKANO Ryousei @ 2007-10-07 5:51 UTC (permalink / raw)
To: ilpo.jarvinen; +Cc: netdev, y-kodama
In-Reply-To: <Pine.LNX.4.64.0710041419560.31129@kivilampi-30.cs.helsinki.fi>
From: "Ilpo Järvinen" <ilpo.jarvinen@helsinki.fi>
Subject: Re: [RFC][PATCH 1/2] TCP: fix lost retransmit detection
Date: Fri, 5 Oct 2007 13:02:07 +0300 (EEST)
> On Thu, 4 Oct 2007, TAKANO Ryousei wrote:
>
> > This patch allows to detect loss of retransmitted packets more
> > accurately by using the highest end sequence number among SACK
> > blocks. Before the retransmission queue is scanned, the highest
> > end sequence number (high_end_seq) is retrieved, and this value
> > is compared with the ack_seq of each packet.
> >
> > Signed-off-by: Ryousei Takano <takano-ryousei@aist.go.jp>
> > Signed-off-by: Yuetsu Kodama <y-kodama@aist.go.jp>
> > ---
> > net/ipv4/tcp_input.c | 14 +++++++++++---
> > 1 files changed, 11 insertions(+), 3 deletions(-)
> >
> > diff --git a/net/ipv4/tcp_input.c b/net/ipv4/tcp_input.c
> > index bbad2cd..12db4b3 100644
> > --- a/net/ipv4/tcp_input.c
> > +++ b/net/ipv4/tcp_input.c
> > @@ -978,6 +978,7 @@ tcp_sacktag_write_queue(struct sock *sk, struct sk_buff *ack_skb, u32 prior_snd_
> > int cached_fack_count;
> > int i;
> > int first_sack_index;
> > + __u32 high_end_seq;
>
> No __-types when not visible to userspace please.
>
I will fix it.
> >
> > if (!tp->sacked_out)
> > tp->fackets_out = 0;
> > @@ -1012,6 +1013,14 @@ tcp_sacktag_write_queue(struct sock *sk, struct sk_buff *ack_skb, u32 prior_snd_
> > if (before(TCP_SKB_CB(ack_skb)->ack_seq, prior_snd_una - tp->max_window))
> > return 0;
> >
> > + /* Retrieve the highest end_seq among SACK blocks. */
> > + high_end_seq = ntohl(sp[0].end_seq);
> > + for (i = 1; i < num_sacks; i++) {
> > + __u32 end_seq = ntohl(sp[i].end_seq);
> > + if (after(end_seq, high_end_seq))
> > + high_end_seq = end_seq;
> > + }
> > +
>
> There's one problem... Net-2.6.24 tree includes SACK block validator
> which is being done in the marking loop. The SACK blocks would not yet be
> validated in that position, yet this code should be protected by the
> validation! My intention is to move the validator earlier anyway (yet to
> be split into smaller logical patches), see:
>
> http://marc.info/?l=linux-netdev&m=119062989408053&w=2
>
I will check the net-2.6.24 tree and your patch.
> > /* SACK fastpath:
> > * if the only SACK change is the increase of the end_seq of
> > * the first block then only apply that SACK block
> > @@ -1161,9 +1170,8 @@ tcp_sacktag_write_queue(struct sock *sk, struct sk_buff *ack_skb, u32 prior_snd_
> > }
> >
> > if ((sacked&TCPCB_SACKED_RETRANS) &&
> > - after(end_seq, TCP_SKB_CB(skb)->ack_seq) &&
> > - (!lost_retrans || after(end_seq, lost_retrans)))
> > - lost_retrans = end_seq;
> > + after(high_end_seq, TCP_SKB_CB(skb)->ack_seq))
> > + lost_retrans = high_end_seq;
>
> Just couple of thoughts, not that this change itself is incorrect...
>
> In case sacktag uses fastpath, this code won't be executed for the skb's
> that we would like to check (those with SACKED_RETRANS set, that are
> below the fastpath_skb_hint). We will eventually deal with the whole queue
> when fastpath_skb_hint gets set to NULL, with the next cumulative ACK that
> fully ACKs an skb at the latest. Maybe there's a need for a larger surgery
> than this to fix it. I think we need additional field to tcp_sock to avoid
> doing a full-walk per ACK:
>
I think the problem occurs in slowpath. For example, in case when the receiver
detects and sends back a new SACK block, the sender may fail to detect loss
of a retransmitted packet.
> Keep minimum of TCP_SKB_CB(skb)->ack_seq of rexmitted segments in
> tcp_sock, when that's exceeded by SACK block, do a full-walk in the
> lost_retrans worker loop like the old code does...
>
>
> In future, please base your work to current development tree instead of
> linus' tree (net-2.6.24 at this point of time, there's also tcp-2.6 but
> it's currently a bit outdated).
>
Thanks for your suggestion. SACK processing has a heavy workload and
it is complex. I agree to make efforts toward a more generic solution.
Your recv_sack_cache patch seems valuable. I will continue to work in
the net-2.6.24 tree, and resend our patches.
Anyway, first of all, I would like to share this problem with kernel
developers.
BTW, what is difference among netdev-2.6, net-2.6 (net-2.6.24), and
tcp-2.6? I am not familiar with linux kernel development process.
>
> --
> i.
Regards,
Ryousei Takano
^ permalink raw reply
* Re: [PATCH 2/5] forcedeth: interrupt handling cleanup
From: Yinghai Lu @ 2007-10-07 4:43 UTC (permalink / raw)
To: Jeff Garzik; +Cc: netdev, Ayaz Abdulla, LKML, Andrew Morton
In-Reply-To: <20071006151400.GB17488@havoc.gtf.org>
On 10/6/07, Jeff Garzik <jeff@garzik.org> wrote:
>
> commit a606d2a111cdf948da5d69eb1de5526c5c2dafef
> Author: Jeff Garzik <jeff@garzik.org>
> Date: Fri Oct 5 22:56:05 2007 -0400
>
> [netdrvr] forcedeth: interrupt handling cleanup
>
> * nv_nic_irq_optimized() and nv_nic_irq_other() were complete duplicates
> of nv_nic_irq(), with the exception of one function call. Consolidate
> all three into a single interrupt handler, deleting a lot of redundant
> code.
>
> * greatly simplify irq handler locking.
>
> Prior to this change, the irq handler(s) would acquire and release
> np->lock for each action (RX, TX, other events).
>
> For the common case -- RX or TX work -- the lock is always acquired,
> making all successive acquire/release combinations largely redundant.
>
> Acquire the lock at the beginning of the irq handler, and release it at
> the end of the irq handler. This is simple, easy, and obvious.
>
> * remove irq handler work loop.
>
> All interesting events emanating from the irq handler either have
> their own work loops, or they poke a timer into action.
>
> Therefore, delete the pointless master interrupt handler work loop.
>
> Signed-off-by: Jeff Garzik <jgarzik@redhat.com>
>
> drivers/net/forcedeth.c | 325 +++++++++++-------------------------------------
> 1 file changed, 77 insertions(+), 248 deletions(-)
>
any chance to create three verion irq handlers for ioapic, msi, msi-x...?
MACRO or inline function?
YH
^ permalink raw reply
* Re: [PATCH] net/core: split dev_ifsioc() according to locking
From: Arnd Bergmann @ 2007-10-07 0:17 UTC (permalink / raw)
To: Jeff Garzik; +Cc: David Miller, netdev, LKML, Andrew Morton
In-Reply-To: <20071006204212.GA32177@havoc.gtf.org>
On Saturday 06 October 2007, Jeff Garzik wrote:
>
> This always bugged me: dev_ioctl() called dev_ifsioc() either inside
> read_lock(dev_base_lock) or rtnl_lock(), depending on the ioctl being
> executed.
>
> This change moves the ioctls executed inside dev_base_lock to a new
> function, dev_ifsioc_locked(). Now the locking context is completely
> clear to the reader.
>
> Signed-off-by: Jeff Garzik <jgarzik@redhat.com>
Great idea!
I've been experimenting with a new compat_dev_ioctl() function along
the lines of what I just posted for the blkdev ioctls. For that, it
would be perfect to streamline dev_ioctl further:
* move the dev_load() and locking into dev_ifsioc{,_locked}
* move the copy_to_user step to a single place at the end of dev_ioctl
After that, we could have very simple dev_ioctl and compat_dev_ioctl
functions calling the same dev_ifsioc{,_locked} functions.
Arnd <><
^ permalink raw reply
* [PATCH] net/core: split dev_ifsioc() according to locking
From: Jeff Garzik @ 2007-10-06 20:42 UTC (permalink / raw)
To: David Miller; +Cc: netdev, LKML, Andrew Morton
This always bugged me: dev_ioctl() called dev_ifsioc() either inside
read_lock(dev_base_lock) or rtnl_lock(), depending on the ioctl being
executed.
This change moves the ioctls executed inside dev_base_lock to a new
function, dev_ifsioc_locked(). Now the locking context is completely
clear to the reader.
Signed-off-by: Jeff Garzik <jgarzik@redhat.com>
---
net/core/dev.c | 88 +++++++++++++++++++++++++++++++++++++--------------------
1 file changed, 58 insertions(+), 30 deletions(-)
diff --git a/net/core/dev.c b/net/core/dev.c
index d998646..ea57527 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -3083,9 +3083,9 @@ int dev_set_mac_address(struct net_device *dev, struct sockaddr *sa)
}
/*
- * Perform the SIOCxIFxxx calls.
+ * Perform the SIOCxIFxxx calls, inside read_lock(dev_base_lock)
*/
-static int dev_ifsioc(struct net *net, struct ifreq *ifr, unsigned int cmd)
+static int dev_ifsioc_locked(struct net *net, struct ifreq *ifr, unsigned int cmd)
{
int err;
struct net_device *dev = __dev_get_by_name(net, ifr->ifr_name);
@@ -3098,25 +3098,15 @@ static int dev_ifsioc(struct net *net, struct ifreq *ifr, unsigned int cmd)
ifr->ifr_flags = dev_get_flags(dev);
return 0;
- case SIOCSIFFLAGS: /* Set interface flags */
- return dev_change_flags(dev, ifr->ifr_flags);
-
case SIOCGIFMETRIC: /* Get the metric on the interface
(currently unused) */
ifr->ifr_metric = 0;
return 0;
- case SIOCSIFMETRIC: /* Set the metric on the interface
- (currently unused) */
- return -EOPNOTSUPP;
-
case SIOCGIFMTU: /* Get the MTU of a device */
ifr->ifr_mtu = dev->mtu;
return 0;
- case SIOCSIFMTU: /* Set the MTU of a device */
- return dev_set_mtu(dev, ifr->ifr_mtu);
-
case SIOCGIFHWADDR:
if (!dev->addr_len)
memset(ifr->ifr_hwaddr.sa_data, 0, sizeof ifr->ifr_hwaddr.sa_data);
@@ -3126,6 +3116,61 @@ static int dev_ifsioc(struct net *net, struct ifreq *ifr, unsigned int cmd)
ifr->ifr_hwaddr.sa_family = dev->type;
return 0;
+ case SIOCGIFSLAVE:
+ err = -EINVAL;
+ break;
+
+ case SIOCGIFMAP:
+ ifr->ifr_map.mem_start = dev->mem_start;
+ ifr->ifr_map.mem_end = dev->mem_end;
+ ifr->ifr_map.base_addr = dev->base_addr;
+ ifr->ifr_map.irq = dev->irq;
+ ifr->ifr_map.dma = dev->dma;
+ ifr->ifr_map.port = dev->if_port;
+ return 0;
+
+ case SIOCGIFINDEX:
+ ifr->ifr_ifindex = dev->ifindex;
+ return 0;
+
+ case SIOCGIFTXQLEN:
+ ifr->ifr_qlen = dev->tx_queue_len;
+ return 0;
+
+ default:
+ /* dev_ioctl() should ensure this case
+ * is never reached
+ */
+ WARN_ON(1);
+ err = -EINVAL;
+ break;
+
+ }
+ return err;
+}
+
+/*
+ * Perform the SIOCxIFxxx calls, inside rtnl_lock()
+ */
+static int dev_ifsioc(struct net *net, struct ifreq *ifr, unsigned int cmd)
+{
+ int err;
+ struct net_device *dev = __dev_get_by_name(net, ifr->ifr_name);
+
+ if (!dev)
+ return -ENODEV;
+
+ switch (cmd) {
+ case SIOCSIFFLAGS: /* Set interface flags */
+ return dev_change_flags(dev, ifr->ifr_flags);
+
+ case SIOCSIFMETRIC: /* Set the metric on the interface
+ (currently unused) */
+ return -EOPNOTSUPP;
+
+ case SIOCSIFMTU: /* Set the MTU of a device */
+ return dev_set_mtu(dev, ifr->ifr_mtu);
+
case SIOCSIFHWADDR:
return dev_set_mac_address(dev, &ifr->ifr_hwaddr);
@@ -3137,15 +3182,6 @@ static int dev_ifsioc(struct net *net, struct ifreq *ifr, unsigned int cmd)
call_netdevice_notifiers(NETDEV_CHANGEADDR, dev);
return 0;
- case SIOCGIFMAP:
- ifr->ifr_map.mem_start = dev->mem_start;
- ifr->ifr_map.mem_end = dev->mem_end;
- ifr->ifr_map.base_addr = dev->base_addr;
- ifr->ifr_map.irq = dev->irq;
- ifr->ifr_map.dma = dev->dma;
- ifr->ifr_map.port = dev->if_port;
- return 0;
-
case SIOCSIFMAP:
if (dev->set_config) {
if (!netif_device_present(dev))
@@ -3172,14 +3208,6 @@ static int dev_ifsioc(struct net *net, struct ifreq *ifr, unsigned int cmd)
return dev_mc_delete(dev, ifr->ifr_hwaddr.sa_data,
dev->addr_len, 1);
- case SIOCGIFINDEX:
- ifr->ifr_ifindex = dev->ifindex;
- return 0;
-
- case SIOCGIFTXQLEN:
- ifr->ifr_qlen = dev->tx_queue_len;
- return 0;
-
case SIOCSIFTXQLEN:
if (ifr->ifr_qlen < 0)
return -EINVAL;
@@ -3290,7 +3318,7 @@ int dev_ioctl(struct net *net, unsigned int cmd, void __user *arg)
case SIOCGIFTXQLEN:
dev_load(net, ifr.ifr_name);
read_lock(&dev_base_lock);
- ret = dev_ifsioc(net, &ifr, cmd);
+ ret = dev_ifsioc_locked(net, &ifr, cmd);
read_unlock(&dev_base_lock);
if (!ret) {
if (colon)
^ permalink raw reply related
* Re: MSI interrupts and disable_irq
From: Jeff Garzik @ 2007-10-06 17:59 UTC (permalink / raw)
To: Yinghai Lu
Cc: Ayaz Abdulla, Manfred Spraul, nedev, Linux Kernel Mailing List,
David Miller, Andrew Morton
In-Reply-To: <86802c440710061043u6e51cd7q468346bd06b08657@mail.gmail.com>
Yinghai Lu wrote:
> On 9/28/07, Jeff Garzik <jgarzik@pobox.com> wrote:
>> Ayaz Abdulla wrote:
>>> I am trying to track down a forcedeth driver issue described by bug 9047
>>> in bugzilla (2.6.23-rc7-git1 forcedeth w/ MCP55 oops under heavy load).
>>> I added a patch to synchronize the timer handlers so that one handler
>>> doesn't accidently enable the IRQ while another timer handler is running
>>> (see attachment 'Add timer lock' in bug report) and for other processing
>>> protection.
>>>
>>> However, the system still had an Oops. So I added a lock around the
>>> nv_rx_process_optimized() and the Oops has not happened (see attachment
>>> 'New patch for locking' in bug report). This would imply a
>>> synchronization issue. However, the only callers of that function are
>>> the IRQ handler and the timer handlers (in non-NAPI case). The timer
>>> handlers use disable_irq so that the IRQ handler does not contend with
>>> them. It looks as if disable_irq is not working properly.
>>>
>>> This issue repros only with MSI interrupt and not legacy INTx
>>> interrupts. Any ideas?
>> (added linux-kernel to CC, since I think it's more of a general kernel
>> issue)
>>
> I wonder if the race is between soft_timer for nv_do_nic_poll from
> different CPUs
Interested parties should try the forcedeth patches I just posted :)
Jeff
^ permalink raw reply
* Re: MSI interrupts and disable_irq
From: Yinghai Lu @ 2007-10-06 17:43 UTC (permalink / raw)
To: Jeff Garzik
Cc: Ayaz Abdulla, Manfred Spraul, nedev, Linux Kernel Mailing List,
David Miller, Andrew Morton
In-Reply-To: <46FDBCB4.9090802@pobox.com>
On 9/28/07, Jeff Garzik <jgarzik@pobox.com> wrote:
> Ayaz Abdulla wrote:
> > I am trying to track down a forcedeth driver issue described by bug 9047
> > in bugzilla (2.6.23-rc7-git1 forcedeth w/ MCP55 oops under heavy load).
> > I added a patch to synchronize the timer handlers so that one handler
> > doesn't accidently enable the IRQ while another timer handler is running
> > (see attachment 'Add timer lock' in bug report) and for other processing
> > protection.
> >
> > However, the system still had an Oops. So I added a lock around the
> > nv_rx_process_optimized() and the Oops has not happened (see attachment
> > 'New patch for locking' in bug report). This would imply a
> > synchronization issue. However, the only callers of that function are
> > the IRQ handler and the timer handlers (in non-NAPI case). The timer
> > handlers use disable_irq so that the IRQ handler does not contend with
> > them. It looks as if disable_irq is not working properly.
> >
> > This issue repros only with MSI interrupt and not legacy INTx
> > interrupts. Any ideas?
>
> (added linux-kernel to CC, since I think it's more of a general kernel
> issue)
>
I wonder if the race is between soft_timer for nv_do_nic_poll from
different CPUs
YH
^ permalink raw reply
* Re: [PATCH 0/5] forcedeth: several proposed updates for testing
From: Jeff Garzik @ 2007-10-06 15:24 UTC (permalink / raw)
To: netdev, Ayaz Abdulla; +Cc: LKML, Andrew Morton
In-Reply-To: <20071006151250.GA17020@havoc.gtf.org>
Jeff Garzik wrote:
> The goals of these changes are:
> * move the driver towards a more sane, simple, easy to verify locking
> setup -- irq handler would often acquire/release the lock twice
> for each interrupt -- and hopefully
s/and hopefully// (it became the next bullet point)
^ permalink raw reply
* Re: [PATCH 0/5] forcedeth: several proposed updates for testing
From: Jeff Garzik @ 2007-10-06 15:17 UTC (permalink / raw)
To: netdev, Ayaz Abdulla; +Cc: LKML, Andrew Morton
In-Reply-To: <20071006151250.GA17020@havoc.gtf.org>
On Sat, Oct 06, 2007 at 11:12:50AM -0400, Jeff Garzik wrote:
>
> The 'fe-lock' branch of
> git://git.kernel.org/pub/scm/linux/kernel/git/jgarzik/netdev-2.6.git fe-lock
It should also be pointed out that these patches were generated on top
of davem's net-2.6.24.git tree.
They -probably- apply to -mm, but you might have to remove the forcedeth
patches -mm already has, before applying.
^ permalink raw reply
* [PATCH 5/5] forcedeth: timer overhaul
From: Jeff Garzik @ 2007-10-06 15:15 UTC (permalink / raw)
To: netdev, Ayaz Abdulla; +Cc: LKML, Andrew Morton
In-Reply-To: <20071006151250.GA17020@havoc.gtf.org>
commit d7c766113ee2ec66ae8975e0acbad086d2c23594
Author: Jeff Garzik <jeff@garzik.org>
Date: Sat Oct 6 10:57:56 2007 -0400
[netdrvr] forcedeth: timer overhaul
* convert stats_poll timer to a delayed-work workqueue stats_task
* protect hw stats update with a lock
* now that recovery is the only remaining use of nv_do_nic_poll(),
rename it to nv_reset_task(), move it from a timer to a workqueue,
and delete all non-recovery-related code from the function.
* kill np->in_shutdown, it mirrors netif_running(). furthermore,
the overwhelming majority of sites that tested np->in_shutdown
were inside rtnl_lock() and guaranteed never to race against shutdown
anyway.
Signed-off-by: Jeff Garzik <jgarzik@redhat.com>
drivers/net/forcedeth.c | 194 ++++++++++++++++++------------------------------
1 file changed, 75 insertions(+), 119 deletions(-)
d7c766113ee2ec66ae8975e0acbad086d2c23594
diff --git a/drivers/net/forcedeth.c b/drivers/net/forcedeth.c
index d6eacd7..a037f49 100644
--- a/drivers/net/forcedeth.c
+++ b/drivers/net/forcedeth.c
@@ -62,6 +62,7 @@
#include <linux/init.h>
#include <linux/if_vlan.h>
#include <linux/dma-mapping.h>
+#include <linux/workqueue.h>
#include <asm/irq.h>
#include <asm/io.h>
@@ -450,7 +451,6 @@ union ring_type {
#define NV_PKTLIMIT_2 9100 /* Actual limit according to NVidia: 9202 */
#define OOM_REFILL (1+HZ/20)
-#define POLL_WAIT (1+HZ/100)
#define LINK_TIMEOUT (3*HZ)
#define STATS_INTERVAL (10*HZ)
@@ -670,7 +670,6 @@ struct fe_priv {
* Locking: spin_lock(&np->lock); */
struct net_device_stats stats;
struct nv_ethtool_stats estats;
- int in_shutdown;
u32 linkspeed;
int duplex;
int autoneg;
@@ -681,7 +680,6 @@ struct fe_priv {
unsigned int phy_model;
u16 gigabit;
int intr_test;
- int recover_error;
/* General data: RO fields */
dma_addr_t ring_addr;
@@ -710,9 +708,9 @@ struct fe_priv {
unsigned int rx_buf_sz;
unsigned int pkt_limit;
struct timer_list oom_kick;
- struct timer_list nic_poll;
- struct timer_list stats_poll;
- u32 nic_poll_irq;
+ struct work_struct reset_task;
+ struct delayed_work stats_task;
+ u32 reset_task_irq;
int rx_ring_size;
/* media detection workaround.
@@ -1388,7 +1386,7 @@ static void nv_mac_reset(struct net_device *dev)
pci_push(base);
}
-static void nv_get_hw_stats(struct net_device *dev)
+static void __nv_get_hw_stats(struct net_device *dev)
{
struct fe_priv *np = netdev_priv(dev);
u8 __iomem *base = get_hwbase(dev);
@@ -1443,6 +1441,16 @@ static void nv_get_hw_stats(struct net_device *dev)
}
}
+static void nv_get_hw_stats(struct net_device *dev)
+{
+ struct fe_priv *np = netdev_priv(dev);
+ unsigned long flags;
+
+ spin_lock_irqsave(&np->lock, flags);
+ __nv_get_hw_stats(dev);
+ spin_unlock_irqrestore(&np->lock, flags);
+}
+
/*
* nv_get_stats: dev->get_stats function
* Get latest stats value from the nic.
@@ -2478,10 +2486,9 @@ static int nv_change_mtu(struct net_device *dev, int new_mtu)
nv_drain_txrx(dev);
/* reinit driver view of the rx queue */
set_bufsize(dev);
- if (nv_init_ring(dev)) {
- if (!np->in_shutdown)
- mod_timer(&np->oom_kick, jiffies + OOM_REFILL);
- }
+ if (nv_init_ring(dev))
+ mod_timer(&np->oom_kick, jiffies + OOM_REFILL);
+
/* reinit nic view of the rx queue */
writel(np->rx_buf_sz, base + NvRegOffloadConfig);
setup_hw_rings(dev, NV_SETUP_RX_RING | NV_SETUP_TX_RING);
@@ -2957,10 +2964,9 @@ static irqreturn_t __nv_nic_irq(struct net_device *dev, bool optimized)
writel(np->irqmask, base + NvRegIrqMask);
pci_push(base);
- if (!np->in_shutdown) {
- np->nic_poll_irq = np->irqmask;
- np->recover_error = 1;
- mod_timer(&np->nic_poll, jiffies + POLL_WAIT);
+ if (netif_running(dev)) {
+ np->reset_task_irq = np->irqmask;
+ schedule_work(&np->reset_task);
}
}
@@ -3061,8 +3067,7 @@ static int nv_napi_poll(struct napi_struct *napi, int budget)
if (retcode) {
spin_lock_irqsave(&np->lock, flags);
- if (!np->in_shutdown)
- mod_timer(&np->oom_kick, jiffies + OOM_REFILL);
+ mod_timer(&np->oom_kick, jiffies + OOM_REFILL);
spin_unlock_irqrestore(&np->lock, flags);
}
@@ -3276,12 +3281,12 @@ static void nv_free_irq(struct net_device *dev)
}
}
-static void nv_do_nic_poll(unsigned long data)
+static void nv_reset_task(struct work_struct *work)
{
- struct net_device *dev = (struct net_device *) data;
- struct fe_priv *np = netdev_priv(dev);
+ struct fe_priv *np = container_of(work, struct fe_priv, reset_task);
+ struct net_device *dev = np->dev;
u8 __iomem *base = get_hwbase(dev);
- u32 mask = 0;
+ u32 mask;
/*
* First disable irq(s) and then
@@ -3290,88 +3295,51 @@ static void nv_do_nic_poll(unsigned long data)
*/
if (!using_multi_irqs(dev)) {
- if (np->msi_flags & NV_MSI_X_ENABLED)
- disable_irq_lockdep(np->msi_x_entry[NV_MSI_X_VECTOR_ALL].vector);
- else
- disable_irq_lockdep(dev->irq);
mask = np->irqmask;
} else {
- if (np->nic_poll_irq & NVREG_IRQ_RX_ALL) {
- disable_irq_lockdep(np->msi_x_entry[NV_MSI_X_VECTOR_RX].vector);
+ mask = 0;
+ if (np->reset_task_irq & NVREG_IRQ_RX_ALL)
mask |= NVREG_IRQ_RX_ALL;
- }
- if (np->nic_poll_irq & NVREG_IRQ_TX_ALL) {
- disable_irq_lockdep(np->msi_x_entry[NV_MSI_X_VECTOR_TX].vector);
+ if (np->reset_task_irq & NVREG_IRQ_TX_ALL)
mask |= NVREG_IRQ_TX_ALL;
- }
- if (np->nic_poll_irq & NVREG_IRQ_OTHER) {
- disable_irq_lockdep(np->msi_x_entry[NV_MSI_X_VECTOR_OTHER].vector);
+ if (np->reset_task_irq & NVREG_IRQ_OTHER)
mask |= NVREG_IRQ_OTHER;
- }
}
- np->nic_poll_irq = 0;
+ np->reset_task_irq = 0;
- if (np->recover_error) {
- np->recover_error = 0;
- printk(KERN_INFO "forcedeth: MAC in recoverable error state\n");
- if (netif_running(dev)) {
- netif_tx_lock_bh(dev);
- spin_lock(&np->lock);
- /* stop engines */
- nv_stop_txrx(dev);
- nv_txrx_reset(dev);
- /* drain rx queue */
- nv_drain_txrx(dev);
- /* reinit driver view of the rx queue */
- set_bufsize(dev);
- if (nv_init_ring(dev)) {
- if (!np->in_shutdown)
- mod_timer(&np->oom_kick, jiffies + OOM_REFILL);
- }
- /* reinit nic view of the rx queue */
- writel(np->rx_buf_sz, base + NvRegOffloadConfig);
- setup_hw_rings(dev, NV_SETUP_RX_RING | NV_SETUP_TX_RING);
- writel( ((np->rx_ring_size-1) << NVREG_RINGSZ_RXSHIFT) + ((np->tx_ring_size-1) << NVREG_RINGSZ_TXSHIFT),
- base + NvRegRingSizes);
- pci_push(base);
- writel(NVREG_TXRXCTL_KICK|np->txrxctl_bits, get_hwbase(dev) + NvRegTxRxControl);
- pci_push(base);
+ printk(KERN_INFO "forcedeth: MAC in recoverable error state\n");
+ if (!netif_running(dev))
+ goto out;
- /* restart rx engine */
- nv_start_txrx(dev);
- spin_unlock(&np->lock);
- netif_tx_unlock_bh(dev);
- }
- }
+ netif_tx_lock_bh(dev);
+ spin_lock(&np->lock);
+ /* stop engines */
+ nv_stop_txrx(dev);
+ nv_txrx_reset(dev);
+ /* drain rx queue */
+ nv_drain_txrx(dev);
+ /* reinit driver view of the rx queue */
+ set_bufsize(dev);
+ if (nv_init_ring(dev))
+ mod_timer(&np->oom_kick, jiffies + OOM_REFILL);
- /* FIXME: Do we need synchronize_irq(dev->irq) here? */
+ /* reinit nic view of the rx queue */
+ writel(np->rx_buf_sz, base + NvRegOffloadConfig);
+ setup_hw_rings(dev, NV_SETUP_RX_RING | NV_SETUP_TX_RING);
+ writel( ((np->rx_ring_size-1) << NVREG_RINGSZ_RXSHIFT) + ((np->tx_ring_size-1) << NVREG_RINGSZ_TXSHIFT),
+ base + NvRegRingSizes);
+ pci_push(base);
+ writel(NVREG_TXRXCTL_KICK|np->txrxctl_bits, get_hwbase(dev) + NvRegTxRxControl);
+ pci_push(base);
+
+ /* restart rx engine */
+ nv_start_txrx(dev);
+ spin_unlock(&np->lock);
+ netif_tx_unlock_bh(dev);
+out:
writel(mask, base + NvRegIrqMask);
pci_push(base);
-
- if (!using_multi_irqs(dev)) {
- if (nv_optimized(np))
- nv_nic_irq_optimized(0, dev);
- else
- nv_nic_irq(0, dev);
- if (np->msi_flags & NV_MSI_X_ENABLED)
- enable_irq_lockdep(np->msi_x_entry[NV_MSI_X_VECTOR_ALL].vector);
- else
- enable_irq_lockdep(dev->irq);
- } else {
- if (np->nic_poll_irq & NVREG_IRQ_RX_ALL) {
- nv_nic_irq_rx(0, dev);
- enable_irq_lockdep(np->msi_x_entry[NV_MSI_X_VECTOR_RX].vector);
- }
- if (np->nic_poll_irq & NVREG_IRQ_TX_ALL) {
- nv_nic_irq_tx(0, dev);
- enable_irq_lockdep(np->msi_x_entry[NV_MSI_X_VECTOR_TX].vector);
- }
- if (np->nic_poll_irq & NVREG_IRQ_OTHER) {
- nv_nic_irq_other(0, dev);
- enable_irq_lockdep(np->msi_x_entry[NV_MSI_X_VECTOR_OTHER].vector);
- }
- }
}
#ifdef CONFIG_NET_POLL_CONTROLLER
@@ -3386,15 +3354,15 @@ static void nv_poll_controller(struct net_device *dev)
}
#endif
-static void nv_do_stats_poll(unsigned long data)
+static void nv_stats_task(struct work_struct *_work)
{
- struct net_device *dev = (struct net_device *) data;
- struct fe_priv *np = netdev_priv(dev);
+ struct delayed_work *work = (struct delayed_work *) _work;
+ struct fe_priv *np = container_of(work, struct fe_priv, stats_task);
+ struct net_device *dev = np->dev;
nv_get_hw_stats(dev);
- if (!np->in_shutdown)
- mod_timer(&np->stats_poll, jiffies + STATS_INTERVAL);
+ schedule_delayed_work(work, STATS_INTERVAL);
}
static void nv_get_drvinfo(struct net_device *dev, struct ethtool_drvinfo *info)
@@ -3840,10 +3808,8 @@ static int nv_set_ringparam(struct net_device *dev, struct ethtool_ringparam* ri
if (netif_running(dev)) {
/* reinit driver view of the queues */
set_bufsize(dev);
- if (nv_init_ring(dev)) {
- if (!np->in_shutdown)
- mod_timer(&np->oom_kick, jiffies + OOM_REFILL);
- }
+ if (nv_init_ring(dev))
+ mod_timer(&np->oom_kick, jiffies + OOM_REFILL);
/* reinit nic view of the queues */
writel(np->rx_buf_sz, base + NvRegOffloadConfig);
@@ -4024,7 +3990,7 @@ static void nv_get_ethtool_stats(struct net_device *dev, struct ethtool_stats *e
struct fe_priv *np = netdev_priv(dev);
/* update stats */
- nv_do_stats_poll((unsigned long)dev);
+ nv_get_hw_stats(dev);
memcpy(buffer, &np->estats, nv_get_sset_count(dev, ETH_SS_STATS)*sizeof(u64));
}
@@ -4322,10 +4288,9 @@ static void nv_self_test(struct net_device *dev, struct ethtool_test *test, u64
if (netif_running(dev)) {
/* reinit driver view of the rx queue */
set_bufsize(dev);
- if (nv_init_ring(dev)) {
- if (!np->in_shutdown)
- mod_timer(&np->oom_kick, jiffies + OOM_REFILL);
- }
+ if (nv_init_ring(dev))
+ mod_timer(&np->oom_kick, jiffies + OOM_REFILL);
+
/* reinit nic view of the rx queue */
writel(np->rx_buf_sz, base + NvRegOffloadConfig);
setup_hw_rings(dev, NV_SETUP_RX_RING | NV_SETUP_TX_RING);
@@ -4473,8 +4438,6 @@ static int nv_open(struct net_device *dev)
nv_txrx_reset(dev);
writel(0, base + NvRegUnknownSetupReg6);
- np->in_shutdown = 0;
-
/* give hw rings */
setup_hw_rings(dev, NV_SETUP_RX_RING | NV_SETUP_TX_RING);
writel( ((np->rx_ring_size-1) << NVREG_RINGSZ_RXSHIFT) + ((np->tx_ring_size-1) << NVREG_RINGSZ_TXSHIFT),
@@ -4579,7 +4542,7 @@ static int nv_open(struct net_device *dev)
/* start statistics timer */
if (np->driver_data & (DEV_HAS_STATISTICS_V1|DEV_HAS_STATISTICS_V2))
- mod_timer(&np->stats_poll, jiffies + STATS_INTERVAL);
+ schedule_delayed_work(&np->stats_task, STATS_INTERVAL);
spin_unlock_irq(&np->lock);
@@ -4594,17 +4557,14 @@ static int nv_close(struct net_device *dev)
struct fe_priv *np = netdev_priv(dev);
u8 __iomem *base;
- spin_lock_irq(&np->lock);
- np->in_shutdown = 1;
- spin_unlock_irq(&np->lock);
netif_stop_queue(dev);
napi_disable(&np->napi);
napi_disable(&np->tx_napi);
synchronize_irq(dev->irq);
del_timer_sync(&np->oom_kick);
- del_timer_sync(&np->nic_poll);
- del_timer_sync(&np->stats_poll);
+ cancel_rearming_delayed_work(&np->stats_task);
+ cancel_work_sync(&np->reset_task);
spin_lock_irq(&np->lock);
nv_stop_txrx(dev);
@@ -4658,12 +4618,8 @@ static int __devinit nv_probe(struct pci_dev *pci_dev, const struct pci_device_i
init_timer(&np->oom_kick);
np->oom_kick.data = (unsigned long) dev;
np->oom_kick.function = &nv_do_rx_refill; /* timer handler */
- init_timer(&np->nic_poll);
- np->nic_poll.data = (unsigned long) dev;
- np->nic_poll.function = &nv_do_nic_poll; /* timer handler */
- init_timer(&np->stats_poll);
- np->stats_poll.data = (unsigned long) dev;
- np->stats_poll.function = &nv_do_stats_poll; /* timer handler */
+ INIT_WORK(&np->reset_task, nv_reset_task);
+ INIT_DELAYED_WORK(&np->stats_task, nv_stats_task);
err = pci_enable_device(pci_dev);
if (err) {
^ permalink raw reply related
* [PATCH 4/5] forcedeth: internal simplification and cleanups
From: Jeff Garzik @ 2007-10-06 15:14 UTC (permalink / raw)
To: netdev, Ayaz Abdulla; +Cc: LKML, Andrew Morton
In-Reply-To: <20071006151250.GA17020@havoc.gtf.org>
commit 39572457a4dfe9a9dc1efd6641e7a6467e5658a1
Author: Jeff Garzik <jeff@garzik.org>
Date: Sat Oct 6 01:21:01 2007 -0400
[netdrvr] forcedeth: internal simplification and cleanups
* remove changelog from source; its kept in git repository
* split guts of RX/TX DMA engine disable into disable portion,
and wait/etc. portions.
* consolidate descriptor version tests using nv_optimized()
* consolidate NIC DMA start, stop and drain into
nv_start_txrx(), nv_stop_txrx(), nv_drain_txrx()
* change nv_poll_controller() to call interrupt handling function
Signed-off-by: Jeff Garzik <jgarzik@redhat.com>
drivers/net/forcedeth.c | 228 +++++++++++++++++-------------------------------
1 file changed, 84 insertions(+), 144 deletions(-)
39572457a4dfe9a9dc1efd6641e7a6467e5658a1
diff --git a/drivers/net/forcedeth.c b/drivers/net/forcedeth.c
index 1c236e6..d6eacd7 100644
--- a/drivers/net/forcedeth.c
+++ b/drivers/net/forcedeth.c
@@ -29,89 +29,7 @@
* along with this program; if not, write to the Free Software
* Foundation, Inc., 59 Temple Place, Suite 330, Boston, MA 02111-1307 USA
*
- * Changelog:
- * 0.01: 05 Oct 2003: First release that compiles without warnings.
- * 0.02: 05 Oct 2003: Fix bug for nv_drain_tx: do not try to free NULL skbs.
- * Check all PCI BARs for the register window.
- * udelay added to mii_rw.
- * 0.03: 06 Oct 2003: Initialize dev->irq.
- * 0.04: 07 Oct 2003: Initialize np->lock, reduce handled irqs, add printks.
- * 0.05: 09 Oct 2003: printk removed again, irq status print tx_timeout.
- * 0.06: 10 Oct 2003: MAC Address read updated, pff flag generation updated,
- * irq mask updated
- * 0.07: 14 Oct 2003: Further irq mask updates.
- * 0.08: 20 Oct 2003: rx_desc.Length initialization added, nv_alloc_rx refill
- * added into irq handler, NULL check for drain_ring.
- * 0.09: 20 Oct 2003: Basic link speed irq implementation. Only handle the
- * requested interrupt sources.
- * 0.10: 20 Oct 2003: First cleanup for release.
- * 0.11: 21 Oct 2003: hexdump for tx added, rx buffer sizes increased.
- * MAC Address init fix, set_multicast cleanup.
- * 0.12: 23 Oct 2003: Cleanups for release.
- * 0.13: 25 Oct 2003: Limit for concurrent tx packets increased to 10.
- * Set link speed correctly. start rx before starting
- * tx (nv_start_rx sets the link speed).
- * 0.14: 25 Oct 2003: Nic dependant irq mask.
- * 0.15: 08 Nov 2003: fix smp deadlock with set_multicast_list during
- * open.
- * 0.16: 15 Nov 2003: include file cleanup for ppc64, rx buffer size
- * increased to 1628 bytes.
- * 0.17: 16 Nov 2003: undo rx buffer size increase. Substract 1 from
- * the tx length.
- * 0.18: 17 Nov 2003: fix oops due to late initialization of dev_stats
- * 0.19: 29 Nov 2003: Handle RxNoBuf, detect & handle invalid mac
- * addresses, really stop rx if already running
- * in nv_start_rx, clean up a bit.
- * 0.20: 07 Dec 2003: alloc fixes
- * 0.21: 12 Jan 2004: additional alloc fix, nic polling fix.
- * 0.22: 19 Jan 2004: reprogram timer to a sane rate, avoid lockup
- * on close.
- * 0.23: 26 Jan 2004: various small cleanups
- * 0.24: 27 Feb 2004: make driver even less anonymous in backtraces
- * 0.25: 09 Mar 2004: wol support
- * 0.26: 03 Jun 2004: netdriver specific annotation, sparse-related fixes
- * 0.27: 19 Jun 2004: Gigabit support, new descriptor rings,
- * added CK804/MCP04 device IDs, code fixes
- * for registers, link status and other minor fixes.
- * 0.28: 21 Jun 2004: Big cleanup, making driver mostly endian safe
- * 0.29: 31 Aug 2004: Add backup timer for link change notification.
- * 0.30: 25 Sep 2004: rx checksum support for nf 250 Gb. Add rx reset
- * into nv_close, otherwise reenabling for wol can
- * cause DMA to kfree'd memory.
- * 0.31: 14 Nov 2004: ethtool support for getting/setting link
- * capabilities.
- * 0.32: 16 Apr 2005: RX_ERROR4 handling added.
- * 0.33: 16 May 2005: Support for MCP51 added.
- * 0.34: 18 Jun 2005: Add DEV_NEED_LINKTIMER to all nForce nics.
- * 0.35: 26 Jun 2005: Support for MCP55 added.
- * 0.36: 28 Jun 2005: Add jumbo frame support.
- * 0.37: 10 Jul 2005: Additional ethtool support, cleanup of pci id list
- * 0.38: 16 Jul 2005: tx irq rewrite: Use global flags instead of
- * per-packet flags.
- * 0.39: 18 Jul 2005: Add 64bit descriptor support.
- * 0.40: 19 Jul 2005: Add support for mac address change.
- * 0.41: 30 Jul 2005: Write back original MAC in nv_close instead
- * of nv_remove
- * 0.42: 06 Aug 2005: Fix lack of link speed initialization
- * in the second (and later) nv_open call
- * 0.43: 10 Aug 2005: Add support for tx checksum.
- * 0.44: 20 Aug 2005: Add support for scatter gather and segmentation.
- * 0.45: 18 Sep 2005: Remove nv_stop/start_rx from every link check
- * 0.46: 20 Oct 2005: Add irq optimization modes.
- * 0.47: 26 Oct 2005: Add phyaddr 0 in phy scan.
- * 0.48: 24 Dec 2005: Disable TSO, bugfix for pci_map_single
- * 0.49: 10 Dec 2005: Fix tso for large buffers.
- * 0.50: 20 Jan 2006: Add 8021pq tagging support.
- * 0.51: 20 Jan 2006: Add 64bit consistent memory allocation for rings.
- * 0.52: 20 Jan 2006: Add MSI/MSIX support.
- * 0.53: 19 Mar 2006: Fix init from low power mode and add hw reset.
- * 0.54: 21 Mar 2006: Fix spin locks for multi irqs and cleanup.
- * 0.55: 22 Mar 2006: Add flow control (pause frame).
- * 0.56: 22 Mar 2006: Additional ethtool config and moduleparam support.
- * 0.57: 14 May 2006: Mac address set in probe/remove and order corrections.
- * 0.58: 30 Oct 2006: Added support for sideband management unit.
- * 0.59: 30 Oct 2006: Added support for recoverable error.
- * 0.60: 20 Jan 2007: Code optimizations for rings, rx & tx data paths, and stats.
+ ****************************************************************************
*
* Known bugs:
* We suspect that on some hardware no TX done interrupts are generated.
@@ -122,7 +40,9 @@
* DEV_NEED_TIMERIRQ from the driver_data flags.
* DEV_NEED_TIMERIRQ will not harm you on sane hardware, only generating a few
* superfluous timer interrupts from the nic.
+ *
*/
+
#define FORCEDETH_VERSION "1.00"
#define DRV_NAME "forcedeth"
@@ -893,6 +813,12 @@ static inline void pci_push(u8 __iomem *base)
readl(base);
}
+static bool nv_optimized(struct fe_priv *np)
+{
+ return (np->desc_ver == DESC_VER_1 || np->desc_ver == DESC_VER_2) ?
+ false : true;
+}
+
static inline u32 nv_descr_getlength(struct ring_desc *prd, u32 v)
{
return le32_to_cpu(prd->flaglen)
@@ -1342,18 +1268,28 @@ static void nv_start_rx(struct net_device *dev)
pci_push(base);
}
-static void nv_stop_rx(struct net_device *dev)
+static void __nv_stop_rx(struct net_device *dev)
{
struct fe_priv *np = netdev_priv(dev);
u8 __iomem *base = get_hwbase(dev);
u32 rx_ctrl = readl(base + NvRegReceiverControl);
- dprintk(KERN_DEBUG "%s: nv_stop_rx\n", dev->name);
if (!np->mac_in_use)
rx_ctrl &= ~NVREG_RCVCTL_START;
else
rx_ctrl |= NVREG_RCVCTL_RX_PATH_EN;
writel(rx_ctrl, base + NvRegReceiverControl);
+}
+
+static void nv_stop_rx(struct net_device *dev)
+{
+ struct fe_priv *np = netdev_priv(dev);
+ u8 __iomem *base = get_hwbase(dev);
+
+ dprintk(KERN_DEBUG "%s: nv_stop_rx\n", dev->name);
+
+ __nv_stop_rx(dev);
+
reg_delay(dev, NvRegReceiverStatus, NVREG_RCVSTAT_BUSY, 0,
NV_RXSTOP_DELAY1, NV_RXSTOP_DELAY1MAX,
KERN_INFO "nv_stop_rx: ReceiverStatus remained busy");
@@ -1377,18 +1313,28 @@ static void nv_start_tx(struct net_device *dev)
pci_push(base);
}
-static void nv_stop_tx(struct net_device *dev)
+static void __nv_stop_tx(struct net_device *dev)
{
struct fe_priv *np = netdev_priv(dev);
u8 __iomem *base = get_hwbase(dev);
u32 tx_ctrl = readl(base + NvRegTransmitterControl);
- dprintk(KERN_DEBUG "%s: nv_stop_tx\n", dev->name);
if (!np->mac_in_use)
tx_ctrl &= ~NVREG_XMITCTL_START;
else
tx_ctrl |= NVREG_XMITCTL_TX_PATH_EN;
writel(tx_ctrl, base + NvRegTransmitterControl);
+}
+
+static void nv_stop_tx(struct net_device *dev)
+{
+ struct fe_priv *np = netdev_priv(dev);
+ u8 __iomem *base = get_hwbase(dev);
+
+ dprintk(KERN_DEBUG "%s: nv_stop_tx\n", dev->name);
+
+ __nv_stop_tx(dev);
+
reg_delay(dev, NvRegTransmitterStatus, NVREG_XMITSTAT_BUSY, 0,
NV_TXSTOP_DELAY1, NV_TXSTOP_DELAY1MAX,
KERN_INFO "nv_stop_tx: TransmitterStatus remained busy");
@@ -1399,6 +1345,18 @@ static void nv_stop_tx(struct net_device *dev)
base + NvRegTransmitPoll);
}
+static void nv_stop_txrx(struct net_device *dev)
+{
+ nv_stop_rx(dev);
+ nv_stop_tx(dev);
+}
+
+static void nv_start_txrx(struct net_device *dev)
+{
+ nv_start_rx(dev);
+ nv_start_tx(dev);
+}
+
static void nv_txrx_reset(struct net_device *dev)
{
struct fe_priv *np = netdev_priv(dev);
@@ -1651,7 +1609,7 @@ static int nv_init_ring(struct net_device *dev)
nv_init_tx(dev);
nv_init_rx(dev);
- if (np->desc_ver == DESC_VER_1 || np->desc_ver == DESC_VER_2)
+ if (!nv_optimized(np))
return nv_alloc_rx(dev);
else
return nv_alloc_rx_optimized(dev);
@@ -1723,7 +1681,7 @@ static void nv_drain_rx(struct net_device *dev)
}
}
-static void drain_ring(struct net_device *dev)
+static void nv_drain_txrx(struct net_device *dev)
{
nv_drain_tx(dev);
nv_drain_rx(dev);
@@ -2158,7 +2116,7 @@ static void nv_tx_timeout(struct net_device *dev)
nv_stop_tx(dev);
/* 2) process all pending tx completions */
- if (np->desc_ver == DESC_VER_1 || np->desc_ver == DESC_VER_2)
+ if (!nv_optimized(np))
nv_tx_done(dev, np->tx_ring_size);
else
nv_tx_done_optimized(dev, np->tx_ring_size);
@@ -2514,12 +2472,10 @@ static int nv_change_mtu(struct net_device *dev, int new_mtu)
netif_tx_lock_bh(dev);
spin_lock(&np->lock);
/* stop engines */
- nv_stop_rx(dev);
- nv_stop_tx(dev);
+ nv_stop_txrx(dev);
nv_txrx_reset(dev);
/* drain rx queue */
- nv_drain_rx(dev);
- nv_drain_tx(dev);
+ nv_drain_txrx(dev);
/* reinit driver view of the rx queue */
set_bufsize(dev);
if (nv_init_ring(dev)) {
@@ -2536,8 +2492,7 @@ static int nv_change_mtu(struct net_device *dev, int new_mtu)
pci_push(base);
/* restart rx engine */
- nv_start_rx(dev);
- nv_start_tx(dev);
+ nv_start_txrx(dev);
spin_unlock(&np->lock);
netif_tx_unlock_bh(dev);
nv_enable_irq(dev);
@@ -3067,7 +3022,7 @@ static int nv_napi_tx_poll(struct napi_struct *napi, int budget)
spin_lock_irqsave(&np->lock, flags);
- if (np->desc_ver == DESC_VER_1 || np->desc_ver == DESC_VER_2)
+ if (!nv_optimized(np))
pkts = nv_tx_done(dev, budget);
else
pkts = nv_tx_done_optimized(dev, budget);
@@ -3096,7 +3051,7 @@ static int nv_napi_poll(struct napi_struct *napi, int budget)
unsigned long flags;
int pkts, retcode;
- if (np->desc_ver == DESC_VER_1 || np->desc_ver == DESC_VER_2) {
+ if (!nv_optimized(np)) {
pkts = nv_rx_process(dev, budget);
retcode = nv_alloc_rx(dev);
} else {
@@ -3214,7 +3169,7 @@ static int nv_request_irq(struct net_device *dev, int intr_test)
if (intr_test) {
handler = nv_nic_irq_test;
} else {
- if (np->desc_ver == DESC_VER_3)
+ if (nv_optimized(np))
handler = nv_nic_irq_optimized;
else
handler = nv_nic_irq;
@@ -3363,12 +3318,10 @@ static void nv_do_nic_poll(unsigned long data)
netif_tx_lock_bh(dev);
spin_lock(&np->lock);
/* stop engines */
- nv_stop_rx(dev);
- nv_stop_tx(dev);
+ nv_stop_txrx(dev);
nv_txrx_reset(dev);
/* drain rx queue */
- nv_drain_rx(dev);
- nv_drain_tx(dev);
+ nv_drain_txrx(dev);
/* reinit driver view of the rx queue */
set_bufsize(dev);
if (nv_init_ring(dev)) {
@@ -3385,8 +3338,7 @@ static void nv_do_nic_poll(unsigned long data)
pci_push(base);
/* restart rx engine */
- nv_start_rx(dev);
- nv_start_tx(dev);
+ nv_start_txrx(dev);
spin_unlock(&np->lock);
netif_tx_unlock_bh(dev);
}
@@ -3398,7 +3350,7 @@ static void nv_do_nic_poll(unsigned long data)
pci_push(base);
if (!using_multi_irqs(dev)) {
- if (np->desc_ver == DESC_VER_3)
+ if (nv_optimized(np))
nv_nic_irq_optimized(0, dev);
else
nv_nic_irq(0, dev);
@@ -3425,7 +3377,12 @@ static void nv_do_nic_poll(unsigned long data)
#ifdef CONFIG_NET_POLL_CONTROLLER
static void nv_poll_controller(struct net_device *dev)
{
- nv_do_nic_poll((unsigned long) dev);
+ struct fe_priv *np = netdev_priv(dev);
+ unsigned long flags;
+
+ local_irq_save(flags);
+ __nv_nic_irq(dev, nv_optimized(np));
+ local_irq_restore(flags);
}
#endif
@@ -3595,8 +3552,7 @@ static int nv_set_settings(struct net_device *dev, struct ethtool_cmd *ecmd)
netif_tx_lock_bh(dev);
spin_lock(&np->lock);
/* stop engines */
- nv_stop_rx(dev);
- nv_stop_tx(dev);
+ nv_stop_txrx(dev);
spin_unlock(&np->lock);
netif_tx_unlock_bh(dev);
}
@@ -3702,8 +3658,7 @@ static int nv_set_settings(struct net_device *dev, struct ethtool_cmd *ecmd)
}
if (netif_running(dev)) {
- nv_start_rx(dev);
- nv_start_tx(dev);
+ nv_start_txrx(dev);
nv_enable_irq(dev);
}
@@ -3746,8 +3701,7 @@ static int nv_nway_reset(struct net_device *dev)
netif_tx_lock_bh(dev);
spin_lock(&np->lock);
/* stop engines */
- nv_stop_rx(dev);
- nv_stop_tx(dev);
+ nv_stop_txrx(dev);
spin_unlock(&np->lock);
netif_tx_unlock_bh(dev);
printk(KERN_INFO "%s: link down.\n", dev->name);
@@ -3767,8 +3721,7 @@ static int nv_nway_reset(struct net_device *dev)
}
if (netif_running(dev)) {
- nv_start_rx(dev);
- nv_start_tx(dev);
+ nv_start_txrx(dev);
nv_enable_irq(dev);
}
ret = 0;
@@ -3859,12 +3812,10 @@ static int nv_set_ringparam(struct net_device *dev, struct ethtool_ringparam* ri
netif_tx_lock_bh(dev);
spin_lock(&np->lock);
/* stop engines */
- nv_stop_rx(dev);
- nv_stop_tx(dev);
+ nv_stop_txrx(dev);
nv_txrx_reset(dev);
/* drain queues */
- nv_drain_rx(dev);
- nv_drain_tx(dev);
+ nv_drain_txrx(dev);
/* delete queues */
free_rings(dev);
}
@@ -3904,8 +3855,7 @@ static int nv_set_ringparam(struct net_device *dev, struct ethtool_ringparam* ri
pci_push(base);
/* restart engines */
- nv_start_rx(dev);
- nv_start_tx(dev);
+ nv_start_txrx(dev);
spin_unlock(&np->lock);
netif_tx_unlock_bh(dev);
nv_enable_irq(dev);
@@ -3946,8 +3896,7 @@ static int nv_set_pauseparam(struct net_device *dev, struct ethtool_pauseparam*
netif_tx_lock_bh(dev);
spin_lock(&np->lock);
/* stop engines */
- nv_stop_rx(dev);
- nv_stop_tx(dev);
+ nv_stop_txrx(dev);
spin_unlock(&np->lock);
netif_tx_unlock_bh(dev);
}
@@ -3988,8 +3937,7 @@ static int nv_set_pauseparam(struct net_device *dev, struct ethtool_pauseparam*
}
if (netif_running(dev)) {
- nv_start_rx(dev);
- nv_start_tx(dev);
+ nv_start_txrx(dev);
nv_enable_irq(dev);
}
return 0;
@@ -4225,8 +4173,7 @@ static int nv_loopback_test(struct net_device *dev)
pci_push(base);
/* restart rx engine */
- nv_start_rx(dev);
- nv_start_tx(dev);
+ nv_start_txrx(dev);
/* setup packet for tx */
pkt_len = ETH_DATA_LEN;
@@ -4304,12 +4251,10 @@ static int nv_loopback_test(struct net_device *dev)
dev_kfree_skb_any(tx_skb);
out:
/* stop engines */
- nv_stop_rx(dev);
- nv_stop_tx(dev);
+ nv_stop_txrx(dev);
nv_txrx_reset(dev);
/* drain rx queue */
- nv_drain_rx(dev);
- nv_drain_tx(dev);
+ nv_drain_txrx(dev);
if (netif_running(dev)) {
writel(misc1_flags, base + NvRegMisc1);
@@ -4346,12 +4291,10 @@ static void nv_self_test(struct net_device *dev, struct ethtool_test *test, u64
writel(NVREG_IRQSTAT_MASK, base + NvRegMSIXIrqStatus);
}
/* stop engines */
- nv_stop_rx(dev);
- nv_stop_tx(dev);
+ nv_stop_txrx(dev);
nv_txrx_reset(dev);
/* drain rx queue */
- nv_drain_rx(dev);
- nv_drain_tx(dev);
+ nv_drain_txrx(dev);
spin_unlock_irq(&np->lock);
netif_tx_unlock_bh(dev);
}
@@ -4392,8 +4335,7 @@ static void nv_self_test(struct net_device *dev, struct ethtool_test *test, u64
writel(NVREG_TXRXCTL_KICK|np->txrxctl_bits, get_hwbase(dev) + NvRegTxRxControl);
pci_push(base);
/* restart rx engine */
- nv_start_rx(dev);
- nv_start_tx(dev);
+ nv_start_txrx(dev);
napi_enable(&np->napi);
napi_enable(&np->tx_napi);
netif_start_queue(dev);
@@ -4621,8 +4563,7 @@ static int nv_open(struct net_device *dev)
* to init hw */
np->linkspeed = 0;
ret = nv_update_linkspeed(dev);
- nv_start_rx(dev);
- nv_start_tx(dev);
+ nv_start_txrx(dev);
napi_enable(&np->napi);
napi_enable(&np->tx_napi);
netif_start_queue(dev);
@@ -4644,7 +4585,7 @@ static int nv_open(struct net_device *dev)
return 0;
out_drain:
- drain_ring(dev);
+ nv_drain_txrx(dev);
return ret;
}
@@ -4666,8 +4607,7 @@ static int nv_close(struct net_device *dev)
del_timer_sync(&np->stats_poll);
spin_lock_irq(&np->lock);
- nv_stop_tx(dev);
- nv_stop_rx(dev);
+ nv_stop_txrx(dev);
nv_txrx_reset(dev);
/* disable interrupts on the nic or we will lock up */
@@ -4680,7 +4620,7 @@ static int nv_close(struct net_device *dev)
nv_free_irq(dev);
- drain_ring(dev);
+ nv_drain_txrx(dev);
if (np->wolenabled) {
writel(NVREG_PFF_ALWAYS|NVREG_PFF_MYADDR, base + NvRegPacketFilterFlags);
@@ -4860,7 +4800,7 @@ static int __devinit nv_probe(struct pci_dev *pci_dev, const struct pci_device_i
dev->open = nv_open;
dev->stop = nv_close;
- if (np->desc_ver == DESC_VER_1 || np->desc_ver == DESC_VER_2)
+ if (!nv_optimized(np))
dev->hard_start_xmit = nv_start_xmit;
else
dev->hard_start_xmit = nv_start_xmit_optimized;
^ permalink raw reply related
* [PATCH 3/5] forcedeth: process TX completions using NAPI
From: Jeff Garzik @ 2007-10-06 15:14 UTC (permalink / raw)
To: netdev, Ayaz Abdulla; +Cc: LKML, Andrew Morton
In-Reply-To: <20071006151250.GA17020@havoc.gtf.org>
commit 57cbfacc00d69be2ba02b65d1021442273b76263
Author: Jeff Garzik <jeff@garzik.org>
Date: Fri Oct 5 23:25:56 2007 -0400
[netdrvr] forcedeth: process TX completions using NAPI
Signed-off-by: Jeff Garzik <jgarzik@redhat.com>
drivers/net/forcedeth.c | 143 +++++++++++++++++++++++++++---------------------
1 file changed, 83 insertions(+), 60 deletions(-)
57cbfacc00d69be2ba02b65d1021442273b76263
diff --git a/drivers/net/forcedeth.c b/drivers/net/forcedeth.c
index 1d1a5f5..1c236e6 100644
--- a/drivers/net/forcedeth.c
+++ b/drivers/net/forcedeth.c
@@ -744,6 +744,7 @@ struct fe_priv {
struct net_device *dev;
struct napi_struct napi;
+ struct napi_struct tx_napi;
/* General data:
* Locking: spin_lock(&np->lock); */
@@ -810,7 +811,6 @@ struct fe_priv {
union ring_type tx_ring;
u32 tx_flags;
int tx_ring_size;
- int tx_stop;
/* vlan fields */
struct vlan_group *vlangrp;
@@ -1763,10 +1763,7 @@ static int nv_start_xmit(struct sk_buff *skb, struct net_device *dev)
empty_slots = nv_get_empty_tx_slots(np);
if (unlikely(empty_slots <= entries)) {
- spin_lock_irq(&np->lock);
netif_stop_queue(dev);
- np->tx_stop = 1;
- spin_unlock_irq(&np->lock);
return NETDEV_TX_BUSY;
}
@@ -1879,10 +1876,7 @@ static int nv_start_xmit_optimized(struct sk_buff *skb, struct net_device *dev)
empty_slots = nv_get_empty_tx_slots(np);
if (unlikely(empty_slots <= entries)) {
- spin_lock_irq(&np->lock);
netif_stop_queue(dev);
- np->tx_stop = 1;
- spin_unlock_irq(&np->lock);
return NETDEV_TX_BUSY;
}
@@ -1987,12 +1981,16 @@ static int nv_start_xmit_optimized(struct sk_buff *skb, struct net_device *dev)
*
* Caller must own np->lock.
*/
-static void nv_tx_done(struct net_device *dev)
+static int nv_tx_done(struct net_device *dev, int limit)
{
struct fe_priv *np = netdev_priv(dev);
u32 flags;
+ int orig_limit = limit;
struct ring_desc* orig_get_tx = np->get_tx.orig;
+ if (unlikely(limit < 1))
+ return 0;
+
while ((np->get_tx.orig != np->put_tx.orig) &&
!((flags = le32_to_cpu(np->get_tx.orig->flaglen)) & NV_TX_VALID)) {
@@ -2039,19 +2037,28 @@ static void nv_tx_done(struct net_device *dev)
np->get_tx.orig = np->first_tx.orig;
if (unlikely(np->get_tx_ctx++ == np->last_tx_ctx))
np->get_tx_ctx = np->first_tx_ctx;
+
+ limit--;
+ if (limit == 0)
+ break;
}
- if (unlikely((np->tx_stop == 1) && (np->get_tx.orig != orig_get_tx))) {
- np->tx_stop = 0;
+
+ if (np->get_tx.orig != orig_get_tx)
netif_wake_queue(dev);
- }
+
+ return orig_limit - limit;
}
-static void nv_tx_done_optimized(struct net_device *dev, int limit)
+static int nv_tx_done_optimized(struct net_device *dev, int limit)
{
struct fe_priv *np = netdev_priv(dev);
u32 flags;
+ int orig_limit = limit;
struct ring_desc_ex* orig_get_tx = np->get_tx.ex;
+ if (unlikely(limit < 1))
+ return 0;
+
while ((np->get_tx.ex != np->put_tx.ex) &&
!((flags = le32_to_cpu(np->get_tx.ex->flaglen)) & NV_TX_VALID) &&
(limit-- > 0)) {
@@ -2075,10 +2082,11 @@ static void nv_tx_done_optimized(struct net_device *dev, int limit)
if (unlikely(np->get_tx_ctx++ == np->last_tx_ctx))
np->get_tx_ctx = np->first_tx_ctx;
}
- if (unlikely((np->tx_stop == 1) && (np->get_tx.ex != orig_get_tx))) {
- np->tx_stop = 0;
+
+ if (np->get_tx.ex != orig_get_tx)
netif_wake_queue(dev);
- }
+
+ return orig_limit - limit;
}
/*
@@ -2149,9 +2157,9 @@ static void nv_tx_timeout(struct net_device *dev)
/* 1) stop tx engine */
nv_stop_tx(dev);
- /* 2) check that the packets were not sent already: */
+ /* 2) process all pending tx completions */
if (np->desc_ver == DESC_VER_1 || np->desc_ver == DESC_VER_2)
- nv_tx_done(dev);
+ nv_tx_done(dev, np->tx_ring_size);
else
nv_tx_done_optimized(dev, np->tx_ring_size);
@@ -2923,6 +2931,7 @@ static irqreturn_t __nv_nic_irq(struct net_device *dev, bool optimized)
u8 __iomem *base = get_hwbase(dev);
u32 events;
int handled = 0;
+ u32 upd_mask = 0;
dprintk(KERN_DEBUG "%s: nv_nic_irq%s\n", dev->name,
optimized ? "_optimized" : "");
@@ -2942,19 +2951,25 @@ static irqreturn_t __nv_nic_irq(struct net_device *dev, bool optimized)
if (!(events & np->irqmask))
goto out;
- if (optimized)
- nv_tx_done_optimized(dev, TX_WORK_PER_LOOP);
- else
- nv_tx_done(dev);
-
if (events & NVREG_IRQ_RX_ALL) {
netif_rx_schedule(dev, &np->napi);
/* Disable furthur receive irq's */
- np->irqmask &= ~NVREG_IRQ_RX_ALL;
+ upd_mask |= NVREG_IRQ_RX_ALL;
+ }
+
+ if (events & NVREG_IRQ_TX_ALL) {
+ netif_rx_schedule(dev, &np->tx_napi);
+
+ /* Disable furthur xmit irq's */
+ upd_mask |= NVREG_IRQ_TX_ALL;
+ }
+
+ if (upd_mask) {
+ np->irqmask &= ~upd_mask;
if (np->msi_flags & NV_MSI_X_ENABLED)
- writel(NVREG_IRQ_RX_ALL, base + NvRegIrqMask);
+ writel(upd_mask, base + NvRegIrqMask);
else
writel(np->irqmask, base + NvRegIrqMask);
}
@@ -3019,7 +3034,7 @@ static irqreturn_t nv_nic_irq_optimized(int foo, void *data)
static irqreturn_t nv_nic_irq_other(int foo, void *data)
{
- struct net_device *dev = (struct net_device *) data;
+ struct net_device *dev = data;
return __nv_nic_irq(dev, true);
}
@@ -3029,45 +3044,48 @@ static irqreturn_t nv_nic_irq_tx(int foo, void *data)
struct fe_priv *np = netdev_priv(dev);
u8 __iomem *base = get_hwbase(dev);
u32 events;
- int i;
- unsigned long flags;
- dprintk(KERN_DEBUG "%s: nv_nic_irq_tx\n", dev->name);
+ events = readl(base + NvRegMSIXIrqStatus) & NVREG_IRQ_TX_ALL;
+ writel(NVREG_IRQ_TX_ALL, base + NvRegMSIXIrqStatus);
- for (i=0; ; i++) {
- events = readl(base + NvRegMSIXIrqStatus) & NVREG_IRQ_TX_ALL;
- writel(NVREG_IRQ_TX_ALL, base + NvRegMSIXIrqStatus);
- dprintk(KERN_DEBUG "%s: tx irq: %08x\n", dev->name, events);
- if (!(events & np->irqmask))
- break;
+ if (events) {
+ netif_rx_schedule(dev, &np->tx_napi);
+ /* disable receive interrupts on the nic */
+ writel(NVREG_IRQ_TX_ALL, base + NvRegIrqMask);
+ pci_push(base);
+ }
+ return IRQ_HANDLED;
+}
- spin_lock_irqsave(&np->lock, flags);
- nv_tx_done_optimized(dev, TX_WORK_PER_LOOP);
- spin_unlock_irqrestore(&np->lock, flags);
+static int nv_napi_tx_poll(struct napi_struct *napi, int budget)
+{
+ struct fe_priv *np = container_of(napi, struct fe_priv, napi);
+ struct net_device *dev = np->dev;
+ u8 __iomem *base = get_hwbase(dev);
+ unsigned long flags;
+ int pkts;
- if (unlikely(events & (NVREG_IRQ_TX_ERR))) {
- dprintk(KERN_DEBUG "%s: received irq with events 0x%x. Probably TX fail.\n",
- dev->name, events);
- }
- if (unlikely(i > max_interrupt_work)) {
- spin_lock_irqsave(&np->lock, flags);
- /* disable interrupts on the nic */
- writel(NVREG_IRQ_TX_ALL, base + NvRegIrqMask);
- pci_push(base);
+ spin_lock_irqsave(&np->lock, flags);
- if (!np->in_shutdown) {
- np->nic_poll_irq |= NVREG_IRQ_TX_ALL;
- mod_timer(&np->nic_poll, jiffies + POLL_WAIT);
- }
- spin_unlock_irqrestore(&np->lock, flags);
- printk(KERN_DEBUG "%s: too many iterations (%d) in nv_nic_irq_tx.\n", dev->name, i);
- break;
- }
+ if (np->desc_ver == DESC_VER_1 || np->desc_ver == DESC_VER_2)
+ pkts = nv_tx_done(dev, budget);
+ else
+ pkts = nv_tx_done_optimized(dev, budget);
+ if (pkts < budget) {
+ /* re-enable receive interrupts */
+ __netif_rx_complete(dev, napi);
+
+ np->irqmask |= NVREG_IRQ_TX_ALL;
+ if (np->msi_flags & NV_MSI_X_ENABLED)
+ writel(NVREG_IRQ_TX_ALL, base + NvRegIrqMask);
+ else
+ writel(np->irqmask, base + NvRegIrqMask);
}
- dprintk(KERN_DEBUG "%s: nv_nic_irq_tx completed\n", dev->name);
- return IRQ_RETVAL(i);
+ spin_unlock_irqrestore(&np->lock, flags);
+
+ return pkts;
}
static int nv_napi_poll(struct napi_struct *napi, int budget)
@@ -4316,7 +4334,8 @@ static void nv_self_test(struct net_device *dev, struct ethtool_test *test, u64
if (test->flags & ETH_TEST_FL_OFFLINE) {
if (netif_running(dev)) {
- netif_stop_queue(dev);
+ netif_tx_disable(dev);
+ napi_disable(&np->tx_napi);
napi_disable(&np->napi);
netif_tx_lock_bh(dev);
spin_lock_irq(&np->lock);
@@ -4375,8 +4394,9 @@ static void nv_self_test(struct net_device *dev, struct ethtool_test *test, u64
/* restart rx engine */
nv_start_rx(dev);
nv_start_tx(dev);
- netif_start_queue(dev);
napi_enable(&np->napi);
+ napi_enable(&np->tx_napi);
+ netif_start_queue(dev);
nv_enable_hw_interrupts(dev, np->irqmask);
}
}
@@ -4603,8 +4623,9 @@ static int nv_open(struct net_device *dev)
ret = nv_update_linkspeed(dev);
nv_start_rx(dev);
nv_start_tx(dev);
- netif_start_queue(dev);
napi_enable(&np->napi);
+ napi_enable(&np->tx_napi);
+ netif_start_queue(dev);
if (ret) {
netif_carrier_on(dev);
@@ -4635,14 +4656,15 @@ static int nv_close(struct net_device *dev)
spin_lock_irq(&np->lock);
np->in_shutdown = 1;
spin_unlock_irq(&np->lock);
+ netif_stop_queue(dev);
napi_disable(&np->napi);
+ napi_disable(&np->tx_napi);
synchronize_irq(dev->irq);
del_timer_sync(&np->oom_kick);
del_timer_sync(&np->nic_poll);
del_timer_sync(&np->stats_poll);
- netif_stop_queue(dev);
spin_lock_irq(&np->lock);
nv_stop_tx(dev);
nv_stop_rx(dev);
@@ -4850,6 +4872,7 @@ static int __devinit nv_probe(struct pci_dev *pci_dev, const struct pci_device_i
dev->poll_controller = nv_poll_controller;
#endif
netif_napi_add(dev, &np->napi, nv_napi_poll, RX_WORK_PER_LOOP);
+ netif_napi_add(dev, &np->tx_napi, nv_napi_tx_poll, RX_WORK_PER_LOOP);
SET_ETHTOOL_OPS(dev, &ops);
dev->tx_timeout = nv_tx_timeout;
dev->watchdog_timeo = NV_WATCHDOG_TIMEO;
^ permalink raw reply related
* [PATCH 2/5] forcedeth: interrupt handling cleanup
From: Jeff Garzik @ 2007-10-06 15:14 UTC (permalink / raw)
To: netdev, Ayaz Abdulla; +Cc: LKML, Andrew Morton
In-Reply-To: <20071006151250.GA17020@havoc.gtf.org>
commit a606d2a111cdf948da5d69eb1de5526c5c2dafef
Author: Jeff Garzik <jeff@garzik.org>
Date: Fri Oct 5 22:56:05 2007 -0400
[netdrvr] forcedeth: interrupt handling cleanup
* nv_nic_irq_optimized() and nv_nic_irq_other() were complete duplicates
of nv_nic_irq(), with the exception of one function call. Consolidate
all three into a single interrupt handler, deleting a lot of redundant
code.
* greatly simplify irq handler locking.
Prior to this change, the irq handler(s) would acquire and release
np->lock for each action (RX, TX, other events).
For the common case -- RX or TX work -- the lock is always acquired,
making all successive acquire/release combinations largely redundant.
Acquire the lock at the beginning of the irq handler, and release it at
the end of the irq handler. This is simple, easy, and obvious.
* remove irq handler work loop.
All interesting events emanating from the irq handler either have
their own work loops, or they poke a timer into action.
Therefore, delete the pointless master interrupt handler work loop.
Signed-off-by: Jeff Garzik <jgarzik@redhat.com>
drivers/net/forcedeth.c | 325 +++++++++++-------------------------------------
1 file changed, 77 insertions(+), 248 deletions(-)
a606d2a111cdf948da5d69eb1de5526c5c2dafef
diff --git a/drivers/net/forcedeth.c b/drivers/net/forcedeth.c
index 49906cc..1d1a5f5 100644
--- a/drivers/net/forcedeth.c
+++ b/drivers/net/forcedeth.c
@@ -2917,208 +2917,110 @@ static void nv_link_irq(struct net_device *dev)
dprintk(KERN_DEBUG "%s: link change notification done.\n", dev->name);
}
-static irqreturn_t nv_nic_irq(int foo, void *data)
+static irqreturn_t __nv_nic_irq(struct net_device *dev, bool optimized)
{
- struct net_device *dev = (struct net_device *) data;
struct fe_priv *np = netdev_priv(dev);
u8 __iomem *base = get_hwbase(dev);
u32 events;
- int i;
+ int handled = 0;
- dprintk(KERN_DEBUG "%s: nv_nic_irq\n", dev->name);
+ dprintk(KERN_DEBUG "%s: nv_nic_irq%s\n", dev->name,
+ optimized ? "_optimized" : "");
- for (i=0; ; i++) {
- if (!(np->msi_flags & NV_MSI_X_ENABLED)) {
- events = readl(base + NvRegIrqStatus) & NVREG_IRQSTAT_MASK;
- writel(NVREG_IRQSTAT_MASK, base + NvRegIrqStatus);
- } else {
- events = readl(base + NvRegMSIXIrqStatus) & NVREG_IRQSTAT_MASK;
- writel(NVREG_IRQSTAT_MASK, base + NvRegMSIXIrqStatus);
- }
- dprintk(KERN_DEBUG "%s: irq: %08x\n", dev->name, events);
- if (!(events & np->irqmask))
- break;
+ spin_lock(&np->lock);
- spin_lock(&np->lock);
- nv_tx_done(dev);
- spin_unlock(&np->lock);
+ if (!(np->msi_flags & NV_MSI_X_ENABLED)) {
+ events = readl(base + NvRegIrqStatus) & NVREG_IRQSTAT_MASK;
+ writel(NVREG_IRQSTAT_MASK, base + NvRegIrqStatus);
+ } else {
+ events = readl(base + NvRegMSIXIrqStatus) & NVREG_IRQSTAT_MASK;
+ writel(NVREG_IRQSTAT_MASK, base + NvRegMSIXIrqStatus);
+ }
- if (events & NVREG_IRQ_RX_ALL) {
- netif_rx_schedule(dev, &np->napi);
+ dprintk(KERN_DEBUG "%s: irq: %08x\n", dev->name, events);
- /* Disable furthur receive irq's */
- spin_lock(&np->lock);
- np->irqmask &= ~NVREG_IRQ_RX_ALL;
+ if (!(events & np->irqmask))
+ goto out;
- if (np->msi_flags & NV_MSI_X_ENABLED)
- writel(NVREG_IRQ_RX_ALL, base + NvRegIrqMask);
- else
- writel(np->irqmask, base + NvRegIrqMask);
- spin_unlock(&np->lock);
- }
- if (unlikely(events & NVREG_IRQ_LINK)) {
- spin_lock(&np->lock);
- nv_link_irq(dev);
- spin_unlock(&np->lock);
- }
- if (unlikely(np->need_linktimer && time_after(jiffies, np->link_timeout))) {
- spin_lock(&np->lock);
- nv_linkchange(dev);
- spin_unlock(&np->lock);
- np->link_timeout = jiffies + LINK_TIMEOUT;
- }
- if (unlikely(events & (NVREG_IRQ_TX_ERR))) {
- dprintk(KERN_DEBUG "%s: received irq with events 0x%x. Probably TX fail.\n",
- dev->name, events);
- }
- if (unlikely(events & (NVREG_IRQ_UNKNOWN))) {
- printk(KERN_DEBUG "%s: received irq with unknown events 0x%x. Please report\n",
- dev->name, events);
- }
- if (unlikely(events & NVREG_IRQ_RECOVER_ERROR)) {
- spin_lock(&np->lock);
- /* disable interrupts on the nic */
- if (!(np->msi_flags & NV_MSI_X_ENABLED))
- writel(0, base + NvRegIrqMask);
- else
- writel(np->irqmask, base + NvRegIrqMask);
- pci_push(base);
+ if (optimized)
+ nv_tx_done_optimized(dev, TX_WORK_PER_LOOP);
+ else
+ nv_tx_done(dev);
- if (!np->in_shutdown) {
- np->nic_poll_irq = np->irqmask;
- np->recover_error = 1;
- mod_timer(&np->nic_poll, jiffies + POLL_WAIT);
- }
- spin_unlock(&np->lock);
- break;
- }
- if (unlikely(i > max_interrupt_work)) {
- spin_lock(&np->lock);
- /* disable interrupts on the nic */
- if (!(np->msi_flags & NV_MSI_X_ENABLED))
- writel(0, base + NvRegIrqMask);
- else
- writel(np->irqmask, base + NvRegIrqMask);
- pci_push(base);
+ if (events & NVREG_IRQ_RX_ALL) {
+ netif_rx_schedule(dev, &np->napi);
- if (!np->in_shutdown) {
- np->nic_poll_irq = np->irqmask;
- mod_timer(&np->nic_poll, jiffies + POLL_WAIT);
- }
- spin_unlock(&np->lock);
- printk(KERN_DEBUG "%s: too many iterations (%d) in nv_nic_irq.\n", dev->name, i);
- break;
- }
+ /* Disable furthur receive irq's */
+ np->irqmask &= ~NVREG_IRQ_RX_ALL;
+ if (np->msi_flags & NV_MSI_X_ENABLED)
+ writel(NVREG_IRQ_RX_ALL, base + NvRegIrqMask);
+ else
+ writel(np->irqmask, base + NvRegIrqMask);
}
- dprintk(KERN_DEBUG "%s: nv_nic_irq completed\n", dev->name);
- return IRQ_RETVAL(i);
-}
+ if (unlikely(events & NVREG_IRQ_LINK))
+ nv_link_irq(dev);
-/**
- * All _optimized functions are used to help increase performance
- * (reduce CPU and increase throughput). They use descripter version 3,
- * compiler directives, and reduce memory accesses.
- */
-static irqreturn_t nv_nic_irq_optimized(int foo, void *data)
-{
- struct net_device *dev = (struct net_device *) data;
- struct fe_priv *np = netdev_priv(dev);
- u8 __iomem *base = get_hwbase(dev);
- u32 events;
- int i;
+ if (unlikely(np->need_linktimer && time_after(jiffies, np->link_timeout))) {
+ nv_linkchange(dev);
+ np->link_timeout = jiffies + LINK_TIMEOUT;
+ }
- dprintk(KERN_DEBUG "%s: nv_nic_irq_optimized\n", dev->name);
+ if (unlikely(events & (NVREG_IRQ_TX_ERR))) {
+ dprintk(KERN_DEBUG "%s: received irq with events 0x%x. "
+ "Probably TX fail.\n",
+ dev->name, events);
+ }
- for (i=0; ; i++) {
- if (!(np->msi_flags & NV_MSI_X_ENABLED)) {
- events = readl(base + NvRegIrqStatus) & NVREG_IRQSTAT_MASK;
- writel(NVREG_IRQSTAT_MASK, base + NvRegIrqStatus);
- } else {
- events = readl(base + NvRegMSIXIrqStatus) & NVREG_IRQSTAT_MASK;
- writel(NVREG_IRQSTAT_MASK, base + NvRegMSIXIrqStatus);
- }
- dprintk(KERN_DEBUG "%s: irq: %08x\n", dev->name, events);
- if (!(events & np->irqmask))
- break;
+ if (unlikely(events & (NVREG_IRQ_UNKNOWN))) {
+ printk(KERN_DEBUG "%s: received irq with unknown events 0x%x. "
+ "Please report\n",
+ dev->name, events);
+ }
- spin_lock(&np->lock);
- nv_tx_done_optimized(dev, TX_WORK_PER_LOOP);
- spin_unlock(&np->lock);
+ if (unlikely(events & NVREG_IRQ_RECOVER_ERROR)) {
+ /* disable interrupts on the nic */
+ if (!(np->msi_flags & NV_MSI_X_ENABLED))
+ writel(0, base + NvRegIrqMask);
+ else
+ writel(np->irqmask, base + NvRegIrqMask);
+ pci_push(base);
- if (events & NVREG_IRQ_RX_ALL) {
- netif_rx_schedule(dev, &np->napi);
+ if (!np->in_shutdown) {
+ np->nic_poll_irq = np->irqmask;
+ np->recover_error = 1;
+ mod_timer(&np->nic_poll, jiffies + POLL_WAIT);
+ }
+ }
- /* Disable furthur receive irq's */
- spin_lock(&np->lock);
- np->irqmask &= ~NVREG_IRQ_RX_ALL;
+ handled = 1;
- if (np->msi_flags & NV_MSI_X_ENABLED)
- writel(NVREG_IRQ_RX_ALL, base + NvRegIrqMask);
- else
- writel(np->irqmask, base + NvRegIrqMask);
- spin_unlock(&np->lock);
- }
- if (unlikely(events & NVREG_IRQ_LINK)) {
- spin_lock(&np->lock);
- nv_link_irq(dev);
- spin_unlock(&np->lock);
- }
- if (unlikely(np->need_linktimer && time_after(jiffies, np->link_timeout))) {
- spin_lock(&np->lock);
- nv_linkchange(dev);
- spin_unlock(&np->lock);
- np->link_timeout = jiffies + LINK_TIMEOUT;
- }
- if (unlikely(events & (NVREG_IRQ_TX_ERR))) {
- dprintk(KERN_DEBUG "%s: received irq with events 0x%x. Probably TX fail.\n",
- dev->name, events);
- }
- if (unlikely(events & (NVREG_IRQ_UNKNOWN))) {
- printk(KERN_DEBUG "%s: received irq with unknown events 0x%x. Please report\n",
- dev->name, events);
- }
- if (unlikely(events & NVREG_IRQ_RECOVER_ERROR)) {
- spin_lock(&np->lock);
- /* disable interrupts on the nic */
- if (!(np->msi_flags & NV_MSI_X_ENABLED))
- writel(0, base + NvRegIrqMask);
- else
- writel(np->irqmask, base + NvRegIrqMask);
- pci_push(base);
+out:
+ spin_unlock(&np->lock);
- if (!np->in_shutdown) {
- np->nic_poll_irq = np->irqmask;
- np->recover_error = 1;
- mod_timer(&np->nic_poll, jiffies + POLL_WAIT);
- }
- spin_unlock(&np->lock);
- break;
- }
+ dprintk(KERN_DEBUG "%s: nv_nic_irq%s completed\n", dev->name,
+ optimized ? "_optimized" : "");
- if (unlikely(i > max_interrupt_work)) {
- spin_lock(&np->lock);
- /* disable interrupts on the nic */
- if (!(np->msi_flags & NV_MSI_X_ENABLED))
- writel(0, base + NvRegIrqMask);
- else
- writel(np->irqmask, base + NvRegIrqMask);
- pci_push(base);
+ return IRQ_RETVAL(handled);
+}
- if (!np->in_shutdown) {
- np->nic_poll_irq = np->irqmask;
- mod_timer(&np->nic_poll, jiffies + POLL_WAIT);
- }
- spin_unlock(&np->lock);
- printk(KERN_DEBUG "%s: too many iterations (%d) in nv_nic_irq.\n", dev->name, i);
- break;
- }
+static irqreturn_t nv_nic_irq(int foo, void *data)
+{
+ struct net_device *dev = data;
+ return __nv_nic_irq(dev, false);
+}
- }
- dprintk(KERN_DEBUG "%s: nv_nic_irq_optimized completed\n", dev->name);
+static irqreturn_t nv_nic_irq_optimized(int foo, void *data)
+{
+ struct net_device *dev = data;
+ return __nv_nic_irq(dev, true);
+}
- return IRQ_RETVAL(i);
+static irqreturn_t nv_nic_irq_other(int foo, void *data)
+{
+ struct net_device *dev = (struct net_device *) data;
+ return __nv_nic_irq(dev, true);
}
static irqreturn_t nv_nic_irq_tx(int foo, void *data)
@@ -3227,79 +3129,6 @@ static irqreturn_t nv_nic_irq_rx(int foo, void *data)
return IRQ_HANDLED;
}
-static irqreturn_t nv_nic_irq_other(int foo, void *data)
-{
- struct net_device *dev = (struct net_device *) data;
- struct fe_priv *np = netdev_priv(dev);
- u8 __iomem *base = get_hwbase(dev);
- u32 events;
- int i;
- unsigned long flags;
-
- dprintk(KERN_DEBUG "%s: nv_nic_irq_other\n", dev->name);
-
- for (i=0; ; i++) {
- events = readl(base + NvRegMSIXIrqStatus) & NVREG_IRQ_OTHER;
- writel(NVREG_IRQ_OTHER, base + NvRegMSIXIrqStatus);
- dprintk(KERN_DEBUG "%s: irq: %08x\n", dev->name, events);
- if (!(events & np->irqmask))
- break;
-
- /* check tx in case we reached max loop limit in tx isr */
- spin_lock_irqsave(&np->lock, flags);
- nv_tx_done_optimized(dev, TX_WORK_PER_LOOP);
- spin_unlock_irqrestore(&np->lock, flags);
-
- if (events & NVREG_IRQ_LINK) {
- spin_lock_irqsave(&np->lock, flags);
- nv_link_irq(dev);
- spin_unlock_irqrestore(&np->lock, flags);
- }
- if (np->need_linktimer && time_after(jiffies, np->link_timeout)) {
- spin_lock_irqsave(&np->lock, flags);
- nv_linkchange(dev);
- spin_unlock_irqrestore(&np->lock, flags);
- np->link_timeout = jiffies + LINK_TIMEOUT;
- }
- if (events & NVREG_IRQ_RECOVER_ERROR) {
- spin_lock_irq(&np->lock);
- /* disable interrupts on the nic */
- writel(NVREG_IRQ_OTHER, base + NvRegIrqMask);
- pci_push(base);
-
- if (!np->in_shutdown) {
- np->nic_poll_irq |= NVREG_IRQ_OTHER;
- np->recover_error = 1;
- mod_timer(&np->nic_poll, jiffies + POLL_WAIT);
- }
- spin_unlock_irq(&np->lock);
- break;
- }
- if (events & (NVREG_IRQ_UNKNOWN)) {
- printk(KERN_DEBUG "%s: received irq with unknown events 0x%x. Please report\n",
- dev->name, events);
- }
- if (unlikely(i > max_interrupt_work)) {
- spin_lock_irqsave(&np->lock, flags);
- /* disable interrupts on the nic */
- writel(NVREG_IRQ_OTHER, base + NvRegIrqMask);
- pci_push(base);
-
- if (!np->in_shutdown) {
- np->nic_poll_irq |= NVREG_IRQ_OTHER;
- mod_timer(&np->nic_poll, jiffies + POLL_WAIT);
- }
- spin_unlock_irqrestore(&np->lock, flags);
- printk(KERN_DEBUG "%s: too many iterations (%d) in nv_nic_irq_other.\n", dev->name, i);
- break;
- }
-
- }
- dprintk(KERN_DEBUG "%s: nv_nic_irq_other completed\n", dev->name);
-
- return IRQ_RETVAL(i);
-}
-
static irqreturn_t nv_nic_irq_test(int foo, void *data)
{
struct net_device *dev = (struct net_device *) data;
^ permalink raw reply related
* [PATCH 1/5] forcedeth: make NAPI unconditional
From: Jeff Garzik @ 2007-10-06 15:13 UTC (permalink / raw)
To: netdev, Ayaz Abdulla; +Cc: LKML, Andrew Morton
In-Reply-To: <20071006151250.GA17020@havoc.gtf.org>
commit 7bfc023b952e8e12c7333efccd2e78023c546a7c
Author: Jeff Garzik <jeff@garzik.org>
Date: Fri Oct 5 20:50:24 2007 -0400
[netdrvr] forcedeth: make NAPI unconditional
Signed-off-by: Jeff Garzik <jgarzik@redhat.com>
drivers/net/Kconfig | 17 -----
drivers/net/forcedeth.c | 149 +-----------------------------------------------
2 files changed, 4 insertions(+), 162 deletions(-)
7bfc023b952e8e12c7333efccd2e78023c546a7c
diff --git a/drivers/net/Kconfig b/drivers/net/Kconfig
index 9c635a2..59eab61 100644
--- a/drivers/net/Kconfig
+++ b/drivers/net/Kconfig
@@ -1430,23 +1430,6 @@ config FORCEDETH
<file:Documentation/networking/net-modules.txt>. The module will be
called forcedeth.
-config FORCEDETH_NAPI
- bool "Use Rx and Tx Polling (NAPI) (EXPERIMENTAL)"
- depends on FORCEDETH && EXPERIMENTAL
- help
- NAPI is a new driver API designed to reduce CPU and interrupt load
- when the driver is receiving lots of packets from the card. It is
- still somewhat experimental and thus not yet enabled by default.
-
- If your estimated Rx load is 10kpps or more, or if the card will be
- deployed on potentially unfriendly networks (e.g. in a firewall),
- then say Y here.
-
- See <file:Documentation/networking/NAPI_HOWTO.txt> for more
- information.
-
- If in doubt, say N.
-
config CS89x0
tristate "CS89x0 support"
depends on NET_PCI && (ISA || MACH_IXDP2351 || ARCH_IXDP2X01 || ARCH_PNX010X)
diff --git a/drivers/net/forcedeth.c b/drivers/net/forcedeth.c
index dae30b7..49906cc 100644
--- a/drivers/net/forcedeth.c
+++ b/drivers/net/forcedeth.c
@@ -123,12 +123,7 @@
* DEV_NEED_TIMERIRQ will not harm you on sane hardware, only generating a few
* superfluous timer interrupts from the nic.
*/
-#ifdef CONFIG_FORCEDETH_NAPI
-#define DRIVERNAPI "-NAPI"
-#else
-#define DRIVERNAPI
-#endif
-#define FORCEDETH_VERSION "0.60"
+#define FORCEDETH_VERSION "1.00"
#define DRV_NAME "forcedeth"
#include <linux/module.h>
@@ -1587,7 +1582,6 @@ static int nv_alloc_rx_optimized(struct net_device *dev)
}
/* If rx bufs are exhausted called after 50ms to attempt to refresh */
-#ifdef CONFIG_FORCEDETH_NAPI
static void nv_do_rx_refill(unsigned long data)
{
struct net_device *dev = (struct net_device *) data;
@@ -1596,41 +1590,6 @@ static void nv_do_rx_refill(unsigned long data)
/* Just reschedule NAPI rx processing */
netif_rx_schedule(dev, &np->napi);
}
-#else
-static void nv_do_rx_refill(unsigned long data)
-{
- struct net_device *dev = (struct net_device *) data;
- struct fe_priv *np = netdev_priv(dev);
- int retcode;
-
- if (!using_multi_irqs(dev)) {
- if (np->msi_flags & NV_MSI_X_ENABLED)
- disable_irq(np->msi_x_entry[NV_MSI_X_VECTOR_ALL].vector);
- else
- disable_irq(dev->irq);
- } else {
- disable_irq(np->msi_x_entry[NV_MSI_X_VECTOR_RX].vector);
- }
- if (np->desc_ver == DESC_VER_1 || np->desc_ver == DESC_VER_2)
- retcode = nv_alloc_rx(dev);
- else
- retcode = nv_alloc_rx_optimized(dev);
- if (retcode) {
- spin_lock_irq(&np->lock);
- if (!np->in_shutdown)
- mod_timer(&np->oom_kick, jiffies + OOM_REFILL);
- spin_unlock_irq(&np->lock);
- }
- if (!using_multi_irqs(dev)) {
- if (np->msi_flags & NV_MSI_X_ENABLED)
- enable_irq(np->msi_x_entry[NV_MSI_X_VECTOR_ALL].vector);
- else
- enable_irq(dev->irq);
- } else {
- enable_irq(np->msi_x_entry[NV_MSI_X_VECTOR_RX].vector);
- }
-}
-#endif
static void nv_init_rx(struct net_device *dev)
{
@@ -2383,11 +2342,8 @@ static int nv_rx_process(struct net_device *dev, int limit)
skb->protocol = eth_type_trans(skb, dev);
dprintk(KERN_DEBUG "%s: nv_rx_process: %d bytes, proto %d accepted.\n",
dev->name, len, skb->protocol);
-#ifdef CONFIG_FORCEDETH_NAPI
netif_receive_skb(skb);
-#else
- netif_rx(skb);
-#endif
+
dev->last_rx = jiffies;
np->stats.rx_packets++;
np->stats.rx_bytes += len;
@@ -2480,28 +2436,14 @@ static int nv_rx_process_optimized(struct net_device *dev, int limit)
dev->name, len, skb->protocol);
if (likely(!np->vlangrp)) {
-#ifdef CONFIG_FORCEDETH_NAPI
netif_receive_skb(skb);
-#else
- netif_rx(skb);
-#endif
} else {
vlanflags = le32_to_cpu(np->get_rx.ex->buflow);
- if (vlanflags & NV_RX3_VLAN_TAG_PRESENT) {
-#ifdef CONFIG_FORCEDETH_NAPI
+ if (vlanflags & NV_RX3_VLAN_TAG_PRESENT)
vlan_hwaccel_receive_skb(skb, np->vlangrp,
vlanflags & NV_RX3_VLAN_TAG_MASK);
-#else
- vlan_hwaccel_rx(skb, np->vlangrp,
- vlanflags & NV_RX3_VLAN_TAG_MASK);
-#endif
- } else {
-#ifdef CONFIG_FORCEDETH_NAPI
+ else
netif_receive_skb(skb);
-#else
- netif_rx(skb);
-#endif
- }
}
dev->last_rx = jiffies;
@@ -3001,7 +2943,6 @@ static irqreturn_t nv_nic_irq(int foo, void *data)
nv_tx_done(dev);
spin_unlock(&np->lock);
-#ifdef CONFIG_FORCEDETH_NAPI
if (events & NVREG_IRQ_RX_ALL) {
netif_rx_schedule(dev, &np->napi);
@@ -3015,16 +2956,6 @@ static irqreturn_t nv_nic_irq(int foo, void *data)
writel(np->irqmask, base + NvRegIrqMask);
spin_unlock(&np->lock);
}
-#else
- if (nv_rx_process(dev, RX_WORK_PER_LOOP)) {
- if (unlikely(nv_alloc_rx(dev))) {
- spin_lock(&np->lock);
- if (!np->in_shutdown)
- mod_timer(&np->oom_kick, jiffies + OOM_REFILL);
- spin_unlock(&np->lock);
- }
- }
-#endif
if (unlikely(events & NVREG_IRQ_LINK)) {
spin_lock(&np->lock);
nv_link_irq(dev);
@@ -3116,7 +3047,6 @@ static irqreturn_t nv_nic_irq_optimized(int foo, void *data)
nv_tx_done_optimized(dev, TX_WORK_PER_LOOP);
spin_unlock(&np->lock);
-#ifdef CONFIG_FORCEDETH_NAPI
if (events & NVREG_IRQ_RX_ALL) {
netif_rx_schedule(dev, &np->napi);
@@ -3130,16 +3060,6 @@ static irqreturn_t nv_nic_irq_optimized(int foo, void *data)
writel(np->irqmask, base + NvRegIrqMask);
spin_unlock(&np->lock);
}
-#else
- if (nv_rx_process_optimized(dev, RX_WORK_PER_LOOP)) {
- if (unlikely(nv_alloc_rx_optimized(dev))) {
- spin_lock(&np->lock);
- if (!np->in_shutdown)
- mod_timer(&np->oom_kick, jiffies + OOM_REFILL);
- spin_unlock(&np->lock);
- }
- }
-#endif
if (unlikely(events & NVREG_IRQ_LINK)) {
spin_lock(&np->lock);
nv_link_irq(dev);
@@ -3248,7 +3168,6 @@ static irqreturn_t nv_nic_irq_tx(int foo, void *data)
return IRQ_RETVAL(i);
}
-#ifdef CONFIG_FORCEDETH_NAPI
static int nv_napi_poll(struct napi_struct *napi, int budget)
{
struct fe_priv *np = container_of(napi, struct fe_priv, napi);
@@ -3288,9 +3207,7 @@ static int nv_napi_poll(struct napi_struct *napi, int budget)
}
return pkts;
}
-#endif
-#ifdef CONFIG_FORCEDETH_NAPI
static irqreturn_t nv_nic_irq_rx(int foo, void *data)
{
struct net_device *dev = (struct net_device *) data;
@@ -3309,54 +3226,6 @@ static irqreturn_t nv_nic_irq_rx(int foo, void *data)
}
return IRQ_HANDLED;
}
-#else
-static irqreturn_t nv_nic_irq_rx(int foo, void *data)
-{
- struct net_device *dev = (struct net_device *) data;
- struct fe_priv *np = netdev_priv(dev);
- u8 __iomem *base = get_hwbase(dev);
- u32 events;
- int i;
- unsigned long flags;
-
- dprintk(KERN_DEBUG "%s: nv_nic_irq_rx\n", dev->name);
-
- for (i=0; ; i++) {
- events = readl(base + NvRegMSIXIrqStatus) & NVREG_IRQ_RX_ALL;
- writel(NVREG_IRQ_RX_ALL, base + NvRegMSIXIrqStatus);
- dprintk(KERN_DEBUG "%s: rx irq: %08x\n", dev->name, events);
- if (!(events & np->irqmask))
- break;
-
- if (nv_rx_process_optimized(dev, RX_WORK_PER_LOOP)) {
- if (unlikely(nv_alloc_rx_optimized(dev))) {
- spin_lock_irqsave(&np->lock, flags);
- if (!np->in_shutdown)
- mod_timer(&np->oom_kick, jiffies + OOM_REFILL);
- spin_unlock_irqrestore(&np->lock, flags);
- }
- }
-
- if (unlikely(i > max_interrupt_work)) {
- spin_lock_irqsave(&np->lock, flags);
- /* disable interrupts on the nic */
- writel(NVREG_IRQ_RX_ALL, base + NvRegIrqMask);
- pci_push(base);
-
- if (!np->in_shutdown) {
- np->nic_poll_irq |= NVREG_IRQ_RX_ALL;
- mod_timer(&np->nic_poll, jiffies + POLL_WAIT);
- }
- spin_unlock_irqrestore(&np->lock, flags);
- printk(KERN_DEBUG "%s: too many iterations (%d) in nv_nic_irq_rx.\n", dev->name, i);
- break;
- }
- }
- dprintk(KERN_DEBUG "%s: nv_nic_irq_rx completed\n", dev->name);
-
- return IRQ_RETVAL(i);
-}
-#endif
static irqreturn_t nv_nic_irq_other(int foo, void *data)
{
@@ -4619,9 +4488,7 @@ static void nv_self_test(struct net_device *dev, struct ethtool_test *test, u64
if (test->flags & ETH_TEST_FL_OFFLINE) {
if (netif_running(dev)) {
netif_stop_queue(dev);
-#ifdef CONFIG_FORCEDETH_NAPI
napi_disable(&np->napi);
-#endif
netif_tx_lock_bh(dev);
spin_lock_irq(&np->lock);
nv_disable_hw_interrupts(dev, np->irqmask);
@@ -4680,9 +4547,7 @@ static void nv_self_test(struct net_device *dev, struct ethtool_test *test, u64
nv_start_rx(dev);
nv_start_tx(dev);
netif_start_queue(dev);
-#ifdef CONFIG_FORCEDETH_NAPI
napi_enable(&np->napi);
-#endif
nv_enable_hw_interrupts(dev, np->irqmask);
}
}
@@ -4910,9 +4775,7 @@ static int nv_open(struct net_device *dev)
nv_start_rx(dev);
nv_start_tx(dev);
netif_start_queue(dev);
-#ifdef CONFIG_FORCEDETH_NAPI
napi_enable(&np->napi);
-#endif
if (ret) {
netif_carrier_on(dev);
@@ -4943,9 +4806,7 @@ static int nv_close(struct net_device *dev)
spin_lock_irq(&np->lock);
np->in_shutdown = 1;
spin_unlock_irq(&np->lock);
-#ifdef CONFIG_FORCEDETH_NAPI
napi_disable(&np->napi);
-#endif
synchronize_irq(dev->irq);
del_timer_sync(&np->oom_kick);
@@ -5159,9 +5020,7 @@ static int __devinit nv_probe(struct pci_dev *pci_dev, const struct pci_device_i
#ifdef CONFIG_NET_POLL_CONTROLLER
dev->poll_controller = nv_poll_controller;
#endif
-#ifdef CONFIG_FORCEDETH_NAPI
netif_napi_add(dev, &np->napi, nv_napi_poll, RX_WORK_PER_LOOP);
-#endif
SET_ETHTOOL_OPS(dev, &ops);
dev->tx_timeout = nv_tx_timeout;
dev->watchdog_timeo = NV_WATCHDOG_TIMEO;
^ permalink raw reply related
page: next (older) | prev (newer) | latest
- recent:[subjects (threaded)|topics (new)|topics (active)]
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox