Netdev List
 help / color / mirror / Atom feed
From: jamal <hadi@cyberus.ca>
To: Baruch Even <baruch@ev-en.org>
Cc: Stephen Hemminger <shemminger@osdl.org>,
	John Heffner <jheffner@psc.edu>,
	"David S. Miller" <davem@davemloft.net>,
	rhee@eos.ncsu.edu, Yee-Ting.Li@nuim.ie, netdev@oss.sgi.com
Subject: Re: netif_rx packet dumping
Date: 08 Mar 2005 17:02:01 -0500	[thread overview]
Message-ID: <1110319320.1078.97.camel@jzny.localdomain> (raw)
In-Reply-To: <422DCB14.1040805@ev-en.org>

[-- Attachment #1: Type: text/plain, Size: 4927 bytes --]

On Tue, 2005-03-08 at 10:56, Baruch Even wrote:
> jamal wrote:
> > 
> > Were the processors tied to NICs? 
> 
> No. These are single CPU machines (with HT).
> 

You have SMP.

> > In other words the whole queue was infact dedicated just for your one
> > flow - thats why you can call this queue a transient burst queue. 
> 
> Indeed, For a router or a web server handling several thousand flows it 
> might be different, but I don't expect it handles a single packet in one 
> ms (or more) as it happens for the current end-system ack handling code.
> 

I think it would be interesting as well to see more than one flow;
maybe go upto 16 (1, 2, 4, 8, 16) and if theres anything to observe you
should see it.

> > Do you still have the data that shows how many packets were dropped
> > during this period. Do you still have the experimental data? I am
> > particulary interested in seeing the softnet stats as well as tcp
> > netstats.
> 
> No, These tests were not run by me, I'll probably rerun similar tests as 
> well to base my work on, send me in private how do I get the stats from 
> the kernel and I'll add it to my test scripts.
> 

I will be more than happy to help. Let me know when you are ready.

> > I think your main problem was the huge amounts of SACK on the writequeue
> > and the resultant processing i.e section 1.1 and how you resolved that.
> 
> That is my main guess as well, the original work was done rather 
> quickly, we are now reorganizing thoughts and redoing the tests in a 
> more orderly fashion.
> 

I have attached the little patch i forgot to last time where you can
adjust. Adjust the lo_cong parameter in /proc to tune the number of
packets in which the congestion valve gets opened. Make sure you 
are within range of the other parameters like no_cong etc.
Or you can change that check to be using no_cong instead.

> > I dont see any issue in dropping ACKs, many of them even for such large
> > windows as you have - TCPs ACKs are cummulative. It is true if you drop
> > "large" enough amounts of ACKS, you will end up in timeouts - but large
> > enough in your case must be in the minimal 1000 packets. And to say you
> > dropped a 1000 packets while processing 300 means you were taking too
> > long processing the 300.
> 
> With the current code SACK processing takes a long time, so it is 
> possible that it happened to drop more than a thousand packets while 
> handling 300. I think that after the fixing of the SACK code, the rest 
> might work without getting to much into the ingress queue. But that 
> might still change when we go to even higher speeds.
> 

Sure. I am not questioning fixing the SACK code - I dont think it was
envisioned that someone was going to have 1000 outstanding SACKs in 
_one_ flow.  But an interesting test case would be to fix the SACK
processing then retest without getting rid of the congestion valve.
Again, shall i repeat you really should be using NAPI? ;->

> > Then what would be really interesting is to see the perfomance you get
> > from multiple flows with and without congestion.
> 
> We'd need to get a very high speed link for multiple high speed flows.
> 

Get a couple of PCs and hook them back to back.

> > I am not against a the benchmarky nature of the single flow and tuning
> > for that, but we should also look at a wider scope at the effect before
> > you handwave based on the result of one testcase.
> 
> I can't say I didn't handwave, but then, there is little experimentation 
> done to see if the other claims are correct and that AFQ is really 
> needed so early in the packet receive stage. There are also voices that 
> say AFQ sucks and causes more damage than good, I don't remember details 
> currently.
> 

Its not totaly AFQ,
The idea is that code is measuring (based on history) how busy the
system is. What Stephen and I were discussing is that you may wanna
totaly punish new flows coming in instead of the one flow that already
is flowing when the going gets tough (i.e congestion detected). However,
this maybe too big unneeded a hack when we have NAPI which will do just
fine for you.

> > So if i was you i would repeat 1.2 with the fix from 1.1 as well as
> > tying the NIC to one CPU. And it would be a good idea to present more
> > detailed results - not just tcp windows fluctuating (you may not need
> > them for the paper, but would be useful to see for debugging purposes
> > other parameters).
> 
> I'd be happy to hear what other benchmarks you would like to see, I 
> currently intend to add some ack processing time analysis and oprofile 
> information. With possibly showing the size of the ingress queue as a 
> measure as well.
> 
> Making it as thorough as possible is one of my goals. Input is always 
> welcome.
> 

Good - and thanks for not being defensive; your work can only get better
this way. Ping me when you have ported to 2.6.11 and are ready to do the
testing.

cheers,
jamal

[-- Attachment #2: cong_p --]
[-- Type: text/plain, Size: 291 bytes --]

--- 2611-mod/net/core/dev.c	2005/03/07 12:26:21	1.1
+++ 2611-mod/net/core/dev.c	2005/03/07 12:30:19
@@ -1742,6 +1742,9 @@
 		if (work >= quota || jiffies - start_time > 1)
 			break;
 
+		if (queue->throttle && work >= lo_cong)
+			queue->throttle = 0;
+
 	}
 
 	backlog_dev->quota -= work;

  reply	other threads:[~2005-03-08 22:02 UTC|newest]

Thread overview: 51+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2005-03-03 20:38 netif_rx packet dumping Stephen Hemminger
2005-03-03 20:55 ` David S. Miller
2005-03-03 21:01   ` Stephen Hemminger
2005-03-03 21:18   ` jamal
2005-03-03 21:21     ` Stephen Hemminger
2005-03-03 21:24       ` jamal
2005-03-03 21:32         ` David S. Miller
2005-03-03 21:54           ` Stephen Hemminger
2005-03-03 22:02             ` John Heffner
2005-03-03 22:26               ` jamal
2005-03-03 23:16                 ` Stephen Hemminger
2005-03-03 23:40                   ` jamal
2005-03-03 23:48                   ` Baruch Even
2005-03-04  3:45                     ` jamal
2005-03-04  8:47                       ` Baruch Even
2005-03-07 13:55                         ` jamal
2005-03-08 15:56                           ` Baruch Even
2005-03-08 22:02                             ` jamal [this message]
2005-03-22 21:55                             ` cliff white
2005-03-03 23:48                   ` John Heffner
2005-03-04  1:42                     ` Lennert Buytenhek
2005-03-04  3:10                       ` John Heffner
2005-03-04  3:31                         ` Lennert Buytenhek
2005-03-04 19:52                 ` Edgar E Iglesias
2005-03-04 19:54                   ` Stephen Hemminger
2005-03-04 21:41                     ` Edgar E Iglesias
2005-03-04 19:49             ` Jason Lunz
2005-03-03 22:01           ` jamal
2005-03-03 21:26 ` Baruch Even
2005-03-03 21:36   ` David S. Miller
2005-03-03 21:44     ` Baruch Even
2005-03-03 21:54       ` Andi Kleen
2005-03-03 22:04         ` David S. Miller
2005-03-03 21:57       ` David S. Miller
2005-03-03 22:14         ` Baruch Even
2005-03-08 15:42         ` Baruch Even
2005-03-08 17:00           ` Andi Kleen
2005-03-08 18:01             ` Baruch Even
2005-03-08 18:09             ` David S. Miller
2005-03-08 18:18               ` Andi Kleen
2005-03-08 18:37                 ` Thomas Graf
2005-03-08 18:51                   ` Arnaldo Carvalho de Melo
2005-03-08 22:16                   ` Andi Kleen
2005-03-08 18:27               ` Ben Greear
2005-03-09 23:57                 ` Thomas Graf
2005-03-10  0:03                   ` Stephen Hemminger
2005-03-10  8:33                   ` Andi Kleen
2005-03-10 14:08                     ` Thomas Graf
2005-03-31 16:33         ` Baruch Even
2005-03-03 22:03   ` jamal
2005-03-03 22:31     ` Baruch Even

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=1110319320.1078.97.camel@jzny.localdomain \
    --to=hadi@cyberus.ca \
    --cc=Yee-Ting.Li@nuim.ie \
    --cc=baruch@ev-en.org \
    --cc=davem@davemloft.net \
    --cc=jheffner@psc.edu \
    --cc=netdev@oss.sgi.com \
    --cc=rhee@eos.ncsu.edu \
    --cc=shemminger@osdl.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox