The Linux Kernel Mailing List
 help / color / mirror / Atom feed
* P4 Xeon summary inquiry
@ 2002-05-06 23:27 J.A. Magallon
  2002-05-07  7:12 ` Jesse Wyant
  0 siblings, 1 reply; 5+ messages in thread
From: J.A. Magallon @ 2002-05-06 23:27 UTC (permalink / raw)
  To: Lista Linux-Kernel

Hi, all

I am trying to make a dual P4-Xeon box (P4DCE, Intel 869) work optimally.
After some search in list archives, I have found this points:

- You need ACPI, and boot with acpismp=force, to have hyperthreading
  (see 4 cpus)
- To balance interrupts, a patch from Ingo is needed
- To balance timers, one other patch was needed, but this in included
  in 2.4.19-pre8 (do not know since which pre is there)

I tried to boot with acpismp=force, but then performance was dog slow,
I could count lines scrolling on a rxvt.
I checked mtrr and look like working. Box has 1Gb of ram, so kernel
is using CONFIG_HIGHMEM4G=y.
Kernel is the standard highmem Mandrake kernel in 8.2
(2.4.18-6mdkenterprise) that is mainly a plain 2.4.18 with some
.19-pre1 fixes.

Any correction ? Any known problem in 2.4.18 about this issue that
has bee corrected in pres for 19 ?

Any idea about the performance loss ?

TIA

-- 
J.A. Magallon                           #  Let the source be with you...        
mailto:jamagallon@able.es
Mandrake Linux release 8.3 (Cooker) for i586
Linux werewolf 2.4.19-pre8-jam1 #1 SMP dom may 5 23:46:04 CEST 2002 i686

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: P4 Xeon summary inquiry
  2002-05-06 23:27 P4 Xeon summary inquiry J.A. Magallon
@ 2002-05-07  7:12 ` Jesse Wyant
  2002-05-07 15:00   ` James Bourne
                     ` (2 more replies)
  0 siblings, 3 replies; 5+ messages in thread
From: Jesse Wyant @ 2002-05-07  7:12 UTC (permalink / raw)
  To: J.A. Magallon; +Cc: Lista Linux-Kernel


I just put together a SuperMicro P4DCE+ (Intel i860/ICH2/CS4299 AC'97 codec
[use the OSS i810_audio or ALSA's intel_8x0 or whatever it is]) with 1GB of
Samsung RDRAM and two 2.0GHz Xeon (Prestonias.)  Currently 2.4.19-pre7 with
Ingo's interrupt patch, and timer balancing patch.  Running on ext3, using
RedHat 7.2.

I also tried the acpismp=force option to get into hyperthreading
(successfully--it appears as though I have 4 CPUs), and the overall system
seems just as fast as with the HT disabled.  Qualitatively: kernel compile
with HT disabled, using 

    'make -j2 bzImage; make -j2 modules; make modules_install; sync'

results in compile times around 3 minutes 45 seconds or so.  With HT enabled,
and using -j4 instead of -j2, my compile time comes down to around 2:57 or
so--a significant improvement.

However, 'dnetc's throughput in RC5 keys/s is much lower with HT enabled: 
it runs 4 clients, and each client chugs through about 720kKeys/s.  
With HT disabled, the two dnetc clients run through 2.8MKeys/s each.  (So
it's around half as fast with HT enabled!)  When I'm finished downloading
RedHat 7.3, I'll reboot into Hyperthreading-enabled mode, and run 
'dnetc --benchmark' to confirm this.

Haven't had a chance to benchmark more than that yet.  But no gross issues
when running with HT enabled.  (Tribes2, Q3A and RTCW feel equally fast 
between the two configurations.)

-jesse


> I am trying to make a dual P4-Xeon box (P4DCE, Intel 869) work optimally.
> After some search in list archives, I have found this points:
> 
> - You need ACPI, and boot with acpismp=force, to have hyperthreading
>   (see 4 cpus)
> - To balance interrupts, a patch from Ingo is needed
> - To balance timers, one other patch was needed, but this in included
>   in 2.4.19-pre8 (do not know since which pre is there)
> 
> I tried to boot with acpismp=force, but then performance was dog slow,
> I could count lines scrolling on a rxvt.
> I checked mtrr and look like working. Box has 1Gb of ram, so kernel
> is using CONFIG_HIGHMEM4G=y.
> Kernel is the standard highmem Mandrake kernel in 8.2
> (2.4.18-6mdkenterprise) that is mainly a plain 2.4.18 with some
> .19-pre1 fixes.
> 
> Any correction ? Any known problem in 2.4.18 about this issue that
> has bee corrected in pres for 19 ?
> 
> Any idea about the performance loss ?
> 

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: P4 Xeon summary inquiry
  2002-05-07  7:12 ` Jesse Wyant
@ 2002-05-07 15:00   ` James Bourne
  2002-05-07 15:58   ` Nicolae P. Costescu
  2002-05-07 18:00   ` Bill Davidsen
  2 siblings, 0 replies; 5+ messages in thread
From: James Bourne @ 2002-05-07 15:00 UTC (permalink / raw)
  To: Jesse Wyant; +Cc: J.A. Magallon, Lista Linux-Kernel

On Tue, 7 May 2002, Jesse Wyant wrote:

[...]
>     'make -j2 bzImage; make -j2 modules; make modules_install; sync'
> 
> results in compile times around 3 minutes 45 seconds or so.  With HT enabled,
> and using -j4 instead of -j2, my compile time comes down to around 2:57 or
> so--a significant improvement.
> 
> However, 'dnetc's throughput in RC5 keys/s is much lower with HT enabled: 
> it runs 4 clients, and each client chugs through about 720kKeys/s.  
> With HT disabled, the two dnetc clients run through 2.8MKeys/s each.  (So
> it's around half as fast with HT enabled!)  When I'm finished downloading
> RedHat 7.3, I'll reboot into Hyperthreading-enabled mode, and run 
> 'dnetc --benchmark' to confirm this.

We have found the same thing.  If you are running single threaded, 2 or
less consecutive on a 2-proc system it's better to disable HT, however,
on a system which will have multiple runable processes looking for CPU
time concurrently, then using HT is a benefit.

We noticed ~2%-5% drop in single process performance with HT
enabled, but >20% with multiple processes with HT enabled.

Thing is, with 2 1.8 GHz CPUs, a 2% drop may be appreciable, but not
generally noticable overall.

The tests were done with a Dell PE4600 (2x1.8GHz, 2GB RAM), 2.4.18, with
Ingos' irqbalance-2.4.17-B1.patch and timer-irq-balance-2.4.18.patch.

The system is now in production with these patches and kernel, and has
been very stable (of course now it *will* crash).

Regards
James Bourne

> 
> Haven't had a chance to benchmark more than that yet.  But no gross issues
> when running with HT enabled.  (Tribes2, Q3A and RTCW feel equally fast 
> between the two configurations.)
> 
> -jesse
> 
> 
> > I am trying to make a dual P4-Xeon box (P4DCE, Intel 869) work optimally.
> > After some search in list archives, I have found this points:
> > 
> > - You need ACPI, and boot with acpismp=force, to have hyperthreading
> >   (see 4 cpus)
> > - To balance interrupts, a patch from Ingo is needed
> > - To balance timers, one other patch was needed, but this in included
> >   in 2.4.19-pre8 (do not know since which pre is there)
> > 
> > I tried to boot with acpismp=force, but then performance was dog slow,
> > I could count lines scrolling on a rxvt.
> > I checked mtrr and look like working. Box has 1Gb of ram, so kernel
> > is using CONFIG_HIGHMEM4G=y.
> > Kernel is the standard highmem Mandrake kernel in 8.2
> > (2.4.18-6mdkenterprise) that is mainly a plain 2.4.18 with some
> > .19-pre1 fixes.
> > 
> > Any correction ? Any known problem in 2.4.18 about this issue that
> > has bee corrected in pres for 19 ?
> > 
> > Any idea about the performance loss ?
> > 
> -
> To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at  http://vger.kernel.org/majordomo-info.html
> Please read the FAQ at  http://www.tux.org/lkml/
> 

-- 
James Bourne, Supervisor Data Centre Operations
Mount Royal College, Calgary, AB, CA
www.mtroyal.ab.ca

******************************************************************************
This communication is intended for the use of the recipient to which it is
addressed, and may contain confidential, personal, and or privileged
information. Please contact the sender immediately if you are not the
intended recipient of this communication, and do not copy, distribute, or
take action relying on it. Any communication received in error, or
subsequent reply, should be deleted or destroyed.
******************************************************************************


^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: P4 Xeon summary inquiry
  2002-05-07  7:12 ` Jesse Wyant
  2002-05-07 15:00   ` James Bourne
@ 2002-05-07 15:58   ` Nicolae P. Costescu
  2002-05-07 18:00   ` Bill Davidsen
  2 siblings, 0 replies; 5+ messages in thread
From: Nicolae P. Costescu @ 2002-05-07 15:58 UTC (permalink / raw)
  To: Jesse Wyant; +Cc: linux-kernel

Jesse,
The mixed performance bag of hyperthreading you are seeing is pretty normal 
from what I understand.

Here's a nice summary article on the topic from anandtech...

http://anandtech.com/showdoc.html?i=1576&p=1

Thanks
Nick


^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: P4 Xeon summary inquiry
  2002-05-07  7:12 ` Jesse Wyant
  2002-05-07 15:00   ` James Bourne
  2002-05-07 15:58   ` Nicolae P. Costescu
@ 2002-05-07 18:00   ` Bill Davidsen
  2 siblings, 0 replies; 5+ messages in thread
From: Bill Davidsen @ 2002-05-07 18:00 UTC (permalink / raw)
  To: Jesse Wyant; +Cc: J.A. Magallon, Lista Linux-Kernel

On Tue, 7 May 2002, Jesse Wyant wrote:

> However, 'dnetc's throughput in RC5 keys/s is much lower with HT enabled: 
> it runs 4 clients, and each client chugs through about 720kKeys/s.  
> With HT disabled, the two dnetc clients run through 2.8MKeys/s each.  (So
> it's around half as fast with HT enabled!)  When I'm finished downloading
> RedHat 7.3, I'll reboot into Hyperthreading-enabled mode, and run 
> 'dnetc --benchmark' to confirm this.

  I believe that what you are seeing is caused by the two threads in each
CPU contending for cache, at least at L1 level, perhaps also L2. Other
than a careful study of the code or a hardware probe, I don't know if you
could even roughly qualtify that, but I'm moderately sure you're beating
the cache to death.

  If you had a Xeon with larger L2 it might be interesting to see if HT
ran at the same speed as the single thread CPU with half the cache. And if
you want to play more, you could use the BIOS to disable the L1 or L2
cache and see how much the performance changes. Doesn't matter, unless you
have another algorithm it doesn't address the behaviour.

-- 
bill davidsen <davidsen@tmr.com>
  CTO, TMR Associates, Inc
Doing interesting things with little computers since 1979.


^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2002-05-07 18:04 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2002-05-06 23:27 P4 Xeon summary inquiry J.A. Magallon
2002-05-07  7:12 ` Jesse Wyant
2002-05-07 15:00   ` James Bourne
2002-05-07 15:58   ` Nicolae P. Costescu
2002-05-07 18:00   ` Bill Davidsen

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox