All of lore.kernel.org
 help / color / mirror / Atom feed
* Ultra AXmp
@ 1998-07-11 18:36 Bob Drzyzgula
  1998-07-12  1:56 ` David S. Miller
                   ` (40 more replies)
  0 siblings, 41 replies; 42+ messages in thread
From: Bob Drzyzgula @ 1998-07-11 18:36 UTC (permalink / raw)
  To: ultralinux


Sun's newly announced Ultra AXmp board is a pretty interesting
platform. Based on early pricing estimates I've seen, it should
be possible to build a quad-processor @ 300+MHz, 2GB rackmount
system for under around $US 20K. I have a lot of Solaris applications
for this board, but I also wonder what capabilities one would
expect ULtraLinux to have on this machine...

http://www.sun.com/microelectronics/SPARCengineUltraAXmp/ 
http://www.sun.com/microelectronics/SPARCengineUltraAXmp/805-5865.pdf
http://www.sun.com/microelectronics/SPARCengineUltraAXmp/s7-mod.pdf
http://www.sun.com/microelectronics/SPARCengineUltraAXmp/cover980706.html
http://www.sun.com/smi/Press/sunflash/9807/sunflash.980706.1.html

--Bob

-- 
==============================
Bob Drzyzgula                             It's not a problem
bob@drzyzgula.org                until something bad happens
==============================

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
@ 1998-07-12  1:56 ` David S. Miller
  1998-07-12  2:21 ` Matthew Jacob
                   ` (39 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: David S. Miller @ 1998-07-12  1:56 UTC (permalink / raw)
  To: ultralinux

   Date: 	Sat, 11 Jul 1998 14:36:06 -0400
   From: Bob Drzyzgula <bob@drzyzgula.org>

   Based on early pricing estimates I've seen, it should be possible
   to build a quad-processor @ 300+MHz, 2GB rackmount system for under
   around $US 20K.

I wonder what the equivalent quad-PPRO system would cost (even with
400MHZ chips...)...  I love the Ultra as everyone knows, but I have to
say, price talks bullshit walks... and Sun has a lot of trouble
learning this especially on the high end.

Later,
David S. Miller
davem@dm.cobaltmicro.com

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
  1998-07-12  1:56 ` David S. Miller
@ 1998-07-12  2:21 ` Matthew Jacob
  1998-07-12  2:29 ` David S. Miller
                   ` (38 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Matthew Jacob @ 1998-07-12  2:21 UTC (permalink / raw)
  To: ultralinux



On Sat, 11 Jul 1998, David S. Miller wrote:

>    Date: 	Sat, 11 Jul 1998 14:36:06 -0400
>    From: Bob Drzyzgula <bob@drzyzgula.org>
> 
>    Based on early pricing estimates I've seen, it should be possible
>    to build a quad-processor @ 300+MHz, 2GB rackmount system for under
>    around $US 20K.
> 
> I wonder what the equivalent quad-PPRO system would cost (even with
> 400MHZ chips...)...  I love the Ultra as everyone knows, but I have to
> say, price talks bullshit walks... and Sun has a lot of trouble
> learning this especially on the high end.
> 

Yes, but at the high end costs of CPUs are irrelevant since it
usually becomes storage and network infrastructure costs that
predominate.

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
  1998-07-12  1:56 ` David S. Miller
  1998-07-12  2:21 ` Matthew Jacob
@ 1998-07-12  2:29 ` David S. Miller
  1998-07-12  5:02 ` Matthew Jacob
                   ` (37 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: David S. Miller @ 1998-07-12  2:29 UTC (permalink / raw)
  To: ultralinux

   Date: Sat, 11 Jul 1998 19:21:14 -0700 (PDT)
   From: Matthew Jacob <mjacob@feral.com>

   Yes, but at the high end costs of CPUs are irrelevant since it
   usually becomes storage and network infrastructure costs that
   predominate.

This argument sounds %100 logical and correct.

But then why doesn't Intel just cause 4 way SMP machines to cost $20k
just like Sun?

Later,
David S. Miller
davem@dm.cobaltmicro.com

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (2 preceding siblings ...)
  1998-07-12  2:29 ` David S. Miller
@ 1998-07-12  5:02 ` Matthew Jacob
  1998-07-12 12:54 ` Douglas Eadline
                   ` (36 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Matthew Jacob @ 1998-07-12  5:02 UTC (permalink / raw)
  To: ultralinux




>    Date: Sat, 11 Jul 1998 19:21:14 -0700 (PDT)
>    From: Matthew Jacob <mjacob@feral.com>
> 
>    Yes, but at the high end costs of CPUs are irrelevant since it
>    usually becomes storage and network infrastructure costs that
>    predominate.
> 
> This argument sounds %100 logical and correct.
> 
> But then why doesn't Intel just cause 4 way SMP machines to cost $20k
> just like Sun?
> 

Good question, although the answer may be "Intel is not in
the system's business".

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (3 preceding siblings ...)
  1998-07-12  5:02 ` Matthew Jacob
@ 1998-07-12 12:54 ` Douglas Eadline
  1998-07-12 23:59 ` Bob Drzyzgula
                   ` (35 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Douglas Eadline @ 1998-07-12 12:54 UTC (permalink / raw)
  To: ultralinux

On Sat, 11 Jul 1998, Bob Drzyzgula wrote:

> 
> Sun's newly announced Ultra AXmp board is a pretty interesting
> platform. Based on early pricing estimates I've seen, it should
> be possible to build a quad-processor @ 300+MHz, 2GB rackmount
> system for under around $US 20K. I have a lot of Solaris applications
> for this board, but I also wonder what capabilities one would
> expect ULtraLinux to have on this machine...
> 
> http://www.sun.com/microelectronics/SPARCengineUltraAXmp/ 
> http://www.sun.com/microelectronics/SPARCengineUltraAXmp/805-5865.pdf
> http://www.sun.com/microelectronics/SPARCengineUltraAXmp/s7-mod.pdf
> http://www.sun.com/microelectronics/SPARCengineUltraAXmp/cover980706.html
> http://www.sun.com/smi/Press/sunflash/9807/sunflash.980706.1.html
> 

While I find such hardware interesting, lately I seem to have
developed an attitude of "If it is not in Computer Shopper"
(or equivalent) I do not consider it commodity and therefore
it is premium priced. Such introductory prices do not mention
service contracts and upgrade prices - which add up.
Getting spare parts at the corner PC shop or discount store
has a certain comfort. 

At the same time, high end hardware usually has better reliability
but, the "off the shelf stuff" is not that bad.  I grew up on
non-Intel platforms, but at the end of the day, reliability, price to 
performance and cost of ownership always seem to go to Intel/Linux. 

This is my opinion - take what works for you - throw the rest away.

Doug Eadline
-------------------------------------------------------------------
Paralogic, Inc.           |     PEAK     |      Voice:+610.861.6960
115 Research Drive        |   PARALLEL   |        Fax:+610.861.8247
Bethlehem, PA 18017 USA   |  PERFORMANCE |    http://www.plogic.com
-------------------------------------------------------------------

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (4 preceding siblings ...)
  1998-07-12 12:54 ` Douglas Eadline
@ 1998-07-12 23:59 ` Bob Drzyzgula
  1998-07-13  6:27 ` Qiru Zhou
                   ` (34 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Bob Drzyzgula @ 1998-07-12 23:59 UTC (permalink / raw)
  To: ultralinux

OK, so now this is really interesting. I had no idea that
my mail would solicit this kind of response, especially
with a complete lack of any attempt to answer the actual
question (i.e. what will this thing do under Linux).

But since y'all raise the point, I think it might be well
to examine it a little.  I'm curious about the accuracy
of the assertion that the AXmp is overpriced compared to
commodity x86 hardware. First of all, vis-a-vis the AXmp,
Is there any "commodity" hardware? There's the Goliath, and
by the end of the year there should be some 450NX boards
around, but... commodity? A Goliath-based configuration
with top-end Pros should easily run into five digits,
hardly chicken feed. If you don't want to build one
yourself, there's always Compaq, HP, ALR, etc., but Sun has
nothing on those guys as far as price is concerned... Price
out a decently-configured 4-way ProLiant 7000 with 200MHz
Pros each at 1MB cache, and I think that you'll find that
it ain't pocket change. So what hardware, advertised in
Computer Shopper or whatever, would you compare the AXmp
to and find that the AXmp is way overpriced? If you want
to compare to, say, four uniprocessor or two dual-processor
systems, I'd say that would be a apples-and-oranges thing,
and you'd probably want to compare to the AXi, which sells
for around $1800, motherboard and processor together,
has dual-channel Symbios SCSI, 10/100 Ethernet and two
independant PCI busses, three slots on each. You can
easily build an 300MHz AXi-based system with a half-gig
of memory for under $5,000, showing SPEC95 ratings in the
mid-to-upper teens.

But back to the AXmp, it was dissed somewhat on the basis
of I/O, which I don't understand. The AXmp has four PCI
busses -- two sets of two, each set bridged into a 72-bit
crossbar channel; each set has a 66MHz bus and a 33MHz
bus. The memory bus operates at 576 bits wide into the
crossbar, and each of two UPA processor complexes pipe in
at 144bits wide and 120MHz with the 360MHz processors. The
board costs around $6000, the 300MHz processors around
$2200 each, and the memory is standard 168-pin DIMMs.

Xeon processors are in the same price range as the
UltraSPARC II processors, and, although the Xeon does
better on integer and the SPARC better on floating point,
the sum of SPECfp95 and SPECint95 for a 400MHz Xeon is
almost exactly the same as that for a 300MHz UltraSPARC
II. I can believe that a 450NX motherboard will cost less
than the AXmp (something that remains to be seen), but I
don't see the I/O capacity in the 450NX PCIset that I see
in the Sun Crossbar switch.

As far as what you hang off the back in the way of RAID
devices, it seems to me that the situation is largely the
same between a 450NX system and an AXmp. The AXmp uses
the 53C876 and the SC450NX, for example (which looks,
I must say, like a *very* nice platform) uses a 53C896
and a 53C810AE.  But the AXmp has more PCI bandwidth
than the SC450NX.  When I need a back end disk device,
I ususally use the Kingston DS-500 9-bay rackmount fully
decked out with dual power supplies, DE-100 SCA hot-swap
carriers & frames, a CMD CRD-5440 SCSI-to-SCSI controller
and a 6V lead-acid battery. Lately we've been populating
them with the Fujitsu MAB3091-SC 9.1 GB drives, which work
well with the CMD, run cool and of which I bought a big
batch recently for $640 each. A 4U rack-mount enclosure
with 45GB of disk in a RAID 5 setup (5 data, 1 parity,
1 hot spare, 1 warm spare) runs somewhere around $10K in
aggregate. You can put this thing on the back of just about
anything you want... Pentium, Xeon, AXmp, UltraEnterprise
6000, Alpha, whatever.

[And BTW, I'm looking at the AXmp not for file servers
but for compute servers, so I could care less about the
back-end disk. This is more of a Beowulf kind of situation,
and wondering about UltraPenguin on the AXmp is related
to this.  I've got users who burn up machines for months
straight doing stochastic simulations and monte-carlo
models.  In these cases, all I care about are local
CPU-memory bandwidth and floating point performance,
and the SPARC machines still beat the pants off an x86
for this; and yes, an Alpha or a PowerPC could probably
do better, but we've been a Sun shop since 1985 and other
RISC architectures are a hard sell...]

As far as maintenance and support, I generally take as
little as possible. Sun's bronze contracts aren't that
much, and you always have the option of buying spare parts
yourself. We mostly use Solaris on our SPARC systems,
but if UltraPenguin became a viable option for the AXmp,
then the Solaris support issue would largely go away
when the application allows it. (Although Sun still
makes you buy a Solaris RTU for the SME motherboards).

So, yean, the UltraEnterprise GigaPlane stuff is
outrageously priced. The 450 is still up there although
not quite as bad.  And yeah, you can build a decent
dual-processor x86 server for $2K (and that's what I do
for NT print servers, for example...), but when you get
into the quad-processor range, the AXmp seems to me to be
one of the most cost-competitive things that Sun has done
in a long time.  Please tell me how I'm wrong about this...

--Bob

-- 
==============================
Bob Drzyzgula                             It's not a problem
bob@drzyzgula.org                until something bad happens
==============================

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (5 preceding siblings ...)
  1998-07-12 23:59 ` Bob Drzyzgula
@ 1998-07-13  6:27 ` Qiru Zhou
  1998-07-13  6:38 ` David S. Miller
                   ` (33 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Qiru Zhou @ 1998-07-13  6:27 UTC (permalink / raw)
  To: ultralinux

After I looked the web page:
http://www.sun.com/microelectronics/SPARCengineUltraAXmp

I think it is an interesting HW and I believe it does have
better throughput and I believe SUN has better design for
highend SMP HW (crossbar, etc.). But I do have the following
questions to be answered if I really serious on it:

1. Operating system cost, stability and performance:
Unlike low end dual PII's, nobody will buy it just to
play with it. Then the question is what OS should I use.
If I use Solaris, what is the cost of the OS+dev environment?
From my last experience, this will be over $3000 per user.
If I use Linux SMP for Sparc, what is the current status of
it? Is it stable enough for a production system? Actually,

2. HW component availability, and cost:
From the spec, I read:
Memory		576 bits wide, 2GB max in 2 banks, 16 slots
Memory type	Fast page mode/EDO, 72 bits, 3.3V DIMMS

Line one means you have to fill at least 8 slot of identical
DIMMS and line two means these DIMM are pretty old and you
need special order. The EDO/FPM DIMMS are extremely non-standard.
Read http://silicon.micron.com/crucial/cart/html/selector.cfm,
then you know what I am talking about. I've in several occassion
to asked vendor make a special assemble of these DIMMs.

If Sun can fill up the DIMMS with good price, then I don't need
to worry about mem upgrade. But from my last experience, it
will cost you 3-4 times higher than a PC (Intel or Alpha).

I am not sure what other PC PCI card that I can put there.
We have SGI O2s with PCI slots. But I almost cannot put
any PCI card from retail store, since it is not compatible
with O2 or there is simply no driver for it (you may develop
your own, if you really want.). The problem is that these
SUN and SGI PCI systems are not open-architecture. And there
market is so small, compare with PC quantity, There are
almost no third party vendors are interested in develop a
driver for them. We saw the same the same problem for Apple,
too.

I like SUN HW, and I really like to see other good CPU to
compete with Intel. But the problem is these guys kill themself.
I guess an open architecture system makes a big difference.


We've been seriously looking for SUN solutions severel times and
we've been pushed away by their price and the performance (Sparc
are constantly at the low end in the CPU war, and in our benchmark
tests, even consider Intel and Intel clones). I hope this time Sun
learned how to compete and makes some interesting offfers. Otherwise
I don't think there are going to be long that Sun drops their
Sparc from workstation line, as SGI did.


================================Qiru Zhou                           qzhou@research.bell-labs.com
2D428 Bell Labs, Lucent Technologies          tel (908) 582-4562
600 Mountain Avenue, Murray Hill, NJ 07974    fax (908) 582-7308
================================

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (6 preceding siblings ...)
  1998-07-13  6:27 ` Qiru Zhou
@ 1998-07-13  6:38 ` David S. Miller
  1998-07-13 10:40 ` Bob Drzyzgula
                   ` (32 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: David S. Miller @ 1998-07-13  6:38 UTC (permalink / raw)
  To: ultralinux

   Date: 	Mon, 13 Jul 1998 02:27:27 -0400
   From: Qiru Zhou <qzhou@research.bell-labs.com>

   I am not sure what other PC PCI card that I can put there.  We have
   SGI O2s with PCI slots. But I almost cannot put any PCI card from
   retail store, since it is not compatible with O2 or there is simply
   no driver for it (you may develop your own, if you really want.).

Whats the problem?  I have off the shelf 3com and Tulip PCI ethernet
cards, and off the shelf Adaptec and NCR scsi cards, in my UltraSPARC
PCI systems here.  (I've also played around with a Matrox M-II card in
them as well) In fact one of my PCI Ultra's is a router into a test
subnet I have, using a 4-port Tulip card.

Just about any existing Linux PCI driver can be made to work in about
15 minutes of hacking done by one of the developers who work on the
port.

All your "open architecture" claims are false as well, you can get
complete documentation on Sun's bus, cpu, and device programming
interfaces on the PCI Ultra's and timing specs are available as well.

As far as I've seen, any PCI card I've thrown into one of my Ultra's
was found on the bus and would work just fine once I did the driver
porting work.  The only part where anything is problematic is where a
card has complex x86 firmware, and this matters really only for a
device which you'd like to boot off of.  But all the Ultra/PCI systems
have on board bootable devices with open-firmware on them so it's a
non-issue.

So don't give the PCI Ultra's and the O2's a bad rap just because the
support list Solaris and IRIX have are so small for off the shelf PC
cards.  It is not an indication of a sub-standard PCI bus
implementation at all, it's just plain lack of OS support from these
vendors, no more no less.

Later,
David S. Miller
davem@dm.cobaltmicro.com

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (7 preceding siblings ...)
  1998-07-13  6:38 ` David S. Miller
@ 1998-07-13 10:40 ` Bob Drzyzgula
  1998-07-13 12:09 ` Xavier Beaudouin
                   ` (31 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Bob Drzyzgula @ 1998-07-13 10:40 UTC (permalink / raw)
  To: ultralinux

On Mon, Jul 13, 1998 at 02:27:27AM -0400, Qiru Zhou wrote:
> After I looked the web page:
> http://www.sun.com/microelectronics/SPARCengineUltraAXmp
> 
> 1. Operating system cost, stability and performance:
> Unlike low end dual PII's, nobody will buy it just to
> play with it. Then the question is what OS should I use.
> If I use Solaris, what is the cost of the OS+dev environment?
> >From my last experience, this will be over $3000 per user.
> If I use Linux SMP for Sparc, what is the current status of
> it? Is it stable enough for a production system? Actually,

To play with, get an AXi. A bootable system with 128MB of memory
and maybe 4GB of disk should be obtainable for under $3K total.
A four-user copy of Solaris Server costs around $600 or less.
If you want a low-cost SMP Sun to play with, there are
a few places you can get remanufactured Suns for pretty
cheap. The Ultra 2s are still holding their price, and sell for
maybe $6-10K used depending on configuration. You can get
a dual-processor SS 20 or even a 1000 for less than $5K, 
if you don't mind messing around with Mbus stuff.

I, too am very interested in how UltraPenguin would do on
the AXmp, and still would like very much to know what
people thought about this...

> Line one means you have to fill at least 8 slot of identical
> DIMMS and line two means these DIMM are pretty old and you
> need special order. The EDO/FPM DIMMS are extremely non-standard.
> Read http://silicon.micron.com/crucial/cart/html/selector.cfm,
> then you know what I am talking about. I've in several occassion
> to asked vendor make a special assemble of these DIMMs.

In my experience, EDO ECC DIMMs are generally
available. From http://www.pricewatch.com I see that
32MB units should cost around $38, meaning you could
populate eight slots with a total of 256MB for around
$300. 128MB modules for this board should be well under
$200 each. Anyone who sells any of the AX* motherboards
should have the correct memory in stock. The procurement
group at my office has had no trouble finding DIMMs such
as these, and they do a lot more PC stuff than Sun stuff.
Offline I can give you the names of some vendors if
you want...

> I hope this time Sun
> learned how to compete and makes some interesting offfers. Otherwise
> I don't think there are going to be long that Sun drops their
> Sparc from workstation line, as SGI did.

Same here. Sun does have, however, a record of being
extremely scrappy, independant and stubborn, often
to a fault. They are also doing reasonably well
in terms of sales. Following the release of the AXi,
the sales of that board quickly cleared out their
inventory and overran their production; there was about
a 30-day period when you pretty much couldn't buy one.
I don't expect them to cave in any time soon.

--Bob

-- 
==============================
Bob Drzyzgula                             It's not a problem
bob@drzyzgula.org                until something bad happens
==============================

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (8 preceding siblings ...)
  1998-07-13 10:40 ` Bob Drzyzgula
@ 1998-07-13 12:09 ` Xavier Beaudouin
  1998-07-13 13:59 ` Robert G. Brown
                   ` (30 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Xavier Beaudouin @ 1998-07-13 12:09 UTC (permalink / raw)
  To: ultralinux

> If Sun can fill up the DIMMS with good price, then I don't need
> to worry about mem upgrade. But from my last experience, it
> will cost you 3-4 times higher than a PC (Intel or Alpha).

This is really a problem... I use 2 alpha and 5 sun sparc boxes... The
main problem is the cost of these memory upgrades since thoses boxes are
second hand...

Seems that some alpha motherboard are now compatible with those cheap PC
memory... That's good, but sun boxes not... This is really a pain I think,
for example in France for a 64Mb upgrade for SS10 I must pay 1000FF 
without the VAT (about $200 !!) a this price I've got about 256 Mb of ram
(okay, in 4 DIMMS) and this price is with the VAT.

Since today, the main problem with Sparc and Alpha boxes is not the CPU (I
don't care, if I need some more CPU power for one usage I add a new Alpha
or Sparc), but the RAM... 

> I am not sure what other PC PCI card that I can put there.
> We have SGI O2s with PCI slots. But I almost cannot put
> any PCI card from retail store, since it is not compatible
> with O2 or there is simply no driver for it (you may develop
> your own, if you really want.). The problem is that these
> SUN and SGI PCI systems are not open-architecture. And there
> market is so small, compare with PC quantity, There are
> almost no third party vendors are interested in develop a
> driver for them. We saw the same the same problem for Apple,
> too.

?? Hugh ? For Apple, I currently use PC video card and some capture card
without any problems, on Linux PPC (this maybe why I don't have any
problems ?). About PCI under non x86 boxes, seems that all my PCI card I
have tested on an Alpha (AlphaStation 4/266) are really working without
any problems : Video (S3, Cirrus, Matrox), Networking (3C900 / 3C590 /
Tulip), Some BUSLOGICS PCI (I don't use them for boot !) and the NCR
compatible PCI SCSI for ASUS...

I think the main problem for PCI card is :

 1- OS drivers for specific hardware... This is the problem
    of the manufacturer _and_ the OS (for example OpenBSD and Linux!)
 2- Some of PCI card use some ROM on it... Good, but allmost time
    this is x86 code... ;-( 
    -> Alpha Box have an x86 bios emulator... What about SGI and SUN ?
       

/Xavier

--
 Cybernet Consulting Associates - I.S.P. & Network Design and Consulting
 Xavier Beaudouin   -   Network Designer & Consultant   -   kiwi@oav.net
 Phone/Fax : +33 (0)1 4734 3366     -      Cellular : +33 (0)6 6026 5108
 Web sites : http://oav.net/  - http://cybernet.fr/  - http://kazar.com/
 [---------------------------------------------------------------------]
There's been no top authority saying what marijuana does to you.  I
really don't know that much about it.  I tried it once but it didn't do
anything to me.
		-- John Wayne

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (9 preceding siblings ...)
  1998-07-13 12:09 ` Xavier Beaudouin
@ 1998-07-13 13:59 ` Robert G. Brown
  1998-07-13 14:29 ` Robert HYATT
                   ` (29 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Robert G. Brown @ 1998-07-13 13:59 UTC (permalink / raw)
  To: ultralinux

On Sat, 11 Jul 1998, Matthew Jacob wrote:

> 
> 
> 
> >    Date: Sat, 11 Jul 1998 19:21:14 -0700 (PDT)
> >    From: Matthew Jacob <mjacob@feral.com>
> > 
> >    Yes, but at the high end costs of CPUs are irrelevant since it
> >    usually becomes storage and network infrastructure costs that
> >    predominate.
> > 
> > This argument sounds %100 logical and correct.
> > 
> > But then why doesn't Intel just cause 4 way SMP machines to cost $20k
> > just like Sun?
> > 
> 
> Good question, although the answer may be "Intel is not in
> the system's business".
> 
> 
> 
> 

...and "manufacturer's who use Intel chips in 4+ CPU SMP machines have to
actually compete in an open market"....

I'm also not sure about the original premise.  At the high end of
multiprocessing systems (or even at the low end:-), I would have
expected the bulk of the cost to be in the design and implementation
of the bus structure that connects the processors to enable shared
access to memory, peripherals, and each other to facilitate high speed
IPC's.  Disk is cheap (especially when packaged for a standard
interface, i.e. -- SCSI UW on a PCI bus) and really NIC's (similarly
packaged) are too.  Getting disk and networks onto a bus (proprietary
or otherwise) and interfaced with lots of processors and memory and
working out all the DMA issues and cache coherence issues and
busmastering issues -- that's expensive.

I thought the real claim to fame of all of Sun's multiprocessing
systems (and the SP2, and the Power Challenge, and the...) has always
been their really fast/expensive IPC bus, not their interface to
peripherals.

    rgb

Robert G. Brown	                       http://www.phy.duke.edu/~rgb/
Duke University Dept. of Physics, Box 90305
Durham, N.C. 27708-0305
Phone: 1-919-660-2567  Fax: 919-660-2525     email:rgb@phy.duke.edu

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (10 preceding siblings ...)
  1998-07-13 13:59 ` Robert G. Brown
@ 1998-07-13 14:29 ` Robert HYATT
  1998-07-13 14:54 ` Matthew Jacob
                   ` (28 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Robert HYATT @ 1998-07-13 14:29 UTC (permalink / raw)
  To: ultralinux


On Mon, 13 Jul 1998, Robert G. Brown wrote:

> 
> ...and "manufacturer's who use Intel chips in 4+ CPU SMP machines have to
> actually compete in an open market"....
> 
> I'm also not sure about the original premise.  At the high end of
> multiprocessing systems (or even at the low end:-), I would have
> expected the bulk of the cost to be in the design and implementation
> of the bus structure that connects the processors to enable shared
> access to memory, peripherals, and each other to facilitate high speed
> IPC's.  Disk is cheap (especially when packaged for a standard
> interface, i.e. -- SCSI UW on a PCI bus) and really NIC's (similarly
> packaged) are too.  Getting disk and networks onto a bus (proprietary
> or otherwise) and interfaced with lots of processors and memory and
> working out all the DMA issues and cache coherence issues and
> busmastering issues -- that's expensive.
> 
> I thought the real claim to fame of all of Sun's multiprocessing
> systems (and the SP2, and the Power Challenge, and the...) has always
> been their really fast/expensive IPC bus, not their interface to
> peripherals.
> 


you are correct.  If you take a high-end machine, like a Cray T90, over
50% of the total cost is the memory interconnection network, because it
is so hard to provide almost a terrabyte of data/second spread over 32
cpus.  :)

but the memory interconnect is always the sticky wicket on SMP machines,
although for the PC, we simply dribble along at (now) 100mhz.  :)

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (11 preceding siblings ...)
  1998-07-13 14:29 ` Robert HYATT
@ 1998-07-13 14:54 ` Matthew Jacob
  1998-07-13 14:56 ` Matthew Jacob
                   ` (27 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Matthew Jacob @ 1998-07-13 14:54 UTC (permalink / raw)
  To: ultralinux



On Mon, 13 Jul 1998, Robert G. Brown wrote:


> ...and "manufacturer's who use Intel chips in 4+ CPU SMP machines have to
> actually compete in an open market"....
> 
> I'm also not sure about the original premise.  At the high end of
> multiprocessing systems (or even at the low end:-), I would have
> expected the bulk of the cost to be in the design and implementation
> of the bus structure that connects the processors to enable shared
> access to memory, peripherals, and each other to facilitate high speed
> IPC's.  Disk is cheap (especially when packaged for a standard
> interface, i.e. -- SCSI UW on a PCI bus) and really NIC's (similarly
> packaged) are too.  Getting disk and networks onto a bus (proprietary
> or otherwise) and interfaced with lots of processors and memory and
> working out all the DMA issues and cache coherence issues and
> busmastering issues -- that's expensive.

I think we're all beginning to talk about different things. I believe
you can build a very reasonable 2 (to maybe 4-way SMP with off the shelf
Intel parts- good both for CPU/thread performance as well as the
attachment of a couple PCI busses. That'll not give you a balanced system,
but a *better* balanced system than a workstation with a single bus.

When you get into bigger than this, that's when, as you say, it
gets very expensive. And single SCSI disks *are* cheap, but the
infrastructure for 100 of them gets quite expensive.

> 
> I thought the real claim to fame of all of Sun's multiprocessing
> systems (and the SP2, and the Power Challenge, and the...) has always
> been their really fast/expensive IPC bus, not their interface to
> peripherals.

wrt Sun: Surely you're kidding (unless you're referring only to the EXXXK
series).

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (12 preceding siblings ...)
  1998-07-13 14:54 ` Matthew Jacob
@ 1998-07-13 14:56 ` Matthew Jacob
  1998-07-13 14:59 ` Matti Aarnio
                   ` (26 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Matthew Jacob @ 1998-07-13 14:56 UTC (permalink / raw)
  To: ultralinux


> but the memory interconnect is always the sticky wicket on SMP machines,
> although for the PC, we simply dribble along at (now) 100mhz.  :)

Umm.. I thought that the new PC motherboards coming out are running
at 400Mhz (no, not CPU speed- I mean internal bus speed...)

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (13 preceding siblings ...)
  1998-07-13 14:56 ` Matthew Jacob
@ 1998-07-13 14:59 ` Matti Aarnio
  1998-07-13 15:24 ` Robert HYATT
                   ` (25 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Matti Aarnio @ 1998-07-13 14:59 UTC (permalink / raw)
  To: ultralinux

One "spelling" error buggers me major way:
... 
> you are correct.  If you take a high-end machine, like a Cray T90, over
> 50% of the total cost is the memory interconnection network, because it
> is so hard to provide almost a terrabyte of data/second spread over 32
> cpus.  :)

	There is no such multiplier prefix as "terra".  The one you were
	thinking about is "tera" (10**12 in Fortran notation).
	That is the multiplier that is meant when "T" is used -- thus
	"almoast a TB of data per second" (of course "TB" means something
	else too, but that goes way beside the point.)

	Although MicroSoft has a system called TerraServer(.microsoft.com),
	they do use the intentional misspelling of the prefix to refer to
	its size, although they also use the Latin "Terra" to mean "Earth"
	to refer to its contents...

> but the memory interconnect is always the sticky wicket on SMP machines,
> although for the PC, we simply dribble along at (now) 100mhz.  :)

Whatever systems we consider to buy/use, unless they are of commodity
type products, they will always cost way more than the commodity stuff.
Of course, nothing beats T90 in memory bandwidth per CPU :-)

/Matti Aarnio <matti.aarnio@sonera.fi>

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (14 preceding siblings ...)
  1998-07-13 14:59 ` Matti Aarnio
@ 1998-07-13 15:24 ` Robert HYATT
  1998-07-13 15:38 ` Robert HYATT
                   ` (24 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Robert HYATT @ 1998-07-13 15:24 UTC (permalink / raw)
  To: ultralinux


no.  that is the cache to cpu speed.  memory is still stuck on the
good old 100mhz bus.


Robert Hyatt                    Computer and Information Sciences
hyatt@cis.uab.edu               University of Alabama at Birmingham
(205) 934-2213                  115A Campbell Hall, UAB Station 
(205) 934-5473 FAX              Birmingham, AL 35294-1170

On Mon, 13 Jul 1998, Matthew Jacob wrote:

> 
> > but the memory interconnect is always the sticky wicket on SMP machines,
> > although for the PC, we simply dribble along at (now) 100mhz.  :)
> 
> Umm.. I thought that the new PC motherboards coming out are running
> at 400Mhz (no, not CPU speed- I mean internal bus speed...)
> 
> 
> 

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (15 preceding siblings ...)
  1998-07-13 15:24 ` Robert HYATT
@ 1998-07-13 15:38 ` Robert HYATT
  1998-07-13 16:56 ` Robert G. Brown
                   ` (23 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Robert HYATT @ 1998-07-13 15:38 UTC (permalink / raw)
  To: ultralinux



On Mon, 13 Jul 1998, Matti Aarnio wrote:

> One "spelling" error buggers me major way:
> ... 
> > you are correct.  If you take a high-end machine, like a Cray T90, over
> > 50% of the total cost is the memory interconnection network, because it
> > is so hard to provide almost a terrabyte of data/second spread over 32
> > cpus.  :)
> 
> 	There is no such multiplier prefix as "terra".  The one you were
> 	thinking about is "tera" (10**12 in Fortran notation).
> 	That is the multiplier that is meant when "T" is used -- thus
> 	"almoast a TB of data per second" (of course "TB" means something
> 	else too, but that goes way beside the point.)


sorry...  I generally get "terabyte" correct.  But I have been a science-
fiction reader for 40 years, and read all the time, and "terra" is a
natural screw-up for those of us so inclined.  :)



> 
> 	Although MicroSoft has a system called TerraServer(.microsoft.com),
> 	they do use the intentional misspelling of the prefix to refer to
> 	its size, although they also use the Latin "Terra" to mean "Earth"
> 	to refer to its contents...
> 
> > but the memory interconnect is always the sticky wicket on SMP machines,
> > although for the PC, we simply dribble along at (now) 100mhz.  :)
> 
> Whatever systems we consider to buy/use, unless they are of commodity
> type products, they will always cost way more than the commodity stuff.
> Of course, nothing beats T90 in memory bandwidth per CPU :-)
> 
> /Matti Aarnio <matti.aarnio@sonera.fi>
> 

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (16 preceding siblings ...)
  1998-07-13 15:38 ` Robert HYATT
@ 1998-07-13 16:56 ` Robert G. Brown
  1998-07-13 17:04 ` Robert G. Brown
                   ` (22 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Robert G. Brown @ 1998-07-13 16:56 UTC (permalink / raw)
  To: ultralinux

On Mon, 13 Jul 1998, Matthew Jacob wrote:

> I think we're all beginning to talk about different things. I believe
> you can build a very reasonable 2 (to maybe 4-way SMP with off the shelf
> Intel parts- good both for CPU/thread performance as well as the
> attachment of a couple PCI busses. That'll not give you a balanced system,
> but a *better* balanced system than a workstation with a single bus.
> 
> When you get into bigger than this, that's when, as you say, it
> gets very expensive. And single SCSI disks *are* cheap, but the
> infrastructure for 100 of them gets quite expensive.

Total agreement.  Although I don't believe that >>most<< 4-processor
system clients are looking for systems that support a terabyte or so of
online, fast hard disk.  Or rather, we'd all probably love to have
one, if somebody else will pay for it...:-)

> > I thought the real claim to fame of all of Sun's multiprocessing
> > systems (and the SP2, and the Power Challenge, and the...) has always
> > been their really fast/expensive IPC bus, not their interface to
> > peripherals.
> 
> wrt Sun: Surely you're kidding (unless you're referring only to the EXXXK
> series).

Half kidding, maybe.  Even the old Sparc 1000 and 2000 systems had
decent IPC/memory busses compared to, well, at the time there wasn't
much to compare to but even now to a dual PPro, and if you could
afford them gave you quite a few processors on a native IPC bus
(definitely better than network IPC's).  I couldn't say what the
current profile of Sun vs IBM vs SGI multiprocessor systems is,
though, because every time I've looked at it in the the past it has
been absurdly far from commodity Intel boxes (in either dual or quad
packaging) in price/performance for our task mix.  So I've stopped
looking, except passively by following posts here and elsewhere.

At $20K for a fully equipped quad system, it sounds like Sun has made
major strides in price performance, but I'm afraid that it is still a
matter of reducing Intel's price/performance advantage for our task
mix from 3 to 2.

   rgb

Robert G. Brown	                       http://www.phy.duke.edu/~rgb/
Duke University Dept. of Physics, Box 90305
Durham, N.C. 27708-0305
Phone: 1-919-660-2567  Fax: 919-660-2525     email:rgb@phy.duke.edu

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (17 preceding siblings ...)
  1998-07-13 16:56 ` Robert G. Brown
@ 1998-07-13 17:04 ` Robert G. Brown
  1998-07-13 17:13 ` Lawrence D. Lopez
                   ` (21 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Robert G. Brown @ 1998-07-13 17:04 UTC (permalink / raw)
  To: ultralinux

On Mon, 13 Jul 1998, Matthew Jacob wrote:

> 
> > but the memory interconnect is always the sticky wicket on SMP machines,
> > although for the PC, we simply dribble along at (now) 100mhz.  :)
> 
> Umm.. I thought that the new PC motherboards coming out are running
> at 400Mhz (no, not CPU speed- I mean internal bus speed...)

The current 440BX systems have a 100 MHz memory bus (50% faster than
the old PPro/PII bus but still slower than one would like).  As I
learned very forcefully recently, their cache speed is half the CPU
clock, so a 400 MHz CPU has a cache running at 200 MHz, making it
finally the equal or superior of the PPro in every respect.  The
motherboards support PC-100 SDRAM, which is measurably faster in than
EDO in some applications -- but your milage will very much depend upon
your application -- its size (does it fit in cache), its locality
(does it cache thrash bytewise or read streaming sequences of data)
and the like.  Unless you are running jobs more than a few MB in size
(including data) you will, frankly, probably not see much advantage in
the SDRAM memory or the memory clock, but folks doing large matrix
multiplies and the like have reported significant speedup and
considerably better SMP scaling of memory access times.

Now, there are a whole slew of things purported to be coming down the
pipe from Intel this fall starting with 450MHz processors, and (as I
recall) we should see "Merced" (their next generation CPU and system)
sometime next year.  In merced I believe we'll see some very
significant capacity/speed upgrades in lots of venues -- you might
check out the Intel website to see what they say.

   rgb

Robert G. Brown	                       http://www.phy.duke.edu/~rgb/
Duke University Dept. of Physics, Box 90305
Durham, N.C. 27708-0305
Phone: 1-919-660-2567  Fax: 919-660-2525     email:rgb@phy.duke.edu

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (18 preceding siblings ...)
  1998-07-13 17:04 ` Robert G. Brown
@ 1998-07-13 17:13 ` Lawrence D. Lopez
  1998-07-13 17:26 ` Mike
                   ` (20 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Lawrence D. Lopez @ 1998-07-13 17:13 UTC (permalink / raw)
  To: ultralinux

I think its like this:

PCI bus 	33 MHZ
Memory bus	100 MHZ
L2 cache bus	200 MHZ
L1 cache bus	400 MHZ
Internal stuff	400 MHZ

So, internal bus speed depends on which internal bus!

Matthew Jacob wrote:
> 
> > but the memory interconnect is always the sticky wicket on SMP machines,
> > although for the PC, we simply dribble along at (now) 100mhz.  :)
> 
> Umm.. I thought that the new PC motherboards coming out are running
> at 400Mhz (no, not CPU speed- I mean internal bus speed...)

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (19 preceding siblings ...)
  1998-07-13 17:13 ` Lawrence D. Lopez
@ 1998-07-13 17:26 ` Mike
  1998-07-14  1:44 ` Shannon
                   ` (19 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Mike @ 1998-07-13 17:26 UTC (permalink / raw)
  To: ultralinux

Robert G. Brown wrote:

> The current 440BX systems have a 100 MHz memory bus (50% faster than
> the old PPro/PII bus but still slower than one would like).  As I
> learned very forcefully recently, their cache speed is half the CPU
> clock, so a 400 MHz CPU has a cache running at 200 MHz, making it
> finally the equal or superior of the PPro in every respect.  

That was so until just over a week ago. The Xeon version of the Pentium
II runs the L2 cache at processor speed. So far as I can see, this has a
minimal effect on improving performance. The only plus is being able to
run 8 of them together.

Mike.

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (20 preceding siblings ...)
  1998-07-13 17:26 ` Mike
@ 1998-07-14  1:44 ` Shannon
  1998-07-14  7:00 ` David S. Miller
                   ` (18 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Shannon @ 1998-07-14  1:44 UTC (permalink / raw)
  To: ultralinux


In message <199807130638.XAA09243@dm.cobaltmicro.com>, "David S. Miller" writes
:

>    Date: 	Mon, 13 Jul 1998 02:27:27 -0400
>    From: Qiru Zhou <qzhou@research.bell-labs.com>
> 
>    I am not sure what other PC PCI card that I can put there.  We have
>    SGI O2s with PCI slots. But I almost cannot put any PCI card from
>    retail store, since it is not compatible with O2 or there is simply
>    no driver for it (you may develop your own, if you really want.).
> 
> As far as I've seen, any PCI card I've thrown into one of my Ultra's
> was found on the bus and would work just fine once I did the driver
> porting work.  The only part where anything is problematic is where a
> card has complex x86 firmware, and this matters really only for a
> device which you'd like to boot off of.  But all the Ultra/PCI systems
> have on board bootable devices with open-firmware on them so it's a
> non-issue.


On this note, does Sun have anything that will emulate x86 BIOS?

My Alphastation runs the BIOS for all my PCI cards via emulation
in the ROM during bootup.

--
csh - shendrix@widomaker.com - http://www.widomaker.com/~shendrix/myresume.html
----------------------------------------------------------------------
"Meddle not in the affairs of Wizards, for thou art crunchy, and
taste good with ketchup."

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (21 preceding siblings ...)
  1998-07-14  1:44 ` Shannon
@ 1998-07-14  7:00 ` David S. Miller
  1998-07-14 18:07 ` Douglas Eadline
                   ` (17 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: David S. Miller @ 1998-07-14  7:00 UTC (permalink / raw)
  To: ultralinux

   Date: Mon, 13 Jul 1998 21:44:51 -0400
   From: Shannon <shendrix@escape.widomaker.com>

   On this note, does Sun have anything that will emulate x86 BIOS?

   My Alphastation runs the BIOS for all my PCI cards via emulation in
   the ROM during bootup.

No they don't, and personally I feel this was a good decision.

We had argued internally inside the UltraPenguin team whether we
should add such a thing to either our bootloader or the kernel
itself.  We balked for two reasons:

1) It's heavily complex, even 'real PC' machines crash on bootup due
   to really strange x86 firmware found on some PCI cards.  It would
   take several long months of non-stop work to get this right.

2) We felt that in no case were users being left totally out of
   luck because we lacked this "feature".

My main reason for agreeing with Sun for not providing such a thing is
simple, it's silly to further encourage PCI card manufacturers to
continue writing CPU-specific firmware.  I know it's a pipe dream to
get them to stop entirely...  (and note IMHO things like OBP are the
right way to go, CPU independant and quite portable)  And yes I
realize how much the way PCI is implemented on PC's drives what the
"PCI standard" really is.

Later,
David S. Miller
davem@dm.cobaltmicro.com

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (22 preceding siblings ...)
  1998-07-14  7:00 ` David S. Miller
@ 1998-07-14 18:07 ` Douglas Eadline
  1998-07-14 23:23 ` Bob Drzyzgula
                   ` (16 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Douglas Eadline @ 1998-07-14 18:07 UTC (permalink / raw)
  To: ultralinux

On Mon, 13 Jul 1998, Mike wrote:

> Robert G. Brown wrote:
> 
> > The current 440BX systems have a 100 MHz memory bus (50% faster than
> > the old PPro/PII bus but still slower than one would like).  As I
> > learned very forcefully recently, their cache speed is half the CPU
> > clock, so a 400 MHz CPU has a cache running at 200 MHz, making it
> > finally the equal or superior of the PPro in every respect.  
> 
> That was so until just over a week ago. The Xeon version of the Pentium
> II runs the L2 cache at processor speed. So far as I can see, this has a
> minimal effect on improving performance. The only plus is being able to
> run 8 of them together.
> 
> Mike.
> 

Well the first news I read about the Xeon and supporting chipset
had some bug(s).  IMO, the Xeon is being positioned away from the PII
(although inside it is a Pentium II with faster cache).  I see the Xeon
is priced VERY expensive because it is intended for the server market.
I think it is Intel's plan to get more $$ for its CPUs that are
used in servers.  Servers always provide manufactures more margin than
desktops, but up until now, servers used "desktop chips" like the PII and
Pentium Pro.  So Dell and Compaq could make much more $$ selling a server,
but Intel made the same amount.  Now that will change.  Look at the quantity
1000 price for the Xeon:

400 MHz - 2 MB cache $4489 !!!!!!!!!!!
400 MHz - 1 MB cache $2836
400 MHz - 512K cache $1124

Some report that the performance difference will be only 10%
compared to a PII of the same speed - which of course depends
on the application.   Other than a faster cache, the only big difference
is that the Xeon can run 4 or even 8 on a motherboard.  So Intel
now has a differentiated product for it's server customers.  Whether
the cost is worth the performance gain for single user SMP or SMP
clusters remains unclear.

Doug Eadline


-------------------------------------------------------------------
Paralogic, Inc.           |     PEAK     |      Voice:+610..861.6960
115 Research Drive        |   PARALLEL   |        Fax:+610.861.8247
Bethlehem, PA 18017 USA   |  PERFORMANCE |    http://www.plogic.com
-------------------------------------------------------------------

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (23 preceding siblings ...)
  1998-07-14 18:07 ` Douglas Eadline
@ 1998-07-14 23:23 ` Bob Drzyzgula
  1998-07-15  3:31 ` Dave Wreski
                   ` (15 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Bob Drzyzgula @ 1998-07-14 23:23 UTC (permalink / raw)
  To: ultralinux

Well, I guess that answers the question of why Intel doesn't
cause quad-processor x86 servers to cost $20K... they just
hadn't gotten around to it yet. :-)

BTW, does anyone happen to know how a Quad-processor
Ultra AXmp system is likely to do under UltraPenguin?
From what I've seen it should be possible to build such
a thing for about $20K. I've seen some discussion of
this board on some Linux mailing lists, but Linux-specific
issues related to this board weren't discussed... :-\

--Bob

On Tue, Jul 14, 1998 at 02:07:25PM -0400, Douglas Eadline wrote:
> On Mon, 13 Jul 1998, Mike wrote:
> 
> > Robert G. Brown wrote:
> > 
> > > The current 440BX systems have a 100 MHz memory bus (50% faster than
> > 
> > That was so until just over a week ago. The Xeon version of the Pentium
> > 
> 
> 400 MHz - 2 MB cache $4489 !!!!!!!!!!!
> 400 MHz - 1 MB cache $2836
> 400 MHz - 512K cache $1124

-- 
==============================
Bob Drzyzgula                             It's not a problem
bob@drzyzgula.org                until something bad happens
==============================

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (24 preceding siblings ...)
  1998-07-14 23:23 ` Bob Drzyzgula
@ 1998-07-15  3:31 ` Dave Wreski
  1998-07-15  8:07 ` Ward Deng
                   ` (14 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Dave Wreski @ 1998-07-15  3:31 UTC (permalink / raw)
  To: ultralinux

>    Based on early pricing estimates I've seen, it should be possible
>    to build a quad-processor @ 300+MHz, 2GB rackmount system for under
>    around $US 20K.
> 
> I wonder what the equivalent quad-PPRO system would cost (even with
> 400MHZ chips...)...  I love the Ultra as everyone knows, but I have to
> say, price talks bullshit walks... and Sun has a lot of trouble
> learning this especially on the high end.

How about two CPUs for $198,400?  From this month's SunExpert:

"The top-of-the-line Enterprise 6500 can hold a maximum of 30 CPUs, two 8.4-Gig
disk boards (each containing two drives) more than 375 GB of rack-mounted
storage in the system cabinet and 30 GB of memory.  A basic configuration with
8.4 GB of internal storage, 256 MB of memory and two CPUs costs $198,400 for a
250-Mhz version, and $204,400 for a 336-Mhz version"

Besides the ability to hold 30 GB of RAM, and 1/3 GB of disk, what could
possibly make it so expensive?  Does the Intel SMP arch support 30 CPUs?

The picture is not very good, but it looks as high as a convention refrigerator.

Dave

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (25 preceding siblings ...)
  1998-07-15  3:31 ` Dave Wreski
@ 1998-07-15  8:07 ` Ward Deng
  1998-07-15 11:14 ` Bob Drzyzgula
                   ` (13 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Ward Deng @ 1998-07-15  8:07 UTC (permalink / raw)
  To: ultralinux

> 
> I thought the real claim to fame of all of Sun's multiprocessing
> systems (and the SP2, and the Power Challenge, and the...) has always
> been their really fast/expensive IPC bus, not their interface to
                                   ^^^^^^^
What do you mean that?

IBM RS6000/SP2 is distributed memory system, or workstations linked with
external high-speed switch. Sun's MP systems are all SMP design. SGI 
Power Challenge or Onyx are SMP too. The newer SGI Origin is SMP with
special memory links -- ccNUMA. They are very different architectures
while SP nodes are no different from IBM's RS6000 workstations.

> peripherals.
> 
>     rgb
> 
> Robert G. Brown	                       http://www.phy.duke.edu/~rgb/
> Duke University Dept. of Physics, Box 90305
> Durham, N.C. 27708-0305
> Phone: 1-919-660-2567  Fax: 919-660-2525     email:rgb@phy.duke.edu
> 
> 
> 
> 

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (26 preceding siblings ...)
  1998-07-15  8:07 ` Ward Deng
@ 1998-07-15 11:14 ` Bob Drzyzgula
  1998-07-15 11:27 ` David S. Miller
                   ` (12 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Bob Drzyzgula @ 1998-07-15 11:14 UTC (permalink / raw)
  To: ultralinux

On Wed, Jul 15, 1998 at 02:07:05AM -0600, Ward Deng wrote:
> > 
> > I thought the real claim to fame of all of Sun's multiprocessing
> > systems (and the SP2, and the Power Challenge, and the...) has always
> > been their really fast/expensive IPC bus, not their interface to
>                                    ^^^^^^^
> What do you mean that?

I belive that, in the case of the UltraSPARC, it would
refer to the Ultra Port Architecture (UPA), which
provides a point-to-point packet switched interconnect
for processors, memory and I/O. The UPA has special
accelerators for the cache coherancy mechanism, and
supports device connections with various bitwidths.
At over four processors, the UPA gets extended into
a more distributed architecture called the gigaplane.

Personally, I would rather that Sun had a distributed
memory machine... the UltraSPARC machines have per-CPU
cache but not per-CPU DRAM. I believe that the
UltraSPARC III specifically does provide for per-CPU
DRAM; I don't know excactly how that would be implemented
in a system architecture or OS-wise.

--Bob

See:
http://www.sun.com/microelectronics/whitepapers/wp95-023.html.
http://wwwwseast2.usec.sun.com/servers/enterprise/10000/Tour/interconnect.html
http://wwwwseast2.usec.sun.com/servers/enterprise/10000/wp/E10000.pdf


-- 
==============================
Bob Drzyzgula                             It's not a problem
bob@drzyzgula.org                until something bad happens
==============================

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (27 preceding siblings ...)
  1998-07-15 11:14 ` Bob Drzyzgula
@ 1998-07-15 11:27 ` David S. Miller
  1998-07-15 15:54 ` Robert G. Brown
                   ` (11 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: David S. Miller @ 1998-07-15 11:27 UTC (permalink / raw)
  To: ultralinux

   Date: 	Wed, 15 Jul 1998 07:14:23 -0400
   From: Bob Drzyzgula <bob@drzyzgula.org>

   I believe that the UltraSPARC III specifically does provide for
   per-CPU DRAM; I don't know excactly how that would be implemented
   in a system architecture or OS-wise.

It can be treated as processor local memory, and thus requires
something along the lines of NUMA to encourage localization.

I would predict that UPA is scrapped in Ultra-III (Cheetah) for a new
processor interconnect with slightly different characteristics.  I
really have high hopes for the memory bandwidth characteristics of
Cheetah in general.  It'll probably be 6 or 8 scalar as well.
(these are all blind predictions, I have no idea what it'll really
 be like)

Later,
David S. Miller
davem@dm.cobaltmicro.com

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (28 preceding siblings ...)
  1998-07-15 11:27 ` David S. Miller
@ 1998-07-15 15:54 ` Robert G. Brown
  1998-07-15 17:11 ` Robert G. Brown
                   ` (10 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Robert G. Brown @ 1998-07-15 15:54 UTC (permalink / raw)
  To: ultralinux

On Tue, 14 Jul 1998, Bob Drzyzgula wrote:

> Well, I guess that answers the question of why Intel doesn't
> cause quad-processor x86 servers to cost $20K... they just
> hadn't gotten around to it yet. :-)

Yeah, but watch competition drop that price like a rock over the next
year.  Intel is just trying to squeeze a bit of extra juice out of the
lemon before its X86 competitors come out with multiprocessing CPUs
and before it comes out with 450 MHz PII's and ultimately merced.  PII
prices have been in slow-but-steady decline from the minute they came
out, and let's be frank -- unless somebody shows a huge performance
advantage for the new processors (and so far it looks, unsurprisingly
to me, like good old clock by itself continues to dominate overall CPU
performance) they are not going to sell if they are priced at 2-4
times single CPU prices.

> > 400 MHz - 2 MB cache $4489 !!!!!!!!!!!
> > 400 MHz - 1 MB cache $2836
> > 400 MHz - 512K cache $1124

These prices are not so bad, really -- I've paid $1100 for 512K 200MHz
PPros!  The only unknown is the cost of a quad motherboard that uses
them.  The Goliath quads have always been expensive, but if quads come
out at no more than (say) $1K and octets at $2K (I can dream, can't
I?) then a four processor system would probably cost around $6K in
components, plus memory; an eight processor system would cost perhaps
$11K, plus memory.  Memory is likely to be a bigger hit -- 128 MB
SDRAM DIMMS are cheap at $200, but 256 MB DIMMS are still rather dear
at ~$800-900.  If the motherboards still come with only four DIMM
slots, getting more than 128 MB/processor will get pretty expensive.

Note that I discount the 1MB cache and 2 MB cache CPU's entirely -- for
MOST beowulfy apps I really don't think that this will matter enough
to be worth it until they drop these prices significantly.  I'd rather
have 2 512K cache CPUs than 0.8 1024K cache CPUs...although real
webserver-type apps might find the bigger caches worth the money.

   rgb

Robert G. Brown	                       http://www.phy.duke.edu/~rgb/
Duke University Dept. of Physics, Box 90305
Durham, N.C. 27708-0305
Phone: 1-919-660-2567  Fax: 919-660-2525     email:rgb@phy.duke.edu

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (29 preceding siblings ...)
  1998-07-15 15:54 ` Robert G. Brown
@ 1998-07-15 17:11 ` Robert G. Brown
  1998-07-15 23:22 ` Luis Ponce de Leao
                   ` (9 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Robert G. Brown @ 1998-07-15 17:11 UTC (permalink / raw)
  To: ultralinux

On Wed, 15 Jul 1998, Ward Deng wrote:

> > 
> > I thought the real claim to fame of all of Sun's multiprocessing
> > systems (and the SP2, and the Power Challenge, and the...) has always
> > been their really fast/expensive IPC bus, not their interface to
>                                    ^^^^^^^
> What do you mean that?
> 
> IBM RS6000/SP2 is distributed memory system, or workstations linked with
> external high-speed switch. Sun's MP systems are all SMP design. SGI 
> Power Challenge or Onyx are SMP too. The newer SGI Origin is SMP with
> special memory links -- ccNUMA. They are very different architectures
> while SP nodes are no different from IBM's RS6000 workstations.

That is what I meant by that.  Perhaps I should have said "IPC
Channels" to avoid suggesting that they were all actually bus-linked.

The point I was trying to make is that "real parallel supercomputer"
systems advertise CPU-to-CPU (or process to process, if you want to be
picky) communication rates that significantly exceed what one can (or
could when the systems originally were released) be obtained from a
conventional network, i.e. a beowulf.  Whether IPC's are attained with
real shared memory on a common, arbitrated bus, dedicated, proprietary
internode switches, or whatever, the "point" of buying expensive "real"
parallel multiprocessing systems instead of building and using a
beowulf is that they supposedly have faster (often significantly, e.g.
order of magnitude faster) IPC's with reduced latencies to match and
hence far better scaling for fine-grained parallel problems.

For coarse to medium grained parallel problems, of course, this extra,
VERY EXPENSIVE speed is wasted, which is why in many cases a per-node
comparison of an SP2, a Power Challenge, and a beowulf comes down to
essentially a direct comparison of the CPU speeds themselves.  A
factor of two CPU speed advantage is worthless if the COST of the SP2
(per processor) is 5-10 times that of a commodity processor.

Re: the other part of the discussion, many of these systems ALSO have
fast mainframe-like interfaces to various peripherals, e.g. disk
arrays, to avoid bottlenecks that obviously can occur there, but
sensible calculation design avoids peripheral communications like the
plague it is -- if possible.  Disk operations are typically VERY
expensive even with the fastest disk, measured in wasted raw CPU time.

I think that bottom line (on which most of us can agree) is that when
considering or comparing parallel systems (or preparing to engineer
one of your own) one cannot be misled by glitter or the dazzling array
of benchmarks of this or that -- one has to consider the problems to
be solved with the system FIRST AND FOREMOST and THEN consider the
optimal parallel technology from all points of view, including
communications topology, price/performance, scaling, the various
minimax involved (it does no good to buy at optimal price/performance
if the best system one can buy is too small to solve your problem).

From what I have seen on the various lists I belong to, I believe that
80-90% of parallelizable problems fall in the coarse-to-medium grain
category that is optimally accessible by a beowulf-type architecture.
10-20% (with some overlap -- the distinctions are not sharp) are
optimally accessible by garden variety "parallel supercomputers" like
the ones discussed above, and 1-2% are accessible only by exotic
systems like current generation Crays or homemade/custom/dedicated
parallel systems.

    rgb

Robert G. Brown	                       http://www.phy.duke.edu/~rgb/
Duke University Dept. of Physics, Box 90305
Durham, N.C. 27708-0305
Phone: 1-919-660-2567  Fax: 919-660-2525     email:rgb@phy.duke.edu

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (30 preceding siblings ...)
  1998-07-15 17:11 ` Robert G. Brown
@ 1998-07-15 23:22 ` Luis Ponce de Leao
  1998-07-21 23:29 ` Ward Deng
                   ` (8 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Luis Ponce de Leao @ 1998-07-15 23:22 UTC (permalink / raw)
  To: ultralinux


Sun multi-processors systems and SGI Challenge are *SMP* systems,
SGI Onyx and Origin as well as Cray T3D/E are distributed
memory systems, CC/NUMA (non-uniform memory access with cache
coherency) and SP2 is a message passing based systems, it's
simply a cluster of RS/6000 workstations connected by a very high speed 
network. 

> On Wed, 15 Jul 1998, Ward Deng wrote:
> 
> > What do you mean that?
> > 
> > IBM RS6000/SP2 is distributed memory system, or workstations linked with
> > external high-speed switch. Sun's MP systems are all SMP design. SGI 
> > Power Challenge or Onyx are SMP too. The newer SGI Origin is SMP with
> > special memory links -- ccNUMA. They are very different architectures
> > while SP nodes are no different from IBM's RS6000 workstations.

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (31 preceding siblings ...)
  1998-07-15 23:22 ` Luis Ponce de Leao
@ 1998-07-21 23:29 ` Ward Deng
  1998-07-25 23:34 ` Ward Deng
                   ` (7 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Ward Deng @ 1998-07-21 23:29 UTC (permalink / raw)
  To: ultralinux

On Thu, 16 Jul 1998, Luis Ponce de Leao wrote:

> Date: Thu, 16 Jul 1998 00:22:53 +0100 (WET DST)
> From: Luis Ponce de Leao <lmmdpl@camoes.rnl.ist.utl.pt>
> To: linux-smp@vger.rutgers.edu
> Cc: ultralinux@vger.rutgers.edu
> Subject: Re: Ultra AXmp
> 
> 
> Sun multi-processors systems and SGI Challenge are *SMP* systems,
> SGI Onyx and Origin as well as Cray T3D/E are distributed
  ^^^^^^^^

Onyx is an SMP system while Onyx-2 is the ccNUMA architecture. Onyx the
workstation version of Challege while Onyx-2 is workstation version of 
Origin. NUMA architectue is pretty much SMP nodes with local memory 
inerlinked together. 

T3D/E, Paragon, CM-5 and n-CUBE are in another category. Traditionally
they are called MPP (Massive Parallel Processor). They distributed 
memory systems with dedicated processor linkage.
 
> memory systems, CC/NUMA (non-uniform memory access with cache
> coherency) and SP2 is a message passing based systems, it's
> simply a cluster of RS/6000 workstations connected by a very high speed 
> network. 
> 
> > On Wed, 15 Jul 1998, Ward Deng wrote:
> > 
> > > What do you mean that?
> > > 
> > > IBM RS6000/SP2 is distributed memory system, or workstations linked with
> > > external high-speed switch. Sun's MP systems are all SMP design. SGI 
> > > Power Challenge or Onyx are SMP too. The newer SGI Origin is SMP with
> > > special memory links -- ccNUMA. They are very different architectures
> > > while SP nodes are no different from IBM's RS6000 workstations.
> 

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (32 preceding siblings ...)
  1998-07-21 23:29 ` Ward Deng
@ 1998-07-25 23:34 ` Ward Deng
  1998-07-26  0:02 ` Ward Fenton
                   ` (6 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Ward Deng @ 1998-07-25 23:34 UTC (permalink / raw)
  To: ultralinux

> 
> >From what I have seen on the various lists I belong to, I believe that
> 80-90% of parallelizable problems fall in the coarse-to-medium grain
> category that is optimally accessible by a beowulf-type architecture.
> 10-20% (with some overlap -- the distinctions are not sharp) are
> optimally accessible by garden variety "parallel supercomputers" like
> the ones discussed above, and 1-2% are accessible only by exotic
> systems like current generation Crays or homemade/custom/dedicated
> parallel systems.

My background is computational mechanics and I have been working in 
fluid dynamics and solid mechnics for many many years. I can say most of
the applications in these fields are not "multi-body" type, ie. Beowulf
type problems. If you guys are interested in the problems/applications,
look into the "blue book" or its website (http://www.hpcc.gov/pubs/blue98). 

It has been hard to communicate computer scientists (:-) since most of them
do not understand what PDE stands for. ;-) To put it in short, the difficulty
is from the view of level of discretization. Almost any problem that needs to 
be simulated numerically is impossible to start from particle level (atoms,
molecules...). Instead we use _continuum_ to discribe it. In mathmatical
formulae, we "smoothen" the physical domain and treat it "infinite dimension" 
(remember Newton's calculus...) then we partition the domain into "finite 
dimension" such as "cells," "grids" and "elements" so we can calculate 
the mathematical formulae to get the approximate solutions under certain
_boundary_ and/or _initial_ conditions. You can just imagine Beowulf problem
consists around 1 million particles of rigid bodies. How many particles 
are involved in one cubic-foot of soil in the foundation of your house? How do
you deal with the particles with irregular shapes with only statistical data
available? How about add plasticity (soil particle will change its shape)
and viscousity (moisture...)?

Of course, we dream someday we can use super-supercomputer to solve our
problems starting from discrete domain but it is not possible in a forseeable
future. Beowulf-type COTS systems are very exciting but the majority of 
numerical problems are not in that category. Searching for low-latency,
high-throughput inter-processor communication should be still listed high
on our agenda. At least this is my view.

Just add a little bit noise into this interesting discussion.

--ward deng

Ward Deng, Ph.D                   |  Kachina Technologies, Inc.
Vice President and COO            |  4708 Douglas MacArthur Road NE,
Tel: 505-888-5934                 |  Albuquerque NM 87110
Fax: 505-888-5902                 |  info@KachinaTech.COM
Email: Ward.Deng@KachinaTech.COM  |  URL: http://www.KachinaTech.COM

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (33 preceding siblings ...)
  1998-07-25 23:34 ` Ward Deng
@ 1998-07-26  0:02 ` Ward Fenton
  1998-07-26  0:45 ` Rich Martin
                   ` (5 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Ward Fenton @ 1998-07-26  0:02 UTC (permalink / raw)
  To: ultralinux

a small comment about the bluebook...

http://www.ccic.gov/pubs/blue99/ should be up in a couple days...

another good reference is on its way in the form of the 
Presidential Advisory Committee for Computing, Information, and
Communications, Information Technology, and the Next Generation
Internet's "Interim report to the President".

I tend to agree with the comments made by Ward Deng.

On Sat, 25 Jul 1998, Ward Deng wrote:

> > 
> > >From what I have seen on the various lists I belong to, I believe that
> > 80-90% of parallelizable problems fall in the coarse-to-medium grain
> > category that is optimally accessible by a beowulf-type architecture.
> > 10-20% (with some overlap -- the distinctions are not sharp) are
> > optimally accessible by garden variety "parallel supercomputers" like
> > the ones discussed above, and 1-2% are accessible only by exotic
> > systems like current generation Crays or homemade/custom/dedicated
> > parallel systems.
> 
> My background is computational mechanics and I have been working in 
> fluid dynamics and solid mechnics for many many years. I can say most of
> the applications in these fields are not "multi-body" type, ie. Beowulf
> type problems. If you guys are interested in the problems/applications,
> look into the "blue book" or its website (http://www.hpcc.gov/pubs/blue98). 
> 
> It has been hard to communicate computer scientists (:-) since most of them
> do not understand what PDE stands for. ;-) To put it in short, the difficulty
> is from the view of level of discretization. Almost any problem that needs to 
> be simulated numerically is impossible to start from particle level (atoms,
> molecules...). Instead we use _continuum_ to discribe it. In mathmatical
> formulae, we "smoothen" the physical domain and treat it "infinite dimension" 
> (remember Newton's calculus...) then we partition the domain into "finite 
> dimension" such as "cells," "grids" and "elements" so we can calculate 
> the mathematical formulae to get the approximate solutions under certain
> _boundary_ and/or _initial_ conditions. You can just imagine Beowulf problem
> consists around 1 million particles of rigid bodies. How many particles 
> are involved in one cubic-foot of soil in the foundation of your house? How do
> you deal with the particles with irregular shapes with only statistical data
> available? How about add plasticity (soil particle will change its shape)
> and viscousity (moisture...)?
> 
> Of course, we dream someday we can use super-supercomputer to solve our
> problems starting from discrete domain but it is not possible in a forseeable
> future. Beowulf-type COTS systems are very exciting but the majority of 
> numerical problems are not in that category. Searching for low-latency,
> high-throughput inter-processor communication should be still listed high
> on our agenda. At least this is my view.
> 
> Just add a little bit noise into this interesting discussion.
> 
> --ward deng
> 
> Ward Deng, Ph.D                   |  Kachina Technologies, Inc.
> Vice President and COO            |  4708 Douglas MacArthur Road NE,
> Tel: 505-888-5934                 |  Albuquerque NM 87110
> Fax: 505-888-5902                 |  info@KachinaTech.COM
> Email: Ward.Deng@KachinaTech.COM  |  URL: http://www.KachinaTech.COM
> 

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (34 preceding siblings ...)
  1998-07-26  0:02 ` Ward Fenton
@ 1998-07-26  0:45 ` Rich Martin
  1998-07-26  3:34 ` Bob Drzyzgula
                   ` (4 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Rich Martin @ 1998-07-26  0:45 UTC (permalink / raw)
  To: ultralinux


I'm not sure this dicussion belongs on this list, but I'll provide an
alternate opinion to liven the debate. 

> It has been hard to communicate computer scientists (:-) since most of them
> do not understand what PDE stands for. ;-) To put it in short, the difficulty

Yup! All we get from the scientists is something like "please solve my 
Ax=b fast. Oh yea, I have a some funny structure in my A and b too". 

> numerical problems are not in that category. Searching for low-latency,
> high-throughput inter-processor communication should be still listed high
> on our agenda. At least this is my view.

Actually, anything that can be expressed in terms of large linear
algebra operations is a good candidate for a Beowulf style system. 
There has been endless work on MPPs and recently, clusters in this area.
Now there are plenty of good latency tolerating algorithms for most 
common operations. 

Dissecting the NAS Parallel Benchmarks on our cluster has convinced me that
these codes can run over a high bandwidth TCP just fine. On our 64 node
UltraSparc/Myrinet cluster, most of the NPB codes spent 5%-15% of the
time in communication. FT (a 3D FFT), the program that stresses the 
network the most, is just a bandwidth hog (high bandwidth alone is not an
interesting CS problem). FT, like all the other NPB codes, can tolerate 
huge software overheads and network latencies. 

Hardly worth investing time or $$$ in a communication network when you're
only spending 15% of the time there! If the NPB aren't representative,
well, that's another problem  ... 

I'm more and more convinced that inter-processor communication is no 
longer a first order issue. Focusing on memory system efficiency 
(effective use of the cache and memory BW) will get you much more bang for
your buck. Even in the parallel case, I would spend money on huge L2
caches and memory bandwidth over processor interconnect. 

My humble $0.02 worth... 

						-Rich

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (35 preceding siblings ...)
  1998-07-26  0:45 ` Rich Martin
@ 1998-07-26  3:34 ` Bob Drzyzgula
  1998-07-26  4:00 ` Ward Deng
                   ` (3 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Bob Drzyzgula @ 1998-07-26  3:34 UTC (permalink / raw)
  To: ultralinux

On Sat, Jul 25, 1998 at 05:45:53PM -0700, Rich Martin wrote:

> I'm more and more convinced that inter-processor communication is no 
> longer a first order issue. Focusing on memory system efficiency 
> (effective use of the cache and memory BW) will get you much more bang for
> your buck. Even in the parallel case, I would spend money on huge L2
> caches and memory bandwidth over processor interconnect. 

(The following is more an amplification than an argument...)

In nearly every iterative or repetitive algorithm, it
is usually possible to farm out individual iterations to
loosely-coupled machines. Many of the problems that crop
up in econometric simulations (my particular concern)
are highly iterative or repetitive in nature. For example,
it is common to design an analytical model and then throw
thousands of real or made-up sets of initial conditions,
assumptions and/or parameters at it; in the end one hopes
to judge the quality of the model by these thousands of
outcomes, often by contrasting them to empirical data. It
seems to me that, as long as a workload is made up of
a large number of independent calculations which can
be mapped out at an initialization stage, then a Beowulf
approach can easily be applied. However, if you have a very
large, single problem that must be calculated in a single
pass, then the Beowulf approach may well be throttled by
the the interconnect, as Ward Deng has suggested.

I enthusiastically agree with the assessment that memory
bandwidth is a major contributing factor to the performance
of a Beowulf, in our case (I work at the Federal Reserve
Board), this extends beyond clusters and into the design
of our general computing facility. We have found that,
even with the Suns, it is not cost-effective for us
to implement machines with greater than two to four
processors. The biggest trouble we run into is that so
many of our tasks saturate the memory channel that it is
impractical to time-share beyond two or three users on
*any* Sun machine, even a gigaplane machine such as an
Ultra 3000. Since our budget does not allow for buying
a hundred gigaplane machines, we must resort to smaller
machines, which in any event seem to work about as well for
our stuff (keeping in mind that much of it is written in
commercial packages such as Mathematica, Matlab, S-plus,
SAS and Troll, and as such it is not always possible to
take advantage of parallelizing compilers).

I am also more interested in memory bandwidth than L2
cache, although at 4MB the cache probably begins to be
helpful. Some of our tasks will grab 100MB or more of
main memory and spin through it repeatedly for weeks
or months at a time, thus rendering the cache largely
irrelevant. The memory bandwidth in the UPA is a step
in the right direction for us, although it is still the
case that several processors all squeezing memory access
through a single memory port can be a drag (the AXmp has a
single 144-bit -- 128 bits of data and 16 bits of ECC --
UPA port for memory; the EDO memory is accessed 576 bits at
a time, but then it is multiplexed through XB9 crossbars
into the single UPA port). Multiple memory ports as on
the gigaplane help, but in the end it seems kind of silly
spending tens of thousands of dollars to get a machine and
then to tune your workload to make it work like a cluster;
I'd still rather have a pile of single or dual processor
machines; I can afford a lot more processors that way,
for one thing.  The biggest problem I have here is
physical space, although I saw a nice 2U chassis at
Linux Expo (DCG Computers), and the manufacturer is working
on modifying it for the AXi motherboard. Too bad the
AXi's memory channel is choked off to Pentium II levels.

Speaking of space, has anyone ever tried using PICMG
single-board computers for performance computing? Some
of these things are starting to look pretty interesting,
for example it is nominally possible to put four of
these: http://www.dtims.com/lbc8522pic.htm onto one
of these: http://www.dtims.com/pbp4x4pic.htm, potentially
resulting in eight 400MHz Pentium IIs in a single 4U
rackmount chassis, or eighty in a single 70" high
19" machine rack. You'd be pretty limited on disk,
it isn't clear that the backplane's power distribution
can handle all those Pentium IIs, and I know those
boards aren't cheap, but they seem pretty interesting.
Until recently, a lot of these kinds of boards used
Realtek 10/100 or just "10" Ethernet chips, which
are no great bandwidth demons, but then these new
DTI boards are using Intel 82558 chips. 

Then again, perhaps something like this:
http://www.jump.de/potd/mops5.jpg. One ought to be
able to get a few hundred of *those* into a single
rack, although the per-node memory is pretty limited :-)

--Bob

-- 
==============================
Bob Drzyzgula                             It's not a problem
bob@drzyzgula.org                until something bad happens
==============================

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (36 preceding siblings ...)
  1998-07-26  3:34 ` Bob Drzyzgula
@ 1998-07-26  4:00 ` Ward Deng
  1998-07-26  4:29 ` Ward Deng
                   ` (2 subsequent siblings)
  40 siblings, 0 replies; 42+ messages in thread
From: Ward Deng @ 1998-07-26  4:00 UTC (permalink / raw)
  To: ultralinux

> 
> Hardly worth investing time or $$$ in a communication network when you're
> only spending 15% of the time there! If the NPB aren't representative,
> well, that's another problem  ... 

This is one side of the view. However, because of the heterogeniety of the
memory allocation, you simply leave the burden to programmers. There are
some algorithms are invented, which is good news. Instead of letting
programmers learn new computer architectures one by one before they know
how to write their programs :-).

I just try to remind people that the current status of Beowulf type system
is not panacea for the majority of computational challenges. This is why
CM-5 users and Cray users are more productive than IBM SP2 users. We still
have long way to go. Also most bandwidth tolerent algorithms have not been
invented (it also take time and $$$) if ever in computations a little bit
more sophisticated than Ax=b or FFT problems. Remember some algorithms
require users to make the dimension of their matrix A dividable by the
number of processors ;-(.

> 
> I'm more and more convinced that inter-processor communication is no 
> longer a first order issue. Focusing on memory system efficiency 
> (effective use of the cache and memory BW) will get you much more bang for
> your buck. Even in the parallel case, I would spend money on huge L2
> caches and memory bandwidth over processor interconnect. 
>

I believe someone feel that way but I have not convinced yet since starting
working on the first generation of IBM SP system. I totally agree that
high system throughput and SMP is very important and welcome. However, lower
lantency inter-processor communication is also important. It is not that
binary.

> My humble $0.02 worth... 
> 

Me too.

--ward

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (37 preceding siblings ...)
  1998-07-26  4:00 ` Ward Deng
@ 1998-07-26  4:29 ` Ward Deng
  1998-07-26 12:29 ` Bob Drzyzgula
  1998-07-26 13:57 ` Douglas Eadline
  40 siblings, 0 replies; 42+ messages in thread
From: Ward Deng @ 1998-07-26  4:29 UTC (permalink / raw)
  To: ultralinux

> 
> I am also more interested in memory bandwidth than L2
> cache, although at 4MB the cache probably begins to be
> helpful. Some of our tasks will grab 100MB or more of
> main memory and spin through it repeatedly for weeks
> or months at a time, thus rendering the cache largely
> irrelevant. The memory bandwidth in the UPA is a step
> in the right direction for us, although it is still the
> case that several processors all squeezing memory access
> through a single memory port can be a drag (the AXmp has a
> single 144-bit -- 128 bits of data and 16 bits of ECC --
> UPA port for memory; the EDO memory is accessed 576 bits at
> a time, but then it is multiplexed through XB9 crossbars
> into the single UPA port). Multiple memory ports as on
> the gigaplane help, but in the end it seems kind of silly
> spending tens of thousands of dollars to get a machine and
> then to tune your workload to make it work like a cluster;
> I'd still rather have a pile of single or dual processor
> machines; I can afford a lot more processors that way,
> for one thing.  The biggest problem I have here is
> physical space, although I saw a nice 2U chassis at
> Linux Expo (DCG Computers), and the manufacturer is working
> on modifying it for the AXi motherboard. Too bad the
> AXi's memory channel is choked off to Pentium II levels.

[snip]

On my datasheets, AXi (Panther) has 144-bit memory path while Pentium II
is only 64-bit. AX (Photon) has 288-bit full-scale UPA. AXmp (Chrico)
has two UPAs with 576-bit memory path. In our tests, even AXi has much
better scalability than PII.

We experienced the same problem as Bob did. A large-scale computational
chemisty program ran dog-slow once the problem scaled up to practical
level on an Alpha-PC system while it ran blazing fast during development
phase. The reason: system throughput was too bad. HAL system was a big
improvement then until UltraSPARC came out. Most real computing jobs won't
fit in your SRAM cache no matter how big you make it.

This is why we are so interested in UltraSPARC architecture and try to
make it available to Linux users. We first started Pentium Pro and got
very disappointed. Then we found OEM Alpha PC (DEC intended for NT users)
was not a good candidate either. We hope UltraSPARC nodes make into
"Beowulf" clusters soon. Pricewise, I think UltraSPARC is very reasonable
comparing with Intel Xeon. Maybe SME should provide a version similar to
Intel Celeron (joke) just for marketing purpose.

--ward

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (38 preceding siblings ...)
  1998-07-26  4:29 ` Ward Deng
@ 1998-07-26 12:29 ` Bob Drzyzgula
  1998-07-26 13:57 ` Douglas Eadline
  40 siblings, 0 replies; 42+ messages in thread
From: Bob Drzyzgula @ 1998-07-26 12:29 UTC (permalink / raw)
  To: ultralinux

On Sat, Jul 25, 1998 at 10:29:46PM -0600, Ward Deng wrote:
>
> On my datasheets, AXi (Panther) has 144-bit memory path while Pentium II
> is only 64-bit. AX (Photon) has 288-bit full-scale UPA. AXmp (Chrico)
> has two UPAs with 576-bit memory path. In our tests, even AXi has much
> better scalability than PII.

We have four of the AXi boards and have similarly
been very happy with them. Nonetheless, the
UltraSPARC IIi chip only has a 72-bit (64/8) window
on the memory subsystem. If you'll take a look at
http://www.sun.com/microelectronics/UltraSPARC-IIi/specs.html,
you'll see that the data lines feeding back to the
processor from the XCVR chip is only 72 bits wide.
On the AXi the path from the memory to the XCVR is 144
bits so that they can assure a full bandwidth feed to the
processor using commodity 60ns buffered EDO or FPM DIMMs.
I think that this is one reason why the AXi does so much
better than the PII on memory access; while both chips
have an 8-byte path to memory, the AXi is designed to keep
it filled. Note that, for example, you can usually run a
PII board with a single DIMM, while the AXi needs them in
multiples of two.

[The specs on the AXi claim only 400MBps for memory
throughput with a 300MHz processor. Since the path is
8 bytes wide, I concluded that the memory interface on
the 300MHz module runs at 50MHz, meaning that the 300MHz
processor is clock-sextupled. Does anyone know if this is
wrong? The block diagram shows the memory interface as
being runnable at 83MHz, I'd guess that the 333MHz chip
is clock-quadrupled?]

In the AXmp, the 576-bit channel is multiplexed down
to 144 bits into the processor; you can see this on
page 3-3 of the OEM manual:
http://www.sun.com/microelectronics/SPARCengineUltraAXmp/805-5865.pdf
(the block diagram on the Web is almost illegible). 
Still, the AXmp has two such 144-bit channels,
each channel shared between up to two processors.
When these 16-byte busses are set to run at 100MHz,
they result in a 1.6GBps memory channel. Running
flat out with four processors, this thing should
be able to pump upward of 3GBps out of memory, which
is to me just an amazing number for something selling
at this price. [Does anyone know if the 360MHz
processor will be running at 120MHz tripled on
the AXmp as it does on the Ultra 60? If so, this
would crank the memory channels up to 1.9GBps]
We have four four-processor AXmps on order. I'm
looking forward to them...

You can see from this that the AXi and the AXmp are
designed with the same kind of proportions. The
AXi processor has a 8-byte-wide channel, so that the
the memory is accessed in 16-byte chunks to assure
a steady feed from commodity memory. In the AXmp,
there are two 16-byte-wide channels, requiring
a full-speed 32-byte-wide feed. Thus the commodity
memory is accessed in 64-byte chunks and multiplexed
down.

> Maybe SME should provide a version similar to
> Intel Celeron (joke) just for marketing purpose.

I suppose that, comparing the the UltraSPARC II
to the UltraSPARC IIi, it could be argued that
the IIi is SME's Celeron-equivalent. It just
happens to be a whole lot less embarrasing. :-)

--Bob

-- 
==============================
Bob Drzyzgula                             It's not a problem
bob@drzyzgula.org                until something bad happens
==============================

^ permalink raw reply	[flat|nested] 42+ messages in thread

* Re: Ultra AXmp
  1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
                   ` (39 preceding siblings ...)
  1998-07-26 12:29 ` Bob Drzyzgula
@ 1998-07-26 13:57 ` Douglas Eadline
  40 siblings, 0 replies; 42+ messages in thread
From: Douglas Eadline @ 1998-07-26 13:57 UTC (permalink / raw)
  To: ultralinux

On Sat, 25 Jul 1998, Ward Deng wrote:
> 
> Of course, we dream someday we can use super-supercomputer to solve our
> problems starting from discrete domain but it is not possible in a forseeable
> future. Beowulf-type COTS systems are very exciting but the majority of 
> numerical problems are not in that category. Searching for low-latency,
> high-throughput inter-processor communication should be still listed high
> on our agenda. At least this is my view.

I think the point is that there are many problems that do fit the
"beowulf" model that previously were very expensive to solve. 
There will always be a case for the very high end (read as expensive)
machines.  But, a class of hardware that was previously unavailable
to many people just a few years ago is now very much in reach. 

We (Paralogic) are helping customers answer the question "will
a beowulf cluster work for me?"  And it all depends on
the application (problem) ....

Doug

-------------------------------------------------------------------
Paralogic, Inc.           |     PEAK     |      Voice:+610.861.6960
115 Research Drive        |   PARALLEL   |        Fax:+610.861.8247
Bethlehem, PA 18017 USA   |  PERFORMANCE |    http://www.plogic.com
-------------------------------------------------------------------

^ permalink raw reply	[flat|nested] 42+ messages in thread

end of thread, other threads:[~1998-07-26 13:57 UTC | newest]

Thread overview: 42+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
1998-07-11 18:36 Ultra AXmp Bob Drzyzgula
1998-07-12  1:56 ` David S. Miller
1998-07-12  2:21 ` Matthew Jacob
1998-07-12  2:29 ` David S. Miller
1998-07-12  5:02 ` Matthew Jacob
1998-07-12 12:54 ` Douglas Eadline
1998-07-12 23:59 ` Bob Drzyzgula
1998-07-13  6:27 ` Qiru Zhou
1998-07-13  6:38 ` David S. Miller
1998-07-13 10:40 ` Bob Drzyzgula
1998-07-13 12:09 ` Xavier Beaudouin
1998-07-13 13:59 ` Robert G. Brown
1998-07-13 14:29 ` Robert HYATT
1998-07-13 14:54 ` Matthew Jacob
1998-07-13 14:56 ` Matthew Jacob
1998-07-13 14:59 ` Matti Aarnio
1998-07-13 15:24 ` Robert HYATT
1998-07-13 15:38 ` Robert HYATT
1998-07-13 16:56 ` Robert G. Brown
1998-07-13 17:04 ` Robert G. Brown
1998-07-13 17:13 ` Lawrence D. Lopez
1998-07-13 17:26 ` Mike
1998-07-14  1:44 ` Shannon
1998-07-14  7:00 ` David S. Miller
1998-07-14 18:07 ` Douglas Eadline
1998-07-14 23:23 ` Bob Drzyzgula
1998-07-15  3:31 ` Dave Wreski
1998-07-15  8:07 ` Ward Deng
1998-07-15 11:14 ` Bob Drzyzgula
1998-07-15 11:27 ` David S. Miller
1998-07-15 15:54 ` Robert G. Brown
1998-07-15 17:11 ` Robert G. Brown
1998-07-15 23:22 ` Luis Ponce de Leao
1998-07-21 23:29 ` Ward Deng
1998-07-25 23:34 ` Ward Deng
1998-07-26  0:02 ` Ward Fenton
1998-07-26  0:45 ` Rich Martin
1998-07-26  3:34 ` Bob Drzyzgula
1998-07-26  4:00 ` Ward Deng
1998-07-26  4:29 ` Ward Deng
1998-07-26 12:29 ` Bob Drzyzgula
1998-07-26 13:57 ` Douglas Eadline

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.