linux-um archives
 help / color / mirror / Atom feed
* [uml-devel] Re: [SYSEMU] New benchmarks results
       [not found] <200406121616.26839.blaisorblade_spam@yahoo.it>
@ 2004-06-13 21:50 ` Laurent Vivier
  2004-06-15  3:54   ` Jeff Dike
  0 siblings, 1 reply; 2+ messages in thread
From: Laurent Vivier @ 2004-06-13 21:50 UTC (permalink / raw)
  To: BlaisorBlade; +Cc: user-mode-linux-devel, roland

[-- Attachment #1: Type: text/plain, Size: 8160 bytes --]

Le sam 12/06/2004 à 16:16, BlaisorBlade a écrit :
> I've decided to do benchmarks to check how much SYSEMU saves in benchmark 
> which also access memory (memLoop.c) and how much could save the 0 context 
> switch idea (provided that segmentation has low cost).
> 
> First, about the benchmark on the Laurent Vivier page: I think that the "60 %" 
> number is meaningless - I guess it is that calculated with "real time", which 
> is not very meaningful IMHO - that is the time from when the process start to 
> when it ends, and counts even time spent by executing other processes. A more 
> meaningful difference is done with the sum of user+system time:
> 
> average time (user+system):
> - without SYSEMU 
> 64.910
> - with SYSEMU
> 51.321
> 
> SYSEMU saves (64.910 - 51.321) / 64.910 * 100 % = 20,9 % of the time without 
> SYSEMU, in this benchmark.

Hello Paolo,

thank you for your comments.

the real question is: how accurate is the command "time" under UML ?

I choose "real" time for several reasons:

- what user feels is the most important (how many time he waits ?)
- I don't really know how is computed "sys" time under UML: is it host +
guest "sys" time ? How "time" takes into account the time of the process
"ptracing" the user process and, thus, the sys time of the guest kernel
?

I made my measurements on a 8 cpus Xeon server with several gigabytes of
memory, with no load and only one user: me.

Host:

real               0m7.920s
user+sys           0m7.930s
real - (user+sys) -0m0.010s 
(mmhhh, a negative value, there is really no load ;-) )

So I didn't really explain the measurements I had :

w/o SYSEMU:

real     6m16.956s 6m17.126s 6m16.461s
user+sys 1m03.712s 1m06.577s 1m04.442s

w/ SYSEMU:

real     3m55.052s 3m56.964s 3m54.179s
user+sys 0m52.347s 0m48.481s 0m53.135s

Could you explain where we lost :

w/o SYSEMU

real - (user+sys) 5m13.144s 5m10.549s 5m12.002s 

w/ SYSEMU

real - (user+sys) 3m02.705s 3m08.483s 3m01.044s

In the TLB flushes ? in the "ptracing" process ? in other processes ?

IMHO, I thought it's in guest kernel, so "real" is more significant than "user+sys".
BUT I think you're the real specialist of UML and I'm not...

> I've re-benchmarked UML with SYSEMU using memLoop.c which tries to measure the 
> effects of accessing memory: it access one byte per page, thus causing the 
> CPU to reload in the TLB the page table entry (PTE) for that page. IMHO, this 
> benchmark shows that most of the gap vs the host is in the 2 remaining CS per 
> syscall: the 2 we save with SYSEMU account for about 25% of the getpid 
> execution, most of the gap is still there.
> 
> In the attached files NPAGES = 64 (see source), but I also posted results with 
> NPAGES = 512. Also, please, don't look at the "elapsed" time: it's 
> meaningless.
> 
> In fact getpidLoop measures only the cost of TLB flushes, while memLoop also 
> measures the cost of TLB misses after the TLB flush, which can be compared 
> against memLoopPure, which runs no syscall and thus never flushes the TLBs.
> 
> To see this, I must be sure that memLoopPure has no TLB fault, i.e. that the 
> PTEs for all pages fit in the TLB; this happen when NPAGES = 64, not when 
> NPAGES=512. In the two cases, we have working sets of 64 * PAGE_SIZE = 128k 
> and of 512 * PAGE_SIZE = 2 M.
> 
> On the host, memLoop and memLoopPure have similar user time, since there is 
> never a TLB flush. When NPAGES = 512, each page access causes a TLB miss, so 
> the user time is always similar, both on the host and the guest, and both 
> with and without syscalls.
> 
> But when NPAGES = 64, on the host the TLB is never flushed (except when 
> another process is executing): it is filled only once and then used.
> 
> On the guest, instead, with NPAGES = 64 the user time of memLoop is double 
> than the memLoopPure one. And since 0.40 s are for the getpid() calls, 
> touch_mem() uses 0.40 s in memLoopPure and 1.20 s in memLoop: 3 times the old 
> time.
> --------
> HOST:
> 
> host $ time ./getpidLoop 1000000
> 
> 0.27user 0.21system 0:00.55elapsed 87%CPU (0avgtext+0avgdata 0maxresident)k
> 0inputs+0outputs (70major+11minor)pagefaults 0swaps
> --------
> With NPAGES = 64:
> 
> host $ time ./memLoop 1000000
> 
> 1.11user 0.23system 0:01.46elapsed 91%CPU (0avgtext+0avgdata 0maxresident)k
> 0inputs+0outputs (79major+75minor)pagefaults 0swaps
> ----
> host $ time ./memLoopPure 1000000
> 0.88user 0.00system 0:00.97elapsed 90%CPU (0avgtext+0avgdata 0maxresident)k
> 0inputs+0outputs (78major+75minor)pagefaults 0swaps
> --------
> With NPAGES = 512
> 
> host $ time ./memLoop 1000000
> 
> 8.93user 0.24system 0:09.84elapsed 93%CPU (0avgtext+0avgdata 0maxresident)k
> 0inputs+0outputs (79major+523minor)pagefaults 0swaps
> ----
> host $ time ./memLoopPure 1000000
> 
> 8.71user 0.01system 0:09.43elapsed 92%CPU (0avgtext+0avgdata 0maxresident)k
> 0inputs+0outputs (78major+523minor)pagefaults 0swaps
> 
> ------------
> On the guest, with SYSEMU:
> 
> guest # /usr/bin/time 
> /mnt/host/home/paolo/Dati/Sorgenti/Varie/C-C++/getpidLoop 1000000
> 
> 0.42user 3.87system 0:16.09elapsed 26%CPU (0avgtext+0avgdata 0maxresident)k
> 0inputs+0outputs (0major+76minor)pagefaults 0swaps
> --------
> With NPAGES = 64:
> ----
> guest # /usr/bin/time /mnt/host/home/paolo/Dati/Sorgenti/Varie/C-C++/memLoop 
> 1000000
> 
> 1.60user 4.00system 0:18.02elapsed 31%CPU (0avgtext+0avgdata 0maxresident)k
> 0inputs+0outputs (0major+146minor)pagefaults 0swaps
> ----
> guest # /usr/bin/time 
> /mnt/host/home/paolo/Dati/Sorgenti/Varie/C-C++/memLoopPure 1000000
> 
> 0.85user 0.05system 0:01.01elapsed 88%CPU (0avgtext+0avgdata 0maxresident)k
> 0inputs+0outputs (0major+146minor)pagefaults 0swaps
> --------
> With NPAGES = 512:
> 
> guest # /usr/bin/time /mnt/host/home/paolo/Dati/Sorgenti/Varie/C-C++/memLoop 
> 1000000
> 
> 9.09user 4.18system 0:28.37elapsed 46%CPU (0avgtext+0avgdata 0maxresident)k
> 0inputs+0outputs (0major+594minor)pagefaults 0swaps
> ----
> guest # /usr/bin/time 
> /mnt/host/home/paolo/Dati/Sorgenti/Varie/C-C++/memLoopPure 1000000
> 
> 8.76user 0.07system 0:11.57elapsed 76%CPU (0avgtext+0avgdata 0maxresident)k
> 0inputs+0outputs (0major+594minor)pagefaults 0swaps
> 
> ----------------
> On the guest, without SYSEMU:
> (we always about 25% increase for system time vs SYSEMU, except for 
> memLoopPure, but equal user time: we don't save the TLB misses)
> 
> # /usr/bin/time /mnt/host/home/paolo/Dati/Sorgenti/Varie/C-C++/getpidLoop 
> 1000000
> 0.42user 5.01system 0:21.08elapsed 25%CPU (0avgtext+0avgdata 0maxresident)k
> 0inputs+0outputs (0major+76minor)pagefaults 0swaps
> ----
> With NPAGES = 64:
> 
> guest # /usr/bin/time 
> /mnt/host/home/paolo/Dati/Sorgenti/Varie/C-C++/memLoopPure 1000000
> (about the same, as expected)
> 
> 0.86user 0.02system 0:00.94elapsed 92%CPU (0avgtext+0avgdata 0maxresident)k
> 0inputs+0outputs (0major+146minor)pagefaults 0swaps
> ----
> guest # /usr/bin/time /mnt/host/home/paolo/Dati/Sorgenti/Varie/C-C++/memLoop 
> 1000000
> (about 25% increase for system time, equal user time: we don't save the TLB 
> misses)
> 1.62user 5.00system 0:26.73elapsed 24%CPU (0avgtext+0avgdata 0maxresident)k
> 0inputs+0outputs (0major+146minor)pagefaults 0swaps
> 
> --------
> 
> With NPAGES = 512
> 
> guest # /usr/bin/time 
> /mnt/host/home/paolo/Dati/Sorgenti/Varie/C-C++/memLoopPure 1000000
> 
> 8.84user 0.02system 0:10.86elapsed 81%CPU (0avgtext+0avgdata 0maxresident)k
> 0inputs+0outputs (0major+594minor)pagefaults 0swaps
> 
> ----
> 
> guest # /usr/bin/time /mnt/host/home/paolo/Dati/Sorgenti/Varie/C-C++/memLoop 
> 1000000
> 9.15user 5.06system 0:36.66elapsed 38%CPU (0avgtext+0avgdata 0maxresident)k
> 0inputs+0outputs (0major+594minor)pagefaults 0swaps
-- 
                   Laurent Vivier
+------------------------------------------------+
     "Any sufficiently advanced technology is 
indistinguishable from magic." -- Arthur C. Clarke
   "Aller les Bleus" - France 2 - 1 Angleterre

[-- Attachment #2: Ceci est une partie de message numériquement signée. --]
[-- Type: application/pgp-signature, Size: 189 bytes --]

^ permalink raw reply	[flat|nested] 2+ messages in thread

* Re: [uml-devel] Re: [SYSEMU] New benchmarks results
  2004-06-13 21:50 ` [uml-devel] Re: [SYSEMU] New benchmarks results Laurent Vivier
@ 2004-06-15  3:54   ` Jeff Dike
  0 siblings, 0 replies; 2+ messages in thread
From: Jeff Dike @ 2004-06-15  3:54 UTC (permalink / raw)
  To: Laurent Vivier; +Cc: BlaisorBlade, user-mode-linux-devel, roland

LaurentVivier@wanadoo.fr said:
> Could you explain where we lost :
> w/o SYSEMU
> real - (user+sys) 5m13.144s 5m10.549s 5m12.002s 
> w/ SYSEMU
> real - (user+sys) 3m02.705s 3m08.483s 3m01.044s 

Somewhere in the host kernel.  Given this, it would seem that real is a better
measurement than user+sys on an unloaded system since user+sys leaves out the
host kernel.

What you really want is time on the host, except covering only the benchmark
period rather than the period from boot to halt.

				Jeff



-------------------------------------------------------
This SF.Net email is sponsored by The 2004 JavaOne(SM) Conference
Learn from the experts at JavaOne(SM), Sun's Worldwide Java Developer
Conference, June 28 - July 1 at the Moscone Center in San Francisco, CA
REGISTER AND SAVE! http://java.sun.com/javaone/sf Priority Code NWMGYKND
_______________________________________________
User-mode-linux-devel mailing list
User-mode-linux-devel@lists.sourceforge.net
https://lists.sourceforge.net/lists/listinfo/user-mode-linux-devel

^ permalink raw reply	[flat|nested] 2+ messages in thread

end of thread, other threads:[~2004-06-15  3:03 UTC | newest]

Thread overview: 2+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
     [not found] <200406121616.26839.blaisorblade_spam@yahoo.it>
2004-06-13 21:50 ` [uml-devel] Re: [SYSEMU] New benchmarks results Laurent Vivier
2004-06-15  3:54   ` Jeff Dike

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox