From mboxrd@z Thu Jan 1 00:00:00 1970 From: Jeremy Fitzhardinge Subject: Re: [Xen-devel] [PATCH 00/10] [PATCH RFC V2] Paravirtualized ticketlocks Date: Mon, 10 Oct 2011 12:51:29 -0700 Message-ID: <4E934CC1.9040804@goop.org> References: <201109282008.17722.stephan.diestelhorst@amd.com> <2707952.s3VYcmPHUN@chlor> <4E8DE7F1.3050108@goop.org> <4E8DEED0.1020909@goop.org> <20111010073214.GB29035@elte.hu> Mime-Version: 1.0 Content-Type: text/plain; charset=ISO-8859-1 Content-Transfer-Encoding: 7bit Cc: Stephan Diestelhorst , Linus Torvalds , "H. Peter Anvin" , Jan Beulich , Jeremy Fitzhardinge , Andi Kleen , Peter Zijlstra , Nick Piggin , the arch/x86 maintainers , "xen-devel@lists.xensource.com" , Avi Kivity , Marcelo Tosatti , KVM , Linux Kernel Mailing List , Konrad Rzeszutek Wilk To: Ingo Molnar Return-path: In-Reply-To: <20111010073214.GB29035@elte.hu> Sender: linux-kernel-owner@vger.kernel.org List-Id: kvm.vger.kernel.org On 10/10/2011 12:32 AM, Ingo Molnar wrote: > * Jeremy Fitzhardinge wrote: > >> On 10/06/2011 10:40 AM, Jeremy Fitzhardinge wrote: >>> However, it looks like locked xadd is also has better performance: on >>> my Sandybridge laptop (2 cores, 4 threads), the add+mfence is 20% slower >>> than locked xadd, so that pretty much settles it unless you think >>> there'd be a dramatic difference on an AMD system. >> Konrad measures add+mfence is about 65% slower on AMD Phenom as well. > xadd also results in smaller/tighter code, right? Not particularly, mostly because of the overflow-into-the-high-part compensation. But its only a couple of extra instructions, and no conditionals, so I don't think it would have any concrete effect. But, as Stephen points out, perhaps locked add is preferable to locked xadd, since it also has the same barrier as mfence but has (significantly!) better performance than either mfence or locked xadd... J