From: Orit Wasserman <owasserm@redhat.com>
To: Peter Lieven <pl@kamp.de>
Cc: Stefan Hajnoczi <stefanha@gmail.com>,
Paolo Bonzini <pbonzini@redhat.com>,
qemu-devel@nongnu.org, quintela@redhat.com
Subject: Re: [Qemu-devel] [PATCHv4 2/9] cutils: add a function to find non-zero content in a buffer
Date: Mon, 25 Mar 2013 11:26:58 +0200 [thread overview]
Message-ID: <51501862.8000909@redhat.com> (raw)
In-Reply-To: <9AC1619C-476D-4EE4-B79B-34F2052313CD@kamp.de>
On 03/25/2013 10:56 AM, Peter Lieven wrote:
>
> Am 25.03.2013 um 09:53 schrieb Orit Wasserman <owasserm@redhat.com>:
>
>> On 03/22/2013 02:46 PM, Peter Lieven wrote:
>>> this adds buffer_find_nonzero_offset() which is a SSE2/Altivec
>>> optimized function that searches for non-zero content in a
>>> buffer.
>>>
>>> due to the optimizations used in the function there are restrictions
>>> on buffer address and search length. the function
>>> can_use_buffer_find_nonzero_content() can be used to check if
>>> the function can be used safely.
>>>
>>> Signed-off-by: Peter Lieven <pl@kamp.de>
>>> ---
>>> include/qemu-common.h | 13 +++++++++++++
>>> util/cutils.c | 45 +++++++++++++++++++++++++++++++++++++++++++++
>>> 2 files changed, 58 insertions(+)
>>>
>>> diff --git a/include/qemu-common.h b/include/qemu-common.h
>>> index e76ade3..078e535 100644
>>> --- a/include/qemu-common.h
>>> +++ b/include/qemu-common.h
>>> @@ -472,4 +472,17 @@ void hexdump(const char *buf, FILE *fp, const char *prefix, size_t size);
>>> #define ALL_EQ(v1, v2) ((v1) == (v2))
>>> #endif
>>>
>>> +#define BUFFER_FIND_NONZERO_OFFSET_UNROLL_FACTOR 8
>>> +static inline bool
>>> +can_use_buffer_find_nonzero_offset(const void *buf, size_t len)
>>> +{
>>> + if (len % (BUFFER_FIND_NONZERO_OFFSET_UNROLL_FACTOR
>>> + * sizeof(VECTYPE)) == 0
>>> + && ((uintptr_t) buf) % sizeof(VECTYPE) == 0) {
>>> + return true;
>>> + }
>>> + return false;
>>> +}
>>> +size_t buffer_find_nonzero_offset(const void *buf, size_t len);
>>> +
>>> #endif
>>> diff --git a/util/cutils.c b/util/cutils.c
>>> index 1439da4..41c627e 100644
>>> --- a/util/cutils.c
>>> +++ b/util/cutils.c
>>> @@ -143,6 +143,51 @@ int qemu_fdatasync(int fd)
>>> }
>>>
>>> /*
>>> + * Searches for an area with non-zero content in a buffer
>>> + *
>>> + * Attention! The len must be a multiple of
>>> + * BUFFER_FIND_NONZERO_OFFSET_UNROLL_FACTOR * sizeof(VECTYPE)
>>> + * and addr must be a multiple of sizeof(VECTYPE) due to
>>> + * restriction of optimizations in this function.
>>> + *
>>> + * can_use_buffer_find_nonzero_offset() can be used to check
>>> + * these requirements.
>>> + *
>>> + * The return value is the offset of the non-zero area rounded
>>> + * down to BUFFER_FIND_NONZERO_OFFSET_UNROLL_FACTOR * sizeof(VECTYPE).
>>> + * If the buffer is all zero the return value is equal to len.
>>> + */
>>> +
>>> +size_t buffer_find_nonzero_offset(const void *buf, size_t len)
>>> +{
>>> + VECTYPE *p = (VECTYPE *)buf;
>>> + VECTYPE zero = ZERO_SPLAT;
>>> + size_t i;
>>> +
>>> + assert(len % (BUFFER_FIND_NONZERO_OFFSET_UNROLL_FACTOR
>>> + * sizeof(VECTYPE)) == 0);
>>> + assert(((uintptr_t) buf) % sizeof(VECTYPE) == 0);
>>> +
>>> + if (*((const long *) buf)) {
>>> + return 0;
>>> + }
>>> +
>>> + for (i = 0; i < len / sizeof(VECTYPE);
>> Why not put len/sizeof(VECTYPE) in a variable?
>
> are you afraid that there is a division at each iteration?
>
> sizeof(VECTYPE) is a power of 2 so i think the compiler will optimize it
> to a >> at compile time.
true, but it still is done every iteration.
>
> I would also be ok with writing len /= sizeof(VECTYPE) before the loop.
I would prefer it :)
Orit
>
> Peter
>
>> Orit
>>> + i += BUFFER_FIND_NONZERO_OFFSET_UNROLL_FACTOR) {
>>> + VECTYPE tmp0 = p[i + 0] | p[i + 1];
>>> + VECTYPE tmp1 = p[i + 2] | p[i + 3];
>>> + VECTYPE tmp2 = p[i + 4] | p[i + 5];
>>> + VECTYPE tmp3 = p[i + 6] | p[i + 7];
>>> + VECTYPE tmp01 = tmp0 | tmp1;
>>> + VECTYPE tmp23 = tmp2 | tmp3;
>>> + if (!ALL_EQ(tmp01 | tmp23, zero)) {
>>> + break;
>>> + }
>>> + }
>>> + return i * sizeof(VECTYPE);
>>> +}
>>> +
>>> +/*
>>> * Checks if a buffer is all zeroes
>>> *
>>> * Attention! The len must be a multiple of 4 * sizeof(long) due to
>>>
>>
>
next prev parent reply other threads:[~2013-03-25 9:25 UTC|newest]
Thread overview: 44+ messages / expand[flat|nested] mbox.gz Atom feed top
2013-03-22 12:46 [Qemu-devel] [PATCHv4 0/9] buffer_is_zero / migration optimizations Peter Lieven
2013-03-22 12:46 ` [Qemu-devel] [PATCHv4 1/9] move vector definitions to qemu-common.h Peter Lieven
2013-03-25 8:35 ` Orit Wasserman
2013-03-22 12:46 ` [Qemu-devel] [PATCHv4 2/9] cutils: add a function to find non-zero content in a buffer Peter Lieven
2013-03-22 19:37 ` Eric Blake
2013-03-22 20:03 ` Peter Lieven
2013-03-22 20:22 ` [Qemu-devel] indentation hints [was: [PATCHv4 2/9] cutils: add a function to find non-zero content in a buffer] Eric Blake
2013-03-23 11:18 ` Peter Maydell
2013-03-25 8:53 ` [Qemu-devel] [PATCHv4 2/9] cutils: add a function to find non-zero content in a buffer Orit Wasserman
2013-03-25 8:56 ` Peter Lieven
2013-03-25 9:26 ` Orit Wasserman [this message]
2013-03-25 9:42 ` Paolo Bonzini
2013-03-25 10:03 ` Orit Wasserman
2013-03-22 12:46 ` [Qemu-devel] [PATCHv4 3/9] buffer_is_zero: use vector optimizations if possible Peter Lieven
2013-03-25 8:53 ` Orit Wasserman
2013-03-22 12:46 ` [Qemu-devel] [PATCHv4 4/9] bitops: use vector algorithm to optimize find_next_bit() Peter Lieven
2013-03-25 9:04 ` Orit Wasserman
2013-03-22 12:46 ` [Qemu-devel] [PATCHv4 5/9] migration: search for zero instead of dup pages Peter Lieven
2013-03-22 19:49 ` Eric Blake
2013-03-22 20:02 ` Peter Lieven
2013-03-25 9:30 ` Orit Wasserman
2013-03-22 12:46 ` [Qemu-devel] [PATCHv4 6/9] migration: add an indicator for bulk state of ram migration Peter Lieven
2013-03-25 9:32 ` Orit Wasserman
2013-03-22 12:46 ` [Qemu-devel] [PATCHv4 7/9] migration: do not sent zero pages in bulk stage Peter Lieven
2013-03-22 20:13 ` Eric Blake
2013-03-25 9:44 ` Orit Wasserman
2013-03-22 12:46 ` [Qemu-devel] [PATCHv4 8/9] migration: do not search dirty " Peter Lieven
2013-03-25 10:05 ` Orit Wasserman
2013-03-22 12:46 ` [Qemu-devel] [PATCHv4 9/9] migration: use XBZRLE only after " Peter Lieven
2013-03-25 10:16 ` Orit Wasserman
2013-03-22 17:25 ` [Qemu-devel] [PATCHv4 0/9] buffer_is_zero / migration optimizations Paolo Bonzini
2013-03-22 19:20 ` Peter Lieven
2013-03-22 21:24 ` Paolo Bonzini
2013-03-23 7:34 ` Peter Lieven
2013-03-25 10:17 ` Peter Lieven
2013-03-25 10:53 ` Paolo Bonzini
2013-03-25 11:26 ` Peter Lieven
2013-03-25 13:02 ` Paolo Bonzini
2013-03-25 13:23 ` Peter Lieven
2013-03-25 13:32 ` Peter Lieven
2013-03-25 14:34 ` Paolo Bonzini
2013-03-25 21:37 ` Peter Lieven
2013-03-26 8:14 ` Peter Lieven
2013-03-26 9:20 ` Paolo Bonzini
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=51501862.8000909@redhat.com \
--to=owasserm@redhat.com \
--cc=pbonzini@redhat.com \
--cc=pl@kamp.de \
--cc=qemu-devel@nongnu.org \
--cc=quintela@redhat.com \
--cc=stefanha@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).