From: Guoqing Jiang <guoqing.jiang@cloud.ionos.com>
To: Yufen Yu <yuyufen@huawei.com>, song@kernel.org
Cc: linux-raid@vger.kernel.org, neilb@suse.com, colyli@suse.de,
xni@redhat.com, houtao1@huawei.com
Subject: Re: [PATCH v3 00/11] md/raid5: set STRIPE_SIZE as a configurable value
Date: Fri, 29 May 2020 00:07:24 +0200 [thread overview]
Message-ID: <f0ab9a2d-696b-b21f-faec-370cd7c3ed3a@cloud.ionos.com> (raw)
In-Reply-To: <20200527131933.34400-1-yuyufen@huawei.com>
On 5/27/20 3:19 PM, Yufen Yu wrote:
> Hi, all
>
> For now, STRIPE_SIZE is equal to the value of PAGE_SIZE. That means, RAID5 will
> issus echo bio to disk at least 64KB when PAGE_SIZE is 64KB in arm64. However,
> filesystem usually issue bio in the unit of 4KB. Then, RAID5 will waste resource
> of disk bandwidth.
Could you explain a little bit about "waste resource"? Does it mean the
chance for
full stripe write is limited because of the incompatible between fs
(4KB bio) and
raid5 (64KB stripe unit)?
> To solve the problem, this patchset provide a new config CONFIG_MD_RAID456_STRIPE_SHIFT
> to let user config STRIPE_SIZE. The default value is 1, means 4096(1<<9).
>
> Normally, using default STRIPE_SIZE can get better performance. And NeilBrown have
> suggested just to fix the STRIPE_SIZE as 4096.But, out test result show that
> big value of STRIPE_SIZE may have better performance when size of issued IOs are
> mostly bigger than 4096. Thus, in this patchset, we still want to set STRIPE_SIZE
> as a configureable value.
I think it is better to define stripe size as 4K if it fits to generally
scenario, and also
aligns with fs.
> In current implementation, grow_buffers() uses alloc_page() to allocate the buffers
> for each stripe_head. With the change, it means we allocate 64K buffers but just
> use 4K of them. To save memory, we try to 'compress' multiple buffers of stripe_head
> to only one real page. Detail shows in patch #2.
>
> To evaluate the new feature, we create raid5 device '/dev/md5' with 4 SSD disk
> and test it on arm64 machine with 64KB PAGE_SIZE.
>
> 1) We format /dev/md5 with mkfs.ext4 and mount ext4 with default configure on
> /mnt directory. Then, trying to test it by dbench with command:
> dbench -D /mnt -t 1000 10. Result show as:
>
> 'STRIPE_SHIFT = 64KB'
>
> Operation Count AvgLat MaxLat
> ----------------------------------------
> NTCreateX 9805011 0.021 64.728
> Close 7202525 0.001 0.120
> Rename 415213 0.051 44.681
> Unlink 1980066 0.079 93.147
> Deltree 240 1.793 6.516
> Mkdir 120 0.004 0.007
> Qpathinfo 8887512 0.007 37.114
> Qfileinfo 1557262 0.001 0.030
> Qfsinfo 1629582 0.012 0.152
> Sfileinfo 798756 0.040 57.641
> Find 3436004 0.019 57.782
> WriteX 4887239 0.021 57.638
> ReadX 15370483 0.005 37.818
> LockX 31934 0.003 0.022
> UnlockX 31933 0.001 0.021
> Flush 687205 13.302 530.088
>
> Throughput 307.799 MB/sec 10 clients 10 procs max_latency=530.091 ms
> -------------------------------------------------------
>
> 'STRIPE_SIZE = 4KB'
>
> Operation Count AvgLat MaxLat
> ----------------------------------------
> NTCreateX 11999166 0.021 36.380
> Close 8814128 0.001 0.122
> Rename 508113 0.051 29.169
> Unlink 2423242 0.070 38.141
> Deltree 300 1.885 7.155
> Mkdir 150 0.004 0.006
> Qpathinfo 10875921 0.007 35.485
> Qfileinfo 1905837 0.001 0.032
> Qfsinfo 1994304 0.012 0.125
> Sfileinfo 977450 0.029 26.489
> Find 4204952 0.019 9.361
> WriteX 5981890 0.019 27.804
> ReadX 18809742 0.004 33.491
> LockX 39074 0.003 0.025
> UnlockX 39074 0.001 0.014
> Flush 841022 10.712 458.848
>
> Throughput 376.777 MB/sec 10 clients 10 procs max_latency=458.852 ms
> -------------------------------------------------------
What is the default io unit size of dbench?
> 2) We try to evaluate IO throughput for /dev/md5 by fio with config:
>
> [4KB randwrite]
> direct=1
> numjob=2
> iodepth=64
> ioengine=libaio
> filename=/dev/md5
> bs=4KB
> rw=randwrite
>
> [64KB write]
> direct=1
> numjob=2
> iodepth=64
> ioengine=libaio
> filename=/dev/md5
> bs=1MB
> rw=write
>
> The fio test result as follow:
>
> + +
> | STRIPE_SIZE(64KB) | STRIPE_SIZE(4KB)
> +----------------------------------------------------+
> 4KB randwrite | 15MB/s | 100MB/s
> +----------------------------------------------------+
> 1MB write | 1000MB/s | 700MB/s
>
> The result show that when size of io is bigger than 4KB (64KB),
> 64KB STRIPE_SIZE has much higher IOPS. But for 4KB randwrite, that
> means, size of io issued to device are smaller, 4KB STRIPE_SIZE
> have better performance.
The 4k rand write performance drops from 100MB/S to 15MB/S?! How about other
io sizes? Say 16k, 64K and 256K etc, it would be more convincing if 64KB
stripe
has better performance than 4KB stripe overall.
Thanks,
Guoqing
next prev parent reply other threads:[~2020-05-28 22:07 UTC|newest]
Thread overview: 30+ messages / expand[flat|nested] mbox.gz Atom feed top
2020-05-27 13:19 [PATCH v3 00/11] md/raid5: set STRIPE_SIZE as a configurable value Yufen Yu
2020-05-27 13:19 ` [PATCH v3 01/11] md/raid5: add CONFIG_MD_RAID456_STRIPE_SHIFT to set STRIPE_SIZE Yufen Yu
2020-05-27 13:54 ` Guoqing Jiang
2020-05-27 23:30 ` John Stoffel
2020-05-28 6:17 ` Yufen Yu
2020-05-27 15:16 ` Xiao Ni
2020-05-28 6:29 ` Yufen Yu
2020-05-27 20:21 ` kbuild test robot
2020-05-28 14:23 ` Song Liu
2020-05-29 8:42 ` Yufen Yu
2020-05-27 13:19 ` [PATCH v3 02/11] md/raid5: add a member of r5pages for struct stripe_head Yufen Yu
2020-05-27 13:19 ` [PATCH v3 03/11] md/raid5: allocate and free pages of r5pages Yufen Yu
2020-05-27 13:19 ` [PATCH v3 04/11] md/raid5: set correct page offset for bi_io_vec in ops_run_io() Yufen Yu
2020-05-27 13:19 ` [PATCH v3 05/11] md/raid5: set correct page offset for async_copy_data() Yufen Yu
2020-05-27 13:19 ` [PATCH v3 06/11] md/raid5: add new xor function to support different page offset Yufen Yu
2020-05-27 13:19 ` [PATCH v3 07/11] md/raid5: add offset array in scribble buffer Yufen Yu
2020-05-27 13:19 ` [PATCH v3 08/11] md/raid5: compute xor with correct page offset Yufen Yu
2020-05-27 13:19 ` [PATCH v3 09/11] md/raid6: let syndrome computor support different " Yufen Yu
2020-05-27 13:19 ` [PATCH v3 10/11] md/raid6: compute syndrome with correct " Yufen Yu
2020-05-27 13:19 ` [PATCH v3 11/11] raid6test: adaptation with syndrome function Yufen Yu
2020-05-28 14:10 ` [PATCH v3 00/11] md/raid5: set STRIPE_SIZE as a configurable value Song Liu
2020-05-28 14:28 ` Song Liu
2020-05-29 9:32 ` Yufen Yu
2020-05-28 22:07 ` Guoqing Jiang [this message]
2020-05-29 11:49 ` Yufen Yu
2020-05-29 12:22 ` Guoqing Jiang
2020-05-30 2:15 ` Yufen Yu
2020-06-01 14:02 ` Guoqing Jiang
2020-06-02 6:59 ` Song Liu
2020-06-04 13:17 ` Yufen Yu
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=f0ab9a2d-696b-b21f-faec-370cd7c3ed3a@cloud.ionos.com \
--to=guoqing.jiang@cloud.ionos.com \
--cc=colyli@suse.de \
--cc=houtao1@huawei.com \
--cc=linux-raid@vger.kernel.org \
--cc=neilb@suse.com \
--cc=song@kernel.org \
--cc=xni@redhat.com \
--cc=yuyufen@huawei.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox