From: Ryan Ding <ryan.ding@oracle.com>
To: ocfs2-devel@oss.oracle.com
Subject: [Ocfs2-devel] [PATCH 0/8] ocfs2: fix ocfs2 direct io code patch to support sparse file and data ordering semantics
Date: Thu, 08 Oct 2015 15:13:15 +0800 [thread overview]
Message-ID: <5616178B.9060508@oracle.com> (raw)
In-Reply-To: <561609A0.5010108@huawei.com>
Hi Joseph,
On 10/08/2015 02:13 PM, Joseph Qi wrote:
> Hi Ryan,
>
> On 2015/10/8 11:12, Ryan Ding wrote:
>> Hi Joseph,
>>
>> On 09/28/2015 06:20 PM, Joseph Qi wrote:
>>> Hi Ryan,
>>> I have gone through this patch set and done a simple performance test
>>> using direct dd, it indeed brings much performance promotion.
>>> Before After
>>> bs=4K 1.4 MB/s 5.0 MB/s
>>> bs=256k 40.5 MB/s 56.3 MB/s
>>>
>>> My questions are:
>>> 1) You solution is still using orphan dir to keep inode and allocation
>>> consistency, am I right? From our test, it is the most complicated part
>>> and has many race cases to be taken consideration. So I wonder if this
>>> can be restructured.
>> I have not got a better idea to do this. I think the only reason why direct io using orphan is to prevent space lost when system crash during append direct write. But maybe a 'fsck -f' will do that job. Is it necessary to use orphan?
> The idea is taken from ext4, but since ocfs2 is cluster filesystem, so
> it is much more complicated than ext4.
> And fsck can only be used offline, but using orphan is to perform
> recovering online. So I don't think fsck can replace it in all cases.
OK, I agree.
>
>>> 2) Rather than using normal block direct io, you introduce a way to use
>>> write begin/end in buffer io. IMO, if it wants to perform like direct
>>> io, it should be committed to disk by forcing committing journal. But
>>> journal committing will consume much time. Why does it bring performance
>>> promotion instead?
>> I use buffer io to write only the zero pages. Actual data payload is written as direct io. I think there is no need to do a force commit. Because direct means "Try to minimize cache effects of the I/O to and from this file.", it does not means "write all data & meta data to disk before write return".
> So this is protected by "UNWRITTEN" flag, right?
Yes.
>
>>> 3) Do you have a test in case of lack of memory?
>> I tested it in a system with 2GB memory. Is that enough?
> What I mean is doing many direct io jobs in case system free memory is
> low.
You use dio or aio+dio?
Thanks,
Ryan
>
> Thanks,
> Joesph
>
>> Thanks,
>> Ryan
>>> On 2015/9/11 16:19, Ryan Ding wrote:
>>>> The idea is to use buffer io(more precisely use the interface
>>>> ocfs2_write_begin_nolock & ocfs2_write_end_nolock) to do the zero work beyond
>>>> block size. And clear UNWRITTEN flag until direct io data has been written to
>>>> disk, which can prevent data corruption when system crashed during direct write.
>>>>
>>>> And we will also archive a better performance:
>>>> eg. dd direct write new file with block size 4KB:
>>>> before this patch:
>>>> 2.5 MB/s
>>>> after this patch:
>>>> 66.4 MB/s
>>>>
>>>> ----------------------------------------------------------------
>>>> Ryan Ding (8):
>>>> ocfs2: add ocfs2_write_type_t type to identify the caller of write
>>>> ocfs2: use c_new to indicate newly allocated extents
>>>> ocfs2: test target page before change it
>>>> ocfs2: do not change i_size in write_end for direct io
>>>> ocfs2: return the physical address in ocfs2_write_cluster
>>>> ocfs2: record UNWRITTEN extents when populate write desc
>>>> ocfs2: fix sparse file & data ordering issue in direct io.
>>>> ocfs2: code clean up for direct io
>>>>
>>>> fs/ocfs2/aops.c | 1118 +++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++------------------------------------------------------------------------------------------
>>>> fs/ocfs2/aops.h | 11 +-
>>>> fs/ocfs2/file.c | 138 +---------------------
>>>> fs/ocfs2/inode.c | 3 +
>>>> fs/ocfs2/inode.h | 3 +
>>>> fs/ocfs2/mmap.c | 4 +-
>>>> fs/ocfs2/ocfs2_trace.h | 16 +--
>>>> fs/ocfs2/super.c | 1 +
>>>> 8 files changed, 568 insertions(+), 726 deletions(-)
>>>>
>>>> .
>>>>
>>
>> .
>>
>
next prev parent reply other threads:[~2015-10-08 7:13 UTC|newest]
Thread overview: 22+ messages / expand[flat|nested] mbox.gz Atom feed top
2015-09-11 8:19 [Ocfs2-devel] [PATCH 0/8] ocfs2: fix ocfs2 direct io code patch to support sparse file and data ordering semantics Ryan Ding
2015-09-11 8:19 ` [Ocfs2-devel] [PATCH 1/8] ocfs2: add ocfs2_write_type_t type to identify the caller of write Ryan Ding
2015-09-11 8:19 ` [Ocfs2-devel] [PATCH 2/8] ocfs2: use c_new to indicate newly allocated extents Ryan Ding
2015-09-11 8:19 ` [Ocfs2-devel] [PATCH 3/8] ocfs2: test target page before change it Ryan Ding
2015-09-11 8:19 ` [Ocfs2-devel] [PATCH 4/8] ocfs2: do not change i_size in write_end for direct io Ryan Ding
2015-09-11 8:19 ` [Ocfs2-devel] [PATCH 5/8] ocfs2: return the physical address in ocfs2_write_cluster Ryan Ding
2015-09-11 8:19 ` [Ocfs2-devel] [PATCH 6/8] ocfs2: record UNWRITTEN extents when populate write desc Ryan Ding
2015-09-11 8:19 ` [Ocfs2-devel] [PATCH 7/8] ocfs2: fix sparse file & data ordering issue in direct io Ryan Ding
2015-09-11 8:19 ` [Ocfs2-devel] [PATCH 8/8] ocfs2: code clean up for " Ryan Ding
2015-09-28 10:20 ` [Ocfs2-devel] [PATCH 0/8] ocfs2: fix ocfs2 direct io code patch to support sparse file and data ordering semantics Joseph Qi
2015-10-08 3:12 ` Ryan Ding
2015-10-08 6:13 ` Joseph Qi
2015-10-08 7:13 ` Ryan Ding [this message]
2015-10-12 6:34 ` Ryan Ding
2015-12-10 7:54 ` Joseph Qi
2015-12-10 8:48 ` Ryan Ding
2015-12-10 10:36 ` Joseph Qi
2015-12-14 5:31 ` Ryan Ding
2015-12-14 10:36 ` Joseph Qi
2015-12-16 1:39 ` Ryan Ding
2015-12-16 2:26 ` Joseph Qi
2015-12-16 3:12 ` Ryan Ding
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=5616178B.9060508@oracle.com \
--to=ryan.ding@oracle.com \
--cc=ocfs2-devel@oss.oracle.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox