From: "Benoît Canet" <benoit@irqsave.net>
To: qemu-devel@nongnu.org
Cc: kwolf@redhat.com, pbonzini@redhat.com,
"Benoît Canet" <benoit@irqsave.net>,
stefanha@redhat.com
Subject: [Qemu-devel] [RFC V5 00/62] QCOW2 deduplication
Date: Wed, 16 Jan 2013 16:47:39 +0100 [thread overview]
Message-ID: <1358351321-4891-1-git-send-email-benoit@irqsave.net> (raw)
This 3 step patchset implements deduplication in QCOW2.
First patchset create the core infrastructure for deduplication and enable it
in QCOW2 image.
It ends at "qcow2: Enable the deduplication feature."
Second patchset implements some metrics in QMP.
It ends at "qapi: Return virtual block device deduplication metrics in QMP"
Third patchset implements asynchronous deduplication.
It's a work in progress patchset that is included in this post so reviewers
can have a grasp of where the feature is heading.
One can compile and install https://github.com/wernerd/Skein3Fish and use the
--enable-skein-dedup configure option in order to use the faster skein HASH.
Images must be created with "-o dedup=[skein|sha256]" in order to activate the
deduplication in the image.
Deduplication is now fast enough to be usable.
Nice side effect is that duplicated writes are faster than native QCOW2:
v5:
Move qemu-io-test dedup patch [Eric]
Reserve some room at the end of the QCOW header extensions. [Eric]
Fix the specification. [Eric]
Now overflow deduplication refcount at 2^16/2 [Stefan]
Implements metrics.
Implement asynchronous deduplication.
Increase L2 table size and deduplication block hash size.
Random cleanups
v4: Fix and complete qcow2 spec [Stefan]
Hash the hash_algo field in the header extension [Stefan]
Fix qcow2 spec [Eric]
Remove pointer to hash and simplify hash memory management [Stefan]
Rename and move qcow2_read_cluster_data to qcow2.c [Stefan]
Document lock dropping behaviour of the previous function [Stefan]
cleanup qcow2_dedup_read_missing_cluster_data [Stefan]
rename *_offset to *_sect [Stefan]
add a ./configure check for ssl [Stefan]
Replace openssl by gnutls [Stefan]
Implement Skein hashes
Rewrite pretty every qcow2-dedup.c commits after Add
qcow2_dedup_read_missing_and_concatenate to simplify the code
Use 64KB deduplication hash block to reduce allocation flushes
Use 64KB l2 tables to reduce allocation flushes [breaks compatibility]
Use lazy refcounts to avoid qcow2_cache_set_dependency loops resultings
in frequent caches flushes
Do not create and load dedup RAM structures when bdrs->read_only is true
v3: make it work barely
replace kernel red black trees by gtree.
Benoît Canet (62):
qcow2: Add deduplication to the qcow2 specification.
qcow2: Add deduplication structures and fields.
qcow2: Add qcow2_dedup_read_missing_and_concatenate
qcow2: Make update_refcount public.
qcow2: Create a way to link to l2 tables when deduplicating.
qcow2: Add qcow2_dedup and related functions
qcow2: Add qcow2_dedup_store_new_hashes.
qcow2: Implement qcow2_compute_cluster_hash.
qcow2: Extract qcow2_dedup_grow_table
qcow2: Add qcow2_dedup_grow_table and use it.
qcow2: Makes qcow2_alloc_cluster_link_l2 mark to deduplicate
clusters.
qcow2: make the deduplication forget a cluster hash when a cluster is
to dedupe
qcow2: Create qcow2_is_cluster_to_dedup.
qcow2: Load and save deduplication table header extension.
qcow2: Extract qcow2_do_table_init.
qcow2-cache: Allow to choose table size at creation.
qcow2: Extract qcow2_add_feature and qcow2_remove_feature.
block: Add qemu-img dedup create option.
qcow2: Add a deduplication boolean to update_refcount.
qcow2: Drop hash for a given cluster when dedup makes refcount >
2^16/2.
qcow2: Remove hash when cluster is deleted.
qcow2: Add qcow2_dedup_is_running to probe if dedup is running.
qcow2: Integrate deduplication in qcow2_co_writev loop.
qcow2: Serialize write requests when deduplication is activated.
qcow2: Add verification of dedup table.
qcow2: Adapt checking of QCOW_OFLAG_COPIED for dedup.
qcow2: Add check_dedup_l2 in order to check l2 of dedup table.
qcow2: Do not overwrite existing entries with QCOW_OFLAG_COPIED.
qcow2: Integrate SKEIN hash algorithm in deduplication.
qcow2: Add lazy refcounts to deduplication to prevent
qcow2_cache_set_dependency loops
qcow2: Use large L2 table for deduplication.
qcow: Set large dedup hash block size.
qemu-iotests: Filter dedup=on/off so existing tests don't break.
qcow2: Add qcow2_dedup_init and qcow2_dedup_close.
qcow2: Add qcow2_co_dedup_resume to restart deduplication.
qcow2: Enable the deduplication feature.
qcow2: Add deduplication metrics structures.
qcow2: Initialize deduplication metrics.
qcow2: Collect unaligned writes missing data reads metric.
qcow2: Collect deduplicated cluster metric.
qcow2: Collect undeduplicated cluster metric.
qcow2: Count QCowHashNode creation metrics.
qcow2: Count QCowHashNode removal from tree for metrics.
qcow2: Count cluster deleted metric
qcow2: Count deduplication refcount overflow metric.
qapi: Add support for deduplication infos in qapi-schema.json.
block: Add deduplication metrics to BlockDriverInfo.
qcow2: Add qcow2_dedup_update_metrics to compute dedup RAM usage.
qcow2: returns deduplication metrics and status via bdrv_get_info()
qapi: Return virtual block device deduplication metrics in QMP
block: Add BlockDriver function prototype to pause and resume
deduplication.
qcow2: Add code to deduplicate cluster flagged with
QCOW_OFLAG_TO_DEDUP.
block: Add bdrv_has_dedup.
block: Add bdrv_is_dedup_running.
block: Add bdrv_resume_dedup.
block: Add bdrv_pause_dedup.
qcow2: Add qcow2_pause_dedup.
qcow2: Add qcow2_resume_dedup.
qcow2: Make dedup status persists.
qerror: Add QERR_DEVICE_NOT_DEDUPLICATED.
qmp: Add block-pause-dedup.
qmp: Add block_resume_dedup.
block.c | 108 +++
block/Makefile.objs | 1 +
block/qcow2-cache.c | 12 +-
block/qcow2-cluster.c | 182 ++++--
block/qcow2-dedup.c | 1492 ++++++++++++++++++++++++++++++++++++++++++
block/qcow2-refcount.c | 175 +++--
block/qcow2.c | 378 +++++++++--
block/qcow2.h | 149 ++++-
blockdev.c | 36 +
configure | 55 ++
docs/specs/qcow2.txt | 104 ++-
include/block/block.h | 18 +
include/block/block_int.h | 5 +
include/qapi/qmp/qerror.h | 3 +
qapi-schema.json | 76 ++-
qmp-commands.hx | 46 ++
tests/qemu-iotests/common.rc | 3 +-
17 files changed, 2708 insertions(+), 135 deletions(-)
create mode 100644 block/qcow2-dedup.c
--
1.7.10.4
next reply other threads:[~2013-01-16 15:48 UTC|newest]
Thread overview: 67+ messages / expand[flat|nested] mbox.gz Atom feed top
2013-01-16 15:47 Benoît Canet [this message]
2013-01-16 15:47 ` [Qemu-devel] [RFC V5 01/62] qcow2: Add deduplication to the qcow2 specification Benoît Canet
2013-01-16 16:43 ` Eric Blake
2013-01-16 15:47 ` [Qemu-devel] [RFC V5 02/62] qcow2: Add deduplication structures and fields Benoît Canet
2013-01-16 16:30 ` Eric Blake
2013-01-16 15:47 ` [Qemu-devel] [RFC V5 03/62] qcow2: Add qcow2_dedup_read_missing_and_concatenate Benoît Canet
2013-01-16 15:47 ` [Qemu-devel] [RFC V5 04/62] qcow2: Make update_refcount public Benoît Canet
2013-01-16 15:47 ` [Qemu-devel] [RFC V5 05/62] qcow2: Create a way to link to l2 tables when deduplicating Benoît Canet
2013-01-16 15:47 ` [Qemu-devel] [RFC V5 06/62] qcow2: Add qcow2_dedup and related functions Benoît Canet
2013-01-16 15:47 ` [Qemu-devel] [RFC V5 07/62] qcow2: Add qcow2_dedup_store_new_hashes Benoît Canet
2013-01-16 15:47 ` [Qemu-devel] [RFC V5 08/62] qcow2: Implement qcow2_compute_cluster_hash Benoît Canet
2013-01-16 15:47 ` [Qemu-devel] [RFC V5 09/62] qcow2: Extract qcow2_dedup_grow_table Benoît Canet
2013-01-16 15:47 ` [Qemu-devel] [RFC V5 10/62] qcow2: Add qcow2_dedup_grow_table and use it Benoît Canet
2013-01-16 15:47 ` [Qemu-devel] [RFC V5 11/62] qcow2: Makes qcow2_alloc_cluster_link_l2 mark to deduplicate clusters Benoît Canet
2013-01-16 15:47 ` [Qemu-devel] [RFC V5 12/62] qcow2: make the deduplication forget a cluster hash when a cluster is to dedupe Benoît Canet
2013-01-16 15:47 ` [Qemu-devel] [RFC V5 13/62] qcow2: Create qcow2_is_cluster_to_dedup Benoît Canet
2013-01-16 15:47 ` [Qemu-devel] [RFC V5 14/62] qcow2: Load and save deduplication table header extension Benoît Canet
2013-01-16 15:47 ` [Qemu-devel] [RFC V5 15/62] qcow2: Extract qcow2_do_table_init Benoît Canet
2013-01-16 15:47 ` [Qemu-devel] [RFC V5 16/62] qcow2-cache: Allow to choose table size at creation Benoît Canet
2013-01-16 15:47 ` [Qemu-devel] [RFC V5 17/62] qcow2: Extract qcow2_add_feature and qcow2_remove_feature Benoît Canet
2013-01-16 15:47 ` [Qemu-devel] [RFC V5 18/62] block: Add qemu-img dedup create option Benoît Canet
2013-01-16 15:47 ` [Qemu-devel] [RFC V5 19/62] qcow2: Add a deduplication boolean to update_refcount Benoît Canet
2013-01-16 15:47 ` [Qemu-devel] [RFC V5 20/62] qcow2: Drop hash for a given cluster when dedup makes refcount > 2^16/2 Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 21/62] qcow2: Remove hash when cluster is deleted Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 22/62] qcow2: Add qcow2_dedup_is_running to probe if dedup is running Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 23/62] qcow2: Integrate deduplication in qcow2_co_writev loop Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 24/62] qcow2: Serialize write requests when deduplication is activated Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 25/62] qcow2: Add verification of dedup table Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 26/62] qcow2: Adapt checking of QCOW_OFLAG_COPIED for dedup Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 27/62] qcow2: Add check_dedup_l2 in order to check l2 of dedup table Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 28/62] qcow2: Do not overwrite existing entries with QCOW_OFLAG_COPIED Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 29/62] qcow2: Integrate SKEIN hash algorithm in deduplication Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 30/62] qcow2: Add lazy refcounts to deduplication to prevent qcow2_cache_set_dependency loops Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 31/62] qcow2: Use large L2 table for deduplication Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 32/62] qcow: Set large dedup hash block size Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 33/62] qemu-iotests: Filter dedup=on/off so existing tests don't break Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 34/62] qcow2: Add qcow2_dedup_init and qcow2_dedup_close Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 35/62] qcow2: Add qcow2_co_dedup_resume to restart deduplication Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 36/62] qcow2: Enable the deduplication feature Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 37/62] qcow2: Add deduplication metrics structures Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 38/62] qcow2: Initialize deduplication metrics Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 39/62] qcow2: Collect unaligned writes missing data reads metric Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 40/62] qcow2: Collect deduplicated cluster metric Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 41/62] qcow2: Collect undeduplicated " Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 42/62] qcow2: Count QCowHashNode creation metrics Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 43/62] qcow2: Count QCowHashNode removal from tree for metrics Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 44/62] qcow2: Count cluster deleted metric Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 45/62] qcow2: Count deduplication refcount overflow metric Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 46/62] qapi: Add support for deduplication infos in qapi-schema.json Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 47/62] block: Add deduplication metrics to BlockDriverInfo Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 48/62] qcow2: Add qcow2_dedup_update_metrics to compute dedup RAM usage Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 49/62] qcow2: returns deduplication metrics and status via bdrv_get_info() Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 50/62] qapi: Return virtual block device deduplication metrics in QMP Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 51/62] block: Add BlockDriver function prototype to pause and resume deduplication Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 52/62] qcow2: Add code to deduplicate cluster flagged with QCOW_OFLAG_TO_DEDUP Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 53/62] block: Add bdrv_has_dedup Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 54/62] block: Add bdrv_is_dedup_running Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 55/62] block: Add bdrv_resume_dedup Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 56/62] block: Add bdrv_pause_dedup Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 57/62] qcow2: Add qcow2_pause_dedup Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 58/62] qcow2: Add qcow2_resume_dedup Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 59/62] qcow2: Make dedup status persists Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 60/62] qerror: Add QERR_DEVICE_NOT_DEDUPLICATED Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 61/62] qmp: Add block-pause-dedup Benoît Canet
2013-01-16 15:48 ` [Qemu-devel] [RFC V5 62/62] qmp: Add block_resume_dedup Benoît Canet
2013-01-16 16:03 ` [Qemu-devel] [RFC V5 00/62] QCOW2 deduplication Eric Blake
2013-01-16 16:26 ` Benoît Canet
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=1358351321-4891-1-git-send-email-benoit@irqsave.net \
--to=benoit@irqsave.net \
--cc=kwolf@redhat.com \
--cc=pbonzini@redhat.com \
--cc=qemu-devel@nongnu.org \
--cc=stefanha@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).