Git development
 help / color / mirror / Atom feed
From: "Pablo Sabater" <pabloosabaterr@gmail.com>
To: "Derrick Stolee" <stolee@gmail.com>,
	"Pablo Sabater" <pabloosabaterr@gmail.com>, <git@vger.kernel.org>
Subject: Re: [PATCH RFC 0/5] Add --dry-run option to git-backfill(1)
Date: Wed, 30 Sep 2026 20:06:21 +0100	[thread overview]
Message-ID: <DLSVWT1QTELK.17U19NG169AW9@gmail.com> (raw)
In-Reply-To: <ed1b9048-d438-4143-a224-fa0e28d4fd42@gmail.com>

On Wed Sep 30, 2026 at 7:14 PM WEST, Derrick Stolee wrote:
> On 9/29/2026 8:21 PM, Pablo Sabater wrote:
>> [Cc'd Derrick Stolee for his work in the backfill(1) command]
>> 
>> This series adds a --dry-run option to git-backfill(1) that reports how
>> many missing blobs would be fetched and, when the remote server
>> supports the object-info capability, their total size:
>> 
>>         $ git backfill --dry-run
>>         After backfill, 48 blobs would be fetched (1.20 KiB).
>> 
>> If the server does not advertise object-info, only the count is shown.
>
> This is a helpful capability, but I'm not sure the size check counts
> as a "dry run" because it involves a network call (and possibly many
> depending on --min-batch-size).
>
> Perhaps a different argument would be better, such as --info=(count|size)
> to make it clear what level of information you want to know in advance
> and thus how much effort are you willing to put in to discover this. 

Makes sense to have it as an --info option.

>> I am not a git-backfill(1) user myself, but it seemed useful for users
>> to know how much data a backfill would bring in before running it.
>
> I'm not sure that we want to add a feature based on speculation. Git
> is a collection of "itches" that the contributors needed scratched.
> The work is motivated by real needs.
>
> While I can see some benefit to curiosity, I'm not sure how much this
> would prevent users from making their decision as to whether they
> should run backfill or not.
>
>> The number of missing blobs is the sum of the number of blobs to be
>> fetched in each batch. The object-info capability lets us ask the server
>> for the size of each blob without downloading it, so summing them gives
>> an estimate of the total.
>> 
>> Note that this is an upper bound rather than the exact disk usage:
>> object-info reports the uncompressed size of each object, while the
>> objects end up stored compressed and possibly deltified in a packfile,
>> so the space actually used on disk will usually be smaller.
>
> I don't think the uncompressed size is a useful metric here, as it is
> likely astronomically larger than what will be downloaded. How will
> this help a user make a decision?

Yes, that's one of the itches I have with the object-info protocol: it
cannot give you a reliable compressed size, and that's why only the
total size is supported.

The object-info protocol could be extended to support the
objectsize:disk attribute, either by having the server know what we
already have and the oids that we want, or by directly having the
server send us the compressed size of its local copy (to avoid too much
work).  Even with the first option, because it goes in batches, it
would still be an estimate, just a closer one.

Given that I don't use backfill, I Cc'd you because I wasn't sure if it
was really useful, and the main motivation was the "I'm going to check
--dry-run before backfilling" case, so it helps to decide.  If it's not
that helpful and seems to end up as a decoration option, it might be
better to drop it.

>
> Thanks,
> -Stolee

Thanks for taking a look,
Pablo


  reply	other threads:[~2026-09-30 19:06 UTC|newest]

Thread overview: 18+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-30  0:21 [PATCH RFC 0/5] Add --dry-run option to git-backfill(1) Pablo Sabater
2026-09-30  0:21 ` [PATCH RFC 1/5] transport-internal: update fetch_object_info comment Pablo Sabater
2026-09-30  0:21 ` [PATCH RFC 2/5] fetch-object-info: add enum for fetch_object_info() statuses Pablo Sabater
2026-09-30 10:56   ` Karthik Nayak
2026-09-30 11:59     ` Pablo Sabater
2026-09-30  0:21 ` [PATCH RFC 3/5] fetch-object-info: return a status instead of dying Pablo Sabater
2026-09-30 17:07   ` Junio C Hamano
2026-09-30 18:03     ` Pablo Sabater
2026-09-30 20:01       ` Junio C Hamano
2026-09-30  0:21 ` [PATCH RFC 4/5] backfill: add --dry-run option Pablo Sabater
2026-09-30 11:06   ` Karthik Nayak
2026-09-30 12:19     ` Pablo Sabater
2026-09-30  0:21 ` [PATCH RFC 5/5] backfill: report total size of missing blobs in --dry-run Pablo Sabater
2026-09-30 17:17   ` Junio C Hamano
2026-09-30 18:31     ` Pablo Sabater
2026-09-30 18:14 ` [PATCH RFC 0/5] Add --dry-run option to git-backfill(1) Derrick Stolee
2026-09-30 19:06   ` Pablo Sabater [this message]
2026-09-30 20:03   ` Junio C Hamano

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=DLSVWT1QTELK.17U19NG169AW9@gmail.com \
    --to=pabloosabaterr@gmail.com \
    --cc=git@vger.kernel.org \
    --cc=stolee@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox