From: "Philip Oakley" <philipoakley@iee.org>
To: "Vitaly Arbuzov" <vit@uber.com>,
"Jeff Hostetler" <git@jeffhostetler.com>
Cc: "Git List" <git@vger.kernel.org>
Subject: Re: How hard would it be to implement sparse fetching/pulling?
Date: Sat, 2 Dec 2017 16:30:24 -0000 [thread overview]
Message-ID: <A2F606CCDE4542D8AAFC41EB50909FE3@PhilipOakley> (raw)
In-Reply-To: bac032c8-b9c2-4520-58e5-d518f4efd9d6@jeffhostetler.com
From: "Jeff Hostetler" <git@jeffhostetler.com>
Sent: Friday, December 01, 2017 2:30 PM
> On 11/30/2017 8:51 PM, Vitaly Arbuzov wrote:
>> I think it would be great if we high level agree on desired user
>> experience, so let me put a few possible use cases here.
>>
>> 1. Init and fetch into a new repo with a sparse list.
>> Preconditions: origin blah exists and has a lot of folders inside of
>> src including "bar".
>> Actions:
>> git init foo && cd foo
>> git config core.sparseAll true # New flag to activate all sparse
>> operations by default so you don't need to pass options to each
>> command.
>> echo "src/bar" > .git/info/sparse-checkout
>> git remote add origin blah
>> git pull origin master
>> Expected results: foo contains src/bar folder and nothing else,
>> objects that are unrelated to this tree are not fetched.
>> Notes: This should work same when fetch/merge/checkout operations are
>> used in the right order.
>
> With the current patches (parts 1,2,3) we can pass a blob-ish
> to the server during a clone that refers to a sparse-checkout
> specification.
I hadn't appreciated this capability. I see it as important, and should be
available both ways, so that a .gitNarrow spec can be imposed from the
server side, as well as by the requester.
It could also be used to assist in the 'precious/secret' blob problem, so
that AWS keys are never pushed, nor available for fetching!
> There's a bit of a chicken-n-egg problem getting
> things set up. So if we assume your team would create a series
> of "known enlistments" under version control, then you could
s/enlistments/entitlements/ I presume?
> just reference one by <branch>:<path> during your clone. The
> server can lookup that blob and just use it.
>
> git clone --filter=sparse:oid=master:templates/bar URL
>
> And then the server will filter-out the unwanted blobs during
> the clone. (The current version only filters blobs; you still
> get full commits and trees. That will be revisited later.)
I'm for the idea that only the in-heirachy trees should be sent.
It should also be possible that the server replies that it is only sending a
narrow clone, with the given (accessible?) spec.
>
> On the client side, the partial clone installs local config
> settings into the repo so that subsequent fetches default to
> the same filter criteria as used in the clone.
>
>
> I don't currently have provision to send a full sparse-checkout
> specification to the server during a clone or fetch. That
> seemed like too much to try to squeeze into the protocols.
> We can revisit this later if there is interest, but it wasn't
> critical for the initial phase.
>
Agreed. I think it should be somewhere 'visible' to the user, but could be
setup by the server admin / repo maintainer if they don't have write access.
But there could still be the catch-22 - maybe one starts with a <commit |
toptree> : <tree> pair to define an origin point (it's not as refined as a
.gitNarrow spec file, but is definative). The toptree option could even
allow sub-tree clones.. maybe..
>
>>
>> 2. Add a file and push changes.
>> Preconditions: all steps above followed.
>> touch src/bar/baz.txt && git add -A && git commit -m "added a file"
>> git push origin master
>> Expected results: changes are pushed to remote.
>
> I don't believe partial clone and/or partial fetch will cause
> any changes for push.
I suspect that pushes could be rejected if the user 'pretends' to modify
files or trees outside their area. It does need the user to be able to spoof
part of a tree they don't have, so an upstream / remote would immediatly
know it was a spoof but locally the narrow clone doesn't have enough detail
about the 'bad' oid. It would be right to reject such attempts!
>
>>
>> 3. Clone a repo with a sparse list as a filter.
>> Preconditions: same as for #1
>> Actions:
>> echo "src/bar" > /tmp/blah-sparse-checkout
>> git clone --sparse /tmp/blah-sparse-checkout blah # Clone should be
>> the only command that would requires specific option key being passed.
>> Expected results: same as for #1 plus /tmp/blah-sparse-checkout is
>> copied into .git/info/sparse-checkout
I presume clone and fetch are treated equivalently here.
>
> There are 2 independent concepts here: clone and checkout.
> Currently, there isn't any automatic linkage of the partial clone to
> the sparse-checkout settings, so you could do something like this:
>
I see an implicit link that clearly one cannot checkout (inflate/populate) a
file/directory that one does not have in the object store. But that does not
imply the reverse linkage. The regular sparse checkout should be available
independently of the local clone being a narrow one.
> git clone --no-checkout --filter=sparse:oid=master:templates/bar URL
> git cat-file ... templates/bar >.git/info/sparse-checkout
> git config core.sparsecheckout true
> git checkout ...
>
> I've been focused on the clone/fetch issues and have not looked
> into the automation to couple them.
>
I foresee that large files and certain files need to be filterable for
fetch-clone, and that might not be (backward) compatible with the
sparse-checkout.
>
>>
>> 4. Showing log for sparsely cloned repo.
>> Preconditions: #3 is followed
>> Actions:
>> git log
>> Expected results: recent changes that affect src/bar tree.
>
> If I understand your meaning, log would only show changes
> within the sparse subset of the tree. This is not on my
> radar for partial clone/fetch. It would be a nice feature
> to have, but I think it would be better to think about it
> from the point of view of sparse-checkout rather than clone.
>
One option maybe by making a marker for the tree/blob to be a first class
citizen. So the oid (and worktree file) has content ".gitNarrowTree <oid>"
or ",gitNarrowBlob <oid>" as required (*), which is safe, and allows a
consistent alter-ego view of the tree contents and hence for git-log et.al.
(*) I keep flip flopping between a single object marker, and distinct object
markers for the types. It partly depends on whether one can know in advance,
locally, what the oid type should be, and how it should be embedded in the
object store - need to re-check the specs.
I'm tending toward distinct types to cope with the D/F conflict in the
worktrees - the directory must be created (holds the name etc), and the
alter-ego content then must be placed in a _known_ sub-file ".gitNarrowTree"
(without the oid in the file name, but included in the content). Presence of
a ".gitNarrowTree" should be standalone in the directory when that part of
the work-tree is clean.
>
>>
>> 5. Showing diff.
>> Preconditions: #3 is followed
>> Actions:
>> git diff HEAD^ HEAD
>> Expected results: changes from the most recent commit affecting
>> src/bar folder are shown.
>> Notes: this can be tricky operation as filtering must be done to
>> remove results from unrelated subtrees.
>
> I don't have any plan for this and I don't think it fits within
> the scope of clone/fetch. I think this too would be a sparse-checkout
> feature.
>
See my note about first class citizens for marker OIDs
>
>>
>> *Note that I intentionally didn't mention use cases that are related
>> to filtering by blob size as I think we should logically consider them
>> as a separate, although related, feature.
>
> I've grouped blob-size and sparse filter together for the
> purposes of clone/fetch since the basic mechanisms (filtering,
> transport, and missing object handling) are the same for both.
> They do lead to different end-uses, but that is above my level
> here.
>
>
>>
>> What do you think about these examples above? Is that something that
>> more-or-less fits into current development? Are there other important
>> flows that I've missed?
>
> These are all good ideas and it is good to have someone else who
> wants to use partial+sparse thinking about it and looking for gaps
> as we try to make a complete end-to-end feature.
>>
>> -Vitaly
>
> Thanks
> Jeff
>
Philip
next prev parent reply other threads:[~2017-12-02 16:30 UTC|newest]
Thread overview: 26+ messages / expand[flat|nested] mbox.gz Atom feed top
2017-11-30 3:16 How hard would it be to implement sparse fetching/pulling? Vitaly Arbuzov
2017-11-30 14:24 ` Jeff Hostetler
2017-11-30 17:01 ` Vitaly Arbuzov
2017-11-30 17:44 ` Vitaly Arbuzov
2017-11-30 20:03 ` Jonathan Nieder
2017-12-01 16:03 ` Jeff Hostetler
2017-12-01 18:16 ` Jonathan Nieder
2017-11-30 23:43 ` Philip Oakley
2017-12-01 1:27 ` Vitaly Arbuzov
2017-12-01 1:51 ` Vitaly Arbuzov
2017-12-01 2:51 ` Jonathan Nieder
2017-12-01 3:37 ` Vitaly Arbuzov
2017-12-02 16:59 ` Philip Oakley
2017-12-01 14:30 ` Jeff Hostetler
2017-12-02 16:30 ` Philip Oakley [this message]
2017-12-04 15:36 ` Jeff Hostetler
2017-12-05 23:46 ` Philip Oakley
2017-12-02 15:04 ` Philip Oakley
2017-12-01 17:23 ` Jeff Hostetler
2017-12-01 18:24 ` Jonathan Nieder
2017-12-04 15:53 ` Jeff Hostetler
2017-12-02 18:24 ` Philip Oakley
2017-12-05 19:14 ` Jeff Hostetler
2017-12-05 20:07 ` Jonathan Nieder
2017-12-01 15:28 ` Jeff Hostetler
2017-12-01 14:50 ` Jeff Hostetler
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=A2F606CCDE4542D8AAFC41EB50909FE3@PhilipOakley \
--to=philipoakley@iee.org \
--cc=git@jeffhostetler.com \
--cc=git@vger.kernel.org \
--cc=vit@uber.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox