* Re: broken post-via-gmane link from https://git-scm.com/community
From: Sandro Santilli @ 2016-10-04 8:54 UTC (permalink / raw)
To: Andrey Rybak; +Cc: git
In-Reply-To: <36c73304-aceb-7c51-2788-5ba4cdbc862f@gmail.com>
On Tue, Oct 04, 2016 at 11:52:18AM +0300, Andrey Rybak wrote:
> Hi,
>
> On 04.10.2016 11:11, Sandro Santilli wrote:
> > The "post via gmane" link on https://git-scm.com/community points
> > to an unexistent server 'post.gmane.org':
> > http://post.gmane.org/post.php?group=gmane.comp.version-control.git
> It would probably be better to address this on git-scm.com
> github page, where the source for the site is hosted:
> https://github.com/git/git-scm.com/issues
Done:
https://github.com/git/git-scm.com/issues/859
--strk;
() Free GIS & Flash consultant/developer
/\ https://strk.kbt.io/services.html
^ permalink raw reply
* Merge conflicts in .gitattributes can cause trouble
From: Lars Schneider @ 2016-10-04 10:19 UTC (permalink / raw)
To: git; +Cc: Jeff King, Johannes.Schindelin, me
Hi,
If there is a conflict in the .gitattributes during a merge then it looks
like as if the attributes are not applied (which kind of makes sense as Git
would not know what to do). As a result Git can treat e.g. binary files
as text and they can end up with changed line endings in the working tree.
After resolving the conflict in .gitattributes all files would be marked
as binary, again, and the user can easily commit the wrongly changed line
endings.
Consider this script on Windows:
$ git init .
$ touch first.commit
$ git add .
$ git commit -m "first commit"
$ git checkout -b branch
$ printf "*.bin binary\n" >> .gitattributes
$ git add .
$ git commit -m "tracking *.bin files"
$ git checkout master
$ printf "binary\ndata\n" > file.dat # <-- Unix line ending!
$ printf "*.dat binary\n" >> .gitattributes # <-- Tell Git to keep Unix line ending!
$ git add .
$ git commit -m "tracking *.dat files"
$ git cat-file -p :file.dat | od -c
0000000 b i n a r y \n d a t a \n
^^^^ ^^^^ <-- Correct!
$ git checkout branch
$ git merge master # <-- Causes merge conflict!
$ printf "*.bin binary\n*.dat binary\n" > .gitattributes # <-- Fix merge conflict!
$ git add .
$ git commit -m "merged"
$ git cat-file -p :file.dat | od -c
0000000 b i n a r y \r \n d a t a \r \n
^^^^^^^^ ^^^^^^^^ <-- Wrong!
Possible solutions:
1. We could print an appropriate warning if we detect a merge conflict
in .gitattributes
2. We could disable all line ending conversions in case of a merge conflict
(I am not exactly sure about all the implications, though)
3. We could salvage what we could of the .gitattributes file,
perhaps by using the version from HEAD (or more likely, the ours stage of
the index) -- suggested by Peff on the related GitHub issue mentioned below
Thoughts?
Thanks,
Lars
PS: I noticed that behavior while working with Git LFS and started a discussion
about it here: https://github.com/github/git-lfs/issues/1544
^ permalink raw reply
* Re: Repeatable Extraction
From: Johannes Schindelin @ 2016-10-04 10:25 UTC (permalink / raw)
To: chris king; +Cc: git
In-Reply-To: <CAJQwtsidixAAJKp7-b2PmXgs=mS+PbT5ebOmKLJU1nEn7UJ2og@mail.gmail.com>
Hi Chris,
On Tue, 27 Sep 2016, chris king wrote:
> Is there a way automate extraction that will repeatably generate the
> same files? Currently, each time I extract git portable many of the
> binaries change slightly. For example, if I extract twice using
>
> PortableGit-2.10.0-32-bit.7z.exe -y -gm2
>
> then Beyond Compare tells me that many of the files in usr\bin have
> changed at offset 0x88 and 0x89. Why is that?
The reason is that you look at 32-bit, where technical limitations force
us to hard-code a certain base address for all of the includede MSYS2 .dll
files (i.e. all libraries that require, or implement, the POSIX emulation
layer called MSYS2).
To avoid clashes with other .dll files, that base address is adjusted via
the post-install.bat script for your particular environment.
If you want to avoid that, you will have to extract the installer via
7-Zip: it is a self-extracting .7z archive (and the self-extractor
automatically executes post-install.bat, which subsequently deletes
itself).
Ciao,
Johannes
^ permalink raw reply
* GL bug: can not commit, reports error on changed submodule directory
From: ern0 @ 2016-10-04 10:40 UTC (permalink / raw)
To: git
When I say:
$ gl commit -m "blah blah"
It reports:
✘ Failed to read file into stream: Is a directory
Reason: I have a submodule which has changes.
$ git status
On branch develop
Your branch is up-to-date with 'origin/develop'.
Changes not staged for commit:
(use "git add/rm <file>..." to update what will be committed)
(use "git checkout -- <file>..." to discard changes in working directory)
modified: remoting (new commits)
no changes added to commit (use "git add" and/or "git commit -a")
Workaround: I should sync the directory...
$ cd remoting
$ git commit -am "yada"
$ cd ..
$ git commit -am "yada yada"
$ git push
$ echo I feel clean now
$ echo "# wow" >> test.py
$ gl commit -m "added wow"
...and it works again.
--
ern0
dataflow evangelist
^ permalink raw reply
* Re: Slow pushes on 'pu' - even when up-to-date..
From: Heiko Voigt @ 2016-10-04 11:18 UTC (permalink / raw)
To: Linus Torvalds; +Cc: Junio C Hamano, Stefan Beller, Git Mailing List
In-Reply-To: <CA+55aFyos78qODyw57V=w13Ux5-8SvBqObJFAq22K+XKPWVbAA@mail.gmail.com>
Hi,
On Mon, Oct 03, 2016 at 02:11:36PM -0700, Linus Torvalds wrote:
> This seems to be because I'm now on 'pu' as of a day or two ago in
> order to test the abbrev logic, but lookie here:
>
> time git ls-remote ra.kernel.org:/pub/scm/linux/kernel/git/torvalds/linux
> .. shows all the branches and tags ..
> real 0m0.655s
> user 0m0.011s
> sys 0m0.004s
>
> so the remote is fast to connect to, and with network connection
> overhead and everything, it's just over half a second. But then:
>
> time git push ra.kernel.org:/pub/scm/linux/kernel/git/torvalds/linux
The reason behind this is when pushing to an address we do not easily
have the remote refs to compare available. When pushing an existing ref
it would be easy and could get a shortcut but it gets more complicated
for new refs. Currently we fall back to walking the whole history since
that is "the most correct way" we have. But obviously it is not a
practical solution in any way.
I mentioned this fact when discussing the current state and my patches
to make this check less painful. So we still need to think about a
solution for this check when passing an address.
IMO: It's definitely not ready to be switched on as default, unless we
find something a lot cheaper for the above case.
My idea of a solution goes like this:
* collect all SHA1's of the remotes refs
* check if we have them locally
* if not we abort and tell the user to fetch them somehow into local
refs or disable the check
* when we have them locally we proceed passing those SHA1's as bases
instead of --remotes=<name>
Cheers Heiko
^ permalink raw reply
* Re: Merge conflicts in .gitattributes can cause trouble
From: Duy Nguyen @ 2016-10-04 11:26 UTC (permalink / raw)
To: Lars Schneider; +Cc: git, Jeff King, Johannes Schindelin, me
In-Reply-To: <248A6E81-8D5C-4183-9756-51A0D5193E3E@gmail.com>
On Tue, Oct 4, 2016 at 5:19 PM, Lars Schneider <larsxschneider@gmail.com> wrote:
> Hi,
>
>
> If there is a conflict in the .gitattributes during a merge then it looks
> like as if the attributes are not applied (which kind of makes sense as Git
> would not know what to do). As a result Git can treat e.g. binary files
> as text and they can end up with changed line endings in the working tree.
> After resolving the conflict in .gitattributes all files would be marked
> as binary, again, and the user can easily commit the wrongly changed line
> endings.
>
> Consider this script on Windows:
>
> $ git init .
> $ touch first.commit
> $ git add .
> $ git commit -m "first commit"
>
> $ git checkout -b branch
> $ printf "*.bin binary\n" >> .gitattributes
> $ git add .
> $ git commit -m "tracking *.bin files"
>
> $ git checkout master
> $ printf "binary\ndata\n" > file.dat # <-- Unix line ending!
> $ printf "*.dat binary\n" >> .gitattributes # <-- Tell Git to keep Unix line ending!
> $ git add .
> $ git commit -m "tracking *.dat files"
> $ git cat-file -p :file.dat | od -c
> 0000000 b i n a r y \n d a t a \n
> ^^^^ ^^^^ <-- Correct!
> $ git checkout branch
> $ git merge master # <-- Causes merge conflict!
> $ printf "*.bin binary\n*.dat binary\n" > .gitattributes # <-- Fix merge conflict!
> $ git add .
> $ git commit -m "merged"
> $ git cat-file -p :file.dat | od -c
> 0000000 b i n a r y \r \n d a t a \r \n
> ^^^^^^^^ ^^^^^^^^ <-- Wrong!
>
> Possible solutions:
>
> 1. We could print an appropriate warning if we detect a merge conflict
> in .gitattributes
This is good regardless, to encourage people to resolve conflicts in
.gitattributes first. A good place for this warning may be "git
status"?
> 2. We could disable all line ending conversions in case of a merge conflict
> (I am not exactly sure about all the implications, though)
>
> 3. We could salvage what we could of the .gitattributes file,
> perhaps by using the version from HEAD (or more likely, the ours stage of
> the index) -- suggested by Peff on the related GitHub issue mentioned below
We already have code to fall back to index version in some cases,
adding "fall back on merge conflicts" (and probably updating the index
lookup code too because it looks for stage 0 now) sounds reasonable
(especially with the warning in #1).
BTW whoever fixes this probably should do the same for .gitignore files.
--
Duy
^ permalink raw reply
* Re: Reference a submodule branch instead of a commit
From: Heiko Voigt @ 2016-10-04 11:36 UTC (permalink / raw)
To: Junio C Hamano; +Cc: Jeremy Morton, git
In-Reply-To: <xmqqfuod6yw2.fsf@gitster.mtv.corp.google.com>
On Mon, Oct 03, 2016 at 12:00:45PM -0700, Junio C Hamano wrote:
> Jeremy Morton <admin@game-point.net> writes:
>
> > At the moment, supermodules must reference a given commit in each of
> > its submodules. If one is in control of a submodule and it changes on
> > a regular basis, this can cause a lot of overhead with "submodule
> > updated" commits in the supermodule. It would be useful of git allows
> > the option of referencing a submodule's branch instead of a given
> > submodule commit. How about adding this functionality?
>
> When somebody downstream fetches from your superproject and grabs
> the set of submodules, how would s/he know what _exact_ state you
> meant to record? When s/he says "I have your superproject commit X,
> which binds submodule's branch Y at path sub/, and it simply does
> not work. Your project is broken", how do you go about reproducing
> the exact state s/he had trouble with to help her/him?
>
> The only thing s/he knows is that the commit used from the submodule
> must be one of the commits that was on branch Y at some point in
> time, hopefully close to the timestamp recorded in the commit in the
> superproject. And your record in the history of the superproject
> does not tell you more than that, so you wouldn't have any idea
> better than what s/he already has to help.
>
> Hence, such a "functionality" will never happen, at least in the
> exact form you are describing.
>
> It is conceivable to add some feature that allows you to squelch the
> report that the submodule recorded in your superproject is not up to
> date from "git status" etc. to help those who thinks it is OK to not
> bind the latest submodule commit to the superproject all the time,
> though.
We already have options to support these kinds of workflows. Look at the
option '--remote' for 'git submodule update'.
You then only have to commit the submodule if you do not want to see it
as dirty locally, but you will always get the tip of a remote tracking
branch when updating.
Cheers Heiko
^ permalink raw reply
* Re: [RFC PATCH] clone: add clone.recursesubmodules config option
From: Heiko Voigt @ 2016-10-04 11:41 UTC (permalink / raw)
To: Stefan Beller
Cc: Jeremy Morton, Chris Packham, git@vger.kernel.org, mara.kim,
Junio C Hamano
In-Reply-To: <CAGZ79kbNVy7VFj31m7VKZYP6xphkV_d9Y1x9Q0_=5PZ+_068HA@mail.gmail.com>
On Mon, Oct 03, 2016 at 10:18:32AM -0700, Stefan Beller wrote:
> On Mon, Oct 3, 2016 at 8:36 AM, Jeremy Morton <admin@game-point.net> wrote:
> > Did this ever get anywhere? Can we recursively update submodules with "git
> > pull" in the supermodule now?
>
> I think the idea is sound.
I am confused there is nothing handling *pull* here? This patch was
about clone. Handling 'pull' is a much bigger topic[1].
Cheers Heiko
[1] https://github.com/jlehmann/git-submod-enhancements/wiki/Recursive-submodule-checkout
^ permalink raw reply
* Re: Slow pushes on 'pu' - even when up-to-date..
From: Jeff King @ 2016-10-04 11:44 UTC (permalink / raw)
To: Heiko Voigt
Cc: Linus Torvalds, Junio C Hamano, Stefan Beller, Git Mailing List
In-Reply-To: <20161004111845.GA20309@book.hvoigt.net>
On Tue, Oct 04, 2016 at 01:18:45PM +0200, Heiko Voigt wrote:
> On Mon, Oct 03, 2016 at 02:11:36PM -0700, Linus Torvalds wrote:
> > This seems to be because I'm now on 'pu' as of a day or two ago in
> > order to test the abbrev logic, but lookie here:
> >
> > time git ls-remote ra.kernel.org:/pub/scm/linux/kernel/git/torvalds/linux
> > .. shows all the branches and tags ..
> > real 0m0.655s
> > user 0m0.011s
> > sys 0m0.004s
> >
> > so the remote is fast to connect to, and with network connection
> > overhead and everything, it's just over half a second. But then:
> >
> > time git push ra.kernel.org:/pub/scm/linux/kernel/git/torvalds/linux
>
> The reason behind this is when pushing to an address we do not easily
> have the remote refs to compare available. When pushing an existing ref
> it would be easy and could get a shortcut but it gets more complicated
> for new refs. Currently we fall back to walking the whole history since
> that is "the most correct way" we have. But obviously it is not a
> practical solution in any way.
>
> I mentioned this fact when discussing the current state and my patches
> to make this check less painful. So we still need to think about a
> solution for this check when passing an address.
>
> IMO: It's definitely not ready to be switched on as default, unless we
> find something a lot cheaper for the above case.
>
> My idea of a solution goes like this:
> * collect all SHA1's of the remotes refs
> * check if we have them locally
> * if not we abort and tell the user to fetch them somehow into local
> refs or disable the check
> * when we have them locally we proceed passing those SHA1's as bases
> instead of --remotes=<name>
As I argued in [1], I think it's not just "this must be cheaper" but
"this must not be enabled if submodules are not in use at all". Most
repositories don't have submodules enabled at all, so anything that
cause any extra traversal, even of a portion of the history, is going to
be a net negative for a lot of people.
I think the only sane default is going to be some kind of heuristic that
says "submodules are probably in use". Something like "is there a
.gitmodules file" is not perfect (you can have gitlink entries without
it), but it's a really cheap constant-time check.
-Peff
[1] Quoted in
http://public-inbox.org/git/xmqqh9aaot49.fsf@gitster.mtv.corp.google.com/
^ permalink raw reply
* Re: Re: Slow pushes on 'pu' - even when up-to-date..
From: Heiko Voigt @ 2016-10-04 12:04 UTC (permalink / raw)
To: Jeff King; +Cc: Linus Torvalds, Junio C Hamano, Stefan Beller, Git Mailing List
In-Reply-To: <20161004114428.4wyq54afd4td3epp@sigill.intra.peff.net>
On Tue, Oct 04, 2016 at 07:44:28AM -0400, Jeff King wrote:
> > My idea of a solution goes like this:
> > * collect all SHA1's of the remotes refs
> > * check if we have them locally
> > * if not we abort and tell the user to fetch them somehow into local
> > refs or disable the check
> > * when we have them locally we proceed passing those SHA1's as bases
> > instead of --remotes=<name>
>
> As I argued in [1], I think it's not just "this must be cheaper" but
> "this must not be enabled if submodules are not in use at all". Most
> repositories don't have submodules enabled at all, so anything that
> cause any extra traversal, even of a portion of the history, is going to
> be a net negative for a lot of people.
>
> I think the only sane default is going to be some kind of heuristic that
> says "submodules are probably in use". Something like "is there a
> .gitmodules file" is not perfect (you can have gitlink entries without
> it), but it's a really cheap constant-time check.
I agree. We are adding convenience for submodules, so we can also say a
checked out ".gitmodules" file is a must to have convenience.
I am not sure if I agree on another layer of options for this as
suggested in your post. More options mean more implementation
complexity and more confusion on the users side.
How about we choose our defaults based on the existence of a checked out
.gitmodules file? So the default would only be --recurse-submodules=check
if there is a .gitmodules file in the worktree. All other users need to
either pass or explicitly configure it.
Cheers Heiko
> [1] Quoted in
> http://public-inbox.org/git/xmqqh9aaot49.fsf@gitster.mtv.corp.google.com/
^ permalink raw reply
* Re: Re: Slow pushes on 'pu' - even when up-to-date..
From: Jeff King @ 2016-10-04 12:07 UTC (permalink / raw)
To: Heiko Voigt
Cc: Linus Torvalds, Junio C Hamano, Stefan Beller, Git Mailing List
In-Reply-To: <20161004120421.GA20701@book.hvoigt.net>
On Tue, Oct 04, 2016 at 02:04:21PM +0200, Heiko Voigt wrote:
> > I think the only sane default is going to be some kind of heuristic that
> > says "submodules are probably in use". Something like "is there a
> > .gitmodules file" is not perfect (you can have gitlink entries without
> > it), but it's a really cheap constant-time check.
>
> I agree. We are adding convenience for submodules, so we can also say a
> checked out ".gitmodules" file is a must to have convenience.
>
> I am not sure if I agree on another layer of options for this as
> suggested in your post. More options mean more implementation
> complexity and more confusion on the users side.
>
> How about we choose our defaults based on the existence of a checked out
> .gitmodules file? So the default would only be --recurse-submodules=check
> if there is a .gitmodules file in the worktree. All other users need to
> either pass or explicitly configure it.
That's OK with me. Though you may end up in the long run wanting some
name for the default behavior (e.g., if people configure something else
and then want to override back to "auto" in some instances), but that
can probably come later.
-Peff
^ permalink raw reply
* Re: [PATCH v8 00/11] Git filter protocol
From: Jeff King @ 2016-10-04 12:11 UTC (permalink / raw)
To: Junio C Hamano
Cc: Lars Schneider, Torsten Bögershausen, git, Stefan Beller,
Jakub Narębski, Martin-Louis Bright, ramsay
In-Reply-To: <xmqqvax974dl.fsf@gitster.mtv.corp.google.com>
On Mon, Oct 03, 2016 at 10:02:14AM -0700, Junio C Hamano wrote:
> The timeout would be good for you to give a message "filter process
> running the script '%s' is not exiting; I am waiting for it". The
> user is still left with a hung Git, and can then see if that process
> is hanging around. If it is, then we found a buggy filter. Or we
> found a buggy Git. Either needs to be fixed. I do not think it
> would help anybody by doing a kill(2) to sweep possible bugs under
> the rug.
I would argue that we should not even bother with such a timeout. This
is an exceptional, buggy condition, and hanging is not at all restricted
to this particular case. If git is hanging, then the right tools are
"ps" or "strace" to figure out what is going on. I know that not all
users are comfortable with those tools, but enough are in practice that
the bugs get ironed out, without git having to carry a bunch of extra
timing code that is essentially never exercised.
-Peff
^ permalink raw reply
* Re: [PATCH 3/3] abbrev: auto size the default abbreviation
From: Jeff King @ 2016-10-04 12:18 UTC (permalink / raw)
To: Junio C Hamano; +Cc: Linus Torvalds, Git Mailing List
In-Reply-To: <xmqqbmz051yp.fsf@gitster.mtv.corp.google.com>
On Mon, Oct 03, 2016 at 06:37:18PM -0700, Junio C Hamano wrote:
> Jeff King <peff@peff.net> writes:
>
> >> OK, as Linus's "count at the point of use" is already in 'next',
> >> could you make it incremental with a log message?
> >
> > Sure. I wasn't sure if you actually liked my direction or not, so I was
> > mostly just showing off what the completed one would look like.
>
> To be quite honest, I am not just unsure if I liked your direction;
> rather I am not sure if I actually understood what you perceived as
> a difference that matters between the two approaches. I wanted to
> hear you explain the difference in terms of "Linus's does this, but
> it is bad in X and Y way, so let's avoid it and do it like Z
> instead". One effective way to extract that out of you was to force
> you to justify the "incremental" update.
>
> And it seems that I succeeded ;-).
>
> I am still not sure if I 100% agree with your first paragraph, but
> at least now I think I see where you are coming from.
For the record, I am OK with Linus's patch as-is. It's mostly "that's
not how I would have done it, and the flow seems confusing to me". But
that's subjective; I don't think there are any functional flaws in it.
> You probably will hear from Ramsay about extern-ness of msb().
Heh. I seem to have a real problem with that lately.
-Peff
^ permalink raw reply
* [PATCH v9 02/14] convert: modernize tests
From: larsxschneider @ 2016-10-04 12:59 UTC (permalink / raw)
To: git; +Cc: ramsay, jnareb, gitster, j6t, tboegi, peff, mlbright,
Lars Schneider
In-Reply-To: <20161004125947.67104-1-larsxschneider@gmail.com>
From: Lars Schneider <larsxschneider@gmail.com>
Use `test_config` to set the config, check that files are empty with
`test_must_be_empty`, compare files with `test_cmp`, and remove spaces
after ">" and "<".
Please note that the "rot13" filter configured in "setup" keeps using
`git config` instead of `test_config` because subsequent tests might
depend on it.
Reviewed-by: Stefan Beller <sbeller@google.com>
Signed-off-by: Lars Schneider <larsxschneider@gmail.com>
Signed-off-by: Junio C Hamano <gitster@pobox.com>
---
t/t0021-conversion.sh | 58 +++++++++++++++++++++++++--------------------------
1 file changed, 29 insertions(+), 29 deletions(-)
diff --git a/t/t0021-conversion.sh b/t/t0021-conversion.sh
index e799e59..dc50938 100755
--- a/t/t0021-conversion.sh
+++ b/t/t0021-conversion.sh
@@ -38,8 +38,8 @@ script='s/^\$Id: \([0-9a-f]*\) \$/\1/p'
test_expect_success check '
- cmp test.o test &&
- cmp test.o test.t &&
+ test_cmp test.o test &&
+ test_cmp test.o test.t &&
# ident should be stripped in the repository
git diff --raw --exit-code :test :test.i &&
@@ -47,10 +47,10 @@ test_expect_success check '
embedded=$(sed -ne "$script" test.i) &&
test "z$id" = "z$embedded" &&
- git cat-file blob :test.t > test.r &&
+ git cat-file blob :test.t >test.r &&
- ./rot13.sh < test.o > test.t &&
- cmp test.r test.t
+ ./rot13.sh <test.o >test.t &&
+ test_cmp test.r test.t
'
# If an expanded ident ever gets into the repository, we want to make sure that
@@ -130,7 +130,7 @@ test_expect_success 'filter shell-escaped filenames' '
# delete the files and check them out again, using a smudge filter
# that will count the args and echo the command-line back to us
- git config filter.argc.smudge "sh ./argc.sh %f" &&
+ test_config filter.argc.smudge "sh ./argc.sh %f" &&
rm "$normal" "$special" &&
git checkout -- "$normal" "$special" &&
@@ -141,7 +141,7 @@ test_expect_success 'filter shell-escaped filenames' '
test_cmp expect "$special" &&
# do the same thing, but with more args in the filter expression
- git config filter.argc.smudge "sh ./argc.sh %f --my-extra-arg" &&
+ test_config filter.argc.smudge "sh ./argc.sh %f --my-extra-arg" &&
rm "$normal" "$special" &&
git checkout -- "$normal" "$special" &&
@@ -154,9 +154,9 @@ test_expect_success 'filter shell-escaped filenames' '
'
test_expect_success 'required filter should filter data' '
- git config filter.required.smudge ./rot13.sh &&
- git config filter.required.clean ./rot13.sh &&
- git config filter.required.required true &&
+ test_config filter.required.smudge ./rot13.sh &&
+ test_config filter.required.clean ./rot13.sh &&
+ test_config filter.required.required true &&
echo "*.r filter=required" >.gitattributes &&
@@ -165,17 +165,17 @@ test_expect_success 'required filter should filter data' '
rm -f test.r &&
git checkout -- test.r &&
- cmp test.o test.r &&
+ test_cmp test.o test.r &&
./rot13.sh <test.o >expected &&
git cat-file blob :test.r >actual &&
- cmp expected actual
+ test_cmp expected actual
'
test_expect_success 'required filter smudge failure' '
- git config filter.failsmudge.smudge false &&
- git config filter.failsmudge.clean cat &&
- git config filter.failsmudge.required true &&
+ test_config filter.failsmudge.smudge false &&
+ test_config filter.failsmudge.clean cat &&
+ test_config filter.failsmudge.required true &&
echo "*.fs filter=failsmudge" >.gitattributes &&
@@ -186,9 +186,9 @@ test_expect_success 'required filter smudge failure' '
'
test_expect_success 'required filter clean failure' '
- git config filter.failclean.smudge cat &&
- git config filter.failclean.clean false &&
- git config filter.failclean.required true &&
+ test_config filter.failclean.smudge cat &&
+ test_config filter.failclean.clean false &&
+ test_config filter.failclean.required true &&
echo "*.fc filter=failclean" >.gitattributes &&
@@ -197,8 +197,8 @@ test_expect_success 'required filter clean failure' '
'
test_expect_success 'filtering large input to small output should use little memory' '
- git config filter.devnull.clean "cat >/dev/null" &&
- git config filter.devnull.required true &&
+ test_config filter.devnull.clean "cat >/dev/null" &&
+ test_config filter.devnull.required true &&
for i in $(test_seq 1 30); do printf "%1048576d" 1; done >30MB &&
echo "30MB filter=devnull" >.gitattributes &&
GIT_MMAP_LIMIT=1m GIT_ALLOC_LIMIT=1m git add 30MB
@@ -207,7 +207,7 @@ test_expect_success 'filtering large input to small output should use little mem
test_expect_success 'filter that does not read is fine' '
test-genrandom foo $((128 * 1024 + 1)) >big &&
echo "big filter=epipe" >.gitattributes &&
- git config filter.epipe.clean "echo xyzzy" &&
+ test_config filter.epipe.clean "echo xyzzy" &&
git add big &&
git cat-file blob :big >actual &&
echo xyzzy >expect &&
@@ -215,20 +215,20 @@ test_expect_success 'filter that does not read is fine' '
'
test_expect_success EXPENSIVE 'filter large file' '
- git config filter.largefile.smudge cat &&
- git config filter.largefile.clean cat &&
+ test_config filter.largefile.smudge cat &&
+ test_config filter.largefile.clean cat &&
for i in $(test_seq 1 2048); do printf "%1048576d" 1; done >2GB &&
echo "2GB filter=largefile" >.gitattributes &&
git add 2GB 2>err &&
- ! test -s err &&
+ test_must_be_empty err &&
rm -f 2GB &&
git checkout -- 2GB 2>err &&
- ! test -s err
+ test_must_be_empty err
'
test_expect_success "filter: clean empty file" '
- git config filter.in-repo-header.clean "echo cleaned && cat" &&
- git config filter.in-repo-header.smudge "sed 1d" &&
+ test_config filter.in-repo-header.clean "echo cleaned && cat" &&
+ test_config filter.in-repo-header.smudge "sed 1d" &&
echo "empty-in-worktree filter=in-repo-header" >>.gitattributes &&
>empty-in-worktree &&
@@ -240,8 +240,8 @@ test_expect_success "filter: clean empty file" '
'
test_expect_success "filter: smudge empty file" '
- git config filter.empty-in-repo.clean "cat >/dev/null" &&
- git config filter.empty-in-repo.smudge "echo smudged && cat" &&
+ test_config filter.empty-in-repo.clean "cat >/dev/null" &&
+ test_config filter.empty-in-repo.smudge "echo smudged && cat" &&
echo "empty-in-repo filter=empty-in-repo" >>.gitattributes &&
echo dead data walking >empty-in-repo &&
--
2.10.0
^ permalink raw reply related
* [PATCH v9 06/14] pkt-line: extract set_packet_header()
From: larsxschneider @ 2016-10-04 12:59 UTC (permalink / raw)
To: git; +Cc: ramsay, jnareb, gitster, j6t, tboegi, peff, mlbright,
Lars Schneider
In-Reply-To: <20161004125947.67104-1-larsxschneider@gmail.com>
From: Lars Schneider <larsxschneider@gmail.com>
Extracted set_packet_header() function converts an integer to a 4 byte
hex string. Make this function locally available so that other pkt-line
functions could use it.
Signed-off-by: Lars Schneider <larsxschneider@gmail.com>
Signed-off-by: Junio C Hamano <gitster@pobox.com>
---
pkt-line.c | 19 +++++++++++++------
1 file changed, 13 insertions(+), 6 deletions(-)
diff --git a/pkt-line.c b/pkt-line.c
index 0a9b61c..e8adc0f 100644
--- a/pkt-line.c
+++ b/pkt-line.c
@@ -97,10 +97,20 @@ void packet_buf_flush(struct strbuf *buf)
strbuf_add(buf, "0000", 4);
}
-#define hex(a) (hexchar[(a) & 15])
-static void format_packet(struct strbuf *out, const char *fmt, va_list args)
+static void set_packet_header(char *buf, const int size)
{
static char hexchar[] = "0123456789abcdef";
+
+ #define hex(a) (hexchar[(a) & 15])
+ buf[0] = hex(size >> 12);
+ buf[1] = hex(size >> 8);
+ buf[2] = hex(size >> 4);
+ buf[3] = hex(size);
+ #undef hex
+}
+
+static void format_packet(struct strbuf *out, const char *fmt, va_list args)
+{
size_t orig_len, n;
orig_len = out->len;
@@ -111,10 +121,7 @@ static void format_packet(struct strbuf *out, const char *fmt, va_list args)
if (n > LARGE_PACKET_MAX)
die("protocol error: impossibly long line");
- out->buf[orig_len + 0] = hex(n >> 12);
- out->buf[orig_len + 1] = hex(n >> 8);
- out->buf[orig_len + 2] = hex(n >> 4);
- out->buf[orig_len + 3] = hex(n);
+ set_packet_header(&out->buf[orig_len], n);
packet_trace(out->buf + orig_len + 4, n - 4, 1);
}
--
2.10.0
^ permalink raw reply related
* [PATCH v9 07/14] pkt-line: add packet_write_fmt_gently()
From: larsxschneider @ 2016-10-04 12:59 UTC (permalink / raw)
To: git; +Cc: ramsay, jnareb, gitster, j6t, tboegi, peff, mlbright,
Lars Schneider
In-Reply-To: <20161004125947.67104-1-larsxschneider@gmail.com>
From: Lars Schneider <larsxschneider@gmail.com>
packet_write_fmt() would die in case of a write error even though for
some callers an error would be acceptable. Add packet_write_fmt_gently()
which writes a formatted pkt-line like packet_write_fmt() but does not
die in case of an error. The function is used in a subsequent patch.
Signed-off-by: Lars Schneider <larsxschneider@gmail.com>
Signed-off-by: Junio C Hamano <gitster@pobox.com>
---
pkt-line.c | 34 ++++++++++++++++++++++++++++++----
pkt-line.h | 1 +
2 files changed, 31 insertions(+), 4 deletions(-)
diff --git a/pkt-line.c b/pkt-line.c
index e8adc0f..56915f0 100644
--- a/pkt-line.c
+++ b/pkt-line.c
@@ -125,16 +125,42 @@ static void format_packet(struct strbuf *out, const char *fmt, va_list args)
packet_trace(out->buf + orig_len + 4, n - 4, 1);
}
+static int packet_write_fmt_1(int fd, int gently,
+ const char *fmt, va_list args)
+{
+ struct strbuf buf = STRBUF_INIT;
+ ssize_t count;
+
+ format_packet(&buf, fmt, args);
+ count = write_in_full(fd, buf.buf, buf.len);
+ if (count == buf.len)
+ return 0;
+
+ if (!gently) {
+ check_pipe(errno);
+ die_errno("packet write with format failed");
+ }
+ return error("packet write with format failed");
+}
+
void packet_write_fmt(int fd, const char *fmt, ...)
{
- static struct strbuf buf = STRBUF_INIT;
va_list args;
- strbuf_reset(&buf);
va_start(args, fmt);
- format_packet(&buf, fmt, args);
+ packet_write_fmt_1(fd, 0, fmt, args);
+ va_end(args);
+}
+
+int packet_write_fmt_gently(int fd, const char *fmt, ...)
+{
+ int status;
+ va_list args;
+
+ va_start(args, fmt);
+ status = packet_write_fmt_1(fd, 1, fmt, args);
va_end(args);
- write_or_die(fd, buf.buf, buf.len);
+ return status;
}
void packet_buf_write(struct strbuf *buf, const char *fmt, ...)
diff --git a/pkt-line.h b/pkt-line.h
index 1902fb3..3caea77 100644
--- a/pkt-line.h
+++ b/pkt-line.h
@@ -23,6 +23,7 @@ void packet_flush(int fd);
void packet_write_fmt(int fd, const char *fmt, ...) __attribute__((format (printf, 2, 3)));
void packet_buf_flush(struct strbuf *buf);
void packet_buf_write(struct strbuf *buf, const char *fmt, ...) __attribute__((format (printf, 2, 3)));
+int packet_write_fmt_gently(int fd, const char *fmt, ...) __attribute__((format (printf, 2, 3)));
/*
* Read a packetized line into the buffer, which must be at least size bytes
--
2.10.0
^ permalink raw reply related
* [PATCH v9 09/14] pkt-line: add packet_write_gently()
From: larsxschneider @ 2016-10-04 12:59 UTC (permalink / raw)
To: git; +Cc: ramsay, jnareb, gitster, j6t, tboegi, peff, mlbright,
Lars Schneider
In-Reply-To: <20161004125947.67104-1-larsxschneider@gmail.com>
From: Lars Schneider <larsxschneider@gmail.com>
packet_write_fmt_gently() uses format_packet() which lets the caller
only send string data via "%s". That means it cannot be used for
arbitrary data that may contain NULs.
Add packet_write_gently() which writes arbitrary data and does not die
in case of an error. The function is used by other pkt-line functions in
a subsequent patch.
Signed-off-by: Lars Schneider <larsxschneider@gmail.com>
Signed-off-by: Junio C Hamano <gitster@pobox.com>
---
pkt-line.c | 16 ++++++++++++++++
1 file changed, 16 insertions(+)
diff --git a/pkt-line.c b/pkt-line.c
index 286eb09..3fd4dc0 100644
--- a/pkt-line.c
+++ b/pkt-line.c
@@ -171,6 +171,22 @@ int packet_write_fmt_gently(int fd, const char *fmt, ...)
return status;
}
+static int packet_write_gently(const int fd_out, const char *buf, size_t size)
+{
+ static char packet_write_buffer[LARGE_PACKET_MAX];
+ const size_t packet_size = size + 4;
+
+ if (packet_size > sizeof(packet_write_buffer))
+ return error("packet write failed - data exceeds max packet size");
+
+ packet_trace(buf, size, 1);
+ set_packet_header(packet_write_buffer, packet_size);
+ memcpy(packet_write_buffer + 4, buf, size);
+ if (write_in_full(fd_out, packet_write_buffer, packet_size) == packet_size)
+ return 0;
+ return error("packet write failed");
+}
+
void packet_buf_write(struct strbuf *buf, const char *fmt, ...)
{
va_list args;
--
2.10.0
^ permalink raw reply related
* [PATCH v9 11/14] convert: make apply_filter() adhere to standard Git error handling
From: larsxschneider @ 2016-10-04 12:59 UTC (permalink / raw)
To: git; +Cc: ramsay, jnareb, gitster, j6t, tboegi, peff, mlbright,
Lars Schneider
In-Reply-To: <20161004125947.67104-1-larsxschneider@gmail.com>
From: Lars Schneider <larsxschneider@gmail.com>
apply_filter() returns a boolean that tells the caller if it
"did convert or did not convert". The variable `ret` was used throughout
the function to track errors whereas `1` denoted success and `0`
failure. This is unusual for the Git source where `0` denotes success.
Rename the variable and flip its value to make the function easier
readable for Git developers.
Signed-off-by: Lars Schneider <larsxschneider@gmail.com>
Signed-off-by: Junio C Hamano <gitster@pobox.com>
---
convert.c | 15 ++++++---------
1 file changed, 6 insertions(+), 9 deletions(-)
diff --git a/convert.c b/convert.c
index 986c239..597f561 100644
--- a/convert.c
+++ b/convert.c
@@ -451,7 +451,7 @@ static int apply_filter(const char *path, const char *src, size_t len, int fd,
*
* (child --> cmd) --> us
*/
- int ret = 1;
+ int err = 0;
struct strbuf nbuf = STRBUF_INIT;
struct async async;
struct filter_params params;
@@ -477,23 +477,20 @@ static int apply_filter(const char *path, const char *src, size_t len, int fd,
return 0; /* error was already reported */
if (strbuf_read(&nbuf, async.out, len) < 0) {
- error("read from external filter '%s' failed", cmd);
- ret = 0;
+ err = error("read from external filter '%s' failed", cmd);
}
if (close(async.out)) {
- error("read from external filter '%s' failed", cmd);
- ret = 0;
+ err = error("read from external filter '%s' failed", cmd);
}
if (finish_async(&async)) {
- error("external filter '%s' failed", cmd);
- ret = 0;
+ err = error("external filter '%s' failed", cmd);
}
- if (ret) {
+ if (!err) {
strbuf_swap(dst, &nbuf);
}
strbuf_release(&nbuf);
- return ret;
+ return !err;
}
static struct convert_driver {
--
2.10.0
^ permalink raw reply related
* [PATCH v9 10/14] pkt-line: add functions to read/write flush terminated packet streams
From: larsxschneider @ 2016-10-04 12:59 UTC (permalink / raw)
To: git; +Cc: ramsay, jnareb, gitster, j6t, tboegi, peff, mlbright,
Lars Schneider
In-Reply-To: <20161004125947.67104-1-larsxschneider@gmail.com>
From: Lars Schneider <larsxschneider@gmail.com>
write_packetized_from_fd() and write_packetized_from_buf() write a
stream of packets. All content packets use the maximal packet size
except for the last one. After the last content packet a `flush` control
packet is written.
read_packetized_to_strbuf() reads arbitrary sized packets until it
detects a `flush` packet.
Signed-off-by: Lars Schneider <larsxschneider@gmail.com>
Signed-off-by: Junio C Hamano <gitster@pobox.com>
---
pkt-line.c | 69 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
pkt-line.h | 8 ++++++++
2 files changed, 77 insertions(+)
diff --git a/pkt-line.c b/pkt-line.c
index 3fd4dc0..8ffde22 100644
--- a/pkt-line.c
+++ b/pkt-line.c
@@ -196,6 +196,47 @@ void packet_buf_write(struct strbuf *buf, const char *fmt, ...)
va_end(args);
}
+int write_packetized_from_fd(int fd_in, int fd_out)
+{
+ static char buf[LARGE_PACKET_DATA_MAX];
+ int err = 0;
+ ssize_t bytes_to_write;
+
+ while (!err) {
+ bytes_to_write = xread(fd_in, buf, sizeof(buf));
+ if (bytes_to_write < 0)
+ return COPY_READ_ERROR;
+ if (bytes_to_write == 0)
+ break;
+ err = packet_write_gently(fd_out, buf, bytes_to_write);
+ }
+ if (!err)
+ err = packet_flush_gently(fd_out);
+ return err;
+}
+
+int write_packetized_from_buf(const char *src_in, size_t len, int fd_out)
+{
+ static char buf[LARGE_PACKET_DATA_MAX];
+ int err = 0;
+ size_t bytes_written = 0;
+ size_t bytes_to_write;
+
+ while (!err) {
+ if ((len - bytes_written) > sizeof(buf))
+ bytes_to_write = sizeof(buf);
+ else
+ bytes_to_write = len - bytes_written;
+ if (bytes_to_write == 0)
+ break;
+ err = packet_write_gently(fd_out, src_in + bytes_written, bytes_to_write);
+ bytes_written += bytes_to_write;
+ }
+ if (!err)
+ err = packet_flush_gently(fd_out);
+ return err;
+}
+
static int get_packet_data(int fd, char **src_buf, size_t *src_size,
void *dst, unsigned size, int options)
{
@@ -305,3 +346,31 @@ char *packet_read_line_buf(char **src, size_t *src_len, int *dst_len)
{
return packet_read_line_generic(-1, src, src_len, dst_len);
}
+
+ssize_t read_packetized_to_strbuf(int fd_in, struct strbuf *sb_out)
+{
+ int packet_len;
+
+ size_t orig_len = sb_out->len;
+ size_t orig_alloc = sb_out->alloc;
+
+ for (;;) {
+ strbuf_grow(sb_out, LARGE_PACKET_DATA_MAX);
+ packet_len = packet_read(fd_in, NULL, NULL,
+ // TODO: explain + 1
+ sb_out->buf + sb_out->len, LARGE_PACKET_DATA_MAX+1,
+ PACKET_READ_GENTLE_ON_EOF);
+ if (packet_len <= 0)
+ break;
+ sb_out->len += packet_len;
+ }
+
+ if (packet_len < 0) {
+ if (orig_alloc == 0)
+ strbuf_release(sb_out);
+ else
+ strbuf_setlen(sb_out, orig_len);
+ return packet_len;
+ }
+ return sb_out->len - orig_len;
+}
diff --git a/pkt-line.h b/pkt-line.h
index 3fa0899..18eac64 100644
--- a/pkt-line.h
+++ b/pkt-line.h
@@ -25,6 +25,8 @@ void packet_buf_flush(struct strbuf *buf);
void packet_buf_write(struct strbuf *buf, const char *fmt, ...) __attribute__((format (printf, 2, 3)));
int packet_flush_gently(int fd);
int packet_write_fmt_gently(int fd, const char *fmt, ...) __attribute__((format (printf, 2, 3)));
+int write_packetized_from_fd(int fd_in, int fd_out);
+int write_packetized_from_buf(const char *src_in, size_t len, int fd_out);
/*
* Read a packetized line into the buffer, which must be at least size bytes
@@ -77,8 +79,14 @@ char *packet_read_line(int fd, int *size);
*/
char *packet_read_line_buf(char **src_buf, size_t *src_len, int *size);
+/*
+ * Reads a stream of variable sized packets until a flush packet is detected.
+ */
+ssize_t read_packetized_to_strbuf(int fd_in, struct strbuf *sb_out);
+
#define DEFAULT_PACKET_MAX 1000
#define LARGE_PACKET_MAX 65520
+#define LARGE_PACKET_DATA_MAX (LARGE_PACKET_MAX - 4)
extern char packet_buffer[LARGE_PACKET_MAX];
#endif
--
2.10.0
^ permalink raw reply related
* [PATCH v9 12/14] convert: prepare filter.<driver>.process option
From: larsxschneider @ 2016-10-04 12:59 UTC (permalink / raw)
To: git; +Cc: ramsay, jnareb, gitster, j6t, tboegi, peff, mlbright,
Lars Schneider
In-Reply-To: <20161004125947.67104-1-larsxschneider@gmail.com>
From: Lars Schneider <larsxschneider@gmail.com>
Refactor the existing 'single shot filter mechanism' and prepare the
new 'long running filter mechanism'.
Signed-off-by: Lars Schneider <larsxschneider@gmail.com>
---
convert.c | 60 ++++++++++++++++++++++++++++++++++--------------------------
1 file changed, 34 insertions(+), 26 deletions(-)
diff --git a/convert.c b/convert.c
index 597f561..71e11ff 100644
--- a/convert.c
+++ b/convert.c
@@ -442,7 +442,7 @@ static int filter_buffer_or_fd(int in, int out, void *data)
return (write_err || status);
}
-static int apply_filter(const char *path, const char *src, size_t len, int fd,
+static int apply_single_file_filter(const char *path, const char *src, size_t len, int fd,
struct strbuf *dst, const char *cmd)
{
/*
@@ -456,12 +456,6 @@ static int apply_filter(const char *path, const char *src, size_t len, int fd,
struct async async;
struct filter_params params;
- if (!cmd || !*cmd)
- return 0;
-
- if (!dst)
- return 1;
-
memset(&async, 0, sizeof(async));
async.proc = filter_buffer_or_fd;
async.data = ¶ms;
@@ -493,6 +487,9 @@ static int apply_filter(const char *path, const char *src, size_t len, int fd,
return !err;
}
+#define CAP_CLEAN (1u<<0)
+#define CAP_SMUDGE (1u<<1)
+
static struct convert_driver {
const char *name;
struct convert_driver *next;
@@ -501,6 +498,29 @@ static struct convert_driver {
int required;
} *user_convert, **user_convert_tail;
+static int apply_filter(const char *path, const char *src, size_t len,
+ int fd, struct strbuf *dst, struct convert_driver *drv,
+ const unsigned int wanted_capability)
+{
+ const char *cmd = NULL;
+
+ if (!drv)
+ return 0;
+
+ if (!dst)
+ return 1;
+
+ if ((CAP_CLEAN & wanted_capability) && drv->clean)
+ cmd = drv->clean;
+ else if ((CAP_SMUDGE & wanted_capability) && drv->smudge)
+ cmd = drv->smudge;
+
+ if (cmd && *cmd)
+ return apply_single_file_filter(path, src, len, fd, dst, cmd);
+
+ return 0;
+}
+
static int read_convert_config(const char *var, const char *value, void *cb)
{
const char *key, *name;
@@ -839,7 +859,7 @@ int would_convert_to_git_filter_fd(const char *path)
if (!ca.drv->required)
return 0;
- return apply_filter(path, NULL, 0, -1, NULL, ca.drv->clean);
+ return apply_filter(path, NULL, 0, -1, NULL, ca.drv, CAP_CLEAN);
}
const char *get_convert_attr_ascii(const char *path)
@@ -872,18 +892,12 @@ int convert_to_git(const char *path, const char *src, size_t len,
struct strbuf *dst, enum safe_crlf checksafe)
{
int ret = 0;
- const char *filter = NULL;
- int required = 0;
struct conv_attrs ca;
convert_attrs(&ca, path);
- if (ca.drv) {
- filter = ca.drv->clean;
- required = ca.drv->required;
- }
- ret |= apply_filter(path, src, len, -1, dst, filter);
- if (!ret && required)
+ ret |= apply_filter(path, src, len, -1, dst, ca.drv, CAP_CLEAN);
+ if (!ret && ca.drv && ca.drv->required)
die("%s: clean filter '%s' failed", path, ca.drv->name);
if (ret && dst) {
@@ -907,7 +921,7 @@ void convert_to_git_filter_fd(const char *path, int fd, struct strbuf *dst,
assert(ca.drv);
assert(ca.drv->clean);
- if (!apply_filter(path, NULL, 0, fd, dst, ca.drv->clean))
+ if (!apply_filter(path, NULL, 0, fd, dst, ca.drv, CAP_CLEAN))
die("%s: clean filter '%s' failed", path, ca.drv->name);
crlf_to_git(path, dst->buf, dst->len, dst, ca.crlf_action, checksafe);
@@ -919,15 +933,9 @@ static int convert_to_working_tree_internal(const char *path, const char *src,
int normalizing)
{
int ret = 0, ret_filter = 0;
- const char *filter = NULL;
- int required = 0;
struct conv_attrs ca;
convert_attrs(&ca, path);
- if (ca.drv) {
- filter = ca.drv->smudge;
- required = ca.drv->required;
- }
ret |= ident_to_worktree(path, src, len, dst, ca.ident);
if (ret) {
@@ -938,7 +946,7 @@ static int convert_to_working_tree_internal(const char *path, const char *src,
* CRLF conversion can be skipped if normalizing, unless there
* is a smudge filter. The filter might expect CRLFs.
*/
- if (filter || !normalizing) {
+ if ((ca.drv && ca.drv->smudge) || !normalizing) {
ret |= crlf_to_worktree(path, src, len, dst, ca.crlf_action);
if (ret) {
src = dst->buf;
@@ -946,8 +954,8 @@ static int convert_to_working_tree_internal(const char *path, const char *src,
}
}
- ret_filter = apply_filter(path, src, len, -1, dst, filter);
- if (!ret_filter && required)
+ ret_filter = apply_filter(path, src, len, -1, dst, ca.drv, CAP_SMUDGE);
+ if (!ret_filter && ca.drv && ca.drv->required)
die("%s: smudge filter %s failed", path, ca.drv->name);
return ret | ret_filter;
--
2.10.0
^ permalink raw reply related
* [PATCH v9 13/14] convert: add filter.<driver>.process option
From: larsxschneider @ 2016-10-04 12:59 UTC (permalink / raw)
To: git; +Cc: ramsay, jnareb, gitster, j6t, tboegi, peff, mlbright,
Lars Schneider
In-Reply-To: <20161004125947.67104-1-larsxschneider@gmail.com>
From: Lars Schneider <larsxschneider@gmail.com>
Git's clean/smudge mechanism invokes an external filter process for
every single blob that is affected by a filter. If Git filters a lot of
blobs then the startup time of the external filter processes can become
a significant part of the overall Git execution time.
In a preliminary performance test this developer used a clean/smudge
filter written in golang to filter 12,000 files. This process took 364s
with the existing filter mechanism and 5s with the new mechanism. See
details here: https://github.com/github/git-lfs/pull/1382
This patch adds the `filter.<driver>.process` string option which, if
used, keeps the external filter process running and processes all blobs
with the packet format (pkt-line) based protocol over standard input and
standard output. The full protocol is explained in detail in
`Documentation/gitattributes.txt`.
A few key decisions:
* The long running filter process is referred to as filter protocol
version 2 because the existing single shot filter invocation is
considered version 1.
* Git sends a welcome message and expects a response right after the
external filter process has started. This ensures that Git will not
hang if a version 1 filter is incorrectly used with the
filter.<driver>.process option for version 2 filters. In addition,
Git can detect this kind of error and warn the user.
* The status of a filter operation (e.g. "success" or "error) is set
before the actual response and (if necessary!) re-set after the
response. The advantage of this two step status response is that if
the filter detects an error early, then the filter can communicate
this and Git does not even need to create structures to read the
response.
* All status responses are pkt-line lists terminated with a flush
packet. This allows us to send other status fields with the same
protocol in the future.
Helped-by: Martin-Louis Bright <mlbright@gmail.com>
Reviewed-by: Jakub Narebski <jnareb@gmail.com>
Signed-off-by: Lars Schneider <larsxschneider@gmail.com>
Signed-off-by: Junio C Hamano <gitster@pobox.com>
---
Documentation/gitattributes.txt | 157 +++++++++++++-
convert.c | 295 +++++++++++++++++++++++++-
t/t0021-conversion.sh | 447 +++++++++++++++++++++++++++++++++++++++-
t/t0021/rot13-filter.pl | 191 +++++++++++++++++
4 files changed, 1080 insertions(+), 10 deletions(-)
create mode 100755 t/t0021/rot13-filter.pl
diff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt
index 7aff940..5868f00 100644
--- a/Documentation/gitattributes.txt
+++ b/Documentation/gitattributes.txt
@@ -293,7 +293,13 @@ checkout, when the `smudge` command is specified, the command is
fed the blob object from its standard input, and its standard
output is used to update the worktree file. Similarly, the
`clean` command is used to convert the contents of worktree file
-upon checkin.
+upon checkin. By default these commands process only a single
+blob and terminate. If a long running `process` filter is used
+in place of `clean` and/or `smudge` filters, then Git can process
+all blobs with a single filter command invocation for the entire
+life of a single Git command, for example `git add --all`. See
+section below for the description of the protocol used to
+communicate with a `process` filter.
One use of the content filtering is to massage the content into a shape
that is more convenient for the platform, filesystem, and the user to use.
@@ -373,6 +379,155 @@ not exist, or may have different contents. So, smudge and clean commands
should not try to access the file on disk, but only act as filters on the
content provided to them on standard input.
+Long Running Filter Process
+^^^^^^^^^^^^^^^^^^^^^^^^^^^
+
+If the filter command (a string value) is defined via
+`filter.<driver>.process` then Git can process all blobs with a
+single filter invocation for the entire life of a single Git
+command. This is achieved by using a packet format (pkt-line,
+see technical/protocol-common.txt) based protocol over standard
+input and standard output as follows. All packets, except for the
+"*CONTENT" packets and the "0000" flush packet, are considered
+text and therefore are terminated by a LF.
+
+Git starts the filter when it encounters the first file
+that needs to be cleaned or smudged. After the filter started
+Git sends a welcome message ("git-filter-client"), a list of
+supported protocol version numbers, and a flush packet. Git expects
+to read a welcome response message ("git-filter-server") and exactly
+one protocol version number from the previously sent list. All further
+communication will be based on the selected version. The remaining
+protocol description below documents "version=2". Please note that
+"version=42" in the example below does not exist and is only there
+to illustrate how the protocol would look like with more than one
+version.
+
+After the version negotiation Git sends a list of all capabilities that
+it supports and a flush packet. Git expects to read a list of desired
+capabilities, which must be a subset of the supported capabilities list,
+and a flush packet as response:
+------------------------
+packet: git> git-filter-client
+packet: git> version=2
+packet: git> version=42
+packet: git> 0000
+packet: git< git-filter-server
+packet: git< version=2
+packet: git> clean=true
+packet: git> smudge=true
+packet: git> not-yet-invented=true
+packet: git> 0000
+packet: git< clean=true
+packet: git< smudge=true
+packet: git< 0000
+------------------------
+Supported filter capabilities in version 2 are "clean" and
+"smudge".
+
+Afterwards Git sends a list of "key=value" pairs terminated with
+a flush packet. The list will contain at least the filter command
+(based on the supported capabilities) and the pathname of the file
+to filter relative to the repository root. Right after these packets
+Git sends the content split in zero or more pkt-line packets and a
+flush packet to terminate content. Please note, that the filter
+must not send any response before it received the content and the
+final flush packet.
+------------------------
+packet: git> command=smudge
+packet: git> pathname=path/testfile.dat
+packet: git> 0000
+packet: git> CONTENT
+packet: git> 0000
+------------------------
+
+The filter is expected to respond with a list of "key=value" pairs
+terminated with a flush packet. If the filter does not experience
+problems then the list must contain a "success" status. Right after
+these packets the filter is expected to send the content in zero
+or more pkt-line packets and a flush packet at the end. Finally, a
+second list of "key=value" pairs terminated with a flush packet
+is expected. The filter can change the status in the second list.
+------------------------
+packet: git< status=success
+packet: git< 0000
+packet: git< SMUDGED_CONTENT
+packet: git< 0000
+packet: git< 0000 # empty list, keep "status=success" unchanged!
+------------------------
+
+If the result content is empty then the filter is expected to respond
+with a "success" status and an empty list.
+------------------------
+packet: git< status=success
+packet: git< 0000
+packet: git< 0000 # empty content!
+packet: git< 0000 # empty list, keep "status=success" unchanged!
+------------------------
+
+In case the filter cannot or does not want to process the content,
+it is expected to respond with an "error" status. Depending on the
+`filter.<driver>.required` flag Git will interpret that as error
+but it will not stop or restart the filter process.
+------------------------
+packet: git< status=error
+packet: git< 0000
+------------------------
+
+If the filter experiences an error during processing, then it can
+send the status "error" after the content was (partially or
+completely) sent. Depending on the `filter.<driver>.required` flag
+Git will interpret that as error but it will not stop or restart the
+filter process.
+------------------------
+packet: git< status=success
+packet: git< 0000
+packet: git< HALF_WRITTEN_ERRONEOUS_CONTENT
+packet: git< 0000
+packet: git< status=error
+packet: git< 0000
+------------------------
+
+If the filter dies during the communication or does not adhere to
+the protocol then Git will stop the filter process and restart it
+with the next file that needs to be processed. Depending on the
+`filter.<driver>.required` flag Git will interpret that as error.
+
+The error handling for all cases above mimic the behavior of
+the `filter.<driver>.clean` / `filter.<driver>.smudge` error
+handling.
+
+In case the filter cannot or does not want to process the content
+as well as any future content for the lifetime of the Git process,
+it is expected to respond with an "abort" status at any point in
+the protocol. Depending on the `filter.<driver>.required` flag Git
+will interpret that as error for the content as well as any future
+content for the lifetime of the Git process but it will not stop or
+restart the filter process.
+------------------------
+packet: git< status=abort
+packet: git< 0000
+------------------------
+
+After the filter has processed a blob it is expected to wait for
+the next "key=value" list containing a command. Git will close
+the command pipe on exit. The filter is expected to detect EOF
+and exit gracefully on its own.
+
+If you develop your own long running filter
+process then the `GIT_TRACE_PACKET` environment variables can be
+very helpful for debugging (see linkgit:git[1]).
+
+If a `filter.<driver>.process` command is configured then it
+always takes precedence over a configured `filter.<driver>.clean`
+or `filter.<driver>.smudge` command.
+
+Please note that you cannot use an existing `filter.<driver>.clean`
+or `filter.<driver>.smudge` command with `filter.<driver>.process`
+because the former two use a different inter process communication
+protocol than the latter one.
+
+
Interaction between checkin/checkout attributes
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
diff --git a/convert.c b/convert.c
index 71e11ff..88581d6 100644
--- a/convert.c
+++ b/convert.c
@@ -3,6 +3,7 @@
#include "run-command.h"
#include "quote.h"
#include "sigchain.h"
+#include "pkt-line.h"
/*
* convert.c - convert a file when checking it out and checking it in.
@@ -490,11 +491,287 @@ static int apply_single_file_filter(const char *path, const char *src, size_t le
#define CAP_CLEAN (1u<<0)
#define CAP_SMUDGE (1u<<1)
+struct cmd2process {
+ struct hashmap_entry ent; /* must be the first member! */
+ unsigned int supported_capabilities;
+ const char *cmd;
+ struct child_process process;
+};
+
+static int cmd_process_map_initialized;
+static struct hashmap cmd_process_map;
+
+static int cmd2process_cmp(const struct cmd2process *e1,
+ const struct cmd2process *e2,
+ const void *unused)
+{
+ return strcmp(e1->cmd, e2->cmd);
+}
+
+static struct cmd2process *find_multi_file_filter_entry(struct hashmap *hashmap, const char *cmd)
+{
+ struct cmd2process key;
+ hashmap_entry_init(&key, strhash(cmd));
+ key.cmd = cmd;
+ return hashmap_get(hashmap, &key, NULL);
+}
+
+static void kill_multi_file_filter(struct hashmap *hashmap, struct cmd2process *entry)
+{
+ if (!entry)
+ return;
+ sigchain_push(SIGPIPE, SIG_IGN);
+ /*
+ * We kill the filter most likely because an error happened already.
+ * That's why we are not interested in any error code here.
+ */
+ close(entry->process.in);
+ close(entry->process.out);
+ sigchain_pop(SIGPIPE);
+ finish_command(&entry->process);
+ hashmap_remove(hashmap, entry, NULL);
+ free(entry);
+}
+
+static int packet_write_list(int fd, const char *line, ...)
+{
+ va_list args;
+ int err;
+ va_start(args, line);
+ for (;;) {
+ if (!line)
+ break;
+ if (strlen(line) > LARGE_PACKET_DATA_MAX)
+ return -1;
+ err = packet_write_fmt_gently(fd, "%s\n", line);
+ if (err)
+ return err;
+ line = va_arg(args, const char*);
+ }
+ va_end(args);
+ return packet_flush_gently(fd);
+}
+
+static struct cmd2process *start_multi_file_filter(struct hashmap *hashmap, const char *cmd)
+{
+ int err;
+ struct cmd2process *entry;
+ struct child_process *process;
+ const char *argv[] = { cmd, NULL };
+ struct string_list cap_list = STRING_LIST_INIT_NODUP;
+ char *cap_buf;
+ const char *cap_name;
+
+ entry = xmalloc(sizeof(*entry));
+ hashmap_entry_init(entry, strhash(cmd));
+ entry->cmd = cmd;
+ entry->supported_capabilities = 0;
+ process = &entry->process;
+
+ child_process_init(process);
+ process->argv = argv;
+ process->use_shell = 1;
+ process->in = -1;
+ process->out = -1;
+ process->wait_on_exit = 1;
+
+ if (start_command(process)) {
+ error("cannot fork to run external filter '%s'", cmd);
+ kill_multi_file_filter(hashmap, entry);
+ return NULL;
+ }
+
+ sigchain_push(SIGPIPE, SIG_IGN);
+
+ err = packet_write_list(process->in, "git-filter-client", "version=2", NULL);
+ if (err)
+ goto done;
+
+ err = strcmp(packet_read_line(process->out, NULL), "git-filter-server");
+ if (err) {
+ error("external filter '%s' does not support filter protocol version 2", cmd);
+ goto done;
+ }
+ err = strcmp(packet_read_line(process->out, NULL), "version=2");
+ if (err)
+ goto done;
+
+ err = packet_write_list(process->in, "clean=true", "smudge=true", NULL);
+
+ for (;;) {
+ cap_buf = packet_read_line(process->out, NULL);
+ if (!cap_buf)
+ break;
+ string_list_split_in_place(&cap_list, cap_buf, '=', 1);
+
+ if (cap_list.nr != 2 || strcmp(cap_list.items[1].string, "true"))
+ continue;
+
+ cap_name = cap_list.items[0].string;
+ if (!strcmp(cap_name, "clean")) {
+ entry->supported_capabilities |= CAP_CLEAN;
+ } else if (!strcmp(cap_name, "smudge")) {
+ entry->supported_capabilities |= CAP_SMUDGE;
+ } else {
+ warning(
+ "external filter '%s' requested unsupported filter capability '%s'",
+ cmd, cap_name
+ );
+ }
+
+ string_list_clear(&cap_list, 0);
+ }
+
+done:
+ sigchain_pop(SIGPIPE);
+
+ if (err || errno == EPIPE) {
+ error("initialization for external filter '%s' failed", cmd);
+ kill_multi_file_filter(hashmap, entry);
+ return NULL;
+ }
+
+ hashmap_add(hashmap, entry);
+ return entry;
+}
+
+static void read_multi_file_filter_status(int fd, struct strbuf *status) {
+ struct strbuf **pair;
+ char *line;
+ for (;;) {
+ line = packet_read_line(fd, NULL);
+ if (!line)
+ break;
+ pair = strbuf_split_str(line, '=', 2);
+ if (pair[0] && pair[0]->len && pair[1]) {
+ if (!strcmp(pair[0]->buf, "status=")) {
+ strbuf_reset(status);
+ strbuf_addbuf(status, pair[1]);
+ }
+ }
+ strbuf_list_free(pair);
+ }
+}
+
+static int apply_multi_file_filter(const char *path, const char *src, size_t len,
+ int fd, struct strbuf *dst, const char *cmd,
+ const unsigned int wanted_capability)
+{
+ int err;
+ struct cmd2process *entry;
+ struct child_process *process;
+ struct stat file_stat;
+ struct strbuf nbuf = STRBUF_INIT;
+ struct strbuf filter_status = STRBUF_INIT;
+ char *filter_type;
+
+ if (!cmd_process_map_initialized) {
+ cmd_process_map_initialized = 1;
+ hashmap_init(&cmd_process_map, (hashmap_cmp_fn) cmd2process_cmp, 0);
+ entry = NULL;
+ } else {
+ entry = find_multi_file_filter_entry(&cmd_process_map, cmd);
+ }
+
+ fflush(NULL);
+
+ if (!entry) {
+ entry = start_multi_file_filter(&cmd_process_map, cmd);
+ if (!entry)
+ return 0;
+ }
+ process = &entry->process;
+
+ if (!(wanted_capability & entry->supported_capabilities))
+ return 0;
+
+ if (CAP_CLEAN & wanted_capability)
+ filter_type = "clean";
+ else if (CAP_SMUDGE & wanted_capability)
+ filter_type = "smudge";
+ else
+ die("unexpected filter type");
+
+ if (fd >= 0 && !src) {
+ if (fstat(fd, &file_stat) == -1)
+ return 0;
+ len = xsize_t(file_stat.st_size);
+ }
+
+ sigchain_push(SIGPIPE, SIG_IGN);
+
+ assert(strlen(filter_type) < LARGE_PACKET_DATA_MAX - strlen("command=\n"));
+ err = packet_write_fmt_gently(process->in, "command=%s\n", filter_type);
+ if (err)
+ goto done;
+
+ err = strlen(path) > LARGE_PACKET_DATA_MAX - strlen("pathname=\n");
+ if (err) {
+ error("path name too long for external filter");
+ goto done;
+ }
+
+ err = packet_write_fmt_gently(process->in, "pathname=%s\n", path);
+ if (err)
+ goto done;
+
+ err = packet_flush_gently(process->in);
+ if (err)
+ goto done;
+
+ if (fd >= 0)
+ err = write_packetized_from_fd(fd, process->in);
+ else
+ err = write_packetized_from_buf(src, len, process->in);
+ if (err)
+ goto done;
+
+ read_multi_file_filter_status(process->out, &filter_status);
+ err = strcmp(filter_status.buf, "success");
+ if (err)
+ goto done;
+
+ err = read_packetized_to_strbuf(process->out, &nbuf) < 0;
+ if (err)
+ goto done;
+
+ read_multi_file_filter_status(process->out, &filter_status);
+ err = strcmp(filter_status.buf, "success");
+
+done:
+ sigchain_pop(SIGPIPE);
+
+ if (err || errno == EPIPE) {
+ if (!strcmp(filter_status.buf, "error")) {
+ /* The filter signaled a problem with the file. */
+ } else if (!strcmp(filter_status.buf, "abort")) {
+ /*
+ * The filter signaled a permanent problem. Don't try to filter
+ * files with the same command for the lifetime of the current
+ * Git process.
+ */
+ entry->supported_capabilities &= ~wanted_capability;
+ } else {
+ /*
+ * Something went wrong with the protocol filter.
+ * Force shutdown and restart if another blob requires filtering.
+ */
+ error("external filter '%s' failed", cmd);
+ kill_multi_file_filter(&cmd_process_map, entry);
+ }
+ } else {
+ strbuf_swap(dst, &nbuf);
+ }
+ strbuf_release(&nbuf);
+ return !err;
+}
+
static struct convert_driver {
const char *name;
struct convert_driver *next;
const char *smudge;
const char *clean;
+ const char *process;
int required;
} *user_convert, **user_convert_tail;
@@ -510,13 +787,15 @@ static int apply_filter(const char *path, const char *src, size_t len,
if (!dst)
return 1;
- if ((CAP_CLEAN & wanted_capability) && drv->clean)
+ if ((CAP_CLEAN & wanted_capability) && !drv->process && drv->clean)
cmd = drv->clean;
- else if ((CAP_SMUDGE & wanted_capability) && drv->smudge)
+ else if ((CAP_SMUDGE & wanted_capability) && !drv->process && drv->smudge)
cmd = drv->smudge;
if (cmd && *cmd)
return apply_single_file_filter(path, src, len, fd, dst, cmd);
+ else if (drv->process && *drv->process)
+ return apply_multi_file_filter(path, src, len, fd, dst, drv->process, wanted_capability);
return 0;
}
@@ -558,6 +837,9 @@ static int read_convert_config(const char *var, const char *value, void *cb)
if (!strcmp("clean", key))
return git_config_string(&drv->clean, var, value);
+ if (!strcmp("process", key))
+ return git_config_string(&drv->process, var, value);
+
if (!strcmp("required", key)) {
drv->required = git_config_bool(var, value);
return 0;
@@ -919,7 +1201,7 @@ void convert_to_git_filter_fd(const char *path, int fd, struct strbuf *dst,
convert_attrs(&ca, path);
assert(ca.drv);
- assert(ca.drv->clean);
+ assert(ca.drv->clean || ca.drv->process);
if (!apply_filter(path, NULL, 0, fd, dst, ca.drv, CAP_CLEAN))
die("%s: clean filter '%s' failed", path, ca.drv->name);
@@ -944,9 +1226,10 @@ static int convert_to_working_tree_internal(const char *path, const char *src,
}
/*
* CRLF conversion can be skipped if normalizing, unless there
- * is a smudge filter. The filter might expect CRLFs.
+ * is a smudge or process filter (even if the process filter doesn't
+ * support smudge). The filters might expect CRLFs.
*/
- if ((ca.drv && ca.drv->smudge) || !normalizing) {
+ if ((ca.drv && (ca.drv->smudge || ca.drv->process)) || !normalizing) {
ret |= crlf_to_worktree(path, src, len, dst, ca.crlf_action);
if (ret) {
src = dst->buf;
@@ -1407,7 +1690,7 @@ struct stream_filter *get_stream_filter(const char *path, const unsigned char *s
struct stream_filter *filter = NULL;
convert_attrs(&ca, path);
- if (ca.drv && (ca.drv->smudge || ca.drv->clean))
+ if (ca.drv && (ca.drv->process || ca.drv->smudge || ca.drv->clean))
return NULL;
if (ca.crlf_action == CRLF_AUTO || ca.crlf_action == CRLF_AUTO_CRLF)
diff --git a/t/t0021-conversion.sh b/t/t0021-conversion.sh
index dc50938..52b7fe9 100755
--- a/t/t0021-conversion.sh
+++ b/t/t0021-conversion.sh
@@ -4,13 +4,77 @@ test_description='blob conversion via gitattributes'
. ./test-lib.sh
-cat <<EOF >rot13.sh
+TEST_ROOT="$(pwd)"
+
+cat <<EOF >"$TEST_ROOT/rot13.sh"
#!$SHELL_PATH
tr \
'abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ' \
'nopqrstuvwxyzabcdefghijklmNOPQRSTUVWXYZABCDEFGHIJKLM'
EOF
-chmod +x rot13.sh
+chmod +x "$TEST_ROOT/rot13.sh"
+
+generate_random_characters () {
+ LEN=$1
+ NAME=$2
+ test-genrandom some-seed $LEN |
+ perl -pe "s/./chr((ord($&) % 26) + ord('a'))/sge" >"$TEST_ROOT/$NAME"
+}
+
+file_size () {
+ cat "$1" | wc -c | sed "s/^[ ]*//"
+}
+
+filter_git () {
+ rm -f rot13-filter.log &&
+ git "$@" 2>git-stderr.log &&
+ sed '/Waiting for/d' git-stderr.log >git-stderr-clean.log &&
+ test_must_be_empty git-stderr-clean.log &&
+ rm -f git-stderr.log git-stderr-clean.log
+}
+
+# Count unique lines in two files and compare them.
+test_cmp_count () {
+ for FILE in $@
+ do
+ sort $FILE | uniq -c | sed "s/^[ ]*//" >$FILE.tmp
+ cat $FILE.tmp >$FILE
+ done &&
+ test_cmp $@
+}
+
+# Count unique lines except clean invocations in two files and compare
+# them. Clean invocations are not counted because their number can vary.
+# c.f. http://public-inbox.org/git/xmqqshv18i8i.fsf@gitster.mtv.corp.google.com/
+test_cmp_count_except_clean () {
+ for FILE in $@
+ do
+ sort $FILE | uniq -c | sed "s/^[ ]*//" |
+ sed "s/^\([0-9]\) IN: clean/x IN: clean/" >$FILE.tmp
+ cat $FILE.tmp >$FILE
+ done &&
+ test_cmp $@
+}
+
+# Compare two files but exclude clean invocations because they can vary.
+# c.f. http://public-inbox.org/git/xmqqshv18i8i.fsf@gitster.mtv.corp.google.com/
+test_cmp_exclude_clean () {
+ for FILE in $@
+ do
+ grep -v "IN: clean" $FILE >$FILE.tmp
+ cat $FILE.tmp >$FILE
+ done &&
+ test_cmp $@
+}
+
+# Check that the contents of two files are equal and that their rot13 version
+# is equal to the committed content.
+test_cmp_committed_rot13 () {
+ test_cmp "$1" "$2" &&
+ "$TEST_ROOT/rot13.sh" <"$1" >expected &&
+ git cat-file blob :"$2" >actual &&
+ test_cmp expected actual
+}
test_expect_success setup '
git config filter.rot13.smudge ./rot13.sh &&
@@ -31,7 +95,10 @@ test_expect_success setup '
cat test >test.i &&
git add test test.t test.i &&
rm -f test test.t test.i &&
- git checkout -- test test.t test.i
+ git checkout -- test test.t test.i &&
+
+ echo "content-test2" >test2.o &&
+ echo "content-test3 - filename with special characters" >"test3 '\''sq'\'',\$x.o"
'
script='s/^\$Id: \([0-9a-f]*\) \$/\1/p'
@@ -279,4 +346,378 @@ test_expect_success 'diff does not reuse worktree files that need cleaning' '
test_line_count = 0 count
'
+test_expect_success PERL 'required process filter should filter data' '
+ test_config_global filter.protocol.process "$TEST_DIRECTORY/t0021/rot13-filter.pl clean smudge" &&
+ test_config_global filter.protocol.required true &&
+ rm -rf repo &&
+ mkdir repo &&
+ (
+ cd repo &&
+ git init &&
+
+ echo "git-stderr.log" >.gitignore &&
+ echo "*.r filter=protocol" >.gitattributes &&
+ git add . &&
+ git commit . -m "test commit 1" &&
+ git branch empty-branch &&
+
+ cp "$TEST_ROOT/test.o" test.r &&
+ cp "$TEST_ROOT/test2.o" test2.r &&
+ mkdir testsubdir &&
+ cp "$TEST_ROOT/test3 '\''sq'\'',\$x.o" "testsubdir/test3 '\''sq'\'',\$x.r" &&
+ >test4-empty.r &&
+
+ S=$(file_size test.r) &&
+ S2=$(file_size test2.r) &&
+ S3=$(file_size "testsubdir/test3 '\''sq'\'',\$x.r") &&
+
+ filter_git add . &&
+ cat >expected.log <<-EOF &&
+ START
+ init handshake complete
+ IN: clean test.r $S [OK] -- OUT: $S . [OK]
+ IN: clean test2.r $S2 [OK] -- OUT: $S2 . [OK]
+ IN: clean test4-empty.r 0 [OK] -- OUT: 0 [OK]
+ IN: clean testsubdir/test3 '\''sq'\'',\$x.r $S3 [OK] -- OUT: $S3 . [OK]
+ STOP
+ EOF
+ test_cmp_count expected.log rot13-filter.log &&
+
+ filter_git commit . -m "test commit 2" &&
+ cat >expected.log <<-EOF &&
+ START
+ init handshake complete
+ IN: clean test.r $S [OK] -- OUT: $S . [OK]
+ IN: clean test2.r $S2 [OK] -- OUT: $S2 . [OK]
+ IN: clean test4-empty.r 0 [OK] -- OUT: 0 [OK]
+ IN: clean testsubdir/test3 '\''sq'\'',\$x.r $S3 [OK] -- OUT: $S3 . [OK]
+ IN: clean test.r $S [OK] -- OUT: $S . [OK]
+ IN: clean test2.r $S2 [OK] -- OUT: $S2 . [OK]
+ IN: clean test4-empty.r 0 [OK] -- OUT: 0 [OK]
+ IN: clean testsubdir/test3 '\''sq'\'',\$x.r $S3 [OK] -- OUT: $S3 . [OK]
+ STOP
+ EOF
+ test_cmp_count_except_clean expected.log rot13-filter.log &&
+
+ rm -f test2.r "testsubdir/test3 '\''sq'\'',\$x.r" &&
+
+ filter_git checkout --quiet --no-progress . &&
+ cat >expected.log <<-EOF &&
+ START
+ init handshake complete
+ IN: smudge test2.r $S2 [OK] -- OUT: $S2 . [OK]
+ IN: smudge testsubdir/test3 '\''sq'\'',\$x.r $S3 [OK] -- OUT: $S3 . [OK]
+ STOP
+ EOF
+ test_cmp_exclude_clean expected.log rot13-filter.log &&
+
+ filter_git checkout --quiet --no-progress empty-branch &&
+ cat >expected.log <<-EOF &&
+ START
+ init handshake complete
+ IN: clean test.r $S [OK] -- OUT: $S . [OK]
+ STOP
+ EOF
+ test_cmp_exclude_clean expected.log rot13-filter.log &&
+
+ filter_git checkout --quiet --no-progress master &&
+ cat >expected.log <<-EOF &&
+ START
+ init handshake complete
+ IN: smudge test.r $S [OK] -- OUT: $S . [OK]
+ IN: smudge test2.r $S2 [OK] -- OUT: $S2 . [OK]
+ IN: smudge test4-empty.r 0 [OK] -- OUT: 0 [OK]
+ IN: smudge testsubdir/test3 '\''sq'\'',\$x.r $S3 [OK] -- OUT: $S3 . [OK]
+ STOP
+ EOF
+ test_cmp_exclude_clean expected.log rot13-filter.log &&
+
+ test_cmp_committed_rot13 "$TEST_ROOT/test.o" test.r &&
+ test_cmp_committed_rot13 "$TEST_ROOT/test2.o" test2.r &&
+ test_cmp_committed_rot13 "$TEST_ROOT/test3 '\''sq'\'',\$x.o" "testsubdir/test3 '\''sq'\'',\$x.r"
+ )
+'
+
+test_expect_success PERL 'required process filter takes precedence' '
+ test_config_global filter.protocol.clean false &&
+ test_config_global filter.protocol.process "$TEST_DIRECTORY/t0021/rot13-filter.pl clean" &&
+ test_config_global filter.protocol.required true &&
+ rm -rf repo &&
+ mkdir repo &&
+ (
+ cd repo &&
+ git init &&
+
+ echo "*.r filter=protocol" >.gitattributes &&
+ cp "$TEST_ROOT/test.o" test.r &&
+ S=$(file_size test.r) &&
+
+ # Check that the process filter is invoked here
+ filter_git add . &&
+ cat >expected.log <<-EOF &&
+ START
+ init handshake complete
+ IN: clean test.r $S [OK] -- OUT: $S . [OK]
+ STOP
+ EOF
+ test_cmp_count expected.log rot13-filter.log
+ )
+'
+
+test_expect_success PERL 'required process filter should be used only for "clean" operation only' '
+ test_config_global filter.protocol.process "$TEST_DIRECTORY/t0021/rot13-filter.pl clean" &&
+ rm -rf repo &&
+ mkdir repo &&
+ (
+ cd repo &&
+ git init &&
+
+ echo "*.r filter=protocol" >.gitattributes &&
+ cp "$TEST_ROOT/test.o" test.r &&
+ S=$(file_size test.r) &&
+
+ filter_git add . &&
+ cat >expected.log <<-EOF &&
+ START
+ init handshake complete
+ IN: clean test.r $S [OK] -- OUT: $S . [OK]
+ STOP
+ EOF
+ test_cmp_count expected.log rot13-filter.log &&
+
+ rm test.r &&
+
+ filter_git checkout --quiet --no-progress . &&
+ # If the filter would be used for "smudge", too, we would see
+ # "IN: smudge test.r 57 [OK] -- OUT: 57 . [OK]" here
+ cat >expected.log <<-EOF &&
+ START
+ init handshake complete
+ STOP
+ EOF
+ test_cmp_exclude_clean expected.log rot13-filter.log
+ )
+'
+
+test_expect_success PERL 'required process filter should process multiple packets' '
+ test_config_global filter.protocol.process "$TEST_DIRECTORY/t0021/rot13-filter.pl clean smudge" &&
+ test_config_global filter.protocol.required true &&
+
+ rm -rf repo &&
+ mkdir repo &&
+ (
+ cd repo &&
+ git init &&
+
+ # Generate data requiring 1, 2, 3 packets
+ S=65516 && # PKTLINE_DATA_MAXLEN -> Maximal size of a packet
+ generate_random_characters $(($S )) 1pkt_1__.file &&
+ generate_random_characters $(($S +1)) 2pkt_1+1.file &&
+ generate_random_characters $(($S*2-1)) 2pkt_2-1.file &&
+ generate_random_characters $(($S*2 )) 2pkt_2__.file &&
+ generate_random_characters $(($S*2+1)) 3pkt_2+1.file &&
+
+ for FILE in "$TEST_ROOT"/*.file
+ do
+ cp "$FILE" . &&
+ "$TEST_ROOT/rot13.sh" <"$FILE" >"$FILE.rot13"
+ done &&
+
+ echo "*.file filter=protocol" >.gitattributes &&
+ filter_git add *.file .gitattributes &&
+ cat >expected.log <<-EOF &&
+ START
+ init handshake complete
+ IN: clean 1pkt_1__.file $(($S )) [OK] -- OUT: $(($S )) . [OK]
+ IN: clean 2pkt_1+1.file $(($S +1)) [OK] -- OUT: $(($S +1)) .. [OK]
+ IN: clean 2pkt_2-1.file $(($S*2-1)) [OK] -- OUT: $(($S*2-1)) .. [OK]
+ IN: clean 2pkt_2__.file $(($S*2 )) [OK] -- OUT: $(($S*2 )) .. [OK]
+ IN: clean 3pkt_2+1.file $(($S*2+1)) [OK] -- OUT: $(($S*2+1)) ... [OK]
+ STOP
+ EOF
+ test_cmp_count expected.log rot13-filter.log &&
+
+ rm -f *.file &&
+
+ filter_git checkout --quiet --no-progress -- *.file &&
+ cat >expected.log <<-EOF &&
+ START
+ init handshake complete
+ IN: smudge 1pkt_1__.file $(($S )) [OK] -- OUT: $(($S )) . [OK]
+ IN: smudge 2pkt_1+1.file $(($S +1)) [OK] -- OUT: $(($S +1)) .. [OK]
+ IN: smudge 2pkt_2-1.file $(($S*2-1)) [OK] -- OUT: $(($S*2-1)) .. [OK]
+ IN: smudge 2pkt_2__.file $(($S*2 )) [OK] -- OUT: $(($S*2 )) .. [OK]
+ IN: smudge 3pkt_2+1.file $(($S*2+1)) [OK] -- OUT: $(($S*2+1)) ... [OK]
+ STOP
+ EOF
+ test_cmp_exclude_clean expected.log rot13-filter.log &&
+
+ for FILE in *.file
+ do
+ test_cmp_committed_rot13 "$TEST_ROOT/$FILE" $FILE
+ done
+ )
+'
+
+test_expect_success PERL 'required process filter with clean error should fail' '
+ test_config_global filter.protocol.process "$TEST_DIRECTORY/t0021/rot13-filter.pl clean smudge" &&
+ test_config_global filter.protocol.required true &&
+ rm -rf repo &&
+ mkdir repo &&
+ (
+ cd repo &&
+ git init &&
+
+ echo "*.r filter=protocol" >.gitattributes &&
+
+ cp "$TEST_ROOT/test.o" test.r &&
+ echo "this is going to fail" >clean-write-fail.r &&
+ echo "content-test3-subdir" >test3.r &&
+
+ test_must_fail git add .
+ )
+'
+
+test_expect_success PERL 'process filter should restart after unexpected write failure' '
+ test_config_global filter.protocol.process "$TEST_DIRECTORY/t0021/rot13-filter.pl clean smudge" &&
+ rm -rf repo &&
+ mkdir repo &&
+ (
+ cd repo &&
+ git init &&
+
+ echo "*.r filter=protocol" >.gitattributes &&
+
+ cp "$TEST_ROOT/test.o" test.r &&
+ cp "$TEST_ROOT/test2.o" test2.r &&
+ echo "this is going to fail" >smudge-write-fail.o &&
+ cp smudge-write-fail.o smudge-write-fail.r &&
+
+ S=$(file_size test.r) &&
+ S2=$(file_size test2.r) &&
+ SF=$(file_size smudge-write-fail.r) &&
+
+ git add . &&
+ rm -f *.r &&
+
+ rm -f rot13-filter.log &&
+ git checkout --quiet --no-progress . 2>git-stderr.log &&
+
+ grep "smudge write error at" git-stderr.log &&
+ grep "error: external filter" git-stderr.log &&
+
+ cat >expected.log <<-EOF &&
+ START
+ init handshake complete
+ IN: smudge smudge-write-fail.r $SF [OK] -- OUT: $SF [WRITE FAIL]
+ START
+ init handshake complete
+ IN: smudge test.r $S [OK] -- OUT: $S . [OK]
+ IN: smudge test2.r $S2 [OK] -- OUT: $S2 . [OK]
+ STOP
+ EOF
+ test_cmp_exclude_clean expected.log rot13-filter.log &&
+
+ test_cmp_committed_rot13 "$TEST_ROOT/test.o" test.r &&
+ test_cmp_committed_rot13 "$TEST_ROOT/test2.o" test2.r &&
+
+ # Smudge failed
+ ! test_cmp smudge-write-fail.o smudge-write-fail.r &&
+ "$TEST_ROOT/rot13.sh" <smudge-write-fail.o >expected &&
+ git cat-file blob :smudge-write-fail.r >actual &&
+ test_cmp expected actual
+ )
+'
+
+test_expect_success PERL 'process filter should not be restarted if it signals an error' '
+ test_config_global filter.protocol.process "$TEST_DIRECTORY/t0021/rot13-filter.pl clean smudge" &&
+ rm -rf repo &&
+ mkdir repo &&
+ (
+ cd repo &&
+ git init &&
+
+ echo "*.r filter=protocol" >.gitattributes &&
+
+ cp "$TEST_ROOT/test.o" test.r &&
+ cp "$TEST_ROOT/test2.o" test2.r &&
+ echo "this will cause an error" >error.o &&
+ cp error.o error.r &&
+
+ S=$(file_size test.r) &&
+ S2=$(file_size test2.r) &&
+ SE=$(file_size error.r) &&
+
+ git add . &&
+ rm -f *.r &&
+
+ filter_git checkout --quiet --no-progress . &&
+ cat >expected.log <<-EOF &&
+ START
+ init handshake complete
+ IN: smudge error.r $SE [OK] -- OUT: 0 [ERROR]
+ IN: smudge test.r $S [OK] -- OUT: $S . [OK]
+ IN: smudge test2.r $S2 [OK] -- OUT: $S2 . [OK]
+ STOP
+ EOF
+ test_cmp_exclude_clean expected.log rot13-filter.log &&
+
+ test_cmp_committed_rot13 "$TEST_ROOT/test.o" test.r &&
+ test_cmp_committed_rot13 "$TEST_ROOT/test2.o" test2.r &&
+ test_cmp error.o error.r
+ )
+'
+
+test_expect_success PERL 'process filter signals abort once to abort processing of all future files' '
+ test_config_global filter.protocol.process "$TEST_DIRECTORY/t0021/rot13-filter.pl clean smudge" &&
+ rm -rf repo &&
+ mkdir repo &&
+ (
+ cd repo &&
+ git init &&
+
+ echo "*.r filter=protocol" >.gitattributes &&
+
+ cp "$TEST_ROOT/test.o" test.r &&
+ cp "$TEST_ROOT/test2.o" test2.r &&
+ echo "error this blob and all future blobs" >abort.o &&
+ cp abort.o abort.r &&
+
+ SA=$(file_size abort.r) &&
+
+ git add . &&
+ rm -f *.r &&
+
+ filter_git checkout --quiet --no-progress . &&
+ cat >expected.log <<-EOF &&
+ START
+ init handshake complete
+ IN: smudge abort.r $SA [OK] -- OUT: 0 [ABORT]
+ STOP
+ EOF
+ test_cmp_exclude_clean expected.log rot13-filter.log &&
+
+ test_cmp "$TEST_ROOT/test.o" test.r &&
+ test_cmp "$TEST_ROOT/test2.o" test2.r &&
+ test_cmp abort.o abort.r
+ )
+'
+
+test_expect_success PERL 'invalid process filter must fail (and not hang!)' '
+ test_config_global filter.protocol.process cat &&
+ test_config_global filter.protocol.required true &&
+ rm -rf repo &&
+ mkdir repo &&
+ (
+ cd repo &&
+ git init &&
+
+ echo "*.r filter=protocol" >.gitattributes &&
+
+ cp "$TEST_ROOT/test.o" test.r &&
+ test_must_fail git add . 2>git-stderr.log &&
+ grep "does not support filter protocol version" git-stderr.log
+ )
+'
+
test_done
diff --git a/t/t0021/rot13-filter.pl b/t/t0021/rot13-filter.pl
new file mode 100755
index 0000000..1a6959c
--- /dev/null
+++ b/t/t0021/rot13-filter.pl
@@ -0,0 +1,191 @@
+#!/usr/bin/perl
+#
+# Example implementation for the Git filter protocol version 2
+# See Documentation/gitattributes.txt, section "Filter Protocol"
+#
+# The script takes the list of supported protocol capabilities as
+# arguments ("clean", "smudge", etc).
+#
+# This implementation supports special test cases:
+# (1) If data with the pathname "clean-write-fail.r" is processed with
+# a "clean" operation then the write operation will die.
+# (2) If data with the pathname "smudge-write-fail.r" is processed with
+# a "smudge" operation then the write operation will die.
+# (3) If data with the pathname "error.r" is processed with any
+# operation then the filter signals that it cannot or does not want
+# to process the file.
+# (4) If data with the pathname "abort.r" is processed with any
+# operation then the filter signals that it cannot or does not want
+# to process the file and any file after that is processed with the
+# same command.
+#
+
+use strict;
+use warnings;
+
+my $MAX_PACKET_CONTENT_SIZE = 65516;
+my @capabilities = @ARGV;
+
+open my $debug, ">>", "rot13-filter.log" or die "cannot open log file: $!";
+
+sub rot13 {
+ my $str = shift;
+ $str =~ y/A-Za-z/N-ZA-Mn-za-m/;
+ return $str;
+}
+
+sub packet_bin_read {
+ my $buffer;
+ my $bytes_read = read STDIN, $buffer, 4;
+ if ( $bytes_read == 0 ) {
+ # EOF - Git stopped talking to us!
+ print $debug "STOP\n";
+ exit();
+ }
+ elsif ( $bytes_read != 4 ) {
+ die "invalid packet: '$buffer'";
+ }
+ my $pkt_size = hex($buffer);
+ if ( $pkt_size == 0 ) {
+ return ( 1, "" );
+ }
+ elsif ( $pkt_size > 4 ) {
+ my $content_size = $pkt_size - 4;
+ $bytes_read = read STDIN, $buffer, $content_size;
+ if ( $bytes_read != $content_size ) {
+ die "invalid packet ($content_size bytes expected; $bytes_read bytes read)";
+ }
+ return ( 0, $buffer );
+ }
+ else {
+ die "invalid packet size: $pkt_size";
+ }
+}
+
+sub packet_txt_read {
+ my ( $res, $buf ) = packet_bin_read();
+ unless ( $buf =~ s/\n$// ) {
+ die "A non-binary line MUST be terminated by an LF.";
+ }
+ return ( $res, $buf );
+}
+
+sub packet_bin_write {
+ my $buf = shift;
+ print STDOUT sprintf( "%04x", length($buf) + 4 );
+ print STDOUT $buf;
+ STDOUT->flush();
+}
+
+sub packet_txt_write {
+ packet_bin_write( $_[0] . "\n" );
+}
+
+sub packet_flush {
+ print STDOUT sprintf( "%04x", 0 );
+ STDOUT->flush();
+}
+
+print $debug "START\n";
+$debug->flush();
+
+( packet_txt_read() eq ( 0, "git-filter-client" ) ) || die "bad initialize";
+( packet_txt_read() eq ( 0, "version=2" ) ) || die "bad version";
+( packet_bin_read() eq ( 1, "" ) ) || die "bad version end";
+
+packet_txt_write("git-filter-server");
+packet_txt_write("version=2");
+
+( packet_txt_read() eq ( 0, "clean=true" ) ) || die "bad capability";
+( packet_txt_read() eq ( 0, "smudge=true" ) ) || die "bad capability";
+( packet_bin_read() eq ( 1, "" ) ) || die "bad capability end";
+
+foreach (@capabilities) {
+ packet_txt_write( $_ . "=true" );
+}
+packet_flush();
+print $debug "init handshake complete\n";
+$debug->flush();
+
+while (1) {
+ my ($command) = packet_txt_read() =~ /^command=([^=]+)$/;
+ print $debug "IN: $command";
+ $debug->flush();
+
+ my ($pathname) = packet_txt_read() =~ /^pathname=([^=]+)$/;
+ print $debug " $pathname";
+ $debug->flush();
+
+ # Flush
+ packet_bin_read();
+
+ my $input = "";
+ {
+ binmode(STDIN);
+ my $buffer;
+ my $done = 0;
+ while ( !$done ) {
+ ( $done, $buffer ) = packet_bin_read();
+ $input .= $buffer;
+ }
+ print $debug " " . length($input) . " [OK] -- ";
+ $debug->flush();
+ }
+
+ my $output;
+ if ( $pathname eq "error.r" or $pathname eq "abort.r" ) {
+ $output = "";
+ }
+ elsif ( $command eq "clean" and grep( /^clean$/, @capabilities ) ) {
+ $output = rot13($input);
+ }
+ elsif ( $command eq "smudge" and grep( /^smudge$/, @capabilities ) ) {
+ $output = rot13($input);
+ }
+ else {
+ die "bad command '$command'";
+ }
+
+ print $debug "OUT: " . length($output) . " ";
+ $debug->flush();
+
+ if ( $pathname eq "error.r" ) {
+ print $debug "[ERROR]\n";
+ $debug->flush();
+ packet_txt_write("status=error");
+ packet_flush();
+ }
+ elsif ( $pathname eq "abort.r" ) {
+ print $debug "[ABORT]\n";
+ $debug->flush();
+ packet_txt_write("status=abort");
+ packet_flush();
+ }
+ else {
+ packet_txt_write("status=success");
+ packet_flush();
+
+ if ( $pathname eq "${command}-write-fail.r" ) {
+ print $debug "[WRITE FAIL]\n";
+ $debug->flush();
+ die "${command} write error";
+ }
+
+ while ( length($output) > 0 ) {
+ my $packet = substr( $output, 0, $MAX_PACKET_CONTENT_SIZE );
+ packet_bin_write($packet);
+ # dots represent the number of packets
+ print $debug ".";
+ if ( length($output) > $MAX_PACKET_CONTENT_SIZE ) {
+ $output = substr( $output, $MAX_PACKET_CONTENT_SIZE );
+ }
+ else {
+ $output = "";
+ }
+ }
+ packet_flush();
+ print $debug " [OK]\n";
+ $debug->flush();
+ packet_flush();
+ }
+}
--
2.10.0
^ permalink raw reply related
* [PATCH v9 01/14] convert: quote filter names in error messages
From: larsxschneider @ 2016-10-04 12:59 UTC (permalink / raw)
To: git; +Cc: ramsay, jnareb, gitster, j6t, tboegi, peff, mlbright,
Lars Schneider
In-Reply-To: <20161004125947.67104-1-larsxschneider@gmail.com>
From: Lars Schneider <larsxschneider@gmail.com>
Git filter driver commands with spaces (e.g. `filter.sh foo`) are hard
to read in error messages. Quote them to improve the readability.
Signed-off-by: Lars Schneider <larsxschneider@gmail.com>
Signed-off-by: Junio C Hamano <gitster@pobox.com>
---
convert.c | 12 ++++++------
1 file changed, 6 insertions(+), 6 deletions(-)
diff --git a/convert.c b/convert.c
index 077f5e6..986c239 100644
--- a/convert.c
+++ b/convert.c
@@ -412,7 +412,7 @@ static int filter_buffer_or_fd(int in, int out, void *data)
child_process.out = out;
if (start_command(&child_process))
- return error("cannot fork to run external filter %s", params->cmd);
+ return error("cannot fork to run external filter '%s'", params->cmd);
sigchain_push(SIGPIPE, SIG_IGN);
@@ -430,13 +430,13 @@ static int filter_buffer_or_fd(int in, int out, void *data)
if (close(child_process.in))
write_err = 1;
if (write_err)
- error("cannot feed the input to external filter %s", params->cmd);
+ error("cannot feed the input to external filter '%s'", params->cmd);
sigchain_pop(SIGPIPE);
status = finish_command(&child_process);
if (status)
- error("external filter %s failed %d", params->cmd, status);
+ error("external filter '%s' failed %d", params->cmd, status);
strbuf_release(&cmd);
return (write_err || status);
@@ -477,15 +477,15 @@ static int apply_filter(const char *path, const char *src, size_t len, int fd,
return 0; /* error was already reported */
if (strbuf_read(&nbuf, async.out, len) < 0) {
- error("read from external filter %s failed", cmd);
+ error("read from external filter '%s' failed", cmd);
ret = 0;
}
if (close(async.out)) {
- error("read from external filter %s failed", cmd);
+ error("read from external filter '%s' failed", cmd);
ret = 0;
}
if (finish_async(&async)) {
- error("external filter %s failed", cmd);
+ error("external filter '%s' failed", cmd);
ret = 0;
}
--
2.10.0
^ permalink raw reply related
* [PATCH v9 00/14] Git filter protocol
From: larsxschneider @ 2016-10-04 12:59 UTC (permalink / raw)
To: git; +Cc: ramsay, jnareb, gitster, j6t, tboegi, peff, mlbright,
Lars Schneider
From: Lars Schneider <larsxschneider@gmail.com>
The goal of this series is to avoid launching a new clean/smudge filter
process for each file that is filtered.
A short summary about v1 to v5 can be found here:
https://git.github.io/rev_news/2016/08/17/edition-18/
This series is also published on web:
https://github.com/larsxschneider/git/pull/13
Patches 1 and 2 are cleanups and not strictly necessary for the series.
Patches 3 to 12 are required preparation. Patch 13 is the main patch.
Patch 14 adds an example how to use the Git filter protocol in contrib.
Thanks a lot to
Ramsay, Jakub, Junio, Johannes, Torsten, and Peff
for very helpful reviews,
Lars
## Major changes since v8
* refactor test cases to make them more robust and more readable
* split main "convert: add filter" patch into three:
1. preparation patch
2. main patch
3. contrib patch
* add run-command "wait_on_exit" flag to fix flaky t0021, see discussion:
http://public-inbox.org/git/xmqq8tubitjs.fsf@gitster.mtv.corp.google.com/
## All changes since v8
### Ramsay
* make async_exit()
http://public-inbox.org/git/78f2bdd0-f6ad-db5c-f9f2-f90528bc4f77@ramsayjones.plus.com/
### Jacub
* improve commit message for:
pkt-line: rename packet_write() to packet_write_fmt
pkt-line: extract set_packet_header()
* remove unnecessary brackets in packet_write_gently
* introduce variable 'packet_size' to make packet_write_gently()
more readable
* replace 'PKTLINE_DATA_MAXLEN' with 'LARGE_PACKET_DATA_MAX'
* rename 'read_packetized_to_buf' to 'read_packetized_to_strbuf'
* use 'ssize_t' as proper return value for write_in_full()
* rename 'paket_len' to 'packet_len'
* remove unnecessary options variable in read_packetized_to_strbuf()
* rename variables in read_packetized_to_strbuf()
* fix +1 in read_packetized_to_strbuf()
* fix length comparison in packet_read()
* replace spaces with tabs in Perl scripts
* ... many smaller changes. See interdiff below.
## Interdiff (v8..v9)
```diff
diff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt
index 946dcad..a182ef2 100644
--- a/Documentation/gitattributes.txt
+++ b/Documentation/gitattributes.txt
@@ -387,9 +387,9 @@ If the filter command (a string value) is defined via
single filter invocation for the entire life of a single Git
command. This is achieved by using a packet format (pkt-line,
see technical/protocol-common.txt) based protocol over standard
-input and standard output as follows. All packets are considered
-text and therefore are terminated by an LF. Exceptions are the
-"*CONTENT" packets and the flush packet.
+input and standard output as follows. All packets, except for the
+"*CONTENT" packets and the "0000" flush packet, are considered
+text and therefore are terminated by a LF.
Git starts the filter when it encounters the first file
that needs to be cleaned or smudged. After the filter started
@@ -403,10 +403,10 @@ protocol description below documents "version=2". Please note that
to illustrate how the protocol would look like with more than one
version.
-After the version negotiation Git sends a list of supported capabilities
-and a flush packet. Git expects to read a list of desired capabilities,
-which must be a subset of the supported capabilities list, and a flush
-packet as response:
+After the version negotiation Git sends a list of all capabilities that
+it supports and a flush packet. Git expects to read a list of desired
+capabilities, which must be a subset of the supported capabilities list,
+and a flush packet as response:
------------------------
packet: git> git-filter-client
packet: git> version=2
@@ -430,7 +430,9 @@ a flush packet. The list will contain at least the filter command
(based on the supported capabilities) and the pathname of the file
to filter relative to the repository root. Right after these packets
Git sends the content split in zero or more pkt-line packets and a
-flush packet to terminate content.
+flush packet to terminate content. Please note, that the filter
+must not send any response before it received the content and the
+final flush packet.
------------------------
packet: git> command=smudge
packet: git> pathname=path/testfile.dat
@@ -451,16 +453,16 @@ packet: git< status=success
packet: git< 0000
packet: git< SMUDGED_CONTENT
packet: git< 0000
-packet: git< 0000 # empty list!
+packet: git< 0000 # empty list, keep "status=success" unchanged!
------------------------
If the result content is empty then the filter is expected to respond
-with a success status and an empty list.
+with a "success" status and an empty list.
------------------------
packet: git< status=success
packet: git< 0000
packet: git< 0000 # empty content!
-packet: git< 0000 # empty list!
+packet: git< 0000 # empty list, keep "status=success" unchanged!
------------------------
In case the filter cannot or does not want to process the content,
@@ -497,10 +499,11 @@ handling.
In case the filter cannot or does not want to process the content
as well as any future content for the lifetime of the Git process,
-it is expected to respond with an "abort" status. Depending on
-the `filter.<driver>.required` flag Git will interpret that as error
-for the content as well as any future content for the lifetime of the
-Git process but it will not stop or restart the filter process.
+it is expected to respond with an "abort" status at any point in
+the protocol. Depending on the `filter.<driver>.required` flag Git
+will interpret that as error for the content as well as any future
+content for the lifetime of the Git process but it will not stop or
+restart the filter process.
------------------------
packet: git< status=abort
packet: git< 0000
diff --git a/contrib/long-running-filter/example.pl b/contrib/long-running-filter/example.pl
index c13a631..f4102d2 100755
--- a/contrib/long-running-filter/example.pl
+++ b/contrib/long-running-filter/example.pl
@@ -3,6 +3,9 @@
# Example implementation for the Git filter protocol version 2
# See Documentation/gitattributes.txt, section "Filter Protocol"
#
+# Please note, this pass-thru filter is a minimal skeleton. No proper
+# error handling was implemented.
+#
use strict;
use warnings;
@@ -18,7 +21,7 @@ sub packet_bin_read {
exit();
}
elsif ( $bytes_read != 4 ) {
- die "invalid packet size '$bytes_read' field";
+ die "invalid packet: '$buffer'";
}
my $pkt_size = hex($buffer);
if ( $pkt_size == 0 ) {
@@ -28,27 +31,27 @@ sub packet_bin_read {
my $content_size = $pkt_size - 4;
$bytes_read = read STDIN, $buffer, $content_size;
if ( $bytes_read != $content_size ) {
- die "invalid packet ($content_size expected; $bytes_read read)";
+ die "invalid packet ($content_size bytes expected; $bytes_read bytes read)";
}
return ( 0, $buffer );
}
else {
- die "invalid packet size";
+ die "invalid packet size: $pkt_size";
}
}
sub packet_txt_read {
my ( $res, $buf ) = packet_bin_read();
- unless ( $buf =~ /\n$/ ) {
- die "A non-binary line SHOULD BE terminated by an LF.";
+ unless ( $buf =~ s/\n$// ) {
+ die "A non-binary line MUST be terminated by an LF.";
}
- return ( $res, substr( $buf, 0, -1 ) );
+ return ( $res, $buf );
}
sub packet_bin_write {
- my ($packet) = @_;
- print STDOUT sprintf( "%04x", length($packet) + 4 );
- print STDOUT $packet;
+ my $buf = shift;
+ print STDOUT sprintf( "%04x", length($buf) + 4 );
+ print STDOUT $buf;
STDOUT->flush();
}
@@ -119,5 +122,6 @@ while (1) {
}
}
packet_flush(); # flush content!
- packet_flush(); # empty list!
+ packet_flush(); # empty list, keep "status=success" unchanged!
+
}
diff --git a/convert.c b/convert.c
index bd66257..88581d6 100644
--- a/convert.c
+++ b/convert.c
@@ -541,7 +541,7 @@ static int packet_write_list(int fd, const char *line, ...)
for (;;) {
if (!line)
break;
- if (strlen(line) > PKTLINE_DATA_MAXLEN)
+ if (strlen(line) > LARGE_PACKET_DATA_MAX)
return -1;
err = packet_write_fmt_gently(fd, "%s\n", line);
if (err)
@@ -573,6 +573,7 @@ static struct cmd2process *start_multi_file_filter(struct hashmap *hashmap, cons
process->use_shell = 1;
process->in = -1;
process->out = -1;
+ process->wait_on_exit = 1;
if (start_command(process)) {
error("cannot fork to run external filter '%s'", cmd);
@@ -588,7 +589,7 @@ static struct cmd2process *start_multi_file_filter(struct hashmap *hashmap, cons
err = strcmp(packet_read_line(process->out, NULL), "git-filter-server");
if (err) {
- error("external filter '%s' does not support long running filter protocol", cmd);
+ error("external filter '%s' does not support filter protocol version 2", cmd);
goto done;
}
err = strcmp(packet_read_line(process->out, NULL), "version=2");
@@ -634,7 +635,7 @@ static struct cmd2process *start_multi_file_filter(struct hashmap *hashmap, cons
return entry;
}
-static void read_multi_file_filter_values(int fd, struct strbuf *status) {
+static void read_multi_file_filter_status(int fd, struct strbuf *status) {
struct strbuf **pair;
char *line;
for (;;) {
@@ -648,6 +649,7 @@ static void read_multi_file_filter_values(int fd, struct strbuf *status) {
strbuf_addbuf(status, pair[1]);
}
}
+ strbuf_list_free(pair);
}
}
@@ -698,17 +700,16 @@ static int apply_multi_file_filter(const char *path, const char *src, size_t len
sigchain_push(SIGPIPE, SIG_IGN);
- err = strlen(filter_type) > PKTLINE_DATA_MAXLEN;
- if (err)
- goto done;
-
+ assert(strlen(filter_type) < LARGE_PACKET_DATA_MAX - strlen("command=\n"));
err = packet_write_fmt_gently(process->in, "command=%s\n", filter_type);
if (err)
goto done;
- err = strlen(path) > PKTLINE_DATA_MAXLEN;
- if (err)
+ err = strlen(path) > LARGE_PACKET_DATA_MAX - strlen("pathname=\n");
+ if (err) {
+ error("path name too long for external filter");
goto done;
+ }
err = packet_write_fmt_gently(process->in, "pathname=%s\n", path);
if (err)
@@ -725,16 +726,16 @@ static int apply_multi_file_filter(const char *path, const char *src, size_t len
if (err)
goto done;
- read_multi_file_filter_values(process->out, &filter_status);
+ read_multi_file_filter_status(process->out, &filter_status);
err = strcmp(filter_status.buf, "success");
if (err)
goto done;
- err = read_packetized_to_buf(process->out, &nbuf) < 0;
+ err = read_packetized_to_strbuf(process->out, &nbuf) < 0;
if (err)
goto done;
- read_multi_file_filter_values(process->out, &filter_status);
+ read_multi_file_filter_status(process->out, &filter_status);
err = strcmp(filter_status.buf, "success");
done:
@@ -753,7 +754,7 @@ static int apply_multi_file_filter(const char *path, const char *src, size_t len
} else {
/*
* Something went wrong with the protocol filter.
- * Force shutdown and restart if another blob requires filtering!
+ * Force shutdown and restart if another blob requires filtering.
*/
error("external filter '%s' failed", cmd);
kill_multi_file_filter(&cmd_process_map, entry);
@@ -786,9 +787,9 @@ static int apply_filter(const char *path, const char *src, size_t len,
if (!dst)
return 1;
- if (!drv->process && (CAP_CLEAN & wanted_capability) && drv->clean)
+ if ((CAP_CLEAN & wanted_capability) && !drv->process && drv->clean)
cmd = drv->clean;
- else if (!drv->process && (CAP_SMUDGE & wanted_capability) && drv->smudge)
+ else if ((CAP_SMUDGE & wanted_capability) && !drv->process && drv->smudge)
cmd = drv->smudge;
if (cmd && *cmd)
diff --git a/pkt-line.c b/pkt-line.c
index a0a8543..b82aaca 100644
--- a/pkt-line.c
+++ b/pkt-line.c
@@ -137,7 +137,7 @@ static int packet_write_fmt_1(int fd, int gently,
const char *fmt, va_list args)
{
struct strbuf buf = STRBUF_INIT;
- size_t count;
+ ssize_t count;
format_packet(&buf, fmt, args);
count = write_in_full(fd, buf.buf, buf.len);
@@ -174,15 +174,15 @@ int packet_write_fmt_gently(int fd, const char *fmt, ...)
static int packet_write_gently(const int fd_out, const char *buf, size_t size)
{
static char packet_write_buffer[LARGE_PACKET_MAX];
+ const size_t packet_size = size + 4;
- if (size > sizeof(packet_write_buffer) - 4) {
+ if (packet_size > sizeof(packet_write_buffer))
return error("packet write failed - data exceeds max packet size");
- }
+
packet_trace(buf, size, 1);
- size += 4;
- set_packet_header(packet_write_buffer, size);
- memcpy(packet_write_buffer + 4, buf, size - 4);
- if (write_in_full(fd_out, packet_write_buffer, size) == size)
+ set_packet_header(packet_write_buffer, packet_size);
+ memcpy(packet_write_buffer + 4, buf, size);
+ if (write_in_full(fd_out, packet_write_buffer, packet_size) == packet_size)
return 0;
return error("packet write failed");
}
@@ -198,7 +198,7 @@ void packet_buf_write(struct strbuf *buf, const char *fmt, ...)
int write_packetized_from_fd(int fd_in, int fd_out)
{
- static char buf[PKTLINE_DATA_MAXLEN];
+ static char buf[LARGE_PACKET_DATA_MAX];
int err = 0;
ssize_t bytes_to_write;
@@ -217,7 +217,7 @@ int write_packetized_from_fd(int fd_in, int fd_out)
int write_packetized_from_buf(const char *src_in, size_t len, int fd_out)
{
- static char buf[PKTLINE_DATA_MAXLEN];
+ static char buf[LARGE_PACKET_DATA_MAX];
int err = 0;
size_t bytes_written = 0;
size_t bytes_to_write;
@@ -347,29 +347,34 @@ char *packet_read_line_buf(char **src, size_t *src_len, int *dst_len)
return packet_read_line_generic(-1, src, src_len, dst_len);
}
-ssize_t read_packetized_to_buf(int fd_in, struct strbuf *sb_out)
+ssize_t read_packetized_to_strbuf(int fd_in, struct strbuf *sb_out)
{
- int paket_len;
- int options = PACKET_READ_GENTLE_ON_EOF;
+ int packet_len;
- size_t oldlen = sb_out->len;
- size_t oldalloc = sb_out->alloc;
+ size_t orig_len = sb_out->len;
+ size_t orig_alloc = sb_out->alloc;
for (;;) {
- strbuf_grow(sb_out, PKTLINE_DATA_MAXLEN+1);
- paket_len = packet_read(fd_in, NULL, NULL,
- sb_out->buf + sb_out->len, PKTLINE_DATA_MAXLEN+1, options);
- if (paket_len <= 0)
+ strbuf_grow(sb_out, LARGE_PACKET_DATA_MAX);
+ packet_len = packet_read(fd_in, NULL, NULL,
+ /* strbuf_grow() above always allocates one extra byte to
+ * store a '\0' at the end of the string. packet_read()
+ * writes a '\0' extra byte at the end, too. Let it know
+ * that there is already room for the extra byte.
+ */
+ sb_out->buf + sb_out->len, LARGE_PACKET_DATA_MAX+1,
+ PACKET_READ_GENTLE_ON_EOF);
+ if (packet_len <= 0)
break;
- sb_out->len += paket_len;
+ sb_out->len += packet_len;
}
- if (paket_len < 0) {
- if (oldalloc == 0)
+ if (packet_len < 0) {
+ if (orig_alloc == 0)
strbuf_release(sb_out);
else
- strbuf_setlen(sb_out, oldlen);
- return paket_len;
+ strbuf_setlen(sb_out, orig_len);
+ return packet_len;
}
- return sb_out->len - oldlen;
+ return sb_out->len - orig_len;
}
diff --git a/pkt-line.h b/pkt-line.h
index 3d873f3..18eac64 100644
--- a/pkt-line.h
+++ b/pkt-line.h
@@ -82,11 +82,11 @@ char *packet_read_line_buf(char **src_buf, size_t *src_len, int *size);
/*
* Reads a stream of variable sized packets until a flush packet is detected.
*/
-ssize_t read_packetized_to_buf(int fd_in, struct strbuf *sb_out);
+ssize_t read_packetized_to_strbuf(int fd_in, struct strbuf *sb_out);
#define DEFAULT_PACKET_MAX 1000
#define LARGE_PACKET_MAX 65520
-#define PKTLINE_DATA_MAXLEN (LARGE_PACKET_MAX - 4)
+#define LARGE_PACKET_DATA_MAX (LARGE_PACKET_MAX - 4)
extern char packet_buffer[LARGE_PACKET_MAX];
#endif
diff --git a/run-command.c b/run-command.c
index b72f6d1..96c54fe 100644
--- a/run-command.c
+++ b/run-command.c
@@ -6,19 +6,6 @@
#include "thread-utils.h"
#include "strbuf.h"
-void check_pipe(int err)
-{
- if (err == EPIPE) {
- if (in_async())
- async_exit(141);
-
- signal(SIGPIPE, SIG_DFL);
- raise(SIGPIPE);
- /* Should never happen, but just in case... */
- exit(141);
- }
-}
-
void child_process_init(struct child_process *child)
{
memset(child, 0, sizeof(*child));
@@ -34,6 +21,9 @@ void child_process_clear(struct child_process *child)
struct child_to_clean {
pid_t pid;
+ char *name;
+ int stdin;
+ int wait;
struct child_to_clean *next;
};
static struct child_to_clean *children_to_clean;
@@ -41,14 +31,35 @@ static int installed_child_cleanup_handler;
static void cleanup_children(int sig, int in_signal)
{
- while (children_to_clean) {
+ int status;
struct child_to_clean *p = children_to_clean;
+
+ /* Close the the child's stdin as indicator that Git will exit soon */
+ while (p) {
+ if (p->wait)
+ if (p->stdin > 0)
+ close(p->stdin);
+ p = p->next;
+ }
+
+ while (children_to_clean) {
+ p = children_to_clean;
children_to_clean = p->next;
+
+ if (p->wait) {
+ fprintf(stderr, _("Waiting for '%s' to finish..."), p->name);
+ while ((waitpid(p->pid, &status, 0)) < 0 && errno == EINTR)
+ ; /* nothing */
+ fprintf(stderr, _("done!\n"));
+ }
+
kill(p->pid, sig);
- if (!in_signal)
+ if (!in_signal) {
+ free(p->name);
free(p);
}
}
+}
static void cleanup_children_on_signal(int sig)
{
@@ -62,10 +73,16 @@ static void cleanup_children_on_exit(void)
cleanup_children(SIGTERM, 0);
}
-static void mark_child_for_cleanup(pid_t pid)
+static void mark_child_for_cleanup(pid_t pid, const char *name, int stdin, int wait)
{
struct child_to_clean *p = xmalloc(sizeof(*p));
p->pid = pid;
+ p->wait = wait;
+ p->stdin = stdin;
+ if (name)
+ p->name = xstrdup(name);
+ else
+ p->name = "process";
p->next = children_to_clean;
children_to_clean = p;
@@ -76,6 +93,13 @@ static void mark_child_for_cleanup(pid_t pid)
}
}
+#ifdef NO_PTHREADS
+static void mark_child_for_cleanup_no_wait(pid_t pid, const char *name, int timeout, int stdin)
+{
+ mark_child_for_cleanup(pid, NULL, 0, 0);
+}
+#endif
+
static void clear_child_for_cleanup(pid_t pid)
{
struct child_to_clean **pp;
@@ -434,8 +458,9 @@ int start_command(struct child_process *cmd)
}
if (cmd->pid < 0)
error_errno("cannot fork() for %s", cmd->argv[0]);
- else if (cmd->clean_on_exit)
- mark_child_for_cleanup(cmd->pid);
+ else if (cmd->clean_on_exit || cmd->wait_on_exit)
+ mark_child_for_cleanup(
+ cmd->pid, cmd->argv[0], cmd->in, cmd->wait_on_exit);
/*
* Wait for child's execvp. If the execvp succeeds (or if fork()
@@ -495,8 +520,9 @@ int start_command(struct child_process *cmd)
failed_errno = errno;
if (cmd->pid < 0 && (!cmd->silent_exec_failure || errno != ENOENT))
error_errno("cannot spawn %s", cmd->argv[0]);
- if (cmd->clean_on_exit && cmd->pid >= 0)
- mark_child_for_cleanup(cmd->pid);
+ if ((cmd->clean_on_exit || cmd->wait_on_exit) && cmd->pid >= 0)
+ mark_child_for_cleanup(
+ cmd->pid, cmd->argv[0], cmd->in, cmd->clean_on_exit_timeout);
argv_array_clear(&nargv);
cmd->argv = sargv;
@@ -647,7 +673,7 @@ int in_async(void)
return !pthread_equal(main_thread, pthread_self());
}
-void NORETURN async_exit(int code)
+static void NORETURN async_exit(int code)
{
pthread_exit((void *)(intptr_t)code);
}
@@ -697,13 +723,26 @@ int in_async(void)
return process_is_async;
}
-void NORETURN async_exit(int code)
+static void NORETURN async_exit(int code)
{
exit(code);
}
#endif
+void check_pipe(int err)
+{
+ if (err == EPIPE) {
+ if (in_async())
+ async_exit(141);
+
+ signal(SIGPIPE, SIG_DFL);
+ raise(SIGPIPE);
+ /* Should never happen, but just in case... */
+ exit(141);
+ }
+}
+
int start_async(struct async *async)
{
int need_in, need_out;
@@ -765,7 +804,7 @@ int start_async(struct async *async)
exit(!!async->proc(proc_in, proc_out, async->data));
}
- mark_child_for_cleanup(async->pid);
+ mark_child_for_cleanup_no_wait(async->pid);
if (need_in)
close(fdin[0]);
diff --git a/run-command.h b/run-command.h
index e7c5f71..f7b9907 100644
--- a/run-command.h
+++ b/run-command.h
@@ -42,7 +42,10 @@ struct child_process {
unsigned silent_exec_failure:1;
unsigned stdout_to_stderr:1;
unsigned use_shell:1;
+ /* kill the child on Git exit */
unsigned clean_on_exit:1;
+ /* close the child's stdin on Git exit and wait until it terminates */
+ unsigned wait_on_exit:1;
};
#define CHILD_PROCESS_INIT { NULL, ARGV_ARRAY_INIT, ARGV_ARRAY_INIT }
@@ -54,8 +57,6 @@ int finish_command(struct child_process *);
int finish_command_in_signal(struct child_process *);
int run_command(struct child_process *);
-void check_pipe(int err);
-
/*
* Returns the path to the hook file, or NULL if the hook is missing
* or disabled. Note that this points to static storage that will be
@@ -141,7 +142,7 @@ struct async {
int start_async(struct async *async);
int finish_async(struct async *async);
int in_async(void);
-void NORETURN async_exit(int code);
+void check_pipe(int err);
/**
* This callback should initialize the child process and preload the
diff --git a/t/t0021-conversion.sh b/t/t0021-conversion.sh
index 210c4f6..52b7fe9 100755
--- a/t/t0021-conversion.sh
+++ b/t/t0021-conversion.sh
@@ -4,13 +4,77 @@ test_description='blob conversion via gitattributes'
. ./test-lib.sh
-cat <<EOF >rot13.sh
+TEST_ROOT="$(pwd)"
+
+cat <<EOF >"$TEST_ROOT/rot13.sh"
#!$SHELL_PATH
tr \
'abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ' \
'nopqrstuvwxyzabcdefghijklmNOPQRSTUVWXYZABCDEFGHIJKLM'
EOF
-chmod +x rot13.sh
+chmod +x "$TEST_ROOT/rot13.sh"
+
+generate_random_characters () {
+ LEN=$1
+ NAME=$2
+ test-genrandom some-seed $LEN |
+ perl -pe "s/./chr((ord($&) % 26) + ord('a'))/sge" >"$TEST_ROOT/$NAME"
+}
+
+file_size () {
+ cat "$1" | wc -c | sed "s/^[ ]*//"
+}
+
+filter_git () {
+ rm -f rot13-filter.log &&
+ git "$@" 2>git-stderr.log &&
+ sed '/Waiting for/d' git-stderr.log >git-stderr-clean.log &&
+ test_must_be_empty git-stderr-clean.log &&
+ rm -f git-stderr.log git-stderr-clean.log
+}
+
+# Count unique lines in two files and compare them.
+test_cmp_count () {
+ for FILE in $@
+ do
+ sort $FILE | uniq -c | sed "s/^[ ]*//" >$FILE.tmp
+ cat $FILE.tmp >$FILE
+ done &&
+ test_cmp $@
+}
+
+# Count unique lines except clean invocations in two files and compare
+# them. Clean invocations are not counted because their number can vary.
+# c.f. http://public-inbox.org/git/xmqqshv18i8i.fsf@gitster.mtv.corp.google.com/
+test_cmp_count_except_clean () {
+ for FILE in $@
+ do
+ sort $FILE | uniq -c | sed "s/^[ ]*//" |
+ sed "s/^\([0-9]\) IN: clean/x IN: clean/" >$FILE.tmp
+ cat $FILE.tmp >$FILE
+ done &&
+ test_cmp $@
+}
+
+# Compare two files but exclude clean invocations because they can vary.
+# c.f. http://public-inbox.org/git/xmqqshv18i8i.fsf@gitster.mtv.corp.google.com/
+test_cmp_exclude_clean () {
+ for FILE in $@
+ do
+ grep -v "IN: clean" $FILE >$FILE.tmp
+ cat $FILE.tmp >$FILE
+ done &&
+ test_cmp $@
+}
+
+# Check that the contents of two files are equal and that their rot13 version
+# is equal to the committed content.
+test_cmp_committed_rot13 () {
+ test_cmp "$1" "$2" &&
+ "$TEST_ROOT/rot13.sh" <"$1" >expected &&
+ git cat-file blob :"$2" >actual &&
+ test_cmp expected actual
+}
test_expect_success setup '
git config filter.rot13.smudge ./rot13.sh &&
@@ -34,7 +98,7 @@ test_expect_success setup '
git checkout -- test test.t test.i &&
echo "content-test2" >test2.o &&
- echo "content-test3 - subdir" >"test3 - subdir.o"
+ echo "content-test3 - filename with special characters" >"test3 '\''sq'\'',\$x.o"
'
script='s/^\$Id: \([0-9a-f]*\) \$/\1/p'
@@ -282,47 +346,6 @@ test_expect_success 'diff does not reuse worktree files that need cleaning' '
test_line_count = 0 count
'
-check_filter () {
- rm -f rot13-filter.log actual.log &&
- "$@" 2> git_stderr.log &&
- test_must_be_empty git_stderr.log &&
- cat >expected.log &&
- sort rot13-filter.log | uniq -c | sed "s/^[ ]*//" >actual.log &&
- test_cmp expected.log actual.log
-}
-
-check_filter_count_clean () {
- rm -f rot13-filter.log actual.log &&
- "$@" 2> git_stderr.log &&
- test_must_be_empty git_stderr.log &&
- cat >expected.log &&
- sort rot13-filter.log | uniq -c | sed "s/^[ ]*//" |
- sed "s/^\([0-9]\) IN: clean/x IN: clean/" >actual.log &&
- test_cmp expected.log actual.log
-}
-
-check_filter_ignore_clean () {
- rm -f rot13-filter.log actual.log &&
- "$@" &&
- cat >expected.log &&
- grep -v "IN: clean" rot13-filter.log >actual.log &&
- test_cmp expected.log actual.log
-}
-
-check_filter_no_call () {
- rm -f rot13-filter.log &&
- "$@" 2> git_stderr.log &&
- test_must_be_empty git_stderr.log &&
- test_must_be_empty rot13-filter.log
-}
-
-check_rot13 () {
- test_cmp "$1" "$2" &&
- ./../rot13.sh <"$1" >expected &&
- git cat-file blob :"$2" >actual &&
- test_cmp expected actual
-}
-
test_expect_success PERL 'required process filter should filter data' '
test_config_global filter.protocol.process "$TEST_DIRECTORY/t0021/rot13-filter.pl clean smudge" &&
test_config_global filter.protocol.required true &&
@@ -332,81 +355,91 @@ test_expect_success PERL 'required process filter should filter data' '
cd repo &&
git init &&
+ echo "git-stderr.log" >.gitignore &&
echo "*.r filter=protocol" >.gitattributes &&
git add . &&
- git commit . -m "test commit" &&
- git branch empty &&
+ git commit . -m "test commit 1" &&
+ git branch empty-branch &&
- cp ../test.o test.r &&
- cp ../test2.o test2.r &&
+ cp "$TEST_ROOT/test.o" test.r &&
+ cp "$TEST_ROOT/test2.o" test2.r &&
mkdir testsubdir &&
- cp "../test3 - subdir.o" "testsubdir/test3 - subdir.r" &&
+ cp "$TEST_ROOT/test3 '\''sq'\'',\$x.o" "testsubdir/test3 '\''sq'\'',\$x.r" &&
>test4-empty.r &&
- check_filter \
- git add . \
- <<-\EOF &&
- 1 IN: clean test.r 57 [OK] -- OUT: 57 . [OK]
- 1 IN: clean test2.r 14 [OK] -- OUT: 14 . [OK]
- 1 IN: clean test4-empty.r 0 [OK] -- OUT: 0 [OK]
- 1 IN: clean testsubdir/test3 - subdir.r 23 [OK] -- OUT: 23 . [OK]
- 1 START
- 1 STOP
- 1 wrote filter header
+ S=$(file_size test.r) &&
+ S2=$(file_size test2.r) &&
+ S3=$(file_size "testsubdir/test3 '\''sq'\'',\$x.r") &&
+
+ filter_git add . &&
+ cat >expected.log <<-EOF &&
+ START
+ init handshake complete
+ IN: clean test.r $S [OK] -- OUT: $S . [OK]
+ IN: clean test2.r $S2 [OK] -- OUT: $S2 . [OK]
+ IN: clean test4-empty.r 0 [OK] -- OUT: 0 [OK]
+ IN: clean testsubdir/test3 '\''sq'\'',\$x.r $S3 [OK] -- OUT: $S3 . [OK]
+ STOP
EOF
+ test_cmp_count expected.log rot13-filter.log &&
- check_filter_count_clean \
- git commit . -m "test commit" \
- <<-\EOF &&
- x IN: clean test.r 57 [OK] -- OUT: 57 . [OK]
- x IN: clean test2.r 14 [OK] -- OUT: 14 . [OK]
- x IN: clean test4-empty.r 0 [OK] -- OUT: 0 [OK]
- x IN: clean testsubdir/test3 - subdir.r 23 [OK] -- OUT: 23 . [OK]
- 1 START
- 1 STOP
- 1 wrote filter header
+ filter_git commit . -m "test commit 2" &&
+ cat >expected.log <<-EOF &&
+ START
+ init handshake complete
+ IN: clean test.r $S [OK] -- OUT: $S . [OK]
+ IN: clean test2.r $S2 [OK] -- OUT: $S2 . [OK]
+ IN: clean test4-empty.r 0 [OK] -- OUT: 0 [OK]
+ IN: clean testsubdir/test3 '\''sq'\'',\$x.r $S3 [OK] -- OUT: $S3 . [OK]
+ IN: clean test.r $S [OK] -- OUT: $S . [OK]
+ IN: clean test2.r $S2 [OK] -- OUT: $S2 . [OK]
+ IN: clean test4-empty.r 0 [OK] -- OUT: 0 [OK]
+ IN: clean testsubdir/test3 '\''sq'\'',\$x.r $S3 [OK] -- OUT: $S3 . [OK]
+ STOP
EOF
+ test_cmp_count_except_clean expected.log rot13-filter.log &&
- rm -f test?.r "testsubdir/test3 - subdir.r" &&
+ rm -f test2.r "testsubdir/test3 '\''sq'\'',\$x.r" &&
- check_filter_ignore_clean \
- git checkout . \
- <<-\EOF &&
+ filter_git checkout --quiet --no-progress . &&
+ cat >expected.log <<-EOF &&
START
- wrote filter header
- IN: smudge test2.r 14 [OK] -- OUT: 14 . [OK]
- IN: smudge testsubdir/test3 - subdir.r 23 [OK] -- OUT: 23 . [OK]
+ init handshake complete
+ IN: smudge test2.r $S2 [OK] -- OUT: $S2 . [OK]
+ IN: smudge testsubdir/test3 '\''sq'\'',\$x.r $S3 [OK] -- OUT: $S3 . [OK]
STOP
EOF
+ test_cmp_exclude_clean expected.log rot13-filter.log &&
- check_filter_ignore_clean \
- git checkout empty \
- <<-\EOF &&
+ filter_git checkout --quiet --no-progress empty-branch &&
+ cat >expected.log <<-EOF &&
START
- wrote filter header
+ init handshake complete
+ IN: clean test.r $S [OK] -- OUT: $S . [OK]
STOP
EOF
+ test_cmp_exclude_clean expected.log rot13-filter.log &&
- check_filter_ignore_clean \
- git checkout master \
- <<-\EOF &&
+ filter_git checkout --quiet --no-progress master &&
+ cat >expected.log <<-EOF &&
START
- wrote filter header
- IN: smudge test.r 57 [OK] -- OUT: 57 . [OK]
- IN: smudge test2.r 14 [OK] -- OUT: 14 . [OK]
+ init handshake complete
+ IN: smudge test.r $S [OK] -- OUT: $S . [OK]
+ IN: smudge test2.r $S2 [OK] -- OUT: $S2 . [OK]
IN: smudge test4-empty.r 0 [OK] -- OUT: 0 [OK]
- IN: smudge testsubdir/test3 - subdir.r 23 [OK] -- OUT: 23 . [OK]
+ IN: smudge testsubdir/test3 '\''sq'\'',\$x.r $S3 [OK] -- OUT: $S3 . [OK]
STOP
EOF
+ test_cmp_exclude_clean expected.log rot13-filter.log &&
- check_rot13 ../test.o test.r &&
- check_rot13 ../test2.o test2.r &&
- check_rot13 "../test3 - subdir.o" "testsubdir/test3 - subdir.r"
+ test_cmp_committed_rot13 "$TEST_ROOT/test.o" test.r &&
+ test_cmp_committed_rot13 "$TEST_ROOT/test2.o" test2.r &&
+ test_cmp_committed_rot13 "$TEST_ROOT/test3 '\''sq'\'',\$x.o" "testsubdir/test3 '\''sq'\'',\$x.r"
)
'
-test_expect_success PERL 'required process filter should clean only and take precedence' '
- test_config_global filter.protocol.clean ./../rot13.sh &&
+test_expect_success PERL 'required process filter takes precedence' '
+ test_config_global filter.protocol.clean false &&
test_config_global filter.protocol.process "$TEST_DIRECTORY/t0021/rot13-filter.pl clean" &&
test_config_global filter.protocol.required true &&
rm -rf repo &&
@@ -416,40 +449,55 @@ test_expect_success PERL 'required process filter should clean only and take pre
git init &&
echo "*.r filter=protocol" >.gitattributes &&
- git add . &&
- git commit . -m "test commit" &&
- git branch empty &&
-
- cp ../test.o test.r &&
-
- check_filter \
- git add . \
- <<-\EOF &&
- 1 IN: clean test.r 57 [OK] -- OUT: 57 . [OK]
- 1 START
- 1 STOP
- 1 wrote filter header
- EOF
+ cp "$TEST_ROOT/test.o" test.r &&
+ S=$(file_size test.r) &&
- check_filter_count_clean \
- git commit . -m "test commit" \
- <<-\EOF
- x IN: clean test.r 57 [OK] -- OUT: 57 . [OK]
- 1 START
- 1 STOP
- 1 wrote filter header
+ # Check that the process filter is invoked here
+ filter_git add . &&
+ cat >expected.log <<-EOF &&
+ START
+ init handshake complete
+ IN: clean test.r $S [OK] -- OUT: $S . [OK]
+ STOP
EOF
+ test_cmp_count expected.log rot13-filter.log
)
'
-generate_test_data () {
- LEN=$1
- NAME=$2
- test-genrandom end $LEN |
- perl -pe "s/./chr((ord($&) % 26) + 97)/sge" >../$NAME.file &&
- cp ../$NAME.file . &&
- ./../rot13.sh <../$NAME.file >../$NAME.file.rot13
-}
+test_expect_success PERL 'required process filter should be used only for "clean" operation only' '
+ test_config_global filter.protocol.process "$TEST_DIRECTORY/t0021/rot13-filter.pl clean" &&
+ rm -rf repo &&
+ mkdir repo &&
+ (
+ cd repo &&
+ git init &&
+
+ echo "*.r filter=protocol" >.gitattributes &&
+ cp "$TEST_ROOT/test.o" test.r &&
+ S=$(file_size test.r) &&
+
+ filter_git add . &&
+ cat >expected.log <<-EOF &&
+ START
+ init handshake complete
+ IN: clean test.r $S [OK] -- OUT: $S . [OK]
+ STOP
+ EOF
+ test_cmp_count expected.log rot13-filter.log &&
+
+ rm test.r &&
+
+ filter_git checkout --quiet --no-progress . &&
+ # If the filter would be used for "smudge", too, we would see
+ # "IN: smudge test.r 57 [OK] -- OUT: 57 . [OK]" here
+ cat >expected.log <<-EOF &&
+ START
+ init handshake complete
+ STOP
+ EOF
+ test_cmp_exclude_clean expected.log rot13-filter.log
+ )
+'
test_expect_success PERL 'required process filter should process multiple packets' '
test_config_global filter.protocol.process "$TEST_DIRECTORY/t0021/rot13-filter.pl clean smudge" &&
@@ -461,42 +509,57 @@ test_expect_success PERL 'required process filter should process multiple packet
cd repo &&
git init &&
- # Generate data that requires 3 packets
- PKTLINE_DATA_MAXLEN=65516 &&
+ # Generate data requiring 1, 2, 3 packets
+ S=65516 && # PKTLINE_DATA_MAXLEN -> Maximal size of a packet
+ generate_random_characters $(($S )) 1pkt_1__.file &&
+ generate_random_characters $(($S +1)) 2pkt_1+1.file &&
+ generate_random_characters $(($S*2-1)) 2pkt_2-1.file &&
+ generate_random_characters $(($S*2 )) 2pkt_2__.file &&
+ generate_random_characters $(($S*2+1)) 3pkt_2+1.file &&
- generate_test_data $(($PKTLINE_DATA_MAXLEN )) 1pkt_1__ &&
- generate_test_data $(($PKTLINE_DATA_MAXLEN + 1)) 2pkt_1+1 &&
- generate_test_data $(($PKTLINE_DATA_MAXLEN * 2 - 1)) 2pkt_2-1 &&
- generate_test_data $(($PKTLINE_DATA_MAXLEN * 2 )) 2pkt_2__ &&
- generate_test_data $(($PKTLINE_DATA_MAXLEN * 2 + 1)) 3pkt_2+1 &&
+ for FILE in "$TEST_ROOT"/*.file
+ do
+ cp "$FILE" . &&
+ "$TEST_ROOT/rot13.sh" <"$FILE" >"$FILE.rot13"
+ done &&
echo "*.file filter=protocol" >.gitattributes &&
- check_filter \
- git add *.file .gitattributes \
- <<-\EOF &&
- 1 IN: clean 1pkt_1__.file 65516 [OK] -- OUT: 65516 . [OK]
- 1 IN: clean 2pkt_1+1.file 65517 [OK] -- OUT: 65517 .. [OK]
- 1 IN: clean 2pkt_2-1.file 131031 [OK] -- OUT: 131031 .. [OK]
- 1 IN: clean 2pkt_2__.file 131032 [OK] -- OUT: 131032 .. [OK]
- 1 IN: clean 3pkt_2+1.file 131033 [OK] -- OUT: 131033 ... [OK]
- 1 START
- 1 STOP
- 1 wrote filter header
+ filter_git add *.file .gitattributes &&
+ cat >expected.log <<-EOF &&
+ START
+ init handshake complete
+ IN: clean 1pkt_1__.file $(($S )) [OK] -- OUT: $(($S )) . [OK]
+ IN: clean 2pkt_1+1.file $(($S +1)) [OK] -- OUT: $(($S +1)) .. [OK]
+ IN: clean 2pkt_2-1.file $(($S*2-1)) [OK] -- OUT: $(($S*2-1)) .. [OK]
+ IN: clean 2pkt_2__.file $(($S*2 )) [OK] -- OUT: $(($S*2 )) .. [OK]
+ IN: clean 3pkt_2+1.file $(($S*2+1)) [OK] -- OUT: $(($S*2+1)) ... [OK]
+ STOP
EOF
- git commit . -m "test commit" &&
+ test_cmp_count expected.log rot13-filter.log &&
rm -f *.file &&
- git checkout -- *.file &&
- for f in *.file
+ filter_git checkout --quiet --no-progress -- *.file &&
+ cat >expected.log <<-EOF &&
+ START
+ init handshake complete
+ IN: smudge 1pkt_1__.file $(($S )) [OK] -- OUT: $(($S )) . [OK]
+ IN: smudge 2pkt_1+1.file $(($S +1)) [OK] -- OUT: $(($S +1)) .. [OK]
+ IN: smudge 2pkt_2-1.file $(($S*2-1)) [OK] -- OUT: $(($S*2-1)) .. [OK]
+ IN: smudge 2pkt_2__.file $(($S*2 )) [OK] -- OUT: $(($S*2 )) .. [OK]
+ IN: smudge 3pkt_2+1.file $(($S*2+1)) [OK] -- OUT: $(($S*2+1)) ... [OK]
+ STOP
+ EOF
+ test_cmp_exclude_clean expected.log rot13-filter.log &&
+
+ for FILE in *.file
do
- git cat-file blob :$f >actual &&
- test_cmp ../$f.rot13 actual
+ test_cmp_committed_rot13 "$TEST_ROOT/$FILE" $FILE
done
)
'
-test_expect_success PERL 'required process filter should with clean error should fail' '
+test_expect_success PERL 'required process filter with clean error should fail' '
test_config_global filter.protocol.process "$TEST_DIRECTORY/t0021/rot13-filter.pl clean smudge" &&
test_config_global filter.protocol.required true &&
rm -rf repo &&
@@ -507,11 +570,10 @@ test_expect_success PERL 'required process filter should with clean error should
echo "*.r filter=protocol" >.gitattributes &&
- cp ../test.o test.r &&
+ cp "$TEST_ROOT/test.o" test.r &&
echo "this is going to fail" >clean-write-fail.r &&
echo "content-test3-subdir" >test3.r &&
- # Note: There are three clean paths in convert.c we just test one here.
test_must_fail git add .
)
'
@@ -526,38 +588,48 @@ test_expect_success PERL 'process filter should restart after unexpected write f
echo "*.r filter=protocol" >.gitattributes &&
- cp ../test.o test.r &&
- cp ../test2.o test2.r &&
+ cp "$TEST_ROOT/test.o" test.r &&
+ cp "$TEST_ROOT/test2.o" test2.r &&
echo "this is going to fail" >smudge-write-fail.o &&
- cat smudge-write-fail.o >smudge-write-fail.r &&
+ cp smudge-write-fail.o smudge-write-fail.r &&
+
+ S=$(file_size test.r) &&
+ S2=$(file_size test2.r) &&
+ SF=$(file_size smudge-write-fail.r) &&
+
git add . &&
- git commit . -m "test commit" &&
rm -f *.r &&
- check_filter_ignore_clean \
- git checkout . \
- <<-\EOF &&
+ rm -f rot13-filter.log &&
+ git checkout --quiet --no-progress . 2>git-stderr.log &&
+
+ grep "smudge write error at" git-stderr.log &&
+ grep "error: external filter" git-stderr.log &&
+
+ cat >expected.log <<-EOF &&
START
- wrote filter header
- IN: smudge smudge-write-fail.r 22 [OK] -- OUT: 22 [WRITE FAIL]
+ init handshake complete
+ IN: smudge smudge-write-fail.r $SF [OK] -- OUT: $SF [WRITE FAIL]
START
- wrote filter header
- IN: smudge test.r 57 [OK] -- OUT: 57 . [OK]
- IN: smudge test2.r 14 [OK] -- OUT: 14 . [OK]
+ init handshake complete
+ IN: smudge test.r $S [OK] -- OUT: $S . [OK]
+ IN: smudge test2.r $S2 [OK] -- OUT: $S2 . [OK]
STOP
EOF
+ test_cmp_exclude_clean expected.log rot13-filter.log &&
- check_rot13 ../test.o test.r &&
- check_rot13 ../test2.o test2.r &&
+ test_cmp_committed_rot13 "$TEST_ROOT/test.o" test.r &&
+ test_cmp_committed_rot13 "$TEST_ROOT/test2.o" test2.r &&
- ! test_cmp smudge-write-fail.o smudge-write-fail.r && # Smudge failed!
- ./../rot13.sh <smudge-write-fail.o >expected &&
+ # Smudge failed
+ ! test_cmp smudge-write-fail.o smudge-write-fail.r &&
+ "$TEST_ROOT/rot13.sh" <smudge-write-fail.o >expected &&
git cat-file blob :smudge-write-fail.r >actual &&
- test_cmp expected actual # Clean worked!
+ test_cmp expected actual
)
'
-test_expect_success PERL 'process filter should not restart in case of an error' '
+test_expect_success PERL 'process filter should not be restarted if it signals an error' '
test_config_global filter.protocol.process "$TEST_DIRECTORY/t0021/rot13-filter.pl clean smudge" &&
rm -rf repo &&
mkdir repo &&
@@ -567,32 +639,36 @@ test_expect_success PERL 'process filter should not restart in case of an error'
echo "*.r filter=protocol" >.gitattributes &&
- cp ../test.o test.r &&
- cp ../test2.o test2.r &&
+ cp "$TEST_ROOT/test.o" test.r &&
+ cp "$TEST_ROOT/test2.o" test2.r &&
echo "this will cause an error" >error.o &&
cp error.o error.r &&
+
+ S=$(file_size test.r) &&
+ S2=$(file_size test2.r) &&
+ SE=$(file_size error.r) &&
+
git add . &&
- git commit . -m "test commit" &&
rm -f *.r &&
- check_filter_ignore_clean \
- git checkout . \
- <<-\EOF &&
+ filter_git checkout --quiet --no-progress . &&
+ cat >expected.log <<-EOF &&
START
- wrote filter header
- IN: smudge error.r 25 [OK] -- OUT: 0 [ERROR]
- IN: smudge test.r 57 [OK] -- OUT: 57 . [OK]
- IN: smudge test2.r 14 [OK] -- OUT: 14 . [OK]
+ init handshake complete
+ IN: smudge error.r $SE [OK] -- OUT: 0 [ERROR]
+ IN: smudge test.r $S [OK] -- OUT: $S . [OK]
+ IN: smudge test2.r $S2 [OK] -- OUT: $S2 . [OK]
STOP
EOF
+ test_cmp_exclude_clean expected.log rot13-filter.log &&
- check_rot13 ../test.o test.r &&
- check_rot13 ../test2.o test2.r &&
+ test_cmp_committed_rot13 "$TEST_ROOT/test.o" test.r &&
+ test_cmp_committed_rot13 "$TEST_ROOT/test2.o" test2.r &&
test_cmp error.o error.r
)
'
-test_expect_success PERL 'process filter should be able to signal an error for all future files' '
+test_expect_success PERL 'process filter signals abort once to abort processing of all future files' '
test_config_global filter.protocol.process "$TEST_DIRECTORY/t0021/rot13-filter.pl clean smudge" &&
rm -rf repo &&
mkdir repo &&
@@ -602,25 +678,27 @@ test_expect_success PERL 'process filter should be able to signal an error for a
echo "*.r filter=protocol" >.gitattributes &&
- cp ../test.o test.r &&
- cp ../test2.o test2.r &&
+ cp "$TEST_ROOT/test.o" test.r &&
+ cp "$TEST_ROOT/test2.o" test2.r &&
echo "error this blob and all future blobs" >abort.o &&
cp abort.o abort.r &&
+
+ SA=$(file_size abort.r) &&
+
git add . &&
- git commit . -m "test commit" &&
rm -f *.r &&
- check_filter_ignore_clean \
- git checkout . \
- <<-\EOF &&
+ filter_git checkout --quiet --no-progress . &&
+ cat >expected.log <<-EOF &&
START
- wrote filter header
- IN: smudge abort.r 37 [OK] -- OUT: 0 [ABORT]
+ init handshake complete
+ IN: smudge abort.r $SA [OK] -- OUT: 0 [ABORT]
STOP
EOF
+ test_cmp_exclude_clean expected.log rot13-filter.log &&
- test_cmp ../test.o test.r &&
- test_cmp ../test2.o test2.r &&
+ test_cmp "$TEST_ROOT/test.o" test.r &&
+ test_cmp "$TEST_ROOT/test2.o" test2.r &&
test_cmp abort.o abort.r
)
'
@@ -636,9 +714,9 @@ test_expect_success PERL 'invalid process filter must fail (and not hang!)' '
echo "*.r filter=protocol" >.gitattributes &&
- cp ../test.o test.r &&
- test_must_fail git add . 2> git_stderr.log &&
- grep "not support long running filter protocol" git_stderr.log
+ cp "$TEST_ROOT/test.o" test.r &&
+ test_must_fail git add . 2>git-stderr.log &&
+ grep "does not support filter protocol version" git-stderr.log
)
'
diff --git a/t/t0021/rot13-filter.pl b/t/t0021/rot13-filter.pl
index 8958f71..1a6959c 100755
--- a/t/t0021/rot13-filter.pl
+++ b/t/t0021/rot13-filter.pl
@@ -26,10 +26,10 @@ use warnings;
my $MAX_PACKET_CONTENT_SIZE = 65516;
my @capabilities = @ARGV;
-open my $debug, ">>", "rot13-filter.log";
+open my $debug, ">>", "rot13-filter.log" or die "cannot open log file: $!";
sub rot13 {
- my ($str) = @_;
+ my $str = shift;
$str =~ y/A-Za-z/N-ZA-Mn-za-m/;
return $str;
}
@@ -38,13 +38,12 @@ sub packet_bin_read {
my $buffer;
my $bytes_read = read STDIN, $buffer, 4;
if ( $bytes_read == 0 ) {
-
# EOF - Git stopped talking to us!
print $debug "STOP\n";
exit();
}
elsif ( $bytes_read != 4 ) {
- die "invalid packet size '$bytes_read' field";
+ die "invalid packet: '$buffer'";
}
my $pkt_size = hex($buffer);
if ( $pkt_size == 0 ) {
@@ -54,27 +53,27 @@ sub packet_bin_read {
my $content_size = $pkt_size - 4;
$bytes_read = read STDIN, $buffer, $content_size;
if ( $bytes_read != $content_size ) {
- die "invalid packet ($content_size expected; $bytes_read read)";
+ die "invalid packet ($content_size bytes expected; $bytes_read bytes read)";
}
return ( 0, $buffer );
}
else {
- die "invalid packet size";
+ die "invalid packet size: $pkt_size";
}
}
sub packet_txt_read {
my ( $res, $buf ) = packet_bin_read();
- unless ( $buf =~ /\n$/ ) {
- die "A non-binary line SHOULD BE terminated by an LF.";
+ unless ( $buf =~ s/\n$// ) {
+ die "A non-binary line MUST be terminated by an LF.";
}
- return ( $res, substr( $buf, 0, -1 ) );
+ return ( $res, $buf );
}
sub packet_bin_write {
- my ($packet) = @_;
- print STDOUT sprintf( "%04x", length($packet) + 4 );
- print STDOUT $packet;
+ my $buf = shift;
+ print STDOUT sprintf( "%04x", length($buf) + 4 );
+ print STDOUT $buf;
STDOUT->flush();
}
@@ -105,7 +104,7 @@ foreach (@capabilities) {
packet_txt_write( $_ . "=true" );
}
packet_flush();
-print $debug "wrote filter header\n";
+print $debug "init handshake complete\n";
$debug->flush();
while (1) {
@@ -175,6 +174,7 @@ while (1) {
while ( length($output) > 0 ) {
my $packet = substr( $output, 0, $MAX_PACKET_CONTENT_SIZE );
packet_bin_write($packet);
+ # dots represent the number of packets
print $debug ".";
if ( length($output) > $MAX_PACKET_CONTENT_SIZE ) {
$output = substr( $output, $MAX_PACKET_CONTENT_SIZE );
Lars Schneider (14):
convert: quote filter names in error messages
convert: modernize tests
run-command: move check_pipe() from write_or_die to run_command
run-command: add wait_on_exit
pkt-line: rename packet_write() to packet_write_fmt()
pkt-line: extract set_packet_header()
pkt-line: add packet_write_fmt_gently()
pkt-line: add packet_flush_gently()
pkt-line: add packet_write_gently()
pkt-line: add functions to read/write flush terminated packet streams
convert: make apply_filter() adhere to standard Git error handling
convert: prepare filter.<driver>.process option
convert: add filter.<driver>.process option
contrib/long-running-filter: add long running filter example
Documentation/gitattributes.txt | 159 ++++++++++-
builtin/archive.c | 4 +-
builtin/receive-pack.c | 4 +-
builtin/remote-ext.c | 4 +-
builtin/upload-archive.c | 4 +-
connect.c | 2 +-
contrib/long-running-filter/example.pl | 127 +++++++++
convert.c | 370 +++++++++++++++++++++---
daemon.c | 2 +-
http-backend.c | 2 +-
pkt-line.c | 148 +++++++++-
pkt-line.h | 12 +-
run-command.c | 72 ++++-
run-command.h | 5 +-
shallow.c | 2 +-
t/t0021-conversion.sh | 505 ++++++++++++++++++++++++++++++---
t/t0021/rot13-filter.pl | 191 +++++++++++++
upload-pack.c | 30 +-
write_or_die.c | 13 -
19 files changed, 1519 insertions(+), 137 deletions(-)
create mode 100755 contrib/long-running-filter/example.pl
create mode 100755 t/t0021/rot13-filter.pl
--
2.10.0
^ permalink raw reply related
* [PATCH v9 08/14] pkt-line: add packet_flush_gently()
From: larsxschneider @ 2016-10-04 12:59 UTC (permalink / raw)
To: git; +Cc: ramsay, jnareb, gitster, j6t, tboegi, peff, mlbright,
Lars Schneider
In-Reply-To: <20161004125947.67104-1-larsxschneider@gmail.com>
From: Lars Schneider <larsxschneider@gmail.com>
packet_flush() would die in case of a write error even though for some
callers an error would be acceptable. Add packet_flush_gently() which
writes a pkt-line flush packet like packet_flush() but does not die in
case of an error. The function is used in a subsequent patch.
Signed-off-by: Lars Schneider <larsxschneider@gmail.com>
Signed-off-by: Junio C Hamano <gitster@pobox.com>
---
pkt-line.c | 8 ++++++++
pkt-line.h | 1 +
2 files changed, 9 insertions(+)
diff --git a/pkt-line.c b/pkt-line.c
index 56915f0..286eb09 100644
--- a/pkt-line.c
+++ b/pkt-line.c
@@ -91,6 +91,14 @@ void packet_flush(int fd)
write_or_die(fd, "0000", 4);
}
+int packet_flush_gently(int fd)
+{
+ packet_trace("0000", 4, 1);
+ if (write_in_full(fd, "0000", 4) == 4)
+ return 0;
+ return error("flush packet write failed");
+}
+
void packet_buf_flush(struct strbuf *buf)
{
packet_trace("0000", 4, 1);
diff --git a/pkt-line.h b/pkt-line.h
index 3caea77..3fa0899 100644
--- a/pkt-line.h
+++ b/pkt-line.h
@@ -23,6 +23,7 @@ void packet_flush(int fd);
void packet_write_fmt(int fd, const char *fmt, ...) __attribute__((format (printf, 2, 3)));
void packet_buf_flush(struct strbuf *buf);
void packet_buf_write(struct strbuf *buf, const char *fmt, ...) __attribute__((format (printf, 2, 3)));
+int packet_flush_gently(int fd);
int packet_write_fmt_gently(int fd, const char *fmt, ...) __attribute__((format (printf, 2, 3)));
/*
--
2.10.0
^ permalink raw reply related
* [PATCH v9 14/14] contrib/long-running-filter: add long running filter example
From: larsxschneider @ 2016-10-04 12:59 UTC (permalink / raw)
To: git; +Cc: ramsay, jnareb, gitster, j6t, tboegi, peff, mlbright,
Lars Schneider
In-Reply-To: <20161004125947.67104-1-larsxschneider@gmail.com>
From: Lars Schneider <larsxschneider@gmail.com>
Add a simple pass-thru filter as example implementation for the Git
filter protocol version 2. See Documentation/gitattributes.txt, section
"Filter Protocol" for more info.
Signed-off-by: Lars Schneider <larsxschneider@gmail.com>
---
Documentation/gitattributes.txt | 4 +-
contrib/long-running-filter/example.pl | 127 +++++++++++++++++++++++++++++++++
2 files changed, 130 insertions(+), 1 deletion(-)
create mode 100755 contrib/long-running-filter/example.pl
diff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt
index 5868f00..a182ef2 100644
--- a/Documentation/gitattributes.txt
+++ b/Documentation/gitattributes.txt
@@ -514,7 +514,9 @@ the next "key=value" list containing a command. Git will close
the command pipe on exit. The filter is expected to detect EOF
and exit gracefully on its own.
-If you develop your own long running filter
+A long running filter demo implementation can be found in
+`contrib/long-running-filter/example.pl` located in the Git
+core repository. If you develop your own long running filter
process then the `GIT_TRACE_PACKET` environment variables can be
very helpful for debugging (see linkgit:git[1]).
diff --git a/contrib/long-running-filter/example.pl b/contrib/long-running-filter/example.pl
new file mode 100755
index 0000000..f4102d2
--- /dev/null
+++ b/contrib/long-running-filter/example.pl
@@ -0,0 +1,127 @@
+#!/usr/bin/perl
+#
+# Example implementation for the Git filter protocol version 2
+# See Documentation/gitattributes.txt, section "Filter Protocol"
+#
+# Please note, this pass-thru filter is a minimal skeleton. No proper
+# error handling was implemented.
+#
+
+use strict;
+use warnings;
+
+my $MAX_PACKET_CONTENT_SIZE = 65516;
+
+sub packet_bin_read {
+ my $buffer;
+ my $bytes_read = read STDIN, $buffer, 4;
+ if ( $bytes_read == 0 ) {
+
+ # EOF - Git stopped talking to us!
+ exit();
+ }
+ elsif ( $bytes_read != 4 ) {
+ die "invalid packet: '$buffer'";
+ }
+ my $pkt_size = hex($buffer);
+ if ( $pkt_size == 0 ) {
+ return ( 1, "" );
+ }
+ elsif ( $pkt_size > 4 ) {
+ my $content_size = $pkt_size - 4;
+ $bytes_read = read STDIN, $buffer, $content_size;
+ if ( $bytes_read != $content_size ) {
+ die "invalid packet ($content_size bytes expected; $bytes_read bytes read)";
+ }
+ return ( 0, $buffer );
+ }
+ else {
+ die "invalid packet size: $pkt_size";
+ }
+}
+
+sub packet_txt_read {
+ my ( $res, $buf ) = packet_bin_read();
+ unless ( $buf =~ s/\n$// ) {
+ die "A non-binary line MUST be terminated by an LF.";
+ }
+ return ( $res, $buf );
+}
+
+sub packet_bin_write {
+ my $buf = shift;
+ print STDOUT sprintf( "%04x", length($buf) + 4 );
+ print STDOUT $buf;
+ STDOUT->flush();
+}
+
+sub packet_txt_write {
+ packet_bin_write( $_[0] . "\n" );
+}
+
+sub packet_flush {
+ print STDOUT sprintf( "%04x", 0 );
+ STDOUT->flush();
+}
+
+( packet_txt_read() eq ( 0, "git-filter-client" ) ) || die "bad initialize";
+( packet_txt_read() eq ( 0, "version=2" ) ) || die "bad version";
+( packet_bin_read() eq ( 1, "" ) ) || die "bad version end";
+
+packet_txt_write("git-filter-server");
+packet_txt_write("version=2");
+
+( packet_txt_read() eq ( 0, "clean=true" ) ) || die "bad capability";
+( packet_txt_read() eq ( 0, "smudge=true" ) ) || die "bad capability";
+( packet_bin_read() eq ( 1, "" ) ) || die "bad capability end";
+
+packet_txt_write("clean=true");
+packet_txt_write("smudge=true");
+packet_flush();
+
+while (1) {
+ my ($command) = packet_txt_read() =~ /^command=([^=]+)$/;
+ my ($pathname) = packet_txt_read() =~ /^pathname=([^=]+)$/;
+
+ packet_bin_read();
+
+ my $input = "";
+ {
+ binmode(STDIN);
+ my $buffer;
+ my $done = 0;
+ while ( !$done ) {
+ ( $done, $buffer ) = packet_bin_read();
+ $input .= $buffer;
+ }
+ }
+
+ my $output;
+ if ( $command eq "clean" ) {
+ ### Perform clean here ###
+ $output = $input;
+ }
+ elsif ( $command eq "smudge" ) {
+ ### Perform smudge here ###
+ $output = $input;
+ }
+ else {
+ die "bad command '$command'";
+ }
+
+ packet_txt_write("status=success");
+ packet_flush();
+ while ( length($output) > 0 ) {
+ my $packet = substr( $output, 0, $MAX_PACKET_CONTENT_SIZE );
+ packet_bin_write($packet);
+ if ( length($output) > $MAX_PACKET_CONTENT_SIZE ) {
+ $output = substr( $output, $MAX_PACKET_CONTENT_SIZE );
+ }
+ else {
+ $output = "";
+ }
+ }
+ packet_flush(); # flush content!
+ packet_flush(); # empty list, keep "status=success" unchanged!
+
+}
--
2.10.0
^ permalink raw reply related
page: next (older) | prev (newer) | latest
- recent:[subjects (threaded)|topics (new)|topics (active)]
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox