* Specifying revisions in the future
From: jpaugh @ 2012-02-04 15:58 UTC (permalink / raw)
To: git
Hello.
Is it possible to specify revisions in the future? The gitrevisions man
page implies otherwise. Alternatively, is there a way to find out the
number of commits between two revs---assuming one is an ancestor of the
other?
I want to do a certain arbitrary operation for each revision between
where I am now and the tip of the branch.
v1.0-a master
\ \
o---o---o---o---o---o---o
|
I am here
I've been using the following to do what I want:
ref=master; \
for i in {5..1}; do \
echo; \
git log --stat $ref~$i^\!; \
read -p 'Full diff? '; \
echo; \
if [[ $REPLY == 'y' ]]; then \
git diff $ref~$i^\!; \
fi; \
done;
which lists the log and diffstat for last 5 commits between master and
where I am (e.g. an older tag/branch) with an optional full diff. I know
implementing revision specifiers to the future is nontrivial. (I
realized that when I considered non-linear histories.) In this case,
I've distilled it to the point that all I need is the number of commits
between two revs. Can this be had without manually inspecting git log?
Or, is there a better way to get detailed diffs like this?
Thanks.
Jonathan Paugh
^ permalink raw reply
* [PATCH] Change include order in two compat/ files to avoid compiler warning
From: Ben Walton @ 2012-02-05 1:08 UTC (permalink / raw)
To: git, gitster; +Cc: Ben Walton
The inet_ntop and inet_pton compatibility wrapper source files
included system headers before git-compat-utils.h. This was causing a
warning on Solaris as _FILE_OFFSET_BITS was being redefined in
git-compat-utils.h. Including git-compat-utils.h first avoids the
warnings.
Signed-off-by: Ben Walton <bwalton@artsci.utoronto.ca>
---
I verified that this re-ordering doesn't affect either the build or the
test suite completion on both i386 and sparc. I think the ordering is
simply the result of placing the git-compat-utils.h include where some
others were removed in da523cc597b1.
compat/inet_ntop.c | 4 +---
compat/inet_pton.c | 4 +---
2 files changed, 2 insertions(+), 6 deletions(-)
diff --git a/compat/inet_ntop.c b/compat/inet_ntop.c
index 60b5a1d..f1bf81c 100644
--- a/compat/inet_ntop.c
+++ b/compat/inet_ntop.c
@@ -15,11 +15,9 @@
* SOFTWARE.
*/
+#include "../git-compat-util.h"
#include <errno.h>
#include <sys/types.h>
-
-#include "../git-compat-util.h"
-
#include <stdio.h>
#include <string.h>
diff --git a/compat/inet_pton.c b/compat/inet_pton.c
index 2ec995e..1d44a5d 100644
--- a/compat/inet_pton.c
+++ b/compat/inet_pton.c
@@ -15,11 +15,9 @@
* WITH THE USE OR PERFORMANCE OF THIS SOFTWARE.
*/
+#include "../git-compat-util.h"
#include <errno.h>
#include <sys/types.h>
-
-#include "../git-compat-util.h"
-
#include <stdio.h>
#include <string.h>
--
1.7.8.3
^ permalink raw reply related
* Re: Specifying revisions in the future
From: Jakub Narebski @ 2012-02-05 2:44 UTC (permalink / raw)
To: Jonathan Paugh; +Cc: git
In-Reply-To: <jgjkk0$qrg$1@dough.gmane.org>
jpaugh@gmx.us writes:
> Hello.
> I want to do a certain arbitrary operation for each revision between
> where I am now and the tip of the branch.
>
> v1.0-a master
> \ \
> o---o---o---o---o---o---o
> |
> I am here
That is the problem X.
> Is it possible to specify revisions in the future? The gitrevisions man
> page implies otherwise. Alternatively, is there a way to find out the
> number of commits between two revs---assuming one is an ancestor of the
> other?
That is your idea of a solution, Y.
You have XY problem. You need to do X, and you think you can use Y to
do X, so you ask about how to do Y.
If you want to list all revsions between v1.0-a and master, use
git rev-list v1.0a..master
or
git rev-list --ancestry-path v1.0a..master
depending on definition of _between_ (see "History simplification" in
git-log(1) manpage for description of `--ancestry-path` option).
>
> I've been using the following to do what I want:
>
> ref=master; \
> for i in {5..1}; do \
> echo; \
> git log --stat $ref~$i^\!; \
> read -p 'Full diff? '; \
> echo; \
> if [[ $REPLY == 'y' ]]; then \
> git diff $ref~$i^\!; \
> fi; \
> done;
>
> which lists the log and diffstat for last 5 commits between master and
> where I am (e.g. an older tag/branch) with an optional full diff. I know
> implementing revision specifiers to the future is nontrivial. (I
> realized that when I considered non-linear histories.) In this case,
> I've distilled it to the point that all I need is the number of commits
> between two revs. Can this be had without manually inspecting git log?
> Or, is there a better way to get detailed diffs like this?
--
Jakub Narebski
^ permalink raw reply
* Re: Specifying revisions in the future
From: Jakub Narebski @ 2012-02-05 3:07 UTC (permalink / raw)
To: Jonathan Paugh; +Cc: git
In-Reply-To: <4F2DEF89.4030302@gmx.us>
Jonathan Paugh wrote:
> > You have XY problem. You need to do X, and you think you can use Y to
> > do X, so you ask about how to do Y.
> >
> > If you want to list all revsions between v1.0-a and master, use
> >
> > git rev-list v1.0a..master
> >
> > or
> >
> > git rev-list --ancestry-path v1.0a..master
> Thanks. This Y' will take me to lot's of exciting destinations---and I
> must confess I haven't messed with the plumbing heretofore.
Of course you can also do
git log --ancestry-path v1.0a..master
--
Jakub Narebski
Poland
^ permalink raw reply
* Re: Git performance results on a large repository
From: Nguyen Thai Ngoc Duy @ 2012-02-05 3:47 UTC (permalink / raw)
To: Joshua Redstone; +Cc: git@vger.kernel.org
In-Reply-To: <243C23AF01622E49BEA3F28617DBF0AD5912CA85@SC-MBX02-5.TheFacebook.com>
On Sun, Feb 5, 2012 at 1:05 AM, Joshua Redstone <joshua.redstone@fb.com> wrote:
> It's also conceivable that, if there were an external interface in git to attach other
> systems to efficiently report which files have changed (e.g., via file-system integration),
> it's possible that we could omit managing the index in many cases.
> I know that would be a big change, but the benefits are intriguing.
The "interface to report which files have changed" is exactly "git
update-index --[no-]assume-unchanged" is for. Have a look at the man
page. Basically you can mark every file "unchanged" in the beginning
and git won't bother lstat() them. What files you change, you have to
explicitly run "git update-index --no-assume-unchanged" to tell git.
Someone on HN suggested making assume-unchanged files read-only to
avoid 90% accidentally changing a file without telling git. When
assume-unchanged bit is cleared, the file is made read-write again.
--
Duy
^ permalink raw reply
* Re: Git performance results on a large repository
From: david @ 2012-02-05 4:30 UTC (permalink / raw)
To: Joshua Redstone; +Cc: git@vger.kernel.org
In-Reply-To: <CB5074CF.3AD7A%joshua.redstone@fb.com>
On Fri, 3 Feb 2012, Joshua Redstone wrote:
> The test repo has 4 million commits, linear history and about 1.3 million
> files. The size of the .git directory is about 15GB, and has been
> repacked with 'git repack -a -d -f --max-pack-size=10g --depth=100
> --window=250'. This repack took about 2 days on a beefy machine (I.e.,
> lots of ram and flash). The size of the index file is 191 MB.
This may be a silly thought, but what if instead of one pack file of your
entire history (4 million commits) you create multiple packs (say every
half million commits) and mark all but the most recent pack as .keep (so
that they won't be modified by a repack)
that way things that only need to worry about recent history (blame, etc)
will probably never have to go past the most recent pack file or two
I may be wrong, but I think that when git is looking for 'similar files'
for delta compression, it limits it's search to the current pack, so this
will also keep you from searching the entire project history.
David Lang
^ permalink raw reply
* Re: [PATCH 3/3] t: mailmap: add simple name translation test
From: Jonathan Nieder @ 2012-02-05 6:17 UTC (permalink / raw)
To: Felipe Contreras; +Cc: git, Junio C Hamano, Marius Storm-Olsen, Jim Meyering
In-Reply-To: <CAMP44s0Z=k6VBfv0HOGHyMBLRcPauK7K5RNvuRDbfq5=5aKVpg@mail.gmail.com>
Felipe Contreras wrote:
> You mean the commit message, you haven't made any comment about the code.
No, for this patch, more important than the absence of any explanation
in the commit message (which is also important) is the code change
that seems unnecessarily invasive.
You've already demonstrated that I do not have the right communication
style to explain such things to you and work towards a fix that
addresses both our concerns. So I give up. I'll just give my
feedback on patches that concern code I care about and an explanation
for the sake of others on the list that are better able to interact
with you. I am willing to work with or answer questions from anyone
including you, though.
Sorry,
Jonathan
^ permalink raw reply
* Build oddities
From: Michael @ 2012-02-05 6:34 UTC (permalink / raw)
To: git
Hi there,
I ran into a build oddity with git today - the environment variable $X is
appended to most binaries. At the time, I had X=last, so I ended up with
gitlast, git-peek-remotelast etc.
Does anyone know why this behaviour exists, and if it is still desired?
Best regards,
Michael
^ permalink raw reply
* Re: Build oddities
From: Nguyen Thai Ngoc Duy @ 2012-02-05 6:46 UTC (permalink / raw)
To: Michael; +Cc: git
In-Reply-To: <loom.20120205T072940-523@post.gmane.org>
On Sun, Feb 5, 2012 at 1:34 PM, Michael <kensington@astralcloak.net> wrote:
> Hi there,
>
> I ran into a build oddity with git today - the environment variable $X is
> appended to most binaries. At the time, I had X=last, so I ended up with
> gitlast, git-peek-remotelast etc.
$X is to append .exe for Windows build so you would get git.exe,
git-peek-remote.exe... We should set X to empty from the beginning.
Patches are welcome.
--
Duy
^ permalink raw reply
* [PATCH 0/3] On compresing large index
From: Nguyễn Thái Ngọc Duy @ 2012-02-05 8:30 UTC (permalink / raw)
To: git; +Cc: Joshua Redstone, Nguyễn Thái Ngọc Duy
I was thinking whether compressing index might help when it contained
~2M files. It turns out that only makes the situation worse. Anyway, I
post the code and some numbers here.
The index is created artifically with the program [1]
$ git init
$ touch foo
$ git hash-object -w foo
$ ./a.out 256 256 32 | git update-index --index-info
That gives ~2M files in index, 209 MB in size.
$ time ~/w/git/git ls-files | head >/dev/null
real 0m4.635s
user 0m4.258s
sys 0m0.329s
$ time ~/w/git/git update-index level-0-0000/foo
real 0m4.593s
user 0m4.264s
sys 0m0.323s
Index is compressed with GIT_ZCACHE=1.
$ GIT_ZCACHE=1 ~/w/git/git update-index level-0-0000/foo
which gives 6.8 MB index (the true number may be less impressive
because compressing rate in my artificial tree is really high). The
only problem with this is git uses more time, not less
$ time ~/w/git/git ls-files | head >/dev/null
real 0m4.970s
user 0m4.675s
sys 0m0.289s
$ time GIT_ZCACHE=1 ~/w/git/git update-index level-0-0000/foo
real 0m4.959s
user 0m4.682s
sys 0m0.273s
My guess is Linux caches the whole index in memory already so I/O time
does not really matter, while we still have to pay for zlib's time. We
need to figure out what git uses 4s user time for.
This series may be useful on OSes that do not cache heavily. Though
I'm not sure if there is any out there nowadays.
Nguyễn Thái Ngọc Duy (3):
read-cache: factor out cache entries reading code
read-cache: reduce malloc/free during writing index
Support compressing index when GIT_ZCACHE=1
cache.h | 1 +
read-cache.c | 172 +++++++++++++++++++++++++++++++++++++++++++++++++---------
2 files changed, 148 insertions(+), 25 deletions(-)
[1]
-- 8< --
#include <stdio.h>
#include <string.h>
int main(int argc, char **argv)
{
const char *prefix = "100644 e69de29bb2d1d6434b8b29ae775ad8c2e48c5391 0\t";
int l1, l2, l3;
int m1, m2, m3;
m1 = atoi(argv[1]);
m2 = atoi(argv[2]);
m3 = atoi(argv[3]);
for (l1 = 0; l1 < m1; l1++) {
printf("%slevel-0-%04d/foo\n", prefix, l1);
for (l2 = 0; l2 < m2; l2++)
for (l3 = 0; l3 < m3; l3++)
printf("%slevel-0-%04d/level-1-%04d/foo-%04d\n",
prefix, l1, l2, l3);
}
return 0;
}
-- 8< --
--
1.7.8.36.g69ee2
^ permalink raw reply
* [PATCH 2/3] read-cache: reduce malloc/free during writing index
From: Nguyễn Thái Ngọc Duy @ 2012-02-05 8:30 UTC (permalink / raw)
To: git; +Cc: Joshua Redstone, Nguyễn Thái Ngọc Duy
In-Reply-To: <1328430605-4566-1-git-send-email-pclouds@gmail.com>
Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>
---
read-cache.c | 26 ++++++++++++++++++--------
1 files changed, 18 insertions(+), 8 deletions(-)
diff --git a/read-cache.c b/read-cache.c
index 2dbf923..7b9a989 100644
--- a/read-cache.c
+++ b/read-cache.c
@@ -1521,12 +1521,19 @@ static void ce_smudge_racily_clean_entry(struct cache_entry *ce)
}
}
-static int ce_write_entry(git_SHA_CTX *c, int fd, struct cache_entry *ce)
+static int ce_prepare_ondisk_entry(struct cache_entry *ce,
+ void **ondisk_p, int *ondisk_size)
{
int size = ondisk_ce_size(ce);
- struct ondisk_cache_entry *ondisk = xcalloc(1, size);
+ struct ondisk_cache_entry *ondisk;
char *name;
- int result;
+
+ if (size <= *ondisk_size)
+ ondisk = *ondisk_p;
+ else {
+ ondisk = *ondisk_p = xrealloc(*ondisk_p, size);
+ *ondisk_size = size;
+ }
ondisk->ctime.sec = htonl(ce->ce_ctime.sec);
ondisk->mtime.sec = htonl(ce->ce_mtime.sec);
@@ -1549,10 +1556,7 @@ static int ce_write_entry(git_SHA_CTX *c, int fd, struct cache_entry *ce)
else
name = ondisk->name;
memcpy(name, ce->name, ce_namelen(ce));
-
- result = ce_write(c, fd, ondisk, size);
- free(ondisk);
- return result;
+ return size;
}
static int has_racy_timestamp(struct index_state *istate)
@@ -1588,6 +1592,8 @@ int write_index(struct index_state *istate, int newfd)
struct cache_entry **cache = istate->cache;
int entries = istate->cache_nr;
struct stat st;
+ void *ce_ondisk = NULL;
+ int ce_ondisk_size = 0;
for (i = removed = extended = 0; i < entries; i++) {
if (cache[i]->ce_flags & CE_REMOVE)
@@ -1612,13 +1618,17 @@ int write_index(struct index_state *istate, int newfd)
for (i = 0; i < entries; i++) {
struct cache_entry *ce = cache[i];
+ int size;
+
if (ce->ce_flags & CE_REMOVE)
continue;
if (!ce_uptodate(ce) && is_racy_timestamp(istate, ce))
ce_smudge_racily_clean_entry(ce);
- if (ce_write_entry(&c, newfd, ce) < 0)
+ size = ce_prepare_ondisk_entry(ce, &ce_ondisk, &ce_ondisk_size);
+ if (ce_write(&c, newfd, ce_ondisk, size) < 0)
return -1;
}
+ free(ce_ondisk);
/* Write extension data here */
if (istate->cache_tree) {
--
1.7.8.36.g69ee2
^ permalink raw reply related
* [PATCH 3/3] Support compressing index when GIT_ZCACHE=1
From: Nguyễn Thái Ngọc Duy @ 2012-02-05 8:30 UTC (permalink / raw)
To: git; +Cc: Joshua Redstone, Nguyễn Thái Ngọc Duy
In-Reply-To: <1328430605-4566-1-git-send-email-pclouds@gmail.com>
Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>
---
cache.h | 1 +
read-cache.c | 118 ++++++++++++++++++++++++++++++++++++++++++++++++++++++---
2 files changed, 112 insertions(+), 7 deletions(-)
diff --git a/cache.h b/cache.h
index 10afd71..112bc52 100644
--- a/cache.h
+++ b/cache.h
@@ -99,6 +99,7 @@ unsigned long git_deflate_bound(git_zstream *, unsigned long);
*/
#define CACHE_SIGNATURE 0x44495243 /* "DIRC" */
+#define ZCACHE_SIGNATURE 0x4452435A /* "DRCZ" */
struct cache_header {
unsigned int hdr_signature;
unsigned int hdr_version;
diff --git a/read-cache.c b/read-cache.c
index 7b9a989..45c1712 100644
--- a/read-cache.c
+++ b/read-cache.c
@@ -1182,12 +1182,17 @@ static struct cache_entry *refresh_cache_entry(struct cache_entry *ce, int reall
return refresh_cache_ent(&the_index, ce, really, NULL, NULL);
}
-static int verify_hdr(struct cache_header *hdr, unsigned long size)
+static int verify_hdr(struct cache_header *hdr, unsigned long size,
+ int *deflated)
{
git_SHA_CTX c;
unsigned char sha1[20];
- if (hdr->hdr_signature != htonl(CACHE_SIGNATURE))
+ if (hdr->hdr_signature == htonl(CACHE_SIGNATURE))
+ *deflated = 0;
+ else if (hdr->hdr_signature == htonl(ZCACHE_SIGNATURE))
+ *deflated = 1;
+ else
return error("bad signature");
if (hdr->hdr_version != htonl(2) && hdr->hdr_version != htonl(3))
return error("bad index version");
@@ -1273,6 +1278,43 @@ static struct cache_entry *create_from_disk(struct ondisk_cache_entry *ondisk)
return ce;
}
+static int inflate_cache_entries(struct index_state *istate,
+ const unsigned char *mmap, size_t mmap_size,
+ unsigned long src_offset)
+{
+ unsigned char buf[sizeof(struct ondisk_cache_entry) + PATH_MAX];
+ struct ondisk_cache_entry *disk_ce;
+ struct cache_entry *ce;
+ struct git_zstream stream;
+ int i, status;
+
+ memset(&stream, 0, sizeof(stream));
+ stream.next_in = (unsigned char*)mmap + src_offset;
+ stream.avail_in = mmap_size - src_offset;
+ stream.next_out = buf;
+ stream.avail_out = sizeof(buf);
+ git_inflate_init(&stream);
+
+ for (i = 0; i < istate->cache_nr; i++) {
+ int remaining;
+ do {
+ status = git_inflate(&stream, Z_FINISH);
+ } while (status == Z_OK);
+
+ disk_ce = (struct ondisk_cache_entry *)buf;
+ ce = create_from_disk(disk_ce);
+ set_index_entry(istate, i, ce);
+
+ remaining = stream.next_out - (buf + ondisk_ce_size(ce));
+ memmove(buf, buf + ondisk_ce_size(ce), remaining);
+ stream.next_out = buf + remaining;
+ stream.avail_out = sizeof(buf) - remaining;
+ }
+ assert(status == Z_STREAM_END);
+ git_inflate_end(&stream);
+ return stream.next_in - mmap;
+}
+
static int read_cache_entries(struct index_state *istate,
const char *mmap, unsigned long src_offset)
{
@@ -1300,6 +1342,7 @@ int read_index_from(struct index_state *istate, const char *path)
struct cache_header *hdr;
void *mmap;
size_t mmap_size;
+ int deflated;
errno = EBUSY;
if (istate->initialized)
@@ -1329,7 +1372,7 @@ int read_index_from(struct index_state *istate, const char *path)
die_errno("unable to map index file");
hdr = mmap;
- if (verify_hdr(hdr, mmap_size) < 0)
+ if (verify_hdr(hdr, mmap_size, &deflated) < 0)
goto unmap;
istate->cache_nr = ntohl(hdr->hdr_entries);
@@ -1337,7 +1380,11 @@ int read_index_from(struct index_state *istate, const char *path)
istate->cache = xcalloc(istate->cache_alloc, sizeof(struct cache_entry *));
istate->initialized = 1;
- src_offset = read_cache_entries(istate, mmap, sizeof(*hdr));
+ if (deflated)
+ src_offset = inflate_cache_entries(istate, mmap, mmap_size,
+ sizeof(*hdr));
+ else
+ src_offset = read_cache_entries(istate, mmap, sizeof(*hdr));
istate->timestamp.sec = st.st_mtime;
istate->timestamp.nsec = ST_MTIME_NSEC(st);
@@ -1594,6 +1641,10 @@ int write_index(struct index_state *istate, int newfd)
struct stat st;
void *ce_ondisk = NULL;
int ce_ondisk_size = 0;
+ struct git_zstream stream;
+ int deflate, status;
+ unsigned char *dbuf_out;
+ unsigned char *dbuf_in;
for (i = removed = extended = 0; i < entries; i++) {
if (cache[i]->ce_flags & CE_REMOVE)
@@ -1607,7 +1658,8 @@ int write_index(struct index_state *istate, int newfd)
}
}
- hdr.hdr_signature = htonl(CACHE_SIGNATURE);
+ deflate = getenv("GIT_ZCACHE") != NULL;
+ hdr.hdr_signature = htonl(deflate ? ZCACHE_SIGNATURE : CACHE_SIGNATURE);
/* for extended format, increase version so older git won't try to read it */
hdr.hdr_version = htonl(extended ? 3 : 2);
hdr.hdr_entries = htonl(entries - removed);
@@ -1616,6 +1668,17 @@ int write_index(struct index_state *istate, int newfd)
if (ce_write(&c, newfd, &hdr, sizeof(hdr)) < 0)
return -1;
+ if (deflate) {
+ dbuf_out = xmalloc(WRITE_BUFFER_SIZE);
+ dbuf_in = xmalloc(WRITE_BUFFER_SIZE);
+ memset(&stream, 0, sizeof(stream));
+ stream.next_out = dbuf_out;
+ stream.avail_out = WRITE_BUFFER_SIZE;
+ stream.next_in = dbuf_in;
+ stream.avail_in = 0;
+ git_deflate_init(&stream, zlib_compression_level);
+ }
+
for (i = 0; i < entries; i++) {
struct cache_entry *ce = cache[i];
int size;
@@ -1625,11 +1688,52 @@ int write_index(struct index_state *istate, int newfd)
if (!ce_uptodate(ce) && is_racy_timestamp(istate, ce))
ce_smudge_racily_clean_entry(ce);
size = ce_prepare_ondisk_entry(ce, &ce_ondisk, &ce_ondisk_size);
- if (ce_write(&c, newfd, ce_ondisk, size) < 0)
- return -1;
+ if (!deflate) {
+ if (ce_write(&c, newfd, ce_ondisk, size) < 0)
+ return -1;
+ continue;
+ }
+
+ if (stream.avail_in)
+ memmove(dbuf_in, stream.next_in, stream.avail_in);
+ memcpy(dbuf_in + stream.avail_in, ce_ondisk, size);
+ stream.next_in = dbuf_in;
+ stream.avail_in += size;
+ do {
+ status = git_deflate(&stream, 0);
+ if (stream.next_out > dbuf_out) {
+ size = stream.next_out - dbuf_out;
+ if (ce_write(&c, newfd, dbuf_out, size) < 0)
+ return -1;
+ stream.next_out = dbuf_out;
+ stream.avail_out = WRITE_BUFFER_SIZE;
+ }
+ } while (status == Z_OK);
}
free(ce_ondisk);
+ if (deflate) {
+ do {
+ status = git_deflate(&stream, Z_FINISH);
+ if (stream.next_out > dbuf_out) {
+ int size = stream.next_out - dbuf_out;
+ if (ce_write(&c, newfd, dbuf_out, size) < 0)
+ return -1;
+ stream.next_out = dbuf_out;
+ stream.avail_out = WRITE_BUFFER_SIZE;
+ }
+ } while (status == Z_OK);
+
+ git_deflate_end(&stream);
+ if (stream.next_out > dbuf_out) {
+ int size = stream.next_out - dbuf_out;
+ if (ce_write(&c, newfd, dbuf_out, size) < 0)
+ return -1;
+ }
+ free(dbuf_in);
+ free(dbuf_out);
+ }
+
/* Write extension data here */
if (istate->cache_tree) {
struct strbuf sb = STRBUF_INIT;
--
1.7.8.36.g69ee2
^ permalink raw reply related
* [PATCH 1/3] read-cache: factor out cache entries reading code
From: Nguyễn Thái Ngọc Duy @ 2012-02-05 8:30 UTC (permalink / raw)
To: git; +Cc: Joshua Redstone, Nguyễn Thái Ngọc Duy
In-Reply-To: <1328430605-4566-1-git-send-email-pclouds@gmail.com>
Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>
---
read-cache.c | 32 ++++++++++++++++++++------------
1 files changed, 20 insertions(+), 12 deletions(-)
diff --git a/read-cache.c b/read-cache.c
index a51bba1..2dbf923 100644
--- a/read-cache.c
+++ b/read-cache.c
@@ -1273,10 +1273,28 @@ static struct cache_entry *create_from_disk(struct ondisk_cache_entry *ondisk)
return ce;
}
+static int read_cache_entries(struct index_state *istate,
+ const char *mmap, unsigned long src_offset)
+{
+ const char *buf = mmap + src_offset;
+ int i;
+
+ for (i = 0; i < istate->cache_nr; i++) {
+ struct ondisk_cache_entry *disk_ce;
+ struct cache_entry *ce;
+
+ disk_ce = (struct ondisk_cache_entry *)buf;
+ ce = create_from_disk(disk_ce);
+ set_index_entry(istate, i, ce);
+ buf += ondisk_ce_size(ce);
+ }
+ return buf - mmap;
+}
+
/* remember to discard_cache() before reading a different cache! */
int read_index_from(struct index_state *istate, const char *path)
{
- int fd, i;
+ int fd;
struct stat st;
unsigned long src_offset;
struct cache_header *hdr;
@@ -1319,17 +1337,7 @@ int read_index_from(struct index_state *istate, const char *path)
istate->cache = xcalloc(istate->cache_alloc, sizeof(struct cache_entry *));
istate->initialized = 1;
- src_offset = sizeof(*hdr);
- for (i = 0; i < istate->cache_nr; i++) {
- struct ondisk_cache_entry *disk_ce;
- struct cache_entry *ce;
-
- disk_ce = (struct ondisk_cache_entry *)((char *)mmap + src_offset);
- ce = create_from_disk(disk_ce);
- set_index_entry(istate, i, ce);
-
- src_offset += ondisk_ce_size(ce);
- }
+ src_offset = read_cache_entries(istate, mmap, sizeof(*hdr));
istate->timestamp.sec = st.st_mtime;
istate->timestamp.nsec = ST_MTIME_NSEC(st);
--
1.7.8.36.g69ee2
^ permalink raw reply related
* Re: Installing git-svn on Linux without root
From: Jakub Narebski @ 2012-02-05 10:11 UTC (permalink / raw)
To: Andrew Keller; +Cc: git
In-Reply-To: <AD682311-372A-4AED-B575-E77EB862ABD8@kellerfarm.com>
Andrew Keller wrote:
> On Feb 4, 2012, at 6:32 AM, Jakub Narebski wrote:
> > Andrew Keller <andrew@kellerfarm.com> writes:
> > > So, the module does exist, but not in a location included by @INC.
> >
> > From the above error message it looks like
> >
> > /homedirs/kelleran/local/lib/perl5/site_perl/5.8.8
> >
> > is in @INC, but
> > /homedirs/kelleran/local/lib64/perl5/site_perl/5.8.8/x86_64-linux-thread-multi/
> >
> > is not. Strange.
> >
> > Do you use local::lib?
>
> No - Where should I use it?
local::lib is a way to install Perl modules / packages e.g. in your
home directory. Please read and follow the instructions on local::lib
manpage ("The bootstrapping technique" section):
http://search.cpan.org/~apeiron/local-lib-1.008004/lib/local/lib.pm
Then you can install Alien::SVN (or any other Perl package) using
'cpan' client (or intstall 'cpanm' / 'App::cpanminus').
HTH
--
Jakub Narebski
Poland
^ permalink raw reply
* [PATCH] Explicitly set X to avoid potential build breakage
From: Michael @ 2012-02-05 10:41 UTC (permalink / raw)
To: git
$X is appended to binary names for Windows builds (ie. git.exe).
Pollution from the environment can inadvertently trigger this behaviour,
resulting in 'git' turning into 'gitwhatever' without warning.
Signed-off-by: Michael <kensington@astralcloak.net>
---
Makefile | 3 +++
1 files changed, 3 insertions(+), 0 deletions(-)
diff --git a/Makefile b/Makefile
index c457c34..380d96f 100644
--- a/Makefile
+++ b/Makefile
@@ -388,6 +388,9 @@ SCRIPT_SH =
SCRIPT_LIB =
TEST_PROGRAMS_NEED_X =
+# Binary suffix used for Windows builds
+X =
+
# Having this variable in your environment would break pipelines because
# you cause "cd" to echo its destination to stdout. It can also take
# scripts to unexpected places. If you like CDPATH, define it for your
--
1.7.8.4
^ permalink raw reply related
* Re: Git performance results on a large repository
From: David Barr @ 2012-02-05 11:24 UTC (permalink / raw)
To: david; +Cc: Joshua Redstone, git@vger.kernel.org
In-Reply-To: <alpine.DEB.2.02.1202042026280.6541@asgard.lang.hm>
On Sun, Feb 5, 2012 at 3:30 PM, <david@lang.hm> wrote:
> On Fri, 3 Feb 2012, Joshua Redstone wrote:
>
>> The test repo has 4 million commits, linear history and about 1.3 million
>> files. The size of the .git directory is about 15GB, and has been
>> repacked with 'git repack -a -d -f --max-pack-size=10g --depth=100
>> --window=250'. This repack took about 2 days on a beefy machine (I.e.,
>> lots of ram and flash). The size of the index file is 191 MB.
>
>
> This may be a silly thought, but what if instead of one pack file of your
> entire history (4 million commits) you create multiple packs (say every half
> million commits) and mark all but the most recent pack as .keep (so that
> they won't be modified by a repack)
>
> that way things that only need to worry about recent history (blame, etc)
> will probably never have to go past the most recent pack file or two
>
> I may be wrong, but I think that when git is looking for 'similar files' for
> delta compression, it limits it's search to the current pack, so this will
> also keep you from searching the entire project history.
I don't know if there is an easy way to determine with the with the
current tools
in git but one useful statistic for tuning packing performance is the
size of the
largest component in the delta-chain graph. The significance of this number is
that the product of window-size and maximum depth need not be larger than it.
I've found that with some older repositories I could have a depth as low as 3
and still get good performance from a moderate window size.
--
David Barr
^ permalink raw reply
* Re: [RFD] Rewriting safety - warn before/when rewriting published history
From: Ben Walton @ 2012-02-05 14:33 UTC (permalink / raw)
To: Jakub Narebski; +Cc: git
In-Reply-To: <201202042045.54114.jnareb@gmail.com>
Excerpts from Jakub Narebski's message of Sat Feb 04 14:45:53 -0500 2012:
Hi Jakub,
These items are as much about UI as anything else, I think. UI that
better helps users to know the state of their commits and branches can
only be a good thing. People that have used git for a while and are
comfortable with it may not see the need/point of these, but I think
they could both really help new users.
> In Mercurial 2.1 there are three available phases: 'public' for
> published commits, 'draft' for local un-published commits and
> 'secret' for local un-published commits which are not meant to be
> published.
How do you envision such a feature in git?
A 'draft' commit (or chain of commits) could be determined from the
push matching definitions and then marked with simple decorations in
log output...This would extend the ability of status to note that your
are X commits ahead of foo. This would see any commit on a branch
that would be pushed automatically decorated with a 'draft' status.
> While default "push matching" behavior makes it possible to have
> "secret" commits, being able to explicitly mark commits as not for
> publishing might be a good idea also for Git.
Do you see using configuration or convention to achieve this?
For example, any branch named private/foo could, by convention, be
un-pushable without a force option? Alternately, a config item
similar to the push matching stuff to allow the users to designate
un-pushable branches could work too.
Please don't take the above implementation possibilities as anything
more than a starting point for discussion as they may be deeply
flawed. I'm just tossing a few things out there as I think this is a
good discussion to have.
Thanks
-Ben
--
Ben Walton
Systems Programmer - CHASS
University of Toronto
C:416.407.5610 | W:416.978.4302
^ permalink raw reply
* Fwd: Breakage in master?
From: Erik Faye-Lund @ 2012-02-05 14:46 UTC (permalink / raw)
To: bug-gnu-gettext; +Cc: msysGit, Git Mailing List
In-Reply-To: <CABPQNSbj8QKqkdY49Y7tpAOQd53t+z6Gc5U-CS0-TZWyNz1WfQ@mail.gmail.com>
Git has recently switched to using gettext for translations, and I
have observed a breakage on Windows due to the way gettext handles
vsnprintf.
On MinGW, vsnprintf and _vsnprintf are two different implementations;
vsnprintf is from MinGW-runtime, and provides a reasonably sane
implementation. _vnsprintf on the other hand is from MSVCRT.dll, and
has some issues with it's return value. Before using gettext, the
MinGW-built version of Git called the version from mingw-runtime, and
everything worked fine. When built with MSVC, a shim was used to fixup
the bogus return value. This shim was injected through a define,
similar to what gettext does.
The shim in gettext lead to issues for Git, both on MinGW and on MSVC.
For MinGW, the problem is that libintl_vsnprintf calls _vsnprintf
rather than vsnprintf, giving us the same, broken return value that we
tried to prevent. This means that our code intended to call the
MinGW-runtime version, but gettext ended up calling the MSVCRT.dll
version. I don't find this very reasonable; a call to vsnprintf ends
up as a call to _vsnprintf.
On MSVC the problem is a bit easier to spot; libgnuintl.h.in contains
the following:
---8<---
#if !(defined vsnprintf && defined _GL_STDIO_H) /* don't override gnulib */
#undef vsnprintf
#define vsnprintf libintl_vsnprintf
extern int vsnprintf (char *, size_t, const char *, va_list);
#endif
---8<---
Uhm, what? Unless we're using Gnulib, our definition of vsnprintf
should simply be ignored?
I'm not saying figuring out what to do here is exactly trivial; but I
think undefining any definitions of vsnprintf that aren't exactly
"_vsnprintf" is dangerous. The forwarded mail below contains a
quick-fix I did locally that seems to side-step the problem for me, by
not using _vsnprintf on MinGW. But perhaps there's something better we
can do?
---------- Forwarded message ----------
From: Erik Faye-Lund <kusmabite@gmail.com>
Date: Sat, Feb 4, 2012 at 10:55 PM
Subject: Re: Breakage in master?
To: Jeff King <peff@peff.net>
Cc: Git Mailing List <git@vger.kernel.org>, msysGit
<msysgit@googlegroups.com>, Ævar Arnfjörð <avarab@gmail.com>
On Fri, Feb 3, 2012 at 1:28 PM, Erik Faye-Lund <kusmabite@gmail.com> wrote:
> On Thu, Feb 2, 2012 at 6:46 PM, Jeff King <peff@peff.net> wrote:
>> On Thu, Feb 02, 2012 at 01:14:19PM +0100, Erik Faye-Lund wrote:
>>
>>> But here's the REALLY puzzling part: If I add a simple, unused
>>> function to diff-lib.c, like this:
>>> [...]
>>> "git status" starts to error out with that same vsnprintf complaint!
>>>
>>> ---8<---
>>> $ git status
>>> # On branch master
>>> # Changes not staged for commit:
>>> # (use "git add <file>..." to update what will be committed)
>>> fatal: BUG: your vsnprintf is broken (returned -1)
>>> ---8<---
>>
>> OK, that's definitely odd.
>>
>> At the moment of the die() in strbuf_vaddf, what does errno say?
>
> If I apply this patch:
> ---8<---
> diff --git a/strbuf.c b/strbuf.c
> index ff0b96b..52dfdd6 100644
> --- a/strbuf.c
> +++ b/strbuf.c
> @@ -218,7 +218,7 @@ void strbuf_vaddf(struct strbuf *sb, const char
> *fmt, va_list ap)
> len = vsnprintf(sb->buf + sb->len, sb->alloc - sb->len, fmt, cp);
> va_end(cp);
> if (len < 0)
> - die("BUG: your vsnprintf is broken (returned %d)", len);
> + die_errno("BUG: your vsnprintf is broken (returned %d)", len);
> if (len > strbuf_avail(sb)) {
> strbuf_grow(sb, len);
> len = vsnprintf(sb->buf + sb->len, sb->alloc - sb->len, fmt, ap);
> ---8<---
>
> Then I get "fatal: BUG: your vsnprintf is broken (returned -1): Result
> too large". This goes both for both failure cases I described. I
> assume this means errno=ERANGE.
>
>> vsnprintf should generally never be returning -1 (it should return the
>> number of characters that would have been written). Since you're on
>> Windows, I assume you're using the replacement version in
>> compat/snprintf.c.
>
> No. SNPRINTF_RETURNS_BOGUS is only set for the MSVC target, not for
> the MinGW target. I'm assuming that means MinGW-runtime has a sane
> vsnprintf implementation. But even if I enable SNPRINTF_RETURNS_BOGUS,
> the problem occurs. And it's still "Result too large".
>
> So I decided to do a bit of stepping, and it seems libintl takes over
> vsnprintf, directing us to libintl_vsnprintf instead. I guess this is
> so it can ensure we support reordering the parameters with $1 etc...
> And aparently this vsnprintf implementation calls the system vnsprintf
> if the format string does not contain '$', and it's using _vsnprintf
> rather than vsnprintf on Windows. _vsnprintf is the MSVCRT-version,
> and not the MinGW-runtime, which needs SNPRINTF_RETURNS_BOGUS.
>
> So I guess I can patch libintl to call vsnprintf from MinGW-runtime instead.
>
Indeed, I just got around to testing this, and doing this on top of
gettext seems to fix the problem for me. For the MSVC, a more
elaborate fix is needed, as it doesn't have a sane vsnprintf.
---
diff --git a/gettext-runtime/intl/printf.c b/gettext-runtime/intl/printf.c
index b7cdc5d..f55023e 100644
--- a/gettext-runtime/intl/printf.c
+++ b/gettext-runtime/intl/printf.c
@@ -192,7 +192,7 @@ libintl_sprintf (char *resultbuf, const char *format, ...)
#if HAVE_SNPRINTF
-# if HAVE_DECL__SNPRINTF
+# if HAVE_DECL__SNPRINTF && !defined(__MINGW32__)
/* Windows. */
# define system_vsnprintf _vsnprintf
# else
^ permalink raw reply related
* Re: [RFD] Rewriting safety - warn before/when rewriting published history
From: Jakub Narebski @ 2012-02-05 15:05 UTC (permalink / raw)
To: Ben Walton; +Cc: git
In-Reply-To: <1328452328-sup-6643@pinkfloyd.chass.utoronto.ca>
On Sun, 5 Feb 2012, Ben Walton wrote:
> Excerpts from Jakub Narebski's message of Sat Feb 04 14:45:53 -0500 2012:
>
> Hi Jakub,
>
> These items are as much about UI as anything else, I think. UI that
> better helps users to know the state of their commits and branches can
> only be a good thing. People that have used git for a while and are
> comfortable with it may not see the need/point of these, but I think
> they could both really help new users.
As I said, 1500+ git users would like to have such feature, according
to latest Git User's Survey.
> > In Mercurial 2.1 there are three available phases: 'public' for
> > published commits, 'draft' for local un-published commits and
> > 'secret' for local un-published commits which are not meant to be
> > published.
>
> How do you envision such a feature in git?
>
> A 'draft' commit (or chain of commits) could be determined from the
> push matching definitions and then marked with simple decorations in
> log output...This would extend the ability of status to note that your
> are X commits ahead of foo. This would see any commit on a branch
> that would be pushed automatically decorated with a 'draft' status.
I think that in its basic form (treating all remotes equally) commits
in 'public' phase would be those reachable from remote-tracking branches.
Otherwise commits would be in 'draft' phase, unless explicitly marked
as 'secret' (it we implement 'secret' phase, that is).
The safety new I think of would (similarly to Mercurial phases) prevent
or warn about amending published commit, and rebasing commits which were
already published (in 'public' phase). That would require modifications
to git-commit and git-amend, I think...
Maybe even Git could refuse or warn on the local side about non
fast-forward update of public branch, to help users of third-party tools.
> > While default "push matching" behavior makes it possible to have
> > "secret" commits, being able to explicitly mark commits as not for
> > publishing might be a good idea also for Git.
>
> Do you see using configuration or convention to achieve this?
>
> For example, any branch named private/foo could, by convention, be
> un-pushable without a force option? Alternately, a config item
> similar to the push matching stuff to allow the users to designate
> un-pushable branches could work too.
I'm not sure, but the config item might be a good solution. Git would
skip publishing 'secret' commits (commits from 'secret' branch) if it
would otherwise publish it due to glob refspec, and refuse (or warn)
publishing 'secret' branches explicitly.
Currently if you use default "push matching", then those branches that
you didn't push explicitly wouldn't be pushed. But that does not prevent
pushing them by accident, and does not give UI to check if branch is
private or not (e.g. to use in git-aware shell prompt).
--
Jakub Narebski
Poland
^ permalink raw reply
* Fw: [RFD] Rewriting safety - warn before/when rewriting published history
From: Philip Oakley @ 2012-02-05 15:10 UTC (permalink / raw)
To: Git List
Oops, forget 'reply all' to the list.
Philip
----- Original Message -----
From: "Philip Oakley" <philipoakley@iee.org>
To: "Jakub Narebski" <jnareb@gmail.com>
Sent: Sunday, February 05, 2012 2:31 PM
Subject: Re: [RFD] Rewriting safety - warn before/when rewriting published
history
> From: "Jakub Narebski" <jnareb@gmail.com>
> Sent: Saturday, February 04, 2012 7:45 PM
>> Git includes protection against rewriting published history on the
>> receive side with fast-forward check by default (which can be
>> overridden) and various receive.deny* configuration variables,
>> including receive.denyNonFastForwards.
>>
>> Nevertheless git users requested (among others in Git User's Survey)
>> more help on creation side, namely preventing rewriting parts of
>> history which was already made public (or at least warning that one is
>> about to rewrite published history). The "warn before/when rewriting
>> published history" answer in "17. Which of the following features would
>> you like to see implemented in git?" multiple-choice question in latest
>> Git User's Survey 2011[1] got 24% (1525) responses.
>>
>> [1]: https://www.survs.com/results/Q5CA9SKQ/P7DE07F0PL
>>
>> So people would like for git to warn them about rewriting history before
>> they attempt a push and it turns out to not fast-forward.
>>
>
> Another area that is implicitly related is that of (lack of) publication
> of sub-module updates. A mechanisms that, in the super project, knows the
> status of the (local) submodules, such as where they would be sourced
> from, i.e. what was last pushed & where, could help in such instances.
>
>>
>> What prompted this email is the fact that Mercurial includes support for
>> tracking which revisions (changesets) are safe to modify in its 2.1
>> latest version:
>>
>> http://lwn.net/Articles/478795/
>> http://mercurial.selenic.com/wiki/WhatsNew
>>
>> It does that by tracking so called "phase" of a changeset (revision).
>>
>> http://mercurial.selenic.com/wiki/Phases
>> http://mercurial.selenic.com/wiki/PhasesDevel
>>
>> http://www.logilab.org/blogentry/88203
>> http://www.logilab.org/blogentry/88219
>> http://www.logilab.org/blogentry/88259
>>
>>
>> While we don't have to play catch-up with Mercurial features, I think
>> something similar to what Mercurial has to warn about rewriting
>> published history (amend, rebase, perhaps even filter-branch) would
>> be nice to have. Perhaps even follow UI used by Mercurial, and/or
>> translating its implementation into git terms.
>>
>> In Mercurial 2.1 there are three available phases: 'public' for
>> published commits, 'draft' for local un-published commits and
>> 'secret' for local un-published commits which are not meant to
>> be published.
>>
>> The phase of a changeset is always equal to or higher than the phase
>> of it's descendants, according to the following order:
>>
>> public < draft < secret
>>
>> Commits start life as 'draft', and move to 'public' on push.
>
> Recording where they wer pushed to would be useful for synchronising
> sub-modules and their super projects. That is, giving remote users a clue
> as to where they might find mising sub-modules.
>
>>
>> Mercurial documentation talks about phase of a commit, which might
>> be a good UI, ut also about commits in 'public' phase being "immutable".
>> As commits in Git are immutable, and rewriting history is in fact
>> re-doing commits, this description should probably be changed.
>>
>> While default "push matching" behavior makes it possible to have
>> "secret" commits, being able to explicitly mark commits as not for
>> publishing might be a good idea also for Git.
>>
>
> Being able to mark temporary, out of sequence or other hacks as Secret
> could be useful, as would recording where Public commits had been sent.
>
>>
>> What do you think about this?
>> --
>> Jakub Narebski
>> Poland
>> --
>
> Philip Oakley
>
^ permalink raw reply
* Re: Git performance results on a large repository
From: Nguyen Thai Ngoc Duy @ 2012-02-05 15:17 UTC (permalink / raw)
To: Tomas Carnecky; +Cc: Joshua Redstone, git@vger.kernel.org
In-Reply-To: <4F2E99C2.7090609@dbservice.com>
On Sun, Feb 5, 2012 at 10:01 PM, Tomas Carnecky <tom@dbservice.com> wrote:
> On 2/4/12 7:53 AM, Nguyen Thai Ngoc Duy wrote:
>>
>> On Fri, Feb 3, 2012 at 9:20 PM, Joshua Redstone<joshua.redstone@fb.com>
>> wrote:
>>>
>>> I timed a few common operations with both a warm OS file cache and a cold
>>> cache. i.e., I did a 'echo 3 | tee /proc/sys/vm/drop_caches' and then
>>> did
>>> the operation in question a few times (first timing is the cold timing,
>>> the next few are the warm timings). The following results are on a
>>> server
>>> with average hard drive (I.e., not flash) and> 10GB of ram.
>>>
>>> 'git status' : 39 minutes cold, and 24 seconds warm.
>>>
>>> 'git blame': 44 minutes cold, 11 minutes warm.
>>>
>>> 'git add' (appending a few chars to the end of a file and adding it): 7
>>> seconds cold and 5 seconds warm.
>>>
>>> 'git commit -m "foo bar3" --no-verify --untracked-files=no --quiet
>>> --no-status': 41 minutes cold, 20 seconds warm. I also hacked a version
>>> of git to remove the three or four places where 'git commit' stats every
>>> file in the repo, and this dropped the times to 30 minutes cold and 8
>>> seconds warm.
>>
>> Have you tried "git update-index --assume-unchaged"? That should
>> reduce mass lstat() and hopefully improve the above numbers. The
>> interface is not exactly easy-to-use, but if it has significant gain,
>> then we can try to improve UI.
>>
>> On the index size issue, ideally we should make minimum writes to
>> index instead of rewriting 191 MB index. An improvement we could do
>> now is to compress it, reduce disk footprint, thus disk I/O. If you
>> compress the index with gzip, how big is it?
>
> If you're not afraid to add filesystem-specific code to git, you could
> leverage the btrfs find-new command (or use the ioctl directly) to quickly
> find changed files since a certain point in time. Other CoW filesystems may
> have similar mechanisms. You could for example store the last generation id
> in an index extension, that's what those extensions are for, right?
Sure they could be stored as index extensions. I'm more concerned of
the index size. I guess fs-specific code, if properly implemented
(e.g. clean, handling repos crossing fs boundaries, moving repos...),
may get Junio's approval. There were also talks of implementing NTFS's
journal (or something) on msysgit for similar goal.
--
Duy
^ permalink raw reply
* Re: Git performance results on a large repository
From: Tomas Carnecky @ 2012-02-05 15:01 UTC (permalink / raw)
To: Nguyen Thai Ngoc Duy; +Cc: Joshua Redstone, git@vger.kernel.org
In-Reply-To: <CACsJy8DkLCK0ZUKNz_PJazsxjsRbWVVZwjAU5n2EAjJfCYtpoQ@mail.gmail.com>
On 2/4/12 7:53 AM, Nguyen Thai Ngoc Duy wrote:
> On Fri, Feb 3, 2012 at 9:20 PM, Joshua Redstone<joshua.redstone@fb.com> wrote:
>> I timed a few common operations with both a warm OS file cache and a cold
>> cache. i.e., I did a 'echo 3 | tee /proc/sys/vm/drop_caches' and then did
>> the operation in question a few times (first timing is the cold timing,
>> the next few are the warm timings). The following results are on a server
>> with average hard drive (I.e., not flash) and> 10GB of ram.
>>
>> 'git status' : 39 minutes cold, and 24 seconds warm.
>>
>> 'git blame': 44 minutes cold, 11 minutes warm.
>>
>> 'git add' (appending a few chars to the end of a file and adding it): 7
>> seconds cold and 5 seconds warm.
>>
>> 'git commit -m "foo bar3" --no-verify --untracked-files=no --quiet
>> --no-status': 41 minutes cold, 20 seconds warm. I also hacked a version
>> of git to remove the three or four places where 'git commit' stats every
>> file in the repo, and this dropped the times to 30 minutes cold and 8
>> seconds warm.
> Have you tried "git update-index --assume-unchaged"? That should
> reduce mass lstat() and hopefully improve the above numbers. The
> interface is not exactly easy-to-use, but if it has significant gain,
> then we can try to improve UI.
>
> On the index size issue, ideally we should make minimum writes to
> index instead of rewriting 191 MB index. An improvement we could do
> now is to compress it, reduce disk footprint, thus disk I/O. If you
> compress the index with gzip, how big is it?
If you're not afraid to add filesystem-specific code to git, you could
leverage the btrfs find-new command (or use the ioctl directly) to
quickly find changed files since a certain point in time. Other CoW
filesystems may have similar mechanisms. You could for example store the
last generation id in an index extension, that's what those extensions
are for, right?
tom
^ permalink raw reply
* Re: [RFD] Rewriting safety - warn before/when rewriting published history
From: Jakub Narebski @ 2012-02-05 16:15 UTC (permalink / raw)
To: Philip Oakley; +Cc: git
In-Reply-To: <CAFA910035B74E56A52A96097E76AC39@PhilipOakley>
Please don't remove git mailing list from Cc... Oh, I see that you
forgot to send to list, but resend your email there.
On Sun, 5 Feb 2012, Philip Oakley wrote:
> From: "Jakub Narebski" <jnareb@gmail.com>
> Sent: Saturday, February 04, 2012 7:45 PM
> > Git includes protection against rewriting published history on the
> > receive side with fast-forward check by default (which can be
> > overridden) and various receive.deny* configuration variables,
> > including receive.denyNonFastForwards.
> >
> > Nevertheless git users requested (among others in Git User's Survey)
> > more help on creation side, namely preventing rewriting parts of
> > history which was already made public (or at least warning that one is
> > about to rewrite published history). The "warn before/when rewriting
> > published history" answer in "17. Which of the following features would
> > you like to see implemented in git?" multiple-choice question in latest
> > Git User's Survey 2011[1] got 24% (1525) responses.
> >
> > [1]: https://www.survs.com/results/Q5CA9SKQ/P7DE07F0PL
> >
> > So people would like for git to warn them about rewriting history before
> > they attempt a push and it turns out to not fast-forward.
>
> Another area that is implicitly related is that of (lack of) publication of
> sub-module updates. A mechanisms that, in the super project, knows the
> status of the (local) submodules, such as where they would be sourced from,
> i.e. what was last pushed & where, could help in such instances.
"Better support for submodules" had almost the same number of requests
in the latest Git User's Survey 2011 (25% which means 1582 responses).
Remembering when to do recursive push and where would be a very nice thing.
[...]
> Recording where they were pushed to would be useful for synchronising
> sub-modules and their super projects. That is, giving remote users a clue as
> to where they might find mising sub-modules.
Is it a matter of correctly writing configuration with current git?
I don't use submodules myself, so I cannot say.
> > Mercurial documentation talks about phase of a commit, which might
> > be a good UI, ut also about commits in 'public' phase being "immutable".
> > As commits in Git are immutable, and rewriting history is in fact
> > re-doing commits, this description should probably be changed.
> >
> > While default "push matching" behavior makes it possible to have
> > "secret" commits, being able to explicitly mark commits as not for
> > publishing might be a good idea also for Git.
> >
>
> Being able to mark temporary, out of sequence or other hacks as Secret could
> be useful, as would recording where Public commits had been sent.
Marking as 'secret' must I think be explicit, but I think 'public' phase
should be inferred from remote-tracking branches. The idea of phases is
to allow UI to ask about status of commits: can we amend / rebase it or
not, can we push it or not.
--
Jakub Narebski
Poland
^ permalink raw reply
* [ANNOUNCE] libgit2 v0.16.0
From: Vicent Marti @ 2012-02-05 16:28 UTC (permalink / raw)
To: libgit2, git, git-dev
Hello everyone,
another minor libgit2 release is here, albeit slightly delayed. This
one ships from Brussels, damn it's cold.
The release has been tagged at:
https://github.com/libgit2/libgit2/tree/v0.16.0
A dist package can be found at:
https://github.com/downloads/libgit2/libgit2/libgit2-0.16.0.tar.gz
Updated documentation can be found at:
http://libgit2.github.com/libgit2/
The full change log follows after the message.
Cheers,
Vicent
===================================
libgit2 v0.16.0 "Dutch Fries"
This lovely and much delayed release of libgit2 ships from the cold city
of Brussels, which is currently hosting FOSDEM 2012.
There's been plenty of changes since the latest stable release, here's a
full summary:
- Git Attributes support (see git2/attr.h)
There is now support to efficiently parse and retrieve information
from `.gitattribute` files in a repository. Note that this
information is not yet used e.g. when checking out files.
- .gitignore support
Likewise, all the operations that are affected by `.gitignore` files
now take into account the global, user and local ignores when
skipping the relevant files.
- Cleanup of the object ownership semantics
The ownership semantics for all repository subparts (index, odb,
config files, etc) has been redesigned. All these objects are now
reference counted, and can be hot-swapped in the middle of
execution, allowing for instance to add a working directory and an
index to a repository that was previously opened as bare, or to
change the source of the ODB objects after initialization.
Consequently, the repository API has been simplified to remove all
the `_openX` calls that allowed setting these subparts *before*
initialization.
- git_index_read_tree()
Git trees can now be read into the index.
- More reflog functionality
The reference log has been optimized, and new API calls to rename
and delete the logs for a reference have been added.
- Rewrite of the References code with explicit ownership semantics
The references code has been mostly rewritten to take into account
the cases where another Git application was modifying a repository's
references while the Library was running.
References are now explicitly loaded and free'd by the user, and
they may be reloaded in the middle of execution if the user suspects
that their values may have changed on disk. Despite the new
ownership semantics, the references API stays the same.
- Simplified the Remotes API
Some of the more complex Remote calls have been refactored into
higher level ones, to facilitate the usual `fetch` workflow of a
repository.
- Greatly improved thread-safety
The library no longer has race conditions when loading objects from
the same ODB and different threads at the same time. There's now
full TLS support, even for error codes. When the library is built
with `THREADSAFE=1`, the threading support must be globally
initialized before it can be used (see `git_threads_init()`)
- Tree walking API
A new API can recursively traverse trees and subtrees issuing callbacks for
every single entry.
- Tree diff API
There is basic support for diff'ing an index against two trees.
- Improved windows support
The Library is now codepage aware under Windows32: new API calls
allow the user to set the default codepage for the OS in order to
avoid strange Unicode errors.
^ permalink raw reply
* [RFC/PATCH] git-new-workdir: add option --rsync
From: Michael Schubert @ 2012-02-05 16:27 UTC (permalink / raw)
To: git
Currently, git-new-workdir doesn't allow you to rather duplicate your
old workdir instead of creating a new one based on a branch. This can be
annoying when you are used to carry untracked helper scripts in your
workdir (e.g. test.sh).
Add a new option -s | --rsync to allow users to "sync" the new workdir.
Signed-off-by: Michael Schubert <mschub@elegosoft.com>
---
Maybe we even want to stop after rsync'ing without doing the checkout -f.
Opinions?
contrib/workdir/git-new-workdir | 31 +++++++++++++++++++++++++++++--
1 file changed, 29 insertions(+), 2 deletions(-)
diff --git a/contrib/workdir/git-new-workdir b/contrib/workdir/git-new-workdir
index 75e8b25..bac75f9 100755
--- a/contrib/workdir/git-new-workdir
+++ b/contrib/workdir/git-new-workdir
@@ -10,9 +10,30 @@ die () {
exit 128
}
-if test $# -lt 2 || test $# -gt 3
+if test $# -lt 2
then
- usage "$0 <repository> <new_workdir> [<branch>]"
+ usage "$0 [-s | --rsync] <repository> <new_workdir> [<branch>]"
+fi
+
+rsync=
+
+while test $# != 0
+do
+ case "$1" in
+ -s|--rsync)
+ rsync=t ;;
+ *)
+ break ;;
+ esac
+ shift
+done
+
+if test -n "$rsync"
+then
+ if ! $(hash rsync 2>/dev/null)
+ then
+ die "cannot find rsync"
+ fi
fi
orig_git=$1
@@ -77,6 +98,12 @@ done
cd "$new_workdir"
# copy the HEAD from the original repository as a default branch
cp "$git_dir/HEAD" .git/HEAD
+
+if test -n "$rsync"
+then
+ rsync --archive --exclude '.*' "../$orig_git/" .
+fi
+
# checkout the branch (either the same as HEAD from the original repository, or
# the one that was asked for)
git checkout -f $branch
--
1.7.9.230.gbd302
^ permalink raw reply related
page: next (older) | prev (newer) | latest
- recent:[subjects (threaded)|topics (new)|topics (active)]
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox