What are git commit generation numbers (hacker news link) and what are their significance?
What are git commit generation numbers (hacker news link) and what are their significance?
Just to add to siri's answer, "Commit Generation Numbers" are:
A commit's generation is its height in the history graph, as measured from the farthest root. It is defined as:
- If the commit has no parents, then its generation is 0.
- Otherwise, its generation is 1 more than the maximum of its parents generations.
Linus Torwald (yester, July 14th):
Ok, so I see that the old discussion about generation numbers has resurfaced.
And I have to say, with six years of git use, I think it's not a coincidence that the notion of generation numbers has come up several times over the years: I think the lack of them is literally the only real design mistake we have.
[...]
It actually came up as early as July 2005, so the "let's use generation numbers in commits" thing is really old.
I think it's entirely reasonable to say that the issue basically boils down to one git question: "can commit X be an ancestor of commit Y" (as a way to basically limit certain algorithms from having to walk all the way down). We've used commit dates for it, and realistically it really has worked very well. But it was always a broken heuristic.
So yes, I personally see generation counters as a way to do the commit date comparisons right. And it would be perfec
The problem (as implied in the thread on git@vger.kernel.org) is that the DAG direction that we trust is counted in the reverse direction, from branch head back through parentage. The generation numbers (even if recorded at commit time) are counted through descendants. Plus we mess with the perceived history often in our different (distributed) repos - hence all the issues.
Just read Linus's latest, apart from his misreading about renames (I think George Spelvin was agreeing with him - do not record renames within the repo, simply take snapshots), he does point out that:
the very basic design of git is all about incomplete DAG traversal. The DAG traversal part is pretty obvious and simple, but the partial thing really is very very important.".
Thus essentially a pre-recorded commit "generation" number would tell you how far (the maximum) you still have to go to the bottom (root), so if you can trust it, then you can make the choice about stopping an incomplete DAG traversal. Without it you would have to go the whole way to the root, which is inefficient.
So I think I've changed my mind now I realise it is a stopping criteria. That's not to say that some (locally calculated) cache might not speed up some searches.