Hm but the date is stored inside of the commit. The only way we can know that a commit's date is authentic... is through its hash. If I can forge commits with any SHA1 hash at will, I can make a repository whose head commit has the same SHA1 as the one in torvalds:
/linux but where any commit was replaced by a malicious commit with the same SHA1 and a fake date. You have no way to detect that my repo is inauthentic other than through a deep history comparison. The whole idea behind a merkle tree is that just checking the hash of the top is sufficient to know the identity of the whole tree.
I don't know what the solution is, but I'm inclined to believe that any repo with a single SHA1 commit is as weak as a repo with all SHA1 commits.
If there is one way enforcement (i.e. there is one point where the last SHA1 commit was signed by first SHA256 commit), I think it should be safe ?
The "commit before" might be compromised, but the git commits refer a snapshot of a tree + a list of previous commit IDs, so the "new" SHA256 commit will not have any files altered
Every file (indirectly) referred to by a SHA256 commit using a SHA1 hash in some tree object can still be spoofed. Fixing that requires rehashing all objects and recreating all tree objects so the tree objects referred to by SHA256 commits are purely made up of object references computed by SHA256.
My reading of the docs is that when fetching/pushing from a SHA-256 repo, all objects are referred to (in the packfile fetched) by their SHA-256 names — then, when fetching, you locally compute the SHA-1 of each object for the translation table.
Presuming there’s some validation that SHA-1 names are unique, then that should be safe — the only way I can see one could do a pre-image attack is either fetching from a SHA-1 server (because then you don’t get the SHA-256 object name), which requires a second pre-image attack on SHA-1 (known to be feasible); or by having a second pre-image attack against both SHA-1 and SHA-256 simultaneously (and SHA-256 is still believed to be secure).
The date that a repo receives a commit is known to that repo. And a repo can stop accepting new SHA1 objects. And a SHA256 object could have a flag that says that no SHA1 objects may ever reference it.
The design of Git, as a Merkle tree, is meant to allow for use cases like this:
* I host a mirror of the Linux git repo.
* You download Linux from my mirror.
* You check out a commit, say fd179f8a05be3ccae366b9b96e176b51fbe54aab, which you know is a genuine commit through some out-of-band mechanism (mailing list, GitHub web interface, a line in a Nix file, whatever).
* You check whether the repository I gave you is legitimate or not by re-computing the hash of the commit which I claimed was fd179f8a05be3ccae366b9b96e176b51fbe54aab. If it comes out to be fd179f8a05be3ccae366b9b96e176b51fbe54aab, you know it's legitimate. If it doesn't, you know it's fake.
This is a completely normal use of Git. People download from mirrors all the time. People rely on commit hashes to identify a specific source tree. People trust that if whatever the mirror gave them hashes to the right value, it's genuine. That way, you don't have to trust the mirror.
If I can forge my own commits to have any hash I want, this whole model breaks down. I can replace some old commit in the repo with my own forged commit with the same hash, and when you download a copy of the Linux repo from my mirror, you'll receive a repo with malicious content, but it'll hash to the same fd179f8a05be3ccae366b9b96e176b51fbe54aab hash as a genuine repo would. This breaks the security model of Git.
> You check out a commit, say fd179f8a05be3ccae366b9b96e176b51fbe54aab, which you know is a genuine commit through some out-of-band mechanism
That's a 160 bit hash, which is SHA-1, which has the security properties of SHA-1.
Suppose you check out a commit with a given SHA-256 hash. That commit object represent the root of a tree where all the edges are hashes (and types, etc). I'm suggesting one of two designs:
a) (Simpler but weaker) If Linus has published that commit, then he is confident that he hasn't pulled in any too-new SHA-1 hashes and that there are no collisions present in what he thinks the tree is. So, by induction on the traversal depth, there is only one actual object identified by each edge, and those objects contain the hashes of their child edges, so those hashes are all correct.
This breaks if there is a malicious collision already in the tree.
b) (Stronger but higher overhead and more complex) There would be an object or objects, discoverable from the root by following only SHA-256 edges, that encode a duplicate-free mapping from SHA-1 hash to SHA-256 hash. The client finds and parses that and then, as it traverses the tree, each time it reads a SHA-1 hash, it computes the SHA-1 and SHA-256 hash of the referenced object, verifies that the pair is in the mapping and also verifies that the SHA-1 hash matches what the edge requires.
I think that (b) is genuinely cryptographically secure in the sense that, if you can construct a commit that has the same SHA-256 hash as an official upstream commit but different contents, then there is necessarily a SHA-256 collision.
For A), I don't understand what the point is? I never mentioned what Linus is confident about, I talked about what you can verify when you pull from my mirror. I could replace a commit from 2010 with a malicious one
For B), I would think this could work, but it's a completely different solution from what you proposed and what I responded to.
I think I stand by my second proposal. I also think it's absurd that, after all these years, upstream git still can't figure out a credible migration plan.
Maybe from your mirror. The collision attack that's been discovered is not a preimage attack though. It generates two colliding objects with the same uncontrollable hash, so you can't replace an existing hashed object, you have to sneak in one half of the pair.
More importantly we just shouldn't use your mirror if we don't trust you. If you're evil you're probably lying about all the tags and branch tips anyway.
I don't know what the solution is, but I'm inclined to believe that any repo with a single SHA1 commit is as weak as a repo with all SHA1 commits.