I'm still not sure how this was done. Here's what I have gathered so far from the post: The author has two images taken unknown location. He scaled them and then marked points on each. Then he matched up points on two images manually. Now he has dx and dy for each point on one image, relative to other. Now he asserts that bigger dx means nearer to camera. So I'm thinking he takes some proportionality constant to get z = c * dx.
But wouldn't that produce pretty arbitrary shape depending on value of c?
I think he used something like z = c * dy since earlier in the post he mentioned the photos were from a similar direction but different elevations. He compared this to your how your eyes compute distance but in this case the difference is vertical.
And don't forget "In this case because the precise location and elevation of the photographers isn't known this is slightly more art than science, but it is still fun!"
A question I had was how he made the correspondences between points in the two bolts. I have a hunch that he just traversed the two lightning bolts separately and said point 1 in A corresponds to point 1 in B, ... Point N in A corresponds to point N in B, etc. if this is the case, you would expect dx and dy to grow as from top to bottom and hence his reconstructed depth to become closer from top to bottom, which is what happens.
I think he just assumed two orthographic projections of the same points. The usual framework for thinking precisely about this kind of situation is https://en.wikipedia.org/wiki/Epipolar_geometry
But wouldn't that produce pretty arbitrary shape depending on value of c?