No offense to the creators and not to discourage them, but the whole process already looks so arduous and antique. We're not there yet but I bet that with GPT-x-video-diffusion we get similar results with MUCH less work rather soon. BTW Google just released the other way: a new video-to-text model today.
But: this is about adapting stable diffusion to video. Pretty cool.