Can you share an example of this happening? I am curious. We can get static videos if our model doesn't recognize it as a face (e.g. an Apple with a face, or sketches). Here is an example: https://toinfinityai.github.io/v2-launch-page/static/videos/...
I would be curious if you are getting this with more normal images.
I got it with a more normal image which was two frames from a TV show[1]; with "crop face" on, your model finds the face and animates it[2] and with crop face off the picture was static... just tried to reproduce to show you and now instead it's animated both faces.
One small issue I've encountered is that sometimes images remain completely static. Seems to happen when the audio is short - 3 to 5 seconds long.