There's an interesting comparison in the creative/literary end of llm output too, they're in my experience, dreadful with anatomy. Like, it knows humans have hands, heads, etc, but often times a seemingly limited concept of how anything is connected, or degrees of freedom. (e.g., Why yes, certainly there are many examples of humans rotating their torsos 180ยบ at the hip, seems perfectly cromulent)
I honestly don't know if an image model would help, or if it might analyze the output and go "13 fingers? ship it!" anyway.
I honestly don't know if an image model would help, or if it might analyze the output and go "13 fingers? ship it!" anyway.