I guess I misunderstood the paper, that natural language proofs are not mastery/understanding because they can be ambiguous, miss-interpreted. Natural language can be misinterpreted enough that it is not a given to be correct.
And AI was misinterpreting what a human meant in the ambiguous natural language.