I wouldn't be so pessimistic about the possibility to produce useful specifications of real-world systems.
One recent successful large-scale verification is the CompCert C compiler. The compiler is verified relative to a formal specification of a large subset of C. This specification is in the order of <2000 lines of Coq code (see http://gallium.inria.fr/~xleroy/publi/Clight.pdf ). So it is an example of a specification that is much smaller than an implementation and that can be verified manually.
> The striking thing about our CompCert results is that the middle-
end bugs we found in all other compilers are absent. As of early 2011,
the under-development version of CompCert is the only compiler we
have tested for which Csmith cannot find wrong-code errors. This is
not for lack of trying: we have devoted about six CPU-years to the
task.
One recent successful large-scale verification is the CompCert C compiler. The compiler is verified relative to a formal specification of a large subset of C. This specification is in the order of <2000 lines of Coq code (see http://gallium.inria.fr/~xleroy/publi/Clight.pdf ). So it is an example of a specification that is much smaller than an implementation and that can be verified manually.
Also, there is good evidence that the specification is correct. From https://www.cs.utah.edu/~regehr/papers/pldi11-preprint.pdf :
> The striking thing about our CompCert results is that the middle- end bugs we found in all other compilers are absent. As of early 2011, the under-development version of CompCert is the only compiler we have tested for which Csmith cannot find wrong-code errors. This is not for lack of trying: we have devoted about six CPU-years to the task.