I believe the branch target buffer on modern x86 hardware predicts and preloads indirect calls too. This analysis claims that Core 2 and later chips predict indirect branches quite well: http://www.agner.org/optimize/microarchitecture.pdf
That's certainly true. But if you are consuming fewer BTB resources by making more direct calls, then the CPU can spend its BTBs accelerating things like predictable virtual function calls and the like.
Yes, this. And the BTB is always slower than a direct call.
The first call always takes a big hit, only then it will be cached and subsequent calls are about the same speed is direct calls, with 1-2 cycs overhead. The address rarely changes, so it will always be in the first level BTB. But still.