One possible issue is that if sampling is even slightly biased, it can incorrectly estimate the relative frequency of different points/functions in tight inner function calls (which can't happen with infrequently called functions with a long runtime)?