You misrepresent the non cs ones by a factor of 5-10, even on cpu. First, you’re calling the prng twice per aes single call. That seems pretty dishonest already.
And using Golang? That ludicrously slow also.
To tell us what a cpu can do, do them using SIMD, you get pipelining and then many values per clock. Now try that with AES. Oh, you cannot, it’s not supported.
And people needing lots at full speed will do them on GPUs.
There is no world where even HW accelerated crypto comes close to non cs prngs on mainstream HPC systems.