I'm curious about the general effectiveness of the latest versions of GHC in optimizing single-threaded code. We know of huge gains with the parallelism libraries, but I'm wondering what the consensus is for code that's more inherently sequential.
When you say inliner, are you referring to LLVM's inliner or something closer to the front end? I know LLVM's is pretty aggressive as is -- are you able to get benefits from it?