I wish the author had explained how he went from identifying that the issue is stalled references to DRAM to the CYCLE_ACTIVITY.STALLS_L3_MISS performance event. Similarly from DRAM_Bound to MEM_LOAD_RETIRED.L3_MISS_PS. My complaint is that there's still a lot of magic here that requires carefully reading the Intel manuals that is elided in this post. That said, thanks to the author for the post---it is still very useful.
Hi, I'm glad you like the article.
The process how I went from TMAM metric to particular event that was used to calculate it is describe in the TMAM metrics table:
https://download.01.org/perfmon/TMA_Metrics.xlsx
In the same row for DRAM_bound metric there is precise event PEBS specified that we can use for locating the issue. Sampling on the precise event will let us detect exact place in the code where we have the most amount of L3 misses.
Edit: the author has another post that covers this information https://dendibakh.github.io/blog/2018/06/01/PMU-counters-and...