If you've got RTX 5090, maybe try ninfer (https://github.com/Neroued/ninfer). Folks over on /r/localllama have been reporting wild prefill/token gen speeds with ninfer (NVFP4; 256k ctx).
Looks like slope benchmarks and results, and as usually, people are mixing MTP numbers with non MTP numbers.
Or just 100 token input benchmarks.
Or just failed ones as actual measures.