The chart showing jxl-rs taking 1209ms to decode a 5456x3632 image in a single thread is a little damning, but is there an equivalent benchmark showing multithreaded decode performance? Clients tend to have an abundance of threads these days.
The JPEG XL report [1] measured between 240-270 Megapixels/s on 6 cores using the C++ implementation (disclosure: I was responsible for its SIMD/threading), about twice as fast as the then-current libaom.
Measuring on a single core is deeply misleading because our code was designed to scale well. I believe AVIF requires tiling in order to parallelize, which causes artifacts at tile boundaries.
It’s a measurement error: the `hyperfine` command uses `jxl_cli --speedtest` which, by default, does a warmup run that ends up in hyperfine’s measured time.
The time would be less, but you still need to power those cores. Single thread performance can be a good indicator of what it'll do the battery on your phone.
CPU power is proportional to frequency^2. Running on 4-6 little/efficiency cores (which are widespread on mobile) is likely faster than one big core, and uses less energy.
jxl-rs is a relatively new implementation, not yet optimised for speed. It’s hard to say how much faster it’ll get, but I assume it would at least get near libjxl’s level.