Oh, and another point: instead of bundling an image with the video stream, they could "splice" in near lossless I-frames into the video stream wherever there is a scene change, then extract them on the decoder end. This would have some significant advantages:
- Full backwards compatibility. The "AI decoder" encoded video would play in any regular video player
- No wasted bitrate encoding the same frames twice.
- No complicated colorspace conversion trickery needed to match the (RGB?) JXL image with the video.
- Theoretically, a full GPU workflow with no CPU needed, though researchers would probably do most work on the CPU if prototyping it with VapourSynth or PySceneDetect.
- The DL network could be formatted/trained in YUV420 (or a colorspace with a similar shape) instead of RGB, which could make it much faster/smaller.
This is not that crazy either. av1an already does something like this in AV1 and HEVC, splitting the video into scenes, running the jobs in parallel and then splicing them together.
- Full backwards compatibility. The "AI decoder" encoded video would play in any regular video player
- No wasted bitrate encoding the same frames twice.
- No complicated colorspace conversion trickery needed to match the (RGB?) JXL image with the video.
- Theoretically, a full GPU workflow with no CPU needed, though researchers would probably do most work on the CPU if prototyping it with VapourSynth or PySceneDetect.
- The DL network could be formatted/trained in YUV420 (or a colorspace with a similar shape) instead of RGB, which could make it much faster/smaller.
This is not that crazy either. av1an already does something like this in AV1 and HEVC, splitting the video into scenes, running the jobs in parallel and then splicing them together.