JVET-AK0170 AHG15: On the compression of depth maps from auxiliary data [J. Sauer, T. Solovyev, E. Alshina (Huawei)]
Sequences in classes G1/G3 of the “Common Test Conditions (CTC) for gaming applications” (JVET-AJ2027) are provided along with auxiliary data, which consists of auxiliary meta data (e.g. camera parameters) and auxiliary video data (e.g. a depth map or a motion map). This contribution is a prelimiary study of compressing the auxiliary depth maps.
Depth maps are converted into an integer representation, which is then compressed as aux picture. PSNR is calculated using the depth maps for global motion compensation between subsequent frames. This cannot capture object motion, or newly appearing content, which is indicated by the fact that this PSNR sometimes saturates at relatively low values.
However, as saturation occurs at around 14-16 bit of depth map integer precision (without compression), this may indicate that FP precision might not be necessary.
Before further investigating compression, it should be clarified by which precision depth information is needed at the receiver end by gaming applications.
As a similar example, in applications of MV-HEVC and 3D-HEVC, depth maps were expected to be needed for view synthesis, for which a certain precision of the depth map was required to achieve a certain quality, and then both video and depth were compressed together.
Further study in an AHG was requested on these aspects.
AHG16: Generative face video (2)
Contributions in this area were discussed at 1225–1320 on Saturday 18 Jan. 2025 (chaired by JRO).
Also refer to section 6.1.4.