JVET-AN0263 AHG15: Transformer based depthmaps reconstruction for gaming content [V. Zakharchenko (Nokia)]
Sequences in classes G1/G3 of the JVET-AJ2027 Common Test Conditions (CTC) for gaming applications are accompanied with auxiliary data, including depth map. This contribution provides a study of derived depth maps precision for gaming content. It is observed that the depth map information can be matched to the results of the preliminary study of compressing the auxiliary depth maps as reported in JVET-AK0170. It is suggested to port the solution to NNVC code base for further study and evaluation.
The model has 21 million parameters, 85 Mbyte, 384 layers
Was the accuracy of the generated depth maps investigated in comparison to the precise ones investigated? A good criterion could be the reconstruction quality of a video generated from it, using a similar approach as in JVET-AN0174.
Might be applied at the encoder end or at the decoder end to avoid transmitting depth maps.
Original camera parameters were used. It would however also be possible to estimate them.
Depth maps were generated in full video resolution (HD). Downsampled depth maps might be generated to reduce complexity.
Multiple frames (three past frames) were used to get temporal consistency.
It was agreed to include a branch in the AHG15 software repository to make the software available to interested experts.
The AHG should investigate the possibility to estimate depth maps and camera parameters for those gaming sequences where they are not available.
It was also commented that the usage of 3D models might be useful to improve the prediction in the gaming sequences which have other motion characteristics than camera captured content.
AHG16: Generative face video (1)
Contributions in this area were discussed during 1400–1420 on Thursday 9 Oct. 2025 (chaired by JRO).
See also section 6.2.14.