Back to Search Document details
25th Meeting: by teleconference, October 2021 2021-10-08 16:37
[AHG11 & AHG6] DOVC: Deep Omnidirectional Video Compression
Abstract
This contribution presents an end-to-end deep omnidirectional video compression (DOVC) framework with convolutional neural networks (CNNs). The proposed DOVC framework includes a bidirectional offset prediction module based on deformable convolution to generate effective motion information and a bidirectional motion compensation (MC) network composed of ordinary convolution and residual blocks for motion compensated prediction. The proposed DOVC framework also reuses the motion information generated by the bidirectional offset prediction module to enhance residual in the quality enhancement (QE) module. Regarding omnidirectional video, the modules in the proposed DOVC framework are jointly optimized by weighted rate distortion to minimize the geometric distortion of the sphere-to-plane projection. Experimental results show that DOVC achieves average 21% reduction in BD-BR and average 0.6217dB gain in BD-WS_PSNR over HM-16.16 (360Lib-5.0) under LDP configuration for encoding omnidirectional videos, and the average encoding time of DOVC is only 0.0234 times that of HM-16.16 (360Lib-5.0) and 0.0064 times that of VTM-11.0 (360Lib-12.0).
JVET-X0043 [AHG11 & AHG6] DOVC: Deep Omnidirectional Video Compression [Q. Qin, C. Jung (Xidian Univ.), Z. Dan, M. Li (OPPO)]

This contribution presents an end-to-end deep omnidirectional video compression (DOVC) framework with convolutional neural networks (CNNs). The proposed DOVC framework includes a bidirectional offset prediction module based on deformable convolution to generate effective motion information and a bidirectional motion compensation (MC) network composed of ordinary convolution and residual blocks for motion compensated prediction. The proposed DOVC framework also reuses the motion information generated by the bidirectional offset prediction module to enhance residual in the quality enhancement (QE) module. Regarding omnidirectional video, the modules in the proposed DOVC framework are jointly optimized by weighted rate distortion to minimize the geometric distortion of the sphere-to-plane projection. Experimental results show that DOVC achieves average 21% reduction in BD-BR and average 0.6217dB gain in BD-WS_PSNR over HM-16.16 (360Lib-5.0) under LDP configuration for encoding omnidirectional videos, and the average encoding time of DOVC is only 0.0234 times that of HM-16.16 (360Lib-5.0) and 0.0064 times that of VTM-11.0 (360Lib-12.0).

ERP and CMP were used as projection formats for pre/post processing.

Roughly 20M parameters were used.

I frames (distance 8) were coded by “BPG”. Why not using end-to-end intra compression (or VVC intra) which should have better performance? Proponents did not think about it. How is the bit rate for the intra frame decided?

Performance was worse than VVC (measured by WS-PSNR): 12% bit rate increase in CMP, 38% in ERP.

Coding currently in RGB 4:4:4, while VVC and HEVC in YUV 4:2:0 (proponents expect better performance when also coding in YUV).

Encoding is reported to be faster than VVC, while decoding is slower (computed on a GPU).

This could be applied to common video (by adjusting the loss function).

Further study would be welcome.

References:
JVET-H1030
PATENTS:
CN118104237A 1.00 2024-05-28 US10887589B2 0.91 2021-01-05 US11153606B2 0.82 2021-10-19 US20180167634A1 0.73 2018-06-14 US11689705B2 0.64 2023-06-27 US20240291981A1 0.55 2024-08-29 EP3656126B1 0.46 2024-09-04 US20200275129A1 0.37 2020-08-27
Decisions
Further study would be welcome.
Citation