Back to Search Document details
32nd Meeting: Hannover, DE, October 2023 2023-10-05 16:14
A Study on Decoder Interoperability of Generative Face Video Compression
Abstract
At the 29th JVET meeting, JVET-AC0088 proposed a generative face video SEI message to allow VVC-coded pictures to be used as base pictures to drive a deep generative network to generate additional images, therefore enabling ultra low rate face video compression. At the 30th JVET meeting, JVET-AD0051 generalized the generative face video SEI message to include syntax elements that correspond to a variety of facial representations, including 2D keypoints, 2D landmarks, 3D keypoints, temporal motion features, facial semantics and other formats. At the 31st JVET meeting, JVET-AE0080, JVET-AE0083 and JVET-AE0280 were proposed to support more common syntax design of generative face video SEI message and define the interface between the decoder (which decodes the base pictures and parses the SEI message) and the generative neural network.
JVET-AF0048 A Study on Decoder Interoperability of Generative Face Video Compression [B. Chen, S. Yin, J. Chen, Y. Ye (Alibaba), S. Wang (CityU HK)]

At the 29th JVET meeting, JVET-AC0088 proposed a generative face video SEI message to allow VVC-coded pictures to be used as base pictures to drive a deep generative network to generate additional images, therefore enabling ultra low rate face video compression. At the 30th JVET meeting, JVET-AD0051 generalized the generative face video SEI message to include syntax elements that correspond to a variety of facial representations, including 2D keypoints, 2D landmarks, 3D keypoints, temporal motion features, facial semantics and other formats. At the 31st JVET meeting, JVET-AE0080, JVET-AE0083 and JVET-AE0280 were proposed to support more common syntax design of generative face video SEI message and define the interface between the decoder (which decodes the base pictures and parses the SEI message) and the generative neural network.

One question that was raised during the discussion of these contributions was the interoperability between different facial representations and their associated decoder networks. This contribution attempts to shed some lights on how the interoperability problem might be solved. Specifically, the contribution further analyses the decoder design of generative face video SEI message, and discusses how to support the decoder interoperability and parameter translatability in the two following schemes:

  1. Scheme 1: Recognizing that the end-to-end decoder consists of two independent modules, optical flow estimation module and generator module, the original end-to-end decoder was retrained as independent modules, where the generator module was fixed and the optical flow estimation module was trained to adapt to different types of parameters.
  2. Scheme 2: Establishing a generalized face parameter translator so that different facial representations can be converted into the required type that can be accepted by a fixed decoder.

Analysis and synthesis don’t have to be jointly optimized, a translator aligns the information needed.

Quality somewhat suffers relative to the jointly optimized case, but still acceptable.

The SEI message would represent 3D keypoints.

The network parameters of the translator could also be sent (1.3 million parameters).

The fixed synthesis model would not be mandatory.

It was commented that rather than supporting various approaches, it might be more appropriate selecting one of them.

Decisions
It was commented that rather than supporting various approaches, it might be more appropriate selecting one of them.
Citation