JVET-AF0234 AHG9: Common text for proposed generative face video SEI message [B. Chen, J. Chen, Y. Ye (Alibaba), S. Wang (CityU), S. McCarthy, P. Yin, G.-M. Su, A. K. Choudhury, W. Husak (Dolby)] [late]
This contribution updates the common text of generative face video (GFV) SEI message proposed in JVET-AE0280 by introducing the parameter translator proposed in JVET-AF0048 which attempts to solve the interoperability issue between different facial representations and their associated decoder networks. The proposed GFV SEI message is intended for ultra-low bitrate face video compression applications such as video conferencing, live entertainment, and face animation.
The SEI message contains
- Keypoints (2D or 3D)
- Matrices of different types (e.g. affine, cov, mouth,…)
- Optionally a translator (using NNR)
Depending on the algorithm at the encoder, not all of these parameters are present in the SEI.
Depending on the parameters received, and the synthesis algorithm available, a certain translator may be necessary.
Translator and synthesizer are NNs.
If a decoder does not have a suitable translator and synthesizer available for the parameters found in the SEI, it would give up.
Basically by this concept, many different formats would be supported.
It was concluded that establishing a new AHG is the best way of studying generative face, providing software, coordinating experiments, investigating which parameters would be needed, how they can be efficiently coded, interoperability aspects, quality, stability, etc.
The AHG should also be mandated to develop a document summarizing the potential approaches.
Chairs were suggested to be Y. Ye, H.-B. Teo, Z. Lyu, and S. McCarthy.
AHG9: Source picture timing information SEI message aspects (6)
Contributions in this area were discussed at 1830–2200 on Monday 16 Oct. 2023 (chaired by JRO).