JVET-AE0080 AHG9: Generative Face Video SEI message [S. McCarthy, P. Yin, G.-M. Su, A. K. Choudhury, W. Husak (Dolby)]
This contribution proposes a generative face video (GFV) SEI message for inclusion in the next version of VSEI for use with VVC. The proposed GFV SEI message conveys information for facial features extracted from source video prior to encoding that can be applied to a previously decoded reference picture to reconstruct a face picture at the current time. The proposed GFV SEI message also specifies a neural network, for example, a generative adversarial network (GAN), that could be used to generate face pictures based on the face feature information conveyed in the GFV SEI message. As noted in previous related JVET contributions JVET-AC0088 and JVET-AD0051, the GFV SEI message is intended to enable face picture reconstruction processes to use VVC-coded pictures as reference pictures. It is also noted that extracting face features on the encoder-side can reduce receiver-side complexity. The present contribution proposes additions and refinements to previously proposed syntax and semantics to address issues noted by experts during previous JVET meetings. The proposed GFV SEI message is intended for ultra low bitrate face video compression applications such as video conferencing, live entertainment, and face animation.
When a receiver ignores the SEI message, the base image would be shown.
It was asked whether this would not rather be something to standardize as a specific standard rather than SEI message. The benefit of the technology is well recognized, very low bit rate for video conferencing.
One expert mentions that this technology has a very limited scope. The proponents say this is intentional.
Is it an advantage to couple this to VVC/HEVC? If somebody is interested only in this functionality, they might not want to implement a full VVC decoder.
Would it not be better to define a profile if the intent is limited scope?
The relationship with VPC is also pointed out, but this is more generic, not only using VVC as base codec. According to last meeting’s report, WG 7 also seems to have some plans about Avatar coding. Unclear if that is based on neural networks or old MPEG-4 technology. Joint meeting?
Usage of other legacy codecs is also important in the opinion of another expert.
Complexity issue? It was pointed out that real-time implementations on smartphone exist, which also fulfills real-time requirements (see https://arxiv.org/abs/2012.00328)
One expert suggested that an exploration experiment might be started.
Interoperability between encoder and decoder could be important in this context – can this be provided by an SEI message? This role could be taken over by application standards who could define a specific encoder/decoder pair.