JVET-AJ0016 JVET AHG report: Generative face video compression (AHG16) [Y. Ye (chair), H.-B. Teo, Z. Lyu, S. McCarthy, S. Wang (vice chairs)]
Regarding the mandate on developing and maintaining the GFVC software, the AHG16 GFVC software tool and accompanying usage instructions and exemplar configurations for experimentation are maintained in the GIT repository at https://vcgit.hhi.fraunhofer.de/jvet-ahg-gfvc. During this AHG period, the AHG16 software was updated by replacing the existing CFTE and FOMM models with the retrained CFTE and FOMM models as provided in JVET-AI0048. Further, master scripts are included to invoke the AHG16 software and the VTM SEI software to encode and decode the key pictures, generate and parse the GFV SEI messages, and generate the final output video. Implementation of the GFV SEI message following the syntax defined in the VSEI TuC draft in JVET-AI2032 is available in the TuC branch of VTM as merge requests, at https://vcgit.hhi.fraunhofer.de/jvet-tuc/VVCSoftware_VTM/-/merge_requests/25.
Regarding the mandates to identify additional test material and study GFVC performance on such test material, JVET-AJ0052 proposes multi-resolution FOMM, FV2V, CFTE, DAC models that can be used to code both 256x256 and 512x512 content, and JVET-AJ0209 proposes to add 512x512 test sequences to the GFVC test conditions, and provides performance results using SEI message to code such 512x512 content following GFVC test configurations.
Regarding coordination with AHG9 to develop the GFV SEI message, JVET-AJ0207 illustrates the SEI overhead reduction effort, and proposes to move GFV and GFVE SEI messages from TuC to VSEI v4. JVET-AJ0135 proposes to add pupil position information to the GFVE SEI message and provides showcase results to illustrate the effectiveness of adding such information. Finally, JVET-AJ0051, JVET-AJ0069, JVET-AJ0101, JVET-AJ0108, JVET-AJ0111 and JVET-AJ0132 are all AHG9 contributions related to various high level syntax aspects of the GFV SEI message.
Related contributions
The following list of input contributions to this meeting were identified as being related to the activities of AHG16:
- JVET-AJ0051, AHG9: On the GFV SEI message [Y.-K. Wang (Bytedance), Y. Li, K. Yang, Y. Xu (SJTU), J. Chen, B. Chen, Y. Ye (Alibaba)]
- JVET-AJ0052, AHG 16: Multi-resolution models for generative face video compression [S. Yin, S. Wang (CityU), B. Chen, Y. Ye (Alibaba), G.Konuko, G. Valenzise (CentraleSupelec)]
- JVET-AJ0069, AHG9: On specifying the output order of GFV-generated pictures [L. Chen, O. Chubach, Y.-W. Huang, S. Lei (MediaTek)]
- JVET-AJ0101 AHG9: Miscellaneous modifications for SEI messages in the TuC for future extensions of VSEI [C. Kim, H. Tan, J. Lee, J. Nam, J. Lim, S. Kim (LGE)]
- JVET-AJ0108, AHG9: On signalling of instance count in generative face video SEI message [H. Tan, J. Lee, J. Nam, C. Kim, J. Lim, S. Kim (LGE)]
- JVET-AJ0111, AHG9: On the presence of translator NN filter in GFV SEI message [H. Tan, J. Nam, J. Lee, C. Kim, J. Lim, S. Kim (LGE)]
- JVET-AJ0132, AHG9: Usage of NNPFC SEI message to define the generator NN of the GFV SEI message [M. M. Hannuksela, F. Cricri (Nokia)]
- JVET-AJ0135, AHG9/AHG16: Refined Methodology for Pupil Position SEI Message for Generative Face Video [F. Ma, A. Trioux, F. Yang (Xidian Univ.), B. Li, F. Xing, Z. Wang (Hisense)]
- JVET-AJ0207, AHG9/AHG16: Performance results and suggestions on generative face video SEI message [B. Chen, Y. Ye, J. Chen, R.-L. Liao (Alibaba), S. Yin, Z. Zhang, S. Wang (CityU), K. Yang, Y. Li, Y. Xu (SJTU), Y.-K. Wang (Bytedance), S. Gehlot, G.-M. Su, P. Yin, G. J. Sullivan, S. McCarthy (Dolby), H.-B. Teo (Panasonic)]
- JVET-AJ0209, AHG9/AHG16: Performance results of generative face video SEI message on high-resolution sequences [B. Chen, Y. Ye, R.-L. Liao, J. Chen (Alibaba), S. Yin, S. Wang (CityU)]
Recommendations
The AHG recommended to:
- Review related contributions;
- To continue AHG16 to study GFVC-related topics.
Project development (37)
AHG1: Development, deployment and advertisement of standards (6)
Contributions in this area were discussed at 1505–1515 on Thursday 7 Nov. 2024 (chaired by JRO), unless noted otherwise.