JVET-AL0147 AHG16: Refined parameter translator of generative face video coding [S. Yin, S. Wang, Z. Zhang (CityUHK), B. Chen, Y. Ye, R.-L. Liao, J. Chen (Alibaba)]
In JVET-AG0048, we investigated the issue of interoperability among different GFVC systems and solved the interoperablity problem by adding face parameter translators to the AHG16 software. With these translators, it was shown that three different types of face parameters from different GFVC algorithms can be effectively adapted to one another while retaining promising coding performance. During 34th to 37th JVET meetings, several updates have been made to the GFVC models in the AHG16 software, including more algorithms, more diverse training data, multi-resolution models etc., which resulted in the existing parameter translators in the latest AHG16 software becoming outdated (e.g. inability to handle 512x512 resolution). In this contribution, we propose to update face parameter translators in the AHG16 software with a refined difference translation scheme. Experimental results show that refined translators can achieve translations between 4 multi-resolution models on both 256×256 and 512×512 resolutions while retaining promising coding performance.
The translation in feature difference domain provides results which are visually closer to the original than with previous translators between different GFVC algorithms. Also objective metrics are clearly improved.
Decision(SW): Add JVET-AL0147 to AHG branch of software, replace existing translation.