Back to Search Document details
39th Meeting: Daejeon, KR, June 2025 2025-07-03 11:02
AHG16: Lightweight Multi-resolution CFTE Model and Color Calibration Post-processing Algorithm for Generative Face Video Compression
Abstract
This contribution proposes replacing the current single-resolution lightweight CFTE model in AHG16 GFVC software with a multi-resolution lightweight CFTE model through architectural (i.e., depthwise separable convolution) and operational (i.e., pruning) optimizations. Additionally, we propose incorporating the color calibration method into the AHG16 GFVC software to address the occasional color shift issues observed in some generated video frames. Experimental results illustrate that our proposed lightweight scheme can reduce the complexity and number of parameters whilst maintaining the coding performance, while color calibration noticeably improves the subjective quality of the generated video frames, as summarized below,
JVET-AM0058 AHG16: Lightweight Multi-resolution CFTE Model and Colour Calibration Post-processing Algorithm for Generative Face Video Compression [Z. Zhang, S. Yin, S. Wang (CityUHK), B. Chen, R.-L. Liao, J. Chen, Y. Ye (Alibaba)]

This contribution proposes replacing the current single-resolution lightweight CFTE model in AHG16 GFVC software with a multi-resolution lightweight CFTE model through architectural (i.e., depthwise separable convolution) and operational (i.e., pruning) optimizations. Additionally, the contribution proposes incorporating the colour calibration method into the AHG16 GFVC software to address the occasional colour shift issues observed in some generated video frames. Experimental results illustrate that our proposed lightweight scheme can reduce the complexity and number of parameters whilst maintaining the coding performance, while colour calibration noticeably improves the subjective quality of the generated video frames, as summarized below,

  • Compared to the current multi-resolution CFTE model, our proposed lightweight approach reduces the number of parameters by 89.7%, and kMACs per pixel by 83.1% and 79.7% at resolutions of 256×256 and 512×512, respectively. Besides, at 256×256 resolution, the overall bitrate savings of the proposed method in terms of Rate-DISTS and Rate-LPIPS are {2.56%, 2.57%}. At 512×512 resolution, the overall bitrate savings of the proposed method in terms of Rate-DISTS and Rate-LPIPS are {5.57%, 9.03%}.
  • Compared to GFVC models without colour calibration, those using it can produce more natural reconstruction results that are closer to the corresponding original video.

kMAC/p reduced to 160 for 256x256, and 80 for 512x512 resolutions.Performance still worse than for the existing light-weight single resolution model.

More sensitive to colour shift than the old method (as other higher-complexity multi-resolution model), which can be reduced by colour calibration as post processing.

Further study was recommended to improve the multi-resolution model such that it becomes better than single-resolution also for the low complexity case (as was previously proven for the higher-complexity networks), and to show that the colour calibration post processing also provides benefit for the higher-complexity multi-resolution networks.

AHG17: CfE preparation (16)

Contributions in this area were discussed during 1610–1810 on Monday 30 June 2025 (chaired by JRO). & Tue 0910-1125 & Wedneday 1450-1555

Some aspects discussed under section 4.4 could also be relevant here.

Decisions
Further study was recommended to improve the multi-resolution model such that it becomes better than single-resolution also for the low complexity case (as was previously proven for the higher-complexity networks), and to show that the colour calibration post processing also provides benefit for the higher-complexity multi-resolution networks.
Citation