Back to Search Document details
39th Meeting: Daejeon, KR, June 2025 2025-08-03 15:52
AHG11: Deep Reference Frame Generation for Inter Prediction Enhancement with Structural Re-parameterization
Abstract
This proposal builds upon the lightweight deep reference frame generation (LDRF) method of JVET-AG0122 by introducing a structural re-parameterized interpolation diverse branch block (Inter-DBB), which further enhances inter prediction efficiency of the original LDRF in NNVC. During training, Inter-DBB employs multi-branch convolutions to capture diverse spatial-channel correlations, while at inference these branches are converted into a single 3×3 convolution via structural re-parameterization, preserving model structure and computational complexity. It provides average BD-rate gains of about −1.67% for luma (Y) and −0.71%/−0.77% for chroma (U/V) components compared to the NNVC-12.0 anchor for the RA configuration. Relative to the baseline DRF model in JVET-AG0122, the proposed method achieves −0.13 % BD-rate for Y and +0.20 %/+0.06 % for U/V, respectively.
JVET-AM0177 AHG11: Deep Reference Frame Generation for Inter Prediction Enhancement with Structural Re-parameterization [W. Zhang, C. Gui, N. Fu, X. Chen, W. Ma, Z. Chen (Wuhan Univ.)]

This proposal builds upon the lightweight deep reference frame generation (LDRF) method of JVET-AG0122 by introducing a structural re-parameterized interpolation diverse branch block (Inter-DBB), which further enhances inter prediction efficiency of the original LDRF in NNVC. During training, Inter-DBB employs multi-branch convolutions to capture diverse spatial-channel correlations, while at inference these branches are converted into a single 3×3 convolution via structural re-parameterization, preserving model structure and computational complexity. It provides average BD-rate gains of about −1.67% for luma (Y) and −0.71%/−0.77% for chroma (U/V) components compared to the NNVC-12.0 anchor for the RA configuration. Relative to the baseline DRF model in JVET-AG0122, the proposed method achieves −0.13 % BD-rate for Y and +0.20 %/+0.06 % for U/V, respectively.

图示

描述已自动生成

The framework of the proposed DRF method

图示

描述已自动生成

The architecture of the DRF networks

2900 K parameters, 69 kMAC/pixel

Slightly less gain than method from EE1-3.2, but also significantly lower complexity.

It was agreed to investigate this in an EE, including integer conversion and training crosscheck.

Decisions
It was agreed to investigate this in an EE, including integer conversion and training crosscheck.
Citation