Back to Search Document details
40th Meeting: Geneva, CH, October 2025 2025-10-06 22:22
AHG17: Description of NNVC-based CfE response
Abstract
This contribution summarizes and analyzes results obtained in the Call for Evidence (CfE) responses which were built as combinations of tools available in the recent NNVC tool-box. Responses are provided in the constrained encoder run time category. The goal of those responses is to demonstrate compression performance vs encoder complexity trade-off provided by learnable technologies. This response also serves as sanity check for technologies trained off-line, since at CfE performance is checked under test conditions substantially different from NNVC training and testing conditions. For target run time ‘0.2’ (which is twice faster than HM at CfE rates) performance of ‘VTM default’ was achieved, while at target run time ‘1’ 11% gain over VTM was demonstrated. Normative changes to the anchor are limited to just tool tools for all (except just one) of operation points.
JVET-AN0212 AHG17: Description of NNVC-based CfE response [D. Kim, S.-C. Lim (ETRI), T. Solovyev, J. Sauer, J. Pardo, P. Jia, A. Karabutov, E. Alshina (Huawei), H. Kwon, H. Ko (HYU), F. Galpin, T. Dumas, E. François (InterDigital), Y. Li, M. Karczewicz (Qualcomm), Z. Xiang, R. Chernyak, S. Liu (Tencent)]

This contribution summarizes and analyzes results obtained in the Call for Evidence (CfE) responses which were built as combinations of tools available in the recent NNVC tool-box. Responses are provided in the constrained encoder run time category. The goal of those responses is to demonstrate compression performance vs encoder complexity trade-off provided by learnable technologies. This response also serves as sanity check for technologies trained off-line, since at CfE performance is checked under test conditions substantially different from NNVC training and testing conditions. For target run time ‘×0.2’ (which is twice faster than HM at CfE rates) performance of ‘VTM default’ was achieved, while at target run time ‘×1’ 11% gain over VTM was demonstrated. Normative changes to the anchor are limited to just tool tools for all (except just one) of operation points.

Development principles of those CfE responses are as follows:

  • Only tools available in NNVC-14.1 were used,
  • SADL implementation for NN-based tools, int 16 quantized models,
  • No multi-pass coding (content adaptive NN-based filters are not used),
  • No pre-/post-processing other than MCTF.

Several encoder-only techniques such as time-distortion optimization (TDO) and filter restriction depending on hierarchical depth, were developed during this CfE response preparation which allow drastically reduce decoder run-time of NNVC tools with minimal performance effect.

Two types of NNVC bitstreams were generated as responses in encoder-constrained category:

  • ‘×5’ and ‘×1’ not including RPR, actual encoder run time is ‘×5.7’ and ‘×1.1’ respectively.
  • ‘×0.2’, ‘×0.5’, ‘×1’, ‘×2’ and ‘×5’ with RPR enabled, actual encoder run time is ‘×0.23’, ‘×0.51’, ‘×1.1’, ‘×2.0’ and ‘×5.4’ respectively.

Uses NN-based LF (LOP/HOP), NN Intra Pred, DRF from NNVC14

As encoder-only tools, TDO (JVET-AN0054) and NNLF usage restriction (JVET-AN0215) are used.

The graph above is average over all sequences from all categories (regardless of LB or RA)

What are expectations on visual quality? Only partial visual inspection was done, e.g. to make sure that no visual artifacts appear. RPR has good benefit on visual quality.

Why is no lowest run time point provided for the non-RPR version? Mainly because also RPR contributes to encoder run time reduction.

Which percentage of pixels is actually filtered? Can roughly be deduced from the columns TDO and max hier depth NNLF in the subsequent table:

Different constraints are applied to RPR and non-RPR cases. For LB, gain of RPR is smaller as A classes are not included where RPR gives most gain.

Complexity:

Decisions
Complexity:
Citation