Back to Search Document details
39th Meeting: Daejeon, KR, March 2025 2025-03-27 09:43
JVET AHG report: Neural network-based video coding (AHG11)
Abstract
This document summarizes the activities of AHG11: Neural network-based video coding between the 37th meeting held in Geneva and 38th on-line meeting.
JVET-AL0011 JVET AHG report: Neural network-based video coding (AHG11) [E. Alshina, F. Galpin, S. Liu, A. Segall (co-chairs), J. Li, Y. Li, R.-L. Liao, M. Santamaria, T. Shao, M. Wien, P. Wu (vice chairs)]

Anchor Encoding

Anchor for the NN-based video coding activity made available though the Git repository used for the AHG activity:

https://vcgit.hhi.fraunhofer.de/jvet-ahg-nnvc/nnvc-ctc/-/blob/master/Anchor%20performance/NNVC-10-VTM_vs_NNVC-12.xlsm

also distributed by AhG14 in JVET-AK0014.

New training materials

BVI-AOM training set was used in EE1-2 (NN-Inter) and EE1-3 (retraining NN-filters) categories. Results of training are analyzed in relevant EE1 contributions. So far, some promising performance improvement is observed for HOP NN-filter. It is suggested to conduct detailed discussion, opinion exchange and brainstorming on training using additional materials.

Dataset access

Training data generation and md5 sums for BVI-AOM been discussed during AhG11/AhG14 teleconference in scope of EE1-3 tests.

Interaction with ECM

Some version of neural network based Intra included into ECM already. Interface for NNVC in-loop filter testing in ECM has been added and some test results reported in incoming AhG11 contribution.

In random access configuration NN-based in-loop filter provides roughly 1%, 2% and 5% BD-rate gain on top of ECM with only 1%, 2% and 8% CPU encoding run-time increment for VLOP, LOP and HOP versions of filter correspondently.

JVET-AL0230

AhG11/AhG12: Performance of the NNVC ILF in ECM

D. Rusanovskyy, K. Panusopone, S. Hong, L. Wang (Nokia)

JVET-AL0228

EE2-related: NNLF interface in ECM

T. Poirier, F. Galpin, G. Boisson (InterDigital)

EE Coordination

The AHG finalized, conducted, and discussed the EE on NN based video coding. A summary report for the EE is available at this meeting as:

JVET-AL0023

EE1: Summary report of exploration experiment on neural network-based video coding

E. Alshina, R. Chang, F. Galpin, Yue Li, Yun Li, M. Santamaria, J. Ström, Z. Xie (EE coordinators)

Teleconferences

The AHG conducted two joint teleconferences with AHG14 and EE1 during the interim period. The teleconferences were held on February 05 and March 05, 2025. In those teleconferences, the following topics were discussed:

  • New version of NN-based filter (LOP5) training report and LOP5 inference cross-check
  • NNVC12.0 software integration status and anchors performance,
  • EE1 description finalization
  • Training on BVI-AOM

Combination of two proposals on LOP JVET-AK0195 (luma input to chroma processing at the later stage of NN filter) on top of JVET-AK0150 (attention mechanism) was verified by training cross-check and adoption from last meeting confirmed during teleconference.

JVET-AL0043 

[AHG11] [AHG14] Teleconference on NNVC

E. Alshina, F. Galpin

Performance Evaluation

The performance and complexity of NN-based tools available in NNVC SW is summarized in the table below. All test data provided by AhG14. Encoding and decoding run time is very dependent on cluster used for simulation. Run time data in this table are all from InterDigital.

In NNVC-12, LOP filter performance improved by 0.6% (in random access test). There is no performance change for other configuration, but there is some speed-up (due to SADL optimization).

Test vs NNVC (configured as VTM)

Random Access cfg

kMAC/pxl

Param (Mprm)

Y

U

V

Enc

Dec

Total

Filter

Intra

SR

Total

Filter

Intra

SR

NN-Intra & LOP filter (2 tools)

NNVC-12.0 (LOP5)

-8.2%

-15.3%

-13.5%

1.2

35

21.4

16.6

4.8

0

1.5

0.247

1.3

0

NNVC-11.0 (LOP4)

-7.6%

-14.3%

-13.2%

1.2

36

21.6

16.8

4.8

0

1.5

0.21

1.3

0

NNVC-10.0 (LOP3)

-7.4%

-13.6%

-11.6%

1.2

33

21.7

16.9

4.8

0

1.5

0.21

1.3

0

NNVC-9.1(LOP3)

-7.3%

-13.1%

-11.3%

1.2

81

24.8

16.9

7.9

0

1.7

0.21

1.5

0

NNVC-8.0(LOP2)

-6.9%

-13.2%

-12.1%

1.2

73

25.0

17.1

7.9

0

1.6

0.05

1.5

0

NNVC-7.1(LOP2)

-6.9%

-13.2%

-12.1%

1.3

86

25.0

17.1

7.9

0

1.6

0.05

1.5

0

NNVC-9.1(LOP2CA)

-8.2%

-16.5%

-15.5%

2.5

69

25.5

17.6

7.9

0

1.6

0.05

1.5

0

NN-Intra & HOP filter (2 tools)

NNVC-10.0 (HOP5)

-14.2%

-19.5%

-19.9%

2.6

1135

471

466

4.8

0

2.7

1.4

1.3

0

NNVC-9.1(HOP4)

-14.1%

-19.2%

-19.6%

2.9

1447

484

476

7.9

0

3.0

1.44

1.5

0

NNVC-8.0(HOP3)

-13.7%

-13.9%

-14.5%

2.5

1092

474

466

7.9

0

2.9

1.40

1.5

0

NNVC-7.0(HOP2)

-13.6%

-12.5%

-14.2%

4.1

2071

485

477

7.9

0

3.0

1.50

1.5

0

NN-Intra & VLOP filter (2 tools)

NNVC-11.0 (VLOP3)

-5.8%

-6.6%

-5.7%

1.2

7

9.9

5.10

4.8

0

1.3

0.07

1.3

0

NNVC-10.0 (VLOP2)

-5.6%

-7.6%

-6.4%

1.1

15

10

5.16

4.8

0

1.3

0.06

1.3

0

NNVC-9.1(VLOP)

-5.3%

-5.4%

-5.2%

1.2

40

13

5.12

7.9

0

1.5

0.02

1.5

0

NN-Intra & LOP filter content adaptive (2 tools)

NNVC-10.0 aLOP3

-8.2%

-16.7%

-15.5%

2.2

34

21.8

16.9

4.8

0

1.5

0.21

1.3

0

NN-Intra & LOP filter & adaptive resolution coding (3 tools)

NNVC-11.0 NNSR

-8.5%

-12.2%

-10.9%

26.3

16.8

4.8

4.7

1.4

0.05

1.3

0.05

NNVC-8.0 RPR

-7.5%

-10.9%

-9.7%

25.0

17.1

7.9

0

1.6

0.05

1.5

0

NNVC-8.0 NNSR

-7.8%

-11.8%

-10.5%

45.3

17.1

7.9

20.3

1.8

0.21

1.5

0.1

More details and analysis for tools and tools combination is expected in AhG14 report.

Architectural changes

Following proposals from 36th and 37th meetings several companies propose framework for E2E AI coded reference picture insertion. Promising gain reported for this approach (4% and 1% in all intra and random-access configuration correspondingly).

In order to support this feature bit-exact reconstruction of E2E AI coded picture is required. One contribution discusses fundamental aspects of bit-exact reproducibility and potential implementation using different platforms (GPU, NPU, CPU, ASIC).

NNVC algorithms description

One contribution submitted to this meeting resolves mismatch between LOP5 implementation in NNVC SW and description.

Input contributions

There are 33 input contributions related to the AHG mandates. The list of input contributions is provided below.

Reporting (5)

JVET-AL0023

EE1: Summary report of exploration experiment on neural network-based video coding

E. Alshina, R. Chang, F. Galpin, Yue Li, Yun Li, M. Santamaria, J. Ström, Z. Xie (EE coordinators)

JVET-AL0043

[AHG11] [AHG14] Teleconference on NNVC

E. Alshina, F. Galpin

JVET-AL0230

AhG11/AhG12: Performance of the NNVC ILF in ECM

D. Rusanovskyy, K. Panusopone, S. Hong, L. Wang (Nokia)

JVET-AL0228

EE2-related: NNLF interface in ECM

T. Poirier, F. Galpin, G. Boisson (InterDigital)

JVET-AL0291

EE1-related: Recommendations for resolving mismatches between LOP5 description and implementation

Nam Le, Francesco Cricri (Nokia)

Architectural change and implementation aspects (4)

JVET-AL0080

AHG11: Bit-exact reconstruction for NN video tools

L. Kerofsky, Y. Li, M. Karczewicz (Qualcomm)

JVET-AL0196

AHG 11: Neural Network Coded Reference Frame for Intra Coding with Residual Coding and Intra Blocks

F. Brand, T. Solovyev, E. Alshina (Huawei)

JVET-AL0203

[AHG11] A Hybrid Framework Integrating End-to-End Learned Image Codec with Conventional Codec

N. Zou, A. Hallapuro, F. Cricri, H. Zhang, A. B. Koyuncu, J. Ahonen, M. M. Hannuksela (Nokia)

JVET-AL0243

[AHG11] Multilayer framework for supporting a hybrid codec using End-to-End Learned Image Codec and Conventional Video Codec

F. Urban, Y. Chen, F. Galpin, E. François (InterDigital)

EE1 contributions (11)

JVET-AL0084

EE1-1.2: LOP5 improvement with parallel 1x3/3x1 Backbone

T. Shao, P. Yin, S. McCarthy (Dolby), J. N. Shingala, A. Shyam, A. Suneja, S. P. Badya (Ittiam)

JVET-AL0085

EE1-3.1: NNVC-LOP5 retraining with additional BVI-AOM dataset

A. Suneja, J. N. Shingala, A. Shyam, S. P. Badya (Ittiam), T. Shao, P. Yin, S. McCarthy (Dolby)

JVET-AL0104

EE1-2.1: Lightweight Multiscale Reference Frame Generation for VVC Inter Coding

P. Li, C. Jung, Q. Qin (Xidian Univ.)

JVET-AL0105

EE1-2.2: RA/LDB Unified Reference Frame Synthesis for VVC Inter Coding

Q. Qin, C. Jung (Xidian Univ.)

JVET-AL0144

EE1-3.1: Retraining LOP4 and LOP5 using extended dataset from BVI-AOM

D. Liu, J. Ström, M. Damghanian, P. Wennersten (Ericsson)

JVET-AL0145

EE1-1.4: Conditional loop-filter

M. Santamaria, F. Cricri (Nokia)

JVET-AL0164

EE1-1.1: Multiscale blocks in LOP5 and VLOP3 filters

R. Yang, M. Santamaria, F. Cricri, H. Zhang, J.Lainema, M. M. Hannuksela (Nokia)

JVET-AL0169

EE1-1.3 Dimension-wise decomposed representation of multiplier for content-adaptive loop filtering

Z. Xu, J. Konieczny, A. Filippov, C. Hollmann, V. Rufitskiy, T. Dong (TCL)

JVET-AL0187

EE1-3.1: NNVC-LOP4 and LOP5 retraining with additional BVI-AOM dataset

T. Dumas, A. Monier, F. Galpin (InterDigital)

JVET-AL0189

EE1-3.2 – NNVC-HOP5 retraining adding BVI-AOM

F. Galpin (InterDigital)

JVET-AL0190


EE1-3.3 – NNVC-VLOP3 retraining adding BVI-AOM

F. Galpin, Z. Ameur (InterDigital), R. Chang, L. Wang, X. Xu, S. Liu (Tencent)

New NNVC tools in AhG11 or EE1 related contributions (7)

JVET-AL0166

EE1 related: Improved VLOP with Attention

Y. Li, L. Kerofsky, M. Karczewicz (Qualcomm

JVET-AL0167

EE1 related: Further simplification of VLOP with Attention

Y. Li, L. Kerofsky, M. Karczewicz (Qualcomm)

JVET-AL0184

EE1-related: Deep Reference Frame Generation for Inter Prediction Enhancement

D. Ding, X. Chen, Z. Chen (Wuhan Univ.)

JVET-AL0121

AHG11: Sample-based adaptive blending weight selection for LOP

H. Kwon, H. Ko (HYU)

JVET-AL0136

AHG11: Over-Parameterized LOP In-Loop Filter

J. Han, C. Jung, Q. Qin (Xidian Univ.)

JVET-AL0137

AHG11: Backbone Block Enhancement of LOP In-Loop Filter with Over-Parameterized Training and Variable Channels

J. Han, C. Jung, Q. Qin (Xidian Univ.)

JVET-AL0250

AHG11: Cross-component enhanced NNSR

T. Yang, W.-X. He, Y.-Q. Zhu, J.-D. Ye, X.-T. Xie, J.-S. Gong, Q. Liu (HUST), Z.-Y. Lv (vivo)

Cross-checks (6 some not yet uploaded at the time report was prepared)

JVET-AL0173

Crosscheck of JVET-AL0164 (EE1-1.1: Multiscale blocks in LOP5 and VLOP3 filters)


Y. Li (Qualcomm)

JVET-AL0185

Crosscheck of JVET-AL0169 (EE1-1.3 Dimension-wise decomposed representation of multiplier for content-adaptive loop filtering)

M. Santamaria (Nokia)

JVET-AL0246

Crosscheck of JVET-AL0084 (EE1-1.2: LOP5 improvement with parallel 1x3/3x1 Backbone)

Y. Li (Bytedance)

JVET-AL0284

Crosscheck of JVET-AL0145 (EE1-1.4: Conditional loop-filter)

J. Strom (Ericsson)

JVET-AL0297

Crosscheck of JVET-AL0104 (EE1-2.1: Lightweight Multiscale Reference Frame Generation for VVC Inter Coding)

L. Murn (Nokia)

JVET-AL0298

Crosscheck of JVET-AL0105 (EE1-2.2: RA/LDB Unified Reference Frame Synthesis for VVC Inter Coding)

L. Murn (Nokia)

Recommendations

The AHG recommends:

  • Review all input contributions.
  • Continue investigating neural network-based video coding tools, including coding performance and complexity.
  • Continue collecting training materials for neural network-based video coding tool development and investigate training stability.
  • Conduct detailed discussion on training, exchange opinions and observations among interested parties (potential BoG)
  • Add study of bit-exact reproducibility to the NNVC complexity assessment

Results indicate that adding BVI-AOM to training is not sufficient for optimizing inter prediction part of EE.

Decisions
Results indicate that adding BVI-AOM to training is not sufficient for optimizing inter prediction part of EE.
Citation