Back to Search Document details
31st Meeting: Geneva, CH, July 2023 2023-07-11 08:11
JVET AHG report: Neural network-based video coding (AHG11)
Abstract
This document summarizes the activities of AHG11: Neural network-based video coding between the 30th meeting (21 – 28 April 2023) held in Antalya, TR and the 31st meeting (11 – 19 July 2023) held in Geneva, CH.
JVET-AE0011 JVET AHG report: Neural network-based video coding (AHG11) [E. Alshina, S. Liu, A. Segall (co chairs), F. Galpin, J. Li, R.-L. Liao, D. Rusanovskyy, T. Shao, M. Wien, P. Wu (vice-chairs)]

Activities

The AHG used the main JVET reflector, jvet@lists.rwth-aachen.de, for email. Forty-one emails were exchanged on the reflector related to the AHG mandates.

Common Test Conditions

Document

The AHG released revised common test conditions as decided at the 30th meeting. The final version was uploaded as document JVET-AD2016 on May 12, 2023.

Anchor Encoding

Anchors for the NN-based video coding activity made available on the Git repository used for the AHG activity: https://vcgit.hhi.fraunhofer.de/jvet-ahg-nnvc/nnvc-ctc/-/tree/master.

EE Coordination

The AHG finalized, conducted and discussed the EE on NN based video coding. The final version of the EE description was uploaded to the document repository on May 15, 2023.

A summary report for the EE is available at this meeting as:

JVET-AE0023

EE1: Summary report of exploration experiment on neural network-based video coding

E. Alshina, F. Galpin, Y. Li, D. Rusanovskyy, M. Santamaria, J. Ström, L. Wang, Z. Xie (EE coordinators)

Teleconferences

The AHG conducted two teleconferences during the interim period. The teleconferences were held on June 14, 2023, and June 28, 2023 and attended by approximately 20 participants.

A summary report for the teleconferences is available at this meeting as:

JVET-AE0042

AhG14 & AHG11: Report on AhG teleconference on high operation point (HOP) unified filter training

E. Alshina, F. Galpin

Workshop on Neural Network-Based Technologies

The AHG contributed to the “Third AG4 Workshop on JPEG and MPEG Emerging Activities” that was conducted by the JPEG and MPEG Collaboration Advisory Group, or ISO/IEC JTC1/SC29/AG4.

The workshop took place on June 1st, 2023, and its agenda was:

Speaker

Topic

Elena Alshina

AI-based Video Coding

Werner Bailer

Neural Network Compression

Marc Antonini

Coding for DNA Storage

Tim Bruylants

Event-based Vision

Arianne Hids

Scene-based Interchange and Support of Immersive Displays

The AHG contributed the presentation titled “AI-based Video Coding”. Slides for the presentation are available in JVET-AE0237.

Performance Evaluation

The performance of the NNVC-5.0 anchor compared to NNVC-4.0 is reported below.

Random access Main 10

BD-rate Over NNVC-4.0

Y-PSNR

U-PSNR

V-PSNR

Y-MSIM

U-MSIM

V-MSIM

EncT

DecT CPU

PSNR Overlap

Class A1

-7.23%

-4.86%

-6.00%

-8.72%

-4.05%

-5.06%

134%

7667%

96%

Class A2

-6.56%

-5.78%

-5.09%

-6.53%

-4.14%

-2.90%

130%

7183%

97%

Class B

-6.29%

-8.19%

-7.53%

-5.97%

-6.06%

-5.18%

132%

7834%

98%

Class C

-6.62%

-10.53%

-9.55%

-6.26%

-7.91%

-5.58%

132%

7422%

98%

Class E

-

-

-

-

-

-

-

-

-

Overall

-6.62%

-7.66%

-7.28%

-6.71%

-5.77%

-4.81%

132%

7557%

97%

Class D

-7.57%

-9.45%

-9.68%

-5.76%

-6.81%

-5.94%

129%

8193%

98%

Class F

-3.07%

-5.88%

-4.94%

-3.05%

-4.76%

-3.54%

148%

9739%

99%

Class H

-

-

-

-

-

-

-

-

-

Low delay B Main 10

BD-rate Over NNVC-4.0

Y-PSNR

U-PSNR

V-PSNR

Y-MSIM

U-MSIM

V-MSIM

EncT

DecT CPU

PSNR Overlap

Class A1

-

-

-

-

-

-

-

-

-

Class A2

-

-

-

-

-

-

-

-

-

Class B

-4.85%

-8.13%

-7.84%

-5.15%

-6.16%

-6.65%

134%

6812%

99%

Class C

-5.24%

-11.13%

-9.00%

-5.61%

-10.74%

-4.02%

126%

6233%

99%

Class E

-4.88%

-4.09%

-5.17%

-5.64%

-2.18%

-3.36%

162%

9877%

98%

Overall

-4.98%

-8.12%

-7.56%

-5.42%

-6.69%

-4.95%

138%

7257%

99%

Class D

-6.54%

-10.76%

-9.77%

-5.75%

-9.48%

-6.83%

125%

6400%

98%

Class F

-2.05%

-4.88%

-4.46%

-2.29%

-3.08%

-4.12%

148%

8941%

99%

Class H

-

-

-

-

-

-

-

-

-

All Intra Main 10

BD-rate Over NNVC-4.0

Y-PSNR

U-PSNR

V-PSNR

Y-MSIM

U-MSIM

V-MSIM

EncT

DecT CPU

PSNR Overlap

Class A1

-8.46%

-7.23%

-8.67%

-10.00%

-8.00%

-7.50%

210%

5134%

97%

Class A2

-6.61%

-8.04%

-7.94%

-6.90%

-6.44%

-5.44%

204%

4416%

98%

Class B

-7.03%

-8.58%

-8.71%

-7.04%

-7.61%

-6.99%

203%

4493%

98%

Class C

-7.34%

-9.76%

-9.92%

-7.61%

-8.38%

-7.64%

190%

3770%

98%

Class E

-10.29%

-9.04%

-8.85%

-10.14%

-7.52%

-6.38%

199%

5149%

96%

Overall

-7.81%

-8.60%

-8.87%

-8.16%

-7.64%

-6.86%

201%

4507%

97%

Class D

-7.57%

-8.11%

-9.33%

-7.16%

-6.81%

-7.56%

180%

3760%

98%

Class F

-4.55%

-5.97%

-5.65%

-4.24%

-4.75%

-4.38%

155%

3961%

98%

Class H

-

-

-

-

-

-

-

-

-

Input contributions

There are 47 input contributions related to the AHG mandates. Ten of the contributions are part of the EE activity, while the remaining 37 contributions are related to AHG11 but not part of the EE. The list of input contributions is provided below.

EE and Related Input Contributions

Reporting

JVET-AE0023

EE1: Summary report of exploration experiment on neural network-based video coding

E. Alshina, F. Galpin, Y. Li, D. Rusanovskyy, M. Santamaria, J. Ström, L. Wang, Z. Xie

EE Technology

JVET-AE0067

EE1-4.1: Neural-network loop filters with further complexity reduction

J. N. Shingala, A. Shyam, A. Suneja, S. Badya (Ittiam), T. Shao, A. Arora, P. Yin, F. Pu, T. Lu, Sean McCarthy (Dolby)

JVET-AE0112

EE1-5.1: Deep Reference Frame Generation for Inter Prediction Enhancement

W. Bao, W. Meng, J. Jia, Y. Zhang, H. Wang, Z. Chen (Wuhan Univ.), Z. Liu, X. Xu, S. Liu (Tencent)

JVET-AE0144

EE1-6.1: neural network-based intra prediction with reduced complexity

T. Dumas, F. Galpin, P. Bordes (Interdigital)

JVET-AE0160

EE1-1.5: Optimization for complexity-performance trade-off of HOP network

R. Chang, L. Wang, X. Xu, S. Liu (Tencent)

JVET-AE0164

EE1-1.2 Complexity-performance tradeoff of decomposition

D. Rusanovskyy, Y. Li, M. Karczewicz (Qualcomm)

JVET-AE0165

EE1-4.4: Low complexity NN filter with design elements of Unified Filter Architecture and EE1-1.2 and EE1-1.3

Y. Li, D. Rusanovskyy, M. Karczewicz (Qualcomm)

JVET-AE0191

AhG11: EE1-0 High Operation Point model

F. Galpin (InterDigital), S. Eadie, D. Rusanovskyy (Qualcomm), Y. Li, J. Li (ByteDance), L. Wang, R. Chang (Tencent), Z. Xie (Oppo), E. Alshina (Huawei)

Cross Checks

JVET-AE0183

Crosscheck of JVET-AE0067 (EE1-4.1: Neural-network loop filters with further complexity reduction)

J. Ström (Ericsson)

JVET-AE0229

Crosscheck of JVET-AE0112(EE1-5.1: Deep Reference Frame Generation for Inter Prediction Enhancement)

X.Jie (OPPO)

Non-EE Input Contributions

Reporting

JVET-AE0011

JVET AHG report: Neural network-based video coding (AHG11)

E. Alshina, S. Liu, A. Segall (co chairs), F. Galpin, J. Li, R.-L. Liao, D. Rusanovskyy, T. Shao, M. Wien, P. Wu (vice chairs)

JVET-AE0042

AhG14 & AHG11: Report on AhG teleconference on high operation point (HOP) unified filter training

E. Alshina, F. Galpin

JVET-AE0237

Presentation of AI-based video Coding in "Third AG4 Workshop on JPEG and MPEG Emerging Activities"

E. Alshina

Loop Filtering

JVET-AE0072

[AHG11] On residual adjustments for NNLF

Z. Dai, Y. Yu, H. Yu, D. Wang (OPPO)

JVET-AE0093

AHG11: Content-adaptive neural network loop-filter

R. Yang, M. Santamaria, F. Cricri, H. Zhang, J. Lainema, R. G. Youvalari, M. M. Hannuksela (Nokia)

JVET-AE0161

AHG11: Input and output rotation of model for NNVC in-loop filter

R. Chang, L. Wang, X. Xu, S. Liu (Tencent)

JVET-AE0171

AHG11: Neural network-based in-loop filter with layer normalization

Y. Li, K. Zhang, L. Zhang (Bytedance)

Post Filtering

JVET-AE0048

AHG9: Miscellaneous VSEI changes on neural-network post-filter SEI messages

M. M. Hannuksela, F. Cricri, M. Santamaria (Nokia)

JVET-AE0049

AHG9: Miscellaneous VVC changes on neural-network post-filter SEI messages

M. M. Hannuksela, F. Cricri, M. Santamaria (Nokia)

JVET-AE0050

AHG9: On NNPF input picture selection

M. M. Hannuksela, F. Cricri (Nokia)

JVET-AE0051

AHG9: On persistent NNPF activation

M. M. Hannuksela, F. Cricri (Nokia)

JVET-AE0052

AHG9: NNPF cascades and alternatives

M. M. Hannuksela, F. Cricri (Nokia)

JVET-AE0053

AHG8/AHG9: Neural-network post-filter regions SEI message

T. Chujoh, Y. Yasugi, T. Ikai (Sharp)

JVET-AE0060

[AHG9] Comments on NNPFC

S. Deshpande (Sharp)

JVET-AE0061

[AHG9] On NNPFC Application Purpose

S. Deshpande (Sharp)

JVET-AE0062

[AHG9] On NNPF for Deinterlacing

A. Sidiya, S. Deshpande (Sharp)

JVET-AE0063

[AHG9] On grouping and operations for multiple NNPFs

L. Chen, O. Chubach, Y.-W. Huang, S. Lei (MediaTek)

JVET-AE0068

AHG9: On extensibility of purpose syntax element in NNPFC SEI message

Hendry, J. Nam, S. Kim, J. Lim (LGE)

JVET-AE0069

AHG9: On the signalling of output pictures in NNPFA SEI message

Hendry, J. Nam, S. Kim, J. Lim (LGE)

JVET-AE0070

AHG9: On input pictures that are not present in the bitstream for NNPF SEI messages

Hendry, J. Nam, S. Kim, J. Lim (LGE)

JVET-AE0101

[AHG2][AHG9] Neural network post filter and phase indication SEI messages for AVC and HEVC

T. Ikai, T. Chujoh (Sharp), Y.-K. Wang, J. Xu, W. Jia (Bytedance)

JVET-AE0106

AHG9: On missing value ranges for some syntax elements in the NNPFC SEI message

C. Lin, Y.-K. Wang, J. Xu, W. Jia, J. Li, Y. Li, K. Zhang, L. Zhang (Bytedance)

JVET-AE0113

AHG9: Extendibility and code word length of nnpfc_purpose

R. Sjöberg, M. Pettersson (Ericsson)

JVET-AE0126

AHG9: NNPF cleanup and editorial changes for VSEI

Y.-K. Wang, W. Jia, J. Xu, C. Lin (Bytedance)

JVET-AE0127

AHG9: NNPF editorial changes for VVC

Y.-K. Wang, W. Jia, J. Xu (Bytedance)

JVET-AE0128

AHG9: On NNPFC extensibility and base filter signalling

Y.-K. Wang (Bytedance)

JVET-AE0134

AHG9: Align the design of NNPF with multiple input pictures to NNPF including picture rate upsampling

J. Xu, Y.-K. Wang (Bytedance)

JVET-AE0135

AHG9: On NNPF picture rate upsampling constraints

J. Xu, Y.-K. Wang (Bytedance)

JVET-AE0141

AHG9: Fix a bug in NNPFC SEI message for colourization

J. Xu, Y.-K. Wang (Bytedance)

JVET-AE0142

AHG9: On derivation of NNPF input pictures and the value of nnpfc_purpose

W. Jia, Y.-K. Wang, J. Xu, L. Zhang (Bytedance)

JVET-AE0155

AHG2/AHG9: Editor commentary on the post-filter hint SEI message semantics

G. J. Sullivan, Y.-K. Wang (Editors)

JVET-AE0173

[AHG9] Clarification and improvements of signalling of NNPF update

Y. Lim (Samsung)

JVET-AE0175

[AHG9] Editorial improvements of nnpfc_mode_idc

Y. Lim (Samsung)

JVET-AE0187

AHG9: A summary of SEI proposals on NNPF

Y.-K. Wang (Bytedance)

JVET-AE0189

AHG9: On the design of nnpfa_target_base_flag in NNPFA SEI message

Hendry (LGE)

Test Conditions

JVET-AE0162

AHG11/AHG14: Fix MS-SSIM calculation for SR

R. Chang, L. Wang, X. Xu, S. Liu (Tencent)

Cross Checks

JVET-AE0230

Crosscheck of JVET-AE0162(AHG11/AHG14 : Fix MS-SSIM calculation for SR)

Z. Xie (OPPO)

JVET-AE0232

Crosscheck of JVET-AE0072 ([AHG11] On residual adjustments for NNLF)

T. Shao (OPPO)

Recommendations

The AHG recommended to:

  • Review all input contributions.
  • Continue investigating neural network-based video coding tools, including coding performance and complexity.
  • Encourage EE1 to continue using a workflow similar to the one in JVET-AE0191 with unified training, clearly documented NNVC technology and transparent training process performed by multiple parties in parallel with regular information sharing.
Decisions
The AHG recommended to: Review all input contributions. Continue investigating neural network-based video coding tools, including coding performance and complexity. Encourage EE1 to continue using a workflow similar to the one in JVET-AE0191 with unified training, clearly documented NNVC technology and transparent training process performed by multiple parties in parallel with regular information sharing.
Citation