JVET-AE0011 JVET AHG report: Neural network-based video coding (AHG11) [E. Alshina, S. Liu, A. Segall (co chairs), F. Galpin, J. Li, R.-L. Liao, D. Rusanovskyy, T. Shao, M. Wien, P. Wu (vice-chairs)]
Activities
The AHG used the main JVET reflector, jvet@lists.rwth-aachen.de, for email. Forty-one emails were exchanged on the reflector related to the AHG mandates.
Common Test Conditions
Document
The AHG released revised common test conditions as decided at the 30th meeting. The final version was uploaded as document JVET-AD2016 on May 12, 2023.
Anchor Encoding
Anchors for the NN-based video coding activity made available on the Git repository used for the AHG activity: https://vcgit.hhi.fraunhofer.de/jvet-ahg-nnvc/nnvc-ctc/-/tree/master.
EE Coordination
The AHG finalized, conducted and discussed the EE on NN based video coding. The final version of the EE description was uploaded to the document repository on May 15, 2023.
A summary report for the EE is available at this meeting as:
EE1: Summary report of exploration experiment on neural network-based video coding | E. Alshina, F. Galpin, Y. Li, D. Rusanovskyy, M. Santamaria, J. Ström, L. Wang, Z. Xie (EE coordinators) |
Teleconferences
The AHG conducted two teleconferences during the interim period. The teleconferences were held on June 14, 2023, and June 28, 2023 and attended by approximately 20 participants.
A summary report for the teleconferences is available at this meeting as:
AhG14 & AHG11: Report on AhG teleconference on high operation point (HOP) unified filter training | E. Alshina, F. Galpin |
Workshop on Neural Network-Based Technologies
The AHG contributed to the “Third AG4 Workshop on JPEG and MPEG Emerging Activities” that was conducted by the JPEG and MPEG Collaboration Advisory Group, or ISO/IEC JTC1/SC29/AG4.
The workshop took place on June 1st, 2023, and its agenda was:
Speaker | Topic |
Elena Alshina | AI-based Video Coding |
Werner Bailer | Neural Network Compression |
Marc Antonini | Coding for DNA Storage |
Tim Bruylants | Event-based Vision |
Arianne Hids | Scene-based Interchange and Support of Immersive Displays |
The AHG contributed the presentation titled “AI-based Video Coding”. Slides for the presentation are available in JVET-AE0237.
Performance Evaluation
The performance of the NNVC-5.0 anchor compared to NNVC-4.0 is reported below.
Random access Main 10 | |||||||||
BD-rate Over NNVC-4.0 | |||||||||
Y-PSNR | U-PSNR | V-PSNR | Y-MSIM | U-MSIM | V-MSIM | EncT | DecT CPU | PSNR Overlap | |
Class A1 | -7.23% | -4.86% | -6.00% | -8.72% | -4.05% | -5.06% | 134% | 7667% | 96% |
Class A2 | -6.56% | -5.78% | -5.09% | -6.53% | -4.14% | -2.90% | 130% | 7183% | 97% |
Class B | -6.29% | -8.19% | -7.53% | -5.97% | -6.06% | -5.18% | 132% | 7834% | 98% |
Class C | -6.62% | -10.53% | -9.55% | -6.26% | -7.91% | -5.58% | 132% | 7422% | 98% |
Class E | - | - | - | - | - | - | - | - | - |
Overall | -6.62% | -7.66% | -7.28% | -6.71% | -5.77% | -4.81% | 132% | 7557% | 97% |
Class D | -7.57% | -9.45% | -9.68% | -5.76% | -6.81% | -5.94% | 129% | 8193% | 98% |
Class F | -3.07% | -5.88% | -4.94% | -3.05% | -4.76% | -3.54% | 148% | 9739% | 99% |
Class H | - | - | - | - | - | - | - | - | - |
Low delay B Main 10 | |||||||||
BD-rate Over NNVC-4.0 | |||||||||
Y-PSNR | U-PSNR | V-PSNR | Y-MSIM | U-MSIM | V-MSIM | EncT | DecT CPU | PSNR Overlap | |
Class A1 | - | - | - | - | - | - | - | - | - |
Class A2 | - | - | - | - | - | - | - | - | - |
Class B | -4.85% | -8.13% | -7.84% | -5.15% | -6.16% | -6.65% | 134% | 6812% | 99% |
Class C | -5.24% | -11.13% | -9.00% | -5.61% | -10.74% | -4.02% | 126% | 6233% | 99% |
Class E | -4.88% | -4.09% | -5.17% | -5.64% | -2.18% | -3.36% | 162% | 9877% | 98% |
Overall | -4.98% | -8.12% | -7.56% | -5.42% | -6.69% | -4.95% | 138% | 7257% | 99% |
Class D | -6.54% | -10.76% | -9.77% | -5.75% | -9.48% | -6.83% | 125% | 6400% | 98% |
Class F | -2.05% | -4.88% | -4.46% | -2.29% | -3.08% | -4.12% | 148% | 8941% | 99% |
Class H | - | - | - | - | - | - | - | - | - |
All Intra Main 10 | |||||||||
BD-rate Over NNVC-4.0 | |||||||||
Y-PSNR | U-PSNR | V-PSNR | Y-MSIM | U-MSIM | V-MSIM | EncT | DecT CPU | PSNR Overlap | |
Class A1 | -8.46% | -7.23% | -8.67% | -10.00% | -8.00% | -7.50% | 210% | 5134% | 97% |
Class A2 | -6.61% | -8.04% | -7.94% | -6.90% | -6.44% | -5.44% | 204% | 4416% | 98% |
Class B | -7.03% | -8.58% | -8.71% | -7.04% | -7.61% | -6.99% | 203% | 4493% | 98% |
Class C | -7.34% | -9.76% | -9.92% | -7.61% | -8.38% | -7.64% | 190% | 3770% | 98% |
Class E | -10.29% | -9.04% | -8.85% | -10.14% | -7.52% | -6.38% | 199% | 5149% | 96% |
Overall | -7.81% | -8.60% | -8.87% | -8.16% | -7.64% | -6.86% | 201% | 4507% | 97% |
Class D | -7.57% | -8.11% | -9.33% | -7.16% | -6.81% | -7.56% | 180% | 3760% | 98% |
Class F | -4.55% | -5.97% | -5.65% | -4.24% | -4.75% | -4.38% | 155% | 3961% | 98% |
Class H | - | - | - | - | - | - | - | - | - |
Input contributions
There are 47 input contributions related to the AHG mandates. Ten of the contributions are part of the EE activity, while the remaining 37 contributions are related to AHG11 but not part of the EE. The list of input contributions is provided below.
EE and Related Input Contributions
Reporting | ||
EE1: Summary report of exploration experiment on neural network-based video coding | E. Alshina, F. Galpin, Y. Li, D. Rusanovskyy, M. Santamaria, J. Ström, L. Wang, Z. Xie | |
EE Technology | ||
EE1-4.1: Neural-network loop filters with further complexity reduction | J. N. Shingala, A. Shyam, A. Suneja, S. Badya (Ittiam), T. Shao, A. Arora, P. Yin, F. Pu, T. Lu, Sean McCarthy (Dolby) | |
EE1-5.1: Deep Reference Frame Generation for Inter Prediction Enhancement | W. Bao, W. Meng, J. Jia, Y. Zhang, H. Wang, Z. Chen (Wuhan Univ.), Z. Liu, X. Xu, S. Liu (Tencent) | |
EE1-6.1: neural network-based intra prediction with reduced complexity | T. Dumas, F. Galpin, P. Bordes (Interdigital) | |
EE1-1.5: Optimization for complexity-performance trade-off of HOP network | R. Chang, L. Wang, X. Xu, S. Liu (Tencent) | |
EE1-1.2 Complexity-performance tradeoff of decomposition | D. Rusanovskyy, Y. Li, M. Karczewicz (Qualcomm) | |
EE1-4.4: Low complexity NN filter with design elements of Unified Filter Architecture and EE1-1.2 and EE1-1.3 | Y. Li, D. Rusanovskyy, M. Karczewicz (Qualcomm) | |
AhG11: EE1-0 High Operation Point model | F. Galpin (InterDigital), S. Eadie, D. Rusanovskyy (Qualcomm), Y. Li, J. Li (ByteDance), L. Wang, R. Chang (Tencent), Z. Xie (Oppo), E. Alshina (Huawei) | |
Cross Checks | ||
Crosscheck of JVET-AE0067 (EE1-4.1: Neural-network loop filters with further complexity reduction) | J. Ström (Ericsson) | |
Crosscheck of JVET-AE0112(EE1-5.1: Deep Reference Frame Generation for Inter Prediction Enhancement) | X.Jie (OPPO) | |
Non-EE Input Contributions
Reporting | ||
JVET AHG report: Neural network-based video coding (AHG11) | E. Alshina, S. Liu, A. Segall (co chairs), F. Galpin, J. Li, R.-L. Liao, D. Rusanovskyy, T. Shao, M. Wien, P. Wu (vice chairs) | |
AhG14 & AHG11: Report on AhG teleconference on high operation point (HOP) unified filter training | E. Alshina, F. Galpin | |
Presentation of AI-based video Coding in "Third AG4 Workshop on JPEG and MPEG Emerging Activities" | E. Alshina | |
Loop Filtering | ||
[AHG11] On residual adjustments for NNLF | Z. Dai, Y. Yu, H. Yu, D. Wang (OPPO) | |
AHG11: Content-adaptive neural network loop-filter | R. Yang, M. Santamaria, F. Cricri, H. Zhang, J. Lainema, R. G. Youvalari, M. M. Hannuksela (Nokia) | |
AHG11: Input and output rotation of model for NNVC in-loop filter | R. Chang, L. Wang, X. Xu, S. Liu (Tencent) | |
AHG11: Neural network-based in-loop filter with layer normalization | Y. Li, K. Zhang, L. Zhang (Bytedance) | |
Post Filtering | ||
AHG9: Miscellaneous VSEI changes on neural-network post-filter SEI messages | M. M. Hannuksela, F. Cricri, M. Santamaria (Nokia) | |
AHG9: Miscellaneous VVC changes on neural-network post-filter SEI messages | M. M. Hannuksela, F. Cricri, M. Santamaria (Nokia) | |
AHG9: On NNPF input picture selection | M. M. Hannuksela, F. Cricri (Nokia) | |
AHG9: On persistent NNPF activation | M. M. Hannuksela, F. Cricri (Nokia) | |
AHG9: NNPF cascades and alternatives | M. M. Hannuksela, F. Cricri (Nokia) | |
AHG8/AHG9: Neural-network post-filter regions SEI message | T. Chujoh, Y. Yasugi, T. Ikai (Sharp) | |
[AHG9] Comments on NNPFC | S. Deshpande (Sharp) | |
[AHG9] On NNPFC Application Purpose | S. Deshpande (Sharp) | |
[AHG9] On NNPF for Deinterlacing | A. Sidiya, S. Deshpande (Sharp) | |
[AHG9] On grouping and operations for multiple NNPFs | L. Chen, O. Chubach, Y.-W. Huang, S. Lei (MediaTek) | |
AHG9: On extensibility of purpose syntax element in NNPFC SEI message | Hendry, J. Nam, S. Kim, J. Lim (LGE) | |
AHG9: On the signalling of output pictures in NNPFA SEI message | Hendry, J. Nam, S. Kim, J. Lim (LGE) | |
AHG9: On input pictures that are not present in the bitstream for NNPF SEI messages | Hendry, J. Nam, S. Kim, J. Lim (LGE) | |
[AHG2][AHG9] Neural network post filter and phase indication SEI messages for AVC and HEVC | T. Ikai, T. Chujoh (Sharp), Y.-K. Wang, J. Xu, W. Jia (Bytedance) | |
AHG9: On missing value ranges for some syntax elements in the NNPFC SEI message | C. Lin, Y.-K. Wang, J. Xu, W. Jia, J. Li, Y. Li, K. Zhang, L. Zhang (Bytedance) | |
AHG9: Extendibility and code word length of nnpfc_purpose | R. Sjöberg, M. Pettersson (Ericsson) | |
AHG9: NNPF cleanup and editorial changes for VSEI | Y.-K. Wang, W. Jia, J. Xu, C. Lin (Bytedance) | |
AHG9: NNPF editorial changes for VVC | Y.-K. Wang, W. Jia, J. Xu (Bytedance) | |
AHG9: On NNPFC extensibility and base filter signalling | Y.-K. Wang (Bytedance) | |
AHG9: Align the design of NNPF with multiple input pictures to NNPF including picture rate upsampling | J. Xu, Y.-K. Wang (Bytedance) | |
AHG9: On NNPF picture rate upsampling constraints | J. Xu, Y.-K. Wang (Bytedance) | |
AHG9: Fix a bug in NNPFC SEI message for colourization | J. Xu, Y.-K. Wang (Bytedance) | |
AHG9: On derivation of NNPF input pictures and the value of nnpfc_purpose | W. Jia, Y.-K. Wang, J. Xu, L. Zhang (Bytedance) | |
AHG2/AHG9: Editor commentary on the post-filter hint SEI message semantics | G. J. Sullivan, Y.-K. Wang (Editors) | |
[AHG9] Clarification and improvements of signalling of NNPF update | Y. Lim (Samsung) | |
[AHG9] Editorial improvements of nnpfc_mode_idc | Y. Lim (Samsung) | |
AHG9: A summary of SEI proposals on NNPF | Y.-K. Wang (Bytedance) | |
AHG9: On the design of nnpfa_target_base_flag in NNPFA SEI message | Hendry (LGE) | |
Test Conditions | ||
AHG11/AHG14: Fix MS-SSIM calculation for SR | R. Chang, L. Wang, X. Xu, S. Liu (Tencent) | |
Cross Checks | ||
Crosscheck of JVET-AE0162(AHG11/AHG14 : Fix MS-SSIM calculation for SR) | Z. Xie (OPPO) | |
Crosscheck of JVET-AE0072 ([AHG11] On residual adjustments for NNLF) | T. Shao (OPPO) | |
Recommendations
The AHG recommended to:
- Review all input contributions.
- Continue investigating neural network-based video coding tools, including coding performance and complexity.
- Encourage EE1 to continue using a workflow similar to the one in JVET-AE0191 with unified training, clearly documented NNVC technology and transparent training process performed by multiple parties in parallel with regular information sharing.