JVET-T0011 JVET AHG report: Neural network-based video coding (AHG11) [A. Alshina, S. Liu, J. Pfaff, M. Wien, P. Wu, Y. Ye]
This AHG report was reviewed in session 6 at 0730 UTC Thursday 8 October (chaired by GJS & JRO).
This document summarizes the activities of AHG11: Neural network-based video coding between the 19th Meeting (22 June –1 July 2020) and the 20th meeting (7–16 Oct. 2020), both held by teleconference.
One teleconference meeting was held on July 30 with total 160 experts participated. Among the 160 participants, more than 140 stayed in the meeting for longer than 1 hour, about 120 participated the full (or nearly full) 2 hr meeting.
In this meeting, the following input document was reviewed and discussed. Valuable inputs were received and a couple of revisions of this document were generated and uploaded after the telco meeting.
- JVET-T0041 Methodology and reporting template for neural network coding tool testing [S. Liu, A. Segall, E. Alshina, J. Boyce]
A collection of training datasets was reviewed in this telco meeting.
The telco meeting report (JVET-T0042) was uploaded to the JVET document web site shortly after the meeting.
Anchor generation
Following the suggestion and agreement reached in the AHG telco meeting, AHG11 should generate anchor results using five QP values (22, 27, 32, 37, 42). B-D rates should be calculated using all five QP values. It was suggested to provide Low Delay 4K anchor results and make it optional for proponents to test and report.
The anchor results were generated and announced through JVET reflector on September 22. Thanks to Tencent and Huawei for anchor generation and crosscheck.
Training data
A couple of datasets were identified to be suitable for AHG11 training process, such as the ones from xiph.org and JVET. In addition, University of Bristol would like to contribute a video training dataset, BVI-DVC, which contains 800 sequences at various spatial resolutions from 270p to 2160p and has been evaluated on ten existing network architectures for four different coding tools (https://arxiv.org/abs/2003.13552). Among these sequences, 104 are sourced by University of Bristol and have been granted copyright permission to JVET. University of Bristol is working on collecting copyright permissions from the third party source owners. The progress was reported positive.
In addition, discussions with respect to quality metrics, such as including MS-SSIM besides the conventional PSNR, and related issues were held among interested experts.
AHG11 related input contributions to this meeting included:
- JVET-T0041, Methodology and reporting template for neural network coding tool testing, S. Liu, A. Segall, E. Alshina, J. Boyce, M. Wien, D. Grois (AHG)
- JVET-T0042, Report of AHG11 Meeting on Neural Network-based Video Coding, S. Liu, E. Alshina, J. Pfaff, M. Wien, P. Wu, Y. Ye (AHG)
- JVET-T0057, AHG11: A case study to reduce computation of a neural network based in-loop filter by pruning, C. Auyeung, W. Wang, W. Jiang, X. Li, S. Liu (Tencent)
- JVET-T0058, AHG11: Information on inter-prediction coding tool with deep neural network, B. Choi, Z. Li, W. Wang, W. Jiang, X. Xu, S. Liu (Tencent)
- JVET-T0069, AHG11: SSIM based CNN model for in-loop filtering, T. Ouyang, F. Liu, H. Zhu, Z. Chen (Wuhan Unvi.), X. Xu, S. Liu (Tencent)
- JVET-T0073, AHG11: Neural Network-based intra prediction with transform selection in VVC, T. Dumas, F. Galpin, P. Bordes, F. Leleannec (Interdigital)
- JVET-T0079, AHG11: Neural Network-based In-Loop Filter, H. Wang, M. Karczewicz, J. Chen, A.M. Kotra (Qualcomm)
- JVET-T0088, AHG11: Convolutional neural networks-based in-loop filter, Y. Li, L. Zhang, K. Zhang, Y. He, J. Xu (Bytedance)
- JVET-T0092, AHG9/AHG11: Neural network based super resolution SEI, T. Chujoh, E. Sasaki, T. Ikai (Sharp)
- JVET-T0094, AHG11: In-loop filtering based on neutral network, T.-C. Ma, W. Chen, X. Xiu, Y.-W. Chen, H.-J. Jhu, C.-W. Kuo, X. Wang (Kwai)
- JVET-T0096, AHG11: Deep neural network for super-resolution, J. Y. Lee, Y. Choi (Sejong University), W. Lim, G. Bang (ETRI)
Summary results of technical proposals
Coding performance and runtime complexity
Doc. # | Category | YUV BD rate (AI/RA/LDB) | Enc% CPU | Dec% CPU | Enc% GPU | Dec% GPU | |||
Complexity reduction | RA | −0.79% | −3.70% | −2.82% | 105% | 3244% | |||
LDB | −0.64% | −4.58% | −4.17% | 106% | 3848% | ||||
LDP | −0.65% | −4.95% | −4.49% | 109% | 4157% | ||||
AI | −1.07% | −2.79% | −2.74% | 108% | 3418% | ||||
Virtual ref. picture | RA (Class B, C, D) | −2.26% | 1.81% | 3.13% | 126% | 46257% | |||
LDB | |||||||||
AI | |||||||||
CNN in-loop filter | RA | −0.66% | −5.04% | −3.89% | 20826% | 1859% | |||
LDB | 0.20% | −7.44% | −5.69% | 143% | 19991% | 1733% | |||
AI | −1.61% | −6.76% | −6.46% | 136% | 26294% | 2331% | |||
CNN in-loop filter | RA | −1.84% | −9.00% | −7.61% | 27499% | 2382% | |||
LDB | −1.37% | −10.00% | −9.04% | 134% | 28450% | 2459% | |||
AI | −2.11% | −7.70% | −7.70% | 129% | 25695% | 2312% | |||
NN intra prediction and LFNST | RA | −1.57% | −0.96% | −1.15% | 175% | 692% | |||
LDB | |||||||||
AI | −3.49% | −3.04% | −3.07% | 455% | 3889% | ||||
CNN in-loop filter | RA | −3.20% | −12.25% | −11.39% | 118% | 13967% | |||
RA (inter refinement) | −4.11% | −13.83% | −13.28% | 196% | 14058% | ||||
LDB | |||||||||
AI | −3.63% | −9.47% | −9.98% | 119% | 7853% | ||||
CNN in-loop filter | RA | ||||||||
LDB | |||||||||
AI | −7.45% | −12.24% | −10.67% | 118% | 15582% | ||||
CNN in-loop filter | RA | −3.92% | −18.09% | −16.93% | 193% | 754013% | |||
LDB | |||||||||
AI | −4.99% | −16.39% | −17.34% | 875% | 775056% | ||||
Super resolution | |||||||||
Network information and complexity
Network information in training and inference stages
Doc. # | Category | Training time | Training data | Total Parameters | Memory Parameter (MB) | Memory Temp. (MB) | MAC (Giga) |
Complexity reduction | 24h | DIV2K | 16,778 | 0.067 | 15238.7 | 129.61 | |
Virtual ref. picture | 7 days | Xiph.org, BVI-DVC (subset) | 125,993,047 | 504 | - | 200 | |
CNN in-loop filter | 57h | DIV2K | 131,299 | 0.5 | 32399.36 | 830 | |
NN intra prediction and LFNST | 8h | ILSVRC2012 DIV2K CLIC2020 | 11,614,376 | 50.3 | - | 10.9 | |
CNN in-loop filter | - | - | 1million | - | - | - | |
CNN in-loop filter | 30h | DIV2K, BVI-DVC | 4,876,181 | 18.6 | 1012.5 | 3.05 | |
CNN in-loop filter | - | DIV2K | - | - | - | - | |
Super resolution | - | - | - | - | - | - |
The AHG recommended:
- To review input contributions;
- To continue investigating neural network-based video coding tools, including coding performance, complexity and quality metrics;
- To further discuss, define, refine and regulate common testing as well as training conditions for neural network-based video coding.