Back to Search Document details
20th Meeting: by teleconference, October 2020 2020-10-08 10:12
JVET AHG report: Neural-network-based video coding (AHG11)
Abstract
This document summarizes the activities of AHG11: Neural-network-based video coding between the 19th Meeting (22 June -1 July 2020) and the 20th meeting (07 - 16 Oct. 2020), both held by teleconference.
JVET-T0011 JVET AHG report: Neural network-based video coding (AHG11) [A. Alshina, S. Liu, J. Pfaff, M. Wien, P. Wu, Y. Ye]

This AHG report was reviewed in session 6 at 0730 UTC Thursday 8 October (chaired by GJS & JRO).

This document summarizes the activities of AHG11: Neural network-based video coding between the 19th Meeting (22 June –1 July 2020) and the 20th meeting (7–16 Oct. 2020), both held by teleconference.

One teleconference meeting was held on July 30 with total 160 experts participated. Among the 160 participants, more than 140 stayed in the meeting for longer than 1 hour, about 120 participated the full (or nearly full) 2 hr meeting.

In this meeting, the following input document was reviewed and discussed. Valuable inputs were received and a couple of revisions of this document were generated and uploaded after the telco meeting.

  • JVET-T0041 Methodology and reporting template for neural network coding tool testing [S. Liu, A. Segall, E. Alshina, J. Boyce]

A collection of training datasets was reviewed in this telco meeting.

The telco meeting report (JVET-T0042) was uploaded to the JVET document web site shortly after the meeting.

Anchor generation

Following the suggestion and agreement reached in the AHG telco meeting, AHG11 should generate anchor results using five QP values (22, 27, 32, 37, 42). B-D rates should be calculated using all five QP values. It was suggested to provide Low Delay 4K anchor results and make it optional for proponents to test and report.

The anchor results were generated and announced through JVET reflector on September 22. Thanks to Tencent and Huawei for anchor generation and crosscheck.

Training data

A couple of datasets were identified to be suitable for AHG11 training process, such as the ones from xiph.org and JVET. In addition, University of Bristol would like to contribute a video training dataset, BVI-DVC, which contains 800 sequences at various spatial resolutions from 270p to 2160p and has been evaluated on ten existing network architectures for four different coding tools (https://arxiv.org/abs/2003.13552). Among these sequences, 104 are sourced by University of Bristol and have been granted copyright permission to JVET. University of Bristol is working on collecting copyright permissions from the third party source owners. The progress was reported positive.

In addition, discussions with respect to quality metrics, such as including MS-SSIM besides the conventional PSNR, and related issues were held among interested experts.

AHG11 related input contributions to this meeting included:

  • JVET-T0041, Methodology and reporting template for neural network coding tool testing, S. Liu, A. Segall, E. Alshina, J. Boyce, M. Wien, D. Grois (AHG)
  • JVET-T0042, Report of AHG11 Meeting on Neural Network-based Video Coding, S. Liu, E. Alshina, J. Pfaff, M. Wien, P. Wu, Y. Ye (AHG)
  • JVET-T0057, AHG11: A case study to reduce computation of a neural network based in-loop filter by pruning, C. Auyeung, W. Wang, W. Jiang, X. Li, S. Liu (Tencent)
  • JVET-T0058, AHG11: Information on inter-prediction coding tool with deep neural network, B. Choi, Z. Li, W. Wang, W. Jiang, X. Xu, S. Liu (Tencent)
  • JVET-T0069, AHG11: SSIM based CNN model for in-loop filtering, T. Ouyang, F. Liu, H. Zhu, Z. Chen (Wuhan Unvi.), X. Xu, S. Liu (Tencent)
  • JVET-T0073, AHG11: Neural Network-based intra prediction with transform selection in VVC, T. Dumas, F. Galpin, P. Bordes, F. Leleannec (Interdigital)
  • JVET-T0079, AHG11: Neural Network-based In-Loop Filter, H. Wang, M. Karczewicz, J. Chen, A.M. Kotra (Qualcomm)
  • JVET-T0088, AHG11: Convolutional neural networks-based in-loop filter, Y. Li, L. Zhang, K. Zhang, Y. He, J. Xu (Bytedance)
  • JVET-T0092, AHG9/AHG11: Neural network based super resolution SEI, T. Chujoh, E. Sasaki, T. Ikai (Sharp)
  • JVET-T0094, AHG11: In-loop filtering based on neutral network, T.-C. Ma, W. Chen, X. Xiu, Y.-W. Chen, H.-J. Jhu, C.-W. Kuo, X. Wang (Kwai)
  • JVET-T0096, AHG11: Deep neural network for super-resolution, J. Y. Lee, Y. Choi (Sejong University), W. Lim, G. Bang (ETRI)

Summary results of technical proposals

Coding performance and runtime complexity

Doc. #

Category

YUV BD rate (AI/RA/LDB)

Enc%

CPU

Dec%

CPU

Enc%

GPU

Dec%

GPU

JVET-T0057

Complexity reduction

RA

−0.79%

−3.70%

−2.82%

105%

3244%

LDB

−0.64%

−4.58%

−4.17%

106%

3848%

LDP

−0.65%

−4.95%

−4.49%

109%

4157%

AI

−1.07%

−2.79%

−2.74%

108%

3418%

JVET-T0058

Virtual ref. picture

RA (Class B, C, D)

−2.26%

1.81%

3.13%

126%

46257%

LDB

AI

JVET-T0069

CNN in-loop filter
(result-1)

RA

−0.66%

−5.04%

−3.89%

20826%

1859%

LDB

0.20%

−7.44%

−5.69%

143%

19991%

1733%

AI

−1.61%

−6.76%

−6.46%

136%

26294%

2331%

CNN in-loop filter
(result-2)

RA

−1.84%

−9.00%

−7.61%

27499%

2382%

LDB

−1.37%

−10.00%

−9.04%

134%

28450%

2459%

AI

−2.11%

−7.70%

−7.70%

129%

25695%

2312%

JVET-T0073

NN intra prediction and LFNST

RA

−1.57%

−0.96%

−1.15%

175%

692%

LDB

AI

−3.49%

−3.04%

−3.07%

455%

3889%

JVET-T0079

CNN in-loop filter

RA

−3.20%

−12.25%

−11.39%

118%

13967%

RA (inter refinement)

−4.11%

−13.83%

−13.28%

196%

14058%

LDB

AI

−3.63%

−9.47%

−9.98%

119%

7853%

JVET-T0088

CNN in-loop filter

RA

LDB

AI

−7.45%

−12.24%

−10.67%

118%

15582%

JVET-T0094

CNN in-loop filter

RA

−3.92%

−18.09%

−16.93%

193%

754013%

LDB

AI

−4.99%

−16.39%

−17.34%

875%

775056%

JVET-T0096

Super resolution

Network information and complexity

Network information in training and inference stages

Doc. #

Category

Training time

Training data

Total Parameters

Memory Parameter (MB)

Memory Temp. (MB)

MAC (Giga)

JVET-T0057

Complexity reduction

24h

DIV2K

16,778

0.067

15238.7

129.61

JVET-T0058

Virtual ref. picture

7 days

Xiph.org, BVI-DVC (subset)

125,993,047

504

-

200

JVET-T0069

CNN in-loop filter

57h

DIV2K

131,299

0.5

32399.36

830

JVET-T0073

NN intra prediction and LFNST

8h

ILSVRC2012

DIV2K

CLIC2020

11,614,376

50.3

-

10.9

JVET-T0079

CNN in-loop filter

-

-

1million

-

-

-

JVET-T0088

CNN in-loop filter

30h

DIV2K, BVI-DVC

4,876,181

18.6

1012.5

3.05

JVET-T0094

CNN in-loop filter

-

DIV2K

-

-

-

-

JVET-T0096

Super resolution

-

-

-

-

-

-

The AHG recommended:

  • To review input contributions;
  • To continue investigating neural network-based video coding tools, including coding performance, complexity and quality metrics;
  • To further discuss, define, refine and regulate common testing as well as training conditions for neural network-based video coding.
Decisions
To further discuss, define, refine and regulate common testing as well as training conditions for neural network-based video coding.
Citation