Back to Search Document details
14th Meeting: Geneva, March 2019 2019-03-21 09:31
CE13: Summary Report on Neural Network based Filter for Video Coding
Abstract
This contribution provides a summary report of Core Experiment 13 on neural network based filter for video coding. 22 tests have been tested in CE13 in between JVET-M and JVET-N meetings, to study and evaluate technologies related to screen content coding. In this report, coding performance and complexity of these tests are reported and analyzed. Crosschecking results for the performed tests are integrated in this contribution.
JVET-N0033 CE13: Summary Report on Neural Network based Filter for Video Coding [Y. Li, S. Liu, K. Kawamura]

This contribution provides a summary report of Core Experiment 13 on neural network based filter for video coding. 22 tests have been tested in CE13 in between JVET-M and JVET-N meetings, to study and evaluate technologies related to screen content coding. In this report, coding performance and complexity of these tests are reported and analyzed. Crosschecking results for the performed tests are integrated in this contribution.

The following tests in the table below are performed in CE13.

Summary of tests performed in CE13

Test #

Description

Document

Tester

Cross checker

CE13-1.1a

Loop filter chain: DBF + SAO + ALF + NN filter. NN filter follows the description in M0159

JVET-N0110

Y.-L. Hsiao (Mediatek)

H. Yin (Intel)

CE13-1.1b

Loop filter chain: NN filter only. NN filter follows the description in M0159

JVET-N0110

Y.-L. Hsiao (Mediatek)

CE13-1.2a

Loop filter chain: DBF + SAO + ALF + NN filter. NN filter follows the description in M0566

JVET-N0480

H. Yin (Intel)

Y.-L. Hsiao (Mediatek)

CE13-1.2b

Loop filter chain: NN filter only. NN filter follows the description in M0566

JVET-N0480

H. Yin (Intel)

CE13-2.1a

DBF + NN filter + SAO + ALF, NN filter follows the description in M0351

JVET-N0169

C. Lin (Hikvision)

Y. Wang (Wuhan Univ.)

CE13-2.1b

NN filter only, NN filter follows the description in M0351

JVET-N0169

C. Lin (Hikvision)

Y. Wang (Wuhan Univ.)

CE13-2.1c

NN filter + ALF, NN filter follows the description in M0351

JVET-N0169

C. Lin (Hikvision)

Y. Wang (Wuhan Univ.)

CE13-2.2a

DBF + NN filter + SAO + ALF, NN filter follows the description in M0508

JVET-N0254

Y. Wang (Wuhan Univ.)

C. Lin (Hikvision)

CE13-2.2b

NN filter only, NN filter follows the description in M0508

JVET-N0254

Y. Wang (Wuhan Univ.)

CE13-2.2c

NN filter + ALF, NN filter follows the description in M0508

JVET-N0254

Y. Wang (Wuhan Univ.)

CE13-2.3a

CTC QP for training, CTC QP for testing (the same as CE13-2.2.a)

JVET-N0254

Y. Wang (Wuhan Univ.)

CE13-2.3b

CTC QP for training, CTC QP + 2 for testing

JVET-N0254

Y. Wang (Wuhan Univ.)

CE13-2.3c

CTC QP for training, CTC QP - 2 for testing

JVET-N0254

Y. Wang (Wuhan Univ.)

CE13-2.4a

DBF + NN filter + SAO + ALF, NN filter follows the description in M0510

JVET-N0513

Y. Dai (USTC)

Y. Wang (Wuhan Univ.)

CE13-2.4b

NN filter only, NN filter follows the description in M0510

JVET-N0513

Y. Dai (USTC)

CE13-2.4c

NN filter + ALF, NN filter follows the description in M0510

JVET-N0513

Y. Dai (USTC)

CE13-2.5a

CTC QP for training, CTC QP for testing (the same as CE13-2.4.a)

JVET-N0513

Y. Dai (USTC)

CE13-2.5b

CTC QP for training, CTC QP + 2 for testing

JVET-N0513

Y. Dai (USTC)

CE13-2.5c

CTC QP for training, CTC QP - 2 for testing

JVET-N0513

Y. Dai (USTC)

CE13-2.6a

DBF + NN filter + SAO + ALF, NN filter follows the description in M0872

JVET-

N710

K. Kawamura (KDDI)

CE13-2.6b

NN filter only, NN filter follows the description in M0872

JVET-

N710

K. Kawamura (KDDI)

CE13-2.6c

NN filter + ALF, NN filter follows the description in M0872

JVET-

N710

K. Kawamura (KDDI)

CE13-2.7a

NN filter adaptive on/off in CTU level

JVET-

N710

K. Kawamura (KDDI)

CE13-2.7b

NN filter always on in CTU level

JVET-

N710

K. Kawamura (KDDI)

It comprises 2 categories,

  • CE 13.1: Topics on top of sequence-adaptive (two-pass) methods
  • CE 13.2: Topics on top of sequence-independent (one-pass) methods

Each category will investigate the following problems:

  • The impact of NN filter position in the filter chain.
  • The benefit of the CTU/block level NN filter adaptive on/off.
  • The generalization capability of the NN filter when the test QP is not the same as the training QP.

Specifically, this CE shall use the following these test conditions:

  • Test Condition1: Common Test Conditions (CTC).
  • Test Condition2: CTC, but short test, which only need to test the first intra period.
  • Test Condition3: Based on Test Condition2, test QP=CTC QP +/- 2

The followings are summary tables of the tests in this CE.

CE13 test results with VTM-4.0+ Test Condition1 (CTC)

 

Test#

Y

U

V

EncT

DecT

Y

U

V

EncT

DecT

CTC Full test

CE13-1.1a

-1.36%

-14.96%

-14.91%

100%

142%

CE13-1.2a

-0.58%

-10.91%

-10.69%

103%

127%

CE13-2.1a

-3.48%

-5.18%

-6.77%

142%

38414%

CE13-2.1b

-4.14%

-5.49%

-6.70%

140%

38411%

CE13-2.1c

-4.65%

-6.73%

-7.92%

139%

37956%

CE13-2.2a

-1.52%

-2.12%

-2.73%

107%

4667%

-1.45%

-4.37%

-4.27%

106%

7156%

CE13-2.4a

-0.87%

-0.44%

-0.56%

106%

1912%

-0.49%

-0.23%

-0.33%

124%

468%

CE13-2.6a

CE13 test results with VTM-4.0+Test Condition2

 

Test#

Y

U

V

EncT

DecT

Y

U

V

EncT

DecT

Test Condition

#2

Filter chain

CE13-2.2a

-1.62%

-4.51%

-4.42%

106%

7559%

CE13-2.2b

1.51%

1.21%

1.80%

105%

9910%

CE13-2.2c

-0.65%

-2.18%

-1.60%

105%

9937%

Filter chain

CE13-2.4a

-0.53%

-0.23%

-0.43%

90%

349%

CE13-2.4b

5.73%

6.09%

6.07%

95%

610%

CE13-2.4c

1.14%

2.52%

2.65%

90%

497%

Filter chain

CE13-2.6a

-1.70%

-2.84%

-4.09%

139%

63166%

CE13-2.6b

2.26%

3.86%

2.09%

138%

66303%

CE13-2.6c

-0.80%

-1.21%

-2.59%

138%

66626%

CTU adaptive on/off

CE13-2.7a

-1.70%

-2.84%

-4.09%

139%

63166%

CE13-2.7b

-1.68%

-2.79%

-4.05%

138%

69946%

CE13 test results with VTM-4.0+Test Condition3

 

Test#

Y

U

V

EncT

DecT

Y

U

V

EncT

DecT

Test Condition

#3

generalization capability (QP)

CE13-2.3a

-1.62%

-4.51%

-4.42%

106%

7559%

CE13-2.3b

-1.88%

-4.81%

-4.33%

107%

8008%

CE13-2.3c

-1.37%

-4.57%

-4.04%

105%

6794%

generalization capability (QP)

CE13-2.5a

-0.53%

-0.23%

-0.43%

90%

349%

CE13-2.5b

-0.52%

-0.31%

-0.48%

98%

378%

CE13-2.5c

-0.46%

-0.33%

-0.40%

113%

422%

For the understanding of network, e.g., complexity, the following information is provided.

CE13 test results with complexity analysis for inference stage

Total Conv. Layers

Total FC Layers

Framework

Param. Num

Param. Precision

Mem. P(MB)

Mem. T(MB)(e.g., 4K input)

CE13-1.1

3

0

NA

1506

6-bit/32-bit (I)

0.001282

0.344 (128x128, CTU based)

CE13-1.2

2

0

NA

692x3 (Luma)

402x3 (Chroma)

6-bit/32-bit (I)

0.0028

0.0448(128x128, CTU based)

CE13-2.1

8

0

Caffe

224960

32-bit (F)

0.86

1846.95(192x152, block based)

CE13-2.2

21

0

PyTorch

22371

32-bit (F)

0.09

196.23(pixels num. < 80000, block based)

CE13-2.4

10

0

Tensorflow

7521

32-bit (F)

0.029

10.06(128*128,CTU based)

CE13-2.6

Generally, the performance/complexity tradeoff indicates that the NN technology currently is not mature enough to be included in a standard. The main purpose of this CE had been to provide answers to following questions:

The impact of NN filter position in the filter chain: The results indicate that with a very complex deep network architecture the same or better objective performance can be achieved even if the notwork is run standalone rather than combining it with the conventional filters. This is shown in CE13-2.1, however only for AI configuration. CE13-2.2, CE13-2.4 and CE13-2.6 which are less complex and also run in RA configuration were not able to fully compensate for the gain of conventional filters if those were disabled. It is asked if this would also translate into subjective benefit. An informative session should be prepared during the meeting. Results from experiment 13-2.2 (and if possible 13-2.4) are of particular interest here, because by comparing AI and RA also effects of pure post processing (for AI) or in-loop processing (for RA) could be compared. Best using QP 32 and 37.

The benefit of the CTU/block level NN filter adaptive on/off. This was investigated in CE13-2.7 (a is adaptive, b is always on). The objective gain of adapive is minor, but it would be useful to confirm if it has subjective benefit. Include 13-2.7 in the informal viewing.

The generalization capability of the NN filter when the test QP is not the same as the training QP: This is investigated in CE13-2.3 and CE13-2.5. a is the case where training and test QP are identical. For example, the network in CE13-2.3 was trained in a way that the actual QP is one input parameter of the network. In test a, the network is then run with inputting the true QP, and in cases b and c the actual QP was higher and lower than the one input to the network. However, the difference was not large. Results indicate that the difference is minor, however there is some tendency that the network performs better when the rate is lower (or the actual quality is worse) than in training.

CE13-2.6/7 were only trained with QP37 examples, and the weight by which the difference signal generated by the network was superimposed is decreased for the cases of lower QP.

It can be concluded that there need to be some mechanisms that guarantee that the network is operating differently for different QP values. 13-2.4 and 13-2.5 trained different CNN for every QP value. The effect of CNN filters is probably largest for high QP. Further investigation on this aspect is necessary, also the complexity impact of QP dependent networks should be studied, e.g. the need to load new filter weights.

Another question that should be deeper investigated is the aspect of in-loop or post filter operation. This requires running RA configuration with both in-loop and post processing, and also more to take into account possible subjective impact.

Continue CE. BoG (Y. Li) to review the two CE related contributions, prepare viewing session, and discuss preparation of the upcoming CE

It is mentioned that there is an MPEG standardization activity on neural network compression. Would there be any need for coordination? Future clarification and study of this was recommended.

Decisions
It is mentioned that there is an MPEG standardization activity on neural network compression. Would there be any need for coordination? Future clarification and study of this was recommended.
Citation