Back to Search Document details
21st Meeting: by teleconference, January 2021 2021-01-07 23:15
AHG11: Neural Network-based Super Resolution
Abstract
This contribution studied the performance of applying a Neural-Network based super-resolution used as upsampling filter in the context of VVC RPR. Prior to encoding, a given picture is downsampled by a factor of 2x using the inbuilt RPR mechanism of VTM11. PSNR of the coded frame is computed by calculating the MSE between the original picture and the up-sampled version of the decoded picture. The upsampled picture is generated by the Neural Network-based up-sampling filter instead of the existing VTM up-sampling filter.
JVET-U0099 AHG11: Neural Network-based Super Resolution [A. M. Kotra, K. Reuzé, J. Chen, H. Wang, M. Karczewicz, J. Li (Qualcomm)]

This contribution studied the performance of applying a Neural-Network based super-resolution filter used as upsampling filter in the context of VVC RPR. Prior to encoding, a given picture is downsampled by a factor of 2x using the in-built RPR mechanism of VTM11. The PSNR of the coded frame is computed by calculating the MSE between the original picture and the upsampled version of the decoded picture. The upsampled picture is generated by the Neural Network-based up-sampling filter instead of the existing VTM upsampling filter. On average, for Class A1 sequences, Luma BD-Rate gains of 5.74% and 8.79% for AI and RA configurations, respectively, were reported.

Number of parameters per model was ~1.3 million, and two models were used in this contribution for different bit rate ranges.

Training was performed in two ways: sequences coded in the original resolution, and sequences coded in the downsampled resolution.

NN training takes YUV 4:4:4 as input by repeating every other chroma sample in the YUV 4:2:0 domain.

CTC QP values 22 to 42 were used in testing. For all cases, VVC RPR filters were used in downsampling, and the reported decoding time does not include upsampling.

The following table shows the performance of half resolution coding compared to full-resolution coding with VTM-11.0, with upsampling performed using RPR filters.

Random access Main 10

Over VTM-11.0 (QP 22,27,32,37,42)

Y

U

V

EncT

DecT

Cass A1 4K

Tango2

-4.41%

0.19%

10.52%

50%

33%

FoodMarket4

-0.57%

10.37%

9.06%

Campfire

-2.56%

95.98%

35.38%

Class A2 4K

CatRobot1

24.07%

44.74%

68.09%

40%

30%

DaylightRoad2

42.46%

18.66%

41.19%

ParkRunning3

-0.06%

297.20%

146.30%

The following table shows coding performance of half-resolution coding compared to full-resolution coding with VTM-11.0, with upsampling performed using NN filters trained using method 1.

Random access Main 10

Over VTM-11.0 (QP 22,27,32,37,42)

Y

U

V

EncT

DecT

Cass A1 4K

Tango2

-7.48%

-11.02%

-1.39%

92%

32%

FoodMarket4

-4.03%

7.27%

5.77%

Campfire

-12.68%

79.71%

-3.04%

Class A2 4K

CatRobot1

11.86%

21.67%

32.83%

99%

37%

DaylightRoad2

25.36%

5.73%

21.00%

ParkRunning3

-4.87%

194.75%

89.13%

The following table shows coding performance of half-resolution coding compared to full-resolution coding with VTM-11.0, with upsampling performed using NN filters trained using method 2.

Random access Main 10

Over VTM-11.0 (QP 22,27,32,37,42)

Y

U

V

EncT

DecT

Cass A1 4K

Tango2

-7.71%

-9.67%

4.70%

90%

31%

FoodMarket4

-5.20%

4.88%

4.81%

Campfire

-13.44%

85.12%

6.07%

Class A2 4K

CatRobot1

7.85%

23.33%

40.51%

74%

31%

DaylightRoad2

21.33%

5.31%

23.25

ParkRunning3

-7.44%

197.45%

93.99%

It was commented that training method 2 seems to be better than training method 1.

It was noted that the NN in JVET-U0053 had fewer parameters.

It was commented that visual quality of video coded at QP 47 (as done in JVET-U0053) should be checked.

It was commented that for some of the sequences where BD rate shows a gain, e.g. Campfire, crossing of the luma RD curves between the reference and the tested method can be observed, where the tested method is better at lower rates but worse at higher rates.

Further study of NN upsampling filters was encouraged, perhaps in the context of an EE. See also the notes under JVET-U0091.

Other coding technologies (4)

Contributions in this area were discussed in Session 15 at 0815 UTC on Tuesday 12 January 2021 (chaired by JRO and GJS).

Decisions
Contributions in this area were discussed in Session 15 at 0815 UTC on Tuesday 12 January 2021 (chaired by JRO and GJS).
Citation