Back to Search Document details
30th Meeting: Antalya, TR, April 2023 2023-04-19 05:32
EE1 related: Performance Improvement of AC0052 filter for RPR-based SR in RA configuration
Abstract
This contribution aims to improve the performance of JVET-AC0052 filter for RPR-based super-resolution (SR) in RA configuration. The LMSDANet model for B-frames is trained with 10% bit-rate matching, and TVD and BVI-DVC datasets are used to generate B-frames. In BVI-DVC dataset, 200 2K and 200 4K video sequences are selected for training, while in TVD dataset all video sequences are used. They are compressed in RA configuration with enabled RPR functionality to generate a B-frame training set. At each GOP level, the encoder can adaptively select a scale factor from 1.0x and 2.0x, and the LMSDANet model is applied to the latter case. Compared with VTM-11.0_NNVC-2.0, the proposed CNN filter achieves overall {-2.01% (Y), -1.16% (U), -0.71% (V)} BD-rate gains in all classes, which shows performance improvement in RA configuration.
JVET-AD0170 EE1 related: Performance Improvement of AC0052 filter for RPR-based SR in RA configuration [S. Huang, C. Jung (Xidian Univ.), Y. Liu, M. Li (OPPO)]

This contribution aims to improve the performance of JVET-AC0052 filter for RPR-based super-resolution (SR) in RA configuration. The LMSDANet model for B-frames is trained with 10% bit-rate matching, and TVD and BVI-DVC datasets are used to generate B-frames. In BVI-DVC dataset, 200 2K and 200 4K video sequences are selected for training, while in TVD dataset all video sequences are used. They are compressed in RA configuration with enabled RPR functionality to generate a B-frame training set. At each GOP level, the encoder can adaptively select a scale factor from 1.0x and 2.0x, and the LMSDANet model is applied to the latter case. Compared with VTM-11.0_NNVC-2.0, the proposed CNN filter achieves overall {-2.01% (Y), -1.16% (U), -0.71% (V)} BD-rate gains in all classes, which shows performance improvement in the RA configuration.

Network architecture of the proposed LMSDANet for Y channel

Network architecture of the proposed LMSDANet for U/V channel.

Network architecture of the modified LMSDAB

Network information for the proposed CNN filter testing in inference stage

Network Information in Inference Stage

Mandatory

HW environment:

Intel Core i9 12900k @3.9GHz

Framework:

LibTorch v1.8

Number of GPUs per Task

0

Number of Parameters (Each Model)

luma up-sampling model: 1.13M/model

chroma up-sampling model: 1.18M/model

Total Parameter Number

23.15M

Parameter Precision (Bits)

32 (F)

Memory Parameter (MB)

20 models in total: 92.6 MB

Multiply Accumulate (MAC)

964kMAC/pixel

Optional

Total Conv. Layers

91 for up-sampling the luma, 92 for up-sampling the chroma

Total FC Layers

0

Batch size:

1

Patch size

Whole frame

Network information for the proposed CNN filter testing in training stage

Network Information in Training Stage

Mandatory

GPU Type

GPU: NVIDIA 3090 24GB

Framework:

PyTorch v1.9

Number of GPUs per Task

1

Epoch:

luma:50, chroma:50

Batch size:

32

Training time:

~26h/model for I, ~50h/model for B

Training data information:

DIV2K, BVI-DVC & TVD

Training configurations for generating compressed training data (if different to VTM CTC):

VTM-11.0_NNVC-4.0 , QP {22, 27, 32, 37, 42}

Loss function:

L2

Optional

Number of iterations

Patch size

128128

Learning rate:

1e-4

Optimizer:

ADAM

Preprocessing:

Other information:

BD rates in RA configuration over VTM-11.0_NNVC-4.0.

Random access Main 10

BD-rate Over NNVC-4.0 NnlfOption=0

Sequences

Y-PSNR

U-PSNR

V-PSNR

EncT

DecT

Class A1

-6.14%

-8.60%

-6.40%

70%

#VALUE!

Class A2

-3.23%

1.21%

3.59%

65%

#VALUE!

Class B

0.00%

0.00%

0.00%

100%

#VALUE!

Class C

0.00%

0.00%

0.00%

100%

#VALUE!

Class E

Average on A1 and A2

-4.69%

-3.70%

-1.40%

68%

#VALUE!

Overall

-1.88%

-1.48%

-0.58%

85%

#VALUE!

Performance improvement of the SR model in NNVC was discussed, with a new model for B pictures. There was an intent to further simplify the model until the next meeting; this was agreed not to be in an EE until then.

Decisions
Performance improvement of the SR model in NNVC was discussed, with a new model for B pictures. There was an intent to further simplify the model until the next meeting; this was agreed not to be in an EE until then.
Citation