JVET-AD0170 EE1 related: Performance Improvement of AC0052 filter for RPR-based SR in RA configuration [S. Huang, C. Jung (Xidian Univ.), Y. Liu, M. Li (OPPO)]
This contribution aims to improve the performance of JVET-AC0052 filter for RPR-based super-resolution (SR) in RA configuration. The LMSDANet model for B-frames is trained with 10% bit-rate matching, and TVD and BVI-DVC datasets are used to generate B-frames. In BVI-DVC dataset, 200 2K and 200 4K video sequences are selected for training, while in TVD dataset all video sequences are used. They are compressed in RA configuration with enabled RPR functionality to generate a B-frame training set. At each GOP level, the encoder can adaptively select a scale factor from 1.0x and 2.0x, and the LMSDANet model is applied to the latter case. Compared with VTM-11.0_NNVC-2.0, the proposed CNN filter achieves overall {-2.01% (Y), -1.16% (U), -0.71% (V)} BD-rate gains in all classes, which shows performance improvement in the RA configuration.
Network architecture of the proposed LMSDANet for Y channel
Network architecture of the proposed LMSDANet for U/V channel.
Network architecture of the modified LMSDAB
Network information for the proposed CNN filter testing in inference stage
Network Information in Inference Stage | ||
Mandatory | HW environment: | Intel Core i9 12900k @3.9GHz |
Framework: | LibTorch v1.8 | |
Number of GPUs per Task | 0 | |
Number of Parameters (Each Model) | luma up-sampling model: 1.13M/model chroma up-sampling model: 1.18M/model | |
Total Parameter Number | 23.15M | |
Parameter Precision (Bits) | 32 (F) | |
Memory Parameter (MB) | 20 models in total: 92.6 MB | |
Multiply Accumulate (MAC) | 964kMAC/pixel | |
Optional | Total Conv. Layers | 91 for up-sampling the luma, 92 for up-sampling the chroma |
Total FC Layers | 0 | |
Batch size: | 1 | |
Patch size | Whole frame | |
Network information for the proposed CNN filter testing in training stage
Network Information in Training Stage | ||
Mandatory | GPU Type | GPU: NVIDIA 3090 24GB |
Framework: | PyTorch v1.9 | |
Number of GPUs per Task | 1 | |
Epoch: | luma:50, chroma:50 | |
Batch size: | 32 | |
Training time: | ~26h/model for I, ~50h/model for B | |
Training data information: | DIV2K, BVI-DVC & TVD | |
Training configurations for generating compressed training data (if different to VTM CTC): | VTM-11.0_NNVC-4.0 , QP {22, 27, 32, 37, 42} | |
Loss function: | L2 | |
Optional | Number of iterations | |
Patch size | 128128 | |
Learning rate: | 1e-4 | |
Optimizer: | ADAM | |
Preprocessing: | ||
Other information: | ||
BD rates in RA configuration over VTM-11.0_NNVC-4.0.
Random access Main 10 | |||||
BD-rate Over NNVC-4.0 NnlfOption=0 | |||||
Sequences | Y-PSNR | U-PSNR | V-PSNR | EncT | DecT |
Class A1 | -6.14% | -8.60% | -6.40% | 70% | #VALUE! |
Class A2 | -3.23% | 1.21% | 3.59% | 65% | #VALUE! |
Class B | 0.00% | 0.00% | 0.00% | 100% | #VALUE! |
Class C | 0.00% | 0.00% | 0.00% | 100% | #VALUE! |
Class E | |||||
Average on A1 and A2 | -4.69% | -3.70% | -1.40% | 68% | #VALUE! |
Overall | 85% | #VALUE! | |||
Performance improvement of the SR model in NNVC was discussed, with a new model for B pictures. There was an intent to further simplify the model until the next meeting; this was agreed not to be in an EE until then.