Search Results for "JVET-AC0052"
Found 5 document(s)
Use number, keyword, author, or MPEG number. Filter by meeting when needed.
JVET-AC0052 EE1-2.4: CNN filter Based on RPR-based SR Combined with GOP Level Adaptive Resolution [S. Huang, C. Jung (Xidian Univ.), Y. Liu, M. Li (OPPO), J. Nam, S. Yoo, J. Lim, S. H. Kim (LGE)]
This contribution reports the EE1-2.4 test results, which is a combination of JVET-AB0093 and JVET-Z0065 test 2.1.1. At each GOP level, the encoder can adaptively select a scale factor from 1.0x and 2.0x and CNN-based super-resolution is utilized for is the latter case. Compared with VTM-11.0-nnvc-2.0, the test 2.4.1 experimental results show {-4.14%(Y), -0.33%(U), -0.22%(V)} and {-3.73%(Y), -3.07%(U), -1.15%(V)} and the test 2.4.2 experimental results show {-4.98%(Y), 0.07%(U), -0.78%(V)} and {-3.73%(Y), -3.07%(U), -1.15%(V)} BD-rate gains on average (A1 and A2 classes), under AI and RA configurations.
As a general comment, the constraint of 10% rate matching and showing PSNR graphs allows much better interpretation of SR results.
It was also commented that it would be beneficial to provide SSIM resuls in SR proposals.
JVET-AC0223 Crosscheck of JVET-AC0052 (EE1-2.4: CNN filter Based on RPR-based SR Combined with GOP Level Adaptive Resolution) [D. Liu (Ericsson)] [late]
JVET-AD0170 EE1 related: Performance Improvement of AC0052 filter for RPR-based SR in RA configuration [S. Huang, C. Jung (Xidian Univ.), Y. Liu, M. Li (OPPO)]
This contribution aims to improve the performance of JVET-AC0052 filter for RPR-based super-resolution (SR) in RA configuration. The LMSDANet model for B-frames is trained with 10% bit-rate matching, and TVD and BVI-DVC datasets are used to generate B-frames. In BVI-DVC dataset, 200 2K and 200 4K video sequences are selected for training, while in TVD dataset all video sequences are used. They are compressed in RA configuration with enabled RPR functionality to generate a B-frame training set. At each GOP level, the encoder can adaptively select a scale factor from 1.0x and 2.0x, and the LMSDANet model is applied to the latter case. Compared with VTM-11.0_NNVC-2.0, the proposed CNN filter achieves overall {-2.01% (Y), -1.16% (U), -0.71% (V)} BD-rate gains in all classes, which shows performance improvement in the RA configuration.
Network architecture of the proposed LMSDANet for Y channel
Network architecture of the proposed LMSDANet for U/V channel.
Network architecture of the modified LMSDAB
Network information for the proposed CNN filter testing in inference stage
Network Information in Inference Stage | ||
Mandatory | HW environment: | Intel Core i9 12900k @3.9GHz |
Framework: | LibTorch v1.8 | |
Number of GPUs per Task | 0 | |
Number of Parameters (Each Model) | luma up-sampling model: 1.13M/model chroma up-sampling model: 1.18M/model | |
Total Parameter Number | 23.15M | |
Parameter Precision (Bits) | 32 (F) | |
Memory Parameter (MB) | 20 models in total: 92.6 MB | |