Search Results for "JVET-AC0052"

Found 5 document(s)

Search documents

Use number, keyword, author, or MPEG number. Filter by meeting when needed.

29th Meeting: by teleconference, DE, January 2023 2023-01-06 13:35
Abstract
This contribution reports the EE1-2.4 test results, which is a combination of JVET-AB0093 and JVET-Z0065 test 2.1.1. At each GOP level, the encoder can adaptively select a scale factor from 1.0x and 2.0x and CNN-based super-resolution is utilized for is the latter case. Compared with VTM-11.0-nnvc-2.0, the test 2.4.1 experimental results show {-4.14%(Y), -0.33%(U), -0.22%(V)} and {-3.73%(Y), -3.07%(U), -1.15%(V)} and the test 2.4.2 experimental results show {-4.98%(Y), 0.07%(U), -0.78%(V)} and {-3.73%(Y), -3.07%(U), -1.15%(V)} BD-rate gains on average (A1 and A2 classes), under AI and RA configurations.
JVET-AC0052 EE1-2.4: CNN filter Based on RPR-based SR Combined with GOP Level Adaptive Resolution [S. Huang, C. Jung (Xidian Univ.), Y. Liu, M. Li (OPPO), J. Nam, S. Yoo, J. Lim, S. H. Kim (LGE)]

This contribution reports the EE1-2.4 test results, which is a combination of JVET-AB0093 and JVET-Z0065 test 2.1.1. At each GOP level, the encoder can adaptively select a scale factor from 1.0x and 2.0x and CNN-based super-resolution is utilized for is the latter case. Compared with VTM-11.0-nnvc-2.0, the test 2.4.1 experimental results show {-4.14%(Y), -0.33%(U), -0.22%(V)} and {-3.73%(Y), -3.07%(U), -1.15%(V)} and the test 2.4.2 experimental results show {-4.98%(Y), 0.07%(U), -0.78%(V)} and {-3.73%(Y), -3.07%(U), -1.15%(V)} BD-rate gains on average (A1 and A2 classes), under AI and RA configurations.

As a general comment, the constraint of 10% rate matching and showing PSNR graphs allows much better interpretation of SR results.

It was also commented that it would be beneficial to provide SSIM resuls in SR proposals.

Decisions
It was also commented that it would be beneficial to provide SSIM resuls in SR proposals.
Citation
29th Meeting: by teleconference, DE, January 2023 2023-01-10 15:51
Authors: Du Liu Ericsson
Abstract
This contribution reports the inference crosscheck results for JVET-AC0052 (JVET-AB-EE1-2.4), which presents neural network-based super resolution. One crosscheck test was carried out to for the super resolution network without rate matching. It is reported that for both RA and AI, the crosscheck results can match the results from JVET-AC0052, with a minor difference which is likely due to floating point calculation.
JVET-AC0223 Crosscheck of JVET-AC0052 (EE1-2.4: CNN filter Based on RPR-based SR Combined with GOP Level Adaptive Resolution) [D. Liu (Ericsson)] [late]
References:
Decisions
JVET-AC0223 Crosscheck of JVET-AC0052 (EE1-2.4: CNN filter Based on RPR-based SR Combined with GOP Level Adaptive Resolution) [D. Liu (Ericsson)] [late]
Citation
30th Meeting: Antalya, TR, April 2023 2023-04-19 05:32
Abstract
This contribution aims to improve the performance of JVET-AC0052 filter for RPR-based super-resolution (SR) in RA configuration. The LMSDANet model for B-frames is trained with 10% bit-rate matching, and TVD and BVI-DVC datasets are used to generate B-frames. In BVI-DVC dataset, 200 2K and 200 4K video sequences are selected for training, while in TVD dataset all video sequences are used. They are compressed in RA configuration with enabled RPR functionality to generate a B-frame training set. At each GOP level, the encoder can adaptively select a scale factor from 1.0x and 2.0x, and the LMSDANet model is applied to the latter case. Compared with VTM-11.0_NNVC-2.0, the proposed CNN filter achieves overall {-2.01% (Y), -1.16% (U), -0.71% (V)} BD-rate gains in all classes, which shows performance improvement in RA configuration.
JVET-AD0170 EE1 related: Performance Improvement of AC0052 filter for RPR-based SR in RA configuration [S. Huang, C. Jung (Xidian Univ.), Y. Liu, M. Li (OPPO)]

This contribution aims to improve the performance of JVET-AC0052 filter for RPR-based super-resolution (SR) in RA configuration. The LMSDANet model for B-frames is trained with 10% bit-rate matching, and TVD and BVI-DVC datasets are used to generate B-frames. In BVI-DVC dataset, 200 2K and 200 4K video sequences are selected for training, while in TVD dataset all video sequences are used. They are compressed in RA configuration with enabled RPR functionality to generate a B-frame training set. At each GOP level, the encoder can adaptively select a scale factor from 1.0x and 2.0x, and the LMSDANet model is applied to the latter case. Compared with VTM-11.0_NNVC-2.0, the proposed CNN filter achieves overall {-2.01% (Y), -1.16% (U), -0.71% (V)} BD-rate gains in all classes, which shows performance improvement in the RA configuration.

Network architecture of the proposed LMSDANet for Y channel

Network architecture of the proposed LMSDANet for U/V channel.

Network architecture of the modified LMSDAB

Network information for the proposed CNN filter testing in inference stage

Network Information in Inference Stage

Mandatory

HW environment:

Intel Core i9 12900k @3.9GHz

Framework:

LibTorch v1.8

Number of GPUs per Task

0

Number of Parameters (Each Model)

luma up-sampling model: 1.13M/model

chroma up-sampling model: 1.18M/model

Total Parameter Number

23.15M

Parameter Precision (Bits)

32 (F)

Memory Parameter (MB)

20 models in total: 92.6 MB

Decisions
Performance improvement of the SR model in NNVC was discussed, with a new model for B pictures. There was an intent to further simplify the model until the next meeting; this was agreed not to be in an EE until then.
Citation
29th Meeting: by teleconference, DE, January 2023 2023-01-18 18:33
Abstract
This document contains the report of the BoG on non-normative optimization for machine. The BoG met on Wednesday 18 January 2023 at 1520-1720UTC, discussing subjects related to non-normative optimization for machine consumption of coded video content, including:
JVET-AC0351 BoG report on non-normative optimization for machine [C. Hollmann, S. Liu (BoG coordinators)] This document contains the report of the BoG on non-normative optimization for machine. The BoG met on Wednesday 18 January 2023 at 1520-1720 UTC, discussing subjects related to non-normative optimization for machine consumption of coded video content, including: Common test conditions Test sequences RA/LD/AI configurations QP points and BD-rate calculation Cross-check procedure Reference software Where to setup the repository What to be included in the initial version reference software Other aspects The BoG recommended to: Align common test conditions with WG 4 VCM (including test sequences, test configurations, QP points, etc.) Further discuss on QP point selection for calculation of the BD-rate used for tool evaluation. Align crosscheck procedure with WG 4 VCM. Further discuss on...
Decisions
It was planned to produce output documents for draft 1 of TR (preliminary WD of WG 5), and CTC.
adopted
Adopt JVET-AC0189 (in CTC, set the PPS flag to 1 for class F and class TGM)
Citation
29th Meeting: by teleconference, DE, January 2023 2023-01-11 17:20
Abstract
This document summarizes the activities of AHG11: Neural network-based video coding between the 28th meeting (20 – 28 October 2022) held in Mainz, DE, and the 29th meeting (11 – 20 January 2023) held by teleconference.
JVET-AC0011 JVET AHG report: Neural network-based video coding (AHG11) [E. Alshina, S. Liu, A. Segall (co-chairs), F. Galpin, J. Li, T. Shao, H. Wang, Z. Wang, M. Wien, P. Wu (vice-chairs)] Common Test Conditions The AHG released revised common test conditions as decided at the 28th meeting. The final version was uploaded as document JVET-AB2016 on December 24, 2022. Anchors for the NN-based video coding activity were made available on the Git repository used for the AHG activity: https://vcgit.hhi.fraunhofer.de/jvet-ahg-nnvc/nnvc-ctc/-/tree/master. EE Coordination The AHG finalized, conducted and discussed the EE on NN based video coding. The final version of the EE description was uploaded to the document repository on November 14, 2022. A summary report for the EE is available at this meeting as: JVET-AC0023 [[E. Alshina, F. Galpin, Y. Li, M. Santamaria, H. Wang, L. Wang, Z. Xie] EE...
Decisions
The AHG recommended to: Review all input contributions. Continue investigating neural network-based video coding tools, including coding performance and complexity.
Citation
New Search