Abstract not available in document
JVET-T0042 Report of AHG11 Meeting on Neural Network-based Video Coding on 2020-07-30 [S. Liu, E. Alshina, J. Pfaff, M. Wien, P. Wu, Y. Ye]
This document reports the AHG11 meeting on Neural Network-based Video coding, held online on Thursday, 2020-07-30, 14:00h UTC.
- The AHG meeting held a discussion on contribution JVET-T0041, which includes:
- Methodology and materials for training
- Methodology and materials for inference and testing
- Anchor and reporting template
- Following steps were discussed.
1a. Methodology for training
A question was raised about whether we should crosscheck training. A related question was asked that, whether crosscheck of training should be mandated. It was commented that in many (most) cases it is challenging to obtain exact match results from training the same neural network using the same training data at different time, due to the use of randomness in training. It was also commented that some companies may not like to release their exact training methods.
Three possible scenarios were suggested: (1) describe training principles and provide a training example in text; (2) provide a training example in software script; (3) disclose the exact training method and provide training script. It was commented that at least (1) should be included in the contribution, (2) is strongly encouraged, (3) is optional but would be helpful.
A related question was asked: which part should be cross checked and demanded to match? The answer was: Inference stage results should (shall) be cross checked and demanded to match.
Regarding the generation of compressed training data, the latest available version of VTM should be used. Since VTM10 is not yet available, VTM9.3 is recommended to be used for generating the compressed training data. Once VTM10 is available, it is desirable to use VTM10 for generating the compressed training data.
1b. Methodology for inference
A question was asked whether we can (should) request reporting CPU runtime, i.e., request inference to be run on CPU. It was commented that we should allow inference to be run on GPU. In was further commented that at this stage we should probably leave proponents the freedom to report CPU runtime, or GPU runtime, or both, with the information of the platform used.
A suggestion was made to include QP=42 for inference and testing. Some R-D results from contribution JVET-S0246 were presented to support this suggestion. Two captured frames from the same content (“Tango”) compressed by VTM8.0 with QP=37 and 42, were presented (attached in Appendix of this document) as well to support this suggestion. A few experts expressed the interests in including QP=42, while keeping all four VTM CTC QP values (22, 27, 32, 37). A consensus was reached that five QP values (22, 27, 32, 37, 42) will be used for inference and testing. The B-D rate will be calculated using all five QP values, i.e., 5-point B-D rate.
It was also suggested to include the Multi-Scale Structural Similarity Metric (MS-SSIM) in additional to PSNR. It was commented that MS-SSIM is used in many other (research, academia) activities and the implementation is available (https://gitlab.com/standards/HDRTools). On the other hand, it was commented that the current VTM is not optimized for MS-SSIM and thus it may not be fair for the anchor. It was commented that other (better) metrics might come out in near future. A consensus was reached that PSNR will be kept as the mandatory metric. MS-SSIM results may be included in the contribution to provide additional helpful information.
It is encouraged to study MS-SSIM and other quality metrics, for NN-based video coding tool testing as well as in general. Some test sequences may also be considered being updated, but this should be done carefully without rush (coordination with AHG4.)
1c. Anchor
To align with inference and testing, anchor will be generated using five QP values (22, 27, 32, 37, 42), with B-D rate calculated using all five QP values, i.e., 5-point B-D rate calculation will be used.
It was suggested to provide Low Delay 4K anchor results and make it optional for proponents to test and report. The tasks of generating anchor data additional to that is specified in the current VTM CTC, i.e. additional QP point and Low Delay 4K, will be coordinated between AHG11 and AHG3.
It was suggested to compare Low Delay 4K performance between VVC and HEVC.
1d. Tool testing and reporting
A suggestion was made to report “number of iterations” instead of “Epoch”. It was commented that Epoch is a well know and commonly used term in neural network related areas. It was agreed to keep Epoch in the reporting template. Proponents may report “number of iterations” in their contributions for further information. Definitions should be provided if some terms which are not described in the reporting template document are reported and discussed in the contribution.
1e. Training material availability check
A list of candidate training video sequence sets (provided in Appendix A of JVET-T0041) were presented. A subset of this list was recommended for study with higher priority due to higher chance of being available (with copyright permissions.)
It was asked how much some UGC and Youtube sequences are compressed already? The answer was that some of the UGC and Youtube sequences are “clean” while some are heavily compressed. Significant amount of effort seems to be needed for screening these data sets sequence by sequence in order to select the video sequences suitable for the task.
It was asked whether SJTU could make other video sequence sets (than just “Campfire Party”) available. The answer is possible, but we need to check.
It is strongly encouraged that all interested participants look at the materials and express opinions and preferences.
2. Additional aspects
a. Continue examining training materials and checking their availabilities (copyrights)
b. Generate anchor results for additional QP point (QP=42) and Low Delay 4K
c. Update reporting template with additional QP
d. Study and collect information of additional quality metrics
This report was not presented separately from the AHG 11 report.