JVET-AE0008 JVET AHG report: Optimization of encoders and receiving systems for machine analysis of coded video content (AHG8) [C. Hollmann, S. Liu, S. Wang, M. Zhou (AHG chairs)]
Activities
The AHG used the main JVET reflector, jvet@lists.rwth-aachen.de, for email discussion. There were a few emails exchanged on the reflector uner the context of AHG 8 during this period. There were more email exchanges and discussions among co-chairs, editors and interested experts related to technical report preparation, common test conditions and anchor generation, etc. There was one joint meeting with WG 4 held on 2023-06-21. There are 12 input contriubtions related to AHG 8 mandates submitted to this meeting. They are listed in Section 3.
Common Test Conditions
Common test conditions (CTC) for optimization of encoders and receiving systems for machine analysis of coded video content, are summarized in output document JVET-AD2031. This document includes detailed descriptions of test datasets, anchor software and configurations, anchor generation processes, machine task networks used, test and training conditions, evaluation methodologies and metrics. A reporting template in Excel format is enclosed in the document package. This output document package was uploaded on 2023-05-11, with updated anchor results, and is also available at https://vcgit.hhi.fraunhofer.de/jvet-ahg-ofm/ofm-ctc.
Test datasets
Two video datasets, referred as SFU-HW and TVD, are included in the CTC and results on both are expected to be reported. In addition, three image datasets, referred as TVD (image), OpenImageV6 and FLIR, are included in the CTC and can be tested using intra-only configuration. The results on the image datasets are optional.
It was decided in the last JVET meeting to add “Kimono” sequence (back to) the SFU-HW dataset, as a separate optional (O) class, and thus the SFU-HW dataset now consist of 14 test video sequences, as shown in the table below. These sequences can be found on ftp://hevc@mpeg.tnt.uni-hannover.de. The annotations are available at https://data.mendeley.com/datasets/hwm673bv4m/1.
Test sequences in SFU-HW dataset
Class | Sequence name | Frame count | Frame rate | Bit depth | Frames skipped | Frames coded | Intra | RA | LD |
A | Traffic | 150 | 30 | 8 | 117 | 33 | M | M | M |
B | ParkScene | 240 | 24 | 8 | 207 | 33 | M | M | M |
B | Cactus | 500 | 50 | 8 | 403 | 97 | M | M | M |
B | BasketballDrive | 500 | 50 | 8 | 403 | 97 | M | M | M |
B | BQTerrace | 600 | 60 | 8 | 471 | 129 | M | M | M |
C | RaceHorses | 300 | 30 | 8 | 235 | 65 | M | M | M |
C | BQMall | 600 | 60 | 8 | 471 | 129 | M | M | M |
C | PartyScene | 500 | 50 | 8 | 403 | 97 | M | M | M |
C | BasketballDrill | 500 | 50 | 8 | 403 | 97 | M | M | M |
D | RaceHorses | 300 | 30 | 8 | 235 | 65 | M | M | M |
D | BQSquare | 600 | 60 | 8 | 471 | 129 | M | M | M |
D | BlowingBubbles | 500 | 50 | 8 | 403 | 97 | M | M | M |
D | BasketballPass | 500 | 50 | 8 | 403 | 97 | M | M | M |
O | Kimono | 240 | 24 | 8 | 207 | 33 | O | O | O |
Note: M – mandatory; O – optional.
It was decided in the last JVET meeting to replace some of the compressed source video sequences in Tencent Video Dataset (TVD) by their uncompressed versions. The updated source video sequences are renamed, as shown in the next table below, and can be downloaded from https://multimedia.tencent.com/resources/tvd with corresponding annotations.
Test sequences in TVD
Sequence name | Frame count | Frame rate | Bit depth | Frames skipped | Frames coded | Intra | Random access | Low delay |
TVD-01-1 | 3000 | 50 | 8 | 1500 | 500 | M | M | M |
TVD-01-2 | 3000 | 50 | 8 | 2000 | 500 | M | M | M |
TVD-01-3 | 3000 | 50 | 8 | 2500 | 500 | M | M | M |
TVD-02-1 | 636 | 50 | 10 | 0 | 636 | M | M | M |
TVD-03-1 | 2334 | 50 | 10 | 0 | 500 | M | M | M |
TVD-03-2 | 2334 | 50 | 10 | 500 | 500 | M | M | M |
TVD-03-3 | 2334 | 50 | 10 | 1000 | 500 | M | M | M |
As the original sequences are available in .mp4 format, they shall be converted to YUV420 using FFmpeg. The format conversion process was cleaned up as follows:
ffmpeg -i {input.mp4} {output.yuv}
The TVD (image) dataset is an image dataset of 166 images of 1920x1080 resolution that have annotations for object detection and instance segmentation. The dataset with corresponding annotations is available at https://multimedia.tencent.com/resources/tvd.
The OpenImages dataset consists of around 9 million images. A subset of the validation set of its version 6 containing 5000 images are selected for testing object detection in this activity. The dataset with corresponding annotations is available at https://storage.googleapis.com/openimages/web/index.html.
The FLIR dataset used in the WG 4 VCM activities is a dataset consisting of 300 infrared images. The images, annotations and the fine-tuned model for thermal images can be found on the MPEG content repository (https://content.mpeg.expert/data/).
More detailed information about these test datasets can be found in JVET-AD2031.
Anchor software and configuration
Version 12.0 of the VTM software is used for generating anchor results. The VTM software is available at https://vcgit.hhi.fraunhofer.de/jvet/VVCSoftware_VTM/.
In addition, version 4.2.2 of the FFmpeg is used for the format conversion. The FFmpeg software is available at https://ffmpeg.org/releases.
Three test conditions are used for experimenting encoder and receiving system optimizations on video coding for machine consumptions, i.e., random-access, low-delay and all-intra. The default configuration files provided with the VTM software are used for anchor generation.
- “Random access” (RA): encoder_randomaccess_vtm.cfg
- “Low delay” (LD): encoder_lowdelay_vtm.cfg
- “All Intra” (AI): encoder_intra_vtm.cfg
It was decided in the last JVET meeting that temporal subsampling is disabled for “All Intra” configuration for this activity, thus the following line
TemporalSubsampleRatio : 8
in “encoder_intra_vtm.cfg” needs to be modified to
TemporalSubsampleRatio : 1
A subset of these configurations might be used for a particular experiment. However, as these test conditions are for video coding experiments, results for at least one of the random access or low delay configurations are expected to be provided.
Anchor generation
The anchor generation process is illustrated in the figure below. The task network Faster R-CNN X101-FPN which is part of Detectron2 (https://github.com/facebookresearch/detectron2) is used to evaluate object detection (SFU-HW dataset and image datasets), the model is available here. The task network JDE-1088x608 which is part of the Towards Realtime MOT framework (https://github.com/Zhongdao/Towards-Realtime-MOT) is used to evaluate object tracking (TVD), the model is available at this Google Drive link. More details about anchor generation process and these machine task networks are referred to JVET-AD2031.
Figure: Anchor generation pipeline
Anchor results were generated following the CTC described in JVET-AD2031 and crosschecked by Ericsson and Tencent.
Evaluation and reporting
Proposed technologies are evaluated based on their compression performance, measured by bitrate, PSNR, mAP and MOTA, as well as encoding and decoding runtime to reflect the complexity of the proposed technology, to some extent. Definitions and detailed descriptions of these metrics can be found in JVET-AD2031.
It is noted that the mean Average Precision (mAP) is used to measure object detection performance, and Multiple Object Tracking Accuracy (MOTA) is used to measure object tracking performance. For the purpose of reporting encoding and decoding running times, the anchor and proposal should be simulated on the same platform, e.g. similar CPU and GPU configuration, to have reliable time comparison. Parallel encoding and decoding as described in JVET-B0036 may be applied for RA configurations.
In addition, relevant inference and training information should be reported if the proposed technology consists of learning-based components, such as network structure, number of network parameters, precision of network parameters, number of multiply–accumulate operations (MAC) per pixel, patch size, batch size, epoch, training time, training datasets, lost function, number of iterations, optimizer, and any pre- and post-processings used. More details can be found in JVET-AD2031.
A reporting template in Excel file format has been prepared to illustrate results following CTC and evaluation methodology described in JVET-AD2031. It is enclosed in JVET-AD2031 package.
Reporting template summary
Technical Report
The draft 2 of the technical report (TR) has been prepared and uploaded as JVET-AD2030 based on discussions in the last JVET meeting on 2023-06-07. More descriptions were added to use cases and applications, pre-processing technologies and encoding technologies. Two software implementation examples were included in the TR draft 2 as Annex.
Git Management
AHG 8 related software and documents can be found at https://vcgit.hhi.fraunhofer.de/jvet-ahg-ofm. This repository contains two projects, one (https://vcgit.hhi.fraunhofer.de/jvet-ahg-ofm/ofm-ctc) containing instrucitons and information for conducting experiements and evaluation, such as evaluation scripts, machine task networks, CTC and reporting template with anchor results, while the other (https://vcgit.hhi.fraunhofer.de/jvet-ahg-ofm/vtm-ofm) containing implementation examples. Two software example packages have been uploaded to the git in separate branches:
- JVET-AB0275: a region of interest-based method that uses adaptive QP to reduce the quality in background areas
- JVET-AC0086: a method that uses a pre-analysis to perform content adaptive machine vision oriented preprocessing
AHG Coordination
Collaborative discussions were held between this AHG and the VCM AHG in WG4 to synchronize common test conditions including test datasets, anchor QP points, encoder configurations, etc. A joint meeting was held on 2023-06-21 0500UTC to discuss metrics, existing issues and potential solutions for machine task performance evaluation. Two contributions were presented and discussed during this meeting:
m63676 - [VCM] Proposed performance evaluation method for VCM CE1 (ETRI, Konkuk Univ., Myongji Univ.)
This contribution proposes a performance evaluation method for VCM CE1.
Regarding the performance evaluation for VCM CE1, there are some difficulties as follows:
- The current BD-rate metric provided by the reporting template is not able to produce a value if either the anchor or a proposal is non-monotonic.
- To compensate for this limitation, the reporting template requires submission of BD-rate for multiple subcases (combination of 4 or 5 points out of 6 total points), but it is still difficult to make a direct comparison between proposals since the metric is not singular.
The following rules were proposed as the performance evaluation method for VCM CE1:
- For each test case, a proponent should submit 4 points that correspond to the high 4 rate points of the VTM anchor.
- Only if the result at the high 4 rate points is non-monotonic, then it is allowed to adjust QP in a range of [-2, +2] to produce a monotonic result.
- If a monotonic result cannot be obtained even with the adjustment of QP in the range, then it is not considered as a candidate technology in this round of CE1 evaluation.
- Evaluate the proposed technologies using the summed scores based on the class-wise (SFU) or sequence-wise (TVD) BD-rate rankings of the proposals.
JVET-AE0107 / m63692 - [VCM] Improvements of the BD-rate model using monotonic curve-fitting method (Tencent)
This contribution provides improvements of the BD-rate model used in current Excel templates. Specifically, the contribution proposes a monotonic fitted curve to replace the PCHIP (Piecewise Cubic Hermit Interpolation) curve for BD-rate and BD-metric calculation, such as BD-PSNR, BD-mAP, BD-MOTA, etc. The fitted curve guarantees monotonicity even if the input points are nonmonotonic, thus makes it possible to calculate BD-rate and BD-metric values in various testing conditions.
This contribution was also submitted to JVET (AHG 8) as document JVET-AE0107.
Input contributions
There were 12 input contributions related to AHG 8 mandates. They are listed below.
Report | ||
JVET AHG report: Optimization of encoders and receiving systems for machine analysis of coded video content (AHG8) | C. Hollmann, S. Liu, S. Wang, M. Zhou (AHG chairs) | |
Proposal | ||
AHG8/AHG9: Neural-network post-filter regions SEI message | T. Chujoh, Y. Yasugi, T. Ikai (Sharp) | |
AHG8/AHG9: Signalling encoder preprocessing and human / machine viewing indications | C. Kim, D. Gwak, Hendry, J. Lim, S. Kim (LGE), M. M. Hannuksela, F. Cricri, H. Zhang (Nokia) | |
AHG8/AHG9: Source picture timing information SEI message | S. McCarthy, G. J. Sullivan, P. Yin (Dolby) | |
[AHG8] De-noising filter as pre-processing for machine task | C. Kim, D. Gwak, J. Lim (LGE) | |
AHG8/AHG9: On machine vision indication | J. Gao, H.-B. Teo, C.-S. Lim, K. Abe (Panasonic) | |
AHG8/AHG9: proposed changes to the candidate new object mask information SEI message | P. de Lagrange, E. François, D. Doyen (InterDigital), J. Chen, S. Wang, Y. Ye (Alibaba) | |
[AHG8] Study on using different VTM versions | C. Hollmann, J. Ström (Ericsson) | |
[AHG8] NNPF and post-filter hint SEI messages for the technical report | C. Hollmann, M. Pettersson, R. Sjöberg, M. Damghanian (Ericsson) | |
AHG8: Improvements of the BD-rate model using monotonic curve-fitting method | H. Wang, X. Pan, Z. Liu, X. Xu, S. Liu (Tencent) | |
AHG8: A spatial resampling algorithm and an exemplar software implementation | S. Wang, B. Li, J. Chen, Y. Ye (Alibaba), S. Wang (CityU) | |
Crosscheck | ||
AHG8: Crosscheck of JVET-AE0107 | Honglei Zhang (Nokia) | |
* Late.
Recommendations
The AHG recommended to:
- Review all input contributions.
- Discuss and refine test conditions, evalution and reporting procedures.
- Discuss the non-monotonic issue and potential solutions.
- Continue investigating non-normative technologies and their suitability for machine analysis applications.
- Continue developing draft technical report on optimization of encoders and receiving systems for machine analysis of coded video content.
- Continue collecting new test materials.
It was pointed out that a potential timeline for the technical report should be discussed. This however may also relate to the level of completeness in terms of using video compression standards for a variety of machine usage applications, which somewhat relies on the availability of test material of sufficient variety.