JVET-AC0015 JVET AHG report: Optimization of encoders and receiving systems for machine analysis of coded video content (AHG15) [C. Hollmann, S. Liu, S. Wang, M. Zhou (AHG chairs)]
There were nine input contriubtions related to AHG15 mandates submitted to this meeting, as described below.
Common Test Conditions
Draft common test conditions (CTC) for AHG15: Optimization of encoders and receiving systems for machine analysis of coded video content, are summarized in document JVET-AC0073. This document includes detailed descriptions of test datasets, anchor software and configurations, anchor generation processes, machine task networks used, test and training conditions, evaluation methodologies and metrics. A reporting template in Excel format and a set of configuration files are also included in the same contribution package.
Test datasets
Two video datasets, referred as SFU-HW and TVD, are included in the CTC and results on both are expected to be reported. In addition, three image datasets, referred as TVD (image), OpenImageV6 and FLIR, are included in the CTC and can be tested using intra-only configuration. The results on the image datasets are optional.
The SFU-HW (SFU-HW-objects-v1) dataset is a video dataset consisting of 17 sequences which are known from previous standardization efforts in JCT-VC and JVET. The sequences can be found on ftp://hevc@mpeg.tnt.uni-hannover.de. The annotations are available at https://data.mendeley.com/datasets/hwm673bv4m/1.
The Tencent Video Dataset (TVD) is a video dataset consisting of three sequences in 1920x1080 resolution used for object tracking with lengths of 3000, 636 and 2334 frames, respectively. An overview of the three sequences can be found in Table 3 of JVET-AC0073. The dataset with corresponding annotations is available at https://multimedia.tencent.com/resources/tvd.
The TVD (image) dataset is an image dataset of 166 images of 1920x1080 resolution that have annotations for object detection and instance segmentation. The dataset with corresponding annotations is available at https://multimedia.tencent.com/resources/tvd.
The OpenImages dataset consists of around 9 million images. A subset of the validation set of its version 6 containing 5000 images are selected for testing object detection in this activity. The dataset with corresponding annotations is available at https://storage.googleapis.com/openimages/web/index.html.
The FLIR dataset used in the VCM group is a dataset consisting of 300 infrared images. The images, annotations and the fine-tuned model for thermal images can be found on the MPEG content repository (https://content.mpeg.expert/data/).
More detailed information about these test datasets can be found in JVET-AC0073.
Anchor software
Version 12.0 of the VTM software is used for generating anchor results. The VTM software is available at https://vcgit.hhi.fraunhofer.de/jvet/VVCSoftware_VTM/.
In additiona, version 4.2.2 of the FFmpeg is used for the format conversion. The FFmpeg software is available at https://ffmpeg.org/releases.
Anchor configuration
Three test conditions are used for experimenting encoder and receiving system optimizations on video coding for machine consumptions, i.e., random-access, low-delay and all-intra. The default configuration files provided with the VTM software are used for anchor generation.
- “All Intra” (AI): encoder_intra_vtm.cfg
- “Random access” (RA): encoder_randomaccess_vtm.cfg
- “Low delay” (LD): encoder_lowdelay_vtm.cfg
A subset of these test conditions might be used for a particular experiment. However, as these test conditions are for video coding experiments, results for at least one of the random access or low delay configurations are expected to be provided.
Anchor generation pipeline
The process to generate anchor results is described in the figure above. Detailed description about anchor generation process is referred to JVET-AC0073. It is worthwhile to mention that for machine consumption, machine task networks are applied to compressed and reconstructed videos to obtain corresponding machine task performance. The task network Faster R-CNN X101-FPN which is part of Detectron2 (https://github.com/facebookresearch/detectron2) is used to evaluate object detection (SFU-HW dataset and image dataset), the model is available here. The task network JDE-1088x608 which is part of the Towards Realtime MOT framework (https://github.com/Zhongdao/Towards-Realtime-MOT) is used to evaluate object tracking (TVD dataset), the model is available at https://drive.google.com/open?id=1nlnuYfGNuHWZztQHXwVZSL_FvfE551pA. More details about these machine task networks and how they are used for AHG15 anchor generation are in JVET-AC0073.
Evaluation methodology and metrics
Proposed technologies are evaluated based on their compression performance, measured by bitrate, PSNR, mAP and MOTA, as well as encoding and decoding runtime to reflect the complexity of the proposed technology, to some extent. Definitions and detailed descriptions of these metrics can be found in JVET-AC0073. It is noted that the mean Average Precision (mAP) is used to measure object detection performance, and Multiple Object Tracking Accuracy (MOTA) is used to measure object tracking performance.
For the purpose of reporting encoding and decoding running times, the anchor and proposal should be simulated on the same platform, e.g. similar CPU and GPU configuration, to have reliable time comparison. Parallel encoding and decoding as described in JVET-B0036 may be applied for RA configurations.
In addition, relevant inference and training information should be reported if the proposed technology consists of learning-based components, such as network structure, number of network parameters, precision of network parameters, number of multiply–accumulate operations (MAC) per pixel, patch size, batch size, epoch, training time, training datasets, lost function, number of iterations, optimizer, and any pre- and post-processings used. More details can be found in JVET-AC0073.
Reporting template
A reporting template in Excel file format had been prepared to illustrate results following CTC and evaluation methodology described in JVET-AC0073. It was attached in the JVET-AC0073 package and has the format shown in the figure below.
Reporting template summary
Anchor Results
Anchor results were generated following the CTC described in JVET-AC0073 and crosschecked by Tencent and Ericsson.
Technical Report Preparation
An initial draft of technical report has been prepared and uploaded as JVET-AC0049.
AHG Coordination
Offline discussions were conducted about harmonization and difference on testing materials and conditions between this AHG and the VCM AHG in WG4.
Input contributions
There were nine input contriubtions related to AHG15 mandates. They are listed below.
Report | ||
JVET AHG report: Optimization of encoders and receiving systems for machine analysis of coded video content (AHG15) | C. Hollmann, S. Liu, S. Wang, M. Zhou (AHG chairs) | |
Proposal | ||
[AHG15] Draft technical report on optimizations for encoders and receiving systems for machine analysis of coded video content | ||
AHG15: On common test conditions for optimization of encoders and receiving systems for machine analysis of coded video content | ||
AHG9/AHG15: On the NNPFC SEI message for machine analysis | M. M. Hannuksela, F. Cricri, J. I. Ahonen, H. Zhang (Nokia) | |
AHG9/AHG15: On bitstreams that are potentially suboptimal for user viewing | M. M. Hannuksela, F. Cricri, H. Zhang (Nokia) | |
[AHG15] Effect of the perceptual QP adaptation (QPA) on machine task performance | C. Kim, D. Gwak, J. Lim (LGE) | |
AHG15: Feature based Encoder-only algorithms for the Video Coding for Machines | ||
AHG15: Investigations on the common test conditions of Video Coding for Machines (VCM) | ||
Information | ||
[AHG15] Information about datasets used in VCM | ||
Recommendations
The AHG recommended to:
- Review all input contributions.
- Discuss and refine test conditions, evalution and reporting procedures.
- Discuss and refine anchor generation processes and results.
- Discuss existing and continue collecting new test materials.
- Discuss and establish software development and experiment environment.
- Continue investigating non-normative technologies and their suitability for machine analysis applications.
- Continue developing a draft technical report on optimization of encoders and receiving systems for machine analysis of coded video content.
Project development (33)
Deployment and advertisement of standards (2)
Contributions in this area were discussed in session 27 at 2145–XXXX UTC on Thursday 19 Jan. 2023 (chaired by JRO).