Back to Search Document details
31st Meeting: Geneva, CH, July 2023 2023-07-11 09:54
JVET AHG report: Optimization of encoders and receiving systems for machine analysis of coded video content (AHG8)
Abstract
This document summarizes the activities of AHG 8: Optimization of encoders and receiving systems for machine analysis of coded video content between the 30th meeting (21 – 28 April 2023) held in Antalya, TR and the 31st meeting (11 – 19 July 2023) held in Geneva, CH.
JVET-AE0008 JVET AHG report: Optimization of encoders and receiving systems for machine analysis of coded video content (AHG8) [C. Hollmann, S. Liu, S. Wang, M. Zhou (AHG chairs)]

Activities

The AHG used the main JVET reflector, jvet@lists.rwth-aachen.de, for email discussion. There were a few emails exchanged on the reflector uner the context of AHG 8 during this period. There were more email exchanges and discussions among co-chairs, editors and interested experts related to technical report preparation, common test conditions and anchor generation, etc. There was one joint meeting with WG 4 held on 2023-06-21. There are 12 input contriubtions related to AHG 8 mandates submitted to this meeting. They are listed in Section 3.

Common Test Conditions

Common test conditions (CTC) for optimization of encoders and receiving systems for machine analysis of coded video content, are summarized in output document JVET-AD2031. This document includes detailed descriptions of test datasets, anchor software and configurations, anchor generation processes, machine task networks used, test and training conditions, evaluation methodologies and metrics. A reporting template in Excel format is enclosed in the document package. This output document package was uploaded on 2023-05-11, with updated anchor results, and is also available at https://vcgit.hhi.fraunhofer.de/jvet-ahg-ofm/ofm-ctc.

Test datasets

Two video datasets, referred as SFU-HW and TVD, are included in the CTC and results on both are expected to be reported. In addition, three image datasets, referred as TVD (image), OpenImageV6 and FLIR, are included in the CTC and can be tested using intra-only configuration. The results on the image datasets are optional.

It was decided in the last JVET meeting to add “Kimono” sequence (back to) the SFU-HW dataset, as a separate optional (O) class, and thus the SFU-HW dataset now consist of 14 test video sequences, as shown in the table below. These sequences can be found on ftp://hevc@mpeg.tnt.uni-hannover.de. The annotations are available at https://data.mendeley.com/datasets/hwm673bv4m/1.

Test sequences in SFU-HW dataset

Class

Sequence name

Frame count

Frame rate

Bit depth

Frames skipped

Frames coded

Intra

RA

LD

A

Traffic

150

30

8

117

33

M

M

M

B

ParkScene

240

24

8

207

33

M

M

M

B

Cactus

500

50

8

403

97

M

M

M

B

BasketballDrive

500

50

8

403

97

M

M

M

B

BQTerrace

600

60

8

471

129

M

M

M

C

RaceHorses

300

30

8

235

65

M

M

M

C

BQMall

600

60

8

471

129

M

M

M

C

PartyScene

500

50

8

403

97

M

M

M

C

BasketballDrill

500

50

8

403

97

M

M

M

D

RaceHorses

300

30

8

235

65

M

M

M

D

BQSquare

600

60

8

471

129

M

M

M

D

BlowingBubbles

500

50

8

403

97

M

M

M

D

BasketballPass

500

50

8

403

97

M

M

M

O

Kimono

240

24

8

207

33

O

O

O

Note: M – mandatory; O – optional.

It was decided in the last JVET meeting to replace some of the compressed source video sequences in Tencent Video Dataset (TVD) by their uncompressed versions. The updated source video sequences are renamed, as shown in the next table below, and can be downloaded from https://multimedia.tencent.com/resources/tvd with corresponding annotations.

Test sequences in TVD

Sequence name

Frame count

Frame rate

Bit depth

Frames skipped

Frames coded

Intra

Random access

Low delay

TVD-01-1

3000

50

8

1500

500

M

M

M

TVD-01-2

3000

50

8

2000

500

M

M

M

TVD-01-3

3000

50

8

2500

500

M

M

M

TVD-02-1

636

50

10

0

636

M

M

M

TVD-03-1

2334

50

10

0

500

M

M

M

TVD-03-2

2334

50

10

500

500

M

M

M

TVD-03-3

2334

50

10

1000

500

M

M

M

As the original sequences are available in .mp4 format, they shall be converted to YUV420 using FFmpeg. The format conversion process was cleaned up as follows:

ffmpeg -i {input.mp4} {output.yuv}

The TVD (image) dataset is an image dataset of 166 images of 1920x1080 resolution that have annotations for object detection and instance segmentation. The dataset with corresponding annotations is available at https://multimedia.tencent.com/resources/tvd.

The OpenImages dataset consists of around 9 million images. A subset of the validation set of its version 6 containing 5000 images are selected for testing object detection in this activity. The dataset with corresponding annotations is available at https://storage.googleapis.com/openimages/web/index.html.

The FLIR dataset used in the WG 4 VCM activities is a dataset consisting of 300 infrared images. The images, annotations and the fine-tuned model for thermal images can be found on the MPEG content repository (https://content.mpeg.expert/data/).

More detailed information about these test datasets can be found in JVET-AD2031.

Anchor software and configuration

Version 12.0 of the VTM software is used for generating anchor results. The VTM software is available at https://vcgit.hhi.fraunhofer.de/jvet/VVCSoftware_VTM/.

In addition, version 4.2.2 of the FFmpeg is used for the format conversion. The FFmpeg software is available at https://ffmpeg.org/releases.

Three test conditions are used for experimenting encoder and receiving system optimizations on video coding for machine consumptions, i.e., random-access, low-delay and all-intra. The default configuration files provided with the VTM software are used for anchor generation.

  • “Random access” (RA): encoder_randomaccess_vtm.cfg
  • “Low delay” (LD): encoder_lowdelay_vtm.cfg
  • “All Intra” (AI): encoder_intra_vtm.cfg

It was decided in the last JVET meeting that temporal subsampling is disabled for “All Intra” configuration for this activity, thus the following line

TemporalSubsampleRatio : 8

in “encoder_intra_vtm.cfg” needs to be modified to

TemporalSubsampleRatio : 1

A subset of these configurations might be used for a particular experiment. However, as these test conditions are for video coding experiments, results for at least one of the random access or low delay configurations are expected to be provided.

Anchor generation

The anchor generation process is illustrated in the figure below. The task network Faster R-CNN X101-FPN which is part of Detectron2 (https://github.com/facebookresearch/detectron2) is used to evaluate object detection (SFU-HW dataset and image datasets), the model is available here. The task network JDE-1088x608 which is part of the Towards Realtime MOT framework (https://github.com/Zhongdao/Towards-Realtime-MOT) is used to evaluate object tracking (TVD), the model is available at this Google Drive link. More details about anchor generation process and these machine task networks are referred to JVET-AD2031.

Figure: Anchor generation pipeline

Anchor results were generated following the CTC described in JVET-AD2031 and crosschecked by Ericsson and Tencent.

Evaluation and reporting

Proposed technologies are evaluated based on their compression performance, measured by bitrate, PSNR, mAP and MOTA, as well as encoding and decoding runtime to reflect the complexity of the proposed technology, to some extent. Definitions and detailed descriptions of these metrics can be found in JVET-AD2031.

It is noted that the mean Average Precision (mAP) is used to measure object detection performance, and Multiple Object Tracking Accuracy (MOTA) is used to measure object tracking performance. For the purpose of reporting encoding and decoding running times, the anchor and proposal should be simulated on the same platform, e.g. similar CPU and GPU configuration, to have reliable time comparison. Parallel encoding and decoding as described in JVET-B0036 may be applied for RA configurations.

In addition, relevant inference and training information should be reported if the proposed technology consists of learning-based components, such as network structure, number of network parameters, precision of network parameters, number of multiply–accumulate operations (MAC) per pixel, patch size, batch size, epoch, training time, training datasets, lost function, number of iterations, optimizer, and any pre- and post-processings used. More details can be found in JVET-AD2031.

A reporting template in Excel file format has been prepared to illustrate results following CTC and evaluation methodology described in JVET-AD2031. It is enclosed in JVET-AD2031 package.

Reporting template summary

Technical Report

The draft 2 of the technical report (TR) has been prepared and uploaded as JVET-AD2030 based on discussions in the last JVET meeting on 2023-06-07. More descriptions were added to use cases and applications, pre-processing technologies and encoding technologies. Two software implementation examples were included in the TR draft 2 as Annex.

Git Management

AHG 8 related software and documents can be found at https://vcgit.hhi.fraunhofer.de/jvet-ahg-ofm. This repository contains two projects, one (https://vcgit.hhi.fraunhofer.de/jvet-ahg-ofm/ofm-ctc) containing instrucitons and information for conducting experiements and evaluation, such as evaluation scripts, machine task networks, CTC and reporting template with anchor results, while the other (https://vcgit.hhi.fraunhofer.de/jvet-ahg-ofm/vtm-ofm) containing implementation examples. Two software example packages have been uploaded to the git in separate branches:

  • JVET-AB0275: a region of interest-based method that uses adaptive QP to reduce the quality in background areas
  • JVET-AC0086: a method that uses a pre-analysis to perform content adaptive machine vision oriented preprocessing

AHG Coordination

Collaborative discussions were held between this AHG and the VCM AHG in WG4 to synchronize common test conditions including test datasets, anchor QP points, encoder configurations, etc. A joint meeting was held on 2023-06-21 0500UTC to discuss metrics, existing issues and potential solutions for machine task performance evaluation. Two contributions were presented and discussed during this meeting:

m63676 - [VCM] Proposed performance evaluation method for VCM CE1 (ETRI, Konkuk Univ., Myongji Univ.)

This contribution proposes a performance evaluation method for VCM CE1.

Regarding the performance evaluation for VCM CE1, there are some difficulties as follows:

  1. The current BD-rate metric provided by the reporting template is not able to produce a value if either the anchor or a proposal is non-monotonic.
  2. To compensate for this limitation, the reporting template requires submission of BD-rate for multiple subcases (combination of 4 or 5 points out of 6 total points), but it is still difficult to make a direct comparison between proposals since the metric is not singular.

The following rules were proposed as the performance evaluation method for VCM CE1:

  1. For each test case, a proponent should submit 4 points that correspond to the high 4 rate points of the VTM anchor.
  2. Only if the result at the high 4 rate points is non-monotonic, then it is allowed to adjust QP in a range of [-2, +2] to produce a monotonic result.
  3. If a monotonic result cannot be obtained even with the adjustment of QP in the range, then it is not considered as a candidate technology in this round of CE1 evaluation.
  4. Evaluate the proposed technologies using the summed scores based on the class-wise (SFU) or sequence-wise (TVD) BD-rate rankings of the proposals.

JVET-AE0107 / m63692 - [VCM] Improvements of the BD-rate model using monotonic curve-fitting method (Tencent)

This contribution provides improvements of the BD-rate model used in current Excel templates. Specifically, the contribution proposes a monotonic fitted curve to replace the PCHIP (Piecewise Cubic Hermit Interpolation) curve for BD-rate and BD-metric calculation, such as BD-PSNR, BD-mAP, BD-MOTA, etc. The fitted curve guarantees monotonicity even if the input points are nonmonotonic, thus makes it possible to calculate BD-rate and BD-metric values in various testing conditions.

This contribution was also submitted to JVET (AHG 8) as document JVET-AE0107.

Input contributions

There were 12 input contributions related to AHG 8 mandates. They are listed below.

Report

JVET-AE0008

JVET AHG report: Optimization of encoders and receiving systems for machine analysis of coded video content (AHG8)

C. Hollmann, S. Liu, S. Wang, M. Zhou (AHG chairs)

Proposal

JVET-AE0053

AHG8/AHG9: Neural-network post-filter regions SEI message

T. Chujoh, Y. Yasugi, T. Ikai (Sharp)

JVET-AE0064

AHG8/AHG9: Signalling encoder preprocessing and human / machine viewing indications

C. Kim, D. Gwak, Hendry, J. Lim, S. Kim (LGE), M. M. Hannuksela, F. Cricri, H. Zhang (Nokia)

JVET-AE0079

AHG8/AHG9: Source picture timing information SEI message

S. McCarthy, G. J. Sullivan, P. Yin (Dolby)

JVET-AE0081

[AHG8] De-noising filter as pre-processing for machine task

C. Kim, D. Gwak, J. Lim (LGE)

JVET-AE0090

AHG8/AHG9: On machine vision indication

J. Gao, H.-B. Teo, C.-S. Lim, K. Abe (Panasonic)

JVET-AE0095*

AHG8/AHG9: proposed changes to the candidate new object mask information SEI message

P. de Lagrange, E. François, D. Doyen (InterDigital), J. Chen, S. Wang, Y. Ye (Alibaba)

JVET-AE0096

[AHG8] Study on using different VTM versions

C. Hollmann, J. Ström (Ericsson)

JVET-AE0099

[AHG8] NNPF and post-filter hint SEI messages for the technical report

C. Hollmann, M. Pettersson, R. Sjöberg, M. Damghanian (Ericsson)

JVET-AE0107

AHG8: Improvements of the BD-rate model using monotonic curve-fitting method

H. Wang, X. Pan, Z. Liu, X. Xu, S. Liu (Tencent)

JVET-AE0143

AHG8: A spatial resampling algorithm and an exemplar software implementation

S. Wang, B. Li, J. Chen, Y. Ye (Alibaba), S. Wang (CityU)

Crosscheck

JVET-AE0234

AHG8: Crosscheck of JVET-AE0107

Honglei Zhang (Nokia)

* Late.

Recommendations

The AHG recommended to:

  • Review all input contributions.
  • Discuss and refine test conditions, evalution and reporting procedures.
  • Discuss the non-monotonic issue and potential solutions.
  • Continue investigating non-normative technologies and their suitability for machine analysis applications.
  • Continue developing draft technical report on optimization of encoders and receiving systems for machine analysis of coded video content.
  • Continue collecting new test materials.

It was pointed out that a potential timeline for the technical report should be discussed. This however may also relate to the level of completeness in terms of using video compression standards for a variety of machine usage applications, which somewhat relies on the availability of test material of sufficient variety.

Decisions
The AHG recommended to: Review all input contributions. Discuss and refine test conditions, evalution and reporting procedures. Discuss the non-monotonic issue and potential solutions. Continue investigating non-normative technologies and their suitability for machine analysis applications. Continue developing draft technical report on optimization of encoders and receiving systems for machine analysis of coded video content. Continue collecting new test materials. It was pointed out that a potential timeline for the technical report should be discussed. This however may also relate to the level of completeness in terms of using video compression standards for a variety of machine usage applications, which somewhat relies on the availability of test material of sufficient variety.
Citation