Back to Search Document details
29th Meeting: by teleconference, DE, January 2023 2023-01-02 13:42
[AHG15] Draft technical report on optimizations for encoders and receiving systems for machine analysis of coded video content
Abstract
This contribution contains a draft for the technical report on optimizations for encoders and receiving systems for machine analysis of coded video content that JVET decided on working on during its 28th meeting.
JVET-AC0049 [AHG15] Draft technical report on optimizations for encoders and receiving systems for machine analysis of coded video content [C. Hollmann (Ericsson), S. Liu (Tencent), S. Wang (Alibaba)]

This contribution contains a proposed draft for the technical report on optimizations for encoders and receiving systems for machine analysis of coded video content that JVET decided on working on during its 28th meeting.

It was suggested to include the concept of keyframe identification and analysis. Visual features conained in keyframes give important clues on the content of a video for machine vision tasks and should not be destroyed by coding. This could also be combined with dynamically changing the frame rate based on amount of temporal changes.

It is suggested to also mention conventional (e.g. VVC interpolation) filters for spatial upsampling rather than bicubic.

It was commented that the current version of the technical report is very high-level, and needs to be filled by more concrete technology implementation for the given purposes, e.g in terms of pre and post processing, and related encoder control.

It would also be important to discuss the video content properties which need to be retained for the machine vision tasks.

It might be good to identify a selection of machine vision tasks and describe how a video needs to be encoded in order to not destroy the performance relative to the encoded video.

The purpose should be to generate a bitstream without knowing which machine vision task would be applied after decoding. Encoding for exactly one task could hypothetically be much more optimized, but might not be a realistic application scenario where compression would be needed.

Which types of metadata are existing in our standards that could be useful for machine vision task? For example, there are many ways of describing camera position, movement, etc.

Our standards also contain ways to code depth maps, for example

A comment was made about possible low-latency requirements.

The report requires a clear focus, and that should come from the application perspective – this focus should be the same for the normative activities of WG 4 and non-normative of the TR. JVET experts have experience about how to use and how to optimize encoders for low-level tools.

It was planned to further discuss these aspects in a joint meeting with WG 4 (see section 7.4).

Decisions
It was planned to further discuss these aspects in a joint meeting with WG 4 (see section 7.4).
Citation