Back to Search Document details
32nd Meeting: Hannover, DE, October 2023 2023-10-17 13:58
AHG8: A temporal resampling algorithm
Abstract
In JVET-AE2030, a pre-processing method based on temporal resampling is proposed, which can be used to optimize video encoders and receiving systems for machine consumption. By skipping certain frames before encoding, the bit rate can be significantly reduced without a strong negative impact on the machine consumption performance. In this contribution, a temporal resampling algorithm is detailed and its software implementation is provided. Based on the common test conditions described in JVET-AE2031, it is reported that the following coding performance can be achieved:
JVET-AF0157 AHG8: A temporal resampling algorithm [D. Ding, X. Zhao, Z. Liu, S. Liu (Tencent)]

In JVET-AE2030, a pre-processing method based on temporal resampling is proposed, which can be used to optimize video encoders and receiving systems for machine consumption. By skipping certain frames before encoding, the bit rate can be significantly reduced without a strong negative impact on the machine consumption performance. In this contribution, a temporal resampling algorithm is detailed and its software implementation is provided. Based on the common test conditions described in JVET-AE2031, it is reported that the following coding performance can be achieved:

For VTM 20:

For object tracking

  • 2x temporal resampling rate: -51.02% (AI), -15.88% (LD), -8.32% (RA)
  • 4x temporal resampling rate: -75.93% (AI), -30.53% (LD), -3.84% (RA)

For object detection

  • 2x temporal resampling rate: -49.91% (AI), -13.18% (LD), -6.72% (RA)
  • 4x temporal resampling rate: -69.53% (AI), 5.53% (LD), 31.91% (RA)

For VTM 12:

For object tracking

  • 2x temporal resampling rate: -50.69% (AI), -9.11% (LD), -4.04% (RA)
  • 4x temporal resampling rate: -75.99% (AI), -24.00% (LD), -3.69%(RA)

For object detection

  • 2x temporal resampling rate: -50.50% (AI), -7.76% (LD), 4.94% (RA)
  • 4x temporal resampling rate: -69.61% (AI), 10.69% (LD), 29.1% (RA)

In the V2 verison the results of RA, as well as the comparison with VTM 12 anchor are updated. In addition, it is also discussed about the configuration of RA to fair compare with anchor. It is shown that the solution is an encoder decision coding method.

In the V3 version, the VTM12 Object Track, RA results were updated.

The results for RA are more realistic than for JVET-AF0060, as the RA period is kept as 1 s.

Complexity? Not exactly known, run on GPU.

It was commented that run time may not be sufficient for assessing complexity.

Comments from crosschecker:

Generally the results matched. In case of LB, there was an outlier park scene with considerably worse results.

Further study was recommended for both JVET-AF0060 and JVET-AF0157

  • To investigate the benefit of sophisticated temporal upsampling algorithms, by comparing them against a solution where the machine vision task is just operated at the downsampled sequence, or against a simpler interpolation method, e.g. based on optical flow.
  • To better document the complexity, e.g. model size, kMAC/px, CPU runtime, etc.

General for AHG8:

The CTC requires a change on the aspect of RA period when subsampling is used.

Decisions
The CTC requires a change on the aspect of RA period when subsampling is used.
Citation