JVET-AL0152 AHG8: Dense QP coding results of combining adaptive temporal resampling, pre-processing, ROI-based adaptive QP, QP offset adjustment for higher temporal layers and post-processing algorithms for machine vision [S. Wang, J. Chen, Y. Ye, B. Li (Alibaba), S. Wang (CityUHK)]
This contribution reports the dense QP coding results of jointly enabling adaptive temporal resampling, pre-processing, ROI-based adaptive QP (ROIAQP), QP offset adjustment for higher temporal layers (QPOA) and post-processing algorithms for machine vision.
According to suggestion of last JVET meeting, two sets of algorithms are evaluated under common test conditions (CTC). Set 3 is the combination of adaptive temporal resampling, pre-processing, ROIAQP and post-processing. Set 4 additionally enabling QPOA on top of set 3. It is reported that set 3 achieves 33.77% (RA), 44.33% (LD) and 80.91% (AI) averaged BD-rate savings for object detection on SFU-HW dataset, and 62.89% (RA), 59.96% (LD), 92.33%(AI) averaged BD-rate savings are achieved for object tracking on TVD dataset. Set 4 achieves 35.84% (RA), 44.45% (LD) and 80.91% (AI) averaged BD-rate savings for object detection on SFU-HW dataset, and 65.03% (RA), 63.36% (LD), 92.33% (AI) averaged BD-rate savings for object tracking on TVD dataset.
This contribution reports the dense QP coding results of jointly enabling adaptive temporal resampling, pre-processing, ROI-based adaptive QP (ROIAQP), QP offset adjustment for higher temporal layers (QPOA) and post-processing algorithms for machine vision.
According to suggestion of last JVET meeting, two sets of algorithms are evaluated under common test conditions (CTC). Set 3 is the combination of adaptive temporal resampling, pre-processing, ROIAQP and post-processing. Set 4 additionally enabling QPOA on top of set 3. It is reported that set 3 achieves 33.77% (RA), 44.33% (LD) and 80.91% (AI) averaged BD-rate savings for object detection on SFU-HW dataset, and 62.89% (RA), 59.96%(LD), -92.33xx%(AI) averaged BD-rate savings are achieved for object tracking on TVD dataset. Set 4 achieves 35.84% (RA), 44.45% (LD) and 80.91% (AI) averaged BD-rate savings for object detection on SFU-HW dataset, and 65.03xx% (RA), 63.36xx%(LD), -92.33xx%(AI) averaged BD-rate savings for object tracking on TVD dataset.
It was requested to extend the table of combinations as follows:
The combination can be enabled via the config file, the code existed before and was not changed (was cross-checked previously independently).
It was pointed out that the QP offset adjustment is in clause 8.2 (not 8.1). This also needs to be corrected in the current table of the TR (JVET-AK2030).
It was agreed to include the yellow highlighted row as 10th combination into the table. It was encouraged to upload the bitstreams for the new combinations to the MPEG content server.
Editors of the TR were asked to prepare a draft of the DoCR to be reviewed on Wednesday.
AHG10: Encoding algorithm optimization (3+1)
Contributions in this area were discussed during 1415–1510 on Monday 31 March 2025 (chaired by JRO).