Back to Search Document details
31st Meeting: Geneva, CH, July 2023 2023-07-13 15:09
AHG8/AHG9: Proposed changes to the candidate new object mask information SEI message
Abstract
This contribution proposes an update to the candidate new object mask information (OMI) SEI message proposed in contribution JVET-AD0175. The OMI SEI message is used in conjunction with auxiliary pictures that code object masks. Objects can be detected from the auxiliary pictures thanks to their ID, signaled in the SEI message. One potential issue is when lossy coding is used for those pictures, in which case the sample values inside the mask auxiliary pictures are not guaranteed to perfectly match the object IDs. The proposal aims at addressing this issue, by adding in the SEI a parameter specifying a tolerance range around the object ID values. In addition, bounding box parameters that delimit the masks’ location are also proposed to be included.
JVET-AE0095 AHG8/AHG9: Proposed changes to the candidate new object mask information SEI message [P. de Lagrange, E. François, D. Doyen (InterDigital), J. Chen, S. Wang, Y. Ye (Alibaba)] [late]

AHG10: Encoding algorithm optimization (4)

Contributions in this area were discussed at 1240–1300 and 1430–1530 on Sunday 16 July 2023 (chaired by JRO).

JVET-AE0095 AHG8/AHG9: proposed changes to the candidate new object mask information SEI message [P. de Lagrange, E. François, D. Doyen (InterDigital), J. Chen, S. Wang, Y. Ye (Alibaba)] [late]

This contribution proposes an update to the candidate new object mask information (OMI) SEI message proposed in contribution JVET-AD0175. The OMI SEI message is used in conjunction with auxiliary pictures that code object masks. Objects can be detected from the auxiliary pictures thanks to their ID, signalled in the SEI message. One potential issue is when lossy coding is used for those pictures, in which case the sample values inside the mask auxiliary pictures are not guaranteed to perfectly match the object IDs. The proposal aims at addressing this issue, by adding in the SEI a parameter specifying a tolerance range around the object ID values. In addition, bounding box parameters that delimit the masks’ location are also proposed to be included.

The purpose of using an alpha map to reconstruct binary object shapes (rather than “grey-level shapes”) via an auxiliary picture is well recognized and important. An open question is whether the method of defining a threshold in case of lossy coding of the map is more efficient (and practical from the encoder view point) than right away doing lossless coding (e.g. of a picture that just contains 0/1 values). For complicated shapes, it might be difficult to determine at which QP lossy coding needs to be performed, to guarantee that no points of the object are lost, and no points outside of the object are interpreted as object points.

To some extent, perfect lossless coding may not be necessary, and post processing of the lossy decoded shape could also be performed.

Further study was recommended, bringing examples/results with real-world examples rather than hand-crafted examples.

This would have potential applications in machine consumption tasks, but even more for general purposes of segmented video, such as video conferencing.

It was pointed out that V-PCC has some similar concepts of coding the occupancy map, using either lossy or lossless compression. Comparison would be beneficial.

Decision: Adopt JVET-AE0095 into TuC for VSEI extension, without the thresholding.

Decisions
JVET-AE0095 AHG8/AHG9: Proposed changes to the candidate new object mask information SEI message [P. de Lagrange, E. François, D. Doyen (InterDigital), J. Chen, S. Wang, Y. Ye (Alibaba)] [late]
adopted
Adopt JVET-AE0095 into TuC for VSEI extension, without the thresholding
Citation