JVET-K0002 JVET AHG report: Draft text and test model algorithm description editing (AHG2) [E. Alshina, B. Bross, J. Chen]
This document reports the work of the JVET ad hoc group on draft text and test model algorithm description editing (AHG2) between the 10th Meeting in San Diego, US (10–20 Apr 2018) and the 11th meeting in Ljubljana, SI (10–18 July 2018).
At the 10th JVET meeting, JVET defined the first draft of Versatile Video Coding (VVC) (JVET-J1001) and the VVC Test Model 1 (VTM1) encoding method (JVET-J1002). It was decided to include a quadtree with nested multi-type tree using binary and ternary splits coding block structure as the initial new coding feature of VVC. Draft reference software to implement the VVC decoding process and VTM1 encoding method has also been developed.
The normative decoding process for Versatile Video Coding is specified in the VVC draft 1 text specification document. This VVC Test Model 1 (VTM 1) Algorithm and Encoder Description document provides an algorithm description as well as an encoder-side description of the VVC Test Model 1, which serves as a tutorial for the algorithm and encoding model implemented in the VTM1.0 software.
Two versions of JVET-J1001 and two versions of JVET-J1002 were published by the Editing AHG between the 10th Meeting in San Diego (10–20 Apr 2018) and the 11th meeting (10–18 July 2018).
JVET-J1001 had been established from scratch and currently contained the following:
- Basic definitions, abbreviations and conventions
- A basic high-level syntax (HLS) with NAL units, SPS, PPS and slice header.
- Block partitioning by a quadtree with nested multi-type tree using binary and ternary splits with:
- CU leaf nodes
- Prediction at CU level
- Transform at CU level
- Minimum CU size with 4x4 luma coding block and corresponding chroma coding blocks (2x2 for 4:2:0)
- Maximum TU size with 64x64 luma transform block and corresponding chroma transform blocks (32x32 for 4:2:0)
- Minimum TU size with 4x4 luma transform block and corresponding chroma transform blocks (2x2 for 4:2:0)
- Single tree for luma and chroma
JVET-J1002 had also been established from scratch. The document generally describes the basic coding architecture, the partitioning of the picture into CTUs, and the partitioning of the CTUs using a quadtree with nested multi-type tree.
For initial testing purposes of the aspects of the design that have not yet been determined, the test model software uses syntax, semantics, and decoding processes that correspond to those in prior well-known video coding designs. However, these aspects are considered only to be “placeholders” for specific design details yet to be determined. The exact details of the binary/ternary/quaternary segmentation tree structure to be used are also yet to be determined. It was noted that this document may contain a description of some such details that should not be considered completely agreed upon.
As agreed in the 10th JVET meeting, the following features that are found in HEVC were not included in the initial VVC test model.
- Special strong boundary smoothing for 32×32 luma block intra prediction
- Boundary smoothing across edges for intra prediction (a horizontal filter for vertical prediction and vice versa, and the first row and column with DC prediction)
- DST-VII style transform in 4×4 intra blocks
- Mode-dependent scan for intra blocks
- Quantization weighting matrices
- Residual sign bit hiding
- VPS and VPS VUI
- Dependent slices
- Tiles
- Wavefronts (entropy coding synchronization)
In terms of the impact of this on specific elements of the design, this includes removal of the following features (and some others):
- Partitioning of a CU into multiple PUs (including asymmetric partitionings)
- Partitioning of a CU into multiple luma blocks for intra prediction (i.e., signalling of multiple luma intra prediction modes for a CU), except for implicit splits when the CU size is too large for the maximum transform size
- The coding unit syntax element part_mode
- Partitioning of a CU into multiple TUs, except for implicit splits when the CU size is too large for the maximum transform size
- Transforms that are applied across prediction block boundaries
- The syntax element split_transform_flag
- Non-aligned luma and chroma transform blocks
- All VPS and VPS VUI syntax
- SPS syntax elements
- log2_min_luma_transform_block_size_minus2 (always use 4x4 luma and corresponding chroma)
- log2_diff_max_min_luma_transform_block_size
- max_transform_hierarchy_depth_inter
- max_transform_hierarchy_depth_intra
- amp_enabled_flag
The AHG recommended to:
- Approve the edited JVET-J1001 and JVET-J1002 documents as the JVET outputs:
- Continue to edit the VVC WD and Test Model documents to ensure that all agreed elements of VVC are fully described.
- Compare the VVC documents with the VVC software and resolve any discrepancies that may exist, in collaboration with the Software AHG.
- Continue to improve the editorial consistency of VVC WD and Test Model documents.
- Ensure that, when considering the addition of new feature to VVC, properly drafted text for addition to the VVC Test Model and/or the VVC Working Draft is made available in a timely manner.