Back to Search Document details
32nd Meeting: Hannover, DE, October 2023 2023-10-13 13:59
EE2: Summary report of exploration experiment on enhanced compression beyond VVC capability
Abstract
This document provides a summary report of Exploration Experiment on Enhanced Compression beyond VVC capability. The tests are categorized as partitioning, intra prediction, inter prediction, transform and coefficient coding, and in-loop filtering.
JVET-AF0024 EE2: Summary report of exploration experiment on enhanced compression beyond VVC capability [V. Seregin, J. Chen, R. Chernyak, K. Naser, J. Ström, F. Wang, M. Winken, X. Xiu, K. Zhang]

This document provides a summary report of Exploration Experiment on Enhanced Compression beyond VVC capability. The tests are categorized as partitioning, intra prediction, inter prediction, transform and coefficient coding, and in-loop filtering.

The software basis for this EE is ECM-10.0, released at https://vcgit.hhi.fraunhofer.de/ecm/ECM/-/tags/ECM-10.0. ECM-10.0 is used as an anchor in the tests.

Software for EE tests is released in the corresponding branches at https://vcgit.hhi.fraunhofer.de/ecm/jvet-ae-ee2/ECM/-/branches.

Test results can be found in input JVET contributions, cross-check results are uploaded to https://vcgit.hhi.fraunhofer.de/ecm/jvet-ae-ee2/simulation-results if cross-check reports are not submitted as they are optional for EE tests.

Tests

Tester

Cross-checker

1 Partitioning

1.1a

Non-square quadtree partitioning

LGE

Y. Ahn

JVET-AF0108

InterDigital

R. Utida

1.1b

ECM with maximum MTT depth increments

LGE

Y. Ahn

JVET-AF0108

InterDigital

R. Utida

2 Intra prediction

2.1a

DIMD merge

Nokia

S. Blasi

JVET-AF0120

Ofinno

P. Andrivon

JVET-AF0124

2.1b

DIMD merge with reduced storage

Nokia

S. Blasi

JVET-AF0120

Ofinno

P. Andrivon

JVET-AF0124

2.2a

DIMD with filtered template

vivo

C. Zhou

JVET-AF0131

Alibaba

J. Chen

2.2b

DIMD with filtered template (without modification of gradient operators)

vivo

C. Zhou

JVET-AF0131

Alibaba

J. Chen

2.3a

Test 2.1a + Test 2.2a

Nokia

S. Blasi

vivo

C. Zhou

JVET-AF0126

Bytedance

W. Yin

JVET-AF0204

2.3b

Test 2.1a + Test 2.2b

Nokia

S. Blasi

vivo

C. Zhou

JVET-AF0126

Bytedance

W. Yin

JVET-AF0204

2.3c

Test 2.1b + Test 2.2a

Nokia

S. Blasi

vivo

C. Zhou

JVET-AF0126

InterDigital

K. Naser

2.3d

Test 2.1b + Test 2.2b

Nokia

S. Blasi

vivo

C. Zhou

JVET-AF0126

InterDigital

K. Naser

2.4

IntraCIIP as additional mode of IntraTMP

InterDigital

K. Naser

JVET-AF0130

OPPO

F. Wang

JVET-AF0214

2.5a

TIMD with IntraTMP/IBC

InterDigital

K. Naser

JVET-AF0136

Ofinno

P. Andrivon

JVET-AF0125

2.5b

Test 2.5a + Test 2.4

InterDigital

K. Naser

JVET-AF0136

OPPO

F. Wang

JVET-AF0215

Nokia

S. Blasi

JVET-AF0243

2.6a

Fractional-pel intraTMP BVs are stored in block vector buffer

OPPO

Y. Yu

Qualcomm

P.-H. Lin

JVET-AF0079

ETRI

W. Lim

JVET-AF0227

2.6b

IntraTMP BVs are stored for HMVP

OPPO

Y. Yu

Qualcomm

P.-H. Lin

JVET-AF0079

ETRI

W. Lim

JVET-AF0227

2.6c

Test 2.6a + Test 2.6b

OPPO

Y. Yu

Qualcomm

P.-H. Lin

JVET-AF0079

ETRI

W. Lim

JVET-AF0227

2.7

Extrapolation filter-based intra prediction mode

OPPO

L. Xu

JVET-AF0080

Kwai

H. -J. Jhu

JVET-AF0230

Alibaba

X. Li

JVET-AF0221

2.8

DBV signalling modification

OPPO

L. Xu

JVET-AF0081

Qualcomm

H. Huang

JVET-AF0219

2.9a

Enable DBV in single tree under CTC

Qualcomm

H. Huang

JVET-AF0066

OPPO

L. Xu

JVET-AF0076

2.9b

Enable DBV in single tree when DualITree is set to 0

Qualcomm

H. Huang

JVET-AF0066

OPPO

L. Xu

JVET-AF0076

2.10

Combination test of Test 2.8 and Test 2.9

Qualcomm

H. Huang

OPPO

L. Xu

JVET-AF0207

VIVO

C. Zhou

2.11a

Test 2.1b + Test 2.5b

InterDigital

K. Naser

Nokia

S. Blasi

vivo

C. Zhou


JVET-AF0229

Xiaomi

R.G. Youvalari

M. Abdoli

JVET-AF0244

2.11b

Test 2.2b + Test 2.5b

InterDigital

K. Naser

Nokia

S. Blasi

vivo

C. Zhou

JVET-AF0229

Xiaomi

R.G. Youvalari

JVET-AF0244

2.11c

Test 2.3c + Test 2.5b

InterDigital

K. Naser

Nokia

S. Blasi

vivo

C. Zhou

JVET-AF0229

Xiaomi

R.G. Youvalari

JVET-AF0244

2.11d

Test 2.3d + Test 2.5b

InterDigital

K. Naser

Nokia

S. Blasi

vivo

C. Zhou

JVET-AF0229

Xiaomi

R.G. Youvalari

JVET-AF0244

3 Inter prediction

3.1a

CCP merge mode for chroma inter coding

MediaTek

M.-S. Chiang

Bytedance

Z. Deng

JVET-AF0073

InterDigital
K. Naser

JVET-AF0240

3.1b

CCP merge mode for chroma inter coding without the additional second type of shifted temporal candidates

MediaTek

M.-S. Chiang

Bytedance

Z. Deng

JVET-AF0073

Qualcomm

Y.-J. Chang

JVET-AF0233

3.2

LIC flag derivation of merge candidates with template costs

Bytedance

N. Zhang

JVET-AF0128

OPPO

Z. Xie

3.3a

Multi-model CCRM

Bytedance

Z. Deng

3.3b

CCP merge mode with the derived candidates with 4 derived candidates

Bytedance

Z. Deng

MediaTek

M.-S. Chiang

JVET-AF0073

vivo

Z. Lv

JVET-AF0242

3.1d

Test 3.1a + Test 3.3b (with 1 derived candidate for low-delay pictures)

MediaTek

M.-S. Chiang

Bytedance

Z. Deng

JVET-AF0073

OPPO

F. Wang

JVET-AF0213

vivo

Z. Lv

JVET-AF0242

3.4a

TM-based subblock motion refinement

Bytedance

L. Zhao

JVET-AF0163

OPPO

Z. Xie

JVET-AF0270

3.4b

Interweaved affine prediction

Bytedance

L. Zhao

JVET-AF0163

OPPO

Z. Xie

JVET-AF0270

3.4c

RMVF candidate derivation with multiple CUs

Bytedance

L. Zhao

JVET-AF0163

OPPO

Z. Xie

JVET-AF0270

3.4d

Test 3.4a + Test 3.4b + Test 3.4c

Bytedance

L. Zhao

JVET-AF0163

OPPO

Z. Xie

JVET-AF0270

3.5a

DMVR with robust MV derivation

Ericsson

K. Andersson

JVET-AF0057

3.5b

DMVREncSelect from VTM (encoder only)

Ericsson

K. Andersson

JVET-AF0057

3.5a*

Test 3.5a with QP 40, 43, 47, and 50

Ericsson

K. Andersson

JVET-AF0057

3.5b*

Test 3.5b with QP 40, 43, 47, and 50

Ericsson

K. Andersson

JVET-AF0057

3.6a

Affine subblock BDOF refinement

Qualcomm

Z. Zhang

JVET-AF0159

Kwai

X. Xiu

3.6b

AMVP-merge mode for affine

Qualcomm

Z. Zhang

JVET-AF0159

Kwai

X. Xiu

3.6c

Test 3.6a + Test 3.6b

Qualcomm

Z. Zhang

JVET-AF0159

Kwai

X. Xiu

4 Reference picture resampling

4.1a

Enabling template-based reordering for scaled reference pictures

Kwai

X. Xiu

JVET-AF0190

Qualcomm

Z. Zhang

4.1b

Test 4.1a + Enabling LIC for scaled reference pictures

Kwai

X. Xiu

JVET-AF0190

Qualcomm

Z. Zhang

4.2a

Filtering applied after motion compensation, the post-processing upsampling is not changed

RWTH Aachen Univ.

T. Claßen

JVET-AF0173

4.2b

Test 4.2a + perform the filtering after reconstruction

RWTH Aachen Univ.

T. Claßen

JVET-AF0173

5 In-loop filtering

5.1a

Dynamic TU scale factor for BIF with LUTs interpolation

Ericsson

V. Shchukin

JVET-AF0112

Bytedance

W. Yin

JVET-AF0203

5.1b

Dynamic TU scale factor for BIF

Ericsson

V. Shchukin

JVET-AF0112

Bytedance

W. Yin

JVET-AF0203

5.2a

Luma-residual tap in chroma-ALF

Bytedance

W. Yin

JVET-AF0197

Ericsson

V. Shchukin

JVET-AF0225

5.2b

Test 5.2a + Luma-residual tap in CCALF

Bytedance

W. Yin

JVET-AF0197

Ericsson

V. Shchukin

JVET-AF0225

6 Entropy coding

6.1a

Spatial CABAC tuning

Nokia

J. Lainema

JVET-AF0109

Qualcomm

P. Nikitin, I. Jumakulyyev

JVET-AF0188

6.1b

Spatial CABAC tuning with reduced memory/latency

Nokia

J. Lainema

JVET-AF0109

Qualcomm

P. Nikitin, I. Jumakulyyev

JVET-AF0188

6.2

Retrain I-slices context model parameters

Alibaba

R.-L. Liao

JVET-AF0133

InterDigital
K. Naser

JVET-AF0239

6.3

Test 6.2 + Test 6.1

Nokia

J. Lainema

Alibaba

R.-L. Liao

JVET-AF0110

Kwai

X. Xiu

JVET-AF0232

Test 1 - Partitioning

Test 1.1: Non-square quadtree partitioning (JVET-AF0108)

In this test, the existed restriction that any non-square block cannot be split with quadtree is removed for 2NxN and Nx2N blocks, where N is greater or equal to 8. The minimum block size depends on picture size and configuration, and the maximum block sizes can be decided according to the maximum block sizes for binary or ternary tree partitioning. Under the common test conditions, the largest block sizes for the non-square quadtree partitioning are 128×64 or 64×128 for Class A/B and 64×32 or 32×64 for other classes.

Fast encoding algorithm is applied in the test, which is related to constraints of partitioning depths considering that non-square quadtree depth is included in MTT depths, and non-square quadtree partitioning is allowed only once in MTT depths under the common test configurations. The second fast encoding algorithm is an early termination method based on CBF of the previous coding block split with BT or TT in rate-distortion optimization process. The third fast encoding algorithm is an early termination method based on RD-cost comparison between NO_SPLIT and BT/TT blocks in RDO process.

Test 1.1a: Non-square quadtree with fast encoding algorithms

Test 1.1b: ECM-10.0 with maximum MTT depth plus 1

Tradeoff with encoding time is far from attractive.

It was reported by cross-checkers that similar benefit can be achieved by an encoder-only change (related contribution JVET-AF0280).

Minimum block size is 8x4/4x8 for luma.

Further study was recommended – no action was taken on this at this meeting.

Test 2 - Intra prediction

Test 2.1: DIMD merge (JVET-AF0120)

In this method, a merged Histogram of Gradient (HoG) is introduced which includes HoG of neighbouring blocks coded with DIMD or DIMD merge modes. When a single DIMD or DIMD Merge neighbouring block is available, then its histogram of gradients is used to form the merged HoG for the current block. If more than one DIMD or DIMD Merge neighbouring blocks are available, the corresponding histograms are combined by means of amplitude averaging to derive the merged HoG. Up to maximum 13 neighbour CUs around the current block are considered.

Finally, the merged HoG is used to compute intra-prediction modes and weights, as in conventional DIMD. The directional modes and their weights corresponding to the five highest amplitudes in the merged HoG are selected, and the corresponding predictors are blended as in conventional DIMD.

Test 2.1a: DIMD Merge.

Test 2.1b: DIMD Merge with reduced storage. Neighbour HoGs are combined by averaging the five highest amplitudes from their respective histograms. Only the five highest amplitudes are stored for a DIMD CU.

Worst case memory analysis is performed, considering that a picture is fully partitioned into 4x4 luma blocks and all blocks are encoded using DIMD.

Class

Number of CUs

Test 2.1a memory

[Bytes]

Test 2.1b memory

[Bytes]

A1/A2

960

515632

38480

B/F

480

258352

19280

C

208

112560

8400

D

104

56816

4240

E

320

172592

12880

It was asserted that in CTC All Intra, the actual memory requirements measured correspond to a 1.1% increase for Test 2.1a, and a 0.2% increase for Test 2.1b.

In CTC Random Access, the actual memory requirements measured correspond to a 0.8% increase for Test 2.1a, and a 0.2% increase for Test 2.1b

Test 2.2: DIMD with filtered template (JVET-AF0131)

In the test, a DIMD mode with filtered template is introduced, where the template is filtered using 3x3 filter before deriving HoG.

The 3x3 filter operator is used to filter the template as follows.

In Test 2.2a, the following 3x2 and 2x3 gradient operators are used for the gradient histogram derivation of the left and top templates respectively.

and

and

In Test 2.2b, the following 3x3 Sobel gradient operators are used, which are the same as in the ECM, while using unfiltered samples for gradient computation at positions where filtered samples are unavailable.

and

Test 2.3: Combination of DIMD related tests (JVET-AF0126)

Test 2.3a: Test 2.1a + Test 2.2a

Test 2.3b: Test 2.1a + Test 2.2b

Test 2.3c: Test 2.1b + Test 2.2a

Test 2.3d: Test 2.1b + Test 2.2b

Test 2.4: IntraCIIP as additional mode of IntraTMP (JVET-AF0130)

In this new intra prediction mode, a regular intra prediction derived by TIMD is blended with IntraTMP prediction using the existing CIIP blending of ECM-10.0 as shown in the next figure.

CIIP Blending

Regular Intra Prediction

IntraTMP prediction

TIMD process

IntraTMP process

Intra mode

Block vectors

A flag is signalled to indicate this mode as a submode of IntraTMP, and IntraCIIP is only allowed when IntraTMP fusion, IntraTMP filtering and IntraTMP subpel are not applied.

Test 2.5: TIMD with IntraTMP/IBC (JVET-AF0136)

In the test, block vectors from neighbouring blocks, which are coded by IntraTMP or IBC, are used to compute predictor candidates for TIMD cost evaluation. In TIMD mode selection process, based on template SATD cost, additional predictors obtained using block vectors from both spatial and non-adjacent neighbours are utilized. The modified TIMD process can therefore result in combining regular and IntraTMP/IBC predictions or combining two different IntraTMP/IBC predictions.

Test 2.5a: TIMD with IntraTMP/IBC.

Test 2.5b: Test 2.5a + Test 2.4

Test 2.6: IntraTMP block vector storing (JVET-AF0079)

In the ECM, IntraTMP BVs are stored for the purpose of coding future IBC blocks in integer pel resolution, even though some of the IntraTMP BVs may have a quarter-pel resolution. In addition, IntraTMP BV is not stored for HMVP. Contrary to IBC that if the current block is coded in IBC mode, the IBC BV is always stored in a quarter-pel resolution and IBC BV is also stored for HMVP.

In this test, IntraTMP BV is stored in a quarter-pel resolution, and quarter-pel IntraTMP BV is also stored for HMVP for coding future blocks.

Test 2.6a: Fractional-pel IntraTMP BVs are stored in block vector buffer.

Test 2.6b: IntraTMP BVs are stored for HMVP.

Test 2.6c: Test 2.6a + Test 2.6b.

Test 2.7: Extrapolation filter-based intra prediction mode (JVET-AF0080)

The extrapolation filter-based intra prediction mode consists of three steps:

  1. The extrapolation filter coefficients are derived from a neighbouring reconstructed area of the current block or filter shape and coefficients are inherited from a previous block coded with this mode. For the latter, a mode flag is signalled, and candidate list is constructed using spatial and non-adjacent blocks coded with this mode, initial list of 12 candidates is reduced to 6 after reordering based on SAD cost measured on the L-shape template of size 1. An index is further signalled to identify a candidate from the list.
  2. The extrapolation process generates predicted signals from the top-left to bottom-right corner within the current block.
  3. An intra prediction angle is derived by analyzing the gradient of the predicted block, and then the corresponding intra mode is used to select MTS, NSPT, and LFNST kernels for transform.

The mode is restricted to blocks with sizes not greater than 32x32 and luma component only.

Three filter shapes are used in this method, a choice of reconstructed area and filter shape shown in the next figure is signalled.

A black and white grid

Description automatically generated

A black screen with white squares

Description automatically generated

The selected filter shape moves in the selected reconstructed area with a one-sample step either horizontally or vertically to collect input samples and output samples, then the filter coefficients are derived using CCCM solver.

The merge mode is also introduced, where the filter shape and the filter coefficients are inherited from the previous decoded blocks that are coded with the tested extrapolation filter-based intra prediction mode or the merge mode based on this mode. In the merge mode, the positions and inclusion order of the spatial adjacent, temporal, non-adjacent, shifted temporal, and history candidates are the same as those defined in ECM-10.0 for the CCP merge prediction candidates. In addition, the merge candidates are reordered by comparing the SAD cost on an L-shape template with column width and row height equal to 1. In the SAD calculation, predictions of the template area by the mode filters are generated only from reconstructed (neighbouring and template) samples, allowing the filters to be applied in parallel rather than sequentially.

In the current block, the recursive predictor is derived from the top-left to the bottom-right position by a diagonal prediction order as shown in the next figure, where the predicted results of the previous diagonal are used.

A diagram of a graph

Description automatically generated

The predicted samples are calculated as follows,

where is the predicted value at (x, y) in the current block, is the coefficient of the selected filter, the index of the coefficients is from 0 to 14, is a reconstructed or a predicted value used for the current position’s prediction, and are the position offsets to the current position along x and y directions, respectively.

Test 2.8: DBV signalling modification (JVET-AF0081)

In this test, the signalling of chroma DM mode and DBV mode is changed. When one of the five corresponding luma positions is coded with either IntraTMP or IBC and DBV mode flag is not true, DM flag is skipped, the binarization of chroma prediction mode for this case is shown in the next table.

intra_chroma_pred_mode

bin string

chroma intra mode

0

11100

list[0]

1

11101

list[1]

2

11110

list[2]

3

11111

list[3]

4

110

DIMD chroma

5

10

DM

65

0

DBV

Test 2.9: Enable DBV in single tree (JVET-AF0066)

In this test, the constraint on DBV flag not being used for single tree case is removed.

Test 2.9a: Enable DBV in single tree under CTC.

Test 2.9b: Enable DBV in single tree when DualITree is set to 0. Single tree is enabled for the anchor and the test.

Test 2.10: Combination test of Test 2.8 and Test 2.9 (JVET-AF0207)

Test 2.10: Test 2.8 + Test 2.9a

Test 2.11: Combination test of Tests 2.1, 2.2, 2.3, 2.4, and 2.5 (JVET-AF0229)

Test 2.11a: Test 2.1b + Test 2.5b

Test 2.11b: Test 2.2b + Test 2.5b

Test 2.11c: Test 2.3c + Test 2.5b

Test 2.11d: Test 2.3d + Test 2.5b

Results (AI/RA)

2.1-2.5 and combination 2.11 (2.11d is the best, combining 4 different elements) is giving only close to 0.2% luma gain, but increases processing needs both at encoder and decoder (also reflected in run times) – no attractive tradeoff. 2.6x is targeting unification of the BV storage between IntraTMP and IBC (IBC already using fractional pel). Though such a design polishing is not overly important in the exploration, it also gives a very small coding gain for screen content. Support for 2.6c was expressed by several experts. Decision: Adopt JVET-AF0079 test 2.6c.

2.7 has higher gain than original proposal. The proponents explain that an EIP merge mode was added. Concern is raised that this is new element proposal that was not originally planned in the EE. Furthermore, the previous proposal (test 2.8 from JVET-AE0024) had an encoder run time increase of 2.5% with 0.16% gain, and it had been requested to reduce the run time. In the new 2.7, the run time has further increased to 3% by adding the new element. Though the gain was also increased to 0.24%, this was not the original intent given by the EE definition. Further study in the next EE was requested, to reduce the encoder run time of both the previous proposal and the extension with the new merge mode. It was also commented that the RA performance looks more attractive with the added element.

The signalling modification from 2.8 does not have significant benefit (even less in the combination 2.10 (2.8/ 2.9), enabling DBV for single tree). Test 2.9, enabling DBV with single tree (in CTC only relevant for RA) is asserted to be straightforward, likely simpler than for dual tree where it is already enabled).

Decision: Adopt JVET-AF0066, test 2.9a.

Test 3 - Inter prediction

Test 3.1/3.3: CCP merge mode for chroma inter coding (JVET-AF0073)

In CCP merge mode, cross-component model parameters of the current chroma block can be inherited from a selected candidate in the CCP merge list comprising spatial adjacent, spatial non-adjacent, history-based, and default candidates. For P- and B- slices, temporal, shifted temporal, and InterCCCM candidates are further added.

In the tests, CCP merge is extended to inter coded blocks, and this mode is indicated by a flag. If applied, a CCP model is implicitly selected from a CCP merge list. The final prediction of the current chroma inter block is formed by combining the motion-compensation predicted signals and the cross-component predicted signals derived using the selected CCP model. The weights for combining predictions are fixed as (wCCP, winter) = (3/4, 1/4).

CCP merge list construction

In addition to the existing candidates of CCP merge mode, the CCP merge list includes a second type of shifted temporal candidates and the CCP models in the list can be inherited from intra and inter blocks. After constructing the CCP merge list, the candidate with the lowest template cost is implicitly selected.

Shifted temporal candidates

While positions of the first type shifted temporal candidates are derived based on neighbouring motion vectors, the positions of the second type are derived based on the current motion vector. As depicted in the next figure, the position of the collocated block and the positions of C0i and C1i, 1 ≤ i ≤ 10, are shifted by a motion vector from the current block, where C0i and C1i are the positions of the temporal candidates.

A screenshot of a game

Description automatically generated

Inheritance of CCP models from inter blocks

CCP models can also be inherited from inter blocks, in addition to models from intra blocks coded in CCLM, MMLM, CCCM, GLM, chroma fusion, and CCP merge modes. For each inter block, a CCP model is stored after the block is coded in inter CCP merge mode or InterCCCM, or retrieved from a reference position located by the motion vector of the inter block following the rules of intra prediction mode propagation.

Derived candidates

A CCP model can also be derived based on the neighbouring reconstructed samples of the current block without accessing the luma reconstructed samples of the current block. When any of the derived CCCM/CCLM model parameters of Cb and Cr is non-zero, such derived candidate is inserted at the beginning of the CCP merge list. At most 4 derived candidates, including single-model CCCM, multi-model CCCM, single-model CCLM, and multi-model CCLM, are inserted to the CCP merge list for an inter CCP merge mode.

Test 3.1b: The CCP merge list consists of all candidates in CCP merge mode for intra (i.e. spatial adjacent, temporal, spatial non-adjacent, history-based, shifted temporal, and default candidates) and no additional candidates are included.

Test 3.1a: Test 3.1b with the second type of shifted temporal candidates.

Test 3.3b: Test 3.1b (without temporal and shifted temporal candidates) with at most 4 derived candidates.

Test 3.1d: Test 3.1a + Test 3.3b with at most 1 derived candidate added in the front of the CCP merge list for low-delay pictures.

Test 3.2: LIC flag derivation for merge candidates with template costs (JVET-AF0128)

In ECM, the LIC flag is inherited for a merge candidate. In the test, the LIC flag is derived for a merge candidate based on template costs by comparing two template costs: a SAD-based template cost, denoted as C0, and a Mean Removal SAD (MRSAD)-based template cost, denoted as C1. The LIC flag is set to be false, if C0 <= C1 and is set to be true, if C0 > C1.

To favour the inherited LIC flag, C0 is multiplied by α if the inherited LIC flag is false while C1 is multiplied by α if the inherited LIC flag is true, where α < 1.

Test 3.4: Enhanced subblock-based motion compensation (JVET-AF0163)

Three methods are tested for subblock-based motion compensation enhancement.

TM-based subblock motion refinement

CPMVs of uni-predicted affine merge candidates and the motion shift of SbTMVP candidates are refined using TM. For a uni-predicted affine merge candidate, a same MV offset is assigned to all the CPMVs, and the TM cost of the affine candidate is calculated accordingly. The optimal CPMV offset with the minimum TM cost is used to refine the corresponding affine candidate. For a SbTMVP candidate, the initial motion shift is refined with TM, and then the refined motion shift will be utilized to derive subblock temporal motion information.

Interweaved affine prediction

With the interweaved prediction, a coding block is divided into subblocks with two different dividing patterns. The first dividing pattern (i.e., pattern 0) is the same as that in ECM, while the second dividing pattern also divides the coding block into 4×4 subblocks but with a 2×2 offset as pattern 1 shown in the next figure.

Then the two auxiliary predictions are generated by affine motion compensation with the two dividing patterns. The final prediction is calculated as a weighted sum of the two auxiliary predictions.

As shown in next figure, weights are positions dependent, an auxiliary prediction sample located at the center of a subblock is associated with a weighting value 3, while an auxiliary prediction sample located at the boundary of a subblock is associated with a weighting value 1.

The subblocks associated with two dividing patterns are respectively refined by PROF unless the width or height is smaller than 4. Besides, subblock boundary deblocking/OBMC for affine mode is disabled.

RMVF candidate derivation with multiple CUs

Additional RMVF affine candidates are derived by taking the motion vector field of multiple CUs as regression input. In the test, two previously affine-coded CUs are simultaneously used to generate a RMVF candidate. The additional RMVF candidates are reordered together with other RMVF candidates through ARMC.

Test 3.4a: TM-based subblock motion refinement.

Test 3.4b: Interweaved affine prediction.

Test 3.4c: RMVF candidate derivation with multiple CUs.

Test 3.4d: Test 3.4a + Test 3.4b + Test 3.4c.

Test 3.5: DMVR with robust MV derivation (JVET-AF0057)

It was asserted that multi-pass DMVR can sometimes produce dislocated subblocks due to use of unreliable motion vectors which leads to subjective artifacts. Two approaches (normative and encoder only) are tested.

Normative approach

Select motion vector that minimize distortion based on both subblock boundary distortion and subblock matching distortion when there is a risk to have unreliable motion vectors based on checks on spatial activity, subblock motion vector differences and boundary differences. The method consists of following steps:

  1. Check spatial activity for the reference subblocks centered at the block motion vector. If it is determined that the spatial activity is lower than a threshold go to step 2.
  2. Each candidate motion vector for the subblock is compared with a neighbouring subblock’s motion vectors across a subblock boundary (if there exists). If the absolute difference between the candidate motion vector and the neighbouring subblock’s motion vector for one component is greater than a threshold go to step 3.
  3. A boundary check which is based on neighbouring samples is made for the top respectively the left subblock boundary. If the boundary check indicates that there not is a true edge along the top subblock boundary or the left subblock boundary continue to step 4.
  4. Determine subblock boundary distortion across left and/or top subblock boundary. For any subblock boundary distortion which is greater than a threshold, add the subblock boundary distortion to the block matching distortion and go to step 5.
  5. Select motion vectors that minimize the total distortion.

Encoder only

In the encoder only solution the RD cost is set to a maximum value when at least one subblock of a coding unit has unreliable motion vector based on steps 1 to 4 as in the normative method following steps without adding the boundary distortion to the selection criteria. Furthermore, the RD cost is not punished for the highest temporal layer and the method is only used for QP equal to or above 33 to reduce BDR impact and still maintain subjective quality.

Test 3.5a: DMVR with robust MV derivation

Test 3.5b: DMVREncSelect from VTM (encoder only)

Test 3.5a*: DMVR with robust MV derivation with QP 40, 43, 47, and 50

Test 3.5b*: DMVREncSelect from VTM (encoder only) with QP 40, 43, 47, and 50

Test 3.6: Affine subblock BDOF refinement (JVET-AF0159)

In ECM, BDOF refinement is not applied to subblocks in affine and SbTMVP modes, and AMVP-Merge mode is also not used with subblocks.

In the test, when BDOF condition is satisfied, BDOF subblock MV refinement and sample adjustment is applied to an affine or SbTMVP coded block with subblock MC.

An affine coded block, e.g. affine regular merge mode, affine BM merge mode, affine AMVP mode, derives MVs for each 4×4 subblock from the affine model. The BDOF process starts with the 4×4 subblocks grouping with identical MVs. The first iteration of BDOF MV refinement is processed in 8x8 subblock grid as in ECM-10.0. When the grouped subblock size is less than 256, the second iteration of BDOF MV refinement is processed in 4×4 subblock grid, and otherwise in 8×8 subblock grid. When the grouped subblock size is 4xN or Nx4, the first iteration of BDOF MV refinement is bypassed.

The BDOF enabling condition is the same as ECM-10.0, e.g., two reference pictures have equal POC distance to the current picture, and equal weight prediction.

AMVP-Merge mode applied to affine blocks is also tested. Like the AMVP-Merge mode in ECM-10.0, it signals reference picture list index, reference index and MVP index to indicate the AMVP predictor. A flag is signalled to indicate AMVP-Merge mode followed by affine mode flag if block size is equal or greater than 8×8 and the current picture is not a low-delay picture.

Test 3.6a: Affine subblock BDOF refinement.

Test 3.6b: AMVP-merge mode for affine.

Test 3.6c: Test 3.6a + Test 3.6b.

Results (RA/LB)

Tests 3.1/3.3: Benefit to be expected for chroma. The methods are somewhat different in selection of candidates. Test 3.1d appears most attractive, though mainly for LB.

It was commented that the gain in chroma was significantly decreased, probably due to the recent adoption of CCRM. However, also the encoder had been significantly higher (several percents), and it had been requested to be reduced, which was achieved.

It was asked what the term “low delay picture” means. It was answered that all pictures in the RPL have a POC that is before the current picture. In the case, no derived candidate is put into the list (which explains that 3.1d has no additional gain over 3.1a in RA case).

Support by non-proponents was expressed (including some of the cross-checkers).

Decision: Adopt JVET-AF0073 test 3.1d.

Test 3.2 is asserted to provide reasonable tradeoff, straightforward implementation.

Decision: Adopt JVET-AF0128 test 3.2.

Tests 3.4: Most of the benefit seems to come from 3.4a, whereas 3.4b/c do not provide much standalone. Also in combination, the additional benefit is small. Some support was expressed by independent experts.

Decision: Adopt JVET-AF0163 Test 3.4a.

Test 3.5: The purpose is avoiding visual artifacts caused by DMVR which had been observed in the context of VTM verification tests (in VTM, that had to be resolved by an encoder modification). The higher QP were tested, because more artifacts would appear in that case. It was commented by proponents that deblocking cannot resolve the problem of edge discontinuity.

It was asked if parallel processing of subblocks was still possible with the normative solution? To some extent yes in terms of the MV search, but the decision has to be done sequentially based on neighbour SB samples.

At this point of exploration, it is not urgent to go for a normative solution on this aspect. In case of a standardization, the normative solution is probably the best way to go (but even better solutions might be possible than the one investigated in test 3.5a).

Decision (SW): Adopt JVET-AF0057 test 3.5b. Not enabled in CTC, but when visual testing with ECM is performed, it should be enabled.

Test 3.6a has a good tradeoff and is straightforward (applying BDOF for affine SBs). Compared to that, the additional benefit of test 3.6b is rather small. The increase in encoding runtime in the combination is likely caused by the fact that BDOF has then also to be executed in the affine merge mode check.

It was commented that no padding is performed in ECM, therefore the combination of BDOF with affine should be straightforward.

Decision: Adopt JVET-AF0159 Test 3.6a.

Test 4 - RPR

This section was discussed 1340 to 1410 on Saturday 14 Oct, 2023, chaired by Y. Ye.

In RPR tests, two PSNR calculations are used:

  • PSNR1: it is measured on the decoded picture size by calculating MSE for each picture, accumulating MSEs for all pictures and calculating average PSNR from the accumulated MSE.
  • PSNR2: it is measured after upsampling of the decoded picture if the decoded picture size is in smaller resolution, the existed ECM interpolation filters are used for the upsampling.

Test 4.1: Enabling template-based reordering and LIC for scaled reference pictures (JVET-AF0190)

In the tests, TM based tools are enabled for RPR.

In Test 4.1a, the template-based inter reordering tools, including ARMC, MMVD/affine MMVD reordering, template-based BCW derivation, reference picture list reordering and MVD sign and magnitude-suffix prediction, are enabled for scaled reference pictures in the RPR.

In Test 4.1b, on top of Test 4.1a LIC is additionally enabled when any of reference pictures is in different resolution comparing to the current picture.

Test 4.1 performance (RA/LB)

It was asked why the impact on encoding/decoding time is so small due to the enabling of TM and/or LIC tools. One expert commented that this is because only 5% of the pictures are impacted by the proposal.

Multiple experts including the cross checker commented that the proposed changes are straightforward, and code change is relatively small, whereas the performance vs. complexity tradeoff is very favourable.

Comparing 4.1a and 4.1b, the latter adds LIC tool on top of the TM tool in the case of RPR. Multiple experts commented that 4.1b provided more favourable gain vs complexity.

Decision: Adopt JVET-AF0190 test 4.1b.

Test 4.2: Weighted edge enhancement filtering (JVET-AF0173)

Additional processing step, adaptive locally weighted filter, is added after the upsampling.

The inputs are the decoded, upscaled picture or block and the side information which is encoded in an adaption parameter set in the bitstream. First, the filter parameters are decoded. This includes information regarding the weighting map function, filter shape, luma and chroma flags, and filter coefficients.

Next, the weighting map function is applied to obtain the local weighting map, which is computed such that the strength of the filter is increased at the location of edges and decreased, or set to zero, in parts of the picture with small local gradient.

Then, the filter is applied to the upscaled picture or block to generate a filtered map. The filtered map is multiplied by the local weighting and added to the upsampled picture as illustrated in next figure.

Filter coefficients are derived by minimizing the difference between the upsampled s and original g (non-donwsampled) signals.

The filters have an adaptive shape following the available shapes shown in the next figure and filter symmetries of ALF. The shape is signalled for each picture in the adaption parameter set.

In Test 4.2a, the adaptive locally weighted filtering is applied to inter prediction only but not the upscaled output for PSNR calculation.

In Test 4.2b, the filtering is also applied after reconstruction which affects PSNR calculation.

Results are not complete, available PSNR2 results for LDB are as follows: RPR 1.5x, test 4.2a

Class C

0.00%

0.00%

0.01%

131%

117%

Class E

-0.02%

0.05%

-0.02%

126%

114%

RPR2x, test 4.2a

Class C

-0.02%

-0.05%

-0.02%

129%

117%

Class E

0.00%

0.00%

0.00%

129%

116%

RPR1.5x, test 4.2b

Class C

-3.12%

-0.42%

-0.43%

132%

151%

Class E

-6.21%

0.29%

0.89%

129%

143%

RPR2x, test 4.2b

Class C

-4.03%

-0.69%

-0.70%

135%

148%

Class E

-12.58%

-1.11%

0.65%

134%

152%

It was commented that currently the normative aspect in 4.2a is not showing much gain, but the non-normative post processing filter in 4.2b on top of 4.2a is showing more significant coding gain. The proponent suggests that this is mainly because the RDO process is not yet very effective in 4.2a, and if the post processing filter in 4.2b would be brought “in-loop”, it would provide more useful reference for RPR prediction, and could further improve the performance from the normative aspect.

The proponent commented that 4.2a has more potential for gains, which they intended to show in a future proposal. Such a contribution would be welcomed.

Futher study was encouraged.

Test 5 - In-loop filtering

Test 5.1: Dynamic scaling of bilateral filter (JVET-AF0112)

In ECM, the bilateral filter offset applied to a sample is calculated using 5x5 diamond filter shape, which is sum of 12 offsets:

S0,2

S-1,1

S0,1

S1,1

S-2,0

S-1,0

S0,0

S1,0

S2,0

S-1,-1

S0,-1

S1,-1

S0,-2

where is based on a 26x16 LUT of 8-bit integers denoted by (for 26 QPs from 17 to 42 and 16 intervals of sample difference). For the innermost positions, i.e., , the LUT is used as-is, i.e., . For the remaining positions, the LUT output it right shifted: .

The TU-based scale factor is defined from minimum block size as follows.

TU scale factor in ECM-10.0

Prediction mode

Inter prediction

Intra prediction

In the test, the LUT and calculation changes are summarized as follows:

  1. The TU scale factor depends on the TU shape size and the mean absolute difference (MAD) of the TU,
  2. The BIF LUTs are interpolated.

Three scale factors () are used to pre-compute three LUTs for three different neighbour distances , i.e.:

Test 5.1a additionally uses the averaging linear interpolation

where and are successive entries of . For chroma, the right shift is decreased from 3 to 2.

The right shift in the formula for BIF offset calculation is increased from 5 to 8 and

where is based on the TU’s shape sizes and is based on the mean absolute difference (MAD) of the TU. Both and are calculated using LUTs. In total, four 64-byte tables and four 16-byte tables are introduced (for luma/chroma component, for intra/inter prediction).

Complexity assessment of the tests against ECM-10.0

BIF version

Bit depth

(maximum)

Summations

(per sample)

Multiplications

(per sample)

LUT lookups

(per sample)

LUTs memory

(bytes)

ECM 10.0

12

18

0

6

832

Test 5.1a

15

25

1

12

2816

Test 5.1b

15

19

1

6

2816

Test 5.1a: Dynamic TU scale factor for BIF with LUTs interpolation.

Test 5.1b: Dynamic TU scale factor for BIF.

Test 5.2: Luma residual taps in chroma ALF and CCALF (JVET-AF0197)

In ECM, chroma ALF uses a diamond 9x9 filter shape with 20 taps in total. CCALF filter uses a cross-liked filter shape with 24 taps in total, and all these taps take the spatial luma reconstruction as input.

In the test, all the spatial based taps in chroma ALF and CCALF are kept unchanged. For chroma ALF, only one luma residual tap is added, which takes down-sampled luma residual as input.

图示

描述已自动生成

For CCALF, five luma residual taps in a cross 3x3 shape are added. The extended taps take the co-located and neighbouring luma residual values as input.

图示

描述已自动生成

Test 5.2a: Luma residual tap in chroma ALF.

Test 5.2b: Test 5.2a + luma residual tap in CCALF.

Test 5.1: It was asked how “MAD_TU” is computed: It is the mean absolute difference between the mean value of TU and the samples in it.

Run time and partial results confirmed by cross-checker.

Test 5.1a is slightly more complex than 5.1b (in terms of processing steps), but it also has more compression benefit.

Decision: Adopt JVET-AF0112 Test 5.1a.

Test 5.2a/b affects only chroma quality. However, it was critizized that separate results for CCALF modification are not provided (test 5.2b only tests combination with 5.1a). Gain of the combination is significantly higher for LB.

Several experts expressed support for 5.2b, emphasizing the benefit in compression which comes at practically no increase in run time. One concern was raised that some optimization at the encoder is frame size dependent (but it was also said that a similar approach also exists in CCSAO).

Considering that the standalone gain of test 5.2a is relatively small compared to the combination 5.2b, it can be concluded that most of the benefit in 5.2b (in particular for LB) comes due to the CCALF modification.

Decision: Adopt from JVET-AF0197 the part of luma residual tap in CCALF test 5.2b.

Test 3 - Entropy coding

Test 6.1: Spatial CABAC tuning (JVET-AF0109)

In ECM, CTUs are scanned in raster scan order, while CUs inside CTUs are scanned in nested tree order. When starting to process the current CTU (CTUC) the processing “jumps” from the bottom-right CU (yellow CU 8) to the top-left CU of the current CTU (blue CU 0) as shown in the next figure. At that stage, the context parameters of the CABAC engine can be expected to be finetuned for the bottom-right area of the previous CTU instead of the ideal case where the context parameters would be finetuned for the actual next CU to be processed.

In the tested method, syntax elements bins of the bottom CUs of the above CTU (the green CUs in the figure) are used to tune the CABAC contexts when starting to process a new CTU.

In the implementation, the certain number of bins are buffered for each CTU of a CTU line, where each entry contains a bin and 10-bit context ID associated with that bin. The buffer size is 512 in Test 6.1a and 256 in Test 6.1b.

The following table summarizes the memory requirement per CTU for the bin buffers.

Test

Bits /
element

Number of
elements

Total bits /
CTU

Total bytes /
CTU

Test 6.1a

11

512

5632

704

Test 6.1b

11

256

2816

352

For each context, no more than 8 bins are stored, so 4-bit counter is used. The number of contexts in ECM-10.0 is 852 making the needed size of these counters 4 * 852 bits = 3408 bits = 426 bytes.

The total worst-case memory requirement then depends on the number of CTUs in a CTU row and summarized as follows:

Class

CTUs /
row

Test 6.1a
memory
[Bytes]

Test 6.1b
memory
[Bytes]

A1

15

10986

5706

A2

15

10986

5706

B

15

10986

5706

C

7

5354

2890

D

4

3242

1834

E

10

7466

3946

F

15

10986

5706

TGM

15

10986

5706

Test 6.1a: Spatial CABAC tuning with 512 buffer.

Test 6.1b: Spatial CABAC tuning with 256 buffer.

Test 6.2: Updating I-slice context model parameters (JVET-AF0133)

The context model parameters have not been retrained since ECM-5.0. In this test, the context model parameters for I-slice are retrained. The training scripts are modified from the one used in VVC’s CE on CABAC initialization. The training bitstream are generated using ECM-10.0 and CTC sequences (including class A1, class A2, class B, class C, class D, class E and class F) with QP 17 to QP 42 in AI configuration.

Test 6.3: Combination of Test 6.1 and Test 6.2 (JVET-AF0110)

Retraining of the contexts Test 6.2 is added first and then the spatial CABAC tuning from Test 6.1 is applied.

Test 6.3a: Test 6.2 + Test 6.1a.

Test 6.3b: Test 6.2 + Test 6.1b.

It was suggested that a regular retraining of CABAC initialization within each meeting cycle would be beneficial. Currently, this was only possible for I slices (as per 6.2). A script is provided, which however is not easy to understand according to crosscheckers.

It was asked why the rate offset parameter was not trained? Because those are shared between intra and inter, such that it would introduce some inconsistency as long as inter initialization is not retrained.

It was agreed that the approach from test 6.2 is a reasonable step forward to perform the re-training, even though futher improvements appear possible and can certainly be investigated in upcoming experiments.

Decision: Adopt the CABAC initialization parameters from JVET-AF0133 Test 6.2. Also the script should be included in the ECM package, such that it can be used by other experts.

The benefit of the 6.1 tests is relatively small, and it has some impact on the memory complexity of the CABAC engine. The combination test with 6.2 had been suggested from the previous meeting to understand possible interaction with an optimized CABAC initialization. This seems not to exist for AI, but some inconsistency is observed e.g. for LB in test 6.3b. Considering that new input is available at this meeting for performing the update of CABAC init also for inter cases, further study in the next EE would be valuable in that context.

EE2 contributions: Enhanced compression beyond VVC capability (24)

There was no presentation or discussion about specific proposals in this category.

For actions decided to be taken, see section 5.2.1, unless otherwise noted.

Decisions
adopted
Adopt the CABAC initialization parameters from JVET-AF0133 Test 6.2. Also the script should be included in the ECM package, such that it can be used by other experts
Citation