Back to Search Document details
31st Meeting: Geneva, CH, July 2023 2023-07-16 19:46
EE2: Summary report of exploration experiment on enhanced compression beyond VVC capability
Abstract
This document provides a summary report of Exploration Experiment on Enhanced Compression beyond VVC capability. The tests are categorized as partitioning, intra prediction, inter prediction, transform and coefficient coding, and in-loop filtering.
JVET-AE0024 EE2: Summary report of exploration experiment on enhanced compression beyond VVC capability [V. Seregin, J. Chen, R. Chernyak, K. Naser, J. Ström, F. Wang, M. Winken, X. Xiu, K. Zhang (EE coordinators)]

This document provides a summary report of Exploration Experiment on Enhanced Compression beyond VVC capability. The tests are categorized as partitioning, intra prediction, inter prediction, transform and coefficient coding, and in-loop filtering.

The software basis for this EE is ECM-9.0, released at https://vcgit.hhi.fraunhofer.de/ecm/ECM/-/tags/ECM-9.0. ECM-9.0 is used as an anchor in the tests.

Software for EE tests is released in the corresponding branches at https://vcgit.hhi.fraunhofer.de/ecm/jvet-ad-ee2/ECM/-/branches.

Test results can be found in input JVET contributions, cross-check results are uploaded to https://vcgit.hhi.fraunhofer.de/ecm/jvet-ad-ee2/simulation-results if cross-check reports are not submitted as they are optional for EE tests.

List of tests

Tests

Tester

Cross-checker

1 Partitioning

1.1a

Partitioning prediction

Canon

G. Laroche

JVET-AE0132

InterDigital

S. Puri

1.1b

Partitioning prediction with unconditional MTT depth increment

Canon

G. Laroche

JVET-AE0132

InterDigital

S. Puri

1.1c

Encoder only partitioning prediction

Canon

G. Laroche

JVET-AE0132

InterDigital

F. Le Léannec

2 Intra prediction

2.1a

Block vector guided CCCM with IBC BV

Nokia

R. Youvalari

B<>com

M. Abdoli

2.1b

Block vector guided CCCM with IBC and IntraTMP BV

Nokia

R. Youvalari

B<>com

M. Abdoli

2.2a

Extended IBC-GPM with two IBC predictions

Kwai
C. Ma

withdrawn

2.2b

Bi-predictive IBC-GPM

KDDI
Y. Kidani

JVET-AE0169

withdrawn

2.2c

Bi-predictive IBC-GPM

Kwai
C. Ma

KDDI
Y. Kidani

JVET-AE0169

Bytedance

W. Yin

JVET-AE0197

2.3a

IBC BVP-merge

KDDI
Y. Kidani

JVET-AE0169

2.3b

Test 2.3a + bi-predictive IBC merge without IBC-GPM

KDDI
Y. Kidani

JVET-AE0169

Ericsson

R. Yu

JVET-AE0205

Bytedance

W. Yin

JVET-AE0197

2.3c

Bi-predictive IBC merge without + Test 2.3a + Test 2.2c

KDDI
Y. Kidani

Kwai
C. Ma

withdrawn

2.4a

IBC MBVD list derivation

Qualcomm

Z. Zhang

JVET-AE0169

Ofinno

D. Ruiz Coll

2.4b

Test 2.4a + Test 2.3a

Qualcomm

Z. Zhang

KDDI
Y. Kidani

withdrawn

2.4c

Test 2.4a + Test 2.3b + Test 2.2c

Qualcomm

Z. Zhang

KDDI
Y. Kidani

Kwai
C. Ma

JVET-AE0169

Ericsson

R. Yu

JVET-AE0205

2.5a

Filtered IBC

Kwai
H.-J. Jhu

JVET-AE0159

withdrawn

2.5b

Filtering IBC predicted blocks

Qualcomm
B. Ray

JVET-AE0159

withdrawn

2.5c

Test 2.5a + Test 2.5b

Kwai
H.-J. Jhu

Qualcomm
B. Ray

JVET-AE0159

OPPO

Y. Yu

L. Zhang

JVET-AE0288

2.6a

IBC with non-adjacent spatial candidates

Kwai

C. Ma

withdrawn

2.6b

Non-adjacent spatial candidates for IBC

Bytedance

Y. Wang

withdrawn

2.6c

IBC with non-adjacent spatial candidates (Test 2.6a + Test 2.6b)

Kwai

C. Ma

 

Bytedance

Y. Wang

JVET-AE0094

Alibaba

J. Chen

2.7

Cross-component merge mode with temporal candidates

MediaTek
H.-Y. Tseng

InterDigital
P. Bordes

2.8

An extrapolation filter-based intra prediction mode

OPPO

L. Xu

Kwai

H.-J. Jhu

JVET-AE0217

2.9

Extended search areas for IntraTMP mode

OPPO

Y. Yu

Xidian

Y. Ma,

H.Zhang

Kwai

X. Xiu

JVET-AE0077

InterDigital

K. Naser

JVET-AE0210

Qualcomm

P.-H Lin

JVET-AE0216

Alibaba

X. Li

JVET-AE0245

2.9b

Extended search areas for IntraTMP mode with scan order #2

OPPO

Y. Yu

Xidian

Y. Ma,

H.Zhang

Kwai

X. Xiu

withdrawn

2.9c

IntraTMP mode with partial extended search areas with scan order #1

OPPO

Y. Yu

Xidian

Y. Ma

H. Zhang

Kwai

X. Xiu

withdrawn

2.9d

IntraTMP mode with partial extended search areas with scan order #2

OPPO

Y. Yu

Xidian

Y. Ma

H. Zhang

Kwai

X. Xiu

withdrawn

2.10a

IBC-LIC extension without large block-size constraint

OPPO

Z. Xie

JVET-AE0078

Bytedance

W. Yin

JVET-AE0157

2.10b

ECM IBC-LIC without large block-size constraint

OPPO

Z. Xie

JVET-AE0078

Bytedance

W. Yin

JVET-AE0157

2.11a

Harmonization between IBC HMVP and IBC-LIC

Bytedance

N. Zhang

JVET-AE0084

Kwai
C. Ma

JVET-AE0222

2.11b

Test 2.11a + Test 2.10

Bytedance

N. Zhang

OPPO

Z. Xie

JVET-AE0084

Kwai
C. Ma

JVET-AE0222

3 Inter prediction

3.1a

Cross-component residual model

Nokia

P. Astola

JVET-AE0059

InterDigital

F. Le Léannec

Ittiam
J. Raj

3.1b

Cross-component residual model with complexity reductions

Nokia

P. Astola

JVET-AE0059

InterDigital

F. Le Léannec

Ittiam
J. Raj

3.2

Bi-predictive GPM

Ericsson

R. Yu

JVET-AE0046

LGE

Y. Ahn

JVET-AE0199

KDDI
Y. Kidani

3.3a

Additional TM refinement for bi-prediction

Bytedance

Y. Wang

JVET-AE0087

Alibaba

J. Chen

JVET-AE0239

3.3b

16-point diamond search pattern for TM

Bytedance

Y. Wang

JVET-AE0087

Alibaba

J. Chen

JVET-AE0239

3.3c

Enabling TM for bi-prediction under DMVR condition

Bytedance

Y. Wang

JVET-AE0087

Alibaba

J. Chen

JVET-AE0239

3.3d

Test 3.3a + Test 3.3c

Bytedance

Y. Wang

JVET-AE0087

Alibaba

J. Chen

JVET-AE0239

3.3e

Test 3.3a + Test 3.3b + Test 3.3c

Bytedance

Y. Wang

JVET-AE0087

Alibaba

J. Chen

JVET-AE0239

3.4

HPel flag and BCW weight usage in OBMC

InterDigital

A. Robert

JVET-AE0196

withdrawn

3.5

Iterative BDOF pass in multi-pass DMVR

Bytedance

M. Salehifar

Alibaba

J. Chen

JVET-AE0065

vivo

Z. Lv

JVET-AE0200

3.6

Affine AMVP mode with one MVD

Qualcomm
H. Huang

withdrawn

3.7a

RPR with new filters, scale factor 1.25x

Sharp
J. Samuelsson-Allendes

JVET-AE0150

Qualcomm

Z. Zhang

JVET-AE0204

3.7b

RPR with new filters, scale factor 1.33x

Sharp
J. Samuelsson-Allendes

JVET-AE0150

Qualcomm

Z. Zhang

JVET-AE0204

3.8

Combination of Test 3.3e and Test 3.5

Bytedance

M. Salehifar

Alibaba

J. Chen

JVET-AE0091

Kwai

C. Ma

JVET-AE0223

4 Transform and coefficient coding

4.1

Shifting quantizer center

InterDigital

M. Balcilar

JVET-AE0125

Qualcomm

M. Coban

JVET-AE0149

4.2

Large NSPT

LGE

M. Koo

JVET-AE0086

Qualcomm

P. Garus

JVET-AE0118

4.3

Context modelling for transform coefficients for LFNST/NSPT

Qualcomm

P. Nikitin

JVET-AE0102

IRT b-com
M.Abdoli

JVET-AE0225

4.4a

InterMTS is enabled for IBC-coded blocks in AMVP mode.

Qualcomm

P. Garus

JVET-AE0116

InterDigital
K. Naser

JVET-AE0211

4.4b

Test 4.4a + IntraTMP using interMTS instead of intraMTS kernels

Qualcomm

P. Garus

JVET-AE0116

InterDigital K. Naser

JVET-AE0211

4.4c

IntraMTS disabled for IntraTMP

Qualcomm

P. Garus

JVET-AE0116

InterDigital K. Naser

JVET-AE0211

5 In-loop filtering

5.1a

CCSAO with temporal history offset

Kwai

C.-W. Kuo

JVET-AE0151

Qualcomm

N. Hu

JVET-AE0209

5.1b

Test 5.1a + extended edge classifier

Kwai

C.-W. Kuo

JVET-AE0151

Qualcomm

N. Hu

JVET-AE0209

5.2a

Applying fixed filters to samples before DBF

Qualcomm

N. Hu

JVET-AE0139

Kwai

C.-W. Kuo

JVET-AE0193

5.2b

Test 5.2a + extended classifiers for fixed filters

Qualcomm

N. Hu

JVET-AE0139

Kwai

C.-W. Kuo

JVET-AE0193

5.2c

Test 5.2b + applying the second fixed filter to outputs of the first fixed filter

Qualcomm

N. Hu

JVET-AE0139

Kwai

C.-W. Kuo

JVET-AE0193

5.3

Combination of Test 5.1b and Test 5.2c

Kwai

C.-W. Kuo

Qualcomm

N. Hu

JVET-AE0152

Alibaba

J. Chen

Category 1: Partitioning prediction

Test 1.1: Partitioning prediction (JVET-AE0132)

In this test, a temporal partitioning is introduced, where for each block, the allowed partitioning splits are predicted according to the minimum QT/MTT split and the average QT/MTT split obtained from a temporal area:

  • if the current QT depth is less than the temporal minimum QT depth minus 1, only the QT split is allowed,
  • if the current QT depth is less than the temporal average QT depth minus 1, no split, QT and TT splits are allowed, and BT split is allowed if TT is selected in a parent node,
  • if the temporal maximum MTT depth and less than the maximum MTT depth of the current block, the depth is decreased except if the current QP is less than the temporal QP,
  • if the temporal maximum MTT depth is large than the maximum MTT depth of the current block and if the current depth is equal to the current QT depth, the maximum MTT depth is incremented,
  • if the maximum QT depth of the current frame is larger than the maximum multi-tree depth and if the temporal maximum MTT depth is equal to the maximum MTT depth of the current frame, and if the temporal QT depth is less than the QT depth of the current frame, the maximum MTT depth of the current block is decreased.

The modifications on the maximum MTT depth are enabled only if the palette mode is disabled.

The temporal area is a collocated area from the same reference frame used for the temporal motion predictor.

JVET-AD0147 has memory requirement analysis where the largest memory which is required to store temporal depths for class A is estimated as 40.5 KB.

Test 1.1a: Partitioning prediction (as described above)

Test 1.1b: Test 1.1a where MTT depth increment is enabled unconditionally. It affects classes A and B, class A1 encoder runtime reaches 105.8%.

Test 1.1c: Use partition prediction as encoder only optimization.

Questions:

  • Is the partitioning only from the reference frame? A: Yes.
  • What is the additional memory requirement? A: Approx. 40 Kbyte per reference picture in 4K.
  • Does CABAC parsing require the temporal partitioning info? A; Yes.

It was commented that the memory requirement is significantly less than for motion vectors.

The cross-checkers pointed that the implicit splitting performed at the picture boundary is useless (no gain when removing it, but more complicated).

It was commented that for LB there is no gain (or small losses).

Even though there was no strong objection, from the questions raised the implementation of the idea, it appears not mature. The gain is relatively low, even though it could be attractive in terms of run time, it introduces additional temporal dependencies that might be undesirable.

No action was taken on this.

Category 2: Intra prediction

Test 2.1: Block vector guided CCCM (JVET-AE0100)

In this test, a co-located luma BV is used to determine the reference area for calculating the CCCM parameters. Then the reference area in luma and corresponding area in chroma channel is used to calculate the CCCM parameters. The prediction uses the calculated model parameters and co-located luma samples to do the CCCM prediction as follows:

predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P(C) + c6P(N) + c7P(S) + c8P(W) + c9P(E)+ c10B,

where spatial components are as shown in the next figure, the nonlinear term P is represented as power of two of the corresponding luma sample, and B is the bias term.

A picture containing text, shoji, crossword puzzle

Description automatically generated

The mode is enabled only in intra slices and is controlled by SPS flag.

Test 2.1a: the mode only uses BVs of the IBC coded blocks from co-located luma area.

Test 2.1b: the mode can use BVs of both IBC and IntraTMP coded blocks from co-located luma area.

Test 2.2: Bi-predictive GPM (JVET-AE0169)

In ECM-9.0, IBC GPM uses uni prediction for one partition and intra mode for the other one. In the test, the core GPM design is kept unchanged which uses 48 GPM modes and IBC merge candidate list, while it enables both partitions being predicted using IBC with different BVs.

Two flags are signalled to indicate the prediction modes of two partitions, the first flag indicates whether the first partition is intra predicted, and if not then the second flag is signalled to indicate whether intra prediction is used for the second partition.

The method is applied to SCC only.

Test 2.3: IBC BVP-merge and bi-predictive IBC merge (JVET-AE0169)

In Test 2.3a, IBC-BVP-merge, which is similar to AMVP-merge, derives one BV from IBC block vector prediction (BVP) and the second BV from IBC merge to form bi-prediction for IBC. Two different indices for the IBC BVP and the IBC merge candidates are signalled.

In Test 2.3b, bi-predictive IBC merge is introduced, and it is enabled together with the existed in the ECM MBVD and uni-merge (currently disabled by the encoder configuration for non-SCC classes). In bi-predictive IBC merge, two BVs from the existing IBC merge candidate list are derived, utilizing two different indices, which are signalled. Bi-predictive IBC merge is applied to IBC regular merge and IBC MBVD. Bi-predictive IBC merge, IBC MBVD, and IBC uni-merge are enabled for non-SCC classes and is tested together with Test 2.3a.

Test 2.4: IBC MBVD list derivation (JVET-AE0169)

In the test 2.4a, adaptive BVD offsets along MVBD directions and enabled for IBC MBVD mode. The MBVD candidates search is a two-step process, which starts with checking template SAD costs of offsets added to BVP along each direction with the interval of 1-pel. The second step of the search checks template SAD costs with 1/4-pel interval for the candidates around the selected candidates from the first step. For the integer MBVD (when existed in the ECM ph_fpel_mbvd_enabled_flag is 0), those intervals are multiplied by 4. The candidates with the lowest TM cost are included into the final MBVD list.

Test 2.4c is a combination of Test 2.4a, Test 2.3b, and Test 2.2c.

Test 2.5: Filtered IBC (JVET-AE0159)

In the test, additional filtered IBC mode is introduced, where a filter is applied to IBC predictor, which is derived by minimizing MSE between current and reference template.

Output of the filter is calculated as follows:

predLumaVal = c0C + c1N + c2S + c3E + c4W + c5P + c6B

The nonlinear term P is represented as power of two of the center sample C and scaled to the sample value range of the content:

P = ( C*C + midVal ) >> bitDepth

The bias term B represents a scalar offset between the input and output and is set to middle luma value (512 for 10-bit content).

This filtered mode is used as an additional mode for non-merge IBC blocks, and it is not used together with IBC-LIC, IBC-CIIP or RR-IBC. For IBC merge modes, this filtering mode is inherited when merge mode list is constructed.

In Test 2.5a, the mode flag is signalled conditioned on the IBC-LIC flag, so when the IBC-LIC flag is true, the flag is signalled and used to indicate whether the tested mode is applied to the current block or not.

In Test 2.5b, the mode flag is signalled before IBC-LIC flag, and the model does not have non-linear term.

In Test 2.5c, the model of Test 2.5a and signalling of Test 2.5b are used, different aspects of encoder design is chosen from both tests considering the best trade-off.

Test 2.6: IBC with non-adjacent spatial candidates (JVET-AE0094)

In the test, several new candidates obtained from the BVs of non-adjacent spatial neighbouring blocks are added to the candidate lists of IBC merge modes and IBC AMVP. A pattern similar to the non-adjacent spatial candidates of regular inter is used to obtain non-adjacent BVs shown in the below figure. The obtained BVs are inserted between the adjacent spatial candidates and the HBVP candidates for both IBC merge and IBC AMVP. Additionally, the restriction that spatial candidates are not used to construct IBC merge candidate list for a 4×4 CU is removed.

A grid with blue squares and black text

Description automatically generated

In Test 2.6a, the same reference area as for non-adjacent regular merge is reused for the IBC, where up to 4 rows/columns of spatial neighbouring blocks (in terms of the CU height/width) of the CU can be used as reference for IBC AMVP, and up to 7 rows/columns of the spatial neighbouring blocks (in terms of the CU height/width) of the CU are used for IBC merge.

In Test 2.6b, for both IBC AMVP and IBC merge, the spatial neighbouring blocks located up to 4 rows/columns of spatial neighbouring blocks (in terms of the CU height/width) of the CU can be used as reference. Additionally, the restriction that spatial candidates are not used for the IBC merge of a 4×4 CU in ECM-9.0 is removed.

Test 2.6c is the combination of Test 2.6a and Test 2.6b, where the sampling method in Test 2.6a is applied. Additionally, the removal of the restriction that spatial candidates cannot be used for IBC merge of a 4x4 CU from Test 2.6b is used.

Test 2.7: Cross-component merge mode with temporal candidates (JVET-AE0043)

In the current CCP merge mode, a CCP merge candidate list contains spatial adjacent, spatial non-adjacent, and history-based candidates coded in CCLM, MMLM, CCCM, GLM, chroma fusion, and CCP merge modes. After including these candidates, default models can be included to fill the remaining empty positions in the merge list if necessary. To remove redundant CCP models in the merge list, pruning operations are applied. The candidates in the list are reordered based on the SAD costs obtained using the neighbouring template of the current block.

In the test, two types of candidates, namely temporal candidates and shifted temporal candidates, are additionally included in the merge list for non-intra slices. Temporal candidates are added to the merge list after the spatial adjacent candidates. Shifted temporal candidates are added after the history-based candidates.

CCP merge mode is allowed to be applied in non-intra slices for chroma blocks with sizes less than or equal to 16.

Temporal candidates

Temporal candidates are selected from the collocated picture, the positions and inclusion order of the temporal candidates, which are the same as those for temporal candidates in inter merge mode of ECM-9.0, is shown in the next figure below.

A screenshot of a game

Description automatically generated

Figure: Positions of the temporal candidates

An inclusion order of the temporal candidates is C01 🡪 C02 🡪 … 🡪 C010. If C0i is outside of the picture/slice boundary and C1i is inside of the picture/slice boundary, C1i will be used instead of C0i, where 1 ≤ i ≤ 10. Otherwise, the next inclusion position is checked.

Shifted temporal candidates

Shifted temporal candidates are also selected from the collocated picture. As depicted in the next figure below, the position of the collocated block is shifted by a selected neighbouring motion vector. Consequently, the positions of C0i and C1i are shifted by the same neighbouring motion vector. The inclusion order of shifted temporal candidates is the same as that of temporal candidates.

The adjacent motion vector is selected from the motion vectors of the neighbouring blocks ordered as follows: L0B1 🡪 L1B1 🡪 L0A1 🡪 L1A1 🡪 L0B0 🡪 L1B0 🡪 L0A0 🡪 L1A0 🡪 L0B2 🡪 L1B2. The first motion vector, which uses the collocated picture as the reference picture, is selected. If no such motion vector is found, no shifted temporal candidate is added.

A screenshot of a game

Description automatically generated A green square with blue text and red arrows

Description automatically generated

Figure: Positions of the shifted candidates (left) and selecting the neighbouring motion vector (right)

Test 2.8: Extrapolation filter-based intra prediction mode (JVET-AE0076)

In the test, extrapolation filter-based intra prediction is introduced with 3 types of reconstructed areas and 3 filter shapes as shown in the next figure below, a choice of reconstructed area and filter shape is signalled.

The size of reconstructed area depends on the min(blockWidth, blockHeight) and the selected filter shape. For example, when the current block is an 8x16 block and the selected filter shape is 4x4. The aboveSize of reconstructed area is min(8, 16) + 4 – 1 = 11, and the leftSize of reconstructed area is min(8, 16) + 4 – 1 = 11.

A black screen with white squares

Description automatically generated

A screenshot of a computer game

Description automatically generated

Figure: Three types of reconstructed areas (top) and three filter shapes (bottom)

The selected filter moves in the selected reconstructed area with a one-sample step to collect input samples and output samples, then the filter coefficients are derived using CCCM solver.

In the current block, the recursive predictor is derived from the top-left to the bottom-right position by a diagonal prediction order as shown in the next figure below, where the predicted results of the previous diagonal are used.

A diagram of a graph

Description automatically generated

Figure: Example of generating predictions for different positions in the current block by a diagonal order

To reduce the prediction error, the min and max values from the neighbouring reconstructed area are applied to restrict the output range of each predicted value.

The predicted samples are calculated as follows,

where is the predicted value at (x, y) in the current block. and offset are derived in the selected reconstructed area, is the coefficient of the derived filter, is reconstructed or predicted value used for prediction in the current position.

The tested mode is applied in I-slices only and is only used for the block sizes less than or equal to 32x32.

Test 2.9: Extended search areas for IntraTMP mode (JVET-AE0077)

In the test, IntraTMP search areas are extended by including the bottom-left R5 and top-right R6 areas relative to the current block as they may also be available as shown in the next figure below. In ECM-9.0, the region scan order is R4, R1, R2, and R3. In the tested method, the region scan order is R4, R5, R6, R1, R2, and R3.

A diagram of a diagram

Description automatically generated

Figure: Extended IntraTMP search areas (R5 and R6)

In addition, the current ECM-9.0 always searches all available positions within the current CTU while it may be beyond the specified search range or width/height. In the test, all search areas are restricted to be within the specified search range or width/height specified by (searchRangeWidth, searchRangeHeight) as illustrated in the next figure below. The maximum horizontal and vertical search sizes, i.e., searchRangeWidth and searchRangeHeight, are set to max(5W, 64) and max(5H, 64) respectively. It is not allowed to search any positions outside the specified search range even those available positions are within the current CTU.

Figure: Search areas restriction

Test 2.10: IBC-LIC extension (JVET-AE0078)

In Test 2.10a, IBC-LIC is extended by including 3 additional modes:

  • Use only left or use only above LIC template in addition to the current L-shape template,
  • Extend the MMLM to IBC-LIC, which allows IBC-LIC to have two linear models in one CU for L-shape template.

Large block-size constraint for ECM IBC-LIC, tested separately in Test 2.10b, is additionally removed in Test 2.10a.In Test 2.10b, large block-size constraint for ECM IBC-LIC is removed, so IBC-LIC can be applied to the CU whose block size is larger than or equal to 32.

Test 2.11: Harmonization of IBC HMVP and IBC-LIC (JVET-AE0084)

In ECM-9.0, the inter LIC flag can be inherited from a HMVP candidate. However, the IBC-LIC flag is not inherited when the motion information from an IBC HMVP candidate is used.

In Test 2.11a, IBC-LIC flag is inherited from an IBC HMVP candidate, which is similar to the inter LIC case.

Test 2.11b: Test 2.11a + Test 2.10a

Results for AI/RA

Test 2.1 is a chroma tool, main benefit on screen content. The gain is small, but it has no impact on run time, and implementation is straightforward according to cross-checkers. Adoption was also supported by independent experts.

Decision: Adopt JVET-AE0100, test 2.1b.

Tests 2.2..2.4 target several aspects of improving IBC in context with GPM and bi-prediction, 2.3b also provides gain on camera-captured content (but had been reported previously to not give gain in case of screen content). In the combination 2.4c (2.2c+2.3b+2.4a), 2.3b is disabled for screen content, and 2.2c is disabled for camera-captured content, whereas 2.4a is enabled for both. gives best gain overall, including screen content. However, the results show that 2.4a does not have benefit for camera capured content when combined with 2.3b, whereas for screen content the gain from 2.2c and 2.4a is somewhat additive.

Decision: Adopt JVET-AE0169 Test 2.2c, in CTC enabled only for screen content.

Decision: Adopt JVET-AE0169 Test 2.3b, in CTC enabled only for camera-captured content.

Decision: Adopt JVET-AE0169 Test 2.4a, in CTC enabled only for screen content.

For 2.5c, a discussion was performed about the benefit of a combination test when the individual results of the elements being combined are not available. Proponents are strongly discouraged to withdraw proposal elements when still including them in a combination, as this will not allow to judge the individual benefits.

Individual results of 2.5a and 2.5b were presented in an update of the summary report on Wednesday July 12.

Gains appear to be somewhat additive, but encoding run time is significantly higher in 2.5b than it is in 2.5c. According to proponents, an encoder with less RDO checks was used in 2.5c. This is also the likely reason for losing a slight bit rate reduction in 2.5c, compared to 2.5b.

It was asked whether the encoder search strategy of 2.5c is the same as 2.5a, such that it could be confirmed that the additional benefit of 2.5c comes by the signalling method of 2.5b, and its non-linear filter term? According to proponents, main gain comes from signalling.

Results were matched according to cross-checker (verbally reported in session, document to be uploaded).

Decision: Adopt JVET-AE0159 Test 2.5c (enabled only for screen content).

Same situation was initially found with 2.6c.I Individual results of 2.6a and 2.6b were presented in an update of the summary report on Wednesday July 12.

Results indicate that the combination of 2.6c provides more gain than the individual parts 2.6a and 2.6b. Gains cannot be expected to be additive, as the basic approach is identical. It can however be concluded that the element of 2.6b (removing the restriction on 4x4 blocks) is also beneficial in combination with 2.6a. The proposal was supported by independent experts (including cross-checkers).

Decision: Adopt JVET-AE0094 Test 2.6c (enabled only for screen content).

Test 2.7 is targeting improvement of chroma, by adding additional temporal candidates in cross-component merge (which had been adopted in the last meeting). Candidates are stored in an 8x8 grid.

It was asked what the benefit of shifted candidates is. According to proponents, it had been 10% of the total gain in the original contribution. It was argued that this element of the proposal might be difficult to implement in hardware. It was however asserted by JVET that this is not of highest important in this stage of exploration, and technology investigated in EE should normally be adopted as proposed.

Decision: Adopt JVET-AE0043 test 2.7.

Test 2.8: It was asked how the mode is used when reconstructed area is not available. It is disabled. It was further pointed out that the gain seems to be very small for small blocks, where the pixel dependency might be critical in real-world implementation. The cross-checker confirms that such dependency is probably present, but would nevertheless support this tools at this stage of exploration.

2.5% increase of encoder run time is pointed out to be not a good tradeoff for 0.16% bit rate reduction in AI. For RA, the gain is less, as it is only applied in I slices currently.

Further study was recommended (in the EE) to reduce encoder run time, and also apply in inter slices.

Reduction of pixel dependency also would be important, but probably is more difficult to achieve.

Test 2.9: Increase in decoder run time by 2.5% by additional search operations. Encoder run time also slightly increased. One cross-checker points out that the increase of decoder run time may be caused by pixel-wise availability checks. It was also pointed out that this is straightforward extension of search range relative to ECM 9, and the tradeoff between gain and run time is still in an attractive range.

Decision: Adopt JVET-AE 0077 test 2.9.

Tests 2.10/2.11 are only beneficial for screen content. From comparing 2.10a/b, it cannot be concluded that removing the large block size constraint is beneficial, as there are no results for the IBC-LIC extension with the large block size constraint of current ECM (such a test had originally been planned, but was not conducted). Proponents are requested to present results of a variant of 2.10a with the constraint, and also in combination 2.11.b to assess whether the gain is additive with 2.11a.

Test 2.11a is a minor change which gives some benefit in class F without impact on run time. However, it should also be enabled for camera-captured content to avoid another high-level flag (according to proponents, no losses would be expected). Therefore, the new combination test 2.11b should be conducted such that the method 2.11a is also used for camera-captured content.

A revised version of JVET-AE0024 was presented on July 16 1700. From the new results, it became obvious that the large block size constraint had no impact on run time, but results are slightly better without the constraint. From the results, 2.10a or 2.11b (combining 2.10a with 2.11a) would be candidates. According to the cross-checker, 2.11a is straightforward to implement. Also the runtime variations are low.

Decision: Adopt JVET-AE0084 Test 2.11b. In the combination, both 2.10a and 2.11a should be enabled for camera captured content, also in CTC.

Inter prediction

Test 3.1: Cross-component residual model for inter prediction (JVET-AE0059)

Cross-component residual model to predict chroma samples from reconstructed luma samples when the block uses inter prediction or intra block copy (IBC) is tested. As illustrated in the next figure below, the cross-component filters are derived using the prediction blocks of luma and chroma. The derived filters are applied to the reconstructed luma block and blended with the prediction blocks of chroma to produce the final chroma prediction blocks. In the blending process the filtered reconstructed luma blocks use blending weight of 0.75 and chroma prediction blocks use blending weight of 0.25.

A diagram of a computer

Description automatically generated

Figure: Cross-component residual model

Model uses 8-tap filter consisting of 6 spatial luma samples shown in the next figure below, a nonlinear term, and a bias term as follows:

predChromaVal = c0 L0+ c1L1 + c2L2 + c3L3 + c4L4 + c5L5 + c6 nonlinear((L0+L3+1) >> 1) + c7 B,

A white grid with black letters and numbers

Description automatically generated

Figure: 2 Luma samples L0,..,L5 in relation to the chroma sample C.

The filter coefficients are derived using ECM’s division-free Gaussian elimination method and the necessary offsets are applied to samples prior to filter derivation.

For filter coefficient derivation at most 256 chroma samples are used.

The mode flag is only signalled if the TU’s luma Cbf is non-zero and the CU’s predMode is either MODE_INTER or MODE_IBC.

Test 3.1a: Cross-component residual model, where offsets for division free operation are obtained by averaging the samples in the block.

Test 3.1b: Cross-component residual model, where offsets for division free operation are obtained by averaging the four points correspond to the top-left, top-right, bottom-left and bottom-right corners of the blocks.

Test 3.2: Bi-predictive GPM (JVET-AE0046)

In the test, the existing uni-predictive GPM is extended to allow usage of bi-predictive motion vectors for generating motion compensated prediction samples for inter partitions. The method consists of the following elements:

The first element conditionally invokes the existing extraction process that extracts uni-predictive motion vectors from the initial list. The extraction process is invoked only for small blocks 8x8, 16x8 and 8x16. For other larger blocks, the extraction process is bypassed, so the initial list (which may contain merged Bi-MVs) is directly used as the final GPM merge list. The generation of the initial list is the same as before (i.e., the normal merge list generation without any candidate reordering) except that when generating the initial list for larger blocks (i.e., blocks with the extraction process bypassed), the motion vector difference threshold for controlling whether a candidate can be added into the initial list is increased to be one full sample distance.

The second element modifies GPM-MMVD to support bi-predictive motion vector as the base vector. For low-delay pictures, the signalled MVD is applied on top of the L0 and L1 motion vector as in the existing merge MMVD design. For non-low-delay pictures, the bi-predictive motion vector is converted into a uni-predictive motion vector first and then the MVD is applied on top.

The third element modifies GPM-TM to also support bi-predictive motion vectors.

The last element is to enable the BDOF based mv refinement as in the multi-pass DMVR on top of the associated bi-predictive motion vectors for each inter partition.

Test 3.3: High-Accuracy template matching (JVET-AE0087)

In the ECM, template matching is a decoder-side MV derivation method to refine the motion information of the current CU by finding the closest match between a template (i.e., top and/or left neighbouring blocks of the current CU) in the current picture and a block (i.e., same size to the template) in a reference picture.

Diagram

Description automatically generated with medium confidence

When TM is used for bi-prediction, the following steps are applied:

  1. The initial motion vector of list 0 (MV0) is refined to derive a refined MV (MV’0) and a TM cost C0 is calculated;
  2. The initial motion vector of list 1 (MV1) is refined to derive a refined MV (MV’1) and a TM cost C1 is calculated;
  3. If C0 is larger than C1, MV’1 is used to derive a further refined MV (MV’’0) by refining MV’0; Otherwise, MV’0 is used to derive a further refined MV (MV’’1) by refining MV’1 .

In high-accuracy template matching, three aspects are tested:

  • In Test 3.3a, an additional refinement is applied to TM for bi-prediction, when MV’0 is refined in step 3), MV’’0 is used to derive MV’’1 by refining MV’1; Otherwise, MV’’1 is used to derive MV’’0 by refining MV’0.
  • In Test 3.3b, the diamond search pattern used in TM is modified from 8-point to 16-point.

图标

描述已自动生成

  • In Test 3.3c, TM for bi-prediction is enabled when DMVR condition is satisfied.

Test 3.3d: Test 3.3a + Test 3.3c.

Test 3.3e: Test 3.3a + Test 3.3b + Test 3.3c.

Test 3.4: OBMC with HPel flag and BCW weights (JVET-AE0196)

In the current OBMC design, the motion compensation of the neighbouring blocks and subblocks performed by OBMC only considers the motion vectors and reference pictures.

In the test, HPel flags and BCW weights of the neighbouring blocks and subblocks are considered in the OBMC. When processing the top and left borders, the neighbouring BCW weights and HPel flags are used in addition to the motion vectors and reference pictures. The neighbouring HPel flag allows setting the correct AMVR index. For the internal subblock boundaries, the BCW and AMVR indexes of the current CU are also used.

It has to be noticed that the integration of adopted contribution JVET-AD0213 (Bi-prediction LIC) turns out to already include main aspect of this test (BCW weights in OBMC) in the ECM-9.0 which was not the case in ECM-8.0.

Test 3.5: Iterative BDOF pass in multi-pass DMVR (JVET-AE0065)

In the ECM, a multi-pass DMVR has three passes of motion refinement. In the first pass, bilateral matching (BM) is applied to the coding block (CB) to get the block level refined MV pair. In the second pass, the block is divided into 16×16 subblocks and on top of block level refined MV pair obtained in the first pass, BM is applied on each 16×16 subblock to get a 16×16 subblock level refined MV pair. In the third pass, BDOF based MV refinement is applied on the 4×4 or 8×8 subblocks depending on the CB size.

In the test, the current multi-pass DMVR is extended by adding another BDOF-based motion refinement pass as the 4th pass of the multi-pass DMVR (i.e., the 2nd round of BDOF-based DMVR).

Also, the subblock size of the proposed round of the BDOF-based DMVR is adaptively selected depending on the area (i.e., width×height) of the coding block. For blocks smaller than 1024 pixels, the subblock size of 4×4, and otherwise 8×8 is used.

Test 3.7: RPR with new filters for scale factor 1.25x and 1.33x (JVET-AE0150)

In the test, additional set of 8-bit precision RPR filters consisting of three tables (respectively 12-tap for luma, 6-tap for chroma, and 10-tap for affine) is introduced with the target to improve the performance for scaling ratios 1.25x and 1.33x. These new set of filters is used for scaling ratios in between 1.1x and 1.35x.

In the tests, two PSNR values are used:

  • PSNR1: it is measured on the decoded picture size by calculating MSE for each picture, accumulating MSEs for all pictures and calculating average PSNR from the accumulated MSE.
  • PSNR2: it is measured after upsampling of the decoded picture if the decoded picture size is in smaller resolution, the existed ECM interpolation filters are used for the upsampling.

Test 3.8: Combination of Test 3.3e and Test 3.5 (JVET-AE0091)

Results for RA/LB

Test 3.1: Increase of encoder run time should be expected in test 3.1b (RA results not complete yet).

It was commented that the method should better be called “cross component chroma prediction in inter”, as it is not the residual that is predicted.

It was also commented that a blending with 0.75/0.25 weights was newly introduced, which according to proponents provides better performance in blending. It was requested to provide quantitative information about the gain by blending

Results were presented Wednesday 8:30. Test 3.1b has slightly less encoder run time (which is explainable by having less checks), and bit rate reduction is almost identical. This was also confirmed by cross-checker.

It was asked why a different cross-component model was used than the one in the ECM. It was answered by proponent that downsampling is avoided. The unification might be desirable, but is not of prior importance at this moment.

Decision: Adopt JVET-AE0059 Test 3.1b.

Test 3.2: 0.2% bit rate reduction in RA with slight increase in encoding time. According to one cross-checker it may increase the worst-case memory access, which may however not be too severe, as the proposal is not used in case of small block size, and multi-hypothesis is also used in other tools of ECM. Several experts (including cross-checkers) supported the proposal. It was however suggested to modify the software such that the memory for motion vector storage is reduced in the encoder.

Decision: Adopt JVET-AE0046 Test 3.2.

Test 3.3x, 3.5, and combination 3.8: Test 3.3c seems to have no benefit standalone, but according to proponents has some benefit in combination with 3.3a. In combination, gains of 3.3e and 3.5 are somewhat additive in RA, whereas only 3.3e contributes gain LB (BDOF of 3.5 not applicable in LB). Several experts supported the combination, and no objection was raised.

Decision: Adopt JVET-AE0091 Test 3.8.

Test 3.4 does not show benefit in compression, might be due to new adoptions in ECM 9. No action was taken on this.

Test 3.7 is about RPR (non-CTC test conditions deduced from RPR specific VTM test conditions which only used 1.5x and 2x). The results indicate that improvement is possible by using different filters for small factors of resolution change. Though it is not of prior importance for the exploration, it may be interesting to provide such implementation in the software for experimentation of interested parties.

Decision (SW): Adopt JVET-AE0150 (not in CTC, not in the ECM description).

Transforms and coefficient coding

Test 4.1: Shifting quantizer center (JVET-AE0125)

In the test, a quantization offset is added to the quantized level, the offset is quantization index dependent, a look-up table is used to derive the offset .

where x is the dequantized coefficient, y and y’ are the quantization indices, Q-1 is the dequantization operation, T is a look-up table

.

Test 4.2: Large NSPT (JVET-AE0086)

The large NSPT kernels are tested for 4x32/32x4 and 8x32/32x8 blocks, of which kernel matrix dimensions are 20x128 and 24x256. Therefore, 20 and 24 transform coefficients are generated by applying the two types of kernel matrices, respectively, which are placed from DC position following scan order. The remaining 108 and 232 positions in each transform block are zeroed-out, respectively.

Large NSPT kernels are applied in the same way as for other block sizes 4x4, 4x8/8x4, 4x16/16x4, 8x8, and 8x16/16x8 and those NSPT kernels are not changed.

Test 4.3: Context modelling for transform coefficients for LFNST/NSPT (JVET-AE0102)

In the ECM, the causal 2D neighbourhood of coefficients is used to model the context to parse the sig_coeff_flag, gt1_flag, gt2_flag as shown in the next figure below. For LFNST, DCT-II coefficients are placed into the coefficient block using diagonal reordering.

A number on a black background

Description automatically generated

Figure: Context modelling for LFNST coefficients in the ECM

In the test, when LFNST/NSPT is applied, the previous 5 coefficients in the coding order are used for context derivation instead of 2D neighbourhood.

A number on a black background

Description automatically generated

Figure: Context modelling for LFNST coefficients in TEST4.3

In the ECM, the lfnstIdx is signalled after all the coefficients in a CU. In the test, lfnstIdx is signalled after all last_sig_coeff_pos syntax elements in a CU since lfsntIdx is required for parsing the transform coefficients.

Test 4.4: InterMTS for IBC and IntraTMP (JVET-AE0116)

InterMTS enabled for IBC and IntraTMP is tested. The following tests are performed:

Test 4.4a: InterMTS is enabled for IBC-coded blocks in AMVP mode. MTS is not enabled for IBC merge.

Test 4.4b: InterMTS is enabled for IBC-coded blocks in AMVP mode as in Test 4.4a. In addition, IntraTMP is utilizing InterMTS instead of IntraMTS. Consequently, the number of MTS candidates for IntraTMP has been fixed to 4 instead of adaptively setting it to 1, 4 or 6 as in IntraMTS.

Test 4.4c: IntraMTS is disabled for IntraTMP coded blocks.

Results for AI/RA (LFNST not enabled in LB)

Test 4.1: Shifting quantizer center provides 0.1% gain without any complexity impact.

Decision: Adopt JVET-AE0125 Test 4.1.

Test 4.2: Reasonable tradeoff runtime vs. bit rate reduction, additional memory consumption of larger kernels not relevant in cntext of exploration.

Decision: Adopt JVET-AE0086 Test 4.2.

Test 4.3: It was commented that the usage of immediately preceding coefficients might have some latency impact (VVC did not use coefficients from same diagonal for that reason), but for the purpose of exploration several experts (including cross-checkers) supported the proposal.

Decision: Adopt JVET-AE0102 Test 4.3.

Test 4.4x: Enabling MTS for IBC does not provide relevant gain to justify the increase of encoding time (caused by additional RDO checks). No action was taken on this.

Loop filtering

Test 5.1: CCSAO with extended edge classifiers and history offsets (JVET-AE0151)

In ECM-9.0, CCSAO uses band and edge classifiers switched at CTB level. Classifier parameters/offsets are derived at encoder and signalled for each slice independently.

In Test 5.1a, to reduce the signalling overhead, CCSAO inheritance scheme is introduced, where the offsets/parameters of some coded pictures are stored at both encoder and decoder which are allowed to be used as the CCSAO offsets/classifiers of future pictures. An index is signalled in SH to indicate which candidate in the CCSAO parameter storage is selected for the current slice. The candidates in the CCSAO parameter storage are updated in a FIFO manner and refreshed at IDR pictures.

In Test 5.1b, on top of Test 5.1a, beside the existing CCSAO edge classification, one new edge classification is added, which is a subset of the original one with less edge range divisions.

where is calculated by comparing a sample difference in one direction with a threshold .

Additionally, the component used for edge classification can be selected from one of all three components. Same to the existing CCSAO design, the selected edge classifier and edge component are decided by encoder and signalled in the SH.

Test 5.2: Improved fixed filters for ALF (JVET-AE0139)

In ECM-9.0, two Laplacian-based classifiers (one for each fixed filter) are applied to a 2x2 block. In each classifier, activity and directionality values are derived based on vertical, horizontal, and diagonal gradients using a window surrounding each 2x2 block. Then a class index is determined based on the activity and directionality values. Two 13x13 diamond shaped fixed filters are selected from the two filter sets by using the derived two class indices. Then the two fixed filters ( with ) are applied to the ALF input samples. Finally, a signalled filter is applied to the ALF input samples, samples before the deblocking filter (DBF), outputs of the two fixed filters, output of a gaussian filter and the residual data.

In Test 5.2a, both fixed filters are applied to samples before DBF and ALF input, where additional diamond 9x9 filter is used for the samples before DBF. The shape of the first fixed filter applied to the ALF input samples is reduced from 13x13 to 9x9, and the shape of the second fixed filter, which is 13x13, applied to ALF input is unchanged as shown in the table below.

ECM-9.0

Proposed method

ALF input

Samples before DBF

ALF input

Fixed filer

13x13

9x9

9x9

Fixed filer

13x13

9x9

13x13

In Test 5.2b, on top of Test 5.2a, the classifiers of the fixed filters are extended. For each 2x2 block, the mean value of a surrounding window is calculated. Then, for each sample of this window, the difference between the sample value and the mean value is calculated. A scaling factor is determined based on the activity value derived from a Laplacian classifier. The square root of the sum of the squared differences is further quantized to by a scaling factor. The value of is an integer between 0 and 7, inclusively. With i=0, 1, let denote the classifier from the classifier of i-th fixed filter in ECM-9.0. Then the proposed class index is derived as

.

The total number of the fixed filters is not changed.

In Test 5.2c, on top of Test 5.2b, fixed filter is applied to outputs of (instead of ALF input) and samples before DBF.

Test 5.3: Combination of Test 5.1b and Test 5.2c (JVET-AE0152)

One expert commented that the changes in 5.1a and 5.1b are straightforward and supported it.

It was report that an additional buffer of approcimately 5Kbyte is necessary for storage.

It was confirmed that the data flow in 5.2x is not changed relative to the current ECM.

It can be expected that the gains of CCSAO modifications and ALF modifications are non-overlapping, and it was also reported that preliminary results of 5.3 indicate additive gains.

Decision: Adopt JVET-AE0151 Test 5.1b.

Decision: Adopt JVET-AE0139 Test 5.2c.

It was commented that the results of 5.3 (combination test) are expected to be provided for information.

EE2 contributions: Enhanced compression beyond VVC capability (24)

There was no presentation or discussion about specific proposals in this category.

For actions decided to be taken, see section 5.2.1, unless otherwise noted.

Decisions
adopted
Adopt JVET-AE0139 Test 5.2c
Citation