Back to Search Document details
40th Meeting: Geneva, CH, October 2025 2025-10-04 22:46
EE2: Summary report of exploration experiment on enhanced compression beyond VVC capability
Abstract
This document provides a summary report of Exploration Experiment on Enhanced Compression beyond VVC capability. The tests are categorized as partitioning, intra prediction, inter prediction, transform and coefficient coding, and in-loop filtering.
JVET-AN0024 EE2: Summary report of exploration experiment on enhanced compression beyond VVC capability [V. Seregin, D. Buğdayci Sansli, J. Chen, R. Chernyak, K. Naser, J. Ström, F. Wang, M. Winken, X. Xiu, K. Zhang (EE coordinators)]

Tests

Tester

Cross-checker

1 Partitioning

1.1a

Restrictions for TT splitting (normative way) with split_cu_flag context change

G.Wang (vivo)

Z.Deng(Bytedance)

Y.Liu(Transsion)

Z.Sun(OPPO)

1.1b

Restrictions for TT splitting (non-normative way)

G.Wang (vivo)

Z.Deng(Bytedance)

Z.Sun(OPPO)

1.1c

Temporal partitioning prediction optimization

G.Wang (vivo)

Z.Sun(OPPO)

1.1d

Test 1.1a + Test 1.1c

G.Wang (vivo)

withdraw

1.1e

Test 1.1b + Test 1.1c

G.Wang (vivo)

withdraw

1.1f

Restrictions for TT splitting (normative way)

G.Wang (vivo)

Y.Liu(Transsion)

1.1g

split_cu_flag context change

G.Wang (vivo)

Y.Liu(Transsion)

2 Intra prediction

2.1

TMRL blend

S. Blasi

(Nokia)

V. Rufitskiy (TCL)

2.3a

Switchable interpolation filter for TIMD side reference

T. Dong

(TCL)

S. Lee, Y. Kim, S. Noh, J. Bang, H. Choi (HNU)

2.3b

Switchable interpolation filter for TIMD template.

T. Dong

(TCL)

M. Abdoli (Xiaomi)

2.3c

Test 2.3a + Test 2.3b

T. Dong

(TCL)

S. Lee, Y. Kim, S. Noh, J. Bang, H. Choi (HNU)

2.4

Longer tap interpolation filtering

V. Rufitskiy

(TCL)

M. Abdoli (Xiaomi)

2.5

Combination of Test 2.1, Test 2.3b, and Test 2.4

V. Rufitskiy

(TCL)

P. Andrivon (Ofinno)

2.6a

Adaptive subsampling filter selection for CCLM/intra-CCCM/ALF-CCCM

Y. Kidani

(KDDI)

D. Buğdayci Sansli

(Nokia)

P.Bordes (InterDigital)

2.6b

Adaptive subsampling filter selection for inter-CCCM

Y. Kidani

(KDDI)

D. Buğdayci Sansli

(Nokia)

2.6c

Test 2.6a + Test 2.6b

Y. Kidani

(KDDI)

D. Buğdayci Sansli

(Nokia)

2.7

Reducing candidate modes in decoder derived CCP

S. Wan(NWPU)

S. Xie(ZTE)

P. Bordes (InterDigital)

2.8

Enhanced CCP Fusion mode

P. Bordes

(InterDigital)

Y.Kidani (KDDI)

N.Zouidi

(Ofinno)

3 Inter prediction

3.1

Generated merge candidates

D. Buğdayci Sansli

(Nokia)

K. Naser

(InterDigital)

R. G. Youvalari (Xiaomi)

3.2a

Joint reordering of GPM with affine prediction

L. Zhang

(OPPO)

Y. Wang

(Bytedance)

3.2b

Modifications to the affine merge list

L. Zhang

(OPPO)

Y. Wang

(Bytedance)

3.2c

Test 3.2a + Test 3.2b

L. Zhang

(OPPO)

Y. Wang

(Bytedance)

3.3

Joint reordering of GPM with intra prediction

Z. Sun

(OPPO)

N.Zouidi

(Ofinno)

X.Li

(Alibaba)

3.4a

Test 3.2a + Test 3.3

Z. Sun

(OPPO)

X. Li

(Alibaba)

C. Ma

(Kwai)

3.4b

Test 3.2a + Test 3.2b + Test 3.3

Z. Sun

(OPPO)

X. Li

(Alibaba)

C. Ma

(Kwai)

4 Transform and coefficients coding

4.1

Linked sign prediction

Y. Zhang

(OPPO)

C. Hollmann (TCL)

4.2a

Shifting quantization center for TSRC under non-CTC (RDOQ on)

Y. Yu

(OPPO)

M. Le Pendu (InterDigital)

P. Nikitin (Xiaomi)

M. Balcilar (Ofinno)

B. Ray (Qualcomm)

4.2b

Shifting quantization center for RRC under non-CTC (DQ disabled, RDOQ on)

Y. Yu

(OPPO)

M. Le Pendu (InterDigital)

P. Nikitin (Xiaomi)

M. Balcilar (Ofinno)

B. Ray (Qualcomm)

4.2c

Test 4.2a + Test 4.2b

Y. Yu

(OPPO)

M. Le Pendu (InterDigital)

P. Nikitin (Xiaomi)

M. Balcilar (Ofinno)

B. Ray (Qualcomm)

4.2d

Shifting quantization center for TSRC under CTC

Y. Yu

(OPPO)

M. Le Pendu (InterDigital)

P. Nikitin (Xiaomi)

4.2e

Shifting quantization center for RRC under CTC

Y. Yu

(OPPO)

M. Le Pendu (InterDigital)

P. Nikitin (Xiaomi)

4.2f

Test 4.2d + Test 4.2e

Y. Yu

(OPPO)

M. Le Pendu (InterDigital)

P. Nikitin (Xiaomi)

4.2g

Shifting quantization center for RRC under non-CTC (DQ disabled, RDOQ disabled)

M. Le Pendu (InterDigital)

Y. Yu

(OPPO)

M. Balcilar (Ofinno)

4.2h

Shifting quantization center for TSRC under non-CTC (DQ disabled, RDOQ disabled)

M. Le Pendu (InterDigital)

Y. Yu

(OPPO)

M. Balcilar (Ofinno)

4.2i

Test 4.2g + Test 3.2h

M. Le Pendu (InterDigital)

Y. Yu

(OPPO)

M. Balcilar (Ofinno)

4.3a

Residual sign prediction restriction

C. Hollmann

(TCL)

4.3b

RSP context modelling

C. Hollmann

(TCL)

4.3c

Test 4.3a with retraining of existing contexts

C. Hollmann

(TCL)

4.3d

Test 4.3a + Test 4.3b

C. Hollmann (TCL)

5 In-loop filtering

5.1a

Regularization of ALF-CCCM

Z. Xie

(OPPO)

M. Jia

(ZTE)

5.1b

Regularization for ALF-CCCM enabled in I slice

Z. Xie

(OPPO)

M. Jia

(ZTE)

5.2a

Non-downsampled ALF-CCCM

L. Xu

(OPPO)

5.2b

Reuse of CU partition in ALF-CCCM

L. Xu

(OPPO)

5.2c

Test 5.2a + Test 5.2b

L. Xu

(OPPO)

5.3a

Remove single model of ALF-CCCM

N. Song

(OPPO)

5.3b

Separate “bad window” condition for Cb and Cr components

N. Song

(OPPO)

5.3c

Test 5.3a + Test 5.3b

N. Song

(OPPO)

5.4

Additional models for ALF-CCCM

F. Wang

(OPPO)

5.5a

Test 5.1 + Test 5.2

L. Xu

Z. Xie

(OPPO)

P.Astola (Nokia)

5.5b

Test 5.1 + Test 5.2 + Test 5.4

L. Xu

Z. Xie

F. Wang

(OPPO)

P.Astola (Nokia)

5.5c

Test 5.1 + Test 5.2 + Test 5.3 + Test 5.4

N. Song

L. Xu

Z. Xie

F. Wang

(OPPO)

P.Astola (Nokia)

5.6

Reuse of TALF control information

Y. Bai

(ZTE)

I.Jumakulyyev (Nokia)

Z.Xie (OPPO)

Partitioning

Test 1.1: On partitioning optimization (JVET-AN0121)

In multi-type tree partitioning, different splitting patterns may result in the same coding block structure. As shown in part (a) of the figure below, a binary tree split in vertical direction followed by a ternary tree split in horizontal direction may have the same coding block structure as a ternary tree split in horizontal direction followed by a binary tree split in vertical direction. A similar case is shown in part (b) of the figure.

A two lines with red x and a red x

AI-generated content may be incorrect.

In the tests, the duplicated partitioning splitting is disabled in a normative way and in encoder only. Additionally, context derivation for split_cu_flag was changed where instead of using width for the above neighbour block and height of the left neighbour block, the area of those blocks is utilized in comparing the current block area with the neighbour blocks area.

In the second part of the test, a prediction method for (No split) mode is introduced. For each CU, if the current depth is smaller than the temporal minimum QT depth minus 1, then (No split) is disallowed. Other split predictions are not changed. ECM has a similar prediction scheme for splitting modes (QT, BT, TT).

Test 1.1a: Restrictions for TT splitting (normative way) with split_cu_flag context change

Test 1.1b: Restrictions for TT splitting (non-normative way).

Test 1.1c: Temporal partitioning prediction optimization.

Test 1.1f: Restrictions for TT splitting (normative way).

Test 1.1g: split_cu_flag context change.

Results (also in runtime) for tests a and b confirmed by crosscheckers (at least partial results available), also confirm reduction in encoding time. Some crosscheckers expressed support for test 1.1a.

It was commented that 1.1g (which is a subset of 1.1a) has slightly better gain and also gives runtime reduction. It is a very small modification of the context definition, not adding context. It is also confirmed by crosschecker (crosscheck doc not registered yet) that also the encoding time is lower at least in RA and LB, and slightly higher gain in LB. After availability of full crosscheck results, the values in terms of gain and run time are confirmed.

Decision: Adopt JVET-AN0121 test 1.1g.

Intra prediction

Test 2.1: TMRL blend (JVET-AN0171)

In the test, a new mode is introduced for blending TMRL predictors. In TMRL, a template search is performed to determine pairs of intra mode and MRL index. In this mode, a template search is performed to determine up to three pairs of intra mode and MRL index. Correspondingly up to three predictors are then blended together to compute a prediction for the current block. Up to two predictors are obtained using directional intra prediction modes. A third non-directional predictor can also be used, obtained using PLANAR or BV prediction.

To differentiate TMRL blend from conventional TMRL, a different template size and a different list of MRL indexes are used in TMRL blend than in conventional TMRL. The template size is adaptively set to 2 or 4 depending on the block size, similarly to the template size used in TIMD. Template costs are computed in such a way that template samples closer to the current block are weighted more than template samples located further away from the current block.

A flag is coded to indicate the mode.

Test 2.3: Interpolation filter unification for unwrapping and TIMD template generation (JVET-AN0120)

In Test 2.3a, a set of candidate interpolation filters (4-tap or 6-tap) is defined for generating side reference (shown in the next figure) samples along the prediction direction. The filter switch is done based on the block size.

In Test 2.3b, switchable interpolation filters (4-tap or 8-tap) are applied for template generation of TIMD. The filter switch is done based on the block size.

Test 2.3a: Switchable interpolation filter for TIMD side reference.

Test 2.3b: Switchable interpolation filter for TIMD template.

Test 2.3c: Test 2.3a + Test 2.3b.

Test 2.4: Longer tap interpolation filtering (JVET-AN0158)

For each predictor a component used in DIMD, an interpolation filter is selected, and the selection of an interpolation filter is performed based on the DIMD blending parameters specified for a block.

The length of the interpolation filters is selected according to the next table.

Weight

[1,7]

[8,24]

[25,31]

[32,48]

>48

Filter length

16

16

16

8

8

Similarly, when TIMD uses blending, an interpolation filter of 8-, 12-, or 16-taps is selected based on the number of modes used in the blending classified into the following categories:

  • angular modes;
  • DC mode;
  • non-angular modes other than DC mode.

The filter length (number of taps) is determined from the number of modes in the categories as follows.

Number of angular modes

2

1

1

1

DC mode

any

1

0

any

Non-angular modes other than DC mode

any

0

0

more than 0

Filter length

12

16

8

12

When the main reference side has length of 8 or smaller then the number of taps of the selected interpolation filter is capped to not exceed 10-tap.

Test 2.4: Switchable interpolation filtering for DIMD and TIMD predictors.

Test 2.5: Combination of TMRL, TIMD, and DIMD interpolation filtering (JVET-AN0158)

In this combination test, TMRL blend is combined with TIMD interpolation filters such that Test 2.3b is used for filtering in the prediction derivation for TIMD template, while Test 2.4 filters are used in the TIMD and DIMD prediction derivation.

Test 2.5: Test 2.1 + Test 2.3b + Test 2.4.

Test 2.6:  Adaptive subsampling filter selection for CCLM/CCCM (JVET-AN0275)

Many 4:2:0 colour format videos are generated from 4:4:4 colour format videos using different subsampling filters. For example, {1,1; 1,1} and {1, 2, 1; 1, 2 ,1} filters are used, which are known as MPEG1 and MPEG2 subsampling filters for 4:2:0 format, respectively. ECM uses only MPEG2 filter for CCLM and CCCM.

To align the subsampling filters between the generation process of 4:2:0 colour format videos and CCLM and CCCM processes, this test introduces an adaptive subsampling filter selection method for CCLM and CCCM. MPEG1 and MPEG2 subsampling filters are adaptively selected for the luma subsampling processes of CCLM and CCCM (including intra-CCCM and ALF-CCCM) by reusing the existing SPS flags in ECM, which specify the video colour format and spatial allocation of luma and chroma samples (i.e., sps_chroma_format_idc, sps_chroma_horizontal_collocated_flag and sps_chroma_vertical_collocated_flag). The values of the existing SPS flags for every test sequence are automatically determined by the encoder-side method based on the analysis of a first picture characteristic.

Similarly, the adaptive filter selection method is also introduced into the filter derivation and application processes of inter-CCCM, where the prediction chroma sample value of inter-CCCM can be obtained using the following formula:

predChromaVal = c0 L0+ c1L2 + c2L3 + c3L5 + c4 nonlinear((L0+L2+L3+L5+2) >> 2) + c5 B,

in addition to the following existing inter-CCCM formula:

predChromaVal = c0 L0+ c1L1 + c2L2 + c3L3 + c4L4 + c5L5 + c6 nonlinear((L0+L3+1) >> 1) + c7 B,

where nonlinear is CCCM’s nonlinear operator and B is bias.

A black background with a black square

AI-generated content may be incorrect.

Test 2.6a: Adaptive subsampling filter selection for CCLM/intra-CCCM/ALF-CCCM.

Test 2.6b: Adaptive subsampling filter selection for inter-CCCM.

Test 2.6c: Test 2.6a + Test 2.6b.

Test 2.6d: Test 2.6a for Gaming_LD_HD class with CTC QPs.

Test 2.7: Reduced candidate modes in DDCCP (JVET-AN0220)

In the test, cross-component models with local boost are removed from the DDCCP candidate list, the candidate list is changed from {CCLM, CCCM, CCCM with LB-CCP, MM-CCCM, MM-CCCM with LB-CCP, GL-CCCM} to {CCLM, CCCM, MM CCCM, GL-CCCM}.

The next table summarizes the reduction for the number of modes.

Number of non-fusion modes

Number of fusion modes

Total Number

ECM DDCCP

6

15

21

Test

4

6

10

Reduction

2

9

11

Test 2.7: Reduced candidate modes in decoder derived CCP.

Test 2.8: Enhanced CCP fusion mode (JVET-AN0168)

In the test, for the decoder-derived CCP merge or CCP merge fusion modes, two CCP predictions using the regressing-based GPM (RGPM) method, which derives the blending weights from the template, are combined. The blending weights are derived jointly for the two chroma components using the same method as in RGPM and SGPM modes. The maximum number of fusion candidates is signalled in SPS (8 in AI and for inter slices, otherwise 12).

Test 2.8a: Enhanced CCP fusion mode.

Test 2.8b: Test 2.7 + Test 2.8a (JVET-AN0173).

For 2.1, an expert comments that the gain is higher in class E than in other classes (lowest in class A). Also, some encoder optimization is performed that could also be applied in anchor. No additional TM is used, but additional processing similar to DIMD is necessary. Among 2.1…2.4, 2.1 has worst tradeoff.

It was commented by crosscheckers that gains of 2.3b and 2.4 are believed be additive, however no results are available on such a combination. Both are straightforward changes using longer interpolation filters. Standalone, 2.4 has a better tradeoff.

Decision: Adopt JVET-AN0158 test 2.4.

On Sunday 12 Oct. Proponents of 2.3b reported additional results on combination with 2.4 (cross-checked in JVET-AN0375) Coding gain 0.01% in AI over 2.4 with no runtime increase, however standalone the gain was 0.02%. No results in RA available yet, but for RA 2.3b standalone had no gain, or very small loss in chroma. No action taken.

Test 2.6 shows significant gain only for two sequences (one from class A, and one from class F), which is natural as the other sequences obviously used the chroma subsampling currently implemented.

It is straightforward to use the chroma subsampling (horizontal and/or vertical shift) information from SPS in CCCM/CCLM/ALF-CCCM (if available), as investigated in test 2.6a. An automatic detection algorithm is also provided.

In 2.6b, a different model is used for inter CCCM. However, inter CCCM does not use downsampling, and the new model does not show benefit

Results are confirmed by cross-check.

Decision: Adopt JVET-AN0090 test 2.6a, including the detection method (the latter disabled by default). In CTC, the per-sequence config can be used to signal (update necessary for cat robot and slide show)

Also some of the sequences in CfE would have additional gain for ECM (see info in JVET-AN0090). Proponents provided updated configuration files in version 4 of JVET-AN0090. It is further noted that for CCLM in VTM no change would be useful, as it only supports vertical shift.

2.7 removes the two local boost modes showing reduction of encoding and decoding time with marginal loss in compression.

2.8a reduces the number of candidates from 12 to 8, which also gives similar reduction in encoding and decoding runtime.

2.8a has better tradeoff than 2.7, and also the combination 2.8b has worse tradeoff than 2.8a.

Decision: Adopt JVET-AN0168 test 2.8a.

Inter prediction

Test 3.1: Generated merge candidates (JVET-AM0059)

In this test, new merge candidates are generated from the existing ones in the list. After the initial merge list is constructed, new candidates are added to the list before the pair-wise average candidates are generated and ARMC is applied.

Candidate generation first creates two separate lists of uni-predictors from existing candidates. If a uni-prediction candidate is found, it is taken as it is and if a bi-prediction candidate is found, motion information from each prediction direction is taken to the related list.

The template cost of each candidate in the list is calculated, the lists are sorted in ascending cost order. From the lowest cost predictors, new uni and bi-predicted merge candidates are generated and added to the list before ARMC stage. The actual number of newly generated candidates depends on the found unique uni-predictors, where upper limit is set to the maximum number of merge candidates.

Test 3.1: Generated merge candidates.

Test 3.2:  Joint reordering of GPM with affine prediction (JVET-AN0091)

In ECM, when the jointly reordering GPM method is applied, a candidate list is constructed with each candidate containing one split mode and the MVs of the two GPM partitions. This candidate list is reordered using a template-based scheme. Only the selected candidate index, instead of the split mode and the motion information of the two partitions, needs to be signalled to the decoder to indicate the split mode and motion vector pair.

In Test 3.2a, a jointly reordering GPM with affine mode is introduced, where a candidate list is constructed with each candidate containing one split mode, two partition indices and two affine mode indicators (isAffine0, isAffine1), the usage of the mode is indicated by flag.

For each GPM partition, the partition index is used to get motion information from regular merge or affine merge lists, and the partition is predicted by affine or non-affine motion compensation. The candidate list is reordered using a template-based scheme, the selected candidate index is signalled to the decoder. Besides, the samples of the current template are mapped to the original domain during the construction of the jointly reordering GPM candidate list when the LMCS is enabled.

To reduce the complexity, the tested mode is only used when the block size is not equal to 128 or when there is at least one adjacent/non-adjacent coded block with affine prediction. When constructing the candidate list, the number of available affine merge candidates is reduced to 15, and the candidate list size is reduced to 10 when the current slice satisfied the LDB condition; otherwise, the candidate list size is set to be 16. No additional RDO is added on the encoder side.

In Test 3.2b, the construction of affine candidates is modified. For each control point, while checking the corresponding neighbouring positions, all the available motion information with a different reference picture is stored and then used to construct the possible affine candidates (temporally derived affine candidates for GPM). This process is carried out by traversing the possible CPMVs combinations. The neighbouring positions and CPMVs combination are the same as in ECM.

Test 3.2a: Joint reordering of GPM with affine prediction.

Test 3.2b: Affine merge list modifications.

Test 3.2c: Test 3.2a + Test 3.2b.

Test 3.3: Joint reordering of GPM with intra prediction (JVET-AN0092)

In the test, intra prediction modes are added as additional candidates into the joint reordering of split modes and partition indices, where the regular GPM modes together with intra prediction modes are reordered jointly, and there up to 6 intra-prediction candidate modes are added to the final list. A flag is introduced to indicate whether the scheme is applied.

The construction of intra prediction candidate modes is the same as in intra MPM list. To reduce the overall complexity for both encoding and decoding, the size of candidate list is set to 16 for slice not satisfying low delay condition; otherwise, the candidate list size is set to 10. The total RDO processes in ECM remain unchanged.

Test 3.3: Joint reordering of GPM with intra prediction.

Test 3.4: Combination for GPM reordering tests (JVET-AN0093)

Test 3.4a: Test 3.2a + Test 3.3.

Test 3.4b: Test 3.2a + Test 3.2b + Test 3.3.

Test 3.1: Results confirmed by cross-checkers, and technology supported by them. Tradeoff is within an acceptable margin.

Decision: Adopt JVET-AN0236 test 3.1.

Results confirmed by cross-checkers, and support was expressed for the combination 3.4a which has the best overall tradeoff. It is noted that 3.2b which standalone provides more interesting gain for LB, loses this advantage when combined with 3.2a and 3.3 (in 3.4b).

Decision: Adopt JVET-AN0093 test 3.4a.

Transform and coefficient coding

Test 4.1: Linked sign prediction (JVET-AN0094)

In the test, a linked sign prediction method is introduced, where the residual contribution from non-sign predicted coefficients is added to the residual contribution of sign predicted coefficient when assessing boundary continuity, the method has three steps:

  1. Non sign predicted coefficients linking process
  2. Residual hypothesis of the linked coefficient group
  3. Sign correction of non sign predicted coefficients.

In the non sign predicted coefficient linking process, those coefficients are classified into two groups: inside (type A) and outside (type B) of the sign predicted coefficient area, shown in the next figure.

A black background with blue squares and crosses

AI-generated content may be incorrect.

If neither type A nor type B coefficients exist, then the method is not used. If only type A or B coefficients exists, these non sign predicted coefficients are linked to the sign predicted coefficient with the smallest coefficient magnitude. Together, these coefficients constitute a linked coefficient group. The linked coefficients are used together to calculate two hypotheses for their combined contribution to the residual.

If both type A and B coefficients exist, type A coefficients are linked to the sign predicted coefficient with the second-smallest coefficient magnitude. Type B coefficients are linked to the sign predicted coefficient with the smallest coefficient magnitude. Each linked group will be used to calculate two hypotheses respectively for their combined contribution to the residual.

An example of linked coefficients grouping is shown in the next figure, where LSP denotes a selected sign predicted coefficient in the linked group.

A black background with blue lines and symbols

AI-generated content may be incorrect.

Similar to the current sign prediction method, each linked group is used to calculate two hypotheses (for positive and negative sign) of its contribution to the residual as follows:

or

where represents the combined residual contribution at the x-th and y-th position of the current block for one hypothesis of the linked group; represents the residual contribution at the x-th and y-th position of the current block from the sign predicted coefficient assuming it has a positive sign; represents the residual contribution at the x-th and y-th position of the current block from the i-th non sign predicted coefficient with signs applied as signalled by their corresponding bypass coded sign bits; and represents the number of non-SP coefficients in the LCG.

The signs for both type A and B coefficients remain bypass coded. However, when they are included into a linked group, the meaning of these signalled sign bits may be modified.

If the sign predicted coefficient associated with the group is correct, then the signs of its linked non sign predicted coefficients are the same as indicated by coded sign bits. Otherwise, the signs of its linked non sign predicted coefficients are interpreted as the reverse of what was signalled by their sign bits

Test 4.1: Linked sign prediction.

Test 4.3: Residual sign prediction restriction (JVET-AN0129)

In Test 4.3a, sign prediction is not applied to blocks that have only one non-zero coefficient in the sign prediction area.

Int Test 4.3b, the number of contexts used for the sign prediction flags is increased from 8 to 16 by adding a dependency on the transform type.

In Test 4.3c, a restriction from Test 4.3a has been applied and the initial values for the original context models have been retrained using the CabacTraining application. Initial values for other contexts have not been retrained.

Test 4.3a: Sign prediction restriction.

Test 4.3b: Increased contexts for sign prediction.

Test 4.3c: Test 4.3a with sign prediction context retraining.

Test 4.3d: Test 4.3a + Test 4.3b.

Test 4.2: Shifting quantization center (JVET-AN0095, JVET-AN0096, JVET-AN0097)

A fix for RDOQ implementation with LFNST and NSPT was submitted to ECM-18.0 in MR 960. In ECM, when RDOQ is used for a block with LFNST or NSPT, only up to 16 or 8 coefficients (depending on the block size) are encoded in RDOQ, instead of using the number of output coefficients of LFNST or NSPT for the current block.

In non-CTC configuration, when DQ and SignPred are disabled, the test results for the fix are as follows.

AI: Y -0.52% / U -0.21% / V -0.20%

RA: Y -0.35% / U -0.22% / V -0.14%

LDB: Y -0.25% / U 0.03% / V -0.09%

In the 4.2 tests, (*) denotes the results when the fix is applied.

In ECM for RRC, the shifting amount in the dequantized coefficients is inversely proportional to the quantization indices of DQ as follows.

where is the dequantized coefficient, is the size of the lookup table, is the dequantized value of the -th of DQ,is the auxiliary quantization index calculated as

A lookup table of T is defined as follows.

In the test, the shift to dequantized values is applied when DQ is disabled (non-CTC configuration). The shift is applied as follows.

where is the dequantized coefficient, is the size of the lookup table, is the dequantized value of the quantization level , and is the auxiliary quantization level that is calculated as

Test 4.2a: Shifting of quantization center applied to TSRC (JVET-AN0095).

Test 4.2b: Shifting of quantization center applied to RRC (JVET-AN0095).

Test 4.2c: Test 4.2a + Test 4.2b (JVET-AN0095).

Test 4.2g: Shifting of quantization center applied to RRC without RDOQ (JVET-AN0097).

Test 4.2h: Shifting of quantization center applied to TSRC without RDOQ JVET-AN0097).

Test 4.2i: Test 4.2h + Test 4.2h (JVET-AN0097).

Test 4.2d: Shifting of quantization center applied to TSRC in CTC (JVET-AN0096)

A modification of the quantization center shift is tested with enabled DQ (used only in RRC in the CTC settings). Since the quantization level from Q0 or Q1 (instead of the DQ quantization index) is coded into a bitstream and the value of quantization level can be used to better reflect the rate, this test uses the quantization level from Q0 and Q1 instead of quantization indices of DQ to derive the shifting amount.

The same quantization level from quantizer Q0 and Q1 will generate a similar rate while the dequantization value of Q1 is smaller than that of Q0 for the same quantization level . Therefore, Q1 and Q0 are further modified to use different table sizes, a lookup table of size 32 and a lookup table of size 48 are used for Q0 and Q1, respectively.

where indicates the Q0 or Q1, is the dequantized coefficient, is the size of the lookup table, is the dequantized value of the quantization level from quantizer Qi, is the auxiliary quantization index calculated as

Furthermore, those tables can be interleaved to form one table as follows.

The dequantization value for quantization level k from Q0 or Q1 is calculated as follows.

where is the size of the lookup table, and is the auxiliary quantization index that is calculated as

Test 4.2e: Modified shifting of quantization center applied to RRC (JVET-AN0096)

Test 4.2f: Test 4.2d + Test 4.2e (JVET-AN0096)

In summary, there are the following aspects:

  • RDOQ fix MR 960 (encoder only)
  • Quantization center shift for TSRC with DQ enabled and disabled.
  • Quantization center shift for RRC with DQ disabled.
  • Modified quantization center shift for RRC with DQ enabled.

Test 4.1 does not have interesting tradeoff, encoding runtime increase too large.

Test 4.3a could be attractive in reducing the number of sign predictions, but this has some more significant loss in LB without large decrease in encoding/decoding time. It was reported by proponents that this is mainly caused by class E. Retraining reduces the chroma losses, but mainly by better gain in other classes.

Test 4.3b has good tradeoff for LB, but increasing the number of contexts does not improve in RA, and also in AI, gain is rather small. When combining tests a and b, RA even shows some loss

Overall, no relevant benefit in tests 4.1 and 4.3.

Test 4.2 investigates shifting of quantization center. This shows benefit using the shift for RRC (4.2b*) but not for TSRC (4.2a*) when DQ and sign prediction are disabled (non-CTC), and some further gain if RDOQ is also disabled (4.2g). Under CTC, application to TSRC gives some gain mainly for screen content (4.2d), and results in 4.2a* and 4.2c* also show that it has almost no impact in the non-CTC case when used for TSRC as well. Another variant is tested using a modified shifting for RRC which shows also some gain for non-screen content (4.2e). It is however not known how this modified shifting would work in the non-CTC case, and how this modified shifting would work for TSRC. Therefore, it is not to be considered at this moment, but further results may be provided in the future.

It was confirmed that sign bit hiding is enabled in all cases.

The normative change related to RRC would be to remove the constraint not using the quantization center when DQ is disabled (non-CTC case). This is asserted to be beneficial according to test 4.2b*.

The normative change related to TSRC would be to enable the shift of quantization center (same as for RRC). This would also affect the CTC case (would be beneficial for screen content, test 4.2d).

In terms of normative changes, test 4.2c* includes both aspects.

Decision: Adopt JVET-AN0095 test 4.2c*. Also encoder modifications are necessary for CTC and non-CTC cases.

In-loop filtering

Test 5.1: On regularization of ALF-CCCM (JVET-AN0098)

A regularization method of EIP is applied to derive of ALF-CCCM coefficients, a regularization parameter is determined by the number of input samples. For the cases where the number of input samples is smaller than {16, 32, 64, 256, or 1024}, the regularization parameter is set to {192, 160, 128, 96, or 64}, respectively. Otherwise, the regularization parameter is always set to 32.

If there is a sample whose correction exceeds the threshold, the corresponding samples using the same ALF-CCCM model will not be filtered. In the test, the threshold is modified from 32 to 16 and 28 for the single-model filter and the multi-model filter, respectively.

The tested method is not applied to large blocks, which regression window is larger than 64x64, in ALF-CCCM.

Test 5.1a: Regularization of ALF-CCCM.

Test 5.1b: Regularization of ALF-CCCM enabled in I-slices.

Test 5.2: On ALF-CCCM (JVET-AN0099)

In ALF-CCCM method, the output samples of luma SAO are used as inputs to the CCCM filtering. To obtain a correction signal, the SAO chroma samples are subtracted from the CCCM output samples. The correction is weighted by 0.5 and added to the ALF chroma output to improve chroma reconstruction samples.

For each CTU, the encoder’s RDO decides the best cross-component model from eight possible models, they differ in the number of samples, non-linear term, and biases. When applying multi-models, if any sample’s correction exceeds the threshold with either model, the corresponding samples using the same ALF-CCCM model will not be filtered.

Two aspects are introduced in this test:

  1. Two new model types are introduced, which are non-downsampled models used in intra CCCM prediction and inter CCCM prediction. A slice level syntax element is signalled to indicate the application of non-downsampled models. A CTU level flag is further signalled to indicate whether the current CTU uses the downsampled ALF-CCCM. The filtering process of the two new model types is summarized in the following equations.

New model 0 (same to intra CCCM):

New model 1 (same to inter CCCM):

A white triangle with a black background

AI-generated content may be incorrect.

  1. The CU partition is introduced into ALF-CCCM. A slice level syntax element is signalled to indicate the application of CU partition. A CTU level flag is further signalled to indicate whether the current CTU is processed according to the CU partition instead of the existing 8 ALF-CCCM block sizes {4x4, 8x2, 2x8, 8x8, 16x16, 32x32, 64x64, 128x128}.

Test 5.2a: Non-downsampled ALF-CCCM.

Test 5.2b: CU partitioning in ALF-CCCM.

Test 5.2c: Test 5.2a + Test 5.2b.

Test 5.3: Updated multi-models usage strategy for ALF-CCCM (JVET-AN0101)

In ALF-CCCM, a single CCCM model is calculated for each block, and then samples within the block are filtered by using this model. If any sample in either the Cb or Cr component of the block has a correction exceeding a threshold, the block will use multi-models instead. When filtered by the multi-models, if there is a sample in either the Cb or Cr component whose correction exceeds the threshold, the corresponding samples that use the same model as this sample will not be filtered.

In Test 5.3a, single model is removed, and only multi-models are applied in a block.

In Test 5.3b, the correction threshold condition is treated separately for the Cb and Cr components. ECM disables the model for both components together.

Test 5.3a: Remove single model of ALF-CCCM.

Test 5.3b: Separate correction condition for Cb and Cr components.

Test 5.3c: Test 5.3a + Test 5.3b.

Test 5.4: On ALF-CCCM model (JVET-AN0103)

In the test, nonlinear terms with the chroma samples are added at additional positions for model 0, model 3, model 4 and model 5, as shown in yellow in the next table.

A white sheet with yellow text

AI-generated content may be incorrect.

Test 5.4: Additional models for ALF-CCCM.

Test 5.5: Combination tests for ALF-CCCM (JVET-AN0106)

Test 5.5a: Test 5.1a + Test 5.2c.

Test 5.5b: Test 5.1a + Test 5.2c + Test 5.4.

Test 5.5c: Test 5.1a + Test 5.2c + Test 5.3a +Test 5.4.

Test 5.5c*: Test 5.5c with encoder optimization.

Test 5.5d: Test 5.1a + Test 5.2c + Test 5.3a.

Test 5.6: Reuse of TALF control information (JVET-AN0199)

In this test, slice and CTB-level TALF control information of the current picture is stored for a potential reuse in subsequent pictures. A flag is introduced to indicate whether the TALF control information of the current picture reuses that of the previously encoded picture.

At the encoder side, a similarity is evaluated by comparing the ALF CTB on/off status of the two pictures. The similarity score is defined as the percentage of CTBs with the same ALF on/off status relative to the total number of CTBs. If the similarity score falls below a threshold of 0.2, the TALF control information from the previous picture is not reused.

Test 5.5: Reuse of TALF control information.

Answers to questions:

  • The regularization of 5.1a is used for both cases of single model and multi model.
  • Decoding time is increased in 5.3a, as the option of single model is removed, and always two models need to be computed.

5.1…5.5 are related to different modifications in ALT-CCCM (and combinations thereof). 5.1.a and 5.3a each have a reasonable standalone tradeoff. However, a combination of 5.1a and 5.3a was not tested. 5.5d is a combination additionally including 5.2c (combination of 5.2a and 5.2b). Gains of all three together can at least for RA be considered as additive (as for RA 5.2c has almost no gain overall with loss in luma and small gain in chroma). Therefore, it might be concluded that the gain only comes from 5.1a and 5.3a, such that their gains would be additive in combination. Further, the runtime is also approximately the addition of the individual encoder runtimes of all three proposals (2% in total), whereas the addition of individual runtimes only from 5.1a and 5.3a would only be slightly more than 1%.

It was further remarked that the combination 5.2c has worse performance than 5.2a in RA, and for LB perhaps small gain, but a different balance balance between luma and chroma (which likely comes from 5.2b). Also the standalone tradeoff of either 5.2a or 5.2b is not reasonable in RA.

Decision: Adopt JVET-AN0098 Test 5.1a.

Decision: Adopt JVET-AN0101 Test 5.3a.

Test 5.6 does not provide gain in RA, and small loss in chroma. The runtime reduction is unreliable – according to proponents, there should be no change in runtime.

EE2 contributions: Enhanced compression beyond VVC capability (24)

There was no presentation or discussion about specific proposals in this category – contributions were discussed in the context of the EE summary report JVET-AN0024. For actions decided to be taken, see section 5.2.1, unless otherwise noted.

Decisions
adopted
Adopt JVET-AN0101 Test 5.3a
Citation