JVET-AE0024 EE2: Summary report of exploration experiment on enhanced compression beyond VVC capability [V. Seregin, J. Chen, R. Chernyak, K. Naser, J. Ström, F. Wang, M. Winken, X. Xiu, K. Zhang (EE coordinators)]
This document provides a summary report of Exploration Experiment on Enhanced Compression beyond VVC capability. The tests are categorized as partitioning, intra prediction, inter prediction, transform and coefficient coding, and in-loop filtering.
The software basis for this EE is ECM-9.0, released at https://vcgit.hhi.fraunhofer.de/ecm/ECM/-/tags/ECM-9.0. ECM-9.0 is used as an anchor in the tests.
Software for EE tests is released in the corresponding branches at https://vcgit.hhi.fraunhofer.de/ecm/jvet-ad-ee2/ECM/-/branches.
Test results can be found in input JVET contributions, cross-check results are uploaded to https://vcgit.hhi.fraunhofer.de/ecm/jvet-ad-ee2/simulation-results if cross-check reports are not submitted as they are optional for EE tests.
List of tests
Tests | Tester | Cross-checker | ||
1 Partitioning | ||||
1.1a | Partitioning prediction | Canon G. Laroche | InterDigital S. Puri | |
1.1b | Partitioning prediction with unconditional MTT depth increment | Canon G. Laroche | InterDigital S. Puri | |
1.1c | Encoder only partitioning prediction | Canon G. Laroche | InterDigital F. Le Léannec | |
2 Intra prediction | ||||
2.1a | Block vector guided CCCM with IBC BV | Nokia R. Youvalari | B<>com M. Abdoli | |
2.1b | Block vector guided CCCM with IBC and IntraTMP BV | Nokia R. Youvalari | B<>com M. Abdoli | |
2.2a | Extended IBC-GPM with two IBC predictions | Kwai | withdrawn | |
2.2b | Bi-predictive IBC-GPM | KDDI | withdrawn | |
2.2c | Bi-predictive IBC-GPM | Kwai KDDI | Bytedance W. Yin | |
2.3a | IBC BVP-merge | KDDI | ||
2.3b | Test 2.3a + bi-predictive IBC merge without IBC-GPM | KDDI | Ericsson R. Yu Bytedance W. Yin | |
2.3c | Bi-predictive IBC merge without + Test 2.3a + Test 2.2c | KDDI Kwai | withdrawn | |
2.4a | IBC MBVD list derivation | Qualcomm Z. Zhang | Ofinno D. Ruiz Coll | |
2.4b | Test 2.4a + Test 2.3a | Qualcomm Z. Zhang KDDI | withdrawn | |
2.4c | Test 2.4a + Test 2.3b + Test 2.2c | Qualcomm Z. Zhang KDDI Kwai | Ericsson R. Yu | |
2.5a | Filtered IBC | Kwai | withdrawn | |
2.5b | Filtering IBC predicted blocks | Qualcomm | withdrawn | |
2.5c | Test 2.5a + Test 2.5b | Kwai Qualcomm | OPPO Y. Yu L. Zhang | |
2.6a | IBC with non-adjacent spatial candidates | Kwai C. Ma | withdrawn | |
2.6b | Non-adjacent spatial candidates for IBC | Bytedance Y. Wang | withdrawn | |
2.6c | IBC with non-adjacent spatial candidates (Test 2.6a + Test 2.6b) | Kwai C. Ma
Bytedance Y. Wang | Alibaba J. Chen | |
2.7 | Cross-component merge mode with temporal candidates | MediaTek | InterDigital | |
2.8 | An extrapolation filter-based intra prediction mode | OPPO L. Xu | Kwai H.-J. Jhu | |
2.9 | Extended search areas for IntraTMP mode | OPPO Y. Yu Xidian Y. Ma, H.Zhang Kwai X. Xiu | InterDigital K. Naser Qualcomm P.-H Lin Alibaba X. Li | |
2.9b | Extended search areas for IntraTMP mode with scan order #2 | OPPO Y. Yu Xidian Y. Ma, H.Zhang Kwai X. Xiu | withdrawn | |
2.9c | IntraTMP mode with partial extended search areas with scan order #1 | OPPO Y. Yu Xidian Y. Ma H. Zhang Kwai X. Xiu | withdrawn | |
2.9d | IntraTMP mode with partial extended search areas with scan order #2 | OPPO Y. Yu Xidian Y. Ma H. Zhang Kwai X. Xiu | withdrawn | |
2.10a | IBC-LIC extension without large block-size constraint | OPPO Z. Xie | Bytedance W. Yin | |
2.10b | ECM IBC-LIC without large block-size constraint | OPPO Z. Xie | Bytedance W. Yin | |
2.11a | Harmonization between IBC HMVP and IBC-LIC | Bytedance N. Zhang | Kwai | |
2.11b | Test 2.11a + Test 2.10 | Bytedance N. Zhang OPPO Z. Xie | Kwai | |
3 Inter prediction | ||||
3.1a | Cross-component residual model | Nokia P. Astola | InterDigital F. Le Léannec Ittiam | |
3.1b | Cross-component residual model with complexity reductions | Nokia P. Astola | InterDigital F. Le Léannec Ittiam | |
3.2 | Bi-predictive GPM | Ericsson R. Yu | LGE Y. Ahn KDDI | |
3.3a | Additional TM refinement for bi-prediction | Bytedance | Alibaba J. Chen | |
3.3b | 16-point diamond search pattern for TM | Bytedance | Alibaba J. Chen | |
3.3c | Enabling TM for bi-prediction under DMVR condition | Bytedance | Alibaba J. Chen | |
3.3d | Test 3.3a + Test 3.3c | Bytedance | Alibaba J. Chen | |
3.3e | Test 3.3a + Test 3.3b + Test 3.3c | Bytedance | Alibaba J. Chen | |
3.4 | HPel flag and BCW weight usage in OBMC | InterDigital | withdrawn | |
3.5 | Iterative BDOF pass in multi-pass DMVR | Bytedance M. Salehifar Alibaba J. Chen | vivo Z. Lv | |
3.6 | Affine AMVP mode with one MVD | Qualcomm | withdrawn | |
3.7a | RPR with new filters, scale factor 1.25x | Qualcomm Z. Zhang | ||
3.7b | RPR with new filters, scale factor 1.33x | Qualcomm Z. Zhang | ||
3.8 | Combination of Test 3.3e and Test 3.5 | Bytedance M. Salehifar Alibaba J. Chen | Kwai C. Ma | |
4 Transform and coefficient coding | ||||
4.1 | Shifting quantizer center | InterDigital M. Balcilar | Qualcomm M. Coban | |
4.2 | Large NSPT | LGE M. Koo | Qualcomm P. Garus | |
4.3 | Context modelling for transform coefficients for LFNST/NSPT | Qualcomm P. Nikitin | IRT b-com | |
4.4a | InterMTS is enabled for IBC-coded blocks in AMVP mode. | Qualcomm P. Garus | InterDigital | |
4.4b | Test 4.4a + IntraTMP using interMTS instead of intraMTS kernels | Qualcomm P. Garus | InterDigital K. Naser | |
4.4c | IntraMTS disabled for IntraTMP | Qualcomm P. Garus | InterDigital K. Naser | |
5 In-loop filtering | ||||
5.1a | CCSAO with temporal history offset | Kwai C.-W. Kuo | Qualcomm N. Hu | |
5.1b | Test 5.1a + extended edge classifier | Kwai C.-W. Kuo | Qualcomm N. Hu | |
5.2a | Applying fixed filters to samples before DBF | Qualcomm N. Hu | Kwai C.-W. Kuo | |
5.2b | Test 5.2a + extended classifiers for fixed filters | Qualcomm N. Hu | Kwai C.-W. Kuo | |
5.2c | Test 5.2b + applying the second fixed filter to outputs of the first fixed filter | Qualcomm N. Hu | Kwai C.-W. Kuo | |
5.3 | Combination of Test 5.1b and Test 5.2c | Kwai C.-W. Kuo Qualcomm N. Hu | Alibaba J. Chen | |
Category 1: Partitioning prediction
Test 1.1: Partitioning prediction (JVET-AE0132)
In this test, a temporal partitioning is introduced, where for each block, the allowed partitioning splits are predicted according to the minimum QT/MTT split and the average QT/MTT split obtained from a temporal area:
- if the current QT depth is less than the temporal minimum QT depth minus 1, only the QT split is allowed,
- if the current QT depth is less than the temporal average QT depth minus 1, no split, QT and TT splits are allowed, and BT split is allowed if TT is selected in a parent node,
- if the temporal maximum MTT depth and less than the maximum MTT depth of the current block, the depth is decreased except if the current QP is less than the temporal QP,
- if the temporal maximum MTT depth is large than the maximum MTT depth of the current block and if the current depth is equal to the current QT depth, the maximum MTT depth is incremented,
- if the maximum QT depth of the current frame is larger than the maximum multi-tree depth and if the temporal maximum MTT depth is equal to the maximum MTT depth of the current frame, and if the temporal QT depth is less than the QT depth of the current frame, the maximum MTT depth of the current block is decreased.
The modifications on the maximum MTT depth are enabled only if the palette mode is disabled.
The temporal area is a collocated area from the same reference frame used for the temporal motion predictor.
JVET-AD0147 has memory requirement analysis where the largest memory which is required to store temporal depths for class A is estimated as 40.5 KB.
Test 1.1a: Partitioning prediction (as described above)
Test 1.1b: Test 1.1a where MTT depth increment is enabled unconditionally. It affects classes A and B, class A1 encoder runtime reaches 105.8%.
Test 1.1c: Use partition prediction as encoder only optimization.
Questions:
- Is the partitioning only from the reference frame? A: Yes.
- What is the additional memory requirement? A: Approx. 40 Kbyte per reference picture in 4K.
- Does CABAC parsing require the temporal partitioning info? A; Yes.
It was commented that the memory requirement is significantly less than for motion vectors.
The cross-checkers pointed that the implicit splitting performed at the picture boundary is useless (no gain when removing it, but more complicated).
It was commented that for LB there is no gain (or small losses).
Even though there was no strong objection, from the questions raised the implementation of the idea, it appears not mature. The gain is relatively low, even though it could be attractive in terms of run time, it introduces additional temporal dependencies that might be undesirable.
No action was taken on this.
Category 2: Intra prediction
Test 2.1: Block vector guided CCCM (JVET-AE0100)
In this test, a co-located luma BV is used to determine the reference area for calculating the CCCM parameters. Then the reference area in luma and corresponding area in chroma channel is used to calculate the CCCM parameters. The prediction uses the calculated model parameters and co-located luma samples to do the CCCM prediction as follows:
predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P(C) + c6P(N) + c7P(S) + c8P(W) + c9P(E)+ c10B,
where spatial components are as shown in the next figure, the nonlinear term P is represented as power of two of the corresponding luma sample, and B is the bias term.
The mode is enabled only in intra slices and is controlled by SPS flag.
Test 2.1a: the mode only uses BVs of the IBC coded blocks from co-located luma area.
Test 2.1b: the mode can use BVs of both IBC and IntraTMP coded blocks from co-located luma area.
Test 2.2: Bi-predictive GPM (JVET-AE0169)
In ECM-9.0, IBC GPM uses uni prediction for one partition and intra mode for the other one. In the test, the core GPM design is kept unchanged which uses 48 GPM modes and IBC merge candidate list, while it enables both partitions being predicted using IBC with different BVs.
Two flags are signalled to indicate the prediction modes of two partitions, the first flag indicates whether the first partition is intra predicted, and if not then the second flag is signalled to indicate whether intra prediction is used for the second partition.
The method is applied to SCC only.
Test 2.3: IBC BVP-merge and bi-predictive IBC merge (JVET-AE0169)
In Test 2.3a, IBC-BVP-merge, which is similar to AMVP-merge, derives one BV from IBC block vector prediction (BVP) and the second BV from IBC merge to form bi-prediction for IBC. Two different indices for the IBC BVP and the IBC merge candidates are signalled.
In Test 2.3b, bi-predictive IBC merge is introduced, and it is enabled together with the existed in the ECM MBVD and uni-merge (currently disabled by the encoder configuration for non-SCC classes). In bi-predictive IBC merge, two BVs from the existing IBC merge candidate list are derived, utilizing two different indices, which are signalled. Bi-predictive IBC merge is applied to IBC regular merge and IBC MBVD. Bi-predictive IBC merge, IBC MBVD, and IBC uni-merge are enabled for non-SCC classes and is tested together with Test 2.3a.
Test 2.4: IBC MBVD list derivation (JVET-AE0169)
In the test 2.4a, adaptive BVD offsets along MVBD directions and enabled for IBC MBVD mode. The MBVD candidates search is a two-step process, which starts with checking template SAD costs of offsets added to BVP along each direction with the interval of 1-pel. The second step of the search checks template SAD costs with 1/4-pel interval for the candidates around the selected candidates from the first step. For the integer MBVD (when existed in the ECM ph_fpel_mbvd_enabled_flag is 0), those intervals are multiplied by 4. The candidates with the lowest TM cost are included into the final MBVD list.
Test 2.4c is a combination of Test 2.4a, Test 2.3b, and Test 2.2c.
Test 2.5: Filtered IBC (JVET-AE0159)
In the test, additional filtered IBC mode is introduced, where a filter is applied to IBC predictor, which is derived by minimizing MSE between current and reference template.
Output of the filter is calculated as follows:
predLumaVal = c0C + c1N + c2S + c3E + c4W + c5P + c6B
The nonlinear term P is represented as power of two of the center sample C and scaled to the sample value range of the content:
P = ( C*C + midVal ) >> bitDepth
The bias term B represents a scalar offset between the input and output and is set to middle luma value (512 for 10-bit content).
This filtered mode is used as an additional mode for non-merge IBC blocks, and it is not used together with IBC-LIC, IBC-CIIP or RR-IBC. For IBC merge modes, this filtering mode is inherited when merge mode list is constructed.
In Test 2.5a, the mode flag is signalled conditioned on the IBC-LIC flag, so when the IBC-LIC flag is true, the flag is signalled and used to indicate whether the tested mode is applied to the current block or not.
In Test 2.5b, the mode flag is signalled before IBC-LIC flag, and the model does not have non-linear term.
In Test 2.5c, the model of Test 2.5a and signalling of Test 2.5b are used, different aspects of encoder design is chosen from both tests considering the best trade-off.
Test 2.6: IBC with non-adjacent spatial candidates (JVET-AE0094)
In the test, several new candidates obtained from the BVs of non-adjacent spatial neighbouring blocks are added to the candidate lists of IBC merge modes and IBC AMVP. A pattern similar to the non-adjacent spatial candidates of regular inter is used to obtain non-adjacent BVs shown in the below figure. The obtained BVs are inserted between the adjacent spatial candidates and the HBVP candidates for both IBC merge and IBC AMVP. Additionally, the restriction that spatial candidates are not used to construct IBC merge candidate list for a 4×4 CU is removed.
In Test 2.6a, the same reference area as for non-adjacent regular merge is reused for the IBC, where up to 4 rows/columns of spatial neighbouring blocks (in terms of the CU height/width) of the CU can be used as reference for IBC AMVP, and up to 7 rows/columns of the spatial neighbouring blocks (in terms of the CU height/width) of the CU are used for IBC merge.
In Test 2.6b, for both IBC AMVP and IBC merge, the spatial neighbouring blocks located up to 4 rows/columns of spatial neighbouring blocks (in terms of the CU height/width) of the CU can be used as reference. Additionally, the restriction that spatial candidates are not used for the IBC merge of a 4×4 CU in ECM-9.0 is removed.
Test 2.6c is the combination of Test 2.6a and Test 2.6b, where the sampling method in Test 2.6a is applied. Additionally, the removal of the restriction that spatial candidates cannot be used for IBC merge of a 4x4 CU from Test 2.6b is used.
Test 2.7: Cross-component merge mode with temporal candidates (JVET-AE0043)
In the current CCP merge mode, a CCP merge candidate list contains spatial adjacent, spatial non-adjacent, and history-based candidates coded in CCLM, MMLM, CCCM, GLM, chroma fusion, and CCP merge modes. After including these candidates, default models can be included to fill the remaining empty positions in the merge list if necessary. To remove redundant CCP models in the merge list, pruning operations are applied. The candidates in the list are reordered based on the SAD costs obtained using the neighbouring template of the current block.
In the test, two types of candidates, namely temporal candidates and shifted temporal candidates, are additionally included in the merge list for non-intra slices. Temporal candidates are added to the merge list after the spatial adjacent candidates. Shifted temporal candidates are added after the history-based candidates.
CCP merge mode is allowed to be applied in non-intra slices for chroma blocks with sizes less than or equal to 16.
Temporal candidates
Temporal candidates are selected from the collocated picture, the positions and inclusion order of the temporal candidates, which are the same as those for temporal candidates in inter merge mode of ECM-9.0, is shown in the next figure below.
Figure: Positions of the temporal candidates
An inclusion order of the temporal candidates is C01 🡪 C02 🡪 … 🡪 C010. If C0i is outside of the picture/slice boundary and C1i is inside of the picture/slice boundary, C1i will be used instead of C0i, where 1 ≤ i ≤ 10. Otherwise, the next inclusion position is checked.
Shifted temporal candidates
Shifted temporal candidates are also selected from the collocated picture. As depicted in the next figure below, the position of the collocated block is shifted by a selected neighbouring motion vector. Consequently, the positions of C0i and C1i are shifted by the same neighbouring motion vector. The inclusion order of shifted temporal candidates is the same as that of temporal candidates.
The adjacent motion vector is selected from the motion vectors of the neighbouring blocks ordered as follows: L0B1 🡪 L1B1 🡪 L0A1 🡪 L1A1 🡪 L0B0 🡪 L1B0 🡪 L0A0 🡪 L1A0 🡪 L0B2 🡪 L1B2. The first motion vector, which uses the collocated picture as the reference picture, is selected. If no such motion vector is found, no shifted temporal candidate is added.
Figure: Positions of the shifted candidates (left) and selecting the neighbouring motion vector (right)
Test 2.8: Extrapolation filter-based intra prediction mode (JVET-AE0076)
In the test, extrapolation filter-based intra prediction is introduced with 3 types of reconstructed areas and 3 filter shapes as shown in the next figure below, a choice of reconstructed area and filter shape is signalled.
The size of reconstructed area depends on the min(blockWidth, blockHeight) and the selected filter shape. For example, when the current block is an 8x16 block and the selected filter shape is 4x4. The aboveSize of reconstructed area is min(8, 16) + 4 – 1 = 11, and the leftSize of reconstructed area is min(8, 16) + 4 – 1 = 11.
Figure: Three types of reconstructed areas (top) and three filter shapes (bottom)
The selected filter moves in the selected reconstructed area with a one-sample step to collect input samples and output samples, then the filter coefficients are derived using CCCM solver.
In the current block, the recursive predictor is derived from the top-left to the bottom-right position by a diagonal prediction order as shown in the next figure below, where the predicted results of the previous diagonal are used.
Figure: Example of generating predictions for different positions in the current block by a diagonal order
To reduce the prediction error, the min and max values from the neighbouring reconstructed area are applied to restrict the output range of each predicted value.
The predicted samples are calculated as follows,
where is the predicted value at (x, y) in the current block. and offset are derived in the selected reconstructed area, is the coefficient of the derived filter, is reconstructed or predicted value used for prediction in the current position.
The tested mode is applied in I-slices only and is only used for the block sizes less than or equal to 32x32.
Test 2.9: Extended search areas for IntraTMP mode (JVET-AE0077)
In the test, IntraTMP search areas are extended by including the bottom-left R5 and top-right R6 areas relative to the current block as they may also be available as shown in the next figure below. In ECM-9.0, the region scan order is R4, R1, R2, and R3. In the tested method, the region scan order is R4, R5, R6, R1, R2, and R3.
Figure: Extended IntraTMP search areas (R5 and R6)
In addition, the current ECM-9.0 always searches all available positions within the current CTU while it may be beyond the specified search range or width/height. In the test, all search areas are restricted to be within the specified search range or width/height specified by (searchRangeWidth, searchRangeHeight) as illustrated in the next figure below. The maximum horizontal and vertical search sizes, i.e., searchRangeWidth and searchRangeHeight, are set to max(5W, 64) and max(5H, 64) respectively. It is not allowed to search any positions outside the specified search range even those available positions are within the current CTU.
Figure: Search areas restriction
Test 2.10: IBC-LIC extension (JVET-AE0078)
In Test 2.10a, IBC-LIC is extended by including 3 additional modes:
- Use only left or use only above LIC template in addition to the current L-shape template,
- Extend the MMLM to IBC-LIC, which allows IBC-LIC to have two linear models in one CU for L-shape template.
Large block-size constraint for ECM IBC-LIC, tested separately in Test 2.10b, is additionally removed in Test 2.10a.In Test 2.10b, large block-size constraint for ECM IBC-LIC is removed, so IBC-LIC can be applied to the CU whose block size is larger than or equal to 32.
Test 2.11: Harmonization of IBC HMVP and IBC-LIC (JVET-AE0084)
In ECM-9.0, the inter LIC flag can be inherited from a HMVP candidate. However, the IBC-LIC flag is not inherited when the motion information from an IBC HMVP candidate is used.
In Test 2.11a, IBC-LIC flag is inherited from an IBC HMVP candidate, which is similar to the inter LIC case.
Test 2.11b: Test 2.11a + Test 2.10a
Results for AI/RA
Test 2.1 is a chroma tool, main benefit on screen content. The gain is small, but it has no impact on run time, and implementation is straightforward according to cross-checkers. Adoption was also supported by independent experts.
Decision: Adopt JVET-AE0100, test 2.1b.
Tests 2.2..2.4 target several aspects of improving IBC in context with GPM and bi-prediction, 2.3b also provides gain on camera-captured content (but had been reported previously to not give gain in case of screen content). In the combination 2.4c (2.2c+2.3b+2.4a), 2.3b is disabled for screen content, and 2.2c is disabled for camera-captured content, whereas 2.4a is enabled for both. gives best gain overall, including screen content. However, the results show that 2.4a does not have benefit for camera capured content when combined with 2.3b, whereas for screen content the gain from 2.2c and 2.4a is somewhat additive.
Decision: Adopt JVET-AE0169 Test 2.2c, in CTC enabled only for screen content.
Decision: Adopt JVET-AE0169 Test 2.3b, in CTC enabled only for camera-captured content.
Decision: Adopt JVET-AE0169 Test 2.4a, in CTC enabled only for screen content.
For 2.5c, a discussion was performed about the benefit of a combination test when the individual results of the elements being combined are not available. Proponents are strongly discouraged to withdraw proposal elements when still including them in a combination, as this will not allow to judge the individual benefits.
Individual results of 2.5a and 2.5b were presented in an update of the summary report on Wednesday July 12.
Gains appear to be somewhat additive, but encoding run time is significantly higher in 2.5b than it is in 2.5c. According to proponents, an encoder with less RDO checks was used in 2.5c. This is also the likely reason for losing a slight bit rate reduction in 2.5c, compared to 2.5b.
It was asked whether the encoder search strategy of 2.5c is the same as 2.5a, such that it could be confirmed that the additional benefit of 2.5c comes by the signalling method of 2.5b, and its non-linear filter term? According to proponents, main gain comes from signalling.
Results were matched according to cross-checker (verbally reported in session, document to be uploaded).
Decision: Adopt JVET-AE0159 Test 2.5c (enabled only for screen content).
Same situation was initially found with 2.6c.I Individual results of 2.6a and 2.6b were presented in an update of the summary report on Wednesday July 12.
Results indicate that the combination of 2.6c provides more gain than the individual parts 2.6a and 2.6b. Gains cannot be expected to be additive, as the basic approach is identical. It can however be concluded that the element of 2.6b (removing the restriction on 4x4 blocks) is also beneficial in combination with 2.6a. The proposal was supported by independent experts (including cross-checkers).
Decision: Adopt JVET-AE0094 Test 2.6c (enabled only for screen content).
Test 2.7 is targeting improvement of chroma, by adding additional temporal candidates in cross-component merge (which had been adopted in the last meeting). Candidates are stored in an 8x8 grid.
It was asked what the benefit of shifted candidates is. According to proponents, it had been 10% of the total gain in the original contribution. It was argued that this element of the proposal might be difficult to implement in hardware. It was however asserted by JVET that this is not of highest important in this stage of exploration, and technology investigated in EE should normally be adopted as proposed.
Decision: Adopt JVET-AE0043 test 2.7.
Test 2.8: It was asked how the mode is used when reconstructed area is not available. It is disabled. It was further pointed out that the gain seems to be very small for small blocks, where the pixel dependency might be critical in real-world implementation. The cross-checker confirms that such dependency is probably present, but would nevertheless support this tools at this stage of exploration.
2.5% increase of encoder run time is pointed out to be not a good tradeoff for 0.16% bit rate reduction in AI. For RA, the gain is less, as it is only applied in I slices currently.
Further study was recommended (in the EE) to reduce encoder run time, and also apply in inter slices.
Reduction of pixel dependency also would be important, but probably is more difficult to achieve.
Test 2.9: Increase in decoder run time by 2.5% by additional search operations. Encoder run time also slightly increased. One cross-checker points out that the increase of decoder run time may be caused by pixel-wise availability checks. It was also pointed out that this is straightforward extension of search range relative to ECM 9, and the tradeoff between gain and run time is still in an attractive range.
Decision: Adopt JVET-AE 0077 test 2.9.
Tests 2.10/2.11 are only beneficial for screen content. From comparing 2.10a/b, it cannot be concluded that removing the large block size constraint is beneficial, as there are no results for the IBC-LIC extension with the large block size constraint of current ECM (such a test had originally been planned, but was not conducted). Proponents are requested to present results of a variant of 2.10a with the constraint, and also in combination 2.11.b to assess whether the gain is additive with 2.11a.
Test 2.11a is a minor change which gives some benefit in class F without impact on run time. However, it should also be enabled for camera-captured content to avoid another high-level flag (according to proponents, no losses would be expected). Therefore, the new combination test 2.11b should be conducted such that the method 2.11a is also used for camera-captured content.
A revised version of JVET-AE0024 was presented on July 16 1700. From the new results, it became obvious that the large block size constraint had no impact on run time, but results are slightly better without the constraint. From the results, 2.10a or 2.11b (combining 2.10a with 2.11a) would be candidates. According to the cross-checker, 2.11a is straightforward to implement. Also the runtime variations are low.
Decision: Adopt JVET-AE0084 Test 2.11b. In the combination, both 2.10a and 2.11a should be enabled for camera captured content, also in CTC.
Inter prediction
Test 3.1: Cross-component residual model for inter prediction (JVET-AE0059)
Cross-component residual model to predict chroma samples from reconstructed luma samples when the block uses inter prediction or intra block copy (IBC) is tested. As illustrated in the next figure below, the cross-component filters are derived using the prediction blocks of luma and chroma. The derived filters are applied to the reconstructed luma block and blended with the prediction blocks of chroma to produce the final chroma prediction blocks. In the blending process the filtered reconstructed luma blocks use blending weight of 0.75 and chroma prediction blocks use blending weight of 0.25.
Figure: Cross-component residual model
Model uses 8-tap filter consisting of 6 spatial luma samples shown in the next figure below, a nonlinear term, and a bias term as follows:
predChromaVal = c0 L0+ c1L1 + c2L2 + c3L3 + c4L4 + c5L5 + c6 nonlinear((L0+L3+1) >> 1) + c7 B,
Figure: 2 Luma samples L0,..,L5 in relation to the chroma sample C.
The filter coefficients are derived using ECM’s division-free Gaussian elimination method and the necessary offsets are applied to samples prior to filter derivation.
For filter coefficient derivation at most 256 chroma samples are used.
The mode flag is only signalled if the TU’s luma Cbf is non-zero and the CU’s predMode is either MODE_INTER or MODE_IBC.
Test 3.1a: Cross-component residual model, where offsets for division free operation are obtained by averaging the samples in the block.
Test 3.1b: Cross-component residual model, where offsets for division free operation are obtained by averaging the four points correspond to the top-left, top-right, bottom-left and bottom-right corners of the blocks.
Test 3.2: Bi-predictive GPM (JVET-AE0046)
In the test, the existing uni-predictive GPM is extended to allow usage of bi-predictive motion vectors for generating motion compensated prediction samples for inter partitions. The method consists of the following elements:
The first element conditionally invokes the existing extraction process that extracts uni-predictive motion vectors from the initial list. The extraction process is invoked only for small blocks 8x8, 16x8 and 8x16. For other larger blocks, the extraction process is bypassed, so the initial list (which may contain merged Bi-MVs) is directly used as the final GPM merge list. The generation of the initial list is the same as before (i.e., the normal merge list generation without any candidate reordering) except that when generating the initial list for larger blocks (i.e., blocks with the extraction process bypassed), the motion vector difference threshold for controlling whether a candidate can be added into the initial list is increased to be one full sample distance.
The second element modifies GPM-MMVD to support bi-predictive motion vector as the base vector. For low-delay pictures, the signalled MVD is applied on top of the L0 and L1 motion vector as in the existing merge MMVD design. For non-low-delay pictures, the bi-predictive motion vector is converted into a uni-predictive motion vector first and then the MVD is applied on top.
The third element modifies GPM-TM to also support bi-predictive motion vectors.
The last element is to enable the BDOF based mv refinement as in the multi-pass DMVR on top of the associated bi-predictive motion vectors for each inter partition.
Test 3.3: High-Accuracy template matching (JVET-AE0087)
In the ECM, template matching is a decoder-side MV derivation method to refine the motion information of the current CU by finding the closest match between a template (i.e., top and/or left neighbouring blocks of the current CU) in the current picture and a block (i.e., same size to the template) in a reference picture.
When TM is used for bi-prediction, the following steps are applied:
- The initial motion vector of list 0 (MV0) is refined to derive a refined MV (MV’0) and a TM cost C0 is calculated;
- The initial motion vector of list 1 (MV1) is refined to derive a refined MV (MV’1) and a TM cost C1 is calculated;
- If C0 is larger than C1, MV’1 is used to derive a further refined MV (MV’’0) by refining MV’0; Otherwise, MV’0 is used to derive a further refined MV (MV’’1) by refining MV’1 .
In high-accuracy template matching, three aspects are tested:
- In Test 3.3a, an additional refinement is applied to TM for bi-prediction, when MV’0 is refined in step 3), MV’’0 is used to derive MV’’1 by refining MV’1; Otherwise, MV’’1 is used to derive MV’’0 by refining MV’0.
- In Test 3.3b, the diamond search pattern used in TM is modified from 8-point to 16-point.
- In Test 3.3c, TM for bi-prediction is enabled when DMVR condition is satisfied.
Test 3.3d: Test 3.3a + Test 3.3c.
Test 3.3e: Test 3.3a + Test 3.3b + Test 3.3c.
Test 3.4: OBMC with HPel flag and BCW weights (JVET-AE0196)
In the current OBMC design, the motion compensation of the neighbouring blocks and subblocks performed by OBMC only considers the motion vectors and reference pictures.
In the test, HPel flags and BCW weights of the neighbouring blocks and subblocks are considered in the OBMC. When processing the top and left borders, the neighbouring BCW weights and HPel flags are used in addition to the motion vectors and reference pictures. The neighbouring HPel flag allows setting the correct AMVR index. For the internal subblock boundaries, the BCW and AMVR indexes of the current CU are also used.
It has to be noticed that the integration of adopted contribution JVET-AD0213 (Bi-prediction LIC) turns out to already include main aspect of this test (BCW weights in OBMC) in the ECM-9.0 which was not the case in ECM-8.0.
Test 3.5: Iterative BDOF pass in multi-pass DMVR (JVET-AE0065)
In the ECM, a multi-pass DMVR has three passes of motion refinement. In the first pass, bilateral matching (BM) is applied to the coding block (CB) to get the block level refined MV pair. In the second pass, the block is divided into 16×16 subblocks and on top of block level refined MV pair obtained in the first pass, BM is applied on each 16×16 subblock to get a 16×16 subblock level refined MV pair. In the third pass, BDOF based MV refinement is applied on the 4×4 or 8×8 subblocks depending on the CB size.
In the test, the current multi-pass DMVR is extended by adding another BDOF-based motion refinement pass as the 4th pass of the multi-pass DMVR (i.e., the 2nd round of BDOF-based DMVR).
Also, the subblock size of the proposed round of the BDOF-based DMVR is adaptively selected depending on the area (i.e., width×height) of the coding block. For blocks smaller than 1024 pixels, the subblock size of 4×4, and otherwise 8×8 is used.
Test 3.7: RPR with new filters for scale factor 1.25x and 1.33x (JVET-AE0150)
In the test, additional set of 8-bit precision RPR filters consisting of three tables (respectively 12-tap for luma, 6-tap for chroma, and 10-tap for affine) is introduced with the target to improve the performance for scaling ratios 1.25x and 1.33x. These new set of filters is used for scaling ratios in between 1.1x and 1.35x.
In the tests, two PSNR values are used:
- PSNR1: it is measured on the decoded picture size by calculating MSE for each picture, accumulating MSEs for all pictures and calculating average PSNR from the accumulated MSE.
- PSNR2: it is measured after upsampling of the decoded picture if the decoded picture size is in smaller resolution, the existed ECM interpolation filters are used for the upsampling.
Test 3.8: Combination of Test 3.3e and Test 3.5 (JVET-AE0091)
Results for RA/LB
Test 3.1: Increase of encoder run time should be expected in test 3.1b (RA results not complete yet).
It was commented that the method should better be called “cross component chroma prediction in inter”, as it is not the residual that is predicted.
It was also commented that a blending with 0.75/0.25 weights was newly introduced, which according to proponents provides better performance in blending. It was requested to provide quantitative information about the gain by blending
Results were presented Wednesday 8:30. Test 3.1b has slightly less encoder run time (which is explainable by having less checks), and bit rate reduction is almost identical. This was also confirmed by cross-checker.
It was asked why a different cross-component model was used than the one in the ECM. It was answered by proponent that downsampling is avoided. The unification might be desirable, but is not of prior importance at this moment.
Decision: Adopt JVET-AE0059 Test 3.1b.
Test 3.2: 0.2% bit rate reduction in RA with slight increase in encoding time. According to one cross-checker it may increase the worst-case memory access, which may however not be too severe, as the proposal is not used in case of small block size, and multi-hypothesis is also used in other tools of ECM. Several experts (including cross-checkers) supported the proposal. It was however suggested to modify the software such that the memory for motion vector storage is reduced in the encoder.
Decision: Adopt JVET-AE0046 Test 3.2.
Test 3.3x, 3.5, and combination 3.8: Test 3.3c seems to have no benefit standalone, but according to proponents has some benefit in combination with 3.3a. In combination, gains of 3.3e and 3.5 are somewhat additive in RA, whereas only 3.3e contributes gain LB (BDOF of 3.5 not applicable in LB). Several experts supported the combination, and no objection was raised.
Decision: Adopt JVET-AE0091 Test 3.8.
Test 3.4 does not show benefit in compression, might be due to new adoptions in ECM 9. No action was taken on this.
Test 3.7 is about RPR (non-CTC test conditions deduced from RPR specific VTM test conditions which only used 1.5x and 2x). The results indicate that improvement is possible by using different filters for small factors of resolution change. Though it is not of prior importance for the exploration, it may be interesting to provide such implementation in the software for experimentation of interested parties.
Decision (SW): Adopt JVET-AE0150 (not in CTC, not in the ECM description).
Transforms and coefficient coding
Test 4.1: Shifting quantizer center (JVET-AE0125)
In the test, a quantization offset is added to the quantized level, the offset is quantization index dependent, a look-up table is used to derive the offset .
where x is the dequantized coefficient, y and y’ are the quantization indices, Q-1 is the dequantization operation, T is a look-up table
.
Test 4.2: Large NSPT (JVET-AE0086)
The large NSPT kernels are tested for 4x32/32x4 and 8x32/32x8 blocks, of which kernel matrix dimensions are 20x128 and 24x256. Therefore, 20 and 24 transform coefficients are generated by applying the two types of kernel matrices, respectively, which are placed from DC position following scan order. The remaining 108 and 232 positions in each transform block are zeroed-out, respectively.
Large NSPT kernels are applied in the same way as for other block sizes 4x4, 4x8/8x4, 4x16/16x4, 8x8, and 8x16/16x8 and those NSPT kernels are not changed.
Test 4.3: Context modelling for transform coefficients for LFNST/NSPT (JVET-AE0102)
In the ECM, the causal 2D neighbourhood of coefficients is used to model the context to parse the sig_coeff_flag, gt1_flag, gt2_flag as shown in the next figure below. For LFNST, DCT-II coefficients are placed into the coefficient block using diagonal reordering.
Figure: Context modelling for LFNST coefficients in the ECM
In the test, when LFNST/NSPT is applied, the previous 5 coefficients in the coding order are used for context derivation instead of 2D neighbourhood.
Figure: Context modelling for LFNST coefficients in TEST4.3
In the ECM, the lfnstIdx is signalled after all the coefficients in a CU. In the test, lfnstIdx is signalled after all last_sig_coeff_pos syntax elements in a CU since lfsntIdx is required for parsing the transform coefficients.
Test 4.4: InterMTS for IBC and IntraTMP (JVET-AE0116)
InterMTS enabled for IBC and IntraTMP is tested. The following tests are performed:
Test 4.4a: InterMTS is enabled for IBC-coded blocks in AMVP mode. MTS is not enabled for IBC merge.
Test 4.4b: InterMTS is enabled for IBC-coded blocks in AMVP mode as in Test 4.4a. In addition, IntraTMP is utilizing InterMTS instead of IntraMTS. Consequently, the number of MTS candidates for IntraTMP has been fixed to 4 instead of adaptively setting it to 1, 4 or 6 as in IntraMTS.
Test 4.4c: IntraMTS is disabled for IntraTMP coded blocks.
Results for AI/RA (LFNST not enabled in LB)
Test 4.1: Shifting quantizer center provides 0.1% gain without any complexity impact.
Decision: Adopt JVET-AE0125 Test 4.1.
Test 4.2: Reasonable tradeoff runtime vs. bit rate reduction, additional memory consumption of larger kernels not relevant in cntext of exploration.
Decision: Adopt JVET-AE0086 Test 4.2.
Test 4.3: It was commented that the usage of immediately preceding coefficients might have some latency impact (VVC did not use coefficients from same diagonal for that reason), but for the purpose of exploration several experts (including cross-checkers) supported the proposal.
Decision: Adopt JVET-AE0102 Test 4.3.
Test 4.4x: Enabling MTS for IBC does not provide relevant gain to justify the increase of encoding time (caused by additional RDO checks). No action was taken on this.
Loop filtering
Test 5.1: CCSAO with extended edge classifiers and history offsets (JVET-AE0151)
In ECM-9.0, CCSAO uses band and edge classifiers switched at CTB level. Classifier parameters/offsets are derived at encoder and signalled for each slice independently.
In Test 5.1a, to reduce the signalling overhead, CCSAO inheritance scheme is introduced, where the offsets/parameters of some coded pictures are stored at both encoder and decoder which are allowed to be used as the CCSAO offsets/classifiers of future pictures. An index is signalled in SH to indicate which candidate in the CCSAO parameter storage is selected for the current slice. The candidates in the CCSAO parameter storage are updated in a FIFO manner and refreshed at IDR pictures.
In Test 5.1b, on top of Test 5.1a, beside the existing CCSAO edge classification, one new edge classification is added, which is a subset of the original one with less edge range divisions.
where is calculated by comparing a sample difference in one direction with a threshold .
Additionally, the component used for edge classification can be selected from one of all three components. Same to the existing CCSAO design, the selected edge classifier and edge component are decided by encoder and signalled in the SH.
Test 5.2: Improved fixed filters for ALF (JVET-AE0139)
In ECM-9.0, two Laplacian-based classifiers (one for each fixed filter) are applied to a 2x2 block. In each classifier, activity and directionality values are derived based on vertical, horizontal, and diagonal gradients using a window surrounding each 2x2 block. Then a class index is determined based on the activity and directionality values. Two 13x13 diamond shaped fixed filters are selected from the two filter sets by using the derived two class indices. Then the two fixed filters ( with ) are applied to the ALF input samples. Finally, a signalled filter is applied to the ALF input samples, samples before the deblocking filter (DBF), outputs of the two fixed filters, output of a gaussian filter and the residual data.
In Test 5.2a, both fixed filters are applied to samples before DBF and ALF input, where additional diamond 9x9 filter is used for the samples before DBF. The shape of the first fixed filter applied to the ALF input samples is reduced from 13x13 to 9x9, and the shape of the second fixed filter, which is 13x13, applied to ALF input is unchanged as shown in the table below.
ECM-9.0 | Proposed method | ||
ALF input | Samples before DBF | ALF input | |
Fixed filer | 13x13 | 9x9 | 9x9 |
Fixed filer | 13x13 | 9x9 | 13x13 |
In Test 5.2b, on top of Test 5.2a, the classifiers of the fixed filters are extended. For each 2x2 block, the mean value of a surrounding window is calculated. Then, for each sample of this window, the difference between the sample value and the mean value is calculated. A scaling factor is determined based on the activity value derived from a Laplacian classifier. The square root of the sum of the squared differences is further quantized to by a scaling factor. The value of is an integer between 0 and 7, inclusively. With i=0, 1, let denote the classifier from the classifier of i-th fixed filter in ECM-9.0. Then the proposed class index is derived as
.
The total number of the fixed filters is not changed.
In Test 5.2c, on top of Test 5.2b, fixed filter is applied to outputs of (instead of ALF input) and samples before DBF.
Test 5.3: Combination of Test 5.1b and Test 5.2c (JVET-AE0152)
One expert commented that the changes in 5.1a and 5.1b are straightforward and supported it.
It was report that an additional buffer of approcimately 5Kbyte is necessary for storage.
It was confirmed that the data flow in 5.2x is not changed relative to the current ECM.
It can be expected that the gains of CCSAO modifications and ALF modifications are non-overlapping, and it was also reported that preliminary results of 5.3 indicate additive gains.
Decision: Adopt JVET-AE0151 Test 5.1b.
Decision: Adopt JVET-AE0139 Test 5.2c.
It was commented that the results of 5.3 (combination test) are expected to be provided for information.
EE2 contributions: Enhanced compression beyond VVC capability (24)
There was no presentation or discussion about specific proposals in this category.
For actions decided to be taken, see section 5.2.1, unless otherwise noted.