JVET-M0030 CE10: Summary report on combined and multi-hypothesis prediction [C.-W. Hsu, M. Winken]
Five sub CEs are created to test different methods of combined predictions, including:
- CE10.1: multi-hypothesis prediction,
- CE10.2: overlapped block motion compensation,
- CE10.3: multiple shape prediction partitions,
- CE10.4: diffusion filtering of inter- and intra-prediction signals,
- CE10.5: local illumination compensation.
There are 13, 4, 1, 2 and 2 tests for each sub CE, respectively.
CE proposals are listed as follows,
Proposal Doc # | Corresponding tests | Author(s) | Title |
CE10.1.1 | M.-S. Chiang, C.-W. Hsu, Y.-W. Huang, S.-M. Lei (MediaTek) | CE10.1.1: Multi-hypothesis prediction for improving non-skip & non-merge inter mode and merge mode | |
CE10.1.2 | M. Winken, H. Schwarz, D. Marpe, T. Wiegand (HHI) | CE10: Multi-hypothesis inter prediction (Test 10.1.2) | |
CE10.1.3 | CE10: Simplification on Combined Inter-Intra Prediction (Test 10.1.3) | ||
CE10.1.4 | M.-S. Chiang, C.-W. Hsu, Y.-W. Huang, S.-M. Lei (MediaTek) | CE10.1.4: Simplification of combined inter and intra prediction | |
CE10.1.5 | W. Xu, B. Wang, H. Yang, J. Chen (Huawei), M.-S. Chiang, C.-W. Hsu, Y.-W. Huang, S.-M. Lei (MediaTek) | CE10: Simplification on Combined Inter-Intra Prediction with size restriction (Test 10.1.5) | |
CE10.2.1 | Z.-Y. Lin, T.-D. Chuang, C.-Y. Chen, C.-W. Hsu, C.-C. Chen, Y.-C. Lin, Y.-W. Huang, S.-M. Lei (MediaTek), X. Xiu, Y. He (InterDigital) | CE10.2.1: Uni-prediction-based CU-boundary-only OBMC | |
CE10.2.2 | Z.-Y. Lin, T.-D. Chuang, C.-Y. Chen, C.-W. Hsu, C.-C. Chen, Y.-C. Lin, Y.-W. Huang, S.-M. Lei (MediaTek), X. Xiu, Y. He (InterDigital) | CE10.2.2: Integer-MV-based CU-boundary-only OBMC | |
CE10.2.3 | Y.-C. Lin, C.-C. Chen, C.-W. Hsu, Z.-Y. Lin, T.-D. Chuang, C.-Y. Chen, Y.-W. Huang, S.-M. Lei (MediaTek) | CE10.2.3: Subblock OBMC with uni-prediction-based OBMC at CU boundaries | |
CE10.2.4 | Y.-C. Lin, C.-C. Chen, C.-W. Hsu, Z.-Y. Lin, T.-D. Chuang, C.-Y. Chen, Y.-W. Huang, S.-M. Lei (MediaTek) | CE10.2.4: Subblock OBMC with integer-MV-based OBMC at CU boundaries | |
CE10.3.1 | CE10.3.1: AMVP mode for triangle prediction | ||
CE10.4.1 CE10.4.2 | J. Rasch, A. Henkel, J. Pfaff,H. Schwarz, D. Marpe, T. Wiegand (HHI) | CE10: Uniform Directional Diffusion Filters For Video Coding | |
CE10.5.2 | CE10: Low pipeline latency LIC (test 10.5.2) | ||
CE10.5.3 | CE10: LIC confined within current CTU (test 10.5.3) |
CE10.1: Multi-hypothesis prediction
# | Supported modes | Signalling of hypothesis | # of additional hypotheses | Block constraint in luma samples | BW reduction technique | Hypothesis inheritance | Reference frame access constraints | Color components |
CE10.1.1.a | AMVP (uni only) | merge index | 1 | >= 8x8 | up to 2 | luma + chroma | ||
CE10.1.1.b | skip/merge (uni or bi) | merge index | 1 or 2 | full-pel for additional hypotheses | no temporal | up to 4 | luma + chroma | |
(luma) | spatial: no CTU constraints | |||||||
CE10.1.2.a | merge | ref index + mvp index + | 1 or 2 | > 8x8 | full-pel for additional hypotheses | no temporal | up to 4 | luma + chroma |
(luma + chroma) | spatial: from left or within CTU | |||||||
CE10.1.2.b | merge | ref index + mvp index + | 1 or 2 | > 8x8 | full-pel for additional hypotheses | not allowed | up to 4 | luma + chroma |
(luma + chroma) | ||||||||
CE10.1.2.c | merge | ref index + mvp index + | 1 or 2 | > 8x8 | full-pel for additional hypotheses | no temporal | up to 2 | luma + chroma |
(luma + chroma) | spatial: from left or within CTU | |||||||
CE10.1.2.d | merge | ref index + mvp index + | 1 or 2 | > 8x8 | full-pel for additional hypotheses | no temporal | up to 4 | luma |
(luma + chroma) | spatial: from left or within CTU |
The test results for this aspect are summarized as follows,
Proposal | Config. | Y | U | V | EncT | DecT | |
CE10.1.1.a | RA | -0.17% | -0.13% | -0.10% | 107% | 100% | |
LB | 0.00% | 0.10% | 0.25% | 107% | 101% | ||
CE10.1.1.b | RA | -0.17% | -0.25% | -0.22% | 103% | 102% | |
LB | -0.12% | -0.19% | -0.17% | 103% | 102% | ||
CE10.1.2.a | RA | -0.29% | -0.16% | -0.12% | 107% | 97% | |
LB | -0.41% | -0.04% | -0.04% | 110% | 99% | ||
CE10.1.2.b | RA | -0.20% | -0.02% | -0.05% | 107% | 96% | |
LB | -0.30% | -0.10% | -0.17% | 111% | 102% | ||
CE10.1.2.c | RA | -0.20% | -0.10% | -0.08% | 105% | 97% | |
LB | -0.20% | -0.19% | -0.24% | 107% | 100% | ||
CE10.1.2.d | RA | -0.29% | 0.03% | 0.04% | 107% | 96% | |
LB | -0.40% | 0.06% | -0.02% | 110% | 99% |
According to the CE description, for investigating the impact of multi-hypothesis inter prediction on cache-related aspects, the cache model compiled into the reference decoder software by setting #define JVET_J0090_MEMORY_BANDWITH_MEASURE to 1 is used, in conjunction with the cache config file provided in JVET-K0451. For each bit stream, the decoder outputs a total hit ratio when using the cache model. The following table compares the average hit ratios (in percentage) of the tests with those of VTM-3.0:
Data quoted from JVET-M0176:
Random Access Main 10 | ||
VTM-3.0 (%) | CE10.1.1.b (%) | |
Class A1 | 99.4735 | 99.4663 |
Class A2 | 99.5509 | 99.5461 |
Class B | 99.5308 | 99.5249 |
Class C | 99.3467 | 99.3435 |
Class E | ||
Overall | 99.4755 | 99.4702 |
Class D | 99.3402 | 99.3401 |
Class F (mandatory) | 98.8734 | 98.8607 |
Low delay B Main10 | ||
VTM-3.0 (%) | CE10.1.1.b (%) | |
Class A1 | ||
Class A2 | ||
Class B | 99.6490 | 99.6414 |
Class C | 99.5511 | 99.5415 |
Class E | 99.5562 | 99.5517 |
Overall | 99.5854 | 99.5782 |
Class D | 99.5090 | 99.496 |
Class F (mandatory) | 99.0382 | 99.0268 |
Data quoted from JVET-M0425:
Random Access | |||||
| VTM-3.0 | CE10.1.2.a | CE10.1.2.b | CE10.1.2.c | CE10.1.2.d |
A1 | 99.4888 | 99.4781 | 99.4878 | 99.4858 | 99.4866 |
A2 | 99.5513 | 99.5474 | 99.5505 | 99.5488 | 99.5505 |
B | 99.5340 | 99.5259 | 99.5330 | 99.5326 | 99.5345 |
C | 99.3629 | 99.3593 | 99.3615 | 99.3630 | 99.3634 |
Overall | 99.4843 | 99.4777 | 99.4832 | 99.4826 | 99.4838 |
D | 99.3416 | 99.3379 | 99.3412 | 99.3414 | 99.3415 |
F | 98.9343 | 98.9258 | 98.9267 | 98.9297 | 98.9291 |
Low Delay B | |||||
| VTM-3.0 | CE10.1.2.a | CE10.1.2.b | CE10.1.2.c | CE10.1.2.d |
B | 99.6547 | 99.6462 | 99.6531 | 99.6539 | 99.6567 |
C | 99.5650 | 99.5548 | 99.5594 | 99.5656 | 99.5628 |
E | 99.5242 | 99.5147 | 99.5187 | 99.5224 | 99.5225 |
Overall | 99.5813 | 99.5719 | 99.5771 | 99.5806 | 99.5807 |
D | 99.5715 | 99.5657 | 99.5683 | 99.5727 | 99.5737 |
F | 99.1095 | 99.0934 | 99.0978 | 99.1081 | 99.1034 |
1.1.a uses two hypotheses in uni prediction and merge. Basically, the same prediction could be invoked by using bi prediction. However, only one AMVP list is generated, and it may save some signalling. Encoder time increases by 7%, no gain in LB.
Are gains of 1.1.a/b additive? They are said to have been almost additive before this CE cycle, but in the current CE the combination was not tested.
Worst case memory BW of 1.1.a is uncritical, of 1.1.b was analysed as roughly 80% of VTM.
Combination of 1.1.a/b would be somewhat similar to 1.2.a, which uses multi hypothesis for both merge and AMVP, and additionally allows kind of different weighting of the hypotheses. 1.2.a is the “full set” of functionality of this proposal, whereas b..d are somewhat simplifications, mainly for the benefit of saving local storage (b), or memory bandwidth (c,d). d still uses up to 4 hypotheses (2 each in bi with >=8x16 block size restriction), but only for luma, which is the reason for worse performance in chroma. It is reported that worst case memory BW is 0.97%
It is also mentioned that potentially gains of 1.1.a might add up to 1.2.x.
Generally, the gain of all 1.1.x and 1.2.x proposals is significantly less than it was over VTM2 (cut by half or even more). In particular, the 1.1.b and 1.2.x proposals add need for more building blocks. 1.1.a is relatively simple and ce re-use existing memory access structures and MC logic, but for getting the gain of 0.17% for RA (no gain for LB), the encoder runtime also increases by 7%.
By doing more encoder checks and increasing runtime by 7%, probably it should be possible to get similar gain without a syntax change.
1.1.3…5 target simplification of VTM3 (CIIP)
For simplification aspects, the tests and corresponding results are summarized as follows, where the current design in VTM-3.0 and the differences between each test and VTM-3.0 are listed
# | Proposal | Supported intra modes | Reference sample smoothing | PDPC | Weights for combined | Notes | |
VTM-3.0 | 4: DC, PL, VER, HOR | Yes | Yes | PL, DC: equal weights | area >=64 | ||
VER, HOR: fixed, position dependent weights | W<128 && H<128 | ||||||
CE10.1.3.a | 1: PL | Yes | Yes | equal weights | |||
CE10.1.3.b | 4: DC, PL, VER, HOR | No for PL* | No for PL | PL, DC: equal weights | |||
VER, HOR: fixed, position dependent weights | |||||||
CE10.1.3.c | 4: DC, PL, VER, HOR | Yes | Yes | equal weights | |||
CE10.1.3.d | 1: PL | No for PL* | No | equal weights (4:4) | CE10.1.3.a + CE10.1.3.b | ||
LD frame: 6:2* | |||||||
CE10.1.4 | VTM constraints + | ||||||
area <=1024 | |||||||
CE10.1.5.a | 1: PL | equal weights | VTM constraints + | CE10.1.3.a + CE10.1.4 | |||
area <=1024 | |||||||
CE10.1.5.b | 1: PL | No | equal weights (4:4) | VTM constraints + | CE10.1.3.a + CE10.1.4 + | ||
LD frame: 6:2* | area <=1024 |
* Not in original CE description.
Test results for this aspect are summarized as follows,
Proposal | Config. | Y | U | V | EncT | DecT | |
CE10.1.3.a | RA | -0.01% | 0.13% | 0.09% | 99% | 100% | |
LB | 0.10% | 0.35% | 0.36% | 99% | 100% | ||
CE10.1.3.b | RA | 0.01% | 0.00% | -0.02% | 100% | 100% | |
LB | 0.01% | -0.01% | -0.07% | 100% | 101% | ||
CE10.1.3.c | RA | 0.05% | 0.12% | 0.10% | 100% | 100% | |
LB | 0.10% | 0.18% | 0.05% | 100% | 100% | ||
CE10.1.3.d | RA | 0.02% | 0.07% | 0.03% | 99% | 100% | |
LB | 0.01% | 0.12% | 0.13% | 99% | 100% | ||
CE10.1.4 | RA | 0.01% | 0.02% | 0.05% | 99% | 102% | |
LB | 0.04% | -0.04% | 0.13% | 100% | 100% | ||
CE10.1.5.a | RA | 0.02% | 0.16% | 0.12% | 98% | 100% | |
LB | 0.11% | 0.23% | 0.29% | 98% | 100% | ||
CE10.1.5.b | RA | 0.06% | 0.07% | 0.03% | 98% | 100% | |
LB | 0.07% | 0.15% | 0.03% | 98% | 99% |
It is commented that 10.1.3.b/d was not originally planned in the CE.
In the context of reviewing 10.1.3.x proposals, the following aspects were identified in CIIP:
- does it need a specific MPM derivation? If there was only one mode (as in 10.1.3a/d) it is not needed at all, and it is mentioned that there are proposals just suggesting fixed length coding
- CIIP does not have any serious latency issue. Therefore, simplifications that remove some processing steps that are otherwise used in intra prediction is not necessary.
- Sample-wise unequal weighting is not an implementation issue, whereas equal weighting would be preferrable, unless it costs compression performance or causes qualty problems.
Beyond the approachs tested in 10.1.3 more study is necessary on these aspects.
10.1.4 requests restricting the maximum block size of CIIP to 32x32 (now it is 64x64). The motivation is saving need for additional memory. As however the maximum size of a inter-only or intra-only prediction block is large anyway, the real need for such a restriction is not obvious.
10.1.5 is combining some 10.1.3.x and 10.1.4.
CE10.2: OBMC
# | Proposal | Applied boundaries | # of blending lines | BW reduction technique | Runtime reduction technique | Cost reduction technique |
CE10.2.1 | CU boundaries | 2: width < 8 (left); height < 8 (top) | apply OBMC to uni-prediction blocks | 1. reuse L shape buffer | CTU row buffer removal | |
3. apply MV merge | ||||||
CE10.2.2 | CU boundaries | 2: width < 8 (left); height < 8 (top) | use integer MV to generate OBMC region | 1. reuse L shape buffer | CTU row buffer removal | |
3. apply MV merge | ||||||
CE10.2.3 | CU boundaries + | CU boundary: | CU boundary: | 1. reuse L shape buffer | CTU row buffer removal | |
(when sub-block OBMC is applied, the number of blending line at CU boudnary is 1) | sub-block: | 3. apply MV merge | ||||
CE10.2.4 | CU boundaries + | CU boundary: | CU boundary: | 1. reuse L shape buffer | CTU row buffer removal | |
(when sub-block OBMC is applied, the number of blending line at CU boudnary is 1) | sub-block: | 3. apply MV merge |
All proposals impose a CU size constraint >= 64 samples. It is requested to provide a more detailed analysis of the memory bandwidth. Likely, for processing on-the-fly, the worst case would be that the current block is 8x8, and all of the top and left neighbours are 4x4. Another option would be to store samples from the current block that were already fetched to perform interpolation in the neighbour blocks. In that case, it should be reported how large the additional local buffer would be, for all four methods.
Additional analysis is shown Saturday afternoon for 10.2.1 and 10.2.2 (to be provided in updated version of JVET-M0178/9). This analysis indicates that for those tw approaches the worst-case memory bandwidth of VVC is not increased for on-the-fly fetch (and even a little less for 10.2.2), or if the pre-generation method is used, the memory BW is even lower, but 1.96/1.28 kByte is necessary as local buffers.
In terms of processing, it is reported that worst case number of sample interpolations is not increased relative to the current bi prediction case. Method 10.2.1 uses uni prediction in a 8x8 block, and additionally needs to interpolate four 4x4 areas with other vectors for OBMC. Method 10.2.2 does not use any interpolations for OBMC. For 10.2.1, weighted superposition requires at most 4 shifts and 2 adds per sample. Furthermore, since the operations are locally varying, some additional logic is necessary. For 10.2.2., the latter numbers duplicate. Furthermore, 10.2.2 is more challenging in terms of local memory acces, as 6 different sources need to be blended.
10.2.1 seems manageable from complexity perspective (but definitely adds some complexity).
The test results are summarized as follows,
# | Proposal | Config. | VTM | ||||
Y | U | V | EncT | DecT | |||
CE10.2.1 | RA | -0.27% | -0.58% | -0.62% | 104% | 103% | |
LB | -0.36% | -0.50% | -0.51% | 106% | 105% | ||
CE10.2.2 | RA | -0.37% | -0.91% | -0.90% | 105% | 104% | |
LB | -0.39% | -0.55% | -0.48% | 106% | 106% | ||
CE10.2.3 | RA | -0.39% | -0.59% | -0.63% | 105% | 104% | |
LB | -0.39% | -0.69% | -0.76% | 107% | 108% | ||
CE10.2.4 | RA | -0.49% | -0.91% | -0.87% | 106% | 105% | |
LB | -0.45% | -0.62% | -0.34% | 108% | 110% | ||
As a general note, the gain in compression performance is lower than it was with VTM2 (approx. half). However, from the results, OBMC even gives more gain for LB than for RA, and LB is in overall performance of VVC still worse than RA, this is assessed to be valuable enough. Some support, and no opposition is raised in Track B against adopting it.
It was initially agreed in Track B to adopt JVET-M0178. Specification text was available. This decision was later reverted in the JVET Sunday plenary (see the notes in section 9.1).
This would have a high-level flag for disabling it.
CE10.3: Multiple shape prediction partitions
In CE10.3, the goal is to test AMVP to be combined with non-rectangular prediction partitions within one CU. The tests and corresponding results are summarized as follows:
# | Proposal | Config. | VTM | Description | ||||
Y | U | V | EncT | DecT | ||||
CE10.3.1 | RA | -0.06% | -0.03% | -0.06% | 123% | 100% | Two triangle prediction units (PUs) for AMVP mode | |
LB | -0.07% | -0.10% | -0.01% | 122% | 103% | Restricted to uni-prediction |
From the results (low gain versus high increase in encoder run time), not worthwhile to consider
CE10.4: Diffusion filters for intra and inter prediction
In CE10.4, the goal is to test prediction to be combined using filtering, where three types of filters, two directional filters and one uniform filter are used. The filters are FIR filters, applied on the prediction signal, not iterative. The 1D filters are 9-tap, symmetric, only 1 multiplication, otherwise shifts. The 2D filter is a 5-tap diamond shape, only shift/add operations. At the boundaries, pixel replication is used.
The tests and corresponding results are summarized as follows,
# | Proposal | Config. | VTM | Description | ||||
Y | U | V | EncT | DecT | ||||
CE10.4.1 | AI | -0.19% | -0.03% | -0.03% | 123% | 101% | Uniform Diffusion filters with encoder speedup and simplified filter masks | |
RA | -0.44% | -0.51% | -0.35% | 113% | 98% | |||
LB | -0.17% | 0.27% | 0.59% | 114% | 98% | |||
LP | -0.86% | -0.26% | -0.11% | 119% | 97% | |||
CE10.4.2 | AI | -0.17% | -0.03% | -0.05% | 119% | 101% | Like CE10.4.1, but | |
RA | -0.42% | -0.54% | -0.42% | 111% | 98% | • Simplified encoder search for intra blocks | ||
LB | -0.17% | 0.41% | 0.43% | 113% | 98% | |||
LP | -0.87% | -0.15% | -0.10% | 118% | 96% |
Generally, it was agreed that this method gives interesting gain, and is straightforward to implement at the decoder. Two concerns are raised:
- As it needs to be run after the prediction signal is generated, it produces additional delay in the intra prediction loop.
- For inter blocks, it can be used for 128x128 CU, which would break concepts of 64x64VPDU.
It is requested to provide additional results without using the method in intra prediction (and the intra part of CIIP), and restrict the largest CU size to 64x64.
It is also reported that a late contribution (JVET-M0848) provides new results for the same method of 10.4.2 with some (small) encoder speedup and slightly increased performance (-0.44% for luma).
Furthermore, it would be desirable to use only the prediction samples (no reconstructed samples from current picture). A version which did that was shown in the previous CE.
This was revisited (Track B Thursday 17 January1630) after new results were made available in JVET-M0042v3. The restrictions were implemented as requested, including not using reconstructed samples from neighbour blocks. Results are as follows:
Random Access Main 10 | |||||
Over VTM-3.0 | |||||
Y | U | V | EncT | DecT | |
Class A1 | -0.27% | -0.45% | -0.19% | 108% | 100% |
Class A2 | -0.35% | -0.19% | -0.16% | 109% | 101% |
Class B | -0.41% | -0.37% | -0.45% | 109% | 99% |
Class C | -0.15% | -0.11% | -0.12% | 110% | 99% |
Class E | |||||
Overall | -0.30% | -0.28% | -0.25% | 109% | 100% |
Class D | -0.11% | -0.18% | -0.20% | 110% | 102% |
Class F (mandatory) | -0.09% | -0.13% | -0.11% | 108% | 101% |
The spec text was available and straightforward.
One expert points out that the multiplication by 6 can be implemented by shift and add.
LB results are not complete yet, but by tendency similar as in the original scheme, lower gain than for RA.
Further study in CE, along with LIC and post rec filters. This CE should also identify how potentially gains add up. The restriction of not using recosntructed samples from neighbouring inter blocks may not be necessary.
CE10.5: Local illumination compensation
In CE10.5, the goal is to test prediction to be combined using linear model derived from reconstructed and reference neighbouring samples. The tests and corresponding results are summarized as follows:
# | Proposal | Config. | VTM | Description | ||||
Y | U | V | EncT | DecT | ||||
CE10.5.2 | RA | -0.56% | -0.31% | -0.37% | 135% | 99% | 1. Remove all encoder and decoder LIC processes other than the reconstruction stage | |
LB | -0.51% | -0.44% | -0.58% | 140% | 101% | 2. Modify bS calculation of DBF for LIC boundary. | ||
CE10.5.3 | RA | -0.44% | -0.31% | -0.34% | 134% | 101% | Based on CE10.5.2 | |
LB | -0.39% | -0.40% | -0.42% | 138% | 100% | LIC is confined to use reference samples only from the current CTU |
The method of 10.5.2 is mainly for the benefit of encoders, and does not solve the latency issue that LIC imposes on decoder pipeline. The problem is that a current block needs to wait for reconstruction of the neighbours, and then it requires a certain number of cycles until the parameters of the linear model are computed. Of particular concern is the fact that many decoder implementations target processing inter and intra coded blocks independently in order to make best benefit of parallel processing. Typically, inter coded blocks of a CTU are decoded first. For this, has a non-negligible complexity impact (in particular for parallel processing) if LIC of an inter coded block uses the reconstruction of intra coded neighbours. If such an approach (disabling LIC from intra coded neighbours, including CPR) would be combined with the approach of 10.5.3 (only using reconstructed inter coded samples from current CTU), the situation would be better, however would still mean that the inter reconstruction within a CTU would need to be sequential (similar as the situation in intra is). Regardless of that, LIC has some additional processing complexity (which doubles in case of bi prediction) which needs to be justified by performance. The processing should also be aligned with the 64x64 VPDU concept. (See further notes on these aspects under JVET-M0873.)
Further study was recommended on these aspects.
A BoG (coordinated by C.-W. Hsu and M. Winken) was established to review CE10 related proposals, and suggest candidates for further study in a CE. See the further notes for the discussion of the BoG report JVET-M0873.
The subsequent notes only contain abstracts copied from the documents. Actions taken are noted above under JVET-M0029.