JVET-L0670 Simplified DMVR for inclusion in VVC [S. Esenlik, A. M. Kotra, B. Wang, H. Gao, J. Chen (Huawei), S. Sethuraman (Ittiam)] [late]
This contribution document reports the result of the combination of CE9 tests to reduce the complexity of DMVR in BMS2.1. It is asserted that the proposed DMVR is algorithm provides significant reduction in complexity while increasing the coding gain. Reported gain of 1.67% in luma BD-rate, with Cb and Cr 1.68% and 1.80% chroma BD-rate gain is reported with 1% increase in encoding time and 6% increase in decoding time compared to the BMS anchor with VTM encoder configuration.
The contribution reports the combination of following CE tests:
- Reference sample padding for eliminating memory BW increase (CE9.2.2)
- Integer based DMVR to eliminate intermediate interpolation filters and buffers (CE9.2.1)
- Use refined MV from top and top-left CU (CE9.1.4)
- Parametric error surface based sub-pel refinement (CE9.2.5)
- Disable DMVR for small blocks and subsampled MRSAD (CE9.2.9f)
- Early-termination based on MV difference between merge candidates(CE9.2.13a)
The proposed modifications are independently tested in CE9.
The proponent proposed to adopt the contribution to VVC.
In the previous CE, the main concerns are internal buffer size. CE9.2.1 solves the problem of the internal buffer problem.
It is commented the performance is interesting; but, due to the adoptions at the current meeting, there might be interactions of the method with the new adoptions (e.g., because this has a symmetric motion assumption that is also used in another technique adopted at the meeting - MMVD).
The maximal CU size is 128x128.
The proponent mentioned that if not using MRSAD the loss is 0.24%.
BoG recommendation: Study in the CE.
In Track B review of the BoG report, the proponent said they had tested the scheme in combination with an MMVD method similar to what was adopted (disabling DMVR for blocks that use MMVD and disabling DMVR for blocks Nx8/8xN/Nx4/4xN blocks), and found a BD rate gain of DMVR 1.1% for RA.
Another participant said they estimated gain as 0.9% and commented that since the proposal was late, they could not check it more thoroughly.
A participant said that since this is applied to large blocks, it would add a requirement of large amount of cache memory capacity (~128k), and suggested considering limiting the scheme to smaller blocks (or forcing a split to smaller blocks) than the proposed size range of up to 128x128.
A different variant of DMVR had been tested with a forced split of this sort (splitting to 32x32) as reported in JVET-L0098, with a small impact on coding efficiency (about 0.07% loss). It was commented that such a split might cause artefacts, since the transform spans the block boundaries.
It was remarked that there are multiple possible approaches to the large block issue that can be studied.
It was agreed to further study this scheme in a CE.