Back to Search Document details
14th Meeting: Geneva, March 2019 2019-03-22 14:52
CE11: Summary Report on Deblocking
Abstract
This contribution provides a summary report of Core Experiment 11 on deblocking filtering. Two categories of proposals are covered by this CE, split into two sub-tests. These sub-tests are 1) long-tap deblocking filters, and 2) deblocking at 4x4 block boundaries.
JVET-N0031 CE11: Summary Report on Deblocking [A. Norkin, A. M. Kotra]

Test#

Description

Document#

11-1.1

Very strong deblocking filtering with conditional activation signalling

JVET-N0595

11-2.1

Deblocking for 4xN, Nx4 and 8xN and Nx8 block boundaries not aligned with 8x8 grids

JVET-N0098

11-2.2

Disable deblocking filter for 4xN for vertical edge and Nx4 for Horizontal edge on 4x4 grid

JVET-N0463

11-2.3

Disabling of sub-pu deblocking

JVET-N0181

CE11-1.1 increases worst-case complexity in case of intra slice:

Luma:

 Tests

 

Samples from block bound. modified

Samples from block bound. deblocking decision

Max num. oper for filtering per line (add/mult/compar/shift)

Max number of oper. for decision for 8-sample boundary (add/mult/compar/shift)

Max number of operations including filtering and decision (Per *Sample*)

(add/mult/compar/shift)

Num. line buffers

Worst case complexity increased (Y/N)

VTM4.0

7+7, 7+5, 5+7, 5+5, 7+3, 3+7, 5+3, 3+5

8+8, 8+6, 6+8, 6+6, 8+4, 4+8, 6+4, 4+6

120 (46,24,28,22)

21 (43,0,24,15)/4

4

N

CE11-1.1

VTM (Inter), 3+3...16+16 (Intra)

VTM

Intra slice: 288 (96/96/64/32), Inter
slice: VTM

VTM

VTM

Intra slice: Y

Inter slice: N

Chroma:

 Tests

 

Samples from block boundary modified

Samples from block boundary for deblocking decision

Max number of operations for filtering per line (add/mult/compar/shift)

Max number of oper. for decision for 8-sample boundary (add/mult/compar/shift)

Max number of operations including filtering and decision (Per *Sample*)

(add/mult/compar/shift)

Number of line buffers

Worst case complexity increased (Y/N)

VTM4.0

3+3

4+4

64 (36,2,12,14)

17 (17, 0, 12, 5)/2 per line

2

Y

CE11-1.1

VTM (Inter), 1+1...8+8 (Intra)

VTM

Intra slice: 144 (48/48/32/16), Inter slice: VTM

VTM

VTM

Intra slice: Y

Inter slice: N

In terms of objective criteria, no benefit.

Some discussion was performed on the approach: It is selectively applied (signalled) on a TU basis, where only the side of the current TU is strong deblocked. The implementation is such that it requires two passes (first strong deblocking is applied, and in a second pass the conventional deblocking is applied on the remaining blocks). For a real implementation, it should only be one pass (which would however have other results)

Viewing results from document JVET-N0835

Results from subjective testing on CE11-1, ALF on:

The subjective viewing did not show any differences for QP34, for QP39 the results were diverging (sometimes better, sometimes worse).

It is claimed by proponents that they believe it has benefit for Sunset Beach (HLG), but it is mentioned by other experts that other tools like the luma adaptive deblocking (not in CTC but in standard) likely helps there as well.

No action.

CE11-2: Deblocking 4x4 boundaries

Luma complexity:

 Tests

 

Samples from block bound. Modified

Samples from block bound. for deblocking decision

Max num. oper for filtering per line (add/mult/compar/shift)

Max number of oper. for decision per line (add/mult/compar/shift)

Max number of operations including filtering and decision (per *sample*)

(add/mult/compar/shift)

Num. line buffers

Worst case complexity increased (Y/N)

CE11-2.1

All 4x4: 1+1

All 4x4: 3+3

All 4x4: 14 (6/2/5/1)

VTM (all 4x4 : 1.25 (11/0/5/4) /4 )

All 4x4: 4.75 (2.1875/0.5/1.5625/0.5)

4

N

CE11-2.2

All 4x4: 0

4xN : 3+3

Other: VTM

All 4x4 : 0

4xN : 3+3

Other: VTM

All 4x4 : 0

4xN : VTM

Other: VTM

All 4x4 : 0

4xN : VTM

Other: VTM

All 4x4 : 0

4xN : VTM

Other : VTM

4

N

CE11-2.3

Same as VTM

Same as VTM

Same as VTM

Same as VTM

Same as VTM

4

N

Chroma complexity:

 Tests

 

Samples from block bound. modified

Samples from block bound. for deblocking decision

Max num. oper for filtering per line (add/mult/compar/shift)

Max number of oper. for decision for 8-sample boundary (add/mult/compar/shift)

Max number of operations including filtering and decision (per *sample*)

(add/mult/compar/shift)

Num. line buffers

Worst case complexity increased (Y/N)

CE11-2.1

VTM

VTM

VTM

VTM

VTM

VTM

N

CE11-2.2

VTM

VTM

VTM

VTM

VTM

VTM

N

CE11-2.3

VTM

VTM

VTM

VTM

VTM

VTM

N

Even though in terms of computation operations no increase over VTM occurs, the worst case number of edges is doubling in CE11-2.1/2, and this may be some concern in hardware implementation. In CE11-2.2, the number of edges is doubling in both luma and chroma.

Objective results with ALF on

Test

AI

RA

Y

U

V

EncT

DecT

Y

U

V

EncT

DecT

CE11-2.1

0.05%

0.00%

0.00%

90%

102%

0.03%

-0.06%

-0.06%

102%

102%

CE11-2.2

-0.05%

0.22%

0.14%

99%

98%

-0.02%

0.28%

0.19%

100%

98%

CE11-2.3

0.00%

0.00%

0.00%

100%

100%

0.04%

-0.05%

0.00%

100%

98%

Test

LD-B

LD-P

Y

U

V

EncT

DecT

Y

U

V

EncT

DecT

CE11-2.1

-0.05%

0.04%

0.17%

95%

102%

-0.15%

-0.26%

-0.20%

99%

101%

CE11-2.2

-0.04%

-0.06%

-0.39%

100%

99%

0.00%

-0.26%

-0.39%

99%

98%

CE11-2.3

0.10%

0.08%

-0.21%

100%

97%

0.06%

-0.05%

-0.04%

100%

98%

Objective results with ALF off

Test

AI

RA

Y

U

V

EncT

DecT

Y

U

V

EncT

DecT

CE11-2.1

-0.03%

0.00%

0.00%

100%

102%

-0.12%

-0.08%

-0.06%

100%

101%

CE11-2.2

-0.01%

-0.27%

-0.47%

100%

98%

0.00%

0.60%

0.47%

100%

99%

CE11-2.3

0.00%

0.00%

0.00%

100%

100%

0.06% 

 -0.01%

0.00% 

 100%

100% 

Test

LD-B

LD-P

Y

U

V

EncT

DecT

Y

U

V

EncT

DecT

CE11-2.1

-0.08%

-0.02%

-0.01%

100%

103%

-0.53%

-0.24%

-0.18%

100%

102%

CE11-2.2

-0.01%

-0.17%

-0.31%

100%

99%

0.02%

-0.17%

-0.31%

100%

99%

CE11-2.3

0.08% 

 -0.11%

-0.13% 

100% 

100% 

0.13% 

 -0.04%

 -0.05%

 100%

100% 

It is pointed out that CE11-2.1 deblocks all subblock transform boundaries, whereas VTM does not deblock even 8x8 subblock transform boundaries, but there is a bug in the implementation such that it is not performed.

Results of subjective testing for CE11-2, ALF on:

Visual tests with ALF on are not fully conclusive. Sometimes CE11-2.3 (which is even doing less deblocking than VTM) behaves better, sometimes one of the other two. There is only one clear case of CE11-2.1 (and a few corner cases of others as well) where a visual benefit over VTM4 could be shown with statistical significance.

It was therefore decided to run another round of subjective tests with ALF off on CE11-2 to see if a more clear tendency can be determined.

From the results above, it becomes more evident that proposals whch perform deblocking at more boundaries (CE11-2.1 and CE11-2.2) have visual advantage over the proposal with less deblocking (CE11-2.3The proposal CE11-2.2 has two out of four cases with better visual quality than VTM4 at QP34, which would still be a reasonable quality in the normal application range. CE11-2.1 ha another case in ALF on configuration where it performed better than VTM4 at QP34.

It is further noted that CE11-2.2 does not deblock SBT and ISP boundaries (which VTM4 also did not do, because it was a kind of bug from the last meeting), whereas CW11-2.1 already corrected this bug, and some aspects of its quality may be due to this.

It is confirmed by proponents of CE11-2.2 that the proposal has full capability of parallel processing, as a 4x4 block would not be deblocked, if the next block is again at a distance 4.

Another aspect is that CE11-2.2 is also deblocking the chroma blocks aligned with the modified boundaries of the luma blocks. CE11-2.1 does not do that and leaves the deblocking of the chroma on a fixed 8x8 grid. It seems to be a reasonable design choice to deblock the chroma aligned with the luma.

It was initially suggested to adopt the JVET-N0098 method for deblocking of 4x4 boundaries, including the bug fix on ISP and SBT, in combination with the aspect from JVET-N0463 of deblocking chroma at boundaries that are aligned with luma.

In a follow-up discussion, it was identified that a straightforward combination of the two methods is not possible, as JVET-N0098 in extreme case would deblock all 4x4 block boundaries (with shorter filter), and the corresponding chroma boundaries would be 2x2. In terms of worst case number of boundaries, JVET-N0098 doubles for luma and keeps chroma unchanged relative to VTM4. JVET-N0463 keeps the worst case number for luma unchanged, but doubles for chroma. Both approaches increase the number, but effectively for the case of 4:2:0 the increase is larger for JVET-N0098, as luma has more samples than chroma. (As suggested in some other discussion by hardware experts, the number of boundaries is of larger concern than the multiplications necessary for length of the filters, as decisions have to determined at each boundary). Further study is necessary to identify whether there is benefit to actually deblock small blocks (JVET-N0463 does not deblock any blocks at their 4xN or Nx4 boundaries, i.e. only the longer side of such blocks is deblocked. It however deblocks larger blocks where the boundary is not aligned with an 8x8 grid). The CE should be continued to better identify

- How much benefit comes due to the 4xN and Nx4 deblocking of JVET-N0098

- How much benefit comes from the additional chroma deblocking of JVET-N0463

Further investigation should also be performed on the aspect if 4:4:4 material should be deblocked differently from 4:2:0, as this would also have impact to decide which of the wo methods is more complex (1080p is sufficient for that).

The additional aspect from JVET-N0098 of adapting the MV threshold should also be independently studied in CE.

Decisions
The additional aspect from JVET-N0098 of adapting the MV threshold should also be independently studied in CE.
Citation