Back to Search Document details
12th Meeting: Macao, October 2018 2018-10-06 05:26
CE14: Summary report on post-reconstruction filtering
Abstract
This contribution provides a summary report of Core Experiment 14 on post-reconstruction filtering methods. The techniques are evaluated according to BD-rate gain, complexity (in both encoder and decoder).
JVET-L0034 CE14: Summary report on post-reconstruction filtering [L. Zhang, S. Ikonin]

This contribution provides a summary report of Core Experiment 14 on post-reconstruction filtering methods. The techniques are evaluated according to BD-rate gain, complexity (in both encoder and decoder).

The software basis for the mandatory test in this CE is BMS-2.0.1, for VTM based comparisons the BMS software is configured to produce VTM-2.0.1. For the optional test, BMS-2.1 is used as the anchor with BMS tools enabled. Test sequences, configurations and test conditions are according to JVET-K1010 [1] for SDR.

Test#

Description

Document#

14.1.a

Bilateral filter turning off the filtering for 4x4 intra and inter blocks. No LUT – linear model

JVET-L0172

14.1.b

Bilateral filter turning off the filtering for 4x4, 4x8 and 8x4 intra/inter blocks. No LUT – linear model.

JVET-L0172

14.1.c

Bilateral filter with LUT rows consisting of more than 16
8-bit values (provided as reference). Turned off for 4x4.

JVET-L0172

14.2.a

In-loop bilateral filter (also operated after block reconstruction,
i.e. affecting subsequent intra prediction),

weights presented with piece-wise linear model (PWL), 10 ranges (pieces)

not applied to blocks with 4x4, and not applied to inter blocks with min(W, H)>8

JVET-L0406

14.2.b

In-loop bilateral filter (also operated after block reconstruction,
i.e. affecting subsequent intra prediction),

weights presented with PWL, 2 ranges (pieces)

not applied to blocks with 4x4, and not applied to inter blocks with min(W, H)>8

JVET-L0406

14.2.c*

In-loop bilateral filter (also operated after block reconstruction,
i.e. affecting subsequent intra prediction),

weights presented with PWL, 2 ranges (pieces)

not applied to blocks with 4x4, and not applied to inter blocks with min(W, H)>16, similar to tests 14.1 and 14.3

JVET-L0406

/JVET-L0584

14.3a

Hadamard Transform Domain Filter with LUT 140 bytes, not applied to 4x4 block, applied for intra and inter

JVET-L0326

14.3b

Hadamard Transform Domain Filter with LUT size 70 bytes, not applied to 4x4 block, applied for intra and inter

JVET-L0326

It was suggested that additional information be added to the CE summary which is a kind of table indicating with checkmarks for which cases which technology is applied in intra and inter.

Test

Filter shape

Comp. complex. per sample*

Precis. of mult

Parallel friendly

Latency

(in clock cycles)

Memory. required

(bytes)

How to derive filter coeffs

Min. and max. filtered

CU size

14.1.a

5 pixel “plus”-shape;

For inter, 5x5 area is used to calculate filter weights.

Intra:

4 mult
9 adds
4 checks

Inter:

4 mult
23 adds
10 checks

Intra:

9×8 and 12×9

Inter:

9×8 and 12×11

yes

At very high clock freq: Intra:10

Inter:

11

Estimation at lower clock freq: 3-4 clock cycles.

63

Intra:

Inter:

Min:

4x8, 8x4

Max:

Intra: 64x64

Inter: 16x64, 64x16

14.1.b

—”—

—”—

—”—

—”—

—”—

—”—

—”—

Min: 8x8

Max: same as

14.1.a

14.2.a,

LUT based

5 pixel “plus”-shape

Inter:
w(x) with NL average

Intra:
2 mult
8 adds
2 checks

Inter:
2 mult

18 ads
5 checks

32 bits registers

yes

Intra: 2

Inter: 3

ROM: 120

CU level: <370*2

Computed prior to CU:

Min:

4x8, 8x4

Max:

Intra: 64x64

Inter: 8x64, 64x8

14.2.b

LUT based

—”—

—”—

—”—

—”—

—”—

ROM: 120

CU level: <210*2

Computed prior to CU:

—”—

14.2.c

LUT based

—”—

—”—

—”—

—”—

—”—

—”—

—”—

Min:

4x8, 8x4

Max:

Intra: 64x64

Inter: 16x64, 64x16

14.2.a**

LUT Free

—”—

Intra:
4 mult
12 adds
13 checks

Inter:
4 mult
22 ads
16 checks

32 bits registers

yes

Intra: 4

Inter: 5

ROM: 120


CU level: 24

—”—

14.2.b**
LUT free

—”—

Intra:
4 mult
12 adds
4 checks

Inter:
4 mult
22 ads
7 checks

—”—

—”—

Intra: 3

Inter: 4

ROM: 120

CU level: 10

—”—

14.2.c **

LUT free

—”—

—”—

—”—

—”—

—”—

—”—

—”—

Min:

4x8, 8x4

Max:

Intra: 64x64

Inter: 16x64, 64x16

14.3.a

3x3

0 mult
20 adds + 4 1-bit add for rounding
6 checks

n/a

yes

1 clock:
@770MHz 16nm

@450MHz 28nm
2 clocks:

@770MHz 28nm

140

(32 7-bit values per qp group)

Precalculated in LUT

Min:

4x8 and 8x4

Max:

Intra: 64x64

Inter: 16x64 or 64x16

14.3.b

—”—

—”—

—”—

—”—

—”—

70

(16 7-bit values per qp group)

—”—

AI

RA

LB

Test#

Y

U

V

EncT

DecT

Y

U

V

EncT

DecT

Y

U

V

EncT

DecT

14.1.a

-0.39%

0.12%

0.16%

103%

105%

-0.67%

-0.10%

-0.14%

105%

102%

-0.64%

0.59%

0.39%

103%

103%

14.1.b

-0.30%

0.09%

0.11%

105%

103%

-0.58%

-0.10%

-0.10%

105%

102%

-0.64%

0.45%

0.51%

105%

101%

14.1.c *

-0.42%

0.14%

0.18%

105%

104%

-0.71%

-0.05%

-0.11%

104%

103%

-0.75%

0.30%

0.38%

103%

104%

14.2.a

-0.43%

0.20%

0.18%

113%

110%

-0.60%

-0.12%

-0.07%

87%**

83%**

-0.75%

0.2%

0.03%

103%

101%

14.2.b

-0.42%

0.13%

0.16%

114%

108%

-0.59%

-0.13%

-0.11%

93%**

103%

-0.75%

0.37%

0.03%

104%

102%

14.2.c

-0.42%

0.13%

0.16%

108%

109%*

-0.71%

-0.21%

-0.17%

106%*

110%*

-0.71%

0.48%

0.38%

92%**

103%

14.3.a

-0.48%

0.28%

0.31%

109%

110%

-0.70%

-0.14%

-0.19%

105%

104%

-0.68%

0.40%

0.67%

104%

104%

14.3.b

-0.47%

0.28%

0.32%

109%

110%

-0.70%

-0.17%

-0.06%

105%

104%

-0.66%

0.14%

0.31%

103%

104%

The most interesting technologies are 14.1a and 14.3b. These are directly competing technologies. Both have roughly same compression gain, and similar increase in encoding/decoding time. What might be of more concern (and might also lead to a decision adopting none of them) is the complexity added at a critical position in the decoding, which could cause latency issues in particular for intra coding. A detailed analysis on this shall be performed, documenting worst case number of operations, cycles, also including the possibility that a different LUT may need to be used for the Hadamard filter for each next CU if the QP is switched to a different range. SIMD complexity aspects should also be addressed.

BoG (L. Zhang) to further investigate, and also look into CE-related contributions.

The issue was raised that the post-reconstruction filters cause an additional complexity problem in requiring inverse transform for RD decision at the encoder. However, as the filters are not requiring low-level signalling, they could be disabled at high level without causing additional rate cost, such that any encoder could choose using them or not.

Decisions
The issue was raised that the post-reconstruction filters cause an additional complexity problem in requiring inverse transform for RD decision at the encoder. However, as the filters are not requiring low-level signalling, they could be disabled at high level without causing additional rate cost, such that any encoder could choose using them or not.
Citation