Back to Search Document details
13th Meeting: Marrakech, January 2019 2019-01-17 13:38
CE5 on arithmetic coding: experiments 5.1.1, 5.1.2, 5.1.3, 5.1.4, 5.1.5, 5.1.6, 5.1.7, 5.1.8, 5.1.10, 5.1.11, 5.1.12, 5.1.13, 5.2, and more
Abstract
This report provides results for the following CE5 experiments: 5.1.1, 5.1.2, 5.1.3, 5.1.4, 5.1.5, 5.1.6, 5.1.7, 5.1.8, 5.1.10, 5.1.11, 5.1.12, 5.1.13 and 5.2. Additional variants of these experiments are considered where constants, such as initial state values and shift amounts, are modified.
JVET-M0453 CE5 on arithmetic coding: experiments 5.1.1, 5.1.2, 5.1.3, 5.1.4, 5.1.5, 5.1.6, 5.1.7, 5.1.8, 5.1.10, 5.1.11, 5.1.12, 5.1.13, 5.2, and more [F. Bossen (Sharp)]

This report provides results for the following CE5 experiments: 5.1.1, 5.1.2, 5.1.3, 5.1.4, 5.1.5, 5.1.6, 5.1.7, 5.1.8, 5.1.10, 5.1.11, 5.1.12, 5.1.13 and 5.2. Additional variants of these experiments are considered where constants, such as initial state values and shift amounts, are modified.

Version 2 of the document provides additional results for experiment 5.2 on throughput, as well as BD-rate results for a configuration that yields no encoder run time increase.

The following table reports the decoder throughput from an actual bin stream (first 10 million context-coded bins from ParkRunning, RA configuration, QP22). Three compilers have been tested: clang 10.0 (Apple LLVM version 10.0.0), gcc 7.4 (Homebrew GCC 7.4.0) and gcc 8.2 (Homebrew GCC 8.2.0). The test bed was configured with the macro NOBRANCH_MPS either on (left half of table) or off (right half of table). Results were obtained on an Intel Xeon W-2140B CPU @ 3.20GHz.

In almost all cases the combination of clang and NOBRANCH_MPS enabled yielded the highest throughput. This configuration is thus considered for comparing the various engines.

The rightmost column in the table below reports the decoder throughput for a different bin stream (first 10 million context-coded bins from BQTerrace, RA configuration, QP22, clang, NOBRANCH_MPS enabled).

Engine

clang

gcc 7

gcc 8

clang

gcc 7

gcc 8

clang
stream #2

AVC/HEVC CABAC

132.03

134.92

128.64

127.27

121.13

124.98

137.88

CE5.1.1

120.17

117.48

118.58

113.54

108.52

109.26

121.51

CE5.1.2

115.00

110.93

113.28

104.01

108.47

102.20

114.67

CE5.1.3

106.63

104.98

104.60

95.89

96.84

96.37

102.79

CE5.1.4

109.09

106.68

107.28

104.77

104.45

104.77

110.55

CE5.1.4 mult

134.85

133.09

131.12

120.09

116.32

117.72

135.13

CE5.1.5

113.46

106.75

107.56

111.67

102.44

104.75

113.02

CE5.1.6 (8+12 bit)

110.75

105.07

103.62

103.64

98.89

100.07

108.50

CE5.1.7

119.00

110.98

112.45

118.51

110.15

107.84

120.23

CE5.1.8

118.80

114.27

107.99

107.71

103.96

100.47

119.39

CE5.1.9 / CE 5.1.10

125.05

122.68

120.02

120.47

113.33

115.96

127.29

CE5.1.11

121.79

117.33

117.14

107.25

113.08

111.22

123.35

CE5.1.12

128.95

123.80

127.78

118.60

119.66

120.22

129.01

CE5.1.13 / CE5.1.3 mult

127.28

121.00

122.57

109.90

109.25

110.67

127.32

Experiments were run a second time, where each test is done 7 times and the highest throughput value is recorded. While the numbers are slightly higher, the same trends persist.

Engine

clang

gcc 7

gcc 8

clang

gcc 7

gcc 8

clang
stream #2

AVC/HEVC CABAC

137.27

137.83

131.01

131.69

124.68

127.93

141.37

CE5.1.1

124.17

121.64

120.26

117.41

112.53

110.42

124.48

CE5.1.2

116.69

115.36

116.85

106.22

110.39

109.24

115.93

CE5.1.3

107.96

106.13

105.99

96.69

98.59

97.68

108.09

CE5.1.4

111.83

109.14

109.93

106.80

106.21

106.51

112.09

CE5.1.4 mult

136.30

133.96

132.36

120.85

116.32

121.01

137.27

CE5.1.5

115.84

109.11

108.83

114.71

105.29

105.21

115.73

CE5.1.6 (8+12 bit)

112.69

106.88

103.99

105.00

102.27

102.21

112.68

CE5.1.7

120.75

112.94

113.24

119.12

106.58

114.97

120.89

CE5.1.8

122.01

116.04

108.10

110.09

107.50

102.54

125.28

CE5.1.9 / CE 5.1.10

128.67

124.20

121.22

122.22

119.20

117.55

129.44

CE5.1.11

124.55

118.73

119.37

114.53

114.48

112.39

125.28

CE5.1.12

130.26

128.17

130.77

121.13

121.84

120.45

131.45

CE5.1.13 / CE5.1.3 mult

128.51

122.63

126.85

112.10

111.23

112.96

128.87

It was agreed that the second table of throughput numbers had much lower probability of having outlier values in it.

The table below summarizes the performance of various arithmetic coding engines, classifying them in three groups:

  • Two states, fixed window sizes
  • Two states, variable (context-adaptive) window sizes
  • Single state, variable (context-adaptive) window sizes

BD rate numbers are reported as combined YUV, where BD rate numbers for Y, U and V are weighted as 90%/5%/5% to obtain a single number. This is done to account for the fact that results for individual color components can sometimes diverge.

Also provided for each engine are encoder/decoder run times relative to VTM3, as well as SW throughput results from experiment 5.2.

Within each group and each test condition the best BD rate and throughput results are highlighted in green. The worst results are highlighted in red.

Experiment numbers with a star indicate a variant, where some constants are modified as described.

Engine

AI

RA

LB

LP

AI

RA

LB

Thruput

2 states, fixed window size

5.1.1

-0.85%

-0.74%

-0.56%

-0.56%

106/103

103/102

101/100

120.17

5.1.5

-0.83%

-0.64%

-0.62%

-0.60%

109/107

105/103

103/103

113.46

5.1.8

-0.91%

-0.69%

-0.48%

-0.50%

106/102

103/101

100/98

118.80

5.1.10

-0.83%

-0.70%

-0.55%

-0.61%

108/104

104/102

103/101

125.05

5.1.10* + new init from 5.1.2

-0.84%

-0.73%

-0.51%

-0.53%

109/107

105/104

105/104

125.05

5.1.10* 5x6 mult

-0.85%

-0.73%

-0.58%

-0.64%

109/107

105/104

104/104

127.86

2 states, variable window size

5.1.2

-1.00%

-0.97%

-0.82%

-0.76%

109/105

104/103

101/101

115.00

5.1.3 clz + 8x8 table

-1.03%

-0.96%

-0.74%

-0.83%

107/108

104/104

104/104

106.63

5.1.3 4x5 mult

-0.99%

-0.91%

-0.71%

-0.83%

109/105

104/102

103/102

127.28

5.1.3* 5x6 mult

-1.07%

-0.99%

-0.79%

-0.91%

108/104

104/103

104/103

127.28

5.1.6 8+12 bit state

-0.90%

-0.76%

-0.72%

-0.69%

111/105

105/103

105/102

110.75

5.1.6 10+14 bit state

-0.97%

-0.82%

-0.67%

-0.73%

112/105

105/102

105/102

110.75

5.1.11

-0.97%

-0.93%

-0.75%

-0.76%

108/103

103/101

102/101

121.79

5.1.11* + new init from 5.1.2

-1.00%

-0.98%

-0.77%

-0.75%

110/105

104/102

104/102

121.79

5.1.13

-0.95%

-0.91%

-0.73%

-0.74%

109/103

104/101

101/98

127.28

5.1.13* +new init from 5.1.2

-0.98%

-0.96%

-0.75%

-0.73%

110/105

104/103

106/103

127.28

5.1.13* 5x6 mult

-1.03%

-0.98%

-0.75%

-0.75%

109/105

104/102

104/104

127.28

1 state, variable window size

5.1.4 clz + 8x8 table

-0.92%

-0.65%

-0.41%

-0.44%

108/107

103/104

104/103

109.09

5.1.4 4x5 mult

-0.83%

-0.58%

-0.37%

-0.41%

108/105

104/102

102/102

132.95

5.1.4* 5x6 mult

-0.93%

-0.67%

-0.46%

-0.49%

107/105

104/102

104/103

132.95

5.1.7

-0.70%

-0.57%

-0.53%

-0.52%

107/107

104/103

103/104

119.00

5.1.12

-0.65%

-0.44%

-0.33%

-0.25%

108/103

102/100

101/99

128.95

The table above was copied from JVET-M0453 and then updated in group discussion to form the modified table below. The each category in the table was reviewed and particular tests that provided the highest coding efficiency in each category with good throughput numbers were highlighted for further focused discussion. The properties, similarities and differences between the tested methods were discussed.

The following table was obtained by updating the table above as follows: 1) adding 5.1.9, 2) replacing throughput numbers with the second throughput number table, and 3) keeping only rows that pertain to the CE tests.

Engine

AI

RA

LB

LP

AI

RA

LB

Thruput

VTM3 w/ HEVC engine

0

0

0

0

100/100

100/100

100/100

137.83

2 states, fixed window size

5.1.1

-0.85%

-0.74%

-0.56%

-0.56%

106/103

103/102

101/100

124.17

5.1.5

-0.83%

-0.64%

-0.62%

-0.60%

109/107

105/103

103/103

115.84

5.1.8

-0.91%

-0.69%

-0.48%

-0.50%

106/102

103/101

100/98

122.01

5.1.9

-0.79%

-0.57%

-0.50%

N/A

128.67

5.1.10

-0.83%

-0.70%

-0.55%

-0.61%

108/104

104/102

103/101

128.67

5.1.10* + new init from 5.1.2

-0.84%

-0.73%

-0.51%

-0.53%

109/107

105/104

105/104

128.67

2 states, variable window size

5.1.2

-1.00%

-0.97%

-0.82%

-0.76%

109/105

104/103

101/101

116.69

5.1.3 clz + 8x8 table (config. 2)

-1.03%

-0.96%

-0.74%

-0.83%

107/108

104/104

104/104

107.96

5.1.3 4x5 mult (config. 1)

-0.99%

-0.91%

-0.71%

-0.83%

109/105

104/102

103/102

128.51

5.1.6 8+12 bit state

-0.90%

-0.76%

-0.72%

-0.69%

111/105

105/103

105/102

112.69

5.1.6 10+14 bit state

-0.97%

-0.82%

-0.67%

-0.73%

112/105

105/102

105/102

112.69*

5.1.11

-0.97%

-0.93%

-0.75%

-0.76%

108/103

103/101

102/101

124.55

5.1.11* + new init from 5.1.2

-1.00%

-0.98%

-0.77%

-0.75%

110/105

104/102

104/102

124.55

5.1.13

-0.95%

-0.91%

-0.73%

-0.74%

109/103

104/101

101/98

128.51

5.1.13* +new init from 5.1.2

-0.98%

-0.96%

-0.75%

-0.73%

110/105

104/103

106/103

128.51

1 state, variable window size

5.1.4 clz + 8x8 table (config. 2)

-0.92%

-0.65%

-0.41%

-0.44%

108/107

103/104

104/103

111.83

5.1.4 4x5 mult (config. 1)

-0.83%

-0.58%

-0.37%

-0.41%

108/105

104/102

102/102

136.30

5.1.7

-0.70%

-0.57%

-0.53%

-0.52%

107/107

104/103

103/104

120.75

5.1.12

-0.65%

-0.44%

-0.33%

-0.25%

108/103

102/100

101/99

130.77

It was suggested to look at the hardware aspect of these different engines, which was provided as part of JVET-M0025 subtest 3.

It was agreed that no hardware problem had been identified for any of the CE tests, and that such a hardware problem could be fixed if/when it was identified.

When looking at different categories, it was remarked that the “1 state, variable window” category has the highest throughput (closest to HEVC engine), and the “2 state, variable window” category has the highest coding performance. Regarding the “2 state, fixed window” category, it was remarked that this category does not need custom window size parameters for different context models, but needs re-training of initialization parameters.

Regarding custom window size parameters, it was remarked that that does not seem to increase complexity to such a degree that it justifies going for a simpler solution (i.e. fixed window size).

It was remarked that some of the coding efficiency gain from the tests that use custom window size parameters may have come from training of custom window size parameters based on the test set.

It was remarked that the initialization parameters were also trained on the test set, which also may have provided some of the coding efficiency gain.

In terms of coding efficiency, the options “5.1.3 4x5 mult (config. 1)”, “5.1.13” and “5.1.13* +new init from 5.1.2” are the most attractive options. Between “5.1.3” and “5.1.13,” the latter has a slight advantage for hardware implementations due to needing to support fewer shift values. And the difference between “5.1.13” and “5.1.13* +new init from 5.1.2” is purely due to training of initialization parameters, with the new initialization parameters providing better coding efficiency.

Decision: Adopt “5.1.13* +new init from 5.1.2”.

References:
JVET-L0335
Decisions
adopted
Adopt “5.1.13* +new init from 5.1.2”
Citation