JVET-M0453 CE5 on arithmetic coding: experiments 5.1.1, 5.1.2, 5.1.3, 5.1.4, 5.1.5, 5.1.6, 5.1.7, 5.1.8, 5.1.10, 5.1.11, 5.1.12, 5.1.13, 5.2, and more [F. Bossen (Sharp)]
This report provides results for the following CE5 experiments: 5.1.1, 5.1.2, 5.1.3, 5.1.4, 5.1.5, 5.1.6, 5.1.7, 5.1.8, 5.1.10, 5.1.11, 5.1.12, 5.1.13 and 5.2. Additional variants of these experiments are considered where constants, such as initial state values and shift amounts, are modified.
Version 2 of the document provides additional results for experiment 5.2 on throughput, as well as BD-rate results for a configuration that yields no encoder run time increase.
The following table reports the decoder throughput from an actual bin stream (first 10 million context-coded bins from ParkRunning, RA configuration, QP22). Three compilers have been tested: clang 10.0 (Apple LLVM version 10.0.0), gcc 7.4 (Homebrew GCC 7.4.0) and gcc 8.2 (Homebrew GCC 8.2.0). The test bed was configured with the macro NOBRANCH_MPS either on (left half of table) or off (right half of table). Results were obtained on an Intel Xeon W-2140B CPU @ 3.20GHz.
In almost all cases the combination of clang and NOBRANCH_MPS enabled yielded the highest throughput. This configuration is thus considered for comparing the various engines.
The rightmost column in the table below reports the decoder throughput for a different bin stream (first 10 million context-coded bins from BQTerrace, RA configuration, QP22, clang, NOBRANCH_MPS enabled).
Engine | clang | gcc 7 | gcc 8 | clang | gcc 7 | gcc 8 | clang |
AVC/HEVC CABAC | 132.03 | 134.92 | 128.64 | 127.27 | 121.13 | 124.98 | 137.88 |
CE5.1.1 | 120.17 | 117.48 | 118.58 | 113.54 | 108.52 | 109.26 | 121.51 |
CE5.1.2 | 115.00 | 110.93 | 113.28 | 104.01 | 108.47 | 102.20 | 114.67 |
CE5.1.3 | 106.63 | 104.98 | 104.60 | 95.89 | 96.84 | 96.37 | 102.79 |
CE5.1.4 | 109.09 | 106.68 | 107.28 | 104.77 | 104.45 | 104.77 | 110.55 |
CE5.1.4 mult | 134.85 | 133.09 | 131.12 | 120.09 | 116.32 | 117.72 | 135.13 |
CE5.1.5 | 113.46 | 106.75 | 107.56 | 111.67 | 102.44 | 104.75 | 113.02 |
CE5.1.6 (8+12 bit) | 110.75 | 105.07 | 103.62 | 103.64 | 98.89 | 100.07 | 108.50 |
CE5.1.7 | 119.00 | 110.98 | 112.45 | 118.51 | 110.15 | 107.84 | 120.23 |
CE5.1.8 | 118.80 | 114.27 | 107.99 | 107.71 | 103.96 | 100.47 | 119.39 |
CE5.1.9 / CE 5.1.10 | 125.05 | 122.68 | 120.02 | 120.47 | 113.33 | 115.96 | 127.29 |
CE5.1.11 | 121.79 | 117.33 | 117.14 | 107.25 | 113.08 | 111.22 | 123.35 |
CE5.1.12 | 128.95 | 123.80 | 127.78 | 118.60 | 119.66 | 120.22 | 129.01 |
CE5.1.13 / CE5.1.3 mult | 127.28 | 121.00 | 122.57 | 109.90 | 109.25 | 110.67 | 127.32 |
Experiments were run a second time, where each test is done 7 times and the highest throughput value is recorded. While the numbers are slightly higher, the same trends persist.
Engine | clang | gcc 7 | gcc 8 | clang | gcc 7 | gcc 8 | clang |
AVC/HEVC CABAC | 137.27 | 137.83 | 131.01 | 131.69 | 124.68 | 127.93 | 141.37 |
CE5.1.1 | 124.17 | 121.64 | 120.26 | 117.41 | 112.53 | 110.42 | 124.48 |
CE5.1.2 | 116.69 | 115.36 | 116.85 | 106.22 | 110.39 | 109.24 | 115.93 |
CE5.1.3 | 107.96 | 106.13 | 105.99 | 96.69 | 98.59 | 97.68 | 108.09 |
CE5.1.4 | 111.83 | 109.14 | 109.93 | 106.80 | 106.21 | 106.51 | 112.09 |
CE5.1.4 mult | 136.30 | 133.96 | 132.36 | 120.85 | 116.32 | 121.01 | 137.27 |
CE5.1.5 | 115.84 | 109.11 | 108.83 | 114.71 | 105.29 | 105.21 | 115.73 |
CE5.1.6 (8+12 bit) | 112.69 | 106.88 | 103.99 | 105.00 | 102.27 | 102.21 | 112.68 |
CE5.1.7 | 120.75 | 112.94 | 113.24 | 119.12 | 106.58 | 114.97 | 120.89 |
CE5.1.8 | 122.01 | 116.04 | 108.10 | 110.09 | 107.50 | 102.54 | 125.28 |
CE5.1.9 / CE 5.1.10 | 128.67 | 124.20 | 121.22 | 122.22 | 119.20 | 117.55 | 129.44 |
CE5.1.11 | 124.55 | 118.73 | 119.37 | 114.53 | 114.48 | 112.39 | 125.28 |
CE5.1.12 | 130.26 | 128.17 | 130.77 | 121.13 | 121.84 | 120.45 | 131.45 |
CE5.1.13 / CE5.1.3 mult | 128.51 | 122.63 | 126.85 | 112.10 | 111.23 | 112.96 | 128.87 |
It was agreed that the second table of throughput numbers had much lower probability of having outlier values in it.
The table below summarizes the performance of various arithmetic coding engines, classifying them in three groups:
- Two states, fixed window sizes
- Two states, variable (context-adaptive) window sizes
- Single state, variable (context-adaptive) window sizes
BD rate numbers are reported as combined YUV, where BD rate numbers for Y, U and V are weighted as 90%/5%/5% to obtain a single number. This is done to account for the fact that results for individual color components can sometimes diverge.
Also provided for each engine are encoder/decoder run times relative to VTM3, as well as SW throughput results from experiment 5.2.
Within each group and each test condition the best BD rate and throughput results are highlighted in green. The worst results are highlighted in red.
Experiment numbers with a star indicate a variant, where some constants are modified as described.
Engine | AI | RA | LB | LP | AI | RA | LB | Thruput |
2 states, fixed window size | ||||||||
5.1.1 | -0.85% | -0.74% | -0.56% | -0.56% | 106/103 | 103/102 | 101/100 | 120.17 |
5.1.5 | -0.83% | -0.64% | -0.62% | -0.60% | 109/107 | 105/103 | 103/103 | 113.46 |
5.1.8 | -0.91% | -0.69% | -0.48% | -0.50% | 106/102 | 103/101 | 100/98 | 118.80 |
5.1.10 | -0.83% | -0.70% | -0.55% | -0.61% | 108/104 | 104/102 | 103/101 | 125.05 |
5.1.10* + new init from 5.1.2 | -0.84% | -0.73% | -0.51% | -0.53% | 109/107 | 105/104 | 105/104 | 125.05 |
5.1.10* 5x6 mult | -0.85% | -0.73% | -0.58% | -0.64% | 109/107 | 105/104 | 104/104 | 127.86 |
2 states, variable window size | ||||||||
5.1.2 | -1.00% | -0.97% | -0.82% | -0.76% | 109/105 | 104/103 | 101/101 | 115.00 |
5.1.3 clz + 8x8 table | -1.03% | -0.96% | -0.74% | -0.83% | 107/108 | 104/104 | 104/104 | 106.63 |
5.1.3 4x5 mult | -0.99% | -0.91% | -0.71% | -0.83% | 109/105 | 104/102 | 103/102 | 127.28 |
5.1.3* 5x6 mult | -1.07% | -0.99% | -0.79% | -0.91% | 108/104 | 104/103 | 104/103 | 127.28 |
5.1.6 8+12 bit state | -0.90% | -0.76% | -0.72% | -0.69% | 111/105 | 105/103 | 105/102 | 110.75 |
5.1.6 10+14 bit state | -0.97% | -0.82% | -0.67% | -0.73% | 112/105 | 105/102 | 105/102 | 110.75 |
5.1.11 | -0.97% | -0.93% | -0.75% | -0.76% | 108/103 | 103/101 | 102/101 | 121.79 |
5.1.11* + new init from 5.1.2 | -1.00% | -0.98% | -0.77% | -0.75% | 110/105 | 104/102 | 104/102 | 121.79 |
5.1.13 | -0.95% | -0.91% | -0.73% | -0.74% | 109/103 | 104/101 | 101/98 | 127.28 |
5.1.13* +new init from 5.1.2 | -0.98% | -0.96% | -0.75% | -0.73% | 110/105 | 104/103 | 106/103 | 127.28 |
5.1.13* 5x6 mult | -1.03% | -0.98% | -0.75% | -0.75% | 109/105 | 104/102 | 104/104 | 127.28 |
1 state, variable window size | ||||||||
5.1.4 clz + 8x8 table | -0.92% | -0.65% | -0.41% | -0.44% | 108/107 | 103/104 | 104/103 | 109.09 |
5.1.4 4x5 mult | -0.83% | -0.58% | -0.37% | -0.41% | 108/105 | 104/102 | 102/102 | 132.95 |
5.1.4* 5x6 mult | -0.93% | -0.67% | -0.46% | -0.49% | 107/105 | 104/102 | 104/103 | 132.95 |
5.1.7 | -0.70% | -0.57% | -0.53% | -0.52% | 107/107 | 104/103 | 103/104 | 119.00 |
5.1.12 | -0.65% | -0.44% | -0.33% | -0.25% | 108/103 | 102/100 | 101/99 | 128.95 |
The table above was copied from JVET-M0453 and then updated in group discussion to form the modified table below. The each category in the table was reviewed and particular tests that provided the highest coding efficiency in each category with good throughput numbers were highlighted for further focused discussion. The properties, similarities and differences between the tested methods were discussed.
The following table was obtained by updating the table above as follows: 1) adding 5.1.9, 2) replacing throughput numbers with the second throughput number table, and 3) keeping only rows that pertain to the CE tests.
Engine | AI | RA | LB | LP | AI | RA | LB | Thruput |
VTM3 w/ HEVC engine | 0 | 0 | 0 | 0 | 100/100 | 100/100 | 100/100 | 137.83 |
2 states, fixed window size | ||||||||
5.1.1 | -0.85% | -0.74% | -0.56% | -0.56% | 106/103 | 103/102 | 101/100 | 124.17 |
5.1.5 | -0.83% | -0.64% | -0.62% | -0.60% | 109/107 | 105/103 | 103/103 | 115.84 |
5.1.8 | -0.91% | -0.69% | -0.48% | -0.50% | 106/102 | 103/101 | 100/98 | 122.01 |
5.1.9 | -0.79% | -0.57% | -0.50% | N/A | 128.67 | |||
5.1.10 | -0.83% | -0.70% | -0.55% | -0.61% | 108/104 | 104/102 | 103/101 | 128.67 |
5.1.10* + new init from 5.1.2 | -0.84% | -0.73% | -0.51% | -0.53% | 109/107 | 105/104 | 105/104 | 128.67 |
2 states, variable window size | ||||||||
5.1.2 | -1.00% | -0.97% | -0.82% | -0.76% | 109/105 | 104/103 | 101/101 | 116.69 |
5.1.3 clz + 8x8 table (config. 2) | -1.03% | -0.96% | -0.74% | -0.83% | 107/108 | 104/104 | 104/104 | 107.96 |
5.1.3 4x5 mult (config. 1) | -0.99% | -0.91% | -0.71% | -0.83% | 109/105 | 104/102 | 103/102 | 128.51 |
5.1.6 8+12 bit state | -0.90% | -0.76% | -0.72% | -0.69% | 111/105 | 105/103 | 105/102 | 112.69 |
5.1.6 10+14 bit state | -0.97% | -0.82% | -0.67% | -0.73% | 112/105 | 105/102 | 105/102 | 112.69* |
5.1.11 | -0.97% | -0.93% | -0.75% | -0.76% | 108/103 | 103/101 | 102/101 | 124.55 |
5.1.11* + new init from 5.1.2 | -1.00% | -0.98% | -0.77% | -0.75% | 110/105 | 104/102 | 104/102 | 124.55 |
5.1.13 | -0.95% | -0.91% | -0.73% | -0.74% | 109/103 | 104/101 | 101/98 | 128.51 |
5.1.13* +new init from 5.1.2 | -0.98% | -0.96% | -0.75% | -0.73% | 110/105 | 104/103 | 106/103 | 128.51 |
1 state, variable window size | ||||||||
5.1.4 clz + 8x8 table (config. 2) | -0.92% | -0.65% | -0.41% | -0.44% | 108/107 | 103/104 | 104/103 | 111.83 |
5.1.4 4x5 mult (config. 1) | -0.83% | -0.58% | -0.37% | -0.41% | 108/105 | 104/102 | 102/102 | 136.30 |
5.1.7 | -0.70% | -0.57% | -0.53% | -0.52% | 107/107 | 104/103 | 103/104 | 120.75 |
5.1.12 | -0.65% | -0.44% | -0.33% | -0.25% | 108/103 | 102/100 | 101/99 | 130.77 |
It was suggested to look at the hardware aspect of these different engines, which was provided as part of JVET-M0025 subtest 3.
It was agreed that no hardware problem had been identified for any of the CE tests, and that such a hardware problem could be fixed if/when it was identified.
When looking at different categories, it was remarked that the “1 state, variable window” category has the highest throughput (closest to HEVC engine), and the “2 state, variable window” category has the highest coding performance. Regarding the “2 state, fixed window” category, it was remarked that this category does not need custom window size parameters for different context models, but needs re-training of initialization parameters.
Regarding custom window size parameters, it was remarked that that does not seem to increase complexity to such a degree that it justifies going for a simpler solution (i.e. fixed window size).
It was remarked that some of the coding efficiency gain from the tests that use custom window size parameters may have come from training of custom window size parameters based on the test set.
It was remarked that the initialization parameters were also trained on the test set, which also may have provided some of the coding efficiency gain.
In terms of coding efficiency, the options “5.1.3 4x5 mult (config. 1)”, “5.1.13” and “5.1.13* +new init from 5.1.2” are the most attractive options. Between “5.1.3” and “5.1.13,” the latter has a slight advantage for hardware implementations due to needing to support fewer shift values. And the difference between “5.1.13” and “5.1.13* +new init from 5.1.2” is purely due to training of initialization parameters, with the new initialization parameters providing better coding efficiency.
Decision: Adopt “5.1.13* +new init from 5.1.2”.