JVET-T0069 AHG11: SSIM based CNN model for in-loop filtering [T. Ouyang, F. Liu, H. Zhu, Z. Chen (Wuhan Univ.), X. Xu, S. Liu (Tencent)]
This contribution was discussed in session 9 at 2100 on Friday 9 October (chaired by GJS & JRO).
This contribution provides a CNNLF (convolutional neural network based in-loop filter) for VVC which focus on improving the subjective quality of encoded videos. The performance is evaluated using VMAF in the VTM-10.0. Simulation results report BD-rate savings for luma, and both chroma components compared with VTM-10.0 under AI, RA and LDB configuration.
The proposed CNNLF is introduced into VTM as an additional filter between deblocking filter (DF) and sample adaptive offset (SAO).
Simulation results report 7.50%, 6.01%, 2.30% BD-rate savings on VMAF under AI, RA and LDB configuration.
The first set of tables below use Loss = 0.8× “SSIM”+0.2×MAD for training, then measured using “VMAF”
“SSIM” = C * (1 − SSIM), with some unknown scaled factor C, not using the multiscale SSIM.
The Y, U, V columns are PSNR.
“VMAF” here means what is done in FFMPEG. There are several VMAFs, different versions and different picture resolutions. A particular VMAF “model” was chosen. The contribution does not say which VMAF was used. Version 0.6.1 was used. The “model” and the version are not the same thing.
It was commented that there is a default model in VMAF.
Experimental results (All Intra)
All Intra Main 10 | |||||||
VMAF | Y | U | V | EncT-CPU | DecT-CPU | DecT-GPU | |
Class A1 | −5.84% | −1.39% | −3.72% | −4.11% | 170% | 41444% | 3363% |
Class A2 | −7.22% | −1.56% | −7.54% | −8.19% | 137% | 21508% | 2276% |
Class B | −6.51% | −0.67% | −7.97% | −6.85% | 131% | 21345% | 2043% |
Class C | −8.11% | −2.03% | −8.53% | −7.34% | 117% | 21533% | 1852% |
Class E | −10.29% | −2.86% | −4.65% | −5.29% | 139% | 37679% | 2804% |
Overall | −7.50% | −1.61% | −6.76% | −6.46% | 136% | 26294% | 2331% |
Class D | −9.26% | −2.70% | −7.64% | −8.16% | 115% | 19974% | 1984% |
Class F | −3.08% | −0.34% | −4.70% | −3.50% | 115% | 20602% | 1826% |
Experimental results (Random Access)
Random access Main 10 | |||||||
VMAF | Y | U | V | EncT-CPU | DecT-CPU | DecT-GPU | |
Class A1 | −4.66% | −0.67% | 3.94% | 4.45% | 145% | 51787% | 4174% |
Class A2 | −7.14% | −1.06% | −8.02% | −9.09% | 27635% | 2639% | |
Class B | −6.02% | −0.47% | −7.98% | −5.34% | 141% | 12545% | 1132% |
Class C | −6.17% | −0.60% | −5.86% | −4.44% | 125% | 16030% | 1450% |
Class E | |||||||
Overall | −6.01% | −0.66% | −5.04% | −3.89% | 20826% | 1859% | |
Class D | −6.58% | −0.51% | −4.58% | −4.66% | 124% | 29333% | 2407% |
Class F | −1.57% | 0.39% | −2.15% | −1.34% | 161% | 40522% | 3314% |
Experimental results (Low Delay B)
Low delay B Main 10 | |||||||
VMAF | Y | U | V | EncT-CPU | DecT-CPU | DecT-GPU | |
Class A1 | |||||||
Class A2 | |||||||
Class B | −2.11% | 0.14% | −8.86% | −6.42% | 137% | 14937% | 1307% |
Class C | −3.07% | 0.55% | −8.36% | −6.76% | 126% | 24803% | 2043% |
Class E | −1.59% | −0.16% | −3.85% | −3.07% | 181% | 24373% | 2228% |
Overall | −2.30% | 0.20% | −7.44% | −5.69% | 143% | 19991% | 1733% |
Class D | −1.82% | 0.49% | −6.83% | −6.70% | 123% | 29599% | 2480% |
Class F | −0.15% | 1.52% | −3.02% | −1.66% | 173% | 66822% | 4844% |
Below tables are for a different loss weighting: Loss=0.2×SSIM+0.8×MAD.
Experimental results (All Intra)
All Intra Main 10 | |||||||
VMAF | Y | U | V | EncT-CPU | DecT-CPU | DecT-GPU | |
Class A1 | −4.68% | −1.79% | −5.88% | −6.16% | 156% | 22366% | 2423% |
Class A2 | −4.41% | −1.94% | −8.81% | −7.77% | 130% | 33088% | 2789% |
Class B | −3.96% | −1.48% | −8.85% | −8.31% | 124% | 23956% | 2163% |
Class C | −4.48% | −2.40% | −8.47% | −8.37% | 115% | 20816% | 1829% |
Class E | −6.68% | −3.28% | −5.47% | −7.27% | 130% | 34118% | 2790% |
Overall | −4.72% | −2.11% | −7.70% | −7.70% | 129% | 25695% | 2312% |
Class D | −6.69% | −3.14% | −7.24% | −9.38% | 112% | 19564% | 1939% |
Class F | −1.82% | −0.78% | −4.59% | −3.92% | 114% | 9311% | 1190% |
Experimental results (Random Access)
Random access Main 10 | |||||||
VMAF | Y | U | V | EncT-CPU | DecT-CPU | DecT-GPU | |
Class A1 | −4.89% | −2.40% | −7.31% | −6.02% | 38157% | 3240% | |
Class A2 | −4.04% | −2.29% | −11.11% | −9.08% | 131% | 25447% | 2202% |
Class B | −3.54% | −1.53% | −10.26% | −8.49% | 128% | 27498% | 2425% |
Class C | −2.97% | −1.49% | −7.12% | −6.61% | 120% | 22800% | 1961% |
Class E | |||||||
Overall | −3.76% | −1.84% | −9.00% | −7.61% | 27499% | 2382% | |
Class D | −3.96% | −1.73% | −6.69% | −7.44% | 120% | 13263% | 1292% |
Class F | −0.76% | −0.37% | −2.79% | −2.47% | 151% | 5186% | 886% |
Experimental results (Low Delay B)
Low delay B Main 10 | |||||||
VMAF | Y | U | V | EncT-CPU | DecT-CPU | DecT-GPU | |
Class A1 | |||||||
Class A2 | |||||||
Class B | −2.61% | −1.03% | −12.13% | −9.81% | 127% | 32380% | 2847% |
Class C | −2.43% | −0.91% | −10.84% | −9.41% | 120% | 33890% | 2804% |
Class E | −2.56% | −2.55% | −5.33% | −7.24% | 171% | 18160% | 1616% |
Overall | −2.54% | −1.37% | −10.00% | −9.04% | 134% | 28450% | 2459% |
Class D | −3.09% | −1.17% | −10.34% | −10.32% | 118% | 20621% | 1812% |
Class F | −0.77% | −0.31% | −4.95% | −3.77% | 154% | 11838% | 1326% |
Thus, the first weighting is better for VMAF impact (in overall average), while the second one does better for MAD. SSIM calculation was in the JEM, and a patch had been offered for the VTM in the MPEG Deep Neural Network Video Coding (DNNVC) activity in June 2020. It was commented that it would be better to modify the VTM to use the SSIM implementation from HDRTools, which should not be difficult since something similar is done for HDR testing.
It was commented that there are different versions of SSIM as well (e.g. windowing type, windowing size, how to handle edge regions).
HDRTools has SSIM (with a Gaussian window) and MS-SSIM (without the Gaussian window); VTM does not.
Chroma was upsampled to feed to the NN and then downsampled at the output of that stage (staying in YUV).
A QP is an input to the model. Training used 5 QP values.
Network parameters are 32-bit float.
There was discussion of potentially just abandoning PSNR as a metric.
It was commented that it would be good to check whether visual quality is aligned with the analysis.
Another participant said that in some testing VMAF was not so stable – e.g., with deblocking filtering being harmful to VMAF despite being beneficial to visual quality. Some work has recently been conducted to develop new VMAF metrics to try to improve the behaviour.
Temporal consistency would also be good to study.
SSIM ignores chroma.
A participant emphasized that there is no 100% reliable metric in this world. VMAF, MS-SSIM, PSNR all have issues. It was suggested not to push or restrict any metric to be studied, but such study should not only include Excel table, but also proper viewing with MOS collection and computing correlation between MOS and all of the candidate metrics. The most difficult parts here would be collecting materials for viewing and conducting the viewing. The rest (metrics computation and correlation) is relatively easy. Without having those numbers in our hands, discussions such as which metric is better could be endless.
There was a question about how important is the “SE block”, and the proponent said they believed it to be quite important.
It was suggested to establish a BoG and to have a joint discussion with the parent bodies on the future NNVC exploration plans. See JVET-T0130 for a BoG report and section 7.2 for notes of the related joint meeting with the parent bodies.
Further study on JVET-T0069 was encouraged.