JVET-M0906 Subjective assessment of CE11 (deblocking filter) proposals [V. Baroncini, A. Norkin, A. M. Kotra, K. Andersson, K. Misra, H. Jang, C. M. Tsai, D. Rusanovskyy]
This report was presented in Track B Thursday January 17 1200-1330 (chaired by JRO).
This contribution provides a report of the subjective test for the proposals in Core Experiment 11 on deblocking filtering. Two set of tests were performed during the Marrakech meeting in accordance with the CE11 description document JVET-L1031. Both tests involved visually evaluating CE proposals versus VTM3 with ALF turned OFF as anchor. ALF was also turned OFF for the CE proposals.
To facilitate further analysis a score representing the number of times the Mean Opinion Score of a CE proposal is better than the anchor, and the anchor lies outside the confidence interval was also computed.
Expert viewing was conducted during the Marrakech meeting for both CE 11.1 (Longer tap filter) category and CE 11.2 (4 x 4 luma sample grid filtering). The sequences and the QP points used are as listed in the core experiment plan description JVET-L1031. The invitation for the subjective viewing participation was issued on the JVET reflector. 12 non-proponent viewers were selected as the test participants for subjective testing for both the categories. The 12 participants were divided into 4 groups containing 3 participants each.
The dates and the actual timings of the viewing sessions are included with the attached spreadsheets along with the contribution. Each coded video was compared with the anchor; the order of presentation of anchor and coded videos was randomly changed. The Anchor used in all the configurations was with ALF (adaptive loop filter) tool turned off. The coded video also had the ALF (adaptive loop filter) tool turned off.
A video (preceded by the latter A) was shown on the screen, followed a second video (preceded by the letter B); once again A video (preceded by the latter A), followed by the second video (preceded by the letter B) is shown on the screen; then the message “Vote N” was presented asking to the viewers to rate if “A was much better than B”, “A was better than B” or “B was better than A” or “B was much better than A”.
The order of presentation of the video clips inside each test session was assigned randomly taking care not to show the same sequence two consecutive times.
For the CE 11.1 category, the following test sequences and conditions were used in the subjective tests. The sequence length is 10 seconds.
Num | Name | Resolution | fps | Frames | Mode | QP |
01 | FoodMarket | 3840x2160p | 60fps | 600 | RA | 39, 34 |
02 | Campfire | 3840x2160p | 30fps | 300 | RA | 39, 34 |
03 | Kimono | 1920x1080p | 24fps | 240 | LDB | 34, 39 |
04 | KristenAndSara | 1280x720p | 60fps | 600 | LDP | 34, 39 |
05 | Red Kayak* | 1920x1080p | 30fps | 300 | RA | 34, 39 |
* = The first 300 frames of the RedKayak sequence were used.
To facilitate analysis, the scores provided by viewers were translated to a fix scale of 1, 2, 3 or 4, with:
1 representing anchor is much better than a CE proposal,
2 representing anchor is better than a CE proposal,
3 representing a CE proposal is better than the anchor, and
4 representing a CE proposal is much better than the anchor.
The anchor used is VTM3 with ALF OFF. Mean Opinion Score (MOS) and Confidence Intervals (CI) were computed for each data-point. For each QP, these were in turn used to compute:
- Number of times (x) a CE proposal beats the anchor i.e. number of times (MOS – CI) is greater than 2.5; and
- Number of times (y) anchor beats the CE proposal i.e. (MOS + CI) is less than 2.5.
The total score for each CE proposal, for a given QP, is then computed as x – y.
Example illustrating Test A wins, Test B tie, Test C loses for a sequence
From such an approach, it was found that in all cases proposals were either better equal, no case was found worse than the anchors. All tests were performed with “ALF off” anchor. When “ALF on” was compared to “ALF off”, it was also found to be visually better than the anchor in 6 out of 10 test cases, whereas some proposals are better than the anchor in 9 of the test cases, as shown by the following table:
Listed below are the total score of the CE proposals in CE11.1. The attached excel contains more detailed results.CE11.1: Long Filters Tests – highest possible total score per QP is +5, lowest possible score per QP is -5:
CE Proposal | Total score for QP34 | Total score for QP39 |
CE11.1.1 | 2 | 4 |
CE11.1.2 | 3 | 4 |
CE11.1.3 | 1 | 2 |
CE11.1.4 | 2 | 2 |
CE11.1.5 | 3 | 4 |
CE11.1.6 | 4 | 5 |
CE11.1.7 | 2 | 5 |
CE11.1.8 | 4 | 5 |
One conclusion that can be drawn is that longer-tap deblocking helps for visual quality at large CU boundaries. However, it can not directly be concluded that the effect would still be visible with similar clarity when ALF was on. From looking at scores of individual sequences, it can be seen that some of the ne deblocking with ALF off are still working better than existing deblocking with ALF on in approximately half of the cases, such that it can be concluded that new deblocking methods give a benefit in terms of subjective quality that is additive to the benefit of ALF. Therefore, a conclusion might be drawn that when both new deblocking and ALF were enabled, a visual benefit would still be visible. To investigate if such a statement is true, additional viewing should be done with a selected proposal, one “best performing” (from the table above, where 11.1.8 seems to be the best candidate, as it has good performance and does not increase the need of line buffering). Additional viewing to be performed to confirm that 11.1.8 with ALF on still gives benefit compared to VTM3 with ALF on.
Decision: Adopt JVET-M0471, version 11.1.8 (specification text available in v2 upload, but needs another small modification for restriction of line buffer, was shortly review in Track B Thursday 17 January 1330), pending on confirmation from the viewing, and the more detailed report on complexity impact.
Informal viewing (3 sessions, around 20 participants, 12 of which were not involved in this CE) was conducted Thursday 17 January evening. It is reported that also in comparison to the ALF on anchor, differences are clearly visible in particular for sequences Campfire, Redkayak, and also slight improvement for Foodmarket and KristenSara.
The complexity analysis in JVET-M0031v4 was presented Friday 18 January in Track B. For the adopted proposal, the worst case number of operations is not increased for luma, and the number of line buffers is kept the same. Additional complexity is the need for switching to another filter mode in the case of large blocks, which is not critical. For chroma, the worst case number of operations is increased by approximately 10, as a separate decision is employed. This increase of complexity appears justified by the fact that it gives considerable visual improvement in cases of sequences where chroma is critical.
For CE11.2, results are not as conclusive. A similar analysis as above provides the following table (where in this case, the highest score would be 4):
CE Proposal | Total Score for QP30 | Total Score for QP34 |
CE11.2.1 | 0 | -1 |
CE11.2.2 | 0 | 1 |
Compared to that, ALF on versus ALF off has a score of 3 and 2, for QP30 and 34, respectively. This indicates that ALF has a clear benefit over any of the proposals, and it can hardly be concluded that they would further improve the quality when combined. At least in this range of bit rates, deblocking on an aligned 4x4 grid does not seem to improve the visual quality. No action necessary from these results. Could be due to the fact that the design of the CE had a flaw in selecting too low QP range. Might be worthwhile to continue the study now in combination with longer filters, and comparing with ALF on as anchor.
The subsequent notes only contain abstracts copied from the documents. Actions taken are noted above under JVET-M0031 and JVET-M0906.