JVET-L0365 MS-SSIM as an additional metric [Y. Zhao, H. Yang, J. Chen (Huawei), M. Pettersson, R. Sjöberg, P. Wennersten (Ericsson)]
This contribution proposes to include the MS-SSIM metric as additional metric in VTM and make MS-SSIM Y mandatory in the CTC for SDR video. A patch for MS-SSIM integrated into VTM 2.0.1 and an updated excel template for the CTC for SDR video are provided. In terms of MS-SSIM Y, VTM2.0.1 achieves -16.28%, -22.86% and -17.47% BD-rate over HM16.18 in AI, RA and LB, respectively. In terms of PSNR Y, the BD-rates over HM16.18 are -18.03%, 23.08% and -18.29% in AI, RA and LB, respectively.
No test results were provided in the contribution. Some test results, e.g., VTM vs. HM were requested.
A revised version of the contribution contains such results.
For most tools, BD-rate measurements using PSNR or MS-SSIM are very similar.
It was remarked that it may be preferable to have an external tool to measure MS-SSIM instead of integrating it into VTM as relying on encoder output is not ideal.
No experts who were present in the room (other than the proponents) expressed the intent to use such metric if it were made available in VTM.
No action was taken on this.
No test results were provided in the contribution. Some test results, e.g., VTM vs. HM were requested.
Plenary meetings, joint Meetings, BoG Reports, and Summary of Actions Taken
Plenary meeting Sunday 7 Oct 0900
Reports of the tracks were given by GJS (Track B) and JRO (Track A) as follows (from the discussions, see also some comments added to the respective sessions)
Track B
- CE2 Adaptive loop filter and related
- L0083 Reduced bit-depth coefficients (complexity reduction, no loss)
- CE2.6.2 Subsampling of classifiers (complexity reduction, text in JVET-L0147, tiny loss)
- L0392 Context initial states (minor correction, a tiny bit of gain)
- Remove signalling of 5x5 as a special case for luma
- CE4 Inter prediction and related
- Adopt 4.1.6.a Simplification on AMVP candidate list construction (complexity reduction, text in JVET-L0271, no loss)
- Adopt 4.1.11.a (complexity reduction, line buffer reduction for affine inherited candidates, text in JVET-L0045, 0.1% loss)
- Adopt 4.2.6.d affine merge as modified in JVET-L0632 (coding eff, pending adaptation of text, 0.66%)
- Adopt 4.2.8 moving ATMVP into the affine merge list (cleanup, assuming ATMVP operates on a subblock basis, pending text, 0.03% penalty)
- Adopt 4.4.12.a (0.38% in RA), merge with average of pari, list size 6 (text in JVET-L0090)
- Merge with MVD variant b with base candidate from merge list (1.29% in RA)
- History method with 6 candidates for merge list (0.58% in RA, text in JVET-L0266)
- BoG reviewing related contributions
- CE9 Decoder motion vector derivation and related
- L0256 BIO (1.28% in RA, text being prepared)
- BoG to review non-CE contributions
- CE10 Multi-hypothesis and combined prediction
- No action (yet)
- HLS
- NAL units, SPSs, tiles, no CTU-level slices
- PPSs or (perhaps repeatable) picture headers
- Maximum capability negotiation header level – something with a persistence scope beyond the CVS
- Tile groups, starting point with raster order (not presumptive status), flag after each tile to indicate whether there are more. Discussion in the plenary resulted in the following Decision:
- CABAC stop bit at the end of each tile.
- No flag to indicate the last CTU in tile.
- Number of tiles minus 1 in header, entry point signalled for each tile
- The syntax of the last tile is no different from any other tile, and there is no additional syntax after that tile in the NAL unit.
- Design goal of extracting VCL NAL units of an MCTS and reposition to another bitstream without substantial difficulty
- Some system interaction issues were identified that would require feedback from systems experts if action is to be taken on them.
- Other
- Decision (CTC): Make Class F mandatory (but not included in the average).
Track A:
CE1: Partitioning
- Picture boundary handling – no action
- Split constraints – restrict “virtual data pipeline” to 64x64 – some loss, deemed useful, BoG
- Separate trees – signalling at CTU level or lower (intra / inter): Not enough benefit, except for “constructed sequences” – no action
CE3: Intra prediction and mode coding
- CE3.1: Multiple reference line prediction (9 tests) – adopt method that uses (with signalling) line 1 or 3 (within CTU only) – 0.4% gain, almost same encoding runtime
- CE3.2: Intra prediction modes (9 tests): Interesting gain from sub-partition (1%) and
(non-)linear weighted prediction (2.5%/1.6%), but encoder/decoder complexity concerns; multiple direct modes 0.2% gain, but more complexity in list construction - CE3.3: Intra reference sample interpolation (7 tests): Promising gain (0.4%), complexity study of different proposals ongoing
- CE3.4: Bidirectional prediction (3 tests) – no action
- CE3.5: Cross-component prediction and separate chroma tree (18 tests): Adopt proposal using only 1 line at CTU boundary for CCLM; adopt new computation for linear model which is significantly less complex; adopt MDLM (switching left/top region), about 2% chroma gain; other methods non action or further study
- CE3.6: Intra mode coding (7 tests) – 6 MPM shown to have benefit – BoG about studying different proposal
CE5: Arithmetic coding engine
- Probability estimation with subtopics higher precision (lin vs. log), multi-probability models (add. 0.3%), customized estimation window (add. 0.3%), total around 1%. Complexity/pipeline throughput study of proposals in BoG
- Precision of engine itself
CE6: Transforms
- CE6-1: Primary transforms (21 tests, 11 proposals): Ideally, the transform should have the following properties:
- Sharing of as much as possible building blocks for different transform types and sizes
- Implementation either as matrix multiply or fast algorithm, independent of specification
- Implementation with 16-bit logic (at least for 10-bit video)
- As low complexity as possible
- Re-usability of legacy building blocks might be desirable
- Specification as matrix multiply, or cascade of matrix multiplies, or other e.g. butterfly
- Extraction of smaller transform sizes from largest size 64 (32 for MTS transforms)
No proposal solves all of these aspects. One proposal adopted which goes back to original HEVC DCT-2 transform (8-bit integer representation of matrix values, with 64-length transform added). Some of the MTS transforms (DCT-8/DST-7) more difficult to implement in fast algorithms due to asymmetry. Further analysis of proposal properties in BoG
- CE6-2: Secondary transform (6 tests, 2 proposals): Interesting reduction of complexity, only 4 transform sets, but results only available in combination with reduced version of MTS (which gives 1% in AI, 0.6% in RA with overall no increase in encoder/decoder time). Proponents requested to provide results of sec. transform itself on top of VTM, estimate it would be close to 2% rate gain in AI, close to 1% in RA.
- CE6-3: Transform combinations and signalling (7 tests, 3 proposals): Combinations of MST and NSST more of BMS style, no action here.
CE7: Quantization and coefficient coding
- CE 7.1: Transform coefficient coding (4 tests): Adopted new binarization for dependent quantization, 0.15% gain
- CE 7.2: Block adaptive quantization / residual coding (7 tests): Subjective quantization no specific action, some bit rate reduction for saving quantization parameter bits (but loss versus CTC)
- CE 7.3: Transform coefficient scanning (3 tests): No benefit
CE8: Current picture referencing:
- CPR (HEVC style) results with different restrictions of reference area. Restriction to current CTU (or current plus left) most promising in terms of memory (avoiding problem of storing entire picture in two versions). Gain around 40% for SCC class (less for class F), but only 0.2% for CTC. Another approach with decoder side template matching uses similar memory, but more complex decoder processing (though largely reduced relative to previous proposals), gives 1% for CTC, but less than CPR for SCC content. It is discussed in the plenary whether CPR with CTU restriction would be mature for VVC adoption. Since CTU is larger in VVC, necessary memory (cache?) could be of concern – could also have relation with the “virtual data pipeline” of CE1.
CE11: Deblocking
- Longer filters for large blocks: Various proposals, all add some processing complexity which may be of less concern for large blocks, but also more line buffer. Further analysis of subjective test results ongoing
- Various aspects: Adopted one proposal for brightness-adaptive strength adaptation (aleady in CfP, shown beneficial in subjective test); expression of adaptation variable, default none.
- Deblocking at 4x4 block boundaries: Seems useful as per subjective results, parallelism is possible from CE contributions, complexity analysis of details ongoing
CE12: Mapping functions: HDR – In-loop and out-of-loop reshaping seem almost identical in results (out-loop seems sufficient). For SDR, in-loop reshaping provides objective (BD) gain 1% for AI, 1.2% for RA. However, the method requires block-level reshaping and inverse reshaping in decoding process, and introduces additional inter-component dependency.
CE14: Post-reconstruction filters: Two approaches – bilinear and Hadamard-domain. Both are block-level building blocks, per se manageable in complexity, but could be critical in intra prediction (additional cycles between reconstruction and prediction. Both give 0.7% in RA, 0.5% in AI. More detailed complexity analyis (BoG), then decide.
CE15: Palette mode: Investigates palette (close to HEVC) in VTM, and another reduced complexity method. Only gives gain for screen content sequences, less than CPR (30% in SCC set), but additional gain (few percent) still are retained when combined with CPR. From the discussion in plenary: As palette mode is a completely different building block specifically beneficial, lower complexity solutions may be desirable. General aspect how to deal with tools that are specific for one type of content to be clarified.
CE1-related topics: Some straightforward cleanups and more flexible tree property signalling were agreed.
Plenary Wednesday 10 October 1400
Track A
- Decision: Adopt JVET-L0053 first aspect / JVET-L0272. Proponents shall check if their text is identical and if not, unify them.
- Decision: Adopt JVET-L0279 CE3-related: Unification of angular intra prediction for square and non-square blocks
- Decision: Adopt JVET-L0095 CE6-related: Simplification on MTS kernel derivation
- Decision (BF/SW): Adopt JVET-L0111 CE6-related: Transform Skip Condition on Transform Block size
- Decision: Adopt JVET-L0628 3.1.4.2 CE3.3: Intra reference sample interpolation (as this filter is used somewhere else in the design)
- Decision (ed./text improvement): Adopt JVET-L0217 Non-CE1: Relation Between QT/BT/TT Split Constraint Syntax Elements (as per v4)
- Decision: Adopt JVET-L0361 CE1-related: Context modeling of CU split modes (version with 22 context models)
- Decision: Adopt JVET-L0678 QT/BT/TT Split Constraint Syntax Elements Signalling Method. The split constraints in CTC were not to be changed, but encoder needed to be modified to signal them.
- Decision: Adopt JVET-L0209 PCM mode with dual tree partition
- Decision: JVET-L0362 Quantization parameter signalling Adopt method 1 (depth is QT depth + BT depth)
- Decision: Adopt first aspect of JVET-L0428 Delta QP and Chroma QP Offset for Separate Tree (use centered position to fetch collocated luma QP).
- Decision (bug fix): JVET-L0553 Adopt second fix to semantics of init_qp_minus26 where +25 is changed to +37
- Decision (SW): adopt JVET-L0181 AHG10: Corrected operation of ALF encoding with perceptually optimized QP adaptation.
- Decision: Adopt JVET-L0165. Text was reviewed in BoG. It is however pointed out that there is an inconsistency in the specification of coding the remaining modes. The software codes them as truncated binary, whereas the text specifies fixed length coding (as was used with 3 MPM before). It was to be confirmed by text editors that the specification had been corrected.
- Decision from BoG on CE1 SubCE2 and related contributions Adopt JVET-L0081 Test 2.1.2
- Decision (agreed in the plenary): Adopt JVET-L0552 to set the initial CABAC states based on the CTC test set, for each major version of the VTM (to also verify with an independent set of test sequences that the deviation of results is not severe on that other set).
Track B
- JVET-L0043 Decisions on agreements in principle:
- It was suggested to agree in principle, as a starting point, to have something at the start of the SPS that indicates properties that cannot be violated in the entire bitstream. Agreed.
- It was suggested to start with presumption that there would be a list of disable flags. Agreed.
- It was suggest to agree in principal that there would a separation between things that affect the syntax or decoding process and things that merely express constraints. Agreed.
- JVET-L0098 The decision was to disable the DMVR for the coding block sizes with either width of height of 128.
- JVET-L0691 BoG recommended adoptions to VTM
- Normative changes
- Unification of affine CPMV, choose JVET-L0047 method 1 or JVET-L0047 method 2 (the same as JVET-L0373)
- Method 1 control point MVs are stored and used only for model inheritance, other places (ATMVP storage, deblocking, motion comp, spatial neighbours for merge list, AMVP list derivation) use subblock MVs calculated from the control point MVs. This method has some extra memory (~768 bytes for hardware implementations) and a small benefit (0.05% average). It was commented that a new contribution JVET-L0666 reported that the peak loss for method 2 for some non-CTC affine-friendly sequences was substantially bigger. Method 2 is reportedly always (a little) worse in coding efficiency. Another late contribution reports a way to reduce the extra memory. If the subblock size is made bigger, method 2 would have some inconsistency in the motion vector field relative to the model. Decision (design cleanup): Adopt method 1 as the more consistent and “clean” design (roughly neutral on coding efficiency 0.01%). Further study of other schemes is anticipated.
- ATMVP modification: use fixed subblock size 8x8 for ATMVP (JVET-L0198, JVET-L0468, JVET-L0104, possibly some others). Currently we’re adaptively using 4x4 or 8x8 subblock size, but this has no benefit. Decision: Agreed (approx. no coding efficiency impact).
- In the plenary, it was noted that deblocking for ATMVP only applies to 8x8 CU boundaries, so the subblock boundaries are not deblocked. It was said that deblocking is applied to 8x8 boundaries for affine prediction, but another participant said this was not the case. It was noted that the residual transform is applied across subblock boundaries, and suggested that the deblocking should depend on whether there is a residual and perhaps whether it has non-DC coefficients. It was commented that CE11.3.2 proposed applying deblocking at ATMVP subblock boundaries.
It was agreed that we need to add text for deblocking into the draft text.
Decision: It was agreed that the draft spec will say deblocking is applied only at 8x8 grid-aligned TU boundaries and when there is no residual, deblocking is applied if there is a motion vector difference above the threshold on an 8x8 grid-aligned position. The software needs to be checked to ensure that this is what it is doing too. Further study of these interactions is needed. The deblocking BoG met later to address details.
- In the plenary, it was noted that deblocking for ATMVP only applies to 8x8 CU boundaries, so the subblock boundaries are not deblocked. It was said that deblocking is applied to 8x8 boundaries for affine prediction, but another participant said this was not the case. It was noted that the residual transform is applied across subblock boundaries, and suggested that the deblocking should depend on whether there is a residual and perhaps whether it has non-DC coefficients. It was commented that CE11.3.2 proposed applying deblocking at ATMVP subblock boundaries.
- ATMVP modification: restrict ATMVP mode to CUs of which both the width and height are larger than or equal to 8 (L0055), note that this is already a part of 4.2.8 which had been adopted. Decision: Agreed (approx. no coding efficiency impact).
- ATMVP modification: check the first spatial neighbouring motion vector and use this as the reference motion vector for the collocated position for motion vector derivation (L0198). Decision (complexity reduction): Agreed (approx. no coding efficiency impact).
- Reset the FIFO table in each CTU row for HMVP (JVET-L0106, JVET-L0158 method 1). Decision (complexity reduction): Agreed (approx. no coding efficiency impact).
- Merge index coding: use one context for the first bin of the full-block merge index and bypass coding of other bins (L0194). After discussion, it was noted that this had not been tested for the subblock merge list and we don’t want different syntax for the two cases. Further study in a CE was planned to test doing the same thing for the subblock merge list. Decision (complexity reduction): Agreed (for the full-block merge index only at this time, approximately no coding efficiency impact).
- Generalized bi-prediction (L0646). About 0.66% gain on RA, about 6% increase in encoding runtime. There was discussion of the alternative of using weighted prediction with multiple weights per picture. Weighted prediction would probably work better than this for fade-in, fade-out, and cross-fades (e.g., since it doesn’t need block-level weight selection and since it can extrapolate as well as interpolate), so this proposed method is not a complete replacement for weighted prediction. It was commented that there had been previous contributions describing a benefit for using weighted prediction with multiple weights, and there is some support in the JM for this sort of usage, but not in the HM. AVC had an extra, implicit mode of weighted prediction that was not adopted into HEVC, but it used POC weighting rather than signalling to establish the weights. A proponent noted that the syntax of weighted prediction does not have a shortcut for biprediction with two weights that add up to 1, so the proposed signalling of what is proposed as “generalized biprediction” for VVC would be more efficient when that constraint is intended to apply (e.g., the equivalent of selecting among 5 pairs of weights that add up to 1 would require selection among many more possibilities). Decision (coding efficiency): Adopt JVET-L0646 (0.66% coding efficiency; weighted prediction should also be put in the draft, but this and weighted prediction would be mutually exclusive at the picture level, when used with OBMC the weights of the neigbours would apply for the neighbour predictors, which is how the BMS software already does it, no consideration in deblocking filter). Further study of alternative approaches is expected and encouraged.
- Prohibit 4x4 bi-prediction for inter CU (JVET-L0104 & JVET-L0371). Decision (complexity reduction): Agreed (negligible effect on coding efficiency). Further study is planned for other related aspects.
- L0265 to set the chroma subblock size to 4x4 instead of 2x2 for affine motion compensation by averaging the MVs of the 4x4 luma subblocks. It was commented that SIMD implementation is feasible for 4x4 but not 2x2. This has a negligible coding efficiency effect. It was noted that we also have 4x4 ordinary CUs, so this doesn’t entirely solve the 2x2 problem, but would leave that as the only case where this occurs. This would apply to both uni an bi-prediction. Decision (complexity reduction): Adopt.
- Unification of affine CPMV, choose JVET-L0047 method 1 or JVET-L0047 method 2 (the same as JVET-L0373)
- Bugfix of VTM software
- Align the software with the draft text regarding ATMVP motion vector clipping (L0257). Decision (change software to match text): Agreed.
- Rounding motion vectors toward zero rather than toward minus infinity for AMVR (L0377). Decision (change software to match text): Agreed.
- L0093 align VTM with draft text regarding the pruning of regular merge list (the same as JVET-L0282). The draft text does not do full pruning for the spatial and TMVP candidates in the merge list. The software does full pruning. It was reported that there is no loss for not doing full pruning. Decision (bug fix): Align software with text.
- Encoder optimization
- Encoder optimization for affine motion estimation (L0260). Decision (software): Adopt (0.3% coding gain, 3% encoding time increase).
- L0694 interaction refinement (JVET-L0045 line buffering for affine model inheritance across CTU boundaries interaction with JVET-L0047 storage of subblock motion vectors). This was further discussed in the plenary, without change of the decision.
- Adopt CE10.1.1.c combined intra/inter with restriction to w×h >= 64 luma samples (0.5% in RA)
- Decision (coding efficiency): Adopt Non-rectangular (triangular) partitions (0.57% in RA, 1.23% in LB), with the JVET-L0208 bug fix, flag after combined intra/inter.
The 360° BoG report JVET-L0647 was reviewed in the plenary.
Closing plenary sessions
The sessions of JVET held on the afternoon of Thursday 11 October and the morning of Friday 12 October. These plenary sessions included BoG review, other review and finalization of meeting outcomes, CE and AHG planning, output document planning, assignment of editorships, future meeting plans, expressions of thanks, and closing of the meeting. The outcomes of these plenary discussions are recorded elsewhere in this report.
Joint meetings
No joint meeting sessions at the parent body level were held at this meeting regarding the work of JVET.