JVET-M0464 Non-CE8: Unified Transform Type Signalling and Residual Coding for Transform Skip [B. Bross, T. Nguyen, P. Keydel, H. Schwarz, D. Marpe, T. Wiegand (HHI)]
This contribution proposes two modifications related to transform type signalling and residual coding for transform skip. The first modification includes aligning the cases in which transform skip (TS) and multiple transform selection (MTS) apply. This limits TS to luma transform blocks as well as extends it to transform block sizes up to 32x32. The second modification constitutes of a modified transform coefficient level coding for the TS residual. Relative to the regular residual coding case, the residual coding for TS includes no signalling of the last x/y position, coded_sub_block_flag coded for every subblock, sig_coeff_flag context modelling with reduced template, a single context model for abs_level_gt1_flag and par_level_flag, context modelling for the sign flag, additional greater than 5, 7, 9 flags, modified Rice parameter derivation for the remainder binarization and a limit for the number of context coded bins per sample. The proposed joint signalling (first modification) provides on average (BD-rate Y, enc. time, dec.time):
- Natural: 0.01%, 103%, 101% (AI), -0.02%, 98%, 100% (RA), -0.07%, 102%, 102% (LB)
- Class F: -1.96%, 103%, 97% (AI), -2.14%, 98%, 100% (RA), -2.79%, 103%, 99% (LB)
- TGM: -8.04%, 106%, 88% (AI), -8.74%, 106%, 96% (RA), -9.21%, 112%, 97% (LB)
Additionally changing the residual coding for transform skip (first modification and second modification) results on average in (BD-rate Y, enc. time, dec.time):
- Natural: -0.16%, 103%, 100% (AI), -0.07%, 98%, 101% (RA), -0.07%, 101%, 101% (LB)
- Class F: -7.16%, 104%, 94% (AI), -5.76%, 98%, 100% (RA), -5.85%, 102%, 99% (LB)
- TGM: -21.03%, 108%, 82% (AI), -15.57%, 105%, 94% (RA), -14.01%, 111%, 95% (LB)
v2 provides updated full frame results, details about the combined TS/MTS encoder search with results for different encoder operation points as well as draft text for unified TS/MTS syntax.
v3 corrects wrongly pasted results to match with the cross-check provided in JVET-M0708.
The aspect on TS/MTS in the proposal signals TS before MTS (and therefore also enables it for blocks up to 32x32), but also changes the MTS binarization
Results of TS modification are on top of VTC, would be desirable to see benefit for SC classes when CPR is on
There is also an aspect of encoder speedup, by another approach of early skip for testing MTS cases.
The modified coding method employs max 8 context coded bins per sample – much higher than current VVC limit.
The way of restricting the number of context coded bins (once the budget is consumed, TS is not used any more) is not elegant.
Question is raised whether it has subjective quality impact?
It is dicussed whether the aspect of signalling TS before MTS (and by that way enabling TS for block sizes up to 32x32, but also restricting it to luma) would rather be a straightforward syntax cleanup, which could be adopted at this meeting (but without modifying the MTS index binarization, which is discussed in CE6).
Later, results with a CPR-on anchor were reported. It was demonstrated that large-block TS still provides a benefit for screen content (1.6%, 4.9% for F/TGM in AI, 1.7%, 5.9% for LD). The adoption is reported elsewhere (from Track A).
JVET-M0464 Unified Transform Type Signalling and Residual Coding for Transform Skip
Signals TS before MTS (and therefore also enables it for blocks up to 32x32), but also changes the MTS binarization
CTC: 0.01%, 103%, 101% (AI), -0.02%, 98%, 100% (RA), -0.07%, 102%, 102% (LB)
Class F: -1.96%, 103%, 97% (AI), -2.14%, 98%, 100% (RA), -2.79%, 103%, 99% (LB)
TGM: -8.04%, 106%, 88% (AI), -8.74%, 106%, 96% (RA), -9.21%, 112%, 97% (LB)
Additionally changing the residual coding for transform skip (first modification and second modification) results on average in (BD-rate Y, enc. time, dec. time):
CTC: -0.16%, 103%, 100% (AI), -0.07%, 98%, 101% (RA), -0.07%, 101%, 101% (LB)
Class F: -7.16%, 104%, 94% (AI), -5.76%, 98%, 100% (RA), -5.85%, 102%, 99% (LB)
TGM: -21.03%, 108%, 82% (AI), -15.57%, 105%, 94% (RA), -14.01%, 111%, 95% (LB)
It is dicussed whether the aspect of signalling TS before MTS (and by that way enabling TS for block sizes up to 32x32, but also restricting it to luma) would rather be a straightforward syntax cleanup, which could be adopted at this meeting (but without modifying the MTS index binarization, which is discussed in CE6). However, results of TS modification are on top of VTC, would be desirable to see benefit for SC classes when CPR is on. This requires further consideration.
There are also other contributions JVET-M0072, JVET-M0269, JVET-M0279 and JVET-M0501 which also target TS.
CE9: Decoder-side motion vector derivation
CE9.1: BDOF design
Adopt JVET-M0487 (solution 9.1.1.b) which uses integer positions to generate the prediction samples in extended region, and uses 8-tap DCTIF filters to generate the prediction samples inside the CU.
No loss, but simplifies BDOF
CE9.2: DMVR design
Not decided yet, but good candidate for adoption: JVET-M0147 shows results without using refined MV for spatial MV prediction and deblocking: Overall -1.01%. This seems a reasonable approach avoiding all major dependency problems that were observed in DMVR before. This is however a variant for which cross-check still needs to be provided; specification text to be made available.
Generally, the investigation on DMVR has led to a point where it might be manageable implementation-wise (not low complex, but still giving around 1% gain)
A problem to be investigated: As currently tested, DMVR and BDOF could be applied sequentially, before finally motion comp and reconstruction can be done. JVET-M0223 considers this issue, still needs to be reviewed.
CE10: Combined and multi-hypothesis prediction
CE10.1:
10.1/10.2: Multi-hypothesis prediction: No action – gains are too low (and became lower than before, cut by half or even more) to justify additional complexity
In the context of reviewing 10.1.3.x proposals, which target simplifications of CIIP, the following aspects were identified:
Does it need a specific MPM derivation? If there was only one mode (as in 10.1.3a/d) it is not needed at all, and it is mentioned that there are proposals just suggesting fixed length coding
CIIP does not have any serious latency issue. Therefore, simplifications that remove some processing steps that are otherwise used in intra prediction is not necessary.
Sample-wise unequal weighting is not an implementation issue, whereas equal weighting would be preferrable, unless it costs compression performance or causes qualty problems.
Beyond the solutions tested in 10.1.3 more study is necessary on these aspects.
CE10.2: OBMC
Among the four proposals, version 10.2.1 seems the only one which is manageable from complexity perspective (but definitely adds some complexity). The test results are summarized as follows:
Proposal
Config.
Y
U
V
EncT
DecT
CE10.2.1
RA
-0.27%
-0.58%
-0.62%
104%
103%
LB
-0.36%
-0.50%
-0.51%
106%
105%
Track B initially suggested adoption of JVET-M0178.
During the JVET plenary, it was questioned whether such small gain was justifying the additional complexity. It is commented by several experts that OBMC may have positive impact on subjective quality. However, no such proof was available (and would probably be difficult to get). Thus, the decision was reverted.
CEs on the previous investigated technologies 10.2.1…10.2.4 should be closed.
CE10.3: Multiple shape prediction partitions: No action
CE10.4: Diffusion filters for intra and inter prediction
The new proposed version uses FIR filters, applied on the prediction signal, switchable, not iterative. Two different 1D filters are 9-tap, symmetric, only 1 multiplication, otherwise shifts. One 2D filter is a 5-tap diamond shape, only shift/add operations. At the boundaries, one sample from neighboured reconstruction is used. The approach provides gains of 0.17%/0.42%/0.17%/0.87% for AI/RA/LB/LP. The method has been brought down to acceptable complexity impact for decoder, at some penalty in performance.
Two concerns are raised:
As it needs to be run after the prediction signal is generated, it produces additional delay in the intra prediction loop.
For inter blocks, it can be used for 128x128 CU, which would break concepts of 64x64VPDU.
It was requested to provide additional results without using the method in intra prediction (and the intra part of CIIP), and restrict the largest CU size to 64x64.
It is also reported that a late contribution (JVET-M0848) provides new results for the same method of 10.4.2 with some (small) encoder speedup and slightly increased performance.
Furthermore, it would be desirable to use only the prediction samples (no reconstructed samples from current picture). A version which did that was shown in the previous CE.
No decision was made yet on this at the time of this plenary meeting.
CE10.5: Local illumination compensation
No action on current technologies, but continuation of CE to fulfill dependency/pipeline latency requirements.
CE11: Deblocking
Sub-tests
long-tap deblocking filters (11.1)
deblocking at 4x4 block boundaries (11.2).
The proposals were encoded according to two test conditions, which correspond to two anchors.
Anchor-1 is VTM-3.0 according to the CTC.
Anchor-2 is VTM-3.0 with ALF switched off (other conditions are the same as in Anchor 2).
CE11.1
Technology-wise understood and complexity-wise analysed, with some additional information on complexity requested.
It was pointed out that some of the proposals 11.1.1-5 did not consider that VTM3 now uses subblock boundary deblocking, whereas the decision of using long filters is based only on CU boundary properties. Therefore, it could happen that a long filter is applied first, and afterwards a subblock boundary within the CU is deblocked again. This would inhibit parallelism. This could be classified as a bug, which however in terms of subjective testing might not be too harmful.
CE11.2
VTM3 operates deblocking on an 8x8 grid. If a CU boundary is not on an 8x8 grid, it is not deblocked. If a CU is on an 8x8 grid, furthermore subblocks within that CU are deblocked (provided they are on the 8x8 grid as well). 4xN blocks are deblocked whenever they are coinciding with the 8x8 grid
11.2.1 deblocks on a 4x4 grid, where a boundary is deblocked with VTM filter when the next block or subblock is more than four samples apart, and with a weak and short filter if it is only samples apart.
11.2.2 uses an 8x8 grid and VTM deblocking filter but disables deblocking whenever the boundary would be only four samples apart.
Viewing in CE11 to assess necessity and benefit was still to be done (being prepared).
General remarks were made on late delivery of text – text needs to be provided.
Joint meeting Tuesday 15 January 0900-1000
In a joint meeting of VCEG and MPEG that focused primarily on new standardization projects in MPEG, an issue affecting JVET work was discussed.
It was suggested and agreed to have coordination of the timing of meetings of JVET consideration of systems-related issues (esp. high-level syntax) and to encourage strong feedback from those of the parent bodies that are focused on systems. It was further agreed to plan to write a separate standard text for parts of the design that should be in common between the video and systems specifications (e.g., SEI messages). Potentially relevant aspects were suggested to include SEI, VUI, metadata associated with post-decoding processing, sub-bitstream extraction and merging, and the byte stream format.
As an initial plan, it was agreed to defer final decisions on such topics of common interest at JVET meetings until the Saturday of each meeting, at which time those focused on systems aspects in the parent bodies would attend as effectively a joint meeting session.
Joint meeting Thursday 17 January 0900-0945
JVET met with MPEG Systems and Requirements.
A contribution discussing a proposed system decoding interface that had been submitted to MPEG was discussed, as noted below.
WG 11 M46578 On the decoding interface for immersive media [Emmanuel Thomas (TNO) Rob Koenen (Tiledmedia) Thomas Stockhammer (Qualcomm)]
This topic has been called “Immersive media access and delivery” in some MPEG work, and there was an MPEG output document N 18071 at the October 2018 MPEG meeting.
This is for a scenario with additional processing that takes place after decoding; not a 1:1 mapping between output of decoder and display of the decoded video by the decoding system.
Example: 360° video tiled streaming using cubemap with each face of the cubemap segmented further into tiles. Using view-port-dependent streaming to only serve the tiles needed for viewing. Problems mentioned:
Some systems having limits on the number of decoder instantiations.
Need for systems coordination of timing.
An example approach is rewriting the bitstream to produce a “packed picture” that is decoded.
The proposed alternative is a decoder interface with a relatively large number of decoders – not necessarily a tile-level decoding interface from the video spec perspective – each decoder could be processing whole pictures from its perspective.
Another example: decoding of a background scene and a foreground object that are later composited by the system into a combined scene.
It is proposed to be able to mix profiles, colour formats, frame rates, picture resolutions, etc., using this tile-level interface.
Issues:
Timing coordination (and ensuring decoding speed)
Reference picture (or picture regions) management
Other associated data used by a decoder
Buffer flow characteristics
The subject was reported to be under consideration in MPEG Systems and Requirements, including as MPEG-I Architecture.
Simulcast “layers” were noted as a way to do that. The proposal is a multi-stream model in which there are multiple independent (but synchronized) bitstreams.
It was commented that from an implementation perspective it is more complicated than a matter of memory capacities and resolution-dependent frame rates.
It was commented that this could involve a conformance point that constrains a combined set of decoders or decoder resources or bitstreams.
Plenary meeting Thursday 17 January 0945-1200
Track A review:
Reference picture management in high-level syntax per JVET-M0128 (modified as noted)
Small bug fixes from JVET-M0265
Fix the software to match the WD to remove clipping of luma MVs before deriving chroma MVs.
Adopt rounding away from zero for MV averages to remove inconsistency for similar averages: Offset = 1 << (F-1); M= S >= 0 ? (S + Offset) >> F: -((-S + Offset) >> F).
Flexible rectangular tile groups per JVET-M0853-v2 (constrained as noted), with software in JVET-M0445 and with loop_filter_across_tile_groups_enabled_flag from JVET-M0160 (confirmed in plenary)
High-level syntax actions noted in discussion of JVET-M0816 BoG report.
Simplification of division operation used in CCLM modelling from JVET-M0064.
Simplifying PDPC linear interpolation to use nearest neighbour on secondary boundary for adjacent angular modes
Bug fix in spec text related to CBF signalling identified in JVET-M0361.
Reduce complexity of 32-length DST-7/DCT-8 using zero-out approach of JVET-M0297 Test 2.
Enable transform skip up to 32x32 block size, with associated syntax approach in JVET-M0464 using tu_mts_idx (substantial gain for Class F / SCC).
Bug fix for quantization group QP signalling to make the size consistent per JVET-M0113 & JVET-M0188 (text in JVET-M0113)
Bug fix for transform skip quantization scaling factor for rectangular block shapes from JVET-M0119
Bug fix for QP with parallel encoding – initialize QP from the bottom left CU of the above CTU row when decoding the first CU of a CTU on the left edge of a tile (text in a revision JVET-M0685).
Non-normative (CTC): Enable transform skip for block sizes up to 32x32 in CTC (no effect on encoder runtime with JVET-M0464 encoder search modifications)
Non-normative: Adopt JVET-M0864 memory bandwidth analysis method.
The reverse coding order part of CE3-1.1.1 intra sub-partitions coding mode and text was discussed in the plenary. It was reported that there was only a 0.04% penalty for not doing the reverse coding order, so it was agreed to adopt the proposed scheme without that aspect; text was made available in a revision of JVET-M0102. Further study was suggested for limiting the sub-partition width to be greater than or equal to 4.
Track A action item: Avoiding 32-point DST (with 64-length DCT2 on the other side) in CE6-4.1a: Only a 0.01% penalty was reported in the plenary Thursday, and avoiding all DST combined with 64-length DCT2 has 0.02% penalty, so no DST (including size 32, 16, 8 and 4) combined with 64-length DCT2.
Track A action item: CE12-2 in-loop remapping function (adoption action likely) – experiment results were discussed in a plenary Thursday 17 January; there was no significant penalty for the additional restrictions.
The encoder algorithm was discussed. It was described in JVET-M0427.
At a previous meeting, a curve-crossing problem had been observed and it had been suggested to do something about low-QP operation. There was said to be a very large amount of code in the encoder optimization, with resolution dependency and a smoothness measure and various thresholds and checks. Some of that code was reportedly related to a different variant (CE12-1) and can be removed. Some of it was for HDR, which was not measured in this test (but is also in-scope for VVC). It was discussed whether the code would be difficult to maintain and might have excessive tuning within the code. It was commented that the complexity had been reduced from previous versions. This has been tested in multiple rounds of CE and appears to provide significant gain if an adequate encoding method is use.
Decision: Adopt (modified as noted).
Further study was requested to study the encoder software and check the behaviour outside of the tested conditions.
Two other Track A action items were left open.
Track A also had recommended to discuss enabling CPR in CTC, at least for Class F.
Track B review:
CE related BoGs (CEs 2, 4, 9, 10) had been reviewed, CE8 & CE11 related had been reviewed in track
CE11 viewing was ready, not reviewed, so there was no conclusion yet.
The BoG on NN technology was not reviewed yet.
Various revisits were still open pending on availability of more information. Most relevant were on CPR/IBC, deblocking, diffusion filters
New decisions:
Adopt DMVR JVET-M0147 with SAD cost function&more details somewhere else (approx. 0.9%, 15% decoder runtime);
Adopt combined merge list for adjacent 4x4 subblocks, JVET-M0170
Adopt variant of CE4.4.3, symmetric MVD, and disable BDOF when used – gives 0.33% in RA, increases encoder time by 5%, JVET-M0444
MMVD with switch (tile group header) to integer distance, benefical for screen content and UHD, JVET-M0255
Adopt the approach of not signalling the triangular prediction mode flag in cases where the combination is not allowed (MMVD, CIIP) – various contributions on that, gives 0.07%
Encoder RD optimization with deblocking knowledge shows 0.58%, 0.71% and 0.66% luma gain with similar encoding and decoding time, in AI, RA and LDB configuration respectively over VTM-3.0 anchor (SW adoption, not CTC unless HM would do the same). Could be used in some CEs as additional option, or mandatory.
Hash-based motion search (JVET-M0253) – provides 7.8%/14.9% for classes F/TGM in RA, version that does not affect encoder run time for natural video, switches back to conventional ME – CTC or CTC only for SC. It was suggested in the plenary to go with the solution of enabling CPR via the SPS flag specifically for class F in CTC, and also manually enabling hash-based search for class F in CTC, but have the automatic switching as non-CTC in SW.
Various adoptions of cleanups, harmonizations, simplifications with minor impact (look under BoG report)
JVET-M0063 Non-CE9: An improvement of BDOF
Generalization of BDOF bit-depth restriction for internal bit-depths other than 10 bit.
No impact on CTC.
8-bit coding scenario: -0.32/-0.32/-0.31% change in BD-rate (Y/CB/CR)
12-bit coding scenario: -0.46/-0.03/0.08% change in BD-rate (Y/CB/CR)
From discussion in Track B: A possible reason for this behaviour might be the wrong interpretation of the gradient in case of other bit depths than 10.
Question if the change would still be supporting the 16 bit SIMD design of software? Proponent confirms that this is the case, may need further checking by SW coordinators.
Decision (BF): Adopt JVET-M0063.
For plenary: What is general support of different bit depths in VVC? Might other tools that were added in recent meetings have similar problems. Definitely, bit depths up to 12 bits should be supported consistently, whereas it is likely that for higher bit depths some more precision might be required.
(Also, flexibility of spec in terms of other extensions, e.g. 4:4:4 would be desirable.)
An action item for editors was given to identify potential actions.
There had already been some detailed review and suggestions of technology to be investigated in ongoing CEs.
There was an intent to start a new CE on NN technology, with the primary intent to get a better understanding of adaptation mechanisms and complexity/compression performance impact, studying various methods that have been proposed with unified conditions and constraints.
Some aspects for ALF (particularly for saving line buffers) were planned to be investigated in CE; this also requires subjective inspection.
It was planned to restart CE investigations on post reconstruction filters (bilateral, Hadamard) but only using for inter. These give around 0.4% bit rate reduction for RA, which is still approximately the same as it was over VTM 2, so the gain seems to be additive with in VTM3. Both methods have quite some impact on complexity, where the new version of the bilateral filter is somewhat reduced relative to the previous version. This could, however, be conflicting with LIC in terms of latency, the latter is also under further investigation with additional constraints, these should be studied in combination.
Closing plenary sessions
In the closing plenary, the meeting notes were scanned for open items, AHG planning and CE planning were conducted, CE plan descriptions were reviewed in preliminary form, future workplans were discussed, and output document plans were established.
BoGs (12)