JVET-O0366 Non-CE4: Simplifications on BCW index derivation process [N. Park, J. Nam, H. Jang, J. Lim, S. Kim (LGE)]
JVET-O0366 (method 2): simplify BCW index derivation by setting it to the BCW index of the first candidate spatial neighbour block in the process of constructing an affine merge candidate.
JVET-O0681 (method 2) Disabled DMVR, BDOF and BCW for CIIP + inherit BCW index
Decision: The range of MVD shall be extended to [−217, 217−1] (no specific document was identified regarding this)
CE5: Loop filtering
Deblocking: Extensive subjective testing
CE5-2 Align deblocking with 4x4 grid
No need in case of chroma
Luma shows visual improvement, complexity evaluation of 2 different methods ongoing (one of them seems slightly better than the other concerning the number of cases of improvement over anchor)
CE5-3 MV difference threshold reduced to ½ pel
Visual benefit proven, adopt 0061
CE5-4 Adaptive loop filter
5-4.2, alternative chroma filters with CTU selection (8 filters), >2% chroma gain in RA and LD modes, adopt O0090
From CE related Deblocking
Adopt JVET-O0159 Deblocking tC table for 10-bit video with the modification such that the downscaling operation would exactly reproduce the current 8 bit table.
From CE related ALF
Various simplifications in APS signalling:
Coding of coefficients without prediction, without position dependency, only EG3
Coding of clipping value for nonlinear ALF by fixed length instead adaptive EGk
Remove floating point style expression of clipping value determination
On ALF APS
Multiple filters for chroma - switching between several chroma filters is desirable (results of CE5 indicate that up to 8 are beneficial)
Following options:
Stay with current APS (luma set + 1 chroma), but allow several APS for chroma in slice and select at CTU which one is best for luma and which one is best for chroma
Extend current APS such that several chroma filters are in, stay with 1 APS selection for chroma at slice, use direct indexing of selected chroma at CTU level rather than APS indexing (solution as suggested in CE5-4)
Have separate APS for chroma, either 1 per filter, or 1 with multiple filters
The most flexible solution would be 2., to have an APS that contains 0…25 luma filters and 0..8(or more?) for chroma. This would allow high flexibility in usage, including the case where an APS just carries a collection of chroma filters.
Decision from plenary: Agreed that the concept 2. mentioned above should be used. This is reflected in the adoption of JVET-O0090 from CE5-4, but this needs to be integrated with other changes of ALF APS.
Storage of APS parameters for local decoding
Results e.g. JVET-O0425 indicate that up to 25 filters may not always be needed.
Reasonable constraints for number per picture need to be identified.
One possibility could be to constrain the amount of memory needed for decoded APS parameters per picture. This aspect was discussed in plenary, and agreed to be further investigated.
Filtering at region boundaries
The current virtual boundary concept of ALF (adopted last meeting) only applies at horizontal boundaries. Slice/tile/brick boundaries and subpicture boundaries can also be vertical but always coincide with CTU boundaries.
When filtering is disabled across the boundaries (regardless if they are horizontal or vertical): The deblocking is turned off at the boundary, whereas SAO and ALF still are operated but require some padding. For SAO, a simple repetition of boundary is sufficient. The question is which padding method to use for the ALF within the current slice/tile/brick.
At the horizontal boundary (except picture boundary), the virtual boundary concept is applied.
At the vertical boundary, something must be defined. Track B recommends using the virtual boundary concept for that as well. As the content could be continuous across the vertical boundary, the simple padding method might cause some artefacts. Complexity-wise, extending the VB concept to vertical boundaries that shall not be filtered across boundary is not critical.
Further, the current spec disables the virtual boundary concept at the lower CTU boundary if that is a slice/tile/brick boundary which is inappropriate and should be changed.
At picture boundary, the sample repetition should be retained. This is a simple coordinate clipping.
In the plenary, it is agreed that the best solution for slice/tile/brick/subpicture boundaries is applying the VB concept of ALF, including vertical and lower CTU boundaries. Decision which relates to an adoption of JVET-O0662, and JVET-O0625.
It is noted that the term “virtual boundary” was introduced in the text spec also for something related to 360° video. A better term should be found for this.
In 360° video, face boundaries could be at any multiple of eight, and not aligned with CTU boundaries. This would require a mechanism where any filtering is disabled at those positions, and for ALF invoke a simple padding mechanism. It was agreed in the plenary that this should be the same repetitive padding as used at the picture boundary.
Decision of plenary: For boundaries of 360 faces, allow disabling filtering, and use the same repetitive padding as for picture boundary. Perhaps call that “virtual picture boundary” to differentiate from the above case.
CE8: Screen content coding tools
On IBC BV coding:
Same coding for MV and BV coding? It is reported that the maximum gain that is currently known when designing BVD coding different from MVD coding is around 1.3% for TGM/AI, and 0.5% for class F/AI. This would be more complex, increasing number of context coded bins.
It is agreed that such gain is not substantial enough to justify deviating in BVD coding from MVD coding, or defining different MVD coding just for the benefit of screen content, and further decreasing the CABAC throughput by increasing number of context coded bins.
On BV validity check:
Many opinions (and contributions) on the importance of encoder and decoder side availability checks, regarding the reference block addressed by BV
Interesting new ideas (e.g. O0127) allow doing it with minimum effort and safely on either side. Most appropriate solution would be a simple definition of bitstream constraint, without making it mandatory to check for invalid bitstreams at decoder (also in the interest of keeping decoder conformance definition simple)
Current restrictions seem over-engineered – additional gain reported when all blocks previously decoded can be addressed
Other aspects of SCC:
Agreed that IBC for dual tree should be disabled – minimum loss, and BV derivation for chroma was complicated in that case
For 4:4:4, single tree provides large gain (6% over dual tree in TGM)
Palette:
Palette mode shows interesting gain for 4:4:4 (tested so far with YUV), and here it provides in particular more gain for the single-tree case, which is likely due to less overhead and ability to utilize the redundancy between the components in a single palette
For 4:2:0, the gain is not large enough (even with the additional tools of CE8-2.2/3/5) to justify a complete different sample coding at CU level. It is noted that the gain of the baseline has even become lower in VTM5. This could be due to the adoption of TS. On the other hand, TS is not applied to chroma so far. Though a CE proposal for TS Chroma did not show benefit for 4:2:0 under CTC, it could better perform for 4:4:4-
Investigation of palette on 4:2:0 should be discontinued in CE
It is suggested to include a palette mode in VVC, which shall only be invoked for coding of 4:4:4 content. This should be the well-understood “base palette mode” from CE8-2.1, which can be used for both single tree and dual tree cases. Further investigation should be performed in CE to verify the gains with more 4:4:4 material, and further study improvements or simplifications. There is no need to stick to the HEVC palette method.
Decision of plenary: Adopt JVET-O0119, the CE8-2.1 “base palette” to VVC in 4:4:4 modes/profiles.
On BDPCM/TS/lossless
BDPCM flag at high level
BDPCM only enabled when TS is enabled
When BDPCM is switched on locally, TS is inferred
Intra prediction mode (h/v) alignment with BDPCM
Some inconsistency with deblocking strength, quantization, etc. -> new AHG
CE9: Decoder motion vector derivation
CE9-1.2 / JVET-O0055 Use integer-distance DMVR cost for BDOF early termination & no early BDOF termination when DMVR disabled (no loss)
CE9-2.2 / JVET-O0108: Disable DMVR & BDOF when CIIP is used (no loss)
During the plenary discussion, proponents of CE9-2.1b again requested adoption of their method, which had been disagreed in track B. However, no agreement was reached in plenary.
From CE related:
Slice level flags to disable DMVR and BDOF (O0504)
Further simplification of DMVR early termination (O0590)
CE10: Neural network based loop filtering
Probably too early for normative action
Findings from CE:
Interesting reduction of complexity, but the performance also drops then
Generally most gain for intra / real in-loop training difficult
Subjective viewing to be conducted
Might better continue as AHG only, no CE
Other coding tools
Adopt JVET-O0525 – Remove PCM mode – agreed, as PCM mode is hardly used, and better concepts such as TS+ for lossless.
Adopt JVET-O0526 – minimum CTU size 32x32 - agreed
Also more discussion on VPDU concepts (in context of CE2 et al.): The latency of luma/chroma should not exceed one quarter of the VPDU size. The VPDU size would be min (CTU size, 64x64).
Adopt JVET-O0537 Weighted intra and inter prediction mode
Pipelining in PDCP/CIIP OK with new concept
0.13% in RA, approx. 0.4% in LD modes, 1% encoder runtime
No consensus was reached in plenary, and this decision was reverted.
This was again brought up in the Friday morning plenary. As the tool is not a simplification, but a modification of CIIP targeting compression benefit and the additional BD gain for RA is low, no action was taken.
Adopt JVET-O0549 Encoder-only GOP-based temporal filter
The method was reportedly tested under VTM-5.0 RA, LDB and LDP common test configurations. The average Y/U/V BD-rates for the common test conditions (CTC) are reported to be −3.49%/−6.96%/−6.53% (RA), −1.00%/−1.24%/−2.24% (LDB) and −1.47%/−1.95%/−2.78% (LDP). All BDR numbers were computed using unfiltered source sequences. The method is not proposed for AI encoding.
The general status of track B was then presented and discussed.
Performance assessment contributions were discussed, except adaptive resolution conversion (section 4.3) 1730-1940.
Miscellany were noted.
Signal processing aspects of
AHG12: high-level parallelism & coded picture regions
AHG8: layered coding & resolution adaptation
Scalability
360 degree video (4) (section 6.16) (Track A) – covered in HLS BoG
Project development (section 4): text & software (1), test conditions (1), performance assessment (5)
Scheduling
Docs waiting for approval
Plenary meeting Thursday 11 July 1530
Discussion of ALF interaction with raster slices – example shown in figure below:
Decision: It was agreed that picture dimensions are to be required to be a multiple of 8 – OK.
Regarding the plan for the next meeting:
The previous schedule plan was to meet Tue 1 Oct – Wed 9 Oct. However, it was noted that SG16 had shifted its planned meeting dates and the volume of material for review has been exceptionally large.
It was agreed to start the meeting the afternoon of 1 Oct.
The agenda of 1 & 2 October was planned to be only HLS, layered coding, resolution adaptivity, GDR, and loop filtering. (It was noted that a JCT-VC meeting starts on Friday 4 October, and there is a Saturday systems coordination meeting planned.)
It was announced that a workshop on "the future of media" is planned on 8 October in preparation for World Standards Day.
It was planned to end the next meeting by 1400 on 11 October.
It was agreed that the config files for HDR common test conditions should use chroma location type 2 for PQ content. Further study was recommended regarding the HLG chroma location.
Closing plenary meeting Friday 12 July
In the closing plenary, there was a general review of open issues, AHG and CE plans, output documents, and future meeting plans. See especially sections 11 to 14 of this report.
BoGs (10)