JVET-V0112 AHG9: On Bitstream Properties Signalling for Decoder Initialization [S. Deshpande (Sharp)]
This contribution was discussed in sesssion 18b at 1530 on Monday 26 April 2021 (chaired by GJS).
It is proposed to signal picture storage and picture format related information about CVSs in the bitstream to help the decoder initialization.
It is proposed to signal picture storage and picture format related information about CVSs in the bitstream to help the decoder initialization. In this document this information is called bitstream properties information.
It is asserted that signalling information which provides maximum required decoder resources for decoding a bitstream allows a decoder to allocate enough resources (e.g. memory/ DPB space) which can be used to decode the entire bitstream. Thus, need to deallocate and reallocate memory when decoding each CVS of a bitstream may be avoided. This can also result in lowering the decoder delay when switching decoding from one CVS to another CVS, including in streaming environments. Additionally, this may result in requiring allocation of less memory compared to the maximum required from the signalled profile-tier-level. JVET-U0083 also provides a summary of issues related to decoder initialization and additionally refers to the related discussions in systems, including in Systems for Video AHG.
Two alternative options are proposed in this document for helping decoder initialization:
- Option 1: It is proposed to signal following information in decoding capability information (DCI) (or in a new SEI):
- Maximum number of luma samples for a picture
- Maximum number of (luma and chroma) samples for a picture
- Maximum number of luma samples multiplied by bit depth for a picture
- Maximum number of (luma and chroma) samples multiplied by bit depth for a picture
- Maximum value for number of luma samples for a picture multiplied by maximum required DPB size
- Maximum value for number of (luma and chroma) samples for a picture multiplied by maximum required DPB size
- Maximum value for number of luma samples for a picture multiplied by maximum required DPB size multiplied by bit depth
- Maximum value for number of (luma and chroma) samples for a picture multiplied by maximum required DPB size multiplied by bit depth
- Maximum picture width
- Maximum picture height
- Maximum value of required DPB size
- Maximum bit depth (minus 8)
- Maximum chroma format
- Option 2: A list of unique sets of storage and format information is specified from all the OLSs in the bitstream.
Relative to option 2, JVET-V0081 also includes level information.
There was discussion of JVET-V0112 with JVET-V0081 in joint discussion Tuesday at 1315:
- Does additional decoder initialization information of some sort (beyond PTL, VPS, DCI, subprofiles, general constraint flags, as they exist) need to be specified?
- If so, should that be carried in the bitstream (e.g., in DCI or SEI)?
The main issue is said to be decoder configuration initialization, to ensure that decoders are able to allocate sufficient resources without a reconfiguration interruption. One participant said the existing syntax should be enough and adding more syntax features may be complicated and difficult to understand and might not be used by decoding systems, also saying that if such a thing is defined, it should be minimal.
A participant said decoders should not need to allocate maximum resources (e.g., due to picture size); that if resources less than necessary for the maximum capability of the level are sufficient, then having such information provided could save resources.
It was questioned whether the amount of resource savings would really be sufficient to justify trying to add more detailed resource descriptions into the bitstream.
Further study is needed to reach agreement on the need; such study is needed among both systems and JVET experts to determine the potential requirement.
Plenary meetings, joint meetings, BoG reports, and liaison communications
JVET plenaries
Some of the discussions and actions at plenary sessions are noted in this section (especially those of Monday 26 April 1300–1500):
- Planning of output documents: Standard parts & DoCs, CE & EE descriptions, verification test plan & report?
- Summary of voting for ISO/IEC 23091-2: m56377 – GJS was to prepare a candidate DoCR
- Liaison output to JPEG about progress in NN video coding – Gary Sullivan and Elena Alshina were asked to coordinate preparation of text to be sent out via SG 16.
- White paper on VVC – it was concluded that it is necessary to delay this to the next meeting, and encouraged to work until then on preparing something. It was mentioned that various IEEE papers have been worked on, including a special issue in TCSVT. These will be referenced in the white paper, anyway. The JVET chairs were given the action item to send links about such information to the JVET email reflector. This was followed up after the meeting with a message to the reflector on 24 May, pointing to tutorial information and other resources newly available on the JVET page of the ITU-T website (https://www.itu.int/en/ITU-T/studygroups/2017-2020/16/Pages/video/jvet.aspx). It was also mentioned that the MC-IF (www.mc-if.org) and HHI web pages also maintain lists of references (https://jvet.hhi.fraunhofer.de/).
- Whether to hold a hybrid meeting in October 2021 An informal poll gave 22% who would consider travelling to participate physically, 50% who would not, and 28% who were undecided.
Information sharing meetings
In addition to the joint meetings listed below, information sharing sessions with other WGs of the MPEG community were held on Monday 26 April 0500–0700 and Wednesday 28 April 0500–0700. The status of the work in the MPEG WGs was reviewed at these information sharing sessions. Additionally, the JVET status was also presented by the JVET chairs at another MPEG information sharing session on Friday 30 April 2100–2300 (after the JVET meeting had closed).
Joint meeting with Q6/16 (VCEG), WG2 MPEG Requirements and WG3 MPEG Systems 1300–1400 Tuesday 27 April
The following topics were discussed in this joint session. See also the notes recorded on these topics in other sections of this document.
- Decoder initialization information for systems (JVET-V0081, JVET-V0112)
- Output of partially decoded pictures (JVET-V0109)
- Possibility of extending and nesting SEI message (esp. post-filter hint JVET-V0058)
- Colour transform information SEI message (JVET-V0108) – this has a similar issue to some prior messages (e.g., FPA and CRI)
- Compositing related SEI message for video decoding interface (VDI) input formatting function
- No contribution to this meeting; potential study may be conducted for a future SEI message that would refer to another standard (similar as with green metadata). Study in the Systems context is needed.
- On other SEI messages – no particular concerns were expressed.
Joint meeting with AG5 MPEG Visual Quality Assessment 1400–1500 Tuesday 27 April
Topics of discussion:
- Visual quality metrics
- m56636 on AI-based visual quality metric; further exploration was encouraged as an AG5 activity
- JVET-V0062 AHG9: Picture quality metrics SEI
- VVC verification testing
- JVET-V0174 registered but not yet available at time of the joint discussion
Visual tests had been conducted over the past month for LD and RA SDR HD and 360° PERP and CMP (GCMP for VVC, PCMP for HM).
The test results only became available during the meeting but were presented.
The configuration for 3 of 4 test sequences had not matched what was planned, in that MCTF had accidentally not been enabled for the HM for these sequences. This was said to be unlikely to have a large effect. It was said that it should not be difficult to re-run these tests to correct this within a few days after the meeting.
It was agreed to do this and include the results as supplemental to the results previously measured.
The testing process had been difficult due to Covid-related restrictions on lab use.
Two labs, more than 50 naïve viewers over 9 days of testing for the HD RA testing. Three test sessions for LD (30, 50 and 60 fps), three for RA (60 fps).
Three test sessions for 360° video.
MOS graphs were shown.
VVenC was also tested in addition to the VTM, and it appeared to show very strong performance. The VTM used CTC settings.
For summary-level results, an RD Plot package available on GitHub (at https://github.com/IENT/RDPlot) was used for some of the summary calculations.
The test coordinators, test labs, and others who helped with bitstreams and tools were thanked for their efforts in preparing and conducting these tests.
Further testing was planned for HDR. Work on testing of scalability, screen content, and 4:4:4 was also desired.
BoGs (0)
No break-out groups were established at this meeting; thus all notes of the meeting discussions were recorded directly in drafts of this document rather than in break-out group reports.
Liaison communications
The JVET did not directly receive or send any liaison statements at its current meeting. However, there was some related liaison communication between ITU-T SG 16, the AGs and WGs of ISO/IEC JTC 1/SC 29, and ITU-T Study Group 12 that were coordinated with JVET.
This included exchange of general status information about JVET work and management arrangements for video coding collaboration between the parent bodies of WP 3/16 and SC 29. SC 29 had sent a liaison statement to SG16 [TD535/Gen / SC29 N 19023], and SG 16 prepared a reply.
ITU-T SG12 had sent a liaison statement m56363 to ISO/IEC JTC 1/SC 29/AG 5 (in reply to SC 29 N 19285, MDS 19755 from October 2020). SG12 expressed their strong interest in exchanging information and collaborating on aspects of video quality assessment. SG12 also welcomed the information received about the status of the verification tests of Versatile Video Coding (ITU-T H.266 and ISO/IEC 23090-3 VVC) conducted with SG16 in the Joint Video Experts Team (JVET). In light of possible future extensions of existing video quality models and the development of new assessment methods, SG12 expressed their particular interest in further exchanges about this and similar assessment campaigns. SC 29/AG 5 (Visual Quality Testing) prepared a reply liaison statement to ITU-T SG12 and ITU-R WP6C as document AG 05 N 23 that included a description of recent verification testing of VVC and further plans for such testing.
SC 29/WG 1 had recently sent several liaison statements to ITU-T SG16 [TD537/Gen / SC 29/WG 1 N 89052 (SC 29 N 19202), TD581/Gen / SC 29/WG 1 N 90081 (SC 29 N 19574), and TD600/Gen / SC 29/WG 1 N 91067] that included discussion of neural network image coding, and Gary Sullivan and Elena Alshina of JVET were given the action item to prepare status information about the neural network video coding exploration in JVET to be included in a liaison letter reply from ITU-T SG16 to SC 29/WG 1.
ITU-T SG13 had sent a liaison statement m56366 to SC 29 with an invitation to review an Artificial Intelligence Standardization Roadmap and provide missing or updated information. SC 29/AG 3 (Liaison and Communications) prepared a reply as document AG 3 N 28 that included a mention of the JVET exploration experiment work on neural network based video coding.
Project planning
Software timeline
VTM13.0 including the adoptions from JVET-V0047, JVET-V0054, JVET-V0056, JVET-V0106: 2021-05-21 (needed for CE).
HM16.24 including the adoption from JVET-V0056: 2021-05-21.
VTM13.1 as appropriate date t.b.d. with remaining adoptions of encoder optimization, SEI messages.
Core experiment and exploration experiment planning
A CE on entropy coding for high bit depths and high bit rates was established, as recorded in output document JVET-U2022.
An EE on neural network-based video coding was established, as recorded in output document JVET-U2023.
An EE on enhanced compression technology beyond VVC capability using techniques other than neural-network technology was also established, as recorded in output document JVET-U2024.
Initial versions of these documents were presented and approved in the plenary on Friday 15 January.
Drafting of specification text, encoder algorithm descriptions, and software
The following agreement has been established: the editorial team has the discretion to not integrate recorded adoptions for which the available text is grossly inadequate (and cannot be fixed with a reasonable degree of effort), if such a situation hypothetically arises. In such an event, the text would record the intent expressed by the committee without including a full integration of the available inadequate text.
Plans for improved efficiency and contribution consideration
The group considered it important to have the full design of proposals documented to enable proper study.
Adoptions need to be based on properly drafted working draft text (on normative elements) and HM/VTM encoder algorithm descriptions – relative to the existing drafts. Proposal contributions should also provide a software implementation (or at least such software should be made available for study and testing by other participants at the meeting, and software must be made available to cross-checkers in EEs).
Suggestions for future meetings included the following generally-supported principles:
- No review of normative contributions without draft specification text
- VTM algorithm description text is strongly encouraged for non-normative contributions
- Early upload deadline to enable substantial study prior to the meeting
- Using a clock timer to ensure efficient proposal presentations (5 min) and discussions
The document upload deadline for the next meeting was planned to be Tuesday 13 April 2021.
As general guidance, it was suggested to avoid usage of company names in document titles, software modules etc., and not to describe a technology by using a company name.
General issues for experiments
It was emphasized that those rules which had been set up or refined during the 12th JVET meeting should be observed. In particular, for some CEs of some previous meetings, results were available late, and some changes in the experimental setup had not been sufficiently discussed on the JVET reflector.
Group coordinated experiments have been planned as follows:
- “Core experiments” (CEs) are the coordinated experiments on coding tools which are deemed to be interesting but require more investigation and could potentially become part of a draft standard by the next meeting or in the near future.
- “Exploration experiments” (EEs) are also coordinated experiments. These are conducted on technology which is not foreseen to become part of a draft standard in near future. Investigating methodology for assessment of such technology can also be an important part of an EE. (Further general rules for EEs, as far as deviating from the CE rules below, should be discussed in a future meeting. For the current meeting, procedures as described in the EE description document are deemed to be sufficient)
- A CE is a test of a specific fully described technology in a specific agreed way. It is not a forum for thinking of new ideas (like an AHG). The CE coordinators are responsible for making sure that the CE description is complete and correct and has adequate detail. Reflector discussions about CE description clarity and other aspects of CE plans are encouraged.
- A description of each experiment is to be approved at the meeting at which the experiment plan is established. This should include the issues that were raised by other experts when the tool was presented, e.g., interference with other tools, contribution of different elements that are part of a package, etc. The experiment description document should provide the names of individual people, not just company names.
- Software for tools investigated in a CE will be provided in one or more separate branches of the software repository. Each CE will have a “fork” of the software, and within the CE there may be multiple branches established by the CE coordinator. The software coordinator will help coordinate the creation of these forks and branches and their naming. All JVET members will have read access to the CE software branches (using shared read-only credentials as described below).
- During the experiment, revisions of the experiment plans can be made, but not substantial changes to the proposed technology.
- The CE description must match the CE testing that is done. The CE description needs to be revised if there has been some change of plans.
- The CE summary report must describe any changes that were made in the process of finalizing the CE.
- By the next meeting it is expected that at least one independent cross-checker will report a detailed analysis of each proposed feature that has been tested and confirm that the implementation is correct. Commentary on the potential benefits and disadvantages of the proposed technology in cross-checking reports is highly encouraged. Having multiple cross-checking reports is also highly encouraged (especially if the cross-checking involves more than confirmation of correct test results). The reports of cross-checking activities may (and generally should) be integrated into the CE report rather than submitted as separate documents.
It is possible to define sub-experiments within particular CEs, for example designated as CEX.a, CEX.b, etc., where X is the basic CE number.
As a general rule, it was agreed that each CE should be run under the same testing conditions using one software codebase, which should be based on the group test model software codebase. An experiment is not to be established as a CE unless there is access given to the participants in (any part of) the CE to the software used to perform the experiments.
The general agreed common conditions for single-layer coding efficiency experiments for SDR video are described in the prior output document JVET-T2010.
Experiment descriptions should be written in a way such that it is understood as a JVET output document (written from an objective “third party perspective”, not a proponent perspective – e.g. not referring to methods as “improved”, “optimized”, etc.). The experiment descriptions should generally not express opinions or suggest conclusions – rather, they should just describe what technology will be tested, how it will be tested, who will participate, etc. Responsibilities for contributions to CE work should identify individuals in addition to company names.
CE descriptions contain a basic description of the technology under test, but should not contain excessively verbose descriptions of a technology (at least not unless the technology is not adequately documented elsewhere). Instead, the CE descriptions should refer to the relevant proposal contributions for any necessary further detail. However, the complete detail of what technology will be tested must be available – either in the CE description itself or in documents that are referenced in the CE description that are also available in the JVET document archive.
Any technology must have at least one cross-check partner to establish a CE – a single proponent is not enough. It is highly desirable have more than just one proponent and one cross-checker.
The CE development workflow is described at:
https://vcgit.hhi.fraunhofer.de/jvet/VVCSoftware_VTM/wikis/Core-experiment-development-workflow
CE read access is available using shared accounts: One account exists for MPEG members, which uses the usual MPEG account data. A second account exists for VCEG members with account information available in the TIES system at:
https://www.itu.int/ifa/t/2017/sg16/exchange/wp3/q06/vceg_account.txt
Some agreements relating to CE activities were established as follows:
- Only qualified JVET members can participate in a CE.
- Participation in a CE is possible without a commitment of submitting an input document to the next meeting. Participation is requested by contacting the CE coordinator.
- All software, results, and documents produced in the CE should be announced and made available to JVET in a timely manner.
- A JVET CE reflector will be established and announced on the main JVET reflector. Discussion of logistics arrangements, exchange of data, minor refinement of the test plans, and preparation of documents shall be conducted on the JVET CE reflector, with subject lines prefixed by “[CEx: ]”, where “x” is the number of the CE. All substantial communications about a CE other than such details shall take place on main JVET reflector. In the case that large amounts of data are to be distributed, it is recommended to send a link to the data rather than the data itself, or upload the data as an input contribution to the next meeting.
General timeline for CEs
T1= 3 weeks after the JVET meeting: To revise the CE description and refine questions to be answered. Questions should be discussed and agreed on JVET reflector. Any changes of planned tests after this time need to be announced and discussed on the JVET reflector. Initially assigned description numbers shall not be changed later. If a test is skipped, it is to be marked as “withdrawn”.
T2 = Test model software release + 2 weeks: Integration of all tools into a separate CE branch of the VTM is completed and announced to JVET reflector.
- Initial study by cross-checkers can begin.
- Proponents may continue to modify the software in this branch until T3.
- 3rd parties are encouraged to study and make contributions to the next meeting with proposed changes
T3: 3 weeks before the next JVET meeting or T2 + 1 week, whichever is later: Any changes to the CE test branches of the software must be frozen, so the cross-checkers can know exactly what they are cross-checking. A software version tag should be created at this time. The name of the cross-checkers and list of specific tests for each tool under study in the CE plan description shall be documented in an updated CE description by this time.
T4: Regular document deadline minus 1 week: CE contribution documents including specification text and complete test results shall be uploaded to the JVET document repository (particularly for proposals targeting to be promoted to the draft standard at the next meeting).
The CE summary reports shall be available by the regular contribution deadline. This shall include documentation about crosscheck of software, matching of CE description and confirmation of the appropriateness of the text change, as well as sufficient crosscheck results to create evidence about correctness (crosscheckers must send this information to the CE coordinator at least 3 days ahead of the document deadline). Furthermore, any deviations from the timelines above shall be documented. The numbers used in the summary report shall not be changed relative to the description document.
CE reports may contain additional information about tests of straightforward combinations of the identified technologies. Such supplemental testing needs to be clearly identified in the report if it was not part of the CE plan.
New branches may be created which combine two or more tools included in the CE document or the VTM (as applicable).
It is not necessary to formally name cross-checkers in the initial version of the CE description document. To adopt a proposed feature at the next meeting, we would like see comprehensive cross-checking done, with analysis that the description matches the software, and recommendation of value of the tool given tradeoffs.
The establishment of a CE does not indicate that a proposed technology is mature for adoption or that the testing conducted in the CE is fully adequate for assessing the merits of the technology, and a favourable outcome of CE does not indicate a need for adoption of the technology into a standard.
Availability of spec text is important to have a detailed understanding of the technology and also to judge what its impact on the complexity of the spec will be. There must also be sufficient time to study it in detail. CE contributions without sufficiently mature draft spec text in the CE input document should not be considered for adoption.
Lists of participants in CE documents should be pruned to include only the active participants. Read access to software will be available to all members.
Establishment of ad hoc groups
The ad hoc groups established to progress work on particular subject areas until the next meeting are described in the table below. The discussion list for all of these ad hoc groups was agreed to be the main JVET reflector (jvet@lists.rwth-aachen.de).
Review of AHG plans was conducted in session 25 on Wednesday 28 April 2021.
Title and Email Reflector | Chairs | Mtg |
Project Management (AHG1)
| J.-R. Ohm, G. J. Sullivan (co-chairs) | N |
Draft text and test model algorithm description editing (AHG2)
| B. Bross, J. Chen, C. Rosewarne (co-chairs), F. Bossen, J. Boyce, S. Kim, S. Liu, J.‑R. Ohm, G. J. Sullivan, A. Tourapis, Y.-K. Wang, Y. Ye (vice-chairs) | N |
Test model software development (AHG3)
| F. Bossen, X. Li, K. Sühring (co-chairs), K. Sharman, V. Seregin, A. Tourapis (vice‑chairs) | N |
Test material and visual assessment (AHG4)
| V. Baroncini, T. Suzuki, M. Wien (co-chairs), E. François, S. Liu, A. Norkin, A. Segall, P. Topiwala, S. Wenger, Y. Ye (vice-chairs) | Tel. 2 weeks notice |
Conformance testing (AHG5)
| J. Boyce and W. Wan (co-chairs), E. Alshina, F. Bossen, I. Moccagatta, K. Kawamura, K. Sühring, X. Xu (vice-chairs) | N |
360° video coding, software and test conditions (AHG6)
| J. Boyce and Y. He (co-chairs), K. Choi, Y. Ye (vice-chairs) | N |
Coding of HDR/WCG material (AHG7)
| A. Segall (chair), E. François, W. Husak, S. Iwamura, D. Rusanovskyy (vice-chairs) | N |
High bit depth, high bit rate, and high frame rate coding (AHG8)
| A. Browne and T. Ikai (co-chairs), D. Rusanovskyy, M. Sarwer, X. Xiu, Y. Yu (vice-chairs) | Tel. 2 weeks notice |
SEI message studies (AHG9)
| J. Boyce, S. McCarthy (co-chairs), C. Fogg, P. de Lagrange, A. Luthra, G. J. Sullivan, A. Tourapis, Y.-K. Wang, S. Wenger (vice-chairs) | N |
Encoding algorithm optimization (AHG10)
| A. Duenas, R. Sjöberg and A. Tourapis (co-chairs) | N |
Neural network-based video coding (AHG11)
| E. Alshina, S. Liu, A. Segall, (co‑chairs), J. Chen, F. Galpin, J. Pfaff, S. S. Wang, Z. Wang, M. Wien, P. Wu, J. Xu (vice‑chairs) | Tel. 2 weeks notice |
Enhanced compression beyond VVC capability (AHG12)
| M. Karczewicz, Y. Ye and L. Zhang (co-chairs), B. Bross, X. Li, K. Naser, H. Yang (vice chairs) | Tel. 2 weeks notice |
It was confirmed that the rules which can be found in document ISO/IEC JTC 1/SC 29/AG 2 N010 “Ad hoc group rules for MPEG AGs and WGs” (available at https://www.mpegstandards.org/adhoc/), are consistent with the operation mode of JVET AHGs. It is however pointed out that JVET does not allow separate AHG reflectors, such that any JVET member is implicitly a member of any AHG. This shall be mentioned in the related WG Recommendations. The list above was also issued as a separate WG 5 document (ISO/IEC JTC 1/SC 29/WG 5 N 45) in order to make it easy to reference.
Output documents
The following documents were agreed to be produced or endorsed as outputs of the meeting. Names recorded below indicate the editors responsible for the document production. Where applicable, dates of planned finalization and corresponding parent-body document numbers are also noted.
It was reminded that in cases where the JVET document is also made available as a WG 5 output document, a separate version under the WG 5 document header should be generated. This version should be sent to GJS and JRO for upload.
The list of JVET ad hoc groups was also issued as a WG 5 output document WG 5 N 45, as noted in section 9.