JVET-W0182 BoG Report: Neural Networks Video Coding Analysis and Planning [A. Segall]
The initial progress in this BoG was reported verbally on Monday 12 July at 0830 UTC, with a follow up discussion presenting v2 of the report in session 16 on Tuesday 13 July at 0630 UTC.
The BoG held the following meeting sessions during the 23rd JVET meeting:
- July 9 – 23:20 – July 10 01:20
- July 12 – 23:20 – July 13 01:20
- July 13 – 21:00 – 23:15
Initial recommendations from the BoG are summarized as:
- Review the summary analysis prepared by the BoG
- Include an indication of current hardware performance when summarizing performance in future EE reporting.
- AhG mandate to study and collect information related to near-term and long-term architectures for neural-network video coding.
- Encourage members of hardware companies to comment.
The BoG was established with the following mandates:
- Perform further analysis about the complexity/compression trade-offs of the various loop filter and super resolution proposals and align the reports with AHG11 reporting conditions
- Perform further analysis between luma and chroma gains in detail
- Suggest candidates for an upcoming EE
- Refine the AHG11 reporting conditions as necessary
The following questions were proposed to be considered in performing the analysis:
- How should we align test results that did not use the AHG11 reporting conditions?
- Used VTM-11.0
- Used a different QP range
- Results were incomplete
- How should we summarize the different luma and chroma performances that are being observed?
- One potential solution – use a weighted combination of the Y, Cb and Cr BD-rates. For example, AHG13 previously used 6:1:1
The BoG agreed that these were the main questions to be answered.
One participant commented the use of VTM-11.0 as an anchor would not affect the all-intra results. Additionally, it was commented that the random-access performance would not be affected for all-intra tools in this case. A second participant agreed with the above.
After subsequent discussion (below), the group decided to remove the random-access results for proposals that used a VTM-11.0 anchor for reporting the summary results.
One participant commented that JVET-W0062 used VTM 12.0 with ALF and SAO disabled as an anchor. It was decided to re-compute the BD-Rate performance compared to the VTM-11.0-nnvc anchor.
One participant suggested to compute the BD-Rate performance of all proposals over the range of QP27, 32, 37, 42. This would just be for 4K only results, as all of the other results used a QP range of QP22, 27, 32, 37, 42.
One participant suggested to prepare an additional plot that attempts to compare the super-resolution and in-loop filtering technologies.
One participant suggested that super-resolution method should be required to bring all CTC results going forward. It was further suggested that the performance of super-resolution methods should be considered after in-loop filters.
The following plan was put forward for reporting:
- Create a summary for all tools using the NNVC CTC. This would exclude proposals that did not provide results for all sequences and QP points in the CTC.
- Create a summary of all tools for 4K using the NNVC CTC but with a QP range of QP27, 32, 37, 42. This could then include more of the proposals, as some of the super-resolution methods did not report results for QP22.
- Create a summary of super-resolution tools for 4K using CTC but with a QP range of QP27, 32, 37, 42, 47
The group decided on the above strawman approach.
The group decided to remove proposals from the summary that had incomplete results.
One participant recommended that the group should inspect the chroma BD-rate curves and identify cases where they were crossing. And, if the curves are crossing – the chroma results should be ignored.
Multiple participants expressed concern on reporting a metric that was averaged over the luma and chroma channels. It was proposed to report the luma and chroma results separately.
Multiple participants expressed support for reporting a combined metric to the group.
The group decided to report set of luma results and a set of chroma results, where the chroma results would be the average of the Cb/Cr channels. In addition, the group will check for cases where the chroma curves are crossing – and create a note in these cases that the results are unreliable.
One participant raised a question about complexity reporting in the case that a proposal used multiple models that are selectively enabled. The question was if the summary should use the worst-case performance or a different metric.
One participant expressed that worst-case performance should be reported.
Multiple participants expressed support for reporting the decoding complexity and average complexity in terms of run-time.
One participant expressed support for capturing more detailed information on complexity. This could include latency.
One participant recommended that the summary output of the BoG should also capture memory size.
One participant recommended that the summary capture if MAC operations use a floating point or integer operation.
Multiple participants recommended that the group study what is a realistic complexity and architecture for the near-term and long-term.
The group recommended that this should be a mandate of the NNVC AhG group. And, that the discussion should be incorporated into the EE design as much as possible.
The group decided to report worst-case complexity in the summary.
The group decided to capture the precision of the calculations in the summary.
The group decided to capture the memory size (number of parameters * the precision of the parameters) of all models being used in the proposals.
Next step – off-line activity to begin preparing the aligned summary results. Following the activity, it was agreed to review and revisit the summary plan.
The plan and summary results and were revised in the 12 July meeting session.
The BoG report also includes a detailed analysis of the different proposals in an attached Excel spreadsheet (including plots on number of operations, amount of memory vs. compression gain, and chroma vs. luma). An example is given below (it is noted that this is just meant as an example from the initial presentation, and may have been subject to further changes in the follow-up analysis of the BoG).
The benefit in terms of BD gain seems to increase roughly linearly with the number of operations. The goal would be to shift the points more to left/top (a better tradeoff in complexity vs. compression).
The chroma analysis for superresolution seems to indicate that the losses are quite non-uniform, e.g. most contributions seem to lose most in park running.
The following issues were proposed to be considered when developing the EE test plans:
- Should the proposals in the EE be in integer precision?
- What is the realistic complexity that should be targeted by the EE?
- What is the near-term and long-term architecture that should be targeted by the EE?
- How should the group work toward a common software?
- What tests should be performed?
- What technologies should be included in the test?
- What is the timeline for the EE?
The group agreed that these were the main topics to be discussed.
One participant suggested that proposals in the EE should be implemented in integer precision.
Multiple participants agreed with the above the statement.
Multiple participants expressed concerns about making integer precision mandatory in the next EE cycle.
It was commented that porting a floating-point implementation to a 32-bit integer implementation may be straightforward.
One participant suggested that the memory size of the proposals should be reported and prioritized to encourage participants to transition to integer implementations.
One participant suggested that the group identify the target resolution and frame rate of the activity.
Multiple participants recommended that UHD@60p and 8K@60p should be key targets.
A summary of some capabilities of available devices was provide in JVET-W0131, with all numbers reportedly based on public information:
- Snapdragon 865: peak 15 trillion operation per second (TOPS)
- Snapdragon 888: peak 26 TOPS
- A14 Bionic: 11 TOPS
- M1X: may manage around 5.2 TFLOPS.
- 2080 TI (Single GPU core): around 11 TFLOPS,
- GeForce RTX 3080(Single GPU core): up to 30 TFLOPS
It was commented that one could expect that device performance will increase in the future. But, at the same time, it would be unlikely that the entire device capability would be available for decoding.
One participant commented that it could be beneficial to capture and plot the device performance as part of result reporting. Multiple participants agreed with this idea.
One participant proposed to capture the performance for UHD@60p and a second for 8K@60p.
The group recommended to include an indication of current hardware performance when summarizing performance in future EE reporting.
One participant introduced a paper that would be relevant for this work: Sun et al., “Summarizing CPU and GPU Design Trends with Product Data”.
One participant suggested that the memory size of current devices should also be captured and plotted when summarizing performance in future EE reporting. Here, the memory size would be the number of bytes needed to store the parameters.
One participant commented that it would be beneficial to discuss with other hardware companies if general GPU/TPU resources can be used in their decoding pipeline.
One participant commented that it was unlikely in their architecture to use GPU/TPU resources if the coding tool is in the loop.
One participant commented that technologies that are completely pre- and post-processing solutions could use GPU/TPU resources.
The group recommended to request an AhG to study and collect information related to near-term and long-term architectures and suggested that members of hardware companies be particularly encouraged to comment.
It was commented that JVET-W0181 was related to this topic. One participant suggested that having a technology that used the JVET-W0181 framework in the EE could be beneficial.
One participant suggested that the common software should be limited to inference.
One participant commented that common software should help with cross checking activities. But, a first issue in performing cross checks is the lack of integer implementations.
The BoG was planning to meet again to refine the analysis and continue on defining next EE, including refinements of reporting and CTC.
Further BoG reporting was conducted on Wednesday at 0830 (JRO & GJS).
The BoG had met further. Some complexity estimates had been refined. Current hardware performance reporting was suggested by the BoG for further work; also reporting to require reports to use uniform reporting. EE planning with cross-checking using a two-step cross-checking had been conducted. There was expected to be more emphasis on cross-checking in the future. In-loop versus post-processing comparative result reporting was planned. About 10 proposals were planned to be tested in the next round of EE.
Recommendations from the BoG were summarized in the BoG report as:
- Summary analysis
- Review the summary analysis prepared by the BoG
- Future reporting
- Include an indication of current hardware performance when summarizing NNVC performance in future EE reporting.
- Require all EE and future non-EE proposals to use the AHG11 test conditions and reporting template.
- Require all EE and future non-EE proposals to report complexity information using the reporting template.
- EE Planning
- Use the following cross-checking process for the NNVC going forward
- Initial cross-check is performed on the inference stage.
- If the technology is considered for “adoption”, then the proponent would provide the necessary scripts/information that was used for training.
- The training step would be cross-checked at that point to confirm that the training can be reproduced. It is anticipated that the training step may not be a bit-exact match and instead may require using some threshold/tolerance for acceptance.
- Cross checks should be highly encouraged for the EE tests. It was recommended that the cross-check reports include information on problems that were encountered when trying to perform the cross check and possible solutions.
- Loop filter and super resolution technologies should be encouraged to provide information as their performance as a post filter.
- Include JVET-W0057, JVET-W0081, JVET-W0099, JVET-W0100, JVET-W0105, JVET-W0113, JVET-W0131, JVET-W0132, and JVET-W0151 in the EE.
- Use the following cross-checking process for the NNVC going forward
- Testing conditions
- All EE and future non-EE proposals should be required to use the AHG11 test conditions and reporting template.
- Clarify that the number of MACs requested in the reporting conditions corresponds to the worst-case complexity.
- Clarify the reporting conditions to capture both the number of parameters for each model as well as the total memory of all parameters used by the proposal.
- Create two columns for the CPU run-time and GPU run-time. And, to label the CPU run-time as mandatory.
- AHG Mandates
- Study and collect information related to near-term and long-term architectures for neural-network video coding. Encourage members of hardware companies to comment.
The recommendations from the BoG (as given above) were reviewed and approved.
EE1 contributions: Neural network-based video coding (7)
Contributions in this area were discussed in session 8 at 0015–0120 UTC on Thursday 8 July 2021 (chaired by JRO).
JVET-W0182 BoG Report: Neural Networks Video Coding Analysis and Planning [A. Segall]
See the notes for this BoG report in section 5.2.1. The recommendations from the BoG were reviewed and approved.