JVET-AJ0023 EE1: Summary report of exploration experiment on neural network-based video coding [E. Alshina, F. Galpin, Y. Li, D. Rusanovskyy, M. Santamaria, J. Ström, R. Chang, Z. Xien (EE coordinators)]
For tests competing with technologies in NNVC-10.0 it was agreed to configure the proposed solution targeting close to existing NNVC tool complexity, but not exceeding it:
- kMAC/pxl of EE1 test ≤ kMAC/pxl NNVC (must),
- Number of Parameters EE1 test ≤ Number of Parameters NNVC (if possible).
The minor violation of this requirement is allowed for the final model, if the reason is keeping the number of channels a multiple of 16 (or at least 8), since it is very helpful for neural network algorithms implementation on wide range of platforms. Reported results should follow the constraints for evaluation purpose but the final model might violate the constraints.
NN architecture provided in this test description should not be changed, beside minor adjustment for parameters (such as channels number) in order to meet recommendation above.
Exact parameters settings were announced by proponents by 2nd AhG11/14 teleconference on October 1.
The teleconferences are conducted in order to discuss the NNVC reference software integration and EE1 tests setup.
The table below provides performance and complexity number of several tools’ combinations in NNVC relatively to VVC anchor.
Performance of NNVC tools combinations relative to VVC
Test | Random Access | All Intra | Total kMAC/pxl | Total Param (Mprm) | ||||||||
Y | U | V | Enc | Dec | Y | U | V | Enc | Dec | |||
NNIntra+LOP | -7.4% | -13.6% | -11.6% | 1.2 | 34 | -8.5% | -14.3% | -14.0% | 1.7 | 22 | 21.7 | 1.5 |
NNIntra+HOP | -14.2% | -19.5% | -19.9% | 2.6 | 1441 | -13.6% | -15.9% | -17.1% | 2.7 | 767 | 471 | 2.7 |
NNIntra+VLOP | -5.6% | -7.6% | -6.4% | 1.1 | 15 | -7.2% | -9.4% | -8.8% | 1.7 | 13 | 10.0 | 1.4 |
NNIntra+LOP+SR | -8.1% | -10.3% | -8.9% | -9.2% | -11.5% | -11.0% | 26.4 | 1.5 | ||||
This round of EE1 tests includes:
- EE1-1: LOP and VLOP in-loop filter
- EE1-1.1 – Partial Convolution and Over-Parameterization
- Report: JVET-AJ0080, cross-check: JVET-AJ0150
- EE1-1.2 – LOP with residual groups and Attention
- Report:JVET-AJ0165, cross-check JVET-AJ0327
- EE1-1.3 - NN in-loop filters using early cropping
- Report: JVET-AJ0054, cross-check: JVET-AJ0133
- EE1-1.4 – Reduced complexity input feature extraction for LOP & VLOP
- Report: JVET-AJ0066 cross-check: JVET-AJ0163
- EE-1.5 - Multiscale blocks in LOP3 and VLOP2 filter
- Report: JVET-AJ0182 cross-check: JVET-AJ0206
- EE1-1.1 – Partial Convolution and Over-Parameterization
- EE1-2: HOP in-loop filter
- EE1-2.1 – Attention block with transformers
- Withdrawn
- EE1-2.2 – Block-size invariant implementation of HOP
- Report:JVET-AJ0166 cross-check: JVET-AJ0177
- EE1-2.3 – Block based QP information in NN loop filters
- Report:JVET-AJ0124 cross-check: JVET-AJ0324
- EE1-2.1 – Attention block with transformers
- EE1-3: NN-inter prediction
- E1-3.1 – RA/LDB Unified Reference Frame Synthesis for VVC Inter Coding
- Report: JVET-AJ0099 cross-check: Nokia
- E1-3.1 – RA/LDB Unified Reference Frame Synthesis for VVC Inter Coding
- EE1-4: NN-based super resolution
- EE1-4.1 – Wavelet transform for super-resolution loss function
- Report:JVET-AJ0056, cross-check: JVET-AJ0291
- EE1-4.1 – Wavelet transform for super-resolution loss function
Test results summary
Details of each test can be found in an attached presentation.
LOP and VLOP filter modifications
In NNVC-10 LOP and VLOP filters architecture are unified. Just number of channels and backbone blocks are different to match two different levels of complexity.
LOP and VLOP filters in NNVC-10
Proponents of tests EE1-1.3, EE1-1.4 and EE1-1.5 provided their solutions for both LOP and VLOP keeping filters unified. Tests EE1-1.1 and EE1-1.2 provided solution and test results only for LOP.
Tests EE1-1: LOP in-loop filter
Test | Random Access | All Intra | kMAC/pxl | Param,M | Source | |||||||||
Y | U | V | Y | U | V | Total | Filter | Intra | Total | Filter | Intra | |||
NNVC-LOP | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 21.7 | 16.9 | 4.8 | 1.5 | 0.21 | 1.3 | ||
EE1-1.1 (int) | -0.1% | 1.2% | 1.2% | -0.1% | 1.0% | 0.9% | 21.7 | 16.9 | 4.8 | 1.5 | 0.20 | 1.3 | ||
EE1-1.1 (fl) | 0.0% | 1.1% | 1.1% | -0.1% | 0.9% | 0.8% | 21.7 | 16.9 | 4.8 | 1.5 | 0.20 | 1.3 | ||
EE1-1.2 | -0.4% | 0.0% | 0.3% | -0.4% | -0.1% | 0.4% | 21.7 | 16.9 | 4.8 | 1.5 | 0.49 | 1.3 | ||
EE1-1.3 | -0.3% | -1.2% | -1.0% | -0.3% | -1.1% | -0.9% | 21.6 | 16.8* | 4.8 | 1.5 | 0.21 | 1.3 | ||
EE1-1.4.1 | 0.0% | 0.8% | 0.4% | 0.0% | 0.6% | 0.6% | 21.7 | 16.9 | 4.8 | 1.5 | 0.21 | 1.3 | ||
EE1-1.5.1 | -0.2% | -0.4% | -0.8% | -0.1% | -0.4% | -0.4% | 21.7 | 16.9 | 4.8 | 1.5 | 0.21 | 1.3 | ||
EE1-1.5.2 | 0.0% | -2.5% | -3.2% | 0.0% | -2.4% | -2.8% | 21.7 | 16.8 | 4.8 | 1.5 | 0.21 | 1.3 | ||
(*) w/o optimal implementation
Tests EE1-1 VLOP in-loop filter
Test | Random Access | All Intra | kMAC/pxl | Param,M | Source | ||||||||
Y | U | V | Y | U | V | Total | Filter | Intra | Total | Filter | Intra | ||
NNVC-VLOP | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 10.0 | 5.2 | 4.8 | 1.5 | 0.06 | 1.3 | |
EE1-1.3 | 0.0% | 0.2% | -0.1% | -0.1% | 0.0% | -0.1% | 9.9 | 5.1 | 4.8 | 1.5 | 0.06 | 1.3 | |
EE1-1.4.1 | -0.1% | 1.3% | 1.0% | -0.1% | 1.4% | 1.3% | 10.0 | 5.2 | 4.8 | 1.5 | 0.06 | 1.3 | |
EE1-1.5.3 | -0.1% | 1.9% | 1.5% | -0.2% | 1.8% | 1.4% | 10.0 | 5.2 | 4.8 | 1.5 | 0.06 | 1.3 | |
EE1-1.5.4 | 0.0% | 0.1% | 0.1% | -0.1% | 0.1% | 0.3% | 10.0 | 5.2 | 4.8 | 1.5 | 0.06 | 1.3 | |
(add some more from slides in JVET-AJ0023 ppt)
EE1-1.x on LOP/VLOP
EE1-1.1 Partial convolution and overparametrization gives some reduction of decoding time for LOP, training crosscheck successful, but training time increased. LOP and VLOP would no longer use same architecture.
EE1-1.2 Simplified residual blocks with attention, provides gain around 0.4%. LOP and VLOP would no longer use same architecture – might need to be combined with 1.1.
EE1-1.3 Early cropping, provides luma gain around 0.3%, but increased decoder runtime, which is likely due to the fact of insufficient SADL runtime optimization of some included elements. Cross-check confirmed, both LOP and VLOP. Candidate for adoption.
EE1-1.4.1 Reduced complexity input feature (without DCT for some of the features) – small gain in luma, small loss in chroma. Some further optimization expected, and retraining on top of 1.3 would be needed.
EE1-1.5 LOP with multi-scale blocks, 0.1%/0.2% luma gain in AI/RA, or . Training crosscheck confirmed. Gain reported had been somewhat higher before.
1.3 and 1.4 are simplifications of the previous architecture.
Decision: Adopt JVET-AJ0054 EE1-1.3 early cropping.
Conditional adoption of EE1-1.4 (conditional on training crosscheck within 2 weeks). In terms of meeting cycles, no chance for any other adoption/retraining on top of that.
EE1-1.1 is another simplification which requires more extensive retraining, could be combined with any of the other (except 1.5).
EE1-1.2 and EE1-1.5 make architectectural changes to residual blocks, where 1.2 from current results would be preferable in terms of performance from current results
EE1-1.2 has not yet been implemented for VLOP which would be desirable using the same architecture.
Further investigate EE1-1.2 also in combination with EE1-1.1 in EE, on top of 1.3/1.4 (i.e.NNVC11).
Further investigate EE1-1.5 in EE, proponents of 1.1, 1.2, 1.5 should consider if an overall combination would be possible and provides benefit.
Furthermore, also EE1-1.1 should be retrained on top of 1.3/1.4 in EE, to be adopted in next meeting in case that other attempts fail.
HOP filter modifications
HOP filter architecture is shown in a figure below. Two out of 23 backbone blocks include attention block.
EE1-2.2 - Block-size invariant implementation of HOP
Only inference is changed. Attention map computation equation modified on a way it doesn’t depend on block size. Minor performance variation (0.0%).
EE1-2.3 - Block based QP information in NN loop filters
One of the inputs for NN-base filter in Slice QP. Under common test conditions all blocks share the same QP. If adaptive QP is enabled there are two options: replace slice QP with average QP inside patch or use block QP instead. It is shown that the second variant (block QP) with proper NN-filter retraining brings 1.6% performance improvement (table below). Proposed to change SliceQP to BlockQP (there will be no change at common test conditions).
Tests EE1-2.3 results with Adaptive Qp enabled in the anchor and test
Test | Random Access | All Intra | ||||
Y | U | V | Y | U | V | |
HOP3 (avg QP) vs HOP3 (slice) | 0.8% | 2.0% | 2.4% | 1.4% | 2.6% | 2.6% |
HOP3 block QP(retrained) vs HOP3 (slice) | -0.8% | -0.9% | -1.3% | -0.9% | -1.2% | -1.3% |
HOP5 (avg QP) vs HOP5 (slice) | 0.6% | 1.2% | 1.3% | 0.7% | 2.0% | 2.6% |
HOP5 block QP(retrained) vs HOP5 (slice) | -1.0% | 0.5% | 0.4% | -1.6% | 0.4% | -0.3% |
HOP filter architecture with attention block
Tests EE1-2 HOP in-loop filter
Test | Random Access | All Intra | kMAC/pxl | Param,M | Source | ||||||||||
Y | U | V | Y | U | V | Total | Filter | Intra | Total | Filter | Intra | ||||
NNVC-HOP | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 471 | 466 | 4.8 | 2.96 | 1.439 | 1.3 | |||
EE1-2.2 | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 471 | 466 | 4.8 | 2.96 | 1.439 | 1.3 | |||
EE1-2.2 is having some small gain (0.03% luma) for RA, and decoder runtime reduction, but only for large resolutions (91/94% for A classes). In case of some sequences very small losses, but gain on average. The model itself is unchanged.
Decision: Adopt JVET-AJ0166 EE1-2.2.
Results on EE1-2.3 indicate that availability of local QP as input to the network could be beneficial. It is however commented that training was performed with the same approach for local QP (as done by RDOQ), and it might not work with common approaches of rate control. It was suggested to perform inference test (without retraining) on a rate control algorithm from VTM. Check if an operational rate control algorithm is available in NNVC software. Report results during the meeting (if possible). Partial crosscheck available, confirmed so far.
Decision: Adopt JVET-AJ0124 (change input in NNVC loop filters from slice QP to local QP), under CTC inputting constant QP from slice (not adopting the re-trained model). It is noted that this modification applies to HOP, LOP, VLOP and SR.
Further investigation in EE on the re-training of the model and investigating other aspects, e.g. granularity of the QP information.
EE1-3.1 - RA/LDB Unified Reference Frame Synthesis for VVC Inter Coding
Tests EE1-3: NN-based Inter prediction
Test | Random Access | Low-delay B | kMAC/pxl | Param (Mprm) | |||||||||||
Y | U | V | Y | U | V | Total | FL | Intra | Inter | Total | FL | Intra | Inter | ||
NNVC-10.0-LOP | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 21.7 | 16.9 | 4.8 | 0 | 1.5 | 0.21 | 1.3 | 0 | |
EE1-3.1.2 | -3.5% | -4.4% | -4.3% | -3.6% | -3.9% | -0.3% | 748.7 | 16.9 | 4.8 | 727 | 5.3 | 0.21 | 1.3 | 3.8 | |
The training set is still Vimeo-90K triplet. No decoder run time reported.
Gain is 3.5…3.6% (random access and low-delay B), but complexity is still very high 727 kMAC/pxl.
Cross-checker reports that there is some chroma loss in some sequences. It was also reported that there is some problem in running the decoder due to HLS incompatibility in the way how the additional reference picture is inserted.
The cross-checker would be able for performing training cross-check also with the Vimeo data set, but the still missing training with BVI/TVD is also requested.
Before adoption, full SADL implementation would be required (currently, partially in PyTorch)
Continue the EE on the remaining aspects.
EE1-4.1 - Wavelet transform for super-resolution loss function
NNVC NNSR filter was retrained with and without DWT in loss function. DWT in loss function provides small chroma gain.
It is asserted that the additional computation of DWT is not significantly slowing down the training. Training was cross-checked. 4.1.2 is preferable due to its small benefit in chroma quality.
Decision: Adopt JVET-AJ0056 EE1-4.1.2 (retrained model for NNSR, disabled by default).