Back to Search Document details
37th Meeting: Geneva, CH, January 2025 2025-01-18 10:56
Neural Network Analysis for Next-Generation Video Compression Standard
Abstract
Neural networks (NNs) have been a hot topic for many applications, like autonomous driving, object recognition, natural language processing, etc. The question comes to the point: will NNs be realistic for video coding applications? Some complexity analysis is the keyhole that allows us to peek into the feasibility for NNs in future video compression standards. Currently, in the standard community, several video coding tools utilizing NNs have been investigated, among them are intra prediction and in-loop filtering (ILF) which are analyzed in this contribution. There are mainly two ways to deploy NNs for video coding applications, one is using GPU/NPU on the market, and the other is using in-house designed ASIC. The details will be illustrated later.
JVET-AK0323 Neural Network Analysis for Next-Generation Video Compression Standard [T. Hsieh, W.-J. Chien, V. Seregin, M. Karczewicz (Qualcomm)] [late]

Neural networks (NNs) have been a hot topic for many applications, like autonomous driving, object recognition, natural language processing, etc. The question comes to the point: will NNs be realistic for video coding applications? Some complexity analysis is the keyhole that allows us to peek into the feasibility for NNs in future video compression standards. Currently, in the standard community, several video coding tools utilizing NNs have been investigated, among them are intra prediction and in-loop filtering (ILF) which are analyzed in this contribution. There are mainly two ways to deploy NNs for video coding applications, one is using GPU/NPU on the market, and the other is using in-house designed ASIC. The details are illustrated in this contribution.

Hardware architectures: External NN processor (two separate SoCs), or combined on one chip.

Analysis for intra prediction and loop filter

It was estimated that chip area for NN computation may be more than 60x compared to conventional video ASIC in the case of intra prediction. A 2-chip solution was said to not be practical due to latency in sequential block dependency (4x4 worst case, requiring 5 cycles for processing). Sparsity by zero weights may not be simple to implement. Intra was asserted to not be practical on a single chip either, e.g. due to power consumption.

It was commented that in current NN intra pred. sparsity is designed in a way such that it is grouped, as a rule 16 zero values go together.

It was asked how intra NN would compare to MIP. This was not analysed.

Loop filter is less critical for two-SoC solution, latency not so critical, but high cost of computation still a problem.

For VLOP2, chip area would be similar to entire video codec (multi-standard).

It is to be noted that the analysis is not based on a real implementation study, but uses many assumptions that might be differently resolved by alternative hardware architectures.

More information about “what is possible, or what will be possible within a certain time frame” would be welcome.

AHG6/AHG12: Enhanced compression beyond VVC capability (77)

Summary and BoG reports

Contributions in this area were discussed at 1825–2010 on Tuesday 14 Jan. 2025 (chaired by JRO), at 0900–1315 and at 1430–1800 on Wednesday 15 Jan. 2025 (chaired by JRO).

Decisions
More information about “what is possible, or what will be possible within a certain time frame” would be welcome.
Citation