Back to Search Document details
40th Meeting: Geneva, CH, October 2025 2025-10-08 13:59
AHG18: Proposed methodology and test conditions for ultra-low latency and packet loss resilience performance evaluation
Abstract
In the past several JVET meetings, several documents discussing aspects related to AHG18 were discussed, the related techniques and test conditions were discussed comprehensively. This contribution proposes an evaluation methodology and suggests test conditions based on the discussion targeting further AHG18 work. Comparing with the current AHG18 test condition, which only contains Wi-Fi connection case, the 5G network constraints are additionally proposed, full VTM tool set is enabled, with a wide bitrate range, LDB and spatial scalability testing conditions are further supported. The contribution also suggests the objective quality metrics calculation methodology, accompanied with an excel template for quality plots construction and summary results collecting.
JVET-AN0079 AHG18: Proposed methodology and test conditions for ultra-low latency and packet loss resilience performance evaluation [S. Ikonin, I. Gribushin, X. Ma, E. Alshina (Huawei), S. Deshpande (Sharp)]

In the past several JVET meetings, several documents discussing aspects related to AHG18 were discussed, the related techniques and test conditions were discussed comprehensively. This contribution proposes an evaluation methodology and suggests test conditions based on the discussion targeting further AHG18 work. Comparing with the current AHG18 test condition, which only contains Wi-Fi connection case, the 5G network constraints are additionally proposed, full VTM tool set is enabled, with a wide bitrate range, LDB and spatial scalability testing conditions are further supported. The contribution also suggests the objective quality metrics calculation methodology, accompanied with an excel template for quality plots construction and summary results collecting.

It was suggested to follow a methodology where the error patterns are randomly selected after bitstreams have been generated. It is also suggested to use several different error patterns and compute the average performance (provided that a reasonable objective metric exists).

Some concern was raised on the aspect of encoder configuration. It was agreed to use reduced runtime 3 for the unicast case (as it simplifies the feedback reaction), and use the default VTM settings for spatial scalability.

At this moment, the section 4 about metrics is left is left open. In the future, various metrics might be used, such as variation of errors over a frame, finding of extremely distorted regions, etc. Some investigation about potential existing metrics might be useful.

On the aspects of possible other feedback mechanisms, and usage of GDR rather than IRAP for refresh, see further notes under JVET-AN0089 and JVET-AN0254.

In an update of the document, the current GDR implementation is mentioned as an alternative with IRAP, and the aspect of metrics is left open.

Formulation for section 4: “One metric that may optionally be reported is PSNR. Reports using other metrics to demonstrate the benefit are welcome.”

Furthermore, the Excel sheet should be modified such that quality over rate for a given latency can be plotted, as well as quality over latency for a given rate point.

It was agreed to use the modified JVET-AN0079 as basis for the first version of CTC for ULL/ER

1610-XXXX Presentation of CfE results by M. Wien:

All at QP 22 (single layer or scalable enhancement layer), delay

Broadcast scenario (scalability)

Crowd run (left) is a sequence where also the base layer had a freeze.

Unicast scenario:

There seems to be at least some coherence with PSNR values by tendency, where often PSNR curves are closer than MOS curves. A scatter plot would be helpful.

One participant reports that he found it difficult to decide whether frame freeze at good quality or a smoothed reconstruction was better. It might be useful not testing unicast and broadcast (scalability) scenario in the same session.

It was suggested that a better methodology might be to make an A/B comparison.

Another option might be to run a DCR test where the uncorrupted decoded video is compared against the concealed version and measure impairment.

Generally, it appears useful for visual testing not going into the higher QP range where larger compression errors occur, as the goal is to judge the benefit of an error resilience/concealment method.

In general, it appears possible to test visual benefit of error resilience/error concealment also with relatively short sequences under the assumption that just the behaviour in corrupted cases is included. It might be useful to test different error patterns to guarantee that enough variability of error cases is given.

Decisions
In general, it appears possible to test visual benefit of error resilience/error concealment also with relatively short sequences under the assumption that just the behaviour in corrupted cases is included. It might be useful to test different error patterns to guarantee that enough variability of error cases is given.
Citation