Back to Search Document details
40th Meeting: Geneva, CH, October 2025 2025-10-09 17:54
AhG17 Sharing experience of conducting JPEG AI Call for Proposals
Abstract
Documents shares information about organization call for proposals on learnable image codec technologies in JPEG. Hopefully this experience can be taken into account and help to resolve several issues observed during Call for Evidence on video compression with capability beyond VVC.
JVET-AN0301 AhG17 Sharing experience of conducting JPEG AI Call for Proposals [E. Alshina (Huawei)] [late]

Documents shares information about organization call for proposals on learnable image codec technologies in JPEG. Hopefully this experience can be taken into account and help to resolve several issues observed during Call for Evidence on video compression with capability beyond VVC.

JPEG AI call for proposals was conducted using multiple steps.

1. Call for proposal was issued. Document describes evaluation methodology, anchors, training set, but doesn’t specify test set. Only CfP responses for which decoding on different devices (namely on GPU and on CPU) result in matching reconstruction were seriously considered, the rest to submissions were taken into account just as information. Among 10 submissions only 5 met this criterion.

2. Registered proponents uploaded decoder package (which includes all model parameters) and then write access was closed (up-date of decoder or model parameter no longer was no longer possible).

3. Several days after decoders freeze evaluation set, selected by experts, not participating in CfP, was specified and announced. This might be too much for video coding standard CfP, it should be enough to built secret test set (announced after decoder freeze) which consist of one video sequence per each class (HDR/SDR, UHD/HD, random access and low-delay coded). Those secret test sequences can be selected by visual test coordinator(s) taking into account how good they are for visual evaluation.

4. Proponents were given limited time to encode secret testing set and match target rates. Different credentials were given to uploaded streams, again write access was closed after deadline. This really helps to resolve a mess with multiple versions of uploaded streams. Limited time for rate matching naturally limits encoder complexity.

5. After both streams and decoders are submitted and frozen, the cross-checkers for each response was assigned by chairs (randomly assigned). Cross-checkers were mandated to decode streams using CPU and GPU and check difference in reconstruction. Distributed cross-check reduced loading; device interoperability cross-check increases the quality of submissions.

6. Submission which failed in a cross-check are considered for information only.

The following is proposed:

  • Close write access after decoder submission deadline.
  • Conduct CfP evaluation using (at least partially) secret test set, which is announced after decoder freeze.
  • Close write access after streams submission deadline.
  • Randomly assign cross-checker to reduce loading. Mandate cross-checker to verify device interoperability.
  • Consider implementation and testing algorithms under consideration on mobile device in parallel with standard development as one of the most realistic ways of complexity evaluation.

It was commented that, if SADL was used for NN implementation, no mismatch typically occurs.

The following was agreed :

An additional set of sequences would be beneficial, likely to be selected by the test coordinator and the JVET chair after delivery of decoder binaries and bitstreams.

In the CfP a statement needs to be added that it is mandatory to report which material was used to train NN or other trained approaches.

The aspect about kMAC, number of parameters, number of layers, etc. should be part of the complexity questionnaire.

It was pointed out that the additional set would need to be large enough to unveil relevant information. It needs more consideration about practicality, or potentially not include it in subjective testing.

During 1900-2000 on Friday 10 October, a subsequent discussion was held about possible other changes to CfP conditions and test set.

Adding higher rate points also for proposals appears useful to also be able investigating the gain in the higher quality range (based on objective metrics, as quality may be close to transparent). For this, it would not be required to have precise rate matching, to be useful e.g. for additional BD rate computation in the higher rate range.

During the joint meeting, some concern had been raised about inappropriateness of certain sequences.

On possible aspects to improve the test set, the following comments were made:

  • Hallway scene has a very low rate, but on the other hand seems good for visual testing, also has noise. Several experts supported keeping it.
  • Fashion lady has a very low rate, due to the fact that fine details are rare (except for some areas); on the other hand, it was commented that it has local gradients which is also important to be properly reconstructed in HDR
  • It might be useful to find a replacement for one of the Gregory sequences (should be replaced by another sequence taken directly from a smartphone camera without compression). If a replacement is found, the scarf sequence which needs higher bitrate should be kept.
  • Gaming sequences are all produced with outdated rendering. Level1 (priority) and GTAV could be candidates for replacement, the latter being too simple. GTAV was however mentioned to be relevant for low latency testing (could be kept in a dedicated category when the functionality of ultra low latency / error resilience would be included in CfP).
  • Office walk might be a candidate for removal (not optimum for viewing), but a good replacement would need to be found.
  • The lowest rate for Beatriz may be too low, otherwise it is a good test sequence.
  • Some 8K sequences (except Waterfall Forest and Chandelier) might better be downsampled to 4K. Chandelier has relatively low rate. On the other hand, concern is raised this would make the 8K category almost empty. No replacement should be made if no better material would be available.
  • It might not be necessary to have 6 sequences in UGC.
  • Several experts commented that higher bit rate sequences would be desirable.

Comments and agreed changes on sequences and conditions were collected during the session on Friday 10 October in an edited version JVET-AN2026 of the previous CfE, with the intent to consider those in generating an initial version of the CfP document during the AHG period.

AHG18 Ultra-low latency and packet loss resilience (8+2)

Contributions in this area were discussed during 1400–1700 on Wednesday 8 Oct. 2025, and during 1110–1300 on Thursday 9 Oct. 2025 (chaired by JRO).

Decisions
Comments and agreed changes on sequences and conditions were collected during the session on Friday 10 October in an edited version JVET-AN2026 of the previous CfE, with the intent to consider those in generating an initial version of the CfP document during the AHG period.
Citation