Back to Search Document details
31st Meeting: Geneva, CH, July 2023 2023-07-16 14:50
AHG9: Common text for proposed generative face video SEI message
Abstract
This contribution proposes common text based on JVET-AE0080 and JVET-AE0083 for a generative face video (GFV) SEI message for inclusion in the next version of VSEI. The proposed GFV SEI message conveys information for facial features extracted from source video prior to encoding that can be applied to a previously decoded reference picture to reconstruct a face picture at the current time. The proposed GFV SEI message is intended for ultra low bitrate face video compression applications such as video conferencing, live entertainment, and face animation.
JVET-AE0280 AHG9: Common text for proposed generative face video SEI message [B. Chen, J. Chen, Y. Ye (Alibaba), S. Wang (CityU), S. McCarthy, P. Yin, G.-M. Su, A. K. Choudhury, W. Husak (Dolby)] [late]

This contribution proposes common text based on JVET-AE0080 and JVET-AE0083 for a generative face video (GFV) SEI message for inclusion in the next version of VSEI. The proposed GFV SEI message conveys information for facial features extracted from source video prior to encoding that can be applied to a previously decoded reference picture to reconstruct a face picture at the current time. The proposed GFV SEI message is intended for ultra low bitrate face video compression applications such as video conferencing, live entertainment, and face animation.

The application of translating a real person into an avatar does not require precise encoder/decoder match, synthesis models would translate the parametric description into the avatar’s movements. Ultimately, a decoder would need to be able to interpret different combinations of parameters.

The application of “video realistic” generation of video would require precise match of encoder/decoder. The examples shown in the demo are not fully reconstructing all movements such as eye closing (at 2 kbit/s), but could be adjusted with increased number of landmarks, etc.

This was further discussed in a joint meeting with WG2/WG4/WG7/VCEG (see section 7.4). The differences and commonalities with avatar technology and implicit NN technology for generative NNs were discussed, and it was agreed to perform further study and investigate requirements.

Non-SEI HLS aspects (0)

Section kept as a template for future use.

Plenary meetings, joint meetings, BoG reports, and liaison communications

JVET plenaries

No intermediate plenaries were held, as document review and decisions were made in single-track mode at this meeting (with some BoG activity as noted). Further detail on scheduling is recorded in section 2.15.

Information sharing sessions with other WGs and AGs of the MPEG community were held on Monday 17 July 0900–1130, Wednesday 19 July 0900–1030, and Friday 21 July 1400–1600.

Joint meetings involving JVET were held as follows:

  • JVET, WG 4, Q6/16 (VCEG) on MPI/MIV & VCM related topics, on Monday 17 July at 1615-1735
  • JVET, MPEG WG2 Requirements, MPEG WG4 Video, MPEG WG7 3D Graphics and Gaptics, and Q6/16 (VCEG) on generative face video, on Monday 17 July at 1715-1800
  • JVET, MPEG AG 5 Visual Quality Assessment & Q6/16 (VCEG) on verification testing, on Monday 17 July at 1815-2000
  • Further detail about these sessions with other groups is provided in the other subsections of this section.

General plenary wrap-up discussions are recorded under sections 8, 9, and 10.

Information sharing meetings

Information sharing sessions with other WGs and AGs of the MPEG community were held on Monday 17 July 0900–1130, Wednesday 19 July 0900–1030, and Friday 21 July 1400–1600. The status and plans for the work in the MPEG WGs and AGs was reviewed at these information sharing sessions.

Joint meeting of JVET, WG 4, Q6/16 (VCEG) on MPI/MIV & VCM related topics

This joint session was held on Monday 17 July at 1615-1735. About 150 people were present in the room and about 180 were connected on Zoom (substantially overlapping with those present in the room).

Agenda

  • MIV (MPEG immersive video) and MPI (multiplane image) relationship
  • VCM

On the MIV and MPI relationship topic, a proposal to support MPI using SEI was submitted to JVET as JVET-AE0066. JVET-AE0066 proposes a multiplane image information (MPII) SEI message for inclusion in the next version of VSEI. The proposed MPI SEI message is intended to enable virtual reality/immersive experiences on existing consumer devices using existing video codecs such as HEVC and VVC that were developed for conventional two-dimensional video.

There was discussion of whether this proposal would provide additional functionality or other improvement relative to what the MIV standard can do and whether implementing MIV rather than the proposed SEI approach would be burdensome.

  • It was reported that multiplane image (MPI) and multisphere (MSI) are supported in the “MIV Extended – restricted geometry” profile or sub-profile.
  • For MIV video transport, it was reported that all necessary information can be put into a single video bitstream and used with a single decoder instantiation.
  • There was discussion of whether MIV supports irregular spacing of depths, and it was reported that it does.
  • There was discussion of how many layers are needed for effective MPI usage and whether the number of layers can be restricted sufficiently for an application. It was said that there are “levels” defined for dealing with this.

Further study in MPEG WG4 was suggested to clarify these issues.

For VCM the following issues were suggested to be discussed

  • Anchors
  • Test sequences, conditions and metrics
  • Architecture of a VCM “normative” (pursued in WG4 with a core codec that may differ somewhat from a generic codec) versus a “non-normative” approach (e.g. in the technical report planned to be developed in JVET)
  • The role of SEI messages (e.g. JVET-AE0064, also submitted to WG4 as m64518).

Further study of these issues was suggested, with WG4 and JVET AHG activity to take place between meetings.

There was discussion about proposed SEI messages (or modifications of existing SEI messages) that are related to machine consumption of video content. It was suggested that SEI messages developed in JVET should be “generic” rather than specialized to machine analysis.

It was suggested to use a different NAL unit type for important data related to system usage.

JVET planned to include JVET-AE0064 in a “technology under consideration” output document (as candidate technology for a future version of VSEI or similar), which does not include discussion of a specific usage procedure – it is basically sending flags indicating whether the video is optimized for machine consumption or human consumption.

WG4 is expected to develop “messages” (or some form of signalling) and specific processing methods for machine usage of video.

Source picture timing indication was also suggested to be potentially relevant to machine usage of video content (as this information could be useful for machine analysis), but it is general purpose, not targeted specifically for machine analysis.

Joint JVET, MPEG WG2 Requirements, MPEG WG4 Video, MPEG WG7 3D Graphics and Gaptics, and Q6/16 (VCEG) on generative face video

This joint session was held on Monday 17 July at 1715-1800, continuing in the same room as the joint session described above, with a similar number of people participating.

Several approaches for technology to synthesize moving faces were discussed.

  • JVET has received proposals for “generative face video” SEI messages including JVET-AC0088, JVET-AD0051, JVET-AE0080, JVET-AE0083, JVET-AE0088, and JVET-AE0280. The concept is to have an ordinary 2D picture and some data that guides the generation of additional 2D pictures from the basis picture.
  • WG7 has been working on “avatar” technology based on 3D graphics technology and parameters extracted from analysis of 2D video content of faces or a full body. This provides parameters that are used for an animated 3D mesh representation. Requirements for avatar technology standardization are under study in WG2.
  • WG4 has been developing “implicit neural network” technology, where a neural network is coded as a neural network representation and coded feature data that can generate animated pictures.

Further study is planned, esp. in MPEG Requirements WG2.

Joint JVET, MPEG AG 5 Visual Quality Assessment & Q6/16 (VCEG) on verification testing

This joint session was held on Monday 17 July at 1815-2000, in a different room than the two sessions described above. About 100 people were present in the room and a similar number were connected on Zoom (substantially overlapping with those present in the room).

This joint session was to discuss verification testing preparations, particularly including test materials for film grain.

JVET-AE0219, reporting results of visual checking of scalable VVC verification testing streams, was discussed. Artefacts related to decoder motion vector refinement (DMVR) were noted for low bit rates, for which a fix called DMVREncMvSelect was available and planned to be used (despite a small loss in average PSNR performance).

JVET-AE0288, reporting results of expert viewing for the spatial scalability category of the VVC multilayer VT, was discussed. It was reported that higher fidelity encoding of a base layer that is upsampled is sometimes visually superior to two-layer coding of 4K content, although the two-layer coding may have better PSNR results.

The problem was reported to be the temporal effect of periodic quality “pumping” in the two-layer coded content. It was commented that single-layer coding can also exhibit “pumping”. Lower-resolution coding would not have such pumping.

It was pointed out in the discussion that upscaling could be done with a neural network, which would further improve the visual quality of the upscaled base layer. Another comment was that neural network post-processing of the multilayer coding could also be used.

It was asked whether chroma phase was being used correctly, which seemed to be the case.

It was commented that using the lowest quality coding in the range may not be appropriate, but the test was designed to span a wide range of quality.

Analysis of the seating positions of the viewers and confidence intervals, and possibly the preferences of individual viewers, was suggested.

It was commented that the amount of bit rate to be allocated to the enhancement layer should perhaps be reduced. However, the enhancement layer was said to already be getting a relatively low percentage of the total bit rate.

It was commented that the content may not have enough real high frequency signal in it.

It was noted that the high resolution content may contain noise that is removed by the downsampling.

Further study of these issues was needed before an effective test can be performed.

Further informal viewing in AG5 and AHG activity was planned.

BoGs (3)

The following break-out groups were established at this meeting to conduct discussion and develop recommendations on particular subjects.

Decisions
The following break-out groups were established at this meeting to conduct discussion and develop recommendations on particular subjects.
Citation