Back to Search Document details
19th Meeting: by teleconference, June 2020 2020-06-26 19:57
AHG16: Performance of a reasonably fast VVC software decoder
Abstract
Performance of a reasonably fast VVC software decoder is provided: it is asserted to run about 30 times faster than VTM on an 8-core processor when processing 4K bitstreams (CTC RA). No claim of optimality is made as the optimized decoder was written from scratch in a short amount of time.
JVET-S0224 AHG16: Performance of a reasonably fast VVC software decoder [F. Bossen (Sharp)]

This contribution was discussed during 2010–2030 on 26 June (chaired by JRO)

Performance of a reasonably fast VVC software decoder is described: it is asserted to run about 30 times faster than the VTM on an 8-core processor when processing 4K bitstreams (CTC RA). No claim of optimality was made, as the optimized decoder was written from scratch in a short amount of time.

The table below lists 4K bitstreams (generated with VTM version 9.0, RA configuration of common test conditions) for which the bit rate is 40Mbps or less (maximum rate permitted by Level 5.1). Experiments were run on a computer featuring a single Xeon W processor with 8 cores (Skylake). “Real time” performance (60 frames per second or more) is achieved in all cases with the optimized decoder, while VTM processes between 2.8 and 4.6 frames per second.

Sequence

QP

Frames

Bit rate [Mbps]

VTM [s]

VTM [fps]

Optimized [fps]

Tango

22

294

27.9

93.792

3.13

95.4

FoodMarket

22

300

17.0

89.978

3.33

103.8

Campfire

27

300

30.7

93.977

3.19

110.8

CatRobot

22

300

33.7

103.032

2.91

86.9

DaylightRoad

27

300

10.3

81.181

3.70

109.3

ParkRunning

32

300

22.2

107.255

2.80

92.1

DayStreet

22

300

20.4

88.508

3.39

103.4

FlyingBirds

22

300

15.9

69.801

4.30

121.6

PeopleInShoppingCenter

22

300

11.6

65.853

4.56

127.2

SunsetBeach

27

300

26.6

80.717

3.72

109.2

The three blocks with highest runtime consumption are inter prediction, ALF and deblocking. Runtime consumption of scaling and inverse transform is very small.

Though the optimized code runs significantly faster, the percentage of runtime consumed by individual tools (as measured by tool off tests) does not change significantly compared to the findings of AHG13.

Compared to a similarly optimized HEVC decoder, the runtime increase is around 1.5-2x.

It was asked if the runtime would be significantly slower for higher bit rates. This would probably be the case.

No specific action was necessary on this.

Profile/level specification (1)

See also section 4.5 for studies of coding tools and operation beyond the bit depth supported in VVC v1 profiles.

These contributions were first discussed in the Tuesday 30 June joint meeting.

References:
JVET-S0013
Decisions
These contributions were first discussed in the Tuesday 30 June joint meeting.
Citation