Main content

Can Layered VVC Reduce the Cost of Streaming?

Can Layered VVC Reduce the Cost of Streaming?

When we use video conferencing services, such as Zoom, under the hood the video compression is often performed in a scalable manner so that the same coded video can be adapted to devices with different display resolutions without the need to decode the highest resolution version. Despite the market adoption in video conferencing, scalable video coding has not gained much popularity in streaming. 

Video streaming services usually store the same content in several versions, such as high-definition (HD) and ultra-high-definition (UHD), so that each viewer can receive a version that fits their screen and network conditions. Scalable video coding offers another approach: encode the video once in layers, where a base layer provides one version and additional layers improve resolution or quality. Recent verification test results for dual-layer Versatile Video Coding (VVC/H.266) suggest that this approach may now be practical for streaming, with little or no visible quality penalty compared with conventional single-layer VVC. The results indicate potential for service providers to achieve substantial capacity savings in content delivery networks and servers.

Adaptive bitrate video streaming

To understand why scalable video coding matters for streaming, it helps to first look at how today’s adaptive streaming systems are typically built. Since the Internet may have dynamically changing network traffic and users have access networks of different capacity, adaptive bitrate (ABR) streaming is used. The same video content is encoded to multiple versions having different bitrates, often referred to as a bitrate ladder. While streaming, the player dynamically selects the bitrate version that best suits the experienced network capacity and the device capabilities, such as screen resolution. Figure 1 illustrates an example where the video content is made available in both the high-definition (HD) and ultra-high-definition (UHD) resolutions. Both versions are stored in the origin server and delivered over the content delivery network (CDN) to edge servers. Players can adaptively choose the version that is streamed.

Figure 1. Example of adaptive bitrate streaming.

Figure 1. Example of adaptive bitrate streaming.

Scalable Video Coding

All modern video coding standards include operation modes for scalable video coding, which enables video clips to be compressed once, and decoded with a picture resolution suitable for a display or delivered with a bitrate suitable for network conditions. The inherent characteristic of a scalable video bitstream is that it can still be decoded even after parts of it have been discarded. Consequently, if the capacity of a network does not allow an entire clip to be transmitted, a subset of the clip can be forwarded and still properly decoded at the receiving end. Likewise, if a device cannot display a scalably coded video clip at the full size intended by the creator, a subset of the clip can be decoded to minimize processing power and conserve battery. Where these restrictions do not apply and the full video clip can be decoded, the video quality is not compromised.

While one encoded video serving many needs sounds attractive, the use of scalable video has not taken off in streaming services. Presumably, one reason has been the inherent bitrate overhead or loss of compression efficiency compared to non-scalable video coding. Another reason could have been that scalable coding of standardized codecs prior to VVC/H.266 were added as extensions rather than integral parts of the base standard. Consequently, the additional implementation burden on top of the base standard could have limited the deployment.

Unlike earlier video coding standards, VVC integrates scalability into the base codec design rather than introducing it as a separate extension. To make the implementation burden as low as possible, the scalability support in VVC involves no additional coding tools compared to non-scalable VVC coding. Scalability requires only high-level syntax support, which is typically implemented in software.

Verification test of scalable VVC

Verification testing concludes a standardization project by putting the standard into a test bench. The verification testing of standardized video codecs is performed through systematic and rigorous subjective viewing. 

The Visual Quality Assessment Working Group of MPEG (Moving Picture Experts Group) recently released a verification test for scalable VVC/H.266 focused on dual-layer coding enabling spatial scalability suitable for streaming services. The test covered a wide range of content, including both high-definition (HD) and ultra-high-definition (UHD) resolutions as well as both standard and high dynamic range. The VVC dual-layer coding was compared with two anchors: non-scalable single-layer VVC and the upscaled base layer of the dual-layer bitstream. 

The results of the verification test highlight the high efficiency of the VVC dual-layer configuration, confirming its suitability to be used for a bitrate ladder in adaptive streaming. Dual-layer VVC provided equal or only slightly lower subjective quality compared to non-scalable VVC at the same bitrate. In practice, viewers could not reliably tell the dual-layer version from the regular VVC version at the same bitrate. Figure 3 provides an example of the results obtained for one of the tested video clips. The results show that the 95% confidence interval of the mean opinion scores (MOS) of the respective bitrate points of the single-layer and dual-layer VVC overlap. Thus, no statistically significant difference is observed between the video quality of dual-layer VVC and non-scalable VVC at the respective bitrates of this video clip.

Figure 2. Verification test results for the BodeMusem test sequence (high-definition, standard dynamic range) (source: MPEG N26406)

Figure 2. Verification test results for the BodeMusem test sequence (high-definition, standard dynamic range) (source: MPEG N26406)

Potential of scalable VVC for service providers and users

The results of the verification test suggest that with scalable VVC two bitrate versions can be maintained in the servers and content delivery networks (CDNs) at the storage size of a single bitrate version. As depicted in the example of Figure 3, a scalable video clip including both the HD and UHD resolutions can be stored in the origin server, delivered over the CDN, and stored in the CDN edge server. Players that only consume the HD resolution can be served by the base layer of the scalable video clip, whereas other streaming clients may stream the complete scalable video clip. The storage and delivery cost of a separate HD version of the video clip is saved.

Figure 3. Example of adaptive bitrate streaming with scalable video coding.

Figure 3. Example of adaptive bitrate streaming with scalable video coding.

Such a reduction in the server and CDN capacity not only benefits the streaming service providers but may also benefit users with fewer video freezes, thanks to reduced likelihood of caching misses in CDNs. What's more, these savings come on top of the approximate 50% bitrate saving that VVC achieves relative to its predecessor, the High Efficiency Video Coding (HEVC/H.265).

Additional Uses for Scalable Video

Scalability not only offers capacity reduction possibilities for streaming, but also an overlaying functionality, where supplementary visual content can surround or be overlaid on top of the main video in the base layer of a scalable video bitstream. While there are other approaches to add overlays on top of video content, scalable video coding excels in achieving frame-accurate timing for overlays and requires no additional processing after video decoding. The verification test proved the feasibility of several user-selectable overlaying scenarios, including textual and graphical overlays (such as game statistics) and sign language video on the side of a main video. In summary, the overlaying functionality enables personalization in video streaming.

An increasing amount of video content is analyzed by machine vision tasks and not watched by humans. In many video application domains, such as intelligent traffic, assembly line, and security monitoring systems, a machine vision task performs continuous monitoring, while a human observer may only investigate the video on an as-needed basis. Scalable video coding may be used in such applications to reduce the continuously transmitted video bitrate so that only a machine-optimized base layer is constantly delivered and a human-optimized enhancement layer is conveyed only on request. Our studies indicated that overall transmission bandwidth savings are achieved when human observation is required less than 80% of the time.

Conclusion

While scalable video coding is broadly deployed in video conferencing, it has not yet seen large-scale adoption in video streaming. The dual-layer VVC verification testing results demonstrate significant benefits of scalable video coding for streaming applications. These results should encourage the streaming industry and the research community to carry out more investigation and large-scale testing of scalable video coding for streaming. If the subsequent studies prove successful, scalable VVC could become a practical way to make streaming infrastructure leaner while keeping the viewing experience intact.

Miska Hannuksela

About Miska Hannuksela

Miska Hannuksela, (M.Sc., Dr. Tech), is the Head of Video Research at Nokia Technologies and a Nokia Bell Labs Fellow. He is an internationally acclaimed expert in video and image compression and end-to-end multimedia systems.

Article tags