Dual-view Pyramid Network for Video Frame Interpolation
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3664647.3681555 ↗
摘要
Video frame interpolation is a critical component of video streaming, a vibrant research area dealing with requests of both service providers and users. However, existing methods cannot handle changing video resolutions while improving user perceptual quality. We aim to unleash the multifaceted knowledge yielded by the hierarchical views at multiple scales in a pyramid network. Specifically, we build a dual-view pyramid network by introducing pyramidal dual-view correspondence matching. It compels each scale to actively seek knowledge in view of both the current scale and a coarser scale, conducting robust correspondence matching by considering neighboring scales. Meanwhile, an auxiliary multi-scale collaborative supervision is devised to enforce the exchange of knowledge among scales and thus reduce error propagation from coarse to fine scales. Based on the robust capture of video dynamics via pyramidal dual-view correspondence matching, we further construct a pyramidal refinement module that formulates frame refinement as progressive latent representation generations by developing flow-guided cross-scale attention for feature fusion among frames. The proposed method is able to improve the perceptual quality on several benchmarks of varying video resolutions, while keeping low distortion and a compact model size.