VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking Paper โข 2303.16727 โข Published Mar 29, 2023 โข 2
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs Paper โข 2406.07476 โข Published Jun 11, 2024 โข 37