T2AV-Compass:迈向文本到音视频生成的统一评估 / T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation
1️⃣ 一句话总结
这篇论文提出了一个名为T2AV-Compass的统一评估基准,用于全面衡量文本生成音视频系统的性能,发现现有模型在真实感和跨模态一致性上仍远不及人类水平,为未来研究指明了改进方向。
Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet its evaluation remains fragmented, often relying on unimodal metrics or narrowly scoped benchmarks that fail to capture cross-modal alignment, instruction following, and perceptual realism under complex prompts. To address this limitation, we present T2AV-Compass, a unified benchmark for comprehensive evaluation of T2AV systems, consisting of 500 diverse and complex prompts constructed via a taxonomy-driven pipeline to ensure semantic richness and physical plausibility. Besides, T2AV-Compass introduces a dual-level evaluation framework that integrates objective signal-level metrics for video quality, audio quality, and cross-modal alignment with a subjective MLLM-as-a-Judge protocol for instruction following and realism assessment. Extensive evaluation of 11 representative T2AVsystems reveals that even the strongest models fall substantially short of human-level realism and cross-modal consistency, with persistent failures in audio realism, fine-grained synchronization, instruction following, etc. These results indicate significant improvement room for future models and highlight the value of T2AV-Compass as a challenging and diagnostic testbed for advancing text-to-audio-video generation.
T2AV-Compass:迈向文本到音视频生成的统一评估 / T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation
这篇论文提出了一个名为T2AV-Compass的统一评估基准,用于全面衡量文本生成音视频系统的性能,发现现有模型在真实感和跨模态一致性上仍远不及人类水平,为未来研究指明了改进方向。
源自 arXiv: 2512.21094