AcoustiTrace diagnostic benchmark for audio–video generation

AcoustiTrace

When Plausible Sound Violates Physics

  • The Acoustic-Process Diagnostic Gap: Current audio–video benchmarks emphasize perceptual quality, semantic alignment, and synchronization, but provide limited support for identifying which acoustic process fails or how severely it departs from a physically expected relation.

  • AcoustiTrace Benchmark: We introduce a diagnostic benchmark for T2AV and I2AV generation that organizes acoustic physical realism into eight measurable dimensions across sound generation, propagation environment, and acoustic reception.

  • Dataset & Validated Evaluators: We curate 11,296 unique real-world A/V clips and 82,828 acoustically annotated RGB-D samples, using them to build and validate targeted evaluators and 605-prompt T2AV / 748-prompt I2AV suites.

  • Findings & Actionability: Across nine joint A/V generators, even strong models still violate fundamental acoustic relations. A range-guided intervention improves the targeted relation in 80.16% of valid samples, showing how diagnostic evidence can guide model refinement.

arXiv Code soonData soon
AcoustiTrace acoustic mirror illustration showing sound generation, propagation, and reception inconsistencies
Shiyang Li1,2Yuewen Cao2Yihao Liu2Yuandong Pu3Baochang Zhang4Xiaofei Li5Changqing Zou1,6
1Zhejiang University2Shanghai Artificial Intelligence Laboratory3Shanghai Jiao Tong University4Beihang University5School of Engineering, Westlake University6Zhejiang Lab

01 / Benchmark construction

From acoustic relations
to measurable suites.

AcoustiTrace asks whether a generated soundtrack follows the acoustic relations implied by visible events and scene conditions. Eight diagnostics trace those relations through sound generation, environmental propagation, and acoustic reception.

AcoustiTrace framework from acoustic process to eight evaluation dimensions
Benchmark framework. The same acoustic relations guide data curation, evaluator validation, and prompt-suite coverage, keeping every test within a measurable operating domain.

Why this matters. A soundtrack can sound plausible while violating one underlying mechanism; relation-specific scores reveal where that failure occurs.

11,296unique real-world A/V clips
82,828annotated RGB-D observations
605unique T2AV prompts
748unique I2AV prompts
Construction of the AcoustiTrace dataset from real-world audio-video anchors and acoustically annotated RGB-D observations
Two complementary evidence sources. 11,296 curated real-world A/V clips anchor observable relations; 82,828 RGB-D observations add material-aware absorption maps and 500 Hz RT60 targets for scene-acoustic supervision.
Sunburst distribution of the AcoustiTrace prompt suites across acoustic stages, dimensions, and scenarios
Eight relations, two task suites. Sector area represents prompt count across acoustic stages, diagnostic dimensions, and representative scenarios: 605 unique prompts for T2AV and 748 for I2AV.

How to read the counts. These totals count unique prompts, not evaluator assignments. Approach Gain and Lateral Stability share the receiver-motion pool; Causality Violation reuses eligible Motion–Loudness and Impact Decay cases; and six Range Attenuation prompts overlap with receiver motion.

02 / Evaluator validation

Three checks.
One diagnostic system.

Across all eight evaluators, we test recovery of expected relations, sensitivity to controlled violations, and agreement with expert judgments. Separate checks probe CPRS complementarity and RT60 transfer.

Controlled perturbation responses for eight evaluators, human-evaluator agreement, and comparison with PhyAVBench CPRS
Evaluator validation. Controlled violations, expert pairwise judgments, and a generic embedding-based baseline probe whether each score captures its intended acoustic relation.
Controlled sensitivity

Eight evaluators move as expected

As each targeted acoustic violation becomes more severe, its corresponding consistency score degrades in the expected direction.

Human alignment

87.9–93.5% agreement

Across 1,920 ratings of 64 generated A/V pairs from 30 expert raters, evaluator preferences closely match which output better satisfies the target relation.

Complementary signal

Diagnosis goes beyond similarity

Weak correlation with CPRS over 300 near–far pairs shows that embedding transitions and relation-specific acoustic diagnosis measure different properties.

Visual RT60 estimator validation on simulated data and audio-visual proxy agreement on STARSS23
Visual RT60 validation. Held-out simulation and STARSS23 compare the image-derived estimate with independent acoustic targets.

Takeaway. The selected Sabine-guided estimator reaches 0.0642 s MAE and Spearman ρ = 0.8698 on 16,563 simulated observations; across 26 valid STARSS23 clapping windows, the audio–visual proxy MAE is 0.0791 s.

BRAS CR3 RGB observation, estimated depth, acoustic absorption map, and measured reverberation comparison
Measured-room case study. BRAS CR3 provides an external check using a real room with an official measured reverberation reference.

Takeaway. The image-derived estimate is 1.1019 s at 500 Hz versus the official room-level T20 of 1.2584 s—a 0.1565 s difference. This is an illustrative single-room case, not an aggregate BRAS result.

03 / Benchmark results

Diagnostics across
the acoustic process.

We compare nine generators using conditional scores on valid outputs. Read each row as a diagnostic profile rather than a single ranking: leadership changes with both the acoustic relation and the generation task.

BestSecond
ModelRange AttenuationApproach GainLateral StabilityMotion–LoudnessImpact DecayCausality ViolationRT60 Consistency
LTX-2.381.674.192.288.590.086.541.1
JavisDiT++39.268.487.551.573.686.669.8
NAVA68.169.984.475.988.693.269.7
Ovi57.170.689.656.498.688.969.1
UniVerse-1
MOVA62.481.186.874.059.183.353.6
Veo 3.1official API71.888.781.875.198.698.673.5
Seedance 2.0official API86.571.376.684.091.499.224.3
Wan 2.7official API74.571.782.082.485.995.733.2
Task-dependent leadership

Four models lead different dimensions

LTX-2.3, Ovi, Veo 3.1, and Seedance 2.0 each lead at least one acoustic relation.

Mechanism-level pattern

Local plausibility is not full physics

Short, visibly explicit relations are often modeled more reliably than onset detail, sustained geometry, and scene-dependent reverberation.

Applicability-aware reporting

40.5–100% valid-output range

Scores are conditional on usable evidence. Median validity is 99.8% across model–evaluator cells, with RT60 showing the lowest coverage.

04 / Qualitative evidence

Qualitative
Acoustic-Physics Results.

Explore representative AcoustiTrace cases spanning all eight acoustic dimensions. Each generated video is paired with its diagnostic readout.

Model 01LTX-2.3
Range Attenuation91.5
Evaluator readoutToggle ↕LTX-2.3 Range Attenuation evaluator readout for example 1
Range Attenuation80.8
Evaluator readoutToggle ↕LTX-2.3 Range Attenuation evaluator readout for example 2
Range Attenuation58.3
Evaluator readoutToggle ↕LTX-2.3 Range Attenuation evaluator readout for example 3
Model 02Seedance 2.0
Range Attenuation99.2
Evaluator readoutToggle ↕Seedance 2.0 Range Attenuation evaluator readout for example 1
Range Attenuation90.4
Evaluator readoutToggle ↕Seedance 2.0 Range Attenuation evaluator readout for example 2
Range Attenuation87.6
Evaluator readoutToggle ↕Seedance 2.0 Range Attenuation evaluator readout for example 3
Model 03Veo 3.1
Range Attenuation90.9
Evaluator readoutToggle ↕Veo 3.1 Range Attenuation evaluator readout for example 1
Range Attenuation84.3
Evaluator readoutToggle ↕Veo 3.1 Range Attenuation evaluator readout for example 2
Range Attenuation73.1
Evaluator readoutToggle ↕Veo 3.1 Range Attenuation evaluator readout for example 3
Model 04MOVA
Range Attenuation93.1
Evaluator readoutToggle ↕MOVA Range Attenuation evaluator readout for example 1
Range Attenuation91.5
Evaluator readoutToggle ↕MOVA Range Attenuation evaluator readout for example 2
Range Attenuation91.5
Evaluator readoutToggle ↕MOVA Range Attenuation evaluator readout for example 3
Model 05NAVA
Range Attenuation86.8
Evaluator readoutToggle ↕NAVA Range Attenuation evaluator readout for example 1
Range Attenuation98.9
Evaluator readoutToggle ↕NAVA Range Attenuation evaluator readout for example 2
Range Attenuation97.6
Evaluator readoutToggle ↕NAVA Range Attenuation evaluator readout for example 3
Model 06Ovi
Range Attenuation95.2
Evaluator readoutToggle ↕Ovi Range Attenuation evaluator readout for example 1
Range Attenuation91.7
Evaluator readoutToggle ↕Ovi Range Attenuation evaluator readout for example 2
Range Attenuation91.5
Evaluator readoutToggle ↕Ovi Range Attenuation evaluator readout for example 3
Model 07Wan 2.7
Range Attenuation82.7
Evaluator readoutToggle ↕Wan 2.7 Range Attenuation evaluator readout for example 1
Range Attenuation76.7
Evaluator readoutToggle ↕Wan 2.7 Range Attenuation evaluator readout for example 2
Range Attenuation70.6
Evaluator readoutToggle ↕Wan 2.7 Range Attenuation evaluator readout for example 3

05 / From diagnosis to intervention

A failed relation
becomes an objective.

AcoustiTrace does more than rank models. A diagnosed Range Attenuation residual can guide audio denoising while the visual trajectory stays fixed.

Range-guided audio denoising method and aggregate results
Evaluator-guided refinement. A Range Attenuation residual supplies targeted guidance during audio denoising while the visual trajectory remains fixed.
Mean R²0.714 → 0.858no guidance → guided
Win rate80.16%126 paired outputs
Paired ΔR²+0.14495% CI [0.100, 0.189]

Single-model, single-relation proof of concept. The result demonstrates actionability without claiming a universal guidance method.