Overview

Cross-modal music generation from text and visual inputs is uniquely challenging due to its subjective, one-to-many nature. Current methods face three major limitations: they are bottlenecked by scarce paired data, lack explicit one-to-many modeling, and isolate generation from retrieval.

MERIC proposes a unified framework for multimodal music generation and retrieval. Its core idea is a semantic embedding space (a music semantic anchor), which decouples multimodal understanding from acoustic synthesis. The paper uses a two-stage training pipeline:

  1. Stage 1: a Flow-DiT decoder learns to synthesize audio from music semantic embeddings using large unpaired music corpora.
  2. Stage 2: the diffusion-based Music Head learns to map frozen Qwen3-VL embeddings into the music semantic space using paired multimodal-music data.

At inference time, the predicted anchor embedding is both a condition for acoustic synthesis and a query for music retrieval. This is the central reason generation, retrieval, one-to-many sampling, and the ablations below belong on one project page rather than separate demos.

Framework overview: multimodal queries (image, video, text) are encoded by a frozen Qwen3-VL-Embedding model. The Music Head maps query embeddings to a music semantic anchor, which serves as the interface for both generation (Flow-DiT Decoder) and retrieval (cosine similarity against pre-indexed music embeddings).

Paper Results

The demo samples below are meant to be heard next to the paper numbers, so the main generation and retrieval tables are surfaced here before the galleries. Lower FAD/KL is better; higher Sem and retrieval recall are better.

Vision Generation 0.127 FADC

MERIC on MelBench, below all listed generation baselines.

Text Generation 0.110 FADC

MERIC on MusicCaps, using the same anchor pipeline.

Retrieval One anchor, two tasks

The predicted semantic embedding is both a synthesis condition and a retrieval query.

Vision-to-Music Generation

Main baselines are evaluated on the full test set. Rows marked † are 2026 song-foundation cascade baselines evaluated on fixed N=500 subsets; their Sem scores are not directly comparable to full-test-set Sem.

MethodMelBenchMuImageEmoMVARIA
FADMFADCFADVKLSemFADMFADCFADVKLSemFADMFADCFADVKLSemFADMFADCFADVKLSem
Art2Mus31.60.94513.70.014539.130.80.90215.30.015145.333.91.10914.60.014744.315.10.5607.020.007745.3
GVMGen3.550.4835.780.007252.93.780.5498.730.008649.63.960.6177.200.005763.64.180.3375.830.005260.0
AudioX16.50.72313.70.012753.813.90.62314.40.013051.84.750.4875.450.007179.020.10.76611.70.012038.5
M2UGen7.530.5587.140.009055.35.190.4218.660.009857.89.510.79410.10.008754.624.00.7444.890.007057.1
Qwen3VL + MusicGen4.420.4976.560.007680.62.990.4188.930.007985.67.080.7367.650.008380.35.090.2832.040.005976.6
Qwen3VL + AudioLDM2-L6.570.4085.920.007877.64.590.2746.080.007987.28.610.6097.410.008678.79.950.2623.660.006977.6
Ours1.550.1271.270.005974.71.790.1693.200.007887.13.060.2532.450.006081.03.300.2472.080.005579.6
2026 song-foundation cascade baselines (N=500 subset) †
MERIC (N=500) †1.960.156——80.22.240.183——88.13.320.272——73.79.230.496——65.2
YuE-7B †12.830.518——61.58.930.384——69.815.210.710——68.619.250.657——64.4
HeartMuLa-3B †11.350.534——51.210.640.705——48.214.610.711——53.016.060.819——60.2

Text-to-Music Generation

MethodMusicCapsARIA
FADMFADCSemFADMFADCSem
MusicGen-Medium3.800.42165.13.000.27271.1
MusicGen-Large3.510.40967.12.550.24872.4
AudioLDM2-Music4.600.26068.04.630.25168.6
AudioLDM2-Large4.170.17876.73.930.21177.5
Stable-Audio-Open3.850.26868.33.500.32457.8
Ours1.870.11074.22.140.15574.5

Vision-Music Retrieval (%)

MelBench: 5,000 one-to-one pairs. ARIA: 1,000 queries → ~9,000 audio candidates (one-to-many). Higher is better.

MethodMelBenchARIA
I→MM→II→MM→I
R@1R@10R@1R@10R@1R@10R@1R@10
ImageBind4.3818.003.6617.000.302.300.131.77
Ours1.348.621.129.540.402.800.392.40

Text-Music Retrieval (%)

CLAP is trained on 630K text-audio pairs. Ours routes text through Qwen3VLE → Music Head → anchor space without any text-audio training. Higher is better.

MethodMusicCapsARIA
T→MM→TT→MM→T
R@1R@10R@1R@10R@1R@10R@1R@10
CLAP4.3119.485.5423.530.172.300.633.73
Ours3.2718.802.8216.420.835.470.736.40

Unified Retrieval

MERIC does not add a separate retrieval model. The Music Head predicts an anchor in the shared semantic embedding space from a visual or text query; generation conditions the Flow-DiT decoder on that anchor, while retrieval ranks pre-indexed music embeddings by cosine similarity to the same anchor.

Image / Text / Video
Qwen3-VL Embedding
Semantic Anchor Space
Generate or Retrieve

Vision-to-Music Retrieval

For images and video frames, the same anchor used to drive synthesis can query an indexed music collection. This is why the demo keeps generation and comparison examples close to the retrieval table instead of treating them as separate systems.

ARIA I→M R@10: 2.80 vs ImageBind 2.30.

Text-to-Music Retrieval

Text follows the same route: text → Qwen3-VL → Music Head → semantic anchor. The paper result is notable because MERIC does this without direct text-audio training, unlike CLAP.

ARIA T→M R@10: 5.47 vs CLAP 2.30.

One-to-Many Generation

The same visual input can inspire diverse yet equally valid compositions. The Music Head introduces controlled stochasticity: different random seeds yield different music semantic predictions, producing diverse audio outputs from a single image. Use Play All to hear each variation sequentially.

1 / 22

Vision-to-Music Comparison

Compare audio generated by our method against baselines across multiple benchmarks. Click any play button to listen; only one track plays at a time. Use Play All to hear each method sequentially.

1 / 20

Text-to-Music Comparison

Compare MERIC against dedicated text-to-music baselines on MusicCaps. MERIC reaches these samples through the shared semantic-anchor route rather than direct text-audio training. Click any play button to listen; only one track plays at a time. Use Play All to hear each method sequentially.

1 / 10

Ablation Studies

Audio comparisons for key design choices. Click any player to listen; compare across conditions for the same input image.

Sampling Steps

The Music Head uses DDIM sampling to iteratively refine noise into a music semantic embedding. A single denoising step (s=1) produces collapsed, low-diversity embeddings; s≥5 are perceptually indistinguishable (p>0.2 in human MOS study), though s=20 best matches the ground-truth distribution.

Human MOS by DDIM steps
Human MOS ratings (n=100 images × 4 steps)
Cosine similarity distribution
Intra-group pairwise cosine similarity
Input image for MERIC sampling-steps ablation
s=1
s=5
s=10
s=20
Input image for MERIC sampling-steps ablation
s=1
s=5
s=10
s=20
Input image for MERIC sampling-steps ablation
s=1
s=5
s=10
s=20
Input image for MERIC sampling-steps ablation
s=1
s=5
s=10
s=20
Input image for MERIC sampling-steps ablation
s=1
s=5
s=10
s=20
Input image for MERIC sampling-steps ablation
s=1
s=5
s=10
s=20
Input image for MERIC sampling-steps ablation
s=1
s=5
s=10
s=20
Input image for MERIC sampling-steps ablation
s=1
s=5
s=10
s=20
Input image for MERIC sampling-steps ablation
s=1
s=5
s=10
s=20
Input image for MERIC sampling-steps ablation
s=1
s=5
s=10
s=20
Input image for MERIC sampling-steps ablation
s=1
s=5
s=10
s=20
Input image for MERIC sampling-steps ablation
s=1
s=5
s=10
s=20
Input image for MERIC sampling-steps ablation
s=1
s=5
s=10
s=20
Input image for MERIC sampling-steps ablation
s=1
s=5
s=10
s=20

Projector Architecture

Five projector architectures map frozen vision-encoder embeddings to the music anchor space. Linear (Contrastive) achieves the best retrieval but catastrophic generation quality (FAD 6.6× worse); Linear (MSE) produces good FAD but collapses retrieval by mapping to the conditional mean; VAE methods strike a moderate balance but share MSE’s deterministic collapse. Only the Music Head (diffusion-based) resolves this trade-off—best FAD with competitive retrieval—because iterative denoising produces predictions that are individually accurate yet collectively diverse.

PairCos vs FAD trade-off
Projectors in PairCos–FAD space: only the Music Head achieves both GT-level diversity and low FAD.
Input image for MERIC projector-architecture ablation
Ground Truth
Music Head
Linear (MSE)
Linear (Contrastive)
Bridge VAE
Cross VAE
Input image for MERIC projector-architecture ablation
Ground Truth
Music Head
Linear (MSE)
Linear (Contrastive)
Bridge VAE
Cross VAE
Input image for MERIC projector-architecture ablation
Ground Truth
Music Head
Linear (MSE)
Linear (Contrastive)
Bridge VAE
Cross VAE
Input image for MERIC projector-architecture ablation
Ground Truth
Music Head
Linear (MSE)
Linear (Contrastive)
Bridge VAE
Cross VAE
Input image for MERIC projector-architecture ablation
Ground Truth
Music Head
Linear (MSE)
Linear (Contrastive)
Bridge VAE
Cross VAE
Input image for MERIC projector-architecture ablation
Ground Truth
Music Head
Linear (MSE)
Linear (Contrastive)
Bridge VAE
Cross VAE
Input image for MERIC projector-architecture ablation
Ground Truth
Music Head
Linear (MSE)
Linear (Contrastive)
Bridge VAE
Cross VAE

Vision Encoder

Replacing CLIP ViT-H-14 (1024-d) with Qwen3-VL-Embedding (2048-d) consistently improves generation fidelity (FADMERT: 2.54 → 1.95 on MelBench), confirming that richer vision-language representations benefit cross-modal translation even with a frozen backbone. Both use the same Music Head projector.

Input image for MERIC vision-encoder ablation
CLIP ViT-H-14
Qwen3VL-Embedding
Input image for MERIC vision-encoder ablation
CLIP ViT-H-14
Qwen3VL-Embedding
Input image for MERIC vision-encoder ablation
CLIP ViT-H-14
Qwen3VL-Embedding
Input image for MERIC vision-encoder ablation
CLIP ViT-H-14
Qwen3VL-Embedding
Input image for MERIC vision-encoder ablation
CLIP ViT-H-14
Qwen3VL-Embedding
Input image for MERIC vision-encoder ablation
CLIP ViT-H-14
Qwen3VL-Embedding

Direct Conditioning

A direct Stage-2 baseline that conditions the decoder on Qwen3VL features without the shared semantic anchor trails MERIC on both FAD-MERT and semantic alignment. More importantly, it is an audio-only generator: it cannot produce a retrieval query in the music-anchor space, so it does not unify generation and retrieval.

SubsetAnchored FADMDirect FADMAnchored SemDirect Sem
ARIA9.1411.4070.569.4
MuImage2.232.2387.783.6
MelBench1.973.0480.069.4

Noise Objective in Music Head

Diffusion and rectified flow are statistically indistinguishable when the anchor-space mapping is held fixed. This suggests the gains come from the anchor design rather than the specific generative formulation.

SubsetRDM FADCFM FADCRDM FADMFM FADM
ARIA0.4960.4949.239.14
MuImage0.1840.1842.242.23
MelBench0.1560.1551.961.97

Music-Anchor Encoder

Replacing the frozen music-semantic encoder shows that music-semantic pretraining is essential. CLAP, a general contrastive audio-text model, is an outlier even on its own aligned FAD-CLAP metric; The default music encoder and MERT perform similarly, confirming that the anchor design matters more than the specific encoder.

MetricMERIC (default)CLAPMERT
FADCLAP (unbiased)0.1780.2860.172
FADMERT (aligned)1.312.351.01
Human Sem (N=50)62.8437.5268.20

Citation

The official ECCV 2026 camera-ready paper is available below. Source code is Apache-2.0; checkpoints, pretrained components, and datasets follow their own upstream terms.

@inproceedings{wang2026meric,
  title     = {Meric: A Unified Framework for Multimodal Music Generation and Retrieval via Representation Space Anchoring},
  author    = {Xihua Wang and Yinbo Wang and Jingchao Zhang and Ruihua Song},
  booktitle = {European Conference on Computer Vision (ECCV)},
  year      = {2026},
  url       = {https://media.eventhosts.cc/Conferences/ECCV2026/pdfs/2169.pdf}
}