AnchorShared semantic embedding space as the music anchor
TrainingStage 1 learns from 630K unpaired music tracks
SamplingOne input can yield many faithful compositions
UnificationThe predicted anchor serves generation and retrieval
Overview
Cross-modal music generation from text and visual inputs is uniquely challenging due to its subjective, one-to-many nature.
Current methods face three major limitations: they are bottlenecked by scarce paired data,
lack explicit one-to-many modeling, and isolate generation from retrieval.
MERIC proposes a unified framework for multimodal music generation and retrieval.
Its core idea is a semantic embedding space (a music semantic anchor), which decouples
multimodal understanding from acoustic synthesis. The paper uses a two-stage training pipeline:
Stage 1: a Flow-DiT decoder learns to synthesize audio from music semantic embeddings using large unpaired music corpora.
Stage 2: the diffusion-based Music Head learns to map frozen Qwen3-VL embeddings into the music semantic space using paired multimodal-music data.
At inference time, the predicted anchor embedding is both a condition for acoustic synthesis and a query
for music retrieval. This is the central reason generation, retrieval, one-to-many sampling, and the
ablations below belong on one project page rather than separate demos.
Paper Results
The demo samples below are meant to be heard next to the paper numbers, so the main generation
and retrieval tables are surfaced here before the galleries. Lower FAD/KL is better; higher Sem
and retrieval recall are better.
Vision Generation0.127 FADC
MERIC on MelBench, below all listed generation baselines.
Text Generation0.110 FADC
MERIC on MusicCaps, using the same anchor pipeline.
RetrievalOne anchor, two tasks
The predicted semantic embedding is both a synthesis condition and a retrieval query.
Vision-to-Music Generation
Main baselines are evaluated on the full test set. Rows marked † are 2026 song-foundation cascade baselines evaluated on fixed N=500 subsets; their Sem scores are not directly comparable to full-test-set Sem.
CLAP is trained on 630K text-audio pairs. Ours routes text through Qwen3VLE → Music Head → anchor space without any text-audio training. Higher is better.
Method
MusicCaps
ARIA
T→M
M→T
T→M
M→T
R@1
R@10
R@1
R@10
R@1
R@10
R@1
R@10
CLAP
4.31
19.48
5.54
23.53
0.17
2.30
0.63
3.73
Ours
3.27
18.80
2.82
16.42
0.83
5.47
0.73
6.40
Unified Retrieval
MERIC does not add a separate retrieval model. The Music Head predicts an anchor in the shared semantic embedding space from
a visual or text query; generation conditions the Flow-DiT decoder on that anchor, while retrieval
ranks pre-indexed music embeddings by cosine similarity to the same anchor.
Image / Text / Video
→
Qwen3-VL Embedding
→
Semantic Anchor Space
→
Generate or Retrieve
Vision-to-Music Retrieval
For images and video frames, the same anchor used to drive synthesis can query an indexed
music collection. This is why the demo keeps generation and comparison examples close to the
retrieval table instead of treating them as separate systems.
ARIA I→M R@10: 2.80 vs ImageBind 2.30.
Text-to-Music Retrieval
Text follows the same route: text → Qwen3-VL → Music Head → semantic anchor. The paper result is
notable because MERIC does this without direct text-audio training, unlike CLAP.
ARIA T→M R@10: 5.47 vs CLAP 2.30.
Image-to-Music Gallery
Curated visual inputs paired with music generated by MERIC. These examples make the anchor-based
image-to-music behavior listenable before the controlled baseline comparisons.
One-to-Many Generation
The same visual input can inspire diverse yet equally valid compositions.
The Music Head introduces controlled stochasticity: different random seeds
yield different music semantic predictions, producing diverse audio outputs from a single image.
Use Play All to hear each variation sequentially.
1 / 22
Text-to-Music Gallery
Text prompts are routed through the same Qwen3-VL → Music Head → semantic anchor path used by visual inputs.
MERIC is not trained as a separate direct text-audio model.
Vision-to-Music Comparison
Compare audio generated by our method against baselines across multiple benchmarks.
Click any play button to listen; only one track plays at a time.
Use Play All to hear each method sequentially.
1 / 20
Text-to-Music Comparison
Compare MERIC against dedicated text-to-music baselines on MusicCaps. MERIC reaches these samples
through the shared semantic-anchor route rather than direct text-audio training.
Click any play button to listen; only one track plays at a time.
Use Play All to hear each method sequentially.
1 / 10
Ablation Studies
Audio comparisons for key design choices. Click any player to listen; compare across conditions for the same input image.
Sampling Steps
The Music Head uses DDIM sampling to iteratively refine noise into a music semantic embedding.
A single denoising step (s=1) produces collapsed, low-diversity embeddings; s≥5 are perceptually
indistinguishable (p>0.2 in human MOS study), though s=20 best matches the ground-truth
distribution.
Human MOS ratings (n=100 images × 4 steps)Intra-group pairwise cosine similarity
s=1
s=5
s=10
s=20
s=1
s=5
s=10
s=20
s=1
s=5
s=10
s=20
s=1
s=5
s=10
s=20
s=1
s=5
s=10
s=20
s=1
s=5
s=10
s=20
s=1
s=5
s=10
s=20
s=1
s=5
s=10
s=20
s=1
s=5
s=10
s=20
s=1
s=5
s=10
s=20
s=1
s=5
s=10
s=20
s=1
s=5
s=10
s=20
s=1
s=5
s=10
s=20
s=1
s=5
s=10
s=20
Projector Architecture
Five projector architectures map frozen vision-encoder embeddings to the music anchor space.
Linear (Contrastive) achieves the best retrieval but catastrophic generation quality
(FAD 6.6× worse); Linear (MSE) produces good FAD but collapses retrieval
by mapping to the conditional mean; VAE methods strike a moderate balance but share MSE’s
deterministic collapse. Only the Music Head (diffusion-based) resolves this
trade-off—best FAD with competitive retrieval—because iterative denoising produces
predictions that are individually accurate yet collectively diverse.
Projectors in PairCos–FAD space: only the Music Head achieves both GT-level diversity and low FAD.
Ground Truth
Music Head
Linear (MSE)
Linear (Contrastive)
Bridge VAE
Cross VAE
Ground Truth
Music Head
Linear (MSE)
Linear (Contrastive)
Bridge VAE
Cross VAE
Ground Truth
Music Head
Linear (MSE)
Linear (Contrastive)
Bridge VAE
Cross VAE
Ground Truth
Music Head
Linear (MSE)
Linear (Contrastive)
Bridge VAE
Cross VAE
Ground Truth
Music Head
Linear (MSE)
Linear (Contrastive)
Bridge VAE
Cross VAE
Ground Truth
Music Head
Linear (MSE)
Linear (Contrastive)
Bridge VAE
Cross VAE
Ground Truth
Music Head
Linear (MSE)
Linear (Contrastive)
Bridge VAE
Cross VAE
Vision Encoder
Replacing CLIP ViT-H-14 (1024-d) with Qwen3-VL-Embedding (2048-d) consistently improves
generation fidelity (FADMERT: 2.54 → 1.95 on MelBench), confirming that richer
vision-language representations benefit cross-modal translation even with a frozen backbone.
Both use the same Music Head projector.
CLIP ViT-H-14
Qwen3VL-Embedding
CLIP ViT-H-14
Qwen3VL-Embedding
CLIP ViT-H-14
Qwen3VL-Embedding
CLIP ViT-H-14
Qwen3VL-Embedding
CLIP ViT-H-14
Qwen3VL-Embedding
CLIP ViT-H-14
Qwen3VL-Embedding
Direct Conditioning
A direct Stage-2 baseline that conditions the decoder on Qwen3VL features without the shared semantic anchor
trails MERIC on both FAD-MERT and semantic alignment. More importantly, it is an audio-only generator:
it cannot produce a retrieval query in the music-anchor space, so it does not unify generation and retrieval.
Subset
Anchored FADM
Direct FADM
Anchored Sem
Direct Sem
ARIA
9.14
11.40
70.5
69.4
MuImage
2.23
2.23
87.7
83.6
MelBench
1.97
3.04
80.0
69.4
Noise Objective in Music Head
Diffusion and rectified flow are statistically indistinguishable when the anchor-space mapping is held fixed.
This suggests the gains come from the anchor design rather than the specific generative formulation.
Subset
RDM FADC
FM FADC
RDM FADM
FM FADM
ARIA
0.496
0.494
9.23
9.14
MuImage
0.184
0.184
2.24
2.23
MelBench
0.156
0.155
1.96
1.97
Music-Anchor Encoder
Replacing the frozen music-semantic encoder shows that music-semantic pretraining is essential.
CLAP, a general contrastive audio-text model, is an outlier even on its own aligned FAD-CLAP metric;
The default music encoder and MERT perform similarly, confirming that the anchor design matters more than the specific encoder.
Metric
MERIC (default)
CLAP
MERT
FADCLAP (unbiased)
0.178
0.286
0.172
FADMERT (aligned)
1.31
2.35
1.01
Human Sem (N=50)
62.84
37.52
68.20
Citation
The official ECCV 2026 camera-ready paper is available below.
Source code is Apache-2.0; checkpoints, pretrained
components, and datasets follow their own upstream terms.
@inproceedings{wang2026meric,
title = {Meric: A Unified Framework for Multimodal Music Generation and Retrieval via Representation Space Anchoring},
author = {Xihua Wang and Yinbo Wang and Jingchao Zhang and Ruihua Song},
booktitle = {European Conference on Computer Vision (ECCV)},
year = {2026},
url = {https://media.eventhosts.cc/Conferences/ECCV2026/pdfs/2169.pdf}
}