DipTych study: AI helps musicians surface differences, not make more accurate judgments

Diptych: Scoped, AI-Interpreted Comparison for Reference Listening in Music Production

Chongjun Zhong, Abhinaba Roy, Archishman Ghosh, Kejun Zhang, Dorien Herremans

cs.HC

2026-09-30

Musicians scope the comparison in DipTych; an LLM reads 59 audio features. 9 of 10 new differences drew at least partial blind-expert support; existing judgments' support held flat.

What problem this solves

Reference listening is a standard move in mixing and mastering: play a passage of your work in progress, play the matching moment in a commercial reference track, switch back and forth, decide what to change. Existing tools, iZotope Ozone and Logic Pro's analysis among them, compare whole tracks and report spectral and loudness differences. Two songs rarely line up, though: the chorus lands two minutes into one track and thirty seconds into the other. Holding your chorus against theirs is exactly what these tools cannot do, so the decision about what to compare gets made for you.

The paper splits reference listening into two layers: choosing what to compare, and reading the difference. The first belongs to the person. The second is harder to support than it looks. Auditory comparison leans on working memory, and while one passage plays, the impression of the other fades. Musicians routinely hear more than they can name. An audio LLM can write a fluent, technically correct account of two tracks while pointed at the wrong comparison, and accuracy does not rescue a mis-scoped verdict.

Method

The design fits in one line: the user sets the scope, the AI interprets within it.

The eight categories come from three independent sources converging: the dimensions mix engineers invoke when describing references, the facets the MIR community organizes content around, and taxonomies from production-knowledge engineering. A feature earns its place by giving the producer something specific and checkable to act on, such as vibrato rate, swing ratio, or integrated LUFS.

Results

Twelve musicians, with experience from one year to over ten, were split across four fixed target-reference pairs. They first recorded judgments with a bare A/B player, then revisited each one with the full system, marking it retained, revised, or withdrawn, and adding new claims. Four expert listeners did blind A/B listening and rated anonymized claims without knowing their source.

MeasureResult
Existing judgments fully supported by both experts (before vs. after)24/38 (63.2%) vs. 23/38 (60.5%)
At least partially supported by both, at each stage32/38 (84.2%)
New judgments after system use9 of 10 at least partially supported; 4 fully
System Usability Scalemean 74.79 (SD 9.26) vs. benchmark 68
Rated the system helpful11/12
Clearer about what to do next12/12
Trusted results very much8/12
Said the AI states uncertain things too confidently4/12

The key result is a split between confidence and accuracy. Of 35 judgments with complete confidence ratings, 15 gained confidence after system use, and 7 of those never reached full support from both experts. Post-use judgments ran longer with more technical vocabulary, citing BPM, stereo width, and transients; expert support did not follow. The failures were concrete: a singing voice and its register reported in instrumental tracks, a shuffle feel and brief key changes the experts did not hear. Only 2 of 12 felt pushed toward the reference.

Why it matters

The value sits mostly in the evaluation design, not the tool. Perceived helpfulness and judgment accuracy are different things, and the study measured them separately and found them decoupled: a SUS of 74.79 and 11 of 12 calling it helpful, while expert support for existing judgments stayed flat. Teams building AI feedback products can lift this evaluation wholesale; high satisfaction does not mean user judgment is improving.

One principle came out of the data: pointing earns more trust than concluding. The features participants valued most were segment markers, time positions, and A/B switching, which are correct as long as they take you to the right spot. A written verdict can be wrong about the music, and that is where every trust problem appeared. Be confident about where to listen and careful about what a difference means. Video editing against source footage and writing against model texts, the analogues the paper names, point the same way. As a system, this is an incremental prototype.

Limitations

From the authors: expert ratings leave room for individual interpretation and measure support within their rubric, not ground truth; the 59 features vary in reliability, were drawn from the literature rather than from formative work with producers, and the study never established which ones carry the comparison; the single-session design evaluates exposure to the full system and cannot isolate the AI component.

Reading closely raises more. With 12 participants and four track pairs, a shift from 63.2% to 60.5% says nothing about direction. Of the 10 new judgments, only 4 were fully supported by both experts and 1 was rated unsupported, so finding more also meant finding more weakly grounded claims, which the authors concede. Nobody used the system to decide anything about their own music; being clearer on next steps is self-report with no behavioral follow-up.

Terms

Source

Related papers

All paper explainers