Three Papers Accepted for ISMIR 2026

We are pleased to announce that three research papers involving Yamaha researchers have been accepted to The 27th International Society for Music Information Retrieval Conference (ISMIR 2026), which will be held in November 2026.

ISMIR is one of the leading international conferences in the field of music information retrieval, bringing together researchers and practitioners working on a wide range of topics at the intersection of music and technology, including music analysis, generation, automatic music transcription, and dataset development.

Through research in music information processing, artificial intelligence, and performance motion analysis, Yamaha aims to deepen the understanding of sound and music and create new value that supports diverse forms of musical expression and enjoyment. By presenting these studies at ISMIR 2026, Yamaha will share its latest research with the global academic community, strengthen connections with researchers and practitioners around the world, and explore future opportunities for collaboration and value creation.

Author names in bold are Yamaha researchers.

Learning Jazz Pianist Style with Cross-Attention Conditioning

Andrew Edwards¹, Akira Maezawa, Simon Dixon¹ (¹ Queen Mary University of London)

This research investigates how the distinctive styles of jazz pianists can be learned and analyzed using advanced music representation models. The proposed approach enables music generation conditioned on a specific performer’s style and demonstrates that performer identity can be captured with high accuracy. The study also introduces a method for identifying the musical passages most characteristic of an individual pianist, providing new tools for the analysis of musical expression and creative applications.

Low-Latency Real-Time Automatic Music Transcription with Band-Split RNN and Temporal Lookahead

Yuta Kusaka, Akira Maezawa,

This research proposes a new real-time automatic music transcription architecture that addresses the long-standing challenge of balancing transcription accuracy and latency. By introducing an RNN-based design that reduces reliance on future information, the method achieves fast and accurate transcription of piano and guitar performances. Experimental results demonstrated an improved accuracy–latency trade-off compared with conventional approaches, supporting practical real-time music applications such as performance support and music production.

SKY Piano: A Multimodal Piano Playing Dataset

Joonhyung Bae², Daewon Park³, Taegyun Kwon², Yoon-Seok Choi³, Hyeon Hur³, Shigeru Kai, Yohei Wada, Satoshi Obata, Yu Takahashi, Akira Maezawa, Jaebum Park³, Jonghwa Park³ (² Korea Advanced Institute of Science and Technology, ³ Seoul National University)

We introduce SKY-Piano, a large-scale multimodal piano performance dataset that combines audio, MIDI, motion capture, multiview video, and score data from both professional and amateur pianists. In addition to providing rich performance data across multiple modalities, the dataset includes tools for fingering annotation and an interactive browser for exploring synchronized recordings. By enabling research on performance analysis, motion generation, and the relationship between human movement and musical expression, SKY-Piano aims to advance music information processing, performance support technologies, and future collaborative research in music and performance science.