Teaching Popular Music Analysis with Source Separation
DOI:
https://doi.org/10.15763/issn.2994-7073.2026.39.59-84Keywords:
artificial intelligence, music analysis, music source separationAbstract
Music source separation is a technology that converts a recorded song into multiple files that contain different components of the original track. Typically, source separation tools split a recording into four discrete components: vocals, percussion, bass, and “other.” Methods for isolating components of a recording have been proposed and experimented with for decades, with varying degrees of success. In the late 2010s, developers began using neural networks, a form of artificial intelligence, to design source separation methods. Since then, the performance and availability of source separation models has rapidly improved.
In this article, I aim to demystify and contextualize source separation to explore the ways in which it could impact the study of music analysis while encouraging productive discussions on the relationship between music analysis and artificial intelligence. Access to high-quality isolated instrumental tracks has several useful applications in the study of popular music: consider for instance the way an instructor could use the technology to design specific class activities that focus on only a portion of a track. Source separation, moreover, can play a central role in discussions on notation, musical structure, and the role of technology in music analysis. More importantly, it is an ideal tool for illustrating some of the ethical considerations—transparency, bias, and accessibility—that arise when using AI-based technologies.
Downloads
References
Abrahams, Rosa. 2021. “Rethinking Music Literacy in the Undergraduate Theory Core.” Journal of Music Theory Pedagogy 35: 81–108. https://digitalcollections.lipscomb.edu/jmtp/vol35/iss1/2/. DOI: https://doi.org/10.71156/2994-7073.1223
Araki, Shoko, Nobutaka Ito, Reinhold Haeb-Umbach, Gordon Wichern, Zhong-Qiu Wang, and Yuki Mitsufuji. 2025. “30+ Years of Source Separation Research: Achievements and Future Challenges.” https://doi.org/10.48550/arxiv.2501.11837. DOI: https://doi.org/10.1109/ICASSP49660.2025.10889006
Barna, Alyssa. 2024. “‘Duet Me’: Music Theory Pedagogy in the Age of Social Media.” Journal of Music Theory Pedagogy 38: 21–44. https://digitalcollections.lipscomb.edu/jmtp/vol38/iss1/3/. DOI: https://doi.org/10.71156/2994-7073.1457
Bittner, Rachel, Justin Salamon, Mike Tierney, Matthias Mauch, Chris Cannam, and Juan Bello. 2014. “MedleyDB: A Multitrack Dataset for Annotation-Intensive MIR Research.” In Proceedings of the 15th International Society for Music Information Retrieval (ISMIR) Conference, 155–60. https://doi.org/10.5281/zenodo.1417889.
Coote, Jonathan. 2024. “‘Stem-separating’ AI is Revolutionising the Music Industry—But at What Cost?” World Intellectual Property Review (WPIR), March 26. https://www.worldipreview.com/future-of-ip/stem-separating-ai-is-revolutionising-the-music-industry-but-at-what-cost.
Danaher, Matthew. 2022. “AI Stem Extraction: A Creative Tool or Facilitator of Mass Infringement?” JTIP Blog, May 3. https://jtip.law.northwestern.edu/2022/05/03/ai-stem-extraction-acreative-tool-or-facilitator-of-mass-infringement/.
de Clercq, Trevor. 2024. “Some Proposed Enhancements to the Operationalization of Prominence: Commentary on Michèle Duguay’s ‘Analyzing Vocal Placement in Recorded Virtual Space’.” Music Theory Online 30, no. 1. https://mtosmt.org/issues/mto.24.30.1/mto.24.30.1.declercq.html. DOI: https://doi.org/10.30535/mto.30.1.2
Devaney, Johanna. 2022. “Digital Audio Processing Tools for Music Corpus Studies.” In The Oxford Handbook of Music and Corpus Studies, edited by Daniel Shanahan, John Ashley Burgoyne, and Ian Quinn. Oxford Academic. https://doi.org/10.1093/oxfordhb/9780190945442.013.11. DOI: https://doi.org/10.1093/oxfordhb/9780190945442.013.11
Duguay, Michèle. 2022. “Analyzing Vocal Placement in Recorded Virtual Space.” Music Theory Online 28, no 4. https://mtosmt.org/issues/mto.22.28.4/mto.22.28.4.duguay.html. DOI: https://doi.org/10.30535/mto.28.4.1
Fabbro, Giorgio, Stefan Uhlich, Chieh-Hsin Lai, Woosung Choi, Marco Martínez-Ramírez, Weihsiang Liao, Igor Gadelha, et al. 2024. “The Sound Demixing Challenge 2023 – Music Demixing Track.” Transactions of the International Society for Music Information Retrieval 7, no. 1: 63–84. https://doi.org/10.5334/tismir.171. DOI: https://doi.org/10.5334/tismir.171
FitzGerald, Derry. 2012. “Vocal Separation Using Nearest Neighbours and Median Filtering.” In 23rd IET Irish Signals and Systems Conference. https://doi.org/10.1049/ic.2012.0225. DOI: https://doi.org/10.1049/ic.2012.0225
Heidemann, Kate. 2016. “A System for Describing Vocal Timbre in Popular Song.” Music Theory Online 22, no. 1. https://www.mtosmt.org/issues/mto.16.22.1/mto.16.22.1.heidemann.html. DOI: https://doi.org/10.30535/mto.22.1.2
Hennequin, Romain, Anis Khlif, Felix Voituret, and Manuel Moussallam. 2020. “Spleeter: A Fast and Efficient Music Source Separation Tool with Pre-Trained Models.” Journal of Open Source Software 5 (50): 2154. https://doi.org/10.21105/joss.02154. DOI: https://doi.org/10.21105/joss.02154
Japanese Breakfast. 2021. “Paprika.” Track 1 on Jubilee. DOC225, CD.
Lavengood, Megan L. 2020. “The Cultural Significance of Timbre Analysis: A Case Study in 1980s Pop Music, Texture, and Narrative.” Music Theory Online 26, no. 3. https://mtosmt.org/issues/ DOI: https://doi.org/10.30535/mto.26.3.3
Duguay – Teaching Popular Music Analysis with Source Separation 83 mto.20.26.3/mto.20.26.3.lavengood.html.
Lee, Jun-Yong, and Hyoung-Gook Kim. 2005. “Music and Voice Separation Using Logspectral Amplitude Estimator Based on Kernel Spectrogram Models Backfitting.” Journal of the Acoustical Society of Korea 34, no. 3: 227–33. https://doi.org/10.7776/ASK.2015.34.3.227. DOI: https://doi.org/10.7776/ASK.2015.34.3.227
Malawey, Victoria. 2020. A Blaze of Light in Every Word: Analyzing the Popular Singing Voice. Oxford University Press. DOI: https://doi.org/10.1093/oso/9780190052201.001.0001
Manabe, Noriko. 2019. “We Gon’ Be Alright? The Ambiguities of Kendrick Lamar’s Protest Anthem.” Music Theory Online 25, no. 1. https://mtosmt.org/issues/mto.19.25.1/mto.19.25.1.manabe.html. DOI: https://doi.org/10.30535/mto.25.1.9
Mauch, Matthias, and Simon Dixon. 2014. “PYIN: A Fundamental Frequency Estimator Using Probabilistic Threshold Distributions.” In 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 659–63. https://doi.org/10.1109/ICASSP.2014.6853678. DOI: https://doi.org/10.1109/ICASSP.2014.6853678
Mitsufuji, Yuki, Giorgio Fabbro, Stefan Uhlich, Fabian-Robert Stöter, Alexandre Défossez, Minseok Kim, Woosung Choi, Chin-Yun Yu, and Kin-Wai Cheuk. 2022. “Music Demixing Challenge 2021.” Frontiers in Signal Processing 1. https://doi.org/10.3389/frsip.2021.808395. DOI: https://doi.org/10.3389/frsip.2021.808395
Monét, Victoria. 2023. “Smoke.” Featuring Lucky Daye. Track 1 on Jaguar II. RCA–19658 84907 2, CD.
Moore, Allan F. 2012. Song Means: Analysing and Interpreting Recorded Popular Song. Ashgate.
Morreale, Fabio, Megha Sharma, and I-Chieh Wei, “Data Collection in Music Generation Training Sets: A Critical Analysis.” 2023. In Proceedings of the 24th Int. Society for Music Information Retrieval (ISMIR) Conference. https://hdl.handle.net/2292/65322.
Moussallam, Manuel, Gaël Richard, and Laurent Daudet. 2012. “Audio Source Separation Informed by Redundancy With Greedy Multiscale Decompositions.” In Proceedings of the 20th European Signal Processing Conference (EUSIPCO 2012), 2644–48.
Nobile, Drew F. 2015. “Counterpoint in Rock Music: Unpacking the ‘Melodic-Harmonic Divorce’.” Music Theory Spectrum 37, no. 2: 189–203. DOI: https://doi.org/10.1093/mts/mtv019
—————. 2022. “Alanis Morissette’s Voices.” Music Theory Online 28, no. 4. https://mtosmt.org/issues/mto.22.28.4/mto.22.28.4.nobile.html. DOI: https://doi.org/10.30535/mto.28.4.6
O’Hara, William. 2025. “‘Let’s Think in Layers’: On Twenty-First-Century Instruments of Public Music Theory.” Music Theory Spectrum 47, no. 1: 83–90. DOI: https://doi.org/10.1093/mts/mtae029
Peeters, Geoffroy, Bruno L. Giordano, Patrick Susini, Nicolas Misdariis, and Stephen McAdams. 2011. “The Timbre Toolbox: Extracting Audio Descriptors from Musical Signals.” The Journal of The Acoustical Society of America 130, no. 5: 2902–16. https://doi.org/10.1121/1.3642604. DOI: https://doi.org/10.1121/1.3642604
Pierard, Tom, and David Lines. 2022. “A Constructivist Approach to Music Education with DAWs.” Teachers and Curriculum 22, no. 2: 135–45. https://doi.org/10.15663/tandc.v22i2.406. DOI: https://doi.org/10.15663/tandc.v22i2.406
Rafii, Zafar, Antoine Liutkus, and Bryan Pardo. 2014. “REPET for Background/Foreground Separation in Audio.” In Blind Source Separation: Advances in Theory, Algorithms, and Applications, edited by Ganesh R. Naik and Wenwu Wang, 395–411. Springer. DOI: https://doi.org/10.1007/978-3-642-55016-4_14
Rafii, Zafar, Antoine Liutkus, Fabian-Robert Stöter, Stylianos Ioannis Mimilakis, and Rachel Bittner. 2017. “The MUSDB18 Corpus for Music Separation.” https://hal.inria.fr/hal-02190845/document.
Rafii, Zafar, Antoine Liutkus, Fabian Robert Stöter, Stylianos Ioannis Mimilakis, Derry FitzGerald, and Bryan Pardo. 2018. “An Overview of Lead and Accompaniment Separation in Music.” IEEE/ACM Transactions on Audio Speech and Language Processing 28, no. 8. 1307–35. https://doi.org/10.48550/arXiv.1804.08300. DOI: https://doi.org/10.1109/TASLP.2018.2825440
Rafii, Zafar, and Bryan Pardo. 2013. “REpeating Pattern Extraction Technique (REPET): A Simple Method for Music/Voice Separation.” IEEE Transactions on Audio, Speech and Language Processing 21, no. 1: 73–84. http://doi.org/10.1109/TASL.2012.2213249. DOI: https://doi.org/10.1109/TASL.2012.2213249
Rouard, Simon, Francisco Massa, and Alexandre Defossez. 2023. “Hybrid Transformers for Music Source Separation.” In ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1–5. https://doi.org/10.1109/ICASSP49357.2023.10096956. DOI: https://doi.org/10.1109/ICASSP49357.2023.10096956
Seetharaman, Prem, Fatemeh Pishdadian, and Bryan Pardo. 2017. “Music/Voice Separation Using the 2D Fourier Transform.” In 2017 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics. https://doi.org/10.1109/WASPAA.2017.8169990. DOI: https://doi.org/10.1109/WASPAA.2017.8169990
Simone, Nina. 1959. “My Baby Just Cares for Me.” Track 6 on Little Girl Blue. Recorded December 1957. Bethlehem BS-6028, 1959, Vinyl LP.
Spicer, Mark. 2017. “Fragile, Emergent, and Absent Tonics in Pop and Rock Songs.” Music Theory Online 23, no. 2. https://mtosmt.org/issues/mto.17.23.2/mto.17.23.2.spicer.html#nobile_2015. DOI: https://doi.org/10.30535/mto.23.2.2
Stöter, Fabian-Robert, Antoine Liutkus, and Nobutaka Ito. 2018. “The 2018 Signal Separation Evaluation Campaign.” In LVA/ICA: Latent Variable Analysis and Signal Separation, 293–305. https://doi.org/10.48550/arXiv.1804.06267. DOI: https://doi.org/10.1007/978-3-319-93764-9_28
Stöter, Fabian-Robert, Stefan Uhlich, Antoine Liutkus, and Yuki Mitsufuji. 2019. “Open-Unmix - A Reference Implementation for Music Source Separation.” Journal of Open Source Software 4, no. 41. https://joss.theoj.org/papers/10.21105/joss.01667. DOI: https://doi.org/10.21105/joss.01667
Temperley, David. 2007. “The Melodic-Harmonic ‘Divorce’ in Rock.” Popular Music 26, no. 2: 23–42. https://doi.org/10.1017/S0261143007001249. DOI: https://doi.org/10.1017/S0261143007001249
Vaneph, Alexandre, Ellie Mcneil, François Rigaud, and Rick Silva. 2016. “An Automated Source Separation Technology and Its Practical Applications.” In Proceedings of the AES 140th International Convention. http://www.aes.org/e-lib/browse.cfm?elib=18182.
Walzer, Daniel. 2020. “Blurred Lines: Practical and Theoretical Implications of a DAW-Based Pedagogy.” Journal of Music, Technology and Education 13, no. 1: 79–94. https://doi.org/10.1386/jmte_00017_1. DOI: https://doi.org/10.1386/jmte_00017_1
White, Christopher W. 2025. The AI Music Problem. Routledge. DOI: https://doi.org/10.4324/9781003587415
Yoo, Noah. 2020. “A Flashy New AI Tool Could Be a Producer’s Dream and a Copyright Nightmare.” Pitchfork, March 5. https://pitchfork.com/thepitch/a-flashy-new-ai-tool-could-be-aproducers-dream-and-a-copyright-nightmare/.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Michèle Duguay (Author)

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.