AI Stem Separation Tools in the Aural Skills Classroom

Authors

  • Fred Hosken Author

DOI:

https://doi.org/10.15763/issn.2994-7073.2026.39.39-58

Keywords:

Artificial Intelligence, music theory pedagogy

Abstract

Recent advancements in artificial intelligence and machine learning, particularly AI-powered stem separation tools, offer transformative potential for aural skills education. These tools can isolate individual audio components, such as vocals, bass, drums, and harmonic instruments, from mixed tracks, providing educators and students with novel ways to engage with complex musical textures. One of the key challenges in aural skills training is developing the ability to parse polyphonic textures, a skill known as auditory streaming. AI stem separation tools enable students to focus on individual musical elements, enhancing their ability to isolate and analyze components of multi-voice performances. This paper explores how separators can scaffold auditory streaming development and improve other aspects of aural skills instruction. It also addresses the ethical implications and responsibilities of using such technologies in the classroom. Through case studies, this work demonstrates how AI-driven tools can enrich both music instruction and student learning, while fostering awareness of technology’s broader impact on pedagogy.

Downloads

Download data is not yet available.

References

Airaj, Mohammed. 2024. Ethical Artificial Intelligence for Teaching-Learning in Higher Education. Education and Information Technologies 29, no. 13: 17145–67. https://doi.org/10.1007/s10639-024-12545-x. DOI: https://doi.org/10.1007/s10639-024-12545-x

Anonymous. 2024. “Nine Inch Nails – Multitracks.” nin.wiki. https://www.nin.wiki/Multitracks.

Atilgan, Huriye, and Jennifer K. Bizley. 2021. “Training Enhances the Ability of Listeners to Exploit Visual Information for Auditory Scene Analysis.” Cognition 208: 104529. https://doi.org/10.1016/j.cognition.2020.104529. DOI: https://doi.org/10.1016/j.cognition.2020.104529

Atilgan, Huriye, Stephen M. Town, Katherine C. Wood, Gareth P. Jones, Ross K. Maddox, Adrian K.C. Lee, and Jennifer K. Bizley. 2018. “Integration of Visual Information in Auditory Cortex Promotes Auditory Scene Analysis through Multisensory Binding.” Neuron 97, no. 3: 640–55. https://doi.org/10.1016/j.neuron.2017.12.034. DOI: https://doi.org/10.1016/j.neuron.2017.12.034

Bregman, Albert. 1990. Auditory Scene Analysis. Cambridge, MA: MIT Press. DOI: https://doi.org/10.7551/mitpress/1486.001.0001

Bregman, Albert. 2015. “Progress in Understanding Auditory Scene Analysis.” Music Perception 33, no. 1: 12–19. https://doi.org/10.1525/mp.2015.33.1.12. DOI: https://doi.org/10.1525/mp.2015.33.1.12

Calcus, Axelle. 2024. “Development of Auditory Scene Analysis: A Mini-Review.” Frontiers in Human Neuroscience 18. https://doi.org/10.3389/fnhum.2024.1352247. DOI: https://doi.org/10.3389/fnhum.2024.1352247

Chenette, Timothy. 2022. Foundations of Aural Skills. https://uen.pressbooks.pub/auralskills/.

Coote, Jonathan. 2024. “‘Stem-separating’ AI is Revolutionising the Music Industry.” Bray & Krais. https://www.brayandkrais.com/stem-separating-ai-is-revolutionising-the-music-industry/.

Dewitt, Lucinda, and Crowder, Robert. 1987. “Tonal Fusion of Consonant Musical Intervals: The Oomph in Stumpf.” Perception & Psychophysics 41, no. 1: 73–84. https://doi.org/10.3758/BF03208216. DOI: https://doi.org/10.3758/BF03208216

Duerksen, Marva. 2009. “Review of Gary S. Karpinski, Manual for Ear Training and Sight Singing, Anthology for Sight Singing, Student Recordings CD-ROM, Instructor’s Dictation Manual, and Instructor’s CD-ROM.” Gamut 2, no. 1: 391–405. DOI: https://doi.org/10.7290/gamutnab6

Hosken, Fred. Forthcoming. “AI Stem Separation & MIR: Assessing the Validity of Onset Analyses That Use AI-Isolated Audio.” Empirical Musicology Review.

Huron, David. 2016. Voice Leading: The Science Behind a Musical Art. The MIT Press. https://doi.org/10.7551/mitpress/9780262034852.003.0001. DOI: https://doi.org/10.7551/mitpress/9780262034852.001.0001

Ishida, Kai, and Hiroshi Nittono. 2024. “Different Voice Part Perceptions in Polyphonic and Homophonic Musical Textures.” Psychology of Music. https://doi.org/10.1177/03057356241271027. DOI: https://doi.org/10.1177/03057356241271027

Karpinski, Gary S. 2000. Aural Skills Acquisition: The Development of Listening, Reading, and Performing Skills in College-Level Musicians. New York: Oxford University Press. DOI: https://doi.org/10.1093/oso/9780195117851.001.0001

Marie, Céline, and Laurel J. Trainor. 2012. “Development of Simultaneous Pitch Encoding: Infants Show a High Voice Superiority Effect.” Cerebral Cortex 23, no. 3: 660–69. https://doi.org/10.1093/cercor/bhs050. DOI: https://doi.org/10.1093/cercor/bhs050

Marie, Céline, and Laurel J. Trainor. 2014. “Early Development of Polyphonic Sound Encoding and the High Voice Superiority Effect.” Neuropsychologia 57: 50–8. https://doi.org/10.1016/j.neuropsychologia.2014.02.023. DOI: https://doi.org/10.1016/j.neuropsychologia.2014.02.023

McDonald, Scott. October 9, 2013. “Moby X BitTorrent: Unlock and Remix Innocents.” BitTorrent. https://www.bittorrent.com/blog/2013/10/09/moby-x-bittorrent-unlock-and-remixinnocents/

McGuire, Jack, David De Cremer, and Tim Van de Cruys. 2024. “Establishing the Importance of Co-Creation and Self-Efficacy in Creative Collaboration with Artificial Intelligence.” Scientific Reports 14 (1): 18525. https://doi.org/10.1038/s41598-024-69423-2. DOI: https://doi.org/10.1038/s41598-024-69423-2

Miller, Arthur I. 2020. The Artist in the Machine. London, England: MIT Press.

Munive Benites, David A., and Philippe Lalitte. 2022. “Employment of Cognitive Science Theories to Improve Aural Training through Digital Tools.” In 15th International Conference of Students of Systematic Musicology (SysMus22), 50–51. Ghent.

Rafii, Zafar, Antoine Liutkus, Fabian-Robert Stöter, Stylianos Ioannis Mimilakis, and Rachel Bittner. 2019. “MUSDB18-HQ – An Uncompressed Version of MUSDB18.” Zenodo. https://doi.org/10.5281/zenodo.3338373.

Rasch, Ruth. 1979. “Synchronization in Performed Ensemble Music.” Acustica 43: 121–31.

Rouard, Simon, Francisco Massa, and Alexandre Défossez. 2022. “Hybrid Transformers for Music Source Separation.” arXiv. https://doi.org/10.48550/arXiv.2211.08553.

Sussman, Elyse S. 2017. “Auditory Scene Analysis: An Attention Perspective.” Journal of Speech, Language, and Hearing Research 60, no. 10: 2989–3000. https://doi.org/10.1044/2017_JSLHR-H-17-0041. DOI: https://doi.org/10.1044/2017_JSLHR-H-17-0041

Sussman, Elyse, and Mitchell Steinschneider. 2009. “Attention Effects on Auditory Scene Analysis in Children.” Neuropsychologia 47, no. 3: 771–785. https://doi.org/10.1016/j.neuropsychologia.2008.12.007. DOI: https://doi.org/10.1016/j.neuropsychologia.2008.12.007

Trainor, Laurel J., Céline Marie, Ian C. Bruce, and Gavin M. Bidelman. 2014. “Explaining the High Voice Superiority Effect in Polyphonic Music: Evidence from Cortical Evoked Potentials and Peripheral Auditory Models.” Hearing Research 308: 60–70. https://doi.org/10.1016/j.heares.2013.07.014. DOI: https://doi.org/10.1016/j.heares.2013.07.014

Downloads

Published

2026-07-08

Issue

Section

Special Symposium: AI and Music Theory Pedagogy