Doctoral Dissertation
Computational Audition with Imprecise Labels
Dissertation Overview & Abstract
Computational audition aims to enable artificial systems to process, analyze, and understand the acoustic world in a manner analogous to human listening. While deep learning has driven remarkable breakthroughs in acoustic analysis, standard supervised learning paradigm depends heavily on massive collections of pristine, precisely time-stamped, and verified ground-truth labels — annotations that are prohibitively expensive, subjective, and scarce in real-world acoustic settings.
This dissertation investigates and develops foundational machine learning frameworks for learning from imprecise acoustic labels, including weakly supervised labels (audio-level tags without onset/offset timestamps), webly scraped annotations, noisy/corrupted labels, partial labels, and complementary supervision. We establish theoretical bounds, design robust loss-weighting and sampling strategies, and validate our methods across large-scale benchmarks (Audioset, DCASE challenges, public safety gunshot diagnostics, and real-time smart surveillance).
Citation & BibTeX
@phdthesis{shah2024computational,
title = {Computational Audition with Imprecise Labels},
author = {Shah, Ankit Parag},
year = {2024},
school = {Carnegie Mellon University},
address = {Pittsburgh, PA, USA},
doi = {10.1184/R1/28422542.v1},
url = {https://kilthub.cmu.edu/articles/thesis/Computational_Audition_with_Imprecise_Labels/28422542}
}