{"slug":"vanderbilt-university-applies-nlp-on-databricks-to-analyze-decades-of-tv-news-archive-for-research","url":"https://findausecase.com/use-cases/vanderbilt-university-applies-nlp-on-databricks-to-analyze-decades-of-tv-news-archive-for-research","title":"Vanderbilt University applies NLP on Databricks to analyze decades of TV news archive for research","description":"The Jean and Alexander Heard Libraries at Vanderbilt University transcribed their Vanderbilt Television News Archive and deployed the Databricks Data + AI Platform, using Unity Catalog and Spark Declarative Pipelines, to let researchers run natural language processing and machine learning experiments pairing the archive with datasets like polling data and public policy documents, including studying the impact of presidential executive orders on public discourse. Since implementation, citations of Vanderbilt scholarship using the archive have increased 2.7x.","company":"Vanderbilt University","industry":"Education","aiCapabilities":["Natural Language Processing","Machine Learning"],"technology":["Unity Catalog","Spark Declarative Pipelines","Delta Lake"],"deployment":"Public Cloud","problemStatement":"The Heard Libraries lacked the capacity and full-text transcripts to let scholars run natural language processing tools on the Vanderbilt Television News Archive, and accessing and manipulating this large-scale data was cumbersome with limited capability to handle complex queries spanning multiple data types and time periods.","solutionApproach":"With a gift funding transcription of all video content, Vanderbilt deployed the Databricks Data + AI Platform to provision secure researcher access, using Unity Catalog for governance and Spark Declarative Pipelines for automatic ingestion of new data, enabling researchers to run NLP techniques and machine learning experiments pairing the TV News Archive with other datasets.","businessValue":"Vanderbilt is now cited approximately 2.7 times more often than before, usage of the Vanderbilt Television News Archive grew from thousands of users toward a projected hundreds of thousands, and researchers without deep data science expertise can now run advanced NLP analysis that was previously not possible.","evidence":{"band":"high"},"sourceUrl":"https://www.databricks.com/customers/vanderbilt-university","dates":{"publishedAt":"2026-09-23T05:48:59.271Z","publishedAtSource":"pipeline","updatedAt":"2026-09-23T05:48:59.271Z"},"license":"Open for reading and citing with a link to https://findausecase.com/use-cases/vanderbilt-university-applies-nlp-on-databricks-to-analyze-decades-of-tv-news-archive-for-research. Bulk republication requires permission."}