|
Ashish Mittal
I'm a research scientist at IBM Research India, where I work on multimodal large language models and extending LLM capabilities for Indic languages. My recent work focuses on aligning speech and text modalities for Granite Speech models. At IBM, I have previously contributed to Watson Speech and developed natural language interfaces for Cognos.
I am nearing completion of my PhD at IIT Bombay, advised by Sunita Sarawagi, Preethi Jyothi, and George Saon. My doctoral research focuses on text-only contextualization of ASR models through retrieval, adaptation and LLM integration.
Email /
CV /
Scholar /
Twitter /
|
|
Research
I'm interested in multimodal language models, with a focus on fusing speech for building accessible interfaces to world knowledge. I also work on adapting LLMs to new and low-resource languages, and on inducing reasoning abilities in multilingual settings.
|
|
|
Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities
George Saon,
Avihu Dekel,
Alexander Brooks,
Tohru Nagano,
Abraham Daniels,
Aharon Satt,
Ashish Mittal,
Brian Kingsbury,
David Haws,
Edmilson Morais,
Gakuto Kurata,
Hagai Aronowitz,
Ibrahim Ibrahim,
Jeff Kuo,
Kate Soule,
Luis Lastras,
Masayuki Suzuki,
Ron Hoory,
Samuel Thomas,
Sashi Novitasari,
Takashi Fukuda,
Vishal Sunder,
Xiaodong Cui,
Zvi Kons
ASRU (Accepted), 2025
paper / models
This paper introduces Granite-speech LLMs, compact and efficient speech-language models (2B and 8B parameters) for English ASR and AST, which outperform competitors trained on more data and maintain pace on English-to-X AST for major languages, while preserving text LLM capabilities and being freely available on HuggingFace.
|
|
|
Mathematics Isn't Culture-Free: Probing Cultural Gaps via Entity and Scenario Perturbations
Aditya Tomar,
Nihar Ranjan Sahoo,
Ashish Mittal,
Rudra Murthy,
Pushpak Bhattacharya
Arxiv, 2025
paper
This paper introduces culturally adapted variants of the GSM8K math benchmark for five non-Western regions, evaluating six LLMs and revealing a consistent performance drop on culturally adapted versions, though models with stronger reasoning capabilities are more resilient.
|
|
|
RECAST: Retrieval-Augmented Contextual ASR via Decoder-State Keyword Spotting
Ashish Mittal,
Sunita Sarawagi,
Preethi Jyothi
EMNLP (Findings), 2025
RECAST is a lightweight, retrieval-augmented method for ASR contextual biasing that uses decoder states to retrieve keywords from large dictionaries, significantly improving rare term recognition and overall WER across diverse languages.
|
|
|
Power doesn't reside in size: A Low Parameter Hybrid Language Model (HLM) for Sentiment Analysis in Code-mixed data
Pavan Sai Balaga,
Nagasamudram Karthik,
Challa Vishwanath,
Raksha Sharma,
Rudra Murthy,
Ashish Mittal
EMNLP (Main), 2025
HLM is a hybrid model that couples a multilingual encoder with a lightweight decoder to deliver efficient, high-accuracy sentiment analysis on code-mixed text.
|
|
|
SKIP-SALSA: Skip Synchronous Fusion of ASR LLM Decoders
Ashish Mittal*,
Darshan Prabhu*,
Preethi Jyothi,
Sunita Sarawagi
Interspeech , 2025 (Oral)
This paper identifies a tokenization mismatch issue (token fertility gap) in SALSA, an ASR-LLM integration method, particularly affecting low-resource languages by starving the LLM of audio context; to address this, they propose SKIP-SALSA, which adaptively skips ASR decoder states to synchronize with the LLM, leading to significant ASR performance improvements in low-resource languages.
|
|
|
INDIC QA BENCHMARK: A Multilingual Benchmark to Evaluate Question Answering capability of LLMs for Indic Languages
Abhishek Singh,
Vishwajeet Kumar,
Rudra Murthy,
Jaydeep Sen,
Ashish Mittal,
Ganesh Ramakrishnan
NAACL (findings), 2025
paper / code
This paper introduces Indic-QA, a new large benchmark for context-grounded question answering in 11 major Indian languages, revealing that multilingual LLMs exhibit an English bias and perform weakly in low-resource settings, while a Translate-Test paradigm proves more effective.
|
|
|
Salsa: Speedy ASR-LLM Synchronous Aggregation
Ashish Mittal*,
Darshan Prabhu*,
Preethi Jyothi,
Sunita Sarawagi
Interspeech, 2024
(Nominated for Best Student Paper)
paper / code
This paper introduces SALSA, a novel method for improving ASR systems for low-resource languages by efficiently coupling the ASR and LLM decoders with a simple projection, addressing tokenizer mismatches through cascading tokenization, and achieving up to 38% WER reduction on the FLEURS benchmark.
|
|
|
Speech-enriched Memory for Inference-time Adaptation of ASR Models to Word Dictionaries
Ashish Mittal,
Preethi Jyothi,
Sunita Sarawagi,
George Saon,
Gakuto Kurata
EMNLP, 2023
paper / data
This paper proposes a novel inference algorithm that enhances rare word prediction in state-of-the-art ASR models by using a nearest-neighbor-based matching on an inference-time word list, indexed in a memory as a trie to prevent spurious matches.
|
|
|
Improving RNN-Transducers with Acoustic LOOKAHEAD
Vinit Unni,
Ashish Mittal,
Preethi Jyothi,
Sunita Sarawagi
Interspeech, 2023
paper
This paper introduces a technique for RNN-Transducer ASR models that improves text representations by looking ahead in the audio input, thereby reducing language model biasing and multi-step hallucination, resulting in a 5%-20% relative reduction in word error rate.
|
|
|
In-Situ Text-Only Adaptation of Speech Models with Low-Overhead Speech Imputations
Ashish Mittal,
Sunita Sarawagi,
Preethi Jyothi
ICLR, 2023
paper
This paper introduces a novel text-only adaptation technique for RNN-T ASR models that imputes speech representations to enable in-situ adaptation, significantly reducing word error rate (up to 35%) and mitigating catastrophic forgetting without runtime overhead or retraining the ASR model from scratch.
|
|
|
Soft Random Sampling: A Theoretical and Empirical Analysis
Xiaodong Cui,
Ashish Mittal,
Songtao Lu,
Wei Zhang,
George Saon,
Brian Kingsbury,
Arxiv, 2022
paper
This paper provides a theoretical and empirical analysis of Soft Random Sampling (SRS), a simple and effective data sampling method for large-scale deep neural network training, demonstrating its strong convergence and generalization, as well as its superior accuracy-efficiency trade-off compared to coreset-based methods in image and speech recognition tasks.
|
|
|
Distributed Gradient Matching-based Data Subset Selection for Compute-Efficient and Robust ASR Training
Ashish Mittal*,
Durga Sivasubramaniam*,
Rishabh Iyer,
Preethi Jyothi,
Ganesh Ramakrishnan,
EMNLP (Findings), 2022
paper
This paper proposes Partitioned Gradient Matching (PGM), a novel distributable data subset selection algorithm for RNN-T ASR systems, which significantly reduces training time (3x-6x speedup) with minimal accuracy degradation (< 1% WER difference) even with noisy data, addressing the high computational cost of training state-of-the-art ASR models.
|
|
|
Adaptive Discounting of Implicit Language Models in RNN-Transducers
Vinit Unni,
Shreya Khare,
Ashish Mittal,
Preethi Jyothi,
Sunita Sarawagi,
ICASSP, 2022
paper
This paper introduces a lightweight adaptive language model discounting technique for RNN-Transducer ASR models, which improves rare word performance by masking prediction network output and dynamically discounting the implicit language model, achieving significant WER and PER reductions on a Hindi-English code-mixed task.
|
|
|
Low Resource ASR: The Surprising Effectiveness of High Resource Transliteration
Shreya Khare*,
Ashish Mittal*,
Anuj Diwan*,
Samarth Bharadwaj,
Preethi Jyothi,
Sunita Sarawagi,
Interspeech, 2021
paper
This paper introduces a novel cross-lingual transfer learning strategy for ASR, pretraining high-resource language speech with transliterated text in the low-resource target language, proving effective even for unrelated language families and significantly reducing Word Error Rate (WER) in low-resource scenarios.
|
|
|
Multilingual and code-switching ASR challenges for low resource Indian languages
Anuj Diwan,
Rakeshk Vaideeswaran,
Sanket Shah,
Ankita Singh,
Srinivasa Raghavan,
Shreya Khare,
Vinit Unni,
Saurabh Vyas,
Akash Rajpuria,
Chiranjeevi Yarra,
Ashish Mittal,
Prasanta Kumar Ghosh,
Preethi Jyothi,
Kalika Bali,
Vivek Seshadri,
Sunayana Sitrama,
Samarth Bharadwaj,
Jai Nanavati,
Raoul Nanavati,
Karthik Sankaranarayanan
Interspeech, 2021
paper / data
This paper details a challenge to build multilingual and code-switching ASR systems for seven Indian languages, providing 600 hours of transcribed data and baseline recipes.
|
|
|
Bootstrapping Chatbot Interfaces to Databases
Ashish Mittal,
Diptikalyan Saha,
Parag Jain,
Jaydeep Sen,
Manasa Jammi
Karthik Sankaranarayanan
Proceedings of the CODS-COMAD, 2021
paper
The paper introduces an automated technique for bootstrapping chatbot interfaces for question answering on relational databases, utilizing natural language classifiers from industrial chatbot platforms to translate natural language into structured queries.
|
|
|
Optimizing Interpretation Generation in Natural Language Query Answering for Real Time End Users
Jaydeep Sen,
Ashish Mittal,
Diptikalyan Saha,
Karthik Sankaranarayanan
Proceedings of the CODS-COMAD, 2021
paper
This paper presents novel algorithms that extend the ATHENA system, using Functional Partitioning of Ontology and Lazy Inclusion to significantly improve the automated generation of precise interpretations for natural language database queries by a factor of at least 400%.
|
|
|
Representation Based Meta-Learning For Few-Shot Spoken Intent Recognition
Ashish Mittal,
Samarth Bharadwaj,
Shreya Khare,
Saneem Chemmengath,
Karthik Sankaranarayanan,
Brian Kingsbury
Interspeech, 2020
paper
A few-shot spoken intent classification method that uses a meta-learning approach to create task-agnostic representations of utterances, enabling accurate classification of new intents with limited examples.
|
|
|
ATHENA++: Natural Language Querying for Complex Nested SQL Queries
Jaydeep Sen,
Chuan Lei,
Abdul Quamar,
Fatma Özcan,
Vasilis Efthymiou,
Ayushi Dalmia,
Greg Stager,
Ashish Mittal,
Diptikalyan Saha,
Karthik Sankaranarayanan
Proceedings of the VLDB Endowment, 2020
paper
A system that translates complex natural language queries into nested SQL queries by combining linguistic patterns with deep domain reasoning using ontologies.
|
|
|
Unified Semantic Parsing with Weak Supervision
Priyanka Agrawal*,
Parag Jain*,
Ayushi Dalmia,
Abhishek Bansal,
Ashish Mittal,
Karthik Sankaranarayanan
Proceedings of ACL, 2019
paper / code
Unified semantic parser trained via weak supervision from multiple domain-specific teachers without requiring domain labels.
|
|
|
Ontology-Based Natural Language Query Interfaces for Data Exploration
Chuan Lei,
Fatma Özcan,
Abdul Quamar,
Ashish Mittal,
Jaydeep Sen,
Diptikalyan Saha,
Karthik Sankaranarayanan
IEEE Data Engineering Bulletin, 2018
paper
An ontology-based two-stage architecture enables NLQ systems to translate user questions first into ontology-level queries and then into executable database queries to support intuitive querying over structured data.
|
|
|
Tooling Framework for Instantiating Natural Language Querying System
Manasa Jammi
Jaydeep Sen,
Ashish Mittal,
Sagar Verma,
Vardaan Pahuja,
Rema Ananthanarayanan,
Pranay Lohia
Hima Karanam,
Diptikalyan Saha,
Karthik Sankaranarayanan
Proceedings of VLDB, 2018
paper
A tooling framework to easily instantiate NLQ systems across diverse structured data sources and formats, enabling natural language querying with minimal human configuration.
|
|
|
Functional Partitioning of Ontologies for Natural Language Query Completion in Question Answering Systems
Jaydeep Sen,
Ashish Mittal,
Diptikalyan Saha,
Karthik Sankaranarayanan
Proceedings of IJCAI, 2018
paper
Proposed the first query completion framework for NLIDB systems using functional partitioning of ontologies to generate accurate and semantically meaningful completions without relying on query logs.
|
|
|
An Ontology based Dialog Interface to Database
Ashish Mittal,
Jaydeep Sen,
Diptikalyan Saha,
Karthik Sankaranarayanan
Proceedings of SIGMOD, 2018
paper
Built a system to answer contextual natural language queries on tabular data by combining semantic parsing with program execution.
|
|
|
Creation and interaction with large-scale domain-specific knowledge bases
Shreyas Bharadwaj,
Laura Chiticariu,
Marina Danilevsky,
Samarth Dhingra,
Samved Divekar,
Arnaldo Carreno-Fuentes,
Himanshu Gupta,
Nitin Gupta,
S D Han,
Mauricio Hernández,
Howard Ho,
Parag Jain,
Salil Joshi,
Hima Karanam,
Saravanan Krishnan,
Rajasekar Krishnamurthy,
Yunyao Li,
Satishkumar Manivannan,
Ashish Mittal,
Fatma Özcan,
Abdul Quamar,
Poornima Raman,
Diptikalyan Saha,
Karthik Sankaranarayanan,
Jaydeep Sen,
Prithiviraj Sen,
Shivakumar Vaithyanathan,
Mitesh Vasa,
Hao Wang,
Huaiyu Zhu
Proceedings of VLDB, 2017
paper
Demonstrated a system for answering natural language questions over enterprise knowledge graphs using semantic parsing and graph querying.
|
|
|
ATHENA: An Ontology-Driven System for Natural Language Querying over Relational Data Stores
Diptikalyan Saha,
Avrilia Floratou,
Karthik Sankaranarayanan,
Umar Farooq Minhas,
Ashish Mittal,
Fatma Özcan
Proceedings of VLDB, 2016
paper
Co‑developed ATHENA, an ontology-driven natural language interface that enables users to query complex relational databases using natural language.
|
|
|
Numerical Relation Extraction with Minimal Supervision
Aman Madaan*,
Ashish Mittal*,
Mausam,
Ganesh Ramakrishnan,
Sunita Sarawagi
AAAI, 2016
paper / code
Built a rule-based system and a graphical model to extract factual triples (e.g., (India, Population, 1.3 billion)) from sentences expressing numeric facts.
|
Ontology driven contextual automated speech recognition
Ashish Mittal,
Samarth Bharadwaj,
Shreya Khare
US12315496B2, 2022
|
Artificial intelligence factsheet generation for speech recognition
Shreya Khare,
Ashish Mittal,
Saneem Chemmengath,
Samarth Bharadwaj,
Karthik Sankaranarayanan
US20230419950A1, 2022
|
Automated domain-specific constrained decoding from speech inputs to structured resources
Ashish Mittal,
Samarth Bharadwaj,
Shreya Khare,
Karthik Sankaranarayanan
US20230215427A1, 2022
|
Active learning for natural language question answering
Jaydeep Sen,
Karthik Sankaranarayanan,
Ashish Mittal
US11971886B2, 2021
|
Domain query execution using user-provided definition
Jaydeep Sen,
Ashish Mittal,
Diptikalyan Saha,
Karthik Sankaranarayanan
US11294907B2, 2020
|
Multi-agent conversational agent framework
Ashish Mittal,
Diptikalyan Saha,
Priyanka Agrawal,
Manasa Jammi
US12086561B2, 2019
|
Natural language interface databases
Jaydeep Sen,
Diptikalyan Saha,
Karthik Sankaranarayanan,
Ashish Mittal,
Manasa Jammi
US11200222B2, 2019
|
Automated generation of test cases for analyzing natural-language-interface-to-database systems
Diptikalyan Saha,
Jaydeep Sen,
Manasa Jammi,
Ashish Mittal
US10977164B2, 2018
|
Ontology-based automatic bootstrapping of state-based dialog systems
Jaydeep Sen,
Parag Jain,
Diptikalyan Saha,
Ashish Mittal
US10977164B2, 2018
|
Ontology refinement based on query inputs
Ashish Mittal,
Diptikalyan Saha,
Karthik Sankaranarayanan,
Jaydeep Sen
US10810246B2, 2017
|
Ontology based query suggestion using eye tracking
Ashish Mittal,
Diptikalyan Saha,
Karthik Sankaranarayanan,
Jaydeep Sen
US20190087486A1, 2017
|
Interactive dialog in natural language using an ontology
Ashish Mittal,
Diptikalyan Saha,
Karthik Sankaranarayanan,
Jaydeep Sen
US10621166B2, 2017
|
Translating Structured Languages to Natural Language Using Domain-Specific Ontology
Ashish Mittal,
Diptikalyan Saha,
Karthik Sankaranarayanan
US20180210879A1, 2017
|
Providing notifications in video streams to control video playback
Vijay Ekambaram,
Ashish Mittal,
Ruhi Sharma Mittal,
Yedendra Shrinivasan
US20180210879A1, 2016
|
|