Hello, I'm
Rajvee Sheth
NLP Researcher
I am a Researcher specializing in Natural Language Processing, code-mixed language processing, and multimodal AI. My work focuses on developing datasets, benchmarks, and evaluation frameworks spanning Hindi-English code-mixing, multilingual LLM evaluation, and culturally grounded vision-language understanding.
My research contributions include the COMI-LINGUA dataset for multitask Hindi-English NLP, the COMMENTATOR annotation framework for code-mixed text, and a comprehensive survey of code-switched NLP in the LLM era.
Research Interests
Natural Language Processing
Developing robust computational frameworks for multilingual text understanding, with emphasis on advanced annotation methodologies and scalable evaluation systems that enable interpretable and secure NLP solutions across diverse linguistic contexts.
Code-Mixing
Redefining code-mixed NLP with curated Hindi-English datasets, robust annotation frameworks and comprehensive evaluation on existing state-of-the-art LLMs to improve language understanding.
Technical Skills
Programming Languages
Web Technologies
Databases & Tools
Libraries
Projects
Curating and constructing benchmarks and development of ML models for low level NLP tasks in Hindi-English code-mixing
Funded by ANRF - This comprehensive project encompasses the development of benchmarks and Annotation Framework for Hindi-English code-mixing research. Key deliverables include:
COMI-LINGUA Dataset: A large-scale, expert-annotated dataset for multitask NLP in Hindi-English code-mixing, covering tasks like Matrix Language Identification, POS Tagging, and Named Entity Recognition. Available at: Hugging Face.
COMMENTATOR Portal: A code-mixed multilingual text annotation framework designed to facilitate the annotation and analysis of Hindi-English texts. This tool supports the creation and management of annotated datasets for research purposes. Available at: GitHub Repo.
Survey on Code-Switching: A comprehensive survey covering recent advances by architecture, training strategy, and evaluation methodology, outlining how LLMs have reshaped code-switching, modeling and identifying the challenges that persist. Available at: ACL Anthology and GitHub Repo.
More about the work: Project Website
Publications
Beyond Monolingual Assumptions: A Survey on Code-Switched NLP in the Era of Large Language Models across Modalities
Sheth, R., Sinha, S. R., Patil, M., Beniwal, H., & Singh, M., "Beyond Monolingual Assumptions: A Survey on Code-Switched NLP in the Era of Large Language Models across Modalities," Proceedings of the 2026 Conference on Association for Computational Linguistics, 2026. [PDF]
COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing
R. Sheth, H. Beniwal, M. Singh, "COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing," Findings of the Association for Computational Linguistics: EMNLP 2025. [PDF]
COMMENTATOR: A Code-mixed Multilingual Text Annotation Framework
R. Sheth, S. Nisar, H. Prajapati, H. Beniwal, M. Singh, "COMMENTATOR: A Code-mixed Multilingual Text Annotation Framework," Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 2024. [PDF]
Not Truly Multilingual: Script Consistency as a Missing Dimension in VLM Evaluation
Singh, P., Pawar, B., Reddiboina, M., & Sheth, R., "Not Truly Multilingual: Script Consistency as a Missing Dimension in VLM Evaluation," arXiv preprint arXiv:2606.17188, 2026. [PDF]
EKA-EVAL : An Evaluation Framework for Low-Resource Multilingual Large Language Models
Sinha, S. R., Sheth, R., Upperwal, A., & Singh, M., "EKA-EVAL : An Evaluation Framework for Low-Resource Multilingual Large Language Models," arXiv preprint arXiv:2507.01853, 2025. [PDF]
News and Updates
April 2026: Survey Paper on "Code-Switching in the LLM Era" accepted at ACL Main 2026.
February 2026: Organizing team member for AI Day @ IITGN, an India AI Impact Pre-Summit 2026 event supported by ANRF Pair, the CSE Department, and the Center for AI-driven Innovations, IIT Gandhinagar.
December 2025: Attended the Sixth Indian Symposium on Machine Learning (IndoML) at BITS Pilani, Hyderabad Campus. Website
December 2025: Participated in a research showcase at the Regional AI Impact Conference, a platform focused on advancing AI-driven solutions and fostering collaboration between academia, industry, and government. Website
November 2025: Showcased research at Amalthea Tech Expo 2025. IITGN’s annual technology summit bringing together students, industry leaders, researchers, and academics to connect, collaborate, and share innovative ideas. Website
November 2025: Promoted as Senior Research Fellow at Lingo Research Labs, IIT Gandhinagar.
September 2025: Contributed to the organization and planning for the workshop "AI for Libraries: Building Applications for Search, Chatbots, and Archiving" at CLSTL 2025. Website
August 2025: Paper on "COMI-LINGUA" dataset accepted at EMNLP Findings 2025.
February 2025: Showcased research at CoLab 2025 | An IITGN Industry Open House. IITGN's flagship event connecting industry and academia to explore research partnerships and foster innovation collaborations.
January 2025: Participated in the Curiosity Carnival for School Children at IIT Gandhinagar, engaging with young minds and promoting STEM education.
October 2024: Paper on "COMMENTATOR" framework accepted at EMNLP 2024.
June - July 2024: Actively volunteered at the ACM INDIA Summer School 2024 on GenAI for Text (June 24 - July 5), contributing to educational initiatives in artificial intelligence.
April 2024: Served as a Programming Technical Assistant for a one-day workshop on Python Programming and AI Applications, supporting hands-on learning experiences.
February 2024: Participated in Science Day and CoLab 2024, engaging in collaborative research and academic discussions.
November 2023: Joined Lingo Labs at IIT Gandhinagar as a Junior Research Fellow.