AI research scientist and PhD candidate at Indiana University Bloomington, studying mechanistic interpretability, multilingual reasoning, and knowledge-grounded AI systems.
I am an AI research scientist and dual-major PhD candidate at Indiana University Bloomington, with doctoral training in Computational Linguistics and Middle Eastern Languages & Cultures and a minor in Computer Science. My research connects linguistic theory with modern AI, from internal model representations to multilingual reasoning and structured knowledge systems.
I have built and evaluated datasets and systems for ellipsis detection, dialect classification, sexism detection, and named entity recognition in Arabic financial news. I have also studied public discourse on Arabic Twitter and analyzed language use in depression narratives. More recently, I have expanded my focus toward AI safety, explainability, and mechanistic interpretability — examining how transformer-based models internally represent linguistic features.
I developed Rasid, a 900M+ word Arabic Twitter corpus organized by year, month, and week, and AMWAL, an ontology-aligned named entity recognition system for Arabic financial news.
Completed a 30-hour professional program covering current AI safety techniques, threat modeling through kill-chain analysis, research gaps, and practical pathways for contributing to safer AI systems.
Probing transformer models to understand how speech and language systems encode linguistic features, nativeness, structure, and meaning.
Modeling entailment, contradiction, and pragmatic inference across languages. Current work examines evaluation reliability and translation-induced label drift.
Building ontology-based knowledge graphs and retrieval pipelines that connect model outputs to structured evidence in biomedical and financial domains.
Creating datasets, models, and evaluation methods for languages and domains that remain underrepresented in modern AI.