Alzheimer's Research Knowledge Graph
Current work on an interactive knowledge graph that connects Alzheimer's disease pathology, genetics, brain regions, cellular processes, symptoms, diagnostics, treatments, and protective factors.
Public tools, datasets, and interactive systems built to support research, exploration, and reproducible workflows.
Current work on an interactive knowledge graph that connects Alzheimer's disease pathology, genetics, brain regions, cellular processes, symptoms, diagnostics, treatments, and protective factors.
A 900M+ word longitudinal corpus and exploration environment supporting frequency analysis, collocation search, geographic visualization, and social-language research.
A text-to-speech reader powered by Cloudflare Workers AI using the free MeloTTS model. Fast, lightweight, and directly integrated into your domain with playback speed control.
A text-to-speech reader powered by OmniVoice on Hugging Face Spaces. Offers 10+ voice presets, multiple language support, and advanced voice personalization options.
A browser-based tool for exporting structured metadata from public YouTube videos and playlists. Users enter their own YouTube Data API key, select the fields they want, and download the results as CSV.
Public materials accompanying research in language, discourse, model evaluation, and social data analysis.
A hybrid drug-mention extraction system combining a 562-term pharmacological lexicon, ModernBERT span detection, and a DeBERTa context classifier for ambiguous slang in Reddit-style text.
An ontology-aligned named entity recognition system for Arabic financial news, covering 21 financial entity types with released code and a ready-to-use spaCy model.
Two corpora of patient narratives describing diagnoses of Major Depressive Disorder and Bipolar Disorder, released with the associated appraisal-analysis research.
A 14-week English-language Twitter dataset for studying pandemic themes, psycholinguistic signals, public perception, and conceptual metaphors.
A 12-week Arabic Twitter dataset supporting thematic, psycholinguistic, stylistic, and geographic analysis of pandemic discourse.
The Arabic sub-corpus of the Hoosier Ellipsis Corpus, created for linguistic analysis and machine-learning experiments on omitted material.
Multilingual pretrained language-model experiments for sexism detection on Twitter in the EXIST 2023 shared task.
Code and materials for country-level dialect classification in tweets, developed for the NADI 2023 shared task.
Reusable preprocessing utilities for cleaning and normalizing Arabic text before corpus analysis and machine-learning workflows.
Supplementary materials and anonymized data for a corpus-based analysis of spousal duties in online fatwa inquiries.
An LSTM-based Arabic text-prediction model demonstrating sequence modeling and next-token generation.