Research
Detecting synthetic language, and knowing when we can't.
My work asks a deceptively simple question: given a piece of text, can we tell whether a person or a language model wrote it — and how far does that answer hold up when someone is trying to fool us?
Research interests
- Machine-generated text detection — provenance, watermarking, and classifier design
- Machine translation for NLP — multilingual modeling and translation quality
- Adversarial robustness of detectors under paraphrase and persona conditioning
- Deception detection and trustworthiness in language
- Human-centered evaluation of NLP systems
Master's thesis
SAFE · 2026
Suspicious vs. Authentic Feedback Evaluation for Detecting LLM-Generated Reviews on E-Commerce Platforms
Advised by Dr. Veronica Perez-Rosas, Texas State University.
SAFE studies whether transformer classifiers can reliably separate genuine Amazon reviews from ones generated by large language models. Using a persona-conditioned generation pipeline, I built evaluation data across five product categories and four prompting strategies, spanning three generators — GPT-4.1, LLaMA-3.1-8B, and Mistral — and benchmarked nine detection methods, including fine-tuned DeBERTa-v3 and RoBERTa.
A central finding: persona-conditioned prompting largely collapses back to a zero-shot detection problem, because models default to formal, uniform vocabulary regardless of the persona they are asked to adopt. The result is both encouraging for detection and a caution about how synthetic text is evaluated.
For collaboration or questions about this work, email har98@txstate.edu.