LLM Agent Safety
Locating and attributing failures in LLM agent reasoning during multi-step, safety-critical tasks, and compiling verified failures into runtime checks.
Computer Science Ph.D. student and Graduate Research Assistant at Wayne State University.
01
Computer Science Ph.D. student and Graduate Research Assistant at Wayne State University specializing in trustworthy AI, LLM agent safety, LLM robustness, and mechanistic interpretability. Experienced in designing reproducible GPU/HPC experiments, analyzing internal model behavior, and building evaluation pipelines for language and agent systems. Author of three arXiv preprints. Seeking a research internship focused on reliable and interpretable AI systems.
02
Locating and attributing failures in LLM agent reasoning during multi-step, safety-critical tasks, and compiling verified failures into runtime checks.
Evaluating reasoning stability under meaning-preserving perturbations and diagnosing performance-confidence mismatches.
Using activation analysis, patching, and ablation to identify internal components associated with model failures.
Building multi-step retrieval systems for literature comparison, contribution analysis, and research-gap identification.
03
Preprint · arXiv:2604.01639
Preprint · arXiv:2606.22151
Preprint · arXiv:2504.20342
04
Wayne State University, Trustworthy AI Lab
Advisors: Dongxiao Zhu and Sooin Kim
Research focus 01
05
Research project 01
Evaluated Mistral-7B, Llama-3-8B, and Qwen2.5-7B on 677 paired GSM8K problems, revealing answer-flip rates of 28.8%-45.1% under meaning-preserving perturbations.
Developed the Mechanistic Perturbation Diagnostics (MPD) framework, integrating logit-lens analysis, activation patching, component ablation, and the novel Cascading Amplification Index (CAI).
Identified localized, distributed, and entangled failure modes, demonstrating substantial architectural differences in failure localization and recoverability.
Evaluated steering vectors and layer fine-tuning as targeted repair methods using Python, PyTorch, Hugging Face Transformers, GPUs, and SLURM.
Research project 02
Built an agentic retrieval system over a 100-paper corpus using six components: query analysis, iterative retrieval, ranking, contribution extraction, comparison, and answer generation.
Designed structured contribution records and a three-pass comparison agent to identify paper-level overlaps, methodological differences, and research gaps.
Developed a problem-method gap matrix that generates citation-grounded evidence for unexplored combinations within a research corpus.
Achieved a mean Precision@5 of 0.980, nDCG@5 of 0.739, and 84.0% schema compliance across ten evaluation queries.
Implemented the system using Python, GPT-4o, FAISS/Chroma, SentenceTransformers, and structured prompting.
06
PyTorch, TensorFlow, Hugging Face Transformers, OpenAI API, prompt design, model evaluation, reasoning analysis
Logit lens, activation patching, component ablation, steering vectors, perturbation analysis, failure diagnosis
RAG, FAISS, Chroma, BM25, SentenceTransformers, ReAct-style agents, structured prompting
Python, Java, JavaScript, TypeScript, SQL, Bash, pandas, NumPy, Matplotlib
SLURM, HPC, CUDA/GPU computing, Linux, Git, Docker, FastAPI, Flask, React, Vercel
07
Detroit, Michigan
Ph.D. in Computer Science
Graduate Research Assistant, Trustworthy AI Lab
Boston, Massachusetts
Master of Science in Computer Science
Taipei, Taiwan
Bachelor of Arts in Political Science
Contact