CV
This is my CV. You can also download it as a PDF.
CV Preview
Basics
| Name | Keyu He |
| Label | Master in NLP@CMU SCS || NLP Researcher |
| keyuhe@cmu.edu | |
| Phone | (213) 713-2973 |
| Url | https://keyu-he.github.io/ |
| Summary | Master of Intelligent Information Systems (MIIS) student at Carnegie Mellon University. Previously earned B.S. from USC with double major in Computer Science and Applied & Computational Mathematics. Passionate about socially intelligent AI systems and human-centered AI evaluation. |
Education
-
2025.08 - 2027.05 Pittsburgh, PA
Master's
Carnegie Mellon University
Master of Intelligent Information Systems (MIIS)
GPA: 4.14/4.33
- Advanced Natural Language Processing (A+)
- Multimodal Machine Learning (A+)
- Prompt Engineering (A+)
- Introduction to Question Answering (A)
- Independent Study in Language Technologies (A)
- Directed Study in Language Technologies (A)
-
2021.08 - 2025.05 Los Angeles, CA
Bachelor's
University of Southern California
Computer Science and Applied & Computational Mathematics
GPA: 3.98/4.00
Minor in Artificial Intelligence Applications. Specializations in Applied Analytics and Video Game Programming (USC Viterbi ITP). Dean's List, USC Dornsife and USC Viterbi (2021-2025).
- Language Models in Natural Language Processing (A)
- Applied Machine Learning for Natural Language Processing (A)
- Applied Neural Networks (A)
- Capstone: Design and Construction of Large Software Systems (A)
- Mathematical Statistics (A)
- Probability Theory (A)
- Numerical Methods (A)
Publications
-
2026.08 Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments
Preprint
Social Gym is a benchmark of 21 multi-agent social games with verifiable, objective outcomes; SPaRTan is a training-free self-play and reflect-transfer loop that produces transferable playbooks improving LLM social reasoning across games and models.
-
2026.07 Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations
ACL 2026
We introduce Visual Fidelity and Contrastiveness -- two explanation quality scores that help users more appropriately rely on vision-language model predictions without seeing the image.
-
2026.07 RECAP: An End-to-End Platform for Capturing, Replaying, and Analyzing AI-Assisted Programming Interactions
ACL 2026 Demo
RECAP captures AI chat and code edits in VS Code, merges them into replayable timelines, and provides analysis modules for studying developer-AI interaction patterns.
-
2025.07 ELI-Why: Evaluating the Pedagogical Utility of LLM Explanations
Findings of ACL 2025
Evaluate the pedagogical utility of LLMs in tailoring explanations to users with different educational backgrounds.
-
2025.04 Attributing Culture-Conditioned Generations to Pretraining Corpora
ICLR 2025
This paper introduces MEMOed, a framework to analyze whether AI generations are driven by memorization or generalization, with a focus on cultural symbols.
-
2023.12 Enhancing Debugging Skills of LLMs with Prompt Engineering
Technical Report
Study on improving LLM debugging through prompt engineering techniques. Evaluated few-shot learning and chain-of-thought approaches on GPT-3.5, revealing limitations in current debugging capabilities.
Work
-
2026.05 - 2026.08 San Jose, CA
Machine Learning Engineer Intern
Adobe
- Shipped 5+ agent skills, including a four-skill family outside the team's own domain, through an agent skill generator merged into the team's shared skill repo, which turns rough developer requests into standards-compliant skills, grounding against internal prior art so existing work isn't rebuilt and self-auditing against the agentskills.io spec.
- Developed the systematic quality measure for the team's agent recommendation pipeline, a reference-free 6-dimension rubric combining LLM-as-judge with deterministic checks, reproducible to 0.01–0.05 per-dimension variance across repeated runs and surfacing two systematic weaknesses for the team to act on.
- Took that measure to production across the org, porting it onto the internal GenAI evaluation service as a reusable scorer and staging it end-to-end for two enterprise customers, on a four-stage Databricks pipeline built to run it.
-
2025.05 - 2025.08 Remote
Data Engineer Intern
CMind
- Built a LangChain-based RAG prototype over an internal analytics DB (chunking, embeddings, FAISS indexing/retrieval, LLM answer synthesis) with metadata filters and notebook/API utilities.
- Designed CEO-speech trait metrics and an auto-evaluation pipeline with custom rubrics, achieving Fleiss' κ of 0.89–0.91 against human annotations.
Teaching Experience
-
2026.08 - Present Pittsburgh, PA
Graduate Teaching Assistant, 11-711 Advanced Natural Language Processing
Carnegie Mellon University, Language Technologies Institute
- Supporting a graduate NLP research course covering modeling and learning algorithms and culminating in an original research project.
-
2022.08 - 2025.05 Los Angeles, CA
Teaching & Grading Assistant
University of Southern California
- Coordinated logistics and grading for two computer science courses (CSCI-102 and CSCI-360), serving approximately 300 students per term.
- Led weekly office hours and discussions for approximately 20 students, clarifying concepts and providing feedback on assignments.
Projects
- 2026.02 - 2026.04
VLM Hallucination Mitigation
Training-free latent reflection for vision-language models, at Carnegie Mellon University.
- Built a training-free latent reflection module that monitors decoding entropy and injects reflection steps when the model is uncertain.
- Benchmarked against DAMO, VCD, OPERA, and DeGF on POPE/RePOPE/MMBench across LLaVA-1.5-7B, InstructBLIP-7B, and Qwen2-VL-7B.
- On RePOPE, raised LLaVA-1.5-7B accuracy from 84.0% to 90.3% with the lowest calibration error (ECE 0.028) among all methods compared.
- 2024.03 - 2024.04
LLM Prompt Recovery
Developed a system to recover user prompts using fine-tuned Mixtral models, achieving top 3.4% in a Kaggle competition.
- Achieved a score of 0.65 using sentence-T5-base and sharpened cosine similarity.
- Ranked 75/2175 globally in the Kaggle competition.
- Published fine-tuned model on Kaggle for broader access.
- 2024.11 - 2024.12
AI-Based Career Advisor
Built an AI tool to assist users in planning career paths based on skills and interests.
- Developed a Streamlit-based interactive UI integrating GPT-4o for job suggestions.
- Implemented cosine similarity search for skill-job matching.
- Integrated Bing AI for real-time job application link retrieval.
- 2023.08 - 2023.11
Enhancing Debugging Skills of LLMs with Prompt Engineering
Improved LLM debugging capabilities through advanced prompt engineering techniques.
- Experimented with various prompting strategies to enhance debugging efficiency.
- Achieved significant improvements in LLM performance on debugging tasks.
- 2023.09 - 2023.12
Automated Hate Speech Detection in Social Media
Developed an advanced ML model for detecting hate speech, achieving 94% accuracy.
- Fine-tuned BERT for classification tasks.
- Enhanced online safety and inclusivity through robust model optimization.
Skills
| Programming | |
| C++ | |
| Python | |
| Java | |
| MySQL | |
| HTML | |
| CSS | |
| JS | |
| TS | |
| x86-64 Assembly |
| Software | |
| PyTorch | |
| Pandas | |
| NumPy | |
| Git | |
| AWS | |
| LaTeX | |
| Mathematica | |
| Matlab |
| Areas of Expertise | |
| Machine Learning | |
| Natural Language Processing (NLP) | |
| Large Language Models (LLMs) | |
| Vision Language Models (VLMs) | |
| Data Science / Data Engineering |
| Languages | |
| Mandarin (native) | |
| English (professional) |
Awards
- 2025.05
USC CURVE Research Fellowship
USC Viterbi School of Engineering
Awarded in four terms (S24, Su24, F24, S25); total funding $6,750.
- 2024.04
Silver Medal, Kaggle Competition
Ranked 75/2175 (Top 3.4%) on the global leaderboard, LLM Prompt Recovery Project.
- 2022.12
USC Academic Achievement Award
Awarded in Fall 2022, Spring 2023, Spring 2024, and Fall 2024. Covered 11 units of tuition costs in total, amounting to approximately $24,000.
- 2022.04
4th Place, USC Integral Bee Competition
Ranked 4th in the USC Integral Bee Competition.
- 2021.07
1st Prize, International Linguistics Olympiad
Senior Level, Individual Open Round, China.
- 2021.07
1st Prize, International Linguistics Olympiad
Senior Level, Team Open Round, China.