Luxi (Lucy) He

prof_pic.jpg

Hi and welcome to my homepage! I’m a CS Ph.D. candidate at Princeton University, where I’m fortunate to be co-advised by Prof. Danqi Chen and Prof. Peter Henderson. My current research focuses on understanding language models and improving their alignment and safety. I have studied model behaviors, the impact of data across the language model lifecycle, and have worked on human-AI collaboration topics. A lot of my work is motivated by real-world impact and insights from both tech and policy.

I am currently a Research Fellow at Anthropic, and I was previously a Student Researcher at Google.

Before Princeton, I obtained my Bachelor’s degree from Harvard with Highest Honors in Computer Science & Mathematics and a concurrent Master’s in Applied Math.

Outside of research, I’m a singer, dancer, traveler, and amateur food blogger.

Email: luxihe at princeton.edu

news

2026-09 Our paper on measuring and suppressing misaligned model behaviors accepted to NeurIPS! Our work making LLM judges more consistent multi-rule interpreters accepted for Oral Presentation at NeurIPS AI4GOOD workshop. More releases coming soon!
2026-04 Gave an invited talk at RedHat AI & MIT-IBM Watson AI Lab.
2025-12 Our workshop on Navigating and Addressing Data Problems for Foundation Models has been accepted to ICLR 2026! The workshop will take place on April 26th, 2026.
2025-10 Gave an oral presentation of our AudioLM evaluation paper at AIES 2025.
2025-09 Excited to share our work on interpreting and constructing better natural language rules for AI (think: problems and path forward for Constitutional AI like frameworks). Don’t miss the accompanying X thread, blog post, and policy brief!

selected publications

Please see my Google Scholar page for the full publication list.

  1. CAI_cover.png
    Statutory Construction and Interpretation for Artificial Intelligence
    Luxi He*, Nimra Nadeem*, Michel Liao, Howard Chen, Danqi Chen , and 2 more authors
    PNAS 2026, NeurIPS 2025 RegML Workshop (Oral),
  2. pearl_overview.png
    Gone for Good? Measuring and Strengthening Misaligned Behaviors Suppression
    Luxi He*, Pengcheng Jiang*, Jifan Zhang, Jiawei Han, and Danqi Chen
    NeurIPS, 2026
    Release coming soon
  3. inconsistent_rates.png
    Aligning Judge Models to be Logically Consistent Multi-Rule Interpreters
    Michel Liao, Nimra Nadeem, Luxi He, and Peter Henderson
    NeurIPS AI4Good Workshop (Oral), 2026
    Release coming soon
  4. dataprophet_cover.png
    DataProphet: Demystifying Supervision Data Generalization in Multimodal LLMs
    Xuan Qi, Luxi He, Dan Roth, and Xingyu Fu
    ICLR, 2026
  5. audiolm_illustration.png
    The Model Hears You: Audio Language Model Deployments Should Consider the Principle of Least Privilege
    Luxi He*, Xiangyu Qi*, Michel Liao, Inyoung Cheong, Prateek Mittal , and 2 more authors
    AIES (Oral), 2025
  6. copycat_cover.png
    Fantastic Copyrighted Beasts and How (Not) to Generate Them
    Luxi He*, Yangsibo Huang*, Weijia Shi*, Tinghao Xie, Haotian Liu , and 5 more authors
    ICLR 2025; ICML GenLaw Workshop (Spotlight), 2025
  7. MeCo_cover.png
    Metadata Conditioning Accelerates Language Model Pre-training
    Tianyu Gao, Alexander Wettig, Luxi He, Yihe Dong, Sadhika Malladi , and 1 more author
    ICML, 2025
  8. benign_data_safety.png
    What is in Your Safe Data? Identifying Benign Data that Breaks Safety
    Luxi He*, Mengzhou Xia*, and Peter Henderson
    Conference on Language Modeling (COLM), ICLR Data Problems in Foundation Model (Best Paper), 2024
  9. charxiv_cover.png
    CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
    Zirui Wang, Mengzhou Xia, Luxi He, Howard Chen, Yitao Liu , and 8 more authors
    NeurIPS Datasets & Benchmarks, 2024
  10. fairfront_cover.png
    Aleatoric and Epistemic Discrimination: Fundamental Limits of Fairness Interventions
    Hao Wang, Luxi He, Rui Gao, and Flavio Calmon
    In NeurIPS (Spotlight) , 2023