Drag to turn the room. Tap the floor to send the robot somewhere.

Robot's view

Inspired by 63days.github.io

About me

I am Phillip Lee, a second year PhD student at KAIST Graduate School of AI, advised by Prof. Minhyuk Sung. I am also a student researcher at Google DeepMind, Mountain View. I received both my Masters and Bachelors degree at KAIST. I am also fortunate to closely collaborate with Prof. Leonidas Guibas. I am a recipient of Qualcomm Innovation Fellowship Korea 2025.

My research interests revolve around analyzing and enhancing the spatial understanding of multimodal foundation models, with the broader goal of building truly autonomous and embodied agents capable of operating in the physical world. Specifically, I have worked on advancing the 3D spatial reasoning capabilities of vision-language models (VLMs) and building more robust evaluation systems for spatial intelligence. Previously, I have also worked on generative models for image and video generation.

I am open to collaboration opportunities! Please feel free to contact me via email. My CV is at Curriculum Vitae (CV).

Work experience

Google

Student Researcher, Google DeepMind (Mountain View, CA)

May 2026 – Current (Host: Leonidas Guibas)

News

Publications

  1. Hidden Sensitivity in Spatial Reasoning Evaluation: Diagnosis and Re-ranking with VSI-Bench

    Phillip Y. Lee, Jin Yoo, Minseo Kim, Leonidas Guibas, Minhyuk Sung

    Combining Theory and Benchmark (CTB) Workshop at ICML 2026

  2. Token Warping Helps MLLMs Look from Nearby Viewpoints

    Phillip Y. Lee*, Chanho Park*, Mingue Park, Seungwoo Yoo, Juil Koo, Minhyuk Sung

    CVPR 2026

  3. Toward Ambulatory Vision: Learning Visually-Grounded Active View Selection

    Juil Koo*, Daehyeon Choi*, Sangwoo Youn*, Phillip Y. Lee, Minhyuk Sung

    Reinforcement Learning from World Feedback (RLxF) Workshop at ICML 2026

  4. DiverseVAR: Balancing Diversity and Quality of Next-Scale Visual Autoregressive Models

    Mingue Park*, Prin Phunyaphibarn*, Phillip Y. Lee, Minhyuk Sung

    ECCV 2026

  5. Unconditional Priors Matter! Improving Conditional Generation of Fine-Tuned Diffusion Models

    Prin Phunyaphibarn, Phillip Y. Lee, Jaihoon Kim, Minhyuk Sung

    WACV 2026

  6. Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation

    Phillip Y. Lee, Jihyeon Je, Chanho Park, Mikaela Angelina Uy, Leonidas Guibas, Minhyuk Sung

    ICCV 2025 / Human to Robot (H2R) Workshop at CoRL 2025

  7. GrounDiT: Grounding Diffusion Transformers via Noisy Patch Transplantation

    Phillip Y. Lee*, Taehoon Yoon*, Minhyuk Sung

    NeurIPS 2024

  8. ReGround: Improving Textual and Spatial Grounding at No Cost

    Phillip Y. Lee, Minhyuk Sung

    ECCV 2024

  9. SyncDiffusion: Coherent Montage via Synchronized Joint Diffusions

    Phillip Y. Lee, Kunho Kim, Hyunjin Kim, Minhyuk Sung

    NeurIPS 2023

Teaching

Academic services

Achievements