Youngjin Shin

Youngjin Shin

Research Intern · KAIST AI

Focused on understanding and improving AI models through internal mechanisms and representations. Interested in representation learning, 3D visual understanding, and how models reason through structured analysis.

yeongjins916@gmail.com

Scroll

Interests

Representation Learning 3D Visual Understanding Vision-Language Models

Work Experience

Research Intern · Computer Vision Lab
KAIST AI, Seoul · Advisor: Seungryong Kim
2025.09 – Present
Research Intern · Multi-dimensional Insight Lab
Yonsei University, Seoul · Advisor: Sanghoon Lee
2025.01 – 2025.08

Publications

2026
MORPHOS teaser
MORPHOS: Autoregressive 4D Generation with Temporal Structured Latents
Minkyung Kwon, Jinhyeok Choi, Youngjin Shin, Jaeyeong Kim, JongMin Lee, Seungryong Kim
CORAL teaser
CORAL: Correspondence Alignment for Improved Virtual Try-On
Jiyoung Kim, Youngjin Shin, Siyoon Jin, Dahyun Chung, Jisu Nam, Tongmin Kim, Jongjae Park, Hyeonwoo Kang, Seungryong Kim
2025
YOTO teaser
You Only Touch Once: One-Touch System for Personalized 3D Music Video Generation
Kyungjune Lee, Youngjin Shin, Jungwoo Huh, Sanghoon Lee

Education

B.S. in Electrical and Electronic Engineering
Yonsei University
Mar 2020 – Mar 2026
  • GPA: 4.1 / 4.3
  • Military service leave: Nov 2021 – May 2023

Personal Playground

Feature Visualization playground
Feature Visualization
An interactive playground reproducing the technique from Olah et al., “Feature Visualization” (Distill, 2017) — explore channels, neurons, layers, logits, probabilities, and joint optimizations.
Attention Visualization playground
Attention Visualization
An interactive 3D playground for DINOv3-L attention — see how each query, key, and value is projected per layer and head, trace a query to its keys and on to their values, and read the full 261×261 attention map with a per-patch heatmap on the image. A second view projects every layer's MLP features to PCA-RGB, showing how each layer's focus shifts with depth.

Analysis Blog

Feature Convergence and Mixing — where generation and understanding meet
A reading of six models — REPA, RAE, VGGT and their variants — as one question: where on the depth axis does each place, and mix, its features? With layer-wise linear probes and real per-layer feature maps, it shows low-level detail, high-level semantics, and geometry actively blending inside the networks — including a denoiser's excursion into semantics that the final head returns, dictated by the decoder that reads the latent next.
The Reconstruction Paradox — how a frozen encoder's discarded detail comes back
A close reading of RAEv2's decoder on DINOv3-L: a frozen encoder throws detail away, yet its decoder reconstructs pixels. The post traces how the decoder mirrors the encoder in frequency and representation, why the latent (K) decides what can be reconstructed, where the reconstruction–generation trade-off comes from, and what it means to build a decoder that mirrors on purpose.