Minki Kang
I am a Ph.D. student at the Machine Learning and Artificial Intelligence Lab at KAIST, advised by Prof. Sung Ju Hwang.
During my Ph.D., I completed research internships at NVIDIA and Microsoft, and I completed three years of alternative mandatory military service as Technical Research Personnel, working full-time as a Research Scientist at AITRICS and KRAFTON. More details are listed in Experience.
I am interested in agentic post-training for language models: enabling them to overcome the limits of their parametric capacity by learning to interact with their environments. My research spans distillation, reinforcement learning, and harness optimization. You can read more in my Research Statement.
I am on the job market in Fall 2026, seeking research positions starting in early 2027.
Updates
Released AXPO: Agent Explorative Policy Optimization for Multimodal Agentic Reasoning from my internship at NVIDIA Research.
Started a research internship at NVIDIA Research Taiwan.
Completed a research internship at Microsoft, resulting in ACON: Optimizing Context Compression for Long-horizon LLM Agents.
Returned to the Ph.D. program at KAIST after completing alternative military service.
Selected Publications
AXPO: Agent Explorative Policy Optimization for Multimodal Agentic Reasoning
ACON: Optimizing Context Compression for Long-horizon LLM Agents
T1: Tool-integrated Self-verification for Test-time Compute Scaling in Small Language Models
Distilling LLM Agent into Small Models with Retrieval and Code Tools
Knowledge-Augmented Reasoning Distillation for Small Language Models in Knowledge-Intensive Tasks
Latent Paraphrasing: Perturbation on Layers Improves Knowledge Injection in Large Language Models
Full Publications
As of Aug. 31, 2026
Memory Transfer Learning: How Memories are Transferred Across Domains in Coding Agents
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs
ACON: Optimizing Context Compression for Long-horizon LLM Agents
T1: Tool-integrated Self-verification for Test-time Compute Scaling in Small Language Models
Distilling LLM Agent into Small Models with Retrieval and Code Tools
SafeRoute: Adaptive Model Selection for Efficient and Accurate Safety Guardrails in Large Language Models
HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard Models
Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting
Face-StyleSpeech: Enhancing Zero-shot Speech Synthesis from Face Images with Improved Face-to-Speech Mapping
Latent Paraphrasing: Perturbation on Layers Improves Knowledge Injection in Large Language Models
Knowledge-Augmented Language Model Verification
Knowledge-Augmented Reasoning Distillation for Small Language Models in Knowledge-Intensive Tasks
ZET-Speech: Zero-shot adaptive Emotion-controllable Text-to-Speech Synthesis with Diffusion and Style-based Models
Grad-StyleSpeech: Any-speaker Adaptive Text-To-Speech Synthesis with Diffusion Models
Sparse Token Transformers with Attention Back Tracking
Self-Distillation for Further Pre-training of Transformers
KALA: Knowledge-Augmented Language Model Adaptation
Edge Representation Learning with Hypergraphs
Learning to Perturb Word Embeddings for Out-of-distribution QA
Accurate Learning of Graph Representations with Graph Multiset Pooling
Neural Mask Generator: Learning to Generate Adaptive Word Maskings for Language Model Adaptation
Episodic Memory Reader: Learning What to Remember for Question Answering from Streaming Data
Journal Publications
Rethinking Reward Models for Multi-Domain Test-Time Scaling
Preprints
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients
TIDE: Proactive Multi-Problem Discovery via Template-Guided Iteration
OmniRetrieval: Unified Retrieval across Heterogeneous Knowledge Sources
AXPO: Agent Explorative Policy Optimization for Multimodal Agentic Reasoning
Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR
PREPING: Building Agent Memory without Tasks
THINKSAFE: Self-Generated Safety Alignment for Reasoning Models
Workshop Publications
Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks
Knowledge-Consistent Dialogue Generation with Knowledge Graphs
*: equal contribution.
Work
Experience
NVIDIA
Research Intern at NVIDIA Research Taiwan
Microsoft
Research Intern at Microsoft 365 Research Team
KRAFTON
Part-time Researcher at Natural Language DL Team
KRAFTON
Applied Research Scientist at Natural Language DL Team
AITRICS
Applied Research Scientist at Virtual Human Team
Machine Learning and Artificial Intelligence Lab, KAIST
Research Intern advised by Prof. Sung Ju Hwang
Kakao
Engineering Intern at Context Department
Training
Education
KAIST
Ph.D. in Kim Jaechul Graduate School of AI
KAIST
M.A. in Kim Jaechul Graduate School of AI
Technical University of Munich
Exchange Student in Electrical Engineering and Information Technology
KAIST
B.S. in Electrical Engineering and Computer Science
