World Action Models
Learning predictive models that connect observations, actions, and future physical states for planning and decision making.
About me
Embodied AI Researcher · B.E. Candidate
I’m Junqi Jing (荆浚淇), a B.E. candidate in Software Engineering at Harbin Institute of Technology, expected to graduate in 2027. My research interests center on Embodied AI and world models, with a particular interest in generative approaches to learning and planning for physical interaction.
I am currently an AI Research Intern at KNOWIN AI, where I work on VLA-based dexterous-hand manipulation and exploratory research on world action models. Previously, I spent a semester as an exchange student at POSTECH, conducting research in the MLV Lab with Prof. Kwang In Kim, and later visited Tsinghua University’s LEAP Lab under the guidance of Prof. Gao Huang.
I am currently seeking research-oriented internship opportunities and look forward to connecting with research teams across industry, especially at leading technology companies working on embodied intelligence and robotics.
Research agenda
My work connects visual generation, temporal reasoning, and robot control. I am especially interested in representations that help an embodied agent anticipate what comes next—and choose what to do.
Learning predictive models that connect observations, actions, and future physical states for planning and decision making.
Building vision-language-action systems that translate high-level intent into reliable behavior on real dexterous hands.
Exploring intermediate representations and generative objectives that make physical reasoning more structured, transferable, and useful.
Real-robot work
This reel is built for real hardware footage. Videos play muted when they enter the viewport, pause off-screen, and retain native controls for sound and full-screen viewing.
A real-robot workspace for studying coordinated manipulation, demonstration collection, and human-in-the-loop control.
PLAYBACK
READY
Reserved for the next real-hardware behavior and its task, method, and experimental context.
PLAYBACK
READY
A flexible slot for future experiments—add an MP4 and optional poster without changing the page layout.
MP4, WebM, or direct video URLs can be added without changing this layout.

Selected publication
MTID introduces latent-space temporal interpolation and task-aware masking to improve action-sequence planning between observed start and goal states.
The method supplies visual mid-state supervision directly in latent space, encouraging temporally coherent plans without relying only on text-level supervision.
Research journey
KNOWIN AI
Dexterous-hand manipulation with VLA models and exploratory work on World Action Models.
Tsinghua University · LEAP Lab
Real-hardware practice, teleoperation, and exploration of world-action modeling.
POSTECH · MLV Lab
Embodied AI and generative-model research with Prof. Kwang In Kim.
Harbin Institute of Technology
Research foundations in embodied intelligence, generative AI, and temporal reasoning.
Beyond the lab
Award-winning cross-platform multiplayer game and HIT Best Project of the Year.
Open, collaborative computer-science course materials used by thousands of students.
Qwen-based travel agent recognized at the Alibaba Large Model University Tour.
Field notes
Collaboration
I am open to research internships, technical conversations, and collaborations with industry research teams working on Embodied AI, World Models, and robot learning.