Assistant Professor Mike Zheng Shou from the Department of Electrical and Computer Engineering is advancing AI's ability to understand and generate video, helping to push video AI from passive analysis towards more capable real-world applications.
In recognition of his contributions to multimodal video intelligence, he has been named an Asian Young Scientist Fellow 2026. The Fellowship will support his continued research on long-form video modelling, Video-Language-Action agents, and video world models for digital and physical environments.
The Asian Young Scientist Fellowship (AYSF) announced 12 Fellows this year, comprising accomplished early-career scientists selected for their exceptional scientific contributions and potential in their respective fields. Awarded annually, the Fellowship supports early-career researchers in the fundamental science disciplines of life sciences, physical sciences, mathematics and computer science. Each Fellow receives US$100,000 over two years to support their research, along with opportunities to participate in the AYSF annual conference, academic activities, and a broader network of young scientists across Asia and globally.
Since joining NUS in 2021 as a Presidential Young Professor, Asst Prof Shou has established a leading research programme in video AI, a critical capability for a wide array of applications, including care robots for the elderly, smart CCTV, AI-assisted video creation, the metaverse, and smart glasses. His research has resulted in nine widely distributed open-source software packages and over 280,000 total model downloads.
His research impact spans:
- Video understanding: Specifically, large-scale video-language pre-training. Unlike prior models trained on third-person web videos, embodied robots and smart glasses require visual comprehension from a first-person perspective—a critical gap addressed by EgoVLP work, the world’s first egocentric video-language pre-trained foundation model.
- Video generation: The Show-1 model was among the earliest open-sourced video diffusion foundation models, significantly reducing GPU memory usage from 72 GB to 15 GB. Tune-A-Video pioneered a new one-shot training paradigm that rapidly adapted pre-trained image diffusion models for video generation, slashing the required training time from weeks to just 10 minutes on a single GPU.
- Unified modelling of understanding and generation: Show-o, a compact AI model which understands images, generates visuals and weaves language all within a single framework, outperforms separate systems ten times its size while accelerating autoregressive model generation by 20x. More about this research can be found at: Show-o, don’t tell: new unified AI model does it all - College of Design and Engineering
In 2025, he was awarded the inaugural Provost’s Innovation Chair Professor Award for his project “AI-powered Video Intelligence and Automation” and was also selected for the Innovation Venture Creation Award, recognising the innovation, translational potential, and real-world impact of his research in multimodal video AI.
The Fellowship will add a further boost to his research trajectory.
Asst Prof Shou said, “With this fellowship, I look forward to advancing our research on Video-Language-Action agents and Video World Models—shifting video AI from passive observer into active agent, such as a digital agent for computer use and a robotic agent for healthcare and scientific discovery."
Asst Prof Shou also holds a joint appointment at the School of Computing.


