About
I am an LLM post-training researcher at StepFun. I received my B.Eng. in Computer Science (Innovation Class) from Shenzhen University. My research centers on efficient inference and context compression for large language models. I am always happy to discuss research and collaboration — feel free to email me.
Biography
My current work centers on agentic LLMs and the aesthetics of their generated artifacts — how agents plan, act, and produce outputs that are not only correct but well-crafted and visually refined.
My previous work spanned large model pre-training, efficient inference, context compression, and multimodal AI companions.
Publications
* denotes equal contribution. Full list on Google Scholar.
-
Proact-VL: A Proactive VideoLLM for Real-Time AI CompanionsA proactive video LLM that responds in real time for AI companion scenarios.ICML 2026 — Poster -
Pretraining Context Compressor for Large Language Models with Embedding-Based MemoryPretraining a context compressor with embedding-based memory for efficient long-context LLM inference.ACL 2025 (Main Conference)[Code] -
WebVR: Benchmarking Multimodal LLMs for WebPage Recreation from Videos via Human-Aligned Visual RubricsA benchmark for recreating webpages from videos with human-aligned visual rubrics.arXiv:2603.13391[Paper]
Experience
- StepFun, Beijing, China — Foundation Model Post-training Intern. Worked on reward modeling and RLHF pipelines for Step series models.
- Microsoft Research Asia, Beijing, China — Research Intern, Social Computing Group. 🎉 Awarded "Star of Tomorrow" Internship Award.
- Tencent, Shenzhen, China — Software Engineering Intern.
- RoboMaster, China (Dongguan / Changsha / Shenzhen) — Vision Team Member. Developed perception and vision modules for autonomous robotic systems in RoboMaster competitions.