I am Jiabo Zhan (詹佳博), an incoming M.Eng. student in Computer Technology at Tsinghua University, advised by Prof. Chun Yuan. I received my B.E. in Software Engineering from Beihang University in 2026, graduating 2nd out of 172 students (top 1.2%) with a GPA of 3.873/4.0.

My research interests include computer vision, multimodal large language models (MLLMs), reinforcement learning, AI-generated content (AIGC), and document intelligence. My recent work focuses on transparent and layered RGBA generation, efficient document parsing, multimodal data construction, and evaluation.

My publications have received 14 Google Scholar citations. Citation counts are updated automatically every day. Google Scholar citations

Contact
WeChat: _Marsquakes

🔥 News

  • 2026.09:  WeVisDoc was released as a technical report on robust end-to-end document parsing, with open-source code and 2B/4B model weights.
  • 2026.08:  PaDoc was released with open-source code and model weights.
  • 2026.08:  🎉 OmniAlpha was accepted to ACM Multimedia 2026 (ACM MM 2026) as an Oral Presentation!
  • 2026.04:  OmniAlpha was substantially revised with multi-task reinforcement learning for transparency-aware generation.
  • 2025.11:  OmniAlpha was released with open-source code and model weights.
  • 2025.07:  AlphaVAE was released with open-source code, data, and models.

📝 Publications

ACM MM 2026ORAL
OmniAlpha examples

OmniAlpha: Aligning Transparency-Aware Generation via Multi-Task Unified Reinforcement Learning

Hao Yu*, Jinglin Wang*, Jiabo Zhan*, Rui Chen, Zile Wang, Huaisong Zhang, Hongyu Li, Xinrui Chen, Yongxian Wei, Chun Yuan

* Equal contribution. ACM Multimedia 2026. ORAL PRESENTATION

Paper / PDF / Code / Model / Scholar

  • A unified framework for transparency-aware generation and manipulation across image matting, object removal, layer decomposition, and RGBA generation.
  • Contribution: Co-designed the multi-task data schema and evaluation pipeline, curated training data, and contributed to multi-task SFT and GRPO-style RL post-training; achieved a 9.07% relative reduction in RGB L1 on layer decomposition.

PaDoc: Layout-Grounded Parallel Decoding for Document Parsing

Hao Yu*, Jiabo Zhan*, Kang Liu, Linnan Zhao, Dongxu Yue, Rui Chen, Jinglin Wang, Chong Sun, Chen Li, Jing Lyu, Chun Yuan

* Equal contribution.

Paper / PDF / Code / Model / Scholar

  • A layout-grounded parser that branches regional content from shared full-page prefixes, improving throughput by 67.4–118% over same-backbone sequential decoding while retaining top-tier parsing quality.

AlphaVAE: Unified End-to-End RGBA Image Reconstruction and Generation with Alpha-Aware Representation Learning

Zile Wang, Hao Yu, Jiabo Zhan, Chun Yuan

Paper / PDF / Code / Model / Scholar

  • An alpha-aware VAE for end-to-end RGBA reconstruction and transparent image generation, accompanied by the ALPHA evaluation benchmark.
  • Contribution: Curated and governed 8,124 high-quality RGBA training samples and helped design the benchmark, evaluation code, and LoRA data iteration workflow.

Technical Reports

WeVisDoc-4B benchmark results on OmniDocBench and PureDocBench

WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing

Hao Yu*, Kang Liu*, Linnan Zhao*, Jiabo Zhan*, Chong Sun, Chen Li, Jing Lyu

* Equal contribution (co-first authors). Technical Report, 2026.

Report / PDF / Project / Code / Model (2B) / Model (4B)

  • A two-stage data-centric framework that expands document coverage, then uses residual-error diagnostics to guide targeted data construction and training-budget allocation for robust end-to-end parsing.
  • WeVisDoc-4B achieves an Overall score of 95.38 on OmniDocBench v1.6 and a mean Overall score of 75.54 across the three PureDocBench tracks, ranking first among the compared end-to-end parsers in all four settings.

🎓 Education

Tsinghua University, Shenzhen International Graduate School
Sep. 2026 – Jun. 2028 (expected)
M.Eng. in Computer Technology · Advisor: Prof. Chun Yuan

Research focus: computer vision, multimodal large language models, and reinforcement learning.

Beihang University, School of Software
Sep. 2022 – Jun. 2026
B.E. in Software Engineering · GPA: 3.873/4.0 · Weighted average: 93.005/100

Ranked 2/172 (top 1.2%) and admitted to Tsinghua University for graduate study through recommendation, exempt from the national entrance examination.

💼 Research & Industry Experience

VFlow AI Video Generation · Tsinghua University, Shenzhen International Graduate School
Dec. 2025 – Jun. 2026
Algorithm Intern · Industry project led by Prof. Chun Yuan
  • Helped build a page-level multimodal pipeline that turns product pages, images, and marketing copy into structured attributes, scripts, and generated videos.
  • Built and cleaned image-text and instruction data, iterated prompt templates, and conducted offline evaluation for visual grounding, structured extraction, and multi-turn generation.
Momenta
Sep. 2025 – Nov. 2025
Algorithm Intern · Autonomous-driving road-condition classification
  • Developed a Qwen3-VL classification approach that combined camera images with structured rosbag signals such as speed and position; built the evaluation set, prompts, and offline evaluation pipeline.
  • Migrated a high-cost Gemini-2.5 Pro workflow to a deployable open-source VLM solution, reducing classification cost by approximately 80% while accepting an accuracy change from 80% to 70%.

🏆 Honors & Awards

  • National Scholarship, 2022–2023 and 2023–2024 academic years.
  • Outstanding Student at Beihang University for three consecutive academic years; Outstanding Communist Youth League Member.
  • Samsung Scholarship, 2024–2025; First Prize in Beijing, Chinese Mathematics Competitions for College Students.

🛠 Skills

Models & Training
VLM, VAE, DiT, multi-task SFT, GRPO-style RL; familiar with DPO and PPO.

Data & Evaluation
Dataset construction and governance, benchmark design, evaluation pipelines, offline evaluation, and error analysis.

Multimodal Understanding
Visual grounding, OCR, layout analysis, key information extraction, and image-text-structured signal modeling.

Engineering
Python, PyTorch, C/C++, and Java; multimodal pipeline engineering and product deployment.

Languages
Chinese (native); English (CET-4: 662, CET-6: 633).

Service
Teaching assistant for College English and Compiler Technology; 250 hours of volunteering in Beijing.