PhD student · ICT, Chinese Academy of Sciences

Haocun Ye叶浩存

Multimodal large models — training, RLVR post-training, and trustworthy evaluation. Advised by Prof. Yiqiang Chen.

“When a benchmark score goes up, did the model actually acquire the ability the benchmark claims to measure?”
Portrait of Haocun Ye

Recent news

  • CRAFTThe Diagnosis Is Already in the Notes — accepted to Findings of EMNLP 2026 (first author).

  • KnowMeBenchV2 accepted to the EMNLP 2026 main conference (co-author).

  • Received the ICT Outstanding Master’s Student Award.

  • Invention patent granted: a training method for large models in ophthalmology.

  • VersaFusion presented at AAAI 2025 (first author).

Selected publications

  1. Conceptual cover of Learning Without Looking

    Learning Without Looking: Image-Dependent Gains from Image-Free RLVR

    Haocun Ye et al.

    In preparation

    Does RLVR teach vision-language models to see, or to better use what they already saw? A train × test visual-visibility study, plus FlipTrack — a counterfactual benchmark of paired images whose answers must flip, so image-blind strategies fail by construction. Code and benchmark will be released with the paper.

  2. CRAFT channel-wise audit overview

    The Diagnosis Is Already in the Notes: A Channel-Wise Audit of Diagnostic Shortcuts in Clinical-Text Benchmarks

    Haocun Ye, Xinlong Jiang, Yuanyuan Hu, Shubai Chen, Tingyu Wu, Zhenkun Zheng, Caijiao Yi, Teng Zhang, Liuxin Wang, Tong Zhao, Heping Xu, Weiwei Dai, Yiqiang Chen

    Findings of EMNLP 2026

    Diagnosis-redacted clinical records still recover the diagnosis at 0.944 macro-AUROC. CRAFT is a portable audit protocol that reports each channel’s shortcut floor, so fusion gains are read against what the text alone already gives away.

  3. KnowMeBenchV2 overview: five-task ladder from information extraction to grounded abstention, with per-task diagnostic results

    KnowMeBenchV2: Evidence-Grounded Person-Centric Long-Video Understanding

    Tingyu Wu, Guangyu Cao, Qizhen Lan, Zhisheng Chen, Ziyan Weng, Bingkun Zhu, Miao Su, Chenglong Li, Zhengwei Xie, Huacan Wang, Sen Hu, Haocun Ye, Nan An, Zaoqu Liu, Ronghao Chen

    EMNLP 2026

    An evaluation-only benchmark for evidence-grounded, person-centric long-video understanding.

  4. VersaFusion editing examples and crop-growth simulation

    VersaFusion: A Versatile Diffusion-Based Framework for Fine-Grained Image Editing and Enhancement

    Haocun Ye, Xinlong Jiang, Chenlong Gao, Bingyu Wang, Wuliang Huang, Yiqiang Chen

    AAAI 2025

    A tuning-free dual-branch diffusion framework — proximal negative-prompt inversion, KV memory injection and classifier guidance — for fine-grained editing that leaves everything else alone. AAAI 39(9), 9445–9453.

Patent · Granted A Training Method for Large Models in Ophthalmology — invention patent behind the ophthalmology LLM training pipeline.

Things I’ve built

Learning Without Looking project cover
01 / 04

Learning Without Looking · FlipTrack

Does RLVR teach vision-language models to see, or to better use what they already saw? A causal train × test visual-visibility study on Qwen2.5-VL, plus FlipTrack — a counterfactual benchmark of paired images whose answers must flip, so any image-blind strategy fails one member of every pair by construction.

  • RLVR
  • GRPO
  • Qwen2.5-VL
  • Benchmark design
CRAFT project cover
02 / 04

CRAFT — The Diagnosis Is Already in the Notes

A channel-wise audit protocol for clinical-text benchmarks: diagnosis-redacted records still recover the diagnosis at 0.944 macro-AUROC, so fusion gains must be reported alongside the text branch’s shortcut floors. Now being engineered into a leakage-audit module for our lab’s data-governance platform.

  • Clinical NLP
  • Shortcut learning
  • Benchmark audit
Ophthalmology LLM training pipeline cover
03 / 04

Ophthalmology LLM Training Pipeline

From 11 open medical QA datasets and 212 OCR’d textbooks to 9.0M ophthalmic SFT samples, then multi-node full-parameter SFT of Qwen2.5-14B and Qwen2-VL-7B with DeepSpeed ZeRO-3 — report-generation CIDEr rose from 0.0004 to 0.557. Industry-sponsored; method granted an invention patent.

  • SFT
  • DeepSpeed ZeRO-3
  • Data pipeline
  • Qwen2.5-14B
VersaFusion and MADA crop-growth simulation cover
04 / 04

VersaFusion & MADA Crop-Growth Simulation

Tuning-free dual-branch diffusion editing (AAAI 2025) — 5.23 s preparation versus 63.56 s for DragDiffusion — deployed on the MADA platform to simulate how a crop will look at later growth stages from a single photo, while keeping the plant’s identity and the scene unchanged.

  • Diffusion models
  • Image editing
  • Stable Diffusion

Beyond the benchmarks

I’m Haocun Ye, a PhD student in artificial intelligence at the Institute of Computing Technology, Chinese Academy of Sciences, advised by Prof. Yiqiang Chen.

I work on multimodal large models — how they are trained, how they are post-trained with reinforcement learning, and how they should be evaluated. The question that drives most of my work is simple to state and hard to answer: when a benchmark score goes up, did the model actually acquire the ability the benchmark claims to measure?

Away from the lab I like experimenting in the kitchen, and I’m happiest on days off spent with family, travelling with friends, or around a board-game table — ideally with a dog nearby. In quieter hours I play single-player games (GTA, The Witcher 3, Sekiro). I know I have my share of flaws, and I try to keep improving — both in how I work with others and in myself.

Education
PhD student, Institute of Computing Technology, Chinese Academy of Sciences (2023 – present)
Advisor
Prof. Yiqiang Chen
Award
ICT Outstanding Master’s Student Award, 2025
Training & post-training
PyTorch · DeepSpeed ZeRO-3 multi-node SFT · GRPO / RLVR (EasyR1) · diffusion inversion & guidance · Qwen2/2.5-VL · InternVL · Gemma
Data & evaluation
LLM-driven data pipelines · OCR-to-QA corpus construction · counterfactual benchmark design · shortcut & leakage audits · LLM-as-judge
Languages & tools
Python · C++ · Linux / Shell · Git · multi-GPU / multi-node training on A800 clusters (RDMA / NCCL)

Get in touch

Open to research collaboration and conversations about multimodal models, post-training, and evaluation. Email is the fastest way to reach me.

Institutional: yehaocun23s@ict.ac.cn