My name is Tunyu Zhang (ๅผ ็ๅฎ). I am a second-year Ph.D. student in the Department of Computer Science at Rutgers University, advised by Professor Dimitris Metaxas. I am also fortunate to collaborate with Professor Hao Wang. My research focuses on trustworthy and efficient language model systems, including uncertainty estimation in LLM reasoning and efficient decoding for diffusion language models. Previously, I obtained my bachelorโs degree from the University of Science and Technology of China (USTC) in 2025.
๐ฅ News
- 2026.10: ย ๐ Three papers accepted to NeurIPS 2026: T3D on few-step diffusion language models, SAUCE on multi-agent uncertainty estimation, and When Debate Helps on effective multi-agent reasoning!
- 2026.07: ย ๐ I joined Red Hat AI Innovation
as a Research Intern.
- 2026.01: ย ๐ Our paper TokUR was accepted to ICLR 2026
- 2025.09: ย ๐ Our paper TokUR on Bayesian LLM reasoning was accepted to the NeurIPS 2025 Workshop FoRLM!
- 2025.08: ย ๐ I will join Professor Dimitris Metaxasโs group to pursue my PhD degree at Rutgers.
๐ Publications
where โ*โ denotes equal contribution

Turbo Harness: Instance-Adaptive Harness Optimization
Tunyu Zhang*, Hao Wang*, Kai Xu, Dimitris N. Metaxas
Paper | Code
- Turbo Harness tailors an agentโs harness to each task instance, turning past optimization experience into better agent execution.
- A lightweight, RL-trained harness editor uses a reusable playbook to patch a globally optimized harness, with just one editor call per task.
- Improves performance across seven benchmarks spanning interactive agents, software engineering, and long-horizon terminal tasks, while often reducing execution steps and costs.

Sequential Probabilistic Uncertainty Estimation for Parallel Multi-Agent Reasoning Systems
Tunyu Zhang, Zihao Zhao, Yusong Zhao, Haizhou Shi, Zhuohang Li, Haoxian Chen, Hao Wang, Dimitris N. Metaxas
Paper | Code | Data
- SAUCE estimates the reliability of multi-agent reasoning by tracking how consensus evolves throughout an interaction.
- Combines agent agreement and generation uncertainty in a training-free sequential estimator, using existing inference traces without extra sampling.
- Improves error detection, selective prediction, and calibration across five model backbones, five benchmarks, and two multi-agent protocols.

Few-Step Diffusion Language Models via Trajectory Self-Distillation
Tunyu Zhang*, Xinxi Zhang*, Ligong Han, Haizhou Shi, Xiaoxiao He, Zhuowei Li, Hao Wang, Kai Xu, Akash Srivastava, Chengzhi Mao, Hao Wang, Vladimir Pavlovic, Dimitris N. Metaxas
Paper | Code | Slides
- T3D enables high-quality text generation in fewer diffusion steps by distilling a full-step teacherโs generative trajectory into a few-step student.
- Trajectory supervision reduces token factorization error, while Direct Discriminative Optimization (DDO) further strengthens reasoning through a mode-seeking objective.
- Substantially closes the quality gap between few-step and full-step decoding on reasoning and code generation while preserving full-step performance.

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning
Tunyu Zhang*, Haizhou Shi*, Yibin Wang, Hengyi Wang, Xiaoxiao He, Zhuowei Li, Haoxian Chen, Ligong Han, Kai Xu, Huan Zhang, Dimitris Metaxas, Hao Wang
Paper | Code
- We propose TokUR, a framework for token-level uncertainty estimation tailored for LLM reasoning.
- TokUR introduces a low-rank stochastic perturbation mechanism to approximate predictive distributions efficiently.
- The framework enables more reliable multi-step reasoning, and provides uncertainty-aware signals for downstream tasks.

When Debate Helps: Proposal Supply and Verification-Aware Readout in Multi-Agent Reasoning
Zihao Zhao, Tunyu Zhang, Haizhou Shi, Yusong Zhao, Xinxi Zhang, Hao Wang
Paper | Code
- Explains when multi-agent debate outperforms majority voting: agents must propose a correct answer, and the final decision must recognize it even when it comes from a minority.
- Introduces Latent Verification Debate (LVD) to model how verification evidence helps recover correct proposals that voting misses.
- Selects agents with complementary answer coverage, improving reasoning accuracy across two model backbones under matched inference budgets.

Multimodal needle in a haystack: Benchmarking long-context capability of multimodal large language models
Hengyi Wang, Haizhou Shi, Shiwei Tan, Weiyi Qin, Wenyuan Wang, Tunyu Zhang, Akshay Nambi, Tanuja Ganu, Hao Wang
- MMNeedle provides a systematic evaluation framework for long-context multimodal understanding.
- It enables controlled benchmarking of retrieval and reasoning over large visual contexts, and reveals robustness challenges in current multimodal LLMs.
Complex Networks
- Study of nonequilibrium phase transitions mechanisms in exclusive network and node model of heterogeneous assignment based on real experimental data of KIF3AC and KIF3CC motors, EPJP 2022
- Physical mechanisms of exit dynamics in microchannels of nonequilibrium transport systems, IJMP 2024
๐ Teaching
Teaching Assistant, Rutgers University
- Fall 2026 โ CS 440: Introduction to Artificial Intelligence
- Spring 2026 โ CS 440: Introduction to Artificial Intelligence
- Fall 2025 โ CS 344: Design and Analysis of Computer Algorithms
๐ Honors and Awards
- 2025.06 Outstanding Undergraduate Thesis Award, University of Science and Technology of China
- 2022.12 Second Prize, Asia and Pacific Mathematical Contest in Modeling (APMCM)
- 2022.05 Outstanding Student Scholarship (Gold Award), University of Science and Technology of China
๐ Educations
- 2025.09 - present, Rutgers University, New Brunswick.
- 2021.09 - 2025.06, Univeristy of Science and Technology of China, Hefei.
๐ฌ Invited Talks
- 2026.02, Few-Step Diffusion Language Models (Red Hat AI Innovation Team, Random Sample Talk) Slides.
๐ป Internships
- 2026.07 - present, Research Intern at Red Hat AI Innovation
- 2024.06 - 2025.08, Research Assistant at Rutgers University
- 2023.06 - 2024.05, Research Assistant at University of Hong Kong (HKU)