Publications
2026
- PreprintTest-Time Optimization of Query Embeddings with Ranking Aware Reward MaximizationarXiv preprint arXiv:2608.12569. (Collaboration with Google DeepMind), 2026Distills ranking rewards into a learned vector in a frozen retriever’s embedding space: up to +8.36 nDCG@10 across 15 MTEB retrieval tasks, generalizing to unseen queries and tasks with no weight updates.
- Preprint
- ICML 2026REAL: Regression-Aware Reinforcement Learning for LLM-as-a-JudgeThe 43rd International Conference on Machine Learning. (Collaboration with Google DeepMind), 2026Regression-aware policy gradients for LLM-as-a-Judge: +8.40 Pearson and +7.20 Spearman over the SFT baseline on Qwen3-32B, with stronger out-of-domain generalization.
- ICLR 2026EdiVal-Agent: An Object-Centric Framework for Automated, Scalable, Fine-Grained Evaluation of Multi-Turn EditingThe 14th International Conference on Learning Representations, 2026An object-centric VLM agent that automates fine-grained, multi-turn image-editing evaluation without human annotators; EdiVal-Bench covers 9 instruction types and 16 editing models.
- ICLR 2026Score Distillation Beyond Acceleration: Generative Modeling from Corrupted DataThe 14th International Conference on Learning Representations, 2026Score distillation is not merely an acceleration technique — it enhances generation quality from corrupted data, both empirically and theoretically.
2025
- NeurIPS 2025Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout ReplayAdvances in Neural Information Processing Systems 2025, 2025Difficulty-targeted online data selection and rollout replay cut GRPO fine-tuning time by 23-62% while matching the original performance.
- NeurIPS 2025
2024
- PreprintEnhancing and Accelerating Diffusion-Based Inverse Problem Solving through Measurements OptimizationSubmitted., 2024
- PNASJoint trajectory inference for single-cell genomics using deep learning with a mixture priorProceedings of the National Academy of Sciences, 2024A VAE with a hierarchical prior offers a comprehensive pipeline for integrating multi-omic data, correcting batch effects, inferring pseudotime, and conducting differential analysis.