PhD Candidate · University of Maryland, College Park
Aakriti
Agrawal
Making large models reason reliably — and verifying every step they take to get there.
I am a final-year PhD candidate advised by Prof. Furong Huang. I have also worked closely with Prof. Dinesh Manocha and Prof. Amrit Singh Bedi. My work spans process reward models, step-level supervision, weak-to-strong generalization, and alignment in language and vision-language models.
01
What I work on
Improving the reasoning ability of LLMs and reducing reward hacking in reasoning models.
Reasoning in diffusion language models, including learning the order in which thoughts unfold.
Weak-to-strong generalization and self-improvement: getting stronger models out of weaker supervisors.
Alignment and reasoning in LLMs and vision-language models, with a focus on reducing hallucinations.
Reinforcement learning, multi-agent systems, and uncertainty estimation.
02
Recent news
Updated 2026Scheduling Thoughts accepted at ICML 2026.
VisAlign and EnsemW2S accepted at ACL 2026.
OC-PRM accepted as a poster at the AFAA Workshop @ ICLR 2026 — with a recommendation of Oral from the AC.
VisAlign accepted as a poster at the MM Intelligence Workshop @ ICLR 2026.
Passed my prelim exam and am officially a PhD candidate. Talk: Towards Reliable Reasoning and Alignment in Large Models.
Paper on uncertainty-aware answer selection across multiple LLMs accepted at EMNLP 2025.
One paper accepted at NeurIPS 2025.
Completed a Fall '24–Spring '25 internship at Capital One on reward hacking in reasoning LLMs.
Completed a summer internship at Dolby on reducing hallucinations in video LLMs.
Amazon internship paper accepted at Interspeech 2023.
03
Publications
VeriGate: Verifier-Gated Step-Level Supervision for GRPO
Scheduling Thoughts: Learning the Order of Thought in Diffusion Language Models
Uncertainty-Aware Answer Selection for Improved Reasoning in Multi-LLM Systems
EnsemW2S: Can an Ensemble of SoTA LLMs be Leveraged to Obtain a Stronger LLM?
Easy2Hard-Bench: Standardized Difficulty Labels for Profiling LLM Performance and Generalization
Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings
WAVES: Benchmarking the Robustness of Image Watermarks
PoisonedParrot: Subtle Data Poisoning Attacks to Elicit Copyright-Infringing Content from Large Language Models
Robustness to Multi-Modal Environment Uncertainty in MARL using Curriculum Learning
Learning When to Trust Which Teacher for Weakly Supervised ASR
Revisiting Parameter Sharing in Multi-Agent Deep Reinforcement Learning
RTAW: An Attention Inspired Reinforcement Learning Method for Multi-Robot Task Allocation in Warehouse Environments
DC-MRTA: Decentralized Multi-Robot Task Allocation and Navigation in Complex Environments
Accurate Estimation of 3D-Repetitive-Trajectories using Kalman Filter, Machine Learning and Curve-Fitting for High-Speed Target Interception
Mid-Flight Propeller Failure Detection and Control of Propeller-Deficient Quadcopter using Reinforcement Learning
A Comparative Study of Noise Cancellation Using LMS Adaptive Filter and RNN Filter
04
Industry research
Research intern working on reward hacking in reasoning LLMs, which led to the PRISM work on process reward models.
Research intern on reducing hallucinations in video and vision-language models.
Research intern on weakly supervised ASR, published at Interspeech 2023.
Let's talk
Reach out if you would like to collaborate, if you are hiring, or if you are a student looking for mentorship. I read every email.