May 2026 - Present
Undergraduate Researcher — Algoverse AI Research Program
Remote- →Latent failure prediction in long-horizon LLM agents — paper in preparation for NeurIPS.
- →Reproduced a published method for predicting whether an AI coding agent will succeed, matching the foundation paper's accuracy within 8%, by rebuilding its value-axis signal from scratch inside the Qwen3-8B model.
- →Extended the signal to long-horizon coding agents, a regime the original paper never evaluated, by engineering a Python pipeline that extracts model activations at every step of 60-step coding tasks.
- →Implemented statistical measures of how reliably a value-axis signal predicts success across LLM agent attempts at fixing real GitHub bugs from SWE-bench.



