Duy (Escanord) Le
I play with LLM inference and post-training.
I am a PhD student at Rice University, advised by Dr. Anshumali Shrivastava, where I make LLMs more efficient, capable, and scalable through the lens of inference.
In practice that means speculative decoding, sparse attention, agent memory, reinforcement learning for post-training, and inference serving infrastructure. I work across that stack, from custom attention kernels to serving systems under real load, and increasingly on co-designing training and inference rather than treating them as separate stages.
I received my BS in Computer Science with a minor in Mathematics from Case Western Reserve University, where I was advised by Dr. Yinghui Wu. I have also been fortunate to work with Dr. Zhaozhuo Xu, and previously with Dr. Xia “Ben” Hu. Before this I worked on graph neural networks, and I still find graph structure a useful lens for thinking about attention mechanism.
Please feel free to reach out at el72 [at] rice [dot] edu if you’d like to talk about any of the above.
Experiences
- Google · Research Intern, Vertex Gemini Inference May 2026 — Aug 2026
- Workato · Research Scientist Intern, reinforcement learning for LLM post-training Feb 2026 — May 2026
- Visa · Staff Research Scientist Intern, loop transformer May 2025 — Aug 2025
- Amazon Web Services · Software Development Engineer Intern, LLM search May 2023 — Aug 2023
news
| Aug 20, 2026 | Forget Without Compromise is accepted to EMNLP 2026 — streaming KV-cache eviction under fixed budgets. |
|---|---|
| May 18, 2026 | Joining the Vertex Gemini Inference team as a Research Intern this summer. |
| May 01, 2026 | FAFO and Scout Before You Attend are both accepted to ICML 2026 — lossy KV-cache compression for lossless inference, and sketch-and-walk sparse attention. |
selected publications
- EMNLPForget Without Compromise: Nexus Sampling for Streaming KV-Cache Eviction Under Fixed BudgetsIn Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026
- ICMLScout Before You Attend: Sketch-and-Walk Sparse Attention for Efficient LLM InferenceIn Proceedings of the 43rd International Conference on Machine Learning (ICML), 2026
- ICMLFAFO: Lossy KV Cache Compression for Lossless Inference Acceleration via Draftless Fumble DecodingIn Proceedings of the 43rd International Conference on Machine Learning (ICML), 2026