Duy (Escanord) Le

I play with LLM inference and post-training.

escanord.jpeg

I am a PhD student at Rice University, advised by Dr. Anshumali Shrivastava, where I make LLMs more efficient, capable, and scalable through the lens of inference.

In practice that means speculative decoding, sparse attention, agent memory, reinforcement learning for post-training, and inference serving infrastructure. I work across that stack, from custom attention kernels to serving systems under real load, and increasingly on co-designing training and inference rather than treating them as separate stages.

I received my BS in Computer Science with a minor in Mathematics from Case Western Reserve University, where I was advised by Dr. Yinghui Wu. I have also been fortunate to work with Dr. Zhaozhuo Xu, and previously with Dr. Xia “Ben” Hu. Before this I worked on graph neural networks, and I still find graph structure a useful lens for thinking about attention mechanism.

Please feel free to reach out at el72 [at] rice [dot] edu if you’d like to talk about any of the above.

Experiences

  • Google · Research Intern, Vertex Gemini Inference May 2026 — Aug 2026
  • Workato · Research Scientist Intern, reinforcement learning for LLM post-training Feb 2026 — May 2026
  • Visa · Staff Research Scientist Intern, loop transformer May 2025 — Aug 2025
  • Amazon Web Services · Software Development Engineer Intern, LLM search May 2023 — Aug 2023

news

Aug 20, 2026 Forget Without Compromise is accepted to EMNLP 2026 — streaming KV-cache eviction under fixed budgets.
May 18, 2026 Joining the Vertex Gemini Inference team as a Research Intern this summer.
May 01, 2026 FAFO and Scout Before You Attend are both accepted to ICML 2026 — lossy KV-cache compression for lossless inference, and sketch-and-walk sparse attention.

selected publications

  1. EMNLP
    Forget Without Compromise: Nexus Sampling for Streaming KV-Cache Eviction Under Fixed Budgets
    Duc Duong*, Hoang Anh Duy Le*, Jianwen Xie, Anshumali Shrivastava, and Zhaozhuo Xu
    In Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026
  2. ICML
    Scout Before You Attend: Sketch-and-Walk Sparse Attention for Efficient LLM Inference
    Hoang Anh Duy Le, Sahil Joshi, Zeyu Yang, Zhaozhuo Xu, and Anshumali Shrivastava
    In Proceedings of the 43rd International Conference on Machine Learning (ICML), 2026
  3. ICML
    FAFO: Lossy KV Cache Compression for Lossless Inference Acceleration via Draftless Fumble Decoding
    Hoang Anh Duy Le*, Shaochen Zhong*, Yifan Lu, Yingtong Dou, Jiayi Yuan, Yu-Neng Chuang, Xiran Fan, Guanchu Wang, Yuzhong Chen, and Xia Hu
    In Proceedings of the 43rd International Conference on Machine Learning (ICML), 2026