publications
* denotes equal contribution. see also my Google Scholar profile.
2026
- EMNLPForget Without Compromise: Nexus Sampling for Streaming KV-Cache Eviction Under Fixed BudgetsIn Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026
- ICMLScout Before You Attend: Sketch-and-Walk Sparse Attention for Efficient LLM InferenceIn Proceedings of the 43rd International Conference on Machine Learning (ICML), 2026
- ICMLFAFO: Lossy KV Cache Compression for Lossless Inference Acceleration via Draftless Fumble DecodingIn Proceedings of the 43rd International Conference on Machine Learning (ICML), 2026
- ACLAutoL2S: Auto Long-Short Reasoning for Efficient Large Language ModelsIn Findings of the Association for Computational Linguistics: ACL 2026, 2026
- arXivWhen Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement LearningarXiv preprint arXiv:2607.07976, 2026
- arXivLearning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement LearningarXiv preprint arXiv:2606.22305, 2026
- arXivSOCKET: SOft Collision Kernel EsTimator for Sparse AttentionarXiv preprint arXiv:2602.06283, 2026
- TMLRThe LLM Data Auditor: A Metric-Oriented Survey on Quality and Trustworthiness in Evaluating Synthetic DataTransactions on Machine Learning Research (TMLR), 2026
- PreprintSweeping Promptable Spoofs under the DirtyRAG: A Practical, Query-Blind RAG Attack Done RightPreprint, 2026
2025
- EMNLPWord Salad Chopper: Reasoning Models Waste A Ton Of Decoding Budget On Useless Repetitions, Self-KnowinglyIn Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2025
- ACLReasonerRank: Redefining Language Model Evaluation with Ground-Truth-Free Ranking FrameworksIn Findings of the Association for Computational Linguistics: ACL 2025, 2025
- PreprintGraph Transformers Get the GIST: Graph Invariant Structural Trait for Refined Graph EncodingPreprint, 2025
2024
- EMNLPKV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable ApproachesIn Findings of the Association for Computational Linguistics: EMNLP 2024, 2024
- ICMLKnowledge Graphs Can be Learned with Just Intersection FeaturesIn Proceedings of the 41st International Conference on Machine Learning (ICML), 2024
- ICMLGNNs Also Deserve Editing, and They Need It More Than OnceIn Proceedings of the 41st International Conference on Machine Learning (ICML), 2024
- ICDEGraphLingo: Domain Knowledge Exploration by Synchronizing Knowledge Graphs and Large Language ModelsIn 2024 IEEE 40th International Conference on Data Engineering (ICDE), 2024