Ramchalam Kinattinkara Ramakrishnan

ML Research @Qualcomm AI Research | MSc @McGill University

profile_pic.jpg

Machine Learning Engineer | Generative AI Efficiency | Model Compression

Generative AI Efficiency (Qualcomm AI Research): My current work focuses on creating efficient solutions for Large and Small Language Models (LLMs/SLMs), multimodal foundation models, diffusion large language models (dLLMs) and other advanced LLM inference techniques like speculative decoding, accelerating on-device training.

Model Compression (Huawei Noah’s Ark Lab): At Noah’s Ark Lab, Montreal, my work entailed developing and implementing neural network model compression techniques to improve computational performance (quantization, pruning, NAS across conv and transformer based architectures).

Academic Research (McGill University): I obtained a thesis-based Master of Science in Machine Learning, where my research was advised by Prof. Mathieu Blanchette, focussing on applying Reinforcement Learning techniques in Bioinformatics.

Passion Project: My passion extends beyond inference to areas like LLM distributed training, reinforcement learning, kernel optimization and other low level op implementation.

Beyond my work, I spend my leisure time playing badminton, ping pong, and passionately following Formula-1 and Manchester United FC.

news

Sep 26, 2025 Paper β€œOmniDraft: A Cross-vocabulary, Online Adaptive Drafter for On-device Speculative Decoding β€œ accepted at NeurIPS 2025. πŸŽ‰βœ¨ [link]
Sep 26, 2024 Paper β€œForward-Forward Algorithm for On-Device Learning β€œ accepted at NeurIPS 2024. πŸŽ‰βœ¨ [link]
Jul 16, 2024 Patent β€œSelective neural network pruning by masking filters using scaling factors β€œ has been granted. πŸŽ‰:smile: [link]
Apr 26, 2022 Paper β€œAn Empirical Study of Low Precision Quantization for TinyML β€œ accepted at tinyML Research Symposium 2022. πŸŽ‰βœ¨ [link]
May 01, 2020 Join Qualcomm (Qualcomm AI Research) as a ML Research Engineer on the Embedded AI team.
Sep 05, 2019 Paper β€œDeep Demosaicing for Edge Implementation β€œ accepted at ICIAR 2019. πŸŽ‰βœ¨ [link]
Oct 18, 2018 Join Noah’s Ark Lab (Huawei), Montreal as a ML Research Engineer on the Accelerated Neural Technology (ANT) team.
Sep 05, 2018 Thesis β€œRlalign: a reinforcement learning approach for multiple sequence alignment β€œ accepted at BIBE 2028. πŸŽ‰βœ¨ [link]
Sep 05, 2018 Paper β€œDetection of Errors in Multi-genome Alignments Using Machine Learning Approaches β€œ accepted at BIBE 2028. πŸŽ‰βœ¨ [link]
Sep 01, 2018 Graduated from McGill University, Montreal, Canada, MSc in CS (Thesis)
Oct 18, 2013 Join Bharat Petroleum Corporation Limited (BPCL), India, Mumbai as a Software Engineer for the LPG team.
Sep 01, 2013 Graduated from Model Engineering College, Thrikkakara, India (CUSAT), B-Tech in CSE

selected publications

2025

  1. OmniDraft: A Cross-vocabulary, Online Adaptive Drafter for On-device Speculative Decoding
    Ramchalam Kinattinkara Ramakrishnan, Zhaocong Yuan, Shaojie Zhuo, and 4 more authors
    In Advances in Neural Information Processing Systems, 2025

2024

  1. Stepping Forward on the Last Mile
    Chen Feng, Shaojie Zhuo, Ramchalam Kinattinkara Ramakrishnan, and 3 more authors
    In Advances in Neural Information Processing Systems, 2024