·
AI & ML interests
None yet
Organizations
view article Understanding Post-Training Quantization with LLM Compressor
view article Accelerating LLM Inference: Fused INT8 Weight-Only Quantization in Pallas
view article Breaking the O(N^2) Bottleneck: Implementing High-Performance Block-Sparse Attention with JAX/Pallas
view article Running Native PyTorch on TPUs with Zero Code Changes
published an article about 1 year ago view article Thinking Outside the Attention Box: Introducing Gated Associative Memory (GAM)
rishiraj
• • 5
published an article about 1 year ago view article Understanding Gemma 3n: How MatFormer Gives You Many Models in One
rishiraj
• • 50
published an article about 1 year ago view article Why Maybe We're Measuring LLM Compression Wrong
rishiraj
• • 17
published an article over 1 year ago view article What changed in the Transformer architecture
rishiraj
• • 18
published an article over 1 year ago view article A simple implementation of the attention mechanism in JAX
rishiraj
• • 2
published an article over 2 years ago view article Indexify: Bringing HuggingFace Models to Real-Time Pipelines for Production Applications
rishiraj
• • 7
published an article over 2 years ago view article Combating Evaluation Data Contamination in LLMs: Strategies for High-Quality Finetuning and Model Merging
published an article almost 3 years ago view article Fine-Tuning LLMs: Supervised Fine-Tuning and Reward Modelling
rishiraj
• • 7
published an article almost 3 years ago view article AutoTrain Advanced now supports Experiment Tracking
published an article almost 3 years ago view article Optimizing Convolutional Neural Networks with Mojo - Part 1
rishiraj
• • 3