Owen Reed
owenreed
AI & ML interests
Efficient LLM inference, KV cache optimization, quantization, speculative decoding, model pruning
Recent Activity
upvoted a paper about 14 hours ago
BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference upvoted a paper about 14 hours ago
Language Models Can Control Their Own Attention upvoted a paper about 14 hours ago
Unlocking Lossless Speedups in LLMs via Discrete DiffusionOrganizations
None yet