Rakesh Kariya
Software Developer

Applied AI engineer at Votal.ai — building agentic systems, fine-tuning models and deploying them on GPU infrastructure (cloud and on-prem), and running the inference layer that serves them at scale: vLLM, batching and KV-cache tuning, quantization, and latency/throughput work. Also LLM security tooling: AI red teaming and guardrails.
Electronics engineer by training. Previously built a logistics platform from scratch at Shipyaari, EMS for 5G VRAN at Tidalwave, and wore the PM hat at Thinkly.
I like to build things.
Writing
A reproducible emoji-style image experiment: the checkpoint that won blind review, the cost of serving it, and why a polished generic smiley is a failed high-five.
Trying to reproduce L1/LCPO on DeepSeek-R1-Distill-Qwen-1.5B. The budget dial never formed across three runs and 31 hours of A100 time — but MATH-500 reasoning compressed 3.2x for 2.0 accuracy points. A writeup of all three runs, including the two that failed.
Things I built
A hands-on benchmarking lab that quantifies the real performance impact of key LLM-serving optimizations — continuous batching, paged KV cache, prefix caching, and FP8/INT8 quantization. Drives a vLLM server through each configuration with an async load harness and produces concrete before/after comparisons of p95 latency and throughput (up to 2.64× speedup on Llama-3.1-8B).
A Claude Code skill that rewrites emails using three composite writing archetypes — Decisive Operator, Structured Proposer, and Warm Connector. Built on top of Anthropic's Claude API; style exemplars are derived from the MTSlive/musk-v-altman-exhibits dataset on Hugging Face (CC-BY-4.0). Paste a draft; Claude picks the right archetype automatically or you name one. Deployed as a drop-in skill — two curl commands, no config.