Rakesh Kariya

Software Developer

Rakesh Kariya

Applied AI engineer at Votal.ai — building agentic systems, fine-tuning models and deploying them on GPU infrastructure (cloud and on-prem), and running the inference layer that serves them at scale: vLLM, batching and KV-cache tuning, quantization, and latency/throughput work. Also LLM security tooling: AI red teaming and guardrails.

Electronics engineer by training. Previously built a logistics platform from scratch at Shipyaari, EMS for 5G VRAN at Tidalwave, and wore the PM hat at Thinkly.

I like to build things.

Writing

What a 16-image LoRA taught us about post-training an image model

A reproducible emoji-style image experiment: the checkpoint that won blind review, the cost of serving it, and why a polished generic smiley is a failed high-five.

Three GRPO runs to give a 1.5B model a thinking-budget dial

Trying to reproduce L1/LCPO on DeepSeek-R1-Distill-Qwen-1.5B. The budget dial never formed across three runs and 31 hours of A100 time — but MATH-500 reasoning compressed 3.2x for 2.0 accuracy points. A writeup of all three runs, including the two that failed.

Things I built

vllm-serving-lab

A hands-on benchmarking lab that quantifies the real performance impact of key LLM-serving optimizations — continuous batching, paged KV cache, prefix caching, and FP8/INT8 quantization. Drives a vLLM server through each configuration with an async load harness and produces concrete before/after comparisons of p95 latency and throughput (up to 2.64× speedup on Llama-3.1-8B).

claude-skill-voiceprint

A Claude Code skill that rewrites emails using three composite writing archetypes — Decisive Operator, Structured Proposer, and Warm Connector. Built on top of Anthropic's Claude API; style exemplars are derived from the MTSlive/musk-v-altman-exhibits dataset on Hugging Face (CC-BY-4.0). Paste a draft; Claude picks the right archetype automatically or you name one. Deployed as a drop-in skill — two curl commands, no config.