Foundations
Floating point, vectors, matrices, neural networks, embeddings.
Machine Learning · LLM Systems · AI Infrastructure
I’m Pratheek Vadla, a machine learning engineer interested in LLMs, inference systems, recommendation systems, GPUs, and the infrastructure required to run AI reliably in production.
Featured series
A structured journey from numbers and neural networks to transformers, inference engines, GPUs, and production serving.
Before FP16, BF16, quantization, and tensor cores, we need to understand how computers represent numbers in the first place.
Floating point, vectors, matrices, neural networks, embeddings.
Attention, tokenization, pretraining, fine-tuning, alignment.
KV cache, batching, PagedAttention, quantization, parallelism.
GPUs, Kubernetes, model serving, observability, cost and scale.
Research
My interests sit at the intersection of machine learning systems, recommendation, retrieval, and large language models.
Research page →Projects
This section will collect hands-on work around inference, retrieval, distributed systems, and model serving.
Projects page →About
I work on production machine learning systems and write to understand difficult ideas from first principles. This site is where I document what I learn and build.