Machine Learning · LLM Systems · AI Infrastructure

I build and write about large-scale machine learning systems.

I’m Pratheek Vadla, a machine learning engineer interested in LLMs, inference systems, recommendation systems, GPUs, and the infrastructure required to run AI reliably in production.

Featured series

LLMs, from first principles

A structured journey from numbers and neural networks to transformers, inference engines, GPUs, and production serving.

Coming soon

01 — How computers represent numbers

Before FP16, BF16, quantization, and tensor cores, we need to understand how computers represent numbers in the first place.

Part I

Foundations

Floating point, vectors, matrices, neural networks, embeddings.

Part II

Transformers & LLMs

Attention, tokenization, pretraining, fine-tuning, alignment.

Part III

LLM Inference

KV cache, batching, PagedAttention, quantization, parallelism.

Part IV

Production Systems

GPUs, Kubernetes, model serving, observability, cost and scale.

Research

Applied research in recommendation and LLM systems.

My interests sit at the intersection of machine learning systems, recommendation, retrieval, and large language models.

Research page →

Projects

Experiments, systems, and open-source work.

This section will collect hands-on work around inference, retrieval, distributed systems, and model serving.

Projects page →

About

Machine learning engineer based in the San Francisco Bay Area.

I work on production machine learning systems and write to understand difficult ideas from first principles. This site is where I document what I learn and build.