← Dinesh

Articles by Dinesh

4 articles.

How a Transformer Really Works: Attention, the KV Cache, and Why Inference Eats Memory

Architecture

A from-scratch tour of what's actually inside an LLM: how a transformer turns tokens into predictions, what Query, Key, and Value really mean, and how generating text one token at a time builds the KV cache — the growing pool of memory that makes inference so expensive.

Jul 21, 202610 min

Every Mask in a Transformer, Untangled

Training Systems

The word "mask" means at least four unrelated things in deep learning — what a token can see, what counts toward the loss, what is hidden to create a task, and what is randomly dropped. One field guide to all of them, with why each exists and what breaks without it.

Jul 20, 20269 min

Intuitive Guide to LoRA: Fine-Tuning a Model by 0.2% of weights

Training Systems

You don't need a massive tech budget or a cluster of high-end GPUs to train your own AI. LoRA allows developers to fine-tune giant models right on a standard laptop. Here is the zero-jargon, first-principles explanation of the clever shortcut that leveled the playing field.

Jul 19, 202612 min

Neural Networks From Zero: From a Single Number to a Billion Parameters

Architecture

A neural network never sees a word, an image, or a sound — only a list of numbers. Starting from that one fact and a single neuron, this guide builds the whole machine: how any input becomes numbers, why weights, biases, and activations each exist, and how neurons stack into layers and layers into a model.

Jul 12, 202614 min