# Build a Small LLM From Scratch: A Tested GPT in PyTorch

You can train a real language model on a laptop CPU during a coffee break. It won't be smart. It will have every part a big one has.

This post builds one. It's a decoder-only transformer in PyTorch, the same family as GPT-2, with 813,440 parameters. It trains on Tiny Shakespeare and ends at a validation loss of 1.78.

All the code ran before it went into this post. The full project is in [tinygpt.zip](https://github.com/pradeep200892/tinygpt): model, training script, sampler, baselines, 25 tests, and the trained checkpoint (`out/best.pt`). Every number below comes from one run on a single CPU core, with Python 3.12.3 and PyTorch 2.14.0.

## What a language model actually does

A language model predicts the next token. That's the whole job.

Feed it `To be or not to b` and it returns a probability for every possible next character. Training pushes those probabilities toward what the text really says. The loss is cross-entropy: the negative log of the probability the model gave the correct character.

Loss also gives you a floor to compare against. A model that guesses uniformly across our 65 characters scores ln(65) = 4.174. An untrained network should land right there, and ours does.

Chat models work the same way, with subword tokens, billions of parameters, and more training stages on top. The core loop doesn't change.

We'll use characters instead of subword tokens. The vocabulary is 65 symbols, so there's no tokenizer to train.

The price is that the model spends capacity learning to spell. GPT-2 used byte pair encoding with a 50,257-token vocabulary instead.

## The architecture

![Mermaid Diagram](https://mermaid.ink/img/pako:eNpV0MFu2zAMBuBX4XjaUPnQrdshhwF1UjcFuh6WbBerB8ZmYiKyaEha26zIuw9WAWE9ih9_ieIrdtozLnDv9LkbKCTYrqwHALhuLS4HCtQlDiB9NBAHmhg-1ga2nyw-QlV9h7q1uNUje-Bxx30v_gAXMGmUJPpf0eLj2711ji1bi7XT7giXRZZZVkU-F1lluSnypchNlqbIVZEmy21rsRFPDu7pxOFBw1g6bnPHurV4L54pwMDUG0jCPSSFNDCk9z8r0XWO3s1RPUh6vxsD377O-0GDI4eRpMfFK6aBx3nTPe_pj0to3iq_KQjtHMe5Z68-NTSKO-ECK5omx1U8xcSjgdqJP_6gbpPPjfpkwOKGD8rw686igZ-606QG1uyeOElHBq6DkDMQyccqcpA9mvzIRv7Os1xeTS94PhvcHZbqNOACPzwPkhjP_wCu96XX?type=png)

Each block has two sublayers. Both are wrapped in a residual connection, and both normalize their input first.

![Mermaid Diagram](https://mermaid.ink/img/pako:eNp1kMFqwzAMhl9F02llziFhg5LDYCuUDdoxWrZL3IOSOK2pYwdbWduVvvtI0uUw2Mmy_g99ts5YuFJhipVxh2JHnmGxkhYA4JhJPErcQBQ9grFxJnFBJ-XfnK8lbgbI2LjPiTmTOKM2kIGgTBURs7KsnR1ZYh7YsoxvbyXeSZxMrq4xuKJlGV-9yT_epM9r02QSl4v3FOJkCuzgIU66I06mI1yb5nd-8kc8irpsbA2zXdv9KTeu2Hd107LEDQqsla9Jl5iekXeq7rZXqopawyiGzid5TblRoWMqZ3lOtTYnTDGipjEqCqfAqhbwbLTdL6lY9_e5syxA4lptnYKPV4kCVi537AS8KPOlWBck4MlrMgIC2RAF5XWFopes9Xf3lvi-OeLlIjDfzpxxHlO8Oew0K7z8AOCul2o?type=png)

Moving LayerNorm to the input of each sublayer is one of the changes the GPT-2 paper made to the original transformer. It's why this design is called pre-LN. The original Transformer paper, [Attention Is All You Need](https://arxiv.org/abs/1706.03762), applied it after.

## Setup

```bash
pip install torch numpy pytest
unzip tinygpt.zip
cd tinygpt
python get_data.py
```

`get_data.py` downloads Tiny Shakespeare from Andrej Karpathy's [char-rnn repo](https://github.com/karpathy/char-rnn). The file is 1,115,394 bytes.

On macOS with a Python from python.org, that download can fail with a `CERTIFICATE_VERIFY_FAILED` error. The script catches this and falls back to `curl`. The FAQ at the end explains the real fix.

Files in the project:

-   `data.py`: character tokenizer, train/validation split, batching
    
-   `model.py`: the GPT itself
    
-   `train.py`: training loop, learning-rate schedule, evaluation
    
-   `sample.py`: text generation from a checkpoint
    
-   `checkpoint.py`: save and load
    
-   `baselines.py`: unigram and bigram losses for comparison
    
-   `tests/test_tinygpt.py`: 25 tests
    

## Step 1: Data

_data.py_

```python
"""Character tokenizer and batching for tinygpt."""
from pathlib import Path

import torch

class CharTokenizer:
    """Maps every distinct character in the corpus to an integer id."""

    def __init__(self, chars):
        self.chars = list(chars)
        self.stoi = {c: i for i, c in enumerate(self.chars)}
        self.itos = {i: c for i, c in enumerate(self.chars)}

    @classmethod
    def from_text(cls, text):
        return cls(sorted(set(text)))

    def start_id(self):
        """Id that starts generation: newline if the corpus has one, else id 0."""
        return self.stoi.get("\n", 0)

    @property
    def vocab_size(self):
        return len(self.chars)

    def encode(self, s):
        return [self.stoi[c] for c in s]

    def decode(self, ids):
        return "".join(self.itos[int(i)] for i in ids)

def load_corpus(path, val_fraction=0.1):
    """Read a text file and return (tokenizer, train_ids, val_ids).

    The split is contiguous: the last `val_fraction` of the file is
    validation data. Shuffling characters would leak train text into val.
    """
    text = Path(path).read_text(encoding="utf-8")
    tok = CharTokenizer.from_text(text)
    ids = torch.tensor(tok.encode(text), dtype=torch.long)
    n_val = int(len(ids) * val_fraction)
    return tok, ids[:-n_val], ids[-n_val:]

def get_batch(data, batch_size, block_size, device="cpu", generator=None):
    """Sample `batch_size` random windows.

    x is a window of block_size tokens. y is the same window shifted one
    position to the right, so y[t] is the token that follows x[t].
    """
    hi = len(data) - block_size - 1
    if hi <= 0:
        raise ValueError(f"need at least {block_size + 2} tokens, got {len(data)}")
    starts = torch.randint(0, hi, (batch_size,), generator=generator)
    x = torch.stack([data[s : s + block_size] for s in starts])
    y = torch.stack([data[s + 1 : s + 1 + block_size] for s in starts])
    return x.to(device), y.to(device)
```

The tokenizer sorts the distinct characters in the file and numbers them. Sorting makes the mapping deterministic. `start_id` returns the newline's id if the corpus has one, and generation starts from it when you give no prompt.

```python
from data import load_corpus

tok, train_ids, val_ids = load_corpus("data/input.txt")
print(tok.vocab_size, len(train_ids), len(val_ids))
print(repr("".join(tok.chars)))
print(tok.encode("First"))
print(tok.decode(tok.encode("First")))
```

```text
65 1003855 111539
"\n !$&',-.3:;?ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz"
[18, 47, 56, 57, 58]
First
```

The vocabulary holds a newline, a space, a handful of punctuation marks, the digit `3`, and both alphabets. That's 65 symbols. The last 10% of the text becomes validation data.

The split is contiguous on purpose. A random split would scatter overlapping windows of the same passages across both sets, and validation loss would stop measuring anything.

`get_batch` samples random windows. The target `y` is the input `x` shifted one position right. It refuses data shorter than `block_size + 2` tokens.

```python
import torch
from data import load_corpus, get_batch

tok, train_ids, val_ids = load_corpus("data/input.txt")
torch.manual_seed(1337)
x, y = get_batch(train_ids, batch_size=1, block_size=8)
print(repr(tok.decode(x[0])))
print(repr(tok.decode(y[0])))
```

```text
"Let's he"
"et's hea"
```

Look at the two strings. From `L` the target is `e`, from `Le` it's `t`, from `Let` it's `'`. One window of 8 characters gives you 8 training examples at once.

## Step 2: The model

_model.py, imports and config_

```python
"""A small GPT: decoder-only transformer with pre-LayerNorm blocks."""
import math
from dataclasses import dataclass, asdict

import torch
import torch.nn as nn
import torch.nn.functional as F
```

```python
@dataclass
class GPTConfig:
    vocab_size: int = 65
    block_size: int = 128   # maximum context length
    n_layer: int = 4
    n_head: int = 4
    n_embd: int = 128
    dropout: float = 0.1

    def to_dict(self):
        return asdict(self)
```

The config is small on purpose: 4 layers, 4 heads, 128 dimensions, and a context of 128 characters. Each head works on 128 / 4 = 32 dimensions.

### Causal self-attention

_model.py, the attention layer_

```python
class CausalSelfAttention(nn.Module):
    def __init__(self, cfg):
        super().__init__()
        assert cfg.n_embd % cfg.n_head == 0, "n_embd must divide by n_head"
        self.n_head = cfg.n_head
        self.qkv = nn.Linear(cfg.n_embd, 3 * cfg.n_embd, bias=False)
        self.proj = nn.Linear(cfg.n_embd, cfg.n_embd, bias=False)
        self.attn_drop = nn.Dropout(cfg.dropout)
        self.resid_drop = nn.Dropout(cfg.dropout)
        # Lower-triangular matrix: position t may look at positions 0..t only.
        mask = torch.tril(torch.ones(cfg.block_size, cfg.block_size))
        self.register_buffer(
            "mask", mask.view(1, 1, cfg.block_size, cfg.block_size), persistent=False
        )

    def forward(self, x):
        B, T, C = x.shape
        hs = C // self.n_head  # head size
        q, k, v = self.qkv(x).split(C, dim=2)
        # (B, T, C) -> (B, n_head, T, hs)
        q = q.view(B, T, self.n_head, hs).transpose(1, 2)
        k = k.view(B, T, self.n_head, hs).transpose(1, 2)
        v = v.view(B, T, self.n_head, hs).transpose(1, 2)

        att = (q @ k.transpose(-2, -1)) * (hs ** -0.5)          # (B, nh, T, T)
        att = att.masked_fill(self.mask[:, :, :T, :T] == 0, float("-inf"))
        att = self.attn_drop(F.softmax(att, dim=-1))
        y = att @ v                                              # (B, nh, T, hs)
        y = y.transpose(1, 2).contiguous().view(B, T, C)         # merge heads
        return self.resid_drop(self.proj(y))
```

One linear layer produces queries, keys and values together, and `split` cuts them apart. The `view` and `transpose` calls reshape each from `(B, T, 128)` to `(B, 4, T, 32)`, so every head gets its own 32 dimensions.

The scores are `q @ k.T`, scaled by `hs ** -0.5`. The scaling comes from the scaled dot-product attention in the Transformer paper. The authors suspect that for large head sizes the dot products grow in magnitude, which pushes softmax into regions with extremely small gradients.

The mask is the part that makes this a language model. Position `t` may look at positions `0..t` and nothing later. Here's the mask for 4 positions, and what softmax does with it:

```python
import torch

mask = torch.tril(torch.ones(4, 4))
print(mask)

torch.manual_seed(0)
scores = torch.randn(4, 4).masked_fill(mask == 0, float("-inf"))
print(torch.softmax(scores, dim=-1).round(decimals=2))
```

```text
tensor([[1., 0., 0., 0.],
        [1., 1., 0., 0.],
        [1., 1., 1., 0.],
        [1., 1., 1., 1.]])
tensor([[1.0000, 0.0000, 0.0000, 0.0000],
        [0.5400, 0.4600, 0.0000, 0.0000],
        [0.4500, 0.0900, 0.4600, 0.0000],
        [0.1300, 0.4100, 0.3600, 0.0900]])
```

Masked scores become `-inf` before softmax, so they turn into exact zeros. The second row spreads its weight across positions 0 and 1 only. Before rounding, every row sums to 1.

PyTorch ships a fused version of this in `scaled_dot_product_attention`, with an `is_causal` flag. We write it by hand so you can see every step. A test later in this post checks that our layer matches PyTorch's output.

### MLP, block, and the full model

_model.py, MLP and block_

```python
class MLP(nn.Module):
    def __init__(self, cfg):
        super().__init__()
        self.fc = nn.Linear(cfg.n_embd, 4 * cfg.n_embd, bias=False)
        self.proj = nn.Linear(4 * cfg.n_embd, cfg.n_embd, bias=False)
        self.drop = nn.Dropout(cfg.dropout)

    def forward(self, x):
        return self.drop(self.proj(F.gelu(self.fc(x))))
```

```python
class Block(nn.Module):
    def __init__(self, cfg):
        super().__init__()
        self.ln1 = nn.LayerNorm(cfg.n_embd)
        self.attn = CausalSelfAttention(cfg)
        self.ln2 = nn.LayerNorm(cfg.n_embd)
        self.mlp = MLP(cfg)

    def forward(self, x):
        x = x + self.attn(self.ln1(x))
        x = x + self.mlp(self.ln2(x))
        return x
```

The MLP expands each position from 128 to 512 dimensions, applies GELU, and projects back. Attention mixes information across positions. The MLP transforms each position on its own.

The block adds each sublayer's output to its input. Those residual paths let gradients flow straight through a deep stack.

_model.py, the GPT class_

```python
class GPT(nn.Module):
    def __init__(self, cfg):
        super().__init__()
        self.cfg = cfg
        self.wte = nn.Embedding(cfg.vocab_size, cfg.n_embd)   # token embeddings
        self.wpe = nn.Embedding(cfg.block_size, cfg.n_embd)   # position embeddings
        self.drop = nn.Dropout(cfg.dropout)
        self.blocks = nn.ModuleList([Block(cfg) for _ in range(cfg.n_layer)])
        self.ln_f = nn.LayerNorm(cfg.n_embd)
        self.lm_head = nn.Linear(cfg.n_embd, cfg.vocab_size, bias=False)
        self.lm_head.weight = self.wte.weight                 # weight tying

        self.apply(self._init_weights)
        # GPT-2 trick: shrink the layers that write into the residual stream,
        # so the stream's variance doesn't grow with depth.
        for name, p in self.named_parameters():
            if name.endswith("attn.proj.weight") or name.endswith("mlp.proj.weight"):
                nn.init.normal_(p, mean=0.0, std=0.02 / math.sqrt(2 * cfg.n_layer))

    @staticmethod
    def _init_weights(module):
        if isinstance(module, (nn.Linear, nn.Embedding)):
            nn.init.normal_(module.weight, mean=0.0, std=0.02)

    def forward(self, idx, targets=None):
        B, T = idx.shape
        assert T <= self.cfg.block_size, f"sequence length {T} > block_size {self.cfg.block_size}"
        pos = torch.arange(T, device=idx.device)
        x = self.drop(self.wte(idx) + self.wpe(pos))
        for block in self.blocks:
            x = block(x)
        logits = self.lm_head(self.ln_f(x))                   # (B, T, vocab)
        loss = None
        if targets is not None:
            loss = F.cross_entropy(logits.view(-1, logits.size(-1)), targets.view(-1))
        return logits, loss

    @torch.no_grad()
    def generate(self, idx, max_new_tokens, temperature=1.0, top_k=None, generator=None):
        assert temperature > 0, "use top_k=1 for greedy decoding"
        was_training = self.training
        self.eval()
        for _ in range(max_new_tokens):
            ctx = idx[:, -self.cfg.block_size :]              # crop to the window
            logits, _ = self(ctx)
            logits = logits[:, -1, :] / temperature           # last position only
            if top_k is not None:
                kth = torch.topk(logits, min(top_k, logits.size(-1))).values[:, [-1]]
                logits = logits.masked_fill(logits < kth, float("-inf"))
            probs = F.softmax(logits, dim=-1)
            nxt = torch.multinomial(probs, num_samples=1, generator=generator)
            idx = torch.cat([idx, nxt], dim=1)
        self.train(was_training)
        return idx

    def num_params(self):
        # parameters() counts the tied embedding/head matrix once
        return sum(p.numel() for p in self.parameters())
```

Three details in this class matter more than they look.

First, the token embedding and the output head share one weight matrix. The input side turns an id into a vector, and the output side turns a vector back into scores over ids. Tying them saves 65 × 128 = 8,320 parameters, and the original Transformer paper shares its embedding and pre-softmax matrices the same way.

Second, initialization: every weight starts from a normal distribution with standard deviation 0.02. The two projections that write into the residual stream, `attn.proj` and `mlp.proj`, get that value divided by `sqrt(2 * n_layer)`.

The GPT-2 paper describes scaling residual layers by 1/√N, where N is the number of residual layers. Each block contributes two, hence the `2 *`. [nanoGPT](https://github.com/karpathy/nanoGPT) implements it the same way.

Third, `generate` crops the context to the last `block_size` tokens. The position embedding table only has 128 rows.

Let's count parameters and check the untrained loss.

```python
import math
import torch
from data import load_corpus, get_batch
from model import GPT, GPTConfig

tok, train_ids, val_ids = load_corpus("data/input.txt")
torch.manual_seed(1337)
model = GPT(GPTConfig(vocab_size=tok.vocab_size))
print(f"{model.num_params():,} parameters")
print("token embedding :", model.wte.weight.numel())
print("position embed  :", model.wpe.weight.numel())
print("one block       :", sum(p.numel() for p in model.blocks[0].parameters()))
print("final LayerNorm :", sum(p.numel() for p in model.ln_f.parameters()))

x, y = get_batch(train_ids, 32, 128)
logits, loss = model(x, y)
print(logits.shape)
print(f"loss {loss.item():.4f}   ln(65) = {math.log(65):.4f}")
```

```text
813,440 parameters
token embedding : 8320
position embed  : 16384
one block       : 197120
final LayerNorm : 256
torch.Size([32, 128, 65])
loss 4.1827   ln(65) = 4.1744
```

The numbers line up with the architecture:

| Part | Parameters |
| --- | --- |
| Token embedding (65 × 128) | 8,320 |
| Position embedding (128 × 128) | 16,384 |
| Each block (4 of them) | 197,120 |
| Final LayerNorm | 256 |
| Total | 813,440 |

A block holds 65,536 attention weights (4 × 128²), 131,072 MLP weights (8 × 128²), and 512 LayerNorm parameters. The untrained loss of 4.1827 sits right on ln(65) = 4.1744.

## Step 3: Test it before you train it

Training a broken transformer wastes an afternoon and doesn't announce itself. The loss still goes down. Three checks catch most bugs first.

The first is the one you just saw: initial loss near ln(vocab size). The second is a leak test. Change a token at position 10 and confirm that nothing before position 10 moves.

_tests/test\_tinygpt.py_

```python
def test_model_cannot_see_the_future():
    torch.manual_seed(0)
    model = GPT(small_cfg()).eval()
    a = torch.randint(0, 20, (1, 16))
    b = a.clone()
    b[0, 10] = (b[0, 10] + 1) % 20          # change the token at position 10
    la, _ = model(a)
    lb, _ = model(b)
    assert torch.allclose(la[:, :10], lb[:, :10], atol=1e-6)   # earlier positions unchanged
    assert not torch.allclose(la[:, 10:], lb[:, 10:])          # position 10 onward changes
```

A model that can see the answer gets near-zero training loss and generates junk. This test is what tells you the mask works.

The third check is overfitting. A healthy model can memorize one small batch.

```python
def test_can_overfit_one_batch():
    torch.manual_seed(0)
    model = GPT(small_cfg())
    opt = torch.optim.AdamW(model.parameters(), lr=3e-3)
    x = torch.randint(0, 20, (4, 16))
    y = torch.randint(0, 20, (4, 16))
    first = model(x, y)[1].item()
    for _ in range(300):
        _, loss = model(x, y)
        opt.zero_grad()
        loss.backward()
        opt.step()
    assert first > 2.5
    assert loss.item() < 0.1
```

If the loss can't reach zero on 4 sequences, something is wrong in the forward pass or the optimizer.

## Step 4: Train

_train.py, schedule, optimizer and evaluation_

```python
def get_lr(step, max_steps, lr, min_lr, warmup):
    """Linear warmup, then cosine decay from lr down to min_lr."""
    if step < warmup:
        return lr * (step + 1) / warmup
    if step >= max_steps:
        return min_lr
    progress = (step - warmup) / (max_steps - warmup)
    return min_lr + 0.5 * (1 + math.cos(math.pi * progress)) * (lr - min_lr)
```

```python
def make_optimizer(model, lr, weight_decay, betas=(0.9, 0.99)):
    """AdamW that applies weight decay to matrices only, not to biases or norms."""
    decay = [p for p in model.parameters() if p.dim() >= 2]
    no_decay = [p for p in model.parameters() if p.dim() < 2]
    groups = [
        {"params": decay, "weight_decay": weight_decay},
        {"params": no_decay, "weight_decay": 0.0},
    ]
    return torch.optim.AdamW(groups, lr=lr, betas=betas)
```

```python
@torch.no_grad()
def estimate_loss(model, splits, batch_size, block_size, eval_iters, device):
    model.eval()
    out = {}
    for name, data in splits.items():
        losses = torch.zeros(eval_iters)
        for i in range(eval_iters):
            x, y = get_batch(data, batch_size, block_size, device)
            losses[i] = model(x, y)[1].item()
        out[name] = losses.mean().item()
    model.train()
    return out
```

The learning rate warms up over 100 steps to 1e-3, then follows a cosine curve down to 1e-4. Warmup keeps the first updates small while Adam's statistics settle.

AdamW uses betas of (0.9, 0.99). The second value matches nanoGPT's Shakespeare config, which raises it because each step sees few tokens.

Weight decay of 0.1 applies to weight matrices only. Biases and LayerNorm parameters are excluded, and a test checks the split.

`estimate_loss` averages over several batches and switches to eval mode, which turns dropout off. It switches back afterward.

_train.py, the training loop (excerpt from main)_

```python
torch.manual_seed(args.seed)
device = pick_device(args.device)
out = Path(args.out)
out.mkdir(parents=True, exist_ok=True)

tok, train_ids, val_ids = load_corpus(args.data)
cfg = GPTConfig(
    vocab_size=tok.vocab_size, block_size=args.block_size, n_layer=args.n_layer,
    n_head=args.n_head, n_embd=args.n_embd, dropout=args.dropout,
)
model = GPT(cfg).to(device)
opt = make_optimizer(model, args.lr, args.weight_decay)
splits = {"train": train_ids, "val": val_ids}
print(f"device={device} params={model.num_params():,} vocab={tok.vocab_size} "
      f"train_tokens={len(train_ids):,} val_tokens={len(val_ids):,}")

sample_gen = torch.Generator(device=device).manual_seed(0)
prompt = torch.tensor([[tok.start_id()]], device=device)
history, samples, best_val = [], [], float("inf")
t0 = time.time()

for step in range(args.max_steps + 1):
    if step % args.eval_interval == 0 or step == args.max_steps:
        losses = estimate_loss(model, splits, args.batch_size, args.block_size,
                               args.eval_iters, device)
        text = tok.decode(model.generate(prompt, 200, generator=sample_gen)[0].tolist())
        history.append({"step": step, **losses, "seconds": round(time.time() - t0)})
        samples.append({"step": step, "text": text})
        print(f"step {step:5d} | train {losses['train']:.4f} | val {losses['val']:.4f} "
              f"| {time.time() - t0:6.0f}s", flush=True)
        if losses["val"] < best_val:
            best_val = losses["val"]
            save_checkpoint(out / "best.pt", model, tok, step, best_val)
        (out / "history.json").write_text(json.dumps(history, indent=2))
        (out / "samples.json").write_text(json.dumps(samples, indent=2))
    if step == args.max_steps:
        break

    lr = get_lr(step, args.max_steps, args.lr, args.min_lr, args.warmup)
    for g in opt.param_groups:
        g["lr"] = lr
    x, y = get_batch(train_ids, args.batch_size, args.block_size, device)
    _, loss = model(x, y)
    opt.zero_grad(set_to_none=True)
    loss.backward()
    torch.nn.utils.clip_grad_norm_(model.parameters(), args.grad_clip)
    opt.step()
```

Each step draws a batch, computes the loss, backpropagates, clips the gradient norm to 1.0, and updates. Every 250 steps it evaluates, generates a 200-character sample, and saves `best.pt` if validation loss improved.

A step processes 32 × 128 = 4,096 characters. Two thousand steps is 8,192,000 characters, or about 8.2 passes over the 1,003,855-character training split.

Run it:

```bash
python train.py --data data/input.txt --out out
```

```text
device=cpu params=813,440 vocab=65 train_tokens=1,003,855 val_tokens=111,539
step     0 | train 4.1835 | val 4.1799 |     13s
step   250 | train 2.3962 | val 2.4186 |    160s
step   500 | train 2.1546 | val 2.1950 |    307s
step   750 | train 1.9662 | val 2.0432 |    450s
step  1000 | train 1.8321 | val 1.9598 |    595s
step  1250 | train 1.7527 | val 1.8790 |    740s
step  1500 | train 1.6836 | val 1.8391 |    891s
step  1750 | train 1.6436 | val 1.8080 |   1044s
step  2000 | train 1.6215 | val 1.7838 |   1188s
best val loss 1.7838 (2.573 bits/char)
```

That took 1,188 seconds on one CPU core. The best validation loss was 1.7838, which is 2.573 bits per character.

The samples that `train.py` writes to `out/samples.json` show the model learning in stages. These are the first 110 characters at four checkpoints:

```text
step 0
t3ZxjmhPPdkARxLsXltUWHoHca ?eghkyy&&NdkIgAlsvSBUMyccRD$BygXOfntw
kVrM h&?j SJR!sOEHelC!BTg3xUhPRCs3Ma--3KypMrW

step 250
Inol3 f pifie-mily mpoennd shay tinte.

The beome igle inkee he them.

KES:
S y tonoong sof themyowangh.
HOn d

step 1000
DORY:
Not fieve my that that
Romest thear not ut .
LUCED:
What the gelivl the. Shou jroses.
Find With you lets

step 2000
Eing Noices you, for villing a by like whose mork hard,
Shall appy'd mired: thou almp the place!

CORIOLANUS:
```

At step 0 it's noise drawn from all 65 symbols. By step 250 it has word-length chunks, sentence-ending periods, and even a stray speaker tag (`KES:`).

By step 1000 it knows that a name in capitals followed by a colon starts a speech. By step 2000 the format is solid, and many of the words are still not real ones.

The training loss ends at 1.6215 and validation at 1.7838. That gap of 0.16 is small, and validation was still falling at the last checkpoint, so overfitting isn't what limits this run.

The learning rate also decays to 1e-4 by step 2,000. That makes the flattening curve a poor sign that the model has run out of things to learn.

Here's a tighter validation estimate on 200 batches instead of 40:

```python
import math
import torch
from checkpoint import load_checkpoint
from data import load_corpus
from train import estimate_loss

model, tok, meta = load_checkpoint("out/best.pt")
_, train_ids, val_ids = load_corpus("data/input.txt")
torch.manual_seed(123)
val = estimate_loss(model, {"val": val_ids}, 32, 128, 200, "cpu")["val"]
print(f"val loss {val:.4f}  perplexity {math.exp(val):.2f}  bits/char {val / math.log(2):.3f}")
```

```text
val loss 1.7899  perplexity 5.99  bits/char 2.582
```

Perplexity of 5.99 means the model is about as uncertain as a fair choice between six characters at each step. The 1.7899 differs from the 1.7838 in the log only because it samples different batches.

## Step 5: Is 1.78 any good?

A loss means nothing without a baseline. Two are easy to compute on the same validation split.

The unigram model predicts from character frequencies alone. The bigram model predicts from the previous character alone.

_baselines.py_

```python
def unigram_loss(train_ids, val_ids, vocab):
    counts = torch.bincount(train_ids, minlength=vocab).float() + 1   # add-one smoothing
    probs = counts / counts.sum()
    return -probs[val_ids].log().mean().item()
```

```python
def bigram_loss(train_ids, val_ids, vocab):
    pairs = train_ids[:-1] * vocab + train_ids[1:]
    counts = torch.bincount(pairs, minlength=vocab * vocab).view(vocab, vocab).float() + 1
    probs = counts / counts.sum(dim=1, keepdim=True)                  # P(next | previous)
    return -probs[val_ids[:-1], val_ids[1:]].log().mean().item()
```

```bash
python baselines.py
```

```text
uniform guess : 4.1744
unigram       : 3.3473
bigram        : 2.4819
```

The bigram table, just 65 × 65 counts, reaches 2.48. The transformer's 1.78 comes from using more than the previous character. That gap is what attention over 128 characters of context buys you.

## Step 6: Generate text

_model.py, the sampling loop inside generate_

```python
ctx = idx[:, -self.cfg.block_size :]              # crop to the window
logits, _ = self(ctx)
logits = logits[:, -1, :] / temperature           # last position only
if top_k is not None:
    kth = torch.topk(logits, min(top_k, logits.size(-1))).values[:, [-1]]
    logits = logits.masked_fill(logits < kth, float("-inf"))
probs = F.softmax(logits, dim=-1)
nxt = torch.multinomial(probs, num_samples=1, generator=generator)
idx = torch.cat([idx, nxt], dim=1)
```

The model outputs scores for every position, and we keep the last one. Dividing by temperature reshapes the distribution. Below 1 it sharpens, above 1 it flattens.

Top-k keeps only the `k` highest scores and masks the rest to `-inf`. Then softmax turns scores into probabilities and `multinomial` draws one character. The code asserts that temperature is positive, so greedy decoding uses `top_k=1` instead.

_sample.py_

```python
"""Generate text from a trained checkpoint.

    python sample.py --ckpt out/best.pt --prompt "ROMEO:" --tokens 400
"""
import argparse

import torch

from checkpoint import load_checkpoint

def main():
    ap = argparse.ArgumentParser()
    ap.add_argument("--ckpt", default="out/best.pt")
    ap.add_argument("--prompt", default=None)
    ap.add_argument("--tokens", type=int, default=400)
    ap.add_argument("--temperature", type=float, default=0.8)
    ap.add_argument("--top-k", type=int, default=40)
    ap.add_argument("--seed", type=int, default=42)
    ap.add_argument("--device", default="cpu")
    args = ap.parse_args()

    model, tok, _ = load_checkpoint(args.ckpt, args.device)
    gen = torch.Generator(device=args.device).manual_seed(args.seed)
    if not args.prompt:
        ids = [tok.start_id()]
    else:
        unknown = sorted({c for c in args.prompt if c not in tok.stoi})
        if unknown:
            raise SystemExit(f"prompt has characters the model has never seen: {unknown}")
        ids = tok.encode(args.prompt)
    idx = torch.tensor([ids], device=args.device)
    out = model.generate(idx, args.tokens, args.temperature, args.top_k, generator=gen)
    print(tok.decode(out[0].tolist()))

if __name__ == "__main__":
    main()
```

Temperature 0.8 with top-k 40 is the default:

```bash
python sample.py --prompt "ROMEO:" --tokens 200
```

```text
ROMEO:
As lord shall, you for the read, one and bed forther
Bear head laist to dreath be that God,
Why so, upon this up enemer fament.

FLAUDE:
I lord, the the fall for not presence forfe,
The cuntracia, bu
```

Temperature 1.2 with no top-k filter (`--top-k 65` keeps all 65 characters):

```bash
python sample.py --prompt "ROMEO:" --tokens 200 --temperature 1.2 --top-k 65
```

```text
ROMEO:
As loud shadford: whose chard's, Statenty no fe thy eread?

First MAUXEqd:
Sorrow that's planly so,
Two to it'u
Let it them could to hinde lilenct, a who and for neWrp
Suanntifhed but nighnie his, bu
```

And greedy decoding, which always takes the most likely character:

```bash
python sample.py --prompt "ROMEO:" --tokens 200 --top-k 1
```

```text
ROMEO:
The shall the soul the soul the soul the soul the son,
And the shall the some the so the soul the stand
The shall the so the so the so the so the soul
The shall the shall the shall the shall the stay
```

The default sample has the shape of a play: speaker names in capitals, colons, short lines, and a few real words like "lord", "head" and "God". It has no grammar and no meaning. At 1.2 it invents words, and greedy decoding falls into a loop.

Greedy decoding always picks the single most likely character, so once a phrase repeats, the repeated context makes the next repeat even more likely. Sampling adds randomness that can break the cycle.

## What this model can and can't do

It learned the surface of Shakespeare: layout, capitalization, common letter patterns. It didn't learn to write sentences.

That's expected. The model has 813,440 parameters and read about a million characters, eight times.

I ran one training run with one seed and didn't tune any hyperparameters. Everything ran on CPU only.

## How to scale it up

The [nanoGPT README](https://github.com/karpathy/nanoGPT) gives useful reference points on the same dataset. Its CPU example (4 layers, 4 heads, 128 dimensions, context 64) takes about 3 minutes and reaches a loss of 1.88.

Its GPU config (6 layers, 6 heads, 384 dimensions, context 256) reaches 1.4697 in about 3 minutes on one A100. The settings differ from ours, so treat those as rough comparisons.

The same README lists GPT-2 at 124M parameters. That's about 150 times larger than tinygpt. It says `train.py` reproduces that model on OpenWebText in about 4 days on one node with 8 A100 40GB GPUs.

To move toward that, change things in this order:

1.  More data. A 1 MB corpus caps everything else.
    
2.  A subword tokenizer. Byte pair encoding shortens sequences, so each token carries more meaning.
    
3.  A bigger model. Raise `n_layer`, `n_head` and `n_embd`, and lengthen `block_size`.
    
4.  A GPU. `train.py` picks CUDA or MPS automatically. I haven't run those paths.
    

Don't do step 3 alone. Hoffmann et al. trained over 400 models and found that model size and training tokens should scale together: double the parameters, double the tokens.

Their 70B model, Chinchilla, beat the 280B Gopher on a range of tasks using the same compute and 4× more data. The paper is [Training Compute-Optimal Large Language Models](https://arxiv.org/abs/2203.15556).

## Tests

```bash
python -m pytest -q
```

```text
.........................                                                [100%]
25 passed in 12.23s
```

The suite covers the tokenizer round trip, batch shifting, the parameter-count formula, tied weights, the causal mask, agreement with PyTorch's attention, overfitting a single batch, reproducible sampling, greedy decoding, the learning-rate schedule, the weight-decay split, checkpoint loading, a two-step training run on a tiny corpus, the sampler's rejection of characters it has never seen, and the download script's fallback when Python can't verify certificates.

I also broke the code on purpose. Deleting the `masked_fill` line from the attention layer makes exactly two tests fail: the leak test and the PyTorch-equivalence test. The other 23 still pass, which shows those two guard that line.

## Gotchas I hit while building it

Split contiguously. Sample validation windows from the tail of the file, not from a shuffle of the whole text.

Decay matrices only. Weight decay on biases and LayerNorm gains does nothing useful. The `dim() >= 2` split handles it.

Crop the context. `generate` must slice to `block_size`, or long generations crash on the position table. The test asks for 40 new tokens with a block size of 16.

Seed your sampler. A `torch.Generator` passed to `generate` makes samples reproducible, and one test checks exactly that.

Don't hard-code a start token. My first version seeded sampling with a newline and crashed on any corpus that had none, which the tiny-corpus test now covers.

Save the vocabulary. The checkpoint stores the character list next to the weights. Without it, ids can't be turned back into text.

## Download and run

The project is [tinygpt.zip](https://github.com/pradeep200892/tinygpt). It contains every file listed above, the trained checkpoint in `out/`, and the training log.

```bash
unzip tinygpt.zip
cd tinygpt
pip install -r requirements.txt
python get_data.py
python sample.py --prompt "ROMEO:"
python -m pytest
```

Retraining with `python train.py` writes to `out/` and overwrites the shipped checkpoint. Copy `out/` somewhere first if you want to keep it.

## FAQ

**Is this really an LLM?** It's a language model built the same way as one. The "large" is about scale, and this is roughly 150 times smaller than the 124M GPT-2 model.

**Can I train it on my own text?** Yes. Run `python train.py --data yourfile.txt`. The vocabulary comes from the file.

With the default 128-character context the file needs at least 1,300 characters. Shorter files stop with a `ValueError` that says how many tokens are missing.

**Why does sampling fail on some prompts?** The tokenizer only knows characters from the training text. `sample.py` stops and lists the ones it doesn't know:

```bash
python sample.py --prompt "ROMEO: ☃"
```

```text
prompt has characters the model has never seen: ['☃']
```

Calling `encode` yourself raises a `KeyError` instead:

```python
from data import CharTokenizer

tok = CharTokenizer.from_text("hello")
try:
    tok.encode("hello!")
except KeyError as e:
    print("KeyError:", e)
```

```text
KeyError: '!'
```

**Why does** `get_data.py` **fail with** `CERTIFICATE_VERIFY_FAILED`**?** Python installed from python.org on macOS often ships without root certificates, so it can't verify any HTTPS server. `get_data.py` detects this and downloads with `curl` instead, which keeps verification on.

To fix Python itself, run `Install Certificates.command` from the `/Applications/Python 3.x/` folder. Don't switch verification off to make the error go away.

**Why does** `pip install torch` **say "No matching distribution"?** Usually PyTorch doesn't publish a wheel for your Python version and Mac chip. I checked Python 3.14 with `pip install --dry-run`: Apple Silicon has wheels, and the Intel Mac tag I tried (`macosx_10_15_x86_64`) has none.

Run `uname -m` to see your chip (`arm64` means Apple Silicon). I tested everything on Python 3.12.3.

**What is the "Failed to initialize NumPy" warning?** PyTorch prints it when NumPy isn't installed. Nothing in tinygpt uses NumPy, and the sampler, the training script and the test suite ran fine without it in my check.

`pip install numpy` silences the message, and the install commands above include it.

**Does it run on a GPU?** The code picks CUDA or MPS if PyTorch finds one. I only tested it on CPU.

**Why is the output nonsense?** The model is small and the data is tiny. Better output comes from more data, a bigger model, and longer training, in that order.

## References

-   Vaswani et al., [Attention Is All You Need](https://arxiv.org/abs/1706.03762), 2017.
    
-   Radford et al., _Language Models are Unsupervised Multitask Learners_ (GPT-2), 2019. Source for pre-LN placement and residual-scaled initialization.
    
-   Hoffmann et al., [Training Compute-Optimal Large Language Models](https://arxiv.org/abs/2203.15556), 2022.
    
-   Andrej Karpathy, [nanoGPT](https://github.com/karpathy/nanoGPT) and [char-rnn](https://github.com/karpathy/char-rnn).
    
-   PyTorch documentation, `scaled_dot_product_attention`.

---

*Published via [ZyVOP](https://zyvop.com/build-a-small-llm-from-scratch-a-tested-gpt-in-pytorch-62ccl?utm_source=hashnode&utm_medium=crosspost&utm_campaign=syndication) — Write once in Markdown, auto-backup to GitHub, and syndicate to Dev.to, Medium & Hashnode in 1 click.*
