Skip to content

Repository files navigation

Inference Engineering

How large language models actually run on real hardware, built up from first principles.

Most material on LLM inference is a list of techniques. This is a curriculum organised around one question: which resource is the bottleneck, why, and which technique attacks it? Every idea is derived, tied to the GPU it runs on, and backed by code you can run on a laptop.

What happens between the moment you send a prompt and the moment the model produces its next token?

Read it online: akbarsheikh-debug.github.io/inference-engineering

Read it online: akbarsheikh-debug.github.io/inference-engineering

Start here

00. The roadmap The five-layer stack, the dependency graph, and reading paths
01. What happens between a prompt and the next token Tokens, prefill, decode, and where the engine fits
02. Why decode is slow on a machine that can do a petaflop Arithmetic intensity, the roofline, tokens-per-second ceilings
03. Why the KV cache exists, and why it fills your GPU Cache arithmetic, GQA and MLA, paging
04. Quantization: spending fewer bytes per number Number formats, symmetric vs asymmetric, granularity, the straight-through estimator
05. Continuous batching: never let a finished request hold a seat Head-of-line blocking, iteration-level scheduling, chunked prefill, the two knobs
06. The GPU memory hierarchy Registers to HBM, coalescing and 32-byte sectors, shared-memory tiling
07. FlashAttention: exact attention without the n-by-n matrix Online softmax, why it is exact, what it saves and what it does not
08. Speculative decoding: more tokens per read of the weights Accept/reject rule, the two-line exactness proof, expected speedup
%% title: Learning roadmap as a dependency graph
flowchart TD
    L0["<b>0 The problem</b><br/>tokens, TTFT, ITL, cost"]
    L1["<b>1 The machine</b><br/>CPU host, GPU, memory hierarchy"]
    L2["<b>2 The model</b><br/>transformer, attention, prefill vs decode"]
    L3["<b>3 Why it is slow</b><br/>roofline, arithmetic intensity"]
    L4["<b>4 Shrink it</b><br/>quantization, FlashAttention, fusion"]
    L5["<b>5 Manage state</b><br/>KV cache, paging, prefix caching"]
    L6["<b>6 Schedule</b><br/>continuous batching, chunked prefill, speculation"]
    L7["<b>7 Scale out</b><br/>TP, PP, EP, CP, disaggregation"]
    L8["<b>8 Engines</b><br/>vLLM, SGLang, TensorRT-LLM"]
    L9["<b>9 Production</b><br/>routing, observability, cost"]
    L0 --> L1
    L0 --> L2
    L1 --> L3
    L2 --> L3
    L2 --> L5
    L3 --> L4
    L3 --> L5
    L4 --> L6
    L5 --> L6
    L6 --> L7
    L6 --> L8
    L7 --> L8
    L8 --> L9
    classDef live fill:#d1fae5,stroke:#059669,color:#064e3b
    classDef planned fill:#f3f4f6,stroke:#9ca3af,color:#374151
    class L0,L1,L2,L3,L4,L5,L6 live
    class L7,L8,L9 planned
Loading

Green means at least one article is published; grey is planned. All diagrams are in the diagram gallery.

Run the labs

Everything under labs/cpu/ needs only Python and NumPy. No GPU.

git clone https://github.com/AkbarSheikh-debug/inference-engineering.git
cd inference-engineering
pip install -e ".[dev]"

python labs/cpu/00_tiny_decoder.py      # generate with and without a KV cache
python labs/cpu/01_roofline.py          # tokens-per-second ceilings for a model on a GPU
python labs/cpu/02_kv_calculator.py     # KV bytes per token, sequences that fit, MHA vs GQA vs MQA
python labs/cpu/03_paged_allocator.py   # contiguous slabs vs paged blocks
python labs/cpu/04_quantization.py      # symmetric vs asymmetric, per-tensor vs per-channel vs group
python labs/cpu/05_continuous_batching.py  # static vs continuous batching, token budgets
python labs/cpu/06_gpu_memory.py        # coalescing efficiency, tiled matrix-multiply traffic
python labs/cpu/07_flash_attention.py   # tiled attention is exact; traffic model
python labs/cpu/08_speculative_decoding.py  # accept/reject rule, exactness, expected speedup
pytest                                  # every number quoted in the articles is checked here

What makes the numbers trustworthy

  • Every number is generated by code in src/ie/ or cited to a primary source (papers.md). A test fails if an article quotes a figure the code does not produce.
  • There are no measured benchmarks yet. Everything quoted is arithmetic from published specifications and is labelled a ceiling, not a forecast. Measurements will live in labs/bench/ with the hardware, driver and library versions recorded.
  • Units are stated once and kept. Vendor capacity is binary (an "80 GB" GPU carries 80 GiB of HBM); bandwidth and FLOP/s are decimal. See src/ie/units.py.

Repository layout

curriculum/   the articles
diagrams/     Mermaid sources in src/, generated SVGs in export/, gallery in README.md
labs/cpu/     runnable labs, one per idea
src/ie/       the arithmetic: hardware, models, KV cache, roofline, quantization, batching, paging, GPU memory, attention, speculation
blog/         a long-form post (Markdown and Substack-ready HTML) with 30 figures
tests/        unit tests, plus checks that articles match the code
tools/        diagram sync, figure generators, documentation checks
papers.md     primary sources and the claims they support
GLOSSARY.md   one definition per term, used consistently

Contributing

Corrections are the most valuable contribution: a wrong number, an unsupported claim, an unclear derivation. See CONTRIBUTING.md.

License and attribution

Released under the MIT License. Copyright (c) 2026 Akbar Arif.

You may use, copy, modify, fork and redistribute everything here, including commercially. The one condition is that the copyright and license notice (the LICENSE file) stays with every copy or substantial portion. If this material helps your work, please cite it (GitHub's "Cite this repository" button reads CITATION.cff) and link back to the source.

About

How LLMs actually run on real hardware, from first principles: roofline, KV cache, quantization, batching. Runnable labs, tested numbers, MIT licensed.

Topics

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages