How large language models actually run on real hardware, built up from first principles.
Most material on LLM inference is a list of techniques. This is a curriculum organised around one question: which resource is the bottleneck, why, and which technique attacks it? Every idea is derived, tied to the GPU it runs on, and backed by code you can run on a laptop.
What happens between the moment you send a prompt and the moment the model produces its next token?
Read it online: akbarsheikh-debug.github.io/inference-engineering
Read it online: akbarsheikh-debug.github.io/inference-engineering
| 00. The roadmap | The five-layer stack, the dependency graph, and reading paths |
| 01. What happens between a prompt and the next token | Tokens, prefill, decode, and where the engine fits |
| 02. Why decode is slow on a machine that can do a petaflop | Arithmetic intensity, the roofline, tokens-per-second ceilings |
| 03. Why the KV cache exists, and why it fills your GPU | Cache arithmetic, GQA and MLA, paging |
| 04. Quantization: spending fewer bytes per number | Number formats, symmetric vs asymmetric, granularity, the straight-through estimator |
| 05. Continuous batching: never let a finished request hold a seat | Head-of-line blocking, iteration-level scheduling, chunked prefill, the two knobs |
| 06. The GPU memory hierarchy | Registers to HBM, coalescing and 32-byte sectors, shared-memory tiling |
| 07. FlashAttention: exact attention without the n-by-n matrix | Online softmax, why it is exact, what it saves and what it does not |
| 08. Speculative decoding: more tokens per read of the weights | Accept/reject rule, the two-line exactness proof, expected speedup |
%% title: Learning roadmap as a dependency graph
flowchart TD
L0["<b>0 The problem</b><br/>tokens, TTFT, ITL, cost"]
L1["<b>1 The machine</b><br/>CPU host, GPU, memory hierarchy"]
L2["<b>2 The model</b><br/>transformer, attention, prefill vs decode"]
L3["<b>3 Why it is slow</b><br/>roofline, arithmetic intensity"]
L4["<b>4 Shrink it</b><br/>quantization, FlashAttention, fusion"]
L5["<b>5 Manage state</b><br/>KV cache, paging, prefix caching"]
L6["<b>6 Schedule</b><br/>continuous batching, chunked prefill, speculation"]
L7["<b>7 Scale out</b><br/>TP, PP, EP, CP, disaggregation"]
L8["<b>8 Engines</b><br/>vLLM, SGLang, TensorRT-LLM"]
L9["<b>9 Production</b><br/>routing, observability, cost"]
L0 --> L1
L0 --> L2
L1 --> L3
L2 --> L3
L2 --> L5
L3 --> L4
L3 --> L5
L4 --> L6
L5 --> L6
L6 --> L7
L6 --> L8
L7 --> L8
L8 --> L9
classDef live fill:#d1fae5,stroke:#059669,color:#064e3b
classDef planned fill:#f3f4f6,stroke:#9ca3af,color:#374151
class L0,L1,L2,L3,L4,L5,L6 live
class L7,L8,L9 planned
Green means at least one article is published; grey is planned. All diagrams are in the diagram gallery.
Everything under labs/cpu/ needs only Python and NumPy. No GPU.
git clone https://github.com/AkbarSheikh-debug/inference-engineering.git
cd inference-engineering
pip install -e ".[dev]"
python labs/cpu/00_tiny_decoder.py # generate with and without a KV cache
python labs/cpu/01_roofline.py # tokens-per-second ceilings for a model on a GPU
python labs/cpu/02_kv_calculator.py # KV bytes per token, sequences that fit, MHA vs GQA vs MQA
python labs/cpu/03_paged_allocator.py # contiguous slabs vs paged blocks
python labs/cpu/04_quantization.py # symmetric vs asymmetric, per-tensor vs per-channel vs group
python labs/cpu/05_continuous_batching.py # static vs continuous batching, token budgets
python labs/cpu/06_gpu_memory.py # coalescing efficiency, tiled matrix-multiply traffic
python labs/cpu/07_flash_attention.py # tiled attention is exact; traffic model
python labs/cpu/08_speculative_decoding.py # accept/reject rule, exactness, expected speedup
pytest # every number quoted in the articles is checked here- Every number is generated by code in
src/ie/or cited to a primary source (papers.md). A test fails if an article quotes a figure the code does not produce. - There are no measured benchmarks yet. Everything quoted is arithmetic from published specifications and is labelled a ceiling, not a forecast. Measurements will live in
labs/bench/with the hardware, driver and library versions recorded. - Units are stated once and kept. Vendor capacity is binary (an "80 GB" GPU carries 80 GiB of HBM); bandwidth and FLOP/s are decimal. See
src/ie/units.py.
curriculum/ the articles
diagrams/ Mermaid sources in src/, generated SVGs in export/, gallery in README.md
labs/cpu/ runnable labs, one per idea
src/ie/ the arithmetic: hardware, models, KV cache, roofline, quantization, batching, paging, GPU memory, attention, speculation
blog/ a long-form post (Markdown and Substack-ready HTML) with 30 figures
tests/ unit tests, plus checks that articles match the code
tools/ diagram sync, figure generators, documentation checks
papers.md primary sources and the claims they support
GLOSSARY.md one definition per term, used consistently
Corrections are the most valuable contribution: a wrong number, an unsupported claim, an unclear derivation. See CONTRIBUTING.md.
Released under the MIT License. Copyright (c) 2026 Akbar Arif.
You may use, copy, modify, fork and redistribute everything here, including commercially. The one condition is that the copyright and license notice (the LICENSE file) stays with every copy or substantial portion. If this material helps your work, please cite it (GitHub's "Cite this repository" button reads CITATION.cff) and link back to the source.