Lednik is a full-cycle framework for distilling large transformer encoders into small, fast embedding models — and for proving, with numbers, how much cheaper they are to run.
Large embedding models like deepvk/USER-bge-m3
(359M parameters) deliver great quality but are expensive to serve: a single mid-range GPU
saturates at a few hundred requests per second, and long inputs make the quadratic attention
cost explode. Lednik attacks this from the model side: it initializes a compact student from
the teacher's own embedding space, distills the teacher's knowledge into it, and ships the
result with a serving stack and a benchmark suite that measures quality (RuMTEB), raw
forward-pass speed, and end-to-end throughput under load.
Three student tiers cover the quality/cost trade-off:
- Static Embeddings — a weighted token lookup table. No attention, no context, near-zero compute; runs comfortably on CPU.
- Lednik Transformer (full attention) — a tiny encoder (RoPE, RMSNorm, Liger SwiGLU/GeGLU, gated attention) with a fully unpadded varlen fast path (Flash-Attention 2 / torch varlen SDPA).
- Lednik Transformer (hybrid) — the same encoder with part of the layers replaced by bidirectional Gated DeltaNet (linear attention on flash-linear-attention kernels): linear instead of quadratic scaling with sequence length.
Full guides live in docs/:
- Model Initialization — create a student from a teacher: the architectures, factory functions, configs, save/reload and inference.
- Training without ClearML — build the config,
instantiate the training module, the collator data format, and a minimal
Trainerloop. - Training with ClearML — the
pipelines/distill/pipeline: checkpoint loading, config/dataset wiring, checkpoint uploads, remote queues, the online validation worker, and MTEB benchmarking. - Usage — using trained models: checkpoint loading with
AutoLednikModel, the LitServe server, the request protocol, scaling knobs and Docker deployment. - JMLC — the repository mapped onto the Junior ML Contest evaluation criteria: engineering, data science, AI tooling, product.
Quality on the Russian subset of MTEB (RuMTEB) versus pure forward-pass speed:
| Model | Params | RuMTEB AvgScore | Forward median | Tokens/sec |
|---|---|---|---|---|
| USER-bge-m3 (teacher, sdpa) | 359.0M | 0.601 | 568.6 ms | 19.1k |
| Lednik Transformer — hybrid GDN (varlen) | 59.0M | 0.497 | 42.3 ms | 252.8k |
| Lednik Transformer — full attention (varlen) | 56.0M | 0.490 | 31.0 ms | 349.4k |
| Static Embeddings | 17.8M | 0.421 | 0.32 ms | 32.3M |
Setup: bfloat16 on a single RTX 3080; batches of 8 sequences of 128–4096 tokens
(~10.9k real tokens per batch); medians reported by triton.testing.do_bench.
The transformer students keep ~82% of the teacher's RuMTEB score at 6× fewer parameters and 13–18× faster forward.
Raw records live in bench/mteb_testing/results/ and
bench/forward_testing/results/.
- Initialization factory — seed a student from a teacher: sweep the vocabulary through the teacher, pool per-token embeddings, PCA them down to the student width, optionally weight by Smooth Inverse Frequency (SIF).
- Distillation — contrastive + regression objective on teacher/student sentence embeddings; either a pure-Python Lightning loop or a full ClearML pipeline with artifact loading, config syncing, remote queues and online validation (Redis + Qdrant).
- Flexible student architecture — the layer stack is a config list mixing
full-attention,gated-delta-net(bidirectional GDN) andmobablocks; gated attention, Liger kernels, fully unpadded varlen inference. AutoLednikModel— one entry point that loads any Lednik checkpoint (HF directory or Lightning.ckpt) by resolving the architecture from its config via the model registry.- Serving — a LitServe-based embedding server with dynamic batching, multiple inference workers and HTTP API processes, ZMQ transport, and a benchmark-oriented request protocol (raw texts or pre-tokenized ids).
- Benchmark suite — RuMTEB quality runs, pure forward-pass measurements
(
do_bench+ VRAM + tokens/sec), and an open-loop/closed-loop HTTP load generator.
lednik/
├── lednik/ # core library
│ ├── initialization/ # model factory (create_* fns), PCA, tokenizer utils
│ ├── models/ # LednikModel, StaticEmbeddingsModel, configs, outputs,
│ │ # AutoLednikModel + model/config registries
│ ├── distill/ # DistillationModule, collator, configs, losses, validation
│ ├── serving/ # LitServe embedding server (lednik.serving.server)
│ ├── emb_utils.py # teacher embedding extraction & pooling
│ ├── dist_utils.py # FSDP/DDP helpers, distributed embedding gather
│ └── path_utils.py # determine_path: ClearML ID / HF repo / local path resolver
├── pipelines/distill/ # ClearML + Lightning distillation pipeline
├── bench/
│ ├── mteb_testing/ # RuMTEB benchmark runner + model wrapper
│ ├── forward_testing/ # pure forward-pass bench (do_bench, VRAM, tokens/sec)
│ └── load_testing/ # open-/closed-loop HTTP load generator
├── docker/ # serving / training / flash-attention builder images
├── docker-compose.yaml # serving, training, qdrant, redis services
├── configs/ # YAML configs (training_settings / worker)
├── eda_utils/ # synthetic data generation utilities
├── figs/ # logo and figures
├── kostyl_toolkit/ # git submodule: the `kostyl` ML toolkit (used throughout)
└── docs/ # documentation
kostyl. Lednik depends on kostyl-toolkit, the author's personal ML toolkit, vendored as thekostyl_toolkit/git submodule and installed askostyl-toolkit[ml]. Everykostyl.*import resolves to it.
# Clone with the kostyl submodule
git clone --recurse-submodules <repo-url>
# or, if already cloned:
git submodule update --init --recursive
# Install with uv (the project's package manager)
uv sync # core + default groups (dev, distill)
uv sync --group flash-attn # optional: Flash-Attention + fla kernels for GPU inference
uv sync --group bench # optional: MTEB benchmarking
uv sync --group serving # optional: LitServe serverPython ≥ 3.13 is required (see pyproject.toml).
Initialize a Lednik Transformer student from a teacher:
from transformers import AutoModel, AutoTokenizer
from lednik.models import LednikConfig
from lednik.initialization.factory import create_lednik_transformer
teacher = AutoModel.from_pretrained("deepvk/USER-bge-m3").to("cuda").eval()
tokenizer = AutoTokenizer.from_pretrained("deepvk/USER-bge-m3")
config = LednikConfig(
hidden_size=384,
num_attention_heads=6,
intermediate_size=1152,
# the layer stack is explicit: mix full attention with linear-attention blocks
layers=["full-attention", "gated-delta-net", "full-attention"],
rope_parameters={"rope_type": "default", "rope_theta": 10000.0},
)
student = create_lednik_transformer(
model=teacher,
tokenizer=tokenizer,
model_config=config,
pooling="mean",
embedding_extraction_batch_size=256,
)
student.save_pretrained("weights/lednik_base")
tokenizer.save_pretrained("weights/lednik_base")Then distill it — see Training without ClearML for the pure-Python loop, or run the ClearML pipeline:
uv run python -m pipelines.distill.run --config-path configs/training_settings.yaml # local
uv run python -m pipelines.distill.run --config-path configs/training_settings.yaml \
--remote-execution-queue q # remote ClearML agentLoad any trained checkpoint back without knowing its concrete class:
from lednik.models import AutoLednikModel
# works for HF-format directories and Lightning .ckpt files alike;
# the class is resolved from the checkpoint config via the model registry
model = AutoLednikModel.from_pretrained(
"weights/lednik_base", weights_prefix="student.", strict_prefix=True
)The serving stack lives in lednik/serving/server.py and is
deployed via Docker Compose:
docker compose --profile serving build serving
docker compose --profile serving up -dConfiguration goes through the SERVING_* variables in .env: the model and tokenizer
references (a local path, a ClearML model ID, or an HF Hub repo id), dynamic batching,
worker and API-server counts, and the sequence-length limit. See Usage
for the full variable list, the request protocol and scaling knobs.
LednikModel— a compact encoder built from an explicit per-layer stack (config.layers):full-attentionblocks (RoPE, optional attention gating, Liger SwiGLU/GeGLU MLP), bidirectionalgated-delta-netblocks (linear attention on fla Triton kernels), and experimentalmobablocks (mixture of block attention). Supportseager,sdpa(torch varlen),flash_attention_2andflash_attention_4implementations; the varlen backends run a fully unpadded path driven bycu_seqlens/max_seqlenderived from the attention mask. Returnslast_hidden_stateand mean-pooledsentence_embeddings.StaticEmbeddingsModel— maps tokens directly to static vectors with per-token SIF weights, RMSNorm and mean pooling; no attention. Ideal for ultra-low-latency / CPU inference. AStaticEmbeddingsForSequenceClassificationhead is also available.AutoLednikModel— resolves the concrete class from a checkpoint'sarchitecturesvia the@register_model/@register_configregistries and loads HF directories or Lightning.ckptfiles (withweights_prefixfiltering for distillation checkpoints).
