Skip to content
#

h200

Here are 16 public repositories matching this topic...

Thermal-aware batch controller for vLLM/TensorRT-LLM. Prevents HBM thermal throttling from killing p99 latency on H100/H200. Monitors nvidia-smi, auto-cuts batch size at 85°C, migrates cold KV to DRAM. Prometheus + Grafana included. 4.2s -> 2.1s p99 at 128K context.

  • Updated Apr 13, 2026
  • Python

Add this topic to your repo

To associate your repository with the h200 topic, visit your repo's landing page and select "manage topics."

Learn more