Validated GLM-5.3 Flash recipe for 2x NVIDIA RTX PRO 6000 Blackwell 96GB: 262K context, EXL3/TR3, adaptive MTP, tools, and vision.
-
Updated
Sep 20, 2026 - Python
Validated GLM-5.3 Flash recipe for 2x NVIDIA RTX PRO 6000 Blackwell 96GB: 262K context, EXL3/TR3, adaptive MTP, tools, and vision.
Workload-aware persistent expert-tile scheduling for irregular MoE inference kernels on NVIDIA Blackwell GPUs
A production-ready Docker setup for ComfyUI that unlocks the full potential of NVIDIA Blackwell GPUs (RTX 50 series) through 4-bit quantization with NVFP4.
RTX 5090 & RTX 5060 Docker container with PyTorch + TensorFlow. First fully-tested Blackwell GPU support for ML/AI. CUDA 12.8, Python 3.11, Ubuntu 24.04. Works with RTX 50-series (5090/5080/5070/5060) and RTX 40-series.
Sample application generated using Opencode and Ollama
LoRA fine-tune and serve NVFP4 models on one DGX Spark (GB10, 128 GB UMA): text backbones via generic-family onboarding, plus VLMs (vision tower, or LLM+tower jointly via --train-target both) validated end-to-end on Pixtral and Nemotron-Omni. Fused Triton dequant; runtime-LoRA and merge serving.
🚀 Accelerate image generation with ComfyUI's Docker for NVIDIA Blackwell GPUs, optimizing speed and memory usage through NVFP4 support.
Disaggregated Prefill-Decode serving engine connecting NVIDIA Blackwell (SM120) and Huawei Ascend 910B2 (CANN) across physical nodes.
Empirical characterization of ICICLE NTT on consumer NVIDIA Blackwell (RTX 5070, sm_120), with a prototype for the digit-reversal bottleneck.
Public, one-way mirror of jgdynamite10/blackwell-agentic-inference-lab. Development, issues, pull requests, releases, and security reporting live in the canonical repository.
Training small language models from scratch and building a high-performance Rust/CUDA inference engine for NVIDIA gpus (optimized for Blackwell).
Curated RTX 5090 benchmarks, model recipes, inference engines, creative AI workflows, and system tools
High-performance, multi-GPU computer vision pipeline for Genetec streams using NVIDIA DeepStream, Triton, and Numba JIT to track personnel, evaluate weapon threats, and perform in-VRAM OCR analytics.
Deploy GLM-5.3 Flash with DFlash2 speculative decoding on dual RTX PRO 6000 Blackwell 96GB GPUs, enabling one-million-token context and 16-image prompts via an OpenAI-compatible API.
High-performance, multi-GPU computer vision pipeline for Genetec streams using NVIDIA DeepStream, Triton, and Numba JIT to track personnel, evaluate weapon threats, and perform in-VRAM OCR analytics.
Reproducible research for measuring agentic AI inference outcomes on NVIDIA Blackwell across cloud environments.
To associate your repository with the nvidia-blackwell topic, visit your repo's landing page and select "manage topics."