doublewordai
Popular repositories Loading
-
control-layer
control-layer PublicThe world’s fastest AI model gateway (450x less overhead than LiteLLM). Unified access to LLMs across endpoints (openAI, self-hosted, etc.) behind a single authentication layer - with API key gener…
-
autobatcher
autobatcher PublicDrop-in AsyncOpenAI replacement that transparently batches requests
-
deepseek-reddit-agent
deepseek-reddit-agent PublicAn example notebook which shows how you can build a LLM agent that scrapes information from Reddit and summarize key bullets using a self-hosted DeepSeek-R1-Distill-Llama-8B deployed with Titan Tak…
-
inference-stack
inference-stack PublicThe Doubleword Inference Stack is the easiest & most performant way to run genAI infrastructure in your private environment.
-
inference-lab
inference-lab PublicHigh-performance LLM inference simulator for analyzing serving systems
Repositories
- sglang Public Forked from sgl-project/sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
- dynamo Public Forked from ai-dynamo/dynamo
A Datacenter Scale Distributed Inference Serving Framework
- control-layer Public
The world’s fastest AI model gateway (450x less overhead than LiteLLM). Unified access to LLMs across endpoints (openAI, self-hosted, etc.) behind a single authentication layer - with API key generation, user management, request logging, and more
- vllm Public Forked from vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
- uccl Public Forked from uccl-project/uccl
UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)
- modelexpress Public Forked from ai-dynamo/modelexpress
Model Express is a Rust-based component meant to be placed next to existing model inference systems to speed up their startup times and improve overall performance.
- silt Public
Top languages
Loading…
Most used topics
Loading…