Distributed 56M-parameter LLM inference across 3 ESP32-S3 boards via ESP-NOW , Split-PLE + KV cache, fully offline.
-
Updated
Jul 30, 2026 - C
Distributed 56M-parameter LLM inference across 3 ESP32-S3 boards via ESP-NOW , Split-PLE + KV cache, fully offline.
A list of production-ready models for resource-constrained devices.
An open and practical guide to Edge AI Engineering.
Official library of pre-optimized Tensorbit models. Ready-to-deploy LLMs and Vision Transformers for edge hardware, optimized via the Tensorbit P-D-Q pipeline.
Ayekoo — an offline farming assistant for Ghanaian farmers. Qwen2.5-0.5B (Q4_K_M) on llama.cpp with local retrieval over a curated Ghanaian agriculture corpus. Runs entirely on an 8GB laptop, no internet. ADTC 2026.
TinyML-based IoT system for electric motor fault detection using Edge Impulse and TensorFlow Lite Micro.
Real-time white box detection using YOLOv8n and LiteRT with edge optimization and false-positive handling.
Edge-GNN: Constraint-aware graph neural networks for biological interaction modeling under edge deployment constraints (computational oncology, PPI networks).
Adaptive Cognitive Budgeting (ACB) dynamically allocates context windows and CPU threads based on query complexity and host memory pressure to prevent catastrophic latency stalls in GPU-free LLM inference.
Vision·FSM pipeline for a Jetson Orin Nano follow robot — YOLO11n TensorRT + CLIP-ReID re-identification, Python to C++ reimplementation (11.1Hz → 29Hz)
Run a 56M-parameter language model across three ESP32-S3 boards using ESP-NOW for distributed inference.
To associate your repository with the edge-ai-models topic, visit your repo's landing page and select "manage topics."