Edge deployment · On-prem inference · Production ML systems
I design and deploy AI systems for environments where latency, privacy, hardware constraints, or cloud dependence matter.
- Edge AI: efficient inference on CPUs, mobile devices, and constrained hardware
- On-prem AI: private model serving, retrieval, and data pipelines
- Production ML: model optimization, deployment, evaluation, and observability
- Agent systems: developer tooling and evidence-driven evaluation
- DEIMKit: a practical Python toolkit for training, inference, ONNX export, and deployment with the DEIM object detector
- tracegrad: evidence-gated system-prompt optimization using traces and LLM-judge results
- firstmate: an agent distro that orchestrates coding-agent crews through isolated worktrees and supervised delivery
- Supercharge Your PyTorch Image Models: benchmark-driven inference optimization with ONNX Runtime and TensorRT
- YOLOv5 with DeepSparse: run object detection at more than 180 FPS on a four-core CPU
- PyTorch at the Edge: deploy more than 900 TIMM models on Android with TorchScript and Flutter
|
|
Accelerate TIMM image models with ONNX Runtime and TensorRT, including optimized preprocessing and runtime configuration. September 30, 2024 |
|
|
PyTorch at the Edge: Deploying Over 964 TIMM Models on Android with TorchScript and Flutter Train with Fastai, export with TorchScript, and deploy TIMM models to Android through Flutter. February 7, 2023 |
|
|
Supercharging YOLOv5: How I Got 182.4 FPS Inference Without a GPU Optimize YOLOv5 for CPU inference with SparseML, pruning, quantization, and DeepSparse. June 7, 2022 |
|
|
Faster than GPU: How to 10x your Object Detection Model and Deploy on CPU at 50+ FPS Convert YOLOX to ONNX and OpenVINO, then quantize it for real-time CPU inference above 50 FPS. April 30, 2022 |
|
I Made It to GitHub Trending - My Open Source Journey How x.infer reached GitHub's trending developers list—and what I learned from building and sharing open source. October 28, 2024 |
|
|
Accelerate TIMM image models with ONNX Runtime and TensorRT, including optimized preprocessing and runtime configuration. September 30, 2024 |
|
|
Celebrating a Milestone in the Top 2% of Global Scientists A reflection on ten years in research, the transition from academia to industry, and recognition in Stanford's 2023 scientist ranking. November 17, 2023 |
|
|
PyTorch at the Edge: Deploying Over 964 TIMM Models on Android with TorchScript and Flutter Train with Fastai, export with TorchScript, and deploy TIMM models to Android through Flutter. February 7, 2023 |
|
|
Supercharging YOLOv5: How I Got 182.4 FPS Inference Without a GPU Optimize YOLOv5 for CPU inference with SparseML, pruning, quantization, and DeepSparse. June 7, 2022 |
|
|
Faster than GPU: How to 10x your Object Detection Model and Deploy on CPU at 50+ FPS Convert YOLOX to ONNX and OpenVINO, then quantize it for real-time CPU inference above 50 FPS. April 30, 2022 |
Included in the 2023 Stanford/Elsevier database of top-cited scientists.
Python · PyTorch · ONNX Runtime · TensorRT · OpenVINO · llama.cpp · vLLM · Docker · TypeScript
I'm the Founder and CTO of NeuralEngine AI. I work with teams that need AI to run privately, reliably, and close to where their data is produced.






