Skip to content
View Air000000's full-sized avatar

Highlights

  • Pro

Block or report Air000000

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Air000000/README.md

李宗宽 / Air

AI Agent & LLM Application Engineer · Master's Candidate in Communication Engineering

Building controllable, recoverable, and evaluable AI Agent systems.
Focused on Agent Runtime / Harness, RAG evaluation, Context Engineering, and reliable backend systems.

Open to 2027 Graduate Roles Email

Engineering Highlights

  • Agent Runtime / Harness — bounded Agent Loop, versioned Tool Contracts, persist-before-execute, experiment governance, crash reconciliation, cancellation, and deterministic result validation.
  • RAG Evaluation — frozen TechQA benchmark with 28,481 documents / 610 answerable queries; Dense Top-100 + rerank improved held-out Recall@5 from 64.4% → 72.5% (+8.1pp) and MRR@10 from 0.519 → 0.561.
  • Backend Engineering — Python / FastAPI / PostgreSQL, async workflows, state machines, database migrations, automated tests, CI, and reproducible local environments.

Featured Projects

A controllable, recoverable, and verifiable runtime for machine-learning experiments. It separates Agent decisions, human approval, physical execution, recovery, and deterministic result validation instead of letting an LLM directly control long-running training jobs.

Engineering evidence: fault-injection validation showed that a Coordinator restart could reconcile the original training Attempt without a second physical launch, with explicit cancellation paths for running and pre-launch work.

Python FastAPI PostgreSQL asyncio Agent Runtime Context Engineering Tool Calling SSE


An enterprise IT support backend combining document lifecycle management, RAG, controlled ticket creation, Human-in-the-loop approval, AgentOps tracing, and an offline evaluation pipeline.

Held-out DEV: Dense Top-100 + qwen3-rerank improved Recall@5 from 64.4% → 72.5% (+8.1pp) and MRR@10 from 0.519 → 0.561 on a frozen TechQA benchmark.

Python FastAPI SQLAlchemy ChromaDB RAG Rerank AgentOps Pytest


A Windows-first local desktop app for understanding computer usage, work-rest rhythm, reminders, and focus signals. It captures foreground-app activity, applies privacy processing and local classification rules, stores data in SQLite, and provides review workflows and desktop-companion surfaces.

Tauri Vue 3 TypeScript Rust SQLite Local-first


A leakage-controlled Agent Skill optimization and evaluation pipeline designed to distinguish reusable Skill improvements from public-task overfitting.

Evaluation design: frozen 60 / 17 / 10 Dev–Holdout–Final-Blind split, explicit no_skill vs with_skill paired runs, execution provenance, regression gates, and generalization checks.

Python Agent Skills BenchFlow SkillsBench Agent Evaluation Docker

Other Projects

  • Codex History Restorer — local Windows tool for recovering Codex Desktop conversations that still exist on disk but no longer appear in the application.
  • Bilibili Content — reusable CLI tool and AI Skill for subtitle extraction, validation, transcription, and structured downstream processing.

Tech Stack

  • Agent / LLM: Context Engineering · Tool / Function Calling · Human-in-the-loop · RAG · Retrieval / Rerank · Agent Evaluation
  • Backend: Python · FastAPI · asyncio · SQLAlchemy · PostgreSQL · SQLite
  • Engineering: Pytest · Alembic · GitHub Actions · Docker Compose · Git · Linux / Shell · Ruff · mypy
  • Product / Desktop: TypeScript · Vue 3 · Rust · Tauri
  • Research: PyTorch · causal discovery · causal representation learning · generative models · counterfactual generation

Research

Master's candidate in Communication Engineering, researching causal structure learning for diffusion-based generative models, including causal discovery, causal representation learning, and counterfactual image generation.

First-author manuscript submitted to Neurocomputing.

Popular repositories Loading

  1. Enterprise-Support-AI-Copilot-API Enterprise-Support-AI-Copilot-API Public

    Python 10

  2. codex-history-restorer codex-history-restorer Public

    Restore local Codex Desktop chats that still exist on disk but no longer appear in the app.

    PowerShell 5

  3. bilibili-content bilibili-content Public

    Extract subtitles and transcribe audio from Bilibili videos

    Python 4

  4. dev-error-notebook-skill dev-error-notebook-skill Public

    Reusable engineering error notebook skill for coding agents.

    3

  5. timeprism timeprism Public

    A Tauri desktop app for focus tracking, reminders, and attention insights.

    Vue 3

  6. Air000000 Air000000 Public

    Causal inference for images · LLM apps · Agent systems · Developer tools

    2