Click an icon to open its window.
class Atharva(Engineer):
role = "AI / ML / GenAI Engineer"
company = "CFOLogic"
focus = ["speech pipelines", "LLM systems", "agents", "interpretable ML"]
speaks = ["python", "c++", "java"]
motto = "most of what I build listens for a while before it decides anything"
def previously(self):
return {
"Emitrr": "cut transcription latency on a live voice pipeline by 30%, "
"rebuilt the TTS path, built an agent that recovers failing calls",
"Research": "non-invasive blood glucose estimation (published)",
}
def off_hours(self):
return "small experiments that poke at the edges of these systems"atharva@wittyos:~$ ls -l projects/ --sort=interesting| Permissions | Name | Modified | What it does |
|---|---|---|---|
drwxr-xr-x |
π guardtheweights/ | 2026β02β28 | Adversarial LLM game: get a cryptic narrator to give up its secrets |
drwxr-xr-x |
π PWC-RAG/ | 2024β03β08 | RAG search over recent ML research |
drwxr-xr-x |
π RAG-over-Audio-Data/ | 2023β12β17 | Ask questions about hours of audio |
drwxr-xr-x |
π Agentic-RAG-Customer-Support/ | 2025β03β02 | Support agent with KB and search tools |
drwxr-xr-x |
π Stutter_Detection/ | 2023β12β24 | Wav2vec-based stutter detection |
drwxr-xr-x |
π WeedWatch/ | 2024β03β26 | CNN weed detection, most-starred β |
drwxr-xr-x |
π AstoGemma/ | 2024β02β27 | Local chat UI for any Ollama model |
drwxr-xr-x |
π In-Memory-DB/ | 2024β12β21 | A small RAM-first database in C++ |
8 directories shown. 66 more β
It's 3:12 AM and this pipeline is blowing its latency SLO. Callers are talking over the bot. Can you fix it before stand-up?
flowchart LR
caller(("π Caller")) -->|audio| vad["VAD +<br/>endpointing"]
vad --> stt["Streaming<br/>STT"]
stt -->|partials| llm["LLM agent"]
llm <-->|tools / RAG| kb[("Knowledge<br/>base")]
llm -->|sentences| tts["TTS"]
tts -->|audio| caller
Every choice is a link that jumps to the next scene. There's nothing to install: just click.
Repos by primary language (39 of my own repos)
Jupyter [||||||||||||||||||| 48.7%]
Python [||||||||||||| 33.3%]
Java [||| 7.7%]
C++ [||| 7.7%]
HTML [| 2.6%]
PID USER PRI CPU% MEM% TIME+ Command
1 atharva 20 38.0 12.4 4y8m speech-pipeline --streaming
29 atharva 20 27.5 9.1 3y2m llm-systems --rag --agents
404 atharva 20 14.2 6.3 2y1m interpretable-ml --readable
3012 atharva 20 9.8 2.0 β curiosity
1500 atharva 39 0.0 0.1 0:00 silence_ms (killed, see oncall.sh)
F1Help F2Setup F3Search F9Kill F10Quit
The language meters are real. The process list is vibes.
atharva@wittyos:~$ apt list --installed2026-10-07 14:49 [deploy] portfolio Personal portfolio website
2026-09-23 17:22 [deploy] neetcode-submissions My NeetCode.io problem submissions
2026-02-28 14:25 [deploy] guardtheweights GuardTheWeights is an adversarial LLM game where a langua...
2026-02-21 13:37 [deploy] DSA Some problems of leetcode on various techniques and conce...
2025-05-05 12:05 [deploy] CreditScoreAI
To: atharva
From: you
Subject: let's build something that listens
Happy to talk about a role, a project, or the finer points of threshold calibration.
atharva@wittyos:~$ ls ~/.trash
it_works_on_my_machine.txt
flaky_test_final_FINAL_v3.py
silence_ms=1500.yaml
humor_found_in_bugs.md # 0 bytes. Never found any, physically or virtually.
atharva@wittyos:~$ rm -rf ~/.trash
rm: cannot remove 'silence_ms=1500.yaml': Device or resource busy (it's load-bearing)β oncall.sh lives below this line. Spoilers ahead if you scroll instead of clicking. β
Callers are talking over the bot and the on-call phone won't stop buzzing. Here's a slow trace:
trace 7f3a91 βββββββββββββββββββββββββββ total 4,210 ms
ββ endpointing (vad) 1,840 ms βββββββββββββββββ
ββ stt.final 620 ms ββββββ
ββ llm.first_token 980 ms βββββββββ
ββ tts.first_audio 770 ms βββββββ
What do you do first?
- π₯οΈ Scale the STT GPU pool to 2Γ
- π Dig into the endpointing span
- πͺΆ Swap the LLM for a smaller, faster model
stt-pool replicas 4 β 8 gpu util 23% β 11%
cost +$1,900 / month
p95 4.21 s β 4.18 s
The GPUs were never the bottleneck. Finance is going to notice this one.
llm.first_token 980 ms β 560 ms
eval: intent accuracy 94.1% β 83.0% β
p95 4.21 s β 3.79 s
It's faster, but "reschedule my appointment" now routes to cancel. You're still three times over the SLO, and the eval suite is red.
The pipeline waits for silence to decide that the caller has finished speaking. Here's the config:
# voice-agent/config/endpointing.yaml
endpointing:
silence_ms: 1500 # bumped during last month's noisy-line incident
min_utterance_ms: 300Every single turn pays 1.5 s of silence before STT even finalizes.
caller: "I'd like to book an appointment for, umβ"
agent: "Sure! What day works for you?"
caller: "...for next Tuesday. Why did you interrupt me?"
p95 latency 4.21 s β 2.68 s
barge-in rate 4% β 17% β
Latency dropped, but the bot now cuts people off mid-thought. A pause isn't the same as being finished.
Wait briefly when the transcript sounds finished, and longer when it doesn't:
def end_of_turn(partial: str, silence_ms: int) -> bool:
"""Decide whether the caller has finished speaking."""
if silence_ms >= 900: # hard ceiling
return True
if silence_ms >= 250 and sounds_complete(partial):
return True # "...next Tuesday at 4." -> go
return False # "...for, um" -> keep listening
def sounds_complete(text: str) -> bool:
# cheap heuristics first, a tiny classifier only for the ambiguous middle
if text.rstrip().endswith(("um", "uh", "and", "for", "the", ",")):
return False
return turn_classifier.predict_proba(text) > 0.85trace b81c02 βββββββββββββββββββββββββββ total 2,310 ms
ββ endpointing (vad) 310 ms βββ
ββ stt.final 480 ms βββββ
ββ llm.first_token 860 ms ββββββββ
ββ tts.first_audio 660 ms ββββββ
barge-in rate 4% β 4% β
Nobody gets interrupted, but 2.3 s is still almost double the SLO. Look at the shape of the trace: each stage waits for the previous one to finish completely.
cached phrases "One moment.", "Sure!", "Could you repeat that?" ... (40)
cache hit rate 18%
p95 latency 2.31 s β 2.19 s
That helps with greetings and fillers, but most replies are unique. The real cost is that every stage still runs one after the other.
Stop handing finished results down the line. Let every stage start on partial input:
- the LLM starts on the stable partial transcript, before
stt.final, and is cancelled if the caller keeps talking; - TTS starts speaking on the first complete sentence, not the whole reply.
async for partial in stt.stream(audio): # simplified
if partial.is_stable and llm_task is None:
llm_task = start_llm(partial.text) # speculative head start
elif llm_task and partial.changes_meaning_of(llm_task.prompt):
llm_task.cancel(); llm_task = None # caller kept talking
if end_of_turn(partial.text, vad.silence_ms):
llm_task = llm_task or start_llm(partial.text)
async for sentence in llm_task.sentences():
await tts.speak(sentence) # first sentence plays immediately
breakgantt
title One caller turn, before and after
dateFormat x
axisFormat %-S.%L s
section Sequential
endpointing : 0, 310
STT final : 310, 790
LLM first token : 790, 1650
TTS first audio : 1650, 2310
section Streamed
endpointing : 0, 310
STT final (partials) : 310, 430
LLM (speculative) : 200, 760
TTS (first sentence) : 760, 1040
p95 latency 2.31 s β 1.04 s β
under SLO
SEV-2 voice-agent latency ........................ RESOLVED
p95 4.21 s β 1.04 s (SLO 1.2 s)
barge-in rate unchanged
cost unchanged
time to fix 74 min
Postmortem, in three lines
- Root cause: a static 1,500 ms silence threshold, plus stages that ran strictly one after another.
- Fix: semantic endpointing, then streaming with a speculative LLM start.
- Lesson: read the trace before buying hardware.
This incident was made up and simplified, but the problems are real ones I've worked on. At Emitrr I cut transcription latency on a live voice pipeline by 30%, rebuilt the TTS path, and built an agent[...]
