Skip to content
View wittyicon29's full-sized avatar
πŸ”­
Daydreaming
πŸ”­
Daydreaming

Block or report wittyicon29

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
wittyicon29/README.md

about.py projects/ oncall.sh htop deploys.log mail trash

Click an icon to open its window.

atharva@wittyos: ~/about.py

class Atharva(Engineer):
    role     = "AI / ML / GenAI Engineer"
    company  = "CFOLogic"
    focus    = ["speech pipelines", "LLM systems", "agents", "interpretable ML"]
    speaks   = ["python", "c++", "java"]
    motto    = "most of what I build listens for a while before it decides anything"

    def previously(self):
        return {
            "Emitrr":   "cut transcription latency on a live voice pipeline by 30%, "
                        "rebuilt the TTS path, built an agent that recovers failing calls",
            "Research": "non-invasive blood glucose estimation (published)",
        }

    def off_hours(self):
        return "small experiments that poke at the edges of these systems"

Close the window and go back to the desktop

files: ~/projects

atharva@wittyos:~$ ls -l projects/ --sort=interesting
Permissions Name Modified What it does
drwxr-xr-x πŸ“ guardtheweights/ 2026‑02‑28 Adversarial LLM game: get a cryptic narrator to give up its secrets
drwxr-xr-x πŸ“ PWC-RAG/ 2024‑03‑08 RAG search over recent ML research
drwxr-xr-x πŸ“ RAG-over-Audio-Data/ 2023‑12‑17 Ask questions about hours of audio
drwxr-xr-x πŸ“ Agentic-RAG-Customer-Support/ 2025‑03‑02 Support agent with KB and search tools
drwxr-xr-x πŸ“ Stutter_Detection/ 2023‑12‑24 Wav2vec-based stutter detection
drwxr-xr-x πŸ“ WeedWatch/ 2024‑03‑26 CNN weed detection, most-starred ⭐
drwxr-xr-x πŸ“ AstoGemma/ 2024‑02‑27 Local chat UI for any Ollama model
drwxr-xr-x πŸ“ In-Memory-DB/ 2024‑12‑21 A small RAM-first database in C++

8 directories shown. 66 more β†’

Close the window and go back to the desktop

oncall.sh: pager

It's 3:12 AM and this pipeline is blowing its latency SLO. Callers are talking over the bot. Can you fix it before stand-up?

flowchart LR
    caller(("πŸ“ž Caller")) -->|audio| vad["VAD +<br/>endpointing"]
    vad --> stt["Streaming<br/>STT"]
    stt -->|partials| llm["LLM agent"]
    llm <-->|tools / RAG| kb[("Knowledge<br/>base")]
    llm -->|sentences| tts["TTS"]
    tts -->|audio| caller
Loading

Every choice is a link that jumps to the next scene. There's nothing to install: just click.

Run oncall.sh

Close the window and go back to the desktop

htop

  Repos by primary language (39 of my own repos)
  Jupyter [|||||||||||||||||||                      48.7%]
  Python  [|||||||||||||                            33.3%]
  Java    [|||                                       7.7%]
  C++     [|||                                       7.7%]
  HTML    [|                                         2.6%]

    PID USER     PRI   CPU%  MEM%  TIME+   Command
      1 atharva   20   38.0  12.4  4y8m    speech-pipeline --streaming
     29 atharva   20   27.5   9.1  3y2m    llm-systems --rag --agents
    404 atharva   20   14.2   6.3  2y1m    interpretable-ml --readable
   3012 atharva   20    9.8   2.0  ∞       curiosity
   1500 atharva   39    0.0   0.1  0:00    silence_ms   (killed, see oncall.sh)

  F1Help  F2Setup  F3Search  F9Kill  F10Quit

The language meters are real. The process list is vibes.

atharva@wittyos:~$ apt list --installed

Languages and ML frameworks
Tooling and infrastructure

More meters

GitHub stats Top languages

GitHub streak

Close the window and go back to the desktop

tail -f /var/log/deploys.log

2026-10-07 14:49  [deploy]  portfolio                    Personal portfolio website
2026-09-23 17:22  [deploy]  neetcode-submissions         My NeetCode.io problem submissions
2026-02-28 14:25  [deploy]  guardtheweights              GuardTheWeights is an adversarial LLM game where a langua...
2026-02-21 13:37  [deploy]  DSA                          Some problems of leetcode on various techniques and conce...
2025-05-05 12:05  [deploy]  CreditScoreAI

Close the window and go back to the desktop

mail: new message

To:       atharva
From:     you
Subject:  let's build something that listens

Happy to talk about a role, a project, or the finer points of threshold calibration.

Portfolio LinkedIn Email Stack Overflow

Close the window and go back to the desktop

trash: 4 items

atharva@wittyos:~$ ls ~/.trash
it_works_on_my_machine.txt
flaky_test_final_FINAL_v3.py
silence_ms=1500.yaml
humor_found_in_bugs.md        # 0 bytes. Never found any, physically or virtually.

atharva@wittyos:~$ rm -rf ~/.trash
rm: cannot remove 'silence_ms=1500.yaml': Device or resource busy (it's load-bearing)

Close the window and go back to the desktop

↓ oncall.sh lives below this line. Spoilers ahead if you scroll instead of clicking. ↓

oncall.sh: SEV-2 incident

🚨 03:12 AM · SEV-2 · voice-agent p95 latency 4.2 s (SLO 1.2 s)

Callers are talking over the bot and the on-call phone won't stop buzzing. Here's a slow trace:

trace 7f3a91 ─────────────────────────── total 4,210 ms
β”œβ”€ endpointing (vad)   1,840 ms  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ
β”œβ”€ stt.final             620 ms  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Œ
β”œβ”€ llm.first_token       980 ms  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š
└─ tts.first_audio       770 ms  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‰

What do you do first?

oncall.sh: SEV-2 incident

πŸ’Έ 03:19 AM Β· You doubled the STT pool

stt-pool     replicas 4 β†’ 8        gpu util 23% β†’ 11%
cost         +$1,900 / month
p95          4.21 s β†’ 4.18 s

The GPUs were never the bottleneck. Finance is going to notice this one.

oncall.sh: SEV-2 incident

πŸͺΆ 03:21 AM Β· You swapped in a smaller model

llm.first_token          980 ms β†’ 560 ms
eval: intent accuracy    94.1% β†’ 83.0%   ❌
p95                      4.21 s β†’ 3.79 s

It's faster, but "reschedule my appointment" now routes to cancel. You're still three times over the SLO, and the eval suite is red.

oncall.sh: SEV-2 incident

πŸ” 03:24 AM Β· The endpointing span

The pipeline waits for silence to decide that the caller has finished speaking. Here's the config:

# voice-agent/config/endpointing.yaml
endpointing:
  silence_ms: 1500        # bumped during last month's noisy-line incident
  min_utterance_ms: 300

Every single turn pays 1.5 s of silence before STT even finalizes.

oncall.sh: SEV-2 incident

βœ‚οΈ 03:27 AM Β· Fast, and rude

caller:  "I'd like to book an appointment for, umβ€”"
agent:   "Sure! What day works for you?"
caller:  "...for next Tuesday. Why did you interrupt me?"
p95 latency      4.21 s β†’ 2.68 s
barge-in rate    4% β†’ 17%   ❌

Latency dropped, but the bot now cuts people off mid-thought. A pause isn't the same as being finished.

oncall.sh: SEV-2 incident

🧠 03:41 AM · Semantic endpointing

Wait briefly when the transcript sounds finished, and longer when it doesn't:

def end_of_turn(partial: str, silence_ms: int) -> bool:
    """Decide whether the caller has finished speaking."""
    if silence_ms >= 900:                              # hard ceiling
        return True
    if silence_ms >= 250 and sounds_complete(partial):
        return True                                    # "...next Tuesday at 4."  -> go
    return False                                       # "...for, um"             -> keep listening

def sounds_complete(text: str) -> bool:
    # cheap heuristics first, a tiny classifier only for the ambiguous middle
    if text.rstrip().endswith(("um", "uh", "and", "for", "the", ",")):
        return False
    return turn_classifier.predict_proba(text) > 0.85
trace b81c02 ─────────────────────────── total 2,310 ms
β”œβ”€ endpointing (vad)     310 ms  β–ˆβ–ˆβ–Š
β”œβ”€ stt.final             480 ms  β–ˆβ–ˆβ–ˆβ–ˆβ–
β”œβ”€ llm.first_token       860 ms  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–Š
└─ tts.first_audio       660 ms  β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ

barge-in rate    4% β†’ 4%   βœ…

Nobody gets interrupted, but 2.3 s is still almost double the SLO. Look at the shape of the trace: each stage waits for the previous one to finish completely.

oncall.sh: SEV-2 incident

πŸ—‚οΈ 03:52 AM Β· A TTS cache

cached phrases     "One moment.", "Sure!", "Could you repeat that?" ... (40)
cache hit rate     18%
p95 latency        2.31 s β†’ 2.19 s

That helps with greetings and fillers, but most replies are unique. The real cost is that every stage still runs one after the other.

oncall.sh: SEV-2 incident

🌊 04:10 AM · Streaming the pipeline

Stop handing finished results down the line. Let every stage start on partial input:

  • the LLM starts on the stable partial transcript, before stt.final, and is cancelled if the caller keeps talking;
  • TTS starts speaking on the first complete sentence, not the whole reply.
async for partial in stt.stream(audio):                 # simplified
    if partial.is_stable and llm_task is None:
        llm_task = start_llm(partial.text)              # speculative head start
    elif llm_task and partial.changes_meaning_of(llm_task.prompt):
        llm_task.cancel(); llm_task = None              # caller kept talking
    if end_of_turn(partial.text, vad.silence_ms):
        llm_task = llm_task or start_llm(partial.text)
        async for sentence in llm_task.sentences():
            await tts.speak(sentence)                   # first sentence plays immediately
        break
gantt
    title One caller turn, before and after
    dateFormat x
    axisFormat %-S.%L s
    section Sequential
    endpointing           : 0, 310
    STT final             : 310, 790
    LLM first token       : 790, 1650
    TTS first audio       : 1650, 2310
    section Streamed
    endpointing           : 0, 310
    STT final (partials)  : 310, 430
    LLM (speculative)     : 200, 760
    TTS (first sentence)  : 760, 1040
Loading
p95 latency      2.31 s β†’ 1.04 s   βœ… under SLO

oncall.sh: SEV-2 incident

βœ… 04:26 AM Β· Resolved

SEV-2 voice-agent latency ........................ RESOLVED
p95             4.21 s β†’ 1.04 s    (SLO 1.2 s)
barge-in rate   unchanged
cost            unchanged
time to fix     74 min

Postmortem, in three lines

  1. Root cause: a static 1,500 ms silence threshold, plus stages that ran strictly one after another.
  2. Fix: semantic endpointing, then streaming with a speculative LLM start.
  3. Lesson: read the trace before buying hardware.

This incident was made up and simplified, but the problems are real ones I've worked on. At Emitrr I cut transcription latency on a live voice pipeline by 30%, rebuilt the TTS path, and built an agent[...]

Back to the desktop Play again Let's talk

Pinned Loading

  1. guardtheweights guardtheweights Public

    GuardTheWeights is an adversarial LLM game where a language model guards hidden lore inside an imaginary world. Players attempt to extract secrets using carefully crafted prompts, while the model d…

    Python

  2. Career-Compass-A-Multi-agent-career-guide Career-Compass-A-Multi-agent-career-guide Public

    Forked from BytefulRashi/Career-Compass

    A multi agent driven career guide providing real time market insights with profile assessment and skill evaluation using CrewAI and AWS

    Python 3 1

  3. PWC-RAG PWC-RAG Public

    A RAG application to search through the recent ML research without going through the Papers with Code manually.

    Python 6 2

  4. RAG-over-Audio-Data RAG-over-Audio-Data Public

    Python-based system designed to transcribe audio files, split the transcripts into manageable chunks, create text embeddings using HuggingFace models, and employ advanced question-answering models …

    Python 6 1

  5. AstoGemma AstoGemma Public

    A completely local chat UI for any LLMs available on ollama using LangSmith and ChainLit

    Python 1