Skip to content

Latest commit

 

History

29 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Eavesdrop

Time-synced lyrics, floating over everything you do.

A transparent, always-on-top, click-through overlay showing the words to whatever is playing. No dock icon, no window chrome, no focus stealing.

It listens to Spotify so you don't have to look at it.

Eavesdrop showing synced lyrics over a terminal

Platform Player Status
macOS 14+ Spotify (ScriptingBridge) v1
Windows any (SMTC) experimental
Android 8+ any (MediaSessionManager) experimental

Architecture

mac/  Swift, AppKit + SwiftUI  ─┐
win/  Rust, Direct2D           ─┼─→  core/ — Rust
android/  Kotlin + Compose     ─┘      lrc      LRC parsing (plain / repeat / word)
                                       sync     the timing engine
mac + android via UniFFI               lyrics   provider chain: local → LRCLIB → …
win links the core directly            netease  NetEase Cloud Music provider
                                       matcher  title/artist normalisation + ranking
                                       romanize offline transliteration fallback
                                       cache    on-disk LRC cache by track id
                                       offsets  per-track timing nudges, persisted
                                       theme    colour / weight / type scale
                                       style    the lyric presets, as data

The core owns everything that isn't a window. Parsing, lookup, matching, caching and timing are shared; the shell owns the overlay, the player integration and rendering. Porting means implementing one interface — PlayerSource — so a new OS is a UI change, not a product change. Lyrics lookup never changes, because LRCLIB keys on artist/title/duration rather than on Spotify.

No polling. Players push on every play, pause, seek and track change. Between pushes the core interpolates against a monotonic clock, and a read every 5s corrects drift. There is no render loop either: next_change_ms returns how long until the display next changes and the host arms a single timer for exactly that moment. The per-frame call returns indices only, never strings, so tick() allocates nothing.

Measured on an M-series Mac while playing: ~1.0% CPU, 17–22 MB (phys_footprint), ~10 MB universal bundle.

Build

Run the core's tests with cd core && cargo test (138 units, no network) or cargo test -- --ignored for live provider integration. EAVESDROP_DEBUG=1 logs the active line index to stderr — and on Windows the app id of the media session the overlay decided to follow, which is the quick way to see why it is showing lyrics for something you didn't expect.

Every build script invokes cargo by explicit path (~/.cargo/bin/cargo), because Homebrew's rust ships only the host target and cannot cross-compile.

macOS

rustup target add aarch64-apple-darwin x86_64-apple-darwin  # optional, for universal
./build.sh release      # -> build/Eavesdrop.app

Falls back to a host-only build without both targets. The app is ad-hoc signed — without it macOS cannot remember permission grants.

Windows

Built to be cross-compiled from WSL.

rustup target add x86_64-pc-windows-gnu
sudo apt install gcc-mingw-w64-x86-64

./build-win.sh release   # -> build/win/Eavesdrop.exe

The HTTP stack (ureq + rustls) needs no OpenSSL, so mingw is the only dependency. If the gnu linker fights the windows crate, use cargo-xwin and target x86_64-pc-windows-msvc instead.

Icons are compiled in by win/build.rs via windres (same mingw package), and a build without it degrades to no icon rather than failing. Run the copy the script leaves on NTFS, not the one in the repo — launching over \\wsl.localhost means launching from what Windows counts as a network share, and Explorer does not read icons out of executables on network paths, so the app shows the generic .exe glyph however well-formed its resources are. Set WIN_INSTALL_DIR to choose where that copy lands.

Android

Needs the Android SDK with an NDK and JDK 17.

rustup target add aarch64-linux-android x86_64-linux-android
cargo install cargo-ndk

./build-android.sh debug            # -> android/app/build/outputs/apk/debug
./build-android.sh debug install    # and push to a connected device

Gradle is the real build; the script just fails usefully when the toolchain is missing. Builds arm64-v8a and x86_64 (the latter only for the emulator).

Trap: UniFFI's Kotlin bindings call through JNA, and R8 renames JNA's classes in release builds unless keep rules are added by hand. They are not in the JNA aar — see android/app/proguard-rules.pro. It fails only in release.

Using it

Hover the overlay (tap-and-hold the handle on Android) and a small dot appears bottom-left.

  • Drag the dot — moves the overlay, no need to focus first.
  • Click the dot — promotes the overlay to a real window: glass background, focus, track title, and previous / play-pause / next.
  • Resize while focused. Panel height is the source of truth for text size, so the lyrics scale with it. Clamped to 340×78 … 2400×403, i.e. 14–72pt.
  • Escape, or click away — back to ghost mode.

Lines scroll rather than crossfade. A track change flashes the name and artist for four seconds; a resume does not. Pausing fades the panel out over ¾s and playing brings it back in a quarter of that.

Position, size, text size and colour persist. Reset size & position in the menu is the escape hatch if the overlay ends up somewhere unusable.

Transport is optional at the protocol level — canControl defaults to false and the buttons are absent for a source that cannot act, since ambient recognition can name a song playing through a wall and has no way to pause it.

The menu bar (tray on Windows) has the rest: Wrong lyrics? Find… to search every provider by hand and pick a different match; Lyrics earlier / later for a ±250 ms nudge saved per track, because each LRC is off by its own amount; Word-by-word highlight for karaoke-style enhanced LRC where a track carries the timings; Show romanized lyrics and Show translation for reading along with a song written in a script you don't read; text colour, size presets, and surrounding lines.

Lyric sources

Checked in order, first synced hit wins:

  1. Your own Artist - Title.lrc files, which beat every network source — this is how you fix a wrong match, not just a missing one.
    • macOS: ~/Music/Eavesdrop Lyrics/ (menu bar → Open lyrics folder…)
    • Android: Android/data/app.eavesdrop/files/Lyrics, wiped on uninstall
  2. LRCLIB — free, no auth, no key.
  3. NetEase Cloud Music — broad catalogue, no auth. Measured against 20 tracks it rescued none of LRCLIB's misses; kept because it only costs a request after LRCLIB has already failed. The local folder is the fallback that works.

Adding a source is one Provider impl in core/src/lyrics.rs.

Reading along

Lyrics in a script you can't read are just shapes, so a track can carry two companions to the words: a romanization and a translation, each a whole timestamped document of its own, aligned to the lyrics by time rather than by line number. That matters more than it sounds — the main track carries blank lines for instrumental gaps and, from NetEase, a block of production credits, and its companions carry neither, so matching on position would slide every line one earlier than the words it belongs to. Both toggles are off by default and each is drawn under the live line only.

Crucially the companions are looked up across the whole provider chain rather than asked of whichever source won the lyrics. LRCLIB answers first and has no romanization field at all, so tying the two together would leave the feature dead for exactly the tracks that need it. Instead each provider contributes whichever half it knows, and the walk stops as soon as everything switched on has been found.

Romanization comes from NetEase's romalrc — hand-checked, so it knows readings no algorithm can derive — which covers most Japanese and Korean releases. Where no source has one, core/src/romanize.rs works it out on the device for the scripts where romanisation is a function of the code point rather than of the vocabulary: Hangul (Revised Romanization), kana (Hepburn), Cyrillic (BGN/PCGN) and Greek. It declines a line it cannot convert in full, which is what keeps it honest — kanji have no algorithmic reading, so a line mixing kanji with kana gets nothing rather than a half-converted guess that reads as a rendering bug. Chinese is absent for the same reason: pinyin needs a dictionary, not an algorithm. Lines already in the Latin alphabet are skipped, so an English chorus inside a Korean song is never printed twice — and a track whose lyrics are entirely Latin never hits the network at all, which is what makes the toggle free for anyone listening in English.

Translation is the weaker half, and worth knowing why. The only automatic source is NetEase's tlyric, and it translates into Chinese whatever the original language, because it is a Chinese service serving Chinese listeners — measured across ten tracks in four languages, ten of ten came back in Chinese. There is no free, no-auth English lyric translation source to point at. So the route to one in a language you choose is a sidecar file: drop Artist - Title.translated.lrc (or .romanized.lrc) beside your own lyrics and it beats the network, the same doctrine the main lookup already follows.

Both companions are cached beside the lyrics, misses included — "nobody has romanized this song" is a permanent fact about most songs, and without remembering it every play would ask again.

On the widget and the always-on screen the extra lines are drawn outside the presets, in a strip reserved off the bottom that the preset then composes above. A preset is a sync mechanic, and there is no general way to hang two more rows off nine different ones without wrecking the composition each was designed as — Tunnel, whose lines fly at the camera and past it, has nowhere to put anything. So the surface reserves the space, the band is painted into it, and every preset keeps its design while there is one implementation to keep honest.

Permissions

  • macOS asks to control Spotify — reading playback position means sending Apple Events. Eavesdrop reads the track name and position only; it records no audio and sends your listening history nowhere.
  • Android needs notification access, and nothing else by default — getActiveSessions() is gated on being an enabled notification listener, and that is the only door Android provides to other apps' media sessions. No notification is ever read. Display over other apps is asked for only if you switch the floating overlay on. Neither is available as a runtime dialog; both are grants made in Settings.

Platform notes

Windows reads any player through SMTC, which forces two decisions the Mac build never has to make. Which player: GetCurrentSession means "whatever last grabbed the transport", routinely a browser tab that autoplayed a video, so Eavesdrop ranks every session itself — a playing Spotify wins outright, then any other player making sound, then a paused Spotify, then the rest. A YouTube singalong still works; it just never outbids Spotify. Tray → Follow Spotify only turns the rest off entirely. Which position: a session's Position is a stamp, not a live reading — it is where the track was at LastUpdatedTime, and Spotify restamps only every ~4.5s. Anchoring the core's clock on the raw value runs the lyrics late and makes the 5s drift poll yank them back each time it fires, so the position is extrapolated from its stamp, exactly as Windows' own media flyout does to keep its progress bar smooth.

It draws over normal windows and borderless/windowed-fullscreen games. Exclusive-fullscreen is out of scope — reaching it means hooking the game's present call via DLL injection, which trips anti-cheat. No taskbar button or Alt-Tab entry; the app lives in the tray.

Android leads with a home screen widget, not the overlay. The overlay is the port of the desktop idea, and a phone is where that idea is weakest: the screen is small, it is almost always showing something you chose, and floating text over it needs a permission the widget does not. So the widget is the default and the overlay is a toggle. Updates are pushed on every line change — the famous 30-minute updatePeriodMillis floor governs only the system waking a provider on a timer, not an app calling updateAppWidget, which is what makes per-line lyrics possible at all.

Presets

The widget and the always-on display are drawn by one of ten presets, picked in the app. A preset is a palette, a typeface and — the part that matters — a sync mechanic: how the surface shows the line arriving. Nine of them are distinct visual worlds; the tenth is plain, the look Eavesdrop shipped with, which is the default so that an update redecorates nobody's home screen.

Mechanic
Supercut hollow outlines slam to solid acid yellow on the beat
Ink italic serif inks in word by word
Terminal a timestamped log types itself in under a magenta caret
Overdrive letters punch in one at a time, RGB-split, over a drifting hue mesh
Halo no panel at all; everything but the live line falls out of focus
Knockout a vermilion bar wipes across the line and knocks the letters out
Read Head the whole song is one amber tape running sideways through a fixed head
Tunnel lines fly at the camera and blow past it
Marginalia riso-printed handwriting with the underline drawing itself

They are defined in core/src/style.rs for the same reason themes are — the core owns the ids and the palettes so they cannot drift between shells — and drawn in android/.../style/, because a canvas is a window and windows are not the core's business. Nothing in the core knows a font name; a preset asks for a role and the shell answers with what it ships. See docs/FONTS.md.

Three consequences worth knowing:

No preset paints a background. A widget sits on a wallpaper the user chose, and an opaque rectangle over it is a hole punched in their home screen — so the type and its mechanic go straight onto the wallpaper, which is what Halo was doing alone and is now the house rule. The presets keep their effects: the highlighter wipe, the read head and its LED grid, the drifting hue mesh, the riso misprint, the void that Tunnel's lines fly through. What they lose is the field behind them, and with it the contrast that field was providing — so the lyrics carry their own, the same doctrine the overlay already follows for the same reason. Light ink gets a dark halo and dark ink a light one, at an opacity that falls off with the square root of the text's own rather than linearly: a context line at a quarter alpha needs its carrier more than the live line does, since pale grey on a pale wallpaper is exactly where text disappears.

The widget is a bitmap, not a view tree. None of these is expressible in the dozen view types RemoteViews permits, so the picture is painted app-side and shipped across as one bitmap with real views laid over it for touch. The touch targets are positioned with setViewPadding, which is remotable where setViewLayoutHeight is API 31 against a minSdk of 26 — so a fixed-size button row can be pushed to wherever the painter put the buttons, and each preset keeps its own transport: a row in the middle of a wide panel with the track beside it behind a rule, a bottom sheet on a tall one, squares for the poster presets, circles for the soft ones, a gradient lozenge for Overdrive, brackets for Terminal, and no casing at all around Halo's. Only the button sizes are shared, because those are the touch targets and they live in XML. Tap the lyrics to reveal previous / play / next; tap them again to put it away. An idle widget opens the app instead, since there is nothing to transport.

A long lyric wraps; it does not shrink. Shrinking to keep one line on one row is what makes a widget set a whole sentence at nine points beside two context lines nobody is reading — the context is decoration and the live line is the product. So the live line takes the rows it needs and the surrounding lines are dropped to pay for them, down to none. Type size only gives way when wrapping cannot help: a single word wider than the panel, or a line still too long after three rows.

Six of the ten move within a line, and that is the only thing here that spends battery on decoration. A wipe crossing the words means redrawing the bitmap several times a second and handing it to the launcher each time, so it is gated on playing, screen on, widget placed and preset-actually-animated, and there is a switch. The other four — including Supercut, the loudest of the set — redraw only on line boundaries and cost exactly what the app cost before they existed. The always-on screen is unaffected either way: frames carry a timing anchor rather than a progress figure, so it animates at its own rate off a frame published once, when the line changed, and the core's scheduling-not-polling model survives intact.

There is no audio visualiser on any of them. Eavesdrop reads a media session's metadata and position and never receives audio, so a meter could only be a random number generator wearing a lab coat. Where the designs called for one, the space shows something actually known — how far through the track you are.

When the overlay is on it is two windows: the lyrics window is FLAG_NOT_TOUCHABLE permanently, and a small sibling window carries the drag handle. That removes the 12 Hz cursor poll both desktop shells need, since a click-through window there receives no mouse events at all.

Over a light background white text vanishes, and an app cannot read the pixels under its own overlay without requesting screen capture — so it carries its own contrast, drawn twice as a dark stroke and a fill on top, the way subtitles have always done it. A blurred drop shadow was tried first and is too diffuse to separate a white glyph from a white page.

No Kotlin Multiplatform. The Rust core is already the multiplatform layer and reaches more platforms than KMP does. What KMP would add is shared UI, and that is the one thing that cannot be shared — iOS has no system overlay, so an iOS port is a different product (Live Activities, Dynamic Island) rather than this one recompiled.

The always-on display is imitated, not used. There is no public API for the real one — neither Samsung nor Google opened that layer, and it is drawn by SystemUI with the panel in a low-power mode no app can enter. Android's own AOD refreshes at about 1 Hz with the SoC mostly asleep; ours is an ordinary fullscreen activity over the keyguard at minimum brightness, drifting to spare the panel. Measured on device it roughly doubles idle drain, so it is opt-in with a brightness slider, and the phone's own AOD should be switched off or the two compete.

Two things make it workable rather than the usual hack. It needs no new permission: launching an activity from the background is blocked on Android 10+ except for apps holding SYSTEM_ALERT_WINDOW, which the overlay already requires. And FLAG_KEEP_SCREEN_ON is a screen wake lock, not a partial one, so Play's excessive-partial-wake-lock metric does not apply — and Eavesdrop is sideloaded regardless.

The route that reaches the real AOD is publishing a MediaSession and putting the lyric in METADATA_KEY_TITLE, which OEM always-on displays render. It was built and measured: OxygenOS binds its AOD card to the app that owns the audio, so ours never won even at top session priority with matching flags. It may still work on Samsung, which reads that key directly. It also puts lyric fragments everywhere media metadata is consumed — car stereo included — so it was dropped.

ACTION_SCREEN_OFF carries no intent: the same broadcast means "the user pressed power" and "the system slept the display". The rule is therefore positional rather than semantic — if our screen was up when it arrived, screen-off ends it; otherwise it starts it. Getting that wrong is how the power button stops working.

Ambient recognition (experimental, macOS)

Lyrics for music you are not playing — a colleague's speakers, a café — via microphone identification. CompositeSource ranks sources by evidence quality and hands the overlay to the best one talking; promotion is instant, demotion waits out a grace period.

Off by default, and it uploads audio. ACRCloud takes the samples themselves rather than an on-device signature, so a few seconds of the room — voices included — goes to a third party per request. The source is isExpensive, so it only ever starts when no better source can answer; during your own music the engine does not run and the mic indicator stays dark. Requests are billed, so identification is sparse: every 10s while hunting, every 30s once a track is known, never on silence.

Recognition misses on quiet rooms, live versions and anything outside ACRCloud's catalogue, and even a correct match anchors from a play_offset_ms inferred through the air, so sync is looser than reading Spotify directly.

Credentials live outside the repo, in ~/Library/Application Support/app.eavesdrop/acrcloud.json:

{
  "host": "identify-<your-region>.acrcloud.com",
  "access_key": "…",
  "access_secret": "…"
}

The host is not boilerplate — a key is bound to its project's region, and a valid key sent to the wrong host fails as 3001 Missing/Invalid Access Key, which reads like a bad key and sends you checking the wrong thing.

ShazamKit would have been the better fit — free, unlimited, on-device signatures — but it needs a paid Apple developer team, without which every match returns error 202. The recogniser is confined to ACRCloudSource, so swapping back is one file.

Known gaps

  • Unsigned. Allow it in System Settings, or sign it yourself.
  • Android: overlay position/size persistence, text-size presets, focus mode and the search window are still to come. Text colour, the widget, the ten lyric presets and the imitated always-on display are in.
  • The presets are Android-only so far. The palettes and mechanics live in the core precisely so macOS and Windows can adopt them, but neither shell draws them yet.
  • Windows: animated line transitions, focus mode and the search window likewise.
  • Ambient recognition is macOS-only.
  • Translations are Chinese or your own. There is no automatic English translation, because no free no-auth source offers one — see Reading along. Romanization has no such gap.
  • The offline romanizer cannot read kanji, so a Japanese line mixing kanji with kana falls back to nothing unless NetEase has a romalrc for the track (it usually does). Chinese has no romanization path at all.

Licence

MIT — see LICENSE.

The ten typefaces the Android presets bundle are SIL Open Font License 1.1 and keep their own terms — see docs/FONTS.md.

Lyrics are fetched from LRCLIB and NetEase Cloud Music. Eavesdrop does not host, store or redistribute lyrics; it caches what you look up locally so the same track isn't fetched twice.

About

Time-synced lyrics that float over everything you do. A transparent, click-through macOS overlay for Spotify. Rust core, native Swift UI, no Electron.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages