self-hosted-spotify

selfhostedmusicnavidromesoulseekbeetsmusicbrainzproxmoxhomelab

Write here. Set draft: false when it’s ready to go live.

What it took to go from a Spotify data export to a self-hosted library with my playlists reproduced 1-to-1 — and, more usefully, the dozen gotchas that ate the most time. Skip to Learnings if you just want the landmines.

The goal

Own my music instead of renting it. Spotify lets you export your data but never the audio, so the plan was: take the export, figure out what I actually listen to, re-acquire the files, tag them properly, and serve them myself — with my playlists intact.

The stack:

Component Role
Navidrome Music server + web/mobile player (Subsonic API)
Bandcamp/ripped CDs Source of music files
beets Tagging/organizing; AcoustID fingerprinting → MusicBrainz IDs
MusicBrainz Canonical track identity (recording MBIDs)
Picard Manual tag editing over SMB

Everything keys on MusicBrainz recording MBIDs where possible; the Spotify data keys on spotify:track: URIs, so the first job is bridging the two.

The pipeline

Spotify export (JSON)
   │  parse listening history + saved/playlists
   ▼
resolve URIs → MusicBrainz MBIDs           (a resolver script)
   │  export "Artist,Title,Album" seed CSVs
   ▼
Bandcamp/ripped CDs - aquire the files  
   │
beets import → tags + organizes into /Artist/Title
   │
generate .m3u playlists → Navidrome imports them
   ▼
play anywhere

Four small Python scripts glue it together (standard-library where possible, each resumable and idempotent):

  • resolver — Spotify URIs → MBIDs, and exports download seeds.
  • music source — reads a seed CSV, searches for the music file, enqueues the best match.
  • seed builder — turns a plain text list, an old .m3u, or an Exportify CSV into a seed CSV.
  • m3u builder — matches a seed against the beets library and writes a Navidrome playlist.

Plus a pl.sh wrapper that runs the whole per-playlist flow so I stop forgetting steps.


Learnings — the stuff that cost me evenings

This is the part worth reading. Each one is a symptom → cause → fix.

1. Spotify’s Web API now requires the app owner to have Premium

The clean way to get a track’s ISRC (and thus its MBID) is the Spotify Web API. On a free account you get 403: "Active premium subscription required for the owner of the app." — it applies to the whole API for that app, including playlist reads. So the entire “use Spotify’s API” plan is dead for free users.

Fix: resolve against MusicBrainz only — two passes per track: an exact Spotify-URL relationship lookup (precise, ~20% coverage), then a fuzzy artist+title recording search gated by MusicBrainz’s relevance score. Flag the ambiguous ones; never fabricate a match.

2. exFAT + Navidrome’s quick scan = new files invisible

New albums wouldn’t show up in Navidrome. Cause: the library drive is exFAT, whose timestamps are coarse, so Navidrome’s quick scan (which keys off folder mtime) silently skips freshly-added folders.

Fix: touch the new folders to bump their mtime, then restart Navidrome (or force a full scan). On a fresh beet import with move: yes the mtime updates anyway, so it mostly self-heals — but when in doubt, full scan.

3. Navidrome trims playlists to indexed tracks — and writes that back

I generated a playlist, Navidrome showed 1 of 4 tracks, and then the .m3u file itself got rewritten down to that 1 track. Navidrome imports .m3us from its PlaylistsPath, keeps only tracks already in its library, and (with AutoImportPlaylists) syncs the trimmed version back to disk.

Fixes (all three matter):

  • Generate playlists after importing the tracks, or they get permanently trimmed.
  • Put .m3us in the configured PlaylistsPath directory.
  • Use absolute paths in the .m3u when the tagger and Navidrome mount the library at different paths (common across containers) — relative paths resolve against the wrong root.

4. beets quirks (several)

  • A malformed file hard-crashes the whole import. chromaprint throws Assertion 'length % m_num_channels == 0' failed. Aborted (SIGABRT) on some files and takes the entire run down. Use beet import -A (no autotag/no fingerprint) for bulk imports — it can’t hit the crash. This content rarely matches AcoustID anyway.
  • Quiet mode declines fuzzy matches. beet import -q only auto-applies strong matches; for comp/remix/edit-heavy libraries that’s near-zero, so set quiet_fallback: asis or everything gets “Skipping.” and nothing imports.
  • -A relies on existing tags. As-is import builds the destination path from the file’s own tags, so untagged files don’t sort into artist folders. Tag them first (Picard) or import without -A so beets fetches tags.
  • chroma needs its deps in the right venv. Installed via pipx? pip install pyacoustid goes to the wrong environment — use pipx inject beets pyacoustid. Also needs ffmpeg (decoder) and libchromaprint-tools (fpcalc).
  • Config edits silently ignored. A plugin block does nothing unless the plugin is in plugins:; a duplicate plugins: key means YAML keeps only the last. Always verify with beet version / beet config.

5. The df-vs-du illusion, and 1.5 TB of orphaned disk images

Inside a container, df reported the drive 98% full but du of the music folder found only 157 GB. df reports the whole underlying filesystem even when only a subdirectory is bind-mounted into the container. The “missing” space was elsewhere on the host.

Chasing it on the Proxmox host turned up two preallocated, never-formatted, unreferenced VM disk images (vm-105-disk-*.raw) eating ~1.5 TB. Confirmed orphaned (pct config didn’t list them; file -s/lsblk showed no filesystem), then reclaimed via pct rescan → remove the unused disks in the UI. Entire space problem gone without deleting a single song.

Lesson: when df and du disagree by a lot, the space is on the host, not in the container — go look there before you prune anything you care about.

6. exFAT ownership + container privilege for writes

exFAT mounts with a fixed uid (mine: 1001) and only that uid can write. For a Samba share so Picard can edit tags, the volume being read-write isn’t enough — Samba must write as uid 1001. Create the SMB user with that uid (-u "music;pass;1001;music;1001"). This works because the container is privileged (uid maps straight through); in an unprivileged container it wouldn’t line up.

7. Picard will desync your library if you let it move files

Picard editing tags is fine; Picard renaming/moving files is not — beets’ database still points at the old paths, which produces dead entries that leak into playlists. Keep Picard in tag-only mode (disable move/rename), and run beet update afterward so beets re-reads the new tags. Let one tool own the file layout — beets — not both.

8. Exportify, and the seed-format trap

Playlists not in your Spotify export: use Exportify (OAuth with your own login — no Premium needed) to export any playlist to CSV. But the raw Exportify CSV is not the pipeline’s format — its columns are Track Name / Artist Name(s), while my scripts expect Artist / Title. Feeding the raw file straight in gives 0 matches and silently downloads nothing. Always run it through the seed-builder (which maps the columns) first. This bit me twice.

9. Little things that still cost minutes

  • PowerShell vs bash: $env:VAR="x" on Windows, export VAR=x on Linux. Pasting the wrong one is a daily occurrence across machines.
  • Windows os.replace flakiness: antivirus/indexer briefly locks a file and atomic writes fail with WinError 5; retry the replace, or run the script on the Linux box.
  • Copy the file to the right machine. Half my “No such file or directory” errors were a CSV sitting on the wrong host. The orchestration script helps, but getting the file onto the box is the one manual step no script removes.

Design principles that paid off

Small rules that made the whole thing debuggable and safe to re-run:

  • Treat the source export as read-only. Never mutate the original data.
  • Resumable + idempotent. Every long step keeps state in a TSV with a status column and skips completed work on re-run. Days-long jobs will get interrupted.
  • Flag, don’t guess. An explicit ambiguous / no_match status beats a fabricated result every time — especially for a music library you’ll live with.
  • Standard library first, one self-test per script. No framework, just assert-based checks so a refactor can’t silently break the matcher.
  • One tool owns each concern. beets owns the file layout; Navidrome owns serving; the scripts own acquisition. Two tools managing the same files is how you get phantom entries.

The operational loop

Once it’s all wired up, adding or refreshing a playlist is:

1. Export the playlist (Exportify) → convert to a seed CSV
2. Source the files matching items from step 1
3. beet import -A -s   (tag + organize)
4. beet update         (pick up any manual Picard edits)
5. generate the .m3u into Navidrome's PlaylistsPath
6. rescan Navidrome

Steps 3–6 are re-runnable as downloads trickle in, so playlists fill up over time. I wrapped it in a two-command script (pl.sh add / pl.sh build) because forgetting step 1 (the seed conversion) cost me the same debugging session three times.

Would I do it again?

Yes — but with the learnings above, in an afternoon instead of a week.

The payoff: my library, my files, my playlists, served from a box I own — and a very thorough understanding of exFAT timestamps I never wanted.

← All posts