self-hosted-spotify
Write here. Set draft: false when it’s ready to go live.
What it took to go from a Spotify data export to a self-hosted library with my playlists reproduced 1-to-1 — and, more usefully, the dozen gotchas that ate the most time. Skip to Learnings if you just want the landmines.
The goal
Own my music instead of renting it. Spotify lets you export your data but never the audio, so the plan was: take the export, figure out what I actually listen to, re-acquire the files, tag them properly, and serve them myself — with my playlists intact.
The stack:
| Component | Role |
|---|---|
| Navidrome | Music server + web/mobile player (Subsonic API) |
| Bandcamp/ripped CDs | Source of music files |
| beets | Tagging/organizing; AcoustID fingerprinting → MusicBrainz IDs |
| MusicBrainz | Canonical track identity (recording MBIDs) |
| Picard | Manual tag editing over SMB |
Everything keys on MusicBrainz recording MBIDs where possible; the Spotify
data keys on spotify:track: URIs, so the first job is bridging the two.
The pipeline
Spotify export (JSON)
│ parse listening history + saved/playlists
▼
resolve URIs → MusicBrainz MBIDs (a resolver script)
│ export "Artist,Title,Album" seed CSVs
▼
Bandcamp/ripped CDs - aquire the files
│
beets import → tags + organizes into /Artist/Title
│
generate .m3u playlists → Navidrome imports them
▼
play anywhere
Four small Python scripts glue it together (standard-library where possible, each resumable and idempotent):
- resolver — Spotify URIs → MBIDs, and exports download seeds.
- music source — reads a seed CSV, searches for the music file, enqueues the best match.
- seed builder — turns a plain text list, an old
.m3u, or an Exportify CSV into a seed CSV. - m3u builder — matches a seed against the beets library and writes a Navidrome playlist.
Plus a pl.sh wrapper that runs the whole per-playlist flow so I stop forgetting
steps.
Learnings — the stuff that cost me evenings
This is the part worth reading. Each one is a symptom → cause → fix.
1. Spotify’s Web API now requires the app owner to have Premium
The clean way to get a track’s ISRC (and thus its MBID) is the Spotify Web API.
On a free account you get 403: "Active premium subscription required for the owner of the app." — it applies to the whole API for that app, including
playlist reads. So the entire “use Spotify’s API” plan is dead for free users.
Fix: resolve against MusicBrainz only — two passes per track: an exact Spotify-URL relationship lookup (precise, ~20% coverage), then a fuzzy artist+title recording search gated by MusicBrainz’s relevance score. Flag the ambiguous ones; never fabricate a match.
2. exFAT + Navidrome’s quick scan = new files invisible
New albums wouldn’t show up in Navidrome. Cause: the library drive is exFAT, whose timestamps are coarse, so Navidrome’s quick scan (which keys off folder mtime) silently skips freshly-added folders.
Fix: touch the new folders to bump their mtime, then restart Navidrome (or
force a full scan). On a fresh beet import with move: yes the mtime updates
anyway, so it mostly self-heals — but when in doubt, full scan.
3. Navidrome trims playlists to indexed tracks — and writes that back
I generated a playlist, Navidrome showed 1 of 4 tracks, and then the .m3u
file itself got rewritten down to that 1 track. Navidrome imports .m3us from
its PlaylistsPath, keeps only tracks already in its library, and (with
AutoImportPlaylists) syncs the trimmed version back to disk.
Fixes (all three matter):
- Generate playlists after importing the tracks, or they get permanently trimmed.
- Put
.m3us in the configuredPlaylistsPathdirectory. - Use absolute paths in the
.m3uwhen the tagger and Navidrome mount the library at different paths (common across containers) — relative paths resolve against the wrong root.
4. beets quirks (several)
- A malformed file hard-crashes the whole import. chromaprint throws
Assertion 'length % m_num_channels == 0' failed. Aborted(SIGABRT) on some files and takes the entire run down. Usebeet import -A(no autotag/no fingerprint) for bulk imports — it can’t hit the crash. This content rarely matches AcoustID anyway. - Quiet mode declines fuzzy matches.
beet import -qonly auto-applies strong matches; for comp/remix/edit-heavy libraries that’s near-zero, so setquiet_fallback: asisor everything gets “Skipping.” and nothing imports. -Arelies on existing tags. As-is import builds the destination path from the file’s own tags, so untagged files don’t sort into artist folders. Tag them first (Picard) or import without-Aso beets fetches tags.- chroma needs its deps in the right venv. Installed via
pipx?pip install pyacoustidgoes to the wrong environment — usepipx inject beets pyacoustid. Also needsffmpeg(decoder) andlibchromaprint-tools(fpcalc). - Config edits silently ignored. A plugin block does nothing unless the
plugin is in
plugins:; a duplicateplugins:key means YAML keeps only the last. Always verify withbeet version/beet config.
5. The df-vs-du illusion, and 1.5 TB of orphaned disk images
Inside a container, df reported the drive 98% full but du of the music
folder found only 157 GB. df reports the whole underlying filesystem even
when only a subdirectory is bind-mounted into the container. The “missing”
space was elsewhere on the host.
Chasing it on the Proxmox host turned up two preallocated, never-formatted,
unreferenced VM disk images (vm-105-disk-*.raw) eating ~1.5 TB. Confirmed
orphaned (pct config didn’t list them; file -s/lsblk showed no filesystem),
then reclaimed via pct rescan → remove the unused disks in the UI. Entire
space problem gone without deleting a single song.
Lesson: when df and du disagree by a lot, the space is on the host, not in
the container — go look there before you prune anything you care about.
6. exFAT ownership + container privilege for writes
exFAT mounts with a fixed uid (mine: 1001) and only that uid can write. For a
Samba share so Picard can edit tags, the volume being read-write isn’t
enough — Samba must write as uid 1001. Create the SMB user with that uid
(-u "music;pass;1001;music;1001"). This works because the container is
privileged (uid maps straight through); in an unprivileged container it
wouldn’t line up.
7. Picard will desync your library if you let it move files
Picard editing tags is fine; Picard renaming/moving files is not — beets’
database still points at the old paths, which produces dead entries that leak
into playlists. Keep Picard in tag-only mode (disable move/rename), and run
beet update afterward so beets re-reads the new tags. Let one tool own
the file layout — beets — not both.
8. Exportify, and the seed-format trap
Playlists not in your Spotify export: use Exportify (OAuth with your own
login — no Premium needed) to export any playlist to CSV. But the raw Exportify
CSV is not the pipeline’s format — its columns are Track Name / Artist Name(s), while my scripts expect Artist / Title. Feeding the raw file
straight in gives 0 matches and silently downloads nothing. Always run it
through the seed-builder (which maps the columns) first. This bit me twice.
9. Little things that still cost minutes
- PowerShell vs bash:
$env:VAR="x"on Windows,export VAR=xon Linux. Pasting the wrong one is a daily occurrence across machines. - Windows
os.replaceflakiness: antivirus/indexer briefly locks a file and atomic writes fail withWinError 5; retry the replace, or run the script on the Linux box. - Copy the file to the right machine. Half my “No such file or directory” errors were a CSV sitting on the wrong host. The orchestration script helps, but getting the file onto the box is the one manual step no script removes.
Design principles that paid off
Small rules that made the whole thing debuggable and safe to re-run:
- Treat the source export as read-only. Never mutate the original data.
- Resumable + idempotent. Every long step keeps state in a TSV with a status column and skips completed work on re-run. Days-long jobs will get interrupted.
- Flag, don’t guess. An explicit
ambiguous/no_matchstatus beats a fabricated result every time — especially for a music library you’ll live with. - Standard library first, one self-test per script. No framework, just
assert-based checks so a refactor can’t silently break the matcher. - One tool owns each concern. beets owns the file layout; Navidrome owns serving; the scripts own acquisition. Two tools managing the same files is how you get phantom entries.
The operational loop
Once it’s all wired up, adding or refreshing a playlist is:
1. Export the playlist (Exportify) → convert to a seed CSV
2. Source the files matching items from step 1
3. beet import -A -s (tag + organize)
4. beet update (pick up any manual Picard edits)
5. generate the .m3u into Navidrome's PlaylistsPath
6. rescan Navidrome
Steps 3–6 are re-runnable as downloads trickle in, so playlists fill up over
time. I wrapped it in a two-command script (pl.sh add / pl.sh build) because
forgetting step 1 (the seed conversion) cost me the same debugging session three
times.
Would I do it again?
Yes — but with the learnings above, in an afternoon instead of a week.
The payoff: my library, my files, my playlists, served from a box I own — and a very thorough understanding of exFAT timestamps I never wanted.