Public customer feedback, turned into a product decision worth testing.

the biggest complaint
was not the best opportunity.

628 pieces of public Spotify feedback, 293 of them real complaints, and a single product question carried end to end: discovery, decision, prototype, experiment, spec.

something to read to
rashi raut
ai-assisted build
opportunity agent
jul / aug 2026
0%
collected 628 minus 54 not in English in english 574 minus 281 praise or noise complaints 293
what survived the filters, and what it cost

I built a pipeline that collects public Spotify feedback, groups similar complaints, remembers them across runs, scores them against a stated objective, and recommends what deserves attention.

Its top recommendation was free-tier friction since it had the most evidence by far, but I rejected it because that friction is intentional. It is part of how Spotify converts users to Premium.

Instead, I followed a smaller, much newer signal around trust in a catalogue increasingly filled with AI-generated music. The complaints were unusually negative, heavily engaged with on Reddit, showed up across sources, and people were already building their own workarounds.

From there, I ran the product loop backwards: prototype first, experiment second, spec last. The experiment uses simulated data, so it demonstrates the method and decision-making process rather than claiming to represent real user behavior.

628 records3 sources 8 stages39 tests several opinions revised
why

Hunger for learning

A friend who runs product at a large company described the system his org is building: an agent that reads every customer channel continuously, sizes the pain points, and hands the PM a ranked list, so the job becomes the deciding rather than the gathering. I could have read essays about that shift. Building it sounded like more fun, and building it is the only way to find out where it creaks.

Spotify was an easy choice. I use it constantly, I care a little too much about music, and I know the product well enough to notice when something feels off. Also, Spotify users complain publicly and enthusiastically. Very convenient for me.

app store reddit google play support, NPS, ... the agent group · size · rank a ranked list PM decides
the pattern I was trying to understand by building a small one
the journey

From 628 comments to one thing worth testing

01
Start with what people are already saying
july 13, then ongoing

I collected 628 (and counting) pieces of public Spotify feedback from Apple's App Store feed, Reddit, and Google Play. Then I automated it so the dataset can keep growing without adding the same review twice.

The first version of this project could have stopped at a giant spreadsheet. I didn't want a giant spreadsheet. A snapshot tells me what happened to be loud when I looked. A recurring pipeline gives me a chance to see whether something is actually changing.

The sources immediately disagreed. App Store users talked a lot about free-tier frustration. Reddit cared about stranger, more specific things. That was my first reminder not to confuse "most comments" with "most important problem."
fetch.py
app store aug
500
app store jul
100
reddit
25
google play
3
02
Make the volume slightly less chaotic
august 4 to 5

I built the clustering three times. The first version used keywords. It was fast and also perfectly capable of treating "I love the ads" and "I hate the ads" as the same thought. So I switched to semantic clustering with a small model running locally on my laptop.

Then came the less glamorous part: remove praise and noise, set non-English feedback aside, merge near-duplicate clusters, and flag giant catch-all groups instead of pretending they are meaningful insights. 281 records were praise or noise. 54 were non-English. I ended with 293 complaints.

I also added memory. Each problem gets a permanent ID and keeps its history across runs. Eventually, that means the system can tell me whether a complaint is shrinking, growing, or refusing to go away.

It is not perfect. One catch-all cluster swallowed 139 complaints, almost half the dataset, and the algorithm could not sensibly split it. I simply barred it from winning. The clusters also never got a proper human naming pass. They are still keyword signatures. This is a working system, not a magic trick.

The first memory design failed silently. Nothing crashed. The output just looked completely believable while being wrong. Much worse. I added a regression test.
analyze_semantic.py quality.py registry.py
03
Let the model make its case
august 5

Each cluster gets ranked on five things:

  • reach, how many people raised it
  • corroboration, how many sources raised it
  • severity, using engagement as a rough signal
  • OKR fit, how closely the complaint matches the objective I gave the system
  • confidence, whether the cluster is actually mostly negative

The weights are visible and editable in config. Nothing is pretending to be objective here. Change the objective or the weights and the ranking changes.

The winner was free-tier friction.
0.627 score
52 records
3 of 3 sources
~3x the mentions of AI content
The recommendation had the numbers backing it.

I still didn't take it. Ads and skip limits annoy people, but that does not automatically make them product defects. Some of that friction is the mechanism that makes Premium worth paying for. The system could tell me which problem looked painful. It could not tell me which pain Spotify would actually want to remove.

AI-content trust came second at 0.484 with only 18 records, 6.1% of the complaint set. Small signal. But 10 of those records were clearly negative and only 3 positive. Five of Reddit's top 25 posts were about it, with roughly 1,070 combined upvotes. People were already building their own AI filters. And unlike ads, recommendations, or free-tier complaints, this did not feel like the same argument Spotify users had been having for ten years. It was new enough to be worth looking at properly.

The ranking was useful because I had to explain why I was not following it.
decide.py decision memo
how each score was built reach corroboration severity okr fit confidence free-tier friction 0.627 AI content 0.484 recommendations 0.416 free-tier friction wins on corroboration and volume. AI content wins on reaction strength, and on being cleanly negative.
segment width = factor score × its weight · totals are the ranking
04
Six days later, Spotify did something interesting
august 11 to 18

Six days after the decision memo, Spotify announced AI Persona badges. Starting in September, AI-generated artist identities would be labeled and excluded from editorial and algorithmic recommendations by default, unless the listener already follows that artist.

To be extremely clear: I did not predict Spotify's launch. They had obviously been working on it long before I decided to spend my August reading angry App Store reviews. What interested me was that the concern had already become visible in public customer feedback before Spotify's response became public.

So I looked at what everyone else was doing. Deezer and Tidal already give listeners some control over AI-generated music. Deezer says 44% of new uploads to its platform are now fully AI-generated. In its research with Ipsos, 45% of streaming users said they wanted the ability to filter AI music.

Spotify's badge addresses one part of the problem: is this artist identity AI-generated? It does not fully answer a different question: was this music generated by AI, and do I want it in my recommendations? That became the product gap I explored.

I checked every external claim against primary sources. Two claims became less dramatic after I checked them. Kept the boring accurate versions.
landscape brief
aug 5 my decision memo mar SongDNA apr Verified by Spotify jul RIAA + IFPI two-tier labels jul · Tidal demonetizes AI aug 11 Spotify announces AI Persona badges sept badges appear 6 days
I did not predict this. The signal was public before the response was.
spotify deezer tidal this proposal labels AI content identity track track consumes listener toggle none yes yes yes blocks playback no no yes no demonetizes no partly yes out of scope
the gap is row two, and it is the only row I set out to fill
05
Build it before I get too attached to it
august 18

I built three interactive flows: a setting that hides fully AI-generated music, search results that explicitly say "3 AI-generated tracks hidden" with an option to show them, and a Now Playing flow that notices repeated skips of AI-generated tracks and offers the filter once.

I deliberately limited the filter to music labeled "AI Generated," not "AI Assisted." The industry already distinguishes between the two. Filtering everything AI-assisted would catch human artists who happened to use AI somewhere in production.

I also chose hide over block. If someone else sends you a playlist containing an AI track, I do not think your personal preference should make their song unplayable. Hide it from your experience. Let you reveal it if you want.

My first trigger required someone to skip three AI tracks in a row. The prototype queue alternated AI and human tracks. So the trigger literally could not happen. This is why I'm glad I did not put it in a PRD first.
open full size
the actual prototype, running · all artists fictional
06
Then I made up 30,000 users
august 18
Synthetic experiment. Method demonstration, not user evidence.

This part needs a giant asterisk. I built a simulated A/B test with 30,000 synthetic users across three arms.

The behavioral assumptions are invented. Baseline retention: invented. AI-averse versus indifferent versus AI-positive segments: invented. Expected effects: invented. Toggle adoption: also invented. I am not using fake users to claim the feature works. What I wanted to demonstrate was how I would structure the decision before real data existed.

The sample size is the one principled number. At a 60% baseline, detecting a +2 percentage-point change with 80% power at alpha 0.05 requires 9,333 users per arm. I used 10,000. Before running it, I also locked the rules: +2.0 points to ship, minimum 5% adoption.

Off by default produced +1.07 points, p=0.12. It failed. Only 8.7% of those synthetic users ever found the toggle. On by default produced +2.34 points, p=0.0007, and cleared the synthetic threshold. Same feature. Different default. Completely different outcome.

The feature was not really the experiment. The default was.
test design experiment.py
off by default
+1.07
on by default
+2.34
same feature, same effect per person, different reach off by default 10,000 in the arm 870 ever found the toggle · 8.7% +1.07 pts · p = 0.12 · does not clear +2.0 on by default 10,000 in the arm 8,500 kept it on · 85% +2.34 pts · p = 0.0007
synthetic users · the dilution is the finding, not the lift
07
Make it possible to prove myself wrong later
ongoing

The experiment result gets written back to the AI-content cluster's history, and the feedback pipeline can keep running every week. If Spotify's September rollout genuinely changes how users feel about AI content, I should eventually see the complaint line move.

I have not run that second real cycle yet. So right now, I have built the mechanism for the feedback loop. I have not demonstrated the loop with real before-and-after evidence.

There is another limitation too. Even if complaints fall, I will not have a clean counterfactual proving Spotify's change caused it. And there is a slightly uncomfortable AI problem here: if I keep feeding my own overrides back into the system, eventually I can build a model that is extremely good at agreeing with me. Not exactly the goal.

Memory is useful. A system that only learns my preferences is not the same thing as a system that learns what is true.
registry.py
collect cluster decide build + test measure arrives weeks after the decision was made the option I did not pick never built, so never measured
the two places this loop does not actually close
08
The PRD came last. That was not the original plan.
august 18

Originally, I was going to do this in the respectable order: discovery, PRD, design, build. Halfway through, I changed my mind and decided to push the PRD to the end.

This was not some grand methodology I had planned from day one. I heard the CPO of Webflow talk about building backwards in a podcast and wanted to try it out. I just became increasingly convinced that writing the spec too early would make me prematurely certain about a product I had not even tried to interact with yet.

The decision memo and experiment design were already written and dated, so I still had a record of what I believed before the later work could conveniently rewrite the story.

By the time I wrote the PRD, it contained what had survived the process instead of what I had guessed at the beginning:

  • the feature
  • the segments
  • eight explicit product decisions
  • the option that lost each one, and why it lost
  • the simulated experiment
  • the rollout plan
  • and the questions I still cannot answer
the prd live dashboard
judgment calls

The part AI could not do for me

AI-content trustover free-tier friction, the model's pick
Free-tier friction had much stronger evidence. It was also partly intentional. The model could rank the pain. I had to decide whether Spotify would actually want to remove it.
A user toggleover labels only
A label tells me what something is. A control lets me decide what happens next.
Hide, don't blockover unplayable tracks
My preference should change my experience, not break a shared playlist for someone else.
Fully-AI onlyover anything AI-touched
AI Generated and AI Assisted are already treated as different categories. I did not want a blunt filter hiding human artists because AI touched one part of production.
Platform labelsover crowd tagging
Letting users label artists sounds democratic until people coordinate to falsely tag an artist they dislike. That can become a harassment tool very quickly.
Off by defaultthen test the flip
My simulated result preferred on by default. My simulated result is also built on assumptions I invented. That makes it a reason to run a real test, not a reason to ship.
Say what's hiddenover silently cleaner results
If the product removes something from what I see, I want it to tell me. Quietly changing the list is cleaner UI and worse trust.
Spec lastover spec first
I changed the process halfway through. The prototype immediately found a logic error that the PRD would probably have documented very confidently.

The PRD keeps the rejected options too. Decisions are much easier to understand when you can see what they beat.

fine print

What I am not claiming

next

The next version needs actual people

Public feedback got me to a product question I can defend. It did not get me to a product I would ship.

With real access, I would sit with listeners first and figure out what they actually mean when they say they do not want AI music. Is the problem the music itself? Deception? Artist identity? Recommendation quality? Consent? Something else entirely?

Then I would test whether Spotify's labels are accurate enough to power a filter, instrument whether people can actually find and understand the control, and run the default experiment with real behavior.

The feature is not ready. The question is.

the split

What AI did. What I still had to do.

AI helped me write code, structure analysis, draft documents, debug things, research the market, and move much faster than I could have alone. The pipeline did the repetitive work: collect, group, score, remember.

I chose Spotify. I chose what deserved another look. I changed the process halfway through. I rejected product directions. I set the constraints. I decided what the product should and should not do.

AI drafted a lot of this project with my direction. I own the decisions.

And I wrote them down so someone else can tell me where I'm wrong.

the whole thing, in one line 01collect 02cluster 03 decide 04research 05prototype 06test 07loop 08spec fetch.py registry memo brief prototype readout history PRD
one decision, carried the whole way