628 pieces of public Spotify feedback, 293 of them real complaints, and a single product question carried end to end: discovery, decision, prototype, experiment, spec.
I built a pipeline that collects public Spotify feedback, groups similar complaints, remembers them across runs, scores them against a stated objective, and recommends what deserves attention.
Its top recommendation was free-tier friction since it had the most evidence by far, but I rejected it because that friction is intentional. It is part of how Spotify converts users to Premium.
Instead, I followed a smaller, much newer signal around trust in a catalogue increasingly filled with AI-generated music. The complaints were unusually negative, heavily engaged with on Reddit, showed up across sources, and people were already building their own workarounds.
From there, I ran the product loop backwards: prototype first, experiment second, spec last. The experiment uses simulated data, so it demonstrates the method and decision-making process rather than claiming to represent real user behavior.
A friend who runs product at a large company described the system his org is building: an agent that reads every customer channel continuously, sizes the pain points, and hands the PM a ranked list, so the job becomes the deciding rather than the gathering. I could have read essays about that shift. Building it sounded like more fun, and building it is the only way to find out where it creaks.
Spotify was an easy choice. I use it constantly, I care a little too much about music, and I know the product well enough to notice when something feels off. Also, Spotify users complain publicly and enthusiastically. Very convenient for me.
I collected 628 (and counting) pieces of public Spotify feedback from Apple's App Store feed, Reddit, and Google Play. Then I automated it so the dataset can keep growing without adding the same review twice.
The first version of this project could have stopped at a giant spreadsheet. I didn't want a giant spreadsheet. A snapshot tells me what happened to be loud when I looked. A recurring pipeline gives me a chance to see whether something is actually changing.
I built the clustering three times. The first version used keywords. It was fast and also perfectly capable of treating "I love the ads" and "I hate the ads" as the same thought. So I switched to semantic clustering with a small model running locally on my laptop.
Then came the less glamorous part: remove praise and noise, set non-English feedback aside, merge near-duplicate clusters, and flag giant catch-all groups instead of pretending they are meaningful insights. 281 records were praise or noise. 54 were non-English. I ended with 293 complaints.
I also added memory. Each problem gets a permanent ID and keeps its history across runs. Eventually, that means the system can tell me whether a complaint is shrinking, growing, or refusing to go away.
It is not perfect. One catch-all cluster swallowed 139 complaints, almost half the dataset, and the algorithm could not sensibly split it. I simply barred it from winning. The clusters also never got a proper human naming pass. They are still keyword signatures. This is a working system, not a magic trick.
Each cluster gets ranked on five things:
The weights are visible and editable in config. Nothing is pretending to be objective here. Change the objective or the weights and the ranking changes.
I still didn't take it. Ads and skip limits annoy people, but that does not automatically make them product defects. Some of that friction is the mechanism that makes Premium worth paying for. The system could tell me which problem looked painful. It could not tell me which pain Spotify would actually want to remove.
AI-content trust came second at 0.484 with only 18 records, 6.1% of the complaint set. Small signal. But 10 of those records were clearly negative and only 3 positive. Five of Reddit's top 25 posts were about it, with roughly 1,070 combined upvotes. People were already building their own AI filters. And unlike ads, recommendations, or free-tier complaints, this did not feel like the same argument Spotify users had been having for ten years. It was new enough to be worth looking at properly.
Six days after the decision memo, Spotify announced AI Persona badges. Starting in September, AI-generated artist identities would be labeled and excluded from editorial and algorithmic recommendations by default, unless the listener already follows that artist.
To be extremely clear: I did not predict Spotify's launch. They had obviously been working on it long before I decided to spend my August reading angry App Store reviews. What interested me was that the concern had already become visible in public customer feedback before Spotify's response became public.
So I looked at what everyone else was doing. Deezer and Tidal already give listeners some control over AI-generated music. Deezer says 44% of new uploads to its platform are now fully AI-generated. In its research with Ipsos, 45% of streaming users said they wanted the ability to filter AI music.
Spotify's badge addresses one part of the problem: is this artist identity AI-generated? It does not fully answer a different question: was this music generated by AI, and do I want it in my recommendations? That became the product gap I explored.
I built three interactive flows: a setting that hides fully AI-generated music, search results that explicitly say "3 AI-generated tracks hidden" with an option to show them, and a Now Playing flow that notices repeated skips of AI-generated tracks and offers the filter once.
I deliberately limited the filter to music labeled "AI Generated," not "AI Assisted." The industry already distinguishes between the two. Filtering everything AI-assisted would catch human artists who happened to use AI somewhere in production.
I also chose hide over block. If someone else sends you a playlist containing an AI track, I do not think your personal preference should make their song unplayable. Hide it from your experience. Let you reveal it if you want.
This part needs a giant asterisk. I built a simulated A/B test with 30,000 synthetic users across three arms.
The behavioral assumptions are invented. Baseline retention: invented. AI-averse versus indifferent versus AI-positive segments: invented. Expected effects: invented. Toggle adoption: also invented. I am not using fake users to claim the feature works. What I wanted to demonstrate was how I would structure the decision before real data existed.
The sample size is the one principled number. At a 60% baseline, detecting a +2 percentage-point change with 80% power at alpha 0.05 requires 9,333 users per arm. I used 10,000. Before running it, I also locked the rules: +2.0 points to ship, minimum 5% adoption.
Off by default produced +1.07 points, p=0.12. It failed. Only 8.7% of those synthetic users ever found the toggle. On by default produced +2.34 points, p=0.0007, and cleared the synthetic threshold. Same feature. Different default. Completely different outcome.
The experiment result gets written back to the AI-content cluster's history, and the feedback pipeline can keep running every week. If Spotify's September rollout genuinely changes how users feel about AI content, I should eventually see the complaint line move.
I have not run that second real cycle yet. So right now, I have built the mechanism for the feedback loop. I have not demonstrated the loop with real before-and-after evidence.
There is another limitation too. Even if complaints fall, I will not have a clean counterfactual proving Spotify's change caused it. And there is a slightly uncomfortable AI problem here: if I keep feeding my own overrides back into the system, eventually I can build a model that is extremely good at agreeing with me. Not exactly the goal.
Originally, I was going to do this in the respectable order: discovery, PRD, design, build. Halfway through, I changed my mind and decided to push the PRD to the end.
This was not some grand methodology I had planned from day one. I heard the CPO of Webflow talk about building backwards in a podcast and wanted to try it out. I just became increasingly convinced that writing the spec too early would make me prematurely certain about a product I had not even tried to interact with yet.
The decision memo and experiment design were already written and dated, so I still had a record of what I believed before the later work could conveniently rewrite the story.
By the time I wrote the PRD, it contained what had survived the process instead of what I had guessed at the beginning:
The PRD keeps the rejected options too. Decisions are much easier to understand when you can see what they beat.
Public feedback got me to a product question I can defend. It did not get me to a product I would ship.
With real access, I would sit with listeners first and figure out what they actually mean when they say they do not want AI music. Is the problem the music itself? Deception? Artist identity? Recommendation quality? Consent? Something else entirely?
Then I would test whether Spotify's labels are accurate enough to power a filter, instrument whether people can actually find and understand the control, and run the default experiment with real behavior.
The feature is not ready. The question is.
AI helped me write code, structure analysis, draft documents, debug things, research the market, and move much faster than I could have alone. The pipeline did the repetitive work: collect, group, score, remember.
I chose Spotify. I chose what deserved another look. I changed the process halfway through. I rejected product directions. I set the constraints. I decided what the product should and should not do.
AI drafted a lot of this project with my direction. I own the decisions.
And I wrote them down so someone else can tell me where I'm wrong.