Skip to main content

Turn Support Noise Into a Ranked Experiment Backlog

The fastest way to fix experiment backlog management is to stop debating and start counting. Your support queue, app-store reviews, and research notes already contain a ranked list of what to test next — you just haven't totaled it up. Cluster that text into themes, count how many times customers in

A
Atticus LiApplied Experimentation Lead at NRG Energy (Fortune 150) · Creator of the PRISM Method
9 min read

Editorial disclosure

This article lives on the canonical GrowthLayer blog path for indexing consistency. Review rules, sourcing rules, and update rules are documented in our editorial policy and methodology.

Fortune 150 experimentation lead100+ experiments / yearCreator of the PRISM Method
A/B TestingExperimentation StrategyStatistical MethodsCRO MethodologyExperimentation at Scale

Key takeaways

  • •Your experiment backlog already exists in your support tickets and reviews — ranking it is a counting exercise, not a research project
  • •Mention counts are observations, not estimates, which makes them harder to argue with than any scoring framework
  • •Cluster text into pain-point themes, count mentions per theme per quarter, and rank by raw volume before applying any judgment
  • •Convert top themes into five-field hypothesis cards — pain point, segment, intervention, metric, expected range — and groom them in the same queue as every other idea
  • •In our corpus, the top theme held 400+ mentions and the second nearly 300 — and the resulting order differed materially from the team's intuition
  • •Re-count quarterly so the backlog tracks reality instead of memory, and store counts, cards, and outcomes where they outlive any single analyst

Turn Support Noise Into a Ranked Experiment Backlog

The fastest way to fix experiment backlog management is to stop debating and start counting. Your support queue, app-store reviews, and research notes already contain a ranked list of what to test next — you just haven't totaled it up. Cluster that text into themes, count how many times customers independently mention each one, and convert the top themes into structured hypothesis cards. The ranking that falls out is cheap, legible to any stakeholder, and grounded in evidence instead of whoever argued loudest in grooming.

Mention-frequency ranking is the practice of ordering an experiment backlog by how often customers independently raise each pain point across support tickets, reviews, and research notes.

I run a testing program north of a hundred experiments a year, and this counting exercise reorders the backlog every single time we do it. Not tweaks it — reorders it. Below is the exact pipeline: where to pull the text, how to cluster it, how to count it, and how to turn counts into test-ready hypotheses that survive prioritization scrutiny.

Why does counting mentions beat team intuition?

Because intuition is a memory exercise, and memory is biased toward the recent, the vivid, and the internally political. The ticket a VP forwarded last week feels urgent. The four hundred tickets nobody forwarded feel invisible.

Scoring frameworks don't fix this — they formalize it. An ICE or PIE score asks you to estimate impact and confidence, and those estimates come from the same biased memory the framework was supposed to replace. I've written before about why ICE scores fail at predicting test impact; the short version is that garbage estimates in produce confident-looking garbage out.

A mention count is different in kind, not just degree. It is an observation, not an estimate. When you tell a room "this theme came up 400+ times, that one nearly 300, and the thing we've been debating for three sprints came up twice," the argument changes shape immediately. Nobody debates the count. They debate what to do about it — which is the conversation you actually want.

The count also travels well upward. Executives who glaze over at minimum detectable effects understand "customers told us this 400 times" instantly. It is the rare prioritization artifact that works in both the team standup and the quarterly business review.

Step 1: Pull the text you already have

You are not commissioning research. You are collecting text that already exists:

  • Support tickets — subject lines and first messages are enough; you don't need full threads
  • App-store and third-party reviews — these skew negative, which is exactly what you want for pain-point mining
  • Post-purchase and cancellation survey verbatims — short, high-intent, brutally honest
  • Sales and success call notes — objections are pain points caught pre-purchase
  • Session-replay observations — if your team logs what they see, those notes count as mentions too

One quarter of data is the right starting window. Less and the counts get noisy; more and you're averaging away real shifts. Pull everything into one place — a spreadsheet with a text column and a source column is genuinely sufficient for the first pass.

Qualitative sources matter here for the same reason they matter in test analysis: the numbers tell you what happened, the verbatims tell you why. I've made the case that the most valuable test finding is never the test result — this pipeline is how you industrialize that principle before a test ever launches.

Step 2: Cluster the noise into themes

Raw text doesn't rank; themes rank. A theme is a recurring pain point stated at the level a test could address: "surprise fees at checkout," "can't log in after password reset," "unclear what plan I'm on." Not "billing" — that's a category, not a pain point. Not "the fee disclosure on step three" — that's a solution wearing a theme's clothes.

The mechanics:

  1. Read a sample of a few hundred items and draft an initial theme list — expect 15 to 25 themes
  2. Tag every item against the list, adding themes when something genuinely new appears
  3. Merge themes that keep co-occurring in the same items; split themes whose items describe different root causes
  4. Keep an "unclear" bucket and review it at the end — it's usually where the next theme is hiding

Modern text-clustering tooling can accelerate the tagging, but don't let a tooling decision delay the exercise. The judgment calls — what level to cluster at, what to merge — are the valuable part, and they stay human either way.

Step 3: Count mentions per theme and rank

Now total the mentions per theme and sort descending. That's it. Resist the urge to weight by severity, revenue, or segment on the first pass — weighting reintroduces the opinions you were trying to escape. Raw counts first; apply judgment after the ranking exists, transparently, as a documented adjustment rather than a hidden input.

When we ran this across our own corpus, one theme cluster held 400+ mentions of errors, surprise fees, and payment failures. Another held nearly 300 mentions of login and interface friction. Both dwarfed the themes the team had been actively testing against. The ranking by raw mention volume produced a materially different test order than the team's intuition had — and intuition here meant a group of experienced practitioners who talk to customers regularly. If our informed gut missed the order that badly, an unranked backlog groomed by vibes has no chance.

Two rules keep the count honest:

  • Count items, not intensity. One furious ticket is one mention. Anger is signal for copywriting, not for prioritization math.
  • Log the window. "400+ mentions" means nothing without "in one quarter of support and review text." Every count carries its denominator when it goes in a deck.

Step 4: Convert top themes into hypothesis cards

A ranked list of pain points is research. A backlog is only born when each top theme becomes a structured, falsifiable test idea. The conversion is mechanical — every theme files a card with the same five fields:

FieldWhat it captures
Pain pointThe theme, in the customer's language
Affected segmentWho hits this, and where in the journey
Proposed interventionThe smallest change that could plausibly reduce the pain
Target metricThe one number that moves if the hypothesis is right
Expected rangeAn honest band, not a promise

The five-field discipline matters because vague submissions burn test capacity. A card that says "improve billing clarity" can't be built, powered, or falsified. A card that says "surprise-fee mentions cluster at checkout; test itemized fee disclosure before the payment step; primary metric is checkout completion" can go straight into grooming. This is the same bar I apply to every intake — see the hypothesis intake standard for the full falsifiability checklist.

Then the crucial move: these cards enter the same queue as every other test idea, groomed by the same rules. No special research lane, no special pleading. The difference is that they arrive with their evidence attached — when prioritization happens, the mention count is sitting right there in the card, and the debate is about data rather than opinions.

What did the counts actually change?

Three things, in our program:

The test order. As above — the mention ranking promoted themes nobody was championing and demoted pet projects with senior sponsors. The 400+-mention billing-and-payment cluster and the nearly-300-mention login-friction cluster both outranked initiatives that had been consuming test capacity for quarters.

The meeting. Backlog grooming stopped being a persuasion contest. Cards arrived pre-grounded, so the room argued about sequencing and design instead of whether a problem was real. We produced dozens of test-ready hypotheses this way without a single brainstorm meeting — the corpus generates the ideas; the team's job is shaping and shipping them.

The repeat-failure rate. When hypothesis cards carry their evidence, they also carry their history. A theme that already produced two inconclusive tests shows that lineage, which stops the quiet re-testing of settled questions — a failure mode I've diagnosed in detail in why CRO teams keep repeating failed tests.

Notice what's absent from this list: a conversion-lift claim. The counting pipeline doesn't win tests by itself. It makes sure the tests you run are aimed at problems customers actually have, at the frequency they actually have them — which is the highest-leverage decision in the entire program, made before a single variant is designed.

How do you keep the backlog honest over time?

Re-count quarterly. The backlog should track reality, not memory, and reality moves: a fixed bug drains a theme, a pricing change births a new one, a product launch reshuffles everything. A count from three quarters ago is intuition with a spreadsheet attached.

The quarterly re-count is also your feedback loop on shipped work. If you tested against the surprise-fee theme and the next quarter's count barely moved, that is a finding — the intervention didn't reach the pain, whatever the test's primary metric said.

This only compounds if the counts, cards, and outcomes live somewhere durable. A ranking that lives in one analyst's spreadsheet dies with their tenure — the institutional-memory failure I've written about in CRO knowledge management. The whole point of evidence-attached cards is that the evidence survives the person who collected it. That belief is why I built GrowthLayer the way I did: a browsable experiment library where every test keeps its hypothesis, its evidence, and its outcome in one place, so next quarter's count lands on top of last quarter's learnings instead of starting from zero.

If you want the backlog half of that system working this week: start free, file your top five themes as hypothesis cards, and run your first grooming session where the evidence is already in the room.

Key Takeaways

  • Your experiment backlog already exists in your support tickets and reviews — ranking it is a counting exercise, not a research project
  • Mention counts are observations, not estimates, which makes them harder to argue with than any scoring framework
  • Cluster text into pain-point themes, count mentions per theme per quarter, and rank by raw volume before applying any judgment
  • Convert top themes into five-field hypothesis cards — pain point, segment, intervention, metric, expected range — and groom them in the same queue as every other idea
  • In our corpus, the top theme held 400+ mentions and the second nearly 300 — and the resulting order differed materially from the team's intuition
  • Re-count quarterly so the backlog tracks reality instead of memory, and store counts, cards, and outcomes where they outlive any single analyst

FAQ

How many mentions are enough to justify a test?

There's no absolute threshold — the power of the method is relative ranking, not magic numbers. A theme with 40 mentions in a small program can outrank one with 400 in a large one. What matters is the gap between themes: when your top cluster holds 400+ mentions and your current test roadmap targets a theme with a dozen, the reallocation argument makes itself.

What tools do I need to run the counting pipeline?

A spreadsheet with a text column, a source column, and a theme column is enough for your first quarter. Research repositories and text-clustering tools speed up tagging at volume, but the highest-value steps — choosing the theme level, merging and splitting clusters, writing falsifiable cards — are judgment calls no tool makes for you. Start counting first; buy tooling when the row count hurts.

How is this different from ICE or PIE prioritization?

ICE and PIE ask you to estimate impact and confidence, and those estimates inherit every bias in the room. Mention counting replaces the estimated inputs with observed ones: instead of "I think this matters an 8," the card says "customers raised this 400+ times last quarter." You can still layer scoring on top for reach and effort — but the evidence base under the score is now something that happened, not something someone felt.

Does mention counting replace user research?

No — it sequences it. Counts tell you where the pain concentrates; they don't tell you why the pain exists or which intervention will reach it. The themes at the top of your ranking are precisely the places where interviews, session replays, and usability work pay off fastest. Think of the count as the map and qualitative research as the terrain.

How often should the backlog be re-ranked?

Quarterly. Faster re-counts chase noise; slower ones let the backlog drift back toward memory and politics. The quarterly cadence also doubles as a scoreboard for shipped experiments — a theme whose count doesn't fall after you've tested against it is telling you the intervention missed the mechanism.

FAQ

How many mentions are enough to justify a test?

There's no absolute threshold — the power of the method is relative ranking, not magic numbers. A theme with 40 mentions in a small program can outrank one with 400 in a large one. What matters is the gap between themes: when your top cluster holds 400+ mentions and your current test roadmap targets a theme with a dozen, the reallocation argument makes itself.

What tools do I need to run the counting pipeline?

A spreadsheet with a text column, a source column, and a theme column is enough for your first quarter. Research repositories and text-clustering tools speed up tagging at volume, but the highest-value steps — choosing the theme level, merging and splitting clusters, writing falsifiable cards — are judgment calls no tool makes for you. Start counting first; buy tooling when the row count hurts.

How is this different from ICE or PIE prioritization?

ICE and PIE ask you to estimate impact and confidence, and those estimates inherit every bias in the room. Mention counting replaces the estimated inputs with observed ones: instead of "I think this matters an 8," the card says "customers raised this 400+ times last quarter." You can still layer scoring on top for reach and effort — but the evidence base under the score is now something that happened, not something someone felt.

Does mention counting replace user research?

No — it sequences it. Counts tell you where the pain concentrates; they don't tell you why the pain exists or which intervention will reach it. The themes at the top of your ranking are precisely the places where interviews, session replays, and usability work pay off fastest. Think of the count as the map and qualitative research as the terrain.

How often should the backlog be re-ranked?

Quarterly. Faster re-counts chase noise; slower ones let the backlog drift back toward memory and politics. The quarterly cadence also doubles as a scoreboard for shipped experiments — a theme whose count doesn't fall after you've tested against it is telling you the intervention missed the mechanism.

About the author

A
Atticus Li

Applied Experimentation Lead at NRG Energy (Fortune 150) · Creator of the PRISM Method

Atticus Li has spent 9+ years in growth and experimentation at Silicon Valley Bank and NRG Energy (Fortune 150), and is the founder of GrowthLayer. He is a CXL-certified CRO practitioner and one of ~1,000 people worldwide certified in behavioral economics and consumer psychology through Mindworx. At NRG he has run 150+ experiments with a 24%+ win rate — in 2025 alone, his testing delivered $30M+ in verified financial impact, including $14M+ in cost savings.

Keep exploring

No spam. Unsubscribe anytime.