⏱️ Reading Time

12 min

📅 Publication Date

📝 Abstract

Concept testing asks which of your options people prefer. It is a good question, but a narrow one — the decision that mattered happened weeks earlier, when somebody made the cut. We don't just test concepts. We generate the ones worth testing.

Concept Testing

Product Innovation

AI Product Innovation

Market Research

Consumer Insights

CPG

Beyond Concept Testing

Ideas are cheap, execution is still expensive.

Concept testing is all the rage today, and understandably so: ideas are cheap, execution is still expensive. But we believe most concept testing projects start the wrong way. Traditional concept testing asks which of a given set of options people prefer. It is a good question, but a narrow one. The real decision that mattered happened weeks or months before: it was the cut.

Somewhere upstream, in a workshop or an internal review that left no record, somebody decided on a set of options.

So while a preference test may not be wrong, it is narrow. It can rank what you bring, but it cannot tell you whether the best option in the given set is good enough, it cannot tell you what should have been on the list instead, or which customers inside the average disagree.

When a decision is cheap to reverse, the gap is survivable. When it is a product with a 48-month lead time and over a billion dollars committed, it is not.

AlgoVerde goes beyond concept testing by turning the option set into an output rather than an input. To do so, it builds a living AI model of your market, and against that model a product team can run simulations and make product decisions. The purpose is not validating an existing set of options. The purpose is validating the right set.

Concept testing for AlgoVerde consists of six steps: explore the space, generate alternatives, stress-test them, see which segments they win and which they lose, improve them, and decide with the reasoning attached. Validation is one outcome of that loop rather than the only purpose of it.

Several Fortune 500 companies are embracing AlgoVerde’s richer simulation method, beyond concept testing. A Global Automaker compressed concept development from over a year to a few months, and third-party research costs by 75%. A Global CPG moved concept work from weeks to days and shipped messaging that lifted add-to-cart 60%.

Industry validation and methodological studies provide clear evidence beyond numbers.

Let’s delve into why and how AlgoVerde goes beyond Concept Testing.

It's more than the concepts in front of you.

When you are thirsty, which drink do you reach for? Probably not Coca-Cola Spiced, launched with fanfare in February 2024 as a permanent addition to the range and pulled from shelves roughly seven months later. Liquid Death did the harder thing: it entered a saturated category, added iced teas in 2023, and closed a funding round in March 2024 that valued the company at $1.4 billion. Neither outcome was settled by the quality of a preference test. Both were settled much earlier, by what was considered in the first place.

Concept testing has been the backbone of consumer research for half a century, and deservedly so. Put two or three options in front of a sample, collect the ratings, back the one that scores highest. Done properly it beats guessing, and no amount of scepticism about its scope changes that.

The scope is the issue. A comparison test grades the options in the room. If the strongest idea for your category never got into the room — because nobody thought of it, or because it was cut in a review three months earlier — the test will not say so. You get a winner. You do not find out what you missed.

What a preference test can and cannot tell you.

A preference test can rank the options in front of it reliably, quantify the gap between them, and give you a defensible reason to back one over another. That is real, and running a larger sample makes it more precise.

It cannot tell you whether the winner is good enough to launch, what should have been on the list instead, which element of the concept is carrying it, or who inside the average is objecting. A larger sample improves none of those four.

Teams underrate this because a test result feels like an answer about the market. In reality, it is an answer about your shortlist, and your shortlist was set by whoever happened to be in a room on a Tuesday.

When the shortlist gets more expensive.

None of this is fatal when a decision is easy to walk back. Most decisions worth testing are not. A car program runs on a 48-month lead time with over a billion dollars committed. A new consumer goods platform locks formulation, supply and packaging years before anything reaches a shelf. By the time a concept test happens, the decision it informs is fixed for a long time.

The reflex is to ask for the same answer faster. Speed genuinely helps, but a faster answer to the shortlist question is still an answer about the shortlist.

But what if the option set were an output of the process rather than an input to it? You are no longer limited to choosing between the concepts in front of you. You can explore, generate, test, improve and decide across the whole possibility space.

What you could answer before
What you can answer now

Which of our concepts wins?

What is the whole possibility space, and where in it do we have a right to win?

Where did these concepts come from?

What else belongs on this list, and can we have it this week?

How does the concept score?

Which element is carrying it, and where does it break?

What is the average score?

Who wins, who loses, and how much volume sits with the people who disagree?

Which one do we back?

How do we make the best of them stronger before we commit?

Which one won?

What do we lead with, what complements it, what do we retire and how can we show our work?

Six steps, one loop.

A global household-care business is deciding what to launch into a category that has been flat for three years. The team arrives with three concepts already chosen and a stage gate in eleven weeks. Here’s what happens.

Six steps that connect in one continuous loop, from exploring the possibility space to deciding what to back — and why.

  1. Explore

See the space before choosing where to play.

Concept testing starts with a set of options. AlgoVerde starts one step earlier: what are all the options worth considering?

Using a living model of your market — combining market, category and competitive intelligence with your own customer and product data — AlgoVerde identifies and ranks the opportunity areas where your business has a credible right to win.

The result is not a score or a concept. It is a map of where you could play, before deciding what to test.

AlgoVerde real-life scenario.

The team arrived with three concepts. AlgoVerde surfaced eleven opportunity areas — four the business had not considered — and showed that all three existing concepts sat within the same one.

  1. Generate

Turn the possibility space into options worth testing.

Concept testing asks you to bring the options. AlgoVerde can help create them: what else should be on the table?

Starting from the opportunity map, AlgoVerde generates distinct concepts, explores different directions and ranks them against the problems and opportunities identified in the market.

The result is not just more ideas. It is a broader, ranked set of options grounded in where the opportunity actually is.

AlgoVerde real-life scenario.

Two independent generation passes produced 24 concepts. After removing overlaps, 17 distinct options remained — five in opportunity areas the business had not previously considered.

A preference test never sees this step. Generation happened in a workshop months earlier, and whatever failed to survive the room is invisible to the test that follows.

  1. Stress-Test

Understand why a concept works, and where it breaks.

GenAI Personas representing your customer segments interrogate each concept, surfacing the strengths and weaknesses of its promise, benefits, claims and reasons to believe.

Dozens of interviews and focus groups execute in minutes and return structured findings rather than raw transcripts. Strengths and weaknesses surface so you learn whether to fix the promise or the proof.

The result is not just a score. It is a diagnosis of what to keep, what to fix and what is not strong enough to move forward.

AlgoVerde real-life scenario.

12 of the 17 concepts failed the stress test and were retired. Five had enough potential to carry forward.

  1. Segment

See who you win, and who you risk losing.

Each Persona reacts independently, so conflicting responses across customer segments remain visible rather than disappearing into a mean.

The result is not just an overall preference. It is a view of the trade-offs behind it — and whether the people you risk losing actually matter to the business.

AlgoVerde real-life scenario.

The concept ranked first overall came third among pet-owning households, a group representing a disproportionate share of category volume.

Disagreement is flagged, not averaged.

When Personas react differently to the same concept, AlgoVerde reports the split rather than resolving it into a mean. Conflicting reactions are surfaced as conflicts, with the verbatims behind them.

A small group with an intense objection shows up as a small group with an intense objection, not as a rounding error. Whether that group matters is then a commercial judgement about volume and value, made by your team with the evidence in front of it.

  1. Improve

Make the best option better before you commit.

Concept testing typically ends with a winner. AlgoVerde keeps going: weak elements are rewritten, alternatives generated and the improved concepts tested again against the same customer segments, without waiting for another research cycle.

The result is not simply a winning concept. It is a better concept than the one that entered the process.

AlgoVerde real-life scenario.

The leading concept was rebuilt around a substantiated safety claim and re-tested. Its panel score moved from 6.1 to 7.4 while improving its performance among pet-owning households.

Whether it’s Single Concept Optimization taking one concept through variant generation, or Multi-Concept Optimization taking a portfolio through parallel testing, both workflows then feed straight back into Stress-testing for the best outcome. A preference test ends here, with a verdict on what you brought it. Nothing in the method makes the winner better.

  1. Decide

Commit, with the reasoning attached.

Concept testing gives you a result. AlgoVerde helps turn the evidence into a decision: what should we lead with, what should we change or retire, and why?

The system brings together concept performance, segment reactions, trade-offs and the evidence behind them into a recommendation. The decision remains with the team accountable for making it.

The result is not just a ranking. It is a defensible recommendation, with the evidence, reasoning and next action attached.

AlgoVerde real-life scenario.

In our example, the team left with one lead concept, one complement and three retired options — and reached the stage gate in 4 days with a concept that had not existed when the process began.

What can you test?

LARGE, SLOW PRODUCT DECISIONS
SMALLER, FASTER PRODUCT DECISIONS

Vehicles and mobility — new vehicle concepts, trims, features, powertrains

Flavour and formulation — candy flavours, beverage variants, ingredient combinations

New products — new propositions, category entries, product platforms

Packaging — formats, claims, colours, naming, pack architecture

Brand and positioning — value propositions, brand territories, campaign directions

Features and details — individual functions, UX elements, accessories, options

Messaging — headlines, product claims, ad copy, benefit statements

Market offers — pricing structures, bundles, ownership or subscription models

Can you trust it?

Trust is the first question in the room, and rightly so. There are two sides to the evidence: how AlgoVerde’s simulations compare with human research, and what happens when the system is deployed in the real-world. AlgoVerde’s GenAI Personas have been tested against human research in different settings.

Our Customers’ validation.

  1. A Global CPG ran AlgoVerde’s Personas against its own purchase-intent panels, using proprietary data and a controlled methodology, and measured 90% alignment.

  2. A second client compared the Persona outputs to three separate research results it already trusted and had paid for, and found them close enough to replace those studies.

  3. An automotive client ran a structured three-day session with its concept testing team. Their assessment afterwards: “the closest we’d gotten to real-life simulation results.”

A blind backtest.

In a blind backtest, working with a leading consulting firm, AlgoVerde reproduced a 2024 US consumer sentiment study using a Synthetic Panel representing the US population aged 18 and over, at the original sample size of 2,100.

96.3% agreement with the human panel

3.7 mean absolute error, in points

0.97 rank correlation across the ten factors measured

The structure held: inflation and the price of necessities dominated both lists, income sufficiency and income stability came next in both, and the same factors sat at the bottom. The Synthetic Panel rated most factors slightly higher in absolute terms.

Defensible by design.

The answer does not sit in any single response. It sits in the architecture.

Business context comes first, with curated data from traceable sources. Then the right Personas, specific to the company and the project. Then structured, repeatable workflows with specialised analytical agents, source attribution and validation. Optimized prompting comes last rather than first.

Subject matter expertise sits alongside AI expertise, and experts stay in the loop. They are the reason the output can be defended, not an afterthought to it.

Who is using it.

AlgoVerde’s signature approach to concept generation and testing has been deployed inside Global 500 companies for more than two years, across automotive, consumer goods, retail, finance, advertising, education, media and insurance — 300 growth decisions simulated across 30 forward deployments. It was built in collaboration with the AI Institute at Harvard Business School.

A Global Automaker
A Global CPG

A Global Automaker compressed concept development from over a year to 3-4 months, and cut third-party research costs by roughly 75%.

A Global CPG moved concept work from months to days, shipped messaging that lifted add-to-cart 60%, and measured 90% alignment between AlgoVerde’s Personas and its own human purchase-intent panels.

Where to start.

AlgoVerde can be configured around your existing process and alongside your team’s preferences. It mirrors the workflows, the logic and the output formats your team already trusts, and the scoring criteria are adapted to the ones you use. That is what makes the decision defensible twice over: you can show the board how you got here, and you can run it the same way next quarter.
It is also what makes this a capability you own rather than a service you buy each time, not yet another pilot.

A first engagement needs just three things from you: a project brief, a description of the customer segments that matter, and agreement on what a good answer looks like. Our forward deployment team will get a system up and running in days and results in a few weeks.