Solution Guides
Operational detail for Diamond 3: building and testing solutions
These are the steps behind Build and Test Solutions, and they run in the same three phases as the two diamonds before them: widen, narrow, then test before you commit. If you have worked through Community Guides and Pain Guides, the shape here will be familiar. That repetition is the method, not a coincidence.
How this page relates to the chapters
Unlike the other two diamonds, Diamond 3 never had guides: its methods were written into the chapters and stayed there. So this page is extracted from the chapters rather than written beside them, and every entry links back to the section it came from. All three phases now carry their procedures. The same procedures in machine-readable form are Block 3 of The Method Layer, Stages 12 to 14.
Everything on this page concerns solutions you have not built yet. Nothing here asks a customer to commit, and nothing here decides whether the venture is worth doing — that question belongs to Is This Worth Doing? and the decision itself to Make the Call.
Diverge — Generate Many Solutions
Use with Ideate Many Solutions, which carries the reasoning: why recombination works, and why a single compelling idea is a low-probability bet that feels like a good one. Three moves, in order.
Move One — Fill the Catalog
Recombination cannot combine what you do not have. Before generating anything, stock the pool: adjacent solutions from other industries, how this difficulty is handled elsewhere, what your people already improvise. A thin catalog produces thin ideas and the cause is invisible at the time.
Move Two — Generate
Brainstorming. Useful, and it has two failure modes that the next method exists to fix: the loudest person anchors the room, and evaluation creeps in while people are still producing.
6-3-5 Brainwriting. The book’s primary generation tool, and the numbers are the method: six participants, three ideas each, five minutes a round. Each round you pass your sheet to your right and use what you received as the prompt for three new ideas.
Run six rounds and a team of six produces around a hundred ideas in about half an hour. Teams do not believe this until they have done it, and that reframing is worth more than the technique: idea generation stops being days of staring and becomes a timed, almost mechanical half hour.
It fixes both brainstorming failures at once. Nobody can dominate, because everyone writes simultaneously. Nobody can evaluate, because everyone is busy producing. What happens instead is building — each sheet arrives carrying someone else’s thinking and you extend it.
- Adjust the numbers to your team and keep the structure. Four participants and four rounds still works.
- Keep every sheet. They are the raw material for recombination and for the next chapter.
- Debrief afterward, never during. The session ends, then you look at what you have.
A round of deliberately bad ideas
Chapter: Deliberately Bad Ideas.
Seed them, do not schedule them. Ask everyone to arrive with three or four ideas they think might work and one they know is terrible, and have the bad ones go onto the sheets in round one with the rest. A bad idea written in round four has two rounds left to be built on; the same idea written in round one has five, and it spends them traveling round the table being answered.
- State the target plainly when you ask for it. The worst thing we could build for these people. Infeasible counts; something that makes the pain worse counts double.
- Keep the timing and the silence of a normal round. It is a generating round, not a discussion.
- Pass and build, exactly as usual. Bad ideas compound enjoyably.
- Do not sort them into a separate pile. They go in the catalog with everything else, because their value is as recombination material.
- Harvest before you move on. For two or three of the worst, ask what would have to be true for this to be good, and invert one attribute. Write down whatever that produces.
What the round tells you
If the team struggles to make the pain worse, stop. Inventing harm requires knowing the mechanism, so a silent room here usually means the pain is understood as a label rather than as a cause. Go back to the experience map before generating further.
Move Three — Recombine
The move teams skip, because it looks like effort and arrives dressed in overalls.
SCAMPER and Systematic Inventive Thinking are the two structured recombination methods. SCAMPER is a checklist and the chapter carries it in full. SIT needs a procedure, and it is below, because it returns in the next phase applied to high-scoring requirements.
Systematic Inventive Thinking
SIT works by a reversal that is the whole method, and skipping it turns the templates into arbitrary vandalism. Ordinary design runs problem first: name a need, then invent something to meet it. SIT runs function follows form. You take the thing in front of you, damage it according to a template, look at the strange object you now have, and only then ask what it might be good for (Boyd and Goldenberg 2013).
Two rules govern it.
The closed world. Work only with components already present in the product or its immediate surroundings. Nothing new may be introduced. The constraint is the engine: ideas that could use anything tend toward the elaborate, and elaborate solutions are the ones that do not get built.
Function follows form. Build the virtual product first, absurd as it looks, and withhold judgment until you have asked what it could do. The instinct to reject arrives before the instinct to interpret, and the whole technique lives in that gap.
The five templates. Apply each to a concept that survived screening.
| template | what you do | the question afterward |
|---|---|---|
| Subtraction | Remove an essential component. Not a trivial one — the one that seems indispensable. | What is the remainder good for? Who would want it because that part is gone? |
| Task unification | Assign an existing component a second, unrelated job. | What does the doubled-up part now let you delete? |
| Multiplication | Copy a component, then change the copy in some qualitative way. | Why would anyone want two of these, one of them altered? |
| Division | Divide the product, functionally or physically, and rearrange the pieces in time or space. | What becomes possible when the parts need not be together? |
| Attribute dependency | Make two attributes that were independent vary with each other. | Which pair, and in which direction, and what does that make possible? |
Running it. Twenty minutes per concept is usually enough.
- List the components of the concept and its immediate surroundings. That list is your closed world; nothing off it is allowed in.
- Pick one template and apply it literally. Do not adjust it to be sensible. If subtraction produces a safety device with no alarm, write down the device with no alarm.
- Describe the virtual product in a sentence, as though it existed.
- Ask what it is for, and only now. Who would want this? What does it now do that the original could not? What did removing the part make possible?
- Judge it last, against feasibility and against the market requirements you have. Most results are discarded. That is the expected yield and it is not a failure of the method.
- Repeat with a different template on the same concept before moving to the next one. The templates find different things.
Four templates worked through on a real concept, with the strange intermediate objects left in, are in the Halo Alert recombination record. The subtraction that removes the phone is the one that produced the surviving concept.
Fixing the absurdity before you have questioned it
The reflex when subtraction produces a device with no alarm is to put the alarm back, or to substitute something for it. Both end the exercise. The absurd object is the point: it is a place in the design space you would never have walked to deliberately, and its value shows up only when you ask what it is good for rather than what is wrong with it.
What this page does not cover
This is enough to run the five templates on a concept you already hold. SIT is a fuller system than that — it has more to say about which template suits which situation, and about applying templates to a market rather than to a product. See Goldenberg et al. (2003) and Boyd and Goldenberg (2013).
Diverge is complete when
Judge the pool on three measures, not one.
- Quantity. The raw count, of distinct concepts. Twenty variants of one idea is one idea.
- Variety. Sort the pool into groups. If everything lands in two or three clusters, you explored a narrow slice and should return to recombination with a different constraint.
- Novelty. How far the pool reaches past the obvious. Keep the seemingly infeasible ones; they are frequently unusable as stated and frequently the parent of something usable.
Large, varied, and reaching is ready for screening. Large but narrow is not: it is one idea wearing many hats, and screening will simply hand that idea back to you.
Converge — Narrow to One
Use with Hypothesize a Good Solution, which carries the judgment: where market requirements actually come from, why the first list is provisional, and why this phase eliminates rather than ranks.
Four filters, cheapest first. Run them in order. Each is designed to be quick, because the expensive instrument should only ever see a short list.
Feasibility filter
Chapter: Feasibility Filter.
Two questions per idea, answered fast:
- Could a team like yours build this, within the money, skills and time you actually have?
- Does it address the validated pain, rather than an adjacent one that happens to be interesting?
No to either, set it aside. This should take minutes across a pool of a hundred, not a meeting. You are not judging quality yet and you will be wrong about a few; the point is to stop the careful instruments from spending time on the impossible.
Set aside, do not delete. Every idea earned its place by widening the pool, and an infeasible idea often contains a feasible mechanism. They go to the refrigerator with the rest.
Done when
What remains is plausible, not good. If you find yourself arguing about which of two survivors is better, you have left the filter and started screening early. Stop and move on.
Dot voting
Chapter: Dot Voting.
- Spread the survivors where everyone can see them at once, one idea per note.
- Give each person a fixed number of dots, three to five. Fewer than the number of ideas by a wide margin, or the exercise does nothing.
- Vote silently and at the same time. Sequential voting is not voting; it is watching the first person vote.
- Vote for what you would want to work on, not for what you predict will win. The information you are collecting is conviction, and predicted-winner voting destroys it.
- Count, and read the spread as well as the top. A clear favorite is useful. So is a dead tie, which says the team is genuinely split and that the split will surface later whatever you decide now.
Expect to go from a hundred-odd ideas to ten or twenty.
What dot voting is evidence of
Your team’s enthusiasm, and nothing else. It is a real input — a team that does not believe in the concept will build it badly — and it is not customer evidence. Do not let a dot count overrule a screening result, and never report it as validation.
Screening matrix
Chapter: Screening Matrix. The method is Pugh’s concept screening, and its first stage is deliberately coarse.
1. Rows: market requirements. Each stated as something the solution must or should do for these people. Record where each one came from — exploration, a rejected concept, or an assumption — because provenance is what tells you which result to doubt when one surprises you.
2. Pick a reference solution. The best thing currently available to your people, whatever it is. In the Halo Alert demonstration it was pepper spray. The reference is not a competitor to beat; it is the zero point that makes the ratings mean something, so choose what these people actually use today rather than the most sophisticated product on the market.
3. Columns: the surviving ideas, from dot voting.
4. Rate each cell against the reference, not against the other ideas:
| mark | meaning |
|---|---|
| + | better than the reference on this requirement |
| = | about the same |
| − | worse |
5. Net score: pluses minus minuses, equals counting zero.
6. Decide, per column. Improve a high scorer. Combine a low scorer that nonetheless beat the reference on something, because that plus is a feature worth transplanting. Drop what beat the reference nowhere.
| Requirement | Reference | Idea A | Idea B | Idea C |
|---|---|---|---|---|
| requirement 1 | = | |||
| requirement 2 | = | |||
| requirement 3 | = | |||
| Net score | 0 | |||
| Decision | reference |
Reading the top score as the best idea
The matrix scores each idea against the reference, so a high net score means “beats pepper spray on more counts than it loses,” which is not the same as “is the right thing to build.” It eliminates confidently and chooses only weakly. Carry two or three forward, not one.
It is also unweighted on purpose. On a first pass your requirements are gates rather than preferences, and weighting gates lets a well-rounded concept outscore one that actually clears them.
Done when
- Every surviving idea has a decision written against it: improve, combine, or drop.
- You can say which requirement did the most eliminating, and you can trace it to a person.
- Two or three concepts go forward. One means you ranked rather than screened.
Systematic Inventive Thinking
Returns here, applied to the requirements your surviving concepts scored well on, to improve what survived rather than merely rank it. Take the pluses from the combine column too: a feature that beat the reference inside a losing concept is exactly what recombination is for. The five templates and the procedure are in the divergence phase above, at Systematic Inventive Thinking.
The second pass
Once, and usually only once, a second pass. Testing generates market requirements faster than exploration does, because rejection needs a concept to push against, so the list you screened against is stale by the time you finish screening. Come back with the objections, add rows, rescore. It needs no new respondents and takes about twenty minutes.
A scoring matrix belongs to that second pass and nowhere earlier: weighted requirements, ratings instead of plus and minus, a weighted total. It is worth building only when the weights came from customers rather than from you, which in practice means after a $100 test.
For the Curious — where weights come from when this is done properly
The formal answer is conjoint analysis, and it is how the numbers in a professionally built scoring matrix are actually produced (Green and Srinivasan 1978, 1990).
The idea is easier than the name. Instead of asking people how important each attribute is, which they answer badly, you show them whole products and let the importance be inferred from their choices. Say you were doing it for Halo Alert. You would write a dozen descriptions of a personal safety device, varying four things independently across them: how discreet it is, whether it confirms the alert was received, whether it works without the phone, and the price. Each respondent sees several and ranks them or picks one. Nobody is ever asked what matters most.
Then you estimate a model in which preference is the outcome and each attribute contributes a coefficient. Those coefficients say how much of people’s choosing each attribute accounted for, and those are the weights. The trick is that the attributes vary independently by design, which is what lets the model separate their effects.
Two things stop it working at this point in an expedition, and both are worth understanding rather than taking on trust.
You cannot estimate seven weights from three concepts. Notice what the dozen designed descriptions were for: the attributes had to vary independently, and there had to be enough combinations to pin each coefficient down. What you hold after screening is three survivors that differ in tangled ways, all at once. There is nothing to regress.
Gates do not have coefficients. A model of this kind is compensatory: it assumes a deficit on one attribute can be offset by a surplus on another, and it estimates the exchange rate between them. That assumption is exactly false for a requirement phrased as must. A solution that can be turned against its owner is not redeemed by being attractive, however the arithmetic comes out.
So conjoint belongs where the attribute set has already settled and the customers are already known: established products, line extensions, pricing an option package. Early in an expedition you have neither.
The $100 test is the field-expedient version of the same idea. It elicits relative value directly instead of inferring it from choices, which costs precision and buys the ability to run it on twenty people next week. That is usually the right trade when what you need is a ranking rather than an exchange rate.
Scoring Matrix
The same table with the coarseness taken out: weights on the requirements, a five-point rating instead of plus and minus, and a weighted total. It eliminates further and it separates the survivors, which screening deliberately will not.
Do not run it on a first pass. Its weights have to come from customers rather than from you, which in practice means after a $100 test, and its ratings assume trade-offs are possible, which is false while every requirement is still a gate.
Keep the same requirements you screened against, corrected by whatever testing has told you since.
Keep a reference solution, and keep a column for it. It can be the same reference as before, or the strongest surviving concept if that is now clearly better than what people use today.
Assign a weight to each requirement, summing to one across the list. This is the step that decides whether the exercise is worth anything. A weight is a claim about how much customers care, so it needs evidence: $100 allocations, ranking data, observed spending. A weight you reasoned your way to is a number wearing a costume.
Rate each concept on each requirement against the reference, one row at a time, on five points.
| Rating | Meaning |
|---|---|
| 1 | Much worse than the reference |
| 2 | Worse than the reference |
| 3 | Equal to the reference |
| 4 | Better than the reference |
| 5 | Much better than the reference |
Rate a row at a time rather than a concept at a time, so you are comparing like with like, and argue the ratings out loud as a team rather than averaging private guesses.
- Score. Each concept’s total is the sum of weight times rating across the requirements.
\[S_j = \sum_i w_i \, r_{ij}\]
where \(w_i\) is the weight on requirement \(i\) and \(r_{ij}\) is concept \(j\)’s rating on it.
| Requirement | Weight | Reference | Concept A | Concept B |
|---|---|---|---|---|
| requirement 1 | 3 | |||
| requirement 2 | 3 | |||
| requirement 3 | 3 | |||
| Weighted total | 1.00 |
Letting the total decide
The highest score is not automatically the idea to build, and the arithmetic is not where the value is. Most of what this exercise gives you happens while you are arguing about weights and ratings, which forces the team to say out loud what it believes about customers and where those beliefs came from.
Treat a surprising total as a question rather than an answer. Check the provenance of the two or three requirements carrying the most weight before you trust it, and check whether a gate has been smuggled in as a heavily weighted preference, because a weighted model will happily let a concept fail one and win anyway.
Test — Test Before You Build
Use with Test Solutions with Prototypes. Procedure extraction in progress; the chapter carries each test in full today.
Seven tests. The first distinction to hold on to is what each one asks: validation faces the customer and asks whether you are solving the right problem for the right people; verification faces the workbench and asks whether the thing can be built and works. One rough prototype often serves both.
| test | what it asks |
|---|---|
| Validation testing | Does this solve the pain, for these people? |
| Verification testing | Can it be built, and does it work as intended? |
| Wow factor test | Is there emotional pull, not merely acceptance? |
| $100 test | Which features do they actually value, given a forced trade-off? |
| Wizard of Oz test | Would they use it, if the manual version behaves like the real one? |
| Smoke test | Will anyone act on the offer before it exists? |
| Profit analytics | At what they would pay, does this earn more than it costs? |
Wizard of Oz is not only a test. It is also an input to several of the others, because it is how you put a solution concept in front of someone convincingly enough that their reaction means something. Reach for it whenever a test needs the customer to experience the solution rather than imagine it.
Validation and Verification
Chapter: Validation and Verification. These are not two tests so much as two questions any prototype can be asked, and one rough model usually answers both.
Validation faces the customer. Who exactly is this person? What is the pain? Does this relieve it? You have answers to the first two from Diamond 2; what is new here is the third.
Verification faces the workbench. What must it do? Can it be built? What is the best way to build it?
Run them in that order and keep them separate in your notes, because they fail differently and the difference matters. A solution can verify perfectly and never validate: it works exactly as designed and nobody wants it.
Done when
- You can say which of the two a given result speaks to. A result that seems to answer both usually answers neither.
- A validation failure sends you back to the concept or to the pain. A verification failure sends you back to the workbench. Write which one you are looking at before you start redesigning.
Wow Factor Test
Chapter: Wow Factor Testing. Measures emotional pull rather than acceptance, which are different and easy to confuse.
- Present the concept, saying plainly what it is for and what it does.
- Ask for a rating, one to ten, of their emotional connection to it.
- Collect the positive: what appeals, what benefit they foresee, when they would use it.
- Collect the negative: what they dislike, what is missing, what would stop them buying.
- Invite improvement: how would you change it, what would make you more likely to buy.
- Summarize back to the group and ask what the main learning was.
Steps four and five are the ones that get cut when time is short, and they carry most of the value. A room that has just been asked what it likes will say generous things; the same room asked what would stop them buying says something usable.
The bar
An average around 7.5 is where a concept starts to look promising — high enough to indicate real pull rather than politeness (Swenson et al. 2013).
Treat it as a threshold and not a score to maximize. A 9 from three people who love you is worth less than a 7 from twenty strangers.
$100 Test
Chapter: $100 Test. Finds which features carry the value, by forcing a trade-off instead of asking for an opinion.
- Present the pain and the solution, in person or in a survey.
- List candidate features generously. Cover function, design, and experience, and include some you expect to lose.
- Give each person $100 to allocate across the features, more to what they value more.
- Read the allocations, not the enthusiasm. Features that consistently draw money are the minimum viable product; features that draw none are feature creep you have not built yet.
Watch for segments. If two groups allocate very differently, you may be looking at two populations rather than one, which is a Diamond 2 finding arriving through a Diamond 3 instrument. Take it back to From Community to People.
It is hypothetical money
Nobody hands anything over, so the totals are softer than a purchase. What is reliable is the ranking and the spread: which features beat which, and whether people concentrate their money or scatter it. Concentration means a clear core; scattering usually means the concept is doing several unrelated jobs.
Wizard of Oz Test
Chapter: Wizard of Oz. A person does by hand what the product would do automatically, so you can test whether anyone wants the behavior before building the machinery.
- Create the illusion. A mock-up, a slide animation, a plain interface — enough that the person believes they are using a working thing.
- Operate it manually. Somebody behind the scenes enters the data, produces the result, sends the message.
- Let them use it as they would a real product, and watch rather than explain.
- Collect feedback on concept, usability and perceived value, and note where they hesitated.
- Let them go off script. The unexpected actions are the most informative part, and the temptation to steer them back is the thing to resist.
It is also an input to other tests, and that is the use most people miss. Wow factor, the $100 test and validation all need the person to experience the solution rather than imagine one. A Wizard of Oz rig is usually the cheapest way to give them that experience.
Done when
You can describe what the person did, not only what they said. If your notes contain opinions and no behavior, the rig was a demonstration rather than a test.
Running a Smoke Test
Five steps. The chapter carries why it works and what it cannot reach; this is what to build.
1. A prototype convincing enough to be believed. Looks-like and works-like far enough that a visitor treats it as a real offer. If it reads as a mock-up they leave, and that is indistinguishable in your numbers from leaving because they do not want it.
2. A landing page that asks your question. Decide the question first. The anatomy below is what the page needs regardless of which question it asks.
3. A way to count. Whatever analytics your site builder offers is enough. What matters is not the tool but that you can count arrivals, scroll depth, clicks on the call to action, and completions — and that you verify it records a test conversion of your own before you spend anything on traffic.
4. Traffic. Two kinds of channel, and the distinction outlasts whichever platforms are current. Search reaches intent: people already looking for a solution, which is the population this test exists to reach. Social and display reach demographics: people who match a profile but were not looking. Prefer intent when you can buy it. If you use social, remember you have partly re-created the recruiting problem the test was meant to remove.
5. Analysis against a threshold you set first. Write the number that would count as a pass before the first visitor arrives.
What a landing page contains
Above the fold, where visitors land before scrolling, and often all they see:
| element | job |
|---|---|
| Headline | The benefit, in a few words. Sell the relief, not the product. |
| Subheader | Reinforce the headline and clarify the offer. |
| Offer | One vivid image or short video showing the life improvement. |
| Call to action | The thing you want them to do, and the thing you are measuring. |
Below it, in order:
- Value proposition. The pain, the relief, and why yours rather than another. Listing features or competing propositions here is what lets you measure which ones draw attention.
- Trust elements. Honest ones only, and as a new venture you will not have all of them: experience, testimonials, who you are and why you care, what you guarantee, what you promise. Promise nothing you cannot keep.
- A second call to action, plus a way to capture contact details. That list is your next round of validation respondents, and they volunteered.
Things that hold regardless of platform
- One topic per page. A page asking two questions answers neither.
- The call to action goes above the fold.
- The value has to land in about ten seconds.
- Directional cues — an arrow, a gaze, an image pointing — move attention where you want it.
- Be succinct. Every extra paragraph is a chance to leave.
- Design for the conversion you are counting, not for how the page looks to you.
Why this guide names no tools
An earlier version of this material specified site builders, analytics packages, and a click-by-click advertising setup. Nearly all of it was wrong within a year: the menus moved, the products were renamed, the support links died. Nothing here is worth learning from a book that cannot be updated as fast as the interface changes.
What does not move is the structure: intent traffic beats demographic traffic, the page asks one question, the threshold is set before the test, and the conversion is verified before you pay for a visitor. Use whichever builder and analytics you like; check the vendor’s own current documentation for the buttons.
Profit analytics is deliberately partial here. The chapter introduces what the test does and when it belongs, then hands off: estimating a demand curve and computing expected profit properly is a book’s worth of method, and that book is Is This Worth Doing?. The overlap is intentional. A reader should leave knowing the test exists and why it comes last, and should not try to run it from what is written here.
Profit Analytics
Chapter: Profit Analytics Testing. Comes last, because it needs a concept concrete enough to price.
Deliberately partial here, and the chapter says so too. What belongs in this book is the shape: customers’ willingness to pay reflects how well the solution relieves the pain, so a demand curve estimated from real responses lets you find the price that maximizes expected profit, and a low or negative result is a reason to reconsider rather than to push on.
Doing it properly, which means eliciting willingness to pay without biasing it, estimating the curve, assembling costs, and finding the threshold at which the answer flips, is a book’s worth of method, and that book is Is This Worth Doing?. The overlap is intentional. Leave here knowing the test exists and why it comes last; do not try to run it from what is written here.