Question and when it applies
This is a product discovery and investment-decision question. The team has received explicit requests, but the existence of a request only proves that a voice exists. It does not establish that target users encounter the problem frequently, that the current alternative is costly enough, that the proposed solution will change behavior, or that a buyer will commit budget. Four engineers working for ten weeks create real opportunity cost, so the product manager must use three weeks to buy evidence capable of changing the decision.
A public 2026 product manager question bank still asks candidates how they would validate a product idea. Recent product-validation guidance emphasizes identifying the riskiest assumptions, matching a prototype or concierge test to the question, setting success criteria in advance, and then using real-user evidence to build, adjust, or test further. The GOV.UK Service Manual likewise recommends understanding users, current behavior, constraints, and the problem before committing to a build. It explicitly treats stopping after discovery as a valid outcome when the evidence supports it.
The task is to control investment under uncertainty. The candidate should turn “build an approval portal” back into a problem to investigate, then separate problem, audience, solution, usability, feasibility, and business sustainability. Three weeks cannot prove durable product-market fit, and the answer should not pretend it can. The validation can eliminate fatal assumptions, reduce uncertainty, and produce an auditable decision about the next investment.
The company, request count, budget, schedule, and later gates are fictional interview inputs, not industry benchmarks. A real team must recalculate them for its segment, buying cycle, baseline, and tolerance for a wrong decision.
What the interviewer evaluates
First, can the candidate turn a prescribed solution back into a problem? Sales forwarded requests for an approval portal, but the underlying job may be reducing wait time, preventing version confusion, preserving an audit trail, or avoiding training external clients on the full system. Testing the portal interface immediately could optimize a prematurely selected solution without proving that it addresses the most important problem.
Second, can the candidate distinguish evidence strength? Saying the idea sounds good, agreeing to try it, leaving an email, bringing a real deliverable into a pilot, using it repeatedly, granting workflow access, and approving spend impose progressively higher commitment costs. They answer different questions. One prototype session can reveal comprehension and usability issues, but it cannot prove retention or payment. A few interviews can explain current behavior, but they cannot estimate population prevalence.
Third, can the candidate test assumptions that would kill the investment if false? Button placement and notification wording should come after these questions: Do target accounts repeatedly suffer approval delays? Will the buyer change an email-and-PDF workflow? Will external clients use a new entry point? Can access, audit, and data boundaries be delivered safely? Is the value large enough to support development and service cost?
Fourth, does each experiment match its assumption? Problem interviews and current-state observation examine pain. A clickable prototype examines comprehension and task completion. A manual concierge pilot examines workflow and outcome. A paid pilot or budget approval examines commercial commitment. A survey cannot validate all of them, and an MVP should not mean building a smaller product before the main risk is understood.
Fifth, can the candidate write the gates before seeing the results? The segment, sample, observation window, pass condition, guardrails, and action after failure should be agreed in advance. Changing the metric after the study turns ordinary positive feedback into apparent success.
Finally, does the recommendation allow more than a binary yes or no? A strong answer can build the smallest scope after the evidence passes, narrow to the only segment with a real problem, pivot when the problem is real but the solution fails, or stop when a fatal assumption or non-compensable risk fails.
Clarifying questions before answering
- Who produced the 12 requests? Deduplicate them by current customer or prospect, account size, industry, buying role, and sales owner. Four retellings from one large account do not constitute four independent needs.
- What happened in the most recent real incident behind each request? Ask what needed approval, how long it took, how many rounds occurred, who was blocked, and what the consequence was. Abstract preference invites courtesy answers; past behavior explains the problem.
- What is the current alternative? Email, shared documents, e-signature, tickets, chat, and manual follow-up carry different costs. If the current method is good enough, switching friction may consume the new feature's value.
- Who are the user, buyer, and risk owner? An internal project manager may start the request, an external client completes it, a procurement or department lead pays, and security or legal may veto deployment. Interviewing only one role misses the adoption chain.
- What business objective does the company serve? Winning a named account, retaining existing customers, creating expansion revenue, and serving a broad market require different segments, evidence, and investment ceilings.
- What does the ten-week estimate include? Identity, granular access, version history, notifications, audit export, retention, support, and incident handling may determine the real cost. The estimate must be checked against those boundaries.
- What experience is allowed during the three weeks? Can the study use de-identified deliverables, manual steps, a clickable prototype, or a controlled sandbox? If real data is not allowed, call it research rather than proof of production feasibility.
- Which red lines are non-compensable? Unresolved access control, sensitive-data exposure, incorrect approval, or missing audit evidence cannot be offset by a high click rate. Each red line needs an owner and proof of closure.
30-second answer framework
“I would not treat 12 requests as validated demand. I would deduplicate them, select one target segment, and split the idea into assumptions about problem severity, behavior change, solution usability, technical and risk feasibility, and commercial value. I would rank them by how fatal it would be to be wrong, how uncertain they are, and how cheaply we can learn. I would reconstruct recent approval incidents first, test a clickable prototype second, then use a manually supported real-workflow pilot and budget commitments for behavioral evidence. Before starting, I would freeze the sample, pass gates, access guardrails, and failure actions. If the core problem, repeated use, feasibility, and commercial commitment pass, I build the smallest scope. If evidence holds in one segment, I narrow. If the problem holds but the portal fails, I pivot. If a fatal assumption or red line fails, I stop.”
This opening states the decision sequence before the methods. It avoids presenting a long research checklist as the answer and establishes that validation ends in an investment decision, not a larger pile of positive comments.
Step-by-step deep answer
Start with a validation decision card. It forces the team to say what it believes, what evidence would overturn that belief, and which action follows each result. A minimal version looks like this:
Target segment: professional-services firms with 50–500 employees and weekly external approvals
Problem to test: whether email and PDF approvals cause measurable delay, rework, or audit risk
Fatal assumption: target accounts will move at least one real approval flow into a controlled pilot
Evidence now: 12 forwarded requests from seven accounts, not yet deduplicated or behaviorally validated
Cheapest valid sequence: current-state interview and observation → clickable prototype → concierge real-work pilot
Pass signal: precommitted problem evidence, repeat use, improved outcome, risk closure, and commercial commitment
Failure action: narrow the segment, test another solution, or stop investmentNext, break the idea into an assumption inventory. For each item, label the damage if it is wrong, the strength of current evidence, and the cost of the next test. High, medium, and low are sufficient; a fabricated precision score adds little.
- Problem and frequency: Do target accounts repeatedly face external approval delay, version confusion, or weak accountability? Can recent cases, workflow records, or support tickets reconstruct the current loss?
- Segment and adoption chain: Which accounts hurt most, who initiates, who approves, who buys, and who can block deployment? Do the seven accounts belong to one serviceable segment?
- Solution and behavior: Is an account-light portal better than email, shared documents, or e-signature? Will people move real work rather than click through a demonstration?
- Usability: Do external approvers understand identity checks, version differences, the meaning of approval, and withdrawal rules? Can internal users see the correct state and handle exceptions?
- Feasibility and risk: Can access isolation, audit history, data retention, notifications, and incorrect approvals be handled safely at an acceptable cost?
- Business sustainability: Does value appear through retention, wins, or expansion? Will a buyer approve a paid pilot, a contract add-on, or an explicit budget? Can the company afford incremental support and risk cost?
Use week one for problem validation without selling the portal. Deduplicate the 12 requests by account and role, then inspect sales notes, churn or loss reasons, support tickets, and current collaboration behavior. Recruit 12 accounts from the proposed segment, including requesters, non-requesters, customers, and prospects. Ask each participant to reconstruct the most recent approval: trigger, deliverable, people, steps, wait, rework, error, and result. Observe the current tools and look for costs already paid through manual chasing, extra meetings, or delayed billing.
Those 12 accounts discover mechanisms; they do not estimate a market percentage. A directional gate for this case could be written before recruiting: at least eight accounts can show a real approval task from the last 30 days, at least six demonstrate a repeated and consequential delay, rework, or audit problem, and the problem clusters in a segment the company can reach. The team must agree on the numbers before recruitment, and a real project must change them for its segment, buying cycle, and cost of error. If all evidence comes from one large account, evaluate it as a custom-account business case rather than a general product need.
Use week two for solution and usability evidence. Map the end-to-end journey, then create a clickable prototype covering only the risky steps: an internal user sends a named version, an external client verifies identity, reviews changes, approves or rejects, and both parties see an unambiguous state. Test initiators and approvers from the same account. Include the wrong version, forwarded links, withdrawn approval, and failed notification. Record task completion, material misunderstanding, moderator rescue, and risk concerns.
Passing the prototype shows only that people can understand and complete the task. It does not establish repeated real use or technical safety. If users actually need clearer email versions and reminders rather than another destination, pivot to that smaller concept. Validation exists to find a sufficient solution to the problem, not to defend the initial portal label.
In parallel, engineering, security, legal or compliance, support, and the commercial owner run a red-line review. List identity and authorization, tenant isolation, audit integrity, retention, approval effect, notification delivery, revocation, and dispute handling. Every item needs an owner, status, and closure evidence. An open severe access or data risk blocks a real-data pilot. Strong demand cannot compensate for it.
Use week three for a minimum-commitment pilot. Select six in-segment accounts that passed risk screening and assemble real approval flows from existing safe capabilities, manual operations, and the limited prototype. Tell participants exactly which steps are manual; do not imitate a finished product. Each account brings real work approved for the study and attempts at least three approvals. Record voluntary repeat use, approval turnaround, rework, exceptions, manual service time, and support burden.
Raise the cost of the commercial signal as well. For current customers, request a paid-pilot or expansion letter that states a price range, procurement path after success, and exit condition. For prospects, require the budget owner to participate and approve the next stage. A letter of intent is not revenue, but it carries more information than “sounds useful.” A fake door or landing page can test the value proposition and early intent only; it must disclose that the product is unavailable and must not charge or mislead users.
Freeze the decision table before reading the result. These are case-specific examples shaped by the stated budget, not general benchmarks:
- Build the smallest scope: One defined segment passes the problem gate; at least four of six pilot accounts complete at least three real approvals without continual team prompting; approvers correctly understand version and outcome; no severe risk remains open; at least three budget owners make a conditional paid commitment; and engineering still confirms delivery within the approved investment.
- Narrow: Evidence holds only in one industry, account size, workflow type, or large-account cohort. Build a business case and scope for that segment without generalizing to every customer.
- Pivot: The problem and willingness to change are real, but portal use is weak for a reason inherent to the solution. If users only need an auditable email approval, test that smaller workflow.
- Test once more: One specific, fixable obstacle that can change the conclusion contaminated the result, such as an identity flow that prevented external users from starting. Fix only that causal obstacle while preserving the original gates and budget cap.
- Stop: The team cannot find a repeated consequential problem; users will not move real work; commercial commitment is weak without a diagnosed buying barrier; a feasibility or risk red line cannot close at reasonable cost; or value depends on conditions unique to one account that sales cannot reproduce.
The deliverable includes an evidence ledger. Every assumption, source, evidence-strength label, counterexample, decision, and next owner remains traceable. Interview quotes and praise during a demo are low-cost signals. Real behavior, repeat use, access grants, and budget approval are higher-cost signals. Preserve contradictions between sources instead of erasing them with an average score.
High-quality sample answer
“My initial recommendation is to approve the three-week validation, not to send four engineers directly into a ten-week build. Twelve requests from seven accounts may contain duplicates and sales-process bias, and they do not show whether buyers, internal users, and external approvers will all change behavior.
I would define a target segment first, such as professional-services firms with 50–500 employees and at least one external deliverable approval per week. I would then split the idea into six assumptions: the problem is repeated and consequential; the segment is reachable; a portal beats current alternatives; both sides can use it correctly; access and audit can be safe; and value can produce budget commitment. I rank them by how fatal an error would be, how uncertain we are, and how cheaply evidence can be obtained.
In week one, I deduplicate the 12 requests, review sales notes, support tickets, and current collaboration behavior, then conduct recent-event interviews and workflow observation with 12 target accounts that include requesters and non-requesters. I would not ask whether they like an approval portal. I would reconstruct an approval from the last 30 days, including tools, wait, rework, error, and loss. Interviews reveal mechanisms, not market prevalence. A sample case gate is that at least eight accounts show a recent task and at least six in one serviceable segment demonstrate a repeated consequential problem.
In week two, a clickable prototype tests sending a named version, external identity, approval or rejection, withdrawal, and audit state. That answers comprehension and usability only. Engineering, security, legal or compliance, and support simultaneously review tenant isolation, access, retention, approval effect, notifications, and dispute handling. I do not start a real-data pilot while a severe risk remains open.
In week three, six screened accounts complete real approvals through a manually supported workflow, with every manual step disclosed. Each account attempts at least three approvals. We observe voluntary repeat use, turnaround, rework, exceptions, and service cost. I also ask budget owners for a paid-pilot commitment with a price range and procurement path rather than treating verbal interest as commercial validation.
I freeze the decision before the pilot. If one segment passes the problem gate, at least four accounts complete three real approvals without continual prompting, no material version or outcome misunderstanding appears, severe risks close, at least three budget owners make conditional paid commitments, and the engineering estimate still holds, I recommend the smallest build. If evidence holds in one segment, I narrow. If the problem holds but the portal fails, I pivot to a smaller email or audit solution. I retest only when one diagnosed obstacle contaminated the evidence. A failed fatal assumption or unclosed red line stops the work.
The final deliverable is an evidence ledger, a decision, and a cap on the next investment. Three weeks cannot prove durable product-market fit, but it can prevent the team from spending ten weeks to answer questions that were cheaper to answer.”
Common mistakes
- Counting 12 requests as 12 votes → Requests may repeat an account, a sales opportunity, or a few influential customers → Deduplicate by account, role, and segment, then reconstruct real recent events.
- Showing the solution before asking whether users like it → The concept anchors the interview and courtesy agreement costs nothing → Study current behavior, alternatives, and existing loss before testing the concept.
- Defining an MVP as a smaller build → Identity, access, and audit may remain expensive while the problem is unproven → Buy evidence with prototypes, concierge workflows, and existing safe capabilities.
- Using one method for every assumption → A survey cannot prove workflow behavior, a prototype cannot prove retention, and an interview cannot prove production safety → Match each assumption with the cheapest method that directly tests it.
- Averaging only positive results → One vetoing buyer, severe authorization risk, or unaffordable service cost may kill the project → Manage non-compensable red lines separately and preserve counterexamples.
- Choosing the success metric after results arrive → Any bright spot can be reframed as a pass → Freeze segment, sample, window, gates, guardrails, and failure actions first.
- Generalizing small-sample ratios to the market → Twelve interviews and six pilots can reveal mechanisms and direction, not penetration → State the evidence boundary and test scale later.
- Calling a heavily prompted pilot adoption → It proves high-touch service can push a task, not that scalable product behavior exists → Track manual effort and require repeat use without continual prompting.
- Adding features after failed validation → More scope can hide a failed problem, segment, or value assumption → Narrow, pivot, retest once, or stop according to the failed layer.
Follow-up questions and responses
Follow-up 1: Sales says the company will lose a $1 million contract without this feature and three weeks is too slow. What do you do?
Evaluate the general product decision and the single-account transaction separately. Confirm the revenue amount, close probability, contract term, custom terms, support obligation, and opportunity cost, then require the customer to document the capability, acceptance, and procurement commitment. If transaction value covers development and long-term maintenance, the team can approve it as a custom or design-partner project with bounded scope, data access, and future support. That contract does not prove broad market demand, so run short parallel interviews to look for a productizable segment. Security and authorization red lines still apply regardless of contract value.
Follow-up 2: All 12 interviewees say they need it, but nobody will join a real pilot. How do you interpret that?
There is a gap between expressed attitude and willingness to incur behavioral cost. Check whether legal, security, or migration work makes the pilot unnecessarily expensive, whether participants have a live approval task, and whether interviewees can actually change the process. Remove friction unrelated to the hypothesis, such as doing configuration for them, while preserving evidence-producing commitments: bring real work, invite the real approver, grant required access, and involve the budget owner. If nobody commits after irrelevant friction is removed, interview enthusiasm does not pass the gate.
Follow-up 3: Prototype testing is strong, but security says granular external access will take at least six months. What next?
The feasibility red line invalidates the current scope. Ask whether a smaller safe boundary exists, such as prelisted approvers, one non-downloadable deliverable, short-lived access, and a complete audit trail. The security owner must confirm it; the product manager should not self-approve the risk. If a smaller scope still cannot close the severe risk, stop the portal solution. Preserve the problem evidence and test an alternative that does not expose external data, such as an approval request with an auditable receipt.
Follow-up 4: Pilot usage is high, but no customer will pay extra. Should the team still build it?
Return to the business objective. If the capability measurably improves renewal, win rate, or core-product use, it may create return through the base package rather than an add-on. Validate that path with comparable renewal risk, sales friction, and behavior while including development, risk, and service cost. If there is neither verifiable retention or win value nor direct willingness to pay, high use only shows that the feature is usable; it does not independently justify the investment.
Follow-up 5: A competitor launched a similar feature and leadership wants approval this week. How do you respond?
The launch increases time pressure and provides research material, but it does not prove that this company's target customers need the same solution. Compress the high-risk checks in parallel: deduplicate requests and interview recent events, test the competitor and a low-fidelity workflow, complete the feasibility red-line review, and request real commitments from budget owners. Give leadership two costed, reversible paths: start a bounded scope immediately or spend one week buying key evidence. If leadership chooses immediate investment, document the untested assumptions, maximum loss, and stop point.
Follow-up 6: With only three weeks, why use 12 interviews and six pilots? Where did those numbers come from?
They are planning parameters for this case budget, intended to cover different accounts and roles and observe several instances of real repeat behavior. They are not universal statistical answers. The real count depends on segment heterogeneity, recruitment speed, buying cycle, baseline variance, and the cost of a wrong decision. A homogeneous segment may support rolling small rounds. Estimating conversion or a small effect requires a larger sample derived from the statistical objective. The candidate should state which decision the number serves and which conclusion the evidence cannot support.
Follow-up 7: Four of six pilots pass, but the two largest customers fail. Do you still build according to the gate?
Do not rely on the aggregate alone. Determine whether the two large customers belong to the target segment and whether failure comes from critical authorization, complex approval chains, procurement limits, or incidental execution. If the commercial strategy depends on large accounts, those counterexamples may invalidate the serviceable-segment or cost assumption even though the numeric gate appears to pass. Split results by segment and adoption chain, then choose one coherent target market. Development proceeds only when both the evidence and business objective hold.