# Rapid prototyping: validating before you invest

> How to validate a system or automation idea in weeks, with a hypothesis, a success criterion and real data, before committing budget and people to it.

- Author: Marlon Trettin
- Published: 2026-07-21 · Updated: 2026-08-29
- Language: en
- Canonical: https://yowpi.com/en/blog/rapid-prototyping-validate-before-investing

---

To validate a system or automation idea before investing, build a prototype that answers a single business question in two to four weeks, with real data and with the people who will operate it day to day. Without a written hypothesis and a success criterion defined up front, the prototype becomes an eternal pilot.

## TL;DR

- Globally, 95% of generative AI pilots produce no measurable impact on the bottom line, according to [an MIT study reported by Fortune](https://finance.yahoo.com/news/mit-report-95-generative-ai-105412686.html). The rare problem is the model; the common one is a pilot born outside the real workflow.
- A prototype is an experiment: one hypothesis, one success criterion, one deadline. Without all three, it is just a small project with no owner.
- Discarding an idea in week three costs little. Discovering in production that it does not work costs the year's budget.
- Tools bought from specialist vendors or built in partnership showed [around a 67% success rate, against a third of that for internal builds](https://finance.yahoo.com/news/mit-report-95-generative-ai-105412686.html). The prototype is also where you test that choice.

## Why do so many pilots die before becoming systems?

Because most are born with no question to answer. The pilot gets approved on enthusiasm, runs alongside the operation, and months later nobody can say whether it worked.

MIT's diagnosis points the same way. In [The GenAI Divide, from the NANDA project](https://finance.yahoo.com/news/mit-report-95-generative-ai-105412686.html), 95% of corporate generative AI pilots produced no measurable effect on revenue or cost. The cause identified was not model quality, it was failed integration: tools that do not learn the company's workflow, and pilots disconnected from the people who operate.

That is an architecture problem in the experiment, not a technology problem. The pilot that fails usually tests the tool. The prototype that works tests a decision: is it worth investing here or not?

## What does a prototype need to answer?

One question, formulated before any configuration. Three elements define a serious prototype and fit on half a written page.

First, the hypothesis. "If we automate order intake, the time between order and invoicing halves." Specific, tied to a metric leadership recognizes, and falsifiable: there has to be a result capable of knocking it down.

Second, the success criterion with a baseline. Measure the current process before building anything. Without indicators agreed from the start, the conversation about scaling becomes a contest of opinions rather than a reading of evidence.

Third, the deadline with a decision booked. Two to four weeks of testing and a meeting on the calendar with three possible exits: scale, adjust and test again, or discard. A pilot with no decision date does not end; it fades.

## How do you build a prototype in weeks, not months?

By cutting scope, not the quality of the answer. The prototype covers a small slice of the process, but covers it for real: real data and real users, inside the flow that exists today. A demo screen with fictional data validates the aesthetics, not the decision.

The technical base helps. Visual development platforms are where a functional flow stands up in days. For a prototype, the platform matters less than the discipline: the same test run in a well-designed spreadsheet sometimes answers the question.

Two rules save months. Involve whoever operates from day one, because that person surfaces the exception that breaks the flow. And resist expanding scope mid-test: every extra feature dilutes the answer to the original question.

There is an MIT figure worth the attention of anyone deciding to build in-house: tools bought from specialist vendors or built in partnership showed [around a 67% success rate, against a third of that for internal builds](https://finance.yahoo.com/news/mit-report-95-generative-ai-105412686.html). The prototype is also the moment to test that choice of path, before it gets expensive.

## How does this work in practice?

A hypothetical scenario, built on a pattern that repeats. A distributor with 40 employees receives orders through messaging apps, and two assistants type everything into the ERP. Leadership is considering an automation project with a six-figure budget, and the innovation manager has to decide whether to defend the investment.

Instead of approving the whole project, he writes the hypothesis: "a structured form integrated with the ERP cuts typing time by 70% and eliminates order errors." He measures the baseline for a week: 4 minutes per order, 6% with an item or quantity error. He builds a no-code flow for a single product line, with one salesperson and one assistant genuinely using it.

Three weeks later, the decision meeting has numbers: time per order, error rate, and the exceptions the form did not cover. If the hypothesis holds, the budget request goes to leadership with proof rather than a promise. If it does not hold, the company spent weeks and a few thousand to avoid a project costing hundreds of thousands.

That logic of narrowing scope shows up in [Nuleite's micro-SaaS](/en/cases/nuleite), a lean tool focused on a single problem — the cost of feeding a dairy herd — built in no-code to reach the people who use it quickly. [UniTrust](/en/cases/unitrust) followed the same principle at another scale: the first version of the system was born in no-code to get into production fast, and the definitive architecture only came after years of real use.

## And when the prototype says no?

That is one of the best possible results. A prototype that knocks down the hypothesis in three weeks returns budget and people to the next initiative in the queue. A culture that treats this as failure pushes managers to stretch dying pilots, and that is how you arrive at MIT's 95%.

Document what the test revealed: the exception nobody had mapped, the data the operation does not collect. That record is an architecture asset. The second prototype on the same process starts from the first one's learning, not from zero.

The sentence that sums it up: you do not have a technology problem, you have an architecture problem. The prototype costs weeks and reveals which one is yours; production costs the budget and reveals the same thing, too late.

## Frequently asked questions

**Won't implementing a prototype take too much of my team's time?**

A well-scoped prototype asks for two to four hours a week from one or two people in the operation, over three or four weeks. If the design requires more than that, the scope is too large for a prototype. The time invested comes back in the decision: it is cheaper than months of pilot with no conclusion.

**How do I measure the result of a prototype if the process was never measured?**

By measuring the baseline before building. A week timing the current process, even by hand, already produces the reference number. Without that prior measurement, any result from the prototype has nothing to compare against, and the conversation with leadership goes back to opinion.

**What if the prototype works in the test but doesn't remove the bottleneck at scale?**

That risk exists and the prototype reduces it rather than eliminating it. That is why the test uses real data and real users, and why the decision to scale sets conditions: which exceptions need coverage, which integrations are missing, what volume the design withstands. Scaling is a second project with the first one's answers, and it is in that crossing that the architecture of the definitive system gets decided.

## The next step

If your queue of initiatives has more ideas than evidence, a well-designed prototype converts one into the other in a few weeks. And if you want help designing that test inside your operation, the [Operational Architecture Diagnostic](/en/contact) is a 30-minute conversation to map which question your next prototype needs to answer.
