# How to prioritize innovation and get past the pilot

> Prioritizing innovation is not ranking ideas by impact and effort. It is deciding which bottleneck unblocks the others before you approve the pilot.

- Author: Marlon Trettin
- Published: 2026-07-14 · Updated: 2026-08-29
- Language: en
- Canonical: https://yowpi.com/en/blog/prioritize-innovation-and-get-past-the-pilot

---

Prioritizing innovation initiatives is not ranking a list of ideas by impact and effort. It is deciding which bottleneck in the operation unblocks the others. The queue stalls when a company picks what is easy to test and easy to show in a meeting, rather than what the operation's architecture requires first.

## TL;DR

- The bottleneck is not in the pilot. It is at scale. Most pilots that work in a demo die on the way into production, because they depend on data nobody governs or a system that does not talk to the rest.
- The right tool on the wrong bottleneck is still waste. [Gartner projects](https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027) that more than 40% of agentic AI projects will be canceled by the end of 2027, on rising cost, unclear business value or inadequate risk control.
- The failure is not bad models. It is systems that do not retain context and do not adapt to the workflow, according to [MIT's NANDA report](https://finance.yahoo.com/news/mit-report-95-generative-ai-105412686.html).
- A prioritized portfolio is a short portfolio. Four to six initiatives with impact, feasibility and an owner defined — not forty rows in a spreadsheet.

## Why doesn't the initiative queue move?

Because it was built as a list of ideas rather than a map of dependencies. Twenty initiatives stacked in a spreadsheet look like twenty opportunities. In practice they are usually a handful of structural bottlenecks appearing under different names.

When you prioritize without seeing that dependency, you approve the initiative that was best presented. It runs, it works in the demo, and it dies on the way into production. It depends on data nobody governs, or on a system that does not talk to the rest.

MIT's NANDA report is direct about this: the failure does not come from bad models, it comes from systems that do not retain context and do not adapt to the workflow ([MIT report](https://finance.yahoo.com/news/mit-report-95-generative-ai-105412686.html)). It is the same sentence said another way: the problem is not the technology. It is the architecture around it.

## Are you prioritizing by bottleneck or by visibility?

That question separates the two groups. The market default is to prioritize by visibility: areas where the test is fast, the risk is low and the result looks good on a slide.

The bottleneck holding the operation back rarely lives where testing is comfortable. It lives in the reconciliation that blocks month-end close and in the customer record four departments fill in four different ways.

There is a second, more expensive deviation: prioritizing by tool. A vendor brings a ready-made agent, leadership gets interested, and the initiative is born backwards, from the solution to the problem. Gartner estimates that only around 130 of the thousands of vendors presenting themselves as agentic have real agentic capability, and uses the term *agent washing* for the practice of rebranding existing products as agents ([Gartner, June 2025](https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027)). Buying the tool before naming the bottleneck is asking for the floor plan after the wall is up.

## Three questions that order the queue

Before scoring impact and effort, answer three questions about each initiative. They are architecture questions, not project management ones.

What does this initiative unblock? If it solves an isolated problem, the gain stops at its own boundary. If it organizes a piece of data or a flow that five other initiatives also need, the return compounds. Prioritize dependency before impact. Automating invoicing looks more attractive than standardizing the customer record, until you notice the first depends on the second.

What does the bottleneck cost today? The current cost, which is observable, not the expected gain, which is a guess. Hours of rework per week, days of delay at close, error rate on orders, time between request and delivery. Without that baseline you cannot prioritize, and afterwards you cannot prove the initiative worked.

Who operates this when the pilot ends? That question eliminates a good part of the queue, and it is good that it does so early. The person's name, the system where the process will live, the data that feeds the decision, who is accountable when it goes wrong. If there is no answer, what you have is an experiment. Experiments are legitimate. They just should not compete for budget with things that need to enter operation.

## The missing criterion: a scale condition defined before the pilot

Most companies define the pilot's success criterion and forget the scale criterion. They are different things. The pilot asks whether it works. Scale asks whether the operation can sustain it.

Write the scale condition before approving the pilot, in one sentence: what has to be true in the operation for this solution to run every day, with everyone, without you nearby. The answer almost always reveals an architecture prerequisite. A piece of data that needs an owner. A process that needs redesigning.

That prerequisite is the initiative that should have been at the top of the queue. And it was not.

## A practical example

The scenario below is hypothetical and does not describe a real client. Its numbers are illustrative premises, and they are stated as such.

A packaging manufacturer with 320 employees has fifteen innovation initiatives in the queue. The one that wins is an AI agent to answer quote requests, because sales is pushing and the vendor's demo impressed leadership.

The pilot runs in six weeks and answers well. Then it meets the real operation. The delivery date promised to the customer depends on plant capacity, which lives in a spreadsheet updated twice a week. The price depends on raw material cost, which the ERP holds with a three-day lag. The agent answers fast and answers wrong. Sales goes back to quoting manually, now with one more system to ignore.

The bottleneck was never quote response time. It was the capacity and cost data, which nobody governed. Had the queue been ordered by dependency, the first initiative would have been to bring capacity and cost into one place. A choice with no glamour at all. And one that unblocks four others on the list.

Our cases show that kind of decision. At [Reatop](/en/cases/reatop), hospital waste control was designed around the people working in the field, with a mobile app that works offline, because hospital basements and disposal areas have no stable connectivity. At [Repap On](/en/cases/repap-on), document management started from the company's process map, not from the software. In both, the operation's real constraint entered the architecture before it became code.

## Frequently asked questions

**I can't justify the investment right now. Where do I start?**

Start with the initiative that has an observable baseline and high dependency, even if the gain looks modest. It is easier to defend a two-day reduction in the accounting close, with before-and-after measurement, than a hypothetical productivity gain. What convinces a board is the quality of the measurement, not the size of the promise.

**The implementation will take too much of my team's time. How do I reduce that?**

Reduce the number of simultaneous initiatives before reducing the scope of each one. A prioritized portfolio has few open fronts, with an owner and a deadline, not forty rows in a spreadsheet. The team is not overloaded by the size of the projects, it is overloaded because they are all open at once and none of them closes.

**What if the initiative doesn't actually remove the bottleneck?**

That risk exists, and it shrinks when you measure the bottleneck before attacking it. If you do not know how many hours a week the team spends on rework, you will not know whether the solution worked either. The NANDA report finds that 95% of companies get no financial return from their generative AI projects ([MIT report](https://finance.yahoo.com/news/mit-report-95-generative-ai-105412686.html)). What separates the remaining 5% has less to do with the model chosen and more to do with the discipline of measuring and redesigning the process around it.

## The queue is a symptom, not an agenda

Fifteen stalled initiatives say less about a lack of execution and more about the architecture holding the operation up. Order by dependency and measure the bottleneck before attacking it. The queue gets shorter because a good part of it was never an initiative: it was a dependency with the wrong name.

If you have a list like that on your desk, the [Operational Architecture Diagnostic](/en/contact) is a 30-minute conversation to map which bottlenecks unblock the others in your operation.
