Knowledge base

The 45-day pilot: how to prove in six weeks what AI delivers

A 45-day pilot takes one management question, records a baseline measurement up front and builds one working application on your own sources in six weeks. On day 45 there is a factual decision: sharpen, stop or scale up. No year long program, no report full of recommendations, but measurable proof.

What is a 45-day pilot?

A 45-day pilot is a bounded experiment with one goal: to show whether AI adds value at your leadership table, measured against a baseline recorded up front. You change nothing in your systems. You connect sources, build one application and test it with real management questions.

Six weeks is not an arbitrary term. It is long enough for two full management cycles and short enough to hold everyone's attention. Anything that runs longer quietly turns into a program.

What does the 45-day pilot look like, phase by phase?

  1. Day 1: decision scan. Ten questions, three minutes, a personal report right away. It surfaces your biggest decision bottlenecks, plus a first estimate of available information, feasibility and value. This is the basis for choosing the question.
  2. Day 2 to 7: diagnosis and baseline measurement. We speak with three to five people, look into the sources and record how things work today. How many days sit between signal and decision? How many hours of digging come before that? How complete and how current are the sources? By the end of this week one question is fixed, along with the standard and the measuring points.
  3. Day 8 to 30: build. One application, one agent, on your own data. We read sources, we write nothing back into them. Halfway through, around day 18, there is a working version the leadership team can already look at.
  4. Day 30 to 44: prove. Real users, real management questions. Every outcome gets tested: is the number right, does it trace back to the source, and would you have known this without the application? Findings that do not hold up come off the list.
  5. Day 45: decide. One two hour session with the leadership team. The outcomes sit next to the baseline measurement. Three possible decisions: sharpen, stop or scale up.

What exactly do you measure?

Four quantities, all recorded up front. What you did not measure beforehand, you cannot claim afterward.

WhatBaseline measurement (example)After 45 days
Decision lead time19 working days from signal to decision6 working days
Research time per decision11 hours, spread across 4 people2 hours
Validated value opportunitiesNot in view€ 54.000 in unbilled additional work
TraceabilityUnknown or not recorded100% of the numbers traced back to the source

Validated value opportunities come in three kinds: margin leaks (work you deliver but do not invoice), missed revenue (opportunities picked up too late or not at all) and avoidable costs (rush shipments, duplicate work, penalties for late delivery). An opportunity only counts once someone responsible puts their name under it. An estimate from a model does not count.

Decision quality is measured differently. You test it by putting three decisions from the past year back to the application. Would you have decided the same way with this information? For two of the three decisions the answer should be ‘yes, but faster and with more certainty’. For at least one decision a new insight should come to the surface that you missed at the time.

When do you stop the pilot?

Stopping is a normal outcome, not a failure. You stop in four situations.

  • The sources are too thin. If hours are structurally logged three weeks late, no application can repair that. Process first, then technology.
  • Decision lead time drops by less than a third. Then the bottleneck is not information but decision making or mandate. That is a governance question.
  • The value found is smaller than the upkeep. If it produces € 15.000 of value a year and monitoring it costs half a day a week, the math does not work.
  • There is no owner. Without a leadership team member who looks at it every week, any application disappears within a quarter.
A pilot that stops honestly on day 45 is cheaper than a program that runs for two years because nobody dares to intervene.

What does embedding mean, and when do you scale up?

Embedding is the dull step that makes the difference. The application gets a fixed place on the leadership team agenda, an owner, a quarterly review moment and an agreement on what happens when a signal comes in. Without those four, the result evaporates.

Opschalen doe je pas als de eerste toepassing minstens één volle cyclus zonder hulp heeft gedraaid. Daarna komt de volgende toepassing op dezelfde digitale basis, niet ernaast. Vanaf ongeveer vijf toepassingen ontstaat een Digital Business Twin: een samenhangende managementlaag over je hele bedrijf.

What does it ask of you and your team?

Reken op ongeveer 14 uur MT-tijd over zes weken: twee uur in week één, een uur per twee weken, en twee uur op dag 45. Daarnaast is er één inhoudelijk betrokken medewerker nodig voor ongeveer drie uur per week, plus iemand van IT voor twee tot vier uur in totaal voor de toegang tot bronnen. Welke afspraken je vóór de start vastlegt, staat in AI at the table, people at the helm. Wat de toepassing feitelijk doet, lees je in What are AI agents, en waarom dit iets anders is dan rapportage in From BI to Management Intelligence.

Frequently asked questions

Why 45 days and not three months?
Six weeks covers two full management cycles, enough to see a pattern and test the result. In practice, longer means less sharpness: the question broadens, attention drops and the decision moment slips. The hard end date is exactly what makes the pilot useful.
Do we have to change anything in our ERP or CRM?
No. During the pilot we read from sources and write nothing back into them. You do not have to migrate, switch vendors or build a data warehouse. If data is missing, that is an outcome of the diagnosis, not a condition up front.
What happens if the pilot delivers nothing?
Then we stop on day 45 and you keep the baseline measurement, the diagnosis and a concrete picture of what is wrong with your information supply. That is valuable in itself: you then know where the real bottleneck sits, and that is often process or mandate rather than technology.
Who from our company has to be involved?
One leadership team member as owner of the question, one employee with subject knowledge who knows the sources and the rules, and someone from IT for access. More people make it slow. Fewer people make it unreliable.

De pilot begint met tien vragen over jouw beslisknelpunten. Start the decision scan: 10 vragen, 3 minuten, direct een persoonlijk rapport.