Eleven Weeks to a Live Decisioning Engine
A week-by-week account of one engagement — including the two weeks we lost, and what we would do differently.

Eleven weeks from kickoff to the first real application scored in production. Not a pilot, not a shadow run: a decision that affected a customer. Here is where the time went.
Weeks 1–2 · Scope and Data Reality
We spent the first fortnight doing two things: writing down the decision in one page, and opening every data source we were promised. Half of them were not what the documentation said. That discovery belongs at the start — a data contract signed in week two is worth a month later.
Weeks 3–4 · The Path to Production, Before the Model
Before any modelling, we built the deployment path end to end with a hard-coded decision: request in, decision out, logged, monitored, reversible. It scored nothing useful, but it proved the plumbing and gave every subsequent change a place to land.
Weeks 5–6 · The Two Weeks We Lost
We built the first scorecard against a training set that had been filtered upstream — the rejected applications were missing. Every metric looked excellent and every metric was meaningless. We caught it when the approval rate in shadow mode came out twenty points above the existing policy.
The lesson was not "check your data", which everyone already knows. It was that we had no automated check comparing the shape of our training population against the live population. We added one that afternoon. It has caught two similar problems since, on other projects.
Weeks 7–8 · Shadow Mode
Two weeks running alongside the incumbent policy on live traffic, deciding nothing. This is where the disagreements surface — and where the credit team stops treating the system as a threat, because they can see every case where it differs and say why.
Weeks 9–10 · Committee and the Pack
The assurance pack was assembled as we went, so this was a review rather than a scramble. One round of questions, one round of changes: a tighter out-of-distribution fallback and a clearer override trail.
Week 11 · Live, on a Slice
We went live on a narrow segment with a low limit, watched it for a week, then widened. Nobody remembers a cautious launch. Everybody remembers a bad one.
What We Would Change
- Run the population-shape check from day one, not week six.
- Start shadow mode a week earlier, even with a worse model — the conversations it starts are worth more than the accuracy.
- Write the override path before the scoring logic. It is the part the business actually asks about.
Shahaf Lavi
Founder, Zero Evoke