ZROEVOKE

Eleven Weeks to a Live Decisioning Engine

A week-by-week account of one engagement — including the two weeks we lost, and what we would do differently.

Six bars stepping down from left to right like a schedule, the third orange and the last cyan, ending at a vertical line marking the launch.

Eleven weeks from kickoff to the first real application scored in production. Not a pilot, not a shadow run: a decision that affected a customer. Here is where the time went.

Weeks 1–2 · Scope and Data Reality

We spent the first fortnight doing two things: writing down the decision in one page, and opening every data source we were promised. Half of them were not what the documentation said. That discovery belongs at the start — a data contract signed in week two is worth a month later.

Weeks 3–4 · The Path to Production, Before the Model

Before any modelling, we built the deployment path end to end with a hard-coded decision: request in, decision out, logged, monitored, reversible. It scored nothing useful, but it proved the plumbing and gave every subsequent change a place to land.

Build the boring path first. A model with nowhere to go is a demo; a path with a placeholder in it is a product waiting for its brain.

Weeks 5–6 · The Two Weeks We Lost

We built the first scorecard against a training set that had been filtered upstream — the rejected applications were missing. Every metric looked excellent and every metric was meaningless. We caught it when the approval rate in shadow mode came out twenty points above the existing policy.

The lesson was not "check your data", which everyone already knows. It was that we had no automated check comparing the shape of our training population against the live population. We added one that afternoon. It has caught two similar problems since, on other projects.

Weeks 7–8 · Shadow Mode

Two weeks running alongside the incumbent policy on live traffic, deciding nothing. This is where the disagreements surface — and where the credit team stops treating the system as a threat, because they can see every case where it differs and say why.

Weeks 9–10 · Committee and the Pack

The assurance pack was assembled as we went, so this was a review rather than a scramble. One round of questions, one round of changes: a tighter out-of-distribution fallback and a clearer override trail.

Week 11 · Live, on a Slice

We went live on a narrow segment with a low limit, watched it for a week, then widened. Nobody remembers a cautious launch. Everybody remembers a bad one.

What We Would Change

  • Run the population-shape check from day one, not week six.
  • Start shadow mode a week earlier, even with a worse model — the conversations it starts are worth more than the accuracy.
  • Write the override path before the scoring logic. It is the part the business actually asks about.
Shahaf Lavi Founder, Zero Evoke
Contact Us →
Continue