Skip to content
AI Workify
Multilateral development bank

Evidence before automation: discovery for a development bank

A multilateral lender contracted a two-month operational and technological discovery. This is the method designed for it: cases before opinions, evidence kept apart from inference, five priorities and no long list.

Work in deliveryDiscovery & operational diagnosisGovernance, security & traceability
5AI solutions to prioritise, each with a full dossierMeasured
2 monthsFixed term for diagnosis, architecture and planMeasured
20–25Practitioner interviews planned, 45–60 minutes eachModelled
6–9Walkthroughs of real case files plannedModelled

Context

A multilateral development bank finances public-sector projects in several countries. By the bank's own account, its loans in execution grew for several years while the teams that prepare and supervise them stayed broadly flat, so the number of operations each project lead carries rose sharply.

The bank's operations leadership wants to absorb further growth without a proportional rise in headcount and without weakening fiduciary quality. It contracted AI Workify for a discovery: an operational diagnosis, a conceptual architecture, security requirements, a prioritised portfolio of five AI solutions and a preliminary implementation plan.

The challenge

Three constraints shaped the design. The first is fiduciary: the bank may not buy speed with process quality, and material decisions stay with accountable people. The second is security: the bank's technology, security and risk functions set their requirements first, as inputs to the design and not as a review at the end.

The third is focus. Conversations before the contract had produced a long list of candidate automations. All of it was hypothesis: we had not yet seen the bank's internal rules, templates or case files. The terms of reference ask for five justified solutions and for reuse of the bank's existing technology before anything new. They exclude any detailed design or build.

Our approach

The method designed for this engagement has rules that do not bend: the contract governs scope, cases come before opinions, facts and hypotheses are never mixed, and existing tools come before new ones. The plan has six moves.

  1. 1Agree the rules at kickoffA 60-minute kickoff records decisions: scope, sponsors, data handling and who validates what. Technology, security and risk attend from day one.
  2. 2Scan wide, at low depthAn opt-in staff pulse, aggregated and never scored by individual, and a first round of interviews test up to eight hypotheses across preparation, supervision and monitoring.
  3. 3Interview from real casesEach interview starts from a recent case and separates rule, practice, workaround and exception. A short synthesis goes back to the interviewee for correction.
  4. 4Walk through the workWalkthroughs of a normal, a complex and an exception case locate where controls sit, what an error costs, and what practitioners say must not be automated.
  5. 5Go deep on fiveFive deep dives, each written up as a dossier: host process, baseline, data, integrations, controls, value, architecture and the basis of any estimate.
  6. 6Probe what could break a quoteTwo or three bounded probes, such as an extraction test on sample documents, answer the unknowns that could multiply cost or time. A probe is not covert development.

Sampling follows a saturation rule, not a quota. Interviewing in an area stops when two consecutive sessions add no materially new variant, and expands when a new dependency could move cost or duration by roughly a fifth.

What the discovery is designed to produce

Evidence kept apart from inference

Every finding is to carry two tags. One records the authority of its source: authoritative record, participant report, our own derivation, or hypothesis. The other records evidence strength, from measured and reproducible down to a gap. Nothing is promoted automatically, and a headline benefit may not rest on the weakest levels alone.

A process atlas instead of wall charts

Exhaustive process notation is out of scope. Each process is instead recorded in a versioned atlas that runs from the institution down to individual rules and fields. Walkthroughs feed it: start event, variants, decisions, roles, systems, controls, hands-on time, waiting and rework. It stays broad at the upper levels and goes deep only on the five candidates. Formal notation is kept for the few sub-processes where a solution would sit.

An internal pilot of the atlas exists as a small web application: our working tool, not a contracted deliverable or an accepted platform.

A master plan that states what is not known

The consolidated document is specified to join the diagnosis, a conceptual architecture aligned with the bank's existing estate, the requirements for running AI models safely, five solution dossiers and a preliminary plan with sequence, dependencies and enabling conditions. Each solution will be placed on a quote-readiness ladder: not quotable, rough range, time-boxed pilot at a fixed price, build with frozen scope, or managed operation.

Where evidence is missing, the gap is not filled with opinion. It becomes one of five explicit outcomes:

  • A versioned assumption, with a named owner.
  • A client obligation, such as system access or a sample of files.
  • A bounded spike or due-diligence task with a price and an exit.
  • A wider cost range with a stated trigger for re-estimating.
  • A recommendation not to proceed.

Where the engagement stands

8 → 5 → 2–3Planned funnel: hypotheses, deep dives, decisive probesModelled

Depth structure set in the discovery plan. Planned figures, not completed work.

3–4Challenge sessions planned to test the shortlistModelled

Planned range from the discovery plan, not a count of sessions held.

4Data-request waves, from non-sensitive to probe samplesMeasured

Count from the data-request protocol. A design artefact, not a record of data received.

There are no outcome figures in this article, and we have not modelled any. A discovery builds and deploys nothing: no process at the bank has been automated through this engagement, and no capacity gain is claimed. The figures above are scope and plan.

What exists today is the signed scope, the method and its gates, the data-request protocol and the dossier template. An accepted deliverable does not yet exist. The impact projection the terms of reference ask for must state the strength of its evidence and may not rest on unmeasured estimates alone.

Governance and risk

The protocol requests data in waves: nothing sensitive first, then representative cases and a minimal baseline, then focal evidence on the five priorities, then probe samples. Every request is logged with purpose, owner, custodian, sensitivity and retention. A do-not-request list rules out full staff rosters, credentials, production data dumps, personal identity or banking data, and intrusive task mining.

A gate is a recorded decision, not a meeting. Process owners validate findings about their own process, and final acceptance is written: silence is not approval. Sensitive identifiers stay out of prompts and exports, and no client data is used to train models.

The work product and its intellectual property belong to the bank, under confidentiality that outlasts the contract. That is why this article stays at the level of method.

What we learned

  • Tag every finding by source and by strength. A figure repeated in three meetings is still a reported figure until someone measures it.
  • Ask for five priorities, not a long list. Each candidate must earn its place with a baseline, an owner and a feasible data path.
  • If a missing variable could multiply cost or time, the honest output is a short, priced risk-reduction phase, not a confident number.

Client identity, locations and identifying details are withheld under confidentiality. Figures are rounded. Measured figures come from engagement records; client-reported figures are attributed, not audited; modelled figures are projections from the engagement's business case and are labelled as such.