Start with the problem your organisation needs to solve. AI can help work move faster, but it can also increase delays or spread errors when decisions are unclear or information is unreliable. Choosing a tool before understanding the problem can make the wrong work easier to repeat.
We investigate with the people who own and experience the problem, then recommend whether to test an AI approach, make a different change or stop.
What one investigation produces
Discovery is priced per problem investigated. For each problem, you keep:
- A problem record: what happens today, who owns the outcome and who is affected.
- A baseline: the available evidence of delay, cost, errors or another relevant outcome, with gaps made explicit.
- The context: relevant knowledge, constraints and dependencies, including what is missing or unreliable.
- Alternatives: plausible explanations and options, including changes that do not involve AI.
- A recommendation: the next test or decision, its owner, what improvement would count and what would make us stop.
A promising hypothesis is not yet a proven return. The recommendation distinguishes what the investigation established from what still needs testing.
We do not have a baseline yet
You can start with work that matters and the people who experience it. We help identify who owns the outcome, what should improve and what can be observed now. An initial sample of real work can show where people wait, correct errors or repeat effort. We record what that sample cannot tell us as well as what it shows.
Agree what acceptable work looks like and which information may be used before testing a change. Choose an observation period and decision criteria that fit the work. You do not need a mature dashboard or a calculated return before asking for help.
An illustrative example
Suppose customer requests repeatedly wait for an answer. AI might help retrieve an approved answer, but the delay could also come from unclear ownership or a policy decision nobody is authorised to make.
We would examine actual requests and where they waited, identify the owner, and compare a retrieval test with an ownership or routing change. This is an example of the investigation, not a claim about a client’s results.
To judge a test, look at whether the customer’s issue was resolved correctly, elapsed time including checking and correction, customer effort and total work or cost. Record information-boundary failures and other harms separately. Time saved does not make those failures acceptable.
How will we know whether the change helped?
Before a test starts, name who will review the result and when. Compare the selected outcome and downstream effects with the starting evidence. Ask whether the work improved enough to continue, needs adjustment, or should stop.
Logins, prompts and token counts describe activity. They do not establish that customers received a better answer or that work reached them sooner. Changes in staffing, demand or the kind of work may also explain a difference; a before-and-after comparison alone cannot isolate AI’s contribution.
Discovery defines the next test and decision. Your named owner carries the later review unless we agree further support explicitly. Ongoing measurement is not an automatic part of this engagement.
What it costs
Per problem, fixed, agreed before work starts. No time tracking. If the engagement fails to meet the agreed standard, we refund the fee.
Where it leads
The evidence may support agentic engineering, another improvement to how work gets done, or a decision to make no further investment. You keep the findings whichever route you choose.
Sometimes the next step is a bounded assistant used by a person. Agree its permitted information, how the work will be checked and who will maintain the context before extending its use. A learning exercise becomes an operational responsibility when others begin relying on it. We scope that follow-through around the problem; it is not a fixed programme every organisation must buy.