Skip to main content

Use case

You suspect AI could help. The question is where.

The question is not whether AI can do something — it often can. It is finding the precise step in your process where it actually changes something, checking that your data allows it, and deciding what happens when the output is wrong.

You may recognise some of this

Situations where AI adds something share a trait: a lot of reading, sorting and searching, and little deciding.

  • Your teams spend time looking for information that already exists somewhere.
  • The useful information sits in documents, emails or PDFs, not in a usable database.
  • Free-text requests have to be sorted before they can be handled.
  • The same questions come back and get re-worded answers every time.
  • Large volumes have to be read to extract a few elements: contracts, reports, minutes.
  • Preparing a decision means gathering scattered pieces before it can even be examined.

What happens if it is done badly

The main risk is not the absence of AI. It is an AI project that consumes budget without ever reaching production.

The prototype impresses and stays a prototype
A demo that works on a few well-chosen examples says nothing about behaviour at real volume, with degraded cases and users in a hurry.
Quality is not measured
Without a reference set, the reliability conversation becomes a matter of impressions. You can neither compare two approaches nor detect a regression.
Responsibility becomes unclear
If nobody has defined who approves what, a system error becomes a problem with no owner. That is what makes a tool get abandoned after a few incidents.
The real cost shows up after go-live
Usage-based calls, higher volumes than expected, manual rework: the running cost gets discovered in production if nobody estimated it beforehand.

Where AI can fit into a process

These are not products but kinds of intervention. Each slots in at a specific place, and each assumes different conditions.

Where AI can fit into a processWhen it is relevantWhat it assumes
Retrieval-augmented searchThe information exists in your documents but nobody finds it fast enough.Accessible, up-to-date documents, and answers that cite their sources so they can be checked.
ClassificationIncoming items have to be routed to the right handling path: requests, tickets, documents.Stable categories and a history of already-classified cases to measure accuracy against.
ExtractionUseful data is locked inside loosely structured documents: invoices, contracts, forms.Consistency checks on the output, and matching against your reference data before anything is written.
Assisted draftingRecurring texts get written by hand from the same elements: replies, minutes, summaries.Your templates and rules as the frame, and human review before anything goes out.
Outlier detectionA significant volume goes through as routine and something has to flag what deserves a look.Accepting that the system flags rather than decides: the output is a queue to check, not a verdict.

Several of these needs are better served without generative AI: extracting fields from a fixed form or classifying on stable criteria is often more reliable and cheaper with a simpler approach.

One assisted step, not a replaced process

The process stays yours. AI comes in where the friction is, and its output goes through a review before continuing on its way.

  1. 01

    Intake

  2. 02

    Qualification

  3. 03

    Assisted preparation

  4. 04

    Human review

  5. 05

    Handling

  6. 06

    Closure

Unchanged stepsAssisted stepReview point

If the assisted step becomes unavailable, the process has to be able to continue as before. That is a condition for going live, not an option.

How we approach it

We start from the process and the data. The technical choice comes last, once the question is framed correctly.

  1. 01

    Find the real friction

    Which step takes time, who does it, how many times a week, and what happens when it is done badly. Without this, you optimise something that did not matter.

  2. 02

    Look at the data you have

    Volume, quality, format, usage rights. Sometimes the conclusion is that the data has to be dealt with first — and that this piece of work is worth more than the AI project itself.

  3. 03

    Build a reference set

    Real cases with the expected answer, validated by your business teams. That is what makes it possible to measure, to compare approaches, and to spot a regression later.

  4. 04

    Evaluate against it

    A measured test, not a demo. The result tells you whether the level is good enough for the intended use, and above which confidence threshold a human check is needed.

  5. 05

    Define the behaviour when it is wrong

    Confidence threshold, review path, fallback to conventional processing, traceability of every output. A system without these is not ready for production.

  6. 06

    Fit it into the existing process

    The assistance appears inside the tools teams already use. A separate interface you have to remember to open gets forgotten within weeks.

  7. 07

    Monitor over time

    Quality measured regularly against the reference set, share of cases sent to review, real running cost. Data drifts and models change versions.

Three questions before adding AI

If any of the three has no answer, the project is not ready — however good the demo was.

  1. 1

    Do we have the right data?

    Accessible, complete enough, up to date, and legally usable for this purpose. A capable model running on inaccurate data just produces errors faster.

  2. 2

    Can we measure whether the output is good enough?

    "Good enough" gets defined before you start, on real cases, with a threshold the business owns. Without that definition, nobody will be able to say whether the project succeeded.

  3. 3

    What happens when the model is wrong?

    Who catches it, at what cost, and what the process does next. An error caught by a review is a planned case; one that reaches a client is an incident.

A prototype is not a product

A prototype that works on twenty examples shows a path exists. It does not show that it will hold in production.

  1. The data changes scale

    At real volume you meet unexpected formats, unreadable documents, and the cases nobody showed during the demo.

  2. Confidentiality becomes a decision

    In production you have to settle where data is processed, what the provider commits not to reuse, and what gets retained. That can change the architecture.

  3. The cost becomes recurring

    A usage-billed call multiplied by the daily volume gives a figure better known before go-live than in the first monthly invoice.

  4. The model version moves

    Providers keep changing their models. Without a reference set and a pinned version, an announced improvement can quietly degrade your specific case.

  5. The fallback has to exist

    Outage, out-of-domain case, quota exceeded: the process has to continue another way. Handling that stops because the AI does not answer is not operable.

What this can look like

Shapes a successful integration takes, always on one specific step rather than the whole process.

  • A search across internal documentation that returns the answer with its sources
  • Automatic sorting of incoming requests, with doubtful cases routed to a human
  • Extraction of data from a document, checked before anything is written
  • Drafting support framed by your templates, reviewed before sending
  • Flagging of unusual cases inside a flow handled as routine
  • A quality dashboard so a drift is seen before users notice it

These describe possible shapes of a solution, not delivered projects presented as references.

Frequently asked questions

Do we need to train our own model?

Rarely, and never as a first step. Approaches built on existing models, combined with your documents and your rules, cover the large majority of business needs. Training a specific model requires a significant volume of annotated data and a continuous upkeep cost — that gets decided after measuring that nothing else is sufficient.

Can we use confidential data?

That depends on the architecture, and it is a decision to take explicitly at scoping. Depending on sensitivity, you use a commercial API with contractual no-reuse commitments, or host the processing in an environment you control. Each option has a cost and constraints that we document before you choose.

How do we measure whether the AI works?

Against real cases whose correct answer has been validated by your business teams, with a threshold agreed in advance. That is the only meaningful measurement: a figure published against a public benchmark says nothing about your situation. That same set then serves to monitor quality over time.

How are model errors handled?

As a planned case. Below a confidence threshold, the case goes to human review rather than passing silently. Decisions that commit the business stay approved by a person. And when the model is unavailable or out of its domain, the process falls back to conventional handling.

Should we start with a POC?

A short evaluation on your real data, yes. A showcase POC on hand-picked examples, no: it proves nothing and consumes time. The difference is the set of cases used, and measuring a result rather than demonstrating a capability.

Let us talk about your process

Tell us which step your teams spend the most time searching or sorting in. That is where the question becomes useful.