Skip to main content
Viithiisys
Generate, retrieve, or neither

Decide what generative AI should write, draft or answer in your business, and what it must not.

The consulting question is not which model. It is which work can tolerate an output that is usually right, who checks it, and what happens the first time it is confidently wrong.

The output is a specification and a test setNineteen years of delivery, across six countries

What a pilot provesAnd what it does not
Production

Whether the organisation can operate it

The pilot

Whether the technology can do the work

Those are different questions

The pilot is not the hard part.

Four situations

Four situations that produce this call.

This is the wrong engagement when the work needs retrieval of a known fact rather than generation of new text.

  1. A demo worked and nobody can say whether to ship it

    Someone built something impressive in an afternoon. The question of whether it is safe to put in front of customers has no owner and no test.

  2. The output is plausible and occasionally wrong

    Which is harder to manage than output that is obviously wrong, because nobody catches it in review.

  3. Everyone is asking for it and nobody has named the job

    The board, the competitors and the team all want generative AI. No specific piece of work has been identified where it would help.

  4. A pilot is live and cannot be evaluated

    It produces drafts, people use some of them, and there is no way to say whether quality is improving, declining or stable.

A lookup problem answered by a generative system is slower, more expensive and less reliable than a lookup.

See LLM Development
What gets settled

The five decisions the engagement settles

Decide which work generative AI can take on safely: the acceptance standard, the evaluation set, the review boundary and what is deliberately excluded.

  1. Where generation earns its place

    Work where producing a first draft is the expensive part, and where a person was always going to review the result. That is the shape that pays.

  2. What acceptable output means, concretely

    Written down well enough that two people would score the same output the same way. Without this, quality becomes a matter of opinion after launch.

  3. Which sources the system may draw on

    Approved, current and permitted. Also what happens when the sources do not cover the request, which is the answer most projects never define.

  4. Who reviews, and what they see

    The reviewer needs enough context to decide quickly. A review queue that requires reconstruction gets rubber-stamped within a fortnight.

  5. What is deliberately excluded

    The work generation will not touch, recorded with the reason. This is the shortest section of the output and the most useful one later.

Three treatments

Generate, retrieve, or neither. The test is what a wrong answer costs and who sees it first.

Most pages treat generative AI as the answer and differ only on the capability list. The better question is which of three treatments the work needs; two are cheaper.

TreatmentWhat a person still does
  1. Producing a first draft somebody was always going to edit

    Treatment: Generate

    What a person still does: Edits and approves, as before

    The draft is the expensive part and review was already in the process

  2. Answering a question with one correct answer

    Treatment: Retrieve, do not generate

    What a person still does: Handles the cases retrieval cannot find

    Generation can paraphrase a fact into something subtly wrong. Retrieval either finds it or does not

  3. Summarising a long document for a decision

    Treatment: Generate, with the source visible

    What a person still does: Reads the source when the stakes are high

    Summarising is genuine value. Its risk is a quiet omission, so the reader must be able to check

  4. Classifying incoming work into defined types

    Treatment: Generate or classify, either works

    What a person still does: Reviews the uncertain cases

    Bounded output, checkable against a list, low consequence if wrong

  5. Anything a customer receives unreviewed

    Treatment: Neither, yet

    What a person still does: Everything, until an evaluated threshold exists

    The first confidently wrong output is a public one, and the cost is not proportional to the saving

  6. A regulated statement, a price or a commitment

    Treatment: Never generate

    What a person still does: Authors it

    The consequence is legal or financial rather than reputational

Each bar is divided by which of the three treatments applies, not by a measured proportion. No figure here is published.

Two of six rows say do not use it, and one says use something simpler. That is roughly the real distribution, and it is the reason this page exists as consulting rather than as a build pitch.

How the work runs

How the work runs, step by step

Five steps, in the order they have to happen.

  1. Start from the work, not the model

    Which task, done by whom, how often, and what the current output costs to produce. A model choice made before this is a guess wearing a specification.

  2. Write the acceptance standard

    What good output looks like, specific enough to score. This is the single highest-value artefact of the engagement and the one most often skipped.

  3. Build the evaluation set from real inputs

    Twenty to fifty actual cases with expected outcomes, including the awkward ones. Without it, no later change can be judged.

  4. Set the review boundary and the stop condition

    What proceeds automatically, what needs approval, and what the system must decline. Made now it is policy; made during a release it is a negotiation.

  5. Route to a build, or recommend against one

    With dependencies named. Where the honest answer is retrieval, or nothing yet, that is the recommendation.

Evidence

Where this work has been done

Production AI is where this discipline is proved, and it is where Viithiisys has built.

  1. Fitelo, evaluation and cost controls in production

    An AI product operating with evaluation and cost controls in place rather than shipped and watched, which is exactly what steps two to four above set up.

    Directly relevant here.

  2. Conscious Chemist, a customer-facing product experience

    A documented 38% faster product-question response.

    Evidence of customer-facing AI work, not a consulting result.

  3. Nineteen years, six countries, since 2007

    Nineteen years of delivery, across six countries.

Pilot to production

The pilot is not the hard part, and that is why so many stop there.

A generative AI pilot is cheap to build and hard to promote into production. The reasons are consistent, and four of the five are decisions rather than engineering problems.

  1. Nobody agreed what good means.

    Why it is not obvious at pilot stage

    During a pilot the builder judges the output, and the builder is favourably disposed. At production scale the reviewer is someone else, with different standards and no context

    What resolves it

    An acceptance standard written before the build, specific enough that two people score the same output identically

  2. There is no way to tell whether a change helped.

    Why it is not obvious at pilot stage

    A pilot is judged on impressions of a handful of outputs. A production system needs to survive a model update, a prompt change and a source change

    What resolves it

    An evaluation set of real inputs with expected outcomes, held by the client

  3. The review step costs more than the generation saved.

    Why it is not obvious at pilot stage

    Pilots skip review, or the builder does it. In production someone senior is checking output, and that is often the most expensive time in the process

    What resolves it

    Design the reviewer's view first: enough context to decide in seconds, not a queue requiring reconstruction

  4. No one owns the failure case.

    Why it is not obvious at pilot stage

    Pilots have no real failures because there are no real users. Production has them in week one

    What resolves it

    A named owner for uncertain and out-of-policy cases, agreed before launch rather than after the first incident

  5. The cost per result was never measured.

    Why it is not obvious at pilot stage

    At pilot volume the cost is negligible and nobody looks. At production volume it is a line item

    What resolves it

    Measure it during the pilot, when it does not matter, so the number exists when it does

A pilot tests whether the technology can do the work; production tests whether the organisation can operate it. This engagement answers the second.

What you keep

What you are left holding.

The output is an acceptance standard and an evaluation set, both of which are portable.

  1. The acceptance standard, written down

    What good output means for this work, in a form a new team member can apply. Portable across any model you later choose.

  2. The evaluation set, in your possession

    Real inputs with expected outcomes. This is the asset that makes every future model change a measured decision rather than a hope.

  3. The decisions, with their reasons

    Including what was excluded and why, so a later reversal is deliberate rather than accidental.

  4. No dependency on a vendor choice

    The output is a specification and a test set, not a configuration inside somebody's platform.

Start here

Tell us the piece of writing or answering your team repeats most.

That is enough to judge whether generation, retrieval or neither is the right treatment. All three answers are common, and the second two are cheaper.

FAQ

Questions we are actually asked

Pick a topic, or ask us directly. We answer every inbound within one business day.

Still have questions?

Talk to a senior engineer, not a bot.

Talk to us