Decide what generative AI should write, draft or answer in your business, and what it must not.
The consulting question is not which model. It is which work can tolerate an output that is usually right, who checks it, and what happens the first time it is confidently wrong.
The output is a specification and a test setNineteen years of delivery, across six countries
Whether the organisation can operate it
Whether the technology can do the work
Those are different questions
The pilot is not the hard part.
Four situations that produce this call.
This is the wrong engagement when the work needs retrieval of a known fact rather than generation of new text.
A demo worked and nobody can say whether to ship it
Someone built something impressive in an afternoon. The question of whether it is safe to put in front of customers has no owner and no test.
The output is plausible and occasionally wrong
Which is harder to manage than output that is obviously wrong, because nobody catches it in review.
Everyone is asking for it and nobody has named the job
The board, the competitors and the team all want generative AI. No specific piece of work has been identified where it would help.
A pilot is live and cannot be evaluated
It produces drafts, people use some of them, and there is no way to say whether quality is improving, declining or stable.
A lookup problem answered by a generative system is slower, more expensive and less reliable than a lookup.
See LLM DevelopmentThe five decisions the engagement settles
Decide which work generative AI can take on safely: the acceptance standard, the evaluation set, the review boundary and what is deliberately excluded.
Where generation earns its place
Work where producing a first draft is the expensive part, and where a person was always going to review the result. That is the shape that pays.
What acceptable output means, concretely
Written down well enough that two people would score the same output the same way. Without this, quality becomes a matter of opinion after launch.
Which sources the system may draw on
Approved, current and permitted. Also what happens when the sources do not cover the request, which is the answer most projects never define.
Who reviews, and what they see
The reviewer needs enough context to decide quickly. A review queue that requires reconstruction gets rubber-stamped within a fortnight.
What is deliberately excluded
The work generation will not touch, recorded with the reason. This is the shortest section of the output and the most useful one later.
Generate, retrieve, or neither. The test is what a wrong answer costs and who sees it first.
Most pages treat generative AI as the answer and differ only on the capability list. The better question is which of three treatments the work needs; two are cheaper.
Producing a first draft somebody was always going to edit
Treatment: Generate
What a person still does: Edits and approves, as before
The draft is the expensive part and review was already in the process
Answering a question with one correct answer
Treatment: Retrieve, do not generate
What a person still does: Handles the cases retrieval cannot find
Generation can paraphrase a fact into something subtly wrong. Retrieval either finds it or does not
Summarising a long document for a decision
Treatment: Generate, with the source visible
What a person still does: Reads the source when the stakes are high
Summarising is genuine value. Its risk is a quiet omission, so the reader must be able to check
Classifying incoming work into defined types
Treatment: Generate or classify, either works
What a person still does: Reviews the uncertain cases
Bounded output, checkable against a list, low consequence if wrong
Anything a customer receives unreviewed
Treatment: Neither, yet
What a person still does: Everything, until an evaluated threshold exists
The first confidently wrong output is a public one, and the cost is not proportional to the saving
A regulated statement, a price or a commitment
Treatment: Never generate
What a person still does: Authors it
The consequence is legal or financial rather than reputational
Each bar is divided by which of the three treatments applies, not by a measured proportion. No figure here is published.
Two of six rows say do not use it, and one says use something simpler. That is roughly the real distribution, and it is the reason this page exists as consulting rather than as a build pitch.
How the work runs, step by step
Five steps, in the order they have to happen.
Start from the work, not the model
Which task, done by whom, how often, and what the current output costs to produce. A model choice made before this is a guess wearing a specification.
Write the acceptance standard
What good output looks like, specific enough to score. This is the single highest-value artefact of the engagement and the one most often skipped.
Build the evaluation set from real inputs
Twenty to fifty actual cases with expected outcomes, including the awkward ones. Without it, no later change can be judged.
Set the review boundary and the stop condition
What proceeds automatically, what needs approval, and what the system must decline. Made now it is policy; made during a release it is a negotiation.
Route to a build, or recommend against one
With dependencies named. Where the honest answer is retrieval, or nothing yet, that is the recommendation.
Where this work has been done
Production AI is where this discipline is proved, and it is where Viithiisys has built.
Fitelo, evaluation and cost controls in production
An AI product operating with evaluation and cost controls in place rather than shipped and watched, which is exactly what steps two to four above set up.
Directly relevant here.
Conscious Chemist, a customer-facing product experience
A documented 38% faster product-question response.
Evidence of customer-facing AI work, not a consulting result.
Nineteen years, six countries, since 2007
Nineteen years of delivery, across six countries.
The pilot is not the hard part, and that is why so many stop there.
A generative AI pilot is cheap to build and hard to promote into production. The reasons are consistent, and four of the five are decisions rather than engineering problems.
Nobody agreed what good means.
Why it is not obvious at pilot stage
During a pilot the builder judges the output, and the builder is favourably disposed. At production scale the reviewer is someone else, with different standards and no context
What resolves it
An acceptance standard written before the build, specific enough that two people score the same output identically
There is no way to tell whether a change helped.
Why it is not obvious at pilot stage
A pilot is judged on impressions of a handful of outputs. A production system needs to survive a model update, a prompt change and a source change
What resolves it
An evaluation set of real inputs with expected outcomes, held by the client
The review step costs more than the generation saved.
Why it is not obvious at pilot stage
Pilots skip review, or the builder does it. In production someone senior is checking output, and that is often the most expensive time in the process
What resolves it
Design the reviewer's view first: enough context to decide in seconds, not a queue requiring reconstruction
No one owns the failure case.
Why it is not obvious at pilot stage
Pilots have no real failures because there are no real users. Production has them in week one
What resolves it
A named owner for uncertain and out-of-policy cases, agreed before launch rather than after the first incident
The cost per result was never measured.
Why it is not obvious at pilot stage
At pilot volume the cost is negligible and nobody looks. At production volume it is a line item
What resolves it
Measure it during the pilot, when it does not matter, so the number exists when it does
A pilot tests whether the technology can do the work; production tests whether the organisation can operate it. This engagement answers the second.
Three pages could answer a generative AI question. Here is which one you want.
This is worth being explicit about, because the distinction is not obvious from the page titles and choosing wrongly wastes a first conversation.
- Should we be using AI at all, and whereAI Consulting & StrategyBroader than generation. Covers automation, retrieval, agents and the case for doing nothing yet
- We know generation is the answer. Build itGenerative AI DevelopmentThe build itself, once the decisions on this page are settled
- Can generation do this work safely, and how would we knowGenerative AI Consulting · you are hereThe generative-specific decision: acceptance standard, evaluation, review boundary and exclusions
- It is live and we cannot tell whether it still worksMLOps & LLMOpsOperating an AI system after release, rather than deciding on one
- The problem is that our information is unusableData EngineeringNo generative decision survives unreliable sources
- We need to decide between prompting, retrieval, tuning or hostingLLM DevelopmentThe model-level decision, ordered by cost and reversibility
- The work needs several specialised agents coordinatingAgentic AI DevelopmentMulti-agent architecture and its failure modes
What you are left holding.
The output is an acceptance standard and an evaluation set, both of which are portable.
The acceptance standard, written down
What good output means for this work, in a form a new team member can apply. Portable across any model you later choose.
The evaluation set, in your possession
Real inputs with expected outcomes. This is the asset that makes every future model change a measured decision rather than a hope.
The decisions, with their reasons
Including what was excluded and why, so a later reversal is deliberate rather than accidental.
No dependency on a vendor choice
The output is a specification and a test set, not a configuration inside somebody's platform.
Tell us the piece of writing or answering your team repeats most.
That is enough to judge whether generation, retrieval or neither is the right treatment. All three answers are common, and the second two are cheaper.
Questions we are actually asked
Pick a topic, or ask us directly. We answer every inbound within one business day.
Still have questions?
Talk to a senior engineer, not a bot.