Skip to main content
Viithiisys
One number, one answer

Data engineering services for teams whose reports nobody quite trusts.

We fix where the numbers come from, so the answer is the same whoever runs the query and whenever they run it.

Building and running software since 2007Definitions agreed in writing before anything is built

Revenue, last quartertwo sources
Finance dashboard£4.12mexcludes refunds
Sales dashboard£4.48mexcludes nothing

Same metric. £360k apart. Both defensible.

Traced back

One definition was written down. The other was not.

How trust goes

Three ways a data platform loses the room.

Data engineering is not the right first step if the underlying systems do not capture what you need. Adding a warehouse on top of missing data produces a faster route to the same gap.

  1. Whose number are we using?

    Heard in the room

    Two dashboards, two answers

    Both are defensible, both are wrong in different ways, and the meeting turns into a discussion about the numbers rather than the business.

  2. We can act on that tomorrow, then.

    Heard in the room

    The number arrives too late to act on

    Yesterday's figure at midday is a report. The same figure at 8am is a decision.

  3. Can someone ask her to pull it?

    Heard in the room

    One person can answer the question

    Every analysis routes through the same analyst, and the queue becomes the constraint on what the business can ask.

What we take on

What we take on.

Six pieces of one job: getting the data right, and keeping it right once people depend on it.

  1. Ingestion and pipelines

    Getting data out of the systems that hold it, on a schedule the business actually needs rather than the one that was easiest to build.

  2. Modelling and definitions

    One agreed definition per metric, written down, so revenue means the same thing in two departments.

  3. Warehouse and storage design

    Structured for the questions people ask, and sized for the volume you have rather than the volume in a vendor's example.

  4. Quality checks and monitoring

    Tests on the data itself, so a broken feed is caught by the pipeline rather than by someone noticing a strange chart.

  5. Reporting and self-service

    Enough structure that a business team can answer its own question without opening a ticket.

  6. Data for AI systems

    Retrieval, feature and evaluation datasets prepared properly, which is where many AI systems stall when their source information is incomplete or unreliable.

Where this page stops. Deploying and monitoring models belongs to MLOps and LLMOps. Moving the platform to a new environment belongs to cloud migration.

See cloud migration
Data or platform

Four questions that tell you whether the problem is the data or the platform.

Most teams arrive asking for a warehouse. These four questions usually reveal whether that is the right purchase, and you can answer them without us.

  1. Do two people define your main metric the same way?

    If the answer is noDefinitions are the problem
    What that meansModelling and agreed definitions come before any new platform.
  2. Is the data being captured at all?

    If the answer is noThe source system is the problem
    What that meansA warehouse will faithfully report the gap. Fix capture first.
  3. Would a fresher number change a decision?

    If the answer is noLatency is not the problem
    What that meansReal-time work is a cost with no return here.
  4. Do people trust the current numbers?

    If the answer is noQuality is the problem
    What that meansChecks and monitoring return more than new storage.

Three of these four have answers that point away from buying a platform. That is the usual outcome and it is worth knowing before a procurement process starts.

How the work runs

How a data engagement runs.

One question answered end to end before the second is added, so the foundation is proven on something small enough to check.

  1. Start from one question the business cannot answer

    Not an inventory of every system. One question with a decision attached to it, so the work has a finish line.

  2. Trace it back to the source

    Where the data is captured, what transforms it on the way, and where it stops being trustworthy.

  3. Agree the definitions in writing

    Signed off by whoever will use the number. This is the step most often skipped and most often the actual problem.

  4. Build the narrow pipeline first

    One question answered end to end, with quality checks, before the second question is added.

  5. Add checks that fail loudly

    Tests on volume, freshness and shape, so a broken feed is caught before it reaches a dashboard.

  6. Widen it, then hand it over

    Further questions added on the same foundation, and your team able to add the next one without us.

Definitions before pipelines. Building first and agreeing the meaning afterwards is how two dashboards end up disagreeing.

Evidence

Where our data work has shown up in an outcome.

We have no published data engineering case study with a measured before and after. The closest evidence is a customer-facing result that depended on product data being correct.

Faster product-question response. Conscious Chemist. Closest published evidence, not equivalent.
38%Faster product-question response. Conscious Chemist. Closest published evidence, not equivalent.
Longest running client integration, still in production. Jewellerybox commerce APIs.
9 yrsLongest running client integration, still in production. Jewellerybox commerce APIs.

The Conscious Chemist figures describe that engagement and are not offered as a projection.

See all case studies
Where this starts

Where data work usually starts.

Situations rather than industries. We hold no sector-specific data engineering evidence and will not imply otherwise.

A laptop on a desk showing an analytics dashboard of charts and figures
  1. A board number that keeps being corrected

    Where the same figure is restated between meetings and confidence has gone.

  2. An AI project that stalled on data

    Where the model was never the constraint and nobody said so early enough.

  3. A reporting queue with a person in it

    Where every question waits for one analyst and the backlog decides what gets asked.

  4. Two systems after a merger

    Where both hold customer records and neither agrees on the count.

What it costs to run

What a data platform costs to run, and where the bill comes from.

We do not publish figures because they depend on your volume and query patterns. We do model these four before the build so the running cost is a decision rather than a discovery.

  1. Every query

    Compute is usually the bill, not storage

    Storage is cheap and predictable. Query and transformation compute is neither, and it scales with how people use the platform.

  2. Every run

    Refresh frequency is a cost decision

    Hourly instead of daily multiplies the transformation cost. It is worth it only where a fresher number changes an action.

  3. As adoption grows

    Self-service moves cost, not removes it

    Letting teams query directly is right, and it shifts spend from analyst time to compute. Both are real.

  4. On every change

    Reprocessing history is the surprise

    Changing a definition means recalculating the past. Worth costing before the definition is agreed rather than after.

What you keep

What you have when the engagement ends.

Including the definitions themselves, with an author and a date, which is what stops the argument starting again.

  1. Pipelines in your repository

    Defined as code you own, reviewable and changeable without us.

  2. A written definition for every metric built

    With the person who agreed it and the date they did.

  3. Quality checks that keep running

    Set up to fail loudly after we leave, which is when it matters.

  4. The four cost models

    Compute, refresh, self-service and reprocessing, filled in for your actual usage.

  5. A written list of what we did not build

    Deferred questions with the reason, so the next decision starts from the record.

From your side

Who we need from your side.

Two of these four are business people rather than engineers, which is usually the surprise.

A team working through notes on a wall in a meeting room

Definitions get agreed by the people who will act on the number.

Whoever uses the number

Definitions cannot be agreed without the person who will act on them.

Someone who can grant access

To source systems, in read-only form, without a three-week approval chain.

Your analyst, if you have one

They know where the current numbers are wrong, which is the fastest available map.

Someone who owns the platform bill

Because refresh frequency and self-service are cost decisions, not technical ones.

Start here

Bring the number your team keeps arguing about.

Tell us which figure nobody trusts. We will trace it back to where it is captured and say plainly if you do not need what you came to buy.

FAQ

What teams ask before they start.

Pick a topic, or ask us directly. We answer every inbound within one business day.

Still have questions?

Talk to a senior engineer, not a bot.

Talk to us