Skip to content
AI Digital Hub

By use case

Data Platform Buildout

One version of the numbers, tested and trusted

Usually owned by

  • Data
  • Engineering
  • Finance
  • Operations
99.5%
Pipeline success rate

SLO with freshness and volume alerting

< 15 min
Freshness for AI-serving datasets
40%
Warehouse spend reduction

Typical after partitioning and query review

The change

What actually differs afterwards

Not a maturity model. The concrete difference in how the work happens.

Today

  • Three systems give three different revenue figures
  • Reports built from manual exports each month
  • A pipeline failed silently and nobody noticed for a fortnight
  • Nobody can say what a column means without asking a person
  • Warehouse spend rising with no attribution

Afterwards

  • One modelled source of truth with tested definitions
  • Reports and AI features read from the same layer
  • Pipelines fail loudly and page someone
  • Every model documented and lineage traceable
  • Spend attributed per domain with anomaly alerts

The symptom that brings people here

Someone asks a simple question — how many active customers do we have — and gets three answers. Finance has one, the product dashboard has another, and the CRM has a third. Each is defensible. None is authoritative.

That is not a tooling problem. It is a definitions problem, and the platform is where definitions get settled and enforced.

Build order

Model the domain first. Start from the entities the business reasons about — customer, order, subscription, invoice. Pipelines built backwards from a single dashboard become unmaintainable the moment a second consumer appears.

Ingest idempotently. Loads that can be safely re-run. A pipeline you are afraid to re-run is a pipeline that will cause a bad night at some point.

Transform with tests that fail the build. Uniqueness, referential integrity, freshness and business rules asserted in CI. A warning in a channel nobody reads is not a test.

Serve both consumers from the same definitions. Analytics and AI retrieval read the same modelled layer. This is the whole point.

Attribute the cost. Spend per domain, visible to the team that generates it. Making it visible is usually sufficient to fix it.

What we will talk you out of

Buying a warehouse you do not need yet. A great many companies with under ten million rows of transactional data are paying enterprise data-platform prices for something a well-modelled Postgres instance would handle, and the migration cost of moving up later is far lower than the accumulated overspend.

Choosing the first domain

The instinct is to start with the domain that causes the most pain, which is usually revenue, which is usually also the one that touches every system you own. That is how a platform programme spends a quarter without shipping anything.

A better first domain has three properties. It has a single authoritative source system, so you are modelling rather than reconciling. It has a named consumer waiting — a specific report or automation that will use it the week it lands, which is what makes the value visible to whoever approved the budget. And it is small enough to finish, meaning a handful of entities rather than a domain map.

Revenue is usually the second or third domain, not the first. By then the ingestion patterns, testing conventions and deployment path exist, and the hard part is the definitional argument rather than the engineering — which is where it should be.

The definitions are the deliverable

The engineering on these programmes is rarely the difficult part. Settling what a word means is.

"Active customer" has a different definition in finance, in product and in sales, and each is defensible for its own purpose. The platform cannot hold all three under one name. Someone has to decide, write it down, and accept that two teams will have to change a number they have been reporting for years.

That decision is not a technical one and we cannot make it for you. What we can do is surface every conflict explicitly during modelling rather than letting it get resolved by accident in a transformation nobody reads. Expect a handful of genuinely uncomfortable meetings — they are the point, not a delay.

Where this shows up by sector

  • Manufacturing — OT and IT integration, where asset identity has to be reconciled before anything else works.
  • Financial services — reporting with lineage, where every figure has to be traceable to source for an inspector.
  • SaaS and technology — the modelled layer under product AI features and usage-based renewal signals.
  • Retail and eCommerce — product, order and stock data reconciled across marketplaces, which almost every other retail automation depends on.

FAQ

Questions we get asked

Do we need this before doing anything with AI?

Not for everything. A retrieval system over documents can ship without a warehouse. But if the automation has to reason over transactional data, it inherits every inconsistency you have — so the modelled layer comes first, or the automation will be confidently wrong.

Warehouse, lakehouse, or just Postgres?

Under a few hundred million rows with no heavy analytical workload, Postgres with good modelling is often enough and dramatically cheaper. We will tell you which side of that line you are on rather than defaulting to the expensive answer.

How do you stop it becoming a graveyard of unused tables?

Ownership and usage tracking. Every model has a named owner, and models nothing has queried in ninety days get flagged for deletion. Without that, every platform accumulates until nobody trusts any of it.

Can the same layer serve both analytics and AI retrieval?

That is the design goal. Treating them as separate projects is how a board report and an AI feature end up disagreeing about revenue. Shared definitions, with the retrieval layer carrying lineage back to source rows.

What does a first engagement cost and how long does it take?

Longer than the automation work, and it is worth being direct about that. A first useful slice — one domain modelled, ingested, tested and serving something real — is six to eight weeks. A platform covering the domains a mid-sized business actually runs on is a two to three quarter programme. Anyone quoting a complete data platform in four weeks is either scoping one table or planning to hand you something you cannot maintain.

We already bought Snowflake or Databricks. Is that money wasted?

Not usually, and we are not going to recommend ripping it out to justify a migration. If it is provisioned and underused, the problem is almost always modelling and ownership rather than the platform, and that is what we would work on. The one case where the spend is genuinely questionable is a small transactional estate on an enterprise warehouse, and even then the fix is often better partitioning and query discipline rather than a move.

Do we need to hire a data team to run this afterwards?

Eventually, if the platform keeps growing. Not immediately. A well-tested platform with documented models and alerting runs with a fraction of the attention that an undocumented one demands, and for many mid-sized businesses that is a part of one engineer's role rather than a team. We will tell you when the surface has grown past that point rather than waiting for you to notice through incidents.

How soon do we see anything at all?

The first modelled domain should be serving a real report or a real automation inside eight weeks, and we sequence to make that true. Data platform programmes that show nothing for two quarters lose their sponsor, and the fix is to pick a first domain narrow enough to finish rather than broad enough to be strategically satisfying. Breadth comes after something is demonstrably working.

What ongoing cost should we budget?

Warehouse or database spend, plus the engineering time to maintain models as source systems change — and the second one is the line most budgets omit. Source schemas change, business definitions get revised, and a platform nobody maintains degrades into one nobody trusts within about a year. We quote the run cost alongside the build cost, because a platform is an operational commitment rather than a project with an end date.

Let's find out what is actually automatable

Bring a process that annoys you. In 30 minutes we will tell you whether AI helps, what it would cost, and where it would fail — even if the answer is don't bother.

Or email [email protected] · we reply within 1 business day