Infrastructure & Ops
Infrastructure Build & Maintenance
Build it right, then keep it that way
- Landing zones built with guardrails from day one
- Patching and upgrades on a schedule, not a crisis
- Backups tested quarterly with a written result
- < 30 days
- Critical patch latency
- 4 hrs
- Recovery time objective
- 100%
- Backups restore-tested quarterly
Standing SLO across the managed estate
Standard target; tighter available by design
A backup that has never been restored is a hypothesis
Two jobs that are usually confused
Building infrastructure is a project with an end date. Maintaining it is a standing responsibility with no end date, and it is the one that gets dropped — because it produces no visible output until the day it does, catastrophically.
Patches slip. Certificates expire on a Saturday. A database fills its disk. Backups have been failing silently for five months. None of this requires sophistication to prevent; it requires someone whose job it is.
What maintenance means concretely
Patching to a schedule. Critical CVEs inside 30 days as a standing target, with an out-of-band path for anything actively exploited. Non-production first, then supervised production windows.
Restore tests, quarterly, with a written result. Not "backups are configured" — an actual restore into a clean environment, validated and timed. The timing matters because it is your real RTO.
Drift detection. Someone will eventually change something in a console at 2am during an incident. That is fine, as long as the drift is detected and either codified or reverted rather than becoming permanent undocumented state.
Capacity ahead of demand. Reviewed against growth, not discovered when a disk fills. Headroom is cheaper than an incident.
Greenfield: get the boundaries right
If you are starting fresh, the decisions that are expensive to reverse are the structural ones — how accounts are separated, how networks are segmented, how identity flows, where logs go, and how cost is attributed. Compute choices can be changed later. These cannot, cheaply.
How it runs
What the engagement looks like
Phases, not a proposal. Each one has an output you can see.
- 1
Design for the guardrails first
Weeks 1-2Account boundaries, network segmentation, identity model, logging and cost controls. These are extremely expensive to change later and almost free to get right at the start.
- 2
Build in code
Weeks 2-6Everything in Terraform with modules you can read, in your repository. No resources created by hand, because hand-made infrastructure cannot be reproduced under pressure.
- 3
Prove recovery
Week 6Restore a database from backup into a clean environment and time it. This step routinely uncovers a backup that was never actually working.
- 4
Maintain on a schedule
OngoingPatching, upgrades, certificate renewal, drift remediation and capacity review, on a calendar with maintenance windows agreed in advance.
Proof
Where we have done this
Freight forwarder and 3PL
Recovering 4.1% of freight spend by auditing every invoice
A mid-sized forwarder audited freight invoices by sampling roughly 5% of them, because auditing the rest by hand was uneconomic — so systematic small overcharges passed unnoticed, and shipment exceptions were routinely discovered when the customer rang to complain.
- 4.1%
- Freight spend recovered in year one
- 100%
- Invoices audited, from a 5% sample
- 74%
- Track-and-trace enquiries auto-resolved
Auto components manufacturer
Cutting quality documentation from 6 hours a shift to 90 minutes
Across three plants, quality engineers spent most of a shift writing non-conformance and inspection documentation by hand, while production reporting arrived a morning late because the MES, the historian and the ERP disagreed about what each machine was called.
- 6 hrs → 90 min
- Quality documentation time per shift
- 71%
- Non-conformance reports drafted automatically
- 1 shift
- Production reporting latency, from next morning
FAQ
Questions we get asked
How is this different from Managed DevOps?
DevOps is the delivery path — pipelines, releases, developer experience. This is the estate underneath: networks, databases, backups, patching, capacity, DR. Many clients take both, and they are priced together in the retainer packages.
Do you take over existing infrastructure?
Regularly. We start with an assessment, import what already exists into Terraform incrementally rather than rebuilding, and prioritise remediation by risk. You get a findings report with severity ratings before anything changes.
What does a restore test actually involve?
We restore a real backup into an isolated environment, run agreed validation queries against it, and record the time taken. That number becomes your actual RTO, which is frequently very different from the one in the plan.
How do you handle patching without breaking things?
Staged windows — non-production first, then production in maintenance windows agreed with you. Automated where the risk is low, scheduled and supervised where it is not. Critical CVEs get an out-of-band path.
Talk to someone who does infrastructure management
Thirty minutes with an engineer who has delivered this, not an account manager. You will get a straight answer on feasibility, rough cost and where it would fail.
Or email [email protected] · we reply within 1 business day