Skip to main content
Running & automating

What a website operations retainer should actually include

Most "maintenance" plans are a bucket of hours spent only when something breaks. A real operations retainer is a defined agreement: monitoring, incident response, tested backups, and a standing review.

A row of round analog pressure dials on an industrial steel manifold in warm-tone monochrome, the gauge faces receding under grazing light with their markings left unreadable.

"Maintenance" is the vaguest word on a digital invoice. It can mean a genuine operating discipline, or it can mean a few hours a month that sit unused until something breaks and then get spent in a panic. The straight version is this: a real operations retainer is not a bucket of hours, it is a defined agreement. You are not buying time to spend when the site goes down. You are buying the assurance that someone whose actual job it is keeps the system healthy, and that when something does slip, it gets caught and handled before your customers are the ones who notice.

The difference matters because a live website or platform is not a finished object. It is a thing that runs, and running things need tending: they drift, they fill up, they depend on parts of the wider internet that change without asking. A retainer worth paying for says, in writing, exactly what that tending consists of. Here is what should be inside one.

What "maintenance" usually hides

The common arrangement is a small monthly fee for some undefined number of "maintenance hours." It sounds prudent and it usually means this: nothing actually happens until something is visibly wrong, at which point the hours get spent firefighting, the bill arrives, and you go back to waiting for the next incident. That is not operations. It is on-call billing with a friendlier name, and it leaves the system unattended for the very long stretches when attention would have been cheapest.

A real retainer is defined by what happens when nothing is wrong, which is most of the time. The work that prevents incidents, watching, updating, testing, reviewing, is the work that quietly makes the firefighting unnecessary. If a plan cannot tell you what it does on a normal week when the site is up and fine, it is selling you reaction and calling it care.

Laid against each other, the two arrangements barely overlap.

DimensionMaintenance hoursOperations retainer
MonitoringNone until a customer reports a problemContinuous checks on the paths customers actually use
Incident responseFind whoever is free when something breaksA named path and a clock, set by a written SLA
UpdatesDone when someone remembers, which is rarelyOn a schedule, with a tested rollback
BackupsTaken, but rarely if ever restoredProven by regular restore drills
Standing reviewNoneWeekly or monthly, with an honest dashboard you can read
PricingMetered hours that reward breakageFixed monthly fee for a defined scope
OwnershipOften unstated, lock-in creeps inYour accounts, standard tooling, the system stays yours

What does monitoring actually watch?

The first thing a real retainer buys is monitoring that watches what your customers actually experience, not just whether the server has power. There is a meaningful gap between "the machine is on" and "a person can load the page, log in, and complete the thing they came to do." Good monitoring checks the second kind: it loads the critical paths the way a customer would, measures how fast they respond, and watches the rate of errors, continuously.

The point of all of it is the direction the bad news travels. Without monitoring, you learn about problems from a customer email, which means the problem has already reached the people you least wanted it to reach. With it, you learn from a dashboard or an alert, while there is still time to act quietly. A retainer should be explicit about what is monitored and how you find out when something is wrong, because that is the line between an incident the customer never saw and a public one.

Incident response and the SLA, in plain words

When something does go wrong, the retainer should tell you what happens next before it happens, and that is what a Service-Level Agreement (SLA) is for. Behind the acronym it is a plain set of promises: how quickly someone responds when an issue is raised, how quickly they aim to have it resolved, and what counts as a drop-everything emergency versus a routine fix that can wait for business hours.

Those promises only mean something if they are written in language you can hold someone to. A good agreement names a small number of severity levels in plain words, a customer-facing outage is not the same as a typo on a footer, and attaches a response window to each. What you are buying is the end of the scramble: when something breaks, there is a named path and a clock, instead of a frantic search for whoever might be free and willing to look.

Updates, backups, and restore drills

Underneath the visible service, two unglamorous disciplines do most of the protecting. The first is updates: the building blocks a site depends on need patching on a schedule, with a tested way to roll back if an update misbehaves, so staying current does not mean gambling every time. A retainer should say how updates are handled, not leave them to be done when someone remembers, which is to say rarely.

The second is backups, and this is where a lot of plans quietly fail. Having backups and being able to restore from them are different facts, and only one of them helps you on a bad day. A backup that has never been restored is a hope, not a backup. A real retainer proves its backups by actually restoring them on a schedule, as a drill, so that the day you need to recover is not the day you discover the recovery never worked.

A backup that has never been restored is a hope, not a backup.

The standing review

The piece that separates a system that is merely surviving from one that is improving is a standing review, weekly or monthly depending on the system. It is a recurring look at what changed: what drifted since last time, what is approaching a limit before it becomes a problem, and which small fix made now prevents a large incident later. Without it, operations becomes pure reaction, and the system slowly accumulates the small neglects that eventually combine into an outage.

The review is also where the owner gets to see the state of the system in plain terms, on a dashboard they can read, rather than having to trust that it is all fine. That visibility is part of the deliverable. The real product of a good retainer is that the system stops living in the back of your mind, and the standing review, with its honest dashboard, is the proof that it is genuinely being looked after rather than merely paid for.

How is an operations retainer priced?

A real retainer is priced as a fixed monthly fee for a defined scope, not as metered hours, and the reason is about incentives as much as accounting. When you pay by the hour for maintenance, the arrangement quietly rewards breakage: the more goes wrong, the more hours get billed. A fixed monthly fee inverts that. The work of preventing incidents is the firm's to do within the fee, so when something breaks, it costs the firm its own time, not your extra budget. The interests line up: both sides now want the system boringly stable.

What the fixed fee should not be is a vague number with vague coverage. It should attach to the specific scope above, the monitoring, the response targets, the update and backup disciplines, the review, so you can see what you are paying for and hold it to account. (The same logic runs underneath buying senior technical judgment by the month rather than as a salaried hire, which we cover in CTO as a service: you pay a defined amount for a defined outcome, sized to what the business actually needs.)

What a good retainer leaves out, and what stays yours

A defined retainer is as clear about what it does not cover as about what it does, because a scope without edges is the thing that turns a fixed fee back into an argument. Keeping a system healthy and building new things on top of it are different kinds of work, and a good retainer says so. Monitoring, incident response, updates, backups, and the review are the operating discipline. A new feature, a redesign, or a major integration is a separate engagement with its own scope and price. Folding new development into a maintenance fee is how retainers quietly become either overstretched or wasted, and a clear boundary protects both sides from that.

The second thing a good retainer should be explicit about is ownership, because an operations arrangement is exactly where lock-in tends to creep in. The system should run on standard tooling the business owns outright, and you should hold your own accounts, credentials, and documentation throughout, rather than borrow access to your own platform from the firm that runs it. The test is simple: if you ended the retainer tomorrow, could a competent team pick the system up and keep it running from what is documented and handed over? A retainer worth signing answers yes without hesitation, because the point of being operated is that the system stays yours, not that it becomes dependent on one provider to survive.

That is the line between a retainer that serves the business and one that quietly captures it. You are paying for the system to be watched and kept well, on terms that would let you walk away with everything still working. Anything less is not an operations retainer. It is a dependency with a monthly invoice.

When does a retainer beat a full-time hire?

Running a live system is a genuine job, but for a smaller company it is often not a full week's job, and that mismatch is exactly what a retainer is for. A full-time operations hire gives you depth and presence, and also a full salary, the search to fill the role, and a single point of failure in the form of the one person who takes holidays and gets sick. A retainer gives you the discipline and the coverage without those, because the responsibility sits with a team and an agreement rather than with one calendar.

The honest test is whether the work fills the week. If keeping the system healthy genuinely needs someone full-time and you can carry the role, hire, and treat the retainer as the bridge that keeps things safe until then. If it does not, a defined retainer is the right answer for longer than most owners expect, because the thing you actually need is not a person at a desk. It is the certainty that the system is being watched, updated, backed up, and reviewed, by someone who answers for it.

The deepest version of this is when running the system stops being something you think about at all. The engagement below is that, made real: a live platform with paying users brought under a defined operating agreement, and twelve months later, not a single customer-visible outage.

Running & automating

More in this topic.

All insights
Common questions

Straight answers

The questions this raises most often, answered plainly.

What should a website operations retainer include?

A real operations retainer is a defined agreement, not a bucket of hours: monitoring that watches what customers actually experience, incident response with a written Service-Level Agreement, scheduled updates with a tested rollback, backups proven by regular restore drills, and a standing weekly or monthly review with an honest dashboard. It also states clearly what it does not cover and that the system stays yours.

What is the difference between maintenance hours and an operations retainer?

A maintenance-hours plan does nothing until something is visibly wrong, then spends the hours firefighting and bills you. That is on-call billing with a friendlier name. A real retainer is defined by what happens when nothing is wrong: the watching, updating, testing, and reviewing that quietly makes the firefighting unnecessary.

How is a website operations retainer priced?

A real retainer is priced as a fixed monthly fee for a defined scope, not as metered hours. Paying by the hour quietly rewards breakage, because the more goes wrong, the more hours get billed. A fixed fee inverts that: preventing incidents is the firm time within the fee, so both sides want the system boringly stable.

Why test backups with restore drills?

Having backups and being able to restore from them are different facts, and only one helps you on a bad day. A backup that has never been restored is a hope, not a backup. A real retainer proves its backups by actually restoring them on a schedule, as a drill, so the day you need to recover is not the day you discover the recovery never worked.

When does a retainer beat a full-time operations hire?

When the work does not fill a full week. A full-time hire gives you depth and presence, plus a full salary, the search to fill the role, and a single point of failure in one person who takes holidays and gets sick. A retainer gives the discipline and the coverage without those, because the responsibility sits with a team and an agreement rather than one calendar.

Bangkok office-building rooftop in warm-tone monochrome: HVAC ducts and metal louvres in sharp geometry, hazy city silhouettes reduced to tonal bands above.

Working through a decision like this?

Tell us about it. The person who would do the work reads your message and replies within one business day.

Tell us about your project