Run · Operations

Operations that tell you when a system needs you and stay quiet when it does not.

Monitoring, alerts, dead-man timers and a morning report, plus incidents written up and fixed for good. This is the layer we run under every product of our own, and the reason a two person studio can operate a portfolio.

GaraMasalaFM, built in this lane
Anything already live that has to stay liveGood for
Monthly retainer, one point of contactShape
Every product on this siteProof
Runbooks you keep, alerts you can readHandover
How it is built

The shape of every operations project we run.

Your systemsSites, apps, pipelines, serversScheduled jobsEverything that should runHealth checksLive pages, queues, diskDead-man timersNotice what did not runTriageNeeds a human, or notAlertOnly when a human is neededIncident and fixWritten up, fixed for goodMorning reportAll is well, or here is why not
The stackHealth checksDead-man timersDisk and queue guardsDeploy verificationMorning reportsIncident write-upsRunbooksHetzner and Cloudflare
What lands in your hands

Six things you get, shown on products we already run.

[01]   OperationsChecks that read the real thing

The live page, the queue depth, the disk, the last successful run. Not a green light from a deploy hook.

[02]   OperationsTimers that notice silence

Every scheduled job has a dead-man check. A job that did not run is an alert, not a mystery next week.

[03]   OperationsAlerts with a threshold

A human is paged only when a human is needed. Everything else waits for the morning report.

[04]   OperationsOne report every morning

What ran, what changed, what needs a decision. Short enough to read with coffee.

[05]   OperationsIncidents written up and fixed for good

Root cause, permanent fix and the lesson, filed the same day.

[06]   OperationsRunbooks you keep

Every recurring task documented so the system does not depend on memory, ours or yours.

01 / 06Scroll sideways or use the arrows
Also in this lane

Where this work usually goes next.

[01]Monitoring and alerts

Live checks on pages, queues, disks and jobs, with alert thresholds set so the phone stays quiet.

[02]Morning reports

One email a day per system, written for the person accountable, not for an engineer.

[03]Incident practice

Same day write-ups with the permanent fix, so the same thing never pages twice.

Questions

Asked before hiring us for Operations.

Do you need access to my servers?

Read access to run checks, and deploy access only if we are the ones fixing. Everything is logged and revocable.

What does the retainer cover?

The watching, the morning report, alerts, and incident handling with a written fix. Larger changes are scoped separately.

What if you are asleep when it breaks?

The system does not sleep. Alerts route to whoever is on, and most incidents are caught by a timer before a customer notices.

Get started

Tell us what you need built.

+ One paragraph from you+ Reply within one business day, Sydney time+ Written scope and a fixed price, or an honest no