Run an on-call pager
A watch checks your service every few minutes, and when it fails, a worker finds the cause, fixes it and logs the incident.
Most outages have a cause you could fix in a minute. Let something be awake for it.
- a health endpoint
- a smoke script after each deploy
- a queue that stops draining
This playbook turns one sentence into an on-call rota of one. A watch runs your health check on a schedule. While the check passes, nothing happens and nothing is spent. When it fails, a worker reads the service log and the latest commits, fixes the cause, confirms the check passes again, writes the incident down, and the chat tells you in one line what happened.
Steps
Our service, shopfront, is a small Python storefront that reads config.json on every request and logs errors to logs/app.log.
- In the service's repository, run
codeaf. - Say what to check, how often, and what to do when it fails, and press Enter:
Be my on-call pager for this service: every 2 minutes run curl -sf http://127.0.0.1:18765/health. Whenever it fails, find out why from logs/app.log and the latest commits, fix the cause (not the check), confirm /health answers 200 again, add an entry to INCIDENT.md (what broke, when, the cause, the fix), and tell me in one line.- A card comes back:
wants to watch for something · on-call pager,every 2 minutes, when the health probe fails · shares the day's $500.00 allowance · checked every 5 minutes,where · for this project. Press 1 (Watch for it). The right column shows◦ 1 standing order.
What you see
In our run a teammate committed config: point at the new catalogue file, which pointed the service at a file that does not exist, and /health started answering 500. The next pass ran the check 12 seconds later and the watch fired. Half a minute after that the chat woke with one line: the commit that broke it, the file that never existed, the path put back, /health back to 200, and the incident logged in INCIDENT.md.
The incident entry had all four parts: what broke, when (to the second of the bad commit), the cause and the fix. The fix is left as an uncommitted edit in the working folder, so you review it with git diff before it goes anywhere.
In our run
From the broken commit to a healthy service took about 45 seconds. Setting up the watch took 43 seconds and $0.02, and the firing cost $0.006.
Make it yours
- Check anything. The check is any shell command: a smoke script after each deploy, a query that counts stuck jobs, a
curlagainst staging. The watch fires when the command fails. - Keep your hands on the fix. Replace "fix the cause" with "write the cause and a proposed fix to INCIDENT.md" and the pager only diagnoses.
- Know its limits. Each firing is held to $5 and ten firings a day, and standing orders are checked by a pass every 5 minutes, so a failure is caught within 5 minutes. Say "stop the on-call pager" to retire it. See Watches.
Go further
Watches: Say what you are waiting for; CodeAF looks on a timer and starts work only when the answer is yes.
Coming soon: describe the org you want and CodeAF builds it: teams, managers, budgets, standing orders.