post: loop engineering without a coding agent — the 11 cron jobs that run my business (EN+ZH) + og/banner images
Deploy / build (push) Successful in 18s
Deploy / build (push) Successful in 18s
This commit is contained in:
Binary file not shown.
|
After Width: | Height: | Size: 101 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 44 KiB |
@@ -1049,6 +1049,26 @@ BANNERS['one-hostname-public-tracker-sso-dashboard'] = {
|
||||
],
|
||||
};
|
||||
|
||||
BANNERS['loop-engineering-without-a-coding-agent'] = {
|
||||
titlebar: 'root@hermes — 11 cron loops, 8 with no model',
|
||||
lines: [
|
||||
{ t: 'prompt', text: '$' }, { t: 'cmd', text: 'hermes cron list → 11 loops running this machine' },
|
||||
{ t: 'prompt', text: 'INFO' }, { t: 'cmd', text: '8 of 11: no LLM · script stdout delivered as the notice' },
|
||||
{ t: 'prompt', text: 'WARN' }, { t: 'err', text: 'agent-mode job: [drift_skip] × 29 runs — silent when broken' },
|
||||
{ t: 'prompt', text: '$' }, { t: 'cmd', text: 'rewrite as plain script · exit 1 = CANNOT MEASURE' },
|
||||
{ t: 'prompt', text: 'INFO' }, { t: 'cmd', text: 'control probe first · 2-fail threshold · 3-run cooldown' },
|
||||
{ t: 'prompt', text: 'INFO' }, { t: 'cmd', text: 'state beside the script: run_no · fails · down_since' },
|
||||
{ t: 'prompt', text: 'INFO' }, { t: 'cmd', text: 'tiers: report only → propose → act inside an allowlist' },
|
||||
{ t: 'prompt', text: '' }, { t: 'ok', text: '→ 1,618 ticks · 0 false alarms · self-healed a hung VM ✓' },
|
||||
],
|
||||
flow: [
|
||||
{ n: '1', label: 'trigger' },
|
||||
{ n: '2', label: 'probe' },
|
||||
{ n: '3', label: 'drift', err: true },
|
||||
{ n: '4', label: 'verify' },
|
||||
{ n: '5', label: 'act ✓' },
|
||||
],
|
||||
};
|
||||
|
||||
// ---------- read frontmatter ----------
|
||||
const postPath = join(ROOT, 'src', 'content', 'posts', `${slug}.md`);
|
||||
|
||||
@@ -321,6 +321,11 @@ TERMINALS['one-hostname-public-tracker-sso-dashboard'] = `
|
||||
<div class="line"><span class="prompt"> </span><span class="err">mode=forward_single: SSO 302 ✓ · every app path 404</span></div>
|
||||
<div class="line"><span class="prompt">$</span><span class="cmd">mode=proxy + internal_host=http://umami:3000</span><span class="fix">→ dashboard 200 ✓</span></div>`;
|
||||
|
||||
TERMINALS['loop-engineering-without-a-coding-agent'] = `
|
||||
<div class="line"><span class="prompt">$</span><span class="cmd">hermes cron list → 11 loops · 8 with no model in the path</span></div>
|
||||
<div class="line"><span class="prompt"> </span><span class="err">agent-mode job: [drift_skip] × 29 runs — silent broken and silent healthy</span></div>
|
||||
<div class="line"><span class="prompt">$</span><span class="cmd">rewrite as a script · exit 1 = CANNOT MEASURE</span><span class="fix">→ 1,618 ticks ✓</span></div>`;
|
||||
|
||||
// ---------- read frontmatter ----------
|
||||
const postPath = join(ROOT, 'src', 'content', 'posts', `${slug}.md`);
|
||||
if (!existsSync(postPath)) {
|
||||
|
||||
@@ -0,0 +1,251 @@
|
||||
---
|
||||
title: "Loop Engineering Without a Coding Agent: The 11 Cron Jobs That Run My Business"
|
||||
description: "Loop engineering is 2026's term for designing the system that prompts and checks an agent. I applied it to a business, not a codebase: 11 cron jobs, 8 with no AI at all."
|
||||
pubDate: 2026-10-01
|
||||
category: case-studies
|
||||
tags: ["loop-engineering", "ai-agents", "cron", "self-hosting", "automation", "monitoring"]
|
||||
ogImage: "/og/loop-engineering-without-a-coding-agent.png"
|
||||
banner: "/banners/loop-engineering-without-a-coding-agent.png"
|
||||
draft: false
|
||||
---
|
||||
|
||||
Every article about "loop engineering" is about a coding agent. Claude Code,
|
||||
Codex, `/goal`, `/loop`. The examples are always the same shape: an agent
|
||||
refactoring a repository overnight while you sleep.
|
||||
|
||||
I don't have that problem. I run a small web studio — websites, hosting,
|
||||
self-hosted infrastructure — and my loops keep the *business* alive, not a
|
||||
codebase. Eleven scheduled jobs run on this machine: an endpoint watchdog
|
||||
that fires every five minutes, mail triage twice a day, a shop-voucher
|
||||
audit, a disk-corruption check, a domain-expiry reminder, a guard that
|
||||
notices when an upgrade silently reset my proxy headers.
|
||||
|
||||
Eight of those eleven contain no AI at all.
|
||||
|
||||
## Why this matters more than the agent does
|
||||
|
||||
A website that is down at 2am is a client who finds out before I do. An
|
||||
expiring domain is a live business that stops receiving mail. A corrupt SSD
|
||||
is a week of my work. None of these announce themselves; they all wait
|
||||
quietly until someone looks.
|
||||
|
||||
A loop is that someone. It turns "I should check on that" into "something
|
||||
checks on that, forever, and only speaks when there is news." The business
|
||||
outcome is not magic autonomy — it is that I stopped being the monitoring
|
||||
system. I get a Telegram message when something is actually wrong, and
|
||||
silence when it isn't.
|
||||
|
||||
That is the whole value proposition, and it is worth being precise about it,
|
||||
because the 2026 hype around loops promises considerably more.
|
||||
|
||||
## What loop engineering actually is
|
||||
|
||||
The term arrived in June 2026. Peter Steinberger posted that you shouldn't
|
||||
be prompting coding agents anymore — you should be designing the loops that
|
||||
prompt them. Boris Cherny, who leads Claude Code at Anthropic, put it even
|
||||
shorter: *"I don't prompt Claude anymore. I have loops running that prompt
|
||||
Claude. My job is to write loops."* Addy Osmani published the essay that
|
||||
gave the practice a name, and Anthropic made it official on June 30 with a
|
||||
taxonomy of four loop types.
|
||||
|
||||
It is the newest ring of a ladder that keeps wrapping itself: **prompt
|
||||
engineering** (2022–24, *did I say it clearly?*), **context engineering**
|
||||
(2025, *does it see the right thing?*), **harness engineering** (early 2026,
|
||||
*does it keep doing it right?*), and **loop engineering** (2026, *does it
|
||||
run without me?*). Nothing replaced anything. I still write prompts, usually
|
||||
inside a script.
|
||||
|
||||
Nobody has to adopt the word. The Register called it the latest buzzword in
|
||||
June and made two fair points: agents have always been loops, and the
|
||||
companies selling tokens are the companies most excited about loops that
|
||||
spend tokens unattended. One widely-read report claimed a $1.3M monthly bill
|
||||
for an always-on loop setup. Even Osmani's own essay ends with the caveat
|
||||
that matters: *"The loop changes the work, it does not delete you from it."*
|
||||
|
||||
My loops cost me almost nothing, and the reason is the first design rule
|
||||
below.
|
||||
|
||||
## The inventory
|
||||
|
||||
| Job | Trigger | What proves it worked | Model? |
|
||||
|---|---|---|---|
|
||||
| Endpoint watchdog (7 targets, self-heal) | every 5 min | HTTP/TCP probe + control probe | no — script |
|
||||
| Disk-corruption watchdog | daily 10:00 | SMART counters vs baseline | no |
|
||||
| Domain-expiry reminder | daily 09:00 | 30-day / 7-day thresholds | no |
|
||||
| Mail triage (`[email protected]`) | 09:30 + 18:30 | unread mail, deduped | **yes** |
|
||||
| Shop voucher audit | daily 09:00 | read-only portal scrape | **yes** |
|
||||
| Real-IP + bot-block regression guard | hourly :20 | config invariants + log ratios | no |
|
||||
| unRaid SSD watchdog | daily 10:00 | counter deltas | no |
|
||||
| Mailbox health | monthly | IMAP/SMTP connect | no |
|
||||
| Memory backup | weekly | archive written | no |
|
||||
| Package version check | daily 09:00 | npm latest vs local | no |
|
||||
| Research checkpoint reminder | 1st + 15th | reads project state, reminds me | **yes** |
|
||||
|
||||
Eleven loops, eight of them with no model anywhere in the path. That ratio
|
||||
is not an accident and it is not modesty — it is the single most useful
|
||||
thing I learned.
|
||||
|
||||
## Rule 1 — if the check is deterministic, do not spend a model on it
|
||||
|
||||
My package version check used to be an agent job. It ran for weeks and
|
||||
failed 29 times in a row before I stopped ignoring the notifications. The
|
||||
cause was almost funny: an agent-mode job records the model it was created
|
||||
with, and when my global default model later changed, the job refused to run
|
||||
at all rather than silently adopt the new one. It printed `[drift_skip]` and
|
||||
exited. Twenty-nine times.
|
||||
|
||||
The task itself is trivial: fetch `registry.npmjs.org/<pkg>/latest`, compare
|
||||
it with the local version, print something only if the remote is newer.
|
||||
There is no judgment in that. So I rewrote it as 125 lines of Python with an
|
||||
explicit exit path for every branch, and set the job to `no-agent` — the
|
||||
script's stdout is delivered as the notification, no LLM in the loop.
|
||||
|
||||
It has run green ever since. Anthropic's own guidance says the same thing in
|
||||
one line: *"Use scripts for deterministic work — running a script is cheaper
|
||||
than reasoning through the steps."* I'd add the stronger version: if the
|
||||
stop condition is a string comparison, a model in that loop can only
|
||||
introduce failure modes, and it will, at 3am, silently.
|
||||
|
||||
The measured difference in my own fleet: 8 of 11 loops run with zero tokens.
|
||||
The three that use a model all need judgment I cannot express as code — is
|
||||
this email important, is this voucher about to overspend, is this project
|
||||
stalled.
|
||||
|
||||
## Rule 2 — the verifier must be able to fail loudly about itself
|
||||
|
||||
The second thing I got wrong for a long time: a watchdog that only speaks
|
||||
when something is broken is indistinguishable from a watchdog that has
|
||||
stopped working. Both are silent.
|
||||
|
||||
So every script in the fleet has one contract: **exit code 0 means
|
||||
"measured, here is what I found" and exit code 1 means "I could not measure"
|
||||
— and exit 1 is never silent.** A missing config file, an unparseable state
|
||||
file, an unwritable state directory, an unexpected exception: all of them
|
||||
print a `CANNOT MEASURE` line and exit 1. The runner wraps `main()` in a
|
||||
blanket handler that catches anything I failed to anticipate, because the
|
||||
one failure mode I cannot tolerate is a checker that has quietly died.
|
||||
|
||||
The endpoint watchdog goes further, and this is my favourite piece of the
|
||||
whole setup. Before it judges anything, it probes a **control target** —
|
||||
`https://one.one.one.one/`. If that fails, the problem is my own uplink or
|
||||
DNS, not my servers, and the watchdog reports instead of acting. Without
|
||||
that gate, a router reboot would look exactly like seven simultaneous
|
||||
outages, and an automated loop would have cheerfully restarted a perfectly
|
||||
healthy VM in the middle of it.
|
||||
|
||||
A verifier that cannot tell "broken" from "I can't see" will eventually take
|
||||
a destructive action during a network outage. I would rather it know the
|
||||
difference.
|
||||
|
||||
The number I like quoting from this loop: it has ticked **1,618 times**
|
||||
without a false alarm, because two consecutive failures are required before
|
||||
anything is called down, a group that stays down is re-reported on a slow
|
||||
cadence rather than every tick, and recovery is announced once with the
|
||||
downtime window.
|
||||
|
||||
## Rule 3 — give it the smallest autonomy that still helps
|
||||
|
||||
The tempting design is "the loop fixes things." The design that survives
|
||||
contact with production has three tiers, and each loop sits in exactly one:
|
||||
|
||||
**Tier 1 — report only.** Most of my loops. They observe and tell me. Nothing they do can break anything, which means I can deploy them on a Friday.
|
||||
|
||||
**Tier 2 — propose, human confirms.** The shop-voucher audit runs daily and is *read-only by contract*. The prompt forbids creating or editing a voucher even if the script prints a proposal, because vouchers on that platform have no draft state: confirming means live and escrowed, with money attached. The loop's job is to hand me a decision I can make in ten seconds.
|
||||
|
||||
**Tier 3 — act, within an allowlist.** Exactly one loop in the fleet acts, and the conditions are narrow: it restarts a VM only when the guest is *unreachable from the hypervisor* (verified by ping and ARP from the host), never when the guest is up but its service is failing — that is a bug to fix, not a machine to bounce. A three-run cooldown prevents a restart loop. And it never runs at all when the control probe failed.
|
||||
|
||||
The pattern generalises: **each tier up must earn its right with a verifier
|
||||
I trust more than the agent.** Am I allowed to walk away? Only if the thing
|
||||
that decides "done" is something I would trust in a post-mortem.
|
||||
|
||||
That third-tier restraint is also what the shops on the receiving end of
|
||||
these loops require. A marketplace audit that "helpfully" fixed its own
|
||||
findings would have been a compliance problem, not a win.
|
||||
|
||||
## Trap 1: a loop that fails forever looks exactly like a loop with nothing to report
|
||||
|
||||
My silent-when-healthy convention is what makes these loops tolerable — no
|
||||
daily status spam, only real news. It also hid a job that had failed **29
|
||||
consecutive times**, because a failed run and a clean run both delivered
|
||||
nothing to my phone.
|
||||
|
||||
Two things fixed it, and they apply to any unattended loop:
|
||||
|
||||
- **A distinct "I could not measure" path** (Rule 2). It would have surfaced the drift on the first run instead of the twenty-ninth.
|
||||
- **Check the runner's own history, not just the inbox.** The job's execution log had the evidence the whole time; I trusted the absence of a notification instead. Now, when I add a loop, I check its first few runs explicitly rather than waiting for it to tell me something.
|
||||
|
||||
If your loop has ever been "quiet for a while", go look at its last ten runs
|
||||
before you conclude it is working.
|
||||
|
||||
## Trap 2: a verifier that deletes production data
|
||||
|
||||
This one still stings. I had an end-to-end test script for a delivery
|
||||
pipeline, and its cleanup step ran an unconditional `DELETE FROM orders`. It
|
||||
was written to clean up after itself and it did — along with the seven real
|
||||
orders that happened to be in the same table, and their generated PDFs, with
|
||||
no backup to restore from.
|
||||
|
||||
The lesson is not "test on a copy" (true, but I had). It is narrower and
|
||||
harder: **a verifier is a program with permissions, and its write path needs
|
||||
the same scrutiny as the thing it is verifying.** A loop I trust to run
|
||||
unattended must only be destructive to what it created itself. Every
|
||||
verification script in the fleet now scopes its cleanup to rows carrying its
|
||||
own marker — `WHERE id IN ($MINE)` or `WHERE recipient LIKE '%@example.com'`
|
||||
— and one of them did get rewritten after that incident for exactly this
|
||||
reason.
|
||||
|
||||
Osmani's point about comprehension debt lands here. The more smoothly the
|
||||
loop runs, the less you read its output — and the day it is wrong, nobody
|
||||
was watching. I read the diffs these loops produce. That habit is the entire
|
||||
remaining job.
|
||||
|
||||
## What I deliberately did not automate
|
||||
|
||||
One of the eleven loops does nothing but remind me, twice a month, that a
|
||||
research project exists and here are the commands to run next. It fetches
|
||||
nothing, writes nothing, opens no browser. A loop's job is to keep *a* loop
|
||||
turning — not necessarily its own.
|
||||
|
||||
That is the honest boundary. These eleven jobs buy back attention and catch
|
||||
failures early; they do not run the business. The client emails still get
|
||||
answered by me. The price still gets confirmed by me. What changed is that I
|
||||
no longer spend any part of my day wondering whether something is broken.
|
||||
|
||||
## What I would do differently
|
||||
|
||||
- **Start with the report-only tier and stay there longer.** I wrote one acting loop before I had the control probe and the cooldown, which is a mistake I got lucky on rather than deserved to avoid.
|
||||
- **Add the "cannot measure" exit path to the first version of every script.** Retrofitting it across eleven jobs took an evening; it is one `try/except` in each.
|
||||
- **Check the runner's execution log on day one**, not when a notification feels overdue.
|
||||
- **Write the state file before you need it.** Every loop here keeps a small JSON next to it — tick number, last-known status, when something went down. A loop with no external state cannot tell you what it knew last time, which is most of what "did this recover?" means.
|
||||
|
||||
## The result
|
||||
|
||||
Eleven loops, eight of them running with no model and therefore no token
|
||||
bill, covering seven monitored endpoints, a mailbox, a shop portal, a RAID
|
||||
array and a set of proxy invariants that a CyberPanel or DSM upgrade can
|
||||
silently reset. The endpoint watchdog has completed 1,618 consecutive ticks
|
||||
and has self-healed a hung VM from unreachable to healthy without me
|
||||
touching it. The most expensive job in the fleet is a daily version check
|
||||
that costs nothing and would have kept failing forever if I hadn't read its
|
||||
history.
|
||||
|
||||
If you are self-hosting anything for money, the useful version of "loop
|
||||
engineering" is not a fleet of agents. It is one cron job with a stop
|
||||
condition you can test, a verifier that says when it cannot see, an exit
|
||||
path for the failure you did not anticipate — and permission to do the
|
||||
smallest thing that helps. Start with the one thing you check manually every
|
||||
week, and make it load-bearing.
|
||||
|
||||
## Want this for your business?
|
||||
|
||||
If you are running a website or an online shop and the answer to "is it up
|
||||
right now?" is "I'd have to check" — I build exactly this: self-hosted
|
||||
watchdogs, uptime and certificate monitoring, automated backups with
|
||||
verification, and mail triage that only pings you when something actually
|
||||
needs you. No SaaS subscription per host, no dashboard you have to remember
|
||||
to open.
|
||||
|
||||
**WhatsApp: [+60 12-797 2969](https://wa.me/60127972969)** · **Email: [[email protected]](mailto:[email protected]?subject=Self-hosted%20monitoring%20and%20automation)** · **[hoelee.com](https://hoelee.com)**
|
||||
|
||||
Website design and development is my main line of work; self-hosted
|
||||
infrastructure, monitoring and automation is the other half of it.
|
||||
@@ -0,0 +1,134 @@
|
||||
---
|
||||
title: "没有 Coding Agent 的 Loop Engineering:撑起我生意的 11 个定时任务"
|
||||
description: "Loop engineering 是 2026 年对「设计让 agent 自己反复跑的系统」的叫法。我把它用在生意上而不是代码库上:11 个 cron,其中 8 个完全没有 AI。"
|
||||
pubDate: 2026-10-01
|
||||
category: case-studies
|
||||
tags: ["loop-engineering", "ai-agents", "cron", "self-hosting", "automation", "monitoring"]
|
||||
ogImage: "/og/loop-engineering-without-a-coding-agent.png"
|
||||
banner: "/banners/loop-engineering-without-a-coding-agent.png"
|
||||
draft: false
|
||||
---
|
||||
|
||||
关于 loop engineering 的文章,几乎全都在讲 coding agent:Claude Code、Codex、`/goal`、`/loop`。例子永远是同一个形状——一个 agent 整晚替你重构某个代码仓库。
|
||||
|
||||
我没有这个问题。我经营一家小型网站公司——建站、托管、自托管基础设施——我的循环维持的是**生意**,不是代码库。这台机器上跑着 11 个定时任务:每五分钟跳一次的端点看门狗、一天两次的邮件分诊、店铺优惠券审计、磁盘损坏检查、域名到期提醒,以及一个专门发现「某次升级偷偷把我的代理头配置重置了」的守卫。
|
||||
|
||||
这 11 个里面,有 8 个完全不含 AI。
|
||||
|
||||
## 为什么这件事比 agent 本身重要
|
||||
|
||||
凌晨两点挂掉的网站,等于客户比我更早知道它挂了。域名到期,等于一门还在收信的生意突然断掉。SSD 损坏,等于我一个星期的工作没了。这些事都不会自己喊出来——它们只是静静地等,等到有人去看。
|
||||
|
||||
循环就是那个「有人」。它把「我该去看一下」变成「永远有东西在看,而且只在有事的时候才开口」。生意上的收益不是什么魔法级自动化,而是我不再充当监控系统本身:真的出问题会有 Telegram 消息,没事的时候就是安静。
|
||||
|
||||
这就是全部的价值主张,值得说清楚,因为 2026 年围绕 loop 的炒作承诺的远不止这些。
|
||||
|
||||
## Loop engineering 到底是什么
|
||||
|
||||
这个说法是 2026 年 6 月出现的。Peter Steinberger 发帖说,你不该再去 prompt coding agent 了,应该去设计「prompt 这些 agent 的循环」。Anthropic 的 Claude Code 负责人 Boris Cherny 说得更短:*"I don't prompt Claude anymore. I have loops running that prompt Claude. My job is to write loops."* 随后 Addy Osmani 发表那篇给这个实践命名的文章,Anthropic 也在 6 月 30 日正式给了它一套四种循环类型的分类。
|
||||
|
||||
它是一条不断向外包一层的阶梯上最新的那圈:**prompt engineering**(2022–24,*我说清楚了吗?*)、**context engineering**(2025,*它看到对的东西了吗?*)、**harness engineering**(2026 年初,*它有没有一直做对?*)、**loop engineering**(2026,*它能不能不靠我在跑?*)。没有哪一层取代了前一层。我到现在还在写 prompt,只不过通常写在脚本里。
|
||||
|
||||
没人非得接受这个名词。The Register 在 6 月就管它叫最新一轮 buzzword,并且提了两点公允的批评:agent 本来就是循环;而最热衷推销「无人值守循环」的,恰恰是那些靠卖 token 赚钱的公司。有一份流传很广的报告声称某套常开的循环方案一个月烧了 130 万美元。连 Osmani 自己那篇文章的结尾都留了最重要的保留意见:*"The loop changes the work, it does not delete you from it."*
|
||||
|
||||
我的循环几乎不花钱,原因就是下面第一条设计规则。
|
||||
|
||||
## 清单
|
||||
|
||||
| 任务 | 触发 | 凭什么算它干成了 | 用模型? |
|
||||
|---|---|---|---|
|
||||
| 端点看门狗(7 个目标,自愈) | 每 5 分钟 | HTTP/TCP 探测 + 控制探测 | 否——纯脚本 |
|
||||
| 磁盘损坏看门狗 | 每天 10:00 | SMART 计数对比基线 | 否 |
|
||||
| 域名到期提醒 | 每天 09:00 | 30 天 / 7 天阈值 | 否 |
|
||||
| 邮件分诊(`[email protected]`) | 09:30 + 18:30 | 未读邮件,按 message-id 去重 | **是** |
|
||||
| 店铺优惠券审计 | 每天 09:00 | 只读抓取后台 | **是** |
|
||||
| Real-IP + 机器人拦截回归守卫 | 每小时 :20 | 配置不变量 + 日志比例 | 否 |
|
||||
| unRaid SSD 看门狗 | 每天 10:00 | 计数增量 | 否 |
|
||||
| 邮箱连通性健康检查 | 每月 | IMAP/SMTP 连接 | 否 |
|
||||
| 记忆备份 | 每周 | 归档文件写出来了 | 否 |
|
||||
| 软件版本检查 | 每天 09:00 | npm 最新版 vs 本地版 | 否 |
|
||||
| 研究项目检查点提醒 | 每月 1、15 号 | 读项目状态,提醒我 | **是** |
|
||||
|
||||
11 个循环,其中 8 个整条路径上没有任何模型。这个比例不是意外,也不是谦虚——它是我学到的最有用的一件事。
|
||||
|
||||
## 规则一:检查是确定性的,就别花模型去跑
|
||||
|
||||
我那个软件版本检查原本是 agent 任务。它跑了好几个星期,连续失败了 **29 次**,我才终于不再忽略那些通知。原因有点好笑:agent 模式的定时任务会记录创建时的模型,等我后来改了全局默认模型,这个任务不再悄悄改用新模型,而是干脆整个不跑,打印一行 `[drift_skip]` 就退出。连续 29 次。
|
||||
|
||||
任务本身其实极其简单:抓 `registry.npmjs.org/<pkg>/latest`,和本地版本比大小,只有远端更新才输出内容。这里面没有任何判断。所以我把它重写成 125 行 Python,每条分支都有明确出口,并把任务改成 `no-agent`——脚本的 stdout 直接作为通知内容,循环里没有 LLM。
|
||||
|
||||
从此一路全绿。Anthropic 自己的建议也是一句话:*"Use scripts for deterministic work — running a script is cheaper than reasoning through the steps."* 我想补一句更强的版本:如果停止条件是字符串比较,那么在这个循环里放一个模型,唯一的作用就是引入新的失败模式——而且它一定会在凌晨三点、无声无息地引入。
|
||||
|
||||
在我自己这套任务里的实测差异:11 个循环里有 8 个完全不花 token。用模型的 3 个,需要的都是我写不成代码的判断——这封邮件重不重要、这张券是不是快要超支、这个项目是不是卡住了。
|
||||
|
||||
## 规则二:验证器必须能大声说出「我自己坏了」
|
||||
|
||||
第二个我错了很久的地方:一个只在出事时才开口的看门狗,和一个已经死掉的看门狗,表现完全一样——都很安静。
|
||||
|
||||
所以这套脚本有一个统一契约:**退出码 0 表示「量过了,这是结果」,退出码 1 表示「我量不了」——而退出码 1 永远不静默。** 配置文件不存在、状态文件解析失败、状态目录写不进去、没预料到的异常,全都打印一行 `CANNOT MEASURE` 并退出 1。runner 又把 `main()` 整个包在一层兜底 handler 里,专门接住我没预料到的情况——因为唯一不能容忍的失败模式,是检查器自己悄悄死了。
|
||||
|
||||
端点看门狗走得更远,这也是整套东西里我最喜欢的一处设计。它在做任何判断之前,先探测一个**控制目标**——`https://one.one.one.one/`。如果这个也失败,那问题出在我自己的上行或 DNS,不是我的服务器,于是看门狗只报告、不动手。没有这道闸门,一次路由器重启看起来就和七个服务同时挂掉一模一样,而一个自动化循环会在这种情况下兴高采烈地去重启一台完全健康的虚拟机。
|
||||
|
||||
一个分不清「坏了」和「我看不见」的验证器,迟早会在某次网络故障中做出破坏性操作。我宁愿它知道区别。
|
||||
|
||||
这个循环里我最喜欢引用的数字:它已经连续跑了 **1,618 次**没有误报——因为要连续两次失败才会判定为 down,持续 down 的组会以很慢的节奏重复提醒而不是每个 tick 都喊一遍,恢复只播报一次并附带宕机时长。
|
||||
|
||||
## 规则三:给它「还能帮上忙」的最小权限
|
||||
|
||||
最诱人的设计是「让循环自己修好」。能活过生产的版本是三层,每个循环只属于其中一层:
|
||||
|
||||
**第一层——只报告。** 我大部分循环在这里。它们只观察、只告诉我。它们做什么都不可能弄坏东西,所以我可以周五下午上线。
|
||||
|
||||
**第二层——提议,人确认。** 店铺优惠券审计每天跑,而且按契约是**只读**的。即使脚本打印出一行建议,prompt 也禁止创建或编辑优惠券,因为那个平台的优惠券没有草稿态:确认就等于上线并冻结押金,背后是真金白银。这个循环的任务是递给我一个十秒钟就能做完的决定。
|
||||
|
||||
**第三层——按白名单动手。** 整套任务里只有一个循环会动手,条件收得很窄:只有在客户机**从宿主机都不可达**时(用宿主机 ping + ARP 验证过)才重启虚拟机;客户机活着但服务挂了的情况从不自动重启——那是要修的 bug,不是要重启的机器。三次运行的冷却期防住了重启循环。而控制探测失败时,它压根不会启动。
|
||||
|
||||
这个模式可以推广:**每上一层,都必须用一个我宁可相信它、也不相信 agent 的验证器来换。** 我能不能走开?只有当那个判定「做完了」的东西是我在复盘时也愿意相信的东西时,才行。
|
||||
|
||||
第三层的克制,也是这些循环另一端的平台要求的。一个会「好心」自己修掉发现问题的市场审计循环,是合规问题,不是战果。
|
||||
|
||||
## 坑一:一个永远失败的循环,和一个无事可报的循环长得一模一样
|
||||
|
||||
「健康则静默」这条约定是这些循环能被忍受的原因——没有每日状态刷屏,只有真消息。但它同时也藏住了一个**连续失败 29 次**的任务,因为失败的运行和干净的运行,往我手机上推的都是空。
|
||||
|
||||
后来有两件事把它修好了,而且适用于任何无人值守的循环:
|
||||
|
||||
- **一条独立的「我量不了」通路**(规则二)。它本该在第 1 次就暴露 drift,而不是拖到第 29 次。
|
||||
- **去看 runner 自己的执行历史,而不只是看收件箱。** 证据一直在那个任务的执行日志里,我却一直相信「没有通知」这件事本身。现在每加一个新循环,我会显式看它头几次的执行记录,而不是等它来告诉我。
|
||||
|
||||
如果你的循环「已经安静一阵子」了,先去翻它最近十次运行,再下结论。
|
||||
|
||||
## 坑二:一个会删生产数据的验证器
|
||||
|
||||
这一条我到现在还疼。我给一条交付流水线写过一个端到端测试脚本,它的清理步骤执行了一条无条件的 `DELETE FROM orders`。它本来是清理自己制造的数据,它也照做了——只是顺带把当时恰好在同一张表里的 7 张真实订单删了,连同生成好的 PDF,而且没有备份可还原。
|
||||
|
||||
教训不是「要在副本上测」(对,但当时我确实是),而是更窄、更硬的一条:**验证器是一个有权限的程序,它的写入路径需要和它要验证的东西接受同等程度的审视。** 一个我敢让它无人值守运行的循环,只能对自己的产物具有破坏性。现在这套里每个验证脚本都会把清理范围限定在自己造的、带标记的行上——`WHERE id IN ($MINE)` 或者 `WHERE recipient LIKE '%@example.com'`——其中有一个脚本正是因为这次事故被重写的。
|
||||
|
||||
Osmani 讲的 comprehension debt 就落在这里。循环跑得越顺,你越不会去读它的输出;而它出错的那天,没人在看。这些循环产生的 diff 我都在读。这个习惯就是剩下那部分工作的全部。
|
||||
|
||||
## 我刻意没有自动化的东西
|
||||
|
||||
这 11 个循环里有一个,唯一的工作就是提醒我:每半个月提醒我某个研究项目还在,以及下一步该跑哪些命令。它不抓任何东西、不写任何东西、不开浏览器。循环的职责是让**某个**循环继续转动——不一定是它自己那个。
|
||||
|
||||
这条边界值得说清楚。这 11 个任务买回的是注意力和早期发现;它们不经营生意。客户邮件仍然是人回的,价格仍然是人确认的。真正的变化是:我一天当中再也不会花任何时间去想「是不是哪里坏了」。
|
||||
|
||||
## 我会做得不一样的地方
|
||||
|
||||
- **从「只报告」这一层开始,并且在里面多待一会儿。** 我在还没有控制探测和冷却期的时候就写过会动手的循环——那是一次靠运气躲过、而不是靠设计躲过的错误。
|
||||
- **每个脚本的第一版就带上「量不了」的出口。** 事后给 11 个任务补上花了我一个晚上;其实每个只要一段 `try/except`。
|
||||
- **第一天就去看 runner 的执行日志**,而不是等某条通知「感觉该来了」。
|
||||
- **在你需要之前就把状态文件写好。** 这里每个循环旁边都有一个小 JSON——第几次运行、上次的状态、什么时候开始 down。没有外部状态的循环,说不出它上一次知道什么,而「它恢复了吗」这件事,大部分答案都在那里。
|
||||
|
||||
## 结果
|
||||
|
||||
11 个循环,其中 8 个不带模型、因此也没有 token 账单,覆盖 7 个被监控的端点、一个邮箱、一个店铺后台、一组 RAID 阵列,以及一组会被 CyberPanel 或 DSM 升级悄悄重置的代理不变量。端点看门狗已经连续完成 1,618 次 tick,并且把一台卡死的虚拟机从不可达自愈回健康状态,全程我没有碰它一下。整套任务里最贵的那个是每天一次、成本为零的版本检查——而如果我当时没有去读它的历史,它会一直失败下去。
|
||||
|
||||
如果你在为赚钱而自托管任何东西,那么「loop engineering」有用的那个版本不是一个 agent 舰队,而是:一个停止条件能被测试的 cron 任务、一个说得出「我看不见」的验证器、一条为没预料到的失败准备的出口,以及「允许它只做最小那件事」的权限。从你每周手动检查的那一件事开始,让它变成承重结构。
|
||||
|
||||
## 想给你的生意也做一套?
|
||||
|
||||
如果你的网站或网店对「现在它是不是还活着?」的回答是「我得去看一下」——我就是做这个的:自托管看门狗、可用性与证书监控、带校验的自动备份,以及只在真的有事时才 ping 你的邮件分诊。没有按主机计费的 SaaS 订阅,也没有一个你需要记得去打开的仪表盘。
|
||||
|
||||
**WhatsApp:[+60 12-797 2969](https://wa.me/60127972969)** · **Email:[[email protected]](mailto:[email protected]?subject=Self-hosted%20monitoring%20and%20automation)** · **[hoelee.com](https://hoelee.com)**
|
||||
|
||||
网站设计开发是我的主业;自托管基础设施、监控与自动化是它的另一半。
|
||||
Reference in New Issue
Block a user