post(loop-engineering): reframe for job search — drop business framing, job-seeking CTA (EN+ZH) + og
Deploy / build (push) Successful in 21s

This commit is contained in:
2026-10-08 05:31:50 +08:00
parent 541b00c958
commit d2b15d684f
3 changed files with 54 additions and 59 deletions
Binary file not shown.

Before

Width:  |  Height:  |  Size: 44 KiB

After

Width:  |  Height:  |  Size: 44 KiB

@@ -1,6 +1,6 @@
---
title: "Loop Engineering Without a Coding Agent: The 11 Cron Jobs That Run My Business"
description: "Loop engineering is 2026's term for designing the system that prompts and checks an agent. I applied it to a business, not a codebase: 11 cron jobs, 8 with no AI at all."
title: "Loop Engineering Without a Coding Agent: The 11 Cron Jobs That Run My Infrastructure"
description: "Loop engineering is 2026's term for designing the system that prompts and checks an agent. I applied it to self-hosted infrastructure, not a codebase: 11 cron jobs, 8 with no AI at all."
pubDate: 2026-10-01
category: case-studies
tags: ["loop-engineering", "ai-agents", "cron", "self-hosting", "automation", "monitoring"]
@@ -13,27 +13,28 @@ Every article about "loop engineering" is about a coding agent. Claude Code,
Codex, `/goal`, `/loop`. The examples are always the same shape: an agent
refactoring a repository overnight while you sleep.
I don't have that problem. I run a small web studio — websites, hosting,
self-hosted infrastructure — and my loops keep the *business* alive, not a
codebase. Eleven scheduled jobs run on this machine: an endpoint watchdog
that fires every five minutes, mail triage twice a day, a shop-voucher
audit, a disk-corruption check, a domain-expiry reminder, a guard that
notices when an upgrade silently reset my proxy headers.
I don't have that problem. I run my own self-hosted infrastructure — a NAS,
a couple of servers, mail, a reverse proxy, a CI runner — and my loops keep
*that* alive, not a codebase. Eleven scheduled jobs run on this machine: an
endpoint watchdog that fires every five minutes, mail triage twice a day, a
marketplace voucher audit, a disk-corruption check, a domain-expiry
reminder, a guard that notices when an upgrade silently reset my proxy
headers.
Eight of those eleven contain no AI at all.
## Why this matters more than the agent does
A website that is down at 2am is a client who finds out before I do. An
expiring domain is a live business that stops receiving mail. A corrupt SSD
is a week of my work. None of these announce themselves; they all wait
quietly until someone looks.
A service that is down at 2am stays down until someone notices. An expiring
domain is mail that silently stops arriving. A corrupt SSD is a week of my
work. None of these announce themselves; they all wait quietly until someone
looks.
A loop is that someone. It turns "I should check on that" into "something
checks on that, forever, and only speaks when there is news." The business
outcome is not magic autonomy — it is that I stopped being the monitoring
system. I get a Telegram message when something is actually wrong, and
silence when it isn't.
checks on that, forever, and only speaks when there is news." The payoff is
not magic autonomy — it is that I stopped being the monitoring system. I get
a Telegram message when something is actually wrong, and silence when it
isn't.
That is the whole value proposition, and it is worth being precise about it,
because the 2026 hype around loops promises considerably more.
@@ -73,7 +74,7 @@ below.
| Disk-corruption watchdog | daily 10:00 | SMART counters vs baseline | no |
| Domain-expiry reminder | daily 09:00 | 30-day / 7-day thresholds | no |
| Mail triage (`[email protected]`) | 09:30 + 18:30 | unread mail, deduped | **yes** |
| Shop voucher audit | daily 09:00 | read-only portal scrape | **yes** |
| Marketplace voucher audit | daily 09:00 | read-only portal scrape | **yes** |
| Real-IP + bot-block regression guard | hourly :20 | config invariants + log ratios | no |
| unRaid SSD watchdog | daily 10:00 | counter deltas | no |
| Mailbox health | monthly | IMAP/SMTP connect | no |
@@ -108,7 +109,7 @@ introduce failure modes, and it will, at 3am, silently.
The measured difference in my own fleet: 8 of 11 loops run with zero tokens.
The three that use a model all need judgment I cannot express as code — is
this email important, is this voucher about to overspend, is this project
this email important, does this voucher need a human, is this project
stalled.
## Rule 2 — the verifier must be able to fail loudly about itself
@@ -150,7 +151,7 @@ contact with production has three tiers, and each loop sits in exactly one:
**Tier 1 — report only.** Most of my loops. They observe and tell me. Nothing they do can break anything, which means I can deploy them on a Friday.
**Tier 2 — propose, human confirms.** The shop-voucher audit runs daily and is *read-only by contract*. The prompt forbids creating or editing a voucher even if the script prints a proposal, because vouchers on that platform have no draft state: confirming means live and escrowed, with money attached. The loop's job is to hand me a decision I can make in ten seconds.
**Tier 2 — propose, human confirms.** The voucher audit runs daily and is *read-only by contract*. The prompt forbids creating or editing a voucher even if the script prints a proposal, because vouchers on that platform have no draft state: confirming means live and escrowed, with money attached. The loop's job is to hand me a decision I can make in ten seconds.
**Tier 3 — act, within an allowlist.** Exactly one loop in the fleet acts, and the conditions are narrow: it restarts a VM only when the guest is *unreachable from the hypervisor* (verified by ping and ARP from the host), never when the guest is up but its service is failing — that is a bug to fix, not a machine to bounce. A three-run cooldown prevents a restart loop. And it never runs at all when the control probe failed.
@@ -158,9 +159,9 @@ The pattern generalises: **each tier up must earn its right with a verifier
I trust more than the agent.** Am I allowed to walk away? Only if the thing
that decides "done" is something I would trust in a post-mortem.
That third-tier restraint is also what the shops on the receiving end of
these loops require. A marketplace audit that "helpfully" fixed its own
findings would have been a compliance problem, not a win.
That restraint is also what the platform on the receiving end requires. An
audit that "helpfully" fixed its own findings would be a terms-of-service
problem, not a win.
## Trap 1: a loop that fails forever looks exactly like a loop with nothing to report
@@ -207,9 +208,9 @@ nothing, writes nothing, opens no browser. A loop's job is to keep *a* loop
turning — not necessarily its own.
That is the honest boundary. These eleven jobs buy back attention and catch
failures early; they do not run the business. The client emails still get
answered by me. The price still gets confirmed by me. What changed is that I
no longer spend any part of my day wondering whether something is broken.
failures early. They never make a decision I haven't already handed them,
and the judgement still comes from me. What changed is that I no longer
spend any part of my day wondering whether something is broken.
## What I would do differently
@@ -221,31 +222,27 @@ no longer spend any part of my day wondering whether something is broken.
## The result
Eleven loops, eight of them running with no model and therefore no token
bill, covering seven monitored endpoints, a mailbox, a shop portal, a RAID
array and a set of proxy invariants that a CyberPanel or DSM upgrade can
silently reset. The endpoint watchdog has completed 1,618 consecutive ticks
and has self-healed a hung VM from unreachable to healthy without me
bill, covering seven monitored endpoints, a mailbox, a marketplace account,
a RAID array and a set of proxy invariants that a CyberPanel or DSM upgrade
can silently reset. The endpoint watchdog has completed 1,618 consecutive
ticks and has self-healed a hung VM from unreachable to healthy without me
touching it. The most expensive job in the fleet is a daily version check
that costs nothing and would have kept failing forever if I hadn't read its
history.
If you are self-hosting anything for money, the useful version of "loop
engineering" is not a fleet of agents. It is one cron job with a stop
If you are running anything you have to keep alive, the useful version of
"loop engineering" is not a fleet of agents. It is one cron job with a stop
condition you can test, a verifier that says when it cannot see, an exit
path for the failure you did not anticipate — and permission to do the
smallest thing that helps. Start with the one thing you check manually every
week, and make it load-bearing.
## Want this for your business?
## Open to work
If you are running a website or an online shop and the answer to "is it up
right now?" is "I'd have to check" — I build exactly this: self-hosted
watchdogs, uptime and certificate monitoring, automated backups with
verification, and mail triage that only pings you when something actually
needs you. No SaaS subscription per host, no dashboard you have to remember
to open.
I'm a full-stack developer and DevOps engineer, and I'm open to remote or
hybrid roles — platform, infrastructure, DevOps, or full-stack. This blog is
the work sample: every post here is something I actually built and then had
to keep running, and the eleven loops above are the least glamorous and most
useful part of it.
**WhatsApp: [+60 12-797 2969](https://wa.me/60127972969)** · **Email: [[email protected]](mailto:[email protected]?subject=Self-hosted%20monitoring%20and%20automation)** · **[hoelee.com](https://hoelee.com)**
Website design and development is my main line of work; self-hosted
infrastructure, monitoring and automation is the other half of it.
**Email: [[email protected]](mailto:[email protected]) · [LinkedIn](https://www.linkedin.com/in/hoelee) · [GitHub](https://github.com/hoelee) · [hoelee.com](https://hoelee.com)**
@@ -1,6 +1,6 @@
---
title: "没有 Coding Agent 的 Loop Engineering:撑起我生意的 11 个定时任务"
description: "Loop engineering 是 2026 年对「设计让 agent 自己反复跑的系统」的叫法。我把它用在生意上而不是代码库上:11 个 cron,其中 8 个完全没有 AI。"
title: "没有 Coding Agent 的 Loop Engineering:撑起我这套基础设施的 11 个定时任务"
description: "Loop engineering 是 2026 年对「设计让 agent 自己反复跑的系统」的叫法。我把它用在自己的自托管基础设施上而不是代码库上:11 个 cron,其中 8 个完全没有 AI。"
pubDate: 2026-10-01
category: case-studies
tags: ["loop-engineering", "ai-agents", "cron", "self-hosting", "automation", "monitoring"]
@@ -11,15 +11,15 @@ draft: false
关于 loop engineering 的文章,几乎全都在讲 coding agent:Claude Code、Codex、`/goal`、`/loop`。例子永远是同一个形状——一个 agent 整晚替你重构某个代码仓库。
我没有这个问题。我经营一家小型网站公司——建站、托管、自托管基础设施——我的循环维持的是**生意**,不是代码库。这台机器上跑着 11 个定时任务:每五分钟跳一次的端点看门狗、一天两次的邮件分诊、店铺优惠券审计、磁盘损坏检查、域名到期提醒,以及一个专门发现「某次升级偷偷把我的代理头配置重置了」的守卫。
我没有这个问题。我自己搭了一套自托管的基础设施——一台 NAS、两台服务器、邮件、反向代理、CI runner——我的循环维持的是**这套东西**,不是代码库。这台机器上跑着 11 个定时任务:每五分钟跳一次的端点看门狗、一天两次的邮件分诊、平台优惠券审计、磁盘损坏检查、域名到期提醒,以及一个专门发现「某次升级偷偷把我的代理头配置重置了」的守卫。
这 11 个里面,有 8 个完全不含 AI。
## 为什么这件事比 agent 本身重要
凌晨两点挂掉的网站,等于客户比我更早知道它挂了。域名到期,等于一门还在收信的生意突然断掉。SSD 损坏,等于我一个星期的工作没了。这些事都不会自己喊出来——它们只是静静地等,等到有人去看。
凌晨两点挂掉的服务,只会一直挂到有人发现。域名到期,等于邮件在无声无息中停掉。SSD 损坏,等于我一个星期的工作没了。这些事都不会自己喊出来——它们只是静静地等,等到有人去看。
循环就是那个「有人」。它把「我该去看一下」变成「永远有东西在看,而且只在有事的时候才开口」。生意上的收益不是什么魔法级自动化,而是我不再充当监控系统本身:真的出问题会有 Telegram 消息,没事的时候就是安静。
循环就是那个「有人」。它把「我该去看一下」变成「永远有东西在看,而且只在有事的时候才开口」。带来的好处不是什么魔法级自动化,而是我不再充当监控系统本身:真的出问题会有 Telegram 消息,没事的时候就是安静。
这就是全部的价值主张,值得说清楚,因为 2026 年围绕 loop 的炒作承诺的远不止这些。
@@ -41,7 +41,7 @@ draft: false
| 磁盘损坏看门狗 | 每天 10:00 | SMART 计数对比基线 | 否 |
| 域名到期提醒 | 每天 09:00 | 30 天 / 7 天阈值 | 否 |
| 邮件分诊(`[email protected]`) | 09:30 + 18:30 | 未读邮件,按 message-id 去重 | **是** |
| 店铺优惠券审计 | 每天 09:00 | 只读抓取后台 | **是** |
| 平台优惠券审计 | 每天 09:00 | 只读抓取后台 | **是** |
| Real-IP + 机器人拦截回归守卫 | 每小时 :20 | 配置不变量 + 日志比例 | 否 |
| unRaid SSD 看门狗 | 每天 10:00 | 计数增量 | 否 |
| 邮箱连通性健康检查 | 每月 | IMAP/SMTP 连接 | 否 |
@@ -59,7 +59,7 @@ draft: false
从此一路全绿。Anthropic 自己的建议也是一句话:*"Use scripts for deterministic work — running a script is cheaper than reasoning through the steps."* 我想补一句更强的版本:如果停止条件是字符串比较,那么在这个循环里放一个模型,唯一的作用就是引入新的失败模式——而且它一定会在凌晨三点、无声无息地引入。
在我自己这套任务里的实测差异:11 个循环里有 8 个完全不花 token。用模型的 3 个,需要的都是我写不成代码的判断——这封邮件重不重要、这张券是不是快要超支、这个项目是不是卡住了。
在我自己这套任务里的实测差异:11 个循环里有 8 个完全不花 token。用模型的 3 个,需要的都是我写不成代码的判断——这封邮件重不重要、这张券要不要人来处理、这个项目是不是卡住了。
## 规则二:验证器必须能大声说出「我自己坏了」
@@ -79,13 +79,13 @@ draft: false
**第一层——只报告。** 我大部分循环在这里。它们只观察、只告诉我。它们做什么都不可能弄坏东西,所以我可以周五下午上线。
**第二层——提议,人确认。** 店铺优惠券审计每天跑,而且按契约是**只读**的。即使脚本打印出一行建议,prompt 也禁止创建或编辑优惠券,因为那个平台的优惠券没有草稿态:确认就等于上线并冻结押金,背后是真金白银。这个循环的任务是递给我一个十秒钟就能做完的决定。
**第二层——提议,人确认。** 优惠券审计每天跑,而且按契约是**只读**的。即使脚本打印出一行建议,prompt 也禁止创建或编辑优惠券,因为那个平台的优惠券没有草稿态:确认就等于上线并冻结押金,背后是真金白银。这个循环的任务是递给我一个十秒钟就能做完的决定。
**第三层——按白名单动手。** 整套任务里只有一个循环会动手,条件收得很窄:只有在客户机**从宿主机都不可达**时(用宿主机 ping + ARP 验证过)才重启虚拟机;客户机活着但服务挂了的情况从不自动重启——那是要修的 bug,不是要重启的机器。三次运行的冷却期防住了重启循环。而控制探测失败时,它压根不会启动。
这个模式可以推广:**每上一层,都必须用一个我宁可相信它、也不相信 agent 的验证器来换。** 我能不能走开?只有当那个判定「做完了」的东西是我在复盘时也愿意相信的东西时,才行。
第三层的克制,也是这些循环另一端的平台要求的。一个会「好心」自己修掉发现问题的市场审计循环,是合规问题,不是战果。
第三层的克制,也是这些循环另一端的平台要求的。一个会「好心」自己修掉发现问题的审计循环,是违反平台条款的问题,不是战果。
## 坑一:一个永远失败的循环,和一个无事可报的循环长得一模一样
@@ -110,7 +110,7 @@ Osmani 讲的 comprehension debt 就落在这里。循环跑得越顺,你越
这 11 个循环里有一个,唯一的工作就是提醒我:每半个月提醒我某个研究项目还在,以及下一步该跑哪些命令。它不抓任何东西、不写任何东西、不开浏览器。循环的职责是让**某个**循环继续转动——不一定是它自己那个。
这条边界值得说清楚。这 11 个任务买回的是注意力和早期发现;它们不经营生意。客户邮件仍然是人回的,价格仍然是人确认的。真正的变化是:我一天当中再也不会花任何时间去想「是不是哪里坏了」。
这条边界值得说清楚。这 11 个任务买回的是注意力和早期发现。它们不会替我做任何我没有预先交给它们的决定,判断仍然来自我。真正的变化是:我一天当中再也不会花任何时间去想「是不是哪里坏了」。
## 我会做得不一样的地方
@@ -121,14 +121,12 @@ Osmani 讲的 comprehension debt 就落在这里。循环跑得越顺,你越
## 结果
11 个循环,其中 8 个不带模型、因此也没有 token 账单,覆盖 7 个被监控的端点、一个邮箱、一个店铺后台、一组 RAID 阵列,以及一组会被 CyberPanel 或 DSM 升级悄悄重置的代理不变量。端点看门狗已经连续完成 1,618 次 tick,并且把一台卡死的虚拟机从不可达自愈回健康状态,全程我没有碰它一下。整套任务里最贵的那个是每天一次、成本为零的版本检查——而如果我当时没有去读它的历史,它会一直失败下去。
11 个循环,其中 8 个不带模型、因此也没有 token 账单,覆盖 7 个被监控的端点、一个邮箱、一个平台账号、一组 RAID 阵列,以及一组会被 CyberPanel 或 DSM 升级悄悄重置的代理不变量。端点看门狗已经连续完成 1,618 次 tick,并且把一台卡死的虚拟机从不可达自愈回健康状态,全程我没有碰它一下。整套任务里最贵的那个是每天一次、成本为零的版本检查——而如果我当时没有去读它的历史,它会一直失败下去。
如果你在为赚钱而自托管任何东西,那么「loop engineering」有用的那个版本不是一个 agent 舰队,而是:一个停止条件能被测试的 cron 任务、一个说得出「我看不见」的验证器、一条为没预料到的失败准备的出口,以及「允许它只做最小那件事」的权限。从你每周手动检查的那一件事开始,让它变成承重结构。
如果你在跑任何必须一直活着的东西,那么「loop engineering」有用的那个版本不是一个 agent 舰队,而是:一个停止条件能被测试的 cron 任务、一个说得出「我看不见」的验证器、一条为没预料到的失败准备的出口,以及「允许它只做最小那件事」的权限。从你每周手动检查的那一件事开始,让它变成承重结构。
## 想给你的生意也做一套?
## 我目前的求职状态
如果你的网站或网店对「现在它是不是还活着?」的回答是「我得去看一下」——我就是做这个的:自托管看门狗、可用性与证书监控、带校验的自动备份,以及只在真的有事时才 ping 你的邮件分诊。没有按主机计费的 SaaS 订阅,也没有一个你需要记得去打开的仪表盘。
我是全栈开发 / DevOps 工程师,正在找 remote 或 hybrid 的岗位——平台、基础设施、DevOps 或全栈方向。这个博客就是我的作品集:这里的每一篇,都是我真正做过、并且必须让它持续跑下去的东西,上面这 11 个循环是其中最不花哨、也最实用的部分。
**WhatsApp:[+60 12-797 2969](https://wa.me/60127972969)** · **Email:[[email protected]](mailto:[email protected]?subject=Self-hosted%20monitoring%20and%20automation)** · **[hoelee.com](https://hoelee.com)**
网站设计开发是我的主业;自托管基础设施、监控与自动化是它的另一半。
**Email:[[email protected]](mailto:[email protected]) · [LinkedIn](https://www.linkedin.com/in/hoelee) · [GitHub](https://github.com/hoelee) · [hoelee.com](https://hoelee.com)**