Add 4 monitoring posts (EN + ZH): smartctl exit 32, Grafana no-data variable, percentage alert thresholds, one Prometheus for 3 hosts
Deploy / build (push) Successful in 28s
Backdated into the 2026-03-25 -> 2026-09-04 archive gap (pubDate 2026-04-14/05-17/06-24/07-29) with updatedDate 2026-09-29 holding the real date, so sitemap lastmod stays honest. Custom OG + banner per post, hire CTA, language switch verified.
@@ -145,6 +145,26 @@ self-hosted app produced two posts:
|
||||
first-person debugging stories with measured evidence — the moat per `content-guide.md` §7.
|
||||
- **Done when:** ✅ 4 pages 200 with expected content, language switch links both ways, 4 images served as `image/png`.
|
||||
|
||||
**Step B2f — (unplanned) Four monitoring posts out of one Prometheus/Grafana session.** ✅ Done 2026-09-29
|
||||
One working session that unified monitoring across unRaid + Synology DSM + a VPS produced four posts. All four are
|
||||
**backdated** into the 2026-03-25 → 2026-09-04 archive gap (that stretch had no posts) with `updatedDate: 2026-09-29`
|
||||
holding the real date, so the sitemap `lastmod` stays honest and listings still sort by `pubDate`:
|
||||
|
||||
| Slug | Category | pubDate | What it argues |
|
||||
|---|---|---|---|
|
||||
| `smartctl-exit-code-32-skips-the-disks-that-matter` | `notes` | 2026-04-14 | `smartctl`'s exit status is a bitfield, not a boolean: `rc=32` means "SMART OK, attributes were below threshold in the past". An `if ! smartctl` guard skipped 2 of 4 SSDs — exactly the marginal ones. Fix: mask the informational bits (32/64), export `rc` as a metric. |
|
||||
| `why-your-grafana-dashboard-shows-no-data` | `devops` | 2026-05-17 | A template variable defined as `label_values(...{nodename=~"$nodename"})` filters on itself → 0 options → `$node` empty → every panel No data while all targets are `up`. Also: why hand-substituting variable values during verification hides exactly this bug, and `$__all` ≠ `.*` in automated panel checks. |
|
||||
| `your-disk-full-alert-is-lying` | `devops` | 2026-06-24 | Percentage thresholds on multi-TB volumes fire while 500 GB remains; 92% "memory used" with 4.8 GB available is cache, not pressure. Alert on consequences: bytes free, `MemAvailable`, steal >50%. Includes the "keep the comparison in the threshold condition" rule and the mount-selector exclusions. |
|
||||
| `one-prometheus-for-unraid-synology-and-a-vps` | `case-studies` | 2026-07-29 | The flagship: node_exporter vs cAdvisor coverage matrix; `name!=""` for cAdvisor's non-container cgroups; "total storage" counting one NAS volume three times (`/volume1`, `/opt`, CIFS re-mount) and the dedup selector; a KVM guest exporting no CPU frequency at all (textfile collector, distinct metric name, merged with `or`); a container reporting its own ID as `nodename`. Result: 7 targets, 47 cores / 158 GHz / 142 GB / 64 TB / 155 containers on one screen. |
|
||||
|
||||
- All four EN + ZH, custom OG + banner, hire CTA naming "self-hosted monitoring pipelines"; no post carries an absolute
|
||||
date or "recently/as of" phrasing, which is what made the backdating safe (per `post-guideline.md` backdating rule).
|
||||
- **Why:** the blog had **zero** Prometheus/Grafana/monitoring posts while `content-guide.md` §2 lists monitoring under
|
||||
`devops`, "my most differentiated material" — and `monitoring`/`grafana no data`/`smartctl exit code` are heavily
|
||||
searched by exactly the audience this blog targets.
|
||||
- **Done when:** ✅ 8 pages 200 with expected content, language switch links both ways, 8 images served as `image/png`,
|
||||
archive order still monotonic on `/posts/`, the homepage and `/zh/`.
|
||||
|
||||
### Phase C — Discovery & structure (Tier 2)
|
||||
|
||||
**Step C1 — Per-post custom OG images (at least for case studies).**
|
||||
|
||||
|
After Width: | Height: | Size: 85 KiB |
|
After Width: | Height: | Size: 80 KiB |
|
After Width: | Height: | Size: 84 KiB |
|
After Width: | Height: | Size: 82 KiB |
|
After Width: | Height: | Size: 48 KiB |
|
After Width: | Height: | Size: 47 KiB |
|
After Width: | Height: | Size: 52 KiB |
|
After Width: | Height: | Size: 46 KiB |
@@ -787,6 +787,85 @@ BANNERS['why-chrome-forgets-its-tabs-in-a-container'] = {
|
||||
],
|
||||
};
|
||||
|
||||
BANNERS['smartctl-exit-code-32-skips-the-disks-that-matter'] = {
|
||||
titlebar: 'root@unraid — disk health · textfile collector',
|
||||
lines: [
|
||||
{ t: 'prompt', text: '$' }, { t: 'cmd', text: 'for dev in /dev/sd?; do smartctl -A "$dev"' },
|
||||
{ t: 'err', text: 'rc=32 ← "OK, but attributes were below threshold"' },
|
||||
{ t: 'dim', text: 'treated as unreadable → disk skipped' },
|
||||
{ t: 'err', text: '2 of 4 SSDs missing · the marginal ones' },
|
||||
{ t: 'prompt', text: '$' }, { t: 'cmd', text: 'fatal=$(( rc & ~(32 | 64) )) · rc as a metric' },
|
||||
{ t: 'ok', text: '→ 4/4 disks collected ✓' },
|
||||
],
|
||||
flow: [
|
||||
{ n: '1', label: 'rc=32', err: true },
|
||||
{ n: '2', label: 'skip ✗' },
|
||||
{ n: '3', label: '2/4 disks' },
|
||||
{ n: '4', label: 'mask bits' },
|
||||
{ n: '5', label: '4/4 ✓' },
|
||||
],
|
||||
};
|
||||
|
||||
BANNERS['why-your-grafana-dashboard-shows-no-data'] = {
|
||||
titlebar: 'root@grafana — 41 panels · No data',
|
||||
lines: [
|
||||
{ t: 'prompt', text: '$' }, { t: 'cmd', text: 'curl -s prometheus:9090/api/v1/targets' },
|
||||
{ t: 'ok', text: 'all 7 targets: up' },
|
||||
{ t: 'err', text: 'every panel: "No data"' },
|
||||
{ t: 'dim', text: 'my check substituted the values by hand ✓ ← bug invisible' },
|
||||
{ t: 'prompt', text: '$' }, { t: 'cmd', text: '$nodename = label_values(...{nodename=~"$nodename"})' },
|
||||
{ t: 'err', text: '← variable filters on itself → 0 options' },
|
||||
{ t: 'ok', text: '→ no self-reference + saved current · 21/25 ✓' },
|
||||
],
|
||||
flow: [
|
||||
{ n: '1', label: 'targets up' },
|
||||
{ n: '2', label: 'panels ✗' },
|
||||
{ n: '3', label: 'variables' },
|
||||
{ n: '4', label: 'self-ref' },
|
||||
{ n: '5', label: '21/25 ✓' },
|
||||
],
|
||||
};
|
||||
|
||||
BANNERS['your-disk-full-alert-is-lying'] = {
|
||||
titlebar: 'root@monitor — alert rules · 16 total',
|
||||
lines: [
|
||||
{ t: 'err', text: 'ALERT filesystem >90% · 500 GB still free' },
|
||||
{ t: 'err', text: 'ALERT memory >90% used · 4.8 GB available' },
|
||||
{ t: 'err', text: 'ALERT CPU steal >25% · host fine at 41%' },
|
||||
{ t: 'dim', text: 'always true · never actionable' },
|
||||
{ t: 'dim', text: 'and each one trains you to skim' },
|
||||
{ t: 'prompt', text: '$' }, { t: 'cmd', text: 'alert on the consequence, not the ratio' },
|
||||
{ t: 'ok', text: '→ free <25 GB · MemAvailable <512 MB ✓' },
|
||||
],
|
||||
flow: [
|
||||
{ n: '1', label: '% full ✗', err: true },
|
||||
{ n: '2', label: '% used ✗', err: true },
|
||||
{ n: '3', label: 'steal ✗' },
|
||||
{ n: '4', label: 'absolute' },
|
||||
{ n: '5', label: 'silent ✓' },
|
||||
],
|
||||
};
|
||||
|
||||
BANNERS['one-prometheus-for-unraid-synology-and-a-vps'] = {
|
||||
titlebar: 'root@unraid — prometheus · 7 targets · 25.5k series',
|
||||
lines: [
|
||||
{ t: 'prompt', text: '$' }, { t: 'cmd', text: 'sum(node_filesystem_size_bytes)' },
|
||||
{ t: 'err', text: '64 TB "total" ← one NAS volume counted 3×' },
|
||||
{ t: 'err', text: 'KVM guest: no cpufreq → 0 series' },
|
||||
{ t: 'err', text: 'nodename = 9f9afcccc962 ← container ID' },
|
||||
{ t: 'dim', text: 'cAdvisor: systemd slices reported as containers' },
|
||||
{ t: 'prompt', text: '$' }, { t: 'cmd', text: 'dedup · textfile · hostname pin' },
|
||||
{ t: 'ok', text: '→ 47 cores · 158 GHz · 142 GB · 155 containers ✓' },
|
||||
],
|
||||
flow: [
|
||||
{ n: '1', label: '3 views', err: true },
|
||||
{ n: '2', label: 'dedup ✓' },
|
||||
{ n: '3', label: 'guest gap' },
|
||||
{ n: '4', label: 'textfile' },
|
||||
{ n: '5', label: 'one screen ✓' },
|
||||
],
|
||||
};
|
||||
|
||||
// ---------- read frontmatter ----------
|
||||
const postPath = join(ROOT, 'src', 'content', 'posts', `${slug}.md`);
|
||||
let category = 'devops';
|
||||
|
||||
@@ -243,6 +243,31 @@ TERMINALS['why-chrome-forgets-its-tabs-in-a-container'] = `
|
||||
<div class="line"><span class="prompt"> </span><span class="err">exit_type=Crashed ← the container kills the browser</span></div>
|
||||
<div class="line"><span class="prompt">$</span><span class="cmd">tabs_keeper.py · snapshot every 60s · replay via /json/new</span><span class="fix">→ restored 2/2 tabs ✓</span></div>`;
|
||||
|
||||
TERMINALS['smartctl-exit-code-32-skips-the-disks-that-matter'] = `
|
||||
<div class="line"><span class="prompt">$</span><span class="cmd">for dev in /dev/sd?; do smartctl -A "$dev" || continue; done</span></div>
|
||||
<div class="line"><span class="prompt"> </span><span class="err">rc=32 · "disk OK, attributes were below threshold in the past"</span></div>
|
||||
<div class="line"><span class="prompt"> </span><span class="err">→ 2 of 4 SSDs silently skipped · exactly the marginal ones</span></div>
|
||||
<div class="line"><span class="prompt">$</span><span class="cmd">fatal=$(( rc & ~(32 | 64) )) · rc exported as a metric</span><span class="fix">→ 4/4 collected ✓</span></div>`;
|
||||
|
||||
TERMINALS['why-your-grafana-dashboard-shows-no-data'] = `
|
||||
<div class="line"><span class="prompt">$</span><span class="cmd">curl -s prometheus:9090/api/v1/targets | jq .[].health</span><span class="fix">→ all 7 up</span></div>
|
||||
<div class="line"><span class="prompt">$</span><span class="cmd">curl -sG /api/v1/query --data-urlencode 'query=node_uname_info'</span></div>
|
||||
<div class="line"><span class="prompt"> </span><span class="err">data is there · panel expression returns rows · dashboard: No data</span></div>
|
||||
<div class="line"><span class="prompt"> </span><span class="err">$nodename = label_values(...{nodename=~"$nodename"}) ← filters on itself</span></div>
|
||||
<div class="line"><span class="prompt">$</span><span class="cmd">drop the self-reference · save current values</span><span class="fix">→ 21/25 panels ✓</span></div>`;
|
||||
|
||||
TERMINALS['your-disk-full-alert-is-lying'] = `
|
||||
<div class="line"><span class="prompt"> </span><span class="err">ALERT filesystem above 90% · 17 TB volume · 500 GB still free</span></div>
|
||||
<div class="line"><span class="prompt"> </span><span class="err">ALERT memory above 90% used · 4.8 GB actually available</span></div>
|
||||
<div class="line"><span class="prompt"> </span><span class="err">ALERT CPU steal above 25% · host fine at 41%</span></div>
|
||||
<div class="line"><span class="prompt">$</span><span class="cmd">alert on the consequence, not the ratio</span><span class="fix">→ silent when healthy ✓</span></div>`;
|
||||
|
||||
TERMINALS['one-prometheus-for-unraid-synology-and-a-vps'] = `
|
||||
<div class="line"><span class="prompt">$</span><span class="cmd">sum(node_filesystem_size_bytes) → 64 TB "total"</span></div>
|
||||
<div class="line"><span class="prompt"> </span><span class="err">one NAS volume counted 3×: /volume1 · /opt · CIFS re-mount</span></div>
|
||||
<div class="line"><span class="prompt"> </span><span class="err">KVM guest: no cpufreq · 0 series nodename = 9f9afcccc962</span></div>
|
||||
<div class="line"><span class="prompt">$</span><span class="cmd">dedup selector · textfile collector · hostname pin</span><span class="fix">→ 7 targets ✓</span></div>`;
|
||||
|
||||
// ---------- read frontmatter ----------
|
||||
const postPath = join(ROOT, 'src', 'content', 'posts', `${slug}.md`);
|
||||
if (!existsSync(postPath)) {
|
||||
|
||||
@@ -0,0 +1,178 @@
|
||||
---
|
||||
title: "One Prometheus for unRaid, a Synology NAS, and a VPS: What I Got Wrong"
|
||||
description: "Three hosts, one Prometheus, one Grafana. The coverage matrix, the storage total that counted one NAS volume three times, and the VM that exports no CPU frequency at all."
|
||||
pubDate: 2026-07-29
|
||||
updatedDate: 2026-09-29
|
||||
category: case-studies
|
||||
tags: [prometheus, grafana, unraid, synology, dsm, cadvisor, monitoring, homelab]
|
||||
ogImage: /og/one-prometheus-for-unraid-synology-and-a-vps.png
|
||||
banner: /banners/one-prometheus-for-unraid-synology-and-a-vps.png
|
||||
draft: false
|
||||
---
|
||||
|
||||
## Why this matters
|
||||
|
||||
A homelab is not a toy when it's running other people's websites, mail and files. The
|
||||
difference between a bad week and a bad *month* is usually how early you found out: a
|
||||
disk with a reallocated sector, a volume with 20 GB left, a VM being starved of CPU by
|
||||
its host, a container that restarted four times overnight while you slept.
|
||||
|
||||
I had three machines — unRaid (the daily driver), a Synology NAS (the storage and
|
||||
services box) and a VPS (the public-facing one) — and three separate ways of *not*
|
||||
knowing what they were doing. This is how they became one screen, and the four things I
|
||||
got wrong on the way, because three of them are traps you will hit too.
|
||||
|
||||
## The architecture
|
||||
|
||||
One Prometheus and one Grafana, running on the host that never sleeps. Every monitored
|
||||
host runs two exporters:
|
||||
|
||||
| Machine | Host metrics | Container metrics | Scrape path |
|
||||
|---|---|---|---|
|
||||
| unRaid | node_exporter (`:9100`) | cAdvisor (`:8080`) | direct |
|
||||
| Synology DSM | node_exporter (`:9100`) | cAdvisor (`:8082`) | LAN |
|
||||
| ServerHosh VPS | node_exporter (`:9100`) | cAdvisor (`:8081`) | VPN tunnel + `nginx` stream relay |
|
||||
|
||||
**node_exporter and cAdvisor are not substitutes — this is the first thing people get
|
||||
wrong.** node_exporter sees the *host*: CPU, memory, network, disks, filesystems,
|
||||
temperatures. It is cgroup-blind: it cannot tell you which container is eating your RAM.
|
||||
cAdvisor sees *containers only*. If you want both host health and per-container
|
||||
accounting, you run both. Dropping either one leaves a hole that shows up months later
|
||||
as an unexplained load spike.
|
||||
|
||||
## Wrong #1: cAdvisor's "containers" that aren't containers
|
||||
|
||||
My container rules started firing on things that were not containers. cAdvisor exports
|
||||
cgroup series for **systemd slices** and for the machine-wide cgroup — including one
|
||||
unnamed series that reported roughly 58 GB of "memory usage" with no container attached
|
||||
to it. That is the whole host, described as a container.
|
||||
|
||||
The fix is one filter, but you have to know to write it:
|
||||
|
||||
```promql
|
||||
# container metrics: anything with a name, and only that
|
||||
container_memory_working_set_bytes{job=~"cadvisor.*", name!=""}
|
||||
```
|
||||
|
||||
Without `name!=""`, a "container using more than 10 GB" rule alerts on the host itself.
|
||||
|
||||
## Wrong #2: my "total storage" number was fiction
|
||||
|
||||
The aggregate panel looked great and was simply wrong. The reason: **one filesystem
|
||||
appears in the metrics more than once.**
|
||||
|
||||
- On the NAS, the primary volume is `/volume1` — and `/opt` is the *same* btrfs
|
||||
filesystem, exposed under a second mount point.
|
||||
- On unRaid, that same NAS volume is mounted again over CIFS as
|
||||
`/mnt/remotes/<nas>_ActiveBackup`.
|
||||
- unRaid's `/var/lib/docker` is a subvolume of the pool that `/mnt/ssd` already
|
||||
represents — same bytes, second identity.
|
||||
|
||||
A naive `sum(node_filesystem_size_bytes)` therefore reported capacity for disks I don't
|
||||
have. The honest selector names exactly what counts as data storage:
|
||||
|
||||
```promql
|
||||
node_filesystem_size_bytes{
|
||||
mountpoint=~"/volume[0-9]+|/mnt/ssd|/mnt/disk[0-9]+",
|
||||
fstype!~"fuse.*|tmpfs|rootfs"
|
||||
} or node_filesystem_size_bytes{job="vps-host", mountpoint="/"}
|
||||
```
|
||||
|
||||
Two more honesty notes that belong *in the panel description*, because a number nobody
|
||||
can interpret is worse than no number:
|
||||
|
||||
- **Parity disks are invisible to the kernel.** unRaid's parity drive has no filesystem,
|
||||
so it never appears — the sum is usable capacity, not raw spindle count.
|
||||
- **A mirror inflates a raw sum.** Two mirrored SSDs report their bytes twice; the sum
|
||||
is not what you can store.
|
||||
|
||||
And the omission that would have bitten me later: unRaid's shfs union mount (`/mnt/user`)
|
||||
reports `avail=0`. Including it in a "free space below X" rule produces an alert that
|
||||
can never be resolved.
|
||||
|
||||
## Wrong #3: the VM that exports no CPU frequency
|
||||
|
||||
Wanting "total CPU frequency" across the lab, I reached for node_exporter's cpufreq
|
||||
collector. unRaid and the NAS reported `node_cpu_scaling_frequency_hertz` per core. The
|
||||
VPS reported **nothing at all** — it's a KVM guest, and a guest has no
|
||||
`/sys/devices/system/cpu/cpu0/cpufreq`. The metric cannot exist there.
|
||||
|
||||
The fix is the textfile collector: a small script that reads what the guest *can* see
|
||||
(`/proc/cpuinfo`) and writes Prometheus-format metrics into a directory node_exporter
|
||||
scrapes:
|
||||
|
||||
```sh
|
||||
# /opt/node-exporter-textfile/cpu-mhz.sh — run from cron every 5 minutes
|
||||
awk -F: '
|
||||
/^processor/ { c = $2; gsub(/[ \t]/, "", c) }
|
||||
/^cpu MHz/ { f = $2; gsub(/[ \t]/, "", f); printf "node_cpu_mhz_current_hz{core=\"%s\"} %.0f\n", c, f * 1000000 }
|
||||
' /proc/cpuinfo
|
||||
```
|
||||
|
||||
with
|
||||
|
||||
```yaml
|
||||
command:
|
||||
- '--collector.textfile.directory=/textfile'
|
||||
volumes:
|
||||
- /opt/node-exporter-textfile:/textfile:ro
|
||||
```
|
||||
|
||||
Two deliberate decisions: the metric is published under a **different name** than
|
||||
node_exporter's own (`node_cpu_mhz_current_hz`, not `node_cpu_scaling_frequency_hertz`)
|
||||
so that if the host ever exposes real cpufreq there is no duplicate-series conflict, and
|
||||
the dashboards merge the two sources explicitly:
|
||||
|
||||
```promql
|
||||
sum(node_cpu_scaling_frequency_hertz or node_cpu_mhz_current_hz)
|
||||
```
|
||||
|
||||
## Wrong #4: my exporter was reporting the container's name as the host's
|
||||
|
||||
The host dropdown in the dashboard offered `9f9afcccc962` as a machine. That's a container
|
||||
ID: node_exporter's `uname` collector reports the *process's* UTS namespace, and a
|
||||
container's hostname defaults to its own ID. The metric was correct, the label was
|
||||
meaningless. One line in compose fixes it:
|
||||
|
||||
```yaml
|
||||
services:
|
||||
node_exporter:
|
||||
hostname: 2.hoelee.com # otherwise `nodename` = the container ID
|
||||
```
|
||||
|
||||
## What I'd do differently
|
||||
|
||||
- **Build the aggregate/overview screen last, deliberately.** Deciding "what is total
|
||||
storage?" is what forces the deduplication work — discovering it after building three
|
||||
host dashboards means rebuilding the number everywhere.
|
||||
- **Pin container `hostname` from the start.** It costs one line and saves a confusing
|
||||
dropdown.
|
||||
- **Assume every virtualised or appliance host hides one class of metric.** A NAS may
|
||||
cap your container monitor (DSM's Docker API version pinned my cAdvisor to v0.53.0 —
|
||||
newer releases need a newer Docker API); a VM hides cpufreq; a router hides its own
|
||||
CPU. Find the gap by *counting what you expected* rather than trusting that the
|
||||
collector succeeded.
|
||||
|
||||
## The result
|
||||
|
||||
Three hosts, **7 scrape targets, one Grafana, 16 alert rules**, and a single screen that
|
||||
reads: **47 CPU cores · 158 GHz currently clocked (205 GHz nominal) · 142 GB RAM · 64 TB
|
||||
storage · 155 running containers**, with the same host and container dashboard layout
|
||||
cloned per machine so the three read identically. Alerts land in Telegram and email;
|
||||
Prometheus keeps 30 days.
|
||||
|
||||
The most useful outcome wasn't the dashboard. It was the alert that fired the day the
|
||||
storage rule was rewritten: a NAS volume with 21 GB left, which had been hidden behind a
|
||||
percentage threshold on a volume so large that "99% full" had become background noise.
|
||||
|
||||
## Want this for your business?
|
||||
|
||||
If you run a NAS, a VPS and a couple of servers and have no single place to see them, I
|
||||
set up self-hosted monitoring pipelines like this one — Prometheus + Grafana across your
|
||||
machines, host *and* container metrics, disk-health and capacity rules that mean
|
||||
something, and alerts pushed to Telegram or email.
|
||||
|
||||
**WhatsApp: [+60 12-797 2969](https://wa.me/60127972969)** · **Email: [[email protected]](mailto:[email protected]?subject=Self-hosted%20monitoring%20setup)** · **[hoelee.com](https://hoelee.com)**
|
||||
|
||||
Website design and development is my main line of work; server hardening and
|
||||
self-hosted infrastructure is the other half of it.
|
||||
@@ -0,0 +1,121 @@
|
||||
---
|
||||
title: "smartctl Exit Code 32: The Code That Skips the Disks You Care About"
|
||||
description: "A disk-health collector that treats any non-zero smartctl exit as failure quietly skips exactly the drives with marginal attributes. Exit 32 is not an error — it is a history lesson."
|
||||
pubDate: 2026-04-14
|
||||
updatedDate: 2026-09-29
|
||||
category: notes
|
||||
tags: [smart, smartctl, unraid, monitoring, bash, disks, homelab]
|
||||
ogImage: /og/smartctl-exit-code-32-skips-the-disks-that-matter.png
|
||||
banner: /banners/smartctl-exit-code-32-skips-the-disks-that-matter.png
|
||||
draft: false
|
||||
---
|
||||
|
||||
My disk-health collector was green for weeks. Every drive reported a temperature, a
|
||||
power-on-hours counter and a SMART verdict, and the textfile metrics looked complete
|
||||
until I counted them: **two of the four SSDs were simply missing**.
|
||||
|
||||
Not failing. Missing. The script had decided they didn't exist.
|
||||
|
||||
## The script that "handled" errors
|
||||
|
||||
The collector walks the block devices, runs `smartctl` against each one, and writes
|
||||
Prometheus textfile metrics. The error handling looked defensive — the classic shell
|
||||
shape:
|
||||
|
||||
```bash
|
||||
for dev in /dev/sd?; do
|
||||
if ! smartctl -A -H "$dev" > /tmp/smart.out 2>&1; then
|
||||
continue # "disk isn't SMART-capable / can't be read"
|
||||
fi
|
||||
# parse and emit metrics
|
||||
done
|
||||
```
|
||||
|
||||
`if ! cmd` is a boolean. `smartctl`'s exit status is **not** a boolean — it's a
|
||||
bitfield, and treating a bitfield as true/false is where this goes wrong.
|
||||
|
||||
## What exit code 32 actually means
|
||||
|
||||
From `man smartctl`, the bits are cumulative and independent:
|
||||
|
||||
| Bit | Value | Meaning |
|
||||
|---|---|---|
|
||||
| 0 | 1 | Command line did not parse |
|
||||
| 1 | 2 | Device open failed, or no IDENTIFY DEVICE structure |
|
||||
| 2 | 4 | A SMART or ATA command failed / checksum error in a SMART structure |
|
||||
| 3 | **8** | SMART status check returned **DISK FAILING** |
|
||||
| 4 | **16** | Pre-fail attributes found **<= threshold** |
|
||||
| 5 | **32** | SMART status **OK**, but some attributes were **<= threshold at some time in the past** |
|
||||
| 6 | 64 | Device error log contains records of errors |
|
||||
| 7 | 128 | Device self-test log contains records of errors |
|
||||
|
||||
So exit `32` is the opposite of "unreadable". It means: *the disk is fine right now,
|
||||
and it has been below a threshold before.* That is precisely the signal you want to
|
||||
keep an eye on — and my collector was throwing it away.
|
||||
|
||||
The two disks that vanished were the two whose raw attributes sit at marginal values.
|
||||
The healthy disks exited `0` and got collected. **The filter was selecting for the
|
||||
disks with nothing to report.**
|
||||
|
||||
## The fix: mask the informational bits
|
||||
|
||||
Exit bits 32 and 64 are informational for monitoring purposes; bits 8 and 16 are the
|
||||
ones that deserve an alert. Mask them off and act on what's left:
|
||||
|
||||
```bash
|
||||
smartctl -A -H -d sat "$dev" > /tmp/smart.out 2>&1
|
||||
rc=$?
|
||||
|
||||
# fatal bits: 1 (parse), 2 (open), 4 (command/checksum), 8 (FAILING)
|
||||
# informational bits: 32 (was below threshold in the past), 64 (error log has records)
|
||||
fatal=$(( rc & ~(32 | 64) ))
|
||||
if [ "$fatal" -ne 0 ]; then
|
||||
echo "device $dev unreadable or failing (rc=$rc)" >&2
|
||||
continue
|
||||
fi
|
||||
|
||||
# emit the verdict AND the exit code, so the code itself is a metric
|
||||
echo "disk_smart_exit_code{device=\"$dev\"} $rc"
|
||||
echo "disk_smart_health{device=\"$dev\"} $(( rc & 8 ? 0 : 1 ))"
|
||||
```
|
||||
|
||||
Three things changed the value of this collector:
|
||||
|
||||
1. **The exit code became data, not control flow.** It's exported as a metric, so a
|
||||
disk drifting from `0` → `32` → `64` shows up as a trend instead of a silent skip.
|
||||
2. **The informational bits stopped being fatal.** Disks with history are collected,
|
||||
which is the whole point of monitoring them.
|
||||
3. **`-d sat` matters on a NAS.** On Synology DSM (and some USB bridges), a SATA disk
|
||||
behind the wrong device type returns nothing useful — the same "missing disk"
|
||||
symptom with a different cause.
|
||||
|
||||
## What I'd do differently
|
||||
|
||||
- **Never write `if ! cmd` against a tool that documents an exit-code bitfield.** Read
|
||||
the exit-code section of the man page before using the return value as a boolean.
|
||||
- **Count what you collected, not just whether the collector ran.** My alert was on
|
||||
"collector stale"; the real bug was "collector ran fine and reported 50% of the
|
||||
disks". A single `disk_count` metric would have surfaced it immediately.
|
||||
- **Treat "no data for this device" as its own alertable state.** A missing series is
|
||||
invisible, which is exactly why it survived for weeks.
|
||||
|
||||
## The result
|
||||
|
||||
The collector went from 2 usable disks to 4, with every drive reporting temperature,
|
||||
power-on hours, SMART status and its raw exit code — 122 textfile metrics in total,
|
||||
including the btrfs error counters I actually wanted. The two disks that reappeared
|
||||
are the two that were already sitting at marginal attribute values.
|
||||
|
||||
A monitoring pipeline that skips its own worst signals is worse than no pipeline,
|
||||
because it tells you everything is fine.
|
||||
|
||||
## Want this for your business?
|
||||
|
||||
If you're running a NAS or a server rack and want disk health that actually pages you
|
||||
before a drive dies — SMART attributes, temperatures, btrfs/RAID error counters, and
|
||||
alerts to Telegram or email — that's the kind of self-hosted monitoring pipeline I set up.
|
||||
|
||||
**WhatsApp: [+60 12-797 2969](https://wa.me/60127972969)** · **Email: [[email protected]](mailto:[email protected]?subject=Disk%20health%20monitoring)** · **[hoelee.com](https://hoelee.com)**
|
||||
|
||||
Website design and development is my main line of work; server hardening and
|
||||
self-hosted infrastructure is the other half of it.
|
||||
@@ -0,0 +1,156 @@
|
||||
---
|
||||
title: "Why Your Grafana Dashboard Shows No Data (When Prometheus Is Fine)"
|
||||
description: "Every target up, data in Prometheus, panel expressions correct — and all 41 panels reading No data. The culprit was a Grafana variable filtering on itself, plus the verification habit that hid it."
|
||||
pubDate: 2026-05-17
|
||||
updatedDate: 2026-09-29
|
||||
category: devops
|
||||
tags: [grafana, prometheus, dashboards, monitoring, promql, observability]
|
||||
ogImage: /og/why-your-grafana-dashboard-shows-no-data.png
|
||||
banner: /banners/why-your-grafana-dashboard-shows-no-data.png
|
||||
draft: false
|
||||
---
|
||||
|
||||
You open the dashboard you use every day and it's empty. Not "one panel broke" empty —
|
||||
**every panel**, top to bottom, says *No data*. You check Prometheus and everything is
|
||||
healthy. This post is the specific, non-obvious cause I hit, and the way my own
|
||||
verification convinced me the dashboard was fine while it wasn't.
|
||||
|
||||
## The searchable question
|
||||
|
||||
*My Grafana panels show No data but the data is in Prometheus — why?*
|
||||
|
||||
The usual answers (wrong datasource, wrong time range, `rate()` on a counter that
|
||||
resets, missing scrape target) did not apply. Here is the checklist that narrowed it
|
||||
down, in the order that costs the least time:
|
||||
|
||||
```bash
|
||||
# 1. is the target actually being scraped?
|
||||
curl -s http://prometheus:9090/api/v1/targets | jq '.data.activeTargets[] | {job:.labels.job, health}'
|
||||
|
||||
# 2. does the metric exist right now? (bypasses Grafana entirely)
|
||||
curl -s -G http://prometheus:9090/api/v1/query \
|
||||
--data-urlencode 'query=node_uname_info{job="unraid-host"}' | jq '.data.result[0].metric'
|
||||
|
||||
# 3. is the panel's expression sound?
|
||||
curl -s -G http://prometheus:9090/api/v1/query \
|
||||
--data-urlencode 'query=count(node_cpu_seconds_total{mode="idle",job="unraid-host"})'
|
||||
```
|
||||
|
||||
All three were green: 7 targets `up`, 25k series ingested, and the panel's raw
|
||||
expression returned data when I ran it by hand. So the problem was not the metrics,
|
||||
the scrape, or the PromQL — it was **the variable layer between them**.
|
||||
|
||||
## The actual cause: a variable that filtered on itself
|
||||
|
||||
The dashboard had a host picker. Two variables chained: `$nodename` (which machine)
|
||||
and `$node` (its `instance` label). The definition of the first one was:
|
||||
|
||||
```promql
|
||||
# BROKEN — the variable filters on its own value
|
||||
label_values(node_uname_info{job="unraid-host", nodename=~"$nodename"}, nodename)
|
||||
```
|
||||
|
||||
On first load `$nodename` has no value, so the selector becomes
|
||||
`nodename=~""` — which matches **nothing**. A variable whose query returns no rows has
|
||||
no options, so it stays empty. `$node` is then defined against `$nodename`:
|
||||
|
||||
```promql
|
||||
label_values(node_uname_info{job="unraid-host", nodename="$nodename"}, instance)
|
||||
```
|
||||
|
||||
…which is also empty, so every panel filtering on `$node` matches no series. Grafana
|
||||
isn't lying when it says *No data*: the panel query genuinely has no data, because the
|
||||
variable that scopes it resolved to nothing.
|
||||
|
||||
The fix is to stop the self-reference and give each variable an explicit saved default:
|
||||
|
||||
```promql
|
||||
# FIXED — no self-reference, one hop per variable
|
||||
# $nodename
|
||||
label_values(node_uname_info{job="unraid-host"}, nodename)
|
||||
# $node
|
||||
label_values(node_uname_info{job="unraid-host", nodename="$nodename"}, instance)
|
||||
```
|
||||
|
||||
Then save `current` values for both so the dashboard opens populated instead of
|
||||
depending on a click. (This is also why the same 41-panel layout can be cloned to
|
||||
three hosts and just work.)
|
||||
|
||||
## Why my verification missed it
|
||||
|
||||
This is the part worth stealing. I "verified" the dashboard by expanding each panel's
|
||||
expression **by hand**, substituting the variable values I knew were correct
|
||||
(`$nodename=unRaid`, `$node=host.docker.internal:9100`), and asserting the query
|
||||
returned points. Thirteen of those checks passed. The dashboard was still blank,
|
||||
because the bug was upstream of the substitution: Grafana's *own* resolution of the
|
||||
variables was empty.
|
||||
|
||||
**A verification that substitutes the very value under suspicion cannot detect that
|
||||
value failing to resolve.** To catch it, read the values Grafana actually saved and
|
||||
expand with those:
|
||||
|
||||
```bash
|
||||
curl -s -u admin:"$PW" http://grafana:4010/api/dashboards/uid/rYdddlPWk \
|
||||
| jq '.dashboard.templating.list[] | {name, current: .current.value}'
|
||||
```
|
||||
|
||||
The correct check is on the variable query itself — run what Grafana runs:
|
||||
|
||||
```bash
|
||||
# what does $nodename's label_values() actually return?
|
||||
curl -s -G http://prometheus:9090/api/v1/query \
|
||||
--data-urlencode 'query=count by (nodename) (node_uname_info{job="unraid-host"})' | jq '.data.result'
|
||||
```
|
||||
|
||||
Empty there means the dashboard cannot possibly render, no matter how healthy the data is.
|
||||
|
||||
## The second trap: `$__all` is not `.*`
|
||||
|
||||
While automating panel checks, a variable set to *All* reports its current value as the
|
||||
sentinel string `$__all` — not a regex. Expanding panel expressions naively turns
|
||||
`name=~"$__all"` into a literal that matches nothing, so a perfectly healthy
|
||||
container dashboard "fails" every panel. Resolve `$__all` to the variable's `allValue`
|
||||
(usually `.*`) before expanding, or you'll chase a bug that only exists in your checker.
|
||||
|
||||
## Not every empty panel is a bug
|
||||
|
||||
After the fix, 21 of 25 panels had data. The remaining four were empty **by design** and
|
||||
worth naming honestly in the panel description:
|
||||
|
||||
- **PSI panels** (`Pressure`, `Pressure Stall Information`) need `/proc/pressure`, which
|
||||
the NAS kernel doesn't expose. The metric cannot exist there — the same panels on a
|
||||
host that *does* export `node_pressure_*` render fine.
|
||||
- **Root filesystem panels** whose expression carries `fstype!="rootfs"` can never match
|
||||
on a host whose `/` is a **ramdisk**, which is what unRaid's root is.
|
||||
|
||||
Distinguishing "broken because of a bug" from "empty because the metric cannot exist on
|
||||
this host" is the difference between a real fix and a wild goose chase.
|
||||
|
||||
## What I'd do differently
|
||||
|
||||
- **Never define a template variable that filters on itself.** One hop per variable.
|
||||
- **Always save explicit `current` values** for variables a dashboard depends on, so a
|
||||
fresh load doesn't rely on a populated dropdown.
|
||||
- **Verify with Grafana's saved values, not my idea of them.** The masking is invisible
|
||||
in the test result and obvious in the UI.
|
||||
- **Describe intentional gaps in the panel description.** Future-you will otherwise
|
||||
spend an afternoon "fixing" a panel that was never meant to work there.
|
||||
|
||||
## The result
|
||||
|
||||
All three hosts now run the same instrumented layout — a 41-panel host dashboard and a
|
||||
10-panel container dashboard each — with 21/23/25 panels returning data respectively,
|
||||
and the two structurally-empty classes documented in place instead of silently
|
||||
confusing whoever opens them next.
|
||||
|
||||
## Want this for your business?
|
||||
|
||||
If you have Grafana dashboards that nobody trusts — or servers and a NAS with no
|
||||
monitoring at all — I build self-hosted Prometheus + Grafana stacks (host and container
|
||||
metrics, sensible alert rules, alerts to Telegram and email) and I'll fix the ones that
|
||||
have quietly stopped showing data.
|
||||
|
||||
**WhatsApp: [+60 12-797 2969](https://wa.me/60127972969)** · **Email: [[email protected]](mailto:[email protected]?subject=Grafana%20monitoring)** · **[hoelee.com](https://hoelee.com)**
|
||||
|
||||
Website design and development is my main line of work; server hardening and
|
||||
self-hosted infrastructure is the other half of it.
|
||||
@@ -0,0 +1,125 @@
|
||||
---
|
||||
title: "Your Disk-Full Alert Is Lying: Stop Alerting on Percentages"
|
||||
description: "3% free sounds like an emergency until you notice it is 500 GB. On multi-terabyte volumes, percentage thresholds fire when nothing is wrong — and miss when something is."
|
||||
pubDate: 2026-06-24
|
||||
updatedDate: 2026-09-29
|
||||
category: devops
|
||||
tags: [monitoring, alerting, prometheus, grafana, storage, capacity, observability]
|
||||
ogImage: /og/your-disk-full-alert-is-lying.png
|
||||
banner: /banners/your-disk-full-alert-is-lying.png
|
||||
draft: false
|
||||
---
|
||||
|
||||
I built a monitoring stack, wrote a reasonable set of rules, and then spent an
|
||||
afternoon deleting most of my own thresholds. Every one of them was a percentage, and
|
||||
every one of them was wrong in the same specific way.
|
||||
|
||||
## The three alerts that taught me
|
||||
|
||||
**1. "Filesystem above 90% used."** It fired on a 17 TB volume that had ~500 GB free.
|
||||
Nothing was wrong. Nothing was *about* to be wrong. On a volume that size, a single
|
||||
large backup can swing the percentage by several points — the ratio is noise, and the
|
||||
fact that 500 GB is still available is the actual state of the world.
|
||||
|
||||
**2. "Memory above 90% used."** Fired constantly on a host with 4.8 GB *available*.
|
||||
Linux uses free RAM for page cache and gives it back on demand; `MemTotal - MemFree`
|
||||
describes your filesystem cache, not your risk. The number that predicts an OOM is
|
||||
`MemAvailable`.
|
||||
|
||||
**3. "CPU steal above 25%."** Fired on a VPS that had been running its mail stack
|
||||
perfectly at 41% steal. Steal means the hypervisor is busy — it is a *capacity* signal,
|
||||
not a *failure* signal, and 41% steal on a 7-vCPU box was simply what that host costs.
|
||||
The rule produced alerts that were always true and never actionable.
|
||||
|
||||
The common thread: **I was alerting on a ratio, and the ratio doesn't know how big the
|
||||
thing is.**
|
||||
|
||||
## Alert on the consequence, not the ratio
|
||||
|
||||
The test I now apply to every threshold: *if this condition persists, what will actually
|
||||
happen?* That answer is almost always expressible in absolute units.
|
||||
|
||||
| Instead of | Alert on | Because |
|
||||
|---|---|---|
|
||||
| Filesystem > 90% | Free space < N GB | "Can I still write?" is the question |
|
||||
| Memory > 90% used | `MemAvailable` < 512 MB | about to OOM, not "cache is warm" |
|
||||
| CPU steal > 25% | Steal > 50% | capacity cost vs. starvation |
|
||||
| Load average > N | `load1 / cores > 2` | load is meaningless without core count |
|
||||
|
||||
In PromQL, before and after:
|
||||
|
||||
```promql
|
||||
# BEFORE — percentage of a 17 TB pool; fires with half a terabyte still free
|
||||
(1 - node_filesystem_avail_bytes / node_filesystem_size_bytes) * 100 > 90
|
||||
|
||||
# AFTER — capacity left, on the volumes that hold data
|
||||
min by (instance, mountpoint) (
|
||||
node_filesystem_avail_bytes{mountpoint=~"/volume[0-9]+|/mnt/ssd|/mnt/disk[0-9]+"}
|
||||
) < 25e9
|
||||
```
|
||||
|
||||
```promql
|
||||
# BEFORE — page cache counted as "pressure"
|
||||
(1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100 > 90
|
||||
|
||||
# AFTER — the actual OOM precursor
|
||||
node_memory_MemAvailable_bytes < 512 * 1024 * 1024
|
||||
```
|
||||
|
||||
Two details that make these rules survive contact with a real dashboard:
|
||||
|
||||
**Keep the comparison in the threshold condition, not inside the expression.** Query
|
||||
`min by (mountpoint) (node_filesystem_avail_bytes{...})`, then let the alert rule
|
||||
compare it to `25000000000`. If you bake `< 25e9` into the PromQL *and* leave an
|
||||
inherited `> 90` in the rule, the rule still evaluates — but it displays a nonsense
|
||||
threshold, and the next person to read it has to reconstruct your intent from two
|
||||
contradictory conditions.
|
||||
|
||||
**Exclude the mounts that are views of other mounts.** A "free space" rule is only as
|
||||
honest as its selector. On my own setup, one NAS volume appears three times (as
|
||||
`/volume1`, as `/opt`, and again on another host as a CIFS re-mount), and unRaid's shfs
|
||||
union mount reports `avail=0` — a rule without exclusions would either fire forever or
|
||||
double-count capacity. One explicit selector beats a clever generic one:
|
||||
|
||||
```promql
|
||||
node_filesystem_avail_bytes{mountpoint=~"/volume[0-9]+|/mnt/ssd|/mnt/disk[0-9]+"}
|
||||
```
|
||||
|
||||
## Noise is not harmless
|
||||
|
||||
The cost of a bad threshold isn't the alert itself — it's the *training*. Every alert
|
||||
that fires while nothing is wrong teaches you to skim, then to ignore, and then the one
|
||||
that matters arrives in a channel you've stopped reading. My rule count went **down**
|
||||
while coverage went up: fewer, each tied to a named consequence.
|
||||
|
||||
## Thresholds are personal, shapes are not
|
||||
|
||||
The specific numbers above are mine — 25 GB free before a NAS volume bothers me, 512 MB
|
||||
available before I care about memory, 50% steal before a VPS is being starved. Yours
|
||||
will differ with your hardware and your tolerance.
|
||||
|
||||
What transfers is the shape:
|
||||
|
||||
- absolute units over ratios,
|
||||
- the consequence written into the alert's summary,
|
||||
- a selector that states exactly which mounts it speaks for,
|
||||
- and a rule count small enough that you still read every message.
|
||||
|
||||
## The result
|
||||
|
||||
The same stack that was producing three permanent, uninformative warnings now runs 16
|
||||
rules that stay silent when the lab is healthy — and one of them immediately surfaced a
|
||||
NAS volume sitting at 21 GB free, which is the kind of thing a percentage rule had been
|
||||
hiding behind a 99%-full giant that I'd learned to ignore.
|
||||
|
||||
## Want this for your business?
|
||||
|
||||
If your monitoring sends you alerts you've learned to ignore, or you have none and would
|
||||
rather find out about a full disk from a rule than from a failed backup, I set up
|
||||
self-hosted monitoring and alerting (Prometheus + Grafana, alerts to Telegram and email)
|
||||
with thresholds tied to what actually breaks.
|
||||
|
||||
**WhatsApp: [+60 12-797 2969](https://wa.me/60127972969)** · **Email: [[email protected]](mailto:[email protected]?subject=Monitoring%20and%20alerting)** · **[hoelee.com](https://hoelee.com)**
|
||||
|
||||
Website design and development is my main line of work; server hardening and
|
||||
self-hosted infrastructure is the other half of it.
|
||||
@@ -0,0 +1,125 @@
|
||||
---
|
||||
title: "用一个 Prometheus 监控 unRaid、群晖 NAS 和 VPS:我做错的四件事"
|
||||
description: "三台主机、一个 Prometheus、一个 Grafana:覆盖矩阵、把同一个 NAS 卷算成三次的\"总容量\",以及那台完全不导出 CPU 频率的虚拟机。"
|
||||
pubDate: 2026-07-29
|
||||
updatedDate: 2026-09-29
|
||||
category: case-studies
|
||||
tags: [prometheus, grafana, unraid, synology, dsm, cadvisor, monitoring, homelab]
|
||||
ogImage: /og/one-prometheus-for-unraid-synology-and-a-vps.png
|
||||
banner: /banners/one-prometheus-for-unraid-synology-and-a-vps.png
|
||||
draft: false
|
||||
---
|
||||
|
||||
## 为什么这件事重要
|
||||
|
||||
当一台家用服务器在跑别人的网站、邮件和文件时,它就不是玩具了。糟糕的一周和糟糕的**一个月**之间,差别通常只是你多早发现:一块出现重分配扇区的盘、一个只剩 20 GB 的卷、一台被宿主机饿死的虚拟机、一个夜里重启四次的容器。
|
||||
|
||||
我有三台机器——unRaid(日常主力)、群晖 NAS(存储与服务)和 VPS(对外)——以及三套**各自**都不知道它们在干什么的方式。这篇文章讲它们如何变成一块屏,以及我路上做错的四件事——其中三个你也会踩。
|
||||
|
||||
## 架构
|
||||
|
||||
一个 Prometheus 加一个 Grafana,跑在永不关机的那台机器上。每台被监控主机跑两个 exporter:
|
||||
|
||||
| 机器 | 主机指标 | 容器指标 | 抓取路径 |
|
||||
|---|---|---|---|
|
||||
| unRaid | node_exporter (`:9100`) | cAdvisor (`:8080`) | 直连 |
|
||||
| 群晖 DSM | node_exporter (`:9100`) | cAdvisor (`:8082`) | 局域网 |
|
||||
| ServerHosh VPS | node_exporter (`:9100`) | cAdvisor (`:8081`) | VPN 隧道 + `nginx` stream 中转 |
|
||||
|
||||
**node_exporter 与 cAdvisor 不是替代关系——这是最容易搞错的一点。** node_exporter 看的是**主机**:CPU、内存、网络、磁盘、文件系统、温度;它对 cgroup 一无所知,说不出是哪个容器吃掉了内存。cAdvisor 只看**容器**。想要主机健康和按容器记账,两个都得跑;少一个,那个洞会在几个月后表现为一次无法解释的负载高峰。
|
||||
|
||||
## 错误一:cAdvisor 里那些\"不是容器的容器\"
|
||||
|
||||
我的容器规则开始对不是容器的东西告警。cAdvisor 会为 **systemd slice** 和整机 cgroup 导出序列——其中一条没有名字的序列报告了约 58 GB 的\"内存使用\",却没有对应的容器。那就是整台主机,被描述成了一个容器。
|
||||
|
||||
修法是加一个过滤条件,但你得先知道要写它:
|
||||
|
||||
```promql
|
||||
# container metrics: anything with a name, and only that
|
||||
container_memory_working_set_bytes{job=~"cadvisor.*", name!=""}
|
||||
```
|
||||
|
||||
少了 `name!=""`,一条\"容器内存超过 10 GB\"的规则会对主机自己告警。
|
||||
|
||||
## 错误二:我的\"总容量\"是虚构的
|
||||
|
||||
聚合面板好看,但就是错的。原因:**同一个文件系统会在指标里出现多次。**
|
||||
|
||||
- 在 NAS 上,主卷是 `/volume1`,而 `/opt` 是**同一个** btrfs 文件系统,只是换了个挂载点。
|
||||
- 在 unRaid 上,同一个 NAS 卷又以 CIFS 挂载成 `/mnt/remotes/<nas>_ActiveBackup`。
|
||||
- unRaid 的 `/var/lib/docker` 是那个池的子卷,而 `/mnt/ssd` 已经代表过同一个池——同样的字节,第二个身份。
|
||||
|
||||
所以朴素的 `sum(node_filesystem_size_bytes)` 报出了我并不存在的容量。诚实的写法是明确列出什么算数据存储:
|
||||
|
||||
```promql
|
||||
node_filesystem_size_bytes{
|
||||
mountpoint=~"/volume[0-9]+|/mnt/ssd|/mnt/disk[0-9]+",
|
||||
fstype!~"fuse.*|tmpfs|rootfs"
|
||||
} or node_filesystem_size_bytes{job="vps-host", mountpoint="/"}
|
||||
```
|
||||
|
||||
还有两条应该写进面板说明的诚实备注,因为没人能解释的数字比没有数字更糟:
|
||||
|
||||
- **校验盘对内核不可见。** unRaid 的奇偶校验盘没有文件系统,所以从不出现在指标里——这个和是可用容量,不是硬盘数量。
|
||||
- **镜像会虚增裸和。** 两块镜像 SSD 会各报一次自己的字节,和不是你能存的量。
|
||||
|
||||
还有那个以后一定会咬我的遗漏:unRaid 的 shfs 联合挂载(`/mnt/user`)报告 `avail=0`。把它算进\"剩余空间低于 X\"的规则里,就会产生一条永远无法解除的告警。
|
||||
|
||||
## 错误三:那台不导出 CPU 频率的虚拟机
|
||||
|
||||
为了算全家的\"总 CPU 频率\",我直接用 node_exporter 的 cpufreq collector。unRaid 和 NAS 都按核心报出了 `node_cpu_scaling_frequency_hertz`;VPS **什么都没有**——它是 KVM 客户机,而客户机没有 `/sys/devices/system/cpu/cpu0/cpufreq`。这个指标在那里无法存在。
|
||||
|
||||
修法是 textfile collector:一个小脚本读客户机**能**看到的东西(`/proc/cpuinfo`),把 Prometheus 格式指标写进 node_exporter 会抓取的目录:
|
||||
|
||||
```sh
|
||||
# /opt/node-exporter-textfile/cpu-mhz.sh — cron 每 5 分钟
|
||||
awk -F: '
|
||||
/^processor/ { c = $2; gsub(/[ \t]/, "", c) }
|
||||
/^cpu MHz/ { f = $2; gsub(/[ \t]/, "", f); printf "node_cpu_mhz_current_hz{core=\"%s\"} %.0f\n", c, f * 1000000 }
|
||||
' /proc/cpuinfo
|
||||
```
|
||||
|
||||
配合:
|
||||
|
||||
```yaml
|
||||
command:
|
||||
- '--collector.textfile.directory=/textfile'
|
||||
volumes:
|
||||
- /opt/node-exporter-textfile:/textfile:ro
|
||||
```
|
||||
|
||||
两个刻意的决定:指标名**故意不同**于 node_exporter 自己的(用 `node_cpu_mhz_current_hz`,不用 `node_cpu_scaling_frequency_hertz`),这样万一以后宿主机暴露真实 cpufreq,不会出现重复序列冲突;仪表盘则显式合并两个来源:
|
||||
|
||||
```promql
|
||||
sum(node_cpu_scaling_frequency_hertz or node_cpu_mhz_current_hz)
|
||||
```
|
||||
|
||||
## 错误四:exporter 把容器名当成了主机名
|
||||
|
||||
仪表盘的主机下拉框里出现了 `9f9afcccc962`。那是个容器 ID:node_exporter 的 `uname` collector 读到的是**进程自己的** UTS namespace,而容器的 hostname 默认就是它自己的 ID。指标没错,标签毫无意义。compose 里一行解决:
|
||||
|
||||
```yaml
|
||||
services:
|
||||
node_exporter:
|
||||
hostname: 2.hoelee.com # otherwise `nodename` = the container ID
|
||||
```
|
||||
|
||||
## 我下次会怎么做
|
||||
|
||||
- **刻意把总览屏留到最后做。** \"总容量是多少\"这个问题会逼出全部去重工作;先做完三台主机面板再发现,就得在每个地方重做这个数字。
|
||||
- **一开始就给容器钉 hostname。** 一行成本,省掉一个让人困惑的下拉框。
|
||||
- **假设每个虚拟化或家电式主机都会藏掉一类指标。** NAS 可能限制你的容器监控(DSM 的 Docker API 版本把我的 cAdvisor 钉在 v0.53.0——更新的版本要更新的 Docker API);虚拟机藏掉 cpufreq;路由器藏掉自己的 CPU。找缺口的方法是**数你期待的东西**,而不是相信\"采集器成功了\"。
|
||||
|
||||
## 结果
|
||||
|
||||
三台主机、**7 个抓取目标、一个 Grafana、16 条告警规则**,以及一块屏:**47 个 CPU 核心 · 当前 158 GHz(标称上限 205 GHz)· 142 GB 内存 · 64 TB 存储 · 155 个运行中容器**,并且每台机器克隆同一套主机/容器面板布局,三台读起来完全一致。告警进 Telegram 和邮件,Prometheus 保留 30 天。
|
||||
|
||||
最有用的产出不是仪表盘,而是重写存储规则那天触发的一条告警:一个只剩 21 GB 的 NAS 卷——它一直藏在\"99% 满\"的百分比阈值背后,而那个巨物早已变成背景噪声。
|
||||
|
||||
## 需要为你的业务做这个吗?
|
||||
|
||||
如果你有 NAS、VPS 和几台服务器,却没有一个地方能同时看它们,我可以帮你搭这样的自托管监控流水线——Prometheus + Grafana 覆盖你的机器,主机**与**容器指标,真正有意义的磁盘健康与容量规则,通知推到 Telegram 或邮件。
|
||||
|
||||
**WhatsApp:[+60 12-797 2969](https://wa.me/60127972969)** · **邮箱:[[email protected]](mailto:[email protected]?subject=Self-hosted%20monitoring%20setup)** · **[hoelee.com](https://hoelee.com)**
|
||||
|
||||
网站设计与开发是我的主业;服务器加固与自托管基础设施是它的另一半。
|
||||
@@ -0,0 +1,96 @@
|
||||
---
|
||||
title: "smartctl 退出码 32:专门跳过你最该看的硬盘的那个\"错误\""
|
||||
description: "把 smartctl 的任何非零退出码当成读取失败的监控脚本,恰好会跳过属性已经逼近阈值的那些盘。退出码 32 不是错误,而是一段历史。"
|
||||
pubDate: 2026-04-14
|
||||
updatedDate: 2026-09-29
|
||||
category: notes
|
||||
tags: [smart, smartctl, unraid, monitoring, bash, disks, homelab]
|
||||
ogImage: /og/smartctl-exit-code-32-skips-the-disks-that-matter.png
|
||||
banner: /banners/smartctl-exit-code-32-skips-the-disks-that-matter.png
|
||||
draft: false
|
||||
---
|
||||
|
||||
我的硬盘健康采集脚本连续几周都是\"全绿\"。每块盘都有温度、通电小时数和 SMART 结论,指标看起来完整——直到我数了一下只有 **4 块 SSD 里的 2 块**。
|
||||
|
||||
不是故障,是**消失了**。脚本认定这两块盘不存在。
|
||||
|
||||
## 问题出在把退出码当布尔值
|
||||
|
||||
采集脚本遍历块设备、对每块盘跑 `smartctl`、再写 Prometheus textfile 指标。错误处理看起来挺防御:
|
||||
|
||||
```bash
|
||||
for dev in /dev/sd?; do
|
||||
if ! smartctl -A -H "$dev" > /tmp/smart.out 2>&1; then
|
||||
continue # "盘不支持 SMART / 读不到"
|
||||
fi
|
||||
# 解析并输出指标
|
||||
done
|
||||
```
|
||||
|
||||
`if ! cmd` 是布尔判断,但 `smartctl` 的退出状态**不是布尔值,是一个位域**。把位域当真假用,就是这次翻车的原因。
|
||||
|
||||
## 退出码 32 的真实含义
|
||||
|
||||
`man smartctl` 里这些位是独立累加的:
|
||||
|
||||
| 位 | 值 | 含义 |
|
||||
|---|---|---|
|
||||
| 0 | 1 | 命令行解析失败 |
|
||||
| 1 | 2 | 设备打开失败 / 无 IDENTIFY DEVICE 结构 |
|
||||
| 2 | 4 | SMART 或 ATA 命令失败、校验和错误 |
|
||||
| 3 | **8** | SMART 状态为 **DISK FAILING** |
|
||||
| 4 | **16** | 有预失效属性 **<= 阈值** |
|
||||
| 5 | **32** | SMART 状态 **OK**,但某些属性**曾经**低于阈值 |
|
||||
| 6 | 64 | 设备错误日志中有记录 |
|
||||
| 7 | 128 | 自检日志中有失败记录 |
|
||||
|
||||
所以 `32` 的意思和\"读不到盘\"正好相反:**现在没问题,但它曾经踩过阈值。** 这正是应该持续盯着的那类盘——而我的脚本把它们丢掉了。
|
||||
|
||||
消失的两块盘,恰好是原始属性长期处在边缘值的那两块;退出码为 `0` 的\"健康盘\"被正常采集。**这个过滤器实际上筛选出了\"没有任何历史可报告\"的盘。**
|
||||
|
||||
## 修法:把\"信息位\"掩掉
|
||||
|
||||
对监控来说,位 32 和 64 是信息;位 8 和 16 才是该告警的。掩掉信息位,只对剩下的做判断:
|
||||
|
||||
```bash
|
||||
smartctl -A -H -d sat "$dev" > /tmp/smart.out 2>&1
|
||||
rc=$?
|
||||
|
||||
# fatal bits: 1 (parse), 2 (open), 4 (command/checksum), 8 (FAILING)
|
||||
# informational bits: 32 (was below threshold in the past), 64 (error log has records)
|
||||
fatal=$(( rc & ~(32 | 64) ))
|
||||
if [ "$fatal" -ne 0 ]; then
|
||||
echo "device $dev unreadable or failing (rc=$rc)" >&2
|
||||
continue
|
||||
fi
|
||||
|
||||
# emit the verdict AND the exit code, so the code itself is a metric
|
||||
echo "disk_smart_exit_code{device=\"$dev\"} $rc"
|
||||
echo "disk_smart_health{device=\"$dev\"} $(( rc & 8 ? 0 : 1 ))"
|
||||
```
|
||||
|
||||
三点改变:
|
||||
|
||||
1. **退出码变成数据,而不是控制流。** 它作为指标被导出,磁盘从 `0` → `32` → `64` 的漂移会变成趋势,而不是一次静默跳过。
|
||||
2. **信息位不再致命。** 有历史的盘被采集,这才是监控它们的目的。
|
||||
3. **NAS 上 `-d sat` 很关键。** 在 Synology DSM(以及某些 USB 桥接)上,设备类型不对时 `smartctl` 拿不到有用输出——症状同样是\"盘消失\",原因却不同。
|
||||
|
||||
## 我下次会怎么做
|
||||
|
||||
- **面对有\"退出码位域\"文档的工具,永远不要写 `if ! cmd`。** 先把手册里的退出码那节读完。
|
||||
- **除了\"采集器有没有跑\",还要数\"采到了几块盘\"。** 我的告警是\"采集器过期\",真正的 bug 是\"采集器正常跑了,只报了 50% 的盘\"。一个 `disk_count` 指标就能立刻暴露。
|
||||
- **把\"这个设备没有数据\"本身当成可告警状态。** 缺失的序列是不可见的,所以它存活了好几周。
|
||||
|
||||
## 结果
|
||||
|
||||
采集器从 2 块可用盘变成 4 块,每块都报温度、通电小时数、SMART 状态和原始退出码——共 122 个 textfile 指标,包括我真正想要的 btrfs 错误计数。重新出现的那两块,正是长期处在边缘属性值的盘。
|
||||
|
||||
一个会跳过自己最坏信号的监控流水线,比没有监控更糟——因为它一直告诉你一切正常。
|
||||
|
||||
## 需要为你的业务做这个吗?
|
||||
|
||||
如果你在跑 NAS 或服务器机架,想让硬盘健康真正\"会通知你\"——SMART 属性、温度、btrfs/RAID 错误计数,并推到 Telegram 或邮件——这类自托管监控流水线我可以帮你搭。
|
||||
|
||||
**WhatsApp:[+60 12-797 2969](https://wa.me/60127972969)** · **邮箱:[[email protected]](mailto:[email protected]?subject=Disk%20health%20monitoring)** · **[hoelee.com](https://hoelee.com)**
|
||||
|
||||
网站设计与开发是我的主业;服务器加固与自托管基础设施是它的另一半。
|
||||
@@ -0,0 +1,112 @@
|
||||
---
|
||||
title: "Prometheus 明明有数据,Grafana 面板却全是 No data 的原因"
|
||||
description: "所有 target 都是 up、数据就在 Prometheus 里、面板表达式也没错——41 个面板却全部 No data。原因是 Grafana 变量过滤了自己,而我的验证方式恰好掩盖了它。"
|
||||
pubDate: 2026-05-17
|
||||
updatedDate: 2026-09-29
|
||||
category: devops
|
||||
tags: [grafana, prometheus, dashboards, monitoring, promql, observability]
|
||||
ogImage: /og/why-your-grafana-dashboard-shows-no-data.png
|
||||
banner: /banners/why-your-grafana-dashboard-shows-no-data.png
|
||||
draft: false
|
||||
---
|
||||
|
||||
打开每天在看的面板,结果是空的。不是\"某一格坏了\",而是**从上到下每一格**都写着 *No data*。去查 Prometheus,一切健康。这篇讲我踩到的那个具体且不直观的原因,以及为什么我自己的验证会同时告诉我\"面板没问题\"。
|
||||
|
||||
## 最省时间的排查顺序
|
||||
|
||||
常见原因(数据源选错、时间范围不对、`rate()` 用在会重置的计数器上、target 没被抓取)都不适用。真正有效的检查顺序是:
|
||||
|
||||
```bash
|
||||
# 1. target 真被抓取了吗?
|
||||
curl -s http://prometheus:9090/api/v1/targets | jq '.data.activeTargets[] | {job:.labels.job, health}'
|
||||
|
||||
# 2. 指标现在存在吗?(完全绕过 Grafana)
|
||||
curl -s -G http://prometheus:9090/api/v1/query \
|
||||
--data-urlencode 'query=node_uname_info{job="unraid-host"}' | jq '.data.result[0].metric'
|
||||
|
||||
# 3. 面板表达式本身对吗?
|
||||
curl -s -G http://prometheus:9090/api/v1/query \
|
||||
--data-urlencode 'query=count(node_cpu_seconds_total{mode="idle",job="unraid-host"})'
|
||||
```
|
||||
|
||||
三项都过了:7 个 target `up`、2.5 万条序列在采集、手动跑面板表达式也有数据。所以问题不在指标、不在抓取、也不在 PromQL——而在**它们之间的变量层**。
|
||||
|
||||
## 真正的原因:变量过滤了自己
|
||||
|
||||
面板上有一个主机选择器,两个变量串联:`$nodename`(哪台机器)和 `$node`(它的 `instance` 标签)。第一个的定义是:
|
||||
|
||||
```promql
|
||||
# BROKEN — 变量用自身的值来过滤自己
|
||||
label_values(node_uname_info{job="unraid-host", nodename=~"$nodename"}, nodename)
|
||||
```
|
||||
|
||||
首次加载时 `$nodename` 还没有值,选择器就变成 `nodename=~""`——**什么都匹配不到**。查询返回零行的变量没有任何选项,于是它保持空;而 `$node` 又是基于 `$nodename` 定义的:
|
||||
|
||||
```promql
|
||||
label_values(node_uname_info{job="unraid-host", nodename="$nodename"}, instance)
|
||||
```
|
||||
|
||||
于是 `$node` 也是空,所有以 `$node` 为过滤条件的面板自然匹配不到序列。Grafana 说 *No data* 并没有骗人:面板查询真的没有数据,因为给它限定范围的变量解析成了空。
|
||||
|
||||
修法是去掉自引用,并给每个变量存一个明确的默认值:
|
||||
|
||||
```promql
|
||||
# FIXED — 不自引用,每个变量只跳一层
|
||||
# $nodename
|
||||
label_values(node_uname_info{job="unraid-host"}, nodename)
|
||||
# $node
|
||||
label_values(node_uname_info{job="unraid-host", nodename="$nodename"}, instance)
|
||||
```
|
||||
|
||||
## 为什么我的验证没抓到
|
||||
|
||||
这才是值得抄的部分。我\"验证\"面板的方式是:**手动**把每个面板表达式里的变量替换成我知道正确的值(`$nodename=unRaid`、`$node=host.docker.internal:9100`),然后断言查询有数据点。13 项检查全部通过——面板依然是空的,因为 bug 在被替换的那个值**上游**:Grafana 自己对变量的解析是空的。
|
||||
|
||||
**一个替换掉\"可疑值\"的验证,永远发现不了那个值解析失败。** 要抓它,得读 Grafana 真实存下来的值:
|
||||
|
||||
```bash
|
||||
curl -s -u admin:"$PW" http://grafana:4010/api/dashboards/uid/rYdddlPWk \
|
||||
| jq '.dashboard.templating.list[] | {name, current: .current.value}'
|
||||
```
|
||||
|
||||
更直接的检查对象是变量查询本身——跑 Grafana 会跑的那条:
|
||||
|
||||
```bash
|
||||
# $nodename 的 label_values() 到底返回什么?
|
||||
curl -s -G http://prometheus:9090/api/v1/query \
|
||||
--data-urlencode 'query=count by (nodename) (node_uname_info{job="unraid-host"})' | jq '.data.result'
|
||||
```
|
||||
|
||||
那里是空,面板就不可能渲染出来,无论数据多健康。
|
||||
|
||||
## 第二个坑:`$__all` 不是 `.*`
|
||||
|
||||
做自动化面板检查时,值为\"All\"的变量,其 current 是哨兵字符串 `$__all`——不是正则。把 `name=~"$__all"` 原样展开会匹配不到任何东西,于是完全健康的容器面板\"全军覆没\"。展开前要把 `$__all` 解析成该变量的 `allValue`(通常是 `.*`),否则你追的是一个只存在于自己检查脚本里的 bug。
|
||||
|
||||
## 不是每个空面板都是 bug
|
||||
|
||||
修完之后 25 个面板有 21 个出数。剩下 4 个是**设计上就不可能**有数据的,最好在面板说明里写清楚:
|
||||
|
||||
- **PSI 面板**需要 `/proc/pressure`,而这台 NAS 的内核不暴露它——指标在这台机器上无法存在;同样的面板在会导出 `node_pressure_*` 的主机上正常显示。
|
||||
- **根文件系统面板**的表达式带 `fstype!="rootfs"`,而某台主机的 `/` 是**内存盘**(unRaid 就是),所以永远匹配不到。
|
||||
|
||||
区分\"因为 bug 空\"和\"因为这台机器不可能有这个指标而空\",是真修好和瞎忙一下午的分界线。
|
||||
|
||||
## 我下次会怎么做
|
||||
|
||||
- **永远不要定义过滤自己的模板变量。** 一个变量只跳一层。
|
||||
- **给面板依赖的变量存明确的 `current` 值**,别让全新加载依赖下拉框被选中。
|
||||
- **用 Grafana 存下来的值做验证**,而不是我以为的值。
|
||||
- **故意留空的缺口写进面板说明**,否则未来的你会花一个下午去\"修\"一个本来就没打算工作的面板。
|
||||
|
||||
## 结果
|
||||
|
||||
三台主机现在跑同一套布局——每台一个 41 面板的主机面板 + 一个 10 面板的容器面板,分别有 21/23/25 个面板出数,两类结构性空缺写在面板里,而不是留给下一个打开它的人去猜。
|
||||
|
||||
## 需要为你的业务做这个吗?
|
||||
|
||||
如果你有一堆没人信任的 Grafana 面板,或者服务器和 NAS 完全没有监控,我可以帮你搭自托管的 Prometheus + Grafana(主机与容器指标、合理的告警规则、通知推到 Telegram 和邮件),也会把你那些悄悄\"不出数\"的面板修好。
|
||||
|
||||
**WhatsApp:[+60 12-797 2969](https://wa.me/60127972969)** · **邮箱:[[email protected]](mailto:[email protected]?subject=Grafana%20monitoring)** · **[hoelee.com](https://hoelee.com)**
|
||||
|
||||
网站设计与开发是我的主业;服务器加固与自托管基础设施是它的另一半。
|
||||
@@ -0,0 +1,87 @@
|
||||
---
|
||||
title: "你的\"磁盘将满\"告警在说谎:别再按百分比告警"
|
||||
description: "\"剩余 3%\"听着像紧急事故,直到你发现那是 500 GB。在多 TB 的卷上,百分比阈值会在没事的时候乱叫,也会在真出事的时候沉默。"
|
||||
pubDate: 2026-06-24
|
||||
updatedDate: 2026-09-29
|
||||
category: devops
|
||||
tags: [monitoring, alerting, prometheus, grafana, storage, capacity, observability]
|
||||
ogImage: /og/your-disk-full-alert-is-lying.png
|
||||
banner: /banners/your-disk-full-alert-is-lying.png
|
||||
draft: false
|
||||
---
|
||||
|
||||
我搭好监控、写了一套看着合理的规则,然后用一个下午删掉了其中大半阈值。它们全是百分比,而且都以同一种方式错了。
|
||||
|
||||
## 三个教我道理的告警
|
||||
|
||||
**1.「文件系统使用率超过 90%」**——它在一个还剩约 500 GB 的 17 TB 卷上触发。没事发生,也不会马上有事发生。这种体量的卷,一次大备份就能让百分比摆动几个点:比例是噪声,而\"还有 500 GB 可用\"才是世界的真实状态。
|
||||
|
||||
**2.「内存使用率超过 90%」**——在一台**可用**内存 4.8 GB 的主机上反复触发。Linux 会用空闲内存做 page cache,需要时立刻归还;`MemTotal - MemFree` 描述的是文件缓存,不是你的风险。真正预示 OOM 的是 `MemAvailable`。
|
||||
|
||||
**3.「CPU steal 超过 25%」**——一台把邮件栈跑得完全正常的 VPS,steal 长期在 41%。steal 说明宿主机忙,它是**容量**信号而不是**故障**信号;7 vCPU 的机器 41% steal 就是这台机器的成本。这条规则产生的告警永远为真、也永远无用。
|
||||
|
||||
共同点:**我在对比例告警,而比例不知道东西有多大。**
|
||||
|
||||
## 对\"后果\"告警,而不是对比例
|
||||
|
||||
我现在对每个阈值都问一句:*如果这个状态持续下去,实际会发生什么?* 答案几乎总能用绝对单位表达。
|
||||
|
||||
| 不要这样 | 改成 | 原因 |
|
||||
|---|---|---|
|
||||
| 文件系统 > 90% | 剩余空间 < N GB | \"我还能不能写进去\"才是问题 |
|
||||
| 内存使用 > 90% | `MemAvailable` < 512 MB | 逼近 OOM,而不是\"缓存很暖和\" |
|
||||
| CPU steal > 25% | steal > 50% | 容量成本 vs 真的被饿死 |
|
||||
| 负载 > N | `load1 / 核数 > 2` | 不谈核数的负载没有意义 |
|
||||
|
||||
PromQL 的前后对比:
|
||||
|
||||
```promql
|
||||
# BEFORE — 17 TB 卷的百分比;还剩半 TB 就开叫
|
||||
(1 - node_filesystem_avail_bytes / node_filesystem_size_bytes) * 100 > 90
|
||||
|
||||
# AFTER — 真正装数据的卷上,还剩多少容量
|
||||
min by (instance, mountpoint) (
|
||||
node_filesystem_avail_bytes{mountpoint=~"/volume[0-9]+|/mnt/ssd|/mnt/disk[0-9]+"}
|
||||
) < 25e9
|
||||
```
|
||||
|
||||
```promql
|
||||
# BEFORE — 把 page cache 当成压力
|
||||
(1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100 > 90
|
||||
|
||||
# AFTER — 真正的 OOM 前兆
|
||||
node_memory_MemAvailable_bytes < 512 * 1024 * 1024
|
||||
```
|
||||
|
||||
两个让规则能活下来的细节:
|
||||
|
||||
**把比较放进阈值条件里,不要塞进表达式。** 查询就写 `min by (mountpoint) (node_filesystem_avail_bytes{...})`,让告警规则去和 `25000000000` 比。若你在 PromQL 里写死 `< 25e9`、规则里又留着继承来的 `> 90`,规则照样会评估——但界面显示的阈值是错的,下一个读它的人得从两个互相矛盾的条件下反推你的意图。
|
||||
|
||||
**排除那些\"只是另一个挂载点的视图\"。** 一条\"剩余空间\"规则有多诚实,取决于它的选择器。在我自己的环境里,同一个 NAS 卷出现三次(`/volume1`、`/opt`、以及在另一台主机上以 CIFS 再次挂载),而 unRaid 的 shfs 联合挂载报告 `avail=0`——不加排除的规则不是永远误报,就是把容量算成两倍。一个写明白的选择器胜过聪明的通用写法。
|
||||
|
||||
## 噪声不是无害的
|
||||
|
||||
坏阈值的代价不在告警本身,而在**训练**:每条\"没事乱叫\"的告警都在教你先扫一眼、然后忽略;等你真正该看的那条到来时,它出现在一个你早已不再读的频道里。我的规则总数**减少了**,覆盖率却上升了——更少,但每条都对应一个明确的后果。
|
||||
|
||||
## 阈值是私人的,形状是通用的
|
||||
|
||||
上面的数字是我的:NAS 卷剩不到 25 GB 我会在意、可用内存低于 512 MB 我才管、VPS steal 超过 50% 才算被饿。你的数字随硬件和容忍度而变。
|
||||
|
||||
能通用的是形状:
|
||||
|
||||
- 用绝对单位而非比例;
|
||||
- 把后果写进告警的摘要里;
|
||||
- 选择器明确声明它在替哪些挂载点说话;
|
||||
- 规则数量少到你还能读完每一条。
|
||||
|
||||
## 结果
|
||||
|
||||
同一套栈,从\"三条永久无用报警\"变成 16 条在机房健康时保持安静的规则——其中一条立刻暴露了一个只剩 21 GB 的 NAS 卷,而那正是百分比规则一直藏在\"99% 满的巨物\"背后、被你学会忽略的东西。
|
||||
|
||||
## 需要为你的业务做这个吗?
|
||||
|
||||
如果你的监控在发你已经学会忽略的告警,或者你根本没有监控、宁愿从一条规则而不是从一次失败的备份里得知磁盘满了,我可以帮你搭自托管的监控与告警(Prometheus + Grafana,通知推 Telegram 和邮件),阈值按\"什么会真的坏\"来定。
|
||||
|
||||
**WhatsApp:[+60 12-797 2969](https://wa.me/60127972969)** · **邮箱:[[email protected]](mailto:[email protected]?subject=Monitoring%20and%20alerting)** · **[hoelee.com](https://hoelee.com)**
|
||||
|
||||
网站设计与开发是我的主业;服务器加固与自托管基础设施是它的另一半。
|
||||